Multimodal transport path selection method based on multi-agent reinforcement learning

By optimizing intermodal transport routes with a multi-agent reinforcement learning method combined with an adaptive large neighborhood search algorithm, the path reconstruction and resource optimization problems of intermodal transport route optimization in a dynamic environment are solved, achieving efficient route planning and resource utilization.

CN120765154AActive Publication Date: 2025-10-10SOUTHWEST JIAOTONG UNIV

Patent Information

Application Number
CN202510946008.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-10
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing multimodal transport route optimization technologies have obvious defects in dealing with service time uncertainty, achieving multi-carrier collaborative decision-making, and completing real-time route reconstruction. In particular, it is difficult to achieve effective route reconstruction and resource optimization in dynamic environments.

Method used

A multimodal transport network environment is constructed by adopting a multi-agent reinforcement learning method combined with an adaptive large neighborhood search algorithm. The initial planning is optimized through the multi-agent reinforcement learning algorithm, unexpected action decisions are generated, and dynamic path reconstruction and resource optimization are achieved.

Benefits of technology

It improves the utilization efficiency of transportation resources and the fairness of task allocation, enhances the system's disturbance self-recovery ability, effectively alleviates the dimensionality explosion problem, and improves the generalization ability of the model and the convergence speed of the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765154A_ABST
    Figure CN120765154A_ABST
Patent Text Reader

Abstract

The invention relates to a multimodal transport path selection method based on multi-agent reinforcement learning. The method comprises the following steps: constructing a multimodal transport network and a system environment; carrying out initial planning by adopting a self-adaptive large neighborhood search algorithm based on a multimodal transport network and a system environment; and when an accident occurs, acquiring accident information, optimizing the initial plan by adopting a multi-agent reinforcement learning algorithm based on the accident information in combination with a cost objective function to obtain an accident action decision, and executing a transportation task based on the accident action decision. By constructing a large neighborhood disturbance response mechanism fused with ALNS, when it is detected that node / path service time offset exceeds a threshold value, path reconstruction is intelligently triggered, a new alternative path is generated, based on an MADDPG multi-agent reinforcement learning framework, all carrier Agents can obtain global information in centralized training and make decisions independently in execution, and therefore the probability that the node / path service time offset exceeds the threshold value is lowered. And the utilization efficiency of transportation resources and the task allocation fairness are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of logistics and transportation technology, and in particular to a multimodal transport path selection method based on multi-agent reinforcement learning. Background Art

[0002] Intermodal transport, a transportation organization consisting of at least two modes of transport (such as road, rail, and water), has become a key component of modern logistics systems due to its significant advantages, including wide coverage, low overall costs, and low carbon emissions. In the current development of logistics, intermodal transport is of great significance for integrating transportation resources, improving transportation efficiency, and enhancing supply chain resilience.

[0003] Currently, existing research on multimodal transport route optimization primarily employs operations research methods and metaheuristic algorithms. Operational research methods include mixed integer programming, robust optimization, and stochastic programming, while metaheuristic algorithms include genetic algorithms, simulated annealing, tabu search, and adaptive large neighborhood search. However, these methods generally rely on pre-set transportation network parameters and traffic environment assumptions. In dynamic environments, they lack the ability to perceive and respond to factors such as service time disturbances and transfer node instability in real time. When faced with unexpected events such as weather disasters, traffic control, and equipment failures that lead to service time uncertainty, traditional methods struggle to achieve effective dynamic route reconstruction, which can easily lead to resource waste, transportation delays, and even route failure.

[0004] Furthermore, while reinforcement learning methods, which have emerged in recent years, have been applied to the field of intermodal transport optimization, hoping to enhance the adaptive capabilities of route planning through learning mechanisms, most reinforcement learning methods are only applicable to single-agent scenarios and cannot effectively handle the collaboration and game-playing relationships between multiple carriers. In an intermodal, multi-carrier environment, each carrier makes route selection decisions based on local information, facing challenges such as high state space dimensionality, incomplete information, and severe policy coupling. Existing multi-agent reinforcement learning methods lack effective path reconstruction and coordination mechanisms when dealing with heterogeneous transport modes, dynamic transport demand, and service disturbances, and are prone to falling into local optimality or experiencing training instability.

[0005] In summary, existing multimodal transport path optimization technology has obvious defects in dealing with service time uncertainty, realizing multi-carrier collaborative decision-making, and completing real-time path reconstruction. It is urgent to develop a new path optimization method with dynamic perception, collaborative learning and global optimization capabilities to meet the actual needs of multimodal transport development. Summary of the Invention

[0006] Based on this, it is necessary to provide a multimodal transport path selection method based on multi-agent reinforcement learning to address the above technical problems.

[0007] In a first aspect, the present application provides a multimodal transport path selection method based on multi-agent reinforcement learning. The method comprises: constructing a multimodal transport network and system environment; initial planning based on the multimodal transport network and system environment using an adaptive large neighborhood search algorithm; when an unexpected event occurs, obtaining unexpected event information, and based on the unexpected event information, combining a cost objective function to optimize the initial planning using a multi-agent reinforcement learning algorithm to obtain an unexpected action decision, and executing a transport task based on the unexpected action decision.

[0008] Optionally, in an embodiment of the present application, the cost objective function includes transport cost, transfer cost, storage cost, waiting cost, carbon tax cost, and delay penalty cost.

[0009] Optionally, in an embodiment of the present application, the obtaining of the unexpected event information comprises: obtaining global state information and affected transport task information after the adaptive large neighborhood search algorithm planning.

[0010] Optionally, in an embodiment of the present application, the optimization of the initial planning based on the unexpected event information and the cost objective function using the multi-agent reinforcement learning algorithm comprises: determining service time constraints based on the affected transport task information; optimizing the initial planning based on the service time constraints and the cost objective function using the multi-agent reinforcement learning algorithm when an unexpected event may occur; collecting relevant information to learn and train the multi-agent reinforcement learning algorithm when the unexpected event ends.

[0011] Optionally, in an embodiment of the present application, the collecting of relevant information to learn and train the multi-agent reinforcement learning algorithm when the unexpected event ends comprises: constructing a state space of the multi-agent reinforcement learning algorithm based on the global state information and the affected transport task information; determining a reward function based on the cost objective function and the delay penalty; performing action selection in an action space based on the state space, calculating a reward value based on the action selection and the reward function, and optimizing network parameters of the state space and the action space based on the reward value.

[0012] In a second aspect, the present application also provides a multimodal transport path selection device based on multi-agent reinforcement learning. The device comprises: an environment construction module for constructing a multimodal transport network and system environment; an initial planning module, configured to perform initial planning based on the multimodal transport network and system environment using an adaptive large neighborhood search algorithm; a re-planning module, configured to, when an unexpected event occurs, acquire unexpected event information, optimize the initial planning based on the unexpected event information and a cost objective function using a multi-agent reinforcement learning algorithm, obtain an unexpected action decision, and execute a transport task based on the unexpected action decision.

[0013] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor executes the steps of the method in each of the above embodiments.

[0014] In a fourth aspect, the present application also provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the method in each of the above embodiments.

[0015] The above-mentioned multimodal transport path selection method based on multi-agent reinforcement learning first constructs a multimodal transport network and system environment, then performs initial planning based on the multimodal transport network and system environment using an adaptive large neighborhood search algorithm, and finally, when an unexpected event occurs, acquires unexpected event information, optimizes the initial planning based on the unexpected event information and a cost objective function using a multi-agent reinforcement learning algorithm, obtains an unexpected action decision, and executes a transport task based on the unexpected action decision. That is, by constructing a large neighborhood disturbance response mechanism that integrates ALNS, when it is detected that the node / path service time offset exceeds a threshold value, the path reconstruction is intelligently triggered and a new candidate path is generated, so that the system has the ability of “disturbance self-recovery”; based on the MADDPG multi-agent reinforcement learning framework, each carrier Agent can obtain global information in centralized training and make independent decisions in execution, realizing an efficient path planning mechanism of “centralized training + distributed execution”, effectively improving the utilization efficiency of transport resources and the fairness of task allocation; a hierarchical action design is proposed (first selecting a transport mode by MARL, and then selecting a specific carrier by ALNS), which reduces the original action space dimension from to , effectively alleviating the dimension explosion problem; in the disturbance reconstruction phase, ALNS as an external heuristic module provides high-quality candidate paths for MARL, which helps to avoid the MARL model falling into local optimum, and at the same time, the best “break-repair” path operation sequence is selected through a dynamic operator weight adjustment mechanism. In the path reconstruction phase, the ALNS algorithm disturbs the original path through the “break-repair” mechanism, thereby generating a set of potential high-quality candidate paths. This process not only accelerates the convergence of the training process, but also effectively improves the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a diagram of an application environment for a multimodal transport path selection method based on multi-agent reinforcement learning in one embodiment; Figure 2 1 is a flow chart of a multimodal transport route selection method based on multi-agent reinforcement learning in one embodiment; Figure 3 Schematic diagram of the interaction between ALNS and MARL in one embodiment; Figure 4 is a schematic diagram of a re-planning process in one embodiment; Figure 5 1 is a flow chart of a multi-agent reinforcement learning algorithm according to one embodiment; Figure 6 This is a structural block diagram of a multimodal transport path selection device based on multi-agent reinforcement learning in one embodiment; Figure 7 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0018] The embodiment of the present application provides a multimodal transport path selection method based on multi-agent reinforcement learning, which can be applied to Figure 1 In the application environment shown, the terminal communicates with the server through the network. The data storage system can store data that the server needs to process. The data storage system can be integrated on the server or placed on the cloud or other network servers. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented as a standalone server or a server cluster consisting of multiple servers.

[0019] In one embodiment, Figure 2 As shown in the figure, a multimodal transport path selection method based on multi-agent reinforcement learning is provided. Figure 1 The following steps are used as an example to illustrate the server in the example: S201: Build a multimodal transport network and system environment.

[0020] In the embodiments of the present application, first, a multimodal transport network and system environment is constructed. A transport network model with multiple transport modes is constructed , denotes a node set, including order departure nodes, arrival nodes, transfer nodes, etc., that is ; denotes an edge set, and each two nodes contain at most three edges, representing different transport modes, and each edge has a different distance.

[0021] S203: Based on the multimodal transport network and system environment, an adaptive large neighborhood search algorithm is used for initial planning.

[0022] In the embodiments of the present application, the adaptive large neighborhood search algorithm is used for initial planning based on the multimodal transport network and system environment, that is, the initial solution is generated by using the adaptive large neighborhood search algorithm ALNS. The steps of generating a solution by ALNS are as follows: a set of predefined destruction operators (such as random removal and worst removal) are used to selectively remove part of the elements in the current feasible solution to form an incomplete solution; then, a set of repair operators (such as greedy insertion and regret value insertion) are introduced to reconstruct the incomplete solution to generate a new candidate solution to expand the search neighborhood. Constant destruction and repair are performed until the termination condition is met to obtain a solution.

[0023] S205: When an unexpected event occurs, obtain unexpected event information, and based on the unexpected event information, combine a cost objective function to use a multi-agent reinforcement learning algorithm to optimize the initial planning to obtain an unexpected action decision, and execute the transport task based on the unexpected action decision.

[0024] In the embodiments of the present application, during the transport process, there is a possibility that an unexpected event occurs, which starts at and ends at . When planning the initial transport plan, the location, start time and end time of the unexpected event are unknown. Due to the existence of the unexpected event, the transport service time of the carrier on the transport path is uncertain. If the transport task is delivered to the node and cannot complete the transport task according to the plan, then the transport task is an affected transport task. If the transport task is planned and executed without considering the unexpected event, the delivery time of the transport task may exceed the agreed time and incur a delay penalty cost. Therefore, in order to avoid the delay penalty, an action needs to be taken to avoid the delay.

[0025] By acquiring unexpected event information, a multi-agent reinforcement learning algorithm MARL is used to make decisions in combination with a target cost function. ALNS sends state information to MARL and provides feasible vehicles for road, rail and water respectively, and then waits for actions from MARL, or takes a waiting action. The above steps are repeated until all transportation tasks have been delivered.

[0026] In an embodiment of the present application, the cost target function includes transportation cost, transfer cost, storage cost, waiting cost, carbon tax cost and delay penalty cost.

[0027] In an embodiment of the present application, the target is to minimize the cost, and the cost target function is as follows: the cost consists of transportation cost , transfer cost , storage cost , waiting cost , carbon tax cost , and delay penalty cost .

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035] wherein, is the total cost, is the transportation cost, is the transfer cost, is the storage cost, is the waiting cost, is the carbon tax cost, is the delay penalty cost, is the set of transportation vehicles; is the set of transportation paths (channels), is the set of pickup arcs, is the set of delivery arcs; is the set of transportation tasks ; and is the transfer point. is the unit cost of different items, , is the transportation cost per container per hour per km; is the handling cost per container; is the storage cost per container per hour; is the waiting cost per hour; is the carbon tax cost per container per hour; is the delay cost per container per hour; is the travel time of vehicle on path to , is the distance (km) of path to , is the quantity of goods (TEU) of transportation task ; is a 0-1 variable, equal to 1 if transportation task uses vehicle to travel on path to , otherwise 0; is a 0-1 variable, equal to 1 if transportation task is transferred from vehicle to vehicle at transshipment node , otherwise 0; and is the service start time of transportation task at node for vehicle , is the service end time of transportation task at node for vehicle , is the pickup start time of transportation task , is the waiting time of vehicle at node , is the carbon tax coefficient per ton of vehicle , is the delay time of transportation task at delivery node.

[0036] ​​​​​​​In this embodiment, a multi-dimensional reward function is designed, which includes transportation costs, delay penalties, carbon emissions, etc., to ensure that the generated path not only meets the timeliness requirements but also has the best cost.

[0037] In one embodiment of the present application, obtaining accident information includes: Obtain the global state information and affected transportation task information after the adaptive large neighborhood search algorithm planning.

[0038] In one embodiment of the present application, when the node or channel In time An accident occurred When first determining which mode of transport Affected by this event. Then for the collection Each vehicle in , check whether the process vehicle passes the node or channel If passed, all nodes in the process are identified There is operation or passing through the channel transportation tasks .right Each transport task in If the service start time of the transportation plan is greater than the event occurrence time , then add the transport task to the set of transport tasks affected by the accident middle.

[0039] In one embodiment of the present application, optimizing the initial plan using a multi-agent reinforcement learning algorithm based on the accident information in combination with a cost objective function includes: S301: Determine service time constraints based on the affected transportation task information.

[0040] S303: When an unexpected event may occur, the initial plan is optimized using a multi-agent reinforcement learning algorithm based on the service time constraint and the cost objective function.

[0041] S305: When the accident ends, relevant information is collected to learn and train the multi-agent reinforcement learning algorithm.

[0042] In one embodiment of the present application, when the service start time is scheduled at the node Previous accident or during transport route time Previous accident The service should be in the event After solving it, continue to start, that is, there is a constraint:

[0043]

[0044] In the constraint, the end time of the accident Not sure, this will affect transport vehicles At the node or channel If the wrong transportation route and vehicle are used, it may result in the transportation time window not meeting the service requirements and Longer waiting times , which resulted in significant delays .

[0045] By adopting an event-triggered mechanism, the framework consists of two stages: Possible re-planning before unexpected events occur, and Evaluation / learning is performed at the end of the unexpected event. By combining replanning and learning, the ALNS model is used to assist multi-agent reinforcement learning. The replanning phase stores information about the time step Because MARL cannot obtain immediate rewards for actions at the time of an event, this allows MARL's multi-agent system to learn when the event ends. When an event begins, the situations faced by all vehicles at different endpoints are stored. When the event ends, reinforcement learning is trained by simulating the situation at the time of the event. Since the duration of the event is known at its end, the rewards of the actions taken can be calculated. During the reinforcement learning process, the ALNS model provides the vast majority of information about the environment. During the replanning phase, MARL provides the ALNS model with replanning action plans.

[0046] In multi-agent reinforcement learning, each agent will perform a series of discrete time steps. China and the Environment In addition to all vehicles and the original transportation plan data, the environment also contains the start time and location of unexpected events on the route and nodes that cause service time uncertainty. Figure 3 As shown, at each time step , for the transportation task at the current time , MARL will receive global status information from ALNS , and provide each agent with an optimal option for each of the three modes of transportation: road, rail, and water (if available). Each agent will locally observe , and chooses an action from the three actions according to its strategy When the event ends, the unexpected events that occur The actual start time and duration are known. ALNS checks the feasibility of the schedule after adding this information and gives each agent in MARL a reward The goal of each agent in MARL is to , choose appropriate actions to maximize the cumulative reward , learning to handle unexpected events in the process. During the learning process, ALNS and MARL run synchronously. When a transport task is affected, it is removed and the relevant state information is sent to the reinforcement learning agent. Multiple agents in MARL then determine the appropriate action to assign a vehicle to the transport task. The cost and delay of the new solution are then calculated and sent as a reward to MARL for learning.

[0047] like Figure 4 FIG. 1 is a schematic diagram of a replanning process in one embodiment, which specifically includes the following steps: First, the system receives input parameters, including the vehicle set , transportation tasks that need to be rescheduled , Task Pool and the current scheduling solution . Then, remove the task from the original scheduling solution Related vehicle arrangements, get updated solutions , and add the task to the task pool to be scheduled .

[0048] Then, the system performs a task on each task in the task pool. Try to reinsert it in turn. In each insertion attempt, iterate over the current vehicle , for the task A feasible insertion location is found and the feasibility of the insertion plan is evaluated using an adaptive large neighborhood search or multi-agent reinforcement learning algorithm. If the insertion is feasible, the insertion operation is retained and the corresponding task is removed from the task pool. If it is not feasible, the insertion is withdrawn, and the scheduling solution remains unchanged. This process continues until the task pool is empty or there are no more feasible insertion operations.

[0049] Finally, the output contains all the scheduling solutions after valid insertion And the remaining task pool , achieving efficient rescheduling of infeasible tasks and intelligent reconfiguration of system resources.

[0050] In one embodiment of the present application, when the unexpected event ends, collecting relevant information to learn and train the multi-agent reinforcement learning algorithm includes: constructing a state space of a multi-agent reinforcement learning algorithm based on the global state information and the affected transportation task information; determining a reward function based on the cost objective function and the delay penalty; performing action selection of an action space based on the state space, calculating a reward value based on the action selection and the reward function, and optimizing network parameters of the state space and the action space based on the reward value.

[0051] In an embodiment of the present application, as shown in Figure 5 the state space at each time step includes global state containing all information of the environment on the model running, specifically composed of transportation task information , transportation network information , transportation vehicle information and environment running information .

[0052] contains all information of the transportation task, a total of 8 features representing is the transportation task through node information, is the task departure node information, is the task arrival node information, is the task transfer node information; is the transportation task time window information, is the task departure time window, is the task arrival time window, is the demand of the transportation task.

[0053] contains all information of the transportation route, a total of 6 features representing is the transportation distance from the departure node to the transfer node of each transportation mode (road, rail and water); is the transportation distance from the transfer node to the arrival node of each transportation mode.

[0054] contains all information of the transportation vehicle, which is the vehicle information that each agent can use and its departure and arrival time window.

[0055] contains information of the current running of the environment, respectively composed of the current time step , the location of the unexpected event , the time of the unexpected event .

[0056] Each agent at time step When the global state of the environment Get the local observation state , the policy network then outputs the agent action based on the input observation state.

[0057] Local observation state Will include the current time step , where the accident occurred , time of accident , time window information of the transport task, time window information of the transport vehicle, path information and available transport vehicle information. Most local observation state features will directly use global state information, and the available transport vehicle information will be encoded using one-hot. For path feature information, in most cases, transport tasks have transfers, so the path information of the local observation state is expressed as or , when there is no transfer, only one agent No. 1 can observe the local state , the local observation states of the remaining agents will be assigned null values.

[0058] Action Space , in one time step In the action set Represents the set of all actions that each carrier agent can currently choose. At each time step, , representing the choice of road transport, rail transport and water transport respectively. If a transport solution does not exist, such as the absence of water transport in the actual transport, a virtual transport solution is provided to maximize the cost of adopting the solution.

[0059] Reward Function The goal of multi-agent reinforcement learning is to minimize cost while minimizing delay time, so the reward is set to the negative of the cost. Specifically, if the action mapping solution can complete the transportation task while meeting the transportation order time window, transshipment requirements, and capacity constraints, the reward is assigned to the negative of the cost; otherwise, the reward is assigned to a minimum value.

[0060]

[0061] It is the cost, which consists of transportation cost, transshipment cost, storage cost, waiting cost, carbon tax cost and delay penalty cost. It's the delay time. It is a maximum value, which is used to impose an extremely severe penalty on the delay time, so that the constraints are forced to be satisfied first in the reinforcement learning training.

[0062] The multi-agent deep deterministic policy gradient algorithm framework is adopted to build the training process, including: initializing the Actor-Critic network of each Agent; environment initialization, obtaining the initial solution through ALNS; randomly generating service time disturbance, removing the affected transportation orders and adding them to the order pool; each Agent selects an action according to the state; the environment feedback state transition and reward; experience is stored in the experience pool; periodically sample experience, update the Actor and Critic network parameters; use the soft update mechanism to synchronize the target network. After training, it is deployed in the multimodal transport scheduling system.

[0063] In the above multi-agent reinforcement learning-based multimodal transport path selection method, first, a multimodal transport network and system environment are constructed; then, an adaptive large neighborhood search algorithm is used based on the multimodal transport network and system environment for initial planning; finally, when an unexpected event occurs, the unexpected event information is obtained, and a multi-agent reinforcement learning algorithm is used based on the unexpected event information and a cost objective function to optimize the initial planning to obtain an unexpected action decision, and the transport task is executed based on the unexpected action decision. That is, by constructing a large neighborhood disturbance response mechanism that integrates ALNS, when the node / path service time deviation exceeds the threshold, the path reconstruction is intelligently triggered and a new candidate path is generated, so that the system has the ability of “disturbance self-recovery”; based on the MADDPG multi-agent reinforcement learning framework, each carrier Agent can obtain global information in centralized training and make independent decisions in execution, realizing an efficient path planning mechanism of “centralized training + distributed execution”, effectively improving the utilization efficiency of transport resources and the fairness of task allocation; a hierarchical action design is proposed (first select the transport mode by MARL, and then select the specific carrier by ALNS), which reduces the original action space dimension from to , effectively alleviating the dimension explosion problem; in the disturbance reconstruction phase, ALNS as an external heuristic module provides high-quality candidate paths for MARL, which helps to avoid the MARL model falling into local optimum, and at the same time selects the best “break-repair” path operation sequence through the dynamic operator weight adjustment mechanism. In the path reconstruction phase, the ALNS algorithm disturbs the original path through the “break-repair” mechanism, thereby generating a set of potential high-quality candidate paths. This process not only accelerates the convergence of the training process, but also effectively improves the generalization ability of the model.

[0064] It should be understood that although the steps in the flowcharts involved in the embodiments described above are shown in sequence according to the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the embodiments described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.

[0065] Based on the same inventive concept, the embodiments of the present application also provide a multi-agent reinforcement learning based multimodal transportation path selection device for implementing the above-mentioned multi-agent reinforcement learning based multimodal transportation path selection method. The implementation scheme for solving problems provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more multi-agent reinforcement learning based multimodal transportation path selection device embodiments provided below can refer to the limitations of the multi-agent reinforcement learning based multimodal transportation path selection method in the above text, and will not be repeated here.

[0066] In one embodiment, as shown in Figure 6 A multi-agent reinforcement learning based multimodal transportation path selection device 600 is provided, comprising an environment construction module 601, an initial planning module 603 and a re-planning module 605, wherein: The environment construction module 601 is configured to construct a multimodal transportation network and system environment.

[0067] The initial planning module 603 is configured to perform initial planning based on the multimodal transportation network and system environment using an adaptive large neighborhood search algorithm.

[0068] The re-planning module 605 is configured to, when an unexpected event occurs, acquire unexpected event information, optimize the initial planning based on the unexpected event information in combination with a cost objective function using a multi-agent reinforcement learning algorithm, obtain an unexpected action decision, and execute a transportation task based on the unexpected action decision.

[0069] In an embodiment of the present application, the cost objective function includes transportation cost, transfer cost, storage cost, waiting cost, carbon tax cost and delay penalty cost.

[0070] In an embodiment of the present application, the acquisition of unexpected event information includes: Obtain global state information and affected transportation task information after the adaptive large neighborhood search algorithm planning.

[0071] In an embodiment of the present application, the re-planning module is further configured to: Determine service time constraints based on the affected transportation task information; Optimize the initial plan based on the service time constraints and a cost objective function using a multi-agent reinforcement learning algorithm when an unexpected event occurs; Collect relevant information to learn and train the multi-agent reinforcement learning algorithm when the unexpected event ends.

[0072] In an embodiment of the present application, the re-planning module is further configured to: Construct a state space of a multi-agent reinforcement learning algorithm based on the global state information and the affected transportation task information; Determine a reward function based on the cost objective function and a delay penalty; Perform action selection in an action space based on the state space, calculate a reward value based on the action selection and the reward function, and optimize network parameters of the state space and the action space based on the reward value.

[0073] The above various modules in the multi-modal transport path selection device based on multi-agent reinforcement learning can be all or partially realized by software, hardware, and combinations thereof. The above various modules can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in the computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the above various modules.

[0074] In an embodiment, a computer device, which can be a terminal, is provided, and an internal structure diagram of the computer device can be as shown in Figure 7As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used for wired or wireless communication with external terminals. Wireless mode can be achieved through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement a multi-agent reinforcement learning-based multimodal transport path selection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad provided on the computer device shell, or an external keyboard, touchpad or mouse, etc.

[0075] Those skilled in the art can understand that, Figure 7 The skilled in the art can understand that,

[0076] In one embodiment, a computer device is provided, including a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0077] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in the above method embodiments.

[0078] In one embodiment, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the steps in the above method embodiments.

[0079] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties.

[0080] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0081] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0082] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A multimodal transport path selection method based on multi-agent reinforcement learning, characterized in that: The method comprises: Build a multimodal transport network and system environment; Based on the multimodal transport network and system environment, an adaptive large neighborhood search algorithm is used for initial planning; When an unexpected event occurs, the unexpected event information is obtained, and the initial plan is optimized using a multi-agent reinforcement learning algorithm based on the unexpected event information combined with a cost objective function to obtain an unexpected action decision, and the transportation task is executed based on the unexpected action decision.

2. The multimodal transport route selection method based on multi-agent reinforcement learning according to claim 1, characterized in that: The cost objective function includes transportation cost, transshipment cost, storage cost, waiting cost, carbon tax cost and delay penalty cost.

3. The multimodal transport route selection method based on multi-agent reinforcement learning according to claim 1, characterized in that: The obtaining of accident information includes: Obtain the global state information and affected transportation task information after the adaptive large neighborhood search algorithm planning.

4. The multimodal transport route selection method based on multi-agent reinforcement learning according to claim 3 is characterized in that: The optimizing the initial plan based on the accident information and the cost objective function by using a multi-agent reinforcement learning algorithm includes: determining a service time constraint based on the affected transportation task information; When an unexpected event occurs, the initial plan is optimized using a multi-agent reinforcement learning algorithm based on the service time constraint and the cost objective function; When the accident ends, relevant information is collected to learn and train the multi-agent reinforcement learning algorithm.

5. The multimodal transport route selection method based on multi-agent reinforcement learning according to claim 4 is characterized in that: When the accident ends, collecting relevant information to learn and train the multi-agent reinforcement learning algorithm includes: Constructing a state space of a multi-agent reinforcement learning algorithm based on the global state information and the affected transportation task information; determining a reward function based on the cost objective function and the delay penalty; An action selection is performed in an action space based on the state space, a reward value is calculated based on the action selection and a reward function, and network parameters of the state space and the action space are optimized based on the reward value.

6. A multimodal transport path selection device based on multi-agent reinforcement learning, characterized in that: The device comprises: Environment construction module, used to build multimodal transport network and system environment; An initial planning module, configured to perform initial planning based on the multimodal transport network and system environment using an adaptive large neighborhood search algorithm; The re-planning module is used to obtain accident information when an accident occurs, optimize the initial plan based on the accident information and the cost objective function using a multi-agent reinforcement learning algorithm, obtain an accident action decision, and execute the transportation task based on the accident action decision.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Multimodal transport dynamic path planning method based on game reinforcement learning

    CN113159681A

  • Emergency material transportation optimization scheduling method based on multimodal transport network

    CN117852802A

  • Multi-agent cooperative scheduling method and system based on maximum entropy reinforcement learning

    CN119623933A

  • Dynamic optimization and real-time decision-making method and device for multi-agent collaborative target search

    CN119828460A

  • Multimodal transport path optimization method for deep reinforcement learning

    CN120069723A

Cited By

  • Robot fast path optimization method based on multi-strategy fusion

    CN121163529A

  • Green smart port rating method based on big data

    CN121303976A

  • Dynamic disturbance-oriented emergency material multimodal transport intelligent scheduling method

    CN121638841A

  • An intelligent scheduling method for emergency material multimodal transportation facing dynamic disturbance

    CN121638841B