Electric unmanned motorcade multi-task joint scheduling method based on multi-agent reinforcement learning
Through multi-agent reinforcement learning methods, the complex coupling problems of charging, order response and vehicle rebalancing in electric driverless fleets were solved, multi-task collaborative optimization was achieved, and system efficiency and service quality were improved.
Patent Information
- Application Number
- CN202510686171.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-12
AI Technical Summary
Existing scheduling methods for electric driverless fleets have difficulty handling the complex coupling relationship between charging, order response, and vehicle rebalancing. They lack multi-task collaborative modeling and decision-making mechanisms, insufficient task value assessment, and a lack of coordination mechanisms among multiple agents, resulting in low resource utilization and scheduling conflicts.
A multi-agent reinforcement learning-based method is adopted to establish a multi-task collaborative optimization framework through predicting travel demand, state construction, task candidate generation, reward function design, value function generation and strategy improvement. It integrates local immediate rewards and global advantage rewards, and combines the Actor-Critic and KM algorithms to achieve the optimal allocation of tasks and vehicles.
It has achieved multi-task collaborative optimization of charging, order scheduling and relocation in electric unmanned vehicle fleets, improved the overall efficiency and service quality of the system, solved the problems of low resource utilization and scheduling conflicts in traditional methods, and significantly improved the overall performance of the fleet.
Smart Images

Figure CN120634099A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a vehicle scheduling method, and in particular to a multi-task joint scheduling method for an electric unmanned vehicle fleet based on multi-agent reinforcement learning. Background Art
[0002] With the development of autonomous driving technology, electrification technology, and shared mobility models, urban transportation systems are undergoing profound changes. As a key form of future smart mobility, electric autonomous vehicle fleets (EAVs) offer advantages such as high efficiency, environmental friendliness, and intelligent scheduling, making them a crucial support for green urban mobility. However, compared to traditional fuel-powered fleets, EAVs face new challenges in scheduling and management, particularly due to limited battery capacity and the need for frequent recharging, which complicates operational scheduling.
[0003] Current research on scheduling optimization for electric ride-hailing or autonomous vehicle fleets primarily focuses on order scheduling and vehicle repositioning (rebalancing), with some studies also incorporating charging scheduling strategies. However, these three tasks are not independent in actual operations: charging behavior affects the number of serviceable vehicles, order response impacts power consumption and subsequent charging requirements, and the spatial distribution of vehicles directly determines the efficiency of order acceptance and recharging. Consequently, these three tasks are intricately and dynamically coupled, with strong interdependence in their decision-making.
[0004] Existing technical methods mostly use step-by-step or heuristic scheduling strategies, which are difficult to handle the collaborative optimization problem between multiple tasks. For example:
[0005] The patent, "A Method and System for Dispatching Online Rides," only optimizes for increasing the success rate of order acceptance, but does not consider vehicle energy consumption or charging requirements;
[0006] The patent, "A Method for Planning Charging and Battery Swapping Routes for Electric Rides," optimizes charging routes but fails to consider them in conjunction with order responses or vehicle distribution, resulting in limited robustness.
[0007] The patent "Electric Taxi Charging Navigation Path Planning Method Based on Deep Reinforcement Learning" takes into account battery loss and power constraints, improving the robustness of the charging path strategy, but does not consider scheduling and rebalancing;
[0008] The patent "A method for dispatching online ride-hailing vehicles based on hybrid hierarchical reinforcement learning" introduces high- and low-level strategies to collaboratively dispatch orders, does not involve charging tasks, and cannot achieve overall scheduling coordination.
[0009] In addition, some papers attempt to introduce reinforcement learning methods to improve the optimization of single tasks. For example, "Research Progress on Electric Vehicle Charging Scheduling Algorithms Based on Deep Reinforcement Learning" focuses on charging path planning, while "An Integrated Optimization-Simulation Framework for Scalable Smart Charging and Relocation of Shared Autonomous Electric Vehicles" optimizes charging and vehicle deployment at the system level. However, these papers generally lack a systematic design for jointly modeling and optimizing the three tasks of charging, scheduling, and rebalancing. "Multi-task dispatch of shared autonomous electric vehicles for Mobility-on-Demand services - combination of deep reinforcement learning and combinatorial optimization method" delves into tasks such as dispatching, relocation, and charging, but does not explore the coupling relationship between the three.
[0010] The existing technologies mainly have the following defects and challenges:
[0011] Lack of multi-task collaborative modeling and decision-making mechanism: Current methods usually only consider the binary collaboration between scheduling-charging or scheduling-rebalancing, and are unable to handle the dynamic coupling and resource competition between multiple tasks, making it difficult to achieve global optimization of the fleet.
[0012] Insufficient task value assessment makes it difficult to guide collaborative decision-making: The three tasks each have different time sensitivities and benefit structures, and existing methods make it difficult to measure the impact of current behavior on the overall benefits of the future system.
[0013] Lack of coordination mechanism among multiple agents: In a large-scale convoy environment, the decisions of each vehicle as an independent agent are often conflicting and redundant, and there is a lack of effective local-global coordination mechanism, which affects the overall performance.
[0014] In summary, the current operational scheduling methods for electric driverless fleets are unable to cope with the complex interactions and collaborative optimization problems between charging, order response, and vehicle rebalancing. There is an urgent need for a multi-task, multi-agent, reinforcement learning-driven joint scheduling method to improve the overall system efficiency and service quality. Summary of the Invention
[0015] Purpose of the invention: The purpose of the present invention is to provide a multi-task joint scheduling method for electric unmanned vehicle fleets based on multi-agent reinforcement learning.
[0016] Technical Solution: The multi-task joint scheduling method for an electric unmanned vehicle fleet based on multi-agent reinforcement learning described in the present invention includes the following steps:
[0017] (1) Forecast travel demand;
[0018] (2) State construction: At each decision moment, the system constructs a Markov state, which includes the following dimensions: vehicle-related state, regional supply and demand state, and charging station state;
[0019] (3) Task candidate generation: generation of charging candidates, dispatching candidates, and scheduling candidates;
[0020] (4) Reward function design: including local immediate reward design and global reward construction;
[0021] (5) Value function generation;
[0022] (6) Strategy improvement and Actor update;
[0023] (7)Environmental execution and iterative training.
[0024] Furthermore, the prediction method in step (1) includes: a graph structure model with urban areas as basic units, each grid node contains historical order volume and time characteristics as input features, and the output is the order demand estimate of each area in the next 15 minutes. The result will be used as the input of the vehicle scheduling strategy to guide the global optimization of charging and relocation decisions. After obtaining the prediction results, the iterative update formula of the vehicle status is:
[0025]
[0026] in, is the estimated number of vehicles in region g at time slice t-1; is the actual number of completed orders for grid g in time slice t-1 (from historical data); STGNN(g, t) is the prediction value of the spatiotemporal graph neural network for the number of new vehicle arrivals in region g in time slice t; is the actual number of arriving vehicles based on real-time scheduling feedback; α and β are learnable parameters (constraint: α + β = 1) used to balance prediction and real-time data.
[0027] Furthermore, in step (2), at each decision moment t, the system constructs a Markov state s t , contains the following dimensions, including the following dimensions:
[0028] (1) Vehicle-related state (for each vehicle i∈{1,...,N}):
[0029] State of charge (normalized):
[0030] Location (discrete code):
[0031] Current status indicator (idle / in service / scheduling):
[0032] (2) Regional supply and demand status (for each region) ):
[0033] Current number of orders:
[0034] Historical average number of orders:
[0035] Demand pressure:
[0036] (3) Charging station status:
[0037] The number of idle charging piles at the cth charging station:
[0038] Total number of charging piles: A c
[0039] Idle ratio:
[0040] The final state vector is expressed as:
[0041]
[0042] Furthermore, step (3) generates three candidate tasks for each idle or low-battery vehicle i:
[0043] (1) Charging candidates: collected from accessible charging stations Select the nearest station:
[0044]
[0045] (2) Order Candidates: Match the nearest orders to be served from the order set The vehicle's power level must meet the service conditions:
[0046]
[0047] (3) Scheduling candidate: Select the target area h with higher demand pressure as the deployment target:
[0048]
[0049] Generate candidate sets:
[0050]
[0051] Furthermore, in step (4):
[0052] (1) Local instant reward design:
[0053] Local rewards are targeted at the atomic tasks (such as order acceptance, charging, and dispatching) performed by each vehicle at a certain moment, with immediate feedback as the main focus. They are designed to be task-specific:
[0054] 1) Dispatch Rewards
[0055] When vehicle i successfully receives order o, its local reward is defined as:
[0056]
[0057] R o :Order base income
[0058] Empty distance to pick up passengers
[0059] Order waiting time
[0060] λ1, λ2: Hyperparameters that control the weight of operating costs
[0061] 2) Charge Reward
[0062] Vehicle i goes to charging station c to recharge:
[0063]
[0064] The time it takes for the vehicle to travel to the charging station
[0065] Waiting time
[0066] Δsoc i : Increase in SOC
[0067] λ3, λ4, λ5: coefficients (e.g., bonus charging efficiency)
[0068] 3) Regional Relocation Rewards (Relocate)
[0069] Vehicle i moves from its current location to the target area h:
[0070]
[0071] Dispatch distance
[0072] Current demand intensity in the target area
[0073] λ6, λ7: coefficients that weigh distance cost and area value
[0074] (2) Construction of Global Reward
[0075] The global reward is not targeted at individual vehicles, but is used to measure the scheduling quality of the entire system. It is calculated once per time step or cycle and is used to stabilize learning objectives or reward attribution during training:
[0076] 1) Response rate indicator (order completion rate)
[0077]
[0078] The number of orders completed in the current time step The number of orders generated in the current time step
[0079] 2) Empty driving rate indicator
[0080]
[0081] Empty distance
[0082] Total distance including service and dispatch
[0083] 3) Vehicle distribution balance
[0084] Assume that the standard deviation of the supply-demand ratio in each region is:
[0085]
[0086] Define rewards as:
[0087]
[0088] Furthermore, in step (5)
[0089] (1) Critic network estimated state value function:
[0090] V φ (s t )=Critic(s t ;φ)
[0091] (2) Estimating the action-value function (single-step approximation):
[0092]
[0093] (3) Advantage function calculation:
[0094]
[0095] (4) Construct KM input matrix:
[0096] Constructing the weight matrix The element is M i,j =A(s t , a ij ), input KM algorithm to get the optimal allocation π * .
[0097] Furthermore, in step (6), (1) the local and global rewards are fused into a reward function, which is used during training:
[0098] where λ local +λ global =1
[0099] (2) Actor strategy loss, using KM output as behavior supervision:
[0100]
[0101] The KM output is not used directly to execute actions, but is used as a policy improvement target to guide the policy distribution learning to move closer to the direction of maximum advantage.
[0102] (3) Critic value loss:
[0103]
[0104] Parameter update:
[0105]
[0106] (4) Hierarchical Hybrid Critic and Partitioned Actor Learning Framework:
[0107] Hierarchical hybrid critic refers to deploying a set of regional critic networks in the central and local areas, and a global critic network deployed in the central area. Responsible for modeling the long-term value function at the city level, whose input is the supply and demand matrix containing the entire city grid Charging station load status and the global state s with timestamp τ global , output the predicted discounted cumulative reward The global reward is distributed in the regional critic network of edge computing nodes through weighted fusion of business indicators. Focuses on geographic sub-regions, where the input is the local supply and demand status and neighboring charging station information by minimizing the regional level timing difference error Optimize short-term scheduling efficiency, including is the sum of the local rewards of vehicles in the region. The global and regional critics achieve collaboration through periodic parameter aggregation (e.g., synchronization every 5 minutes), and the update rule is Where β is a trade-off coefficient that ensures the compatibility of city-level goals with regional equilibrium.
[0108] Furthermore, in step (7), the system submits the action output by the Actor to the environment for execution, the environment advances to t+Δt, and collects new states and rewards for iterative optimization.
[0109] Compared with the prior art, the present invention has the following beneficial effects:
[0110] (1) Multi-task collaborative optimization mechanism for power replenishment, order scheduling, and relocation
[0111] The present invention innovatively treats the three tasks of recharging, order scheduling, and relocation in the operation of electric online ride-hailing vehicles as a joint optimization problem, and designs a unified decision-making framework for collaborative processing. Traditional methods often only focus on one or two tasks, ignoring the coupling relationship between tasks, which can easily lead to problems such as low resource utilization and scheduling conflicts. By establishing a linkage mechanism between multiple tasks, the present invention can dynamically allocate the optimal task type based on real-time vehicle conditions, order distribution, and power status, thereby maximizing overall revenue and ensuring service continuity.
[0112] (2) Task allocation strategy based on the combination of Actor-Critic and KM algorithms
[0113] To address the complex multi-vehicle, multi-task scheduling problem, this paper designs a task value learning mechanism based on the actor-critic architecture of reinforcement learning, combined with the KM algorithm to achieve optimal allocation of tasks and vehicles. The actor-critic algorithm continuously optimizes the value assessment of various tasks under different states, while the KM algorithm ensures global optimality in task matching. This combination enables the system to both learn and make optimal task allocation decisions in real time, significantly improving scheduling efficiency and fleet profitability.
[0114] (3) Multi-agent joint learning framework integrating local and global rewards
[0115] This paper constructs a multi-agent reinforcement learning framework that integrates local immediate rewards with global advantage rewards. This framework guides each vehicle to optimize its own short-term interests while also considering the overall benefits of the system. This mechanism addresses the performance degradation caused by the "independent" nature of traditional reinforcement learning, effectively improving the level of collaboration among agents. In large-scale fleet environments, this approach can maintain stable policy convergence and scheduling coordination, improving overall service quality and responsiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0116] Figure 1 Flow chart of the method of the present invention;
[0117] Figure 2 Schematic diagram of the control device structure of the present invention. DETAILED DESCRIPTION
[0118] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be further described below.
[0119] This example uses 300 electric online taxis in a city as the target, builds a "charging-dispatching-scheduling" joint dispatching system, and deploys it on the cloud dispatching platform. The overall system structure is as follows: Figure 2 As shown, the scheduling control process is as follows Figure 1 shown.
[0120] 1. Forecasting travel demand
[0121] This graph model uses urban areas as the basic unit. Each grid node contains historical order volume and time characteristics as input features. The output is an estimate of order demand in each area within the next 15 minutes. This result will serve as input to the vehicle scheduling strategy to guide the global optimization of charging and relocation decisions. After obtaining the prediction results, the iterative update formula of the vehicle status is:
[0122]
[0123] in, is the estimated number of vehicles in region g at time slice t-1; is the actual number of completed orders for grid g in time slice t-1 (from historical data); STGNN(g, t) is the prediction value of the spatiotemporal graph neural network for the number of new vehicle arrivals in region g in time slice t; is the actual number of arriving vehicles based on real-time scheduling feedback; α and β are learnable parameters (constraint: α + β = 1) used to balance prediction and real-time data.
[0124] (2) State Construction
[0125] At each decision moment t, the system constructs a Markov state s t , which contains the following dimensions:
[0126] (1) Vehicle-related state (for each vehicle i∈{1,...,N}):
[0127] State of charge (normalized):
[0128] Location (discrete code):
[0129] Current status indicator (idle / in service / scheduling):
[0130] (2) Regional supply and demand status (for each region) ):
[0131] Current number of orders:
[0132] Historical average number of orders:
[0133] Demand pressure:
[0134] (3) Charging station status:
[0135] The number of idle charging piles at the cth charging station:
[0136] Total number of charging piles: A c
[0137] Idle ratio:
[0138] The final state vector is expressed as:
[0139]
[0140] (3) Task candidate generation
[0141] For each idle or low-battery vehicle i, three candidate tasks are generated:
[0142] (1) Charging candidates: collected from accessible charging stations Select the nearest station:
[0143]
[0144] (2) Order Candidates: Match the nearest orders to be served from the order set The vehicle's power level must meet the service conditions:
[0145]
[0146] (3) Scheduling candidate: Select the target area h with higher demand pressure as the deployment target:
[0147]
[0148] Generate candidate sets:
[0149]
[0150] (4) Reward Function Design
[0151] (1) Local instant reward design:
[0152] Local rewards are targeted at the atomic tasks (such as order acceptance, charging, and dispatching) performed by each vehicle at a certain moment, with immediate feedback as the main focus. They are designed to be task-specific:
[0153] 1) Dispatch Rewards
[0154] When vehicle i successfully receives order o, its local reward is defined as:
[0155]
[0156] R o :Order base income
[0157] Empty distance to pick up passengers
[0158] Order waiting time
[0159] λ1, λ2: Hyperparameters that control the weight of operating costs
[0160] 2) Charge Reward
[0161] Vehicle i goes to charging station c to recharge:
[0162]
[0163] The time it takes for the vehicle to travel to the charging station
[0164] Waiting time
[0165] Δsoc i : Increase in SOC
[0166] λ3, λ4, λ5: coefficients (e.g., bonus charging efficiency)
[0167] 3) Regional Relocation Rewards (Relocate)
[0168] Vehicle i moves from its current location to the target area h:
[0169]
[0170] Dispatch distance
[0171] Current demand intensity in the target area
[0172] λ6, λ7: coefficients that weigh distance cost and area value
[0173] (2) Construction of Global Reward
[0174] The global reward is not targeted at individual vehicles, but is used to measure the scheduling quality of the entire system. It is calculated once per time step or cycle and is used to stabilize learning objectives or reward attribution during training:
[0175] 1) Response rate indicator (order completion rate)
[0176]
[0177] The number of orders completed in the current time step The number of orders generated in the current time step
[0178] 2) Empty driving rate indicator
[0179]
[0180] Empty distance
[0181] Total distance including service and dispatch
[0182] 3) Vehicle distribution balance
[0183] Assume that the standard deviation of the supply-demand ratio in each region is:
[0184]
[0185] Define rewards as:
[0186]
[0187] (V) Value Function Generation
[0188] (1) Critic network estimated state value function:
[0189] V φ (s t )=Critic(s t ;φ)
[0190] (2) Estimating the action-value function (single-step approximation):
[0191]
[0192] (3) Advantage function calculation:
[0193]
[0194] (4) Construct KM input matrix:
[0195] Constructing the weight matrix The element is M i,j =A(s t , a ij ), input KM algorithm to get the optimal allocation π * .
[0196] (6) Strategy Improvement and Actor Update
[0197] (1) Fusion of local and global rewards
[0198] Fusion reward function (used during training):
[0199] where λ local +λ global =1
[0200] (2) Actor strategy loss (using KM output as behavior supervision):
[0201]
[0202] The KM output is not used directly to execute actions, but is used as a policy improvement target to guide the policy distribution learning to move closer to the direction of maximum advantage.
[0203] (3) Critic value loss:
[0204]
[0205] Parameter update:
[0206]
[0207] (4) Hierarchical Hybrid Critic and Partitioned Actor Learning Framework
[0208] In the multi-agent reinforcement learning framework for large-scale urban electric vehicle fleet scheduling, the design of hierarchical hybrid critics and partitioned actors is the key to achieving efficient real-time decision-making. They achieve the decoupling of global goals and local optimization. Hierarchical hybrid critics refer to the deployment of a set of regional critic networks in the central and local areas. The global critic network deployed in the central Responsible for modeling the long-term value function at the city level, whose input is the supply and demand matrix containing the entire city grid Charging station load status and the global state s with timestamp τ global , output the predicted discounted cumulative reward The global reward is achieved through weighted fusion of business indicators. The regional critic network distributed in the edge computing nodes Focuses on geographic sub-regions, where the input is the local supply and demand status and neighboring charging station information by minimizing the regional level timing difference error Optimize short-term scheduling efficiency, including is the sum of the local rewards of vehicles in the region. The global and regional critics achieve collaboration through periodic parameter aggregation (e.g., synchronization every 5 minutes), and the update rule is Where β is a trade-off coefficient that ensures the compatibility of city-level goals with regional equilibrium.
[0209] (7) Environmental execution and iterative training;
[0210] The system submits the actions output by the actor to the environment (such as SUMO) for execution. The environment advances to t+Δt, collects new states and rewards, and iterates optimization. The resulting policy network can be directly called in the actual environment. By inputting the environment state information, the action to be executed is obtained.
[0211] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.
Claims
1. A multi-task joint scheduling method for electric unmanned vehicle fleet based on multi-agent reinforcement learning, characterized by: The steps include: (1) Forecast travel demand; (2) State construction: At each decision moment, the system constructs a Markov state, which includes the following dimensions: vehicle-related state, regional supply and demand state, and charging station state; (3) Task candidate generation: generation of charging candidates, dispatching candidates, and scheduling candidates; (4) Reward function design: including local immediate reward design and global reward construction; (5) Value function generation; (6) Strategy improvement and Actor update; (7)Environment execution and iterative training.
2. The multi-task joint scheduling method for electric unmanned vehicle fleet based on multi-agent reinforcement learning according to claim 1 is characterized in that: Step (1): A graph structure model with urban areas as basic units, each grid node contains historical order volume and time characteristics as input features, and the output is the order demand estimate of each area in the next 15 minutes. This result will serve as the input of the vehicle scheduling strategy to guide the global optimization of charging and relocation decisions. After obtaining the prediction results, the iterative update formula of the vehicle status is: in, is the estimated number of vehicles in region g at time slice t-1; is the actual number of completed orders for grid g in time slice t-1; STGNN(g, t) is the predicted value of the spatiotemporal graph neural network for the number of new vehicle arrivals in region g in time slice t; is the actual number of arriving vehicles based on real-time scheduling feedback; αβ is a learnable parameter with the constraint of α+β=1, which is used to balance prediction and real-time data.
3. The multi-task joint scheduling method for electric unmanned vehicle fleet based on multi-agent reinforcement learning according to claim 1 is characterized in that: In step (2), at each decision moment t, the system constructs a Markov state s t , contains the following dimensions, including the following dimensions: (1) Vehicle-related state (for each vehicle i∈{1,...,N}): State of charge (normalized): Location (discrete code): Current status indicator (idle / in service / scheduling): (2) Regional supply and demand status (for each region) ): Current number of orders: Historical average number of orders: Demand pressure: (3) Charging station status: The number of idle charging piles at the cth charging station: Total number of charging piles: A c Idle ratio: The final state vector is expressed as:
4. The multi-task joint scheduling method for electric unmanned vehicle fleet based on multi-agent reinforcement learning according to claim 1 is characterized in that: Step (3) generates three candidate tasks for each idle or low-battery vehicle i: (1) Charging candidates: collected from accessible charging stations Select the nearest station: (2) Order Candidates: Match the nearest orders to be served from the order set The vehicle's power level must meet the service conditions: (3) Scheduling candidate: Select the target area h with higher demand pressure as the deployment target: Generate candidate sets:
5. The multi-task joint scheduling method for electric unmanned vehicle fleet based on multi-agent reinforcement learning according to claim 1 is characterized in that: In the step (4): (1) Local instant reward design: Local rewards are targeted at the atomic tasks performed by each vehicle at a certain moment, with immediate feedback as the main focus, and are designed to be task-specific: 1) Dispatch When vehicle i successfully receives order o, its local reward is defined as: R o :Order base income Empty distance to pick up passengers Order waiting time λ1, λ2: Hyperparameters that control the weight of operating costs 2) Charge Vehicle i goes to charging station c to recharge: The time it takes for the vehicle to travel to the charging station Waiting time Δsoc i : Increase in SOC λ3, λ4, λ5: coefficients, reward charging efficiency 3) Regional Dispatch Rewards Relocate Vehicle i moves from its current location to the target area h: Dispatch distance Current demand intensity in the target area λ6, λ7: coefficients that weigh distance cost and area value (2) Construction of Global Reward The global reward is not targeted at individual vehicles, but is used to measure the scheduling quality of the entire system. It is calculated once per time step or cycle and is used to stabilize learning objectives or reward attribution during training: 1) Response rate indicator, order completion rate The number of orders completed in the current time step The number of orders generated in the current time step 2) Empty driving rate indicator Empty distance Total distance including service and dispatch 3) Vehicle distribution balance Assume that the standard deviation of the supply-demand ratio in each region is: Define rewards as:
6. The multi-task joint scheduling method for electric unmanned vehicle fleet based on multi-agent reinforcement learning according to claim 1 is characterized in that: In the step (5) (1) Critic network estimated state value function: V φ (s t )=Critic(s t ;φ) (2) Estimating the action-value function (single-step approximation): (3) Advantage function calculation: (4) Construct KM input matrix: Constructing the weight matrix The element is M i,j =A(s t , a ij ), input KM algorithm to get the optimal allocation π * .
7. The multi-task joint scheduling method for electric unmanned vehicle fleet based on multi-agent reinforcement learning according to claim 1 is characterized in that: In the step (6): (1) Fusion of local and global rewards Fusion reward function, used during training: among them local +λ global =1 (2) Actor strategy loss, using KM output as behavior supervision: KM output is not used directly to execute actions, but is used as a strategy improvement target to guide the strategy distribution learning to move closer to the direction of maximum advantage. (3) Critic value loss: Parameter update: (4) Hierarchical Hybrid Critic and Partitioned Actor Learning Framework: Hierarchical hybrid critic refers to deploying a set of regional critic networks in the central and local areas, and a global critic network deployed in the central area. Responsible for modeling the long-term value function at the city level, whose input is the supply and demand matrix containing the entire city grid Charging station load status and the global state s with timestamp τ global , output the predicted discounted cumulative reward The global reward is distributed in the regional critic network of edge computing nodes through weighted fusion of business indicators. Focuses on geographic sub-regions, where the input is the local supply and demand status and neighboring charging station information by minimizing the regional level timing difference error Optimize short-term scheduling efficiency, including is the sum of the local rewards of vehicles in the region. The global and regional critics achieve collaboration through periodic parameter aggregation. The update rule is: Where β is a trade-off coefficient that ensures the compatibility of city-level goals with regional equilibrium.
8. The multi-task joint scheduling method for electric unmanned vehicle fleet based on multi-agent reinforcement learning according to claim 1 is characterized in that: In step (7), the system submits the action output by the Actor to the environment for execution, and the environment advances to t+Δt, and collects new states and rewards for iterative optimization.
9. A control device for a multi-task joint scheduling method for an electric unmanned vehicle fleet based on multi-agent reinforcement learning according to any one of claims 1 to 8, characterized in that: It includes five main modules, each of which interacts with information through the data bus: (1) State perception module (100) This module receives data from the simulation environment and outputs a standardized state tensor as input to the policy module; (2) Joint Strategy Module (200) Adopting the Actor-Critic reinforcement learning structure, combining high-level decision-making task type selection with low-level execution target location selection, a set of action instructions task and target are generated for each vehicle. This module supports multi-agent parallel learning. (3) Prediction module (300) The time series prediction network trained with historical order data outputs the regional order intensity changes in the next T steps, assisting the joint strategy module in more accurately evaluating future value. (4) Task execution module (400) Distribute policy instructions to each vehicle in the SUMO simulation platform, control them to perform corresponding operations, including route planning, order docking, charging operations, etc., and provide feedback on actual execution results; (5) Reward Feedback and Learning Module (500) Combining environmental feedback with vehicle execution results, multi-granularity reward values are calculated, locally / globally, and used to update the neural network of the joint policy module, gradually improving policy performance; in: The state perception module (100) inputs the environmental state into the joint strategy module (200) and the prediction module (300); The forecasting module (300) provides future demand estimates to the joint strategy module (200); The joint strategy module (200) outputs the strategy to the task execution module (400); The task execution module (400) feeds back the execution result to the reward module (500); The reward module (500) optimizes the parameters of the strategy module (200).
Citation Information
Cited By
Anti-interference control method and system for Mecanum wheel omnidirectional robot
CN121043160A
Automatic driving taxi dynamic scheduling system for mixed traffic flow and collaborative decision-making method
CN121481816A
Control method and system under combined transportation scene of steelmaking travelling crane and trolley
CN121500840A
Method and device for joint dispatching and charging of unmanned taxi team
CN121638843A
Mobile crowd sensing fair unit selection method based on multi-agent reinforcement learning
CN121723855A