Mobile vehicle charging and storage dynamic scheduling method and system based on reinforcement learning

By constructing a 3D grid map and using a multi-agent algorithm based on deep reinforcement learning, the problems of poor dynamic environment adaptability and low multi-vehicle coordination efficiency in parking lot charging scheduling are solved. This enables the economical use of peak-valley electricity prices and enhances the grid interaction capability, providing efficient charging services.

CN120930964APending Publication Date: 2025-11-11SHANGHAI TONGYI TECH DEV CO LTD

Patent Information

Application Number
CN202510474635.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing parking lot charging scheduling technologies suffer from poor adaptability to dynamic environments, low efficiency of multi-vehicle coordination, insufficient economic utilization of peak-valley electricity pricing, and weak grid interaction capabilities.

Method used

By constructing a 3D grid map based on UWB positioning system and LiDAR point cloud data, and combining deep reinforcement learning and multi-agent algorithms, dynamic environment modeling and multi-source data fusion are achieved. A multi-objective decision-making model is designed, and a multi-agent deep deterministic policy gradient algorithm and an improved contract net protocol are adopted to achieve efficient task allocation and conflict resolution.

Benefits of technology

It improves the efficiency of multi-vehicle coordination, optimizes the economic utilization of peak-valley electricity pricing, enhances grid interaction capabilities, and enables efficient and flexible charging services within parking lots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930964A_ABST
    Figure CN120930964A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based mobile vehicle charging and storage dynamic scheduling method and system, and solves the problems of insufficient scheduling flexibility and low peak-valley electricity price utilization rate of a fixed charging facility of an existing parking lot. A dynamic environment model is constructed, the real-time SOC of the mobile charging and storage vehicle, the position topological relation and the charging demand space-time distribution are integrated, and a deep reinforcement learning algorithm is adopted to train an intelligent body to generate a multi-dimensional collaborative optimization strategy. According to the method, a charging / discharging time sequence, a task path and energy distribution are autonomously planned, a reward function mechanism fusing dynamic path cost and energy constraint is innovatively designed, a multi-vehicle asynchronous collaborative decision framework is established, and dual targets of charging demand response efficiency and operation cost optimization are achieved. According to the method, an MCSV hardware embedded system which supports an ROS2 communication protocol and has a real-time sensor data processing capability is deployed, so that effective transition from a theoretical strategy to actual application is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent charging scheduling and distributed energy management technology, specifically involving a mobile charging and storage vehicle dynamic scheduling method and system based on reinforcement learning, which is particularly suitable for optimizing the charging demand of new energy vehicles in highly dynamic parking lot scenarios. Background Technology

[0002] With the increasing popularity of new energy vehicles, the demand for charging in parking lots is growing daily. Traditional charging methods typically require vehicles to be fixed at charging station locations, which not only limits parking flexibility but can also lead to a shortage of charging stations. To address this issue, mobile charging technology has emerged. Mobile charging vehicles, as a new type of charging equipment, can autonomously move to the location of the vehicle for charging, greatly improving charging convenience.

[0003] 1. Patent CN112909928B proposes a wireless sensor network mobile charging vehicle path planning and charging scheduling algorithm. By decomposing energy redistribution, traveling salesman problem and time optimization sub-problems, a scheduling sequence is generated using linear programming and a greedy algorithm. The shortcomings of this method are: (1) It does not integrate dynamic electricity price signals. The scheduling strategy only aims to minimize energy loss and time span. It lacks an economic response to the difference between peak and valley electricity prices and cannot achieve operational cost optimization through valley electricity storage and peak electricity discharge; (2) It lacks a multi-vehicle coordination mechanism. The scheme is designed for a single mobile charging vehicle and does not consider path conflicts, task overlap and energy replenishment coordination problems in multi-vehicle scenarios, resulting in limited resource utilization in actual deployment; (3) It lacks environmental adaptability. The static wireless sensor network model is difficult to directly transfer to the dynamic parking lot scenario. It does not model the randomness of new energy vehicle charging demand and the spatial heterogeneity of parking space distribution.

[0004] 2. Patent CN115759505B discloses a task-oriented multi-mobile charging vehicle scheduling method. It maximizes the monitoring utility by constructing a charging utility function and an auxiliary graph model, combined with greedy selection and traveling salesman algorithm. Its technical defects include: (1) The economic model is simple. Although the charging energy consumption cost is introduced, the linkage mechanism between electricity price fluctuation and energy storage charging and discharging revenue is not designed, so peak-valley arbitrage cannot be realized; (2) The dynamic constraint processing is insufficient. The fixed candidate charging location discretization strategy is adopted. The vehicle SOC status and charging demand changes are not updated in real time, which can easily lead to the scheduling strategy being lagging behind; (3) The collaborative optimization is limited. The multi-vehicle scheduling takes utility maximization as the core. It does not integrate the real-time traffic topology and charging station load balancing in the path planning, which may cause charging congestion in local areas. Summary of the Invention

[0005] This invention provides a mobile charging and storage vehicle scheduling method based on reinforcement learning, aiming to solve the problems of poor dynamic environment adaptability, low multi-vehicle coordination efficiency, insufficient economic utilization of peak and valley electricity prices, and weak grid interaction capability in existing parking lot charging scheduling technologies.

[0006] To achieve the above objectives, the specific technical solution of the present invention is as follows: S1. Based on UWB positioning system and LiDAR point cloud data, construct a 3D grid map of the parking lot, perform dynamic environment modeling and multi-source data fusion, and set its grid resolution to 0.5m×0.5m; The path accessibility status is dynamically updated, and the grid cells containing temporary obstacles (such as vehicles blocking the road) are marked as impassable. The detour path weight matrix is ​​calculated.

[0007] S2. Obtain real-time data of vehicles requesting charging via V2X communication, including: location coordinates (x, y, z), accuracy ±10cm; remaining battery charge (SOC); and parking space topology.

[0008] Generate the spatiotemporal distribution matrix of charging demand: ; in The urgency of the i-th demand at time t is represented by the following formula: ; S3. Establish a multi-objective decision-making model based on deep reinforcement learning: S3.1 Define the state space: Mobile charging and storage vehicle status: SOC (normalized to [0,1]), representing the current available energy and energy storage capacity; current location coordinates; current task queue length; Environment: Charging demand matrix Q(t); Topological distance matrix D, which is a weighted adjacency matrix generated based on the location of vehicles with charging demand. The weights are the weighted sum of path distance and dynamic obstacle avoidance cost. S3.2 defines the action space as a multidimensional continuous-discrete hybrid space: Continuous motion control determines charging and discharging power, outputting a ternary set (charging power percentage). Discharge power ratio Standby flag ), satisfying constraints and hour Path planning and decision-making: Based on the reward function mechanism, output the coordinates (x', y') of the next target node and the path curvature parameters. Energy allocation decision: This involves allocating the current State of Charge (SOC) to charging services in a proportional manner for continuous actions, and outputting a power allocation ratio vector for each charging request in the current task. ,satisfy .

[0009] S4. The total reward function is set based on the path efficiency reward, peak-valley arbitrage benefit, and SOC balance penalty.

[0010] Path efficiency bonus: calculated based on task completion time cost, using the following formula. ,in This is the difference between the actual arrival time and the expected arrival time. This is the adjustment coefficient; Peak-valley arbitrage revenue item: Calculate the difference between discharge revenue and charging cost. , Electricity price sensitive factor; SOC Balancing Penalty: Imposing a negative reward for behaviors where SOC falls below a safe threshold (e.g., 20%) or energy allocation exceeds limits. The formula is as follows: , This is the penalty coefficient.

[0011] Task urgency reward: A positive reward is applied to the action based on the task urgency coefficient and the completion effect weight. The formula is as follows: ; The constraints include: discharge amount ≤ current SOC capacity, charging amount ≤ remaining energy storage space; the time cost of adjacent nodes in path planning < the remaining waiting threshold of demanding vehicles. The constraints are embedded into the reinforcement learning objective function using the Lagrange multiplier method.

[0012] S5. A multi-agent deep deterministic strategy gradient algorithm is adopted, combined with an improved contract network protocol and dynamic weight allocation strategy. Each charging and storage vehicle is an independent agent, realizing efficient task allocation and conflict resolution among multiple mobile charging and storage vehicles.

[0013] A two-layer network structure is adopted: Actor network: a 3-layer fully connected neural network (256-128-64 nodes), the input is the local observation state, i.e., SOC value, current position, and task queue, and the output is a continuous action vector, i.e., charging and discharging power, travel direction angle, and energy allocation ratio. The activation function is ReLU, and the output layer uses the Tanh function to constrain the action range; Critic network: a 4-layer fully connected neural network (256-128-64-32 nodes), the input is global state information, i.e., all charging and storage vehicle positions, task distribution, and grid load, and the output is Q-value evaluation. L2 regularization is used to prevent overfitting. Set up an independent experience replay pool (capacity) Local sampling by each agent Empirical data (state, action, reward, next state) is used to asynchronously update network parameters through a central coordinator, addressing the issue of environmental non-stationarity. Simultaneously, relevant key training parameters are appropriately set.

[0014] S5.1 Task Allocation Strategy: Adopting an improved contract network protocol, task packets are broadcast through the dispatch center, containing structured data such as demand location coordinates, power demand, and urgency.

[0015] S5.2 Bidding Weight Calculation: Each charging and storage vehicle calculates its task weight based on its own power level, location, load, and other status.

[0016] ; in, The current battery status of charging vehicle j; The Euclidean distance between the charging vehicle and the target point; The task queue load rate for charging vehicle j (total current task time / maximum carrying time); To replenish the remaining available energy of vehicle j; This is the adjustment coefficient.

[0017] S5.3 Dynamic Bidding Decision: Select the charging / storage vehicle with the highest weight as the winner; if the difference ΔW between the second-highest and highest weights is less than 10, trigger the combined task allocation mode: split the task into two sub-tasks, charging service and route navigation, which are completed collaboratively by different charging / storage vehicles; if all charging / storage vehicles... Initiate a load balancing strategy: forcibly assign tasks to the nearest charging vehicle and increase its power supply priority.

[0018] S6. Spacetime Conflict Resolution Mechanism: To eliminate multi-vehicle path conflicts, a two-layer detection strategy is designed: 1. Path intersection prediction: Based on the path planning results of each charging and storage vehicle, calculate the spatiotemporal trajectory intersection points and construct a conflict probability matrix. ; 2. Perform dynamic priority adjustment: when When necessary, the passage order will be adjusted according to the following rules: Charging and storage vehicles performing emergency tasks (Urgency > 0.7) have priority; charging and storage vehicles with remaining battery power below 30% have priority; in other cases, a random backoff algorithm is used, with a delay time... .

[0019] This application also provides a dynamic scheduling system for mobile charging and storage vehicles, including: Environmental perception module: integrates UWB positioning device (positioning accuracy ±10cm), vehicle vision sensor (RGB-D camera) and V2X communication unit to build a three-dimensional topology map of the parking lot in real time and update the distribution of charging demand; Collaborative Decision Engine: Deploys PPO reinforcement learning models trained asynchronously using multiple threads, supporting distributed updates of the policy network (Actor-Critic architecture) and the value network; Path planner: It adopts an independently improved algorithm, embeds a dynamic obstacle probability map into the global path, and uses model predictive control to achieve real-time obstacle avoidance in local paths; The hardware embedded system, located inside the charging and storage vehicle, supports the ROS2 communication protocol and real-time sensor data processing. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the principle of the present invention. Figure 2 A framework diagram for optimizing reinforcement learning strategies based on a multi-agent simulation model; Figure 3 A schematic diagram illustrating the energy replenishment scheduling of mobile charging vehicles in a parking lot; Figure 4 This is a pseudocode diagram for the deep reinforcement learning training. Specific implementation measures

[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings: In one specific embodiment, the system consists of the following hardware modules: Mobile charging and storage vehicle: Equipped with a lithium battery energy storage system (capacity 100kWh, SOC range 20%-95%), a bidirectional V2G converter (rated power 50kW, efficiency ≥95%), a UWB positioning module (accuracy ±10cm) and an on-board computing unit (NVIDIA Jetson AGX Xavier).

[0024] Environmental perception module: Deploys RGB-D cameras (1920×1080 resolution) and LiDAR (detection range 50m) to build a 3D topology map of the parking lot; supports V2X communication unit with 5G and DSRC protocols to receive charging demand information and grid electricity price data in real time.

[0025] Central scheduling server: Equipped with a multi-threaded GPU cluster (NVIDIA A100×4) for multi-agent model training and global state synchronization.

[0026] In one specific embodiment, the training process of the multi-agent reinforcement learning model is as follows: Figure 1 As shown, relevant data is collected to construct a dynamic environment model, an agent with a multi-dimensional action space is built based on deep reinforcement learning, a reward function that integrates path and energy constraints is designed, multiple agents are trained using an asynchronous distributed architecture, and a dynamic scheduling strategy is generated.

[0027] In one specific embodiment, key training parameters are set, including the learning rate: Actor network = Critic network = Discount factor Soft update coefficient Initialize the experience replay pool capacity = Each agent is initialized to independently explore the environment and store experience data.

[0028] In one specific embodiment, the central server randomly samples 512 sets of data from the playback pool every 1000 steps and calculates the target Q value: ; in and For the target network parameters, through soft update ( )synchronous; In one specific embodiment, the Adam optimizer is used to update the Critic network, and the policy gradient is used to update the Actor network.

[0029] In one specific embodiment, the reward function used in this reinforcement learning is: ; The scheduling strategy for mobile charging and storage vehicles is specifically defined as a vector containing multi-dimensional decision variables, i.e. ,in yes The optimal solution found before time step [time]. yes Perform actions at all times The new solution obtained afterwards; ; The above items represent, in order, movement path penalty, peak-valley arbitrage reward, low battery penalty, and task urgency reward. These are the weighting coefficients. It is the topological distance from the mobile charging vehicle to the kth charging demand point. The maximum allowed path length, These are the charge amount and the discharge amount, respectively. This is the low battery limit for mobile charging vehicles.

[0030] In one specific embodiment, the constraints used in this reinforcement learning are: ;

[0031] For discharge power, For charging power, For time step, The remaining battery power of the mobile charging vehicle. Maximum battery capacity for mobile charging vehicles; ; From node Time cost The remaining waiting threshold for vehicles in demand; ; It is the first The amount of a constraint violation, The Lagrange penalty coefficient is the corresponding constraint. In one specific embodiment, the exploration strategy employs Ornstein-Uhlenbeck noise with an initial noise figure of The weights decay linearly to 0.1; using the Adam optimizer, the weight decay is... .

[0032] ; Set the parameters as follows .

[0033] In one specific embodiment, the dispatch center broadcasts a task package containing structured data including: the coordinates of the demand location (x, y, z), the power demand ΔE (kWh), and the urgency level.

[0034] ; Where: k is the urgency slope coefficient.

[0035] In one specific embodiment, the calculation weight of each charging and storage vehicle is calculated: ; If the weight difference This triggers the allocation of combined tasks.

[0036] The vehicle with the highest weight is selected for charging and storage. If the weights are similar, the task is split (e.g., charging and navigation are separated). The task allocation results are broadcast to the relevant vehicles via V2X.

[0037] In one specific embodiment, spatiotemporal conflict resolution is performed. The spatiotemporal trajectory intersections of each vehicle's path are calculated. If the time windows overlap and the spatial distance is less than a safety threshold (…), then… ), marked as conflict: ; in, For vehicle speed, For a safe distance, This is the sensitivity coefficient.

[0038] In one specific embodiment, dynamic priority adjustment is performed on the following vehicles: Emergency vehicles (Urgency > 0.7) have priority; low-battery vehicles (SOC < 30%) have second priority; other vehicles will be subject to a backoff algorithm with a delay time of: .

[0039] In one specific embodiment, a globally self-improving algorithm is adopted, and the cost function is: ; Set weights .

[0040] In one specific embodiment, the speed and steering angle are calculated in real time based on the Dynamic Window Method (DWA): ; In one specific embodiment, the reinforcement learning training process described above can be described as state input, reward calculation, and final scheduling, as detailed below. Figure 2 As shown.

[0041] In one specific embodiment, a sudden charging request from an electric vehicle (location distribution as follows) Figure 3 As shown in the figure, the power grid is in peak period (electricity price 0.9 yuan / kWh).

[0042] After the dispatch center broadcasts the task, the weights of charging and storage vehicles 1, 2, and 3 are calculated to be 0, 82, and 0, respectively. The charging and storage vehicle 2 won the bid for task T001 (highest weight), and A and B continue to complete the current task.

[0043] In another specific embodiment, after the dispatch center broadcasts the task, the weights of charging and storage vehicles 1, 2, and 3 are calculated to be 85, 78, and 92, respectively; charging and storage vehicle 3 wins task T001 (highest weight), and vehicles 1 and 2 compete for task T002 (weight difference). The combined allocation is triggered: A is responsible for charging, and B is responsible for navigation guidance. A conflict is predicted between the paths of charging and storage vehicles A and B at node N5. According to the priority rules, A (performing an emergency task) has priority passage, and B detours via route L3. Calculations show that completing 3 charging services results in a total discharge of 45 kWh, with an arbitrage profit of 45 × (0.9 - 0.3) = 27 yuan; the average task response time is 8.5 minutes, and the empty driving mileage is reduced by 36%.

[0044] In one specific embodiment, the specific algorithm for reinforcement learning described above can be described by pseudocode, such as... Figure 4 As shown.

[0045] In one specific embodiment, the process can be described as follows: The system first acquires a real-time 3D topological map of the parking lot and charging demand distribution data through an environmental perception module, and transmits this data to the collaborative decision engine. The collaborative decision engine, based on a PPO reinforcement learning model, rapidly generates the optimal scheduling strategy through multi-threaded asynchronous training. The Actor-Critic architecture policy network is responsible for generating specific scheduling instructions, while the value network evaluates the long-term benefits of these instructions to ensure the global optimality of the scheduling strategy. Based on the scheduling strategy generated by the collaborative decision engine, the path planner uses an improved algorithm for global path planning and embeds a dynamic obstacle probability map into the path to cope with dynamic obstacles that may appear in the parking lot. Local path planning uses an MPC model to adjust the vehicle's driving path in real time to ensure safe obstacle avoidance.

Claims

1. A dynamic scheduling method for mobile charging and storage vehicles based on reinforcement learning, characterized in that, Includes the following steps: S1. Construct a dynamic environment model for the parking lot, and collect real-time data on the SOC status of mobile charging and storage vehicles, the spatiotemporal distribution of vehicles with charging needs, and the topological relationship of parking spaces in the parking lot. S2. Construct an intelligent agent based on a deep reinforcement learning algorithm, and map the state space of the dynamic environment model into a multi-dimensional action space that includes charging and discharging decisions, path planning, and energy allocation. S3. Design a reward function that integrates dynamic path cost and energy constraint, wherein the reward function includes a path efficiency reward term, a peak-valley electricity price arbitrage benefit term, and a SOC balance penalty term; S4. Train multiple agents using an asynchronous distributed architecture, establish a collaborative decision-making mechanism for multiple mobile charging and storage vehicles, and generate dynamic scheduling strategies by jointly optimizing charging task allocation and path conflict resolution. S5. Based on the real-time updated charging demand, dynamically adjust the charging and discharging sequence, travel path, and energy storage allocation ratio of the mobile charging and storage vehicle.

2. The method according to claim 1, characterized in that, The construction of the parking lot dynamic environment model in step S1 includes the following key elements: mobile charging and storage vehicles, mobile charging and storage vehicle refueling stations, and fixed parking spaces.

3. The method according to claim 1, characterized in that, The state space modeling in step S2 includes: converting the real-time SOC value of the mobile charging vehicle into a normalized energy state vector, and generating a topological distance matrix by combining the location coordinates of the charging demand vehicle. The topological distance matrix includes dynamic obstacle avoidance weights.

4. The method according to claim 1, characterized in that, The reward function in step S3 satisfies the following constraints: For a single mobile charging and storage vehicle, its discharge capacity must not exceed the capacity corresponding to the current SOC and its charging capacity must not exceed the remaining energy storage space; the time cost between adjacent task nodes in the path planning must be less than the remaining waiting threshold of the demanding vehicle.

5. The method according to claim 1, characterized in that, The multi-vehicle collaborative decision-making mechanism in step S4 includes: Based on the contract network protocol, the task bidding-tender process allocates the nearest neighbor mobile charging vehicle for high-urgency charging needs through a priority competition algorithm, and uses a spatiotemporal conflict detection algorithm to eliminate the risk of deadlock caused by path crossing.

6. The method according to claim 1, characterized in that, The dynamic adjustment strategy in step S5 includes: when a new charging demand is detected, based on the current task queue length and remaining power of each mobile charging vehicle, a weighted decision tree algorithm is used to reallocate tasks and generate local path replanning instructions.

7. A dynamic scheduling system for mobile charging and storage vehicles implementing the method of any one of claims 1-6, characterized in that, include: The environmental sensing module is used to collect real-time information on parking space status and vehicle charging requests. A collaborative decision-making engine deploys deep reinforcement learning models and performs asynchronous training and policy generation for multiple agents. The path planner generates global paths and real-time local obstacle avoidance paths with dynamic obstacle avoidance based on an improved algorithm. The energy management unit controls the charging and discharging power of the mobile charging and storage vehicle and the SOC threshold of the energy storage battery according to the scheduling strategy. An embedded system deployed on MCSV hardware that supports the ROS2 communication protocol and real-time sensor data processing.

8. The system according to claim 7, characterized in that, The environmental perception module integrates a UWB positioning device and an on-board vision sensor to construct a three-dimensional topological map of the parking lot and update the set of passable paths in real time.

Citation Information

Patent Citations

  • A wireless sensor network-based mobile charging vehicle path planning and charging scheduling algorithm

    CN112909928B

Cited By

  • Wharf vehicle and loading and unloading equipment collaborative operation method and device

    CN121119977A

  • A method and device for coordinating the operation of a terminal vehicle and a handling device

    CN121119977B

  • Urban charging stop cooperative scheduling method and scheduling equipment

    CN121860323A

  • Mobile storage and charging system scheduling and path planning method based on multi-objective optimization

    CN122022313A