Large-scale electric heavy truck cluster charging management method based on model predictive control and deep reinforcement learning combined architecture

By combining model predictive control and deep reinforcement learning, the problem of power grid dynamic response and long-term uncertainty in the charging management of large-scale electric heavy-duty truck clusters is solved, realizing the safe and stable operation of the power grid and efficient utilization of resources, and improving the intelligent management level of electric heavy-duty truck cluster charging.

CN121608631AInactive Publication Date: 2026-03-06HUANENG POWER INT INC YINGKOU POWER PLANT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511829460.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-03-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing electric vehicle charging management methods are unable to dynamically respond to changes in grid conditions when dealing with large-scale electric heavy truck cluster charging, leading to grid overload and insufficient resource utilization. Furthermore, they lack the ability to adapt to medium- and long-term uncertainties, affecting the safe and stable operation of the grid and load balance.

Method used

By adopting a combined architecture of Model Predictive Control (MPC) and Deep Reinforcement Learning (DRL), and through multi-source data acquisition and hybrid optimization control strategies, intelligent collaborative management of electric heavy-duty truck clusters is achieved. The charging power allocation is optimized by combining real-time grid status and battery status, and the strategy is adjusted through deep reinforcement learning to cope with long-term uncertainties.

Benefits of technology

It improves the dynamic response capability of the power grid to the charging of electric heavy truck clusters, reduces the risk of power grid failure, enhances the regional load balancing capability, extends battery life, reduces maintenance costs, and realizes intelligent management and efficient utilization of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121608631A_ABST
    Figure CN121608631A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale electric heavy truck cluster charging management method based on a model predictive control and deep reinforcement learning combined architecture, and the method comprises the steps: collecting multi-source data, generating a hybrid optimization control strategy, and issuing an instruction. The method specifically comprises the following steps: firstly, collecting power grid operation parameters, electric heavy truck charging parameters and battery real-time state parameters; then an optimal charging power distribution scheme meeting system constraints is solved in a finite time window based on model predictive control (MPC), deep reinforcement learning (DRL) takes long-term cumulative income maximization as a target, a charging scheme output by the model predictive control is adjusted, and an adjusted charging control instruction is issued to each charging pile for execution; and finally, the deep reinforcement learning revenue function comprehensively considers the power grid peak-valley difference and the user satisfaction degree, and the adjusted charging control instruction is issued to each charging pile for execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system control and electric vehicle charging management technology, specifically to a large-scale electric heavy-duty truck cluster charging management method based on a combined architecture of model predictive control and deep reinforcement learning. Background Technology

[0002] With the large-scale promotion and application of electric vehicles, the centralized charging of large-scale electric vehicle clusters has become a new and typical electricity consumption behavior. Its load agglomeration effect in the spatiotemporal dimensions has brought unprecedented pressure to the safe and stable operation of the power system, regional load balance, and power quality assurance. In current practice, electric vehicle charging management mostly adopts rule-based control strategies or relies on a single optimization algorithm. These methods are gradually showing their inherent limitations when dealing with high uncertainty and complex system constraints, mainly in the following aspects:

[0003] First, rule-based control methods are typically based on preset thresholds or simple heuristic conditions, such as timed start-stop or segmented power regulation. While easy to implement, they lack the ability to dynamically respond to real-time grid conditions, uncertainties in user behavior, and fluctuations in distributed energy resources. Their control logic is relatively rigid, making it difficult to achieve fine-grained scheduling in variable scenarios, which can easily lead to low system operating efficiency and even cause local grid overload or insufficient resource utilization.

[0004] Second, single optimization algorithms, such as model predictive control, can achieve local optimization under constraints within a finite time domain and have good real-time adjustment capabilities. However, these methods rely heavily on the accurate mathematical model of the system and lack the ability to predict and adapt to medium- and long-term uncertainties (such as the random arrival behavior of electric vehicle users, the strong fluctuations in wind and solar power output, and the dynamic changes in electricity price signals). As the prediction time domain extends, the accumulation of uncertainties will lead to a significant degradation in optimization performance and cannot guarantee long-term economic efficiency and robustness.

[0005] Third, although reinforcement learning methods can learn long-term optimization strategies in uncertain scenarios through autonomous interaction with the environment, avoiding dependence on accurate models, they typically handle system constraints in an unconstrained or penalized manner, lacking explicit guarantee mechanisms for the safe operation boundaries of the power system (such as node voltage limits, line capacity constraints, transformer load capacity, etc.). This can lead to the trained strategies generating infeasible solutions during actual control processes, seriously affecting the engineering applicability and reliability of the methods.

[0006] Therefore, in order to address the multi-dimensional challenges brought about by large-scale electric vehicle charging, especially the phenomenon of large-scale heavy truck cluster charging, it is urgent to develop a new intelligent collaborative management method that can both embed the physical constraints of the system to ensure operational feasibility and have the learning and optimization capabilities to cope with long-term uncertainties. This will enable the coordinated achievement of multiple objectives, such as load stabilization and renewable energy consumption, while ensuring the safe and stable operation of the power grid. Summary of the Invention

[0007] To address this, the present invention provides a large-scale electric heavy-duty truck cluster charging management method based on a combined architecture of model predictive control and deep reinforcement learning, in order to solve the problems of poor grid adaptability and difficulty in matching the dynamic changes of grid load in the prior art, which can easily lead to safety hazards such as overload.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] A method for charging management of large-scale electric heavy-duty truck clusters based on a combined architecture of model predictive control and deep reinforcement learning includes the following steps:

[0010] Step 1: Multi-source data acquisition: Collect power grid operation parameters, electric heavy truck charging demand parameters, and real-time battery status parameters;

[0011] Step 2: Generation of Hybrid Optimization Control Strategy: Based on Model Predictive Control (MPC), the optimal charging power allocation scheme that satisfies the system constraints is solved within a finite time window; Deep Reinforcement Learning (DRL) aims to maximize long-term cumulative benefits, adjusts the charging scheme output by Model Predictive Control, and sends the adjusted charging control commands to each charging pile for execution.

[0012] Step 3: The deep reinforcement learning reward function comprehensively considers the peak-valley difference of the power grid and user satisfaction, and sends the adjusted charging control command to each charging pile for execution.

[0013] Furthermore, the power grid's own operating parameters include power grid load level, power quality, electricity price and policy signals, power grid topology and access point information; the battery's real-time status includes core power parameters, temperature parameters, voltage and current parameters.

[0014] Furthermore, the process for collecting the power grid load level is as follows:

[0015] A1. Calculate the real-time power grid load sequence Noise is eliminated using a sliding window filter, and the filtered load is obtained, which is calculated using the following formula:

[0016] ;

[0017] Where K is the sliding window length (ranging from 5 to 10), and T is the total data acquisition time. Original grid load;

[0018] A2. Calculate the power grid load fluctuation coefficient, which reflects load stability:

[0019] ;

[0020] Where W is the fluctuation calculation window, Filtered grid load.

[0021] Furthermore, the process for collecting the charging demand parameters of the electric heavy-duty tractor is as follows:

[0022] B1. Estimated departure time of the data collection vehicle With the current moment The effective charging time is calculated using the following formula:

[0023] ;

[0024] In the formula, This represents the charging buffer time (values ​​range from 0.5 to 1 hour), where i is the vehicle number. Let i be the estimated departure time of the i-th vehicle. The current moment;

[0025] B2. Collect vehicle targets With the present The required charging amount can be calculated using the following formula:

[0026] ;

[0027] in, For the battery's rated capacity, For charging efficiency, The target remaining battery level for vehicle i. The current remaining battery power of vehicle i.

[0028] Furthermore, the process for acquiring the real-time battery status parameters is as follows:

[0029] C1. Collect the voltage of individual battery cells The voltage equalization degree is calculated using the following formula:

[0030] ;

[0031] in, Let i be the total number of battery cells in the i-th vehicle. Voltage of the j-th battery cell in the i-th vehicle;

[0032] C2. Collect battery pack temperature The temperature safety factor is calculated using the following formula:

[0033] ;

[0034] in, The battery pack temperature of vehicle i.

[0035] Furthermore, the optimal charging power allocation scheme is based on a preprocessed data-driven MPC optimization model, which solves for a multi-constraint charging power scheme within the prediction time domain N. The MPC multi-objective optimization function is:

[0036] ;

[0037] In the formula, M represents the total number of electric vehicles. Weighting coefficients (satisfying) ), The charging power for the i-th vehicle at time t.

[0038] Furthermore, the dynamic power upper limit constraint of the MPC is as follows:

[0039] ;

[0040] ;

[0041] ;

[0042] in, For SoC correction functions, Apparent power at the grid connection point For the power factor of the power grid, For other load power.

[0043] Furthermore, the MPC constraints also include:

[0044] (1) Charging power limit ;

[0045] (2) Power grid capacity limitations ;

[0046] (3) Battery state equation ;

[0047] In the formula, Let be the charging power of the i-th electric vehicle at time t; Let be the minimum charging power of the i-th electric vehicle; The maximum charging power of the i-th electric vehicle; N represents the total number of electric vehicles. Let t be the maximum allowable power injection into the power grid at time t; Let be the remaining battery charge (0-1) of the i-th electric vehicle at time t. The charging efficiency of the i-th electric vehicle; To control the cycle length (hours); Let be the total battery capacity (kWh) of the i-th electric vehicle.

[0048] Furthermore, the implementation process of the deep reinforcement learning (DRL) is as follows:

[0049] D1. Optimization goal of deep reinforcement strategy:

[0050] ;

[0051] in, For the policy function, The discount factor is (0.9-0.98). Total decision-making time, The trajectory is a state-action-reward pattern.

[0052] D2 and DRL power adjustment outputs power adjustment via dual Critic and Actor networks:

[0053] ;

[0054] in, For Actor network parameters, For adjustment coefficients;

[0055] D3. Calculate the dynamic weighting coefficients:

[0056] ;

[0057] In the formula, For model prediction of control scheme weights, Weights for deep reinforcement learning schemes;

[0058] D4. Calculate charging power:

[0059] .

[0060] Furthermore, the charging pile control signal is generated by converting the charging power into the charging pile's PWM duty cycle.

[0061] ;

[0062] In the formula, This is the effective value (V) of the grid line voltage. The rated current of the charging pile, This refers to the power factor of the charging pile.

[0063] The present invention has the following advantages:

[0064] 1. By acquiring real-time multi-source data (such as grid load level and power quality) and using sliding window filtering technology, the system dynamically captures changes in grid status and eliminates noise interference. MPC optimizes charging power allocation within a short-term prediction window, significantly smoothing load fluctuations and reducing the peak-valley difference in the grid. This method enhances the grid's dynamic response capability to electric heavy-duty truck cluster charging, reduces the rigidity problems caused by traditional rule-based control, lowers the risk of grid faults, and improves regional load balancing capabilities.

[0065] 2. By quantifying the effective charging time and required charging amount, personalized charging planning is achieved. The buffer time setting enhances scheduling flexibility and avoids vehicle delays due to insufficient charging. User vehicles can leave on time with a full charge, reducing waiting time and improving the experience. At the same time, by avoiding overcharging, battery life and energy efficiency are optimized.

[0066] 3. By monitoring the battery status in real time and using the temperature safety factor, the battery safety boundary is dynamically assessed. The temperature safety factor formula automatically limits the power at low or high temperatures to prevent safety accidents caused by thermal runaway or voltage imbalance, thus extending battery life and reducing maintenance costs. The system reliability is improved, avoiding charging interruptions or safety accidents caused by battery failure.

[0067] 4. The DRL part learns environmental changes (such as random user behavior and fluctuations in wind and solar power output) through the Actor-Critic network and outputs power adjustment amounts, avoiding the optimization degradation of MPC in the long time domain due to model dependence. The system robustness is significantly improved, and it can cope with medium and long-term uncertain scenarios (such as sudden changes in electricity prices or dense vehicle access), reducing the need for manual intervention and realizing intelligent management.

[0068] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. Attached Figure Description

[0069] To more intuitively illustrate the prior art and this application, exemplary drawings are provided below. It should be understood that the specific shapes and structures shown in the drawings should not generally be regarded as limiting conditions for implementing this application; for example, based on the technical concept disclosed in this application and the exemplary drawings, those skilled in the art are able to easily make conventional adjustments or further optimizations to the addition / reduction / classification, specific shapes, positional relationships, connection methods, size ratios, etc. of certain units (components).

[0070] Figure 1This document presents an implementation flowchart of a large-scale electric heavy-duty truck cluster charging management method based on a combined architecture of model predictive control and deep reinforcement learning, as provided in an embodiment of this application.

[0071] Figure 2 This is a flowchart of the MPC+DRL application in the charging of electric vehicle clusters according to the present invention.

[0072] Figure 3 This is a block diagram illustrating the implementation of Deep Reinforcement Learning (DRL) in this invention. Detailed Implementation

[0073] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these embodiments are merely for further explanation of the present invention and should not be construed as limiting the scope of protection of the present invention. Technical engineers in the field can make some non-essential improvements and adjustments to the present invention based on the above-described content. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0074] Please see Figure 1-3 A method for charging management of large-scale electric heavy-duty truck clusters based on a combined architecture of model predictive control and deep reinforcement learning includes the following steps:

[0075] Step 1: Multi-source data acquisition: Collect power grid operating parameters, electric heavy truck charging demand parameters, and real-time battery status parameters; power grid operating parameters include power grid load level, power quality, electricity price and policy signals, power grid topology and access point information; real-time battery status includes core power parameters, temperature parameters, voltage and current parameters.

[0076] Step 2: Generation of Hybrid Optimization Control Strategy: Based on Model Predictive Control (MPC), the optimal charging power allocation scheme that satisfies the system constraints is solved within a finite time window; Deep Reinforcement Learning (DRL) aims to maximize long-term cumulative benefits, adjusts the charging scheme output by Model Predictive Control, and sends the adjusted charging control commands to each charging pile for execution.

[0077] Step 3: The deep reinforcement learning reward function comprehensively considers the peak-valley difference of the power grid and user satisfaction, and sends the adjusted charging control command to each charging pile for execution.

[0078] In a further preferred embodiment of the present invention, the process of collecting power grid load levels is as follows:

[0079] A1. Calculate the real-time power grid load sequence Noise is eliminated using a sliding window filter, and the filtered load is obtained, which is calculated using the following formula:

[0080] ;

[0081] Where K is the sliding window length (ranging from 5 to 10), and T is the total data acquisition time. Original power grid load.

[0082] A2. Calculate the power grid load fluctuation coefficient, which reflects load stability:

[0083] ;

[0084] Where W is the fluctuation calculation window, Filtered grid load.

[0085] In this embodiment, by real-time monitoring and mathematical filtering, changes in the power grid state are dynamically captured, improving data accuracy, eliminating random noise interference, and enabling the load fluctuation coefficient to reflect the stability of the power grid in real time. This provides a reliable input for subsequent MPC and DRL optimization, avoids control deviations caused by data distortion, and enhances the adaptability of the power grid.

[0086] In a further preferred embodiment of the present invention, the process for collecting charging demand parameters of the electric heavy-duty tractor is as follows:

[0087] B1. Estimated departure time of the data collection vehicle With the current moment The effective charging time is calculated using the following formula:

[0088] ;

[0089] In the formula, This represents the charging buffer time (values ​​range from 0.5 to 1 hour), where i is the vehicle number. Let i be the estimated departure time of the i-th vehicle. The current moment;

[0090] B2. Collect vehicle targets With the present The required charging amount can be calculated using the following formula:

[0091] ;

[0092] in, For the battery's rated capacity, For charging efficiency, The target remaining battery level for vehicle i. The current remaining battery power of vehicle i.

[0093] In this embodiment, by combining time constraints and power demand, the charging urgency of each vehicle is quantified to achieve personalized charging planning, avoid overcharging or undercharging, enhance scheduling flexibility and improve user satisfaction (such as ensuring that vehicles leave on time with a full charge) by setting buffer time, while reducing the peak load pressure on the power grid.

[0094] In a further preferred embodiment of the present invention, the process of acquiring real-time battery status parameters is as follows:

[0095] C1. Collect the voltage of individual battery cells The voltage equalization degree is calculated using the following formula:

[0096] ;

[0097] in, Let i be the total number of battery cells in the i-th vehicle. Voltage of the j-th battery cell in the i-th vehicle;

[0098] C2. Collect battery pack temperature The temperature safety factor is calculated using the following formula:

[0099] ;

[0100] in, The battery pack temperature of vehicle i.

[0101] In this embodiment, the battery safety boundary is dynamically assessed through physical parameter monitoring to prevent safety accidents caused by battery overheating or voltage imbalance, thereby extending battery life. The temperature safety factor is integrated as a constraint into subsequent optimizations to ensure the charging process operates within a safe range, improving system reliability.

[0102] In a further preferred embodiment of the present invention, the optimal charging power allocation scheme is constructed based on preprocessed data to build an MPC optimization model, and the charging power scheme satisfying multiple constraints is solved within the prediction time domain N. The MPC multi-objective optimization function is:

[0103] ;

[0104] In the formula, M represents the total number of electric vehicles. Weighting coefficients (satisfying) ), The charging power for the i-th vehicle at time t.

[0105] In this embodiment, MPC achieves real-time optimization within a short window, effectively smoothing grid load fluctuations (such as reducing peak-to-valley differences) and avoiding local overloads. Hard constraints ensure system safety and improve charging efficiency (such as reducing power fluctuation losses), and the weighting coefficients are adjustable to adapt to different scenario requirements.

[0106] In a further preferred embodiment of the present invention, the dynamic power upper limit constraint of MPC is:

[0107] ;

[0108] ;

[0109] ;

[0110] in, For SoC correction functions, Apparent power at the grid connection point For the power factor of the power grid, For other load power.

[0111] In a further preferred embodiment of the present invention, the MPC constraint conditions also include:

[0112] (1) Charging power limit ;

[0113] (2) Power grid capacity limitations ;

[0114] (3) Battery state equation ;

[0115] In the formula, Let be the charging power of the i-th electric vehicle at time t; Let be the minimum charging power of the i-th electric vehicle; The maximum charging power of the i-th electric vehicle; N represents the total number of electric vehicles. Let t be the maximum allowable power injection into the power grid at time t; Let be the remaining battery charge (0-1) of the i-th electric vehicle at time t. The charging efficiency of the i-th electric vehicle; To control the cycle length (hours); Let be the total battery capacity (kWh) of the i-th electric vehicle.

[0116] In a further preferred embodiment of the present invention, the implementation process of deep reinforcement learning (DRL) is as follows:

[0117] D1. Optimization goal of deep reinforcement strategy:

[0118] ;

[0119] in, For the policy function, The discount factor is (0.9-0.98). Total decision-making time, The trajectory is a state-action-reward pattern.

[0120] D2 and DRL power adjustment outputs power adjustment via dual Critic and Actor networks:

[0121] ;

[0122] in, For Actor network parameters, For adjustment coefficients;

[0123] D3. Calculate the dynamic weighting coefficients:

[0124] ;

[0125] In the formula, For model prediction of control scheme weights, Weights for deep reinforcement learning schemes;

[0126] D4. Calculate charging power:

[0127] ;

[0128] In this embodiment, machine learning is used to adapt to environmental changes and avoid model dependency defects. DRL improves the system's adaptability to medium- and long-term uncertainties (such as random user behavior and wind and solar fluctuations) and maximizes cumulative benefits (such as reducing total electricity costs and increasing renewable energy consumption). The dynamic weighting mechanism balances short-term security and long-term economics, enhances robustness, and avoids optimization degradation of MPC in the long time domain.

[0129] In a further preferred embodiment of the present invention, the principle of generating the charging pile control signal is to convert the charging power into the PWM duty cycle of the charging pile:

[0130] ;

[0131] In the formula, This is the effective value (V) of the grid line voltage. The rated current of the charging pile, This refers to the power factor of the charging pile.

[0132] In this embodiment, digital commands are converted into analog control signals through power electronic conversion, achieving high-precision power delivery, reducing conversion losses, improving charging efficiency, and ensuring real-time execution of commands through PWM modulation to avoid overshoot or undershoot, thereby ensuring user satisfaction (such as timely full charge) and maintaining the power quality of the power grid.

[0133] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A large-scale electric heavy truck cluster charging management method based on a model predictive control and deep reinforcement learning combined architecture, characterized in that, Comprising the following steps: Step one, multi-source data collection: collect power grid operation parameters, electric heavy truck charging demand parameters and real-time battery state parameters; Step two, hybrid optimization control strategy generation: based on the model predictive control (MPC), the optimal charging power allocation scheme that meets the system constraints is solved within a limited time window; the deep reinforcement learning algorithm (DRL) adjusts the charging scheme output by the model predictive control with the goal of maximizing long-term cumulative revenue, and the adjusted charging control instructions are issued to each charging pile for execution; Step three, the deep reinforcement learning revenue function considers the peak-valley difference of the power grid and user satisfaction, and the adjusted charging control instructions are issued to each charging pile for execution.

2. The method of claim 1, wherein the method is based on a model predictive control and deep reinforcement learning combined architecture. The power grid's own operation parameters include power grid load level, power quality, electricity price and policy signals, power grid topology and access point information; the real-time state of the battery includes core power parameters, temperature parameters, voltage and current parameters.

3. The method of claim 2, wherein, The collection process of the power grid load level is as follows: A1, calculate real-time power grid load sequence The noise is eliminated by using a sliding window filter to obtain a filtered load, which is calculated by the following formula: ; Wherein, K is the length of the sliding window (value 5~10), T is the total length of data collection, Raw grid load; A2, calculate the power grid load fluctuation coefficient to reflect the load stability: ; wherein W is a wave calculation window, Filtered grid load.

4. The method of claim 1, wherein, The collection process of the electric heavy truck charging demand parameters is as follows: B1, collecting vehicle estimated departure time with the current time The effective charging duration is calculated by the following formula: ; In the formula, is the charging buffer time (value 0.5-1h), i is the vehicle number, is the expected departure time of the ith vehicle, is the current time; B2, collecting vehicle targets with the current The required charging amount is calculated by the following equation: ; wherein, is the battery rated capacity, is the charging efficiency, is the ith vehicle target remaining electric quantity, is the ith vehicle current remaining electric quantity.

5. The method of claim 1, wherein, The collection process of the real-time battery state parameters is as follows: C1, collect cell voltage The voltage equalization degree is calculated by the following formula: ; wherein, is the total number of battery cells for the ith vehicle, is the voltage of the jth section of the ith vehicle. C2, collecting battery pack temperature The temperature safety factor is calculated by the following equation: ; wherein, the ith vehicle battery pack temperature.

6. The method of claim 1, wherein, The optimal charging power allocation scheme is based on the preprocessed data to construct a predictive control model (MPC) optimization model, which solves the charging power scheme that meets multiple constraints within a prediction time domain N, and the multi-objective optimization function of the predictive control model (MPC) is: ; In the formula, M is the total number of electric vehicles, is a weight coefficient (satisfying ), is the charging power of the ith vehicle at time t.

7. The method of claim 6, wherein the method is based on a model predictive control and deep reinforcement learning combined architecture. The dynamic power upper limit constraint of the predictive control model (MPC) is: ; ; ; wherein, is a SoC correction function, is a grid point of view apparent power, is a grid power factor, is a further load power.

8. The method of claim 7, wherein, The constraint conditions of the predictive control model (MPC) also include: (1) Charging power limit ; (2) Power grid capacity limitation ; (3) Battery state equation ; wherein, is the charging power of the i-th electric vehicle at time t; is the minimum charging power of the i-th electric vehicle; is the maximum charging power of the i-th electric vehicle; N represents the total number of electric vehicles is the maximum injection power allowed by the grid at time t; is the remaining battery level of the i-th electric vehicle at time t (0-1); is the charging efficiency of the i-th electric vehicle; is the control period length (hours); is the total battery capacity of the i-th electric vehicle (kWh).

9. The method of claim 1, wherein, The implementation process of the deep reinforcement learning algorithm (DRL) is as follows: D1, the optimization goal of deep reinforcement strategy: ; wherein, is a policy function, is a discount factor (0.9-0.98), is a total decision length, is a state-action-reward trajectory; D2, the DRL power adjustment output, the power adjustment output is output through the double Critic network and the Actor network: ; wherein, is an Actor network parameter, is an adjustment coefficient; D3, calculate the dynamic weight coefficient: ; wherein is a model predictive control scheme weight, is a deep reinforcement learning scheme weight; D4, calculate the charging power: 。 10. The method of claim 9, wherein, The generation principle of the charging pile control signal is to convert the charging power into the charging pile PWM duty cycle: ; In the formula, is the grid line voltage effective value (V), is the charging pile rated current, is the charging pile power factor.