Flexible charging power distribution method, system and equipment for electric vehicle and medium
By collecting charging power demand information from charging terminals in real time and using an LSTM-enhanced deep reinforcement learning model for reasonable allocation, the problem of insufficient charging pile configuration capacity is solved, achieving efficient and reasonable power allocation during peak charging periods and improving the user charging experience.
Patent Information
- Application Number
- CN202510964034.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-14
AI Technical Summary
In large supercharging stations, the power module configuration capacity of the charging pile is usually less than the sum of the maximum charging power of the charging terminals. As a result, not all charging terminals can be allocated sufficient charging power during peak charging periods. How to reasonably allocate the limited charging pile configuration capacity has become an urgent problem to be solved.
The system collects charging power demand information from charging terminals in real time, processes it using an LSTM-enhanced deep reinforcement learning model, outputs an initial power allocation, and performs a balance optimization when the initial power allocation exceeds the demand, ensuring that charging queuing satisfaction and charging delay satisfaction are maximized.
By leveraging dynamic response capabilities, resource waste is avoided, ensuring that user needs are met to the maximum extent possible within the limitations of charging pile capacity, reducing efficiency losses, increasing overall throughput, and lowering user waiting time.
Smart Images

Figure CN120942072A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electric vehicle charging technology, and in particular to a flexible charging power distribution method, system, device and medium for electric vehicles. Background Technology
[0002] With the development of electric vehicle charging technology, flexible charging architectures centered on charging piles enhance the flexibility of electric vehicle charging by allowing charging facility operators to adjust flexible power allocation control switches. However, current charging power allocation methods typically assume that the power configuration capacity of the charging pile power modules is not less than the charging power demand of the electric vehicle, meaning that the charging power demand of the electric vehicle can always be met, and there is no optimization requirement. However, in many large supercharging stations, the power configuration capacity of the charging pile power modules is usually less than the sum of the maximum charging power of the charging terminals, resulting in not all charging terminals being allocated sufficient charging power during peak charging periods.
[0003] Therefore, how to rationally allocate the limited charging pile configuration capacity to each charging terminal has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] This invention provides a method, system, device, and medium for flexible charging power allocation of electric vehicles, solving the problem of how to rationally allocate the limited charging pile configuration capacity to each charging terminal.
[0005] To address the aforementioned technical problems, the first aspect of this invention provides a flexible charging power allocation method for electric vehicles, comprising:
[0006] Real-time collection of charging power demand information of electric vehicles connected to each charging terminal, and quantification of the total charging power demand of all electric vehicles connected to all charging terminals based on the charging power demand information.
[0007] When the total demand for charging power exceeds the configuration capacity of the charging pile, the charging power demand information of each terminal is input into an LSTM-enhanced deep reinforcement learning model for processing, and the initial allocated power of each terminal is output; wherein, the objective of the LSTM-enhanced deep reinforcement learning model is to maximize the charging queuing satisfaction and the charging delay satisfaction.
[0008] When the initial allocated power of each charging terminal is greater than the required charging power of the electric vehicle connected to each charging terminal, the excess allocated power of each charging terminal is balanced and optimized, and the optimized power allocation scheme is executed.
[0009] A second aspect of the present invention provides a flexible charging power distribution system for electric vehicles, comprising:
[0010] The power quantization module is used to collect the charging power demand information of electric vehicles connected to each charging terminal in real time, and quantify the total charging power demand of all electric vehicles connected to all charging terminals based on the charging power demand information.
[0011] The model processing module is used to input the charging power demand information of each terminal into an LSTM-enhanced deep reinforcement learning model for processing when the total demand charging power exceeds the configuration capacity of the charging pile, and output the initial allocated power of each terminal; wherein, the objective of the LSTM-enhanced deep reinforcement learning model is to maximize the charging queuing satisfaction and charging delay satisfaction.
[0012] The scheme execution module is used to balance and optimize the excess power of each charging terminal when the initial allocated power of each terminal is greater than the required charging power of the electric vehicle connected to each charging terminal, and then execute the optimized power allocation scheme.
[0013] A third aspect of the present invention provides an electronic device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the flexible charging power allocation method for electric vehicles as described above.
[0014] A fourth aspect of the present invention provides a computer-readable storage medium comprising a stored computer program, wherein when the device containing the computer-readable storage medium executes the computer program, it implements the electric vehicle flexible charging power allocation method as described above.
[0015] Compared with the prior art, the beneficial effects of the embodiments of the present invention are as follows:
[0016] By collecting real-time charging power demand information from electric vehicles connected to each charging terminal, the total demand for charging power can be accurately quantified. When the total demand exceeds the charging pile capacity, an optimization mechanism is triggered. This dynamic response capability avoids resource waste or overload problems that may occur with traditional static allocation methods, ensuring that the charging pile maximizes user demand within capacity constraints. A deep reinforcement learning model enhanced with LSTM (Long Short-Term Memory) network is introduced, which can efficiently process time-series data and output the initial allocated power. The model optimizes charging queue satisfaction and charging delay satisfaction, balancing user waiting time and charging speed. When the initial allocated power exceeds the actual demand of some terminals, the excess power is balanced and optimized to ensure the rationality of power allocation and the safety and reliability of the charging process. Furthermore, by optimizing power allocation, the solution can reduce efficiency losses caused by insufficient power or unreasonable allocation of the charging pile, thereby improving overall throughput. Attached Figure Description
[0017] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a flexible charging power allocation method for electric vehicles provided in a certain embodiment of the present invention;
[0019] Figure 2 This is a diagram of a flexible charging architecture provided in a certain embodiment of the present invention;
[0020] Figure 3 This is an architecture diagram of an LSTM-based mobile network provided in a certain embodiment of the present invention;
[0021] Figure 4 This is an architecture diagram of an LSTM-based evaluation network provided in a certain embodiment of the present invention;
[0022] Figure 5 This is a flowchart illustrating the process of a deep reinforcement learning model enhanced with LSTM, provided in a certain embodiment of the present invention, processing charging power demand information.
[0023] Figure 6 This is a flowchart of the training process of an LSTM-enhanced deep reinforcement learning model provided in a certain embodiment of the present invention;
[0024] Figure 7 This is a flowchart illustrating a charging service simulation provided in one embodiment of the present invention;
[0025] Figure 8 This is a structural diagram of a flexible charging power distribution system for electric vehicles provided in a certain embodiment of the present invention;
[0026] Figure 9 This is a structural diagram of an electronic device provided in a certain embodiment of the present invention;
[0027] Figure label:
[0028] Among them, 10 is the power quantization module; 20 is the model processing module; 30 is the scheme execution module; 5000 is the electronic equipment; 5001 is the processor; 5002 is the bus; 5003 is the memory; and 5004 is the transceiver. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings and examples. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0030] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0031] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal communication between two components. The terms "vertical," "horizontal," "left," "right," "upper," "lower," and similar expressions used herein are for illustrative purposes only and do not indicate or imply that the system or component referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0032] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is merely for describing specific embodiments and is not intended to limit the invention. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0033] In many large supercharging stations, the power capacity of the charging piles is typically slightly less than the sum of the maximum charging power of the charging terminals. This means that during peak charging periods, not all charging terminals can receive sufficient charging power. For example, a supercharging station for electric vehicles with six 300kW rated charging terminals typically only has a 1440kW rated capacity charging pile (considering a charging concurrency rate of 0.8). In this case, if six electric vehicles with a rated charging power of 300kW are charging simultaneously, their average charging power will only be 240kW. Therefore, rationally allocating the limited charging pile capacity to each charging terminal is crucial for improving the charging experience for electric vehicle users and reducing their waiting time.
[0034] Based on this, in one embodiment, such as Figure 1 As shown, the first aspect of the present invention provides a flexible charging power allocation method for electric vehicles, comprising:
[0035] S1. Real-time collection of charging power demand information of electric vehicles connected to each charging terminal, and quantification of the total charging power demand of all electric vehicles connected to all charging terminals based on the charging power demand information; wherein, the charging power demand information includes rated charging power and remaining charging demand.
[0036] Specifically, this invention employs a flexible charging architecture centered on a charging pile to control a flexible switch, thereby adjusting the power distribution to each charging terminal to improve the flexibility of electric vehicle charging; the flexible charging architecture diagram is shown below. Figure 2 As shown, the data acquisition device in the charging terminal is used to obtain information such as the connection time, rated charging power, battery capacity, current battery state of charge, remaining charging demand, and total charging demand of electric vehicles as the charging power demand information for each electric vehicle connected to the charging terminal. In addition, it is also necessary to collect charging station environmental information in real time, including current time, day of the week, weather conditions, ambient temperature, number of vehicles in the queue, and charging status of each vehicle.
[0037] In one embodiment, quantifying the total charging power demand of all electric vehicles connected to all charging terminals based on the charging power demand information includes:
[0038] The actual power demand of each electric vehicle is determined based on the rated charging power and the remaining charging demand.
[0039] The total required charging power is determined based on the actual power demand of each item.
[0040] Specifically, this invention utilizes real-time collected information to determine whether power allocation optimization is needed: based on the rated charging power and remaining charging demand of each electric vehicle, the actual power demand of that electric vehicle is determined. By summing the actual power demands of all electric vehicles, the total charging power demand of all electric vehicles connected to the current charging terminal can be calculated. However, it should be noted that if an electric vehicle has very little remaining charging demand, the demand power may be less than the rated power. Therefore, the actual power demand of each electric vehicle is min(rated power, remaining charging demand / time granularity).
[0041] S2. When the total demand for charging power exceeds the configuration capacity of the charging pile, the charging power demand information of each terminal is input into an LSTM-enhanced deep reinforcement learning model for processing, and the initial allocated power of each charging terminal is output. Specifically, when the total demand for charging power does not exceed the configuration capacity of the charging pile, the power is allocated according to the demand (each charging terminal receives its demand power, which is the electric vehicle's), that is, the charging power allocated to each charging terminal by each flexible switch is the demand charging power of the electric vehicle connected to it; otherwise, the pre-trained LSTM-enhanced deep reinforcement learning model is activated to intelligently allocate the power of each charging terminal, reducing the average additional waiting time for users and improving charging satisfaction. The objective of the LSTM-enhanced deep reinforcement learning model is to maximize charging queuing satisfaction and charging delay satisfaction; the objective of the LSTM-enhanced deep reinforcement learning model also includes minimizing timeout penalties and queue length penalties.
[0042] Wherein, the charging queuing satisfaction is the ratio of the queuing time of an electric vehicle before connecting to a charging terminal to the charging time required at its rated charging power, which is expressed by the following formula:
[0043]
[0044] In the formula, The charging queue satisfaction of electric vehicles that leave charging terminal i at time t; Queuing time for electric vehicle i; The initial charging demand of the electric vehicle connected to charging terminal i; The rated charging power of the electric vehicle connected to charging terminal i.
[0045] In other words, charging queue satisfaction is the ratio of the queuing time of an electric vehicle when it leaves a charging station to the charging time required at rated power. It can be negative (because the queuing time may be shorter than the charging time at rated power).
[0046] The charging delay satisfaction is the ratio of the charging time expended by an electric vehicle at a charging terminal to the charging time required at its rated charging power, and is expressed by the following formula:
[0047]
[0048] In the formula, Let t be the charging delay satisfaction of the electric vehicle that leaves the charging terminal i at time t; The actual charging time required for the electric vehicle connected to charging terminal i at time t.
[0049] In other words, charging delay satisfaction is the ratio of the actual charging time of an electric vehicle (from the start of charging to full charge) to the charging time required at rated power.
[0050] Considering the significant impact of excessively long waiting times on user satisfaction, a penalty is considered: any additional charging time and queuing time exceeding the required charging time at rated power will incur substantial satisfaction costs; among these,
[0051] The timeout penalty is a linear penalty when the total charging time of the electric vehicle exceeds a preset multiple of the rated charging time, and it is expressed by the following formula:
[0052]
[0053] In the formula, The timeout penalty for an electric vehicle that leaves charging terminal i at time t; c PQ Cost coefficient for satisfaction penalty item; δ represents the queuing time of the electric vehicle connected to charging terminal i at time t; δ is a linear multiple.
[0054] In other words, the timeout penalty is applied if the total charging time of an electric vehicle (queuing + charging) exceeds a preset multiple (δ+1) of the rated charging time, and the excess portion is penalized (multiplied by a coefficient c). PQ ).
[0055] Furthermore, considering that supercharging stations typically occupy a small area, excessive queuing of electric vehicles could cause queue overflow, leading to traffic congestion around the charging facilities. Therefore, a queue length penalty is implemented, which is the square of the penalty when the number of vehicles in the queue exceeds the maximum capacity, expressed by the following formula:
[0056]
[0057] In the formula, c is the queue length penalty at time t; PN The cost coefficient for the queue length penalty term; Let t be the number of electric vehicles queuing in the charging facility at time t; This represents the maximum number of queues that the charging facility can accommodate.
[0058] In other words, the queue length penalty is a penalty applied to the excess number of vehicles when the number of vehicles in the queue exceeds the maximum capacity (multiplied by a coefficient c). PN ).
[0059] Therefore, the objective of the LSTM-enhanced deep reinforcement learning model is to take action A at the current time t (including the charging power allocated to each charging terminal at time t). The expected benefits obtained, i.e., the expected user satisfaction, are maximized. The benefits and penalty costs of the supercharging station at time t are expressed by the following formula:
[0060]
[0061] In the formula, R t , These represent the revenue and penalty cost gained by the supercharging station at time t after taking action A; Let N be the set of electric vehicles that leave the charging station at time t; EVP This refers to the number of charging terminals within a supercharging station.
[0062] In addition, in the initial stage of supercharging station operation, queuing satisfaction and charging delay satisfaction are set with equal weights. Subsequently, the weights of the two can be determined based on user preference surveys or simulation scenario trade-off optimization methods. The weights of the two on the operating revenue of the charging station can be dynamically adjusted in combination with the actual operating data of the charging station and simulation results to obtain the optimal user satisfaction and station operating efficiency.
[0063] Therefore, the expected revenue of a supercharging station during the entire peak charging period from 0 to T is calculated using the following formula:
[0064]
[0065] In the formula, J t Let be the expected return at time t; γ be the reward discount rate.
[0066] In one embodiment, the LSTM-enhanced deep reinforcement learning model includes an LSTM-based action network and an LSTM-based evaluation network; wherein, the LSTM-based action network is used to process temporal state features and output normalized power allocation actions; the LSTM-based evaluation network, also known as the evaluation network, is used to evaluate the long-term satisfaction benefits of the power allocation strategy; the architecture diagram of the LSTM-based action network is shown below. Figure 3 As shown, the architecture diagram of the LSTM-based evaluation network is as follows: Figure 4As shown, in the charging pile power allocation problem, the same environmental conditions (such as 10 AM, the same number of queuing vehicles) do not guarantee the same scenario. The changing trend of charging power demand information over a previous period will affect subsequent environmental information such as the number of arriving vehicles, exhibiting non-Markovian process characteristics. Traditional reinforcement learning assumes Markovian properties, making it difficult to fully consider long-range dependencies. To balance real-time decision-making and historical dependency modeling while maintaining the Markovian properties of state transitions, this invention innovatively introduces an LSTM architecture into the action network and evaluation network to encode historical arrival sequences, queue lengths, and power allocation trajectories, learn long- and short-term dependencies, and capture subsequent traffic change trends.
[0067] The step of inputting the charging power demand information of each of the above into an LSTM-enhanced deep reinforcement learning model for processing, and outputting the initial allocated power of each of the charging terminals, includes:
[0068] The charging power demand information is converted into a real-time state vector, and a state information vector is constructed based on the real-time state vector and the cell state and hidden state information output by the LSTM at the previous time step.
[0069] The state information vector is input into the LSTM-based action network for processing, and the initial allocated power of each charging terminal is output.
[0070] Specifically, the flowchart of the LSTM-enhanced deep reinforcement learning model's processing of charging power demand information is as follows: Figure 5 As shown, the present invention encodes the charging power demand information of each charging terminal and the charging station environment information according to certain rules and formats, and converts them into a real-time state vector. The encoding process can be to quantize each charging power demand information into a value, and then arrange them into a vector in a certain order, with each element representing a specific charging power demand information.
[0071] The real-time state vector is combined with the cell state output by the LSTM at the previous time step and the hidden state information to construct a state information vector. This vector contains the charging power demand information of the charging scenario at the current time point, and together with the long short-term memory historical state information, it provides a rich information foundation for subsequent model processing.
[0072] The constructed state information vector is input into an LSTM-based actor network for processing. The actor network uses the LSTM within the network to process time-series data, capturing long-term dependencies in the data. This transforms the current state information and the long short-term memory historical state information (i.e., the cell state and hidden state information output by the LSTM at the previous time step) into a higher-level feature representation. The features are then further processed and mapped to finally output a normalized action value. This action value is then multiplied by the rated power of the charging terminal to obtain the power allocated by the flexible switch to each charging terminal, thus yielding the power allocation strategy.
[0073] The power allocation policy and real-time state vector output by the action network are input into an LSTM-based evaluation network for processing. The evaluation network, or critic network, is used to assess the expected return of a given power allocation policy under the current real-time state. The LSTM layer in the evaluation network jointly processes the power allocation policy and state information vector, learns the relationship between them, and obtains the expected return as the objective for training the action network.
[0074] This invention constructs a state information vector that considers both the real-time state at the current moment and historical information from long short-term memory. This allows the model to fully utilize historical data, better understand the dynamic changes in electric vehicle charging behavior, and thus make more reasonable power allocation decisions, improving the accuracy and adaptability of power allocation. It also utilizes LSTM to enhance model performance, improving the accuracy and effectiveness of the power allocation strategy. Furthermore, it outputs the power allocation strategy through an LSTM-based action network, and then combines this with an LSTM-based evaluation network to assess the expected return of the strategy. When the expected return reaches the target, the initial power allocation is determined. This decision-making mechanism continuously optimizes the power allocation scheme, enabling the charging system to achieve higher returns or better performance indicators, such as reducing charging costs and improving charging efficiency, while meeting the charging needs of electric vehicles.
[0075] In one embodiment, the LSTM-based mobile network includes an input layer, an LSTM layer, a fully connected layer, and an output layer; wherein,
[0076] The step of inputting the state information vector into the LSTM-based action network for processing and outputting a power allocation strategy includes:
[0077] The state information vector is input to the LSTM layer through the input layer;
[0078] The long short-term state dependency features in the state information vector are extracted by the LSTM layer to obtain the hidden state features at the current time and input to the fully connected layer.
[0079] The hidden state features are transformed by the fully connected layer to obtain low-dimensional features, which are then processed by the output layer to output normalized power allocation action values.
[0080] The power allocation strategy is generated based on the normalized power allocation action value and the rated power of each charging terminal.
[0081] Specifically, this invention uses the constructed state information vector as input data, which is passed to the LSTM layer through the input layer. After receiving the input state information vector, the LSTM layer processes the data using its internal forget gate, input gate, and output gate mechanisms. The forget gate determines which information to forget from the previous time step's cell state, the input gate determines the information to be updated, and the tanh function generates a new candidate value vector. Through the collaborative work of these gating structures, the LSTM layer can effectively extract long-short-term state dependency features from the state information vector to obtain hidden state features, which are then input to the fully connected layer. The fully connected layer receives the hidden state features output from the LSTM layer and performs a fully connected transformation on them. Each neuron in the fully connected layer is connected to all neurons in the previous layer. The input features are linearly combined using weight matrices and bias terms, and then nonlinearly transformed using activation functions (such as ReLU). After the transformation by the fully connected layer, the hidden state features are converted into low-dimensional features. These low-dimensional features are a further refinement and abstraction of the key information in the original state information vector, making them easier for subsequent output layers to process and make decisions. The output layer receives the low-dimensional features output from the fully connected layer, processes them using an appropriate activation function (such as the sigmoid function), and outputs a normalized power allocation action value. Finally, this action value is multiplied by the rated power of the charging terminal, which is the power allocated by the flexible switch to each charging terminal, generating the power allocation strategy at the current moment. This strategy includes the allocated power values for each charging terminal.
[0082] This invention utilizes the unique gating structure in Long Short-Term Memory (LSTM) networks to capture long-short-term state dependency features, enabling the action network to better understand the state change patterns of charging stations at different points in time. The fully connected layer performs a fully connected transformation on the hidden state features output by the LSTM layer, reducing data dimensionality, lowering computational complexity, and extracting more representative and crucial features. This helps subsequent output layers process data more efficiently, while avoiding overfitting and improving the model's generalization ability. The power allocation strategy generation method based on the LSTM deep learning model can dynamically adjust according to the real-time state and historical information of the charging station. Compared to traditional fixed power allocation methods, this scheme can better adapt to the complex and ever-changing environment of charging stations, such as fluctuations in charging demand at different times and the impact of weather changes on charging demand, improving the adaptability and flexibility of power allocation.
[0083] In one embodiment, the LSTM-enhanced deep reinforcement learning model further includes an LSTM-based evaluation network; wherein the training process of the LSTM-enhanced deep reinforcement learning model includes:
[0084] Obtain historical charging power demand information of electric vehicles connected to each of the charging terminals, and convert each of the historical charging power demand information into a historical demand state vector as a state space;
[0085] The normalized power allocation action value of each charging terminal is used as the action space, and a reward function is constructed based on the objective of the LSTM-enhanced deep reinforcement learning model.
[0086] Initialize the parameters of the experience replay pool, the LSTM-based action network, and the LSTM-based evaluation network, and simulate the operation of the charging station from the start of the charging peak to obtain the current state;
[0087] At each time step, based on the current state, the current action is determined by the LSTM-based action network to perform a charging service process simulation, and the current reward and the next state are calculated according to the reward function and the LSTM-based evaluation network, and stored in the experience replay pool together with the current state and the current action.
[0088] A batch of datasets is sampled from the experience replay pool, and the LSTM-based action network and the LSTM-based evaluation network are updated using the Adam optimizer based on the sampled datasets.
[0089] Based on the updated LSTM-based action network and the updated LSTM-based evaluation network, the simulated training process and network update process are repeatedly executed at each time step until the preset time step is reached, resulting in a trained LSTM-enhanced deep reinforcement learning model.
[0090] Specifically, the flowchart of the training process of the LSTM-enhanced deep reinforcement learning model is as follows: Figure 6 As shown, this invention collects historical charging power demand information of electric vehicles connected to various charging terminals. This information includes, but is not limited to, charging time, charging power, battery charge, vehicle dwell time, and historical environmental information of charging stations. This historical information is encoded and converted into a historical demand state vector, and a state space is constructed based on it.
[0091] The normalized power allocation action value of each charging terminal is used as the action space. Each action in the action space corresponds to a set of normalized power allocation values, which are used to represent the power allocation ratio of each charging terminal under different conditions.
[0092] The reward function is constructed based on the objective of the deep reinforcement learning model, and the training objective is to maximize the cumulative expected return.
[0093] Initialize the experience replay pool, which stores data such as the model's state, actions, rewards, and next state during training; and initialize the parameters of the LSTM-based action network and the LSTM-based evaluation network, which can be done randomly, assigning initial values to the network weights and biases; then simulate the operation of the charging station starting from the beginning of the charging peak, and obtain the current state, which contains relevant information about the charging station at that moment, such as the charging status of electric vehicles and the queuing situation.
[0094] The simulation training process then begins: at each time step, the current state is input into the LSTM-based action network. The action network determines the current action, i.e., a set of normalized power allocation action values, based on its internal structure and parameters. A charging service process simulation is then executed, the flowchart of which is shown below. Figure 7 As shown, the system allocates power to the charging terminal based on the current action to simulate the charging process; it calculates the current reward and the next state based on the reward function and the LSTM-based evaluation network; the evaluation network evaluates the merits of the action based on the current state and the current action, gives the corresponding reward value, and calculates the next state based on the operating rules of the charging station and the current state; the current state, the current action, the current reward, and the next state are stored together in the experience replay pool.
[0095] A batch of datasets is sampled from the experience replay pool. The sampling can be random to ensure data independence. Based on the sampled datasets, the parameters of the LSTM-based action network and the LSTM-based evaluation network are updated using the Adam optimizer. The Adam optimizer adaptively adjusts the learning rate based on the gradient information of the data and updates the network weights and biases through the backpropagation algorithm, making the network output closer to the optimal solution.
[0096] Based on the updated LSTM-based action network and the updated LSTM-based evaluation network, the simulation training process and network update process are repeatedly executed at each time step; training is iterated continuously until the preset time step is reached. During training, the model's performance gradually improves, ultimately resulting in a fully trained LSTM-enhanced deep reinforcement learning model that can output a reasonable power allocation strategy based on the real-time status of the charging station.
[0097] For the training process of the LSTM action network, please refer to [reference needed]. Figure 3 , the state vector S t Hidden state H t-1 Unit state C t-1The input layer of the LSTM-based action network is fed into the LSTM layer; the cell state and hidden state store long-term and short-term information in the charging scenario, respectively; the LSTM layer includes a forget gate, an input gate, and an output gate; the forget gate is based on S... t and H t-1 Calculate a forgetting vector and C t-1 Multiplication determines which historical information to forget; the input gate determines the information to be updated, while the tanh function generates a new candidate value vector. Combining these two pieces of information updates the unit state C. t The output gate is based on the updated cell state C. t And the current input S t And the hidden state H from the previous moment t-1 This determines which information will be output to the hidden state H. t Output through the output layer; hide state H t The input is fed into a fully connected layer and transformed to obtain low-dimensional features, which are then input into the activation function σ of the output layer to obtain a normalized action value output. This action value is then multiplied by the rated power of the charging terminal to obtain the power allocated by the flexible switch to each charging terminal, thus generating the power allocation strategy A for the current moment. t This strategy includes the allocated power values for each charging terminal. Furthermore, the LSTM-based action network can be updated using backpropagation, with its loss function maximizing the output value of the LSTM-based evaluation network, i.e., maximizing the expected average user satisfaction value, expressed as follows:
[0098]
[0099] In the formula, M is the number of samples in the training; The output value of the LSTM-based evaluation network; θ represents the output value of the LSTM-based mobile network. A These are the network parameters for an LSTM-based mobile network.
[0100] For the training process of LSTM-based evaluation networks, please refer to [link / reference]. Figure 4 The LSTM-based evaluation network takes state information vectors, power allocation strategies, hidden states, and cell states as inputs. Its structure is similar to that of an LSTM-based action network. It combines cell states and hidden states to store long-term and short-term information about the charging scenario, respectively, and outputs the expected benefit of the supercharger station under that state and power allocation strategy. Alternatively, this network can be updated using backpropagation. Its loss function is the difference between the expected average user satisfaction index and the expected benefit output by the evaluation neural network, expressed as follows:
[0101]
[0102] In the formula, L is the loss function of the LSTM-based evaluation network; θ R The network parameters are for the LSTM-based evaluation network.
[0103] It is important to note that if the expected return obtained during training does not reach the target, the power allocation strategy needs to be adjusted. This can be achieved through policy update methods in reinforcement learning, such as SGD and Adagrad, to adjust the parameters of the action network and generate a new power allocation strategy. This new strategy is then input into the evaluation network for assessment until the expected return reaches the target. Furthermore, during training, this invention uses a simulated environment (charging service process simulation) to generate data; however, during actual execution, this invention uses real-world environment data and allocates power in real time.
[0104] This invention fully utilizes historical data, enabling the model to learn the operational patterns and characteristics of charging stations under different historical scenarios. Clearly defined action and reward mechanisms provide effective feedback for model training, guiding the model towards optimization goals. The use of an experience replay pool breaks the temporal correlation between data points, ensuring the independence of each sample taken from the pool, thus enhancing model stability and generalization ability. Efficient optimization of network parameters accelerates model convergence and improves training efficiency. By simulating real-world scenarios for model training, the complexities of charging stations in actual operation can be more realistically reflected. Furthermore, during the simulation, the model can continuously adjust and learn based on different states and actions, accumulating rich experience to better cope with various emergencies and different charging demands in practical applications.
[0105] S3. When the initial allocated power of each of the charging terminals is greater than the required charging power of the electric vehicles connected to each of the charging terminals, the excess allocated power of each of the charging terminals is balanced and optimized, and the optimized power allocation scheme is executed.
[0106] In one embodiment, step S3 includes:
[0107] The charging terminal whose initial allocated power is greater than the charging power required by the connected electric vehicle is designated as the target terminal, and the excess allocated power of the target terminal is evenly distributed to the remaining charging terminals to update the initial allocated power of the remaining charging terminals.
[0108] The relationship between the updated initial power allocation of the remaining charging terminals and the required charging power of the connected electric vehicles is detected in order to update the target terminal.
[0109] Based on the updated target terminals, the power allocation process and the target terminal update process are repeated until the final allocated power of all charging terminals is less than the charging power required by the electric vehicles they are connected to, and an optimized power allocation scheme is obtained and executed.
[0110] Specifically, all charging terminals are traversed, and charging terminals whose initial allocated power is greater than the charging power required by the connected electric vehicle are marked as target terminals. The total excess allocated power of all target terminals is calculated. The excess allocated power is the difference between the required charging power and the initial allocated power. The total excess allocated power is evenly distributed to the remaining charging terminals (i.e., non-target terminals). The initial allocated power of the remaining charging terminals is updated. The updated initial allocated power is equal to the original initial allocated power plus the allocated extra power.
[0111] The relationship between the updated initial power allocation of other charging terminals and the required charging power of the connected electric vehicles is detected. If some non-target terminals have an updated initial power allocation that is greater than the required charging power of the connected electric vehicles, these terminals are also added to the target terminal set to update the target terminals.
[0112] Based on the updated target terminal set, the power allocation process (i.e., calculating excess power and distributing it evenly) and the target terminal update process are repeated. In each repetition, the target terminals are re-determined and the power allocation is adjusted to make the power allocation more reasonable. The above process is repeated until the final allocated power of all charging terminals is less than the charging power required by the electric vehicles they are connected to (in practical applications, considering a certain error range, the difference between the allocated power and the required charging power can be set to be within an allowable range). At this point, the optimized power allocation scheme is obtained and executed to allocate power to each charging terminal for charging service.
[0113] For cases where the initial allocated power is not greater than the charging power required by the connected electric vehicle, the flexible switch is controlled to allocate power according to the initial allocated power.
[0114] This invention improves charging efficiency by redistributing excess power to other charging terminals in need, thus fully utilizing the power resources of the entire charging station. The use of an average distribution method for excess power ensures fairness and rationality. The optimized power allocation scheme better matches the power allocated to each charging terminal with the charging power demand of the connected electric vehicles, enhancing the stability of the charging station system. This scheme can dynamically adjust based on the relationship between the initial allocated power and the charging power demand of the electric vehicles, reducing the average additional waiting time for users and improving the user charging experience. In actual charging, the charging demand of electric vehicles may change; this dynamic power allocation optimization mechanism can respond promptly to these changes, ensuring that the charging station always operates efficiently and rationally.
[0115] This application proposes a flexible charging power allocation method for electric vehicles (EVs) to address the problem of how to rationally allocate the limited charging pile capacity to each charging terminal. This method utilizes historical data and charging process simulations to train an LSTM-enhanced deep reinforcement learning model. During the actual operation of the flexible charging station, when the charging power demand of the EV connected to a charging terminal exceeds the charging pile capacity, the LSTM-enhanced deep reinforcement learning model optimizes the charging power allocated to each charging terminal. The EV flexible charging power allocation method based on deep reinforcement learning technology provided by this invention optimizes the power allocation among charging terminals in a flexible charging station, reduces the average additional waiting time for users, and improves the user charging experience.
[0116] It should be noted that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order requirement for the execution of these steps, and they can be executed in other orders.
[0117] In another embodiment, such as Figure 8 As shown, a second aspect of the present invention provides a flexible charging power distribution system for electric vehicles, comprising:
[0118] The power quantization module 10 is used to collect the charging power demand information of electric vehicles connected to each charging terminal in real time, and quantify the total charging power demand of all electric vehicles connected to all charging terminals based on the charging power demand information.
[0119] The model processing module 20 is used to input the charging power demand information of each terminal into an LSTM-enhanced deep reinforcement learning model for processing when the total demand charging power exceeds the configuration capacity of the charging pile, and output the initial allocated power of each terminal; wherein, the objective of the LSTM-enhanced deep reinforcement learning model is to maximize the charging queuing satisfaction and charging delay satisfaction.
[0120] The scheme execution module 30 is used to balance and optimize the excess power of each charging terminal when the initial allocated power of each terminal is greater than the required charging power of the electric vehicle connected to each charging terminal, and then execute the optimized power allocation scheme.
[0121] It should be noted that each module in the aforementioned flexible charging power distribution system for electric vehicles can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module. For specific limitations regarding the flexible charging power distribution system for electric vehicles, please refer to the limitations regarding the flexible charging power distribution method for electric vehicles described above; both have the same function and role, and will not be repeated here.
[0122] A third aspect of the present invention provides an electronic device comprising:
[0123] Processor, memory, and bus;
[0124] The bus is used to connect the processor and the memory;
[0125] The memory is used to store operation instructions;
[0126] The processor is configured to execute instructions by calling the operation instructions, causing the processor to perform operations corresponding to a flexible charging power allocation method for electric vehicles as shown in the first aspect of this application.
[0127] In one alternative embodiment, an electronic device is provided, such as Figure 9 As shown, Figure 9 The illustrated electronic device 5000 includes a processor 5001 and a memory 5003. The processor 5001 and the memory 5003 are connected, for example, via a bus 5002. Optionally, the electronic device 5000 may also include a transceiver 5004. It should be noted that in practical applications, the transceiver 5004 is not limited to one type, and the structure of this electronic device 5000 does not constitute a limitation on the embodiments of this application.
[0128] Processor 5001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 5001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0129] Bus 5002 may include a path for transmitting information between the aforementioned components. Bus 5002 may be a PCI bus or an EISA bus, etc. Bus 5002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0130] The memory 5003 may be a ROM or other type of static storage device capable of storing static information and instructions, RAM or other type of dynamic storage device capable of storing information and instructions, or it may be an EEPROM, CD-ROM or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.
[0131] The memory 5003 is used to store application code that executes the scheme of this application, and its execution is controlled by the processor 5001. The processor 5001 is used to execute the application code stored in the memory 5003 to implement the content shown in any of the foregoing method embodiments.
[0132] Among them, electronic devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers.
[0133] The fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a flexible charging power allocation method for electric vehicles as shown in the first aspect of this application.
[0134] Another embodiment of this application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.
[0135] Furthermore, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0136] In summary, this invention relates to the field of electric vehicle charging technology, and discloses a flexible charging power allocation method, system, device, and medium for electric vehicles. It collects charging power demand information of electric vehicles connected to each charging terminal in real time to quantify the total charging power demand of all electric vehicles connected to all charging terminals. When the total charging power demand exceeds the charging pile configuration capacity, the charging power demand information is input into an LSTM-enhanced deep reinforcement learning model for processing, outputting the initial allocation power for each charging terminal. The objective of the LSTM-enhanced deep reinforcement learning model is to maximize charging queuing satisfaction and charging delay satisfaction. When the initial allocation power is greater than the charging power demand of the electric vehicles connected to its charging terminal, the excess allocation power of each charging terminal is balanced and optimized and executed. Based on deep reinforcement learning technology, the limited charging pile configuration capacity is rationally allocated to each charging terminal, improving resource utilization.
[0137] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0138] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.
Claims
1. A flexible charging power allocation method for electric vehicles, characterized in that, include: Real-time collection of charging power demand information of electric vehicles connected to each charging terminal, and quantification of the total charging power demand of all electric vehicles connected to all charging terminals based on the charging power demand information. When the total demand for charging power exceeds the configuration capacity of the charging pile, the charging power demand information of each terminal is input into an LSTM-enhanced deep reinforcement learning model for processing, and the initial allocated power of each terminal is output; wherein, the objective of the LSTM-enhanced deep reinforcement learning model is to maximize the charging queuing satisfaction and the charging delay satisfaction. When the initial allocated power of each charging terminal is greater than the required charging power of the electric vehicle connected to each charging terminal, the excess allocated power of each charging terminal is balanced and optimized, and the optimized power allocation scheme is executed.
2. The method for flexible charging power allocation of electric vehicles according to claim 1, characterized in that, The charging power demand information includes the rated charging power and the remaining charging demand; wherein... The process of quantifying the total charging power demand of all electric vehicles connected to all charging terminals based on the charging power demand information of each of the aforementioned charging terminals includes: The actual power demand of each electric vehicle is determined based on the rated charging power and the remaining charging demand. The total required charging power is determined based on the actual power demand of each item.
3. The method for flexible charging power allocation of electric vehicles according to claim 1, characterized in that, The objectives of the LSTM-enhanced deep reinforcement learning model also include minimizing timeout penalties and queue length penalties; wherein, the charging queue satisfaction is the ratio of the queuing time of an electric vehicle before connecting to a charging terminal to the charging time required at its rated charging power; the charging delay satisfaction is the ratio of the charging time paid by the electric vehicle at the charging terminal to the charging time required at its rated charging power; the timeout penalty is a linear penalty when the total charging time of the electric vehicle exceeds a preset multiple of the rated charging time; and the queue length penalty is a quadratic penalty when the number of vehicles in the queue exceeds the maximum capacity.
4. The method for flexible charging power allocation of electric vehicles according to claim 3, characterized in that, The LSTM-enhanced deep reinforcement learning model includes an LSTM-based action network; wherein, The step of inputting the charging power demand information of each of the above into an LSTM-enhanced deep reinforcement learning model for processing, and outputting the initial allocated power of each of the charging terminals, includes: The charging power demand information is converted into a real-time state vector, and a state information vector is constructed based on the real-time state vector and the cell state and hidden state information output by the LSTM at the previous time step. The state information vector is input into the LSTM-based action network for processing, and the initial allocated power of each charging terminal is output.
5. The method for flexible charging power allocation of electric vehicles according to claim 4, characterized in that, The LSTM-based mobile network comprises an input layer, an LSTM layer, a fully connected layer, and an output layer; wherein, The step of inputting the state information vector into the LSTM-based action network for processing and outputting a power allocation strategy includes: The state information vector is input to the LSTM layer through the input layer; The long short-term state dependency features in the state information vector are extracted by the LSTM layer to obtain the hidden state features at the current time and input to the fully connected layer. The hidden state features are transformed by the fully connected layer to obtain low-dimensional features, which are then processed by the output layer to output normalized power allocation action values. The power allocation strategy is generated based on the normalized power allocation action value and the rated power of each charging terminal.
6. The method for flexible charging power allocation of electric vehicles according to claim 5, characterized in that, The LSTM-enhanced deep reinforcement learning model also includes an LSTM-based evaluation network; wherein, The training process of the LSTM-enhanced deep reinforcement learning model includes: Obtain historical charging power demand information of electric vehicles connected to each of the charging terminals, and convert each of the historical charging power demand information into a historical demand state vector as a state space; The normalized power allocation action value of each charging terminal is used as the action space, and a reward function is constructed based on the objective of the LSTM-enhanced deep reinforcement learning model. Initialize the parameters of the experience replay pool, the LSTM-based action network, and the LSTM-based evaluation network, and simulate the operation of the charging station from the start of the charging peak to obtain the current state; At each time step, based on the current state, the current action is determined by the LSTM-based action network to perform a charging service process simulation, and the current reward and the next state are calculated according to the reward function and the LSTM-based evaluation network, and stored in the experience replay pool together with the current state and the current action. A batch of datasets is sampled from the experience replay pool, and the LSTM-based action network and the LSTM-based evaluation network are updated using the Adam optimizer based on the sampled datasets. Based on the updated LSTM-based action network and the updated LSTM-based evaluation network, the simulated training process and network update process are repeatedly executed at each time step until the preset time step is reached, resulting in a trained LSTM-enhanced deep reinforcement learning model.
7. The method for flexible charging power allocation of electric vehicles according to claim 1, characterized in that, When the initial allocated power of each charging terminal is greater than the charging power required by the electric vehicles connected to each charging terminal, the excess allocated power of each charging terminal is balanced and optimized, and the optimized power allocation scheme is executed, including: The charging terminal whose initial allocated power is greater than the charging power required by the connected electric vehicle is designated as the target terminal, and the excess allocated power of the target terminal is evenly distributed to the remaining charging terminals to update the initial allocated power of the remaining charging terminals. The relationship between the updated initial power allocation of the remaining charging terminals and the required charging power of the connected electric vehicles is detected in order to update the target terminal. Based on the updated target terminals, the power allocation process and the target terminal update process are repeated until the final allocated power of all charging terminals is less than the charging power required by the electric vehicles they are connected to, and an optimized power allocation scheme is obtained and executed.
8. A flexible charging power distribution system for electric vehicles, characterized in that, include: The power quantization module is used to collect the charging power demand information of electric vehicles connected to each charging terminal in real time, and quantify the total charging power demand of all electric vehicles connected to all charging terminals based on the charging power demand information. The model processing module is used to input the charging power demand information of each terminal into an LSTM-enhanced deep reinforcement learning model for processing when the total demand charging power exceeds the configuration capacity of the charging pile, and output the initial allocated power of each terminal; wherein, the objective of the LSTM-enhanced deep reinforcement learning model is to maximize the charging queuing satisfaction and charging delay satisfaction. The scheme execution module is used to balance and optimize the excess power of each charging terminal when the initial allocated power of each terminal is greater than the required charging power of the electric vehicle connected to each charging terminal, and then execute the optimized power allocation scheme.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the electric vehicle flexible charging power allocation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the device containing the computer-readable storage medium executes the computer program, it implements the electric vehicle flexible charging power allocation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Charging and energy supply optimization method and device for charging management system
CN112874369A
Ordered charging method and device for charging pile and storage medium
CN120073838A
Control device, program, and power management system
JP2023027620A
Electric vehicle charging optimization based on predictive analytics utilizing machine learning
US20220055496A1
Demand flexibility optimizing scheduler for ev charging and controlling appliances
US20220144121A1
Cited By
Reinforcement learning-based charging and swapping grid distribution method for charging and swapping energy storage cabinet
CN121939599A