User-oriented electric vehicle intelligent charging method, system, device and medium

CN120552668BActive Publication Date: 2026-08-11BEIJING TENGINEER AIOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,现有的智能充电方法普遍存在灵活性不足、优化效率低或未充分考虑用户个性化需求等问题,此外,现有研究多集中于降低充电成本,而对用户历史充电习惯、太阳能自消纳能力等因素考虑不足,其奖励设置较为固定、单一,难以全面反映用户历史充电行为与实际需求之间的匹配度

Benefits of technology

[0035] This invention presents a user-oriented intelligent charging method for electric vehicles. It constructs a multi-objective function with the goals of minimizing electricity costs, maximizing user charging behavior consistency, and maximizing photovoltaic power generation self-consumption. The intelligent charging process is modeled as a Markov decision process, and the multi-objective function is transformed into a reward function suitable for training the agent's action strategy. A deep Q-network algorithm is then used to train the model, allowing the agent to learn the optimal intelligent charging strategy. By introducing user charging behavior consistency and photovoltaic power generation self-consumption as optimization objectives, this method constructs an adaptive reward adjustment mechanism based on historical user charging behavior and the real-time charging environment. This enables the agent to learn an optimal charging strategy that balances electricity costs, user charging preferences, and photovoltaic power generation self-consumption, better aligning with user charging habits and promoting photovoltaic power generation self-consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120552668B_ABST
    Figure CN120552668B_ABST
Patent Text Reader

Abstract

This invention discloses a user-oriented intelligent charging method, system, device, and medium for electric vehicles. The method constructs a multi-objective function with the goals of minimizing electricity costs, maximizing user charging behavior consistency, and maximizing photovoltaic power generation self-consumption. The intelligent charging process is modeled as a Markov decision process, and the multi-objective function is transformed into a reward function suitable for training the agent's action strategy. A deep Q-network algorithm is used to train the model, allowing the agent to learn the optimal intelligent charging strategy. By introducing user charging behavior consistency and photovoltaic power generation self-consumption as optimization objectives, an adaptive reward adjustment mechanism based on historical user charging behavior and the real-time charging environment is constructed. This enables the optimal charging strategy learned by the agent to balance electricity costs, user charging propensity, and photovoltaic power generation self-consumption, better aligning with user charging habits and promoting photovoltaic power generation self-consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electric vehicle charging technology, and in particular, to a user-oriented intelligent charging method and system for electric vehicles, an electronic device, and a computer-readable storage medium. Background Technology

[0002] With a high proportion of distributed energy resources being integrated into residential loads, such as electric vehicles, the operational pressure and uncertainty of the distribution network will increase significantly. Traditional disordered charging patterns for electric vehicles often fail to effectively utilize time-varying electricity prices and the volatility of renewable energy, leading to excessively high peak grid loads and energy waste. Furthermore, existing smart charging strategies primarily rely on mathematical programming, heuristic algorithms, or machine learning methods. However, these methods generally suffer from insufficient flexibility, low optimization efficiency, or inadequate consideration of personalized user needs. Moreover, current research largely focuses on reducing charging costs, neglecting factors such as users' historical charging habits and solar energy self-consumption capacity. Reward settings are relatively fixed and singular, failing to comprehensively reflect the match between users' historical charging behavior and actual demand. Therefore, there is an urgent need for a dynamic smart charging control method that integrates real-time market signals, historical charging data, and renewable energy utilization rates to better meet user needs and optimize grid resource allocation. Summary of the Invention

[0003] This invention provides a user-oriented intelligent charging method and system for electric vehicles, an electronic device, and a computer-readable storage medium, which can balance three factors: electricity cost, user charging tendency, and photovoltaic power generation self-consumption, making it more in line with users' charging habits and promoting photovoltaic self-consumption.

[0004] According to one aspect of the present invention, a user-oriented intelligent charging method for electric vehicles is provided, comprising the following:

[0005] A multi-objective function is constructed with the objectives of minimizing electricity costs, maximizing the consistency of user charging behavior, and maximizing the self-consumption of photovoltaic power generation. Constraints are set, where user charging behavior consistency refers to the consistency between the charging period recommended by the smart charging strategy of electric vehicles and the historical charging period of electric vehicles, and photovoltaic power generation self-consumption refers to the amount of photovoltaic power generation used for charging electric vehicles.

[0006] The intelligent charging process of electric vehicles is modeled as a Markov decision process, and the multi-objective function is transformed into a reward function suitable for training the action policy of the agent.

[0007] Historical charging data of electric vehicles were acquired to construct a training dataset, and the model was trained using a deep Q-network algorithm.

[0008] Collect current environmental state data related to electric vehicle charging, input it into the trained model, and execute the corresponding charging operation according to the action strategy output by the model.

[0009] Furthermore, the expression for the multi-objective function is:

[0010]

[0011] Among them, f obj Represents a multi-objective function, C E U represents the electricity cost for electric vehicle users. ch This indicates that the user's charging behavior is consistent. λ1, λ2, and λ3 represent the self-consumption of photovoltaic power generation, and λ1, λ2, and λ3 represent weighting factors.

[0012] Furthermore, the agent's cumulative reward function is:

[0013]

[0014] Among them, R all Let T represent the cumulative reward of the agent, and r represent one day. t Indicates the relationship with the current state s t Related rewards, a t This indicates the action at time t, specifically whether the electric vehicle is charging at time t. This indicates sub-rewards related to electricity costs. Indicates daytime charging rewards. Let β represent the nighttime charging reward, and let β represent the activation variable, which is 1 when electric vehicle users tend to charge during the day and 0 when they tend to charge at night. Sub-rewards indicating incentives for promoting photovoltaic self-generation. This refers to a reward based on the range of electric vehicles. This is a sub-reward that strongly penalizes the charging strategy when the daily cumulative activation count of the battery exceeds a preset threshold.

[0015] Furthermore, the calculation formulas for daytime charging rewards and nighttime charging rewards are as follows:

[0016]

[0017] Among them, f EV (t) represents the probability of the vehicle charging at time t, and τ and ω represent the reward factors for auxiliary training.

[0018] Furthermore, the calculation formulas for the sub-rewards related to electricity costs and the sub-rewards for promoting photovoltaic self-generation are as follows:

[0019]

[0020] Among them, pr t This represents the real-time electricity price at time t. Δt represents the charging power of the electric vehicle at time t, and Δt represents the preset time step. Table t shows the photovoltaic power generation at time t.

[0021] Furthermore, the formula for calculating the sub-rewards of the strongly penalized charging strategy is as follows:

[0022]

[0023] in, This indicates the daily cumulative number of battery activations. This indicates the preset daily cumulative activation threshold. η represents the maximum allowed number of activations per day. p This represents the penalty factor.

[0024] Furthermore, the state space of a Markov decision process can be represented as:

[0025]

[0026] Among them, S t pr represents the state vector at time t. t This represents the real-time electricity price at time t. Table t shows the photovoltaic power generation at time t. This represents the load of all household appliances at time t, excluding those consumed by electric vehicles. SOC represents the cumulative energy consumption of the electric vehicle at time t. t This represents the remaining battery charge of the electric vehicle at time t.

[0027] In addition, the present invention also provides a user-oriented intelligent charging system for electric vehicles, comprising:

[0028] The objective function construction module is used to construct a multi-objective function with the objectives of minimizing electricity costs, maximizing the consistency of user charging behavior, and maximizing the self-consumption of photovoltaic power generation, and to set constraints. Among them, the consistency of user charging behavior refers to the consistency between the charging period recommended by the smart charging strategy of electric vehicles and the historical charging period of electric vehicles, and the self-consumption of photovoltaic power generation refers to the amount of photovoltaic power generation used for charging electric vehicles.

[0029] The Markov modeling module is used to model the intelligent charging process of electric vehicles as a Markov decision process and transform multi-objective functions into reward functions suitable for training agent action policies.

[0030] The model training module is used to acquire historical charging data of electric vehicles to build a training dataset and to train the model using a deep Q-network algorithm.

[0031] The intelligent charging control module is used to collect current environmental status data related to electric vehicle charging, input it into the trained model, and execute corresponding charging operations according to the action strategy output by the model.

[0032] In addition, the present invention also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the processor executes the steps of the method described above by calling the computer program stored in the memory.

[0033] In addition, the present invention also provides a computer-readable storage medium for storing a computer program for user-oriented intelligent charging of electric vehicles, wherein the computer program executes the steps of the method described above when running on a computer.

[0034] The present invention has the following beneficial effects:

[0035] This invention presents a user-oriented intelligent charging method for electric vehicles. It constructs a multi-objective function with the goals of minimizing electricity costs, maximizing user charging behavior consistency, and maximizing photovoltaic power generation self-consumption. The intelligent charging process is modeled as a Markov decision process, and the multi-objective function is transformed into a reward function suitable for training the agent's action strategy. A deep Q-network algorithm is then used to train the model, allowing the agent to learn the optimal intelligent charging strategy. By introducing user charging behavior consistency and photovoltaic power generation self-consumption as optimization objectives, this method constructs an adaptive reward adjustment mechanism based on historical user charging behavior and the real-time charging environment. This enables the agent to learn an optimal charging strategy that balances electricity costs, user charging preferences, and photovoltaic power generation self-consumption, better aligning with user charging habits and promoting photovoltaic power generation self-consumption.

[0036] In addition, the user-oriented intelligent charging system for electric vehicles of the present invention also has the above-mentioned advantages.

[0037] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The invention will now be described in further detail with reference to the figures. Attached Figure Description

[0038] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0039] Figure 1This is a flowchart illustrating a user-oriented intelligent charging method for electric vehicles according to a preferred embodiment of this application.

[0040] Figure 2 This is a schematic diagram of the module structure of a user-oriented intelligent charging system for electric vehicles, according to another embodiment of this application. Detailed Implementation

[0041] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0042] Reference Figure 1 A preferred embodiment of this application provides a user-oriented intelligent charging method for electric vehicles, including the following:

[0043] Step S1: Construct a multi-objective function with the objectives of minimizing electricity costs, maximizing the consistency of user charging behavior, and maximizing the self-consumption of photovoltaic power generation, and set constraints. Among them, the consistency of user charging behavior refers to the consistency between the charging period recommended by the smart charging strategy of electric vehicles and the historical charging period of electric vehicles, and the self-consumption of photovoltaic power generation refers to the amount of photovoltaic power generation used for charging electric vehicles.

[0044] Step S2: Model the intelligent charging process of electric vehicles as a Markov decision process, and transform the multi-objective function into a reward function suitable for training the agent's action policy;

[0045] Step S3: Obtain historical charging data of electric vehicles to construct a training dataset, and use a deep Q-network algorithm to train the model;

[0046] Step S4: Collect current environmental state data related to electric vehicle charging, input it into the trained model, and execute the corresponding charging operation according to the action strategy output by the model.

[0047] It is understood that the user-oriented intelligent charging method for electric vehicles in this embodiment constructs a multi-objective function with the goals of minimizing electricity costs, maximizing user charging behavior consistency, and maximizing photovoltaic power generation self-consumption. The intelligent charging process is modeled as a Markov decision process, and the multi-objective function is transformed into a reward function suitable for training the agent's action strategy. Then, a deep Q-network algorithm is used to train the model, allowing the agent to learn the optimal intelligent charging strategy for electric vehicles. This method introduces user charging behavior consistency and photovoltaic power generation self-consumption as optimization objectives, constructing an adaptive reward adjustment mechanism based on users' historical charging behavior and the real-time charging environment. This enables the optimal charging strategy learned by the agent to balance electricity costs, user charging tendencies, and photovoltaic power generation self-consumption, better aligning with users' charging habits and promoting photovoltaic power generation self-consumption.

[0048] In step S1, the user-oriented intelligent charging strategy for electric vehicles of this invention considers not only minimizing electric vehicle charging costs but also the historical electricity consumption patterns of electric vehicle users and the amount of solar photovoltaic self-generated power. Specifically, a multi-objective function is constructed with the objectives of minimizing electricity costs, maximizing the consistency of user charging behavior, and maximizing the self-consumption of photovoltaic power generation. The expression of the multi-objective function is as follows:

[0049]

[0050] Among them, f obj Represents a multi-objective function, C E U represents the electricity cost for electric vehicle users. ch This indicates that the user's charging behavior is consistent. λ1, λ2, and λ3 represent the self-consumption of photovoltaic (PV) power generation, and λ3 represent weighting factors. User charging behavior consistency refers to the consistency between the charging periods recommended by the smart charging strategy for electric vehicles and the historical charging periods of electric vehicles. It measures the degree to which users' charging habits remain unchanged under the suggestions of the smart charging strategy. PV self-consumption refers to the amount of PV power generated for charging electric vehicles. Furthermore, when PV power generation is insufficient to meet the load of all households, it is assumed that the PV power generation facilities will first meet the electricity demand of electric vehicles.

[0051] In addition, the following constraints are imposed on the electric vehicle charging process:

[0052] 1) Capacitor capacitance constraint, which can be expressed as: SOC min ≤SOC t ≤SOC max , Among them, SOC min State of Charge (SOC) indicates the minimum charge required for the battery to operate. tState of Charge (SOC) represents the remaining battery charge of the electric vehicle at time t. max This indicates the battery's maximum rated capacity, and T represents one day.

[0053] 2) The initial charge constraint can be expressed as: Among them, SOC init η represents the initial charge of the battery. ch This represents the charging efficiency of an electric vehicle charging station, where 0 ≤ η. ch ≤1, This represents the daily energy consumption of electric vehicle users in historical data, expressed in kWh. The nominal capacity of the electric vehicle battery is expressed in kWh. The start and end times of charging are unknown for different users; therefore, the state of charge (SOC) distribution of the electric vehicle battery is unknown. This invention primarily considers electric vehicle users who experience one complete charging cycle per day. Therefore, the aforementioned constraints require the electric vehicle to complete charging at the end of the set time step.

[0054] 3) Charging power constraint, which can be expressed as: in, This represents the charging power of the electric vehicle at time t. Indicates the nominal power rating. This indicates a relatively low charging rating, i.e. Less than SoC lim Indicates the SOC state boundary value, for example, SoC lim =0.9. When the battery's SOC is not higher than 0.9, a relatively high nominal power rating is used for charging; when the battery's SOC is higher than 0.9, a relatively low charging rating is used for charging.

[0055] Furthermore, in step S2, the electric vehicle charging process is modeled as a finite discrete-time Markov decision process (hereinafter referred to as MDP) with a step size of Δt minutes. The step size Δt can be adjusted according to the time precision required for charging control; generally, Δt is set to 15 minutes. That is, when the electric vehicle is connected to the charging station, the agent will decide whether to charge in the next time period with a time granularity of 15 minutes. In deep reinforcement learning, the agent interacts with the external environment in the form of a neural network, learning independently from past experiences through a "trial and error" process. At each time step, the agent receives the current state of the environment, and the neural network, under a specified training framework, is responsible for approximating the expected reward (i.e., Q-value) for each possible action. The Q-value represents the expected future reward that the agent can obtain after taking each possible action in the current state. Finally, the agent selects an action and transitions to a new state, receiving a positive or negative reward based on the value of the selected action. Through this iterative process, the agent's goal is to select the action that maximizes the accumulated future reward.

[0056] Specifically, MDP is defined as a quadruple (S, A, R, P), where S represents the state space, A represents the action space, R represents the set of rewards that the agent can obtain at each time step t, and P represents the transition probability of the agent from one state s to another state s'.

[0057] The state space S can be represented as: S t pr represents the state vector at time t. t This represents the real-time electricity price at time t. Table t shows the photovoltaic power generation at time t. This represents the load of all household appliances at time t, excluding those consumed by electric vehicles. If applied to non-household electric vehicle charging, this value is 0. SOC represents the cumulative energy consumption of the electric vehicle at time t. t This represents the remaining battery charge of the electric vehicle at time t.

[0058] In addition, the action space can be represented as: Among them, a t This indicates the action at time t, specifically whether the electric vehicle is charging at time t. This indicates the charging state at time t, where 1 indicates charging and 0 indicates no charging.

[0059] Furthermore, this invention improves upon traditional deep reinforcement learning models by introducing an adaptive reward adjustment mechanism based on users' historical charging behavior and the real-time charging environment. The aim is to propose an optimal intelligent charging strategy for electric vehicles based on user preferences. Therefore, the multi-objective function constructed in step S1 is transformed into a reward function suitable for training the agent's action strategy. The agent's cumulative reward function is:

[0060]

[0061] Among them, R all Let T represent the cumulative reward of the agent, and r represent one day. t Indicates the relationship with the current state s t Related rewards, a t This indicates the action at time t, specifically whether the electric vehicle is charging at time t. This indicates sub-rewards related to electricity costs. Indicates daytime charging rewards. Let β represent the nighttime charging reward, and let β represent the activation variable, which is 1 when electric vehicle users tend to charge during the day and 0 when they tend to charge at night. Sub-rewards indicating incentives for promoting photovoltaic self-generation. This refers to a reward based on the range of electric vehicles. This is a sub-reward that strongly penalizes the charging strategy when the daily cumulative activation count of the battery exceeds a preset threshold.

[0062] The calculation formulas for daytime charging rewards and nighttime charging rewards are as follows:

[0063]

[0064] Among them, f EV (t) represents the probability of the vehicle charging at time t, and τ and ω represent the reward factors for auxiliary training, which are adjusted according to actual applications.

[0065] It is understood that this invention uses the variable β in the reward function to distinguish between reward systems with a daytime charging tendency or a nighttime charging tendency, and sub-rewards. and The goal is to shift the charging process to the user's historical charging time period, penalize charging behavior during periods when the user has a low tendency to charge, and make the charging action strategy more in line with the user's charging habits, thus realizing user-oriented intelligent charging control.

[0066] In addition, the calculation formulas for the sub-rewards related to electricity costs and the sub-rewards for promoting photovoltaic self-generation are as follows:

[0067]

[0068]

[0069] Among them, pr t This represents the real-time electricity price at time t. Δt represents the charging power of the electric vehicle at time t, and Δt represents the preset time step. Table t shows the photovoltaic power generation at time t.

[0070] It is understood that the present invention also introduces sub-rewards related to promoting photovoltaic self-generation and electricity costs into the reward function, so that the optimal charging strategy learned by the agent can balance the three factors of electricity costs, user charging tendency and photovoltaic self-consumption, which is more in line with users' charging habits and at the same time conducive to promoting photovoltaic self-consumption.

[0071] Furthermore, the formula for calculating the sub-rewards of the strongly penalized charging strategy is as follows:

[0072]

[0073] in, This indicates the cumulative number of battery activations per day, representing the total number of times the battery transitions from a non-charging state to a charging state within a day. This indicates the preset daily cumulative activation threshold. η represents the maximum allowed number of activations per day. p This represents the penalty factor, and its value is adjusted according to the actual application.

[0074] It is understood that the present invention also introduces a sub-reward to penalize multiple charging in the reward function. When the cumulative number of battery activations per day exceeds a preset threshold, a penalty is imposed to encourage the charging strategy to reduce the number of charging times, which is beneficial to improving battery life.

[0075] In addition, electric vehicle range sub-rewards This sub-reward incentivizes the agent to fully charge the electric vehicle before the end of each day. The agent decides whether to charge based on how close the electric vehicle's current State of Charge (SOC) is to its ideal SOC. The daily power consumption of the electric vehicle is also considered. Compared with user historical power consumption Similar, generally with a deviation of ±5%, that is k represents the penalty factor, the value of which is adjusted according to the actual application. E target This indicates the target power consumption to meet the user's electricity demand.

[0076] In addition, in step S3, historical charging data of electric vehicles is acquired, including the time when the vehicle connects to the charging pile, the SoC level of the vehicle when connecting to the charging pile, the charging power curve during the period when the vehicle is connected to the charging pile, the time when the electric vehicle leaves the charging pile, the SoC level of the vehicle when leaving the charging pile, real-time electricity price data, etc., and a training dataset is constructed. The reward factor in the reward function of the Markov decision process model constructed in step S2 above needs to be adjusted and determined through training. Therefore, this invention uses the Deep Q-Network algorithm (DQN) to learn the intelligent charging strategy of electric vehicles. The DQN method is a value-based algorithm that aims to estimate the expected value (i.e., Q value) of each state-action pair. Based on the impact of the action on the future environmental state, the Q value is defined as the sum of all possible new states s′, specifically the product of the discounted reward R(s, a, s′) for transitioning from state s to state s′ after taking action a and the transition probability R(s, a, s′), which can be expressed as: Q(s, a)=∑P(s, a, s′)[R(s, a, s′)+γV π [(s′)], Q() represents the action value function, V π () represents the state value function, γ represents the discount factor used to determine the importance of future rewards, and π represents the action policy. Q(s′) describes the maximum reward obtainable in state s′; therefore, the goal of DQN is to maximize... r represents the immediate reward obtained after taking action a from state s and transitioning to state s′.

[0077] The training process of the DQN algorithm for the Markov decision process model is as follows:

[0078] (A) Initialize the Q network parameters θ, the parameters of the target Q network Q′(s, a; θ′) (θ′=θ), the experience replay pool D, which is used to store the experience gained by the agent in interacting with the environment, the learning rate α, the discount factor γ, the exploration rate ò (used for the balance between exploration and exploitation), the minimum batch size B (the number of samples drawn from the experience pool), and the frequency of updating the target Q network (usually once every certain number of steps).

[0079] (B) In each training round, perform the following steps:

[0080] a) At the start of the round, set the initial state s0;

[0081] b) At each time t, repeat the following steps:

[0082] b1) Action Selection: Select an action based on the following strategy:

[0083]

[0084] b2) Perform the action: Perform the selected action in the environment. t And observe the next state s t+1 and reward r t ;

[0085] b3) Storing experience: storing experience (s) t a t r t s t+1 Stored in the experience replay pool D;

[0086] b4) Sampling from the experience replay pool: Randomly sample a small batch B of experience (s) i a i r i s i+1 );

[0087] b5) Calculate the target Q value: For each sampled sample, calculate the target Q value:

[0088]

[0089] b6) Update the Q-network: Minimize the loss function and update the parameters of the Q-network. The loss function is:

[0090]

[0091] b7) Update the target Q network: Every certain number of steps, copy the parameters θ of the Q network to the parameters θ′ of the target Q network.

[0092] (C) As training progresses, the exploration rate is gradually reduced and the utilization rate is increased, thereby enabling the agent to rely more on the learned strategies.

[0093] (D) When training is complete, the agent should have learned the optimal smart electric vehicle charging strategy that maximizes cumulative rewards and balances factors such as cost, user charging propensity, and spontaneous consumption of solar energy.

[0094] Additionally, in step S4, real-time environmental state data related to electric vehicle charging is collected, such as current battery level, electricity cost, household solar power generation, and user charging preferences. This data is input into a pre-trained Q-network, which outputs a Q-value for each possible action. Based on the Q-value, the agent selects an action (e.g., starting charging or remaining idle) to maximize future rewards, thus performing the corresponding charging operation. For example, if starting charging is selected, the agent decides whether to charge efficiently or when electricity costs are low, aligning with the user's historical charging habits. If remaining idle is selected, charging is paused, waiting for the next time period.

[0095] In addition, such as Figure 2 As shown, another embodiment of the present invention also provides a user-oriented intelligent charging system for electric vehicles, preferably employing the user-oriented intelligent charging method for electric vehicles as described above, comprising:

[0096] The objective function construction module is used to construct a multi-objective function with the objectives of minimizing electricity costs, maximizing the consistency of user charging behavior, and maximizing the self-consumption of photovoltaic power generation, and to set constraints. Among them, the consistency of user charging behavior refers to the consistency between the charging period recommended by the smart charging strategy of electric vehicles and the historical charging period of electric vehicles, and the self-consumption of photovoltaic power generation refers to the amount of photovoltaic power generation used for charging electric vehicles.

[0097] The Markov modeling module is used to model the intelligent charging process of electric vehicles as a Markov decision process and transform multi-objective functions into reward functions suitable for training agent action policies.

[0098] The model training module is used to acquire historical charging data of electric vehicles to build a training dataset and to train the model using a deep Q-network algorithm.

[0099] The intelligent charging control module is used to collect current environmental status data related to electric vehicle charging, input it into the trained model, and execute corresponding charging operations according to the action strategy output by the model.

[0100] It is understood that the user-oriented intelligent charging system for electric vehicles in this embodiment constructs a multi-objective function with the goals of minimizing electricity costs, maximizing user charging behavior consistency, and maximizing photovoltaic power generation self-consumption. The intelligent charging process is modeled as a Markov decision process, and the multi-objective function is transformed into a reward function suitable for training the agent's action strategy. Then, a deep Q-network algorithm is used to train the model, allowing the agent to learn the optimal intelligent charging strategy for electric vehicles. This method introduces user charging behavior consistency and photovoltaic power generation self-consumption as optimization objectives, constructing an adaptive reward adjustment mechanism based on users' historical charging behavior and the real-time charging environment. This enables the optimal charging strategy learned by the agent to balance electricity costs, user charging tendencies, and photovoltaic power generation self-consumption, better aligning with users' charging habits and promoting photovoltaic power generation self-consumption.

[0101] In addition, another embodiment of the present invention provides an electronic device including a processor and a memory, wherein the memory stores a computer program, and the processor executes the steps of the method described above by calling the computer program stored in the memory.

[0102] In addition, another embodiment of the present invention provides a computer-readable storage medium for storing a computer program for user-oriented intelligent charging of electric vehicles, wherein the computer program executes the steps of the method described above when run on a computer.

[0103] Common computer-readable storage media include: floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tape, any other physical media with perforated patterns, random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), flash erasable programmable read-only memory (FLASH-EPROM), any other memory chips or cartridges, or any other media readable by a computer. Instructions may further be transmitted or received by a transmission medium. The term transmission medium can include any tangible or intangible medium used to store, encode, or carry instructions for machine execution, and includes digital or analog communication signals or intangible media that facilitate communication of such instructions. Transmission media include coaxial cables, copper wires, and optical fibers, which contain conductors for transmitting a bus of computer data signals.

[0104] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0105] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.

[0106] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0107] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0108] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0109] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

[0110] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A user-oriented intelligent charging method for electric vehicles, characterized in that, Includes the following: A multi-objective function is constructed with the objectives of minimizing electricity costs, maximizing the consistency of user charging behavior, and maximizing the self-consumption of photovoltaic power generation. Constraints are set, where user charging behavior consistency refers to the consistency between the charging periods recommended by the intelligent charging strategy for electric vehicles and the historical charging periods of electric vehicles, and photovoltaic power generation self-consumption refers to the amount of photovoltaic power generated for charging electric vehicles. The expression of the multi-objective function is as follows: ; in, f obj Represents a multi-objective function. C E This indicates the electricity cost for electric vehicle users. U ch This indicates that the user's charging behavior is consistent. This indicates the amount of photovoltaic power generated for self-consumption. Indicates the weighting factor; The intelligent charging process of electric vehicles is modeled as a Markov decision process, and the multi-objective function is transformed into a reward function suitable for training the agent's action policy; the agent's cumulative reward function is: ; in, R all This represents the cumulative reward of the intelligent agent. T Indicates one day. r t Indicates the current state s t Related rewards, a t express t The action at a moment, that is t Whether the electric vehicle is charging at any time. This indicates sub-rewards related to electricity costs. Indicates daytime charging rewards. This indicates a nighttime charging reward. β This represents the activation variable, which is 1 when electric vehicle users tend to charge during the day and 0 when they tend to charge at night. Sub-rewards indicating incentives for promoting photovoltaic self-generation. This refers to a reward based on the range of electric vehicles. This represents a sub-reward for the strong penalty charging strategy when the daily cumulative activation count of the battery exceeds a preset threshold; the calculation formula for the sub-reward of the strong penalty charging strategy is as follows: ; in, This indicates the daily cumulative number of battery activations. This indicates the preset daily cumulative activation threshold. This indicates the maximum allowed number of activations per day. Indicates the penalty factor; Historical charging data of electric vehicles were acquired to construct a training dataset, and the model was trained using a deep Q-network algorithm. Collect current environmental state data related to electric vehicle charging, input it into the trained model, and execute the corresponding charging operation according to the action strategy output by the model.

2. The user-oriented intelligent charging method for electric vehicles as described in claim 1, characterized in that, The formulas for calculating daytime charging rewards and nighttime charging rewards are as follows: ; ; in, f EV ( t ) indicates that the vehicle is in t The probability of charging at any given moment. and This represents the reward factor used to assist training.

3. The user-oriented intelligent charging method for electric vehicles as described in claim 1, characterized in that, The calculation formulas for the sub-rewards related to electricity costs and the sub-rewards for promoting photovoltaic self-generation are as follows: ; ; in, pr t express t Real-time electricity price at any given moment Indicates electric vehicles t Charging power at any time Indicates the preset time step. surface t The power generation capacity of photovoltaics at any given time.

4. The user-oriented intelligent charging method for electric vehicles as described in claim 1, characterized in that, The state space of a Markov decision process can be represented as: ; in, S t express t The state vector at time t, pr t express t Real-time electricity price at any given moment surface t The power generation capacity of photovoltaics at all times. express t All household appliance loads, excluding those consumed by electric vehicles, at all times. express t The cumulative energy consumption of electric vehicles at all times. SOC t Indicates electric vehicles t The remaining battery power at any given time.

5. A user-oriented intelligent charging system for electric vehicles, employing the user-oriented intelligent charging method for electric vehicles as described in any one of claims 1 to 4, characterized in that, include: The objective function construction module is used to construct a multi-objective function with the objectives of minimizing electricity costs, maximizing the consistency of user charging behavior, and maximizing the self-consumption of photovoltaic power generation, and to set constraints. Among them, the consistency of user charging behavior refers to the consistency between the charging period recommended by the smart charging strategy of electric vehicles and the historical charging period of electric vehicles, and the self-consumption of photovoltaic power generation refers to the amount of photovoltaic power generation used for charging electric vehicles. The Markov modeling module is used to model the intelligent charging process of electric vehicles as a Markov decision process and transform multi-objective functions into reward functions suitable for training agent action policies. The model training module is used to acquire historical charging data of electric vehicles to build a training dataset and to train the model using a deep Q-network algorithm. The intelligent charging control module is used to collect current environmental status data related to electric vehicle charging, input it into the trained model, and execute corresponding charging operations according to the action strategy output by the model.

6. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the steps of the method as described in any one of claims 1 to 4 by calling the computer program stored in the memory.

7. A computer-readable storage medium for storing a user-oriented intelligent charging program for electric vehicles, characterized in that, The computer program, when run on a computer, performs the steps of the method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Electric vehicle fast and slow synchronous orderly charging scheduling method and electric quantity settlement method

    CN112488444A

  • Electric vehicle scheduling method considering various demands of users in mixed scene

    CN112907153A