Electric vehicle marketization regulation and control method and device based on reinforcement learning
By using a reinforcement learning-based market regulation method for electric vehicles, and training an agent with the TD3 algorithm to optimize electric vehicle charging strategies, this approach solves the multi-objective collaborative optimization problem faced by traditional methods when dealing with the uncertainty of charging demand and the intermittency of renewable resources. It achieves a balance between user demand, economic benefits, and energy consumption.
Patent Information
- Application Number
- CN202510768171.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-11-07
AI Technical Summary
Existing electric vehicle charging regulation methods are difficult to simultaneously optimize multiple objectives such as green electricity consumption and aggregator revenue when faced with the uncertainty of charging demand and the intermittency of renewable resources. Traditional methods have significant limitations.
A market-based regulation method for electric vehicles based on reinforcement learning is adopted. The agent is trained using the TD3 algorithm, and it judges whether the regulation conditions are met based on the charging demand of electric vehicles. Based on the judgment results, charging is performed, and the charging process is optimized using electric vehicle charging power strategies, including power constraints, time constraints, and SOC constraints.
This achieves the goal of meeting users' travel needs while also taking into account the economic interests of market players, the stability of the power distribution network, and the efficient consumption of renewable energy.
Smart Images

Figure CN120902566A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric vehicle charging regulation, and particularly relates to an electric vehicle marketization regulation method and device based on reinforcement learning. BACKGROUND
[0002] With the popularization of electric vehicles, electric vehicles have become an important direction for sustainable development in the transportation field due to their clean and efficient characteristics. Electric vehicle charging regulation is an important link to ensure the stable operation of the power system and meet user demand.
[0003] In the existing electric vehicle charging regulation, operations research or optimization algorithms such as particle swarm optimization (PSO) and genetic algorithm (GA) are commonly used. With further research, multi-objective collaborative optimization has gradually become a trend. However, traditional methods have certain limitations in terms of charging demand uncertainty and intermittent renewable resources, and existing research has focused on a single or a small number of objectives, making it difficult to consider multi-objective collaborative optimization such as green power consumption and aggregator revenue. SUMMARY
[0004] In order to overcome the above defects, the present application provides an electric vehicle marketization regulation method and device based on reinforcement learning.
[0005] In a first aspect, an electric vehicle marketization regulation method based on reinforcement learning is provided, which comprises:
[0006] determining whether the charging vehicle meets the regulation condition based on the charging demand of the electric vehicle;
[0007] charging the electric vehicle based on the determination result.
[0008] Preferably, the regulation condition is as follows:
[0009]
[0010] In the above formula, t is the charging time, t arr is the electric vehicle charging start time, t dep is the electric vehicle charging end time, P c is the maximum charging power, Δt is the time step, and E exp is the charging demand of the electric vehicle.
[0011] Preferably, the charging of the electric vehicle based on the determination result comprises:
[0012] If the charging vehicle meets the regulation condition, the current observation state of the electric vehicle is taken as an input of the pre-trained agent, an electric vehicle charging power strategy of an output of the pre-trained agent is obtained, and the electric vehicle is charged by using the electric vehicle charging power strategy, otherwise, the electric vehicle is charged at a maximum charging power.
[0013] The pre-trained agent is trained by using a TD3 algorithm.
[0014] Further, the Markov decision process corresponding to the agent is: (S t ,A t ,S t+1 ,R t ), wherein S t is an observation state of the electric vehicle at t time, A t is a charging and discharging power of the electric vehicle at t time, S t+1 is an observation state of the electric vehicle at t+1 time, and R t is a reward function at t time.
[0015] Further, the observation state of the electric vehicle at t time is: [L arr ,L dep ,C arr ,C dep ,t arr ,t dep ], wherein L arr is a power grid load value when the electric vehicle starts charging, L dep is a power grid load value when the electric vehicle ends charging, C arr is an electricity price when the electric vehicle starts charging, C dep is an electricity price when the electric vehicle ends charging, t arr is a charging start time of the electric vehicle, and t dep is a charging end time of the electric vehicle.
[0016] Further, the reward function at t time is as follows:
[0017] r=-[w1f1+w2f2+w3f3+w4f4]
[0018] In the above formula, r is the reward function at t time, w1, w2, w3 and w4 are first, second, third and fourth weight coefficients respectively, f1 is an aggregated merchant benefit, f2 is a user charging cost, f3 is a power grid load fluctuation, and f4 is a renewable energy consumption.
[0019] Further, the aggregated merchant benefit is as follows:
[0020]
[0021] The user charging cost is as follows:
[0022]
[0023] The power grid load fluctuation is as follows:
[0024]
[0025] The renewable energy consumption is as follows:
[0026]
[0027] In the above formula, N is the number of charging vehicles, t arr,i is the charging start time of the i-th electric vehicle, t dep,i is the charging end time of the i-th electric vehicle, p u (t) is the user payment electricity price at time t, p a (t) is the wholesale electricity price of the aggregator from the electricity market at time t, P c,i (t) is the charging power of the i-th electric vehicle at time t, C t is the charging electricity price at time t, P d,i (t) is the discharging power of the i-th electric vehicle at time t, D t is the discharging electricity price at time t, L t is the original power grid load value at time t, T is the charging duration, μ is the power grid load average, γ(t) is the new energy consumption ratio of the charging station at time t, P i,t is the charging and discharging power of the i-th electric vehicle at time t.
[0028] Further, the charging and discharging power of the i-th electric vehicle at time t corresponds to the constraint conditions, including power constraint, charging time constraint, vehicle SOC constraint, and charging and discharging constraint.
[0029] Further, the power constraint is as follows:
[0030] -P d ≤P i,t ≤P c
[0031] The charging time constraint is as follows:
[0032]
[0033] The vehicle SOC constraint is as follows:
[0034]
[0035] The charging and discharging constraint is as follows:
[0036]
[0037] In the above formula, P d is the maximum discharge power, P c is the maximum charging power, SOC i (t) is the state of charge of the ith electric vehicle at time t, SOC i (t-1) is the state of charge of the ith electric vehicle at time t-1, P i,t-1 is the charging and discharging power of the ith electric vehicle at time t-1, η(P i ) is a charging and discharging efficiency function varying with the direction of power, C i is the battery capacity of the ith electric vehicle, η c is the charging efficiency, η d is the discharging efficiency, SOCi i (t arr,i ) is the state of charge of the ith electric vehicle when it arrives at the charging station, η is the charging and discharging efficiency, SOC max is the maximum allowable state of charge of the electric vehicle, E exp is the charging demand of the electric vehicle.
[0038] In a second aspect, a device for market-oriented regulation of electric vehicles based on reinforcement learning is provided, which comprises:
[0039] a judging module configured to judge whether a charging vehicle meets a regulation condition based on a charging demand of the electric vehicle;
[0040] a regulation module configured to charge the electric vehicle based on a judgment result.
[0041] Preferably, the regulation condition is as follows:
[0042]
[0043] In the above formula, t is the charging time, t arr is the start time of charging of the electric vehicle, t dep is the end time of charging of the electric vehicle, P c is the maximum charging power, Δt is the time step, E exp is the charging demand of the electric vehicle.
[0044] Preferably, the regulation module is specifically configured to:
[0045] if the charging vehicle meets the regulation condition, take the current observation state of the electric vehicle as an input of a pre-trained agent to obtain an output of the pre-trained agent, i.e., a charging power strategy of the electric vehicle, and charge the electric vehicle using the charging power strategy, otherwise, charge the electric vehicle at the maximum charging power;
[0046] The pre-trained agent is trained using a TD3 algorithm.
[0047] Further, the Markov decision process corresponding to the agent is: (S t ,A t ,S t+1 ,R t ), wherein S t is the observation state of the electric vehicle at time t, A t is the charging and discharging power of the electric vehicle at time t, S t+1 is the observation state of the electric vehicle at time t+1, and R t is the reward function at time t.
[0048] Further, the observation state of the electric vehicle at time t is: [L arr ,L dep ,C arr ,C dep ,t arr ,t dep ], wherein L arr is the grid load value when the electric vehicle starts charging, L dep is the grid load value when the electric vehicle finishes charging, C arr is the electricity price when the electric vehicle starts charging, C dep is the electricity price when the electric vehicle finishes charging, t arr is the charging start time of the electric vehicle, and t dep is the charging end time of the electric vehicle.
[0049] Further, the reward function at time t is as follows:
[0050] r = - [w1f1 + w2f2 + w3f3 + w4f4]
[0051] In the above formula, r is the reward function at time t, w1, w2, w3, and w4 are first, second, third, and fourth weight coefficients, respectively, f1 is the aggregator's revenue, f2 is the user's charging cost, f3 is the grid load fluctuation, and f4 is the renewable energy consumption.
[0052] Further, the aggregator's revenue is as follows:
[0053]
[0054] The user's charging cost is as follows:
[0055]
[0056] The grid load fluctuation is as follows:
[0057]
[0058] The renewable energy accommodation is as follows:
[0059]
[0060] In the above formula, N is the number of charging vehicles, t arr,i is the charging start time of the i-th electric vehicle, t dep,i is the charging end time of the i-th electric vehicle, p u (t) is the user payment electricity price at time t, p a (t) is the wholesale electricity price of the aggregator from the electricity market at time t, P c,i (t) is the charging power of the i-th electric vehicle at time t, C t is the charging electricity price at time t, P d,i (t) is the discharging power of the i-th electric vehicle at time t, D t is the discharging electricity price at time t, L t is the original grid load value at time t, T is the charging duration, μ is the grid load average, γ(t) is the new energy accommodation ratio of the charging station at time t, P i,t is the charging and discharging power of the i-th electric vehicle at time t.
[0061] Further, the charging and discharging power of the i-th electric vehicle at time t corresponds to the constraint conditions, including power constraint, charging time constraint, vehicle SOC constraint, and charging and discharging constraint.
[0062] Further, the power constraint is as follows:
[0063] -P d ≤P i,t ≤P c
[0064] The charging time constraint is as follows:
[0065]
[0066] The vehicle SOC constraint is as follows:
[0067]
[0068] The charging and discharging constraint is as follows:
[0069]
[0070] In the above formula, P d is the maximum discharging power, P c is the maximum charging power, SOC i (t) is the state of charge of the i-th electric vehicle at time t, SOC i(t-1) is the state of charge of the i-th electric vehicle at time t-1, P i,t-1 is the charging and discharging power of the i-th electric vehicle at time t-1, η(P i ) is a charging and discharging efficiency function varying with the direction of power, C i is the battery capacity of the i-th electric vehicle, η c is the charging efficiency, η d is the discharging efficiency, SOC i (t arr,i ) is the state of charge of the i-th electric vehicle when it arrives at the charging station, η is the charging and discharging efficiency, SOC max is the maximum allowed state of charge of the electric vehicle, E exp is the charging demand of the electric vehicle.
[0071] In a third aspect, a computer device is provided, comprising: one or more processors;
[0072] the processor is configured to store one or more programs;
[0073] When the one or more programs are executed by the one or more processors, the method for market-oriented regulation and control of electric vehicles based on reinforcement learning is implemented.
[0074] In a fourth aspect, a computer readable storage medium is provided, having a computer program stored thereon, which, when executed, implements the method for market-oriented regulation and control of electric vehicles based on reinforcement learning.
[0075] The above one or more technical solutions of the present application have at least one or more of the following advantages:
[0076] The present application provides a method and device for market-oriented regulation and control of electric vehicles based on reinforcement learning, comprising: determining whether the charging vehicle meets the regulation and control conditions based on the charging demand of the electric vehicle; charging the electric vehicle based on the determination result; if the charging vehicle meets the regulation and control conditions, taking the current observation state of the electric vehicle as the input of the pre-trained agent to obtain the electric vehicle charging power strategy output by the pre-trained agent, and charging the electric vehicle using the electric vehicle charging power strategy, otherwise, charging the electric vehicle with the maximum charging power; the technical solution provided by the present application not only meets the travel demand of users, but also takes into account the economic benefits of market entities, the safety and stability of the power distribution network and the efficient consumption of renewable energy. BRIEF DESCRIPTION OF DRAWINGS
[0077] Figure 1 is the main step flowchart of the method for market-oriented regulation and control of electric vehicles based on reinforcement learning of the embodiment of the present application. DETAILED DESCRIPTION
[0078] The specific embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0079] To make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0080] As disclosed in the background, with the popularity of electric vehicles, electric vehicles have become an important direction for sustainable development in the transportation field due to their clean and efficient characteristics. Electric vehicle charging regulation is an important link to ensure the stable operation of the power system and meet user demand.
[0081] In the existing electric vehicle charging regulation, operations research or optimization algorithms such as particle swarm optimization (PSO) and genetic algorithm (GA) are commonly used. With further research, multi-objective collaborative optimization has gradually become a trend. However, traditional methods have certain limitations in terms of charging demand uncertainty and intermittent renewable resources, and existing research has focused on a single or a small number of objectives, making it difficult to consider multi-objective collaborative optimization such as green power consumption and aggregator revenue.
[0082] To improve the above problems, the present application provides a reinforcement learning-based electric vehicle market regulation method and device, which includes: determining whether the charging vehicle meets the regulation condition based on the charging demand of the electric vehicle; charging the electric vehicle based on the determination result; if the charging vehicle meets the regulation condition, taking the current observation state of the electric vehicle as the input of the pre-trained agent to obtain the electric vehicle charging power strategy output by the pre-trained agent, and charging the electric vehicle using the electric vehicle charging power strategy, otherwise, charging the electric vehicle at the maximum charging power; the technical solution provided by the present application not only meets the travel demand of users, but also takes into account the economic benefits of market subjects, the safety and stability of distribution networks and the efficient consumption of renewable energy.
[0083] The above scheme will be described in detail below.
[0084] Embodiment 1
[0085] Referring to the accompanying Figure 1 , Figure 1 is the main step flowchart of the reinforcement learning-based electric vehicle market regulation method of an embodiment of the present application. As shown in Figure 1 , the reinforcement learning-based electric vehicle market regulation method in the embodiment of the present application mainly includes the following steps:
[0086] Step S101: judging whether the electric vehicle meets the regulation condition based on the charging demand of the electric vehicle;
[0087] Step S102: charging the electric vehicle based on the judgment result.
[0088] In the embodiment, the regulation condition is as follows:
[0089]
[0090] In the above formula, t is the charging time, t arr is the starting time of charging the electric vehicle, t dep is the ending time of charging the electric vehicle, P c is the maximum charging power, Δt is the time step, and E exp is the charging demand of the electric vehicle.
[0091] In the embodiment, the charging of the electric vehicle based on the judgment result comprises:
[0092] If the electric vehicle meets the regulation condition, the current observation state of the electric vehicle is taken as the input of the pre-trained agent, the output of the pre-trained agent is obtained, the charging power strategy of the electric vehicle is obtained, and the electric vehicle is charged by using the charging power strategy of the electric vehicle, otherwise, the electric vehicle is charged at the maximum charging power;
[0093] The pre-trained agent is trained by using a TD3 algorithm.
[0094] In one embodiment, the Markov decision process corresponding to the agent is (S t , A t , S t+1 , R t ), wherein S t is the observation state of the electric vehicle at time t, A t is the charging and discharging power of the electric vehicle at time t, S t+1 is the observation state of the electric vehicle at time t+1, and R t is the reward function at time t.
[0095] In one embodiment, the observation state of the electric vehicle at time t is: [l arr , l dep , C arr , C dep , t arr , t dep ], wherein L arr is the power grid load value when the electric vehicle starts charging, L depC is the grid load value when the electric vehicle ends charging, arr C is the electricity price when the electric vehicle starts charging, dep C is the electricity price when the electric vehicle ends charging, arr t is the charging start time of the electric vehicle, dep t is the charging end time of the electric vehicle.
[0096] In one embodiment, the reward function at time t is as follows:
[0097] r = - [w1f1 + w2f2 + w3f3 + w4f4]
[0098] In the above formula, r is the reward function at time t, w1, w2, w3, and w4 are the first, second, third, and fourth weight coefficients respectively, f1 is the aggregator revenue, f2 is the user charging cost, f3 is the grid load fluctuation, and f4 is the renewable energy consumption.
[0099] In one embodiment, the aggregator revenue is as follows:
[0100]
[0101] The user charging cost is as follows:
[0102]
[0103] The grid load fluctuation is as follows:
[0104]
[0105] The renewable energy consumption is as follows:
[0106]
[0107] In the above formula, N is the number of charging vehicles, t arr,i t is the charging start time of the i-th electric vehicle, dep,i t is the charging end time of the i-th electric vehicle, u p(t) is the user payment electricity price at time t, a P(t) is the wholesale electricity price of the aggregator purchasing electricity from the electricity market at time t, c,i C(t) is the charging power of the i-th electric vehicle at time t, t C(t) is the charging electricity price at time t, d,i D(t) is the discharging power of the i-th electric vehicle at time t, t L(t) is the discharging electricity price at time t, t T(t) is the original grid load value at time t, T is the charging duration, μ is the grid load mean value, and γ(t) is the new energy consumption ratio of the charging station at time t, i,tThe charging and discharging power of the ith electric vehicle at time t.
[0108] In an embodiment, the constraint condition corresponding to the charging and discharging power of the ith electric vehicle at time t includes a power constraint, a charging time constraint, a vehicle SOC constraint, and a charging and discharging constraint.
[0109] In an embodiment, the power constraint is as follows:
[0110] -P d ≤P i,t ≤P c
[0111] The charging time constraint is as follows:
[0112]
[0113] The vehicle SOC constraint is as follows:
[0114]
[0115] The charging and discharging constraint is as follows:
[0116]
[0117] In the above formula, P d is the maximum discharging power, P c is the maximum charging power, SOC i (t) is the state of charge of the ith electric vehicle at time t, SOC i (t-1) is the state of charge of the ith electric vehicle at time t-1, P i,t-1 is the charging and discharging power of the ith electric vehicle at time t-1, η(P i ) is a charging and discharging efficiency function varying with the direction of power, C i is the battery capacity of the ith electric vehicle, η c is the charging efficiency, η d is the discharging efficiency, SOC i (t arr,i ) is the state of charge of the ith electric vehicle when it arrives at the charging station, η is the charging and discharging efficiency, SOC max is the maximum allowable state of charge of the electric vehicle, E exp is the charging demand of the electric vehicle.
[0118] In the process of training the pre-trained agent by using the TD3 algorithm, the application adopts a simulation form for research, constructs a travel probability model of an electric vehicle user, thereby deducing a charging demand distribution of the electric vehicle user, and simultaneously accurately sets experimental parameters. The charging demand distribution of the electric vehicle user includes a time distribution of the last travel end of the user in a day, a time distribution of the start of travel of the user in a day, and an electric quantity distribution at the travel end time, etc.; the experimental parameters include an electric vehicle charging and discharging price, vehicle parameters, load prediction, and TD3 algorithm experimental environment configuration, etc.
[0119] For the charging demand distribution of the electric vehicle user, a Monte Carlo method is adopted to simulate and generate the charging behavior of the electric vehicle user, including an electric vehicle arrival time, a leaving time, and a charging demand.
[0120] A neural network model is adopted to perform grid load prediction, and accurately simulate the electric load change of an EV charging area.
[0121] Specifically, first, vehicle state information is extracted in discrete time steps, and the time is compared with the predicted leaving time of the user to determine whether scheduling optimization is needed. If the condition is met, the agent adds noise to the action output by the deterministic policy to explore, thereby selecting a suitable charging action. Subsequently, the double Q network is used to update the Q value corresponding to the state-action pair, and the target Q value takes the smaller value of the output of the two target networks, so as to reduce the overestimation bias. When the training round number reaches a set threshold or the policy converges, the algorithm outputs the final charging scheduling scheme; otherwise, the vehicle continues to charge at the rated power.
[0122] Embodiment 2
[0123] Based on the same inventive concept, the application further provides an electric vehicle marketization regulation and control device based on reinforcement learning, which comprises:
[0124] A judgment module is configured to judge whether the charging vehicle meets a regulation and control condition based on the charging demand of the electric vehicle.
[0125] A regulation and control module is configured to charge the electric vehicle based on the judgment result.
[0126] Preferably, the regulation and control condition is as follows:
[0127]
[0128] In the above formula, t is a charging time, t arr is an electric vehicle charging start time, t dep is an electric vehicle charging end time, P c is a maximum charging power, Δt is a time step, and E expFor the charging demand of electric vehicles.
[0129] Preferably, the regulation module is specifically used for:
[0130] If the charging vehicle meets the regulation condition, the current observation state of the electric vehicle is taken as the input of the pre-trained agent, the output of the pre-trained agent is obtained, the charging power strategy of the electric vehicle, and the electric vehicle is charged by using the charging power strategy of the electric vehicle, otherwise, the electric vehicle is charged with the maximum charging power;
[0131] Wherein, the pre-trained agent is trained by TD3 algorithm.
[0132] Further, the Markov decision process corresponding to the agent is: (S t ,A t ,S t+1 ,R t ), wherein S t is the observation state of the electric vehicle at time t, A t is the charging and discharging power of the electric vehicle at time t, S t+1 is the observation state of the electric vehicle at time t+1, and R t is the reward function at time t.
[0133] Further, the observation state of the electric vehicle at time t is: [L arr ,L dep ,C arr ,C dep ,t arr ,t dep ], wherein L arr is the load value of the power grid when the electric vehicle starts charging, L dep is the load value of the power grid when the electric vehicle ends charging, C arr is the electricity price when the electric vehicle starts charging, C dep is the electricity price when the electric vehicle ends charging, t arr is the starting time of charging of the electric vehicle, and t dep is the end time of charging of the electric vehicle.
[0134] Further, the reward function at time t is as follows:
[0135] r=-[w1f1+w2f2+w3f3+w4f4]
[0136] In the above formula, r is the reward function at time t, w1, w2, w3 and w4 are the first, second, third and fourth weight coefficients respectively, f1 is the aggregated merchant income, f2 is the user charging cost, f3 is the power grid load fluctuation, and f4 is the renewable energy consumption.
[0137] Further, the aggregator's revenue is as follows:
[0138]
[0139] The user charging cost is as follows:
[0140]
[0141] The grid load fluctuation is as follows:
[0142]
[0143] The renewable energy consumption is as follows:
[0144]
[0145] In the above formula, N is the number of charging vehicles, t arr,i is the charging start time of the i-th electric vehicle, t dep,i is the charging end time of the i-th electric vehicle, p u (t) is the user's electricity price at time t, p a (t) is the wholesale electricity price of the aggregator from the electricity market at time t, P c,i (t) is the charging power of the i-th electric vehicle at time t, C t is the charging electricity price at time t, P d,i (t) is the discharging power of the i-th electric vehicle at time t, D t is the discharging electricity price at time t, L t is the original grid load value at time t, T is the charging duration, μ is the grid load average, γ(t) is the new energy consumption ratio of the charging station at time t, P i,t is the charging and discharging power of the i-th electric vehicle at time t.
[0146] Further, the charging and discharging power of the i-th electric vehicle at time t corresponds to the following constraint conditions: power constraint, charging time constraint, vehicle SOC constraint, charging and discharging constraint.
[0147] Further, the power constraint is as follows:
[0148] -P d ≤P i,t ≤P c
[0149] The charging time constraint is as follows:
[0150]
[0151] The vehicle SOC constraint is as follows:
[0152]
[0153]
[0154] The charge-discharge constraints are as follows:
[0155]
[0156] In the above formula, P d is the maximum discharge power, P c is the maximum charge power, SOC i (t) is the state of charge of the ith electric vehicle at time t, SOC i (t-1) is the state of charge of the ith electric vehicle at time t-1, P i,t-1 is the charge-discharge power of the ith electric vehicle at time t-1, η(P i ) is a charge-discharge efficiency function that varies with the direction of power, C i is the battery capacity of the ith electric vehicle, η c is the charge efficiency, η d is the discharge efficiency, SOC i (t arr,i ) is the state of charge of the ith electric vehicle when it arrives at the charging station, η is the charge-discharge efficiency, SOC max is the maximum allowable state of charge of the electric vehicle, E exp is the charging demand of the electric vehicle.
[0157] Embodiment 3
[0158] Based on the same inventive concept, the application further provides a computer device, which comprises a processor and a memory. The memory is used to store a computer program, and the computer program comprises program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor is the computing core and control core of the terminal, and is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions in the computer storage medium to implement a corresponding method flow or corresponding function, so as to implement the steps of the above-mentioned embodiment of the electric vehicle marketization regulation method based on reinforcement learning.
[0159] Embodiment 4
[0160] Based on the same inventive concept, the present application further provides a storage medium, specifically a computer readable storage medium (Memory), which is a memory device in a computer device, used for storing programs and data. It can be understood that the computer readable storage medium herein can include an internal storage medium in the computer device, and of course can also include an extended storage medium supported by the computer device. The computer readable storage medium provides a storage space, which stores an operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and the instructions can be one or more computer programs (including program codes). It should be noted that the computer readable storage medium herein can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. One or more instructions stored in the computer readable storage medium can be loaded and executed by the processor to implement the steps of the above-mentioned embodiment of the electric vehicle marketization regulation method based on reinforcement learning.
[0161] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0162] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of the flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The means for performing the functions specified in one or more flows and / or blocks.
[0163] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the flow Figure 1 The functions specified in the flow or flows and / or blocks Figure 1 The functions specified in the flow or flows and / or blocks
[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow Figure 1 The functions specified in the flow or flows and / or blocks Figure 1 The functions specified in the flow or flows and / or blocks
[0165] Finally, it should be noted that the above-mentioned embodiments are merely used to illustrate the technical solutions of the present application, rather than limit the present application. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and any modification or replacement without departing from the spirit and scope of the present application should be covered in the protection scope of the claims of the present application.
Claims
1. A method for market regulation of electric vehicles based on reinforcement learning, characterized in that, The method comprises: determining whether the charging vehicle meets the regulation condition based on the charging demand of the electric vehicle; charging the electric vehicle based on the determination result.
2. The method of claim 1, wherein, The regulation condition is as follows: In the above equation, t is the charging time, t arr is the start time of charging for the electric vehicle, t dep is the end time of charging for the electric vehicle, P c is the maximum charging power, Δt is the time step, E exp is the charging demand of the electric vehicle.
3. The method of claim 1, wherein, The charging of the electric vehicle based on the determination result comprises: if the charging vehicle meets the regulation condition, taking the current observation state of the electric vehicle as the input of the pre-trained agent to obtain the charging power strategy of the electric vehicle output by the pre-trained agent, and charging the electric vehicle by using the charging power strategy of the electric vehicle, otherwise, charging the electric vehicle at the maximum charging power; wherein the pre-trained agent is trained by using a TD3 algorithm.
4. The method of claim 3, wherein, The Markov decision process corresponding to the intelligent agent is: (S t , A t , S t+1 , R t ), wherein S t is the observation state of the electric vehicle at time t, A t is the charging and discharging power of the electric vehicle at time t, S t+1 is the observation state of the electric vehicle at time t+1, and R t is the reward function at time t.
5. The method of claim 4, wherein, The observed state of the electric vehicle at the time t: [L arr , dep , arr , dep , arr , dep ], wherein L arr is the power grid load value when the electric vehicle starts charging, L dep is the power grid load value when the electric vehicle ends charging, C arr is the electricity price when the electric vehicle starts charging, C dep is the electricity price when the electric vehicle ends charging, t arr is the charging start time of the electric vehicle, and t dep is the charging end time of the electric vehicle.
6. The method of claim 5, wherein, The reward function at the t time is as follows: r=-[w1f1+w2f2+w3f3+w4f4] In the above formula, r is the reward function at the t time, w1, w2, w3 and w4 are respectively the first, second, third and fourth weight coefficients, f1 is the aggregator revenue, f2 is the user charging cost, f3 is the power grid load fluctuation, and f4 is the renewable energy consumption.
7. The method of claim 6, wherein, The aggregator revenue is as follows: The user charging cost is as follows: The power grid load fluctuation is as follows: The renewable energy consumption is as follows: N is the number of charging vehicles, t arr,i is the charging start time of the i-th electric vehicle, t dep,i is the charging end time of the i-th electric vehicle, p u (t) is the user's electricity price at time t, p a (t) is the wholesale electricity price of the aggregator from the electricity market at time t, P c,i (t) is the charging power of the i-th electric vehicle at time t, C t is the charging electricity price at time t, P d,i (t) is the discharging power of the i-th electric vehicle at time t, D t is the discharging electricity price at time t, L t is the original grid load value at time t, T is the charging duration, μ is the grid load average, γ(t) is the new energy consumption ratio of the charging station at time t, P i,t is the charging and discharging power of the i-th electric vehicle at time t.
8. The method of claim 7, wherein, The constraint condition corresponding to the charging and discharging power of the i-th electric vehicle at the t time comprises: power constraint, charging time constraint, vehicle SOC constraint and charging and discharging constraint.
9. The method of claim 8, wherein, The power constraint is as follows: - P d ≤ P i,t ≤ P c The charging time constraint is as follows: The vehicle SOC constraint is as follows: The charging and discharging constraint is as follows: In the above formula, P d is the maximum discharge power, P c is the maximum charging power, SOC i (t) is the state of charge of the i-th electric vehicle at time t, SOC i (t-1) is the state of charge of the i-th electric vehicle at time t-1, P i,t-1 is the charge-discharge power of the i-th electric vehicle at time t-1, η(P i ) is a charge-discharge efficiency function that varies with the direction of power, C i is the battery capacity of the i-th electric vehicle, η c is the charging efficiency, η d is the discharging efficiency, SOC i (t arr,i ) is the state of charge of the i-th electric vehicle when it arrives at the charging station, η is the charge-discharge efficiency, SOC max is the maximum allowable state of charge of the electric vehicle, E exp is the charging demand of the electric vehicle.
10. A device for regulating and controlling the marketization of electric vehicles based on reinforcement learning, characterized in that, The device comprises: a determination module configured to determine whether the charging vehicle meets the regulation condition based on the charging demand of the electric vehicle; a regulation module configured to charge the electric vehicle based on the determination result.
11. The apparatus of claim 9, wherein, The regulation condition is as follows: In the above equation, t is the charging time, t arr is the start time of charging for the electric vehicle, t dep is the end time of charging for the electric vehicle, P c is the maximum charging power, Δt is the time step, E exp is the charging demand of the electric vehicle.
12. The apparatus of claim 9, wherein, The regulation module is specifically configured to: if the charging vehicle meets the regulation condition, take the current observation state of the electric vehicle as the input of the pre-trained agent to obtain the charging power strategy of the electric vehicle output by the pre-trained agent, and charge the electric vehicle by using the charging power strategy of the electric vehicle, otherwise, charge the electric vehicle at the maximum charging power; wherein the pre-trained agent is trained by using a TD3 algorithm.
13. The apparatus of claim 12, wherein, The Markov decision process corresponding to the intelligent agent is: (S t , A t , S t+1 , R t ), wherein S t is the observation state of the electric vehicle at time t, A t is the charging and discharging power of the electric vehicle at time t, S t+1 is the observation state of the electric vehicle at time t+1, and R t is the reward function at time t.
14. The apparatus of claim 13, wherein, The observed state of the electric vehicle at the time t: [L arr , L dep , C arr , C dep , t arr , t dep ] wherein L arr is the grid load value when the electric vehicle starts charging, L dep is the grid load value when the electric vehicle ends charging, C arr is the electricity price when the electric vehicle starts charging, C dep is the electricity price when the electric vehicle ends charging, t arr is the electric vehicle charging start time, t dep is the electric vehicle charging end time.
15. The apparatus of claim 14, wherein, The reward function at the t time is as follows: r=-[w1f1+w2f2+w3f3+w4f4] In the above formula, r is the reward function at the t time, w1, w2, w3 and w4 are respectively the first, second, third and fourth weight coefficients, f1 is the aggregator revenue, f2 is the user charging cost, f3 is the power grid load fluctuation, and f4 is the renewable energy consumption.
16. The apparatus of claim 15, wherein, The aggregator revenue is as follows: The user charging cost is as follows: The power grid load fluctuation is as follows: The renewable energy consumption is as follows: N is the number of charging vehicles, t arr,i is the charging start time of the i-th electric vehicle, t dep,i is the charging end time of the i-th electric vehicle, p u (t) is the user's electricity price at time t, p a (t) is the wholesale electricity price of the aggregator from the electricity market at time t, P c,i (t) is the charging power of the i-th electric vehicle at time t, C t is the charging electricity price at time t, P d,i (t) is the discharging power of the i-th electric vehicle at time t, D t is the discharging electricity price at time t, L t is the original grid load value at time t, T is the charging duration, μ is the grid load average, γ(t) is the new energy consumption ratio of the charging station at time t, P i,t is the charging and discharging power of the i-th electric vehicle at time t.
17. The apparatus of claim 16, wherein, The constraint condition corresponding to the charging and discharging power of the i-th electric vehicle at the t time comprises: power constraint, charging time constraint, vehicle SOC constraint and charging and discharging constraint.
18. The apparatus of claim 17, wherein, The power constraint is as follows: - P d ≤ P i,t ≤ P c The charging time constraint is as follows: The vehicle SOC constraint is as follows: The charging and discharging constraint is as follows: In the above formula, P d is the maximum discharge power, P c is the maximum charging power, SOC i (t) is the state of charge of the i-th electric vehicle at time t, SOC i (t-1) is the state of charge of the i-th electric vehicle at time t-1, P i,t-1 is the charge-discharge power of the i-th electric vehicle at time t-1, η(P i ) is a charge-discharge efficiency function that varies with the direction of power, C i is the battery capacity of the i-th electric vehicle, η c is the charging efficiency, η d is the discharging efficiency, SOC i (t arr,i ) is the state of charge of the i-th electric vehicle when it arrives at the charging station, η is the charge-discharge efficiency, SOC max is the maximum allowable state of charge of the electric vehicle, E exp is the charging demand of the electric vehicle.
19. A computer device, comprising: comprise: one or more processors; the processor is configured to execute one or more programs; When the one or more programs are executed by the one or more processors, the method for regulating marketization of electric vehicles based on reinforcement learning as claimed in any one of claims 1 to 9 is implemented.
20. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed, the method for regulating marketization of electric vehicles based on reinforcement learning as claimed in any one of claims 1 to 9 is implemented.