Optical storage and charging integrated scheduling method and system based on cyclic near-end strategy optimization algorithm
Through an integrated photovoltaic storage and charging scheduling model based on a cyclic proximal strategy optimization algorithm and long short-term memory network training, the problem of the impact of electric vehicle charging on the power grid is solved, efficient use of renewable energy and optimized resource allocation are achieved, and operating costs and carbon emissions are reduced.
Patent Information
- Application Number
- CN202510177843.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-09-26
AI Technical Summary
Large-scale decentralized charging of electric vehicles causes problems such as peak-to-valley load differences and harmonic pollution in the power grid. Existing technologies make it difficult to effectively utilize renewable energy and optimize the operation of integrated photovoltaic, storage and charging systems.
A cyclic proximal strategy optimization algorithm is adopted. By building a Markov decision process for an integrated photovoltaic, storage and charging system and combining it with a long-short-term memory network, an integrated photovoltaic, storage and charging scheduling model is trained to optimize the charging and discharging strategy to reduce costs and improve efficiency.
It improves the utilization rate of renewable energy, reduces dependence on traditional energy, optimizes resource allocation, reduces operating costs, enhances system stability and flexibility, and reduces dependence on the power grid and carbon emissions.
Smart Images

Figure CN120710080A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system dispatching, and in particular to a photovoltaic storage and charging integrated dispatching method and system based on a cyclic proximal strategy optimization algorithm. Background Art
[0002] With the development of photovoltaic power generation and the increasing popularity of electric vehicles, charging stations have entered a period of rapid development and construction. Due to the randomness and uncontrollability of EV charging, the widespread and decentralized access of EVs to the power grid will inevitably impact the grid, such as increased peak-to-valley load variations and severe harmonic pollution. Integrated photovoltaic (PV)-storage-charging systems effectively integrate EV charging stations with photovoltaic power generation systems and energy storage systems. While meeting the charging load requirements of EVs, they also significantly reduce the load pressure on the grid caused by EV charging. In this way, renewable energy can be fully and efficiently utilized. The synergistic effect of vehicles not only improves the utilization rate of PV power generation but also reduces the impact of PV power generation costs on the PV grid. Therefore, developing appropriate algorithms to control and optimize the operation of PV-storage-charging systems is crucial to promoting the smooth integration of EV units into current power grids. Summary of the Invention
[0003] In view of the above problems in the prior art, the present invention is proposed.
[0004] This paper provides a method for integrated solar-storage-charging scheduling based on a cyclic proximal strategy optimization algorithm. The system's structure, basic components, and operational strategy are proposed. The optimal scheduling of this system under grid-connected conditions is investigated, and a method for integrated solar-storage-charging scheduling based on a cyclic proximal strategy optimization algorithm is proposed. The primary goal is to provide a common environment for charging and discharging electric vehicles under various disturbances (e.g., weather conditions, pricing models, random arrival and departure times of electric vehicles, and random battery charge states upon arrival). Consequently, by training repeatedly within these generated environments, the controller will understand and master the underlying charging dynamics and exploit them to effectively achieve its objectives.
[0005] As a preferred solution of the integrated photovoltaic, storage and charging scheduling method based on the cyclic proximal strategy optimization algorithm described in the present invention, the physical architecture of the photovoltaic energy storage and charging station is built, and the scheduling process of the integrated photovoltaic, storage and charging system is modeled as a Markov decision process; and the cyclic proximal strategy optimization algorithm is used to train the integrated photovoltaic, storage and charging scheduling model.
[0006] As a preferred solution of the integrated photovoltaic, storage and charging scheduling method based on the cyclic proximal strategy optimization algorithm described in the present invention, the physical architecture of the photovoltaic energy storage charging station includes a group of multiple charging points, a photovoltaic power generation system, an external power grid that provides electricity at a specific price, and a vehicle-to-grid operation mode. When the available energy is insufficient, the charging station purchases electricity from the power grid with fluctuating electricity prices. The available electricity of the power station is divided into two types: energy stored in the car that can be used under V2G operation, and energy generated by photovoltaic power generation. The controller in the photovoltaic energy storage charging station determines the charge and discharge rate of each charging point based on the current charging demand, renewable energy power generation, and grid electricity price. The goal of the economic operation strategy is to achieve the lowest cost while meeting the charging demand.
[0007] As a preferred solution of the optical storage and charging integrated scheduling method based on the cyclic proximal strategy optimization algorithm described in the present invention, the Markov decision process includes: By the policy function Decision, after state transfer, the second state By the state transition function The decision, strategy and state transition functions are conditional probability distribution functions. In the process of interaction between the agent and the environment, the uncertainty of behavior and state leads to the randomness of MDP. ; Among them, A represents the action space, S represents the state space, and P represents the conditional probability distribution function; From the initial state Initially, the agent executes Action , until the completion signal, the agent experiences a complete event that conforms to the Markov assumption, i.e. , in the cycle Afterwards, the reward decay factor Rewards To indicate cumulative rewards : ; in, Expressed as the reward function at time t, Expressed as the reward function at time t+1, Action-value function is the conditional expectation of reward for a given policy , Describes the current state The quality of the next action, ; State Value Function It's an action set The action-value function expectations, it is only related to the strategy and status Related, describes the Downward Strategy Function The pros and cons: ; Then the state value function Guide the agent to choose the optimal strategy in the current state ; The action value function and the state value function obey the Bellman equation. The Bellman equation calculates the previous variable from the subsequent variable. For the current state , the environment receives an operation To get the next state , the return of the operation is expressed as , the Bellman equation of the action value function and state value function is: ; In each time period The status of the solar energy storage charging station is defined as: ; in: is the current solar radiation value, represents the solar radiation forecast for the next three hours; is the current price charged by the utility company to the power station for the required amount of electricity, is the predicted electricity price for the next three hours; Indicates time period The state of charge of the electric vehicle at the first charging point; Indicates the time that electric vehicles stay at charging stations. and Indicates the physical state of the solar energy storage charging station. If the charging point exist is empty, then and All are 0.
[0008] As a preferred solution of the optical storage and charging integrated scheduling method based on the cyclic proximal strategy optimization algorithm described in the present invention, the Markov decision process also includes using multiple continuous variables To define the charge and discharge rate of each vehicle charging point, it is constrained in the interval [−1,1], and each vehicle In the time step The charge and discharge power is defined as: ; in, For every electric car At the moment The charging status, For electric vehicle battery capacity, For maximum charging power, For the best charge and discharge efficiency, the actions taken by the controller are based on the time step The above calculation of the demand for each charging point, if the action is positive, then is positive, if the action is negative, If it is negative, the action The value of directly affects demand; When performing an action When the vehicle is in the state of charge, the main change in the environment is the state of charge of the electric vehicle: ; The goal of the controller of the solar energy storage charging station is to adopt a scheduling strategy to minimize the cost of purchasing electricity for the solar energy storage charging station, taking into account the penalty for electric vehicles that are not fully charged. The equation of the formula is expressed as, ; in, Indicates the total charging demand that the solar energy storage charging station requests from the utility company; electricity price Follow the different bill profiles provided by the utility company to the solar-storage charging station at each time period.
[0009] As a preferred solution of the optical storage and charging integrated scheduling method based on the cyclic proximal strategy optimization algorithm described in the present invention, the long short-term memory network uses the previous state , Update hidden state and memory status , the current observation status is: ; in, For status; is in hidden state; It is the memory state; They are input, forget gate, output, and candidate unit state gate respectively; are weight and bias respectively; is the activation function.
[0010] As a preferred solution of the optical storage and charging integrated scheduling method based on the cyclic proximal strategy optimization algorithm described in the present invention, the long short-term memory network also includes: After that, the strategy is: ; in, For strategic deviation, Strategy weight, is the activation function; The value function is: ; Consider the PPO objective of the recursive state: ; As in standard PPO, the odds estimation is usually performed using the generalized odds estimation: ; in, is the TD error, Both are discount factors; Strategy The update is done by gradient ascent: .
[0011] As a preferred solution of the integrated photovoltaic, storage and charging scheduling method based on the cyclic proximal policy optimization algorithm described in the present invention, wherein: the training integrated photovoltaic, storage and charging scheduling model collects data by running the interaction between the intelligent agent, that is, the charging station and the environment; at each time step, the RNN strategy outputs the action probability based on the current state, and uses the reward obtained during the data collection process to calculate the advantage function; the advantage function reflects the degree of advantage of each action compared with the expected value under the state; calculating the PPO target, including calculating the ratio between the new strategy probability and the old strategy probability, and taking the minimum value with a trimmed version of the ratio; using the optimization algorithm to update the RNN policy parameters according to the PPO target; repeating the process of data collection, advantage estimation, PPO target calculation and optimization to improve the strategy over time; implementing a training loop in which the intelligent agent interacts with the environment, collects data, estimates the advantage, calculates the PPO target, and iteratively updates the LSTM policy parameters.
[0012] Another objective of the present invention is to provide an integrated photovoltaic, storage, and charging scheduling system based on a cyclic proximal strategy optimization algorithm. This system utilizes a carefully designed physical architecture, including photovoltaic power generation systems and energy storage devices, through a structural building module. This system not only maximizes the utilization of renewable energy and reduces reliance on traditional energy sources, but also optimizes resource allocation and improves the overall stability and reliability of the system. This effective structural design plays a key role in reducing energy waste and maintenance costs, thereby reducing operating costs in the long term. The scheduling model training module utilizes an intelligent scheduling model trained using the cyclic proximal strategy optimization algorithm to provide managers with accurate and efficient decision support. This model exhibits excellent adaptability and flexibility, capable of self-adjusting strategies based on real-time data and environmental changes, ensuring efficient operation under ever-changing market and environmental conditions. Furthermore, the module maximizes cost-effectiveness by optimizing the charging and discharging processes, which is particularly important in situations where electricity prices fluctuate significantly. Furthermore, more efficient utilization of renewable energy and reduced reliance on the power grid contribute to reducing carbon emissions and other environmental impacts.
[0013] As a preferred solution of the photovoltaic storage and charging integrated scheduling system based on the cyclic proximal strategy optimization algorithm described in the present invention, it includes: a structure building module and a scheduling model training module; The structure building module builds the physical architecture of the solar energy storage charging station and models the scheduling process of the integrated solar energy storage and charging system into a Markov decision process; The scheduling model training module uses a cyclic proximal strategy optimization algorithm to train the integrated photovoltaic storage and charging scheduling model.
[0014] A computer device includes a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, it implements the steps of any one of the methods of the optical-storage-charging integrated scheduling method based on the cyclic proximal strategy optimization algorithm.
[0015] A computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of any one of the methods of the optical-storage-charging integrated scheduling method based on a cyclic proximal strategy optimization algorithm are implemented.
[0016] Beneficial effects of the present invention: The method of the present invention significantly improves the utilization efficiency of photovoltaic power generation and energy storage resources, while reducing dependence on traditional energy sources, through a carefully designed physical architecture and modeling the scheduling process as a Markov decision process. It can maximize cost-effectiveness while meeting charging needs, and determine the most economical charging and discharging rate through an intelligent controller. The scheduling model trained using the cyclic proximal policy optimization algorithm provides efficient decision support for the photovoltaic energy storage charging station, enhancing the accuracy and efficiency of scheduling. The application of long short-term memory networks further improves the adaptability and flexibility of the model, enabling it to maintain efficient operation under changing market and environmental conditions. By optimizing the charging and discharging process of photovoltaic power generation and electric vehicles, it helps to reduce dependence on the power grid, thereby reducing carbon emissions and other environmental impacts. Continuous optimization and self-improvement can be achieved through continuous data collection and strategy updates. It not only improves the operating efficiency and intelligence level of the photovoltaic energy storage charging station, reduces operating costs and reduces environmental impact, but also provides an innovative and efficient solution for sustainable energy and smart grid management. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 A schematic flow chart of a method for integrated PV-storage-charging scheduling based on a cyclic proximal strategy optimization algorithm is provided in accordance with an embodiment of the present invention.
[0019] Figure 2 A schematic flow chart of a photovoltaic storage and charging integrated scheduling system based on a cyclic proximal strategy optimization algorithm is provided in accordance with an embodiment of the present invention. DETAILED DESCRIPTION
[0020] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0021] Example 1 Reference Figure 1 , which is the first embodiment of the present invention, provides a method for integrated optical storage and charging scheduling based on a cyclic proximal strategy optimization algorithm, including: The purpose of this invention is to provide a method for integrated solar-storage-charging scheduling based on a cyclic proximal strategy optimization algorithm. The system's structure, basic components, and operational strategy are proposed. The optimal scheduling of this system under grid-connected conditions is investigated, and a method for integrated solar-storage-charging scheduling based on a cyclic proximal strategy optimization algorithm is proposed. The primary goal is to provide a generic environment for EV charging and discharging under various disturbances (e.g., weather conditions, pricing models, random arrival and departure times, and random battery charge states upon arrival). Thus, by training repeatedly within these generated environments, the controller will understand and master the underlying charging dynamics and exploit them to effectively achieve its objectives.
[0022] S1: Build the physical structure of the solar-storage charging station and model the scheduling process of the integrated solar-storage-charging system as a Markov decision process; Specifically, S1 builds the physical structure of the solar-storage charging station and models the scheduling process of the integrated solar-storage-charging system as a Markov decision process, including the following steps: S1-1: Build a physical model of the solar energy storage charging station.
[0023] The architecture of a solar-energy storage charging station includes: a group of multiple charging points, a photovoltaic power generation system, an external power grid that provides electricity at a specific price, and a "vehicle-to-grid" (V2G) operating mode. This operating mode adds vehicle-to-grid functionality to the charging station, allowing the use of energy stored in electric vehicles to charge other electric vehicles when necessary. When the available energy is insufficient, the charging station purchases electricity from the grid with fluctuating electricity prices. The available electricity of the power station (in addition to the grid) is divided into two types: energy stored in the car that can be used under V2G operation, and energy generated by photovoltaic power generation. The controller in the solar-energy storage charging station determines the charge and discharge rate of each charging point based on factors such as current charging demand, renewable energy generation, and grid electricity prices. The goal of its economic operation strategy is to achieve the lowest cost while meeting charging demand.
[0024] Step S1-2: Define the state space, action space, and reward function in the integrated solar-storage-charging scheduling process to build a Markov decision process.
[0025] The Markov Decision Process (MDP) is the mathematical foundation of deep reinforcement learning. Its common elements include an agent, an environment, a state, an action, and a reward. After the agent performs an action in its current state, it receives a certain reward, and its state changes based on the action. The goal of deep reinforcement learning is to calculate the maximum cumulative reward based on the reward at each step.
[0026] action By the policy function Decision, after state transition, the next state By the state transition function The decision,strategy and state transition functions are conditional probability distribution functions.,During the interaction between the agent and the environment, the uncertainty of,behavior and state leads to the randomness of MDP.
[0027] ; From the initial state Initially, the agent executes Action , until the completion signal. The agent experiences a complete event that conforms to the Markov assumption, that is, In the cycle Afterwards, the reward decay factor Rewards To indicate cumulative rewards : ; Action-value function is the conditional expectation of reward. For a given policy , Describes the current state The quality of the next action.
[0028] ; State Value Function It's an action set The action-value function expectations, it is only related to the strategy and status Related, describes the Downward Strategy Function The pros and cons: ; Therefore, the state value function Can guide the agent to choose the optimal strategy in the current state .
[0029] The action value function and the state value function obey the Bellman equation, which calculates the previous variable from the subsequent variable. , the environment receives an operation To get the next state , the return of this operation is expressed as The Bellman equations of the action value function and state value function are: ; In each time period The status of the solar energy storage charging station is defined as: ; This vector contains four types of information: (i) is the current solar radiation value, represents the solar radiation forecast for the next three hours; (ii) is the current price charged by the utility company to the power station for the required amount of electricity, is the predicted electricity price for the next three hours. Although the price changes dynamically throughout the day, simulating tariffs, it is not a function of supply and demand. Therefore, the price is dynamic but has nothing to do with the request amount; (iii) Indicates time period The state of charge of the electric vehicle at the first charging point; (iv) Indicates the time that the electric vehicle stays at the charging station. The last two states, namely and , which can be considered as the physical state of the solar energy storage charging station. exist is empty, then and are all 0; all states in the formula are normalized between 0 and 1.
[0030] A charging station consists of several charging piles, each of which can charge or discharge the connected electric vehicle. Therefore, multiple continuous variables are used. To define the charge and discharge rate of each vehicle charging point, it is constrained to be in the interval [−1,1]. In the time step The charge and discharge power is defined as: ; in For every electric car At the moment The charging status, 、 、 are the battery capacity, maximum charging power, and charging and discharging efficiency of the electric vehicle respectively. Therefore, the above three equations describe the actions taken by the controller at the time step The demand for each charging point is calculated. If the action is positive, then is positive (charging mode), and if the action is negative, Then it is negative (discharge mode). The value of directly affects the demand, as shown above. The second and third formulas describe the constraints on the maximum charge and discharge energy that can be allocated within a time step.
[0031] When performing an action When the vehicle is in the state of charge, the main change in the environment is the state of charge of the electric vehicle: ; The main goal of the controller of the solar energy storage charging station is to adopt a scheduling strategy to minimize the cost of purchasing electricity from the grid. The observed reward function is the electricity bill paid by the solar energy storage charging station to the utility company. However, to provide a more realistic and complete description, an additional term is added to ensure that the controller will effectively utilize the available resources and meet the defined requirements. Further considering the penalty for electric vehicles that are not fully charged, the equation describing this specific formula is as follows: ; in represents the total charging demand that the solar energy storage charging station requests from the utility company as described above; the electricity price Follow the different bill profiles provided by the utility company to the solar energy storage charging station in each time period, the unit is The second term in the formula is related to the state of charge of the EVs expected to depart in the next hour. The goal of the reward function is to fully charge the EVs that depart. However, in future implementations, EV owners can choose their desired state of charge when departing (to reduce charging costs), making this formula more realistic.
[0032] S2: Use the cyclic proximal strategy optimization algorithm to train the integrated solar-storage-charging scheduling model.
[0033] S2-1: Combining the proximal policy optimization algorithm in deep reinforcement learning with the recurrent proximal policy optimization algorithm obtained from the long short-term memory network, and applying it to the integrated photovoltaic storage and charging scheduling problem.
[0034] The Recurrent Proximal Policy Optimization (R-PPO) algorithm is an implementation of the Proximal Policy Optimization (PPO) algorithm that supports recurrent policies. By introducing hidden states, the algorithm enables memory when observations change over time. Aside from supporting recurrent policies, its behavior is identical to that of the PPO algorithm. Recurrent policies utilize Long Short-Term Memory (LSTM) networks. LSTM networks are a special type of recurrent neural network (RNN). When dealing with environments where current observations do not provide complete information about the system state, R-PPO can maintain a "memory" of past observations, which is crucial for making informed decisions.
[0035] LSTM network uses the previous state , Update its hidden state and memory status , its current observation state is: ; in, For status; is in hidden state; It is the memory state; They are input, forget gate, output, and candidate unit state gate respectively; are weight and bias respectively; is the activation function.
[0036] At a given hidden state After that, the strategy is: ; The value function is: ; Consider the PPO objective of the recursive state: ; As in standard PPO, the odds estimation is usually performed using the generalized odds estimation (GAE): ; Where, is the TD error, Both are discount factors.
[0037] Strategy The update is done by gradient ascent:
[0038] S2-2: Use the cyclic proximal strategy optimization algorithm to train a model for integrated solar-storage-charging scheduling.
[0039] The specific implementation steps are as follows: Data Collection: Data is collected by running interactions between the agent (charging station) and the environment. At each time step, the RNN policy outputs the action probability based on the current state.
[0040] Advantage evaluation: The advantage function is calculated using the rewards obtained during data collection. The advantage function reflects the degree of advantage of each action compared to the expected value in that state.
[0041] PPO objective: Calculating the PPO objective involves calculating the ratio between the new policy probability and the old policy probability, and then minimizing a clipped version of that ratio. This is done to prevent drastic policy updates.
[0042] Optimization: Use an optimization algorithm to update the RNN policy parameters according to the PPO objective. Common optimization methods include gradient descent-based algorithms.
[0043] Iterative process: Repeating the process of data collection, advantage estimation, PPO target calculation, and optimization to improve the strategy over time.
[0044] Training loop: Implement a training loop in which the agent interacts with the environment, collects data, estimates the advantage, calculates the PPO target, and iteratively updates the LSTM policy parameters. The ability of LSTM to capture temporal dependencies helps improve the performance of the agent over time.
[0045] Testing and Evaluation: Once a policy has been trained, its performance in the environment can be tested and evaluated by running experiments with simulated or real-world scenarios.
[0046] The R-PPO algorithm integrates LSTM to enhance the agent's decision-making ability in the electric vehicle charging station scenario by taking into account past observations. This allows the agent to make more informed decisions about charging and discharging rates to minimize electricity costs while efficiently utilizing solar energy.
[0047] Example 2 As a second embodiment of the present invention, a photovoltaic storage and charging integrated scheduling method based on a cyclic proximal strategy optimization algorithm is provided. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0048] Assume that each EVCS setup contains 10 charging points and a set of photovoltaic power generation systems. The SOC range of electric vehicles arriving at the EVCS is arrive The arrival time is between 0 and 22 hours, the shortest departure time is 2 hours after arrival, and the longest stay time can be up to the next day. The battery characteristics of all electric vehicles are the same. The charge and discharge efficiency of all electric vehicle batteries is the same. for , battery capacity for The maximum charging power of the charging point for Each charge and discharge determines the amount of electricity consumed in the next hour. The price EVCS pays for electricity from the grid is a peak-valley fixed price. The price is also the purchase price of electricity. The peak and valley electricity price is set to 0-7 hours, 8-20 hours, The environmental simulation platform was built using the Pytorch framework on a Linux server.
[0049] The training hyperparameters are set as follows: the number of episodes per training is set to 24, that is, one training session per day. is 0.99. The learning rate of the neural network is , the number of hidden layer units of the fully connected neural network is 400 and 300, the inner layer activation function is ReLU, and the output layer activation function is Tanh. The variance bounds selected This setting of 3 was determined based on initial experiments and statistical evaluation of gradient dispersion.
[0050] The trained algorithm was tested on ten different days, and the test results are shown in the table. The table compares the charging and discharging strategies based on rules, PPO, and R-PPO. The rule-based strategy refers to a control strategy based on rules. The centralized controller of the charging station checks each charging point and collects the departure time of each connected electric vehicle. If the electric vehicle If the vehicle leaves within 1 hour, the power station will fully charge the specific electric vehicle, otherwise the power station will use the current solar energy to charge the electric vehicle. .
[0051] Table 1. Test results of charge and discharge control strategies of solar-storage-charging systems under different methods
[0052] The data shown in the table reflects that in the charging and discharging scheduling of solar-storage-charging systems, the R-PPO algorithm demonstrates significant advantages in reducing costs, improving charging efficiency, and optimizing the accuracy of battery state of charge compared to rule-based strategies and the PPO algorithm. Specifically, the R-PPO algorithm has the lowest average cost, indicating that it is more economically efficient; its charging usage is also the lowest among the three strategies, suggesting its optimization in energy use; and in terms of SOC deviation, the R-PPO algorithm also shows a smaller error, which means that it is more accurate in ensuring that the battery state of charge reaches the predetermined target. Overall, the R-PPO algorithm has effectively improved the operational efficiency of electric vehicle charging stations through refined control and optimization. This may be due to its reinforcement learning characteristics, especially its ability to handle sequential decision-making problems, which enables it to better predict and adjust charging strategies based on historical data and the current environment.
[0053] Example 3 The third embodiment of the present invention is different from the first two embodiments in that: If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0054] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0055] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in another suitable manner as necessary, and then storing it in a computer memory.
[0056] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0057] Example 4 refer to Figure 2 , which is the fourth embodiment of the present invention, provides an integrated photovoltaic storage and charging scheduling system based on a cyclic proximal strategy optimization algorithm, characterized by: including a structure building module, a Markov decision module, an application module and a scheduling model training module; The structural construction module is used to build the physical system structure of the solar energy storage charging station; The Markov decision module defines the state space, action space, and reward function in the integrated solar-storage-charging scheduling process to build a Markov decision process; The application module combines the proximal policy optimization algorithm in deep reinforcement learning with the recurrent proximal policy optimization algorithm obtained from the long short-term memory network, and applies it to the integrated solar-storage-charging scheduling problem; The scheduling model training module uses a cyclic proximal strategy optimization algorithm to train a model for integrated solar-storage-charging scheduling.
[0058] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A photovoltaic storage and charging integrated scheduling method based on a cyclic proximal strategy optimization algorithm, characterized by: include, Build the physical architecture of the solar energy storage charging station; Define the state space, action space, and reward function in the integrated solar-storage-charging scheduling process to build a Markov decision process; A recurrent proximal policy optimization algorithm derived from deep reinforcement learning and a long short-term memory network is applied to the integrated solar-storage-charging scheduling problem. A model for integrated solar-storage-charging scheduling is trained using a cyclic proximal strategy optimization algorithm.
2. The integrated photovoltaic storage and charging scheduling method based on the cyclic proximal strategy optimization algorithm according to claim 1 is characterized in that: The physical architecture of the solar energy storage charging station includes a set of multiple charging points, a photovoltaic power generation system, an external power grid that provides electricity at a specific price, and a vehicle-to-grid operation mode. When the available energy is insufficient, the charging station purchases electricity from the grid with fluctuating electricity prices.
3. The integrated photovoltaic storage and charging scheduling method based on the cyclic proximal strategy optimization algorithm according to claim 2 is characterized in that: The Markov decision process includes the actions By the policy function Decision, after state transfer, the second state By the state transition function The decision, strategy and state transition functions are conditional probability distribution functions. In the interaction process between the agent and the environment, that is, the interaction process between the charging station and the environment, the uncertainty of behavior and state leads to the randomness of MDP. ; Among them, A represents the action space, S represents the state space, and P represents the conditional probability distribution function; From the initial state Initially, the agent executes Action , until the completion signal, the agent experiences a complete event that conforms to the Markov assumption, i.e. , in the cycle Afterwards, the reward decay factor Rewards To indicate cumulative rewards : ; in, Expressed as the reward at time t, Expressed as the reward at time t+1, Action-value function is the conditional expectation of reward for a given policy , Describes the current state The quality of the next action, ; State Value Function It's an action set The action-value function The expectation of the state value function is only related to the policy and status Related, describes the Downward Strategy Function The pros and cons: ; State Value Function Guide the agent to choose the optimal strategy in the current state ; The action value function and the state value function obey the Bellman equation. The Bellman equation calculates the previous variable from the subsequent variable. For the current state , the environment receives an operation To get the next state , the return of the operation is expressed as , the Bellman equation of the action value function and state value function is: ; In each time period The status of the solar energy storage charging station is defined as: ; in: is the current solar radiation value, represents the solar radiation forecast for the next three hours; is the current price charged by the utility company to the power station for the required amount of electricity, is the predicted electricity price for the next three hours; Indicates time period The state of charge of the electric vehicle at the first charging point; Indicates the time that electric vehicles stay at charging stations. and Indicates the physical state of the solar energy storage charging station. If the charging point exist is empty, then and are all 0; all states in the formula are normalized between 0 and 1.
4. The integrated photovoltaic storage and charging scheduling method based on the cyclic proximal strategy optimization algorithm according to claim 3 is characterized by: The Markov decision process also includes using multiple continuous variables To define the charge and discharge rate of each vehicle charging point, it is constrained in the interval [−1,1], and each vehicle In the time step The charge and discharge power is defined as: ; in, For every electric car At the moment The charging status, For electric vehicle battery capacity, For maximum charging power, For the best charge and discharge efficiency, the actions taken by the controller are based on the time step The above calculation of the demand for each charging point, if the action is positive, then is positive, if the action is negative, If it is negative, the action The value of directly affects the demand, is the actual charge and discharge amount of the i-th car at time t, is the maximum capacity that can be charged or discharged, is the charging and discharging rate of the i-th vehicle charging point at time t. When the vehicle is in the state of charge, the main change in the environment is the state of charge of the electric vehicle: ; The goal of the controller of the solar energy storage charging station is to adopt a scheduling strategy to minimize the cost of purchasing electricity for the solar energy storage charging station, taking into account the penalty for electric vehicles that are not fully charged. The equation of the formula is expressed as, ; in, Indicates the total charging demand that the solar energy storage charging station requests from the utility company; electricity price Follow the different bill profiles provided by the utility company to the solar-storage charging station at each time period.
5. The integrated PV-storage-charging scheduling method based on the cyclic proximal strategy optimization algorithm according to claim 4 is characterized in that: The long short-term memory network utilizes the previous state , Update hidden state and memory status , the current observation status is: ; in, For status; is in hidden state; It is the memory state; They are input, forget gate, output, and candidate unit state gate respectively; are weight and bias respectively; is the activation function, is the hidden state at time t-1, is the input gate weight, is the forget gate weight, is the output gate weight, is the candidate unit state gate weight, is the memory state at time t-1, is the input gate bias term, is the forget gate bias term, is the output gate bias term, is the candidate unit state gate bias term.
6. The integrated photovoltaic storage and charging scheduling method based on the cyclic proximal strategy optimization algorithm according to claim 5 is characterized in that: The long short-term memory network also includes, in a given hidden state After that, the strategy is: ; in, For strategic deviation, is the strategy weight, is the activation function; The value function is: ; Among them, V is the state value function, W val is the weight of the value function, b val is the bias of the value function, Represented as policy network parameters; Consider the PPO objective of the recursive state: ; in, is the objective function used to optimize the strategy in the PPO algorithm, Select actions for the new strategy The ratio of the probability of choosing the action to the probability of choosing the action under the old strategy, To limit the function to prevent the update step from being too large, for Minimum value, for Maximum value, is the estimate of the advantage function, For The old policy in the state, As in standard PPO, the odds estimation is performed using the generalized odds estimation: ; in, is the TD error used to estimate the update of the state value function, are discount factors, is the TD error at time t+1, is the TD error at time t+2; Strategy The update is done by gradient ascent: ; in, Expressed as policy network parameters, Expressed as parameter update change.
7. The integrated PV-storage-charging scheduling method based on the cyclic proximal strategy optimization algorithm according to claim 6 is characterized in that: The model for training integrated solar-storage-charging scheduling collects data by running interactions between intelligent agents, namely charging stations, and the environment. At each time step, the RNN strategy outputs an action probability based on the current state, and uses the rewards obtained during data collection to calculate the advantage function. The advantage function reflects the degree of advantage of each action compared to the expected value in the state. The PPO objective is calculated by calculating the ratio between the new strategy probability and the old strategy probability, and taking the minimum value using a trimmed version of the ratio. Use an optimization algorithm to update the RNN policy parameters according to the PPO objective; repeat the process of data collection, advantage estimation, PPO objective calculation, and optimization to improve the policy over time; Implement a training loop where the agent interacts with the environment, collects data, estimates the advantage, computes the PPO objective, and iteratively updates the LSTM policy parameters.
8. A system based on the PV-storage-charging integrated scheduling method based on the cyclic proximal strategy optimization algorithm according to any one of claims 1 to 7, characterized in that: Including structure building module, Markov decision module, application module and scheduling model training module; The structural construction module is used to build the physical system structure of the solar energy storage charging station; The Markov decision module defines the state space, action space, and reward function in the integrated solar-storage-charging scheduling process to build a Markov decision process; The application module combines the proximal policy optimization algorithm in deep reinforcement learning with the recurrent proximal policy optimization algorithm obtained from the long short-term memory network, and applies it to the integrated solar-storage-charging scheduling problem; The scheduling model training module uses a cyclic proximal strategy optimization algorithm to train a model for integrated solar-storage-charging scheduling.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.