A scheduling and control method for a photovoltaic-storage-direct-current flexible system

By using the pulse empirical playback TD3 algorithm combined with the priority mechanism of pulse neural network and empirical playback in the optical storage direct and flexible system, the problems of inefficient energy efficiency and insufficient sample distribution in the existing technology are solved, and economic scheduling and efficient operation of the optical storage direct and flexible system are achieved.

CN119209680BActive Publication Date: 2025-05-27STATE GRID JIANGXI ELECTRIC POWER CO LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411747215.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-27
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

When the prior art deals with complex continuous control problems, the inefficient energy efficiency of deep neural networks and insufficient robustness of sample distribution lead to problems such as dimensional explosion and reduced training results in the scheduling of optical storage direct-flexible system.

Method used

The pulse experience playback reinforcement learning algorithm is adopted, which combines the priority mechanism of pulse neural network and experience playback. The optimization scheduling model of the optical storage direct and flexible system is optimized through the pulse experience playback TD3 algorithm to realize economic scheduling of the optical storage direct and flexible system.

Benefits of technology

It effectively reduces the dependence of the optical storage direct and flexible system on the external network, reduces the system operation cost and light abandonment rate, and improves the photovoltaic absorption efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119209680B_ABST
    Figure CN119209680B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of photovoltaic-storage-direct-current-soft (PV-ESS-DC-Soft) systems, and particularly to a dispatching and control method for a PV-ESS-DC-Soft system, comprising the following steps: establishing an operation architecture model of the PV-ESS-DC-Soft system; using the DC bus voltage deviation as an adjustment signal to formulate an electricity consumption power control strategy for the PV-ESS-DC-Soft system; establishing an optimal dispatching model of the PV-ESS-DC-Soft system based on the operation architecture model and the electricity consumption power control strategy of the PV-ESS-DC-Soft system; optimizing the optimal dispatching model of the PV-ESS-DC-Soft system based on the pulsed experience replay reinforcement learning algorithm to obtain an optimal dispatching reinforcement learning model of the PV-ESS-DC-Soft system, and dispatching the PV-ESS-DC-Soft system through the optimal dispatching reinforcement learning model of the PV-ESS-DC-Soft system. The present invention uses the pulsed experience replay TD3 algorithm to solve and make decisions, which can effectively reduce the dependence of the PV-ESS-DC-Soft system on the external network and reduce the operation cost and light curtailment rate of the PV-ESS-DC-Soft system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photovoltaic-storage-direct-current-flexible (PV-SD-DC-F) systems, and particularly to a scheduling control method for a PV-SD-DC-F system. Background Art

[0002] As one of the main bodies of energy consumption, buildings have great potential for load regulation. Therefore, exploring the electricity flexibility of buildings plays an important role in improving the regulation ability of the power system.

[0003] The PV-SD-DC-F system is a new concept for intelligent buildings. At present, some domestic and foreign scholars have studied the optimal scheduling of intelligent buildings. Utilizing flexible resources in multi-energy buildings to relieve the power supply pressure can effectively reduce the system operation cost. Existing research has shown that orderly charging and discharging management of electric vehicles can effectively improve the economy and renewable energy utilization rate of buildings with distributed power sources. With the high proportion of new energy sources such as wind and light and DC equipment connected to the distribution system, DC distribution has advantages such as small transmission loss and easy adjustment of load power compared with AC distribution. The PV-SD-DC-F system is proposed under this background to coordinately optimize the control of components such as sources, storage, and loads in the system. Through deep reinforcement learning for energy consumption control, while ensuring user satisfaction, it effectively improves the PV power consumption and reduces the dependence on energy storage batteries. Designing a control strategy to guide the scheduling of each flexible resource in the PV-SD-DC-F system is the key to realizing its zero-carbon economic operation.

[0004] In recent years, reinforcement learning, as an efficient method, has been widely used to handle complex sequential decision-making problems. At present, some reinforcement learning techniques have been successfully applied to this field. However, some related research can only handle discrete action spaces, and for complex systems, the problem of dimensional explosion is likely to occur, resulting in a decrease in the accuracy of training results. Therefore, more and more scholars adopt reinforcement learning algorithms with continuous decision-making capabilities in optimization scheduling. As the energy system becomes more complex, the Deep Deterministic Policy Gradient (DDPG) algorithm, which has problems such as overestimation of the state-action function value (Q function value), unstable policy update, and insufficient exploration, can no longer meet the requirements. Therefore, some scholars have applied the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm with double Q functions and target policy smoothing to low-carbon economic scheduling. However, traditional deep neural networks often consume a large amount of computing resources when dealing with complex high-dimensional state spaces and are difficult to make full use of sparse information. Spiking Neural Networks (SNNs), as the third generation of neural networks, can utilize the event-driven mechanism and population coding characteristics to significantly reduce the amount of computation and energy consumption, and are particularly suitable for real-time decision-making tasks.

[0005] To solve the problem of low energy efficiency of the traditional TD3 algorithm in dealing with complex continuous control problems and improve the robustness of the sample distribution in Experience Replay (ER) simultaneously, it is necessary to design a scheduling control method for a photovoltaic-storage-direct-current-soft (PV-Storage-DC-Soft) system. Summary of the Invention

[0006] The purpose of the present invention is to solve at least one of the technical problems existing in the prior art, and to provide a scheduling control method for a PV-Storage-DC-Soft system.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is as follows: A scheduling control method for a PV-Storage-DC-Soft system, comprising the following steps:

[0008] Step 1, establish an operation architecture model of the PV-Storage-DC-Soft system;

[0009] Step 2, use the DC bus voltage deviation as an adjustment signal to formulate an electricity consumption power control strategy for the PV-Storage-DC-Soft system;

[0010] Step 3, based on the operation architecture model and the electricity consumption power control strategy of the PV-Storage-DC-Soft system, establish an optimal scheduling model of the PV-Storage-DC-Soft system;

[0011] Step 4, combine the priority mechanism of the spiking neural network and Experience Replay, propose a spiking experience replay reinforcement learning algorithm, optimize the optimal scheduling model of the PV-Storage-DC-Soft system based on the spiking experience replay reinforcement learning algorithm, and schedule the PV-Storage-DC-Soft system through the optimized optimal scheduling model of the PV-Storage-DC-Soft system.

[0012] Further, in Step 1, it specifically includes:

[0013] In the PV-Storage-DC-Soft system, the flexible adjustable load model responds to the DC bus voltage signal as follows:

[0014] (1);

[0015] (2);

[0016] In the formula, , T represents the total time period of power adjustment, , N represents the total number of load types; is the load flexibility coefficient; are respectively the PV-Storage-DC-Soft system The adjusted power and the pre-adjustment power of the type load in the time period; is the DC bus voltage deviation of the PV-Storage-DC-Soft system at time period; is the Power of the type load; Respectively the Lower limit and upper limit of the power of the type load;

[0017] It is expressed by the following formula:

[0018] (3);

[0019] In the formula: Is the rated power when the corresponding flexible load does not perform power regulation; Respectively represent the maximum value, minimum value and rated voltage of the DC bus voltage;

[0020] In the optical storage DC flexible system, the energy storage battery BES is described by the state of charge SOC:

[0021] (4);

[0022] In the formula, Is the state of charge of BES at Time; Is the output of BES at Time, the value greater than 0 is discharging, and the value less than 0 is charging; Is the rated capacity of BES; Is the self-loss coefficient of BES; And Are the charge and discharge coefficients of BES respectively;

[0023] The corresponding model of the charge and discharge power of BES with the DC bus voltage signal is as follows:

[0024] (5);

[0025] The charge and discharge power constraints of BES are as follows:

[0026] (6);

[0027] The charge and discharge ramp constraints of BES are as follows:

[0028] (7);

[0029] The state of charge constraints of BES are as follows:

[0030] (8);

[0031] In the formula Respectively are The adjusted charge and discharge power and rated charge and discharge power of the energy storage battery during the period; Is the flexible coefficient of the energy storage battery; The landslide limit value and the climbing limit value of the energy storage battery, respectively; The maximum discharge power, the maximum charge power of the energy storage battery, and The working power of the energy storage battery during the time period; The set lower capacity limit and upper capacity limit of the energy storage battery, respectively;

[0032] In the EV charging model, the charging load of the electric vehicle EV at the building charging station is fitted. The charging load is related to the time when the electric vehicle arrives at the building charging station, the daily driving mileage of the electric vehicle, and the charging duration.

[0033] The time when the electric vehicle arrives at the building charging station follows a normal distribution, and its probability density function is:

[0034] (9);

[0035] In the formula: is the time when the electric vehicle arrives at the charging station; is the standard deviation of the fitting function of the arrival time of the electric vehicle at the station; is the mean value of the fitting function of the arrival time of the electric vehicle at the station;

[0036] The daily driving mileage of the electric vehicle also follows a normal distribution, and its probability density function is:

[0037] (10);

[0038] In the formula: is the daily driving mileage; is its standard deviation; is its mean value;

[0039] For the charging duration of the electric vehicle , it is determined by the state of charge of the electric vehicle before and after charging:

[0040] (11);

[0041] In the formula: are the states of charge of the battery before and after charging, respectively; represents the battery capacity of the electric vehicle; is the charging pile power; is the charging efficiency;

[0042] Among them, the actual state of charge after charging is related to the driving mileage and the initial state of charge of the electric vehicle:

[0043] (12);

[0044] In the formula: is the power consumption per 100 kilometers of the EV;

[0045] The Monte Carlo method is used to simulate the charging load in the disordered charging state in the PV - battery - DC - AC system;

[0046] In the PV - battery - DC - AC system, when the state of charge of the electric vehicle does not reach the minimum capacity set by the EV user, the charging pile charges the electric vehicle at the rated power; when it is higher than the set capacity, the charging pile and the electric vehicle are regarded as a storage battery, and the minimum capacity set by the EV user is determined by the Monte Carlo simulation result, that is:

[0047] (13);

[0048] In the formula, is the state of charge of the th electric vehicle at time; is the charging and discharging power of the th electric vehicle at time, the value greater than 0 means discharging, and the value less than 0 means charging; is the rated capacity of the EV; and are the charging and discharging coefficients of the EV respectively; , M represents the total number of electric vehicles;

[0049] The charging and discharging power of the electric vehicle responds to the DC bus voltage signal as follows:

[0050] (14);

[0051] In the formula, is the rated charging and discharging power of the th electric vehicle, is the required state of charge set by the user of the th electric vehicle;

[0052] The charging and discharging power constraints of the EV are as follows:

[0053] (15);

[0054] The charging and discharging ramp constraints of the EV:

[0055] (16);

[0056] The state of charge constraints of the EV are as follows:

[0057] (17);

[0058] Wherein, are respectively the maximum charging and discharging powers of the th vehicle and the operating power at moment; are respectively the landslide limit value and the climbing limit value of the battery of the th electric vehicle; are respectively the lowest and highest state of charge of the th vehicle and the state of charge at moment.

[0059] Furthermore, in step 2, it specifically includes:

[0060] Taking the DC bus voltage deviation as the adjustment signal for the power of adjustable devices in the PV-storage-DC-AC system: when , the power on the energy supply side is greater than the power on the demand side, and the power consumption of each device is increased; when , the power on the supply side is insufficient, and the power consumption of each device is reduced; therefore, by solving a suitable DC bus voltage signal, the economic operation of the PV-storage-DC-AC system is realized;

[0061] The powers of various types of loads, the charge and discharge powers of the battery and the electric vehicle adjust their own powers according to the DC bus signal and the corresponding device operation strategy;

[0062] The power balance equation of the PV-storage-DC-AC system is shown as the following formula:

[0063] (18);

[0064] Wherein, is the total power of all loads at moment, is the PV power at moment, is the power taken from the power grid by the PV-storage-DC-AC system at moment, is the PV curtailment power at moment, is the adjustment power of the flexible device;

[0065] The adjustment power of the flexible device is:

[0066] (19);

[0067] Wherein, is the parameter determined by the load connected to the DC bus of the PV-storage-DC-AC system at moment;

[0068] To improve PV accommodation and reduce the dependence on power taken from the external power grid, the adjustment power should be equal to the demand power of the PV-storage-DC-AC system, that is:

[0069] (20);

[0070] Wherein, is the required power of the photovoltaic-storage-direct-current-soft system during period. When the operating conditions of the photovoltaic-storage-direct-current-soft system satisfy Equation (20), that is, the total adjustment power of various flexible loads and flexible devices is equal to the required adjustment power of the photovoltaic-storage-direct-current-soft system. At this time, by combining Equation (19) and Equation (20), we get:

[0071] (21);

[0072] (22);

[0073] Wherein, is the system voltage gain coefficient. The building during this period is regarded as a zero-carbon building, and the curtailment of light is 0;

[0074] When the adjustment power of the photovoltaic-storage-direct-current-soft system cannot satisfy Equation (20), Equation (22) will not hold. At this time, a voltage difference signal needs to be found to guide the flexible devices of the photovoltaic-storage-direct-current-soft system to adjust power, so as to achieve the optimal economic operation of the photovoltaic-storage-direct-current-soft system; the voltage difference signal of the photovoltaic-storage-direct-current-soft system is given by Equation (21). By optimizing at each scheduling moment, the power adjustment scheme of the flexible devices of the photovoltaic-storage-direct-current-soft system is obtained, and the economic operation of the photovoltaic-storage-direct-current-soft system is realized.

[0075] Further, in step 3, it specifically includes:

[0076] The economic scheduling objective function of the optimized scheduling model of the photovoltaic-storage-direct-current-soft system is as follows:

[0077] (23);

[0078] (24);

[0079] Wherein: is the total operating cost of the photovoltaic-storage-direct-current-soft system during period; The operating cost of the photovoltaic-storage-direct-current-soft system during period; is the power scheduling cost of the flexible devices of the photovoltaic-storage-direct-current-soft system during period; is the penalty cost for curtailment of light during period; is the cost of purchasing electricity during period; are respectively the scheduling cost coefficient and the adjustment power of the th load during are respectively the maintenance cost coefficient and regulation power of the energy storage battery for a certain period; are respectively the scheduling cost coefficient of the electric vehicle for a certain period and the charging and discharging power of the electric vehicle; is the penalty coefficient for wind and light abandonment; is the amount of light abandonment for a certain period; is the electricity purchase price for a certain period; is the building electricity purchase for a certain period

[0080] The constraint conditions of the optimized scheduling model of the PV-storage-direct-flexibility system include the flexible load and equipment regulation power constraints in the operation architecture model of the PV-storage-direct-flexibility system, and the power balance equation of the PV-storage-direct-flexibility system in the electricity consumption power control strategy.

[0081] Furthermore, in step 4, it specifically includes:

[0082] Step 4.1, action space design:

[0083] The action space represents the set of change amounts of action variables at each time step;

[0084] (25);

[0085] is the action space, satisfying the constraint conditions of the optimized scheduling model of the PV-storage-direct-flexibility system;

[0086] Step 4.2, state space design:

[0087] The state space is as follows:

[0088] (26);

[0089] is the state space, which are respectively the electrical load, PV prediction value, electricity price, electricity purchase quantity, and the action quantity at the previous moment;

[0090] Step 4.3, reward function design:

[0091] Reinforcement learning continuously updates its own policy through the agent to maximize the expected reward value; therefore, the setting of the reward function should be related to the objective function in the scheduling model, and its reward function is as follows:

[0092] (27);

[0093] In the formula: They are respectively the reward value at a moment and the proportion coefficient of the total operating cost of the PV-storage-DC-AC system; It is a parameter for the reward function to return to a positive value;

[0094] Step 4.4, reinforcement learning algorithm:

[0095] Solve the reinforcement learning problem in the continuous action space through the pulsed experience replay TD3 algorithm; the policy update formula of TD3 is:

[0096] (28);

[0097] In the formula: is the TD3 algorithm policy; is the state quantity; is the state of the agent under; is the action value function; is the experience replay buffer;

[0098] Input historical data into the spiking neural network SNN. The spiking neural network encodes the historical data through LIF neurons, simulating the change of the membrane potential of neurons after receiving stimuli; in this way, the spiking neural network can capture the temporal correlation of state changes in continuous control problems, enabling better utilization of historical data during the experience replay process. The formula for experience replay is:

[0099] (29);

[0100] In the formula: The moment of corresponds to the moment of the scheduling model, is the state at the next moment; is the time constant, used to control the attenuation speed of the pulse signal. The larger the time constant, the longer the pulse influence time; is the initial moment; is the input pulse signal at the moment;

[0101] After the experience replay mechanism introduces the spiking neural network, the experience replay is weighted according to the frequency and intensity of pulse firing; the formula for priority is:

[0102] (30);

[0103] In the formula: represents the priority of the th sample, and this priority determines the frequency at which this sample is selected during replay; represents the The TD error of a sample. The larger the TD error, the higher the importance of the sample; is the control parameter for priority sampling, which controls the degree of influence of priority. represents the normalization factor, which is used to ensure that the sum of priorities is 1;

[0104] Step 4.5, The algorithm flow of Pulse Experience Replay TD3:

[0105] Step 4.51, Initialization: Initialize the Actor network and Critic network of the spiking neural network, set the policy network and value function network required by the TD3 algorithm, and initialize the spiking experience replay buffer;

[0106] Step 4.52, Data preprocessing: Perform spiking population coding on the state features in the offline dataset, and use normalized feature transformation to reduce data distribution differences;

[0107] Step 4.53, Spiking experience replay: Sample a batch of experiences from the spiking experience replay buffer, and weight the experience data according to the Q-value error output by the spiking neural network;

[0108] Step 4.54, Update the value function: Use the sampled spiking-coded experiences to update the value function based on the Critic network of TD3, and calculate the Q-value of the state-action pair;

[0109] Step 4.55, Update the policy: Decode the spiking signal through the Actor network to generate actions, and use the loss function with a spiking regularization term to update the policy network to minimize the behavior error;

[0110] Step 4.56, Update the target network: Periodically update the target value function and target policy network, and synchronize the parameters of the spiking neural network;

[0111] Step 4.57, Evaluate performance: Evaluate the performance of the policy at fixed intervals, and test the control effect of the Pulse Experience Replay TD3 algorithm through the simulation environment;

[0112] Step 4.58, Save the training model: Save the trained Pulse Experience Replay TD3 algorithm model for subsequent testing and application;

[0113] Step 4.6, The optimization scheduling process based on Pulse Experience Replay TD3:

[0114] Step 4.61, Input the photovoltaic historical data and load historical data into the reinforcement learning environment, and initialize the parameters of each network in the reinforcement learning;

[0115] Step 4.62: Conduct offline scheduling training, and obtain the state variables at the current moment from the reinforcement learning environment: electrical load, PV power prediction value, electricity price, electricity purchase quantity, and the action variable at the previous moment.

[0116] Step 4.63: The agent gives an action based on the input state.

[0117] Step 4.64: Execute the action in the reinforcement learning environment to obtain the state at the next moment.

[0118] Step 4.65: When the set number of training times is reached, save the optimal scheduling model of the PV-storage-DC-AC system.

[0119] Step 4.66: Input the real-time data into the saved training model, and the optimal scheduling model of the PV-storage-DC-AC system gives the real-time scheduling strategy.

[0120] As can be seen from the above description of the present invention, compared with the prior art, the present invention includes the following beneficial effects: Based on the direct coupling relationship between the bus voltage and the load power in the DC system, the present invention guides the change of the DC bus through the source-load power difference in the system to schedule various flexible loads and devices in the system; and uses the proposed Pulse Experience Replay TD3 algorithm to achieve the economic scheduling of the PV-storage-DC-AC system. The present invention first establishes a mathematical model of the PV-storage-DC-AC system considering PV, energy storage, electric vehicles, and various flexible loads, and at the same time converts this mathematical model into a reinforcement learning environment, and uses the Pulse Experience Replay TD3 algorithm to solve and make decisions, which can effectively reduce the dependence of the PV-storage-DC-AC system on the external power grid and reduce the system operation cost and PV curtailment rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] Figure 1 It is a step flowchart of a scheduling control method for a PV-storage-DC-AC system in a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0122] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0123] Refer to Figure 1 As shown, in a preferred embodiment of the present invention, a scheduling control method for a PV-storage-DC-AC system includes the following steps:

[0124] Step 1: Establish an operation architecture model of the PV-storage-DC-AC system.

[0125] Step 2: Taking the DC bus voltage deviation as the adjustment signal, formulate the power consumption control strategy for the optical storage and direct current flexible (OSDCF) system;

[0126] Step 3: Based on the operation architecture model and the power consumption control strategy of the OSDCF system, establish the optimal scheduling model of the OSDCF system;

[0127] Step 4: Combine the pulse neural network and the priority mechanism of experience replay, propose the pulse experience replay reinforcement learning algorithm, optimize the optimal scheduling model of the OSDCF system based on the pulse experience replay reinforcement learning algorithm, and schedule the OSDCF system through the optimized optimal scheduling model of the OSDCF system.

[0128] As a preferred embodiment of the present invention, it may further have the following additional technical features:

[0129] In this embodiment, in Step 1, it specifically includes:

[0130] In the OSDCF system, the flexible adjustable load model responds to the DC bus voltage signal as follows:

[0131] (1);

[0132] (2);

[0133] In the formula, , T represents the total time period of power adjustment, , N represents the total number of load types; is the load flexibility coefficient; are respectively the adjusted power and the pre-adjustment power of the -th type of load in the OSDCF system during the time period; is the DC bus voltage deviation of the OSDCF system during the time period; is the power of the -th type of load; are respectively the lower limit and the upper limit of the power of the -th type of load;

[0134] It is represented by the following formula:

[0135] (3);

[0136] In the formula: is the rated power when the corresponding flexible load does not perform power adjustment; respectively represent the maximum value, the minimum value and the rated voltage of the DC bus voltage;

[0137] In the photovoltaic-storage-direct-current-soft system, the state of charge (SOC) is used to describe the energy storage battery (BES):

[0138] (4);

[0139] In the formula, is the state of charge of the BES at moment; is the output of the BES at moment. A value greater than 0 means discharging, and a value less than 0 means charging; is the rated capacity of the BES; is the self-loss coefficient of the BES; and are the charging and discharging coefficients of the BES respectively;

[0140] The corresponding model of the charging and discharging power of the BES with the DC bus voltage signal is as follows:

[0141] (5);

[0142] The constraints on the charging and discharging power of the BES are as follows:

[0143] (6);

[0144] The ramping constraints on the charging and discharging of the BES are as follows:

[0145] (7);

[0146] The constraints on the state of charge of the BES are as follows:

[0147] (8);

[0148] In the formula are respectively the adjusted charging and discharging power and the rated charging and discharging power of the energy storage battery during the period; is the flexibility coefficient of the energy storage battery; are respectively the sliding limit and the ramping limit of the energy storage battery; are respectively

[0149] In the EV charging model, the charging load of electric vehicles (EVs) at the building charging station is fitted. The charging load is related to the time when the EVs arrive at the building charging station, the daily driving mileage of the EVs, and the charging duration;

[0150] The time when the EVs arrive at the building charging station follows a normal distribution, and its probability density function is:

[0151] (9);

[0152] Where: is the time when the electric vehicle arrives at the charging station; is the standard deviation of the fitting function of the arrival time of the electric vehicle at the station; is the mean value of the fitting function of the arrival time of the electric vehicle at the station;

[0153] The daily driving mileage of the electric vehicle also follows a normal distribution, and its probability density function is:

[0154] (10);

[0155] Where: is the daily driving mileage; is its standard deviation; is its mean value;

[0156] For the charging duration of the electric vehicle, it is determined by the state of charge of the electric vehicle before and after charging:

[0157] (11);

[0158] Where: are the states of charge of the battery before and after charging, respectively; represents the battery capacity of the electric vehicle; is the power of the charging pile; is the charging efficiency;

[0159] Among them, the actual state of charge after charging is related to the driving mileage and the initial state of charge of the electric vehicle:

[0160] (12);

[0161] Where: is the power consumption per 100 kilometers of the EV;

[0162] The Monte Carlo method is used to simulate the charging load in the disordered charging state in the PV-storage-DC-AC system.

[0163] In the PV-storage-DC-AC system, when the state of charge of the electric vehicle does not reach the minimum capacity set by the EV user, the charging pile charges the electric vehicle at the rated power; when it is higher than the set capacity, the charging pile and the electric vehicle are regarded as an energy storage battery, and the minimum capacity set by the EV user is determined by the Monte Carlo simulation results, that is:

[0164] (13);

[0165] Where, is the state of charge of the th electric vehicle at moment; is the charging and discharging power of the th electric vehicle at moment. A value greater than 0 indicates discharging, and a value less than 0 indicates charging; is the rated capacity of the EV; and are the charging and discharging coefficients of the EV respectively; , M represents the total number of electric vehicles;

[0166] The charging and discharging power of the electric vehicle responds to the DC bus voltage signal as follows:

[0167] (14);

[0168] In the formula, is the rated charging and discharging power of the th electric vehicle, is the required state of charge set by the user of the th electric vehicle;

[0169] The charging and discharging power constraints of the EV are as follows:

[0170] (15);

[0171] The charging and discharging ramp constraints of the EV:

[0172] (16);

[0173] The state of charge constraints of the EV are as follows:

[0174] (17);

[0175] In the formula, are the maximum charging and discharging power and the working power at moment of the th vehicle respectively; are the slide limit value and the ramp limit value of the battery of the th electric vehicle respectively; are the lowest and highest state of charge and the state of charge at moment of the th vehicle respectively.

[0176] In this embodiment, in step 2, it specifically includes:

[0177] The DC bus voltage deviation As the adjustment signal for adjusting the power of adjustable devices in the photovoltaic-storage-direct-current-soft (PV-Storage-DC-Soft) system: When the power on the energy supply side is greater than the power on the demand side, increase the power consumption of each device; when the power supply on the supply side is insufficient, reduce the power consumption of each device; Therefore, by solving a suitable DC bus voltage signal, the economic operation of the PV-Storage-DC-Soft system can be achieved;

[0178] The power of each type of load, the charge and discharge power of the battery and electric vehicle adjust their own power according to the DC bus signal and the corresponding device operation strategy;

[0179] The power balance equation of the PV-Storage-DC-Soft system is shown as follows:

[0180] (18);

[0181] In the formula, is the total power of all loads at time is the photovoltaic power at time is the power taken from the power grid by the PV-Storage-DC-Soft system at time is the curtailment power at time is the adjustment power of the flexible device;

[0182] The adjustment power of the flexible device is:

[0183] (19);

[0184] In the formula, is the parameter determined by the load connected to the DC bus of the PV-Storage-DC-Soft system at time

[0185] To improve the photovoltaic accommodation and reduce the dependence on power taken from the external power grid, the adjustment power should be equal to the demand power of the PV-Storage-DC-Soft system, that is:

[0186] (20);

[0187] In the formula, is the power required by the PV-Storage-DC-Soft system in time period. When the operating conditions of the PV-Storage-DC-Soft system satisfy formula (20), that is, the total adjustment power of various flexible loads and flexible devices is equal to the demand adjustment power of the PV-Storage-DC-Soft system. At this time, by combining formula (19) and formula (20), we get:

[0188] (21);

[0189] (22);

[0190] In the formula, is the system voltage gain coefficient. During this period, the building is regarded as a zero-carbon building and the curtailment of solar power is 0;

[0191] When the regulation power of the PV-ESS-DR system cannot meet Equation (20), Equation (22) will not hold. At this time, a voltage difference signal needs to be found to guide the flexible equipment of the PV-ESS-DR system to adjust power, so as to achieve the optimal economic operation of the PV-ESS-DR system; the voltage difference signal of the PV-ESS-DR system is given by Equation (21). By optimizing at each scheduling moment, the power adjustment plan of the flexible equipment of the PV-ESS-DR system is obtained, and the economic operation of the PV-ESS-DR system is realized.

[0192] In this embodiment, in step 3, it specifically includes:

[0193] The economic scheduling objective function of the PV-ESS-DR system optimization scheduling model is as follows:

[0194] (23);

[0195] (24);

[0196] In the formula: is the total operating cost of the PV-ESS-DR system during the PV-ESS-DR system operating cost during the is the power scheduling cost of the flexible equipment of the PV-ESS-DR system during the is the curtailment of solar power penalty cost during the is the cost of purchasing electricity during the are respectively the scheduling cost coefficient and regulation power of the th type of load during the are respectively the maintenance cost coefficient and regulation power of the energy storage battery during the are respectively the scheduling cost coefficient of the electric vehicle and the charging and discharging power of the th electric vehicle during the is the curtailment of solar power during the is the electricity purchase price during the is the electricity purchase of the building during the

[0197] The constraint conditions of the optimal scheduling model for the photovoltaic-storage-direct-current-flexible system include the flexible load in the operation architecture model of the photovoltaic-storage-direct-current-flexible system, the equipment regulation power constraint, and the power balance equation of the photovoltaic-storage-direct-current-flexible system in the power consumption control strategy.

[0198] In this embodiment, in step 4, it specifically includes:

[0199] Step 4.1, action space design:

[0200] The action space represents the set of change amounts of action variables at each time step;

[0201] (25);

[0202] is the action space, satisfying the constraint conditions of the optimal scheduling model for the photovoltaic-storage-direct-current-flexible system;

[0203] Step 4.2, state space design:

[0204] The state space is as follows:

[0205] (26);

[0206] is the state space, which are the electrical load, photovoltaic power prediction value, electricity price, electricity purchase quantity, and the action quantity at the previous moment, respectively;

[0207] Step 4.3, reward function design:

[0208] Reinforcement learning continuously updates its own policy through the agent to maximize the expected reward value; therefore, the setting of the reward function should be related to the objective function in the scheduling model, and its reward function is as follows:

[0209] (27);

[0210] In the formula: are respectively the reward value at the moment and the total operation cost ratio coefficient of the photovoltaic-storage-direct-current-flexible system; is the parameter for the reward function to return to a positive value;

[0211] Step 4.4, reinforcement learning algorithm:

[0212] Solve the reinforcement learning problem in the continuous action space through the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm; the policy update formula of TD3 is:

[0213] (28);

[0214] In the formula: It is the TD3 algorithm strategy; is the state variable; is the state of the agent under this state; is the action value function; is the experience replay buffer;

[0215] Input historical data into the spiking neural network (SNN). The SNN encodes the historical data through LIF neurons to simulate the change of the membrane potential of neurons after receiving stimuli. In this way, the SNN can capture the temporal correlation of state changes in continuous control problems, enabling better utilization of historical data during the experience replay process. The formula for experience replay is:

[0216] (29);

[0217] In the formula: The moment of corresponds to the moment of the scheduling model, is the state at the next moment; is the time constant, which is used to control the decay speed of the pulse signal. The larger the time constant, the longer the pulse influence time; is the input pulse signal at the moment of

[0218] After introducing the experience replay mechanism into the spiking neural network, the experience replay is weighted according to the frequency and intensity of spike firing. The formula for priority is:

[0219] (30);

[0220] In the formula: represents the priority of the th sample, and this priority determines the frequency at which this sample is selected during replay; represents the TD error of the th sample. The larger the TD error, the higher the importance of this sample; is the regulation parameter for priority sampling, which controls the influence degree of priority, represents the normalization factor, which is used to ensure that the sum of priorities is 1;

[0221] Step 4.5, the algorithm process based on the spiking experience replay TD3:

[0222] Step 4.51, initialization: Initialize the Actor network and Critic network of the spiking neural network, set the policy network and value function network required by the TD3 algorithm, and initialize the spiking experience replay buffer;

[0223] Step 4.52, Data preprocessing: Perform pulse population coding on the state features in the offline dataset, and use normalized feature transformation to reduce the data distribution difference;

[0224] Step 4.53, Pulse experience replay: Sample a batch of experiences from the pulse experience replay buffer, and weight the experience data according to the Q-value error output by the pulse neural network;

[0225] Step 4.54, Update value function: Use the sampled pulse-coded experiences to update the value function based on the Critic network of TD3, and calculate the Q-value of the state-action pair;

[0226] Step 4.55, Update policy: Decode the pulse signal through the Actor network to generate actions, and use the loss function with pulse regularization terms to update the policy network to minimize the behavior error;

[0227] Step 4.56, Update target network: Periodically update the target value function and the target policy network, and synchronize the parameters of the pulse neural network;

[0228] Step 4.57, Evaluate performance: Evaluate the performance of the policy at fixed intervals, and test the control effect of the pulse experience replay TD3 algorithm through the simulation environment;

[0229] Step 4.58, Save the training model: Save the trained pulse experience replay TD3 algorithm model for subsequent testing and application;

[0230] Step 4.6, Optimization scheduling process based on pulse experience replay TD3:

[0231] Step 4.61, Input the photovoltaic historical data and load historical data into the reinforcement learning environment, and initialize the parameters of each network in the reinforcement learning;

[0232] Step 4.62, Conduct offline scheduling training, and obtain the state variables at the current moment from the reinforcement learning environment: electrical load, photovoltaic prediction value, electricity price, electricity purchase quantity, and the action quantity at the previous moment;

[0233] Step 4.63, The agent gives an action according to the input state;

[0234] Step 4.64, Execute the action in the reinforcement learning environment to obtain the state at the next moment;

[0235] Step 4.65, When the set number of training times is reached, save the optimization scheduling model of the integrated energy system with flexible DC collection;

[0236] Step 4.66, Input the real-time data into the saved training model, and the optimization scheduling model of the integrated energy system with flexible DC collection gives the real-time scheduling strategy.

[0237] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its improved concept, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.

Claims

1. A scheduling control method for a PV-storage direct-flexible system, characterized in that: The following steps are involved: Step 1: Establish an operation architecture model of the PV-storage direct-flexible system; Step 2, using the DC bus voltage deviation as a regulation signal, formulate a power control strategy for the solar-storage direct-flexible system; Step 3: Based on the operation architecture model of the PV-storage direct-flexible system and the power consumption control strategy, an optimization scheduling model of the PV-storage direct-flexible system is established; Step 4, combining the pulse neural network and the priority mechanism of experience replay, proposing a pulse experience replay reinforcement learning algorithm, optimizing the optimal scheduling model of the photovoltaic storage direct-flexible system based on the pulse experience replay reinforcement learning algorithm, and scheduling the photovoltaic storage direct-flexible system through the optimized optimal scheduling model of the photovoltaic storage direct-flexible system; In step 1, specifically include: In the PV-storage DC-flexible system, the flexible adjustable load model responds to the DC bus voltage signal as follows: (1); (2); In the formula, , T represents the total period of power adjustment, , N represents the total number of load types; is the load flexibility coefficient; Solar storage direct and flexible system t Time period n The regulated power and unregulated power of the type of load; For the solar storage direct and flexible system t DC bus voltage deviation during the time period; For the n Power of type load; Respectively n The lower and upper limits of power for the type of load; It is expressed by the following formula: (3); Where: It is the rated power when no power adjustment is performed for the corresponding flexible load; Respectively represent the maximum value, minimum value and rated voltage of the DC bus voltage; In the solar-storage direct-flexible system, the energy storage battery BES uses the state of charge SOC to describe: (4); In the formula, For BES t The state of charge at the moment; For BES t Output at the moment, a value greater than 0 indicates discharging, and a value less than 0 indicates charging; is the rated capacity of BES; is the self-loss coefficient of BES; and are the BES charge and discharge coefficients, respectively; In step 2, specifically including: The DC bus voltage deviation As a regulating signal for the power of adjustable devices in the PV-storage direct-flexible system: When the power on the energy supply side is greater than the power on the demand side, the power consumption of each device increases; when When the power on the supply side is insufficient, the power consumption of each device is reduced; therefore, the economic operation of the photovoltaic storage direct-flexible system is achieved by solving a suitable DC bus voltage signal; The power of various types of loads, batteries and electric vehicle charging and discharging powers adjust their own power according to the DC bus signal and the corresponding equipment operation strategy; The power balance equation of the PV-storage direct-flexible system is as follows: (18); In the formula, for t The total power of all loads at the moment, for t Photovoltaic power at all times, for t The solar-storage direct-flexible system draws power from the grid at all times. for t Always discard optical power. regulating power for flexible devices; The regulated power of the flexible device is: (19); In the formula, for t Parameters determined by the load connected to the DC bus of the solar-storage-DC-flexible system at any given moment; for t Moment m The charging and discharging power of an electric vehicle. A value greater than 0 indicates discharging, and a value less than 0 indicates charging. , M Expressed as the total number of electric vehicles; In order to improve photovoltaic consumption and reduce dependence on external power grid, the regulated power should be equal to the required power of the photovoltaic storage direct-flexible system, that is: (20); In the formula, For the solar storage direct and flexible system t The power required for the time period is: when the operating conditions of the PV-storage direct-flexible system meet equation (20), that is, the total regulated power of various flexible loads and flexible equipment is equal to the required regulated power of the PV-storage direct-flexible system, then the combined equations (19) and (20) yield: (21); (22); In the formula, is the system voltage gain coefficient. The building during this period is considered a zero-carbon building, and the abandoned solar power is 0; When the PV-storage direct-flexible system cannot adjust the power to meet equation (20), equation (22) will not hold. At this time, it is necessary to find a voltage difference signal to guide the flexible equipment of the PV-storage direct-flexible system to adjust the power and achieve the optimal economic operation of the PV-storage direct-flexible system. The voltage difference signal of the PV-storage direct-flexible system is given by equation (21). By optimizing the , and obtain the power adjustment scheme of the flexible equipment of the photovoltaic storage direct-flexible system to achieve the economic operation of the photovoltaic storage direct-flexible system.

2. The scheduling control method of a PV-storage direct-flexible system according to claim 1, characterized in that: In step 1, specifically include: The corresponding model of BES charging and discharging power and DC bus voltage signal is as follows: (5); BES charging and discharging power constraints are as follows: (6); BES charging and discharging ramp constraints are as follows: (7); The BES state of charge constraints are as follows: (8); In the formula, They are t The adjusted charging and discharging power and rated charging and discharging power of the time period energy storage battery; is the flexibility coefficient of the energy storage battery; They are the landslide limit and climbing limit of the energy storage battery respectively; are the maximum discharge power, maximum charging power and t Working power of time storage battery; The lower and upper capacity limits are set for the energy storage battery respectively; In the EV charging model, the charging load of the electric vehicle EV at the building charging station is fitted. The charging load is related to the time when the electric vehicle arrives at the building charging station, the daily mileage of the electric vehicle and the charging time. The time it takes for electric vehicles to arrive at a building charging station follows a normal distribution, and its probability density function is: (9); Where: The time it takes for electric vehicles to reach the charging station; is the standard deviation of the fitted function of the electric vehicle arrival time; is the mean of the fitting function of the electric vehicle arrival time; The daily mileage of electric vehicles also follows a normal distribution, and its probability density function is: (10); Where: is the daily mileage; is its standard deviation; is its mean; Charging time for electric vehicles , determined by the state of charge of the electric vehicle before and after charging: (11); Where: They are the state of charge of the battery before and after charging; Indicates the battery capacity of an electric vehicle; is the charging pile power; For charging efficiency; The actual state of charge after charging is related to the mileage and initial state of charge of the electric vehicle: (12); Where: The power consumption per 100 kilometers of the EV; The Monte Carlo method is used to simulate the charging load under disordered charging state in the solar-storage direct-flexible system. In the solar-storage direct-flexible system, when the state of charge of the electric vehicle does not reach the minimum capacity set by the EV user, the charging pile charges the electric vehicle at the rated power; when it is higher than the set capacity, the charging pile and the electric vehicle are regarded as a storage battery, and the minimum capacity set by the EV user is determined by the Monte Carlo simulation results, that is: (13); In the formula, For the m Electric vehicles in t The state of charge at the moment; for t Moment m The charging and discharging power of an electric vehicle. A value greater than 0 indicates discharging, and a value less than 0 indicates charging. is the rated capacity of the EV; is the self-loss coefficient of EV; and are the EV charging and discharging coefficients respectively; , M Expressed as the total number of electric vehicles; The charging and discharging power of electric vehicles responds to the DC bus voltage signal as follows: (14); In the formula, For the m The rated charging and discharging power of an electric vehicle, For the m The required state of charge set by the user of an electric vehicle; The EV charging and discharging power constraints are as follows: (15); EV charging and discharging ramp constraints: (16); The EV state of charge constraints are as follows: (17); In the formula, Respectively m The maximum charging and discharging power of the vehicle and t Working power at all times; Respectively m The landslide limit and climbing limit of electric vehicle batteries; Respectively m The vehicle's minimum and maximum state of charge and t State of charge at the moment.

3. The scheduling control method of a PV-storage direct-flexible system according to claim 2 is characterized in that: In step 3, it specifically includes: The economic dispatch objective function of the optimal dispatch model of the PV-storage direct-flexible system is as follows: (23); (24); Where: for t Total operating cost of the solar-storage direct-flexible system during the time period; Solar storage direct and flexible system Period operating costs; For the solar storage direct and flexible system t Power dispatch cost of time-period flexible equipment; Forsaken Light t Time period penalty cost; To purchase electricity t Period cost; They are t Time period n The dispatch cost coefficient and regulation power of each load; They are t Maintenance cost coefficient and regulation power of time-slot energy storage batteries; They are t The dispatch cost coefficient of electric vehicles in the time period and the m The charging and discharging power of an electric vehicle; for t The penalty cost of abandoning light during a certain period of time; is the penalty coefficient for abandoning light; for t Amount of abandoned light during the time period; for t The electricity purchase price during the period; For Architecture t The amount of electricity purchased during the time period; The constraints of the optimal scheduling model of the PV-storage-direct-flexible system include the flexible load in the PV-storage-direct-flexible system operation architecture model, the equipment adjustment power constraints, and the power balance equation of the PV-storage-direct-flexible system in the power consumption control strategy.

4. The scheduling control method of a PV-storage direct-flexible system according to claim 3 is characterized in that: In step 4, specifically including: Step 4.1, action space design: The action space represents the set of changes in the action variables at each time step; (25); is the action space, satisfying the constraints of the optimal scheduling model of the PV-storage direct-flexible system; Step 4.2, state space design: The state space is as follows: (26); is the state space, They are the electric load and photovoltaic prediction values ​​respectively; Step 4.3, reward function design: Reinforcement learning is to maximize the expected reward value by continuously learning and updating the agent's own strategy; therefore, the setting of the reward function should be related to the objective function in the scheduling model, and its reward function is as follows: (27); Where: They are t The bonus value at the time and the weight coefficient of the total operating cost of the PV-storage direct-flexible system; It is a parameter that regresses the reward function to a positive value; Step 4.4, reinforcement learning algorithm: The pulse experience replay TD3 algorithm is used to solve the reinforcement learning problem in the continuous action space; the strategy update formula of TD3 is: (28); Where: It is the TD3 algorithm strategy; is the state quantity; Yes Status The strategy of the next agent; is the action-value function; It is the experience replay buffer; is a state-action pair; Represents the sampling in the experience replay buffer D Find the expected value of the state-action pair; The historical data is input into the spiking neural network SNN, which encodes the historical data through LIF neurons to simulate the changes in the membrane potential of neurons after receiving stimulation. In this way, the spiking neural network can capture the time correlation of state changes in continuous control problems, so that historical data can be better utilized in the experience playback process. The formula for experience playback is: (29); Where: t The time of corresponds to the time of the scheduling model, It is the state of the next moment; It is the time constant, which is used to control the decay speed of the pulse signal. The larger the time constant, the longer the pulse influence time. is the initial moment; yes t Input pulse signal at the moment; After the experience replay mechanism is introduced into the spiking neural network, the experience replay is weighted according to the frequency and intensity of the pulse emission; the priority calculation formula is: (30); Where: Indicates The priority of a sample, which determines how often the sample is selected during playback; Indicates The TD error of a sample. The larger the TD error, the more important the sample is. is the control parameter of priority sampling, which controls the influence of priority. Represents a normalization factor, used to ensure that the sum of priorities is 1; Step 4.5, the algorithm flow of TD3 based on pulse experience playback: Step 4.51, Initialization: Initialize the Actor network and Critic network of the pulse neural network, set the strategy network and value function network required by the TD3 algorithm, and initialize the pulse experience playback buffer; Step 4.52, data preprocessing: pulse population encoding is performed on the state features in the offline data set, and normalized feature transformation is used to reduce data distribution differences; Step 4.53, pulse experience playback: sample a batch of experience from the pulse experience playback buffer, and weight the experience data according to the Q value error output by the pulse neural network; Step 4.54, update the value function: use the sampled pulse coding experience to update the value function based on the TD3 Critic network and calculate the Q value of the state-action pair; Step 4.55, update strategy: decode the pulse signal through the Actor network to generate actions, and use the loss function with pulse regularization term to update the policy network to minimize the behavior error; Step 4.56, update the target network: periodically update the target value function and the target policy network, and synchronize the pulse neural network parameters; Step 4.57, evaluate performance: evaluate the performance of the strategy at fixed intervals and test the control effect of the pulse experience playback TD3 algorithm through the simulation environment; Step 4.58, save the training model: save the trained pulse experience playback TD3 algorithm model for subsequent testing and application; Step 4.6, optimized scheduling process of TD3 based on pulse experience playback: Step 4.61, inputting photovoltaic historical data and load historical data into the reinforcement learning environment, and initializing various reinforcement learning network parameters; Step 4.62, conduct offline scheduling training, and obtain the current state quantity from the reinforcement learning environment: power load, photovoltaic forecast value, electricity price, power purchase amount, and action quantity at the previous moment; Step 4.63, the agent takes an action based on the input state; Step 4.64, perform actions in the reinforcement learning environment to obtain the state at the next moment; Step 4.65, when the set number of training times is reached, the optimization scheduling model of the PV-storage-direct-flexible system is saved; Step 4.66, input the real-time data into the saved training model, and the solar-storage direct-flexible system optimization scheduling model gives a real-time scheduling strategy.

Citation Information

Patent Citations

  • Optical storage and charging integrated station operation method oriented to power grid electric quantity balance and new energy consumption

    CN115940289A

  • Battery energy storage system control method based on improved virtual synchronous machine

    CN116014772A