New energy and load collaborative optimization decision-making method and system based on reinforcement learning
By constructing a reinforcement learning-based decision-making method for the coordinated optimization of new energy sources and loads, and by utilizing penalty factors and the probability distribution of photovoltaic output to optimize the health status of energy storage, the problem of coordinated optimization of new energy sources and loads is solved, thereby improving the efficiency and quality of power grid dispatch.
Patent Information
- Application Number
- CN202512026285.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies cannot effectively respond to the random fluctuations of new energy sources and are difficult to achieve efficient coordinated optimization between new energy sources and loads, resulting in a decline in grid dispatch performance.
By constructing a penalty factor reflecting the user-side load switching delay and a photovoltaic output probability distribution, and combining the energy storage health status, the state space, action space, and reward function of the reinforcement learning model are defined, and iterative training is performed to optimize decision-making.
It has improved the dispatch efficiency and quality of the power system, enhanced its robustness to fluctuations in new energy sources, and achieved efficient and coordinated optimization between new energy sources and loads.
Smart Images

Figure CN121863559A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power grid dispatching technology, and in particular to a new energy source and load collaborative optimization decision-making method and system based on reinforcement learning. Background Technology
[0002] With the increasing penetration of distributed photovoltaic and wind power and other new energy sources on the distribution network, and the continuous rise in the proportion of flexible residential loads (such as electric vehicles and smart home appliances), the traditional "one-way power reception" model is gradually shifting towards a multi-directional interaction of "source-grid-load". To reduce electricity costs and alleviate peak grid pressure, it is necessary to coordinate the dispatching of distributed new energy sources and diverse loads.
[0003] Existing technologies mainly employ rule-based sorting and model predictive control (MPC) methods for load dispatching. However, these dispatching methods have significant limitations. They cannot respond promptly to the random fluctuations of new energy sources and are difficult to adapt to uncertain environments, leading to errors in prediction results and consequently, a decline in dispatching performance or even failure.
[0004] Therefore, how to effectively implement power grid dispatch and achieve efficient synergistic optimization between new energy sources and loads has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] This invention provides a method and system for collaborative optimization decision-making between new energy sources and loads based on reinforcement learning, which solves the problem of how to make collaborative optimization decisions between new energy sources and loads to improve the dispatch efficiency and quality of the power system.
[0006] To address the aforementioned technical problems, embodiments of the present invention provide a reinforcement learning-based method for collaborative optimization decision-making between new energy sources and loads, comprising: A penalty factor reflecting the user-side load switching delay is constructed based on historical load scheduling data. The penalty factor and the user-side load operating status are combined to form a parameterized definition, resulting in a user-side load parameter set. The probability distribution of photovoltaic output is calculated based on historical photovoltaic irradiance data, and a correlation mapping between energy storage charge and discharge cycles and internal energy storage losses is established. The health status of energy storage is determined by the correlation mapping. The user-side load parameter set, the photovoltaic output probability distribution, and the energy storage health status definition are used to define a reinforcement learning model containing a state space, action space, and reward function, and then iteratively trained. The system acquires real-time operating data of the target distribution network, inputs the real-time operating data into the trained reinforcement learning model for online scheduling decisions, outputs target control commands, and executes them.
[0007] Furthermore, the penalty factor reflecting user-side load switching delay is constructed based on historical load scheduling data. This penalty factor, combined with the user-side load operating status, is then parameterized to obtain a user-side load parameter set, including: Extract the switching delay data, load energy efficiency loss data, operating time interval, duration period and rated power of each user-side load under the daily scheduling cycle from the historical load scheduling data. The switching delay data and the load energy efficiency loss data are quantitatively analyzed to establish a penalty factor that reflects the correlation between load energy consumption loss and switching delay. The user-side load parameter set is determined by parameterizing the operating time interval, the duration period, the rated power, and the penalty factor.
[0008] Furthermore, establishing a correlation mapping between energy storage charge-discharge cycles and internal energy storage losses, and using this correlation mapping to determine the energy storage health status, includes: Real-time acquisition of internal temperature data, depth of discharge, and internal chemical parameters of the target energy storage device during a single charge-discharge cycle; The internal temperature data, the depth of discharge, and the internal chemical parameters are input into a preset lifetime test model for calculation to obtain the equivalent lifetime loss of the target energy storage device under the current charge-discharge cycle. The energy storage health status of the target energy storage device is updated using the equivalent lifetime loss.
[0009] Furthermore, the reinforcement learning model that coordinates the user-side load parameter set, the photovoltaic output probability distribution, and the energy storage health state definition includes a state space, an action space, and a reward function, including: The state space is defined by the energy storage health status, the photovoltaic output probability distribution, and the user-side load parameter set; The action space is defined by control commands for user-side loads, photovoltaics, and energy storage within the target scheduling cycle; The reward function is designed with the goal of minimizing the total operating cost under the target scheduling cycle. The total operating cost includes load delay cost and energy storage lifetime loss cost.
[0010] Furthermore, the iterative training of the reinforcement learning model includes: The Q-learning algorithm is introduced to train the reinforcement learning model; During training, a greedy strategy is used to search for actions and drive the reward function to calculate immediate rewards in order to update the Q value; The training process continues iteratively until the convergence condition is met.
[0011] Furthermore, the step of inputting the real-time running data into the trained reinforcement learning model for online scheduling decisions includes: The reinforcement learning model is driven to query the internal policy table based on the input real-time running data; The optimal control command is issued to each load, photovoltaic and energy storage device in the scheduling window based on the internal strategy table; the optimal control command is used to control the start-up and shutdown and switching of load, photovoltaic output or energy storage charging and discharging power.
[0012] Furthermore, the calculation of the photovoltaic output probability distribution based on historical photovoltaic irradiance data includes: The historical solar irradiance data of the target distribution network is divided by season, and the irradiance data for the same hour within each season is determined based on the division results. Calculate the mean and standard deviation of irradiance data within the same hourly time period; The mean and the standard deviation are input into a preset probability distribution function to obtain the irradiance probability distribution; The irradiance probability distribution is converted in conjunction with the photovoltaic module parameters to obtain the photovoltaic output probability distribution.
[0013] Another embodiment of the present invention provides a new energy and load collaborative optimization decision system based on reinforcement learning, comprising: The parameter establishment module is used to construct a penalty factor reflecting the user-side load switching delay based on historical load scheduling data, and to define the user-side load parameter set by combining the penalty factor with the user-side load operating status. The photovoltaic-storage analysis module is used to calculate the probability distribution of photovoltaic output based on historical photovoltaic irradiance data, and to establish a correlation mapping between energy storage charge-discharge cycles and internal energy storage losses, thereby determining the energy storage health status using the correlation mapping. The model training module is used to coordinate the user-side load parameter set, the photovoltaic output probability distribution, and the energy storage health state to define a reinforcement learning model containing a state space, action space, and reward function, and to perform iterative training. The decision-making module is used to acquire real-time operating data of the target distribution network, input the real-time operating data into the trained reinforcement learning model for online scheduling decisions, output target control commands and execute them.
[0014] Another embodiment of the present invention provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the reinforcement learning-based new energy and load co-optimization decision-making method as described above.
[0015] In another embodiment of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, wherein when the device containing the computer-readable storage medium executes the computer program, it implements the reinforcement learning-based new energy and load collaborative optimization decision-making method as described above.
[0016] Compared with the prior art, the beneficial effects of the embodiments of the present invention are at least one of the following: This invention improves the accuracy of optimization objectives by establishing a user-side load parameter set that integrates penalty factors characterizing load switching delays, thereby quantifying and analyzing equipment energy efficiency losses caused by load delays. Secondly, it calculates the output probability distribution based on historical photovoltaic irradiance data and establishes an energy storage health status model. This probabilistic model accurately analyzes the randomness of photovoltaic power generation and quantifies energy storage cycle losses as health status indicators, enhancing the robustness of scheduling strategies against new energy fluctuations. Finally, it iteratively learns by designing a reinforcement learning model with a state-action space and reward function, and uses the trained model for online scheduling and decision-making, achieving efficient collaborative optimization of new energy and load. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the process of a reinforcement learning-based new energy and load collaborative optimization decision-making method in one embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a reinforcement learning-based new energy and load collaborative optimization decision system in one embodiment of the present invention; Figure 3 This is a structural block diagram of a preferred embodiment of a computer device provided by the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0020] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal communication between two components. The terms "vertical," "horizontal," "left," "right," "upper," "lower," and similar expressions used herein are for illustrative purposes only and do not indicate or imply that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0021] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is for the purpose of describing specific embodiments only and is not intended to limit the invention. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0022] One embodiment of the present invention provides a reinforcement learning-based method for collaborative optimization decision-making between new energy sources and loads. For details, please refer to [link to relevant documentation]. Figure 1 , Figure 1 The diagram shown is a flowchart of a reinforcement learning-based new energy and load collaborative optimization decision-making method according to one embodiment of the present invention, including the following steps: S1. Construct a penalty factor reflecting the user-side load switching delay based on historical load scheduling data. Combine the penalty factor with the user-side load operating status to define the parameter set of the user-side load.
[0023] In this embodiment, in order to fully consider the delay behavior of user-side load during the switching process in subsequent decision-making, the stability and economy of power system dispatching and decision-making are guaranteed, and the user experience is improved.
[0024] It should be understood that controllable loads include washing machines, dryers, dishwashers, heating and ventilation systems, air conditioners, water heaters, electric vehicles, and battery chargers for consumer electronics, which can be turned on at any time within a given time interval. Controllable loads can be further divided into uninterruptible loads and interruptible loads. Once an uninterruptible load is turned on, it will continue to operate for a specified duration. Once an interruptible load is turned on, it can be turned off multiple times within a specified time interval. Based on this, for controllable loads, the delay in their switching process, i.e., the switching process of the power supply (such as switching from the grid to photovoltaic), is further introduced as a penalty factor to participate in subsequent reinforcement learning decisions.
[0025] Specifically, this embodiment acquires historical load scheduling data through an information collection system deployed on the user side, such as through smart metering devices (e.g., smart meters) or user interaction interfaces. From the historical load scheduling data, it extracts switching delay data, load energy efficiency loss data, operating time intervals, durations, and rated power for each user-side load during the daily scheduling cycle. The switching delay data includes the number of load operations and the number of consecutive delays on that day. These indicators, along with penalty factors, are used to define parameters and establish a five-tuple to form the user-side load parameter set.
[0026] For example, suppose the user side includes m A load. Define a timeline where a day consists of 24 equal time intervals, each interval being... k The time period from 0 AM to 1 AM is represented as k =1, as the first time period, and k =24 indicates the last time period. Then each individual load... Represented by a parameterized quintuple: in,[ , ] indicates load The running time interval, This indicates the start time slot of the time interval during which the user-side load is allowed to begin operation; This indicates the end of the time interval during which the user-side load is allowed to start running; This indicates the total duration for which the load remains on. This is the rated power measured in kilowatts; if the first j A load in When activated, it will be in k = arrive k = + It remains on during the period of -1. For example, =(5,11,3,4) represents a load with a rated power of 4 kilowatts. It will be activated between the 5th and 11th time slots, and will run for three time slots.
[0027] It should be understood that load switching is directly related to the energy efficiency quality on the load side, which affects the user experience. Therefore, this embodiment performs quantitative analysis on switching delay data and load energy loss data to establish a penalty factor reflecting the correlation between load energy loss and switching delay, expressed by the following formula: in, To penalize the delay factor, it dynamically adjusts the switching delay data (number of consecutive delays, number of runs per day) to achieve a quantitative mapping between the 'switching delay level' and the 'energy efficiency loss increment', providing an objective basis for scheduling decisions; The penalty is the initial value of the delay, which is the baseline energy efficiency loss determined by the rated operating conditions of the equipment under the ideal operating state of the load without switching delay. This includes, but is not limited to, the energy consumption per unit time when the load is running at its rated speed, and the temperature fluctuation loss threshold when the equipment starts and stops without delay. For load j Number of consecutive postponements on the same day; For load j The actual number of times it has been run that day; α β represents the weighting of energy efficiency loss amplified by the number of consecutive delays, while β represents the weighting of loss offset by the number of runs already completed. This can be set... α =0.15, β =0.10.
[0028] Regarding the settings In fact, it characterizes the load energy efficiency loss caused by abnormal load switching, such as energy efficiency loss caused by temperature and power deviation, while The initial value of the penalty factor is used to characterize the baseline energy efficiency loss of the load under ideal operating conditions.
[0029] S2. Calculate the probability distribution of photovoltaic output based on historical photovoltaic irradiance data, and establish a correlation mapping between energy storage charge and discharge cycles and internal energy storage losses, so as to determine the health status of energy storage.
[0030] It should be understood that due to the uncertainty of solar irradiance, the output of renewable distributed generation units such as photovoltaics is random. When conducting simulation studies, it is necessary to model hourly solar irradiance while taking into account the associated uncertainties. Based on this, this embodiment constructs a stochastic model to account for these random behaviors. Specifically, historical solar irradiance data can be obtained through sensor data and data recording units deployed in the target distribution network.
[0031] The historical solar irradiance data of the target distribution network is divided by season. Based on the division results, the irradiance data for the same hour within each season is determined. Then, the mean and standard deviation of the irradiance data for the same hour are calculated. Finally, the mean and standard deviation are input into a preset probability distribution function to obtain the irradiance probability distribution.
[0032] For example, hourly solar irradiance data for a year are extracted. A year is divided into four seasons, and a day is further divided into 24-hour segments. Considering a month has 30 days (3 months x 30 days), there will be 90 irradiance data points per hour within a season. For each hour, the mean and standard deviation of solar irradiance are calculated based on these data and used as inputs to the probability distribution function. This embodiment preferably uses the beta probability density function (Beta PDF), and the specific calculation process is as follows: Where s is solar irradiance (unit: kW / m²). It is the beta distribution function of s; and These are the parameters of the beta distribution function.
[0033] and It is calculated based on the mean (μ) and standard deviation (σ) of the random variable s. The specific calculation process is as follows: Since the power generation capacity of photovoltaic (PV) modules depends on solar irradiance, ambient temperature, and module characteristics, a further conversion is performed between the irradiance probability distribution and PV module parameters to obtain the PV output probability distribution. The specific conversion process is as follows: Among them, the environmental and irradiance related parameters are: s is the solar irradiance, that is, the solar radiation power received per unit area; The ambient temperature refers to the temperature of the external air where the photovoltaic module is located, which is usually collected by on-site sensors. The operating temperature of a photovoltaic (PV) module, i.e., the actual chip temperature during PV power generation, needs to consider the balance between irradiation heating and ambient heat dissipation. Inherent electrical parameters of PV modules: This is the open-circuit voltage, which is the maximum voltage between the positive and negative terminals of a photovoltaic module when it is in a "no-load, open-circuit" state. This is the short-circuit current, which is the maximum output current of the photovoltaic module under the condition of "direct short circuit between the positive and negative terminals"; This is the voltage at the maximum power point; This is the maximum power point current. Photovoltaic module performance coefficient: FF The fill factor, which is the ratio of the actual maximum power of the module to the ideal maximum power, is a key indicator for measuring the power output efficiency of the module. is the voltage temperature coefficient, determined by the PN junction characteristics of the module material; N is the ideality factor, a parameter characterizing the "non-ideality" of the PN junction of the photovoltaic module, reflecting the influence of carrier recombination on the current-voltage characteristics, and determined by the module manufacturing process. Intermediate calculation parameters: This is the corrected short-circuit current; This is the corrected open-circuit voltage; This is the reverse saturation current.
[0034] Next, we analyze the energy storage. To avoid the drawbacks of overuse and a sharp drop in cycle life during subsequent decision-making, this embodiment calculates the equivalent lifetime consumption per cycle based on a cycle depth and temperature-corrected lifetime testing model. Specifically, firstly, we acquire in real-time the internal temperature data, the depth of discharge generated in a single charge-discharge cycle, and the internal chemical parameters of the target energy storage device. The internal chemical parameters include the activation energy and the gas constant.
[0035] In this embodiment, the Arrhenius model is preferred as the lifetime testing model. Internal temperature data, depth of discharge, and internal chemical parameters are input into the preset lifetime testing model for calculation. The calculation process is expressed by the following formula: in, For the first k Time period j The next cycle for energy storage i The resulting equivalent lifespan loss; For temperature; Temperature correction factor; DOD is depth of discharge. It is the activation energy; R It is the gas constant; The coefficient representing the influence of depth of discharge on lifespan loss is determined by the type of energy storage battery. ≈1.8-2.2.
[0036] Based on the above formula, the equivalent lifetime loss of the target energy storage device under the current charge-discharge cycle is obtained. Then, the energy storage health state (SOH) of the target energy storage device is updated using the equivalent lifetime loss, as expressed by the following formula: Where t refers to the time step. As can be seen from the above formula, the real-time health state (SOH) of energy storage actually represents the percentage of remaining lifetime of the energy storage. The lower the SOH value, the shorter the remaining lifetime of the energy storage, which will be used as a state variable input into the subsequent reinforcement learning model for decision-making.
[0037] S3 ~ S4 First, a reinforcement learning model containing state space, action space and reward function is defined, based on the user-side load parameter set, photovoltaic output probability distribution and energy storage health status definition, and then iteratively trained.
[0038] This embodiment preferably uses the Q-learning algorithm to drive reinforcement learning decision-making. First, the state space, action space, and reward function need to be defined. Specifically, the state space is defined by the energy storage health status, the photovoltaic output probability distribution, and the user-side load parameter set; the action space is defined by the control commands for user-side load, photovoltaic, and energy storage within the target scheduling cycle.
[0039] state space x Contains [ x (1), x (2)], where x (1) indicates the current time slot. x (2) Indicates the total on-time duration of the load. In each state, the action to be taken is to decide whether to turn the load on (from the grid / PV / storage) or turn it off.
[0040] Action space modified to A ={0, 1, 2}. For shutting down the load, the action is represented as: a =0. When using mains power to turn on the device, the action... a The value is 1. When the load is a distributed power source, the action is... a =2, this action will result in available photovoltaic output during that specific hour. P pv Reduce, photovoltaic power output updated to P pvnew The state space is updated to It can be expressed by the following formula: The reward function is designed with the objective of minimizing the total operating cost under the target scheduling cycle. The total operating cost includes load delay cost and energy storage lifetime loss cost. The reward function is specifically expressed as follows: The five reward functions correspond to the following scenarios: basic operating costs for actions a=1 and a=2; penalty costs for insufficient photovoltaic output; switching delay costs for shutting down loads; penalty costs for poor operating conditions; and comprehensive scheduling costs (including energy storage lifespan loss).
[0041] in, Indicates the first k In the first time slot i The cost of one unit of energy from one energy source. Depends on energy source i The properties of photovoltaic solar energy or the power grid; g The reward function does not include a penalty for energy storage lifetime; For a reward function that includes a penalty for energy storage lifetime; For energy storage i The equivalent replacement cost; ∈[0,1] represents the weights, and an adaptive strategy can be adopted during the training phase: when When <0.8, Automatically boosted by 50% to prevent deep aging. It's important to note that after load dispatch, the photovoltaic availability for this time slot will be updated. P pvnew If the available power is less than r j If so, then a penalty is assigned as the reward function.
[0042] The state transition function is then defined as: Here, 'd' represents a binary identifier variable indicating the real-time operating status of the load equipment. Its core function is to quantify the equipment status through action selection results. If the load equipment is in a closed state, then d=0, i.e., a=0; if the equipment is turned on through the grid or photovoltaic system, then d=1, i.e., a=1 or a=2. The value of 'd' is determined by its definition. , A fixed update logic has been established. During the training phase, the agent will update based on its real-time state. Select action 'a', and then complete the process using the defined initialization state transition function. x arrive The actual update.
[0043] This embodiment introduces the Q-learning algorithm to train the reinforcement learning model. During training, a greedy strategy is used to search for actions and drive the reward function to calculate immediate rewards to update the Q-value. The training process is then repeated iteratively until the convergence condition is met, resulting in a well-trained reinforcement learning model and generating the optimal internal policy Q-table. After training, real-time operating data of the target distribution network is acquired and input into the trained reinforcement learning model for online scheduling decisions, outputting and executing target control commands.
[0044] Specifically, reinforcement learning-based Q Learning comprises two phases: learning and execution. During the learning phase, different state-action pairs are updated.Q Value. In this case, the status includes information on load metrics, time period, and duration of operation.
[0045] At the same time, provide for each load =( s j , f j , l j , r j , udc j,k Data. For each load Q The value is initialized to zero, that is Q 0 ( x , a )=0, where the state x =[ x (1) x (2)], a ={0, 1, 2}. The photovoltaic power during the time period was simulated using a beta probability density function that considered the stochasticity of photovoltaic power. x (1)= s j and x (2) = 0, starting from the availability of the photovoltaic power, following ε - A greedy strategy takes an action / makes a decision; based on the action taken, the state transitions. Q The value and availability of photovoltaic power are also updated accordingly; after a sufficient number of iterations, the learning phase converges (this is achieved by...). Q (The amount of value updates is negligible), training is complete. Then, during the runtime phase, the converged values are queried. Q The value table yields the optimal action or power switching for each load at different times, i.e., the target control command.
[0046] Understandably, the optimal decision output by Q-learning is essentially "the action with the largest Q value in the current state", which aligns with the collaborative goal of "lowest load delay cost, highest photovoltaic utilization rate, and minimum energy storage lifespan loss".
[0047] The reinforcement learning model queries its internal policy table based on the input real-time operational data. This real-time operational data includes the operating status of each load during the current time period (e.g., operating time interval, operating duration), the current State of Harmony (SOH) value of the energy storage, and the real-time power output of the photovoltaic inverter. The internal policy table is used for each item within the scheduling window (i.e., the running time interval). , The system issues optimal control commands to the loads, photovoltaic and energy storage devices within the system. The optimal control commands are used to control the start-up and shutdown of the loads and switching, as well as the output of the photovoltaic system or the charging and discharging power of the energy storage system.
[0048] This embodiment provides the following example to analyze the above-described collaborative scheduling process in detail: To study the performance of the proposed scheduling algorithm, a user is considered, comprising dispatchable and undispatchable appliances as well as photovoltaic power generation. Typical dispatchable loads on the user side include washing machines, dryers, dishwashers, vacuum cleaners, irons, and rice cookers. These appliances have different operating intervals, operating durations, and rated power, while undispatchable loads (such as lighting) operate in fixed time slots. Furthermore, the penalty factor for delays in the operation of dispatchable loads is reflected by selecting a unit delay cost value. For example, dishwashers are assigned a higher unit delay cost value, indicating that if their operation is delayed, subsequent washing cycles will require higher heating power or longer operating times to achieve the same effect, resulting in significant additional energy consumption and lower load efficiency. In this example, a typical set of undispatchable loads (lighting) needs to operate continuously for 3 hours from time slots 18 to 20.
[0049] Based on the above, simulations were performed using the Binary Particle Swarm Optimization (BPSO) algorithm. The simulation results were compared with those obtained using reinforcement learning algorithms, and the comparison results are shown in Table 1.
[0050] Table 1 Performance Comparison with BPSO As can be seen, the scheduling cost obtained in both cases is the same, both being the minimum cost of 45.45. However, compared to the BPSO algorithm (16.51 seconds), the reinforcement learning algorithm has a shorter computation time (1.34 seconds). This is because the computation time of the BPSO algorithm depends on the size of the binary particles, resulting in a longer convergence time when dealing with photovoltaic randomness and complex load efficiency.
[0051] If the photovoltaic power generation on the user side is 5 kW, Table 2 shows the appliance dispatch plan including distributed generation. In this dispatch table, the on-state of loads drawing power from photovoltaic power is represented as 2, and the on-state of loads drawing power from the grid is represented as 1. It can be seen that, in this case, for lighting loads that are not dispatchable and need to operate during the time period {18–20}, power is drawn from the grid because photovoltaic power is unavailable during this time period. In addition, dishwasher loads are allocated a higher unit delay cost to avoid dispatch delays. Therefore, it is on during the time period {14–15}. With distributed generation included, the energy cost is reduced to 16.25.
[0052] Table 2 contains the appliance dispatch plan for distributed generation. In summary, this embodiment designs a penalty factor with user-side load switching delay as the research object and defines it in a coordinated and parameterized manner with various load scheduling and operation indicators to construct a parameter set that reflects the characteristics of user-side load. For new energy equipment, especially photovoltaic equipment, it deeply considers the factor of solar irradiance and quantitatively analyzes the output probability distribution of photovoltaics. It also determines the correlation model reflecting the lifespan and health status of energy storage by analyzing the relationship between the charging and discharging cycle of energy storage and internal losses. By combining the load set, the photovoltaic output probability distribution, and the correlation model, a reinforcement learning framework is defined. Using this reinforcement learning framework to make coordinated optimization decisions between new energy and load can meet the requirements of real-time dispatch of the power system.
[0053] One embodiment of the present invention provides a new energy and load collaborative optimization decision system based on reinforcement learning. For details, please refer to [link to relevant documentation]. Figure 2 , Figure 2 The diagram shown is a schematic representation of a reinforcement learning-based new energy and load collaborative optimization decision-making system according to one embodiment of the present invention, comprising: The parameter establishment module M1 is used to construct a penalty factor reflecting the user-side load switching delay based on historical load scheduling data, and to define the user-side load parameter set by combining the penalty factor and the user-side load operating status. The photovoltaic-storage analysis module M2 is used to calculate the probability distribution of photovoltaic output based on historical photovoltaic irradiance data, and to establish a correlation mapping between energy storage charge and discharge cycles and internal energy storage losses, and to determine the health status of energy storage based on the correlation mapping. The model training module M3 is used to coordinate the user-side load parameter set, the photovoltaic output probability distribution, and the energy storage health state definition to form a reinforcement learning model containing a state space, action space, and reward function, and to perform iterative training. The decision module M4 is used to acquire real-time operating data of the target distribution network, input the real-time operating data into the trained reinforcement learning model for online scheduling decision-making, output target control commands and execute them.
[0054] like Figure 3 As shown, this embodiment of the invention also provides a computer device. Figure 3 This is a structural block diagram of a preferred embodiment of a computer device provided by the present invention. The computer device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method described above.
[0055] Preferably, the computer program can be divided into one or more modules / units (such as computer program 1, computer program 2, ...), and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the computer device.
[0056] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can be any conventional processor. The processor is the control center of the terminal device, connecting various parts of the terminal device through various interfaces and lines.
[0057] The memory mainly includes a program storage area and a data storage area. The program storage area can store the operating system, applications required for at least one function, etc., while the data storage area can store related data, etc. Furthermore, the memory can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard drive, a SmartMedia Card (SMC), a Secure Digital (SD) card, and a Flash Card, or other volatile solid-state storage devices.
[0058] It should be noted that the aforementioned terminal devices may include, but are not limited to, processors and memory, as will be understood by those skilled in the art. Figure 3The structural block diagram is merely an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown, or combine certain components, or use different components. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0059] Accordingly, embodiments of the present invention provide a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the steps in the method of the above embodiments, for example... Figure 1 Steps S1 to S4 as described above.
[0060] The technical features and effects of the reinforcement learning-based new energy and load co-optimization decision-making system proposed in this invention are the same as those of the reinforcement learning-based new energy and load co-optimization decision-making method proposed in this invention, and will not be repeated here.
[0061] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A new energy and load collaborative optimization decision-making method based on reinforcement learning, characterized in that, include: A penalty factor reflecting the user-side load switching delay is constructed based on historical load scheduling data. The penalty factor and the user-side load operating status are combined to form a parameterized definition, resulting in a user-side load parameter set. The probability distribution of photovoltaic output is calculated based on historical photovoltaic irradiance data, and a correlation mapping between energy storage charge and discharge cycles and internal energy storage losses is established. The health status of energy storage is determined by the correlation mapping. The user-side load parameter set, the photovoltaic output probability distribution, and the energy storage health status definition are used to define a reinforcement learning model containing a state space, action space, and reward function, and then iteratively trained. The system acquires real-time operating data of the target distribution network, inputs the real-time operating data into the trained reinforcement learning model for online scheduling decisions, outputs target control commands, and executes them.
2. The reinforcement learning-based new energy and load collaborative optimization decision-making method as described in claim 1, characterized in that, The penalty factor reflecting user-side load switching delay is constructed based on historical load scheduling data. This penalty factor, combined with the user-side load operating status, is then parameterized to obtain a user-side load parameter set, including: Extract the switching delay data, load energy efficiency loss data, operating time interval, duration period and rated power of each user-side load under the daily scheduling cycle from the historical load scheduling data. The switching delay data and the load energy efficiency loss data are quantitatively analyzed to establish a penalty factor that reflects the correlation between load energy consumption loss and switching delay. The user-side load parameter set is determined by parameterizing the operating time interval, the duration period, the rated power, and the penalty factor.
3. The reinforcement learning-based new energy and load collaborative optimization decision-making method as described in claim 1, characterized in that, The step of establishing a correlation mapping between energy storage charge-discharge cycles and internal energy storage losses, and using the correlation mapping to determine the energy storage health status, includes: Real-time acquisition of internal temperature data, depth of discharge, and internal chemical parameters of the target energy storage device during a single charge-discharge cycle; The internal temperature data, the depth of discharge, and the internal chemical parameters are input into a preset lifetime test model for calculation to obtain the equivalent lifetime loss of the target energy storage device under the current charge-discharge cycle. The energy storage health status of the target energy storage device is updated using the equivalent lifetime loss.
4. The reinforcement learning-based new energy and load collaborative optimization decision-making method as described in claim 1, characterized in that, The reinforcement learning model that coordinates the user-side load parameter set, the photovoltaic output probability distribution, and the energy storage health state definition includes a state space, an action space, and a reward function, including: The state space is defined by the energy storage health status, the photovoltaic output probability distribution, and the user-side load parameter set; The action space is defined by control commands for user-side loads, photovoltaics, and energy storage within the target scheduling cycle; The reward function is designed with the goal of minimizing the total operating cost under the target scheduling cycle, where the total operating cost is the load delay cost and the energy storage lifetime loss cost.
5. The reinforcement learning-based new energy and load collaborative optimization decision-making method as described in claim 1, characterized in that, The iterative training of the reinforcement learning model includes: The Q-learning algorithm is introduced to train the reinforcement learning model; During training, a greedy strategy is used to search for actions and drive the reward function to calculate immediate rewards in order to update the Q value; The training process continues iteratively until the convergence condition is met.
6. The reinforcement learning-based new energy and load collaborative optimization decision-making method as described in claim 5, characterized in that, The step of inputting the real-time running data into the trained reinforcement learning model for online scheduling decisions includes: The reinforcement learning model is driven to query the internal policy table based on the input real-time running data; The optimal control command is issued to each load, photovoltaic and energy storage device in the scheduling window based on the internal strategy table; the optimal control command is used to control the start-up and shutdown and switching of load, photovoltaic output or energy storage charging and discharging power.
7. The reinforcement learning-based new energy and load collaborative optimization decision-making method as described in claim 1, characterized in that, The calculation of the photovoltaic power output probability distribution based on historical photovoltaic irradiance data includes: The historical solar irradiance data of the target distribution network is divided by season, and the irradiance data for the same hour within each season is determined based on the division results. Calculate the mean and standard deviation of irradiance data within the same hourly time period; The mean and the standard deviation are input into a preset probability distribution function to obtain the irradiance probability distribution; The irradiance probability distribution is converted in conjunction with the photovoltaic module parameters to obtain the photovoltaic output probability distribution.
8. A new energy and load collaborative optimization decision-making system based on reinforcement learning, characterized in that, include: The parameter establishment module is used to construct a penalty factor reflecting the user-side load switching delay based on historical load scheduling data, and to define the user-side load parameter set by combining the penalty factor with the user-side load operating status. The photovoltaic-storage analysis module is used to calculate the probability distribution of photovoltaic output based on historical photovoltaic irradiance data, and to establish a correlation mapping between energy storage charge-discharge cycles and internal energy storage losses, thereby determining the energy storage health status using the correlation mapping. The model training module is used to coordinate the user-side load parameter set, the photovoltaic output probability distribution, and the energy storage health state to define a reinforcement learning model containing a state space, action space, and reward function, and to perform iterative training. The decision-making module is used to acquire real-time operating data of the target distribution network, input the real-time operating data into the trained reinforcement learning model for online scheduling decisions, output target control commands and execute them.
9. A computer device, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the reinforcement learning-based new energy and load co-optimization decision-making method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the device containing the computer-readable storage medium executes the computer program, it implements the reinforcement learning-based new energy and load collaborative optimization decision-making method as described in any one of claims 1 to 7.