An energy management method for a photovoltaic energy storage system
Through real-time data collection, edge computing and deep reinforcement learning technologies, the charging and discharging strategies and distributed energy resource scheduling of photovoltaic energy storage systems are dynamically adjusted, and the problems of slow response speed, weak prediction capabilities and inflexible scheduling in the existing technology are solved, achieving more efficient and economical energy management.
Patent Information
- Application Number
- CN202411730223.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-11-29
AI Technical Summary
The existing photovoltaic energy storage systems have slow response speed, weak prediction capabilities and inflexible scheduling in complex supply and demand environments, making it difficult to effectively manage grid stability and energy efficiency.
By collecting the operating status and environmental requirements data of photovoltaic modules in real time, using edge computing devices for preliminary processing and lightweight machine learning model analysis to generate prediction data. Then, deep reinforcement learning algorithm is used to train the agent, dynamically adjust the charging and discharging strategies of the power electronic converter, and dispatch distributed energy resources in the microgrid through an adaptive supply and demand balance algorithm.
It significantly improves the energy management efficiency and economy of the photovoltaic energy storage system, improves the response speed and prediction capabilities, and achieves more flexible and intelligent power scheduling.
Smart Images

Figure CN119231653B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of energy management, and particularly to an energy management method for a photovoltaic energy storage system. Background Art
[0002] As an important part of the renewable energy field, photovoltaic energy storage systems have developed rapidly in recent years with the surging global demand for clean energy. Early photovoltaic energy storage systems mainly relied on simple charge-discharge control strategies. That is, when the photovoltaic power generation exceeded the immediate load, the excess electric energy was stored, and when the power generation was insufficient, the stored energy was released to fill the demand gap. However, this primary energy storage management method cannot effectively cope with the complex and changing energy supply-demand relationship. Especially in the scenario of large-scale grid connection, it poses higher requirements for grid stability and energy efficiency.
[0003] In recent years, with the maturity of advanced technologies such as the Internet of Things, edge computing, machine learning, and deep reinforcement learning, the energy management of photovoltaic energy storage systems has undergone a revolutionary change. The wide application of Internet of Things sensors makes it possible to collect the operating status and environmental parameters of photovoltaic modules in real time, providing basic data support for refined management. The introduction of edge computing not only improves the speed and efficiency of data processing but also enhances the system's response ability and intelligent decision-making level. Combining machine learning, especially deep reinforcement learning technology, photovoltaic energy storage systems can learn and optimize complex energy scheduling strategies, realizing dynamic adjustment of the charge-discharge strategies of power electronic converters, thereby greatly improving energy utilization efficiency and economic benefits. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides an energy management method for a photovoltaic energy storage system to solve the problems of slow response speed, weak prediction ability, and inflexible scheduling in the existing technology under complex supply-demand environments.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, an embodiment of the present invention provides an energy management method for a photovoltaic energy storage system, which includes: collecting operation status and environmental demand data of photovoltaic modules in real time, and transmitting the data to an edge computing device for preliminary processing; the edge computing device uses a lightweight machine learning model to analyze the operation status and environmental demand data in real time to generate prediction data; based on the prediction data of the edge computing device, a deep reinforcement learning algorithm is used to train an agent to learn an optimal energy management strategy and dynamically adjust the charge and discharge strategy of a power electronic converter; through an adaptive supply-demand balance algorithm, a microservices architecture is used to flexibly dispatch distributed energy resources in a microgrid to generate an optimized dispatch strategy; combining the charge and discharge strategy of the power electronic converter and the optimized dispatch strategy, continuously collecting system state snapshot data; performing real-time monitoring on the system state snapshot data, and if a significant change inconsistent with the prediction is detected, re-evaluating the supply-demand balance state and adjusting the working mode of the power electronic converter.
[0008] As a preferred solution of the energy management method for the photovoltaic energy storage system of the present invention, wherein: the step of collecting the operation status and environmental demand data of the photovoltaic modules in real time and transmitting the data to the edge computing device for preliminary processing is specifically as follows:
[0009] Deploy a comprehensive environmental and performance monitoring sensor network and a device integrating weather station functions to form a network covering the photovoltaic module area;
[0010] Use the performance monitoring sensor network to capture the operation status and environmental parameters of the photovoltaic modules in real time, convert them into JSON format, and create a data packet;
[0011] Apply the moving average method to the measured values in the data packet on the edge computing device to identify abnormal data points deviating from the normal range and mark them in the data packet;
[0012] Verify the marked data packet to complete the preliminary processing of the data.
[0013] As a preferred solution of the energy management method for the photovoltaic energy storage system of the present invention, wherein: the step of applying the moving average method to the measured values in the data packet on the edge computing device to identify abnormal data points deviating from the normal range and mark them in the data packet is specifically as follows:
[0014] The edge computing device receives the data packet from the sensor network, and at the same time unpacks the received JSON format data packet to extract metadata and payload data;
[0015] Parse the measured values of each metadata and payload data point, set the moving average window size N, and calculate the moving average value for each measured value. The expression is:
[0016] ;
[0017] Among them, MA t represents the moving average at time point t, and X i represents the i-th measurement value, where i is an index variable;
[0018] The expression for calculating the standard deviation SD based on the moving average at time point t is:
[0019] ;
[0020] Combining the moving average at time point t and the corresponding standard deviation SD, the expression for defining the standard deviation multiple threshold Td is:
[0021] ;
[0022] Among them, k is the selected standard deviation multiple;
[0023] Compare each measurement value with the threshold Td. If the measurement value exceeds the threshold Td range, it is marked as an outlier;
[0024] In the original data packet, mark the identified outliers.
[0025] As a preferred scheme of the energy management method of the photovoltaic energy storage system described in the present invention, wherein: the
[0026] Generate prediction data, and the specific steps are as follows:
[0027] Receive the data packet of the operating state of the photovoltaic module and the environmental requirements collected and preliminarily processed by the Internet of Things sensor, and perform data cleaning;
[0028] Apply multi-scale wavelet transform to the data packet after data cleaning for feature extraction, and output the operating state and environmental requirement data;
[0029] Based on the operating state and environmental requirement data extracted by the features, construct a real-time prediction expression through polynomial regression and exponential smoothing:
[0030] ;
[0031] Among them, P t represents the predicted value at time point t, β 0 is the intercept term, β i and γ i are the coefficients of the i-th feature and the polynomial order respectively, is the i-th eigenvalue, E t-1 is the error term at the previous moment, α is the smoothing factor, δ is the decay rate controlling exponential smoothing, n is the number of polynomial features, and e is the base of the natural logarithm.
[0032] As a preferred solution of the energy management method of the photovoltaic energy storage system described in the present invention, wherein: the charging and discharging strategy of the dynamic adjustment power electronic converter, the specific steps are as follows:
[0033] Add the predicted value P to the original state representation S t , and generate the extended state S' expression as: ;
[0034] Define the intelligent agent learning framework based on the extended state S', and establish the comprehensive evaluation function F expression as:
[0035] ;
[0036] Wherein, represents the energy efficiency at time point t, C(S', A) is the economic cost C of taking any action A under the extended state S', T is the time starting point, Δt is the time interval of integration, dt is the time differential element, and w is to adjust the predicted value P t The weight coefficient of the relative importance in the comprehensive evaluation function F;
[0037] Combine the comprehensive evaluation function F with the immediate reward R to generate the adjusted immediate reward R adj The expression is:
[0038] ;
[0039] Based on the immediate reward R adj , combine with the deep Q network to train the intelligent agent. In each learning iteration, the intelligent agent updates the Q value, and the update expression is:
[0040] ;
[0041] Wherein, represents the update amount of the Q value after executing the action a in the current state s, a' represents the possible action in the next state s', α is the learning rate, γ is the discount factor, θ represents the parameters of the main network, is the parameter of the target network.
[0042] As a preferred solution of the energy management method of the photovoltaic energy storage system described in the present invention, wherein: the generation of the optimal scheduling strategy, the specific steps are as follows:
[0043] Define a standardized interface for distributed energy resources and build it into a microservice with independent business logic;
[0044] Use Docker containers to encapsulate the microservices and manage and schedule the containerized microservices through the Kubernetes cluster;
[0045] Integrate the charging and discharging strategy of the converter with the dynamic balance algorithm of the market mechanism, and encourage distributed energy resources to actively participate in supply and demand regulation through real-time electricity price strategies. The optimized expression for supply and demand balance is generated as follows:
[0046] ;
[0047] where P sy,i is the total supply power at the i-th time point, P dd,i is the total demand power at the i-th time point, P DER is the power output vector of distributed energy resources, C DER,i is the operating cost of distributed energy resources at the i-th time point, and λ is the cost weight factor;
[0048] Data exchange between distributed energy resource microservices is carried out through RESTful APIs, and at the same time, message queues are used to transmit supply and demand signals;
[0049] In the distributed energy resource microservice, the Raft consensus algorithm is integrated for event listening;
[0050] Taking the output of the supply and demand balance optimization expression as the input, use mathematical programming to generate a scheduling strategy, and send the scheduling strategy to the distributed energy resource microservice to guide its adjustment of the operation mode;
[0051] Collect the actual responses of distributed energy resources, compare them with the expectations, and adjust the strategy until the best balance state is reached.
[0052] As a preferred solution of the energy management method of the photovoltaic energy storage system described in the present invention, among them: combining the charging and discharging strategy of the power electronic converter and the optimized scheduling strategy, continuously collecting system state snapshot data, the specific steps are as follows:
[0053] Define the deep reinforcement learning strategy as the decision rule;
[0054] Define the supply and demand balance algorithm as the adjustment rule for power demand and supply;
[0055] Integrate the real-time evaluation results of the decision rule and the adjustment rule to construct a decision framework;
[0056] According to the current system state and target priority, assign weights to the decision and adjustment rules through the decision framework;
[0057] Combine the weighted rule outputs to generate specific control instructions;
[0058] Send the control instructions to the corresponding power electronic equipment for execution, and output the real-time response status of the power electronic equipment;
[0059] Utilize the existing sensor network, additionally deploy sensors for monitoring the status of power electronic converters, and collect key power data;
[0060] Fuse the newly collected key power data with status data, environmental data, and prediction data to form comprehensive system status snapshot data.
[0061] As a preferred solution of the energy management method for the photovoltaic energy storage system described in the present invention, wherein: the
[0062] Monitor the system status snapshot data in real time. If a significant change inconsistent with the prediction is detected, re-evaluate the supply-demand balance status and adjust the operating mode of the power electronic converter. The specific steps are as follows:
[0063] Use a real-time data stream processing framework to perform real-time aggregation on the system status snapshot data collected by the sensor network to form a unified data stream;
[0064] Perform denoising, anomaly detection, and normalization processing on the aggregated data;
[0065] Apply a statistical process control chart to monitor the statistical characteristics in the data stream and identify points beyond the control limits;
[0066] Feed the detected points beyond the control limits back to the edge computing device to trigger a re-evaluation process;
[0067] Update the decision rule of the deep reinforcement learning agent, and re-train the agent with the new supply-demand balance status and system snapshot data;
[0068] Generate new control instructions according to the updated decision rule, and adjust the charge and discharge strategy of the power electronic converter;
[0069] Execute the control instructions and monitor the response status of the power electronic converter;
[0070] Implement a closed-loop control mechanism to continuously monitor the performance of the power electronic converter and the system status.
[0071] In a second aspect, an embodiment of the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and wherein: when the computer program is executed by the processor, it implements any step of the energy management method for the photovoltaic energy storage system as described in the first aspect of the present invention.
[0072] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and wherein: when the computer program is executed by the processor, it implements any step of the energy management method for the photovoltaic energy storage system as described in the first aspect of the present invention.
[0073] The beneficial effects of the present invention are as follows: Through the comprehensive application of Internet of Things sensors, edge computing, lightweight machine learning, deep reinforcement learning, and adaptive supply-demand balance algorithms, the present invention realizes more accurate data processing, more efficient energy prediction, more intelligent adjustment of the operating mode of power electronic converters, and more flexible scheduling of distributed energy resources. Specifically, the present invention collects the operation status and environmental demand data of photovoltaic modules in real time, uses edge computing devices to identify and process abnormal data points, combines lightweight machine learning models for high-precision prediction, adopts deep reinforcement learning algorithms to train agents to dynamically adjust the charge and discharge strategies of power electronic converters, and flexibly schedules distributed energy resources in the microgrid through adaptive supply-demand balance algorithms, thereby significantly improving the energy management efficiency and economy of the photovoltaic energy storage system and solving the problems existing in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0075] Figure 1 It is a flowchart of the energy management method for the photovoltaic energy storage system in Embodiment 1.
[0076] Figure 2 It is a flowchart of the preliminary processing by the edge computing device in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention in conjunction with the drawings of the specification.
[0078] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0079] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that mutually excludes other embodiments.
[0080] Embodiment 1, refer to Figure 1 and Figure 2, which is the first embodiment of the present invention. This embodiment provides an energy management method for a photovoltaic energy storage system, including the following steps:
[0081] S1. Collect the operating status and environmental requirement data of photovoltaic modules in real time, and transmit the data to the edge computing device for preliminary processing.
[0082] Furthermore, deploy a comprehensive environmental and performance monitoring sensor network, including sensors such as temperature, humidity, light intensity, current, and voltage, as well as devices integrating weather station functions. At least one node is deployed per 100 square meters to form a network covering the photovoltaic module area;
[0083] It should be noted that the devices integrating weather station functions mainly include, but are not limited to, the following sensors and components:
[0084] Temperature sensor: used to measure the temperature of air, soil, or water body.
[0085] Humidity sensor: monitors the relative humidity in the air.
[0086] Wind speed sensor: measures the wind speed, usually used in conjunction with a wind direction sensor.
[0087] Wind direction sensor: determines the wind direction.
[0088] Barometric pressure sensor: observes the atmospheric pressure, which helps predict weather changes.
[0089] Precipitation sensor: includes a rain gauge, used to detect rainfall or snowfall.
[0090] Illuminance sensor: measures the light intensity, which is particularly important for photovoltaic power generation.
[0091] Ultraviolet sensor: monitors the ultraviolet intensity, which has an impact on skin health and material aging.
[0092] Radiation sensor: measures solar radiation and other types of radiation.
[0093] Soil temperature and humidity sensor: used to monitor the temperature and humidity of the soil, which is valuable for agricultural and ecological research.
[0094] PM2.5 / PM10 sensor: monitors the particulate matter concentration in the air and evaluates the air quality.
[0095] Carbon dioxide (CO2) sensor: measures the concentration of carbon dioxide in the air, which has an impact on the greenhouse effect and plant growth.
[0096] In addition to sensors, the devices integrating weather station functions also include:
[0097] Data collector: Responsible for collecting raw data from various sensors and performing preliminary data processing.
[0098] Communication module: Used to transmit data to a central database or cloud platform. Common communication methods include wireless networks (such as LoRaWAN, GPRS, Wi-Fi), satellite communication, or wired networks.
[0099] Power management unit: Such as solar panels and storage batteries, which supply power to the weather station to ensure continuous operation.
[0100] Protection device: Such as protective covers and lightning rods, which protect the equipment from bad weather and external damage.
[0101] Data storage and processing software: Software built-in or in the cloud, used to store, analyze, and visualize meteorological data.
[0102] Furthermore, a performance monitoring sensor network is used to capture the operating status and environmental parameters of photovoltaic modules in real time, convert them into JSON format, and create data packets. Each data packet contains a timestamp, precise location coordinates, specific measurement values, and a checksum for verifying data integrity.
[0103] It should be noted that the specific implementation steps of "using a performance monitoring sensor network to capture the operating status and environmental parameters of photovoltaic modules in real time, convert them into JSON format, and create data packets" are as follows:
[0104] First, the sensors start working and continuously collect the operating status of photovoltaic modules (such as current, voltage, temperature) and environmental parameters (such as light, temperature, humidity, wind speed).
[0105] Second, the sensors convert physical signals (such as light, heat, electricity) into electrical signals, and then the analog-to-digital converter (ADC) converts them into digital signals for subsequent processing.
[0106] Finally, the integrity and reasonableness of the data are preliminarily checked, invalid or abnormal values are removed to ensure data quality, and the processed digital signals are encoded according to a predetermined data format. Common data formats include CSV, XML, JSON, etc. Among them, JSON is selected because of its lightweight and easy-to-parse characteristics; the encoded data is packed into data packets at regular time intervals or event-triggering mechanisms. Each data packet may contain data from multiple sensors and acquisition timestamps.
[0107] Apply the moving average method to the measurement values in the data packet on the edge computing device, identify abnormal data points deviating from the normal range, and mark them in the data packet.
[0108] Furthermore, the edge computing device receives data packets from the sensor network and unpacks the received data packets in JSON format to extract metadata and payload data;
[0109] Among them, metadata: usually includes information describing the context of the data packet, such as timestamps and location information. These information help to understand the temporal and spatial background of the data and are very important for data analysis and subsequent processing; Payload data: refers to the actually measured or collected data, such as measured values of temperature, humidity, pressure, etc. These data are the main outputs of the sensor network and are used to monitor the environmental state or the operation of the device.
[0110] Parse the measured values of each metadata and payload data point, such as current, voltage, temperature, etc., and set the moving average window size N, which depends on the volatility of the data and is generally 5 to 20 data points;
[0111] Calculate the moving average value for each measured value. The expression is:
[0112] ;
[0113] Among them, MA t represents the moving average value at time point t, X i represents the i-th measured value, and i is the index variable;
[0114] During the parsing process, as new data points are added, the moving average value is updated, the oldest data point is removed, and the latest data point is added;
[0115] Calculate the standard deviation SD based on the moving average value at time point t to measure the deviation degree of the data point from the average value. The expression is:
[0116] ;
[0117] Combine the moving average value at time point t and the corresponding standard deviation SD to define the standard deviation multiple threshold Td (for example, 2 or 3 times) to identify outliers. The expression is:
[0118] ;
[0119] Among them, k is the selected standard deviation multiple;
[0120] Compare each measured value with the threshold Td. If the measured value exceeds the threshold Td range, it is marked as an outlier;
[0121] In the original data packet, mark the identified outliers. For example, a boolean field isAnomaly or an outlier type label can be added.
[0122] Preferably, the "moving average method" is used as a tool for outlier detection, mainly due to several significant advantages in processing time series data. These advantages make the moving average method an ideal choice for real-time monitoring and data analysis, especially in the monitoring of the operating status of photovoltaic modules and environmental parameters. The moving average method can smooth short-term fluctuations and highlight long-term trends by calculating the average value of consecutive N data points in a data sequence, which is particularly effective for filtering out outliers caused by random noise. In a photovoltaic system, environmental factors such as light intensity and temperature may be affected by short-term disturbances, which do not reflect the true state but are the result of measurement errors or temporary interferences. By applying the moving average method, we can eliminate such short-term disturbances, make the data more stable, and thus more easily identify true outlier patterns, such as equipment failures or extreme weather events. In addition, the moving average method has high computational efficiency and is suitable for real-time data stream processing. It can be quickly executed on edge computing devices without waiting for a large amount of data to accumulate, which is crucial for real-time monitoring and rapid response. Therefore, the moving average method can not only help us extract useful information from complex data streams but also achieve efficient data preprocessing on resource-constrained edge computing devices, ensuring that subsequent in-depth analysis and decision-making are based on accurate and clean data.
[0123] Verify the marked data packet to ensure that the structure of the marked data packet is complete, all necessary information is correctly encoded, and the data stream is ready to be transmitted to the next processing stage.
[0124] It should be noted that "ensuring that the structure of the marked data packet is complete and all necessary information is correctly encoded" mainly means that
[0125] Data packet integrity check: Check whether the data packet contains all required fields, such as timestamps, location information, measurement values, and outlier marks; confirm that the data types and formats of all fields meet the expectations to avoid decoding failures caused by incorrect data formats.
[0126] Data packet structure verification: Verify whether the structure of the data packet is consistent with the expected format of the receiver, including the correctness of JSON key-value pairs; if the data packet contains nested structures, ensure that the data at all levels is correct.
[0127] Outlier mark confirmation: Confirm that the outlier data points have been correctly marked without omission or mislabeling; ensure that each outlier data point is accompanied by sufficient information, such as outlier type, outlier severity, etc.
[0128] Finally, if any issues are found during the integrity check, reorganize the data packet to ensure all information is accurate and check the encoding again to ensure all data is correctly encoded, especially special characters and numerical data.
[0129] S2. The edge computing device uses a lightweight machine learning model to analyze the operating status and environmental requirement data in real time and generate prediction data.
[0130] Furthermore, receiving the data packet of the operating status and environmental requirements of the photovoltaic module collected and preliminarily processed by the Internet of Things sensors and performing data cleaning, including removing noise and handling missing values, is a key step to ensure data quality.
[0131] Apply multi-scale wavelet transform to the data packet after data cleaning for feature extraction and output the operating status and environmental requirement data.
[0132] Preferably, applying multi-scale wavelet transform for feature extraction is an important part of the lightweight model. The purpose of feature extraction is to extract the key information most helpful for prediction from the original data, reduce the computational burden while maintaining the prediction performance.
[0133] It should be noted that applying multi-scale wavelet transform for feature extraction shows unique beneficial effects compared with other traditional methods, such as Fourier transform or windowed time-frequency analysis methods. Multi-scale wavelet transform can capture the local features of data in both time and frequency domains simultaneously and is especially good at dealing with non-stationary signals, which is crucial for analyzing the operating status and environmental requirement data of photovoltaic modules because these data often contain transient and trend components. By decomposing the signal into different time scales, wavelet transform can separate high-frequency noise or mutations and low-frequency long-term trends, thus more precisely extracting the features valuable for prediction. In addition, the multi-scale property allows us to observe the data at different resolutions, which is extremely beneficial for identifying patterns at different time scales in the photovoltaic system (such as daily light changes, seasonal fluctuations). Therefore, compared with other methods, multi-scale wavelet transform can provide richer and more specific feature information, enhance the accuracy and robustness of the prediction model, and thus improve the energy management efficiency of the entire photovoltaic energy storage system.
[0134] Furthermore, based on the operating status and environmental requirement data extracted by features, construct a real-time prediction expression through polynomial regression and exponential smoothing as follows:
[0135] ;
[0136] where P t represents the predicted value at time point t, β 0 is the intercept term, β i and γ iThey are the coefficient and polynomial order of the \(i\)-th feature respectively, is the \(i\)-th eigenvalue, \(E\) t-1 is the error term at the previous moment, \(\alpha\) is the smoothing factor, \(\delta\) is the decay rate controlling exponential smoothing, \(n\) is the number of polynomial features, and \(e\) is the base of the natural logarithm.
[0137] Preferably, the intercept term \(\beta\) 0 is estimated by minimizing the residual sum of squares (RSS), that is, by finding a set of parameters \(\beta\) (including \(\beta\) 0 ) such that the gap between the predicted value and the actual observed value is minimized. In actual operation, we use numerical optimization algorithms such as gradient descent or quasi-Newton method to solve the model parameters, including the intercept term. Finally, the model parameters we obtain provide optimal energy prediction ability, where \(\beta\) 0 represents the expected energy level when all independent variables are zero.
[0138] It should be noted that determining the order of the polynomial in the polynomial regression model is a crucial step, which directly affects the fitting ability and generalization ability of the model. Usually, choosing the appropriate polynomial order requires balancing the complexity of the model and the fitting effect of the data to avoid overfitting or underfitting. The following are several common methods to determine the polynomial order:
[0139] First, conduct a visual inspection: Use a scatter plot to observe the data distribution and try to identify whether there is an obvious curve trend in the data. If the data shows a non-linear relationship, polynomial regression can be considered. By plotting the scatter plots of polynomial fitting curves of different orders and the original data, visually judge which order of the model can best reflect the true trend of the data.
[0140] Secondly, conduct residual analysis: Check the residual plot, that is, the change of the difference between the predicted value and the actual value with the predicted value. Ideally, the residuals should be randomly distributed and the mean is close to 0. If the residuals show a certain pattern (such as a curve), it may indicate that the model is too simple and the order of the polynomial needs to be increased; if the variance of the residuals increases with the increase of the predicted value, other types of models or data transformation may need to be considered. Analyze the residual sum of squares (RSS) at different orders and select the polynomial order that minimizes the RSS.
[0141] Then, conduct cross-validation: Use cross-validation (such as \(k\)-fold cross-validation) to evaluate the performance of models of different orders on unseen data. Usually select the polynomial order with the lowest average cross-validation error. Cross-validation can help evaluate the generalization ability of the model and avoid selecting a model that performs well on the training data but poorly on new data.
[0142] Furthermore, perform the information criterion: Use the Akaike Information Criterion (AIC) or the Bayesian Information Criterion (BIC) to evaluate the model. These criteria consider the complexity and goodness-of-fit of the model and tend to select a model that can fit the data well without being overly complex. Both AIC and BIC penalize the complexity of the model, with BIC having a greater penalty and being more suitable for model selection in large sample cases.
[0143] Finally, use stepwise regression to increase the order of the polynomial. It should also be noted that stepwise regression is a statistical method that finds the optimal model by automatically adding or deleting variables. In the context of polynomial regression, the order of the polynomial can be gradually increased until the improvement of the model is no longer significant.
[0144] Preferably, combining polynomial regression with exponential smoothing technology and applying it to the energy management of a photovoltaic energy storage system can give full play to the advantages of both. Polynomial regression can capture the non-linear trends in the data and accurately model complex time series patterns, while exponential smoothing is good at handling short-term fluctuations and providing more stable prediction results. This combination not only improves the prediction accuracy but also enhances the adaptability of the model to real-time data changes, enabling the energy management system to more intelligently predict power demand and supply, optimize the scheduling strategy, and thus significantly improve the overall performance and economic benefits of the system.
[0145] It should be noted that this process describes the application process of a lightweight machine learning model in predicting the operating state of photovoltaic modules and environmental requirements, including data preprocessing, feature extraction, and the construction and application of the prediction model. The entire process aims to achieve lightweight, efficient, and real-time prediction goals, which is in line with the design concept of the lightweight machine learning model.
[0146] It should also be noted that edge computing devices using lightweight machine learning models to analyze the operating state and environmental requirement data of photovoltaic systems in real time can quickly generate accurate prediction data, effectively improve energy management efficiency, achieve instant response and optimized scheduling, and significantly enhance the intelligent level and economic benefits of photovoltaic energy storage systems.
[0147] S3. Based on the prediction data of the edge computing device, use a deep reinforcement learning algorithm to train an agent to learn the optimal energy management strategy and dynamically adjust the charging and discharging strategies of the power electronic converter.
[0148] Furthermore, add the predicted value P to the original state representation S t , and generate the extended state S′ expression as: ;
[0149] Preferably, "obtaining the extended state representation S′" means that the agent not only considers the current operating state and environmental parameters but also the future predicted values when making decisions, enhancing the agent's predictability.
[0150] Define the agent learning framework based on the extended state S′, and establish the comprehensive evaluation function F with the expression:
[0151] ;
[0152] where, represents the energy efficiency at time point t, C(S′, A) is the economic cost C of taking any action A in the extended state S′, T is the starting time, Δt is the time interval of integration, dt is the time differential element, and w is the weight coefficient that adjusts the relative importance of the predicted value P t in the comprehensive evaluation function F;
[0153] It should be noted that the "economic cost C" specifically includes the following specific cost types:
[0154] Operating cost: This may include the operating energy consumption cost of power electronic converters and other equipment, as well as any maintenance and depreciation expenses.
[0155] Opportunity cost: If the system does not store or release energy at a certain moment, it may miss market opportunities, such as selling electricity during high electricity price periods or charging during low electricity price periods.
[0156] Loss cost: Energy loss occurs during the charge and discharge processes of energy storage devices, and the energy cost of this part of the loss also needs to be included.
[0157] Equipment life cost: Frequent charge and discharge cycles may shorten the service life of the battery. Therefore, frequent charge and discharge strategies may lead to higher equipment replacement costs.
[0158] Market transaction cost: If the system participates in the electricity market transaction, then the transaction costs of buying and selling electricity should also be considered, which may include transaction fees, imbalance fees, etc.
[0159] Environmental cost: In some regions, taxes may be levied on carbon emissions or pollutant emissions. If the energy storage system causes an increase in the indirect or direct carbon footprint, then this part of the cost should also be included.
[0160] Reserve capacity cost: The system needs to maintain a certain reserve capacity to cope with emergencies, and this unused capacity is also a kind of cost.
[0161] It should be noted that in the context of deep reinforcement learning, an "agent" refers to an entity that can perceive the environment and make decisions based on the perceived information. The goal of the agent is usually to perform a series of actions through interaction with the environment to achieve a certain long-term goal or maximize a certain reward signal. In specific application scenarios, the agent can be a software program, a robot, a virtual character, etc. They can observe the state of the environment, make decisions based on these states, and execute corresponding actions;
[0162] It should be noted that through the above steps, the predicted value P t is incorporated into the decision-making process of the agent, which not only adjusts the learning framework and evaluation function of the agent, but also enables the agent to dynamically adjust the charging and discharging strategies of the power electronic converter based on future prediction data, thereby achieving more efficient and more forward-looking energy management. This strategy not only improves the flexibility and efficiency of energy management, but also enhances the overall stability of the system.
[0163] Combining the comprehensive evaluation function F and the immediate reward R to generate the adjusted immediate reward R adj The expression is:
[0164] ;
[0165] Preferably, in this way, the following ΔQ(s, a) update formula directly reflects the influence of F(S′, A). When learning, the agent not only considers the immediate reward R, but also considers the long-term benefits characterized by F(S′, A), so that the decision-making is more inclined to those actions that can bring higher comprehensive benefits.
[0166] Based on the immediate reward R adj , the agent is trained by combining with the deep Q-network. In each learning iteration, the agent updates the Q value, and the update expression is:
[0167] ;
[0168] Among them, represents the update amount of the Q value after executing the action a in the current state s. a′ represents the possible action in the next state s′. α is the learning rate, which determines the change range of the weight in each update. γ is the discount factor, and θ represents the parameters of the main network. is the parameter of the target network, which is used to calculate the target Q value to help stabilize the training process;
[0169] Preferably, DQN is a reinforcement learning algorithm based on deep learning. It can learn complex strategies from high-dimensional inputs. At the same time, in order to improve the learning efficiency and stability, a double-network structure and a prioritized experience replay mechanism will be adopted.
[0170] It should also be noted that during the learning process of the Deep Q-Network (DQN), the agent plays the following roles: perceiving the environment: the agent can receive information about the current state S of the environment, which may include visual images, sensor readings, system status, etc.; decision-making: based on the current state S, the agent selects an action a from the set of available actions A to execute. This decision is made based on the agent's internal policy, usually guided by a policy function or a value function (such as the Q-function); executing the action: the agent executes the selected action a, which causes the environment to change to a new state S′ and may generate an immediate reward R; learning and adaptation: the agent updates its internal policy or value function according to the immediate reward R and the new state S′. In DQN, this usually means updating the Q-value so that the agent's decision can gradually approach the optimal policy, that is, the action taken in a given state can maximize the expected cumulative reward.
[0171] It should be noted that the specific meanings of the "main network and target network" are as follows: The main network is a network used to predict the Q-value of taking a certain action in the current state. During the learning process, the agent selects actions and predicts Q-values based on the network with parameter θ. Whenever the agent interacts with the environment and receives new experience samples, the parameter θ is updated according to these samples to make the network's prediction more accurate; The role of the target network is to provide a stable target Q-value for updating the parameter θ of the main network. The parameter θ- of the target network is not updated frequently, but is copied from the parameter θ of the main network periodically. This is done to increase the stability of the learning process and avoid unstable phenomena caused by rapid changes in parameters during the training process.
[0172] It should also be noted that after being trained, the agent will be able to dynamically adjust the charge and discharge strategy of the power electronic converter according to the real-time prediction data of the edge computing device. In this step, the agent will evaluate the current state in real time, predict future energy demands and market conditions to optimize the timing of energy storage and release. The agent's decision will be directly converted into control signals for the power electronic converter to achieve automated and intelligent energy management.
[0173] After the Deep Q-Network (DQN) training is completed, the agent will be able to dynamically adjust the charge and discharge strategy of the power electronic converter based on the learned policy.
[0174] It should be noted that the following are the specific steps for how the agent uses the learned knowledge to optimize energy management after training:
[0175] The agent first observes the current environmental state S, which includes but is not limited to information such as the output power of photovoltaic modules, the state of charge of the energy storage system, the market electricity price, and the weather forecast. This information will be converted into a feature vector that the agent can understand and process, serving as the representation of the current state S.
[0176] Based on the current state S, the agent will use its trained Q-network (represented by the parameter θ) to evaluate all possible actions A. For each action a ∈ A, the agent calculates the corresponding Q-value Q(s, a; θ), which reflects the cumulative reward expected to be obtained by taking this action in the current state.
[0177] The agent selects the action a∗ with the highest Q-value, and the expression is:
[0178] ;
[0179] Once the agent selects the optimal action a*, it will directly or indirectly be converted into a control instruction for the power electronic converter. For example, if a* indicates charging, the agent will send a signal to the power electronic converter, commanding it to absorb power from the grid or photovoltaic panels and store it in the battery. On the contrary, if a* indicates discharging, the agent will command the power electronic converter to release the stored power to meet the load demand.
[0180] After executing the action a*, the agent will observe the feedback from the environment, including the new state S′ and the immediate reward R. The new state S′ will serve as the basis for the next decision, and the immediate reward R can be used to evaluate the effect of the action. Even after training is completed, the agent can continue to collect data for the fine-tuning and optimization of future strategies.
[0181] The above process will be carried out in a loop. The agent continuously observes, evaluates, selects actions, executes, and receives feedback to dynamically adjust the charging and discharging strategies of the power electronic converter to achieve optimal energy management and resource allocation.
[0182] Finally, in this way, the agent can flexibly adjust the charging and discharging strategies according to real-time environmental conditions and prediction information to maximize energy utilization efficiency and economic benefits while ensuring the stable operation of the system.
[0183] S4. Through the adaptive supply-demand balance algorithm, flexibly dispatch distributed energy resources within the microgrid using the microservices architecture to generate an optimized dispatch strategy.
[0184] Furthermore, define a standardized interface for distributed energy resources (DERs);
[0185] It should be noted that this standardized interface should cover functions such as identity authentication of energy resources, status reporting, acceptance and execution of control instructions, etc., to facilitate seamless integration across systems and platforms.
[0186] Each distributed energy resource is defined as a microservice with independent business logic, such as energy generation, storage, consumption, or trading;
[0187] Use Docker containers to encapsulate microservices to ensure environmental consistency, and manage and schedule containerized microservices through a Kubernetes cluster to ensure high availability and elastic scaling;
[0188] Integrate the charging and discharging strategy of the converter (such as the power electronic converter in a battery energy storage system) with the dynamic balancing algorithm of the market mechanism, and use real-time electricity price strategies to incentivize distributed energy resources (DERs) to actively participate in supply and demand regulation, thereby optimizing the supply and demand balance problem. The goal is to minimize the supply and demand difference in the power system while considering economic costs, and the supply and demand balance optimization expression is generated as:
[0189] ;
[0190] Among them, P sy,i is the total supply power at the i-th time point, P dd,i is the total demand power at the i-th time point, P DER is the power output vector of distributed energy resources DERs, C DER,i is the operating cost of distributed energy resources DERs at the i-th time point, and λ is the cost weight factor, which balances power mismatch and operating cost.
[0191] Preferably, the Real-Time Pricing (RTP) strategy is not a new concept. It has been studied and applied in the power industry for some time, especially in the fields of smart grid and Demand Side Management (DSM). The core of the real-time electricity price strategy lies in dynamically adjusting the electricity price according to the real-time changes in power supply and demand, so as to reflect the true cost and value of electricity. This strategy aims to encourage electricity consumers to reduce electricity consumption during peak hours and increase electricity consumption during off-peak hours, in order to balance power supply and demand and improve the efficiency and reliability of the power grid.
[0192] It should be noted that the "supply and demand balance optimization expression" is closely related to the previous step of "using the prediction data of edge computing devices, training an agent using a deep reinforcement learning algorithm, learning the optimal energy management strategy, and dynamically adjusting the charging and discharging strategy of the power electronic converter". The relevance between the two concepts is reflected in the following points:
[0193] The goal of the "Supply-Demand Balance Optimization Expression" is to minimize the supply-demand imbalance and operating costs by adjusting the output of Distributed Energy Resources (DERs), which is consistent with the goal of the deep reinforcement learning agent - learning the optimal policy to optimize energy management; the deep reinforcement learning agent learns through trial and error, continuously adjusting its policy to optimize its actions in a specific state (i.e., the charge and discharge strategy of the power electronic converter), which directly acts on P in the above expression. DER , thus affecting the supply-demand balance and operating costs; the learning process of the deep reinforcement learning agent depends on real-time and predictive data, which may include weather forecasts, power demand forecasts, market prices, etc., all of which are collected and preprocessed by edge computing devices. The quality and accuracy of this data directly affect the optimization effect of the agent's policy; the deep reinforcement learning agent can dynamically adjust its policy according to environmental changes (such as real-time power demand, supply conditions), which is consistent with adjusting P in the expression. DER to cope with real-time supply-demand changes.
[0194] In summary, the learning and decision-making process of the deep reinforcement learning agent directly serves the objective function in the optimization expression, that is, to minimize the supply-demand imbalance and operating costs. Through the dynamic policy adjustment of the agent, the optimal scheduling of DERs can be effectively achieved, thereby realizing the optimization of smart grid energy management in practical applications.
[0195] Furthermore, the Distributed Energy Resources DERs microservices interact through RESTful APIs to ensure compatibility between heterogeneous systems, and use message queues (such as RabbitMQ or Kafka) to transmit supply-demand signals to ensure reliable message transmission and sequential processing;
[0196] It should be noted that this step ensures the coordination within the DERs system, and its output is a stable data stream between microservices and a real-time transmission mechanism for supply-demand signals.
[0197] In the Distributed Energy DERs Resources microservices, the Raft consensus algorithm is integrated to listen for events, such as price changes, supply-demand imbalances, etc., and automatically trigger a response mechanism to ensure the consistency of all nodes' responses to events;
[0198] Preferably, in such a distributed energy management system, the "Raft consensus algorithm" is particularly suitable for the event listening and response mechanism under the DERs (Distributed Energy Resources) microservice architecture mainly due to its significant advantages in understandability, efficiency of the election process, and fault tolerance. Different from other complex and theoretically abstract consensus algorithms such as Paxos, Raft simplifies the core process and decomposes the state machine replication problem into three parts: leader election, log replication, and security, making the algorithm not only easy to understand and implement but also perform well in a dynamically changing network environment. Especially in the DERs scenario, the states of energy resources change frequently, and price signals under the market mechanism need to be quickly responded to. The Raft algorithm can quickly select a leader to ensure timely and consistent responses to events such as price changes or supply-demand imbalances in the distributed system, while ensuring that the system can still maintain consistency and high availability even in the case of network partitions or node failures. This is an extremely critical feature for a power system with high real-time requirements and high reliability. Therefore, with its intuitive design, efficient election mechanism, and strong fault tolerance, the Raft consensus algorithm becomes an ideal choice for maintaining decision consistency and system stability among DERs microservices.
[0199] Furthermore, taking the output of the supply-demand balance optimization expression as the input, using mathematical programming to generate a scheduling strategy, evaluating the performance of the strategy under different scenarios to ensure the robustness and effectiveness of the strategy, and sending the scheduling strategy to the DERs (Distributed Energy Resources) microservice of the distributed energy resources to guide its adjustment of the operation mode. The output of this step is a real-time updated scheduling strategy and continuously improved system performance, ultimately realizing the efficient and economic operation of the DERs (Distributed Energy Resources) system;
[0200] Collect the actual responses of the distributed energy resources DERs, compare them with the expectations, and adjust the strategy until the best balance state is reached;
[0201] It should be noted that "collecting the actual responses of the distributed energy resources DERs, comparing them with the expectations, and adjusting the strategy until the best balance state is reached" is a key part of optimizing the scheduling strategy, which is used to ensure that the behavior of DERs (Distributed Energy Resources) conforms to the established goals, such as supply-demand balance, cost minimization, etc. The following is a set of specific steps to implement this process:
[0202] S4.1. First, it is necessary to monitor the operation status and performance indicators of DERs in real time. This includes but is not limited to the charging state, discharging state, actual power output, environmental conditions (such as light intensity, temperature), and any external factors that may affect the behavior of DERs. These data should be automatically collected through sensors and intelligent metering devices and sent to the central management system through the network.
[0203] S4.2. Clean, preprocess, and analyze the collected raw data to convert it into effective information that can be used for comparison and decision-making. For example, compare the actual power output of DERs with the expected output in the scheduling strategy to evaluate whether the DERs are executing as planned.
[0204] S4.3. This may include differences in power output, lags in response speed, or failure to achieve the expected cost efficiency, etc.
[0205] S4.4. Further analyze the reasons for the deviation. Possible reasons include the limitations of DERs themselves, changes in the external environment, inaccuracies in the prediction model, or deficiencies in the scheduling strategy, etc. This step helps to adjust the strategy targeted rather than blindly correcting the deviation.
[0206] S4.5. Adjust the scheduling strategy according to the results of deviation analysis and cause diagnosis. This may involve modifying the working mode of DERs, updating the prediction model, adjusting the cost weight factor λ, or optimizing the coordination mechanism between DERs, etc.
[0207] S4.6. Redistribute the adjusted strategy to the DERs and monitor their new responses. This stage also requires collecting data to verify whether the adjusted strategy is effective, that is, whether it is closer to or reaches the expected goal.
[0208] Optimization is a continuous process that requires continuous monitoring, analysis, adjustment, and verification. Therefore, steps S4.1 to S4.6 form a closed loop, and each cycle aims to gradually improve the scheduling strategy until the best balance state is achieved.
[0209] S5. Combine the charging and discharging strategy of the power electronic converter and the optimized scheduling strategy, and continuously collect system state snapshot data.
[0210] Furthermore, define the deep reinforcement learning strategy as a decision rule to specify the charging and discharging behavior of the power electronic converter under different system states;
[0211] Define the supply-demand balance algorithm as the adjustment rule for power demand and supply;
[0212] Integrate the real-time evaluation results of the decision rule and the adjustment rule to construct a decision framework to ensure that the framework has the ability to handle immediate demands and long-term goals;
[0213] Preferably, the decision framework needs to be able to comprehensively consider long-term interests (deep reinforcement learning strategy) and immediate demands (supply-demand balance algorithm).
[0214] According to the current system state and goal priorities, allocate weights to the decision and adjustment rules through the decision framework; for example, emphasize supply-demand balance during peak hours and emphasize the deep learning strategy during off-peak hours;
[0215] It should be noted that this weight allocation can be dynamically adjusted according to pre-set rules or determined by another optimization algorithm (such as fuzzy logic or genetic algorithm);
[0216] Combined with the weighted rule output, generate specific control instructions, such as adjusting the PWM frequency of the power electronic converter, controlling the battery charge and discharge current, etc.;
[0217] It should be noted that the control instructions should take into account the physical limitations and safety ranges of power electronic devices;
[0218] Send the control instructions to the corresponding power electronic devices for execution and output the real-time response status of the power electronic devices;
[0219] Utilize the existing sensor network, additionally deploy sensors for monitoring the status of the power electronic converter, and collect key power data;
[0220] Preferably, the key power data includes key parameters such as voltage, current, power conversion efficiency, etc.;
[0221] Fuse the newly collected key power data with the status data, environmental data, and prediction data to form comprehensive system status snapshot data.
[0222] It should be noted that by combining the charge and discharge strategy of the power electronic converter with the optimal scheduling strategy and continuously collecting system status snapshot data, the efficiency and stability of the distributed energy resources (DERs) system can be significantly improved. Its beneficial effects are reflected in the following aspects: First, by real-time monitoring and analyzing the operating status of DERs, the system can quickly identify potential imbalances or inefficiencies, and then timely adjust the charge and discharge strategy of the power electronic converter to ensure the accuracy of DERs in responding to market signals and system demands. Second, the dynamic adjustment of the optimal scheduling strategy, based on the collected snapshot data, can more accurately predict and respond to supply and demand changes, promote supply and demand balance, reduce operating costs, and at the same time improve the flexibility and resilience of the entire power system. Finally, continuous data collection provides rich empirical evidence for strategy optimization. Through advanced technologies such as machine learning, the system can self-learn and evolve, continuously improve the intelligence level of decision-making, and ultimately achieve the efficient allocation and sustainable operation of DERs resources. This closed-loop feedback mechanism is the key to realizing the vision of the smart grid, and it will help the power industry move towards a greener, smarter, and more efficient future.
[0223] S6. Real-time monitor the system status snapshot data. If a significant change inconsistent with the prediction is detected, re-evaluate the supply and demand balance status and adjust the working mode of the power electronic converter.
[0224] Furthermore, use a real-time data stream processing framework (such as Apache Kafka Streams or Apache Flink) to perform real-time aggregation on the system state snapshot data collected by the sensor network to form a unified data stream;
[0225] Preferably, the real-time data stream processing framework can quickly process a large amount of continuously arriving data to ensure the system's response speed to the latest data, which is crucial for real-time monitoring.
[0226] Perform denoising, anomaly detection, and normalization processing on the aggregated data to eliminate invalid or incorrect data and ensure data quality;
[0227] Apply statistical process control (SPC) charts, such as the mean-range chart (X-bar and R chart), to monitor the statistical characteristics in the data stream and identify points beyond the control limits; this usually indicates a significant change in the system state;
[0228] It should be noted that applying statistical process control (SPC) charts, especially the mean-range chart (X-bar and R chart), is a commonly used quality control and process monitoring method for monitoring the statistical characteristics in the data stream to identify points beyond the control limits. The following are the specific implementation steps:
[0229] Regularly extract sample groups from the data stream. Each sample group usually contains the same number of observations (for example, 5 or 10 data points). These observations should be collected under the same conditions to reflect the immediate state of the process, and at the same time record all the observations of each sample group;
[0230] Calculate the average of the observations in each sample group, find the maximum and minimum values in each sample group, and then calculate the difference between them;
[0231] Calculate the average of the means and the average of the ranges for all sample groups;
[0232] Calculate the upper control limit (UCL) and the lower control limit (LCL). For the X-bar chart, the control limit expression is:
[0233] UCL = population mean + A2 * average range;
[0234] LCL = population mean - A2 * average range;
[0235] Among them, A2 is a constant that depends on the size of the sample group and can be found in the SPC chart constant table;
[0236] For the R chart, the control limit calculation expression is:
[0237] UCL_R = D4 * average range;
[0238] LCL_R = D3 * Average range;
[0239] Among them, D4 and D3 are also constants found from the SPC chart constant table;
[0240] Mark the numbers of the sample groups on the horizontal axis, plot the mean value of each sample group on the vertical axis, and draw the center line (overall mean) and the upper and lower control limits;
[0241] On another graph, plot the range of each sample group, and also draw the center line (average range) and the upper control limit;
[0242] Check whether the points on the X-bar chart and R chart fall within the control limits. If the points fall outside the control limits or show a non-random pattern (such as trends, periodicity, systematic offsets, etc.), it indicates that the process may be out of control or there is special cause variation;
[0243] Conduct in-depth analysis on the points outside the control limits or showing non-random patterns, determine the reasons, and take corrective measures;
[0244] Based on the analysis results, make necessary adjustments to the process to eliminate special cause variation and restore the stability of the process;
[0245] Even if the process is adjusted to a controlled state, continuously monitor the SPC chart to ensure the stability of the process and continuous improvement.
[0246] Preferably, through the SPC chart, abnormal changes in the system state can be detected early, which helps to take measures in advance to avoid potential failures. At the same time, SPC helps to maintain the stability of the process and ensure that the production or operation process is in a controlled state.
[0247] Furthermore, feedback the detected points outside the control limits to the edge computing device to trigger a re-evaluation process;
[0248] It should be noted that in the energy management of the photovoltaic energy storage system, triggering the re-evaluation process is usually closely related to the dynamic characteristics of the system and changes in the external environment. When the system detects points outside the control limits, this usually means that the system state has changed significantly from the expected or predicted situation, which may affect the stability and efficiency of the system. The following are several specific situations that may trigger the re-evaluation process:
[0249] Changes in environmental conditions: Sudden changes in weather, such as sharp changes in light intensity and temperature, will affect the power generation efficiency of photovoltaic modules and the working state of the energy storage system, thus requiring a re-evaluation of the energy management and scheduling strategies.
[0250] Load demand fluctuations: If there are unexpected large fluctuations in the load demand of the power grid, such as a sudden increase or decrease in industrial activities, or a sudden change in the electricity consumption habits in residential areas, this may require the system to immediately adjust the charging and discharging strategies and scheduling plans.
[0251] Market price fluctuations: Changes in real-time electricity prices, especially when there are large fluctuations in market prices, require the system to respond quickly to optimize costs and revenues.
[0252] Equipment failures or performance degradation: If the performance of power electronic converters or other key equipment is monitored to be below normal levels or a failure occurs, the system needs to be re-evaluated to avoid further losses or safety risks.
[0253] Energy storage state changes: If the SOC (State of Charge) of the battery reaches the warning limit, whether it is too high or too low, the charging and discharging strategies need to be re-evaluated to protect the battery and optimize its use.
[0254] Market trading opportunities: When new trading opportunities appear in the market, such as changes in the prices of renewable energy certificates (RECs) or carbon credits, the system may need to re-evaluate its strategies for participating in the market.
[0255] System security considerations: If the system monitors any abnormalities that may affect security, such as overheating, overvoltage, or short-circuit risks, it must be immediately re-evaluated and necessary preventive measures taken.
[0256] When any of the above situations is detected, the system will automatically trigger a re-evaluation process, which may include the following steps:
[0257] Data collection and analysis: Collect the latest system status data, including environmental data, equipment status, market prices, etc., and analyze this data to determine the current system status.
[0258] Model update: Update the prediction model based on the latest data, such as adjusting the parameters of machine learning models, to more accurately predict future demand and supply.
[0259] Strategy optimization: Based on the updated model, recalculate the optimal charging and discharging strategies and scheduling plans to adapt to the current system conditions and market environment.
[0260] Control instruction generation and execution: Generate new control instructions, adjust the operating mode of power electronic converters, update the scheduling strategies of distributed energy resources, and monitor the impact of these changes on system performance.
[0261] Performance evaluation and feedback: Evaluate the implementation effect of the new strategy, compare it with the expected goals, and if necessary, further adjust the strategy until the system returns to the optimal operating state.
[0262] Preferably, the edge computing device is close to the data source, can quickly process data and respond, reduce latency, and at the same time, once an anomaly is detected, immediately initiate a re-evaluation process to ensure that the system can adjust the strategy in a timely manner.
[0263] Update the decision rules of the deep reinforcement learning agent, and retrain the agent with the new supply-demand balance state and system snapshot data to adapt to the current system conditions;
[0264] Generate new control instructions according to the updated decision rules, and adjust the charge and discharge strategies of the power electronic converter;
[0265] Execute the control instructions, monitor the response status of the power electronic converter to ensure that the instructions are executed and the system state is stable;
[0266] Implement a closed-loop control mechanism to continuously monitor the performance of the power electronic converter and the system state; ensure that any subsequent significant changes can be captured and processed in a timely manner;
[0267] It should be noted that the real-time monitoring of the system state snapshot data, combined with real-time data stream processing frameworks such as Apache Kafka Streams or Apache Flink, realizes instant insight and agile response to the operating state of the power electronic converter. This mechanism improves the data quality through efficient aggregation, denoising, and standardization processing, ensuring the reliability of decision-making. Applying SPC charts to monitor statistical characteristics can detect system state anomalies at an early stage and trigger the rapid re-evaluation process of the edge computing device, thus accelerating strategy adjustment. By updating the deep reinforcement learning model, the system can flexibly change the working mode of the power electronic converter according to the latest supply-demand balance state, enhancing the adaptability and robustness of the system. The implementation of the closed-loop control mechanism ensures continuous monitoring and timely correction, maintaining the stable operation of the system in a complex environment and significantly improving the intelligent level and economic benefits of power management.
[0268] This embodiment also provides a computer device applicable to the energy management method of the photovoltaic energy storage system, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the energy management method of the photovoltaic energy storage system as proposed in the above embodiment.
[0269] The computer device may be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0270] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the energy management method for a photovoltaic energy storage system as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0271] In summary, the present invention activates the Internet of Things sensors, obtains and preprocesses the operating status and environmental data of photovoltaic modules in real time, ensures the data quality, and lays a solid foundation for subsequent analysis; uses lightweight machine learning and deep reinforcement learning technologies to accurately predict the behavior of the photovoltaic system and dynamically adjust the working mode of the power electronic converter, greatly improving the energy utilization efficiency and system response speed; uses the combination of an adaptive supply-demand balance algorithm and a microservices architecture to achieve flexible scheduling of distributed energy resources, enhancing the economy and stability of the system. In summary, the present invention effectively solves the problems of slow response speed, weak prediction ability, and inflexible scheduling in the prior art under complex supply-demand environments, and provides a more advanced and intelligent management solution for the photovoltaic energy storage system.
[0272] Example 2. Referring to Table 1, this is the second example of the present invention. To further verify the advancement of the present invention, experimental simulation data of the energy management method for the photovoltaic energy storage system is given.
[0273] In a large photovoltaic energy storage power station located in North China of China, the energy management method of the present invention was implemented. The power station has approximately 50 MW of photovoltaic modules, is equipped with an advanced battery energy storage system and a series of Internet of Things sensors. The test was carried out in summer for 25 days to fully test the performance of the system under high light and high load conditions.
[0274] First, a high-performance sensor network including temperature sensors, humidity sensors, radiometers, and anemometers, as well as devices integrating weather station functions, was deployed to form a network covering the photovoltaic module area; a performance monitoring sensor network was used to capture the operating status and environmental parameters of the photovoltaic modules in real time, convert them into JSON format, and create data packets; the moving average method was applied to the measured values in the data packets on the edge computing device to identify abnormal data points deviating from the normal range and mark them in the data packets; the marked data packets were verified to complete the preliminary processing of the data.
[0275] Second, the edge computing device used a lightweight machine learning model to analyze the operating status and environmental demand data in real time to generate prediction data; multi-scale wavelet transform was applied for feature extraction, combined with polynomial regression and exponential smoothing to construct a real-time prediction model.
[0276] Then, based on the prediction data of the edge computing device, a deep reinforcement learning algorithm was used to train the agent to learn the optimal energy management strategy and dynamically adjust the charge and discharge strategy of the power electronic converter; through continuous learning, the agent adjusted the charge and discharge strategy in real time according to the system status and prediction results to achieve efficient energy utilization and cost savings.
[0277] Furthermore, a microservices architecture was used to flexibly schedule distributed energy resources in the microgrid to generate an optimized scheduling strategy; data exchange between distributed energy resource microservices was carried out through RESTful APIs, and at the same time, message queues were used to transmit supply and demand signals to ensure fast response and efficient collaboration.
[0278] Finally, in combination with the charge and discharge strategy and the optimal scheduling strategy of the power electronic converter, continuously collect the system state snapshot data; utilize the existing sensor network and additionally deploy sensors for monitoring the state of the power electronic converter to collect key power data, forming comprehensive system state snapshot data; conduct real-time monitoring on the system state snapshot data. If significant changes inconsistent with the prediction are detected, re-evaluate the supply-demand balance state and adjust the working mode of the power electronic converter; execute a closed-loop control mechanism to continuously monitor the performance of the power electronic converter and the system state.
[0279] As shown in Table 1 below:
[0280] Table 1 Experimental Record Table
[0281] Date Light intensity (W / m²) Ambient temperature (°C) Predicted electricity consumption (kWh) Actual electricity consumption (kWh) Electricity deviation (%) System efficiency (%) Cost savings ($) 2024 / 8 / 5 1000 30 4500 4480 0.44 82.5 320 2024 / 8 / 10 1100 32 5000 4975 0.5 83 350 2024 / 8 / 15 950 29 4200 4180 0.48 82 310 2024 / 8 / 20 1050 31 4700 4675 0.53 82.8 330 2024 / 8 / 25 980 30 4300 4270 0.65 82.3 325
[0282] During the 25-day experiment, the energy management method of the photovoltaic energy storage system demonstrated significant advantages, especially in terms of power prediction accuracy, system efficiency, and cost savings. During the experiment, the system efficiency reached an average of 82.5%, at least 3 percentage points higher than the traditional method. This is mainly due to the intelligent agent of the deep reinforcement learning technology learning the optimal energy management strategy and the precise scheduling of the adaptive supply-demand balance algorithm; through the optimized scheduling strategy, the cost saved per month reaches 325 - 325 - 350, compared with the traditional method that saves an average of 200 - 200 - 250 per month, and the cost-saving ability of the present invention has increased by at least 30%.
[0283] In summary, these data clearly prove the innovation and practicality of the energy management method of the photovoltaic energy storage system. It not only improves the prediction accuracy and system efficiency but also significantly reduces the operating cost, playing an important role in promoting the sustainable development of the photovoltaic energy storage industry.
[0284] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. An energy management method for a photovoltaic energy storage system, characterized in that: include, Collect the operating status and environmental demand data of photovoltaic modules in real time, and transmit the data to edge computing devices for preliminary processing; Edge computing devices use lightweight machine learning models to analyze operating status and environmental demand data in real time and generate predictive data; Based on the prediction data of edge computing devices, the deep reinforcement learning algorithm is used to train the intelligent agent, learn the optimal energy management strategy, and dynamically adjust the charging and discharging strategy of the power electronic converter; Through the adaptive supply and demand balancing algorithm, the distributed energy resources in the microgrid are flexibly dispatched using the microservice architecture to generate an optimized dispatching strategy; Combine the charging and discharging strategy of the power electronic converter with the optimization scheduling strategy to continuously collect snapshot data of the system status; Monitor the system status snapshot data in real time. If changes that are inconsistent with the forecast are detected, re-evaluate the supply and demand balance and adjust the power electronic converter working mode. The specific steps of generating prediction data are as follows: Receive the data packets of photovoltaic module operation status and environmental requirements collected and preliminarily processed by IoT sensors, and perform data cleaning; For the cleaned data packets, multi-scale wavelet transform is applied to extract features and output the operation status and environmental demand data; Based on the operating status and environmental demand data extracted from the features, the real-time prediction expression is constructed through polynomial regression and exponential smoothing: ; Among them, P t represents the predicted value at time point t, β0 is the intercept term, β i and γ i are the coefficient and polynomial order of the i-th feature, respectively. is the i-th eigenvalue, E t-1 is the error term at the previous moment, α is the smoothing factor, δ is the decay rate that controls exponential smoothing, n is the number of polynomial features, and e is the base of the natural logarithm.
2. The energy management method of the photovoltaic energy storage system according to claim 1, characterized in that: The real-time collection of the operating status and environmental demand data of the photovoltaic components and the transmission of the data to the edge computing device for preliminary processing are as follows: Deploy a comprehensive network of environmental and performance monitoring sensors, as well as equipment with integrated weather station functions, to form a network covering the PV panel area; Use the performance monitoring sensor network to capture the operating status and environmental parameters of the PV panels in real time, convert them into JSON format, and create data packets; Apply a moving average method to the measurements in the data packet on the edge computing device to identify abnormal data points that deviate from the normal range and mark them in the data packet; Verify the marked data packets and complete the preliminary processing of the data.
3. The energy management method of the photovoltaic energy storage system according to claim 2, characterized in that: The moving average method is applied to the measured values in the data packet on the edge computing device to identify abnormal data points that deviate from the normal range and mark them in the data packet. The specific steps are as follows: The edge computing device receives data packets from the sensor network and unpacks the received JSON format data packets to extract metadata and payload data; Parse the measurement value of each metadata and payload data point, set the moving average window size N, and calculate the moving average of each measurement value. The expression is: ; Among them, MA t represents the moving average at time point t, X i represents the i-th measurement value, i is the index variable; The standard deviation SD expression based on the moving average at time point t is: ; Combining the moving average value at time point t and the corresponding standard deviation SD, the standard deviation multiple threshold Td expression is defined as: ; Where k is the selected standard deviation multiple; Compare each measured value with the threshold Td. If the measured value exceeds the threshold Td, it is marked as an outlier. In the original data packet, the identified outliers are marked.
4. The energy management method of the photovoltaic energy storage system according to claim 3, characterized in that: The specific steps of dynamically adjusting the charging and discharging strategy of the power electronic converter are as follows: Add the predicted value P to the original state representation S t , the expression of the extended state S′ is generated as: ; Based on the extended state S′, the agent learning framework is defined and the comprehensive evaluation function F is established as follows: ; in, represents the energy efficiency at time point t, C(S′, A) is the economic cost C of taking any action A under the extended state S′, T is the starting time, Δt is the time interval of integration, dt is the time differential element, and w is the adjusted prediction value P t The weight coefficient of relative importance in the comprehensive evaluation function F; Combine the comprehensive evaluation function F with the immediate reward R to generate the adjusted immediate reward R adj The expression is: ; Based on the immediate reward R adj , combined with the deep Q network to train the agent, in each learning iteration, the agent updates the Q value, and the update expression is: ; in, It represents the update amount of Q value after executing action a in the current state s, a′ represents the possible action to be taken in the next state s′, α is the learning rate, γ is the discount factor, and θ represents the parameters of the main network. are the parameters of the target network.
5. The energy management method of the photovoltaic energy storage system according to claim 4, characterized in that: The specific steps of generating the optimized scheduling strategy are as follows: Define standardized interfaces for distributed energy resources and build them into microservices with independent business logic; Use Docker containers to encapsulate microservices, and manage and schedule containerized microservices through Kubernetes clusters; The charging and discharging strategy of the converter is integrated with the dynamic balance algorithm of the market mechanism, and the distributed energy resources are encouraged to actively participate in the supply and demand regulation through the real-time electricity price strategy, and the supply and demand balance optimization expression is generated as follows: ; Among them, P sy,i is the total power supplied at the i-th time point, P dd,i is the total power demand at time i, P DER is the power output vector of distributed energy resources, C DER,i is the operating cost of distributed energy resources at the i-th time point, λ is the cost weight factor; Distributed energy resource microservices exchange data through RESTful APIs and use message queues to transmit supply and demand signals; In the distributed energy resource microservice, the Raft consensus algorithm is integrated for event monitoring; The output of the supply-demand balance optimization expression is used as input, and a scheduling strategy is generated using mathematical programming. The scheduling strategy is then sent to the distributed energy resource microservice to guide it to adjust its operation mode. Collect the actual response of distributed energy resources, compare it with the expected response, and adjust the strategy until the optimal balance is achieved.
6. The energy management method of the photovoltaic energy storage system according to claim 5, characterized in that: The charging and discharging strategy and the optimization scheduling strategy of the power electronic converter are combined to continuously collect system status snapshot data. The specific steps are as follows: Define deep reinforcement learning policies as decision rules; The supply-demand balancing algorithm is defined as the adjustment rule for electricity demand and supply; Integrate the real-time evaluation results of decision rules and adjustment rules to build a decision framework; Assign weights to decisions and adjustment rules through a decision framework based on the current system state and target priorities; Combine the weighted rule outputs to generate specific control instructions; Send control instructions to corresponding power electronic equipment for execution, and output the real-time response status of the power electronic equipment; Leverage existing sensor networks to deploy additional sensors to monitor the status of power electronic converters and collect key power data; Combine newly collected critical power data with status data, environmental data, and forecast data to form a comprehensive system status snapshot.
7. The energy management method of the photovoltaic energy storage system according to claim 6, characterized in that: Said Monitor the system status snapshot data in real time. If significant changes that are inconsistent with the forecast are detected, re-evaluate the supply and demand balance and adjust the power electronic converter working mode. The specific steps are as follows: The real-time data stream processing framework is used to aggregate the system status snapshot data collected by the sensor network in real time to form a unified data stream; Perform denoising, anomaly detection and standardization on the aggregated data; Apply statistical process control charts to monitor statistical characteristics in data streams and identify points outside control limits; Feedback detected points beyond control limits to edge computing devices to trigger a reassessment process; Update the decision rules of the deep reinforcement learning agent and retrain the agent based on the new supply and demand balance state and system snapshot data; Generate new control instructions based on the updated decision rules to adjust the charging and discharging strategy of the power electronic converter; Execute control instructions and monitor the response status of power electronic converters; Implement a closed-loop control mechanism to continuously monitor the performance of the power electronic converter and the system status.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the energy management method for the photovoltaic energy storage system described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the energy management method for a photovoltaic energy storage system according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
One-pile multi-connection electric vehicle ordered charging method based on deep reinforcement learning
CN116001624A
Wind turbine generator vibration monitoring and fault diagnosis method
CN118327909A