Photovoltaic cleaning decision-making method and system based on time sequence feature fusion and marginal benefit optimization
Through multi-source data fusion and deep reinforcement learning model, combined with edge-cloud collaborative computing architecture, dynamic optimization of photovoltaic power station cleaning decisions is achieved, solving the problems of inefficient and high error rates in the existing technology, and improving power generation efficiency and benefits.
Patent Information
- Application Number
- CN202510484991.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
The existing cleaning decision-making methods for photovoltaic power plants have problems such as low efficiency in fixed-cycle mode, high error rate of simple prediction model, and insufficient data fusion capability, resulting in loss of power generation efficiency and waste of resources.
Multi-source heterogeneous data acquisition and deep reinforcement learning model are combined with edge-cloud collaborative computing architecture, and through time-series feature fusion and marginal benefit optimization, economic constraint reward functions are designed to achieve dynamic cleaning strategy optimization.
The cleaning frequency is significantly reduced by 35%-40%, the average annual net income is increased by 18%-22%, and the misjudgment rate is less than 10% in extreme weather. The edge-cloud collaborative computing architecture supports real-time scheduling of GigaW power stations to reduce operation and maintenance costs.
Smart Images

Figure CN120338549A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of smart energy, and particularly relates to a photovoltaic panel cleaning decision-making system integrating time series data analysis and reinforcement learning, and is particularly applicable to the operation and maintenance management of large-scale photovoltaic power stations in sandstorm-prone areas. Background Technique
[0002] In the operation and maintenance management of photovoltaic power stations, the impact of panel surface fouling on power generation efficiency has become a key concern in the industry. According to the statistics of the International Energy Agency, the average annual power generation loss of uncleaned photovoltaic panels can reach 15%-25%, and even exceed 35% in sandstorm-prone areas. At present, the mainstream cleaning solutions in the industry mainly rely on fixed-cycle cleaning modes and simple prediction models, but these methods have significant defects. The fixed-cycle cleaning mode ignores the dynamic impact of weather changes on the fouling rate, which may lead to cleaning delays and power generation losses in arid and high-dust environments, while in areas with frequent rainy seasons, it may cause resource waste due to natural rainfall washing. In addition, the simple prediction models in the existing technologies are usually based on linear regression or time series analysis, and are difficult to handle the complex non-linear relationship between meteorological parameters and fouling, and their prediction error rates are as high as 30%-40%. At the same time, these models often only judge whether cleaning is needed without considering the cost differences and marginal benefits of different cleaning methods, resulting in a lack of economic constraints in decision-making.
[0003] There are also deficiencies in the data fusion capabilities of current patented technologies. Meteorological data and operation and maintenance data are usually processed in isolation, and a cross-dimensional correlation model cannot be effectively constructed. For example, the separation of short-term meteorological fluctuations and long-term fouling trends makes it impossible to incorporate key information such as the high probability of future rainfall into real-time decision-making, and there is also a lack of robust processing methods for sensor data missing or outliers. Through the statistical analysis of photovoltaic cleaning-related patents in the past five years, it is found that more than 80% of the patents focus on the improvement of the mechanical structure of cleaning equipment, and only 12% involve cleaning strategy optimization and mostly rely on rule engines rather than data-driven models. In addition, the existing methods are insufficient in multi-objective collaborative optimization. They often single-mindedly pursue the maximization of power generation efficiency or the minimization of cleaning costs, and do not establish a quantitative balance mechanism between the two, resulting in negative-profit operations where the cleaning cost is higher than the power generation benefit in some scenarios. At the same time, high-precision prediction models often cannot meet the real-time decision-making requirements due to calculation delays, and the edge computing solutions are limited by model simplification, and the prediction error rate exceeds 20%, presenting a contradiction between real-time performance and prediction accuracy.
[0004] For the above industry problems, there is an urgent need for a photovoltaic cleaning decision-making method that can integrate multi-source heterogeneous data, capture non-linear relationships, and achieve dynamic cost-benefit balance. This method needs to combine short-term high-precision weather forecasts, long-term historical fouling attenuation laws, and equipment status data through a deep reinforcement learning framework to construct a spatio-temporal joint feature matrix to improve prediction accuracy. At the same time, a reward function including economic constraints needs to be designed to achieve the optimal choice of the marginal benefit of cleaning actions, and an edge-cloud collaborative computing architecture needs to be considered to balance the response speed and algorithm iteration requirements. Such technological innovation will provide important support for the intelligent operation and maintenance of photovoltaic power plants, significantly improving power generation efficiency and reducing cleaning costs. Summary of the Invention
[0005] Aiming at the defects of the existing photovoltaic cleaning decision-making methods, such as low efficiency in the fixed-cycle mode, high error rate in simple prediction models, and insufficient data fusion ability, the present invention adopts the following technical solutions:
[0006] The present invention provides a photovoltaic cleaning decision-making method and system based on temporal feature fusion and marginal benefit optimization. The method realizes the optimization of dynamic cleaning strategies through the collection and preprocessing of multi-source heterogeneous data, the construction and training of a deep reinforcement learning model, and real-time decision-making output. Further, the method combines an edge-cloud collaborative computing architecture, deploys a lightweight inference model at the edge to meet real-time requirements, and completes incremental training of the model through the cloud to improve the algorithm iteration efficiency.
[0007] The multi-source heterogeneous data includes meteorological data, historical power generation efficiency data, and equipment status data. Among them, the meteorological data obtains the solar radiation intensity, temperature, humidity, PM2.5 concentration, and rainfall probability in the target area for the next 72 hours through the meteorological bureau API, with a sampling interval of 15 minutes; the historical data covers the power generation efficiency records of photovoltaic zones, the cleaning operation time, and the cost-benefit data of cleaning methods, covering at least 3 years; the equipment status data is collected by a sensor network, including the inclination angle of photovoltaic panels, the surface fouling detection value, and the operating status of robot cleaning equipment.
[0008] Furthermore, the data preprocessing uses the sliding window method to generate the time series feature matrix of meteorological data. The window size is 24 hours and the step size is 1 hour to capture the short-term meteorological change trend. For missing or abnormal data, a data augmentation model based on the Generative Adversarial Network (GAN) is used for completion. Specifically, the generator uses an LSTM network to simulate the time series correlation of meteorological data, and the discriminator introduces an attention mechanism to detect abnormal phenomena such as sudden changes in PM2.5 concentration. The generator and the discriminator are continuously optimized through adversarial learning, and finally, the completed data that meets the physical constraint conditions is generated. For example, the humidity value is restricted between 0 and 100, and the PM2.5 concentration is non-negative. All input data is normalized and mapped to the [0,1] interval to eliminate the influence of dimensional differences on model training.
[0009] The present invention constructs a deep reinforcement learning model, and its state space is defined as S = [current power generation efficiency, future 24-hour meteorological prediction vector, cumulative number of uncleaned days, panel inclination angle]. Among them, the current power generation efficiency is normalized to the [0,1] interval, the future 24-hour meteorological prediction is encoded as a 24-dimensional time series vector, the cumulative number of uncleaned days reflects the cumulative effect of fouling, and the panel inclination angle affects the dust deposition rate. The action space is defined as A = {no cleaning, dry brush cleaning, water wash cleaning}, corresponding to different cleaning methods and their costs and effects. In particular, the reward function is designed as R = Δ power generation revenue - cleaning cost + penalty term, where Δ power generation revenue is the revenue brought by the improved power generation efficiency after cleaning, the cleaning cost is calculated according to the cleaning method, and the penalty term is -0.1×D, where D is the cumulative number of uncleaned days, reflecting the marginal increasing effect of efficiency loss.
[0010] Furthermore, the Double DQN algorithm is used for model training, and the learning effect of high-reward decision-making samples is strengthened through the prioritized experience replay mechanism. The target network parameters are updated once every 1000 iterations, and the update formula is θ′ = τθ + (1 - τ)θ′, where τ is the update coefficient and its value is 0.01. The initial exploration rate is set to 0.5 and decays exponentially according to the formula ε = ε initial × γt to 0.1, where γ is the decay coefficient and its value is 0.995, and t is the number of training steps.
[0011] In the real-time decision-making stage, the current environmental state is encoded as a state vector S and input into the trained DRL model for inference. For example, when the rainfall probability in the next 48 hours is higher than 60%, the model recommends delaying cleaning until after the rain; when the PM10 concentration exceeds 80 μg / m 3 for 6 consecutive hours, the model recommends dry brush cleaning. The decision output includes the optimal cleaning time window, the recommended cleaning method, and the expected net revenue. Among them, the cleaning time window is the next 6 to 12 hours, the recommended cleaning method is dry brush cleaning or water wash cleaning, and the expected net revenue is the difference between the cleaning cost and the power generation revenue.
[0012] The innovative points of the technical solution of the present invention lie in multi-source data fusion and non-linear modeling. By fusing short-term weather forecasts, long-term fouling attenuation laws, and equipment status data, a deep reinforcement learning model is used to capture their complex non-linear relationships, solving the problem of high prediction error rates in linear models in the prior art. In particular, the dynamic cost-benefit balance mechanism realizes the optimal choice of marginal benefits of cleaning operations by designing a reward function that includes economic constraints, avoiding negative revenue operations such as "cleaning costs higher than power generation revenues" in traditional methods.
[0013] Furthermore, the edge-cloud collaborative computing architecture deploys a lightweight DRL model at the edge for real-time inference, while using the cloud to complete incremental model training, taking into account both response speed and algorithm iteration requirements. The hardware at the edge uses a low-power AI computing module, with a single decision-making time less than 500 milliseconds, and the output result is compatible with the work order instructions in JSON format. The cloud aggregates the data of each edge node every month to retrain the model and verify its performance, and uses differential privacy technology to protect the power plant data to ensure the security of model sharing.
[0014] In addition, the present invention uses the Kalman filter algorithm to perform real-time state estimation on PM2.5 concentration and humidity sensor data, generates a comprehensive fouling index and incorporates it into the state space, further improving the data fusion ability and robustness. When the model output conflicts with the preset business rules, an artificial review process is started and the review results are fed back to the model for online learning. In particular, the federated learning framework is used to aggregate the local model parameters of multiple power plants to generate a globally optimized model, and each power plant only shares the model parameters instead of the original data, ensuring data security and model generality.
[0015] Through the above technical solutions, the present invention achieves a technical effect of reducing the cleaning frequency by 35%-40% and increasing the average annual net income by 18%-22%. In extreme weather such as sandstorms, the model misjudgment rate is less than 10%, which is significantly better than the threshold warning system. The edge-cloud collaborative computing architecture supports the real-time scheduling of gigawatt-level power plants, with a single decision-making time less than 1 second, and the average annual income continues to increase by about 2%-3%. In remote area photovoltaic power plants, the prediction error is reduced by 40%-50%, and the frequency of manual intervention is less than 5%, significantly reducing the operation and maintenance labor cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of the overall system architecture of the present invention;
[0017] Figure 2 It is a schematic diagram of the multi-source data acquisition module of the present invention;
[0018] Figure 3 It is a schematic diagram of the data preprocessing process of the present invention;
[0019] Figure 4Schematic diagram of the deep reinforcement learning model structure of the present invention;
[0020] Figure 5 Schematic diagram of the edge-cloud collaborative computing architecture of the present invention;
[0021] Figure 6 Schematic diagram of the real-time decision output process of the present invention;
[0022] Figure 7 Schematic diagram of the Kalman filter data fusion process of the present invention;
[0023] Figure 8 Schematic diagram of the hybrid decision-making mechanism process of the present invention.
[0024] The reference numerals are as follows:
[0025] 1. Meteorological Bureau API interface; 2. Photovoltaic inverter; 3. Fouling sensor; 4. Sliding window method module; 5. Generative adversarial network (GAN) module; 6. Normalization processing module; 7. Deep reinforcement learning model; 8. Edge computing terminal; 9. Cloud server; 10. Cleaning execution interface; 11. Kalman filter algorithm module; 12. Comprehensive fouling index generation module; 13. Rule engine module; 14. Manual review process module; 15. Federated learning framework module. Detailed implementation manners
[0026] The present invention provides a photovoltaic cleaning decision-making method and system based on temporal feature fusion and marginal benefit optimization. Combining with FIGS. Figure 1 to FIGS. Figure 8 , its specific implementation manners are described in detail. This solution is applicable to the operation and maintenance management of large-scale photovoltaic power stations in sandstorm-prone areas. Through key technical modules such as multi-source heterogeneous data collection, deep reinforcement learning model construction, edge-cloud collaborative computing architecture, and real-time decision output, dynamic cleaning strategy optimization is realized.
[0027] The overall system architecture is as shown in Figure 1 . The system includes a Meteorological Bureau API interface 1, a photovoltaic inverter 2, a fouling sensor 3, a sliding window method module 4, a generative adversarial network (GAN) module 5, a normalization processing module 6, a deep reinforcement learning model 7, an edge computing terminal 8, a cloud server 9, a cleaning execution interface 10, a Kalman filter algorithm module 11, a comprehensive fouling index generation module 12, a rule engine module 13, a manual review process module 14, and a federated learning framework module 15. Each module works together in the system to ensure the efficient operation of the entire process from data collection to decision output.
[0028] The specific implementation process of S1 data collection is as follows: First, high-precision meteorological data for the next 72 hours in the target area is obtained through the meteorological bureau API interface 1, including sunshine intensity, temperature, humidity, PM2.5 concentration, and rainfall probability, with a sampling interval of 15 minutes. These data are subjected to time series feature extraction through the sliding window method module 4. The window size is set to 24 hours, and the step size is 1 hour, generating a feature matrix containing short-term meteorological change trends. For example, the sunshine intensity sequence for the next 24 hours is encoded as a 24-dimensional vector to capture the change trend of sunshine intensity. At the same time, historical power generation efficiency records are obtained through the photovoltaic inverter 2, including zonal power generation efficiency, cleaning operation time, and cost-benefit data of the cleaning method, covering at least 3 years to reflect seasonal change patterns. In addition, the fouling sensor 3 collects equipment status data, including the inclination angle of the photovoltaic panel, surface fouling detection values, and the operating status of the robotic cleaning equipment. All data is stored in the cloud server 9 for subsequent preprocessing and model training.
[0029] The detailed steps of S2 data preprocessing are as Figure 3 shown. The generative adversarial network (GAN) module 5 is used to complete missing or abnormal data. The generator is composed of an LSTM network, which simulates the time series correlation of meteorological data and generates estimated values of missing data. The discriminator introduces an attention mechanism to enhance the detection ability of PM2.5 concentration mutations, ensuring that the completed data meets physical constraint conditions. For example, the humidity data predicted by the generator is restricted between 0 and 100, and the PM2.5 concentration is restricted to non-negative values. The generator and the discriminator are continuously optimized through adversarial learning, and finally high-quality completed data is generated. Subsequently, the normalization processing module 6 performs a normalization operation on all input data, mapping it to the [0,1] interval to eliminate the dimensional differences between different data dimensions. For example, the normalization formula for sunshine intensity is X′ = (X - Xmin) / (Xmax - Xmin), where Xmin and Xmax are the minimum and maximum values in the historical data respectively.
[0030] The construction and training of the deep reinforcement learning model in S3 are as Figure 4As shown in the figure, the state space S is defined as [current power generation efficiency, future 24-hour meteorological prediction vector, cumulative uncleaned days, panel inclination angle]. The current power generation efficiency is normalized to the interval [0, 1]. The future 24-hour meteorological prediction is encoded as a 24-dimensional time series vector. The cumulative uncleaned days reflect the cumulative effect of fouling, and the panel inclination angle affects the dust deposition rate. The action space is defined as A = {no cleaning, dry-brush cleaning, water-washing cleaning}, corresponding to different cleaning methods, their costs, and effects respectively. The reward function is designed as R = Δ power generation revenue - cleaning cost + penalty term, where Δ power generation revenue is the revenue brought by the improved power generation efficiency after cleaning, the cleaning cost is calculated according to the cleaning method, and the penalty term is -0.1×D, where D is the cumulative uncleaned days. The Double DQN algorithm is used for model training. The prioritized experience replay mechanism assigns higher sampling weights to historical high-reward decision samples. The target network parameters are updated every 1000 iterations, and the update formula is θ′ = τθ + (1 - τ)θ′, where τ is the update coefficient with a value of 0.01. The initial exploration rate is set to 0.5 and decays exponentially to 0.1 according to the formula ε = ε initial × γt, where γ is the decay coefficient with a value of 0.995 and t is the number of training steps.
[0031] The specific implementation of the S4 edge-cloud collaborative computing architecture is as Figure 5 shown. A lightweight DRL model is deployed on the edge computing terminal 8. The hardware uses a low-power AI computing module, and the single decision-making time is less than 500 milliseconds, meeting the real-time requirements. The edge side receives sensor data and generates cleaning suggestions. The output result is compatible with the work order instructions in JSON format. The decision-making instructions are converted into robot control signals through the cleaning execution interface 10 and support docking with the SCADA system. The cloud server 9 is responsible for model incremental training and parameter optimization. It aggregates the data of each edge node every month to retrain the model and verify the performance, and uses differential privacy technology to protect the power station data to ensure the security of model sharing.
[0032] The process of S5 real-time decision output is as Figure 6 shown. The current environmental state is encoded as a state vector S and input into the trained DRL model for inference. For example, when the rainfall probability in the next 48 hours is higher than 60%, the model recommends delaying cleaning until after the rain; when the PM10 concentration exceeds 80 μg / m 3 for 6 consecutive hours, the model recommends dry-brush cleaning. The decision output includes the optimal cleaning time window, the recommended cleaning method, and the expected net revenue. Among them, the cleaning time window is the next 6 to 12 hours, the recommended cleaning method is dry-brush cleaning or water-washing cleaning, and the expected net revenue is the difference between the cleaning cost and the power generation revenue.
[0033] The specific implementation of S6 Kalman filter data fusion is as Figure 7As shown, the Kalman filter algorithm module 11 performs real-time state estimation on the PM2.5 concentration and humidity sensor data, generates a comprehensive fouling index, and incorporates it into the state space. The comprehensive fouling index is calculated by the comprehensive fouling index generation module 12, which can reduce noise interference and improve data fusion ability and robustness. For example, when the readings of the PM2.5 concentration sensor fluctuate greatly, the Kalman filter algorithm module 11 generates a more stable comprehensive fouling index through state estimation, avoiding misjudgment caused by noise.
[0034] The implementation of the S7 hybrid decision-making mechanism is as Figure 8 shown. The rule engine module 13 presets extreme condition response strategies, such as forcibly starting dry-brush cleaning during sandstorms. The DRL model makes autonomous decisions in a normal environment and evaluates the feasibility of actions through Q values. When the cleaning method output by the model conflicts with the preset business rules, the manual review process module 14 is activated, and the review results are fed back to the model for online learning. For example, when the model recommends performing water washing cleaning at night, the rule engine module 13 will block this decision and trigger the manual review process module 14 to ensure that the decision conforms to the business logic.
[0035] The implementation of the S8 federated learning framework is as Figure 5 shown. The federated learning framework module 15 is used to aggregate the local model parameters of multiple power stations to generate a globally optimized model. Each power station only shares the model parameters instead of the original data, ensuring data security and model generality. For example, multiple photovoltaic power stations distributed in different regions can share model parameters through the federated learning framework module 15 to generate a globally optimized model suitable for various climate conditions, further improving the prediction accuracy and generalization ability of the model.
[0036] Through the above specific implementation manners, the present invention achieves technical effects of reducing the cleaning frequency by 35% - 40% and increasing the annual average net income by 18% - 22%. In extreme weather such as sandstorms, the misjudgment rate of the model is less than 10%, which is significantly better than the threshold warning system. The edge-cloud collaborative computing architecture supports the real-time scheduling of gigawatt-level power stations, with the single decision-making time less than 1 second, and the annual average income continuously increasing by about 2% - 3%. In remote photovoltaic power stations, the prediction error is reduced by 40% - 50%, and the frequency of manual intervention is less than 5%, significantly reducing the operation and maintenance labor cost.
Claims
1. A photovoltaic cleaning decision-making method based on temporal feature fusion and marginal benefit optimization, characterized in that It includes the following steps: Obtain the meteorological data of the target area for the next 72 hours through the meteorological bureau API interface (1), with a sampling interval of 15 minutes; obtain the historical power generation efficiency data through the photovoltaic inverter (2); collect the equipment status data through the fouling sensor (3); use the sliding window method module (4) to generate the time series feature matrix of the meteorological data, with a window size of 24 hours and a step size of 1 hour; use the generative adversarial network (GAN) module (5) to complete the filling of missing or abnormal data; normalize all input data to the interval [0,1] through the normalization processing module (6); construct a deep reinforcement learning model (7) and define the state space, action space and reward function; deploy a lightweight inference model on the edge computing terminal (8) and output the real-time decision result; complete the incremental training of the model through the cloud server (9).
2. The method according to claim 1, wherein The meteorological data includes solar radiation intensity, temperature, humidity, PM2.5 concentration and rainfall probability.
3. The method according to claim 2, wherein The generator of the generative adversarial network (GAN) module (5) uses an LSTM network to simulate the time series correlation of meteorological data, and the discriminator introduces an attention mechanism to detect the mutation of PM2.5 concentration.
4. The method according to claim 1, characterized in that The state space is defined as S = [current power generation efficiency, meteorological prediction vector for the next 24h, cumulative number of uncleaned days, panel tilt angle], where the cumulative number of uncleaned days reflects the fouling accumulation effect, and the panel tilt angle affects the dust deposition rate.
5. The method according to claim 4, characterized in that The action space is defined as A = {no cleaning, dry brush cleaning, water wash cleaning}, corresponding to different cleaning methods and their costs and effects respectively.
6. The method according to claim 1, characterized in that The reward function is designed as R = Δ power generation revenue - cleaning cost + penalty term, where the penalty term is -0.1×D, and D is the cumulative number of uncleaned days.
7. The method according to claim 1, wherein The edge computing terminal (8) uses a low-power AI computing module, with a single decision-making time less than 500 milliseconds, and outputs a work order instruction in JSON format.
8. The method according to claim 1, wherein The Kalman filter algorithm module (11) performs real-time state estimation on the PM2.5 concentration and humidity sensor data, generates a comprehensive fouling index and incorporates it into the state space.
9. The method according to claim 1, wherein When the model output conflicts with the preset business rules, start the manual review process module (14), and feedback the review result to the model for online learning.
10. The method according to claim 1, characterized in that The federated learning framework module (15) is used to aggregate the local model parameters of multiple power stations to generate a globally optimized model, and each power station only shares the model parameters rather than the original data.