Photovoltaic module fault early warning method and system based on string current data mining

Through string current data mining and CNN-BiLSTM model prediction, combined with sliding window technology, the accuracy and adaptability of photovoltaic module fault warning are solved, and efficient early warning and operation and maintenance efficiency of photovoltaic module faults are achieved. It is suitable for GW-level power stations.

CN120498381APending Publication Date: 2025-08-15GUODIAN QUANZHOU POWER GENERATION CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510707082.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing photovoltaic module fault warning methods have problems such as idealized physical models, poor generalization of data-driven, rigid thresholds and insufficient multimodal utilization, resulting in high false alarm rates, high missed alarm rates, and low operation and maintenance efficiency, making it difficult to meet the intelligent needs of GW-level super-large-scale power stations.

Method used

Through string current data mining, fusion of environmental parameters such as temperature and irradiance, current prediction is performed using the CNN-BiLSTM model, and an alarm threshold mechanism is established in combination with the idea of sliding window to achieve early warning of photovoltaic module failures.

Benefits of technology

It improves the accuracy and adaptability of photovoltaic module fault warning, reduces the false alarm rate, improves operation and maintenance efficiency, and supports unmanned operation and maintenance of GW-level power stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498381A_ABST
    Figure CN120498381A_ABST
Patent Text Reader

Abstract

The invention discloses a photovoltaic module fault early warning method and system based on string current data mining. On the basis of the Internet of Things and the embedded technology, an ammeter is used for collecting string current data of a photovoltaic module, a string current prediction model is built around a bidirectional long-short-term memory network, and feature extraction is optimized by adopting a convolutional neural network. Then modeling is carried out on three weather conditions of a sunny day, a cloudy day and a rainy day, training is carried out, experimental analysis is carried out according to the model, finally, an alarm threshold mechanism is established according to the calculated deviation rate index in combination with a sliding window method, and timely early warning of the photovoltaic module before fault alarm pushing is achieved. According to the invention, accurate judgment of the fault of the photovoltaic module can be realized only through historical string current data, operation and maintenance personnel can be helped to discover abnormity in time and carry out predictive maintenance work before the alarm information is reported, the operation efficiency of the system is improved to a greater extent, and thus the economic benefit of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data analysis, and in particular relates to a photovoltaic module fault early warning method and system based on string current data mining. Background Art

[0002] As my country's photovoltaic power generation capacity continues to expand, intelligent and efficient operation and maintenance (O&M) of photovoltaic systems has become a key focus of the industry. Traditional manual O&M is evolving towards intelligent decision-making and data mining based on the Internet of Things, cloud platforms, and big data. Early warning systems are playing an increasingly critical role in O&M: by proactively identifying potential faults or anomalies (e.g., string current, voltage, discrete rate, full-load hours, and other indicators), preventive maintenance measures can be implemented to reduce equipment failure rates and minimize economic losses.

[0003] Currently, there are many commonly used methods for warning of abnormal faults in photovoltaic modules. Among them, the threshold method determines anomalies by setting thresholds (such as full-time hour deviation and discrete rate threshold); the physical characteristic method relies on the IV characteristic model of the photovoltaic system, calculates the difference between ideal power and actual power, and combines single diode model parameters (photocurrent, dark saturation current, etc.) to evaluate performance. In addition, data-driven methods are divided into statistical methods and machine learning methods. Statistical methods analyze parameter residuals or significant differences through least squares methods, t-tests, etc., and combine fuzzy inference systems (such as shading and irradiance data) to predict status; define the normal operating boundaries of the system, and determine anomalies if the boundaries are exceeded. Machine learning methods are currently the mainstream technology, combining SVM (support vector machine) and SVR (support vector regression) combined models to identify anomalies by the difference between predicted power generation and actual values (such as SVR estimates expected power generation, SVM determines anomalies). Both methods judge anomalies based on "comparison between actual values and predicted values." However, these methods have their own shortcomings:

[0004] Limitations of the physical model approach:

[0005] Existing physical models (such as the single-diode model) assume that all photovoltaic modules have consistent performance parameters (such as photocurrent and dark saturation current), ignoring the differences in module characteristics caused by aging, microcracks, hot spot effects, and manufacturing tolerances in actual operation. This idealized assumption can introduce systematic errors, especially in the later stages of power plant operation, when module performance becomes increasingly discrete, significantly reducing model prediction accuracy.

[0006] Solving parameters based on the IV characteristic curve relies on iterative nonlinear equations (such as the Newton-Raphson method), which is computationally complex and prone to convergence failure, making it difficult to meet real-time monitoring requirements. For example, parallel computing for thousands of strings in a large-scale power plant can overload cloud platform resources.

[0007] Physical models have limited ability to respond to dynamic environmental factors (such as instantaneous shadows, dust accumulation, and irradiance fluctuations), and it is difficult to quantify the coupling effects of multiple factors (such as the impact of efficiency reduction caused by temperature increase and the superposition of dust obstruction).

[0008] Bottlenecks of data-driven approaches:

[0009] Statistical methods (such as t-tests and residual analysis) and traditional machine learning models (SVM / SVR) rely heavily on the quality and completeness of historical data. For new fault types (such as PID effects and backplane cracking) or scenarios with small sample sizes (such as new power plants), the model's generalization ability drops dramatically, leading to missed or false positives.

[0010] Single models, such as SVR, are insufficiently capable of representing the complex nonlinear relationships in photovoltaic systems. For example, rapid fluctuations in sunlight intensity during cloudy weather can lead to non-stationary time series characteristics in power generation, making it difficult for traditional regression models to capture these high-frequency variations.

[0011] Existing methods rely on manually defined features (such as discrete rate and full-power hours), and are unable to fully explore implicit associations in high-dimensional data (such as the potential connection between alarm codes in inverter logs and string current anomalies). In addition, the dynamic characteristics of time series data (such as seasonal decay trends) are not effectively modeled.

[0012] Static and adaptive defects of the threshold method:

[0013] Fixed thresholds (e.g., triggering an alarm when the dispersion ratio exceeds 10%) cannot adapt to dynamic scenarios such as component performance degradation and long-term environmental changes (e.g., differences in dust accumulation rates during the rainy season). For example, the same power deviation in winter with low irradiance may be a normal fluctuation, while in summer it may indicate a fault, but static thresholds cannot distinguish between these scenarios.

[0014] Existing threshold methods typically set thresholds independently for a single indicator (such as current or voltage), lacking a mechanism for joint analysis of multiple indicators. For example, a current drop accompanied by abnormal voltage fluctuations could indicate a string disconnection, but a single threshold method struggles to identify such complex patterns.

[0015] Insufficient utilization of multimodal data:

[0016] Environmental sensor data (irradiance, temperature), electrical performance data (current, voltage), infrared thermal imaging data, and maintenance logs are stored in a decentralized manner, lacking cross-modal correlation analysis. For example, a current drop caused by dust accumulation may be coupled with a local temperature rise in an infrared image, but existing methods lack such cross-modal feature fusion mechanisms.

[0017] The topology of the photovoltaic array (such as the string-parallel relationship) is not effectively modeled, making it impossible to identify fault propagation paths (such as a branch fault causing an overload in an adjacent string). Furthermore, the temporal evolution of faults (such as the progressive deterioration of hot spot effects) has not been fully explored.

[0018] In summary, the shortcomings of existing methods directly lead to two major risks for photovoltaic power plants: 1. Economic risk: A high false alarm rate (approximately 15%-20%) causes unnecessary downtime and maintenance, resulting in power generation losses; missed alarms lead to the spread of faults (such as hot spots causing fires), doubling repair costs. 2. Technical risk: Relying on manual experience to adjust thresholds and model parameters leads to low O&M efficiency and is unable to support the intelligent needs of gigawatt-scale power plants.

[0019] Precisely because of the shortcomings of the existing technology, the technical solution of the present invention can be proposed. Through technical breakthroughs such as physical-data fusion modeling, dynamic threshold optimization, and multimodal deep learning, a high-precision, highly adaptable, and scalable photovoltaic intelligent early warning system is constructed to achieve an upgrade of the operation and maintenance mode from "post-fault processing" to "pre-risk control". Summary of the Invention

[0020] In response to the core defects of the existing technology, such as idealized physical models, poor data-driven generalization, rigid thresholds, and insufficient multi-modal utilization, the present invention proposes a photovoltaic module fault warning method and system based on string current data mining.

[0021] The present invention uses only one basic parameter, string current, and integrates environmental parameters such as temperature, irradiance, and humidity to predict the PV string current based on the CNN-BiLSTM model. It also models different weather conditions separately, compares the predicted results with the actual values reported in real time by the PV strings, calculates the current deviation rate index, and establishes an alarm threshold mechanism based on the sliding window concept, so as to carry out early warning and maintenance before the fault alarm actually occurs.

[0022] In a first aspect, the present invention provides a photovoltaic module fault early warning method based on string current data mining, comprising the following steps:

[0023] Step 1. Collect historical data including string current and related environmental parameters;

[0024] Step 2. Preprocess the collected historical data, including filtering the data set under specific weather conditions, normalizing the data, and dividing it into training and test sets;

[0025] Step 3. Use convolutional neural networks to perform feature compression and feature extraction on the preprocessed data. The convolution layer extracts spatial feature information, and the pooling layer performs information filtering and feature compression.

[0026] Step 4. Build a deep learning model and input the data processed by the convolutional neural network into a bidirectional long short-term memory network to extract temporal feature information from the data. Use the gating units of the bidirectional long short-term memory network to extract the temporal relationship in the data and establish a bidirectional time fitting relationship.

[0027] Step 5. Compare the prediction results of the deep learning model with the actual values reported in real time by the PV strings to calculate the current deviation rate indicator;

[0028] Step 6. Analyze the deviation rate indicator based on the sliding window concept, and perform fault warning processing of varying degrees according to the statistics of the deviation rate within the window.

[0029] In a second aspect, the present invention provides a photovoltaic module fault warning system based on string current data mining, comprising:

[0030] Data acquisition module, used to collect historical data including string current and related environmental parameters;

[0031] The data preprocessing module is used to preprocess the collected historical data, including filtering the data set under specific weather conditions, normalizing the data, and dividing it into training and test sets;

[0032] The feature extraction module is used to perform feature compression and feature extraction on the preprocessed data using a convolutional neural network. The convolution layer extracts spatial feature information and the pooling layer performs information filtering and feature compression.

[0033] The model building module is used to build a deep learning model. It inputs the data processed by the convolutional neural network into the bidirectional long short-term memory network, extracts the temporal feature information of the data, extracts the temporal relationship in the data through the gating unit of the bidirectional long short-term memory network, and establishes a bidirectional time fitting relationship.

[0034] The prediction and comparison module is used to compare the prediction results of the deep learning model with the actual values reported in real time by the photovoltaic strings and calculate the current deviation rate indicator;

[0035] The decision module is used to analyze the deviation rate indicator based on the sliding window idea, and perform fault warning processing of different degrees according to the statistics of the deviation rate within the window.

[0036] Beneficial effects of the present invention:

[0037] The present invention uses historical string current data to predict current by constructing a CNN-BiLSTM model, and compares it with real-time PV string data to calculate the deviation rate for early warning. This can help operation and maintenance personnel promptly identify problems and quickly resolve them, thereby effectively reducing the probability of equipment failure.

[0038] The present invention uses the CNN-BiLSTM model to predict current, which combines the sensitivity of convolutional neural networks to local features and the modeling ability of bidirectional long short-term memory networks to long-term temporal dependencies. It can simultaneously capture short-term fluctuations and long-term trends in current data, improving the comprehensiveness and accuracy of the prediction.

[0039] The present invention uses a CNN-BiLSTM model to predict current. The convolutional layer of the CNN supports parallel processing and can quickly extract the spatial features of the input data. Combined with BiLSTM, through phased training or modular design, it optimizes the allocation of computing resources, significantly shortens the model training time, and is more suitable for processing massive historical photovoltaic current data.

[0040] The present invention uses the CNN-BiLSTM model to predict current. The convolution kernel filtering and pooling operations of the CNN can effectively suppress random noise interference in the original current data, and the BiLSTM smoothes abnormal fluctuations through the temporal memory mechanism, doubly ensuring the stability of the model output.

[0041] The present invention uses a CNN-BiLSTM hybrid model to predict current, which can avoid the limitations of a single model and improve the overall prediction accuracy by integrating the advantages of both models. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flow chart of the photovoltaic module fault early warning method of the present invention;

[0043] Figure 2 This is the flow chart of the CNN-BiLSTM model prediction algorithm;

[0044] Figure 3 This is a comparison chart of the predictions of the three models for sunny conditions;

[0045] Figure 4 This is a comparison chart of the predictions of the three models for cloudy conditions;

[0046] Figure 5 This is a comparison chart of the predictions of the three models for rainy days. DETAILED DESCRIPTION

[0047] In order to make the objectives, technical solutions and beneficial effects of the present invention more clearly described, the present invention will be further described in detail below with reference to the accompanying drawings and specific examples.

[0048] like Figure 1 This is a flow chart of a photovoltaic module fault warning method based on string current data mining according to an embodiment of the present application, which includes the following steps:

[0049] S1: Collect historical string current value data.

[0050] First, the historical data provided by the system is collected, which includes the PV string current from 7:00 am to 7:00 pm in a certain period of time, to form a data set. The attributes contained in the data set are temperature, irradiance, humidity, and the PV string current at the corresponding time.

[0051] In a preferred example:

[0052] This implementation uses a 1490 kW photovoltaic power station in Jiangsu as an example. Historical string current data from January 1, 2023, to December 31, 2023, is collected from the system. The data set for PV string current prediction includes attributes such as temperature, irradiance, humidity, and the corresponding PV string current. The specific time range is selected from 7:00 AM to 7:00 PM, with a total of 144 sampling points per day, resulting in 22,530 sets of data.

[0053] S2: Filter the sunny, cloudy and rainy day datasets for preprocessing.

[0054] The datasets of sunny, cloudy and rainy days were selected for separate training. The datasets were normalized and divided into training set and test set in a ratio of 8:2.

[0055] In a preferred example:

[0056] S21: Filter the data sets of sunny, cloudy and rainy days for the historical string current data.

[0057] S22: Normalize the data set.

[0058] Input data normalization: This process maps the input data values to a specific region to facilitate further analysis of the data's characteristics. Data normalization can reduce the impact of large ranges in the original data, improving model training speed and prediction accuracy. This example uses min-max normalization to keep the input data between [0, 1].

[0059]

[0060] Where, X * is the normalized data; x max and x min are the minimum and maximum values of the sample data set, and are the original sample data.

[0061] S22: Divide the processed data set into a training set and a test set in a ratio of 8:2.

[0062] S3 uses the CNN algorithm to perform feature compression and feature extraction on the data set.

[0063] The CNN algorithm is adopted, and the CNN convolution layer is used to extract spatial feature information of the data. The CNN pooling layer is used to filter the information. The maximum pooling method is used to output the largest data in the regional grid to achieve feature compression.

[0064] In a preferred example:

[0065] S31: Use CNN convolutional layer to extract spatial feature information from data.

[0066] Convolutional Neural Networks (CNNs) are feedforward neural networks. Unlike traditional fully connected neural networks, CNNs extract local features by applying convolution operations to input data and automatically learn the parameters of these convolution operations through training. With their excellent nonlinear fitting and feature extraction capabilities, CNNs are capable of effectively extracting information from two-dimensional images and high-dimensional data, enabling them to solve complex problems.

[0067] Convolutional layers extract features through convolution operations. Each convolutional layer typically includes multiple convolution kernels, each of which performs a convolution operation on the input data to produce a feature map. Convolution kernels also use a rectified linear unit (ReLU) as an activation function to perform nonlinear calculations. The performance of a convolutional layer is determined by the number and size of convolution kernels, as well as the step size. The calculation formula is as follows:

[0068]

[0069] ReLU(x)=max(0,x)

[0070] in is the j-th feature map output of layer l, f is the activation function, and s is the number of input feature maps; is the j-th mapping vector of the l-1th layer, is a trainable convolution kernel, For bias.

[0071] S32: Use CNN pooling layer to filter information.

[0072] The pooling layer reduces the size of the feature map by downsampling, enhancing the robustness of the model and its feature extraction capabilities. Max pooling is used to extract the maximum value of feature points within a region, that is, to extract the most obvious features, thereby achieving feature compression.

[0073] S4: Build CNN-BiLSTM current model.

[0074] The information processed by the CNN is fed into a single-layer BiLSTM network through a Dropout layer, which transfers it according to the time step (a one-dimensional vector). Temporal features are extracted from the data. The network's unique gating units extract temporal relationships within the data and learn and establish a bidirectional temporal fitting relationship to prevent overfitting. The Dropout layer connects to the Dense layer, which strengthens the data features. The training checks whether convergence conditions are met. If not, iterations continue; otherwise, training terminates. The model outputs the current prediction results.

[0075] In a preferred example:

[0076] S41: Set the initialization parameters of the CNN-BiLSTM network model as shown in the following table:

[0077]

[0078]

[0079] S42: The information processed by CNN is transferred to a single-layer BiLSTM network through the Dropout layer according to the time step (one-dimensional vector). The temporal feature information of the data is extracted. The unique gating unit in the network is used to extract the temporal relationship in the data and learn and establish a bidirectional temporal fitting relationship to prevent overfitting.

[0080] The Dropout layer maps features into a high-dimensional feature space and then performs classification or regression using a softmax function. During training, a backpropagation algorithm is typically used to update network parameters to minimize the loss function. In regression tasks, the output of the Dropout layer can be a continuous value. By adjusting the weight matrix and bias term, the Dropout layer can learn the relationship between the input features and the regression results.

[0081] LSTM is a deep learning model suitable for time series data. It has memory units to capture long-term dependencies and is suitable for processing data with time correlation.

[0082] LSTM primarily consists of an input gate, a forget gate, and an output gate. The input gate analyzes the output of the previous time step and combines it with the current input to select the content that needs to be updated in the cell state. The forget gate determines the information to be forgotten in the cell state based on the previous output and the current input value, which is controlled by a sigmoid function. The output gate selects the information to be output in the cell state, which is updated according to the following equation. The improvement of LSTM lies in the addition of an additional memory unit to the original network, which can remember and store past information. Furthermore, the output of the memory module at time t in the LSTM model is determined by the output gate and the cell state.

[0083] f t =σ(W f ·[h t-1 ,x t ]+b f )

[0084] i t =σ(W f ·[h t-1 ,x t ]+b f )

[0085] a t =tanh(W c ·[h t-1 ,x t ]+b c )

[0086] C t =f*C t-1 +i t *a t

[0087] o t =σ(W o ·[h t-1 ,x t ]+b o )

[0088] h t =o t *tanh(C t )

[0089] Where x t is the input vector; h t is the output vector; i represents the input gate; o represents the output gate; f represents the forget gate; C t Indicates the current state; C t-1 Indicates the memory information status of the previous moment; h t-1 represents the output of the hidden layer unit at the previous moment; ht represents the output of the hidden layer unit; σ represents the singmoid activation function; tanh represents the tangent function; W represents the weight matrix; b represents the bias vector.

[0090] Bi-directional Long Short-Term Memory (BiLSTM) is an extended LSTM. Traditional LSTM can only capture contextual information before the current moment, while BiLSTM can more comprehensively understand and model sequence data by considering both historical context and future context.

[0091] BiLSTM implements bidirectional modeling by introducing a forward LSTM and a backward LSTM between time steps. One extracts long-term dependency features from the input sequence forward and backward, while the other extracts long-term dependency features from the input sequence backward and forward. It has the ability to recurse and feedback into the past and future. During the training phase, the forward and backward LSTMs process the input sequence separately and concatenate their hidden states, which are then passed as input to subsequent layers or used as output for tasks. During the prediction phase, the forward and backward LSTMs process the input sequence separately and concatenate their hidden states to perform various prediction tasks.

[0092] S42: Connect the Dense layer through the Dropout layer and strengthen the data information features through the Dense layer to perform final classification and regression on the features extracted in the previous layer.

[0093] S43: Check whether the convergence condition is met. If not, continue iteration; otherwise, stop training.

[0094] To train a CNN-BiLSTM network model, loop through the following steps until convergence or the maximum number of iterations is reached: First, input the training data into the CNN-BiLSTM for forward propagation and calculate the predicted value. Then calculate the loss function and calculate the gradient through backpropagation. Then update the network parameters based on the gradient and learning rate. Finally, check whether the convergence conditions are met. If not, continue iterations. Otherwise, stop training. Figure 2 .

[0095] S44: Output the current prediction result of the model.

[0096] The test set is input into CNN-BiLSTM for forward propagation prediction to obtain the current prediction result.

[0097] S5: Calculate the deviation rate between the predicted string current and the real-time current to provide fault warning.

[0098] The system compares the predicted current with the actual values reported in real time by the PV strings to calculate the current deviation rate. Using a sliding window, the system divides the entire time series into multiple fixed-size windows. Feature extraction and analysis are performed within each window, and fault warnings of varying degrees are applied based on the deviation rate of the current series within a given window. A level 1 fault warning is issued if the deviation rate of all 12 current series within a given window exceeds 0.2, while a level 2 fault warning is issued if the deviation rate exceeds 0.4.

[0099] In a preferred example:

[0100] S51: Calculate the current deviation rate index based on the current prediction result and the actual value reported in real time by the photovoltaic string.

[0101] S52: Use sliding windows to split the entire time series into multiple fixed-size windows, and perform feature extraction and analysis in each window.

[0102] S52: When the deviation rates of the 12 current sequences within a certain time window all exceed 0.2, a first-level fault warning is issued; when the deviation rates all exceed 0.4, a second-level fault warning is issued.

[0103] The embodiment of this application uses a Dropout layer to prevent overfitting. During model training, a Dropout layer is added after each layer, with the corresponding parameter value set to 0.25. The Dropout layer is widely used in each hidden stage, and during each forward and backward propagation, some nodes are randomly deactivated with a certain probability and the parameters are updated, thereby improving the generalization ability and stability of the model.

[0104] The application also uses the mean absolute error (MAE), root mean square error (RMSE) and complex determination coefficient (R 2 ) to evaluate the model.

[0105] The performance evaluation indicators are mean absolute error (MAE), root mean square error (RMSE) and correlation coefficient (R 2 ) are used to measure the performance of the algorithm, which are three important indicators for evaluating the model.

[0106] MAE is the average of the absolute differences between the predicted values and the true values. A smaller MAE value indicates a smaller difference between the predicted results and the actual observed values, i.e., a higher prediction accuracy; a larger MAE value indicates a larger difference between the predicted results and the actual observed values, i.e., a lower prediction accuracy. The calculation formula is as follows:

[0107]

[0108] RMSE is the average size of the measurement error, which is the square root of the average of the squared differences between the calculated value and the true value. It is more intuitive in terms of magnitude. The smaller the result of the indicator, the better the model calculation effect. The calculation formula is as follows:

[0109]

[0110] Among them, M is the number of calculations, y is the true value, Calculate values for the model

[0111]

[0112] For the same data set, a higher R 2 Indicates that the difference between the observed data and the fitted value is small; R 2 The value range is from -1 to +1. Values closer to +1 indicate a stronger linear relationship between the two datasets. Values closer to -1 indicate an inverse linear relationship between the two datasets, and values closer to 0 indicate a weaker relationship between the two datasets.

[0113] Verification example:

[0114] This experiment uses the optimal parameters of the above process to establish a CNN-BiLSTM model, and compares the predictions of the LSTM model and the CNN-LSTM model. The predictions are divided into three conditions: sunny, cloudy, and rainy. Figure 3 、 Figure 4 and Figure 5 .

[0115]

[0116]

[0117] The table above compares the evaluation indicators of the three models under sunny, cloudy and rainy conditions. The MAE and RMSE values obtained by the CNN-BiLSTM model are the smallest compared to the other two single models. 2 The CNN-BiLSTM prediction model has good prediction accuracy under all three weather conditions.

[0118] The summary analysis of prediction errors under the three different weather conditions above shows that the CNN-BiLSTM model performs well in predicting PV string currents, with low prediction errors. It also demonstrates strong high-dimensional feature extraction capabilities and excellent time series prediction capabilities, reflecting the inherent connections and patterns in the operation of string currents in PV systems. The predictions for the three weather conditions above show that the CNN-BiLSTM model performs well across different weather conditions, with the impact of these errors being negligible and within acceptable limits, indicating that the experimental results met expectations.

[0119] The prediction model for photovoltaic string current data is established under the premise of stable component operation. In the model prediction constructed under three different weather conditions: sunny, cloudy, and rainy, it can be found that when the component is in normal working condition, the error between the actual current value collected by the smart gateway and the model's predicted value is very small. Therefore, by comparing the two values, the abnormal trend of the string current can be displayed, thereby providing real-time warning before the photovoltaic component fails. The embodiment of the present application proposes a deviation rate indicator dev for photovoltaic string current, which is calculated as follows:

[0120]

[0121] Y in the formula pred is the predicted value of the PV string current output by the model, Y true It is the actual value collected by the PV inverter.

[0122] When the predicted component is operating normally, the smaller the error between the predicted and actual values, the lower the calculated deviation rate. However, when the predicted component is operating abnormally, the difference between the actual and predicted values becomes more pronounced, and the deviation rate increases. In this case, a real-time fault warning needs to be reported.

[0123] The core idea of the traditional deviation rate algorithm is to compare the different performance indicators inside the device with the pre-set threshold parameters, and make corresponding early warning judgments based on the comparison results. If it is lower than the threshold, it is normal, and if it exceeds it, it is abnormal. However, if the cloud platform pushes an early warning once the string current deviation rate calculated by the cloud platform exceeds the early warning threshold, the burr effect caused by meteorological conditions will be ignored. Environmental factors in severe weather conditions are likely to cause burr jitter on the deviation rate of normal components, and even exceed the set secondary fault early warning threshold.

[0124] The embodiment of the present application uses a sliding window method to optimize the warning effect. The sliding window method can be used to analyze time series data. When analyzing a certain current section, the sliding window can divide the entire time series into multiple fixed-size windows and perform feature extraction and analysis in each window. The sliding window mainly consists of two properties: window size and step size. Each time the window length is determined, it will slide back by the length of the step size and divide the new window again. The same current information may appear in multiple windows.

[0125] The system's intelligent gateway transmits string current data for each branch detected by the PV inverter every five minutes. Therefore, the deviation rate calculation interval is also set to five minutes. To prevent a warning from being issued as soon as the string current deviation rate exceeds the warning threshold, an optimization based on a sliding window concept is implemented, treating the current data for the previous hour, including the current moment, as a sequence. Since calculations are performed every five minutes, each window contains 12 sets of sequence data. The calculated string current deviation rate for each branch is stored in the database.

[0126] If the deviation rate of all 12 current sequences within a given window exceeds 0.2, a Level 1 fault warning is issued; if the deviation rate exceeds 0.4, a Level 2 fault warning is issued. This optimization can significantly reduce the false alarm rate caused by glitches and jitter caused by meteorological conditions at a given moment.

[0127] In summary, through the above-mentioned technological breakthroughs, the present invention has achieved a leap in photovoltaic operation and maintenance from "threshold alarm" to "prediction and early warning", and from "manual analysis" to "intelligent decision-making", providing core support for unmanned operation and maintenance of GW-level power stations.

[0128] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A photovoltaic module fault warning method based on string current data mining, characterized in that: The following steps are involved: Step 1. Collect historical data including string current and related environmental parameters; Step 2. Preprocess the collected historical data, including filtering the data set under specific weather conditions, normalizing the data, and dividing it into training and test sets; Step 3. Use convolutional neural networks to perform feature compression and feature extraction on the preprocessed data. The convolution layer extracts spatial feature information, and the pooling layer performs information filtering and feature compression. Step 4. Build a deep learning model and input the data processed by the convolutional neural network into a bidirectional long short-term memory network to extract temporal feature information from the data. Use the gating units of the bidirectional long short-term memory network to extract the temporal relationship in the data and establish a bidirectional time fitting relationship. Step 5. Compare the prediction results of the deep learning model with the actual values reported in real time by the PV strings to calculate the current deviation rate indicator; Step 6. Analyze the deviation rate indicator based on the sliding window concept, and perform fault warning processing of varying degrees according to the statistics of the deviation rate within the window.

2. The photovoltaic module fault early warning method based on string current data mining according to claim 1 is characterized in that: The environmental parameters include temperature, irradiance and humidity.

3. The photovoltaic module fault early warning method based on string current data mining according to claim 1 or 2, characterized in that: The specific weather conditions include sunny days, cloudy days and rainy days. In the data preprocessing step, normalization processing is performed on the data sets under different weather conditions.

4. The photovoltaic module fault early warning method based on string current data mining according to claim 1 is characterized in that: The convolution layer of the convolutional neural network includes multiple convolution kernels, each of which performs a convolution operation on the input data and performs nonlinear calculations through an activation function; the pooling layer uses a maximum pooling method to downsample the feature map.

5. The photovoltaic module fault early warning method based on string current data mining according to claim 1 or 4, characterized in that: The bidirectional long short-term memory network includes a forward LSTM and a reverse LSTM, which extract long-term dependency features from the input sequence from front to back and from back to front respectively, and splice the forward and reverse hidden states together for subsequent processing.

6. The photovoltaic module fault early warning method based on string current data mining according to claim 5 is characterized in that: The size of the sliding window is 1 hour, the step length is 5 minutes, and each window contains 12 sets of current sequence data.

7. The photovoltaic module fault early warning method based on string current data mining according to claim 6 is characterized in that: The fault warning process includes: when the deviation rates of the 12 groups of current sequences in a certain time window all exceed 0.2, a first-level fault warning is issued; when the deviation rates all exceed 0.4, a second-level fault warning is issued.

8. The photovoltaic module fault early warning method based on string current data mining according to claim 1, characterized in that: During the training process of the deep learning model, mean absolute error, root mean square error and complex determination coefficient are used as performance evaluation indicators.

9. The photovoltaic module fault early warning method based on string current data mining according to claim 1, characterized in that: The current deviation rate indicator dev is calculated as follows: Among them, Y pred is the predicted value of the photovoltaic string current output by the model, Y true is the actual value collected by the PV inverter.

10. A photovoltaic module fault warning system based on string current data mining, characterized in that: include: Data acquisition module, used to collect historical data including string current and related environmental parameters; The data preprocessing module is used to preprocess the collected historical data, including filtering the data set under specific weather conditions, normalizing the data, and dividing it into training and test sets; The feature extraction module is used to perform feature compression and feature extraction on the preprocessed data using a convolutional neural network. The convolution layer extracts spatial feature information and the pooling layer performs information filtering and feature compression. The model building module is used to build a deep learning model. It inputs the data processed by the convolutional neural network into the bidirectional long short-term memory network, extracts the temporal feature information of the data, extracts the temporal relationship in the data through the gating unit of the bidirectional long short-term memory network, and establishes a bidirectional time fitting relationship. The prediction and comparison module is used to compare the prediction results of the deep learning model with the actual values reported in real time by the photovoltaic strings and calculate the current deviation rate indicator; The decision module is used to analyze the deviation rate indicator based on the sliding window idea, and perform fault warning processing of different degrees according to the statistics of the deviation rate within the window.

Citation Information

Cited By

  • Photovoltaic string reflux identification method and system based on current characteristics

    CN121071567A