Photovoltaic inverter anomaly detection method based on LSTM-DAGMM

By preprocessing and training photovoltaic inverter data through the LSTM-DAGMM model, the problem of high-dimensional data processing was solved, real-time anomaly detection of photovoltaic inverters was achieved, and maintenance costs were reduced.

CN120654146APending Publication Date: 2025-09-16CHINA NUCLEAR POWER OPERATION TECH CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510736439.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively process high-dimensional photovoltaic inverter data and capture its characteristics, resulting in a lack of real-time fault understanding in photovoltaic power generation system maintenance and high maintenance costs.

Method used

A LSTM-DAGMM-based method is used to collect historical monitoring data of photovoltaic inverters, pre-process and train the model to obtain the anomaly detection threshold and detect anomalies in real time.

Benefits of technology

It achieves accurate health status judgment of photovoltaic inverters, ensures normal operation of the system, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654146A_ABST
    Figure CN120654146A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of anomaly detection of photovoltaic power generation equipment, and aims to solve the problems of anomaly detection and state evaluation of an inverter of a photovoltaic power generation system and the problems that a machine learning model and a statistical method model cannot process high-dimensional data and cannot well capture data features. The invention discloses a photovoltaic inverter anomaly detection method based on LSTM-DAGMM. The method comprises the steps of collecting historical monitoring data of a photovoltaic inverter, preprocessing the collected historical monitoring data, training an LSTM-DAGMM model, obtaining an anomaly detection threshold value, performing anomaly detection on the photovoltaic inverter in real time, feeding back a result and updating the LSTM-DAGMM model. According to the invention, the health or fault condition of the photovoltaic inverter can be accurately and effectively judged, so that the normal operation of the photovoltaic power generation system can be ensured, and the maintenance and operation cost can be greatly saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of photovoltaic power generation equipment anomaly detection, and in particular relates to a photovoltaic inverter anomaly detection method based on LSTM-DAGMM. Background Art

[0002] Energy drives the progress of human civilization and is crucial to national security and livelihoods. It not only impacts human survival and development but also promotes economic and social development. Solar energy resources are abundant and distributed throughout the world. Photovoltaic power generation technology is an emerging renewable energy technology that offers the advantages of safety, cleanliness, and efficiency.

[0003] A photovoltaic power generation system primarily consists of three parts: the input section (including the photovoltaic array and DC combiner box), the inverter section (primarily the inverter), and the output section (primarily the AC distribution box). The photovoltaic inverter primarily converts the DC power generated by the photovoltaic array into AC power for output, while also protecting the circuit and maximizing the performance of the solar cells. Photovoltaic inverters are categorized as centralized inverters, string inverters, and distributed inverters. Each type of inverter has its own unique advantages: centralized inverters primarily benefit from large-scale applications; string inverters benefit from multi-channel MPPT controllers; and distributed inverters, which fully consider the advantages of centralization and clustering, decentralized control, and centralized grid connection, are widely used.

[0004] Anomaly detection is a key application in photovoltaic power station research. The operation and maintenance of photovoltaic inverters typically relies on scheduled, fixed-point maintenance and post-failure maintenance. This leaves maintenance personnel with little way of understanding the inverter's real-time fault status. Summary of the Invention

[0005] The main purpose of this application is to provide a photovoltaic inverter anomaly detection method based on LSTM-DAGMM to solve the anomaly detection and status assessment problems of photovoltaic power generation system inverters.

[0006] Another purpose of this application is to provide a photovoltaic inverter anomaly detection method based on LSTM-DAGMM, which aims to solve the problem that classic machine learning models and statistical method models cannot process high-dimensional data and fail to capture data features well, and provide guidance for the actual operation of photovoltaic stations.

[0007] In order to achieve the above objectives, this application provides the following technical solutions:

[0008] A photovoltaic inverter anomaly detection method based on LSTM-DAGMM, comprising:

[0009] Step 1: Collect historical monitoring data of photovoltaic inverters;

[0010] Step 2: Preprocess the collected historical monitoring data;

[0011] Step 3: Train the LSTM-DAGMM model to obtain the anomaly detection threshold;

[0012] Step 4: Real-time anomaly detection of photovoltaic inverters;

[0013] Step 5: Feedback the results and update the LSTM-DAGMM model.

[0014] In some embodiments, the photovoltaic inverter historical monitoring data includes three-phase voltage, three-phase current, line voltage, inverter conversion efficiency, inverter power factor, inverter chassis temperature, input power, and output power.

[0015] In some embodiments, step 2 includes cleaning of erroneous values, filling of missing data, and data normalization.

[0016] In some embodiments, a box plot algorithm is proposed to clean erroneous values.

[0017] In some embodiments, the step of identifying abnormal data based on the box plot includes:

[0018] Divide a well-ordered data sample into 4 parts, and obtain 3 data points Q1, Q2, Q3, etc. at each equal position;

[0019] Arrange a set of data in ascending order to obtain the sorted data sample X={x1,x2,......,x n}, calculate the median Q2:

[0020]

[0021] Calculate the lower quartile Q1 and upper quartile Q3:

[0022] When n = 2k (k = 1, 2, ...), the median Q2 divides the original data X into two parts, Q2 is not included in it, and the medians of these two parts are calculated according to the following formula, and Q1 < Q3;

[0023] When n=4k+3, k=0,1,2,..., then:

[0024]

[0025] When n=4k+1, k=0,1,2,..., we have:

[0026]

[0027] Get the interquartile range:

[0028] I QR =Q3-Q1

[0029] According to the interquartile range, the inner limit of abnormal data of the original data X can be determined:

[0030] [F1,F u ]=[Q1-λI QR ,Q3+λI QR ]

[0031] The interquartile range distributes the data into four parts and sorts them from low to high, with each part containing the same number of samples.

[0032] In some embodiments, a linear interpolation method is used to fill in missing data.

[0033] In some embodiments, Z-Score standardization is used to perform data normalization.

[0034] In some embodiments, the preprocessed historical monitoring data is used for training the LSTM-DAGMM model and is divided into a training data set and a test data set in a ratio of 70%:30%, respectively.

[0035] In some embodiments, step 3 includes data set partitioning, model training, and anomaly detection threshold acquisition.

[0036] In some embodiments, real-time monitoring data of the inverter is obtained, input into a trained LSTM-DAGMM model, and the model output is compared with an anomaly detection threshold; if the model output exceeds the threshold, it is identified as an anomaly; if the model output does not exceed the threshold, it is identified as normal.

[0037] Compared with the existing technology, the photovoltaic inverter anomaly detection method based on LSTM-DAGMM provided in this application has the following beneficial effects:

[0038] This application can accurately and effectively determine the health or fault status of the photovoltaic inverter to ensure the normal operation of the photovoltaic power generation system, while also significantly saving maintenance and operation costs.

[0039] Furthermore, we used specific data for calculations. For a certain inverter model at a certain electric field, we selected one year of historical data and trained an LSTM-DAGMM model after preprocessing. The trained model was used for real-time data anomaly detection. The model's anomaly detection performance over a week was analyzed, and a binary confusion matrix was constructed as follows:

[0040]

[0041] Based on this calculation, the accuracy is 0.9032, the recall is 0.9180, and the F1 score is 0.9105, which meets the actual engineering needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solution of this application, the following is a brief introduction to the drawings required for the technical description.

[0043] Figure 1 Flowchart of the photovoltaic inverter anomaly detection method based on LSTM-DAGMM provided in this application;

[0044] Figure 2 Data flow diagram of the photovoltaic inverter anomaly detection method based on LSTM-DAGMM provided in this application;

[0045] Figure 3 This is the LSTM-DAGMM model structure diagram. DETAILED DESCRIPTION

[0046] The following is further detailed description through specific implementation methods.

[0047] like Figure 1 and Figure 2 As shown, the present application provides a photovoltaic inverter anomaly detection method based on LSTM-DAGMM. The method combines a long short-term memory network (LSTM) and a deep autoencoder Gaussian mixture model (DAGMM) to improve the accuracy and efficiency of photovoltaic inverter anomaly detection. The method includes:

[0048] Step 1: Collect historical monitoring data of PV inverters.

[0049] Key monitoring parameters for PV inverters include three-phase voltage, three-phase current, line voltage, inverter conversion efficiency, inverter power factor, inverter chassis temperature, and inverter input and output power. Based on actual site conditions, collect at least one year of historical inverter monitoring data to ensure comprehensiveness and representativeness.

[0050] Specifically, these data are:

[0051] Three-phase voltage: reflects the voltage at the input end of the photovoltaic inverter;

[0052] Three-phase current: reflects the current at the input end of the photovoltaic inverter;

[0053] Line voltage: the voltage between the internal lines of the photovoltaic inverter;

[0054] Inverter conversion efficiency: measures the inverter's ability to convert DC power to AC power;

[0055] Inverter power factor: reflects the phase relationship between the inverter output current and voltage;

[0056] Chassis temperature: monitor the internal temperature of the inverter to prevent overheating;

[0057] Input power and output power: represent the electrical energy received and output by the inverter respectively.

[0058] Step 2: Historical data preprocessing.

[0059] Data preprocessing involves cleaning, filling in missing values, and normalizing complex, chaotic data with missing values ​​to make it easier to analyze and process. This application proposes to use a box plot algorithm to clean erroneous values, a linear interpolation method to fill in missing data, and Z-score standardization for data normalization.

[0060] Specifically, it includes the following:

[0061] Error value cleaning: Use the box plot algorithm to identify and clean outliers to ensure data accuracy and prevent outliers from interfering with subsequent analysis. The box plot is an effective statistical tool for identifying outliers in a data set.

[0062] Missing data filling: linear interpolation method is used to fill missing data to maintain data continuity and avoid the adverse effects of missing values ​​on subsequent analysis;

[0063] Data normalization: Z-score standardization method was used to convert the data into a standard normal distribution for subsequent analysis. Data normalization was used to convert the data into the same scale for subsequent analysis.

[0064] Step 3: Train the LSTM-DAGMM model to obtain the anomaly detection threshold.

[0065] The preprocessed historical data can be used to train the LSTM-DAGMM model. It is divided into a training dataset and a test dataset in a 70%:30% ratio. The training dataset is used to train the LSTM-DAGMM model, while the test dataset is used to test the model's effectiveness and obtain anomaly detection thresholds.

[0066] Specifically, it includes the following:

[0067] Dataset partitioning: The preprocessed historical data is divided into a training dataset and a test dataset in a ratio of 70%:30%. The purpose of the 70%:30% partitioning is to ensure that the training set has enough data for the model to learn, and the test set has enough data to evaluate the model performance.

[0068] Model training: Use the training dataset to train the LSTM-DAGMM model. This includes: inputting the training dataset into the LSTM-DAGMM model, which iteratively learns based on the input data and continuously adjusts its internal parameters (such as weights and biases) to minimize the loss function. During the training process, optimization algorithms can be used to accelerate training and improve model performance.

[0069] Anomaly detection threshold acquisition: Test the model performance using a test data set and determine a reasonable anomaly detection threshold. This threshold is used to determine whether the data is abnormal in subsequent real-time anomaly detection. Specifically, it includes:

[0070] Run the trained LSTM-DAGMM model on the test dataset and calculate the anomaly score (or reconstruction error, probability density, etc.) for each data point;

[0071] According to the distribution of anomaly scores, an appropriate threshold is selected to distinguish normal data from abnormal data.

[0072] Step 4: Real-time anomaly detection of PV inverters.

[0073] Obtain real-time monitoring data from the inverter and input it into the trained LSTM-DAGMM model. Compare the model output to the anomaly detection threshold. If the model output exceeds the threshold, it is identified as an anomaly; if it does not, it is identified as normal.

[0074] Data acquisition: Real-time collection of monitoring data of photovoltaic inverters; monitoring data includes voltage, current, power, temperature, etc.;

[0075] Model input: input the collected data into the LSTM-DAGMM model;

[0076] Anomaly judgment: Compare the model output with the anomaly detection threshold to determine whether the data is abnormal.

[0077] Before inputting the collected data into the LSTM-DAGMM model, data preprocessing is required. Data preprocessing includes data cleaning (noise removal, missing value filling, etc.), data transformation (such as normalization, standardization, etc.), and data feature extraction to improve the accuracy and stability of the model.

[0078] In anomaly detection, after processing input data, the LSTM-DAGMM model outputs an anomaly score or probability value to indicate the degree of anomaly. A higher anomaly score indicates a higher probability of anomaly. The anomaly detection threshold is determined based on the model's training results and test data to distinguish between normal and abnormal data. The model output is compared with the anomaly detection threshold. If the model output exceeds the threshold, the data is considered anomaly; otherwise, the data is considered normal. The anomaly detection result can trigger appropriate alarms or processing mechanisms, for example, to promptly respond to and handle abnormal events.

[0079] Step 5: Result feedback and model update.

[0080] On-site operation and maintenance personnel periodically collect statistics on the warning results of the anomaly detection method, analyze the warning effect of the method, and conduct model update training if necessary. Specifically including:

[0081] Result feedback: Feedback the abnormality detection results to the on-site operation and maintenance personnel so that timely measures can be taken to deal with the abnormality;

[0082] Model update: On-site operation and maintenance personnel analyze the warning effects and methods and conduct model update training if necessary to improve the accuracy and efficiency of detection.

[0083] Example

[0084] like Figure 2 As shown in the figure, in this embodiment, the photovoltaic inverter anomaly detection method based on LSTM-DAGMM mainly consists of five steps: inverter historical monitoring data collection, historical data preprocessing, LSTM-DAGMM model training and anomaly detection threshold acquisition, inverter real-time anomaly detection, result feedback and model update.

[0085] In this embodiment, the photovoltaic inverter error value cleaning adopts the box plot algorithm. The abnormal data identification steps based on the box plot are as follows:

[0086] (1) Divide a well-ordered data sample into four equal parts. Thus, we obtain three data points Q1, Q2, Q3, etc. at each equal position.

[0087] (2) Arrange a set of data in ascending order to obtain the sorted data sample X = {x1, x2, ..., x n}, calculate the median Q2:

[0088]

[0089] (3) Calculate the lower quartile Q1 and the upper quartile Q3.

[0090] When n=2k (k=1,2,...), the median Q2 divides the original data X into two parts, Q2 is not included in it. According to the following formula, the medians of these two parts are calculated respectively, namely Q1 and Q3, and Q1<Q3

[0091] When n=4k+3, k=0,1,2,..., then:

[0092]

[0093] When n=4k+1, k=0,1,2,..., we have:

[0094]

[0095] Then, we can get the interquartile range:

[0096] I QR =Q3-Q1

[0097] According to the interquartile range, the inner limit of abnormal data of the original data X can be determined:

[0098] [F1,F u ]=[Q1-λI QR ,Q3+λI QR ]

[0099] The interquartile range distributes the data into four parts and sorts them from low to high, with each part containing the same number of samples. The first quartile (Q1) is the value of the data point on the boundary. The same applies to Q2 and Q3. The interquartile range (IQR) is the data point in the middle of the two parts (representing 50% of the data). The interquartile range includes all data points above Q1 and below Q3. If the point is above Q3 + (1.5 x IQR), it indicates that there is a high-value outlier. If it is Q1 - (1.5 x IQR), there is a low-value outlier.

[0100] In this embodiment, the photovoltaic inverter data is normalized using Z-Score standardization. After Z-Score processing, the average value of the data is equal to 0, and the standard deviation of the data is equal to 1. The formula is as follows:

[0101]

[0102] Among them, x * is the mean value of all data of the PV inverter before normalization, and σ is the standard deviation of all data of the inverter before normalization.

[0103] In this embodiment, the LSTM-DAGMM model is used to fit the normal behavior of the photovoltaic inverter. The LSTM-DAGMM model structure is as follows: Figure 3As shown in the figure, it mainly consists of two parts: the compression network and the evaluation network. The compression network is used to extract high-dimensional data features, and the evaluation network is used to obtain data feature indicators. The compression network reduces the dimensionality of the input sample through the autoencoder, extracts low-dimensional representations and reconstructs error features. The evaluation network uses these features to predict the possibility of abnormality of the sample under the framework of the Gaussian mixture model. By comparing the energy of the sample with the preset threshold, it can be determined whether the sample is abnormal. Among them, the features of the compression network include two parts: the low-dimensional feature z learned by the deep autoencoder and the low-dimensional feature z. c and the low-dimensional features z of the reconstruction error r , then form z, provide it to the subsequent evaluation network, and finally get the output value of the model through multiple layers of full connection It contains the category probability after softmax

[0104]

[0105] Where, is the mixing probability of the kth individual in the model, is the mean of the kth individual, is the variance of the kth individual, E(z) is the energy value of the sample;

[0106] After obtaining the output of the model, the current energy can be obtained according to the multivariate Gaussian probability density related formula and the energy evaluation formula shown, and high-energy samples can be predicted as abnormal through the pre-selected threshold. The DAGMM model training loss function is:

[0107]

[0108] Wherein, the first term is the reconstruction error of the autoencoder network, the second term is the likelihood function of the Gaussian mixture model, the third term is the sample energy value, and the fourth term is to prevent the diagonal value of the covariance matrix from being 0.

[0109] In this embodiment, in order to determine the threshold of the monitoring index, it is necessary to perform statistical analysis on the monitoring index of the normal sample data. This embodiment adopts the non-parametric estimation method kernel density estimation to analyze the probability density distribution of the monitoring index of the normal sample data according to the characteristics of the data itself. Assume that the monitoring index {t k The true probability density function of} is, then the fixed bandwidth probability density function based on kernel density estimation is:

[0110]

[0111] Among them, t kis the monitoring index of the kth sample data point in the sample data, h is the bandwidth parameter, the number of sample data n is m-β+1, and K(·) is the kernel function. The essence of the kernel function is a weight function, which represents the contribution of each sample point to the density. In this embodiment, the Gaussian function is used as the kernel function for kernel density estimation:

[0112]

[0113] According to the probability density function of the detection index of the normal operation data of the inverter, the threshold value of the inverter abnormality detection under the confidence level is

[0114]

[0115] In this embodiment, the operation and maintenance personnel periodically count the warning results of the anomaly detection method and analyze the warning effect of the proposed method by counting the number of correct alarms, the number of false alarms, and the number of missed alarms. In this embodiment, abnormal samples are considered as positive classes and normal samples as negative classes, and a confusion matrix for the binary classification problem is constructed. This embodiment uses three indicators based on the confusion matrix: precision, recall, and F1 score to evaluate the performance of the time series data anomaly detection model:

[0116]

[0117] Among them, TP (True Positive) represents samples that are actually positive and correctly detected as positive; TN (True Negative) represents samples that are actually negative and correctly detected as negative; FP (False Positive) represents samples that are actually negative but mistakenly detected as positive; FN (False Negative) represents samples that are actually positive but mistakenly detected as negative.

[0118] The above description is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed in this application should be covered by the scope of protection of the present application.

Claims

1. A photovoltaic inverter anomaly detection method based on LSTM-DAGMM, characterized in that: include: Step 1: Collect historical monitoring data of photovoltaic inverters; Step 2: Preprocess the collected historical monitoring data; Step 3: Train the LSTM-DAGMM model to obtain the anomaly detection threshold; Step 4: Real-time anomaly detection of photovoltaic inverters; Step 5: Feedback the results and update the LSTM-DAGMM model.

2. The photovoltaic inverter anomaly detection method based on LSTM-DAGMM according to claim 1 is characterized in that: In step 1, the historical monitoring data of the photovoltaic inverter includes three-phase voltage, three-phase current, line voltage, inverter conversion efficiency, inverter power factor, inverter chassis temperature, input power and output power.

3. The photovoltaic inverter anomaly detection method based on LSTM-DAGMM according to claim 1 is characterized in that: Step 2 includes cleaning of erroneous values, filling of missing data, and data normalization.

4. The photovoltaic inverter anomaly detection method based on LSTM-DAGMM according to claim 3 is characterized in that: The box plot algorithm is used to clean up the erroneous values.

5. The photovoltaic inverter anomaly detection method based on LSTM-DAGMM according to claim 4 is characterized in that: The steps for identifying abnormal data based on box plots include: Divide a well-ordered data sample into 4 parts, and obtain 3 data points Q1, Q2, Q3, etc. at each equal position; Arrange a set of data in ascending order to obtain the sorted data sample X={x1,x2,......,x n }, calculate the median Q2: Calculate the lower quartile Q1 and upper quartile Q3: When n = 2k (k = 1, 2, ...), the median Q2 divides the original data X into two parts, Q2 is not included in it, and the medians of these two parts are calculated according to the following formula, and Q1 < Q3; When n=4k+3, k=0,1,2,..., then: When n=4k+1, k=0,1,2,..., we have: Get the interquartile range: I QR =Q3-Q1 According to the interquartile range, the inner limit of abnormal data of the original data X can be determined: [F1,F u ]=[Q1-λI QR ,Q3+λI QR ] The interquartile range distributes the data into four parts and sorts them from low to high, with each part containing the same number of samples.

6. The photovoltaic inverter anomaly detection method based on LSTM-DAGMM according to claim 3 is characterized in that: Missing data were filled using linear interpolation method.

7. The photovoltaic inverter anomaly detection method based on LSTM-DAGMM according to claim 3 is characterized in that: Z-Score standardization was used for data normalization.

8. The photovoltaic inverter anomaly detection method based on LSTM-DAGMM according to claim 1 is characterized in that: In step 3, the preprocessed historical monitoring data is used to train the LSTM-DAGMM model and is divided into a training data set and a test data set in a ratio of 70%:30%, respectively.

9. The photovoltaic inverter anomaly detection method based on LSTM-DAGMM according to claim 1, characterized in that: Step 3 includes dataset partitioning, model training, and anomaly detection threshold acquisition.

10. The photovoltaic inverter anomaly detection method based on LSTM-DAGMM according to claim 1, characterized in that: In step 4, the real-time monitoring data of the inverter is obtained and input into the trained LSTM-DAGMM model. The model output is compared with the anomaly detection threshold. If the model output exceeds the threshold, it is identified as an anomaly; if the model output does not exceed the threshold, it is identified as normal.

Citation Information

Patent Citations

  • Photovoltaic inverter anomaly detection method based on DAGMM network model

    CN117473341A

  • Photovoltaic station data interpolation method and system based on sunset law

    CN119448224A