Photovoltaic array small sample fault diagnosis method based on neural network

The WGAN-GP neural network generates false samples to expand the photovoltaic array fault data set, and combines the LSTM neural network to solve the problem of insufficient photovoltaic array fault samples, achieving efficient fault diagnosis and accurate identification.

CN120354130APending Publication Date: 2025-07-22WUXI LONGMAX TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510424357.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Insufficient photovoltaic array fault samples lead to low diagnostic accuracy or overfitting of neural networks, and it is difficult for the existing technology to efficiently collect sufficient fault data, increasing time and labor costs.

Method used

The WGAN-GP neural network is used to generate high-quality false samples, expand the photovoltaic array fault data set, and combine it with the LSTM neural network for fault diagnosis.

Benefits of technology

The photovoltaic array fault data set is effectively expanded, the accuracy and efficiency of fault diagnosis is improved, the data collection cost is reduced, and high-performance diagnosis is ensured under small sample conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354130A_ABST
    Figure CN120354130A_ABST
Patent Text Reader

Abstract

The invention discloses a photovoltaic array small sample fault diagnosis method based on a WGAN-GP neural network. The method comprises the following steps: firstly, extracting data samples of a photovoltaic array under different environmental parameters and preprocessing the data samples; training a WGAN-GP neural network by using a small number of fault data samples, and evaluating a training effect of the WGAN-GP by using a density index; after all fault types obtain a generator network with good performance, a large amount of false data is generated and mixed with real data, and then the LSTM neural network is used to carry out fault diagnosis on the photovoltaic array. According to the method, the false samples with relatively good quality can be generated, and the small sample data set is expanded into the data set capable of supporting the neural network to complete training, so that the cost required for collecting the fault data of the photovoltaic array is saved, and the LSTM neural network can be applied to a photovoltaic power station lacking enough fault samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of photovoltaic power generation array fault diagnosis, and specifically to a method for using a WGAN-GP neural network to augment photovoltaic array fault samples and achieve small-sample fault diagnosis. Background Art

[0002] As the core component of a photovoltaic power generation system, the photovoltaic array is prone to various faults due to being exposed outdoors all year round and being affected by changing environments, which affects its service life. At the same time, faults will directly lead to a reduction in the power generation of the photovoltaic power station, irreversible damage to the components, and even cause fires in severe cases.

[0003] Generally, protection equipment is installed on the DC side of the photovoltaic array. However, when light or moderate faults occur or when faults occur under low irradiance conditions, traditional protection equipment cannot act in time due to maximum power point tracking and too low fault current. Therefore, researchers have proposed using a long short-term memory (LSTM) neural network for real-time fault diagnosis of photovoltaic arrays. LSTM is a time-recursive neural network that can learn long-term dependencies in time series, and there are loops inside the model to maintain the continuity of information.

[0004] Constructing an LSTM photovoltaic array fault diagnosis model can be divided into three stages: the training stage, the testing stage, and the actual operation stage. Among them, the LSTM neural network in the training stage can accurately identify faults by learning a large number of fault samples. However, when the number of fault samples is insufficient, the neural network may show low diagnostic accuracy in the testing stage, or overfitting may occur in the training stage, resulting in the model being unable to complete training.

[0005] In real-world photovoltaic power stations, there are strict operation and maintenance management systems and a large number of protection devices. With the help of these measures, the fault incidence rate of the photovoltaic array is relatively low. Therefore, it takes a large amount of time and labor costs to collect sufficient fault samples. Summary of the Invention

[0006] Aiming at the problem of insufficient fault samples, the present invention provides a method for using a WGAN-GP neural network to augment a small number of photovoltaic array fault samples. The WGAN-GP neural network used in the present invention is a generative adversarial network (GAN) improved by using the Wasserstein distance and the gradient penalty term (GP), and includes two parts: a generator network and a discriminator network. After training, the generator network can generate high-quality fake samples, so as to achieve the fault diagnosis function with few samples.

[0007] The present invention adopts the following technical solution. The small-sample fault diagnosis method for a photovoltaic array based on a neural network includes the following steps:

[0008] Step 1: Acquisition and processing of photovoltaic array output data:

[0009] Sampling is carried out during the lighting time. The sampling time is B hours, the sampling frequency is 1 time per second, and the number of sampling days is A. That is, the collected data is A groups of time series with a length of 3600*B time points. One time point represents the output data of the photovoltaic array per second; the data samples contain 4 characteristic parameters, namely: open-circuit voltage, short-circuit current, maximum power point voltage, and maximum power point current; the 6 working states of the photovoltaic array are included in the A groups of time series, namely: normal operation, open-circuit fault, short-circuit fault, aging fault, shadow shielding fault, and fault with both short-circuit and shadow occurring simultaneously;

[0010] A sliding window with a length of C time points is made, and 3600*B*A / C data samples are intercepted from the A groups of time series, where A>10, B>10, and C>200; the fault type at the last time point of the data sample is selected as the sample label, and the original data samples of various working states are finally obtained. Use to represent the original sample set of normal operation, to represent the original sample set of open-circuit fault, to represent the original sample set of short-circuit fault, to represent the original sample set of aging fault, to represent the original sample set of shadow shielding, to represent the original sample set of both short-circuit and shadow occurring simultaneously;

[0011] The original data samples are normalized, and the parameters at each time point in the time series are normalized to the range of [-1,1]. The set is correspondingly transformed into sets X1 to X6;

[0012] Next, in order to obtain the training data set and validation data set required during the training of the neural network, a certain number of samples X n (n = 1,..., 6) are extracted from X, n ′, and they are respectively composed of the training data set T = [X1′...X n ′] and the validation data set V = [X1 - X1′...X n - X n ′];

[0013] Step 2: Train the WGAN-GP neural network and generate false samples to expand the data set:

[0014] When a certain type of data set X nWhen the number of samples of X' is small, use the WGAN-GP neural network to learn the sample distribution of X', so as to obtain a generator network that can transform noise into available data samples; after obtaining the corresponding generator network from X', generate a sufficient number of fake samples with labels and mix them with the real samples X' to obtain a mixed dataset X''. And the set of all mixed datasets X'' and the dataset X' with sufficient samples is the fault diagnosis dataset, denoted as T; n The objective function L of the WGAN-GP neural network is as follows: n where x is the data sample after normalization in step one, is the real data, is the generated data, P and P respectively represent the probability distributions of the real data and the generated data, E[…] represents the mathematical expectation, x~P means that x follows the P distribution, means that follows the P distribution, D(·) is the output of the discriminator, the gradient penalty term λ is the penalty coefficient, is the random interpolation sampling of the generated sample and the real sample, ε represents a random value in the interval [0,1), and is the gradient of the discriminator; n Step three: Train the LSTM neural network and perform fault diagnosis: n Use the LSTM network layer, Dropout network layer and two fully connected layers to form a photovoltaic array fault diagnosis model. Set the input feature dimension of the LSTM in the model to 4, the feature dimension of the hidden layer to 64, and the output feature dimensions of the fully connected layers to 32 and 6 respectively. After setting the model structure and parameters, use the fault diagnosis dataset T obtained in step two as a new training set and send it into the model for training; save the model with the smallest validation loss value during the iteration process, and finally obtain a neural network that can be used for photovoltaic array fault diagnosis. n Specifically, in step one, continuous sampling is carried out from 6 am to 6 pm, and the sampling time is 12 hours. n ; new ;

[0015] The objective function L of the WGAN-GP neural network is: GP For:

[0016]

[0017] Among them, x is the data sample after normalization in step one, is the real data, is the generated data, P r and P g respectively represent the probability distributions of the real data and the generated data, E[…] represents the mathematical expectation, x~P r means that x follows P r distribution, means follows P g distribution, D(·) is the output of the discriminator, the gradient penalty term λ is the penalty coefficient, is the random interpolation sampling of the generated sample and the real sample, ε represents a random value in the interval [0,1), is the gradient of the discriminator;

[0018] Step three: Train the LSTM neural network and perform fault diagnosis:

[0019] Use the LSTM network layer, Dropout network layer and two fully connected layers to form a photovoltaic array fault diagnosis model. Set the input feature dimension of the LSTM in the model to 4, the feature dimension of the hidden layer to 64, and the output feature dimensions of the fully connected layers to 32 and 6 respectively. After setting the model structure and parameters, use the fault diagnosis dataset T obtained in step two new as a new training set and send it into the model for training; save the model with the smallest validation loss value during the iteration process, and finally obtain a neural network that can be used for photovoltaic array fault diagnosis.

[0020] Specifically, in step one, continuous sampling is carried out from 6 am to 6 pm, and the sampling time is 12 hours.

[0021] Specifically, in Step 2, the WGAN-GP neural network is used to learn the sample distribution of X n ′, and finally a generator network that can transform noise into available data samples is obtained. The specific steps include:

[0022] Step 2.1: Set the number of iterations, sample batches, and parameters of the optimization algorithm required for training WGAN-GP;

[0023] Step 2.2: Exclude samples in the normal working state, and select a type of fault sample for training the WGAN-GP neural network;

[0024] Step 2.3: Generate the same number of Gaussian white noises z according to the number of samples in one training batch;

[0025] Step 2.4: The noise passes through the generator network to generate as fake samples;

[0026] Step 2.5: Mix the input samples and fake samples in the same batch, send them into the discriminant model to obtain the loss value, and use this as a reference to update the discriminator network;

[0027] Step 2.6: Generate fake samples again Use the updated discriminant model to judge the samples, output the loss value; update the generator network;

[0028] Step 2.7: Use two indicators, the density indicator and the generator loss value indicator, to judge whether the sample generator network meets the standard. If it meets the standard, save the model;

[0029] Step 2.8: Repeat Step 2.2 to Step 2.7 to obtain the corresponding generator network for each type of fault.

[0030] Specifically, in Step 2.7, the density indicator density uses the neighborhood sphere of real samples to calculate whether the fake samples meet the standard, and can evaluate whether the generated data is the same type of fault data. Its calculation formula is:

[0031]

[0032] where k represents the proximity distance, M represents the number of fake samples, N represents the number of real samples, 1... represents the indicator function, represents the generated fake samples, B(x, NND k (X n ′)) represents the spherical region with a radius of NND k (X n ′) around the real data x, X nLet \(X'\) be a set composed of a certain type of real data \(x\), NND k (X' n ) represents the \(k\) - th nearest - neighbor distance from a single real data \(x\) to other real data \(X'\) of the same type n in the set

[0033] The advantages of the present invention are as follows:

[0034] The present invention uses a WGAN - GP neural network to generate characteristic parameters of a photovoltaic array in different faults, which can save the labor and time costs required for collecting fault data, and solves the problem of insufficient fault data of the current photovoltaic array. Using a density index to evaluate the performance of the generator can ensure that the fault samples generated by the model have good quality, enabling the fault diagnosis model to exhibit good fault diagnosis accuracy in the test set

[0035] The present invention can generate false samples with good quality, expand a small - sample data set into a data set that can support a neural network to complete training, enabling the LSTM neural network to be applied to a photovoltaic power station lacking sufficient fault samples BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic diagram of data acquisition of a photovoltaic array

[0037] Figure 2 is a schematic diagram of the operation principle of a WGAN - GP neural network

[0038] Figure 3 is a structural diagram of a WGAN - GP neural network

[0039] Figure 4 is the overall flow chart of the present invention

[0040] Figure 5 is a structural diagram of an LSTM neural network

[0041] Figure 6 is a curve graph of the loss value of training an LSTM model using small - sample data

[0042] Figure 7 is a curve graph of the loss value of training an LSTM model using mixed data DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The present invention will be further described in detail below in combination with the drawings and specific embodiments, so that the present invention can be more easily understood by those skilled in the art

[0044] As Figure 4 shown, the implementation steps of the present invention are as follows

[0045] Step 1: Data acquisition and processing

[0046] 1. Data acquisition.

[0047] In the embodiment, the photovoltaic array data comes from a photovoltaic power station with an installed capacity of 5 GW. The sampling time is set from 6:00 am to 6:00 pm, and the sampling frequency is 1 second / time. The sensor is used to collect the output data of the photovoltaic array for 12 days. That is, the original data is 12 time series with a length of 43,200 time points. One time point represents the output data of the photovoltaic array for one second. There are 4 characteristic parameters in the sample, namely: open circuit voltage, short circuit current, maximum power point voltage, and maximum power point current. The schematic diagram of data acquisition is as Figure 1 shown.

[0048] These 12 time series include but are not limited to 6 working states of the photovoltaic array, namely: normal operation, open circuit fault, short circuit fault, aging fault, shadow shielding, and the fault of short circuit and shadow occurring simultaneously; if new fault samples appear in subsequent implementations, they will also be marked according to the same structure. In the following embodiments, the above 6 working states are considered for illustration.

[0049] Taking five minutes as the standard, a sliding window with a length of 300 time points is made to segment the time series, and 1,728 data samples are intercepted from the 12 time series; the working state (fault type) at the last time point of each data sample is selected as the sample label. Finally, the data samples of various working states (including normal working state and fault working state) are as follows:

[0050]

[0051] Among them, represents the original data sample with a length of 300 time points, represents the original sample set of normal operation, represents the original sample set of open circuit fault, represents the original sample set of short circuit fault, represents the original sample set of aging fault, represents the original sample set of shadow shielding, represents the original sample set of short circuit and shadow occurring simultaneously.

[0052] 2. Data preprocessing.

[0053] The original data sample is normalized, and the parameters at each time point within are normalized to the range of [-1, 1]. The formula for data normalization is as follows:

[0054]

[0055] Among them, represents the original data sample, and x represents the new data sample after normalization. represents the maximum value among all data samples. represents the minimum value among all data samples. represents the average value among all data samples.

[0056] After normalization, the data samples of various working states are as follows:

[0057]

[0058] Among them, x a-b (a = 1, …, 6; b = 1, …, 1586) represents the data sample with 300 time points in length. X1 represents the sample set of normal operation after normalization, X2 represents the sample set of open - circuit fault after normalization, X3 represents the sample set of short - circuit fault after normalization, X4 represents the sample set of aging fault after normalization, X5 represents the sample set of shadow - shielding fault after normalization, and X6 represents the sample set of the co - occurrence of short - circuit and shadow after normalization.

[0059] 3. Divide the training set and the validation set.

[0060] It can be seen that the number of normal operation samples of the photovoltaic array is 1586, while the total number of 5 types of fault samples is only 142. And among all the fault samples, the number of fault samples with the co - occurrence of short - circuit and shadow is the least, only 17.

[0061] When making the training set, considering that a part of the samples of each working state needs to be left for comparative verification, and to prevent training errors caused by unbalanced sample numbers, the number of various fault samples is preferably kept the same. Therefore, 12 samples are selected for each type of fault.

[0062] Considering that the number of normal samples is much larger than that of fault samples, and the number of various samples should be as balanced as possible, 60 normal samples and 60 fault samples are randomly selected to form a small - sample training set, and its expression formula is as follows:

[0063]

[0064] Among them, X1′ represents the sample set composed of 60 normal operation samples, X i ′(i = 2, …, 6) represents the sample set composed of 12 fault samples, and * represents the non - existence of the sample.

[0065] Then all the remaining samples of the 6 - type data are composed into the validation set V to reflect the state of the model during the training process.

[0066] Step 2: Train the WGAN - GP neural network and generate a fault diagnosis data set.

[0067] When the number of a certain type of dataset X n ' in the original training set T is small, the present invention uses a WGAN-GP neural network to learn the sample distribution of X n ', so as to obtain a generator network that can transform noise into available data samples; after obtaining the corresponding generator network for X n ', generate a sufficient number of fake samples with labels, and mix them with the real samples X n ' to obtain a mixed dataset X″ n , and the set T n of all the mixed datasets X″ n and the dataset X′ new with sufficient sample quantity is the fault diagnosis dataset.

[0068] Prepare a WGAN-GP neural network, and its working principle is as Figure 2 shown. Except for the loss function, the working principle of WGAN-GP is the same as that of the traditional generative adversarial network, both using noise to generate fake samples, and judging whether the input samples are real samples or fake samples in the discriminator, and updating the network parameters according to the results.

[0069] The structures of the generator and discriminator of the WGAN-GP neural network are both composed of 5-layer one-dimensional convolutional neural networks, as Figure 3 shown. The convolutional kernels of the one-dimensional convolutional neural network only perform convolution along the time step sequence, which can effectively extract the change features of the time series. At the same time, the parameters of the generator and discriminator are the same, ensuring the consistent performance of the generator and discriminator, and preventing training failure due to the overly strong performance of one side.

[0070] The activation function of the last layer of the generator network is the hyperbolic tangent function, which can activate the eigenvalue extracted by the convolutional layer to between [-1, 1], consistent with the upper and lower limits of the normalized values.

[0071] The objective function L 2P of the WGAN-GP neural network adopted by the present invention is:

[0072]

[0073] where x is the real data, is the generated data, P r and P g respectively represent the probability distributions of the real data and the generated data, E[…] represents the mathematical expectation, x~P r means that x follows the P r distribution, represents follows P gDistribution, D(·) is the output of the discriminator D, and the gradient penalty term λ is the penalty coefficient, is the random interpolation sampling of the generated samples and the real samples, ε represents a random value within the interval [0, 1), is the gradient of the discriminator D.

[0074] Use the above WGAN-GP neural network to learn the sample data distribution of the dataset X n ' with a small number of samples, and finally obtain a generator network that can transform noise into available data samples. The specific steps include:

[0075] Step 2.1: Set the number of iterations, sample batches, parameters of the optimization algorithm, etc. required when training WGAN-GP.

[0076] In the embodiment, the number of iterations of the model is set to 3000, the training batch is 4 samples each time, the Adam optimization algorithm is selected, and its parameters are set as learning rate = 0.0001, beta1 = 0, beta2 = 0.9, and other parameters are default values.

[0077] Step 2.2: Without considering the samples in the normal working state, take a fault sample dataset with a small number of samples of one type as the input sample for the training of WGAN-GP;

[0078] Step 2.3: Generate the same number of Gaussian white noises z according to the number of samples in one training batch;

[0079] Step 2.4: The noise passes through the generator network to generate false samples:

[0080] Step 2.5: Mix the input samples and false samples of the same batch, send them into the discriminant model, obtain the loss value, and use this as a reference to update the discriminator network;

[0081] Step 2.6: Generate false samples again Use the updated discriminant model to judge the samples, output the loss value; update the generator network;

[0082] Step 2.7: Use two indicators, density and generator loss value, to judge whether the sample generator network meets the standard. If it meets the standard, save the model;

[0083] Step 2.8: Repeat steps 2.2 to 2.7 to obtain the corresponding generator network for each type of fault. In the embodiment, after the sample sets of all types of fault data are trained in the WGAN-GP network, 5 generator networks corresponding to the photovoltaic array fault types are obtained.

[0084] Then, using the trained model, 100 false samples are generated for each type of fault. Then, calculate the density values of the false samples and the true samples of the same type of fault. Set the proximity distance k = 5, the number of false samples M = 100, and the number of true samples N = 12. The density index (density) formula used in the present invention is as follows:

[0085]

[0086] The density index (density) uses the neighborhood sphere of the true samples to calculate whether the false samples meet the standard, and can evaluate whether the generated data is data of the same type of fault. Among them, k represents the proximity distance, M represents the number of false samples, N represents the number of true samples, 1… represents the index function. represents the false sample, B(x, NND k (X′ i )) represents the spherical region with a radius of NND k (X′ i ) around the true data x, X i ′ is the set composed of a certain type of true sample x, and NND k (X i ′) represents the kth nearest neighbor distance from a single true data x to other true data X i ′ of the same type.

[0087] Substitute the numerical values of the embodiments:

[0088]

[0089] Set the density threshold to 0.9. When the false samples of a certain type of fault cannot reach the threshold, retrain the WGAN-GP neural network of this type of fault. After obtaining a new generator network, perform density calculation again until all fault samples have obtained a qualified generator network.

[0090] The final density values of each type of generator network are shown in Table 1.

[0091] Table 1

[0092]

[0093] After obtaining the corresponding generator network for all types of faults, generate sufficient false samples with clear labels, and mix them with the true samples to obtain a mixed sample, then a fault diagnosis data set with sufficient sample quantity can be obtained.

[0094] In the embodiment, 100 false samples are generated for each type of fault. After making the labels, the false samples and the original samples are combined into a training set, and its expression formula is as follows:

[0095]

[0096] Among them, X′1 represents a sample set composed of 112 normally operating samples, and X″ i represents a mixed sample set composed of 12 faulty samples and 100 false samples.

[0097] Step 3: Train the LSTM neural network and perform fault diagnosis on the photovoltaic array.

[0098] The long short-term memory neural network (LSTM) is a special type of recurrent neural network (RNN). The LSTM neural network model is a time-recurrent neural network that can learn the long-term dependencies of time series, and there are loops inside the model to maintain the continuity of information.

[0099] By studying the output data of the photovoltaic array, it can be found that there are many time points in this data and they are related to each other in chronological order. Therefore, LSTM is used as the fault diagnosis model for the photovoltaic array. The fault diagnosis model for the photovoltaic array in the present invention is composed of an LSTM network layer, a Dropout network layer, and two fully connected layers.

[0100] The architecture and parameters of the photovoltaic array fault diagnosis model are as Figure 5 shown. The model uses the LSTM neural network to extract the data features of the photovoltaic array, and integrates the features through the Dropout and fully connected layers to achieve fault diagnosis. Since the input of LSTM is related to the sample variables and the output is related to the hidden layer parameters, and there are 4 parameter variables and 6 working states in the data samples used in the present invention. Therefore, the input feature dimension of LSTM is set to 4, the feature dimension of the hidden layer is 64, Dropout = 0.25, the output feature dimensions of the two fully connected layers are 32 and 6 respectively, the number of iterations of the model is 600, the training batch is 64 samples each time, the validation batch is 32 samples each time, the optimization algorithm selects Adam, and its parameters are set as the learning rate = 0.0001, and other parameters are default values.

[0101] After determining the model structure and parameters, the mixed data set T obtained in Step 2 new is used as the training set, and the validation set V is obtained in Step 1. Then, these two data sets are fed into the model for training. Saving the model with the smallest validation loss value during the iteration process can obtain a neural network that can be used for fault diagnosis of the photovoltaic array.

[0102] The effectiveness of the method of the present invention is proved by analyzing the training loss value of the model and the test accuracy in the software.

[0103] First, use the small sample training set T and the validation set V to train the LSTM neural network, and its loss value curve is asFigure 6 As shown. Obviously, under the condition of small samples, the validation loss value cannot converge, and the model shows overfitting phenomenon.

[0104] Then, the new training set T new and the validation set V are input into the LSTM fault diagnosis model for training. The loss value curve obtained by the model is as Figure 7 shown. It can be seen from the trend of the loss values of the training set and the validation set in the figure that the mixed samples can help the photovoltaic array fault diagnosis model complete the training smoothly, avoiding the overfitting phenomenon that may be caused by the small sample training set.

[0105] The trained photovoltaic array fault diagnosis model is imported into the platform software, run for 10 days and record its diagnosis results, as shown in Table 2. It can be seen from the diagnosis results that the model has a diagnosis accuracy rate of more than 95% for all working states of the photovoltaic array. The present invention can help the photovoltaic power station obtain a fault diagnosis model with high performance under the condition of small samples.

[0106] Table 2

[0107]

Claims

1. A small-sample fault diagnosis method for a photovoltaic array based on a neural network, characterized in that Including the following steps: Step 1, acquisition and processing of photovoltaic array output data: Sampling is carried out during the lighting time. The sampling time is B hours, the sampling frequency is 1 time per second, and the number of sampling days is A. That is, the acquired data is A groups of time series with a length of 3600*B time points. One time point represents the photovoltaic array output data for one second. The data samples contain 4 characteristic parameters, namely: open-circuit voltage, short-circuit current, maximum power point voltage, and maximum power point current. The 6 working states of the photovoltaic array are included in the A groups of time series, namely: normal operation, open-circuit fault, short-circuit fault, aging fault, shadow shielding fault, and fault where short-circuit and shadow occur simultaneously; Create a sliding window with a length of C time points, and extract 3600 * B * A / C data samples from the time series in Group A, where A > 10, B > 10, and C > 200; select the fault type at the last time point of the data samples as the sample label, and finally obtain the original data samples of various working states. Use to represent the original sample set of normal operation, to represent the original sample set of open circuit faults, to represent the original sample set of short circuit faults, to represent the original sample set of aging faults, to represent the original sample set of shadow occlusion, to represent the original sample set where short circuit and shadow occur simultaneously; Normalize the original data samples, and normalize the parameters at each time point in the time series to the range of [-1, 1], and the set is correspondingly transformed into sets X1 to X6; Next, in order to obtain the training dataset and validation dataset required during the training of the neural network, a certain number of samples X' are extracted from X n (n = 1, …, 6), and the training dataset T = [X1' … X' n and the validation dataset V = [X1 - X1' … X n - X' n are respectively formed; n ​ Step 2, training the WGAN-GP neural network and generating false samples to expand the dataset: When the number of samples in a certain type of dataset X' in the original training set T n is small, use the WGAN-GP neural network to learn the sample distribution of X' n to obtain a generator network that can transform noise into available data samples; after obtaining the corresponding generator network for X' n , generate a sufficient number of fake samples with labels and mix them with the real samples X' n to obtain a mixed dataset X n ″, and the set of all mixed datasets X n ″ and datasets X' with sufficient sample numbers n is the fault diagnosis dataset, denoted as T new ; The objective function \(L\) of the WGAN-GP neural network GP is as follows: Among them, x is the data sample after normalization in step one, which is the real data, is the generated data, P r and P g respectively represent the probability distributions of the real data and the generated data, E[…] represents the mathematical expectation, x~P r means that x follows P r distribution, means follows P g distribution, D(·) is the output of the discriminator, and the gradient penalty term λ is the penalty coefficient, is the random interpolation sampling of the generated sample and the real sample, ε represents a random value in the interval [0,1), is the gradient of the discriminator; Step 3, training the LSTM neural network and performing fault diagnosis: A photovoltaic array fault diagnosis model is composed of an LSTM network layer, a Dropout network layer, and two fully connected layers. Set the input feature dimension of the LSTM in the model to 4, the feature dimension of the hidden layer to 64, and the output feature dimensions of the fully connected layers to 32 and 6 respectively. After setting the model structure and parameters, use the fault diagnosis data set T obtained in step 2 new as the new training set and send it into the model for training; save the model with the smallest validation loss value during the iteration process, and finally obtain a neural network that can be used for photovoltaic array fault diagnosis.

2. The photovoltaic array small-sample fault diagnosis method based on neural network according to claim 1, characterized in that In Step 1, continuous sampling is carried out from 6 am to 6 pm, and the sampling time is 12 hours.

3. The small-sample fault diagnosis method for a photovoltaic array based on a neural network according to claim 1, characterized in that, Step 2: Use the WGAN-GP neural network to learn the sample distribution of X′ n Finally, obtain a generator network that can transform noise into available data samples. The specific steps are as follows: Step 2.1, set the parameters of the number of iterations, sample batches, and optimization algorithm required for training the WGAN-GP; Step 2.2, without considering the samples in the normal working state, take one type of fault sample for the training of the WGAN-GP neural network; Step 2.3, generate the same number of Gaussian white noises z according to the number of samples in one training batch; Step 2.4, the noise passes through the generator network to generate false samples; Step 2.5, mix the input samples and false samples in the same batch and send them into the discriminant model to obtain the loss value, and use this as a reference to update the discriminator network; Step 2.

6. Generate fake samples again Use the updated discriminant model to judge the samples, output the loss value; update the generator network; Step 2.7, use two indicators, the density index and the generator loss value index, to judge whether the sample generator network meets the standard. If it meets the standard, save the model; Step 2.8, repeat Step 2.2 to Step 2.7, and each type of fault obtains the corresponding generator network.

4. The small-sample fault diagnosis method for a photovoltaic array based on a neural network according to claim 1, characterized in that, In Step 2.7, the density index density uses the neighborhood sphere of the real samples to calculate whether the false samples meet the standard, and can evaluate whether the generated data is the same type of fault data. Its calculation formula is: Among them, k represents the proximity distance, M represents the number of false samples, N represents the number of true samples, 1... represents the indicator function, denotes the generated false samples, B(x, NND k (X′ n )) represents the spherical region with a radius of NND around the true data x k (X′ n ), X′ n is the set composed of a certain type of true data x, and NND k (X′ n ) represents the k-th nearest neighbor distance from a single true data x to other true data X′ n of the same class.