Small sample fault identification method of distributed photovoltaic system and related device

By collecting and processing current, voltage, and image data in a distributed photovoltaic system, and using conditions to generate adversarial networks and binary SVM models, the problem of low rare fault recognition accuracy is solved, and high-precision small sample fault recognition is achieved.

CN120386982APending Publication Date: 2025-07-29HUANENG JIANGXI CLEAN ENERGY GENERATION CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510484760.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art has low accuracy for rare fault recognition in distributed photovoltaic systems, especially in small sample scenarios, which leads to high missed detection rates.

Method used

By collecting the current, voltage and image data of photovoltaic components, preprocessing and feature fusion, using the conditional generation adversarial network to generate an expanded data set, and using a binary classification support vector machine model, setting high penalty weights, and identifying rare faults.

Benefits of technology

It improves the recognition accuracy of rare faults, enhances the generalization ability of the model under complex operating conditions, and reduces the missed detection rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386982A_ABST
    Figure CN120386982A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample fault identification method of a distributed photovoltaic system and a related device, and belongs to the technical field of fault diagnos.The method comprises the steps that current time sequence data, voltage time sequence data and image data of a photovoltaic module are collected, and fault category labels are extracted; preprocessing the current time sequence data, the voltage time sequence data and the image data, and performing feature fusion to obtain a multi-dimensional feature vector; inputting the multi-dimensional feature vector and the fault category label into a pre-trained conditional generative adversarial network to obtain an expanded data set; and inputting the expanded data set into a pre-trained dichotomy SVM sub-model, and performing identification to obtain a rare fault result and other fault results, thereby realizing small sample fault identification. According to the invention, the problem of low rare fault identification precision in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fault diagnosis, and particularly relates to a small-sample fault identification method and related device for a distributed photovoltaic system. Background Art

[0002] In the actual operation of a distributed photovoltaic power station, the types of faults are diverse and the occurrence frequencies vary significantly. For some specific faults, due to their slow development and hidden evolution process, the sample data thereof is scarce. Traditional machine learning and deep learning models rely on large-scale and balanced sample data for sufficient training to capture fault features and construct accurate classification boundaries. However, in the case of small samples, especially when the samples of rare fault types are insufficient, the model is prone to overfitting, that is, overlearning the noise or local features in the training data and failing to effectively extract general fault patterns. This will directly lead to a decline in the generalization ability of the model, and it is difficult to accurately identify rare faults when facing new samples or complex working conditions in actual diagnosis, resulting in missed or misjudged cases, which seriously affects the safe operation and maintenance efficiency of the photovoltaic power station.

[0003] Taking the fault of the encapsulation material aging of photovoltaic modules as an example, its fault characteristics are weak and the development cycle is long, resulting in difficult collection of fault samples. During the training process of traditional models, due to the lack of sufficient aging fault samples, it is difficult to learn the unique feature patterns of this type of fault, making the model prone to misjudging the aging fault as other similar faults or completely missing this fault type during diagnosis. In addition, the unbalanced data distribution will further exacerbate the bias of the model towards the majority-class faults, further reducing the recognition accuracy of rare faults.

[0004] Therefore, the existing technology relies on large-scale data sets. In the case of small samples, the model overfits noise, has poor generalization ability, and the data imbalance causes the classification hyperplane to be biased towards the majority class, resulting in a high missed detection rate of rare faults. For rare faults such as the aging of the encapsulation material of photovoltaic modules, the sample is too small to extract robust features. Summary of the Invention

[0005] The purpose of the present invention is to provide a small-sample fault identification method and related device for a distributed photovoltaic system to solve the problem of low recognition accuracy of rare faults in the existing technology.

[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, a small-sample fault identification method for a distributed photovoltaic system includes the following steps: Collect the current time-series data, voltage time-series data, and image data of photovoltaic modules, and extract the fault category labels; Preprocess the current time-series data, voltage time-series data, and image data, and perform feature fusion to obtain a multi-dimensional feature vector; Input the multi-dimensional feature vector and the fault category label into a pre-trained conditional generative adversarial network to obtain an augmented data set; Input the augmented data set into a pre-trained binary SVM sub-model to identify rare fault results and other fault results, realizing small-sample fault identification.

[0007] In some embodiments, the steps of preprocessing the current time-series data, voltage time-series data, and image data and performing feature fusion to obtain a multi-dimensional feature vector specifically include: Use the Isolation Forest algorithm to remove outliers from the current time-series data and voltage time-series data, and align the timestamps of the current time-series data, voltage time-series data, and image data; Extract the statistical features, frequency-domain features, and time-frequency features of the current time-series data and voltage time-series data, and fuse them with the image features extracted from the image data, and normalize them to obtain a multi-dimensional feature vector.

[0008] In some embodiments, the image features of the image data are extracted by DCNN, and the image features include hidden crack image features, aging image features, and hot spot image features.

[0009] In some embodiments, the steps of pre-training the conditional generative adversarial network include: Construct a generator network with an input of the concatenation of a random noise vector and the fault category label and an output of a virtual sample; Construct a discriminator network with an input of the multi-dimensional feature vector or virtual sample and the corresponding fault category label and an output of the true probability of the sample; Train the CGAN. During the process, use the WGAN-GP loss function to calculate the Wasserstein distance between the multi-dimensional feature vector and the virtual sample, and constrain the gradient output by the discriminator. Then, evaluate the quality of the generated multi-dimensional feature vector through FID and screen the generated effective multi-dimensional feature vectors to augment the data set until the training is completed.

[0010] In some embodiments, the steps of pre-training the binary SVM sub-model include: Select the radial basis function as the kernel function, and initialize the penalty parameter and kernel parameter to construct a binary SVM sub-model; Input the augmented data set into the binary SVM sub-model for training. During the training process, set a high misclassification penalty weight for the rare fault category until the training is completed.

[0011] In some embodiments, the steps of inputting the augmented data set into a pre-trained binary SVM sub-model to identify rare fault results and other fault results further include: Fusion weights each of the rare fault results and other fault results identified by the binary SVM sub-models to obtain the final fault identification result, achieving small-sample fault identification.

[0012] In a second aspect, a small-sample fault identification system for a distributed photovoltaic system includes: A data acquisition module for collecting current time-series data, voltage time-series data, and image data of photovoltaic modules, and extracting fault category labels; A feature extraction module for preprocessing the current time-series data, voltage time-series data, and image data, and performing feature fusion to obtain a multi-dimensional feature vector; A small-sample data augmentation module for inputting the multi-dimensional feature vector and the fault category labels into a pre-trained conditional generative adversarial network to obtain an augmented data set; A small-sample fault identification module for inputting the augmented data set into a pre-trained binary SVM sub-model to identify rare fault results and other fault results, achieving small-sample fault identification.

[0013] In a third aspect, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable in the processor. When the processor executes the computer program, the steps of the small-sample fault identification method for a distributed photovoltaic system are implemented.

[0014] In a fourth aspect, a computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the small-sample fault identification method for a distributed photovoltaic system are implemented.

[0015] In a fifth aspect, a computer program product includes a computer program. When the computer program is executed by a processor, the steps of the small-sample fault identification method for a distributed photovoltaic system are implemented.

[0016] Compared with the prior art, the present invention has the following beneficial effects: The present invention preprocesses current time-series data, voltage time-series data, and image data, and performs feature fusion to obtain a multi-dimensional feature vector. Among them, current, voltage, and image data reflect fault characteristics from different dimensions (such as electrical characteristics and physical appearance), and more comprehensive fault patterns can be captured after fusion. The multi-dimensional feature vector and fault category labels are input into a pre-trained conditional generative adversarial network to obtain an augmented data set, which can solve the problem of model overfitting in small-sample scenarios. By generating rare fault samples, the training data distribution is balanced, and the bias of the model towards majority-class faults is reduced. Moreover, the samples generated by the CGAN cover a wider range of fault patterns, improving the generalization ability of the model under complex working conditions. Therefore, the present invention can improve the recognition accuracy of small-sample faults.

[0017] Furthermore, the present invention constructs a binary SVM sub-model using the RBF kernel function, sets a high misclassification penalty weight for rare fault categories, and improves the recognition accuracy by weighted fusion of the results of multiple sub-models. The high penalty weight makes the model pay more attention to rare fault samples and avoids them being overwhelmed by majority-class samples, which is suitable for small-sample scenarios.

[0018] Furthermore, the present invention weighted-fuses the rare fault results and other fault results identified by each of the binary SVM sub-models to obtain the final fault recognition result, realizing small-sample fault recognition, which can reduce the accidental error of a single model and improve the recognition robustness. Description of the Drawings

[0019] Figure 1 is a flowchart of a method for identifying small-sample faults in a distributed photovoltaic system provided in this embodiment; Figure 2 is a structural diagram of a system for identifying small-sample faults in a distributed photovoltaic system provided in this embodiment. Detailed Embodiments

[0020] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the drawings. The content is an explanation of the present invention rather than a limitation.

[0021] It should be noted that the terms "including" and "having" and any variations thereof in the description and claims of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, systems, products, or devices.

[0022] Such as Figure 1As shown in the figure, this embodiment provides a small-sample fault identification method for a distributed photovoltaic system, including the following steps: S1: Data collection and preprocessing S1.1: Sensor data collection: Timing data of the current, voltage, temperature, and irradiance of the photovoltaic modules, as well as the inverter power output data, are collected in real time through sensors; Image data collection: Infrared thermal imaging cameras are used to obtain the hot spot images of the photovoltaic modules, and electroluminescence (EL) imaging devices are used to collect the crack images and aging images; Label data collection: Fault category labels are extracted from the operation and maintenance logs, covering multiple types of faults such as packaging aging, cracks, hot spots, and PID effects.

[0023] S1.2: Data cleaning and alignment. The isolation forest algorithm is used to remove the outliers in the above timing data and inverter power output data, and based on the inverter power output data, the acquisition time of the image data is synchronized to ensure that the timestamps of the above sensor data and image data are consistent.

[0024] The above current timing data and voltage timing data are corrected by the temperature, irradiance, and inverter power output data to improve the data quality.

[0025] S1.3: Timing feature extraction: Statistical features are extracted from the current timing data and voltage timing data, and the mean and variance of the current and voltage are calculated; Frequency domain features are extracted, and the power spectral density is extracted through the fast Fourier transform; Time-frequency features are extracted, and the energy entropy is extracted using wavelet transform.

[0026] Image feature extraction: High-dimensional features of the crack images and aging images in the EL (Electroluminescence, EL) images are extracted through a pre-trained deep convolutional neural network (DCNN), and the texture features of the hot spot images of the photovoltaic modules in the infrared images are calculated; The timing features and image features are fused and normalized to obtain a multi-dimensional feature vector.

[0027] The dataset of the multi-dimensional feature vectors is divided into a training set, a validation set, and a test set according to the fault category labels, and all rare fault samples are ensured to be retained in the training set.

[0028] S2. Sample generation is performed through a conditional generative adversarial network (CGAN); S2.1 Construct a CGAN model Construct a generator network whose input is the concatenation of a random noise vector and a fault class label. The network structure includes fully connected layers, upsampling layers, and convolutional layers, gradually decoding the input vector into a virtual sample with the same dimension as the real sample.

[0029] Construct a discriminator network whose input is a real sample (the multi-dimensional feature vector in S1) or a virtual sample generated by the generator, along with the corresponding fault class label. The network structure includes convolutional layers and fully connected layers, and the output sample is the probability of being real.

[0030] S2.2 CGAN Model Training During the training process, the generator and the discriminator perform adversarial learning. The generator attempts to generate realistic virtual samples to deceive the discriminator, while the discriminator endeavors to distinguish between real samples and virtual samples. The WGAN-GP (Wasserstein GAN with Gradient Penalty) loss function is adopted. This loss function calculates the Wasserstein distance between real samples and virtual samples and combines a gradient penalty term to improve the diversity and authenticity of the generated samples. Set reasonable training parameters, including the learning rate, batch size, and number of training epochs, to ensure the stable convergence of the model.

[0031] S2.3, Generated Sample Screening and Dataset Expansion During the training process, regularly evaluate the quality of the generated samples using FID (Fréchet Inception Distance). FID measures the authenticity of the generated samples by comparing the distribution differences between the generated samples and the real samples in the feature space. Screen the generated samples with lower FID as valid samples to expand the dataset. Combine the selected valid generated samples with the original training set to construct an expanded balanced dataset of multi-dimensional feature vectors. Ensure that the number of samples for each type of fault in the expanded dataset is balanced to avoid the impact of data imbalance on the model performance.

[0032] S3, Train a Support Vector Machine (SVM) S3.1 SVM Model Construction Select the radial basis function (RBF) as the kernel function of the SVM to construct a non-linear classification hyperplane. The RBF kernel function can handle non-linear classification problems, and the complexity and generalization ability of the model can be controlled by adjusting the kernel parameters. Initialize the SVM model parameters, including the penalty parameter C and the kernel parameter γ. The penalty parameter C controls the penalty strength of the model for misclassified samples, and the kernel parameter γ controls the width of the RBF kernel function.

[0033] S3.2 SVM Model Training The SVM model is trained using the cross - validation method. The augmented dataset is divided into multiple training subsets and validation subsets. Each subset is used as the validation set in turn, and the remaining subsets are used as the training set for model training to ensure that rare - fault samples are retained in the training set according to the fault - category labels.

[0034] During the cross - validation process, grid search is performed on the penalty parameter C and the kernel parameter γ to find the optimal parameter combination. By comparing the performance of the model on the validation set under different parameter combinations, the parameter combination with the strongest generalization ability is selected.

[0035] During the model training process, a higher misclassification penalty weight is set for rare - fault categories. By adjusting the class - weight parameter, the model pays more attention to rare - fault samples and improves its recognition sensitivity to rare faults.

[0036] S3.3: For each rare - fault category, a binary - classification SVM sub - model is trained. Each sub - model is responsible for distinguishing the corresponding rare fault from all other faults. By combining the recognition results of multiple sub - models (the recognition results of rare faults and other faults), multi - fault - type recognition is achieved. The trained multiple binary - classification SVM sub - models are fused to form the final multi - classification model. When performing fault recognition, the recognition results of each sub - model are integrated, and the final fault - recognition result is obtained through a weighted - fusion method such as voting or weighted voting.

[0037] As Figure 2 shown, this embodiment provides a small - sample fault - recognition system for a distributed photovoltaic system, including: A data - acquisition module, which is used to acquire the current - time - series data, voltage - time - series data, and image data of photovoltaic modules and extract fault - category labels; A feature - extraction module, which is used to pre - process the current - time - series data, voltage - time - series data, and image data and perform feature fusion to obtain a multi - dimensional feature vector; A small - sample data - augmentation module, which is used to input the multi - dimensional feature vector and fault - category labels into a pre - trained conditional generative adversarial network to obtain an augmented dataset; A small - sample fault - recognition module, which is used to input the augmented dataset into a pre - trained binary - classification SVM sub - model to identify rare - fault results and other - fault results, thereby realizing small - sample fault recognition.

[0038] The division of modules in the embodiments of the present invention is illustrative, merely a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present invention, the functional modules may be integrated in one processor, may exist separately physically, or two or more modules may be integrated in one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0039] In this embodiment, a computer device is further provided. The computer device includes a processor and a memory. The memory is used to store a computer program (in this embodiment, the computer program includes a calculation component and an iteration component, capable of performing model calculation and model update). The computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiments of the present invention may be used for the operation of a small-sample fault identification method for a distributed photovoltaic system.

[0040] This embodiment also provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by the processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of a small-sample fault identification method for a distributed photovoltaic system in the above embodiment.

[0041] This embodiment also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the corresponding steps of a small-sample fault identification method for a distributed photovoltaic system in the above embodiment.

[0042] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0043] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0044] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.

[0045] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.

[0046] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A small-sample fault identification method for a distributed photovoltaic system, characterized in that, The following steps are involved: Collect current time series data, voltage time series data, and image data of photovoltaic modules, and extract fault category labels; Preprocessing the current time series data, voltage time series data and image data, and performing feature fusion to obtain a multi-dimensional feature vector; Inputting the multidimensional feature vector and the fault category label into a pre-trained conditional generative adversarial network to obtain an expanded data set; The expanded data set is input into the pre-trained two-class SVM sub-model to identify rare fault results and other fault results, thereby realizing small sample fault identification.

2. The small-sample fault identification method for a distributed photovoltaic system according to claim 1, characterized in that, The steps of preprocessing the current time series data, the voltage time series data and the image data and fusing the features to obtain a multi-dimensional feature vector specifically include: An isolation forest algorithm is used to remove outliers from the current time series data and the voltage time series data, and to align timestamps of the current time series data, the voltage time series data, and the image data; Statistical features, frequency domain features and time-frequency features of the current time series data and the voltage time series data are extracted, and fused with the image features extracted from the image data, and normalized to obtain a multidimensional feature vector.

3. The small sample fault identification method for a distributed photovoltaic system according to claim 2, wherein Image features of image data are extracted through DCNN, and the image features include hidden crack image features, aging image features and hot spot image features.

4. The small-sample fault identification method for a distributed photovoltaic system according to claim 1, wherein The steps of pre-training the conditional generative adversarial network include: Constructing a generator network whose input is the concatenation of a random noise vector and the fault category label and whose output is a virtual sample; Constructing a discriminator network whose input is the multidimensional feature vector or virtual sample and the corresponding fault category label, and whose output is the true probability of the sample; The CGAN is trained by using the WGAN-GP loss function to calculate the Wasserstein distance between the multidimensional feature vector and the virtual sample, and the gradient of the discriminator output is constrained. The quality of the generated multidimensional feature vector is then evaluated by FID and the generated valid multidimensional feature vectors are screened to expand the data set until the training is completed.

5. A small sample fault identification method for a distributed photovoltaic system according to claim 1, characterized in that, The steps of pre-training the binary SVM sub-model include: Select radial basis function as kernel function and initialize penalty parameters and kernel parameters to build a binary classification SVM sub-model; The expanded data set is input into the binary classification SVM sub-model for training, and high misclassification penalty weights are set for rare fault categories during the training process until the training is completed.

6. The small sample fault identification method for a distributed photovoltaic system according to claim 1, characterized in that, The step of inputting the expanded data set into the pre-trained binary classification SVM sub-model to identify rare fault results and other fault results also includes: The rare fault results and other fault results obtained by each of the two-classification SVM sub-models are weightedly fused to obtain a final fault identification result, thereby realizing small sample fault identification.

7. A small-sample fault identification system for a distributed photovoltaic system, characterized in that, include: The data acquisition module is used to collect the current time series data, voltage time series data and image data of the photovoltaic module and extract the fault category label; A feature extraction module is used to pre-process the current time series data, voltage time series data and image data, and obtain a multi-dimensional feature vector after performing feature fusion; A small-sample data augmentation module for inputting the multi-dimensional feature vector and the fault category label into a pre-trained conditional generative adversarial network to obtain an augmented data set; A small-sample fault recognition module for inputting the augmented data set into a pre-trained binary SVM sub-model to recognize rare fault results and other fault results, thereby realizing small-sample fault recognition.

8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable in the processor. When the processor executes the computer program, the steps of the small-sample fault recognition method for a distributed photovoltaic system according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the small-sample fault recognition method for a distributed photovoltaic system according to any one of claims 1 to 6 are implemented.

10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the small-sample fault recognition method for a distributed photovoltaic system according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Hydroelectric generating set bolt looseness monitoring method and device

    CN121253142A

  • Method and device for monitoring bolt loosening of a hydroelectric unit

    CN121253142B