Photovoltaic fault diagnosis method and device based on convolutional neural network and maximum mean difference

By aligning the cross-domain feature distribution of photovoltaic arrays using the CNN-MMD framework, the problem of insufficient generalization ability of diagnostic models caused by differences in photovoltaic array topology is solved. This enables accurate fault identification and improved migration capabilities under unlabeled data conditions, making it suitable for intelligent monitoring and management of photovoltaic systems.

CN120934455APending Publication Date: 2025-11-11HUANENG CLEAN ENERGY RES INST +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510853997.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing photovoltaic array fault diagnosis models have limited generalization and transfer capabilities when faced with photovoltaic arrays of different topologies, and it is difficult to obtain high-quality labeled data on newly deployed photovoltaic arrays, resulting in insufficient diagnostic accuracy and adaptability.

Method used

An unsupervised domain adaptive deep learning framework based on convolutional neural networks (CNN) and maximum mean difference (MMD) loss function is adopted to construct a cross-photovoltaic array fault diagnosis model by aligning cross-domain feature distributions. Feature extraction and classification are performed using source domain labeled data and target domain unlabeled data.

Benefits of technology

Even in the absence of labeled data in the target domain, this method achieves accurate identification and cross-domain generalization of photovoltaic array faults, improving the accuracy and adaptability of fault diagnosis and making it suitable for intelligent monitoring and management of photovoltaic systems deployed on a large scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120934455A_ABST
    Figure CN120934455A_ABST
Patent Text Reader

Abstract

The invention provides a photovoltaic fault diagnosis method and device based on a convolutional neural network and maximum mean difference, and belongs to the technical field of photovoltaic array fault diagnosis of a cross-topological structure. The method comprises the steps that source domain labeled photovoltaic array electrical time sequence data and target domain unlabeled photovoltaic array electrical time sequence data are acquired, the source domain and the target domain correspond to photovoltaic arrays of different topological structures, and a sample data set of a cross-photovoltaic array fault diagnosis model is formed; performing normalization processing on each piece of electrical time sequence data in the sample data set; using test data in the sample data set to construct a cross-photovoltaic array fault diagnosis model based on a CNN-MMD unsupervised domain adaptive deep learning framework; and inputting the residual target domain unlabeled data for testing into the CNN-MMD model for fault diagnosis so as to accurately classify the fault type of the target domain to-be-tested data. The method can accurately identify normal, open circuit, short circuit and local shadow categories under the condition that the target domain lacks category labeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of photovoltaic array fault diagnosis technology across topologies, specifically relating to a photovoltaic fault diagnosis method and apparatus based on convolutional neural networks and maximum mean difference. Background Technology

[0002] Photovoltaic (PV) power generation technology, with its environmental friendliness and sustainability advantages, has become an important component of the modern energy system. Driven by the global consensus on addressing climate change, the deployment scale of PV systems continues to expand. Common faults in PV array operation, such as short-circuit faults, local shading faults, and open-circuit faults, can not only cause permanent damage to the solar panels but also potentially trigger safety accidents such as fires, resulting in significant economic losses. Therefore, real-time and accurate fault diagnosis is crucial for ensuring the safe and economical operation of PV systems. With the large-scale deployment and intelligent upgrading of PV systems, diagnostic models urgently need to address the challenges of differences in new array topologies and the scarcity of historical fault data in the target domain. Developing more adaptable cross-array fault diagnosis solutions has become an industry necessity.

[0003] Current machine learning methods generally rely on the assumption that training and test data are independently and identically distributed. This assumption is difficult to meet in practical engineering. Specifically, when diagnosing photovoltaic (PV) arrays with different topologies, fault data must be collected independently for each array, and a training set must be constructed. In practice, the same fault type in PV arrays with different topologies results in heterogeneous data distributions due to differences in parameters such as load. This severely limits the generalization and transfer capabilities of existing diagnostic models when dealing with heterogeneous PV arrays. Furthermore, obtaining high-quality labeled data from newly deployed PV arrays is often very difficult, further restricting the practical application of traditional supervised learning methods. Therefore, achieving accurate fault diagnosis of the target domain under the condition of sufficient source domain labeling and only unlabeled data in the target domain has become a significant technical bottleneck in the intelligent operation and maintenance of PV arrays.

[0004] To address the aforementioned issues, this invention proposes an unsupervised cross-domain adaptive cross-photovoltaic array fault diagnosis method. This method extracts discriminative features using a convolutional neural network (CNN) and aligns the feature distributions of the source and target domains using the maximum mean difference (MMD) loss function. Ultimately, it achieves accurate identification of multiple types of faults even when the target domain lacks labels. Summary of the Invention

[0005] This disclosure aims to at least solve one of the technical problems existing in the prior art, and to provide a photovoltaic fault diagnosis method and apparatus based on convolutional neural networks and maximum mean difference.

[0006] One aspect of this disclosure provides a photovoltaic fault diagnosis method based on convolutional neural networks and the difference between the maximum mean and the maximum mean, the method comprising:

[0007] S110. Obtain electrical timing data of tagged photovoltaic arrays in the source domain and electrical timing data of untagged photovoltaic arrays in the target domain. The source domain and the target domain correspond to photovoltaic arrays with different topologies, forming a sample dataset for a cross-photovoltaic array fault diagnosis model.

[0008] S120. Normalize each electrical timing data in the sample dataset;

[0009] S130. Using the test data in the sample dataset, construct a cross-photovoltaic array fault diagnosis model based on the CNN-MMD unsupervised domain adaptive deep learning framework;

[0010] S140. Input the remaining unlabeled target domain data for testing into the trained CNN-MMD model for fault diagnosis, so as to accurately classify the fault types of the target domain test data.

[0011] Optionally, in step S110, the source domain tagged photovoltaic array electrical timing data is obtained from the photovoltaic array of the first topology by a photovoltaic inverter;

[0012] The electrical timing data of the target domain unlabeled photovoltaic array is collected from the target photovoltaic array of the second topology using a photovoltaic inverter.

[0013] Optionally, in step S110, the source domain tagged photovoltaic array electrical timing data includes voltage and current timing signals and fault category tags;

[0014] The electrical timing data of the target domain unlabeled photovoltaic array includes voltage and current timing signals.

[0015] Optionally, in step S120, the two reference photovoltaic panels and the arranged photovoltaic array are placed in the same working environment. Based on the open-circuit voltage and short-circuit current collected by the reference photovoltaic panels, each electrical timing data in the source and target domains is preprocessed using normalization. The method is as follows:

[0016]

[0017] Among them, V NORM and I NORM For the normalized data, V PVA I represents the total array voltage. PVA This represents the total array current. (V) ROC I represents the open-circuit voltage of the reference photovoltaic panel. RSC This represents the short-circuit current of the reference photovoltaic panel.

[0018] Optionally, in step S130, using the test data in the sample dataset, a cross-photovoltaic array fault diagnosis model based on the CNN-MMD unsupervised domain adaptive deep learning framework is constructed, including:

[0019] Based on the test data in the sample dataset, a convolutional neural network feature extraction network is used to extract high-dimensional features from the normalized source and target domain electrical time-series data; then, a label classifier is used to predict the category of the source domain training data based on the source domain data features, and the label classification loss is output.

[0020] Calculate the maximum mean difference loss of the feature distributions of the source domain data and the target domain data selected for training, measure the difference in the distributions of the source domain and the target domain in the reproducing kernel Hilbert space, and output the MMD loss.

[0021] Based on the label classification loss and the MMD loss, a total loss function for the CNN-MMD model is constructed, and the model is trained to obtain a cross-photovoltaic array fault diagnosis model that maintains the source domain classification capability aligned with the cross-domain feature distribution.

[0022] Optionally, the feature extraction network consists of five layers of one-dimensional convolutional layers, normalization layers, pooling layers, and projection embedding layers;

[0023] The label classifier consists of two fully connected layers, and the output layer uses the Softmax activation function to generate the class probability distribution.

[0024] Optionally, the total loss function for constructing the CNN-MMD model includes two parts: the source domain classification loss function L. cls and MMD loss function λ mmd The components are combined using weighted coefficients to form the total loss function L. total ;in,

[0025] The total loss function is:

[0026] L total =L cls +λ mmd L mmd

[0027] Where, λ mmd These are the weighting coefficients of the MMD loss function.

[0028] Optionally, in step S140, the fault type includes any one of normal, open circuit, short circuit, and partial shadowing fault.

[0029] In another aspect, this disclosure proposes a photovoltaic fault diagnosis device based on convolutional neural networks and the maximum mean difference, characterized in that the device comprises: a sample dataset acquisition unit, a normalization processing unit, a model building unit, and a fault diagnosis unit; wherein,

[0030] The sample dataset acquisition unit is used to acquire electrical timing data of labeled photovoltaic arrays in the source domain and electrical timing data of unlabeled photovoltaic arrays in the target domain. The source domain and the target domain correspond to photovoltaic arrays with different topologies, forming a sample dataset for a cross-photovoltaic array fault diagnosis model.

[0031] The normalization processing unit is used to normalize each electrical timing data in the sample dataset;

[0032] The model building unit is used to construct a cross-photovoltaic array fault diagnosis model based on the CNN-MMD unsupervised domain adaptive deep learning framework using the test data in the sample dataset.

[0033] The fault diagnosis unit is used to input the remaining unlabeled data of the target domain used for testing into the trained CNN-MMD model for fault diagnosis, so as to accurately classify the fault type of the test data in the target domain.

[0034] Optionally, the model building unit includes: a label classification loss output module, an MMD loss output module, and a training module; wherein,

[0035] The label classification loss output module is used to extract high-dimensional features from the normalized source and target domain electrical time-series data based on the test data in the sample dataset through a convolutional neural network feature extraction network; then, a label classifier predicts the category of the source domain training data based on the source domain data features and outputs the label classification loss.

[0036] The MMD loss output module is used to calculate the maximum mean difference loss of the feature distributions of the source domain data and the target domain data selected for training, measure the distribution difference between the source domain and the target domain in the regenerating kernel Hilbert space, and output the MMD loss.

[0037] The training module is used to construct the total loss function of the CNN-MMD model based on the label classification loss and the MMD loss, and train the model to obtain a cross-photovoltaic array fault diagnosis model that maintains the source domain classification ability aligned with the cross-domain feature distribution.

[0038] This disclosure proposes a photovoltaic (PV) fault diagnosis method and apparatus based on convolutional neural networks and maximum mean difference (MMD). The method includes: acquiring labeled electrical time-series data of a source domain PV array and unlabeled electrical time-series data of a target domain PV array, where the source and target domains correspond to PV arrays with different topologies, forming a sample dataset for a cross-PV array fault diagnosis model; normalizing each electrical time-series data in the sample dataset; constructing a cross-PV array fault diagnosis model based on a CNN-MMD unsupervised domain adaptive deep learning framework using test data from the sample dataset; and inputting the remaining unlabeled target domain data used for testing into the trained CNN-MMD model for fault diagnosis to accurately classify the fault type of the target domain test data. This method integrates labeled electrical time-series data from the source domain and unlabeled electrical time-series data from the target domain to construct a sample set, learns high-dimensional feature representations through a CNN-based feature extraction network and label classifier, and combines MMD to accurately measure and align the feature distributions of the source and target domains, improving the model's cross-domain generalization performance and transfer capability to the target domain. This method can accurately identify normal, open circuit, short circuit, and partial shadow categories even when the target domain lacks category labels, and is suitable for intelligent monitoring and fault management of photovoltaic systems under large-scale deployment. Attached Figure Description

[0039] Figure 1 This is a flowchart illustrating a photovoltaic fault diagnosis method based on convolutional neural networks and the difference between the maximum mean and the maximum mean, according to a specific embodiment of this disclosure.

[0040] Figure 2 This is a structural diagram of the CNN-MMD model in Embodiment 1 of this disclosure;

[0041] Figure 3 This is a schematic diagram of a photovoltaic fault diagnosis device based on convolutional neural network and maximum mean difference according to Embodiment 1 of this disclosure. Detailed Implementation

[0042] To enable those skilled in the art to better understand the technical solutions of this disclosure, the disclosure will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain this disclosure and represent a part of the embodiments of this disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the protection scope of this disclosure.

[0043] like Figure 1 and Figure 2 As shown, one aspect of this disclosure provides a photovoltaic fault diagnosis method S100 based on the difference between a convolutional neural network and the maximum mean, specifically including the following steps S110 to S140:

[0044] S110. Obtain the electrical timing data of the source domain labeled photovoltaic array and the electrical timing data of the target domain unlabeled photovoltaic array. The source domain and the target domain correspond to photovoltaic arrays with different topologies. Together, they form a sample dataset for a cross-photovoltaic array fault diagnosis model.

[0045] In step S110, the source-domain tagged photovoltaic array electrical timing data is acquired from the photovoltaic array of the first topology by the photovoltaic inverter. The source-domain tagged photovoltaic array electrical timing data includes voltage and current timing signals and fault category tags.

[0046] In step S110, the electrical timing data of the target domain untagged photovoltaic array is collected from the target photovoltaic array of the second topology using a photovoltaic inverter. The electrical timing data of the target domain untagged photovoltaic array includes voltage and current timing signals.

[0047] It should be understood that the target domain data should be divided into a training subset and a test subset. The training subset is used for model training, and the test subset is used to test the diagnostic accuracy of the model.

[0048] S120. Normalize each electrical timing data in the sample dataset to obtain normalized voltage and current timing data, and then concatenate the two to form a two-dimensional data matrix.

[0049] In step S120, the two reference photovoltaic panels and the arranged photovoltaic array are placed in the same working environment. Based on the open-circuit voltage and short-circuit current collected by the reference photovoltaic panels, the electrical timing data of each source domain and target domain are normalized and preprocessed. The method is as follows:

[0050]

[0051] Among them, V NORM and I NORM For the normalized data, V PVA I represents the total array voltage. PVA This represents the total array current. (V) ROC I represents the open-circuit voltage of the reference photovoltaic panel. RSC This represents the short-circuit current of the reference photovoltaic panel.

[0052] S130. Using the test data in the sample dataset, construct a cross-photovoltaic array fault diagnosis model based on the CNN-MMD unsupervised domain adaptive deep learning framework.

[0053] In step S130, the model construction process includes the following steps:

[0054] S1301. Based on the test data in the sample dataset, high-dimensional features are extracted from the normalized source and target domain electrical time-series data through a convolutional neural network feature extraction network; then, a label classifier is used to predict the category of the source domain training data based on the source domain data features, and the label classification loss is output.

[0055] In step S1301, the cross-photovoltaic array fault diagnosis model adopts a modular design, including two parts: a feature extraction network and a label classifier. The feature extraction network consists of five layers of one-dimensional convolutional layers, normalization layers, pooling layers, and projection embedding layers, which can extract high-dimensional features with task relevance from the input samples and are used to extract high-dimensional features from the normalized data. The label classifier consists of two fully connected layers, and the output layer uses the Softmax activation function to generate the class probability distribution.

[0056] S1302. Calculate the maximum mean difference loss of the feature distributions of the source domain data and the target domain data selected for training, measure the difference in distribution between the source domain and the target domain in the reproducing kernel Hilbert space, and output the MMD loss.

[0057] In step S1302, the maximum mean difference (MMD) is introduced as a distribution distance metric to accurately calculate the feature distribution difference between the source and target domain feature extraction network outputs in the regenerating kernel Hilbert space (RKHS). The MMD loss is applied to the output layer of the feature extractor, achieving feature space alignment by minimizing the inter-domain distribution difference, thereby improving the model's generalization ability and diagnostic accuracy in the target domain.

[0058] S1303. Based on the label classification loss and the MMD loss, construct the total loss function of the CNN-MMD model, and train the model to obtain a cross-photovoltaic array fault diagnosis model that maintains the source domain classification capability aligned with the cross-domain feature distribution.

[0059] In step S1303, the constructed CNN-MMD total loss function consists of two parts: the source domain classification loss function L. cls and MMD loss function λ mmd The various components are combined using weighted coefficients to form a unified optimization objective function, the total loss function L. total for:

[0060] L total =L cls +λ mmd L mmd

[0061] Where, λ mmdThese are the weighting coefficients of the MMD loss function. Optimizing this joint loss allows the model to maintain high classification accuracy in the source domain while making the target domain samples closer to the source domain sample distribution in the feature space. This achieves the dual goals of feature consistency learning and discriminative enhancement, thereby improving the diagnostic accuracy of the target domain data.

[0062] In step S1303, the total loss function of the CNN-MMD model is constructed by weighted fusion of the label classification loss and the MMD loss. The label classification loss focuses on optimizing the accuracy of fault identification, while the MMD loss aims to achieve fine-grained feature alignment. Through an end-to-end joint training mechanism, this model can maintain high classification accuracy while improving cross-domain adaptability, ultimately achieving efficient transfer of knowledge from the source domain to the target domain.

[0063] S140. Input the remaining unlabeled target domain data for testing into the trained CNN-MMD model for fault diagnosis, so as to accurately classify the fault types of the target domain test data.

[0064] In step S140, the proposed CNN-MMD method can improve the generalization ability of cross-array fault knowledge transfer and diagnostic model under the condition that only a small amount of unlabeled target domain data is used for training.

[0065] This invention provides a cross-photovoltaic array fault diagnosis method based on the CNN-MMD unsupervised domain adaptive deep learning framework. This method can construct a cross-photovoltaic array fault diagnosis model even when the target domain lacks labels. Compared with existing algorithms, this method can effectively alleviate the domain offset problem caused by topological changes, improve the accuracy of cross-photovoltaic array fault diagnosis, and has strong practicality and broad application prospects.

[0066] It should be noted that the diagnostic method provided in this disclosure can be executed by the unsupervised domain adaptive cross-photovoltaic array fault diagnosis device provided in this disclosure, or by electronic devices, wherein the electronic devices may include, but are not limited to, terminal devices such as desktop computers and tablet computers.

[0067] It should be understood that electronic devices include one or more processors, one or more memories, input devices, output devices, etc., and these components are interconnected through bus systems and / or other forms of connection mechanisms.

[0068] The processor may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0069] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, which a processor may execute to implement the client functions (implemented by the processor) in the embodiments of this disclosure described below, and / or other desired functions. Various applications and various data may also be stored in the computer-readable storage medium, such as various data used and / or generated by the applications.

[0070] The input device can be a device used by a user to input instructions, and may include one or more of the following: keyboard, mouse, microphone, and touch screen.

[0071] The output device can output various information (such as images or sounds) to the outside (e.g., a user) and may include one or more of a display, speaker, etc.

[0072] like Figure 2 and Figure 3 As shown, another aspect of this disclosure proposes a photovoltaic fault diagnosis device 200 based on convolutional neural networks and maximum mean difference. This device includes: a sample dataset acquisition unit 210, a normalization processing unit 220, a model building unit 230, and a fault diagnosis unit 240. The sample dataset acquisition unit 210 acquires electrical time-series data of labeled photovoltaic arrays in the source domain and unlabeled photovoltaic arrays in the target domain, where the source and target domains correspond to photovoltaic arrays with different topologies, forming a sample dataset for a cross-photovoltaic array fault diagnosis model. The normalization processing unit 220 normalizes each electrical time-series data in the sample dataset. The model building unit 230 uses test data from the sample dataset to construct a cross-photovoltaic array fault diagnosis model based on the CNN-MMD unsupervised domain adaptive deep learning framework. The fault diagnosis unit 240 inputs the remaining unlabeled target domain data used for testing into the trained CNN-MMD model for fault diagnosis, accurately classifying the fault type of the target domain test data.

[0073] Furthermore, the sample dataset acquisition unit 210 acquires source domain data from the photovoltaic array of the first topology structure and target domain data from the target photovoltaic array of the second topology structure through the photovoltaic inverter.

[0074] Furthermore, the normalization processing unit 220 performs normalization preprocessing on each electrical timing data of the source and target domains based on the open-circuit voltage and short-circuit current collected by the reference photovoltaic panels, in an environment where the two reference photovoltaic panels and the arranged photovoltaic array are placed in the same working environment. The method is as follows:

[0075]

[0076] Among them, V NORM and I NORM For the normalized data, V PVA I represents the total array voltage. PVA This represents the total array current. (V) ROC I represents the open-circuit voltage of the reference photovoltaic panel. RSC This represents the short-circuit current of the reference photovoltaic panel.

[0077] Further, the model building unit 230 includes: a label classification loss output module 231, an MMD loss output module 232, and a training module 233; wherein, the label classification loss output module 231 is used to extract high-dimensional features from the normalized source domain and target domain electrical time-series data based on the test data in the sample dataset through a convolutional neural network feature extraction network; then, a label classifier predicts the category of the source domain training data based on the source domain data features, and outputs the label classification loss; the MMD loss output module 232 is used to calculate the maximum mean difference loss of the feature distribution between the source domain data and the target domain data selected for training, measure the distribution difference between the source domain and the target domain in the regenerating kernel Hilbert space, and output the MMD loss; the training module 233 is used to construct the total loss function of the CNN-MMD model according to the label classification loss and the MMD loss, and train the model to obtain a cross-photovoltaic array fault diagnosis model whose source domain classification ability is aligned with the cross-domain feature distribution.

[0078] In the training module, the constructed CNN-MMD total loss function consists of two parts: the source domain classification loss function L... cls and MMD loss function λ mmd The various components are combined using weighted coefficients to form a unified optimization objective function, the total loss function L. total for:

[0079] L total =L cls +λ mmd L mmd

[0080] Where, λ mmdThese are the weighting coefficients of the MMD loss function. Optimizing this joint loss allows the model to maintain high classification accuracy in the source domain while making the target domain samples closer to the source domain sample distribution in the feature space. This achieves the dual goals of feature consistency learning and discriminative enhancement, thereby improving the diagnostic accuracy of the target domain data.

[0081] The cross-photovoltaic array fault diagnosis model in this embodiment adopts a modular design, comprising a feature extraction network and a label classifier. The feature extraction network consists of five layers: a one-dimensional convolutional layer, a normalization layer, a pooling layer, and a projection embedding layer. It can extract high-dimensional features with task relevance from the input samples and is used for high-dimensional feature extraction on the normalized data. The label classifier consists of two fully connected layers, and the output layer uses the Softmax activation function to generate the class probability distribution.

[0082] In this implementation, the MMD loss output module introduces Maximum Mean Difference (MMD) as a distribution distance metric to accurately calculate the feature distribution difference between the source and target domain feature extraction network outputs in the Regenerating Kernel Hilbert Space (RKHS). The MMD loss is applied to the feature extractor output layer, achieving feature space alignment by minimizing the inter-domain distribution difference, thereby improving the model's generalization ability and diagnostic accuracy in the target domain.

[0083] It should be noted that the system embodiments described in this disclosure are merely illustrative. For example, the division of the units can be a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0084] The following will further illustrate the cross-photovoltaic array fault diagnosis method based on the CNN-MMD unsupervised domain adaptive deep learning framework with specific embodiments:

[0085] Example 1

[0086] This example demonstrates a cross-photovoltaic array fault diagnosis method based on the CNN-MMD unsupervised domain adaptive deep learning framework, including the following steps:

[0087] Step S1: Obtain electrical timing data of tagged photovoltaic arrays in the source domain and electrical timing data of untagged photovoltaic arrays in the target domain. The source domain and the target domain correspond to photovoltaic arrays with different topologies, thereby constructing a sample dataset for a cross-photovoltaic array fault diagnosis model.

[0088] Step S2: Normalize each electrical timing data in the sample dataset;

[0089] Step S3: Construct a cross-photovoltaic array fault diagnosis model based on the CNN-MMD unsupervised domain adaptive deep learning framework. It extracts high-dimensional features from the electrical time-series data of the source and target domains through a convolutional neural network feature extraction network. Then, it predicts the category of the source domain training data based on the source domain data features through a label classifier and outputs the classification loss.

[0090] Step S4: Calculate the maximum mean difference (MMD) loss of the feature distributions of the source domain data and the target domain data selected for training, which measures the difference in distribution between the source domain and the target domain in the reproducing kernel Hilbert space (RKHS).

[0091] Step S5: Jointly optimize the label classification loss and MMD loss to construct the total loss function of the CNN-MMD model, and train the model to achieve the alignment of source domain classification ability with cross-domain feature distribution;

[0092] Step S6: Input the remaining unlabeled target domain data for testing into the trained CNN-MMD model for fault diagnosis to accurately classify the fault types of the target domain test data, including normal, open circuit, short circuit and local shadow faults.

[0093] In step S1 of this embodiment, data is acquired from the photovoltaic array via a photovoltaic inverter, and voltage and current time series are collected by sensors and a host computer to construct a dataset; the photovoltaic array is equipped with GL-M100 solar cells; the inverter is a GoodWe GW8000-MS-C30. The source domain array is a 2×6 array, and the target domain array is a 3×6 array.

[0094] In this embodiment, all samples from the source domain and 20% of the unlabeled samples from the target domain are used to train the model, while the remaining 80% of the target domain samples are used as the test set for model testing. The distribution of the data samples is shown in Table 1.

[0095] Table 1 Distribution of Data Samples

[0096]

[0097] The model classification performance is shown in Table 2. ERM stands for Empirical Risk Minimization, which minimizes the sum of empirical risks from the source domain samples. In this patent, its optimization function is set as the cross-entropy loss function for comparison. As can be seen from Table 2, the CNN-MMD model proposed in this patent achieves an accuracy of 97.18% in the set task, which is superior to the compared ERM model.

[0098] Table 2. Classification accuracy (%) of each model

[0099]

[0100]

[0101] Table 2 shows that the ERM model has an accuracy of 66.77%, indicating that the proportion of correctly classified samples in the overall classification task is relatively low. The CNN-MMD model achieves an accuracy of 97.18%, indicating that it has high overall accuracy in identifying different fault states (including normal states), correctly classifying the vast majority of samples with few misclassifications. The ERM model has a recall of 66.88%, meaning that 66.88% of the actual positive samples were correctly identified by the model, with a considerable number of positive samples not being correctly identified. The CNN-MMD model has a recall of 97.19%, indicating that it can capture various types of samples that actually exist (including fault and normal states) well with few omissions. The ERM model has an F1 score of 59.07, indicating that it performs poorly in balancing precision and recall. The CNN-MMD model, with an F1 score of 97.19, achieves a better balance between accurately identifying samples (high precision) and identifying as many relevant samples as possible (high recall), resulting in superior overall performance. In conclusion, the CNN-MMD model exhibits better classification performance and can more accurately identify samples in various states (normal and different fault types) under different array specifications.

[0102] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0103] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0104] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0105] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0106] It is understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of this disclosure, and this disclosure is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and substance of this disclosure, and these modifications and improvements are also considered to be within the scope of protection of this disclosure.

Claims

1. A photovoltaic fault diagnosis method based on convolutional neural networks and maximum mean difference, characterized in that, The method includes: S110. Obtain electrical timing data of tagged photovoltaic arrays in the source domain and electrical timing data of untagged photovoltaic arrays in the target domain. The source domain and the target domain correspond to photovoltaic arrays with different topologies, forming a sample dataset for a cross-photovoltaic array fault diagnosis model. S120. Normalize each electrical timing data in the sample dataset; S130. Using the test data in the sample dataset, construct a cross-photovoltaic array fault diagnosis model based on the CNN-MMD unsupervised domain adaptive deep learning framework; S140. Input the remaining unlabeled target domain data used for testing into the trained CNN-MMD model for fault diagnosis, so as to accurately classify the fault types of the target domain test data.

2. The method according to claim 1, characterized in that, In step S110, the source domain tagged photovoltaic array electrical timing data is obtained from the photovoltaic array of the first topology by a photovoltaic inverter; The electrical timing data of the target domain unlabeled photovoltaic array is collected from the target photovoltaic array of the second topology using a photovoltaic inverter.

3. The method according to claim 1, characterized in that, In step S110, the source domain tagged photovoltaic array electrical timing data includes voltage and current timing signals and fault category tags; The electrical timing data of the target domain unlabeled photovoltaic array includes voltage and current timing signals.

4. The method according to claim 1, characterized in that, In step S120, the two reference photovoltaic panels and the arranged photovoltaic array are placed in the same working environment. Based on the open-circuit voltage and short-circuit current collected by the reference photovoltaic panels, the electrical timing data of each source domain and target domain are normalized and preprocessed. The method is as follows: Among them, V NORM and I NORM For the normalized data, V PVA I represents the total array voltage. PVA V represents the total array current. ROC I represents the open-circuit voltage of the reference photovoltaic panel. RSC This represents the short-circuit current of the reference photovoltaic panel.

5. The method according to claim 1, characterized in that, In step S130, using the test data in the sample dataset, a cross-photovoltaic array fault diagnosis model based on the CNN-MMD unsupervised domain adaptive deep learning framework is constructed, including: Based on the test data in the sample dataset, a convolutional neural network feature extraction network is used to extract high-dimensional features from the normalized source and target domain electrical time-series data; then, a label classifier is used to predict the category of the source domain training data based on the source domain data features, and the label classification loss is output. Calculate the maximum mean difference loss of the feature distributions of the source domain data and the target domain data selected for training, measure the difference in the distributions of the source domain and the target domain in the reproducing kernel Hilbert space, and output the MMD loss. Based on the label classification loss and the MMD loss, a total loss function for the CNN-MMD model is constructed, and the model is trained to obtain a cross-photovoltaic array fault diagnosis model that maintains the source domain classification capability aligned with the cross-domain feature distribution.

6. The method according to claim 5, characterized in that, The feature extraction network consists of five layers: a one-dimensional convolutional layer, a normalization layer, a pooling layer, and a projection embedding layer. The label classifier consists of two fully connected layers, and the output layer uses the Softmax activation function to generate the class probability distribution.

7. The method according to claim 5, characterized in that, The total loss function for constructing the CNN-MMD model consists of two parts: the source domain classification loss function L. cls and MMD loss function λ mmd The components are combined using weighted coefficients to form the total loss function L. total ;in, The total loss function is: L total L cls +λ mmd L mmd Where, λ mmd These are the weighting coefficients of the MMD loss function.

8. The method according to claim 1, characterized in that, In step S140, the fault type includes any one of normal, open circuit, short circuit, and partial shadowing fault.

9. A photovoltaic fault diagnosis device based on convolutional neural networks and maximum mean difference, characterized in that, The device includes: a sample dataset acquisition unit, a normalization processing unit, a model building unit, and a fault diagnosis unit; wherein... The sample dataset acquisition unit is used to acquire electrical timing data of labeled photovoltaic arrays in the source domain and electrical timing data of unlabeled photovoltaic arrays in the target domain. The source domain and the target domain correspond to photovoltaic arrays with different topologies, forming a sample dataset for a cross-photovoltaic array fault diagnosis model. The normalization processing unit is used to normalize each electrical timing data in the sample dataset; The model building unit is used to construct a cross-photovoltaic array fault diagnosis model based on the CNN-MMD unsupervised domain adaptive deep learning framework using the test data in the sample dataset. The fault diagnosis unit is used to input the remaining unlabeled data of the target domain used for testing into the trained CNN-MMD model for fault diagnosis, so as to accurately classify the fault type of the test data in the target domain.

10. The apparatus according to claim 9, characterized in that, The model building unit includes: a label classification loss output module, an MMD loss output module, and a training module; wherein... The label classification loss output module is used to extract high-dimensional features from the normalized source and target domain electrical time-series data based on the test data in the sample dataset through a convolutional neural network feature extraction network; then, a label classifier predicts the category of the source domain training data based on the source domain data features and outputs the label classification loss. The MMD loss output module is used to calculate the maximum mean difference loss of the feature distributions of the source domain data and the target domain data selected for training, measure the distribution difference between the source domain and the target domain in the regenerating kernel Hilbert space, and output the MMD loss. The training module is used to construct the total loss function of the CNN-MMD model based on the label classification loss and the MMD loss, and train the model to obtain a cross-photovoltaic array fault diagnosis model that maintains the source domain classification ability aligned with the cross-domain feature distribution.