Fault Diagnosis Method, Device and Equipment for Photovoltaic Array Based on Improved Label Propagation
By improving label propagation method and cost-sensitive learning, the label probability adjacency matrix is constructed, which solves the problem of indistinguishable similar fault types in photovoltaic array fault diagnosis, improves detection accuracy and efficiency, and performs better especially under low irradiance conditions.
Patent Information
- Application Number
- CN202311215799.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-09-19
AI Technical Summary
Existing machine learning algorithms have limited generalization capabilities in photovoltaic array fault diagnosis, making it difficult to accurately distinguish similar fault types, and the detection accuracy is reduced under low irradiance conditions.
The improved label propagation method is used to construct the label probability adjacency matrix, and the cost-sensitive learning is used to model the photovoltaic array fault sample data. The rules are learned by a small amount of labeled data and the label is propagated to unlabeled samples, and the risk is optimized and predicted in combination with Bayesian risk theory.
It improves the accuracy of similar fault detection in photovoltaic arrays, enhances the efficiency of fault diagnosis, and performs better in low irradiance conditions.
Smart Images

Figure CN117171690B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of photovoltaic array fault diagnosis, and particularly to a photovoltaic array fault diagnosis method, device and equipment based on improved label propagation. Background Technique
[0002] Since a large amount of manpower and material resources are required for manual marking of the fault types of large-scale photovoltaic arrays, and the proportion of marked data in the actual data is small. To solve the problem of fault diagnosis for large sample size data with a low proportion of labeled data, in recent years, with the rapid development of artificial intelligence, machine learning algorithms, as an important branch of artificial intelligence, have also been widely used in solving the problem of photovoltaic array fault diagnosis. When using machine learning algorithms for fault diagnosis, it is necessary to input fault sample data, establish a model through intelligent learning, and predict the fault types of photovoltaic arrays. Since the method based on artificial intelligence machine learning relies on establishing a mapping relationship by learning the characteristics of data samples, rather than establishing a physical or mathematical model, it is also called a data model. The machine learning methods for diagnosing photovoltaic arrays generally can be divided into three categories: supervised learning, semi-supervised learning, and unsupervised learning.
[0003] The data set used for photovoltaic array fault diagnosis based on supervised learning is a labeled data set, and data that needs to be manually marked is used as training data to learn the relationship between features and target parameters, so as to realize the fault diagnosis of new unlabeled samples. Wang Yuanzhang et al. As the input data of the neural network, a photovoltaic array fault diagnosis model based on the BP neural network was proposed, realizing the identification of four faults: short circuit, open circuit, shading, and aging. Spataru S. et al. proposed a method for detecting shading faults, abnormal aging faults, and potential induced degradation affecting photovoltaic arrays by analyzing the changes in the I-V curve of photovoltaic arrays. The disadvantage of this method is that it can only be carried out under conditions of higher irradiance, and the detection accuracy is significantly reduced at low irradiance.
[0004] The research on photovoltaic array fault diagnosis based on semi-supervised learning means that only part of the data is marked in the input data, and the fault types of unmarked data can be identified through the mutual propagation between the marks. Zhao Y. et al. proposed a photovoltaic array fault diagnosis model based on a graph, connected special points, and identified the data deviating from the connection line of the fault samples as faults. This model only uses a small amount of marked training data and normalizes these data to improve the visualization effect.
[0005] The basic idea of the photovoltaic array fault diagnosis method based on unsupervised learning is to cluster the fault sample data of the photovoltaic array through the clustering method and then perform label assignment. Lin P. et al. proposed a photovoltaic array fault diagnosis method based on density peak clustering, and obtained the clustering results through the connectivity between data samples for fault identification. Ding H. et al. proposed a fault diagnosis method for photovoltaic systems based on local outlier factor, using the outlier detection algorithm (LOF) to compare the density of sample points with the density of surrounding labeled sample points to judge the fault type. Zhu H. et al. proposed a photovoltaic array fault diagnosis method based on unsupervised sample clustering and probability neural network model, analyzed the output characteristics and electrical feature vector distribution of the photovoltaic array under typical fault conditions, introduced the unit method and Gaussian kernel function into the fuzzy C-means algorithm, and improved the applicability of unsupervised screening to various fault samples and the fuzzy clustering ability.
[0006] To sum up, although the current intelligent diagnosis algorithms based on machine learning have been widely applied, most of the algorithms only perform fault diagnosis based on a single classifier, still having deficiencies, with limited generalization ability. It is difficult to select a sufficiently accurate clustering center for noisy data. Moreover, different fault types are similar, the parameter discrimination degree is not high, it is difficult to accurately distinguish and cluster, and the prediction accuracy still needs to be improved. Summary of the Invention
[0007] Based on this, in view of the above technical problems, it is necessary to provide a photovoltaic array fault diagnosis method, device and equipment based on improved label propagation that can improve the detection accuracy of similar faults in photovoltaic arrays.
[0008] A photovoltaic array fault diagnosis method based on improved label propagation, the method includes:
[0009] Obtain training sample data through the photovoltaic array. The training sample data is photovoltaic electric energy, where the photovoltaic electric energy includes labeled samples and unlabeled samples.
[0010] Input the training sample data into the fault diagnosis model for training, and use the improved label propagation method to obtain the propagation probability of the labeled samples propagating to the unlabeled samples adjacent to the labeled samples.
[0011] Construct a label probability adjacency matrix according to the number of training sample data. In the label probability adjacency matrix, the labeled nodes determine the new labels of the unlabeled nodes adjacent to the labeled nodes according to the propagation probability.
[0012] Obtain the label information of the unlabeled samples by comparing the labels of the labeled nodes with the information of the new labels.
[0013] Traverse each node in the label probability adjacency matrix, and use cost-sensitive learning to process the label information of unlabeled samples to obtain the fault diagnosis result.
[0014] In one embodiment, it further includes: obtaining current, voltage, power, and temperature as input photovoltaic power characteristics through a photovoltaic array, and using the fault types of photovoltaic panels as the output label information of the fault diagnosis model to construct training sample data.
[0015] In one embodiment, it further includes: inputting the training sample data into the fault diagnosis model for training, and constructing a kernel function using the improved label propagation method:
[0016] ;
[0017] Wherein, is the training sample data, is the labeled training sample data, is a constant set according to the adjacent nodes of the nodes in the training sample data.
[0018] Calculate the edge weight of the labeled samples in the training sample data propagating to the unlabeled samples adjacent to the labeled samples according to the kernel function:
[0019] ;
[0020] Wherein, is the edge weight between the labeled sample and the unlabeled sample adjacent to the labeled sample, is the labeled sample, is the unlabeled sample adjacent to the labeled sample, is the total number of edges of the nodes in the training sample data, is the sample variance, is the average distance from node i to node j.
[0021] Calculate the propagation probability of the unlabeled samples adjacent to the labeled samples propagating to the labeled samples according to the edge weight:
[0022] ;
[0023] Wherein, is the propagation probability of the unlabeled samples adjacent to the labeled samples propagating to the labeled samples, is the edge weight between the labeled sample and the unlabeled sample adjacent to the labeled sample, is the number of labeled samples, is a constant set according to the adjacent nodes of the nodes in the training sample data.
[0024] In one embodiment, it further includes: constructing a label probability adjacency matrix based on the number of labeled samples and unlabeled samples in the training sample data, where the labeled nodes in the label probability adjacency matrix perform weighted summation on the propagation probabilities of the unlabeled nodes adjacent to the labeled nodes according to the propagation probability to obtain a weighted probability. The unlabeled nodes adjacent to the labeled nodes are labeled with the label information corresponding to the maximum value of the weighted probability to obtain new labels for the unlabeled nodes.
[0025] In one embodiment, it further includes: if there are multiple maximum values of the weighted probability, randomly select the label information corresponding to one of the maximum values of the weighted probability to label the unlabeled nodes adjacent to the labeled nodes, and update the new labels of the unlabeled nodes.
[0026] In one embodiment, it further includes: by comparing the labels of the labeled nodes with the information of the new labels of the unlabeled nodes, if the labels of the labeled nodes are the same as the information of the new labels of the unlabeled nodes, and the information of the new labels of multiple unlabeled nodes adjacent to the labeled nodes is the same, then obtain the information of the new labels of multiple unlabeled nodes adjacent to the labeled nodes to obtain the label information of the unlabeled samples.
[0027] In one embodiment, the label information includes: shading fault, abnormal aging fault, hot spot fault, open circuit fault, and normal.
[0028] In one embodiment, it further includes: traversing each node in the label probability adjacency matrix, and constructing a total misclassification cost function for each node's label information using cost-sensitive learning:
[0029] ;
[0030] where is the cost matrix corresponding to the label information of the unlabeled samples, is the proportion of the label information of node i being predicted as the label information of node j among the samples where the label information of the unlabeled samples is predicted incorrectly, is the proportion of the label information of node j being predicted as the label information of node i among the samples where the label information of the unlabeled samples is predicted incorrectly;
[0031] Process the label information of the unlabeled samples according to the total misclassification cost function, and optimize the prediction risk using Bayesian risk theory to obtain the fault diagnosis result.
[0032] A photovoltaic array fault diagnosis device based on improved label propagation, the device includes:
[0033] A sample data acquisition module, configured to obtain training sample data through a photovoltaic array. The training sample data is photovoltaic electric energy, where the photovoltaic electric energy includes labeled samples and unlabeled samples.
[0034] A propagation probability acquisition module, which is used to input training sample data into a fault diagnosis model for training, and adopts an improved label propagation method to obtain the propagation probability of labeled samples propagating to unlabeled samples adjacent to the labeled samples.
[0035] A new label acquisition module, which is used to construct a label probability adjacency matrix according to the quantity of training sample data, and the labeled nodes in the label probability adjacency matrix determine the new labels of the unlabeled nodes adjacent to the labeled nodes according to the propagation probability.
[0036] A label information comparison module, which is used to obtain the label information of unlabeled samples by comparing the labels of labeled nodes with the information of the new labels.
[0037] A fault diagnosis result acquisition module, which is used to traverse each node in the label probability adjacency matrix, and uses cost-sensitive learning to process the label information of unlabeled samples to obtain the fault diagnosis result.
[0038] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0039] Training sample data is obtained through a photovoltaic array. The training sample data is photovoltaic electric energy, where the photovoltaic electric energy includes labeled samples and unlabeled samples.
[0040] The photovoltaic electric energy is input into a fault diagnosis model for training, and an improved label propagation method is adopted to obtain the propagation probability of labeled samples propagating to unlabeled samples adjacent to the labeled samples.
[0041] A label probability adjacency matrix is constructed according to the quantity of training sample data, and the labeled nodes in the label probability adjacency matrix determine the new labels of the unlabeled nodes adjacent to the labeled nodes according to the propagation probability.
[0042] The label information of unlabeled samples is obtained by comparing the labels of labeled nodes with the information of the new labels.
[0043] Each node in the label probability adjacency matrix is traversed, and cost-sensitive learning is used to process the label information of unlabeled samples to obtain the fault diagnosis result.
[0044] The above-mentioned photovoltaic array fault diagnosis method, device and equipment based on improved label propagation adopt a semi-supervised learning algorithm to model and identify photovoltaic array fault sample data. Learning the rules from a small amount of labeled data and then predicting unlabeled samples is beneficial to increasing the capacity of the data training set, and also greatly improves the model accuracy compared with unsupervised learning due to the inclusion of a small amount of labeled data. The improved label propagation algorithm is used to utilize a large number of unlabeled samples of the photovoltaic array, and the labels of a small amount of labeled data are propagated to similar samples to determine the fault types of unlabeled data. Aiming at the problem of high misjudgment rate of similar faults in the label propagation method, cost-sensitive learning is used to improve the fault diagnosis model, further improving the detection accuracy of similar fault types and accelerating the fault diagnosis efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 FIG. is a schematic flow chart of a photovoltaic array fault diagnosis method based on improved label propagation in an embodiment;
[0046] Figure 2 FIG. is a schematic flow chart of the steps of a photovoltaic array fault diagnosis based on the label propagation method in an embodiment;
[0047] Figure 3 FIG. is a schematic diagram of the diagnosis result of a fault diagnosis model based on the label propagation method in an embodiment, where Figure 3 (a) is the label of unlabeled data, Figure 3 (b) is the actual label;
[0048] Figure 4 FIG. is a schematic diagram of the fault diagnosis result of a photovoltaic array based on the improved label propagation method in an embodiment;
[0049] Figure 5 FIG. is a structural block diagram of a photovoltaic array fault diagnosis device based on improved label propagation in an embodiment;
[0050] Figure 6 FIG. is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0052] In one embodiment, as Figure 1 shown, a photovoltaic array fault diagnosis method based on improved label propagation is provided, including the following steps:
[0053] Step 102: Obtain training sample data through a photovoltaic array. The training sample data is photovoltaic electric energy, where the photovoltaic electric energy includes labeled samples and unlabeled samples.
[0054] Specifically, obtain the current, voltage, power, and temperature as the input photovoltaic electric energy characteristics through the photovoltaic array, and use the fault type of the photovoltaic panel as the output label information of the fault diagnosis model to construct the training sample data. In addition, the label information is used to mark the fault type of the photovoltaic array, and its fault types include: shading fault, abnormal aging fault, hot spot fault, open circuit fault, and normal.
[0055] Step 104: Input the training sample data into the fault diagnosis model for training, and use the improved label propagation method to obtain the propagation probability of the labeled samples propagating to the unlabeled samples adjacent to the labeled samples.
[0056] Specifically, input the training sample data into the fault diagnosis model for training, where the training sample data contains labeled samples and unlabeled samples. Each sample point is used as a node, set the number of iterations t = 1, and calculate the weight of each edge connecting the nodes using the weight formula:
[0057] ;
[0058] Among them, is the edge weight between the labeled sample and the unlabeled sample adjacent to the labeled sample, is the labeled sample, is the unlabeled sample adjacent to the labeled sample, is the total number of edges of the nodes in the training sample data, is the sample variance, is the average distance from node i to node j. represents the weight of the edge between two nodes i and j. The smaller the distance between the two nodes, the the larger the value. Each labeled node propagates to the remaining unlabeled nodes through adjacent points, and the node with a larger weight is more likely to affect the adjacent unlabeled nodes.
[0059] Furthermore, calculate the propagation probability from node j to i according to the obtained weight :
[0060] ;
[0061] Among them, is the propagation probability of the unlabeled sample adjacent to the labeled sample propagating to the labeled sample, is the edge weight between the labeled sample and the unlabeled sample adjacent to the labeled sample, where \(n\) is the number of labeled samples and \(l\) is the number of unlabeled samples. is a constant set according to the adjacent nodes of the nodes in the training sample data. is the edge weight between the constant \(k\) node and node \(j\).
[0062] Step 106: Construct a label probability adjacency matrix according to the number of training sample data. In the label probability adjacency matrix, the labeled nodes determine the new labels of the unlabeled nodes adjacent to the labeled nodes according to the propagation probability.
[0063] Specifically, a label probability adjacency matrix \(Y\) is generated according to the number of labeled samples and unlabeled samples in the training sample data, and its size is , the -th row represents the probability distribution of node . Add the labeled probability values propagated by its surrounding nodes to each node according to the propagation probability and weight, and update them to the probability distribution of the node. Then, select the highest probability to determine the new label of the unlabeled node. If there are multiple labels with the highest probability at the same time, randomly select one of them as the new label.
[0064] Step 108: Obtain the label information of the unlabeled samples by comparing the labels of the labeled nodes with the information of the new labels.
[0065] Step 110: Traverse each node in the label probability adjacency matrix, and use cost-sensitive learning to process the label information of the unlabeled samples to obtain the fault diagnosis result.
[0066] Specifically, if the label of each node appears most frequently among its adjacent nodes, the training stops. Otherwise, set , recalculate the weight probability, and iterate continuously until the model converges to obtain the final result: the fault type of each unlabeled node (i.e., the unlabeled fault sample); compare the model prediction result with the true fault category of the photovoltaic array, and calculate the fault detection rate.
[0067] Furthermore, different "costs" are set for different label information (i.e., fault types) for cost-sensitive learning. Traverse each node in the label probability adjacency matrix, and use cost-sensitive learning to construct the total misclassification cost function for the label information of each node:
[0068] ;
[0069] where is the cost matrix corresponding to the label information of the unlabeled samples, is the proportion that the label information of node \(i\) is predicted as the label information of node \(j\) among the samples in which the label information of the unlabeled samples is predicted incorrectly, It is the proportion that the label information of node j is predicted as the label information of node i among the samples whose label information of the unlabeled samples is predicted incorrectly.
[0070] Process the label information of the unlabeled samples according to the total misclassification cost function, and optimize the prediction risk using the Bayesian risk theory to obtain the fault diagnosis result.
[0071] The above photovoltaic array fault diagnosis method based on improved label propagation uses a semi-supervised learning algorithm to model and identify the photovoltaic array fault sample data. Learning the rules from a small amount of labeled data and then predicting the unlabeled samples is beneficial to increasing the capacity of the data training set, and also greatly improves the model accuracy compared with unsupervised learning because it contains a small amount of labeled data. The improved label propagation algorithm is used to utilize a large number of unlabeled samples of the photovoltaic array, propagate the labels of a small amount of labeled data to similar samples, and determine the fault types of the unlabeled data. Aiming at the problem of high misjudgment rate of similar faults in the label propagation method, the fault diagnosis model is improved using cost-sensitive learning, further improving the detection accuracy of similar fault types and accelerating the fault diagnosis efficiency.
[0072] In one embodiment, the current, voltage, power and temperature are obtained from the photovoltaic array as the input photovoltaic power characteristics, and the fault types of the photovoltaic panels are used as the output label information of the fault diagnosis model to construct the training sample data.
[0073] In one embodiment, the training sample data is input into the fault diagnosis model for training, and an improved label propagation method is used to construct the kernel function:
[0074] ;
[0075] Among them, is the training sample data, is the labeled training sample data, is a constant set according to the adjacent nodes of the nodes in the training sample data.
[0076] Calculate the edge weight that the labeled samples in the training sample data propagate to the unlabeled samples adjacent to the labeled samples according to the kernel function:
[0077] ;
[0078] Among them, is the edge weight between the labeled sample and the unlabeled sample adjacent to the labeled sample, is the labeled sample, is the unlabeled sample adjacent to the labeled sample, is the total number of edges of the nodes in the training sample data, is the sample variance, is the average distance from node i to node j.
[0079] Calculate the propagation probability that an unlabeled sample adjacent to a labeled sample propagates to the labeled sample according to the edge weight:
[0080] ;
[0081] where is the propagation probability that an unlabeled sample adjacent to a labeled sample propagates to the labeled sample, is the edge weight between a labeled sample and an unlabeled sample adjacent to the labeled sample, is the number of labeled samples, is a constant set according to the adjacent nodes of the nodes in the training sample data.
[0082] In one embodiment, construct a label probability adjacency matrix according to the number of labeled samples and unlabeled samples in the training sample data. In the label probability adjacency matrix, the labeled nodes perform weighted summation on the propagation probabilities of the unlabeled nodes adjacent to the labeled nodes according to the propagation probability to obtain the weighted probability. Take the label information corresponding to the maximum value of the weighted probability to label the unlabeled nodes adjacent to the labeled nodes, and obtain the new labels of the unlabeled nodes.
[0083] In one embodiment, if there are multiple maximum values of the weighted probability, randomly select the label information corresponding to one of the maximum values of the weighted probability to label the unlabeled nodes adjacent to the labeled nodes, and update the new labels of the unlabeled nodes.
[0084] In one embodiment, by comparing the labels of the labeled nodes with the information of the new labels of the unlabeled nodes, if the labels of the labeled nodes are the same as the information of the new labels of the unlabeled nodes, and the information of the new labels of multiple unlabeled nodes adjacent to the labeled nodes is the same, then obtain the information of the new labels of multiple unlabeled nodes adjacent to the labeled nodes to obtain the label information of the unlabeled samples.
[0085] In one embodiment, the label information includes: shading fault, abnormal aging fault, hot spot fault, open circuit fault, and normal.
[0086] In one embodiment, traverse each node in the label probability adjacency matrix, and use cost-sensitive learning to construct a total misclassification cost function for the label information of each node:
[0087] ;
[0088] where is the cost matrix corresponding to the label information of the unlabeled sample, For the samples in which the label information of the unlabeled samples is predicted incorrectly, the ratio of the label information of node i being predicted as the label information of node j For the samples in which the label information of the unlabeled samples is predicted incorrectly, the ratio of the label information of node j being predicted as the label information of node i
[0089] Process the label information of the unlabeled samples according to the total misclassification cost function, and optimize the prediction risk using Bayesian risk theory to obtain the fault diagnosis result
[0090] It should be noted that from the fault diagnosis results of the label propagation method, it can be seen that the actual fault types with misjudgment are shading fault, abnormal aging fault, and hot spot fault {3, 4, 5}, and the prediction results have four situations: normal, open circuit fault, shading fault, abnormal aging fault, and hot spot fault {2, 3, 4, 5}. Therefore, a cost matrix is defined for the misjudged fault types in the photovoltaic array to improve the accuracy of the misreported fault types and change the weight selection of adjacent nodes. The misclassification costs of different fault types in the photovoltaic array are shown in Table 1
[0091] Table 1 Misclassification costs of fault types in the photovoltaic array
[0092]
[0093] Therefore, in fault diagnosis, the misclassification function is defined as
[0094] ;
[0095] where is the probability that fault type i is misclassified as type j . The key to cost-sensitive learning is that the setting of the penalty can be determined according to the proportion of classification errors. Among the actually misreported fault types of the photovoltaic array based on the label propagation method, the number of samples actually being shading fault is 24, the number of samples actually being abnormal aging fault is 48, and the number of samples actually being hot spot fault is 24. The total number of misclassified samples is 101. Therefore, the initial misclassification proportion of each fault type can be used as a reference for weight setting, and the misclassification proportion of each fault type is
[0096] ;
[0097] ;
[0098] ;
[0099] ;
[0100] ;
[0101] ;
[0102] ;
[0103] 。
[0104] Therefore, set the cost matrix as follows:
[0105] 。
[0106] The fault diagnosis results based on the label propagation method are post - processed using cost - sensitive learning, and the Bayesian risk theory is used to adjust the results to achieve the minimum total cost loss. For example, for a sample , the original model believes that the probability that the fault type of sample is is , then the prediction risk of the Bayesian function that sample belongs to the fault type is:
[0107] ;
[0108] In one embodiment, as shown in Figure 2 , the specific steps of the photovoltaic array fault diagnosis method based on the label propagation method are as follows:
[0109] S1. Input the training data, including labeled samples and unlabeled samples. Each sample point is regarded as a node, and set the number of iterations t = 1.
[0110] S2. Calculate the similarity between data. Calculate the weight of each edge connecting nodes using the weight formula:
[0111] ;
[0112] Among them, is the edge weight between the labeled sample and the unlabeled sample adjacent to the labeled sample, is the labeled sample, is the unlabeled sample adjacent to the labeled sample, is the total number of edges of the nodes in the training sample data, is the sample variance, is the average distance from node i to node j. represents the weight of the edge between two nodes . The smaller the distance between the two nodes, The greater. Each labeled node propagates to the remaining unlabeled nodes through adjacent nodes. Nodes with larger weights are more likely to influence adjacent unlabeled nodes.
[0113] S3. Define a label matrix of size . Its th row represents the probability distribution of node . Add the labeled probability values propagated by its surrounding nodes to this node's probability distribution according to the propagation probability, and update the probability distribution of this node.
[0114] S4. Select the new label of the node according to the highest probability. If there are multiple labels with the highest probability at the same time, randomly select one of them as the label.
[0115] S5. If the label of each node appears most frequently among its adjacent nodes, the algorithm stops. Otherwise, set , return to step (3), and iterate continuously until the model converges to obtain the final result: the fault type of each unlabeled node (i.e., the unlabeled fault sample); compare the model prediction result with the true fault category of the photovoltaic array, and calculate the fault detection rate.
[0116] In one embodiment, for the imbalanced dataset, the SMOTE sampling method is used to expand the number of open - circuit and short - circuit faults to 141 each, copy and expand the number of normal samples to 100, and then perform de - labeling processing on the dataset. The data composition after processing is shown in Table 2, which is used as the original dataset of photovoltaic array fault samples for input:
[0117] Table 2 Data composition before and after de - labeling
[0118]
[0119] First, use the fault diagnosis model based on the label propagation method to train the photovoltaic array simulation fault sample data. The training time of the fault diagnosis model based on the label propagation method is 0.61836 s, and the fault diagnosis time is 0.01322 s. The training time is relatively long because the model training set contains two thousand fault samples. It is necessary to first learn the labeled dataset and then predict the labels of the unlabeled data. Since the scale of the training set is large, the training time increases, but the training is still completed within 1 second, and there is no obvious difference in the actual running time. Compare the predicted labels of the unlabeled data with the actual labels. The specific results of the fault diagnosis are as Figure 3 shown, Figure 3 (a) is the label of the unlabeled data, Figure 3(b) is the actual label. It can be seen that the abnormal aging faults are misjudged as open circuit faults, shading faults, and hot spot faults respectively. Among them, the possibility of being misjudged as a hot spot fault is the greatest. The specific situations of the misdiagnosed faults are shown in Table 3:
[0120] Table 3 Specific situations of misjudgment of the fault diagnosis model based on the label propagation method
[0121]
[0122] In addition, as Figure 4 shown, when the labeled samples only account for 1 / 6, while the diagnostic accuracy of the fault diagnosis model is above 90%, the model training and prediction times are also relatively fast. For the misjudgments of abnormal aging faults and hot spot faults, cost-sensitive learning is used to improve the model to enhance the overall fault recognition rate. The calculation results of the evaluation indexes of the photovoltaic array fault diagnosis model based on the improved label propagation method and the traditional label propagation method are shown in Table 4:
[0123] Table 4 Evaluation indexes of the fault diagnosis model based on the improved label propagation method
[0124]
[0125] It can be seen from the fault diagnosis results that although the overall accuracy rate of the photovoltaic array fault diagnosis model based on the label propagation method reaches above 0.9, the misjudgment rates for the three types of faults of shading, abnormal aging, and hot spots are very high, which are 0.9118, 0.8235, and 0.8885 respectively. This means that the fault diagnosis model based on the traditional label propagation method cannot accurately distinguish similar faults of the photovoltaic array, and the misjudgment rate is relatively high. The fault detection rate of the fault diagnosis model improved by cost-sensitive learning reaches 100%, which means that all faults have been detected. The detection rates of the three types of faults of shading, abnormal aging, and hot spots have been significantly improved, and the fault detection rate reaches above 0.96. The accuracy rate of the improved diagnosis model has also been increased to 98.14%. However, due to the fact that the output characteristics of some fault samples of shadow faults, abnormal aging faults, and hot spot faults are very similar in some cases, there are still misjudgments in the improved model for these three types of faults.
[0126] It should be understood that although Figure 1 - Figure 2 the steps in the flowchart of Figure 1 - Figure 2At least a part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0127] In one embodiment, as Figure 5 shown, a photovoltaic array fault diagnosis device based on improved label propagation is provided, including: a sample data acquisition module 502, a propagation probability acquisition module 504, a new label acquisition module 506, a label information comparison module 508, and a fault diagnosis result acquisition module 510, where:
[0128] The sample data acquisition module 502 is configured to obtain training sample data through a photovoltaic array. The training sample data is photovoltaic electric energy, where the photovoltaic electric energy includes labeled samples and unlabeled samples.
[0129] The propagation probability acquisition module 504 is configured to input the training sample data into a fault diagnosis model for training, and use the improved label propagation method to obtain the propagation probability of the labeled samples propagating to the unlabeled samples adjacent to the labeled samples.
[0130] The new label acquisition module 506 is configured to construct a label probability adjacency matrix according to the quantity of the training sample data. In the label probability adjacency matrix, the labeled nodes determine the new labels of the unlabeled nodes adjacent to the labeled nodes according to the propagation probability.
[0131] The label information comparison module 508 is configured to obtain the label information of the unlabeled samples by comparing the labels of the labeled nodes with the information of the new labels.
[0132] The fault diagnosis result acquisition module 510 is configured to traverse each node in the label probability adjacency matrix, and use cost-sensitive learning to process the label information of the unlabeled samples to obtain a fault diagnosis result.
[0133] For the specific limitations of the photovoltaic array fault diagnosis device based on improved label propagation, reference can be made to the limitations of the photovoltaic array fault diagnosis method based on improved label propagation in the above text, which will not be elaborated here. Each module in the above photovoltaic array fault diagnosis device based on improved label propagation can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0134] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in Figure 6 . The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a photovoltaic array fault diagnosis method based on improved label propagation. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0135] Those skilled in the art can understand that Figure 5 - Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0136] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:
[0137] Obtain training sample data through a photovoltaic array. The training sample data is photovoltaic electric energy, where the photovoltaic electric energy includes labeled samples and unlabeled samples.
[0138] Input the photovoltaic electric energy into a fault diagnosis model for training, and use the improved label propagation method to obtain the propagation probability of the labeled samples propagating to the unlabeled samples adjacent to the labeled samples.
[0139] Construct a label probability adjacency matrix according to the quantity of the training sample data. In the label probability adjacency matrix, the labeled nodes determine the new labels of the unlabeled nodes adjacent to the labeled nodes according to the propagation probability.
[0140] By comparing the labels of the labeled nodes with the information of the new labels, obtain the label information of the unlabeled samples.
[0141] Traverse each node in the label probability adjacency matrix, and use cost-sensitive learning to process the label information of the unlabeled samples to obtain a fault diagnosis result.
[0142] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0143] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0144] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A photovoltaic array fault diagnosis method based on improved label propagation, characterized in that The method includes: Obtaining training sample data through a photovoltaic array; the training sample data is photovoltaic electric energy, where the photovoltaic electric energy includes labeled samples and unlabeled samples; Inputting the training sample data into a fault diagnosis model for training, and using an improved label propagation method to obtain the propagation probability of the labeled samples propagating to the unlabeled samples adjacent to the labeled samples; inputting the training sample data into a fault diagnosis model for training, and using an improved label propagation method to construct a kernel function: ; Among them, is the training sample data, is the labeled training sample data, is a constant set according to the adjacent nodes of the nodes in the training sample data; Calculating the edge weights of the labeled samples in the training sample data propagating to the unlabeled samples adjacent to the labeled samples according to the kernel function: ; Among them, is the edge weight between the labeled sample and the unlabeled sample adjacent to the labeled sample, is the labeled sample, is the unlabeled sample adjacent to the labeled sample, is the total number of edges of the nodes in the training sample data, is the sample variance, is the average distance from node i to node j; Calculating the propagation probability of the unlabeled samples adjacent to the labeled samples propagating to the labeled samples according to the edge weights: ; Among them, is the propagation probability that the unlabeled samples adjacent to the labeled samples are propagated to the labeled samples, is the edge weight between the labeled samples and the unlabeled samples adjacent to the labeled samples, is the number of the labeled samples, is a constant set according to the adjacent nodes of the nodes in the training sample data; Constructing a label probability adjacency matrix according to the quantity of the training sample data, and determining new labels of the unlabeled nodes adjacent to the labeled nodes by the labeled nodes in the label probability adjacency matrix according to the propagation probability; Obtaining the label information of the unlabeled samples by comparing the labels of the labeled nodes with the information of the new labels; Traversing each node in the label probability adjacency matrix, and processing the label information of the unlabeled samples by using cost-sensitive learning to obtain a fault diagnosis result.
2. The method according to claim 1, wherein Obtaining training sample data through a photovoltaic array, including: Obtaining current, voltage, power and temperature as input photovoltaic electric energy features through a photovoltaic array, and constructing training sample data with the fault types of photovoltaic panels as the output label information of the fault diagnosis model.
3. The method according to claim 1, characterized in that, Constructing a label probability adjacency matrix according to the quantity of the training sample data, and determining new labels of the unlabeled nodes adjacent to the labeled nodes by the labeled nodes in the label probability adjacency matrix according to the propagation probability, including: Constructing a label probability adjacency matrix according to the quantity of the labeled samples and the quantity of the unlabeled samples in the training sample data, and performing weighted summation on the propagation probabilities of the unlabeled nodes adjacent to the labeled nodes by the labeled nodes in the label probability adjacency matrix according to the propagation probability to obtain a weighted probability; Labeling the unlabeled nodes adjacent to the labeled nodes with the label information corresponding to the maximum value of the weighted probability to obtain the new labels of the unlabeled nodes.
4. The method according to claim 3, characterized in that, After the step of labeling the unlabeled nodes adjacent to the labeled nodes with the label information corresponding to the maximum value of the weighted probability to obtain the new labels of the unlabeled nodes, it further includes: If there are multiple maximum values of the weighted probability, randomly selecting one of the label information corresponding to the maximum value of the weighted probability to label the unlabeled nodes adjacent to the labeled nodes, and updating the new labels of the unlabeled nodes.
5. The method according to claim 4, wherein Obtaining the label information of the unlabeled samples by comparing the labels of the labeled nodes with the information of the new labels, including: By comparing the labels of the labeled nodes with the information of the new labels of the unlabeled nodes, if the labels of the labeled nodes are the same as the information of the new labels of the unlabeled nodes, and the information of the new labels of multiple unlabeled nodes adjacent to the labeled nodes is the same, then obtain the information of the new labels of multiple unlabeled nodes adjacent to the labeled nodes to obtain the label information of the unlabeled samples.
6. The method according to any one of claims 1 to 5, characterized in that The label information includes: shading fault, abnormal aging fault, hot spot fault, open circuit fault, and normal.
7. The method according to claim 6, wherein Traverse each node in the label probability adjacency matrix, and use cost-sensitive learning to process the label information of the unlabeled samples to obtain a fault diagnosis result, including: Traverse each node in the label probability adjacency matrix, and use cost-sensitive learning to construct a total misclassification cost function for the label information of each node: ; Among them, is the cost matrix corresponding to the label information of the unlabeled sample, is the proportion of the label information of node i being predicted as the label information of node j among the samples with the label information of the unlabeled sample being predicted wrongly, is the proportion of the label information of node j being predicted as the label information of node i among the samples with the label information of the unlabeled sample being predicted wrongly; Process the label information of the unlabeled samples according to the total misclassification cost function, and optimize the prediction risk using Bayesian risk theory to obtain a fault diagnosis result.
8. A photovoltaic array fault diagnosis device based on improved label propagation, which is used to implement the method described in any one of claims 1 to 7, and is characterized in that, The device includes: A sample data acquisition module, configured to acquire training sample data through a photovoltaic array; the training sample data is photovoltaic electric energy, where the photovoltaic electric energy includes labeled samples and unlabeled samples; A propagation probability acquisition module, configured to input the training sample data into a fault diagnosis model for training, and use an improved label propagation method to obtain the propagation probability of the labeled samples propagating to the unlabeled samples adjacent to the labeled samples; A new label acquisition module, configured to construct a label probability adjacency matrix according to the quantity of the training sample data, and determine the new labels of the unlabeled nodes adjacent to the labeled nodes by the labeled nodes in the label probability adjacency matrix according to the propagation probability; A label information comparison module, configured to obtain the label information of the unlabeled samples by comparing the labels of the labeled nodes with the information of the new labels; A fault diagnosis result acquisition module, configured to traverse each node in the label probability adjacency matrix, and use cost-sensitive learning to process the label information of the unlabeled samples to obtain a fault diagnosis result.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Artificial immune detector training method based on label propagation
CN113378997A
Adaptive manifold probability distribution-based bearing fault diagnosis method
US20220373430A1