Solar photovoltaic array fault diagnosis method and device

By combining IV curves and additional parameter information, using sequence floating forward selection algorithm and weighted ensemble learning model, the photovoltaic array fault diagnosis is optimized, and the problem of insufficient real-time and accuracy of photovoltaic system fault diagnosis in the existing technology is solved, achieving efficient and reliable fault identification and classification.

CN120508875APending Publication Date: 2025-08-19CHINA THREE GORGES CORPORATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510589874.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing photovoltaic system fault diagnosis methods have shortcomings in real-time, accuracy and reliability, especially in complex environmental conditions, and it is difficult to effectively classify and identify fault types.

Method used

The IV curve and additional parameter information of the photovoltaic array are obtained, and the data set is optimized through feature extraction and sequence floating forward selection algorithm, and the weighted ensemble learning model of logistic regression, support vector machine and naive Bayes classifier are combined for fault diagnosis. The genetic algorithm is used to optimize the classifier weight to build a multi-level fault diagnosis model.

Benefits of technology

It improves the real-time, accuracy and reliability of photovoltaic array fault diagnosis, can more comprehensively analyze different types of faults and their severity, adapt to dynamic changes, and improves the accuracy and meticulousness of fault classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508875A_ABST
    Figure CN120508875A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of photovoltaic power generation, and discloses a solar photovoltaic array fault diagnosis method and device, and the method comprises the steps: obtaining an IV curve and additional parameter information of a target photovoltaic array, carrying out the feature extraction based on the IV curve and the additional parameter information, and obtaining an initial data set; performing feature selection on the initial data set by using a sequence floating forward selection algorithm to obtain an optimized data set; training a weighted ensemble learning model by using the optimized data set to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model comprises a logistic regression classifier, a support vector machine classifier and a naive Bayes classifier; and performing fault diagnosis on the solar photovoltaic array by using the trained weighted integration model to obtain a fault diagnosis result of the solar photovoltaic array. The method improves the real-time performance, accuracy and reliability of fault diagnosis of the solar photovoltaic array, and is of great significance to development and application of fault diagnosis of a photovoltaic system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photovoltaic power generation, and in particular to a solar photovoltaic array fault diagnosis method and device. Background Art

[0002] With the continuous growth of global energy demand, solar photovoltaic technology has been widely used as a clean and renewable energy form. However, photovoltaic systems will inevitably encounter various electrical faults during long-term operation, such as abnormal current and voltage, component damage, poor connection, etc. These faults not only affect the efficiency and stability of the system, but may also pose a serious threat to safety. Therefore, timely and accurate diagnosis and classification of faults in photovoltaic systems are crucial to ensuring the normal operation of photovoltaic systems, improving energy utilization, and reducing maintenance costs.

[0003] Due to the expansion of photovoltaic system scale and the increase of complexity, solar photovoltaic array fault diagnosis methods still face challenges in terms of real-time performance, accuracy and reliability. Summary of the Invention

[0004] In view of this, the present invention provides a solar photovoltaic array fault diagnosis method and device to solve the problems of low real-time performance, accuracy and reliability of the solar photovoltaic array fault diagnosis method.

[0005] In a first aspect, the present invention provides a solar photovoltaic array fault diagnosis method, the method comprising:

[0006] Obtaining the IV curve and additional parameter information of the target photovoltaic array, performing feature extraction based on the IV curve and additional parameter information, and obtaining an initial data set;

[0007] The sequence floating forward selection algorithm is used to perform feature selection on the initial data set to obtain the optimized data set;

[0008] The weighted ensemble learning model is trained using the optimized data set to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model includes a logistic regression classifier, a support vector machine classifier, and a naive Bayes classifier;

[0009] The trained weighted ensemble model is used to diagnose the faults of the solar photovoltaic array and obtain the fault diagnosis results of the solar photovoltaic array.

[0010] This embodiment provides a solar photovoltaic array fault diagnosis method, which obtains an IV curve and additional parameter information of a target photovoltaic array, performs feature extraction based on the IV curve and the additional parameter information to obtain an initial data set; uses a sequential floating forward selection algorithm to perform feature selection on the initial data set to obtain an optimized data set; uses the optimized data set to train a weighted ensemble learning model to obtain a trained weighted ensemble model; and uses the trained weighted ensemble model to perform fault diagnosis on the solar photovoltaic array, thereby improving the real-time, accuracy, and reliability of solar photovoltaic array fault diagnosis, and has important significance for the development and application of photovoltaic system fault diagnosis.

[0011] In an optional embodiment, feature extraction is performed based on the IV curve and additional parameter information to obtain an initial data set, including:

[0012] Select multiple characteristic points on the IV curve of the target photovoltaic array;

[0013] Perform feature extraction on multiple feature points and additional parameter information to obtain multiple photovoltaic array feature data;

[0014] An initial dataset is constructed based on multiple photovoltaic array characteristic data.

[0015] This embodiment provides a solar photovoltaic array fault diagnosis method that extracts features from multiple feature points and additional parameter information to obtain multiple photovoltaic array feature data, thereby accurately extracting photovoltaic array features and constructing an initial data set based on the photovoltaic array feature data, so that the initial data set can better reflect the characteristics of the photovoltaic array.

[0016] In an optional embodiment, the initial data set is subjected to feature selection using a sequential floating forward selection algorithm to obtain an optimized data set, including:

[0017] Normalize the initial dataset;

[0018] A sequential forward selection algorithm is used to select features from the normalized initial data set, and the selected features are added to the initial feature subset to obtain a candidate feature subset; wherein the initial feature subset is an empty set;

[0019] If the number of features in the candidate feature subset is greater than the preset threshold, the features in the candidate feature subset are deleted in sequence using the sequential backward selection algorithm to obtain multiple current feature subsets;

[0020] Perform classification accuracy evaluation on multiple current feature subsets, and screen the multiple current feature subsets based on the classification accuracy evaluation results to obtain the optimal feature subset;

[0021] If the number of features in the optimal feature subset is equal to the preset dimension, the optimal feature subset is used as the optimized data set.

[0022] A solar photovoltaic array fault diagnosis method provided in this embodiment uses an appropriate normalization method to normalize an initial data set, thereby ensuring the consistency of features of different dimensions, optimizing data distribution, reducing algorithm computational complexity, and accelerating the convergence process of a weighted ensemble model. A sequential floating feature selection algorithm is used on the normalized initial data set to screen high-dimensional features. By dynamically updating feature subsets and weighing redundancy and correlation, core features closely related to photovoltaic fault classification are ultimately screened out to form an optimized data set. This reduces the dimension of the data set, simplifies the computational process, and improves the training efficiency of the weighted ensemble learning model.

[0023] In an optional embodiment, the weighted ensemble learning model is trained using the optimized data set to obtain a trained weighted ensemble model, including:

[0024] The optimized data set is input into the logistic regression classifier, support vector machine classifier and naive Bayes classifier respectively to obtain multiple classification prediction results;

[0025] The weighted voting algorithm is used to combine multiple classification prediction results to obtain the fault classification prediction result;

[0026] If the fault classification prediction result is different from the fault diagnosis result in the optimized data set, the genetic algorithm is used to iteratively optimize the weights of the logistic regression classifier, support vector machine classifier, and naive Bayes classifier until the fault classification prediction result is the same as the fault diagnosis result, and the trained weighted ensemble model is obtained.

[0027] This embodiment provides a solar photovoltaic array fault diagnosis method, in which a weighted integrated model combines the advantages of a logistic regression classifier, a support vector machine classifier, and a naive Bayes classifier. The weight of each classifier is optimized through a genetic algorithm, thereby improving the overall accuracy of the weighted integrated model, optimizing resource allocation, improving the fault diagnosis efficiency of the weighted integrated model, and adapting to the dynamic changes of fault problems. It can achieve a high accuracy rate with a small amount of training data. The trained weighted integrated model can effectively classify and detect various types of faults and their severity.

[0028] In an optional embodiment, the method further includes:

[0029] The trained weighted ensemble learning model is verified and evaluated, and the trained weighted ensemble learning model is optimized based on the verification and evaluation results.

[0030] The solar photovoltaic array fault diagnosis method provided in this embodiment realizes comprehensive verification and evaluation of the trained weighted ensemble learning model, further improving the accuracy and reliability of fault diagnosis of the weighted ensemble learning model.

[0031] In an optional embodiment, performing verification and evaluation on the trained weighted ensemble learning model, and optimizing the trained weighted ensemble learning model based on the verification and evaluation results, includes:

[0032] The feasibility of the trained weighted ensemble learning model is verified by using K-fold cross validation combined with confusion matrix to obtain the feasibility verification results of the model;

[0033] Obtaining performance evaluation indicators, and using the performance evaluation indicators to perform performance evaluation on the trained weighted ensemble learning model to obtain performance evaluation results;

[0034] The trained weighted ensemble learning model is optimized based on the model feasibility verification results and performance evaluation results.

[0035] This embodiment provides a solar photovoltaic array fault diagnosis method that uses K-fold cross-validation technology or a confusion matrix to verify the feasibility of a weighted ensemble learning model, and uses performance evaluation indicators to perform performance evaluation on the trained weighted ensemble learning model, thereby further improving the fault diagnosis accuracy of the weighted ensemble learning model.

[0036] In a second aspect, the present invention provides a solar photovoltaic array fault diagnosis device, the device comprising:

[0037] A feature extraction module is used to obtain the IV curve and additional parameter information of the target photovoltaic array, perform feature extraction based on the IV curve and additional parameter information, and obtain an initial data set;

[0038] The feature selection module is used to select features of the initial data set using the sequential floating forward selection algorithm to obtain an optimized data set;

[0039] A training module is used to train the weighted ensemble learning model using the optimized data set to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model includes a logistic regression classifier, a support vector machine classifier, and a naive Bayes classifier;

[0040] The fault diagnosis module is used to perform fault diagnosis on the solar photovoltaic array using the trained weighted integrated model to obtain the solar photovoltaic array fault diagnosis result.

[0041] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the solar photovoltaic array fault diagnosis method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0042] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the solar photovoltaic array fault diagnosis method of the first aspect or any corresponding embodiment thereof.

[0043] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the solar photovoltaic array fault diagnosis method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 is a schematic flow chart of a solar photovoltaic array fault diagnosis method according to an embodiment of the present invention;

[0046] Figure 2 is a schematic diagram of a workflow for fault diagnosis using an integrated learning model according to an embodiment of the present invention;

[0047] Figure 3 is a schematic diagram of a fault diagnosis result of a first-layer solar photovoltaic array according to an embodiment of the present invention;

[0048] Figure 4 is a schematic diagram of a fault diagnosis result of a second-layer solar photovoltaic array according to an embodiment of the present invention;

[0049] Figure 5 is a schematic diagram of a fault diagnosis result of a third-layer solar photovoltaic array according to an embodiment of the present invention;

[0050] Figure 6 is a schematic diagram of a fault diagnosis result of a fourth-layer solar photovoltaic array according to an embodiment of the present invention;

[0051] Figure 7 is a schematic diagram of a fault diagnosis result of a fifth-layer solar photovoltaic array according to an embodiment of the present invention;

[0052] Figure 8 is a schematic diagram of a fault diagnosis result of a sixth-layer solar photovoltaic array according to an embodiment of the present invention;

[0053] Figure 9 is a flow chart of another solar photovoltaic array fault diagnosis method according to an embodiment of the present invention;

[0054] Figure 10 is a schematic diagram of characteristic points of a target photovoltaic array IV curve according to an embodiment of the present invention;

[0055] Figure 11 is a schematic flow chart of a sequential forward selection algorithm according to an embodiment of the present invention;

[0056] Figure 12 is a flow chart of another solar photovoltaic array fault diagnosis method according to an embodiment of the present invention;

[0057] Figure 13 is a flow chart of another solar photovoltaic array fault diagnosis method according to an embodiment of the present invention;

[0058] Figure 14 is a schematic diagram of the working process of a solar photovoltaic array fault diagnosis method according to an embodiment of the present invention;

[0059] Figure 15 This is a structural block diagram of a solar photovoltaic array fault diagnosis device according to an embodiment of the present invention;

[0060] Figure 16 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0061] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0062] At present, photovoltaic system fault diagnosis has become a research hotspot. The photovoltaic system fault diagnosis methods mainly include manual inspection method, physical model method, data-driven method and intelligent optimization algorithm; among them, data-driven method includes machine learning method and weighted ensemble learning method.

[0063] Although the above methods have achieved certain success in fault diagnosis, with the expansion of the scale and increase in complexity of photovoltaic systems, these methods still face challenges in terms of real-time performance, accuracy, and reliability. For example, complex environmental conditions (such as shadows, dust, and temperature changes) and diverse fault types place higher demands on diagnostic models. In addition, most of the above diagnostic methods rely on fixed feature sets and algorithms, lacking sufficient adaptability and generalization capabilities.

[0064] To solve the above technical problems, an embodiment of the present invention provides a solar photovoltaic array fault diagnosis method, which combines the advantages of a data-driven method and an intelligent optimization algorithm, and proposes a comprehensive multi-layer model for detecting, classifying and identifying the severity of electrical faults in photovoltaic systems. The comprehensive multi-layer model can more comprehensively analyze different types of faults and refine the classification and severity of the faults. Among them, a genetic algorithm is used to optimize the weights of the three classifiers in the weighted ensemble learning algorithm, thereby improving the overall accuracy of the model. In each layer, a sequential floating forward selection algorithm is used for feature selection, thereby reducing the dimension of the data set, simplifying the calculation process, and improving the training efficiency. In general, a new diagnostic technology is provided to improve the accuracy of solar photovoltaic array fault diagnosis and the precision and meticulousness of fault classification, and has achieved remarkable success in experimental verification. It is of great significance to the development and application of photovoltaic system fault diagnosis, and points out the direction for future photovoltaic fault diagnosis research.

[0065] An embodiment of the present invention provides a method for diagnosing solar photovoltaic array faults. It should be noted that the method for diagnosing solar photovoltaic array faults provided by the embodiment of the present invention may be executed by a device for diagnosing solar photovoltaic array faults. The device for diagnosing solar photovoltaic array faults may be implemented as part or all of an electronic device through software, hardware, or a combination of software and hardware. The electronic device may be a server or a terminal. The server in the embodiment of the present application may be a single server or a server cluster composed of multiple servers. The terminal in the embodiment of the present application may be a smart phone, a personal computer, a tablet computer, a wearable device, an intelligent robot, or other intelligent hardware devices. In the following method embodiments, the execution subject is an electronic device as an example for explanation.

[0066] According to an embodiment of the present invention, an embodiment of a solar photovoltaic array fault diagnosis method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0067] In this embodiment, a solar photovoltaic array fault diagnosis method is provided, which can be used for the above-mentioned electronic equipment. Figure 1 FIG. 1 is a flow chart of a solar photovoltaic array fault diagnosis method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0068] Step S101 : obtaining an IV curve and additional parameter information of a target photovoltaic array, performing feature extraction based on the IV curve and the additional parameter information, and obtaining an initial data set.

[0069] Specifically, five representative statistical features (such as the peak values of current and voltage, slope change points, etc.) are selected from the IV (Intensity of Current Voltage) curve of the target photovoltaic array, and additional parameter information such as the brand parameters of the photovoltaic array are integrated as the basis for feature extraction to construct the initial data set.

[0070] Step S102 : performing feature selection on the initial data set using a sequential floating forward selection algorithm to obtain an optimized data set.

[0071] Specifically, the Sequential Floating Forward Selection (SFFS) algorithm is used to screen high-dimensional features. By dynamically updating feature subsets and weighing redundancy and correlation, the core features closely related to photovoltaic fault classification are finally screened out to form an optimized dataset.

[0072] Furthermore, the SFFS algorithm is used for feature selection. The SFFS algorithm is a wrapper feature selection technology that combines sequential forward selection (SFS) and sequential backward selection (SBS). Features are added through SFS and deleted through SBS (features float, not constantly increase or decrease).

[0073] Step S103, using the optimized data set to train the weighted ensemble learning model to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model includes a logistic regression classifier, a support vector machine classifier and a naive Bayes classifier.

[0074] Specifically, based on the optimized dataset, the prediction results of the logistic regression classifier, the support vector machine classifier and the naive Bayes classifier were combined, and the classifier weights were optimized using a weighted voting scheme and a genetic algorithm to obtain a trained weighted ensemble model.

[0075] Step S104 , performing fault diagnosis on the solar photovoltaic array using the trained weighted integrated model to obtain a solar photovoltaic array fault diagnosis result.

[0076] Specifically, if Figure 2 As shown in the figure, a weighted integrated model is used to determine whether the solar photovoltaic array has a current-based fault. If it is not a current-based fault, it is further determined whether it is a normal situation. If it is not a normal situation, the fault needs to be classified as a series attenuation fault or an open circuit fault. If it is determined to be a current-based fault, it is necessary to further determine the specific fault type, including line-to-line fault, line-to-ground fault, and array attenuation fault. Among them, line-to-line fault and line-to-ground fault need to be further classified according to the specific mismatch ratio.

[0077] Furthermore, the weighted integrated model is a multi-layer model used to detect, classify, and identify the severity of electrical faults in photovoltaic systems. Photovoltaic system electrical faults include: LL (Line-Line), LG (Line-Ground), OC (Open Circuited), and string array attenuation faults. The multi-layer structure of the weighted integrated model can more comprehensively analyze different types of faults and refine the fault classification and severity.

[0078] Furthermore, the multi-level structure of the weighted integration model includes: Figure 3 As shown in Figure 1, the first layer: Fault Current-Based Separator (FCBS) uses a weighted integrated model to classify data samples into two categories: 1) low-current faults, including normal conditions, open circuit (OC) faults, and series attenuation faults; 2) high-current faults, including array attenuation faults and line-to-line (LL) and line-to-ground (LG) faults; Figure 4 As shown, for the low current fault identified by the first layer ("No" event), the second layer distinguishes between normal conditions and low current faults; Figure 5 As shown in , the third layer further classifies low current faults into open circuit (OC) faults and series decay faults; Figure 6 As shown in FIG, a weighted integrated model is used at the fourth level to classify high current faults (“yes” events) into LL faults, LG faults, and array attenuation faults; Figure 7 As shown in , the fifth layer uses a weighted integrated model to identify the severity of LL faults, which are divided into 10%, 20% and more than 20% mismatch levels; Figure 8 As shown, the sixth layer uses a weighted integrated model to classify the severity of LG faults into 50%, 40% and less than 40% mismatch levels.

[0079] This embodiment provides a solar photovoltaic array fault diagnosis method, which obtains an IV curve and additional parameter information of a target photovoltaic array, performs feature extraction based on the IV curve and the additional parameter information, and obtains an initial data set; uses a sequential floating forward selection algorithm to perform feature selection on the initial data set to obtain an optimized data set; uses the optimized data set to train a weighted ensemble learning model to obtain a trained weighted ensemble model; uses the trained weighted ensemble model to perform fault diagnosis on the solar photovoltaic array, accurately obtains solar photovoltaic array fault diagnosis results, improves the real-time, accuracy, and reliability of solar photovoltaic array fault diagnosis, and has important significance for the development and application of photovoltaic system fault diagnosis.

[0080] In this embodiment, a solar photovoltaic array fault diagnosis method is provided, which can be used in the above-mentioned mobile terminals, such as mobile phones, tablet computers, etc. Figure 9 FIG. 1 is a flow chart of a solar photovoltaic array fault diagnosis method according to an embodiment of the present invention. Figure 9 As shown, the process includes the following steps:

[0081] Step S901 : obtaining an IV curve and additional parameter information of a target photovoltaic array, performing feature extraction based on the IV curve and the additional parameter information, and obtaining an initial data set.

[0082] Specifically, the above step S901 includes:

[0083] Step S9011: Select multiple feature points on the IV curve of the target photovoltaic array.

[0084] Specifically, if Figure 10 As shown in the figure, the five characteristic points selected on the IV curve of the target photovoltaic array are A(0,I sc )、B(V oc / 2,I Voc / 2 )、C(V MPP ,I MPP )、D(V Isc / 2 ,I sc / 2) and E(V oc ,0), where I sc Indicates short-circuit current, V oc Indicates the open circuit voltage, I Voc / 2 Indicates the current corresponding to half the open circuit voltage, V MPP and I MPP Represent the voltage and current at the maximum power point, V Isc / 2 Indicates the voltage corresponding to half the short-circuit current, I sc / 2 means half of the short-circuit current.

[0085] Step S9012: extract features from the multiple feature points and additional parameter information to obtain multiple photovoltaic array feature data.

[0086] Specifically, the additional parameter information includes I sc(STC) 、V oc(STC) , I MPP(STC) and V MPP(STC) , where I sc(STC) 、V oc(STC) are the short-circuit current and open-circuit voltage under standard test conditions, I MPP(STC) 、V MPP(STC) They are respectively the voltage and current at the maximum power point under standard test conditions (STC).

[0087] Furthermore, feature extraction is performed based on the five feature points selected above and the additional parameter information, and a total of 16 photovoltaic array feature data are extracted. The 16 photovoltaic array feature data are as follows:

[0088] f1=I sc / I sc(STC) (1)

[0089] f2=V oc / V oc(STC) (2)

[0090] f3=V MPP / V MPP(STC) (3)

[0091] f4=I MPP / I MPP(STC) (4)

[0092]

[0093] f 10 =(I MPP -I sc ) / (V MPP ) (10)

[0094] f 11 =(-I MPP ) / (V oc -V MPP ) (11)

[0095]

[0096]

[0097] In the above formula, f1 is the normalized ratio of the short-circuit current to the standard test conditions, which is used to evaluate the attenuation of the short-circuit current and may be affected by light intensity, temperature, component aging or shading; f2 is the normalized ratio of the open-circuit voltage to the standard test conditions, which is used to monitor the impact of temperature or aging on the open-circuit voltage; f3 is the maximum power point voltage (V MPP ) is the normalized ratio relative to the standard test conditions and is used to analyze the performance changes of the photovoltaic array. Component aging or internal loss may cause the MPP (Maximum Power Point) voltage drop; f4 is the current at the MPP (I MPP ) relative to the normalized ratio under standard test conditions, combined with f3, can be used to analyze the changes in the maximum power output capacity of the photovoltaic array under different environmental conditions; f5 is the ratio of the current at half-open circuit voltage to the standard short-circuit current, which is used to analyze the changes in the series resistance of the photovoltaic cell. If f5 decreases, it may mean that there is a hidden crack inside the cell or the series resistance increases; f6 is the ratio of the voltage corresponding to the half-short circuit current to the standard open circuit voltage, which is used to monitor the changes in the nonlinear characteristics of the photovoltaic cell. If the ratio is abnormal, there may be a fault or aging inside the component; f7 is the ratio of the normalized MPP current to the MPP voltage, which is used to reflect the relative change between the current and voltage at the maximum power point; f8 is the ratio of the normalized MPP voltage to the normalized open circuit voltage. By comparing the ratio of the maximum power point voltage to the open circuit voltage in actual operation with the corresponding ratio under standard test conditions, it is possible to diagnose whether the internal parameters of the photovoltaic array are abnormal; f9 is the ratio of the normalized MPP current to the normalized short-circuit current. A low value may indicate shadow or mismatch loss; f 10 f is the current-voltage slope from the maximum power point to the short-circuit point, describing the linear characteristics of the IV curve near the short-circuit point. Abnormal slope may be caused by series resistance or diode characteristic degradation; 11 f is the negative current-voltage slope from the maximum power point to the open circuit point, reflecting the nonlinear characteristics of the IV curve near the open circuit point. Abnormal values may indicate a parallel resistor failure or hot spot effect. 12 f is the average slope from the half-open-circuit voltage point to the short-circuit point, which is used to evaluate the current decay rate in the medium-voltage area. Abnormalities may be caused by local shadows or cell mismatch; 13 f is the average slope from the maximum power point to the half-open circuit voltage point, reflecting the steepness of the IV curve near the maximum power point. A steep change may indicate component aging or temperature abnormality; 14 f is the average slope from the half-short-circuit current point to the maximum power point, describing the voltage change rate in the medium current region. Abnormalities may be caused by increased series resistance or cell damage; 15 It is the negative slope from the half-short-circuit current point to the open-circuit point, reflecting the current attenuation characteristics in the high-voltage region. Abnormal values may indicate parallel resistance degradation or leakage current problems. 16is the normalized fill factor, which represents the ratio of the actual fill factor to the STC, and comprehensively measures the output efficiency of the component; FF is the fill factor under actual operating conditions; FF (STC) is the fill factor under standard test conditions.

[0098] Step S9013: constructing an initial data set based on multiple photovoltaic array characteristic data.

[0099] Specifically, the five characteristic points and additional parameter information in the collected IV curve of each target photovoltaic array are subjected to feature extraction according to the above feature extraction process to obtain an initial data set.

[0100] Step S902 : performing feature selection on the initial data set using a sequential floating forward selection algorithm to obtain an optimized data set.

[0101] Specifically, if Figure 11 As shown, the above step S902 includes:

[0102] Step S9021: normalize the initial data set.

[0103] Specifically, the data normalization process adopts the maximum-minimum normalization method, and the normalization processing formula is as follows:

[0104]

[0105] In the above formula, a′ represents the normalized value of feature B in the initial dataset, a represents the initial value of feature B in the initial dataset, and max B and min B They represent the maximum and minimum values of feature B respectively.

[0106] Step S9022: Select features from the normalized initial data set using a sequential forward selection algorithm, and add the selected features to the initial feature subset to obtain a candidate feature subset; wherein the initial feature subset is an empty set.

[0107] Specifically, the initial feature subset k is set to 0, that is, the number of currently selected features is zero; a new feature is selected using the sequential forward selection algorithm and added to the initial feature subset, and then k is set to k+1, that is, the number of features in the selected feature subset is increased by one.

[0108] Step S9023: If the number of features in the candidate feature subset is greater than a preset threshold, the features in the candidate feature subset are deleted in sequence using a sequential backward selection algorithm to obtain multiple current feature subsets.

[0109] Step S9024: perform classification accuracy evaluation on the multiple current feature subsets, and screen the multiple current feature subsets based on the classification accuracy evaluation results to obtain the optimal feature subset.

[0110] Specifically, if k≥2, the classification accuracy of the current candidate feature subset is calculated, and then each feature in the candidate feature subset is removed in turn, and the classification accuracy of the candidate feature subset after removal is calculated. If the classification accuracy is improved or remains basically unchanged after deleting a certain feature, it is determined to delete the feature and set k=k-1, that is, the number of selected features is reduced by one, otherwise the feature is retained.

[0111] Step S9025: If the number of features of the optimal feature subset is equal to the preset dimension, the optimal feature subset is used as the optimized data set.

[0112] Specifically, check whether the number k of currently selected features is equal to the required dimension d. If k=d, the process stops; if not, continue to select features.

[0113] Furthermore, it is evaluated whether the candidate feature subset after each addition or subtraction of a feature is the best subset so far, that is, the classification accuracy of the candidate feature subset after each addition or subtraction of a feature is calculated. If the classification accuracy is the highest, the candidate feature subset is considered to be optimal, the features in the candidate feature subset are retained and the forward step is continued; if the classification accuracy is not the highest, the conditionally excluded features are added back to the candidate feature subset; if the features are added back to the candidate feature subset, the algorithm restarts the forward selection step and continues to iterate until the number k of currently selected features is equal to the required dimension d, or the expected goal is achieved, the iteration is stopped, and the optimized data set is output.

[0114] Step S903: Use the optimized data set to train the weighted ensemble learning model to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model includes a logistic regression classifier, a support vector machine classifier, and a naive Bayes classifier. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.

[0115] Step S904: Use the trained weighted integrated model to perform fault diagnosis on the solar photovoltaic array to obtain the fault diagnosis result of the solar photovoltaic array. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.

[0116] This embodiment provides a solar photovoltaic array fault diagnosis method, which ensures the consistency of features of different dimensions by normalizing an initial data set, adopts an appropriate normalization method to optimize data distribution, reduces algorithm computational complexity, and accelerates the convergence process of a subsequent weighted integration model. Furthermore, by performing feature extraction on multiple feature points and additional parameter information, multiple photovoltaic array feature data are obtained, thereby achieving accurate extraction of photovoltaic array features. An initial data set is constructed based on the photovoltaic array feature data, so that the initial data set can better reflect the characteristics of the photovoltaic array, laying a foundation for training the weighted integration model.

[0117] In this embodiment, a solar photovoltaic array fault diagnosis method is provided, which can be used for the above-mentioned electronic equipment. Figure 12 FIG. 1 is a flow chart of a solar photovoltaic array fault diagnosis method according to an embodiment of the present invention. Figure 12 As shown, the process includes the following steps:

[0118] Step S1201: Obtain the IV curve and additional parameter information of the target photovoltaic array, perform feature extraction based on the IV curve and additional parameter information, and obtain an initial data set. Figure 9 Step S901 of the illustrated embodiment will not be described in detail here.

[0119] Step S1202: Use the Sequential Floating Forward Selection algorithm to perform feature selection on the initial dataset to obtain an optimized dataset. Figure 9 Step S902 of the illustrated embodiment will not be described in detail here.

[0120] Step S1203: Use the optimized data set to train the weighted ensemble learning model to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model includes a logistic regression classifier, a support vector machine classifier, and a naive Bayes classifier.

[0121] Specifically, the above step S1203 includes:

[0122] In step S12031, the optimized data set is input into the logistic regression classifier, the support vector machine classifier and the naive Bayes classifier respectively to obtain multiple classification prediction results.

[0123] Specifically, the weighted ensemble learning model includes three classifiers: Logistic Regression (LR) classifier, Support Vector Machine (SVM) classifier and Naive Bayes (NB) classifier. The optimized data set is input into the above three classifiers for prediction, and the classification prediction results corresponding to each classifier are generated.

[0124] Step S12032: A weighted voting algorithm is used to combine multiple classification prediction results to obtain a fault classification prediction result.

[0125] Specifically, based on the voting probability in the weighted voting algorithm, the classification prediction results of the three classifiers are integrated to obtain the fault classification prediction result. The calculation formula is as follows:

[0126]

[0127] Where x represents the input data of the WEL (Weighted Ensemble Learning) model, y represents the fault classification prediction result output by the WEL model, j represents the number of category labels, n represents the number of classifiers, and g represents the number of classifiers. ji represents the classification prediction result of the i-th classifier based on probability, w i Represents the weight assigned to the i-th classifier, that is, the weight w=(w1,w2,w3) in the WEL model, w i ∈[0, 1];

[0128] In step S12033, if the fault classification prediction result is different from the fault diagnosis result in the optimized data set, the genetic algorithm is used to iteratively optimize the weights of the logistic regression classifier, the support vector machine classifier, and the naive Bayes classifier until the fault classification prediction result is the same as the fault diagnosis result, thereby obtaining a trained weighted integrated model.

[0129] Specifically, the specific steps of using the genetic algorithm (GA) to optimize the weights of the three classifiers include: 1) initializing the population: using the genetic algorithm to generate a group of random chromosomes, called the initial population; 2) chromosome representation: each chromosome contains a real number array, and the real numbers in the real number array are between 0 and 1, which are used to represent the weight assigned to each classifier; 3) fitness evaluation: evaluating the quality of each chromosome through the fitness function, that is, evaluating the accuracy of the WEL model; 4) selecting high-quality chromosomes: based on the results of the fitness evaluation, selecting the most accurate chromosome to generate the next generation; 5) applying genetic operations: using mutation and crossover operations to generate a new generation of chromosomes, and the specific ratio is determined by the mutation rate and crossover rate; 6) termination condition: the cycle continues until a specific termination condition is met (that is, the fault classification prediction result is the same as the fault diagnosis result), and the trained weighted integration model is obtained.

[0130] Step S1204: Use the trained weighted integrated model to perform fault diagnosis on the solar photovoltaic array to obtain the fault diagnosis result of the solar photovoltaic array. Figure 9Step S904 of the illustrated embodiment will not be described in detail here.

[0131] This embodiment provides a solar photovoltaic array fault diagnosis method using a weighted ensemble model that combines the advantages of a logistic regression classifier, a support vector machine classifier, and a naive Bayesian classifier. By optimizing the weights of each classifier using a genetic algorithm, the overall accuracy of the weighted ensemble model is improved. This optimizes resource allocation, enhances the fault diagnosis efficiency of the weighted ensemble model, and adapts to dynamic changes in fault problems. This method achieves high accuracy with minimal training data. The trained weighted ensemble model can effectively classify and detect various fault types and their severity.

[0132] In this embodiment, a solar photovoltaic array fault diagnosis method is provided, which can be used for the above-mentioned electronic equipment. Figure 13 FIG. 1 is a flow chart of a solar photovoltaic array fault diagnosis method according to an embodiment of the present invention. Figure 13 As shown, the process includes the following steps:

[0133] Step S1301: Obtain the IV curve and additional parameter information of the target photovoltaic array, perform feature extraction based on the IV curve and additional parameter information, and obtain an initial data set. Figure 12 Step S1201 of the illustrated embodiment will not be described in detail here.

[0134] Step S1302: Use the Sequential Floating Forward Selection algorithm to perform feature selection on the initial dataset to obtain an optimized dataset. Figure 12 Step S1202 of the illustrated embodiment will not be described in detail here.

[0135] Step S1303: Use the optimized data set to train the weighted ensemble learning model to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model includes a logistic regression classifier, a support vector machine classifier, and a naive Bayes classifier. Figure 12 Step S1203 of the illustrated embodiment will not be described in detail here.

[0136] Step S1304: Use the trained weighted integrated model to perform fault diagnosis on the solar photovoltaic array to obtain the solar photovoltaic array fault diagnosis result. Figure 12 Step S1204 of the illustrated embodiment will not be described in detail here.

[0137] Step S1305 , verify and evaluate the trained weighted ensemble learning model, and optimize the trained weighted ensemble learning model based on the verification and evaluation results.

[0138] Specifically, the K-fold cross-validation technique is used in conjunction with evaluation indicators such as the confusion matrix to comprehensively verify the stability and generalization ability of the model; at the same time, the false detection rate and missed detection rate of the weighted ensemble learning model are evaluated by calculating key performance indicators such as the classifier's accuracy, precision, and recall rate.

[0139] The above step S1305 includes:

[0140] Step S13051: Use K-fold cross validation combined with confusion matrix to verify the feasibility of the trained weighted ensemble learning model to obtain a model feasibility verification result.

[0141] Specifically, the optimized dataset is divided into a training set (80%) and a validation set (20%), and then the training set is divided into K subsets (K=10), where each subset contains the same number of data samples, and K-fold cross-validation is performed.

[0142] Furthermore, the specific steps of K-fold cross-validation include: 1) start iteration: perform K iterations, select a subset as the validation set and the remaining K-1 subsets as the training set in each iteration; 2) weighted ensemble learning model training: in each iteration, use the training set (i.e., K-1 subsets) to train the weighted ensemble learning model; 3) model verification: use the reserved validation set (i.e., the selected subset) to verify the feasibility of the trained weighted ensemble learning model, and calculate the accuracy of the current stage; 4) repeat iteration: repeat the above steps 1) to step 3) until each subset is used as a validation set once; 5) collect accuracy: in each iteration stage, record the accuracy of the validation set of the current stage; 6) calculate the final accuracy: take the average of the accuracy obtained from K iterations as the final accuracy of weighted ensemble learning.

[0143] Furthermore, the format of the confusion matrix elements is "prediction-actual-true / false", where the first number of each matrix element represents the output category of the classification prediction result, the second number represents the category to which the sample actually belongs, and the last number represents a Boolean parameter used to verify whether the predicted category is consistent with the actual category.

[0144] Step S13052: Obtain a performance evaluation index, and use the performance evaluation index to perform performance evaluation on the trained weighted ensemble learning model to obtain a performance evaluation result.

[0145] Specifically, in order to determine the feasibility of the classification algorithm, the accuracy A is measured by deploying the elements in the confusion matrix. The calculation formula of the accuracy A is as follows:

[0146]

[0147] The numerator is the sum of the diagonal elements of the confusion matrix, and the denominator is the number of all samples.

[0148] Furthermore, recall and precision are both used to indicate the proportion of correctly predicted samples. Recall indicates the proportion of samples that actually belong to the studied category, while precision indicates the proportion of samples that are predicted to belong to a certain category. The calculation formulas for recall and precision are as follows:

[0149]

[0150] Among them, the numerator uuT is the number of samples that the weighted ensemble learning model correctly predicts category u and actually belongs to category u; the denominator ouF is the number of samples that actually belong to category u but are incorrectly predicted to be category o (o≠u); ∑ u≠o (ouF) is the total number of samples of category u that are misclassified as other categories o; T is the abbreviation for True (correct) and F is the abbreviation for False (wrong).

[0151] Step S13053: Optimize the trained weighted ensemble learning model based on the model feasibility verification results and performance evaluation results.

[0152] This embodiment provides a solar photovoltaic array fault diagnosis method that uses K-fold cross-validation technology or a confusion matrix to verify the feasibility of a weighted ensemble learning model, and uses performance evaluation indicators to perform performance evaluation on the trained weighted ensemble learning model, thereby further improving the fault diagnosis accuracy of the weighted ensemble learning model.

[0153] The following describes the specific steps of a solar photovoltaic array fault diagnosis method through a specific embodiment.

[0154] Example 1:

[0155] like Figure 14 As shown in FIG, the specific steps of the solar photovoltaic array fault diagnosis method include:

[0156] 1) The five statistical features and additional parameter information selected on the IV curve are used as the data basis for feature extraction, and the initial data set is established: the five feature points selected on the IV curve are: A(0, I sc )、B(V oc / 2,I Voc / 2 )、C(V MPP , I MPP )、D(V Isc / 2 , I sc / 2)、E(V oc , 0); the additional parameter information required is: I sc (STC), V oc(STC), I MPP (STC), V MPP (STC); The five selected feature points and additional parameter information are used for feature extraction, and a total of 16 features are extracted. The 16 features are shown in Table 1 below:

[0157] Table 1:

[0158]

[0159]

[0160] 2) The initial data set is normalized to improve the convergence speed of the algorithm. The normalized initial data set is shown in Table 2 below:

[0161] Table 2:

[0162] Minimum Maximum majority median average value Standard deviation <![CDATA[f1]]> 0 1 0.412483 0.496771 0.533007 0.2871 <![CDATA[f2]]> 0 1 0.973071 0.963948 0.86266 0.220243 <![CDATA[f3]]> 0 1 0.944752 0.86811 0.746508 0.247224 <![CDATA[f4]]> 0 1 0.755252 0.475786 0.503236 0.264531 <![CDATA[f5]]> 0 1 0.982822 0.471833 0.508662 0.271918 <![CDATA[f6]]> 0 1 0.955711 0.928292 0.805421 0.241809 <![CDATA[f7]]> 0 1 0.100976 0.059322 0.074529 0.082449 <![CDATA[f8]]> 0 1 0.780984 0.753157 0.649398 0.27197 <![CDATA[f9]]> 0 1 0.962644 0.956955 0.926084 0.080595 <![CDATA[f 10 ]]> 0 1 0.941936 0.968543 0.935964 0.100037 <![CDATA[f 11 ]]> 0 1 0.820984 0.914536 0.899351 0.095542 <![CDATA[f 12 ]]> 0 1 0.995288 0.982615 0.956832 0.094805 <![CDATA[f 13 ]]> 0 1 0.94949 0.973533 0.925734 0.116306 <![CDATA[f 14 ]]> 0 1 0.903309 0.945913 0.927995 0.083694 <![CDATA[f 15 ]]> 0 1 0.790066 0.885201 0.87768 0.106138 <![CDATA[f 16 ]]> 0 1 0.922078 0.864141 0.795116 0.174157

[0163] 3) Use the sequential floating forward selection algorithm to perform feature selection on the initial data set, filter out the features most relevant to fault classification, and obtain the optimized data set;

[0164] 4) Using the optimized data set to train the weighted ensemble learning model to obtain a trained weighted ensemble model;

[0165] 5) Using the trained weighted integrated model to perform fault diagnosis on the solar photovoltaic array to obtain the fault diagnosis results of the solar photovoltaic array;

[0166] 6) Verify and evaluate the trained weighted ensemble learning model, and optimize the trained weighted ensemble learning model based on the verification and evaluation results: Use K-fold cross-validation combined with confusion matrix to verify the feasibility of the trained weighted ensemble learning model, and obtain the model feasibility verification result; Use performance evaluation indicators to evaluate the performance of the trained weighted ensemble learning model, and obtain performance evaluation results. Optimize the trained weighted ensemble learning model based on the model feasibility verification result and the performance evaluation result; Among them, the performance evaluation indicators include: accuracy, training accuracy, and test accuracy; The training accuracy and test accuracy of each layer in the comprehensive multi-layer model are shown in Table 3 below:

[0167] Table 3:

[0168] Number of layers Training accuracy Test accuracy 1 95.32% 96.25% 2 99.67% 100% 3 100% 100% 4 96.17% 97.33% 5 99.06% 100% 6 100% 100%

[0169] As shown in Table 3 above, the training and testing accuracy of the multi-layer model for detecting, identifying, and classifying various types of electrical faults in photovoltaic systems are high, and the model is very feasible.

[0170] In the above-mentioned embodiment 1, the weighted ensemble learning model is a comprehensive multi-layer model comprising six layers, which is used for the diagnosis, classification and severity identification of photovoltaic array electrical faults. The multi-layer structure can more comprehensively analyze different types of faults and refine the classification and severity of the faults; wherein, a weighted ensemble learning (WEL) algorithm is adopted in each layer of the weighted ensemble learning model. The algorithm combines the advantages of three classifiers: support vector machine (SVM), naive Bayes (NB) and logistic regression (LR). The weight of each classifier is optimized by genetic algorithm (GA), thereby improving the overall accuracy of the model. In each layer, the sequential floating forward selection (SFFS) algorithm is adopted for feature selection, thereby reducing the dimension of the data set, simplifying the calculation process, and improving the training efficiency.

[0171] This embodiment also provides a solar photovoltaic array fault diagnosis device, which is used to implement the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0172] This embodiment provides a solar photovoltaic array fault diagnosis device, such as Figure 15 As shown, the device includes:

[0173] The feature extraction module 1501 is used to obtain the IV curve and additional parameter information of the target photovoltaic array, perform feature extraction based on the IV curve and the additional parameter information, and obtain an initial data set.

[0174] The feature selection module 1502 is used to perform feature selection on the initial data set using a sequential floating forward selection algorithm to obtain an optimized data set.

[0175] The training module 1503 is used to train the weighted ensemble learning model using the optimized data set to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model includes a logistic regression classifier, a support vector machine classifier, and a naive Bayes classifier.

[0176] The fault diagnosis module 1504 is used to perform fault diagnosis on the solar photovoltaic array using the trained weighted integrated model to obtain a solar photovoltaic array fault diagnosis result.

[0177] In some optional implementations, the feature extraction module 1501 includes:

[0178] The selection unit is used to select multiple feature points on the IV curve of the target photovoltaic array.

[0179] The feature extraction unit is used to extract features from multiple feature points and additional parameter information to obtain multiple photovoltaic array feature data.

[0180] The construction unit is used to construct an initial data set based on a plurality of photovoltaic array characteristic data.

[0181] In some optional implementations, the feature selection module 1502 includes:

[0182] The normalization processing unit is used to perform normalization on the initial data set.

[0183] The feature selection unit is used to select features from the normalized initial data set using a sequential forward selection algorithm, and add the selected features to the initial feature subset to obtain a candidate feature subset; wherein the initial feature subset is an empty set.

[0184] The screening unit is used to delete the features in the candidate feature subset in sequence using a sequential backward selection algorithm if the number of features in the candidate feature subset is greater than a preset threshold, so as to obtain multiple current feature subsets.

[0185] The first evaluation unit is used to perform classification accuracy evaluation on multiple current feature subsets, and screen the multiple current feature subsets based on the classification accuracy evaluation results to obtain the optimal feature subset.

[0186] As a unit, if the number of features of the optimal feature subset is equal to the preset dimension, the optimal feature subset is used as the optimized data set.

[0187] In some optional implementations, the training module 1503 includes:

[0188] The input unit is used to input the optimized data set into the logistic regression classifier, the support vector machine classifier and the naive Bayes classifier respectively to obtain multiple classification prediction results.

[0189] The synthesis unit is used to synthesize multiple classification prediction results using a weighted voting algorithm to obtain a fault classification prediction result.

[0190] Iterative optimization unit, if the fault classification prediction result is different from the fault diagnosis result in the optimized data set, the genetic algorithm is used to iteratively optimize the weights of the logistic regression classifier, support vector machine classifier and naive Bayes classifier until the fault classification prediction result is the same as the fault diagnosis result, and the trained weighted integration model is obtained.

[0191] In some optional embodiments, the method further includes:

[0192] The evaluation module is used to verify and evaluate the trained weighted ensemble learning model, and optimize the trained weighted ensemble learning model based on the verification and evaluation results.

[0193] In some optional embodiments, the evaluation module includes:

[0194] The feasibility verification unit performs model feasibility verification on the trained weighted ensemble learning model by using K-fold cross validation combined with a confusion matrix to obtain a model feasibility verification result.

[0195] The second evaluation unit is used to obtain a performance evaluation index, and use the performance evaluation index to perform performance evaluation on the trained weighted ensemble learning model to obtain a performance evaluation result.

[0196] An optimization unit is used to optimize the trained weighted ensemble learning model based on the model feasibility verification result and the performance evaluation result.

[0197] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0198] In this embodiment, a solar photovoltaic array fault diagnosis device is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0199] The embodiment of the present invention also provides a computer device having the above Figure 15 A solar photovoltaic array fault diagnosis device is shown.

[0200] See also Figure 16 , Figure 16 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 16 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of a GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories. Equally, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 16 A processor 10 is taken as an example.

[0201] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0202] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0203] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0204] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0205] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 16 The bus connection is taken as an example.

[0206] The input device 30 can receive input digital or character information and generate key signal input related to user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0207] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0208] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0209] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A solar photovoltaic array fault diagnosis method, characterized in that: The method comprises: Acquiring an IV curve and additional parameter information of a target photovoltaic array, performing feature extraction based on the IV curve and the additional parameter information, and obtaining an initial data set; Performing feature selection on the initial data set using a sequential floating forward selection algorithm to obtain an optimized data set; Using the optimized data set to train a weighted ensemble learning model to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model includes a logistic regression classifier, a support vector machine classifier, and a naive Bayes classifier; The trained weighted integrated model is used to perform fault diagnosis on a solar photovoltaic array to obtain a solar photovoltaic array fault diagnosis result.

2. The method according to claim 1, characterized in that The extracting features based on the IV curve and the additional parameter information to obtain an initial data set includes: Selecting a plurality of characteristic points on the IV curve of the target photovoltaic array; Performing feature extraction on the plurality of feature points and the additional parameter information to obtain a plurality of photovoltaic array feature data; The initial data set is constructed based on the plurality of photovoltaic array characteristic data.

3. The method according to claim 1, characterized in that The method of performing feature selection on the initial data set using a sequential floating forward selection algorithm to obtain an optimized data set includes: performing normalization processing on the initial data set; Selecting features from the normalized initial data set using a sequential forward selection algorithm, and adding the selected features to the initial feature subset to obtain a candidate feature subset; wherein the initial feature subset is an empty set; If the number of features in the candidate feature subset is greater than a preset threshold, the features in the candidate feature subset are deleted in sequence using a sequential backward selection algorithm to obtain multiple current feature subsets; Performing classification accuracy evaluation on the multiple current feature subsets, and screening the multiple current feature subsets based on the classification accuracy evaluation results to obtain an optimal feature subset; If the number of features of the optimal feature subset is equal to the preset dimension, the optimal feature subset is used as the optimized data set.

4. The method according to claim 1, wherein The method of training the weighted ensemble learning model using the optimized data set to obtain the trained weighted ensemble model includes: Inputting the optimized data set into the logistic regression classifier, the support vector machine classifier and the naive Bayes classifier respectively to obtain multiple classification prediction results; Using a weighted voting algorithm to combine the multiple classification prediction results to obtain a fault classification prediction result; If the fault classification prediction result is different from the fault diagnosis result in the optimized data set, the weights of the logistic regression classifier, the support vector machine classifier and the naive Bayes classifier are iteratively optimized using a genetic algorithm until the fault classification prediction result is the same as the fault diagnosis result, thereby obtaining the trained weighted integrated model.

5. The method according to claim 1, wherein Also includes: The trained weighted ensemble learning model is verified and evaluated, and the trained weighted ensemble learning model is optimized based on the verification and evaluation results.

6. The method according to claim 5, characterized in that The verifying and evaluating the trained weighted ensemble learning model, and optimizing the trained weighted ensemble learning model based on the verification and evaluation results, includes: The feasibility of the weighted ensemble learning model after training is verified by using K-fold cross validation combined with confusion matrix to obtain a model feasibility verification result; Obtaining a performance evaluation indicator, and using the performance evaluation indicator to perform a performance evaluation on the trained weighted ensemble learning model to obtain a performance evaluation result; The trained weighted ensemble learning model is optimized based on the model feasibility verification result and the performance evaluation result.

7. A solar photovoltaic array fault diagnosis device, characterized in that: The device comprises: a feature extraction module, configured to obtain an IV curve and additional parameter information of a target photovoltaic array, and perform feature extraction based on the IV curve and the additional parameter information to obtain an initial data set; A feature selection module is used to perform feature selection on the initial data set using a sequential floating forward selection algorithm to obtain an optimized data set; A training module, configured to train a weighted ensemble learning model using the optimized data set to obtain a trained weighted ensemble model; wherein the weighted ensemble learning model includes a logistic regression classifier, a support vector machine classifier, and a naive Bayes classifier; The fault diagnosis module is used to perform fault diagnosis on the solar photovoltaic array using the trained weighted integrated model to obtain a solar photovoltaic array fault diagnosis result.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the solar photovoltaic array fault diagnosis method according to any one of claims 1 to 6 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the solar photovoltaic array fault diagnosis method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to enable a computer to execute the solar photovoltaic array fault diagnosis method according to any one of claims 1 to 6.