Photovoltaic system fault diagnosis method based on Aasso-IKOA-XGBoost
By adopting the ALasso-IKOA-XGBoost method in photovoltaic system fault diagnosis, multiple types of characteristics of photovoltaic data are extracted and model parameters are optimized, and the problem of insufficient adaptability of photovoltaic system fault diagnosis methods in the existing technology is solved, achieving more efficient and accurate fault diagnosis and intelligent operation and maintenance.
Patent Information
- Application Number
- CN202510010725.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The existing photovoltaic system fault diagnosis methods are weak in adaptability when processing high-dimensional and dynamically changing photovoltaic data, and the recognition rate is affected by the degree of sample equalization, making it difficult to effectively diagnose the faults of photovoltaic arrays and inverters.
The method based on ALasso-IKOA-XGBoost is adopted to extract the waveform time-frequency feature components through wavelet analysis, and multiple types of feature components of fault data are extracted using the adaptive Lasso method. The model parameters are optimized by combining the IKOA algorithm with Logistic, Sine and Tent chaotic mapping, and finally a fault diagnosis model is built with XGBoost.
The efficiency, accuracy and intelligent operation and maintenance level of photovoltaic system fault diagnosis have been improved, ensuring the reliable, efficient and sustainable operation of the system.
Smart Images

Figure CN120074369A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of photovoltaic system fault diagnosis, and particularly relates to a photovoltaic system fault diagnosis method based on ALasso-IKOA-XGBoost. Background Technique
[0002] With the transformation and upgrading of the global energy structure, renewable energy, as a clean and environmentally friendly form of green energy, is developing rapidly. As an important part of renewable energy, the photovoltaic system has become a key solution to address energy challenges and climate change. With the continuous progress of photovoltaic technology, the installed capacity of the photovoltaic system has been increasing, making a great contribution to sustainable energy supply. However, in the actual operation process of the photovoltaic system, there are problems such as unstable power generation efficiency and difficult operation and maintenance, resulting in major challenges to its reliability and performance. Among them, the photovoltaic array and the inverter, as the core components of the system, their failures will directly lead to a decrease in the power generation efficiency of the system and affect the utilization of renewable energy. Therefore, establishing an automated fault diagnosis and intelligent operation and maintenance system for the photovoltaic array and the inverter is of great significance for ensuring the efficient and stable operation of the photovoltaic system.
[0003] After retrieving the existing literature, it is found that currently, the traditional fault diagnosis methods for photovoltaic power generation systems mainly focus on signal monitoring methods, mathematical model methods, signal analysis methods, statistical analysis methods, and manual experience methods.
[0004] Literature [1]: "Intelligent Fault Detection Method for Photovoltaic Arrays Based on Support Vector Machine Algorithm" (Zhou Yunfeng, Liu Guangyu, Li Huajun, etc. Intelligent Fault Detection Method for Photovoltaic Arrays Based on Support Vector Machine Algorithm [J]. Manufacturing Automation, 2021, 43(06): 45-8.) proposed a fault detection method for photovoltaic arrays based on SVM. This method has a good recognition rate for each fault state by combining the external environment and the internal electrical characteristics of the photovoltaic system. However, the training process of SVM requires solving quadratic programming problems and is difficult to process the massive data generated by the photovoltaic system.
[0005] Reference [2]: "Research on IGBT Open - Circuit Fault Diagnosis and Efficiency Loss Evaluation Model of Photovoltaic Grid - Connected Inverters" (Zhao Zilinglong. Research on IGBT Open - Circuit Fault Diagnosis and Efficiency Loss Evaluation Model of Photovoltaic Grid - Connected Inverters [D]. Hefei University of Technology, 2021.) constructed a fault diagnosis model of photovoltaic inverters based on SVM. This model uses the superior optimization and classification characteristics of SVM to diagnose and classify the feature quantities extracted from 22 types of IGBT open - circuit faults. The fault diagnosis model adopted in this literature is the same as the previous one, and the difference lies in the diagnosis object. Therefore, the deficiencies here can be described as follows: When dealing with 22 types of IGBT open - circuit faults, the training complexity of SVM will increase significantly. Another way of handling is to describe and summarize the deficiencies of this type of algorithm together in the way of the above - mentioned literature after introducing these two literatures.
[0006] Reference [3]: "Fault Diagnosis of Photovoltaic Systems Based on Principal Component Optimized Neural Networks" (Han Yuchen. Fault Diagnosis of Photovoltaic Systems Based on Principal Component Optimized Neural Networks [D]. Shanghai Dianji University, 2021.) constructed a GA - BP neural network model, which can effectively diagnose different open - circuit type faults of NPC inverters. However, the BP neural network itself relies on the gradient descent method for parameter update and is easily restricted by local optima. In addition, the BP neural network is a static model and is difficult to capture time - dynamic characteristics.
[0007] Generally speaking, the existing machine - learning methods have weak adaptability to the high - dimensional and dynamically changing data of photovoltaic systems, and the recognition rate is affected by the degree of sample balance. Therefore, how to establish an intelligent diagnosis system that can process high - dimensional and unbalanced photovoltaic data and perform automatic fault recognition is one of the key problems in current research.
[0008] XGBoost is integrated by using CART as the base classifier through the method of gradient boosting and has excellent classification and prediction performance. Reference [4]: "Research on Photovoltaic Array Fault Diagnosis Method Based on XGBoost" (Liu Xingxing, Pazilaimu·Mahemuti, Cheng Zhijiang, etc. Research on Photovoltaic Array Fault Diagnosis Method Based on XGBoost [J]. Electronic Measurement Technology, 2023, 46(12): 8 - 14.) constructed a fault diagnosis model based on XGBoost. This model improves the fault classification accuracy and ensures the generalization performance by extracting fault features under different fault states of the photovoltaic array. Although XGBoost designs a feature importance evaluation mechanism, it does not specifically screen or reduce the dimension of redundant features. This literature relies on the method of manual parameter tuning, which is time - consuming and may not find the global optimal parameter combination.
[0009] Literature [5]: "Photovoltaic Power Generation Anomaly Detection Based on Improved VMD-XGBoost-BiLSTM Combined Model" (Zhao Bochao, Ma Jiajun, Cui Lei, etc. Photovoltaic Power Generation Anomaly Detection Based on Improved VMD-XGBoost-BiLSTM Combined Model [J]. Computer Engineering, 2024, 50(03): 306-16.) also achieved effective diagnosis of internal faults in photovoltaic power generation units based on the XGBoost ensemble learning model. However, such models still face the risk of overfitting for small datasets, and the fault diagnosis accuracy is affected by model parameters. For the monitoring data with high dimensions in the photovoltaic system, there are redundant and noise features, and the lasso feature selection method can be used to improve the diagnosis effect.
[0010] Literature [6]: "Variable Selection Method Based on Spatiotemporal Group Lasso and Hierarchical Bayesian Spatiotemporal Model" (Wang Ling, Kang Zihao. Variable Selection Method Based on Spatiotemporal Group Lasso and Hierarchical Bayesian Spatiotemporal Model [J]. Journal of Geo-Information Science, 2023, 25(07): 1312-24.) constructed a hierarchical Bayesian spatiotemporal group Lasso variable selection model. This model fully considers spatiotemporal correlation, performs variable selection through spatiotemporal group Lasso, and then uses the hierarchical Bayesian spatiotemporal model to verify the effect of variable selection, accurately selecting the variable subset that has the greatest impact on the dependent variable, thereby improving the prediction effect. However, in the photovoltaic fault diagnosis system, there is a high degree of linear relationship or correlation between independent variables, resulting in unstable or difficult-to-interpret coefficient estimation of the model. And when facing highly correlated features, Lasso usually only selects one of the features and ignores other relevant features. Summary of the Invention
[0011] In view of the deficiencies of the existing technical literature, the present invention proposes a photovoltaic system fault diagnosis method based on ALasso-IKOA-XGBoost. First, to fully mine fault features and solve the problem of multicollinearity, wavelet analysis is used to extract waveform time-frequency feature components, and Alasso is used to extract multi-class feature components of fault data. Then, aiming at the problems of slow convergence speed, poor robustness and easy to fall into local optimum of KOA, an IKOA algorithm integrating Logistic, Sine and Tent chaotic mappings is proposed to generate a better initial solution distribution, parameter setting and position update method. Then, it is combined with XGBoost to build an IKOA-XGBoost fault diagnosis model to improve the learning expression and classification accuracy of the photovoltaic fault diagnosis model under complex high-dimensional data conditions. This method can effectively improve the diagnosis efficiency, accuracy and intelligent operation and maintenance level of the photovoltaic system, and ensure the reliable, efficient and sustainable operation of the system.
[0012] The technical solution adopted by the present invention is as follows:
[0013] A fault diagnosis method for a photovoltaic system based on ALasso-IKOA-XGBoost includes the following steps:
[0014] Step 1: Extract the time-frequency feature components of the waveform through wavelet analysis, and use the adaptive Lasso method to extract multiple types of feature components of the fault data;
[0015] Step 2: Propose an IKOA algorithm that fuses Logistic, Sine, and Tent chaotic mappings to generate a better initial solution distribution, parameter setting, and position update method;
[0016] Step 3: Combine the IKOA algorithm in Step 2 with XGBoost to build an IKOA-XGBoost fault diagnosis model to diagnose the faults of the photovoltaic system.
[0017] In the above Step 1, first, wavelet analysis is used to replace the infinitely long trigonometric function basis in the Fourier transform with a finitely long and decaying wavelet basis function, extract the different frequency features of the signal, retain important information, remove noise, improve the signal quality and readability. The wavelet analysis extracts the time-frequency features of the three-phase inverter output current, which can be expressed as:
[0018]
[0019] In the formula: WT f (a, b) represents the wavelet transform result of the signal f(t); f(t) is the current signal to be analyzed; is the conjugate of; a and b are the scaling scale and translation amount of the wavelet basis function ψ(t), respectively.
[0020] Then, an adaptive Lasso method is proposed to solve the irrationality problems caused by the same penalty for all coefficients in the Lasso method. Its principle is to assign different weights to different penalty terms on the basis of the Lasso method, and its expression is as follows:
[0021]
[0022] In the formula: represents the regression coefficient estimate value; λ 2 is the adaptive Lasso penalty weight; β j is the penalty term of the model parameter β, which can shrink the model regression coefficient towards zero;
[0023]
[0024] In the formula: represents the adaptive weight, which is used to adjust the importance of the jth regression coefficient in the penalty term; j represents the index of the feature; p represents the number of features; represents the initial estimated value; represents the coefficient estimated value obtained by the least squares method; X j =(X 1j , X 2j , …, X nj ) T represents the predictor variable, where j = 1, 2, …, p.
[0025] Y=(Y 1 , Y 2 , …, Y n ) T represents the response variable, Y 1 , Y 2 , …, Y n respectively represent the true output values of the 1st to nth samples; T represents the transpose of the vector.
[0026] Its weight expression is as follows:
[0027]
[0028] In the formula: represents the adaptive weight matrix; respectively represent the weight estimated values assigned to the 1st to pth features; γ is the adaptive Lasso penalty parameter, γ ≥ 0.
[0029] In step 1, the adaptive Lasso method is used to extract multi-class feature components of the fault data; specifically as follows:
[0030] The faults of photovoltaic modules mainly include the following four types: module short circuit, module open circuit, partial shadow covering, and panel aging and damage. The IV curve characteristics of module faults are as Figure 7 shown. In the case of short circuit, the branch voltage U, the maximum power point current Im, the voltage Um, and the open circuit voltage Uoc are selected as feature components. In the case of open circuit, the branch current I, the short circuit current Isc, and the maximum power point current Im are selected as feature components. In the case of aging, the maximum power point current Im and the voltage Um are selected as feature quantities. In the case of partial shadow, the maximum power point current Im, the voltage Um, the open circuit voltage Uoc, and the number of local maximum power points are selected as feature quantities.
[0031] The output characteristics of different types of inverter open circuit faults are as Figure 8 shown. When a single transistor is open, the phase current will be missing half a cycle of the waveform, and harmonic generation will occur in other phase currents, resulting in waveform distortion and a low output voltage. When two transistors in the same phase fail, the phase current will be missing. When two transistors in different phases fail, two phases will be missing half a cycle of the waveform, and the waveforms of other phases will be distorted. Therefore, the three-phase output current is selected as the feature quantity.
[0032] In step 2, the chaos strategy mainly utilizes the properties of the chaotic system to generate a better search strategy and avoid falling into local optimal solutions.
[0033] 1) Introduce the Logistic map to initialize the population, replacing completely random initialization while retaining the randomness of the initial value to make the initial solutions more evenly distributed;
[0034]
[0035] In the formula: represents the Logistic map value of the j-th variable at the i-th iteration; μ represents the parameter of the Logistic map, and its value range is (0, 4]; represents the Logistic map value of the j-th variable at the (i - 1)-th iteration; represents the updated value of the j-th decision variable of the i-th planet; represents the lower bound of the j-th variable of the i-th planet; represents the random sequence value generated by the Logistic map, used to map the variable to the range between its upper and lower bounds; represents the upper bound of the j-th variable of the i-th planet; i represents the index of the planet; j represents the index of the decision variable; N represents the number of candidate solutions in the search space; d represents the dimension of the problem to be optimized.
[0036] 2) Introduce a control factor based on the Logistic map into the gravitational force F to finely adjust the amplitude of the change in the planet's position to balance exploration and exploitation;
[0037]
[0038] In the formula: r 1 represents the value of the Logistic map sequence L t ; L t and L t-1 represent the results of the Logistic map at the t-th and (t - 1)-th iterations respectively; represents the gravitational strength received by the i-th planet (solution) at the t-th iteration; μ represents; e i represents the control factor; μ(t) represents the gravitational constant, which usually gradually decreases as the number of iterations t increases; represents the normalized value of the solar mass; represents the normalized mass of the i-th planet; represents the normalized distance between the current planet and the sun; ε represents a very small positive number.
[0039] 3) Introduce the Sine map to replace the orbital eccentricity ei The quality random value retains the initial value S 1 The randomness can increase the randomness and oscillation range of the parameter value, which is beneficial to enhancing the population diversity and the uncertainty of the algorithm.
[0040]
[0041] In the formula: S i represents the Sine mapping value at the i-th iteration; a represents a constant used to adjust the amplitude of the Sine mapping value; S i-1 represents the Sine mapping value at the (i - 1)-th iteration; e i represents the orbital eccentricity;
[0042]
[0043] In the formula: r 2 is equivalent to S t , that is, the value generated by the Sine mapping; S t represents the random sequence value at the t-th iteration; S t-1 represents the value of the Sine mapping in the previous round of iteration; M s represents the weight factor related to the Sine mapping value S t ; fit s (t) represents the meaning; fit k (t) represents the fitness value of the t-th individual; worst(t) represents the fitness value of the worst individual in the current population; N here represents the population size.
[0044] 4) Introduce a perturbation term based on the Tent mapping to fully disrupt the sequence, as a global perturbation operator to enhance the global search ability of the algorithm.
[0045]
[0046] In the formula: represents the output vector of the Tent mapping; represents the random sequence vector generated by the Tent mapping at the t-th iteration; represents the output vector of the Tent mapping at the (t - 1)-th iteration; T t-1 represents the output value of the Tent mapping at the (t - 1)-th iteration; β represents the piecewise parameter of the Tent mapping; V i (t) represents the velocity of planet i at time t; and represent system parameters used to adjust and the influence degree on the velocity; r 4 represents the random factor; Represents the current position vector of the $i$-th candidate solution; Represents the current position vector of the optimal solution in the population; Represents the average position vector in the population; $R$ i-nom $(t)$ represents the normalized value of the $i$-th solution; and Represents an additional perturbation term used to increase the diversity of solutions and prevent the algorithm from falling into local optima; and Represent the upper and lower bounds of the $i$-th solution respectively, used to limit the search range; $L$ represents the control variable used to strengthen the influence on the solution, reflecting the global search ability; $U$ 2 Represents a perturbation term used to adjust the scope of global search; $r$ 3 Represents a random factor for adjusting the weight.
[0047] The IKOA algorithm that fuses Logistic, Sine, and Tent chaotic maps is as follows:
[0048] The present invention adopts three chaotic strategies to improve the KOA algorithm to enhance the global search ability of the algorithm and better balance the relationship between exploration and exploitation.
[0049] The IKOA algorithm that fuses Logistic, Sine, and Tent chaotic maps includes equations (1) to (5);
[0050] First, use formula (1) to introduce the Logistic map to initialize the population, replacing completely random initialization. At the same time, use equation (2) to introduce a control factor based on the Logistic map into the gravitational force $F$ to finely adjust the amplitude of the planet position change;
[0051] Second, use equations (3) and (4) to introduce the Sine map to replace the random values in the orbital eccentricity and mass respectively;
[0052] Finally, use equation (5) to introduce a perturbation term based on the Tent map and use it as a global perturbation operator to fully shuffle the sequence to enhance the global search ability of the algorithm.
[0053] In step 3, a photovoltaic system fault diagnosis model based on XGBoost is constructed:
[0054] First, the XGBoost proposed by the present invention is based on CART. By the gradient boosting method, multiple weak learners CART are combined to construct a powerful prediction classification model. In the iterative training, each learner corrects the error of the previous round through gradient descent to minimize the loss function. In addition, the second-order Taylor formula is used to optimize the loss function to improve the calculation accuracy, and L1 and L2 regularization terms are introduced to limit the depth of the tree and the number of nodes, adding the complexity of the tree model to the regularization term to penalize the tree structure with multiple leaf nodes. In addition, sample shrinkage and feature sub-sampling are combined to avoid overfitting and simplify the model, thereby improving the generalization ability of the model.
[0055] 1) Regularized objective function:
[0056] For a data set with n samples and m features x i represents the input feature vector of the i-th sample; y i represents the true value of the i-th sample; n represents the number of samples; represents the dimension of the feature space; represents the range of the target value y i ; The final prediction output of K CARTs is defined as:
[0057]
[0058] In the formula: represents the predicted value of the i-th sample; φ(x i ) represents the prediction function of the model; f k (x i ) represents the predicted value of the k-th decision tree for the sample x i ; K represents the total number of CART decision trees used in the model;
[0059] represents the space of CART regression trees; f(x) represents the prediction function of the decision tree; q represents the structure vector; w represents the leaf weight; represents a vector space composed of T real numbers; each function f k corresponds to an independent tree structure vector q and leaf weight w; q points from the sample to the corresponding leaf label, and each leaf node of each CART corresponds to a continuous fractional value, that is, the weight; the score of the i-th node is w i ; G is the number of leaf nodes; w q(x) is the score for the sample x, that is, the model predicted value.
[0060] For each sample, each CART classifies it into a leaf node according to different classification rules, and the final prediction result is obtained by accumulating the scores w of the corresponding leaves.
[0061] Objective function It is composed of the loss function \(l\) and the regularization term \(\Omega\):
[0062] The loss function measures the error between the predicted result and the true result, and usually constrains the loss function by minimizing the error. The specific expression of the loss function is
[0063] The regularization term evaluates the complexity of the XGBoost model to avoid overfitting or underfitting problems. As shown in the following formula:
[0064]
[0065] In the formula: \(\mathcal{L}(\theta)\) represents the objective function of the model; \(l(y_i, \hat{y}_i)\) represents the loss function of a single sample; \(\gamma\) represents the regularization parameter that controls the complexity of the tree structure; \(\lambda\) represents the regularization parameter that controls the size of the leaf node weights; \(\mathbf{w}\) represents the leaf node weight vector of the tree.
[0066] The additional regularization term is used to penalize the complexity of the XGBoost model, smooth the finally learned weights, and avoid overfitting.
[0067] 2) Gradient tree boosting:
[0068] The XGBoost model is trained in an additive manner. Assume \(\hat{y}_i^{(t)}\) represents the prediction of the \(i\)-th sample in the \(t\)-th iteration, and add \(f_t(x_i)\) t to minimize the following objective;
[0069]
[0070] In the formula: \(\mathcal{L}^{(t)}\) represents the objective function at the \(t\)-th iteration; \(\hat{y}_i^{(t - 1)}\) represents the predicted value of sample \(i\) at the \((t - 1)\)-th iteration; \(f_t(x_i)\) t (x i _i t ) represents the predicted value of the \(t\)-th tree for sample \(i\); \(\Omega(f_t)\)
[0071] To quickly optimize the objective function: Perform a second-order Taylor expansion on the objective function, and we can get:
[0072]
[0073] In the formula: \(g_i\) and \(h_i\) are the first-order and second-order gradient statistics on the loss function. \(l(y_i, \hat{y}_i^{(t - 1)})\) represents the current loss value of sample \(i\).
[0074] To simplify the objective of the \(t\)-th training, the constant term is removed:
[0075]
[0076] Define \(I\) j =\(\{i|q(x i ) = j\}\) as the sample set of leaf \(j\), where \(q(x i )\) represents the structure function of the decision tree, which is used to assign the input sample \(x i \) to leaf node \(j\), and expand \(\Omega\):
[0077]
[0078] where: \(w j \) represents the prediction weight of the \(j\)-th leaf node; \(I j \) represents the sample set on leaf node \(j\); \(i\) represents the index of the sample.
[0079] For a fixed structure \(q(x)\), the optimal weight of leaf \(j\) is defined as follows:
[0080]
[0081] The corresponding optimal value is used as a scoring function to measure the quality of the tree structure \(q\), which is defined as follows:
[0082]
[0083] where: represents the scoring function of the quality of the tree structure \(q(x)\) in the \(t\)-th iteration.
[0084] 3). XGBoost adopts an approximate greedy algorithm, weighing computational accuracy and model complexity, and then finding the best split point. First, candidate split points are determined based on the percentiles of the feature distribution, and then the continuous features are mapped into the buckets formed by these candidate points, and the statistical data are aggregated.
[0085] Usually, the percentiles of the features are used to evenly distribute the candidates over the data. Formally, use the multi-set to represent the \(k\)-th feature value and the second-order gradient statistic of each training sample, \(x ik \) represents the feature value of the \(i\)-th sample on feature \(k\); \(h i \) represents the second-order gradient of the \(i\)-th sample.
[0086] Define a rank function \(r k :\) as:
[0087]
[0088] where: r k (z) represents the weighted proportion of samples with feature k value less than z; x represents the feature value of the sample; h represents the second-order gradient of the sample; z represents the value of the current candidate split point; indicates that the sample (x, h) belongs to the sample set corresponding to the k-th feature
[0089] r k (z) represents the proportion of instances where the feature value k is less than z . The goal is to find the candidate split points {s k1 , s k2 , …, s kl}},
[0090] s k1 , s k2 , …, s kl respectively represent the 1st to l-th candidate split points of feature k.
[0091] such that:
[0092]
[0093] where: r k (s k,j ) represents the proportion of samples of feature k on the left side of the candidate split point s k,j ; r k (s k,j+1 ) represents the proportion of samples of feature k on the left side of the candidate split point s k,j+1 ; s k1 represents the 1st candidate split point of feature k; s kl represents the l-th candidate split point of feature k; x ik represents the value of the i-th sample on the k-th feature.
[0094] ε represents an approximation factor. Intuitively, the number of candidate split points is inversely proportional to ε, which means there are approximately 1 / ε candidate points. Here each data point is weighted by h i . Therefore, rewrite the formula:
[0095]
[0096] as a weighted squared loss with label g i / h i and weight h i :
[0097]
[0098] where: represents the optimization objective function in the t-th iteration; ft (x i ) represents the predicted value of the t-th tree for the sample x i ); Ω(f t ) represents the regularization term of the t-th tree; constant represents a constant term independent of f t (x i ).
[0099] In step 3, the IKOA algorithm is used to optimize the key parameters in the XGBoost-based photovoltaic system fault diagnosis model, including the following steps:
[0100] 1) Collect the fault data of photovoltaic modules and inverters, and perform preprocessing such as cleaning and screening; then, use wavelet transform to extract time-frequency domain features, and perform dimensionality reduction through the adaptive Lasso method to select key features, which are divided into a test set and a training set.
[0101] 2) Set the initial XGBoost model hyperparameters and the optimization interval, set the number of iterations of the IKOA algorithm, and generate a population of a given scale through the IKOA algorithm.
[0102] In the present invention, the IKOA algorithm is used to optimize the following 7 hyperparameters in the XGBoost model: the maximum depth of the tree, the learning rate, the number of weak learners (trees), and the sampling ratio when training each tree. Among them, the optimization interval of the maximum depth of the tree is [3, 100], the optimization interval of the learning rate is [0.001, 0.1], the optimization interval of the number of weak learners (trees) is [50, 1000], the optimization interval of the sampling ratio when training each tree is [0.5, 1], the number of iterations is set to 200 generations, and the mathematical model for generating the population is as follows: the same as formula (1):
[0103]
[0104] 3) In the iteration period, obtain the fitness of the multi-fold model through 5-fold cross-validation of the dataset; then, execute the elitist strategy to find and record the sun, that is, the optimal solution; specifically as follows:
[0105] By executing the elitist strategy, ensure the optimal positions of the planets and the sun. As shown in the following formula:
[0106]
[0107] In the formula: represents the position of the i-th planet at the (t + 1)-th iteration after update; represents the candidate position of the i-th planet at the (t + 1)-th iteration; represents the current position of the i-th planet at the (t + 1)-th iteration; represents the position The corresponding fitness value; t represents the number of iterations.
[0108] 4) Calculate the parameters related to the planets in the IKOA algorithm, and update the positions of the planets and the distances from the sun; repeat step 3) until the maximum number of iterations is reached;
[0109] Update the position of the planet through the following formula:
[0110]
[0111] In the formula: represents the scaling factor; represents the velocity of the i-th planet in the t-th iteration; F gi (t) represents the gravitational force acting on the i-th planet; |r| represents a random number; represents the unit vector; represents the central position of the system.
[0112] Update the distance from the sun through the following formula:
[0113]
[0114] In the formula: represents the position of the a-th reference planet; represents the position of the b-th reference planet; represents the adaptive factor that controls the distance between the sun and the current planet at time t; η = (a 2 -1) × r 4 +1 represents a linearly decreasing factor from 1 to -2. Among them, represents a loop control parameter, which gradually decreases from -1 to -2 during T iterations in the whole optimization process.
[0115] 5) Use the obtained sun as the optimal parameters of the XGBoost model;
[0116] The results of the optimal parameters are as follows: the maximum depth of the tree is 50, the learning rate is 0.0959, the number of weak learners (trees) is 174, and the sampling ratio when training each tree is 0.81.
[0117] And evaluate the diagnostic effect based on the following metrics through the test set;
[0118] The metrics include accuracy, recall, precision, and F1-score, and the formulas are expressed as follows:
[0119]
[0120] In the formula: TP and TN represent the correctly predicted positive and negative samples respectively, and FP and FN represent the incorrectly predicted positive and negative samples respectively.
[0121] The present invention relates to a fault diagnosis method for a photovoltaic system based on ALasso-IKOA-XGBoost, and the technical effects are as follows:
[0122] 1) The main advantages of step 1 of the present invention are summarized as follows:
[0123] a. Precise feature selection: The most critical fault features are adaptively screened out through the ALasso method, avoiding the interference of irrelevant features.
[0124] b. Improved diagnostic efficiency: The reduction of the feature dimension significantly reduces the computational complexity of the subsequent diagnostic model.
[0125] c. Enhanced diagnostic accuracy and robustness: Focusing on the features related to faults improves the generalization ability of the model. In addition, compared with the traditional Lasso method, ALasso has the following innovations in feature selection:
[0126] a. ALasso introduces an adaptive weight for each feature, dynamically adjusting the regularization strength according to the importance of the feature, avoiding the limitation of the traditional Lasso that imposes the same constraints on all features.
[0127] b. Through weight adjustment, ALasso can avoid missing key features while reducing the interference of irrelevant features.
[0128] c. There are often multiple fault modes in the data of the photovoltaic system, and the adaptive weight can better cope with the differences in feature importance under different fault modes.
[0129] 2) In step 2 of the present invention, IKOA is improved on the basis of the traditional Kepler optimization algorithm, and its advantages are summarized as follows:
[0130] a. Enhanced randomness: By introducing chaotic optimization strategies such as Logistic mapping and Tent mapping, the algorithm is more random in the process of generating the initial population and iterative update, avoiding premature convergence.
[0131] b. Improved exploration diversity: The improved randomization mechanism increases the diversity of candidate solutions, ensuring that the algorithm can comprehensively cover the search space.
[0132] c. Strong stability: In the face of possible noise or data anomalies in the photovoltaic system, the chaotic mechanism of IKOA can improve the robustness of parameter optimization.
[0133] 3) In step 3 of the present invention, first, the XGBoost algorithm is used for photovoltaic system fault diagnosis. a. By introducing a regularization mechanism (L1 and L2 regularization), overfitting is effectively suppressed, making the model perform more evenly on the training set and the test set. b. XGBoost uses an approximate greedy algorithm to efficiently search for split points, and combined with the optimization in step 2, the model training speed is further improved.
[0134] Secondly, the hyperparameters of the XGBoost model are optimized by IKOA, significantly improving the model performance. The optimized model has the following advantages: a. Precise parameter optimization: Traditional parameter search methods (such as grid search and random search) are often less efficient, while IKOA can quickly converge to the optimal parameter combination. b. Improved diagnostic accuracy: By optimizing the key parameters of XGBoost, the model has a stronger ability to learn the fault characteristics of the photovoltaic system, thus improving the accuracy of the diagnostic results. c. Reduced hyperparameter tuning cost: The automated optimization process greatly reduces the labor cost of parameter debugging and avoids the cumbersome manual adjustment. Finally, the innovation lies in the synergistic effect with the previous two steps. Step 1 selects important features, and step 2 optimizes the parameters, ensuring that both the input data and the hyperparameters of the XGBoost model are in the best state. Combining the previous two steps, a complete and efficient photovoltaic system fault diagnosis process is constructed, significantly improving the accuracy and efficiency of diagnosis. Brief Description of the Drawings
[0135] The present invention will be further described below in conjunction with the drawings and examples;
[0136] Figure 1 It is the overall flowchart of the photovoltaic system fault diagnosis based on ALasso-IKOA-XGBoost proposed by the present invention.
[0137] Figure 2 It is the IV curve characteristic diagram of the photovoltaic module fault proposed by the present invention.
[0138] Figure 3 It is the output characteristic diagram of the inverter open circuit fault proposed by the present invention.
[0139] Figure 4 It is the fault diagnosis result of the photovoltaic module proposed by the present invention.
[0140] Figure 5 It is the fault diagnosis result of the inverter proposed by the present invention.
[0141] Figure 6 It is the schematic diagram of the overall evaluation index of different diagnostic models proposed by the present invention.
[0142] Figure 7 It is the schematic diagram of the IV curve characteristic of the component fault.
[0143] Figure 8 Schematic diagram of the output characteristics of open - circuit faults of different types of inverters. Specific implementation manners
[0144] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.
[0145] Figure 1 This is the overall flowchart of the photovoltaic system fault diagnosis based on ALasso - IKOA - XGBoost proposed by the present invention. First, to fully extract fault features and solve the problem of multicollinearity, wavelet analysis is used to extract the time - frequency feature components of the waveform, and Alasso is used to extract various types of feature components of the fault data. Then, aiming at the problems of slow convergence speed, poor robustness and easy falling into local optimum of KOA, an IKOA algorithm integrating Logistic, Sine and Tent chaotic maps is proposed to generate a better initial solution distribution, parameter setting and position update method. Then, the IKOA algorithm is combined with XGBoost to build an IKOA - XGBoost fault diagnosis model to improve the learning expression and classification accuracy of the photovoltaic fault diagnosis model under complex high - dimensional data conditions.
[0146] Figure 2 This is the IV curve characteristics of the photovoltaic module faults of the method proposed by the present invention. The faults of the photovoltaic module mainly include the following four types: module short - circuit, module open - circuit, local shadow covering, and panel aging and damage. It can be concluded that in the case of short - circuit, the current of the fault branch still remains positive, the branch voltage decreases, and its volt - ampere characteristic diagram presents multiple local maximum power points. After the short - circuit fault occurs, Isc hardly changes, while the changes of Uoc and at the maximum power point are more obvious. When the fault branch is open - circuit, the output current of the fault branch is 0, the current of the parallel branch decreases, and the total output current decreases. When aging occurs, the volt - ampere characteristics change. The more serious the aging degree is, the more the maximum power point decreases, and the changes of Uoc and Isc are not significant. Shadows may cause multi - peak and multi - knee phenomena, that is, multiple local maximum power points, and the global maximum power point will decrease with the increase of the shadow degree, and the open - circuit voltage will also change slightly, and Isc changes little.
[0147] Figure 3 This is the output characteristics of the open - circuit fault of the inverter of the method proposed by the present invention. It can be concluded that when a single - tube is open - circuit, the phase current will be missing half - cycle waveform, while the other phase currents will generate harmonics resulting in waveform distortion, and the output voltage is low. When two tubes in the same phase fail, the phase current will be missing. When two tubes in different phases fail, two phases will be missing half - cycle waveforms, and the waveforms of other phases will be distorted.
[0148] Figure 4 This is the fault diagnosis result of the photovoltaic module of the present invention. Figure 4It can be seen that in the aging and shadow conditions set by the present invention, some aging is misdiagnosed as normal and shadow, some shadow is misdiagnosed as aging and normal, and some component short circuits and open circuits are also misdiagnosed as shadow.
[0149] Figure 5 This is the fault diagnosis result of the inverter proposed by the present invention. From Figure 5 it can be seen that there is no misdiagnosis between large categories, but there is misdiagnosis in the small category fault diagnosis. This is because the waveform differences between large category faults are obvious, but between small category faults, the extracted frequency domain signals may not be very different.
[0150] Figure 6 This is the overall evaluation index of different diagnostic models proposed by the present invention. From Figure 6 it can be seen that the precision, Recall and F1-score of the model proposed by the present invention reach 0.97, 0.96 and 0.96, and it still has an advantage in performance comparison of algorithms. Specifically, the method proposed by the present invention has the highest average accuracy of 96.15%, indicating that it performs the best in terms of overall prediction accuracy. The recall rate represents the proportion of the number of positive samples successfully predicted by the model among all positive samples. And the precision represents the proportion of true positive samples among the samples predicted as positive samples by the model. The method proposed by the present invention also performs the best, which means that it has the best effect in detecting true positive samples.
[0151] Table 1 Comparative analysis table of fault diagnosis accuracy of different diagnostic models
[0152]
[0153]
[0154] Table 1 is the comparative analysis table of fault diagnosis accuracy of different diagnostic models proposed by the present invention. From Table 1, it can be seen that the fault diagnosis accuracy of the Alasso-IKOA-XGBoost algorithm model proposed by the present invention is the highest, reaching 96.25% and 96.13%, which are 4.12% and 3.45% higher than that of KOA respectively, verifying the effectiveness of the multi-chaos strategy to optimize KOA in this paper. Further compared with the commonly used IGWO
[14] at present, it is also 2.99% and 0.89% higher respectively, indicating that the method proposed in this paper effectively improves the global search ability of KOA, and then significantly improves the accuracy of fault diagnosis.
Claims
1. Photovoltaic system fault diagnosis method based on ALasso-IKOA-XGBoost, characterized by The following steps are involved: Step 1: Extract the time-frequency characteristic components of the waveform through wavelet analysis, and use the adaptive Lasso method to extract multiple types of characteristic components of the fault data; Step 2: Propose an IKOA algorithm that integrates Logistic, Sine and Tent chaotic mapping to generate better initial solution distribution, parameter setting and position update method; Step 3: Combine the IKOA algorithm in step 2 with XGBoost to build an IKOA-XGBoost fault diagnosis model.
2. The photovoltaic system fault diagnosis method based on ALasso-IKOA-XGBoost according to claim 1, characterized in that: In step 1, first, wavelet analysis is used to extract the time-frequency characteristics of the output current of the three-phase inverter, which can be expressed as: Where: WT f (a, b) represent the wavelet transform results of signal f(t); f(t) is the current signal to be analyzed; for The conjugate of; a, b are the scaling and translation of the wavelet basis function ψ(t); Next, the adaptive Lasso method is proposed. Its principle is to assign different weights to different penalty items based on the Lasso method. Its expression is as follows: Where: represents the estimated value of regression coefficient; λ2 is the adaptive Lasso penalty weight; β j is the penalty term of the model parameter β, which can shrink the model regression coefficient to zero; Where: represents the adaptive weight, which is used to adjust the importance of the jth regression coefficient in the penalty term; j represents the index of the feature; p represents the number of features; represents the initial estimate; represents the coefficient estimate obtained by the least squares method; X j =(X 1j ,X 2j ,…,X nj ) T represents the predictor variable, where j = 1, 2, …, p; Y=(Y1,Y2,…,Y n ) T represents the response variable, Y1,Y2,…,Y n Respectively represent the true output values of samples 1 to n; T represents the transposition of the vector; The weight expression is as follows: Where: represents the adaptive weight matrix; They represent the estimated weights of the 1st to pth feature assignments respectively; γ is the adaptive Lasso penalty parameter, γ ≥ 0.
3. The photovoltaic system fault diagnosis method based on ALasso-IKOA-XGBoost according to claim 2 is characterized in that: The adaptive Lasso method is used to extract multi-class feature components of fault data; the details are as follows: There are four types of photovoltaic module failures: module short circuit, module open circuit, local shadow covering, and aging and damage of the solar panel. In case of short circuit, the branch voltage U, maximum power point current Im, voltage Um, and open circuit voltage Uoc are selected as characteristic components; in case of open circuit, the branch current I, short circuit current Isc, and maximum power point current Im are selected as characteristic components; in case of aging, the maximum power point current Im and voltage Um are selected as characteristic quantities; in case of local shadow, the maximum power point current Im voltage Um, open circuit voltage Uoc, and the number of local maximum power points are selected as characteristic quantities. In the output characteristics of open-circuit faults of different types of inverters, when a single tube is open, the phase current will be missing half a cycle of the waveform, and other phase currents will have harmonics generated, resulting in waveform distortion, and the output voltage is low; when two tubes in the same phase fail, it will cause the phase current to be missing, and when two tubes in different phases fail, two phases will be missing half a cycle of the waveform, and the waveforms of other phases will be distorted; therefore, the three-phase output current is selected as the characteristic quantity.
4. The photovoltaic system fault diagnosis method based on ALasso-IKOA-XGBoost according to claim 1, characterized in that: In the step 2, 1) Introduce Logistic mapping to initialize the population, replacing completely random initialization, but retaining the initial value The randomness of makes the initial solution more evenly distributed; Where: represents the Logistic mapping value of the jth variable at the i-th iteration; μ represents the parameter of the Logistic mapping; Represents the Logistic mapping value of the j-th variable at the i-1-th iteration; represents the updated value of the jth decision variable of the i-th planet; represents the lower bound of the jth variable of the ith planet; Represents the random sequence value generated by the Logistic map, which is used to convert the variable Mapped to the range between its upper and lower bounds; represents the upper bound of the jth variable of the i-th planet; i represents the index of the planet; j represents the index of the decision variable; N represents the number of candidate solutions in the search space; d represents the dimension of the problem to be optimized; 2) Introducing a control factor based on logistic mapping into the gravity F to fine-tune the amplitude of the planet position change to balance exploration and utilization; Where: r1 represents the Logistic mapping sequence L t The value of L t and L t-1 Respectively represent the results of Logistic mapping at the tth and t-1th iterations; represents the gravitational strength of the i-th planet (solution) at the t-th iteration; μ represents; e i represents the control factor; μ(t) represents the gravitational constant, which usually decreases as the number of iterations t increases; represents the normalized value of the sun's mass; represents the normalized mass of the ith planet; Indicates the normalized distance between the current planet and the sun; ε indicates a positive number; 3) Introducing Sine mapping to replace orbital eccentricity e i , quality random value, retaining the randomness of the initial value S1; Where: S i represents the Sine mapping value at the i-th iteration; a represents a constant used to adjust the amplitude of the Sine mapping value; S i-1 represents the Sine mapping value at the i-1th iteration; e i represents the orbital eccentricity; Where: r2 is equivalent to S t , which is the value generated by the Sine map; S t represents the random sequence value at the tth iteration; S t-1 Represents the value of the Sine map in the previous iteration; M s Represents the Sine mapping value S t Related weight factors; fit s (t) means; fit k (t) represents the fitness value of the tth individual; worst(t) represents the fitness value of the worst individual in the current population; N here represents the size of the population; 4) Introduce a perturbation term based on Tent mapping to fully disrupt the sequence as a global perturbation operator; Where: Represents the output vector of Tent mapping; Represents the random sequence vector of the tth iteration generated by the Tent mapping; represents the output vector of the Tent mapping at the t-1th iteration; T t-1 Represents the output value of the Tent mapping at the t-1th iteration; β represents the segmentation parameter of Tent mapping; V i (t) represents the velocity of planet i at time t; l and Indicates system parameters, used to adjust and The degree of influence on speed; r4 represents the random factor; Represents the current position vector of the i-th candidate solution; Represents the current position vector of the optimal solution in the group; represents the average position vector in the group; R i-nom (t) represents the normalized value of the i-th solution; and Represents an additional disturbance term, which is used to increase the diversity of solutions and prevent the algorithm from falling into the local optimum; and Respectively represent the upper and lower bounds of the ith solution, which are used to limit the search range; Indicates the control amount, used to strengthen The impact on the solution reflects the global search capability; U2 represents the disturbance term, which is used to adjust the scope of the global search; r3 represents the random factor, which adjusts The weight of .
5. The photovoltaic system fault diagnosis method based on ALasso-IKOA-XGBoost according to claim 4 is characterized in that: The IKOA algorithm that integrates Logistic, Sine and Tent chaotic mapping is as follows: The IKOA algorithm integrating Logistic, Sine and Tent chaotic mapping includes equations (1) to (5); First, the Logistic mapping is introduced using formula (1) to initialize the population, replacing the completely random initialization. At the same time, a control factor based on the Logistic mapping is introduced into the gravity F using formula (2) to fine-tune the amplitude of the planet position change. Secondly, the Sine map is introduced using equations (3) and (4) to replace the random values in orbital eccentricity and mass respectively; Finally, the disturbance term based on Tent mapping is introduced using formula (5) and used as a global perturbation operator to fully disrupt the sequence to enhance the global search ability of the algorithm.
6. The photovoltaic system fault diagnosis method based on ALasso-IKOA-XGBoost according to claim 1, characterized in that: In step 3, a photovoltaic system fault diagnosis model based on XGBoost is constructed, including: 1) Regularization objective function: For a dataset with n samples and m features x i represents the input feature vector of the i-th sample; y i represents the true value of the i-th sample; n represents the number of samples; Represents the dimension of the feature space; Represents the target value y i The value range of K-tree CART final prediction output is defined as: Where: represents the predicted value of the i-th sample; φ(x i ) represents the prediction function of the model; f k (x i ) represents the kth decision tree for sample x i The predicted value of; K represents the total number of CART decision trees used in the model; represents the space of CART regression tree; f(x) represents the prediction function of decision tree; q represents the structure vector; w represents the leaf weight; Denotes a vector space consisting of T real numbers; Each function f k Corresponding to an independent tree structure vector q and leaf weight w; q points to the corresponding leaf label from the sample, and each leaf node of each CART tree corresponds to a continuous score value, i.e., weight; the score of the i-th node is w i ; G is the number of leaf nodes; w q(x) is the score of sample x, that is, the model prediction value; For each sample, each CART classifies it into leaf nodes according to different classification rules, and obtains the final prediction result by accumulating the scores w of the corresponding leaves; Objective Function It consists of the loss function l and the regular term Ω: The loss function measures the error between the predicted result and the actual result, and constrains the loss function by minimizing the error. The specific expression of the loss function is The regularization term evaluates the complexity of the XGBoost model to avoid overfitting or underfitting problems; as shown in the following formula: Where: represents the objective function of the model; Represents the loss function of a single sample; γ represents the regularization parameter that controls the complexity of the tree structure; λ represents the regularization parameter that controls the weight of the leaf node; w represents the leaf node weight vector of the tree; 2) Gradient Tree Enhancement: The XGBoost model is trained in an additive way; represents the prediction of the i-th sample in the t-th iteration, adding f t To minimize the following objective; Where: represents the objective function at the tth iteration; represents the predicted value of sample i at the t-1th iteration; f t (x i ) represents the predicted value of the t-th tree for sample i; Ω(f t ) represents the regularization term of the t-th tree; n represents the total number of samples; To quickly optimize the objective function: Performing a second-order Taylor expansion on the objective function, we can obtain: Where: are the first-order and second-order gradient statistics on the loss function; Represents the current loss value of sample i; In order to simplify the objective of the t-th training, remove the constant term: Definition I j ={i|q(x i )=j} as the sample set of leaf j, q(x i ) represents the structural function of the decision tree, which is used to transform the input sample x i Assign to leaf node j and expand Ω: Where: w j Represents the prediction weight of the jth leaf node; I j represents the sample set on leaf node j; i represents the index of the sample; For a fixed structure q(x), the optimal weight of leaf j is The definition is as follows: The corresponding optimal value is used as a scoring function to measure the quality of the tree structure q, which is defined as follows: Where: The scoring function representing the quality of the tree structure q(x) in the tth iteration; 3). XGBoost uses an approximate greedy algorithm to balance computational accuracy and model complexity to find the best split point. First, candidate split points are determined based on the percentiles of feature distribution, and then continuous features are mapped to buckets formed by these candidate points, and statistical data are clustered. The percentiles of the features are used to evenly distribute the candidates over the data; multiple sets are used represents the kth eigenvalue and second-order gradient statistics of each training sample, x ik represents the eigenvalue of the i-th sample on feature k; h i represents the second-order gradient of the i-th sample; Define a rank function r k : for: Where: r k (z) represents the weighted proportion of samples whose feature k value is less than z; x represents the feature value of the sample; h represents the second-order gradient of the sample; z represents the value of the current candidate split point; Indicates that the sample (x, h) belongs to the sample set corresponding to the kth feature r k (z) means that the eigenvalue k is less than z The goal is to find the candidate segmentation point {s k1 ,s k2 ,…,s kl },s k1 ,s k2 ,…,s kl Respectively represent the 1st to 1st candidate splitting points of feature k; So that: Where: r k (s k,j ) indicates that feature k is at the candidate split point s k,j The sample proportion on the left; r k (s k,j+1 ) indicates that feature k is at the candidate split point s k,j+1 The sample proportion on the left; s k1 represents the first candidate split point of feature k; s kl represents the lth candidate split point of feature k; x ik Represents the value of the i-th sample on the k-th feature; ε represents an approximation factor; intuitively, the number of candidate split points is inversely proportional to ε, which means there are 1 / ε candidate points; here each data point is represented by h i weighted; therefore, the formula: Rewrite with label g i / h i and weight h i The weighted square loss of: Where: represents the optimization objective function in the tth iteration; f t (x i ) represents the t-th tree for sample x i The predicted value of Ω(f t ) represents the regularization term of the tth tree; constant represents the regularization term of f t (x i ) is an irrelevant constant term.
7. The photovoltaic system fault diagnosis method based on ALasso-IKOA-XGBoost according to claim 6, characterized in that: In step 3, the IKOA algorithm is used to optimize the key parameters in the XGBoost-based photovoltaic system fault diagnosis model, including the following steps: 1) Collect PV module and inverter fault data and perform preprocessing such as cleaning and screening; then, use wavelet transform to extract time and frequency domain features, and use adaptive Lasso method to reduce dimension, select key features, and divide them into test set and training set; 2) Initialize the XGBoost model hyperparameters and optimization interval, set the number of IKOA algorithm iterations, and generate a population of a given size through the IKOA algorithm; The IKOA algorithm is used to optimize the seven hyperparameters of the maximum depth of the tree, the learning rate, the number of weak learners (trees), and the sampling ratio when training each tree in the XGBoost model; the optimization interval of the maximum depth of the tree is [3, 100], the optimization interval of the learning rate is [0.001, 0.1], the optimization interval of the number of weak learners (trees) is [50, 1000], the optimization interval of the sampling ratio when training each tree is [0.5, 1], the number of iterations is set to 200 generations, and the mathematical model of the generated population is as follows: Same as formula (1): 3) During the iteration cycle, the multi-fold model fitness is obtained through a 5-fold cross-validation data set; then, the elite strategy is executed to find and record the sun, which is the best solution; the details are as follows: By implementing an elitist strategy, the optimal positions of the planets and the Sun are ensured; as shown below: Where: represents the position of the ith planet at the t+1th iteration after update; represents the candidate position of the ith planet at the t+1th iteration; represents the current position of the ith planet at the t+1th iteration; Indicates location The corresponding fitness value; t represents the number of iterations; 4) Calculate the planet-related parameters in the IKOA algorithm and update the planet position and the distance from the sun; repeat operation 3) until the maximum number of iterations is reached; Update the planet's position by: Where: represents the scaling factor; represents the velocity of the ith planet in the tth iteration; F gi (t) represents the gravitational force on the i-th planet; |r| represents a random number; represents a unit vector; Indicates the central location of the system; Update the distance to the sun by: Where: represents the position of the ath reference planet; represents the position of the bth reference planet; represents the adaptive factor controlling the distance between the sun and the current planet at time t; η=(a2-1)×r4+1 represents a linear decreasing factor from 1 to -2; where, represents a loop control parameter, which gradually decreases from -1 to -2 in T iterations during the entire optimization process; 5) The obtained results are used as the optimal parameters of the XGBoost model; and the diagnostic effect is evaluated based on the test set using the following indicators; the indicators include accuracy, recall, precision and F1-score, and the formula is as follows: Where TP and TN represent the correctly predicted positive and negative samples, respectively, and FP and FN represent the incorrectly predicted positive and negative samples, respectively.
Citation Information
Patent Citations
System and method for checking insulativity of power-off cable
CN104569747A
Power transmission line fault identification method and system based on fault feature matrix and IPSO-WNN
CN116008859A
Frequency converter fault identification method and monitoring device based on XGBoost fault diagnosis model
CN117421677A
Fault diagnosis method and system of photovoltaic array, electronic equipment and storage medium
CN117473245A
Characteristic waveform signals decomposing method for extracting dynamic information of machinery
CN1376905A