Method for realizing high-precision intelligent diagnosis of photovoltaic array fault based on integrated learning model CatBoost
By improving the IV curve correction and combining the CatBoost model with the EKSSA optimization algorithm, the accuracy and efficiency issues of photovoltaic array fault diagnosis under complex operating conditions are solved, realizing high-precision photovoltaic array fault diagnosis, which is suitable for intelligent operation and maintenance of distributed and centralized photovoltaic systems.
Patent Information
- Application Number
- CN202510009730.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-12-30
AI Technical Summary
Existing photovoltaic array fault diagnosis methods suffer from low accuracy, insufficient efficiency, and poor adaptability under complex operating conditions. Traditional methods rely on a large amount of manpower and resources and have limited ability to diagnose hidden faults or diverse fault types in complex environments.
By combining an improved IV curve correction algorithm and the CatBoost model with the EKSSA optimization algorithm, high-precision real-time diagnosis of photovoltaic array faults can be achieved by correcting the IV curve data and optimizing the hyperparameters of the CatBoost model.
It achieves high-precision classification of photovoltaic array faults under different environmental conditions, improves diagnostic efficiency and robustness, reduces operation and maintenance costs, and is suitable for intelligent operation and maintenance of small-scale distributed and large-scale centralized photovoltaic systems.
Smart Images

Figure CN121234239A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of intelligent diagnosis and optimization of photovoltaic systems, and specifically relates to a photovoltaic array fault diagnosis method based on an integrated learning model CatBoost and an improved sparrow search algorithm. By introducing an I-V curve correction algorithm, the influence of environmental variables (such as irradiance and temperature) on diagnostic accuracy is addressed. The CatBoost model is used to achieve high-precision real-time fault diagnosis under small sample conditions. The improved sparrow search algorithm (EKSSA) that integrates elite reverse learning strategy and Cauchy Gaussian mutation strategy is used to optimize model hyperparameters, thereby improving diagnostic efficiency and model performance. The application is suitable for real-time monitoring, fault diagnosis and predictive maintenance of photovoltaic arrays, and has significant advantages in improving the reliability and operating efficiency of photovoltaic power generation systems. BACKGROUND
[0002] Photovoltaic arrays operate in complex outdoor environments and are susceptible to environmental factors such as temperature, irradiance, wind speed, and component aging, leading to open-circuit, short-circuit, aging, and shadow shading faults. These faults can significantly reduce system power generation efficiency and may affect the safe operation of the system, increasing operation and maintenance costs. Currently, traditional fault diagnosis methods include manual inspection, infrared imaging detection, and mathematical modeling based on equivalent circuits. These methods typically rely on a large amount of manpower and resources, have long diagnostic cycles, low efficiency, and limited diagnostic capabilities for hidden faults or diverse fault types in complex environments. In addition, some techniques only use partial feature information from the I-V curve, failing to fully exploit the potential fault information, resulting in some faults not being detected in a timely manner. With the introduction of machine learning technology, data-driven photovoltaic array fault diagnosis methods have gradually emerged. However, these methods are highly dependent on large-scale labeled data and are susceptible to sample size limitations during model training. Additionally, the models have high complexity and parameter adjustment difficulties. In actual small sample or complex working conditions, the diagnostic accuracy and efficiency of existing methods still need to be improved. SUMMARY
[0003] The application relates to the field of photovoltaic array fault diagnosis and proposes a photovoltaic array fault diagnosis method and system based on a deep optimization integrated learning model CatBoost, aiming to solve the problems of low diagnostic accuracy, insufficient efficiency, and poor adaptability of existing technologies in complex working conditions. Photovoltaic arrays, as an important form of clean energy generation, are exposed to complex outdoor environments for a long time and are affected by factors such as temperature, irradiance, wind speed, and humidity, making them prone to open-circuit, short-circuit, aging, and shadow shading faults. These faults not only reduce power generation efficiency but also may cause safety hazards. Therefore, developing an intelligent, efficient, and robust photovoltaic fault diagnosis method is of great significance for improving the operational reliability of photovoltaic systems.
[0004] The application proposes an improved photovoltaic array fault diagnosis method, combining the physical characteristics of photovoltaic arrays and intelligent algorithm optimization. The main technical innovations include the following aspects:
[0005] The I-V curve of a photovoltaic array is a key diagnostic indicator that contains a wealth of information about the health status of photovoltaic modules. However, this characteristic curve will change significantly due to environmental variables such as temperature and irradiance, thereby interfering with the accurate extraction of fault features. The application proposes an improved I-V curve correction algorithm based on the IEC 60891 standard, which standardizes the actual collected I-V curve to a uniform test condition (STC: 25℃, 1000W / m 2 ) by dynamically adjusting the temperature and irradiance. This improved correction algorithm not only improves the accuracy of feature extraction but also has stronger adaptability and can be widely applied to photovoltaic array fault diagnosis under different environmental conditions.
[0006] In the selection of fault diagnosis algorithms, the application uses the CatBoost integrated learning model based on gradient boosting. This model performs well in handling small sample data and high-dimensional nonlinear data, can automatically handle data missing problems, and reduces the complexity of data preprocessing. By inputting the corrected I-V curve data into the CatBoost model, the application realizes real-time high-precision classification of various fault types of photovoltaic arrays such as open circuit, short circuit, aging, and shadow shading. In addition, to verify the model performance, the application uses the confusion matrix to calculate the accuracy, precision, recall, and F1 score of the diagnosis, and achieves excellent diagnosis effect in both simulated data and actual field data.
[0007] The diagnostic performance of the CatBoost model is closely related to the selection of its key hyperparameters such as tree depth and learning rate. To further improve the performance of the model, the application introduces an improved sparrow search algorithm (EKSSA) to optimize the hyperparameters of the CatBoost model. EKSSA enhances the optimization ability of the traditional sparrow search algorithm through the following innovative strategies: elite reverse learning strategy: select individuals with lower fitness in each iteration, expand the search space by generating reverse solutions, and select better solutions to improve population diversity and avoid falling into local optima. Cauchy Gaussian mutation strategy: mutate the optimal individual, introduce random disturbance to increase the diversity of the search, and dynamically adjust the step size parameter to balance global exploration and local development capabilities, thereby accelerating algorithm convergence.
[0008] The CatBoost model optimized by EKSSA performs significantly in terms of diagnostic accuracy and robustness under complex working conditions, effectively solving the problems of complex parameter adjustment and low efficiency of traditional methods.
[0009] The method has strong real-time performance and robustness. Through rapid acquisition and analysis of I-V curve data, the application can realize real-time diagnosis of photovoltaic array faults, ensure timely discovery of problems and provide diagnosis results. Due to the use of improved algorithm design, the diagnosis performance of the application remains stable under different geographical environments and complex operating conditions, and can adapt to various practical application scenarios.
[0010] The application is not only suitable for small-scale distributed photovoltaic systems, but also can be popularized to intelligent operation and maintenance of large-scale centralized photovoltaic power stations. By improving the accuracy and efficiency of fault detection, the application can significantly reduce the operation and maintenance cost of the photovoltaic power generation system, prolong the service life of the system, and improve the overall economic benefit. In addition, the improved algorithm and diagnosis framework proposed by the application also provide reference value for fault diagnosis in other energy fields, and have wide popularization and application potential.
[0011] The application solves the limitations of traditional photovoltaic fault diagnosis technology in terms of diagnosis accuracy, efficiency and adaptability through innovative I-V curve correction algorithm, application of CatBoost model and EKSSA optimization strategy. Experimental results show that the method can realize high-precision fault classification under different scenarios, and the optimized model has higher stability and popularization value. Through the intelligent diagnosis process, the application provides a comprehensive solution for real-time monitoring, fault diagnosis and predictive maintenance of photovoltaic arrays.
[0012] In summary, the application proposes an intelligent diagnosis method combining physical modeling and machine learning, breaks through the technical bottleneck of traditional diagnosis methods, realizes high-precision, high-efficiency and wide adaptability of photovoltaic array fault detection, and provides strong support for intelligent operation and maintenance of photovoltaic power generation systems. BRIEF DESCRIPTION OF DRAWINGS
[0013] In order to better understand the technical solutions of the application, the drawings involved in the application will be briefly described below. The described drawings are only part of the embodiments of the application. For those skilled in the art, other possible embodiments can be derived from these drawings without creative work and adjustment.
[0014] Figure 1 is a photovoltaic array fault diagnosis flowchart of the application.
[0015] Figure 2 is a photovoltaic array fault diagnosis flowchart of the application.
[0016] Figure 3 is a flowchart of the improved sparrow search algorithm of the application.
[0017] Figure 4Is the elite reverse strategy + Cauchy Gaussian variation strategy improved sparrow search algorithm effect diagram of the application.
[0018] Figure 5 Is the CatBoost training set loss function curve diagram of the application.
[0019] Figure 6 Is the test set confusion matrix diagram of the application.
[0020] Figure 7 Is the analysis result diagram of the model performance metric under different operating conditions of the application.
[0021] Specific implementation steps
[0022] To achieve high-precision intelligent diagnosis of photovoltaic array faults, the application proposes a complete implementation process, covering key links such as data acquisition and preprocessing, I-V curve correction, model construction and optimization, and performance verification. The following are the specific implementation steps:
[0023] Step 1: Data acquisition and preprocessing and analysis of I-V curve
[0024] The original data including electrical data, environmental data and fault labels are collected from the operation of photovoltaic array. The electrical data mainly includes the characteristic points of current-voltage (I-V) curve, such as open circuit voltage (U OC ), short circuit current (I SC ), maximum power point voltage (U MP ) and current (I MP ). Environmental data covers temperature, irradiance and other information, which is used to correct the fluctuations caused by environmental factors. Fault labels include normal operation data and labeled information of known fault modes (such as open circuit, short circuit, aging, shadow shielding). Illumination and temperature are two important factors affecting the performance of photovoltaic array, and their influence on I-V curve is significant. As shown in Figure 1 , when G increases, short circuit current I sc and open circuit voltage U oc both show an upward trend; when the temperature rises, U oc decreases and I sc increases. Therefore, when analyzing the I-V curve of photovoltaic array, the influence of illumination and temperature on the performance of photovoltaic cell must be considered
[0025] Step 2: I-V curve correction and data preprocessing
[0026] According to the analysis of I-V curve in step 1, the I-V characteristic curve of photovoltaic array is often significantly affected by irradiance and temperature. Therefore, when diagnosing the fault of photovoltaic array, the influence of different environmental conditions on the performance of the same fault type feature must be considered to ensure the accuracy and reliability of the diagnostic model. In order to eliminate the influence of environmental factors on the feature curve and make the diagnostic model more focused on capturing information related to fault types, the I-V curve is corrected to the same environmental conditions, here taking STC (25℃, 1000W / m 2 ) as the target condition. The reference I-V curve correction method is program 1 and program 2 in IEC 60891 standard. However, these methods have limited performance when there is a fault. Therefore, the present invention proposes an improved correction method, which only requires actual current, voltage and corresponding environmental condition parameters, and is easy to implement and operate in practical applications, as shown in the following two formulas:
[0027]
[0028] U STC =U measured +(U oc,STC -U oc,measured )
[0029] In order to ensure that the diagnostic model can quickly iterate training and the training results are true and effective, the Z-score normalization method is used to standardize the I-V curve and environmental data, eliminate the dimension difference and improve the convergence of the model. Noise removal and filtering: use statistical analysis method to filter abnormal values and noise signals to ensure the reliability of data quality. Data distribution: divide the data set into training set and test set according to the ratio of 8:2, which are used for model training, performance test and final verification respectively.
[0030] Step 3: Model construction
[0031] 3.1 Selection and training of diagnostic model
[0032] The present invention uses CatBoost model as the core classification algorithm. CatBoost model is based on gradient boosting framework, which can handle high-dimensional nonlinear data, automatically adapt to missing values and category features, and improve the stability and generalization ability of diagnosis. The training set data is used to initially train the CatBoost model, and the basic parameters (such as the number of iterations, learning rate, etc.) are set. Monitor the loss function value of the model during training to evaluate the preliminary classification ability of the model.
[0033] 3.2 Selection of optimization algorithm
[0034] To further improve the performance of the model, a sparrow search algorithm (SSA) is proposed to optimize the hyperparameters of the CatBoost model, such as tree depth, learning rate, sampling rate, etc. The optimization process of the sparrow search algorithm is as follows: let the target be to minimize the loss function L(θ) of the classification task, where θ is the combination of hyperparameters of the CatBoost model, and define the hyperparameter space S, where each hyperparameter θ i has a defined domain D i . In each iteration, according to the evaluation results of the objective function, the sparrow selects a group of hyperparameters for capture. The sparrow adjusts the hyperparameters according to the current optimal solution θ best , and the specific selection rules are as follows.
[0035] θ prey (t) = θ best (t) + α × rand() × (P max -P min )
[0036] 3.3 Introducing the elite reverse learning strategy to optimize and improve the sparrow search algorithm
[0037] Elite reverse learning strategy: generate reverse solutions for individuals with low fitness, expand the search space and select better solutions to enhance the global search ability of the algorithm. It solves the problem of uneven population distribution and reduced population diversity in the SSA algorithm. As Figure 3 shown, the core idea is to select individuals with low fitness in each generation for special processing. In the hyperparameter optimization process of the CatBoost model, the elite reverse learning strategy can be applied to the hyperparameter population of each generation. Suppose N individuals with low fitness are selected, then these individuals are specially processed, such as re-initialization or introduction of new search strategies, to increase the diversity of the search. The principle is to find the inverse solution of a problem and evaluate the original solution and the inverse solution to select the better solution as the next generation individual.
[0038] 3.4 Introducing the Cauchy Gaussian mutation strategy to optimize and improve the sparrow search algorithm
[0039] Randomly mutate the current optimal solution to increase search diversity and avoid the algorithm falling into local optima. The sparrow search algorithm finds a local optimal solution near the initial solution, and it is easy to fall into it and cannot find a better solution. Especially in high-dimensional space, due to the complexity of the search space, the algorithm may take longer to converge to the global optimal solution. As Figure 3 To solve the problem of the sparrow search algorithm falling into local optima, the Cauchy Gaussian mutation strategy is introduced. The individual is subjected to mutation operation, introducing randomness to increase the diversity of the search. With the increase of the number of iterations, the search step is gradually reduced to improve the search accuracy and ensure that the algorithm converges to the global optimal solution.
[0040] 3.5 Optimization process and effect analysis
[0041] Population initialization: A set of candidate solutions containing CatBoost hyperparameters is randomly generated; Fitness evaluation: The loss function value of each candidate solution is calculated using validation set data; Population update: The population is iteratively updated according to the elite inverse learning strategy and the Cauchy-Gaussian mutation strategy; Output results: The candidate solution with the highest fitness is selected as the optimal hyperparameter configuration of the CatBoost model. As shown in Table 1, to further verify the improvement effect of the elite inverse learning strategy and the Cauchy-Gaussian mutation strategy on the SSA algorithm compared with other strategies, a series of different benchmark functions were considered to evaluate the convergence speed and the quality performance of the optimal solution of various strategy combinations. Each optimization algorithm is initialized by running iteratively multiple times, with a maximum of 500 iterations and a population size of 30. Figure 4 As shown, in each iteration, the convergence curve is recorded, and the performance of different strategy combinations is compared by plotting the convergence curve results.
[0042] Table 1 Optimization Strategies
[0043]
[0044] Step 4: Model Validation and Performance Comparison Analysis
[0045] The optimized CatBoost model was applied to the validation set to test its ability to classify various fault types. The accuracy, precision, recall, and F1 score of the model were calculated using the confusion matrix to comprehensively evaluate its performance. The impact and effectiveness of this invention on the CatBoost diagnostic model were analyzed by comparing the loss functions before and after the optimization. Figure 5 As shown, the loss functions of all optimization algorithms gradually converge and eventually stabilize with increasing iterations. However, the CatBoost model optimized by the EKSSA algorithm exhibits faster convergence and a smaller minimum loss function, demonstrating significant improvements in both convergence speed and error reduction, thus resulting in a more pronounced performance enhancement. Figure 6 As shown, the diagnostic results for the five states in the test set are visualized using a confusion matrix, and evaluation metrics are calculated based on the confusion matrix. Figure 7As shown, for both open-circuit and short-circuit conditions, the three models achieve the same diagnostic performance, reaching 100%. Analysis of recall before and after optimization reveals that the CatBoost model optimized with the SSA algorithm shows improvement compared to the unoptimized model. Its diagnostic performance is significantly improved in normal, aging fault, and shadowed conditions. The CatBoost model optimized with the improved EKSSA algorithm achieves the best diagnostic performance, with only one false diagnosis and improved diagnostic performance for complex faults. Analysis of accuracy shows that, except for the open-circuit and short-circuit conditions where the three models perform similarly, the EKSSA-CatBoost model performs best in other conditions. Comparison of F1 scores indicates that the EKSSA-CatBoost model shows significant improvement in diagnostic performance under normal, aging, and shadowed conditions.
[0046] The above-described specific implementations can be partially adjusted by those skilled in the art in different ways without departing from the principles and purpose of the present invention. The scope of protection of the present invention is defined by the claims and is not limited to the above-described specific implementations. All implementation schemes within the scope of the claims are bound by the present invention.
Claims
1. A photovoltaic array fault diagnosis method based on an ensemble learning model, characterized in that, The method comprises the following steps: data acquisition: collecting current-voltage characteristic curve (I-V curve), environmental temperature and light intensity data from photovoltaic array; Data preprocessing: normalize the collected I-V curve, and eliminate the influence of temperature and irradiance on fault characteristics based on the improved environmental correction algorithm; Model training: using CatBoost model to train the processed data to generate a preliminary fault diagnosis model; super parameter optimization: optimizing the key super parameters of CatBoost model based on the improved sparrow search algorithm, including tree depth and learning rate; fault classification: classifying the fault data through the optimized CatBoost model and outputting the fault type.
2. The method of claim 1, wherein, The environmental correction step employs an improved I-V curve correction algorithm based on the IEC 60891 standard, which corrects the curve to the standard test conditions (STC: 25°C, 1000 W / m 2 ) for the dynamic changes of the curve in different environmental conditions, thus ensuring the accuracy of the characteristic data. Only the actual current, voltage and corresponding environmental condition parameters are required, which is easy to implement and operate in practical applications. The temperature coefficient a is an adjustable parameter that can be adjusted according to the actual measurement data or the parameters provided by the manufacturer, so it can adapt to different types and specifications of photovoltaic modules. V STC = V measured + (V oc,STC - V oc,measured ).
3. The method of claim 1, wherein, The data preprocessing further comprises: normalization operation: using Z-score standardization method to normalize the data, removing the dimension effect and ensuring the consistency of the data; noise filtering: using statistical analysis and filtering algorithm to remove abnormal values and noise signals in the collected data to improve the reliability of the model.
4. The method of claim 1, wherein, The CatBoost model is optimized through the following steps: initializing the model: setting the basic super parameters, including the number of trees, splitting strategy and objective function; incremental training: gradually constructing decision trees based on gradient boosting algorithm to improve the classification accuracy; super parameter adjustment: dynamically adjusting the key super parameters of the model, such as learning rate, tree depth and tree number, by combining sparrow search algorithm.
5. The method of claim 1, wherein, Using complete I-V curve data as model input, through embedded feature learning mechanism, the curve shape change trend is captured, so that high precision results are obtained in the diagnosis of complex fault types such as short circuit, open circuit, aging and shadow shielding.
6. A photovoltaic array fault diagnostic system characterized by, The method comprises the following modules: data acquisition module, used for acquiring I-V curve data and environmental parameters of photovoltaic array in real time; Data preprocessing module, used for environmental correction, normalization and noise filtering of data; model calculation module, embedded with CatBoost model, responsible for photovoltaic array fault diagnosis; optimization module, based on improved sparrow search algorithm to optimize model super parameters; Output display module, used for visualizing the diagnosis results, including fault type, diagnosis confidence and recommended treatment scheme.
7. The method of claim 6, wherein, The sparrow search algorithm optimizes the super parameters of the CatBoost model by introducing an elite reverse learning strategy, which comprises the following steps: selecting the individuals with poor performance in each iteration, and generating new solutions through reverse learning; calculating the fitness of the reverse solution and the original solution, retaining the solution with better fitness to form a new population; improving the diversity and global search ability of the population to avoid falling into local optimal solution.
8. The method of claim 6, wherein, The sparrow search algorithm optimizes the super parameters of the CatBoost model by introducing Cauchy Gaussian mutation strategy, which comprises the following steps: randomly mutating the individual with the best performance in the current population according to Cauchy distribution and Gaussian distribution to generate a new solution; determining whether to accept the mutated solution according to the fitness difference between the mutated solution and the original solution; balancing the local development and global exploration ability by dynamically adjusting the mutation amplitude, improving the search efficiency and optimization performance.
9. A method for diagnosing faults in a photovoltaic array based on the system of claim 6, characterized by, Achieved by the following steps: data input: collect photovoltaic array data on site, input to diagnostic system; model analysis: analyze data through CatBoost model, generate diagnostic results including open circuit, short circuit, aging and shadowing fault; result verification: analyze the classification accuracy of the model using confusion matrix, and output performance evaluation indicators (such as accuracy, recall rate and F1 score).
10. An operation and maintenance optimization method based on a photovoltaic array fault diagnosis model, characterized in that, By combining the diagnostic results and historical operation and maintenance data, a predictive maintenance plan for photovoltaic arrays is developed, including: fault priority ranking, ranking maintenance tasks according to diagnostic confidence and fault impact; dynamic operation and maintenance scheduling, using diagnostic system recommendations to optimize photovoltaic array power generation efficiency and operation and maintenance resource allocation; fault recurrence learning, through the online learning mechanism of new fault cases, gradually improving the adaptability and robustness of the diagnostic system.