A method and system for photovoltaic device cluster benchmarking and latent fault early warning

By combining real-time data acquisition, cluster analysis, and the LSTM-XGBoost model, the problems of incomplete data and limitations of benchmarking analysis in photovoltaic equipment fault early warning are solved, enabling accurate fault early warning and operation and maintenance optimization of photovoltaic equipment clusters, and improving equipment health status assessment and power generation efficiency.

CN121261330BActive Publication Date: 2026-05-01SHANDONG ARTAPLAY INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG ARTAPLAY INTELLIGENT TECH CO LTD
Filing Date
2025-12-03
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing photovoltaic equipment fault early warning methods suffer from incomplete data collection and a lack of benchmarking analysis from a cluster perspective. This results in untimely detection of hidden faults and a high false alarm rate. The early warnings are not suitable for complex operating conditions and the entire life cycle of photovoltaic equipment, and they do not quantify the loss of power generation revenue.

Method used

The system collects basic parameters and implicit correlation data of photovoltaic equipment in real time through sensors, performs data preprocessing using interpolation and outlier removal methods, constructs a group health benchmark using clustering algorithms and performs individual difference calibration, combines an LSTM-XGBoost fusion model to predict fault risks, generates implicit fault warnings, and prioritizes them based on revenue loss assessment.

Benefits of technology

It enables precise benchmarking analysis and latent fault early warning for photovoltaic equipment clusters, reduces the failure rate, improves the accuracy and timeliness of fault early warning, optimizes the operation and maintenance management of photovoltaic power plants, and improves equipment lifespan and power generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121261330B_ABST
    Figure CN121261330B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of photovoltaic equipment monitoring fault early warning, and relates to a photovoltaic equipment cluster benchmarking analysis and implicit fault early warning method and system. The system collects basic parameters and implicit associated data of photovoltaic equipment through a data acquisition module to ensure data comprehensiveness and timeliness. Through a data fusion module, the collected data is processed and fused to generate a standardized data sequence. Through a cluster benchmarking module, the photovoltaic equipment is divided into different clusters, a group health benchmark is constructed, and the model accuracy is improved through individual difference calibration. Through a dynamic early warning module, early warning information is generated, and a basis is provided for subsequent fault troubleshooting. Through an intelligent operation and maintenance module, early warning is prioritized and an operation and maintenance scheme is generated, which is pushed to operation and maintenance personnel for optimization guidance. The present application can realize accurate benchmarking analysis and implicit fault early warning of photovoltaic equipment clusters, and reduce the failure rate.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters Technical Field

[0001] This invention relates to the field of photovoltaic equipment monitoring and fault early warning technology, specifically to a method and system for photovoltaic equipment cluster benchmarking analysis and latent fault early warning. Background Technology

[0002] With the continuous expansion of photovoltaic power generation, the operation and maintenance management of photovoltaic equipment clusters faces enormous challenges. In existing technologies, photovoltaic equipment fault early warning methods often have a single data collection dimension and lack benchmarking analysis from a cluster perspective, resulting in untimely detection of hidden faults and a high false alarm rate.

[0003] To this end, invention patent CN119834735A discloses a remote fault diagnosis method for photovoltaic power generation equipment, characterized in that the method includes:

[0004] The system collects photovoltaic (PV) equipment attribute information from distributed PV devices within the target area to obtain a PV equipment attribute set. This attribute set includes a PV equipment model set, a performance parameter set, an installed capacity set, and a location coordinate set. Based on predetermined power generation influencing factors and the location coordinate set, the system calls regional environmental monitoring logs to perform environmental change analysis. Cluster analysis is then performed on the distributed PV devices based on the environmental change feature set, the PV equipment model set, and the performance parameter set to identify multiple PV equipment sets. Under preset monitoring nodes, multiple power generation sets from these PV equipment sets are obtained. Based on these power generation sets and the installed capacity set, abnormal power generation is identified to determine abnormal PV equipment sets. Finally, a first abnormal PV device is selected from these abnormal PV equipment sets. The associated monitoring array of the first abnormal photovoltaic device is activated to perform photovoltaic device status monitoring and environmental monitoring, obtaining status monitoring data and environmental monitoring data. The first abnormal photovoltaic device is any one of the abnormal photovoltaic devices in the abnormal photovoltaic device set. Based on the photovoltaic device model, performance parameters, and environmental monitoring data of the first abnormal photovoltaic device, an associated fault feature library and an associated fault prediction model are determined. The status monitoring data are input into the associated fault feature library and the associated fault prediction model respectively. The predicted fault type is determined by fusing the output results. A first fault diagnosis result is generated based on the predicted fault type and the location coordinates of the first abnormal photovoltaic device. A photovoltaic device fault diagnosis report for the target area is generated based on the first fault diagnosis result.

[0005] The above-mentioned technical solutions have made progress in photovoltaic equipment clustering and abnormal state identification, but the following technical problems still exist: data collection does not cover implicit fault-related factors such as the aging degree of photovoltaic equipment and the state of partial shading, resulting in incomplete data collection; benchmarking analysis is only based on environmental and photovoltaic equipment attribute clustering and lacks the construction of group health benchmarks and individual difference calibration, resulting in limitations in benchmarking analysis; fault early warning is not adapted to complex operating conditions and the entire life cycle of photovoltaic equipment, and the lack of quantification of power generation revenue loss leads to unclear early warning priorities; thus, fault early warning is not suitable.

[0006] In view of this, it is very necessary to provide a method and system for benchmarking analysis of photovoltaic equipment clusters and early warning of hidden faults in order to solve the above-mentioned defects in the prior art. Summary of the Invention

[0007] The purpose of this invention is to solve the problems of incomplete data collection, limitations in benchmarking analysis, and incompatibility of fault early warning. In response to the technical defects of the above-mentioned existing technologies, this invention provides a method and system for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters to solve the above-mentioned technical problems.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] In a first aspect, the present invention provides a method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters, comprising the following steps:

[0010] Step S1: Data acquisition step, which involves collecting and uploading the basic parameter data and implicit correlation data of the photovoltaic equipment in real time through sensors;

[0011] Step S2: The data fusion step involves preprocessing data from different sources using interpolation and outlier removal methods, and unifying them into a standard format.

[0012] Step S3: The cluster benchmarking analysis step involves using clustering algorithms to divide similar photovoltaic devices into clusters based on the operating parameters and historical data of the photovoltaic equipment, constructing a group health benchmark, and calibrating according to the differences in photovoltaic equipment.

[0013] Step S4: The dynamic early warning step uses the LSTM-XGBoost fusion model, combined with the operating data of photovoltaic equipment and cluster benchmarks, to predict the future failure risk of photovoltaic equipment and generate latent fault early warning.

[0014] Step S5: The steps of intelligent operation and maintenance are as follows: based on the fault warning and the assessment of revenue loss, the warning information is sorted according to priority and an operation and maintenance decision plan is generated and pushed to the operation and maintenance personnel.

[0015] Secondly, the present invention also provides a photovoltaic equipment cluster benchmarking analysis and latent fault early warning system, comprising:

[0016] The data acquisition module is used to collect basic parameters and implicit correlation data of photovoltaic equipment;

[0017] The data fusion module is used to process and fuse the collected data to generate standardized data sequences.

[0018] The cluster benchmarking module is used to divide photovoltaic equipment into different clusters, build a group health benchmark, and improve the model accuracy through individual difference calibration.

[0019] The dynamic early warning module uses a deep learning model to predict the failure risk of photovoltaic equipment and generate early warning information.

[0020] The intelligent operation and maintenance module prioritizes early warnings based on early warning information and economic loss assessment, generates operation and maintenance plans, pushes them to operation and maintenance personnel, and provides optimization guidance.

[0021] The modules work together to achieve accurate benchmarking analysis and early warning of hidden faults in photovoltaic equipment clusters. Through intelligent data collection, analysis and prediction, the operation and maintenance management of photovoltaic power plants is optimized, and the failure rate and downtime are reduced.

[0022] The beneficial effects of this invention are as follows:

[0023] This invention uses sensors to comprehensively collect data from photovoltaic equipment, ensuring the comprehensiveness and timeliness of the data. It supplements the shortcomings of traditional data collection and provides a solid foundation for subsequent analysis and fault early warning, ensuring real-time monitoring and accurate analysis of photovoltaic equipment.

[0024] This invention effectively overcomes the limitations of traditional benchmarking analysis by employing cluster benchmarking analysis and individual difference calibration techniques. Through cluster analysis based on photovoltaic (PV) equipment model, batch, and installation region, PV equipment can be rationally assigned to corresponding clusters, establishing a group health benchmark. Addressing individual differences among different PV devices, individual difference calibration eliminates misjudgments and false alarms caused by these differences, improving the accuracy of PV equipment health status assessment. Especially in large-scale PV power plants, this benchmarking analysis helps identify and analyze potential problems, reduces misjudgments due to individual differences, and improves the accuracy of analysis results.

[0025] This invention combines the operational data of photovoltaic (PV) equipment with cluster benchmarks using an LSTM-XGBoost fusion model to accurately predict future failure risks. Through multi-dimensional data fusion and comparison with historical data, this invention can identify potential faults in advance and generate latent fault warnings, significantly improving the accuracy and timeliness of fault warnings. The warnings do not rely solely on a single data source but comprehensively consider various aspects of PV equipment information, ensuring the warning system can adapt to changes in different operating conditions and lifecycle stages. This predictive capability helps maintenance personnel identify PV equipment problems early, avoid sudden failures, and improve the stability and economy of the power plant.

[0026] This invention provides priority ranking and decision support for operation and maintenance personnel by comprehensively analyzing fault early warning and revenue loss assessment. Based on the potential economic losses and urgency of photovoltaic equipment failures, operation and maintenance personnel can allocate resources more rationally, prioritizing high-risk, high-loss photovoltaic equipment failures. This improves resource allocation efficiency, reduces power plant downtime and revenue losses caused by untimely handling of photovoltaic equipment failures, and optimizes the overall operation and maintenance management of photovoltaic power plants. Through intelligent decision-making, the operational efficiency of photovoltaic power plants is significantly improved, and operation and maintenance costs are effectively controlled.

[0027] This invention improves the monitoring accuracy and fault early warning capability of photovoltaic equipment clusters, ensuring that the operating status of photovoltaic equipment can be accurately monitored, and identifies potential fault risks through real-time data analysis, thereby reducing the failure rate of photovoltaic equipment and extending the service life of photovoltaic equipment.

[0028] Therefore, it is evident that the present invention has outstanding substantive features and significant progress compared with the prior art, and the beneficial effects of its implementation are also obvious. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0030] Figure 1 is a flowchart of a method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters;

[0031] Figure 2 is a schematic diagram of a photovoltaic equipment cluster benchmarking analysis and latent fault early warning system. Detailed Implementation

[0032] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The following embodiments are explanations of the present invention, but the present invention is not limited to the following implementation methods.

[0033] Example 1:

[0034] As shown in Figure 1, this embodiment provides a method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters, including the following steps:

[0035] Step S1: Data acquisition step, which involves collecting and uploading the basic parameter data and implicit correlation data of the photovoltaic equipment in real time through sensors;

[0036] Step S2: The data fusion step involves preprocessing data from different sources using interpolation and outlier removal methods, and unifying them into a standard format.

[0037] Step S3: The cluster benchmarking analysis step involves using clustering algorithms to divide similar photovoltaic devices into clusters based on the operating parameters and historical data of the photovoltaic equipment, constructing a group health benchmark, and calibrating according to the differences in photovoltaic equipment.

[0038] Step S4: The dynamic early warning step uses the LSTM-XGBoost fusion model, combined with the operating data of photovoltaic equipment and cluster benchmarks, to predict the future failure risk of photovoltaic equipment and generate latent fault early warning.

[0039] Step S5: The steps of intelligent operation and maintenance are as follows: based on the fault warning and the assessment of revenue loss, the warning information is sorted according to priority and an operation and maintenance decision plan is generated and pushed to the operation and maintenance personnel.

[0040] In step S1:

[0041] In a photovoltaic (PV) equipment cluster, multi-dimensional data from the PV devices is collected in real time to provide comprehensive data support for subsequent analysis and early warning. This process includes basic parameter collection, implicit correlation data collection, and data transmission.

[0042] The data collected from photovoltaic (PV) equipment includes basic parameter data and implicit correlation data. Basic parameter data includes electrical information data, environmental data, factory parameters, and historical operation and maintenance data. Implicit correlation data includes aging data and localized shading data. Within the basic parameter data, the electrical information data includes the output voltage, output current, and power factor of the PV string; the input power, output power, conversion efficiency, module temperature, and PV string backplane temperature of the inverter; and the DC bus voltage, branch current, and branch temperature of the combiner box. The environmental data includes irradiance, ambient temperature, wind speed, and precipitation.

[0043] Step S11, the steps for collecting basic parameters:

[0044] The step of collecting basic parameter data, which includes electrical information data and environmental data of the photovoltaic equipment, as well as the photovoltaic equipment's factory parameters and historical operation and maintenance data, is as follows:

[0045] By setting string-level current and voltage sensors at both ends of the photovoltaic string, installing an energy acquisition module inside the inverter, and setting up an environmental monitoring station at the power station site, the electrical information data and environmental data of the photovoltaic equipment can be collected in real time.

[0046] The electrical information data of photovoltaic (PV) equipment includes the output voltage, output current, and power factor of the PV string; the input power, output power, conversion efficiency, module temperature, and PV string backplane temperature of the inverter; and the DC bus voltage, branch current, and branch temperature of the combiner box. Environmental data for PV equipment includes irradiance, ambient temperature, wind speed, and precipitation. Historical operation and maintenance records and factory parameter data of the PV equipment also need to be collected to provide a complete data foundation for subsequent analysis. A PV string refers to a group of PV modules connected in series, which work together to generate electricity.

[0047] Specifically, the sampling frequency for photovoltaic string voltage and inverter power is 10 seconds per sampling; the temperature of the combiner box branch and the backplane temperature of the photovoltaic string are collected using surface-mount temperature sensors. These sensors transmit the collected photovoltaic equipment operating parameters and environmental parameters in real time via communication protocols such as RS485, LoRa, and 4G / 5G, with a sampling frequency of 30 seconds per sampling. This data acquisition provides the real-time electrical operating status and environmental parameters of the photovoltaic equipment, offering fundamental data for subsequent analysis.

[0048] Among them, RS485 (Recommended Standard 485) is a commonly used serial communication protocol used for data transmission between photovoltaic devices such as photovoltaic string sensors and inverter acquisition modules and the data transmission layer; LoRa (Long Range Radio) is a low-power wide-area network communication technology, usually used as one of the communication methods for transmitting photovoltaic device operating parameters and environmental parameters from the data acquisition layer to the data transmission layer of a photovoltaic power station; the 4th Generation Mobile Communication Technology (4G) is a mobile communication standard used for photovoltaic power station local gateways to upload collected data to cloud data centers, ensuring the real-time performance and stability of data transmission; the 5th Generation Mobile Communication Technology (5G) is a new generation of mobile communication technology, which, together with 4G, serves as a network transmission method for the data transmission layer, offering higher bandwidth and lower latency compared to 4G, and adapting to the needs of high-frequency data transmission.

[0049] Step S12, the steps for collecting implicit association data:

[0050] The implicit correlation data of photovoltaic (PV) equipment includes the collection of PV equipment aging data and partial shading data. Capacitor sensors are installed inside the inverter to collect the equivalent series resistance value of the capacitors. This equivalent series resistance is then tested using the IV curve test according to the IEC 61215 standard to obtain the PV equipment degradation rate data, which is the PV equipment aging data. Partial shading data is collected by capturing images from high-definition cameras installed in the power station. Image recognition algorithms identify the shading area and degree of shading. Simultaneously, shading data is collected using dust thickness sensors, with sampling frequencies of every 15 minutes and every hour, respectively. This data can promptly detect implicit faults such as PV equipment aging and partial shading, providing strong support for accurate analysis and fault prediction.

[0051] Among them, IEC 61215 (International Electrotechnical Commission 61215) is an international standard for photovoltaic module performance testing. The document is used to guide the monitoring of photovoltaic module degradation rate. The module degradation rate data is obtained through the IV curve test under this standard. The IV curve refers to the current-voltage characteristic curve of the photovoltaic module. Through the IV curve test, key parameters such as the module's output power, open-circuit voltage, and short-circuit current can be obtained, which are used to calculate the module degradation rate.

[0052] Step S13, data transmission steps:

[0053] All collected photovoltaic equipment operating parameters and environmental parameters are transmitted in real time to the data transmission layer via communication protocols such as RS485, LoRa, and 4G / 5G to ensure timely data upload and subsequent processing. The sampling frequency is typically 1 to 5 minutes per sampling to ensure the real-time performance and accuracy of data transmission.

[0054] The data transmission layer consists of a local gateway at the power station and a cloud platform. The local gateway is responsible for aggregating the data collected by various sensors and photovoltaic devices, performing simple format conversion, and then uploading it to the cloud data center via Ethernet or 4G / 5G network. It also has a local data caching function to avoid data loss due to network interruption.

[0055] By implementing this step, high-frequency, accurate, and multi-dimensional data collection can be achieved, comprehensively monitoring the real-time status and potential problems of photovoltaic equipment, and providing a solid data foundation for fault early warning, photovoltaic equipment health assessment, and subsequent decision-making.

[0056] In step S2: After data acquisition, the acquired data is cleaned and fused to address noise, missing data, and data from different sources and frequencies, in order to generate a standardized data sequence. This step includes data cleaning and data fusion.

[0057] Step S21, data cleaning steps:

[0058] To address potential outliers and missing values ​​in the collected data, an improved preprocessing algorithm was employed for cleaning. This improved algorithm includes: using outlier removal methods, and combining the 3σ criterion with the physical characteristics of photovoltaic equipment to construct a dual-verification mechanism of rules and statistics, avoiding the erroneous removal of valid data. The physical characteristics of photovoltaic equipment include, for example, that voltage cannot be negative and that power is positively correlated with sunlight. For missing values, different frequency data interpolation methods were used: linear interpolation was used for high-frequency data, while interpolation based on the similarity of adjacent time periods was used for low-frequency data. These processing methods improved data integrity and ensured the accuracy of subsequent analysis.

[0059] Step S22, the data fusion steps:

[0060] A weighted fusion algorithm is employed to transform data from different sources and frequencies into a standardized 1-minute / instance data sequence. For example, capacitor lifespan data is mapped to minute-level data through time interpolation, while occlusion severity data is correlated with contemporaneous electrical information data, assigning weights to the occlusion impact of the electrical information data. This achieves spatiotemporal alignment and semantic fusion of data from different dimensions. This process integrates different types of data into a unified standard format, facilitating subsequent analysis and modeling.

[0061] By implementing this step, data from different sources and frequencies can be effectively cleaned and integrated to generate a unified and standardized data sequence, providing accurate input data for subsequent photovoltaic equipment health assessments, fault prediction, and early warning models.

[0062] In step S3:

[0063] Based on data collection and cleaning, the cluster benchmarking layer provides crucial support for the health assessment and fault early warning of photovoltaic equipment through the clustering of photovoltaic devices, the construction of a group health benchmark, and the calibration of individual differences. This step includes the photovoltaic device clustering step, the group health benchmark construction step, and the individual difference calibration step.

[0064] Step S31, the steps for clustering photovoltaic devices:

[0065] Based on the K-means clustering algorithm, photovoltaic (PV) equipment is clustered according to its model, batch, and installation region. Specifically, taking inverters as an example, clustering features include the rated power of the PV equipment, its manufacturing date, and the average solar irradiance at the installation location. By calculating similarity, the clustering algorithm groups PV equipment with a similarity ≥ 90% into the same cluster, ensuring that the PV equipment within a cluster has similar basic characteristics. This grouping of PV equipment ensures that subsequent performance analysis and benchmark evaluation of each PV equipment group can be targeted.

[0066] Step S32, the steps for constructing a population health benchmark:

[0067] For each photovoltaic (PV) equipment cluster, PV devices that have maintained stable performance and been fault-free for the past three months are selected as health samples. A health sample can be further described as having a fluctuation of ≤5% within those three months. By statistically analyzing the multi-dimensional parameters of these health samples, standard values ​​for each parameter are calculated, and a cluster health benchmark model is constructed based on these values. These health samples represent the normal operating status of the PV equipment cluster and serve as a reference benchmark, aiding in subsequent health assessments and fault predictions for other PV devices.

[0068] Step S33, Individual Difference Calibration Steps:

[0069] To address individual differences in photovoltaic (PV) equipment during actual operation (such as installation angle and module compatibility), individual calibration coefficients are constructed. Taking PV strings as an example, the deviation rate between the operating parameters of the PV equipment during its historical healthy period and the group health benchmark is used as input, and the coefficients are calculated using the formula... The individual calibration coefficient is calculated, and then the real-time parameters of the photovoltaic equipment to be evaluated are multiplied by the individual calibration coefficient and compared with the group health benchmark. For example, if the power of a string of photovoltaic equipment is consistently 3% lower than the group health benchmark during the healthy period, the calibration coefficient is 0.97. By multiplying the real-time parameters of the photovoltaic equipment to be evaluated by the individual calibration coefficient and then comparing them with the group health benchmark during real-time monitoring, misjudgments caused by individual differences can be effectively eliminated, ensuring a more accurate health assessment of the photovoltaic equipment.

[0070] By implementing this step, accurate health assessments of photovoltaic equipment can be conducted based on cluster characteristics and individual differences, a more scientific benchmark model can be built, and more accurate basis can be provided for fault early warning and performance optimization.

[0071] In step S4:

[0072] After completing data collection and cluster benchmarking, the dynamic early warning layer performs multi-dimensional fault risk prediction and early warning for photovoltaic equipment. This layer ensures comprehensive monitoring of the health status of photovoltaic equipment and improves the accuracy of early warnings through three collaborative steps: operating condition adaptation, life cycle stage identification, and fault risk prediction. These steps include operating condition adaptation, life cycle stage identification, and fault risk prediction.

[0073] Step S41, the steps for adapting to operating conditions:

[0074] A working condition classification model based on the random forest algorithm categorizes the operating environment of photovoltaic (PV) equipment into nine combinations based on light intensity and ambient temperature. Specifically, light intensity is divided into: ≤200W / ㎡ (weak light), 200-800W / ㎡ (medium light), and ≥800W / ㎡ (strong light); ambient temperature is divided into: ≤15℃, 15-35℃, and ≥35℃. The PV equipment is then classified according to these two parameters. For each working condition, a corresponding warning threshold is trained based on PV equipment data. For example, under weak light conditions, the inverter's power warning threshold is 70% of the overall health baseline; under strong light conditions, the warning threshold is 85%. This adaptive working condition warning system dynamically adjusts the warning strategy, improving warning accuracy.

[0075] Step S42, the steps for lifecycle stage identification:

[0076] Using a gradient boosting tree model, the system automatically identifies the current lifecycle stage of photovoltaic (PV) equipment based on historical O&M data, electrical information data, and aging data. Based on these data, the PV equipment lifecycle is divided into three stages: break-in period (0-1 year), stable period (1-8 years), and aging period (≥8 years). The performance characteristics of the PV equipment differ at each stage, therefore different early warning strategies are applied based on the stage identification results. This process ensures more accurate health assessments and fault warnings for PV equipment at different lifecycle stages.

[0077] Step S43, Fault Risk Prediction Steps:

[0078] An LSTM-XGBoost fusion model is used, combining multi-dimensional input data and cluster benchmarking results to predict the future failure probability of photovoltaic (PV) equipment. When the risk probability reaches or exceeds 60%, it is identified as a latent fault warning, and the specific fault type is output. The judgment is based on key parameters including PV string backsheet temperature, power output, and inverter capacitor equivalent series resistance. For example, if the module backsheet temperature is 5°C higher than the group health benchmark and the power output is 10% lower than the calibrated benchmark, it is identified as a hot spot risk; if the inverter capacitor's equivalent series resistance is 20% higher than the group health benchmark, it is identified as an IGBT performance degradation risk.

[0079] By implementing this step, photovoltaic equipment can be effectively monitored under various environmental conditions, ensuring that potential fault risks can be accurately identified at every stage of photovoltaic equipment operation, thereby triggering early warnings in a timely manner, reducing the risk of photovoltaic equipment downtime, and extending the service life of photovoltaic equipment.

[0080] In step S5:

[0081] In the intelligent operation and maintenance phase of photovoltaic (PV) equipment clusters, a complete operation and maintenance handling mechanism is formed by combining revenue loss quantification, early warning prioritization, and operation and maintenance decision output. This ensures timely and accurate response to PV equipment fault warnings and the formulation of reasonable operation and maintenance plans. This step includes revenue loss quantification, early warning prioritization, and operation and maintenance decision output.

[0082] Step S51, the step of quantifying gains and losses:

[0083] An anomaly-revenue loss correlation formula is established to calculate the revenue loss value, quantifying the potential economic loss caused by photovoltaic equipment failure. Specifically, the formula for the revenue loss value 's' is: Where p is the rated power of the equipment, t is the duration of the abnormality, and m is the electricity price. This represents the power attenuation rate under abnormal conditions.

[0084] For example, if an inverter with a rated power of 500kW experiences a power degradation rate of 30% under abnormal conditions, and this abnormal condition lasts for 24 hours, with an electricity price of 0.4 yuan / kWh, then the revenue loss is... This step allows for the quantification of economic losses caused by abnormal conditions of each photovoltaic device, providing an important reference for operation and maintenance decisions.

[0085] Step S52, the steps for prioritizing early warnings:

[0086] Based on revenue loss and fault urgency, an analytic hierarchy process (AHP) is used to comprehensively rank the various early warning messages. Revenue loss is weighted at 60%, and fault urgency at 40%. The warning level is determined by calculating a comprehensive score. A score ≥ 80 is classified as a Level 1 warning; 60 ≤ score < 80 is classified as a Level 2 warning; and a score < 60 is classified as a Level 3 warning. This step ensures the prioritization of fault warnings, facilitating resource and maintenance personnel to prioritize the handling of high-risk photovoltaic equipment.

[0087] Step S53, the steps for outputting operation and maintenance decisions:

[0088] Based on the warning type and the actual location of the photovoltaic equipment, a specific operation and maintenance plan is automatically generated. Specifically, when a photovoltaic device experiences a "Level 1 Hot Spot Warning for Modules," the system outputs, "It is recommended that maintenance personnel be dispatched to Area B within 24 hours, equipped with a thermal imager, to investigate and replace the abnormal module." Simultaneously, the maintenance decision is displayed on web and mobile platforms, including a list of warning priorities and a maintenance route plan, and is pushed to the maintenance personnel's terminal photovoltaic equipment. This step ensures the timeliness and accuracy of maintenance decisions, enabling maintenance personnel to respond quickly to fault warnings and reduce losses to photovoltaic equipment.

[0089] By implementing the above steps, we can achieve collaborative operation based on revenue loss quantification, early warning priority ranking, and intelligent operation and maintenance decision-making, further improving the efficiency and accuracy of photovoltaic equipment fault early warning and operation and maintenance response, and effectively avoiding the impact of photovoltaic equipment faults on photovoltaic power generation revenue.

[0090] Furthermore, this embodiment can add an edge node computation step, using the Tiny-CNN lightweight machine learning model to replace the original random forest operating condition classification model. The model parameters are compressed to below 5MB, ensuring efficient operation of the edge nodes. Simultaneously, a data synchronization mechanism between the edge and the cloud is established, uploading local processing results to the cloud every hour to update the cluster baseline model and perform in-depth fault risk prediction, ensuring data consistency. Some lightweight processing of the data fusion layer and the operating condition classification model of the dynamic early warning layer are migrated to the edge nodes, realizing a collaborative mode of local real-time preprocessing and cloud-based in-depth analysis. For example, the edge node can complete a preliminary judgment of anomalies in the electrical information data of a single inverter within 1 second. If a significant anomaly is detected, a local alarm is triggered directly without waiting for cloud analysis. The early warning response time is shortened from 1 minute in Embodiment 1 to within 10 seconds, suitable for emergency handling of sudden faults in photovoltaic power plants.

[0091] Furthermore, this embodiment can add a fault self-healing linkage in step S5. By establishing a self-healing feasibility assessment model, inputting features such as fault type, photovoltaic equipment status, and remote control permissions, and using a logistic regression algorithm to determine the self-healing success rate, if it is ≥80%, self-healing is executed. At the same time, a self-healing monitoring threshold is set. If the fault risk does not decrease within 30 minutes after self-healing, it is immediately upgraded to a manual operation and maintenance warning to prevent the fault from escalating due to self-healing failure. For some hidden faults that can be handled remotely, the remote control module of the power station is linked to realize fault self-healing. For example, when the dynamic warning module determines that an inverter is overheating due to insufficient fan speed, the fault self-healing linkage module can remotely send a command to increase the fan speed from 2000rpm to 2500rpm, monitor the temperature change in real time, and if the temperature drops back to the normal range, the warning is lifted without manual intervention.

[0092] Example 2:

[0093] As shown in Figure 1, taking a 100MW large-scale photovoltaic power station as an example, this photovoltaic power station includes 200 inverters and 10,000 photovoltaic strings. This embodiment provides a method for photovoltaic equipment cluster benchmarking analysis and latent fault early warning, including the following steps:

[0094] Step S1: Data acquisition step, which involves collecting and uploading the basic parameter data and implicit correlation data of the photovoltaic equipment in real time through sensors;

[0095] Step S2: The data fusion step involves preprocessing data from different sources using interpolation and outlier removal methods, and unifying them into a standard format.

[0096] Step S3: The cluster benchmarking analysis step involves using clustering algorithms to divide similar photovoltaic devices into clusters based on the operating parameters and historical data of the photovoltaic equipment, constructing a group health benchmark, and calibrating according to the differences in photovoltaic equipment.

[0097] Step S4: The dynamic early warning step uses the LSTM-XGBoost fusion model, combined with the operating data of photovoltaic equipment and cluster benchmarks, to predict the future failure risk of photovoltaic equipment and generate latent fault early warning.

[0098] Step S5: The steps of intelligent operation and maintenance are as follows: based on the fault warning and the assessment of revenue loss, the warning information is sorted according to priority and an operation and maintenance decision plan is generated and pushed to the operation and maintenance personnel.

[0099] In step S1:

[0100] In a photovoltaic (PV) equipment cluster, multi-dimensional data from the PV devices is collected in real time to provide comprehensive data support for subsequent analysis and early warning. This process includes basic parameter collection, implicit correlation data collection, and data transmission.

[0101] The data collected from photovoltaic (PV) equipment includes basic parameter data and implicit correlation data. Basic parameter data includes electrical information data, environmental data, factory parameters, and historical operation and maintenance data. Implicit correlation data includes aging data and localized shading data. Within the basic parameter data, the electrical information data includes the output voltage, output current, and power factor of the PV string; the input power, output power, conversion efficiency, module temperature, and PV string backplane temperature of the inverter; and the DC bus voltage, branch current, and branch temperature of the combiner box. The environmental data includes irradiance, ambient temperature, wind speed, and precipitation.

[0102] Step S11, the steps for collecting basic parameters:

[0103] The step of collecting basic parameter data, which includes electrical information data and environmental data of the photovoltaic equipment, as well as the photovoltaic equipment's factory parameters and historical operation and maintenance data, is as follows:

[0104] By setting string-level current and voltage sensors at both ends of the photovoltaic string, installing an energy acquisition module inside the inverter, and setting up an environmental monitoring station at the power station site, the electrical information data and environmental data of the photovoltaic equipment can be collected in real time.

[0105] The electrical information data of photovoltaic (PV) equipment includes the output voltage, output current, and power factor of the PV string; the input power, output power, conversion efficiency, module temperature, and PV string backplane temperature of the inverter; and the DC bus voltage, branch current, and branch temperature of the combiner box. Environmental data for PV equipment includes irradiance, ambient temperature, wind speed, and precipitation. Historical operation and maintenance records and factory parameter data of the PV equipment also need to be collected to provide a complete data foundation for subsequent analysis. A PV string refers to a group of PV modules connected in series, which work together to generate electricity.

[0106] Step S12, the steps for collecting implicit association data:

[0107] The implicit correlation data of photovoltaic (PV) equipment includes the collection of PV equipment aging data and partial shading data. Capacitor sensors are installed inside the inverter to collect the equivalent series resistance value of the capacitors. This equivalent series resistance is then tested using the IV curve test according to the IEC 61215 standard to obtain the PV equipment degradation rate data, which is the PV equipment aging data. Partial shading data is collected by capturing images from high-definition cameras installed in the power station. Image recognition algorithms identify the shading area and degree of shading. Simultaneously, shading data is collected using dust thickness sensors, with sampling frequencies of every 15 minutes and every hour, respectively. This data can promptly detect implicit faults such as PV equipment aging and partial shading, providing strong support for accurate analysis and fault prediction.

[0108] Step S13, data transmission steps:

[0109] All collected photovoltaic equipment operating parameters and environmental parameters are transmitted in real time to the data transmission layer via LoRa and 5G networks to ensure that the data can be uploaded in a timely manner and processed subsequently.

[0110] In step S2: After data acquisition, the acquired data is cleaned and fused to address noise, missing data, and data from different sources and frequencies, in order to generate a standardized data sequence. This step includes data cleaning and data fusion.

[0111] Step S21, data cleaning steps:

[0112] To address potential outliers and missing values ​​in the collected data, an improved preprocessing algorithm was employed for cleaning. This improved algorithm includes: using an outlier removal method, and combining the 3σ criterion with the physical characteristics of photovoltaic equipment to construct a dual-verification mechanism of rules and statistics, thus avoiding the erroneous removal of valid data. The physical characteristics of photovoltaic equipment include, for example, that voltage cannot be negative and that power is positively correlated with sunlight. For missing values, different frequency data interpolation methods were used: linear interpolation was used for high-frequency data, while interpolation based on the similarity of adjacent time periods was used for low-frequency data.

[0113] S22, Data Fusion

[0114] A weighted fusion algorithm is used to convert data from different sources and frequencies into a standardized 1-minute / time data sequence. For example, capacitor life data is mapped to minute-level data through time interpolation, and occlusion level data is correlated with concurrent electrical information data, assigning weights to the occlusion impact of the electrical information data.

[0115] In step S3: Based on the data collection and cleaning, the cluster benchmarking layer divides the 200 inverters into 4 clusters according to their model and installation area, builds a group health benchmark for each cluster, and performs individual calibration on an inverter located in the western area (cluster 2) to obtain a calibrated power benchmark of 480kW; this provides important support for the health assessment and fault early warning of photovoltaic equipment.

[0116] The specific steps are as follows:

[0117] Step S31, the steps for clustering photovoltaic devices:

[0118] Based on the K-means clustering algorithm, 200 inverters were clustered into 4 clusters according to the model, batch and installation area of ​​the photovoltaic equipment, ensuring that the photovoltaic equipment in the clusters have similar basic characteristics.

[0119] Step S32, the steps for constructing a population health benchmark:

[0120] Photovoltaic equipment with stable performance and no faults within the past three months was selected as a healthy sample. A healthy sample can be further described as having a fluctuation of ≤5% within three months. By statistically analyzing the multi-dimensional parameters of these healthy samples, the standard values ​​of each parameter were calculated, and a population health benchmark model was constructed.

[0121] Step S33, Individual Difference Calibration Steps:

[0122] Individual calibration coefficients are calculated by comparing the deviations of photovoltaic equipment's operating parameters during its historical healthy period with the population health benchmark. The formula for the individual calibration coefficient is: By multiplying the real-time parameters of the photovoltaic equipment to be evaluated by an individual calibration coefficient and then comparing them with a group health benchmark, misjudgments caused by individual differences are eliminated, and a calibrated power benchmark of 480kW is finally obtained.

[0123] In step S4: After completing data acquisition and cluster benchmarking, the dynamic early warning layer performs multi-dimensional fault risk prediction and early warning for photovoltaic equipment. This step includes the steps of operating condition adaptation, life cycle stage identification, and fault risk prediction.

[0124] Step S41, the steps for adapting to operating conditions:

[0125] A working condition classification model based on the random forest algorithm categorizes the working environment of photovoltaic equipment into nine combinations of working conditions according to light intensity and ambient temperature, and trains corresponding warning thresholds. In this embodiment, the current working condition is identified as strong light (1000W / ㎡) and ambient temperature (38℃) (high temperature condition); under the strong light condition, the warning threshold is 85%. This adaptive working condition warning can dynamically adjust the warning strategy and improve the accuracy of the warning.

[0126] Step S42, the steps for lifecycle stage identification:

[0127] The system automatically identifies the current lifecycle stage of photovoltaic (PV) equipment using a gradient boosting tree model. Based on historical operation and maintenance data and aging data of the PV equipment, the lifecycle is divided into a break-in period, a stable period, and an aging period, and different early warning strategies are applied according to the stage identification results.

[0128] Step S43, Fault Risk Prediction Steps:

[0129] Using the LSTM-XGBoost fusion model, the inverter is predicted to have a 75% probability of failure in the next 3 days, which is identified as an early warning of IGBT performance degradation in the inverter.

[0130] In step S5:

[0131] In the intelligent operation and maintenance phase of photovoltaic (PV) equipment clusters, a complete operation and maintenance handling mechanism is formed by combining revenue loss quantification, early warning prioritization, and operation and maintenance decision output. This ensures timely and accurate response to PV equipment fault warnings and the formulation of reasonable operation and maintenance plans. This step includes revenue loss quantification, early warning prioritization, and operation and maintenance decision output.

[0132] Step S51, the step of quantifying gains and losses:

[0133] Based on the types and probabilities of photovoltaic equipment failures, the potential losses caused by untimely handling of failures are quantified. The quantification shows that if a failure is not addressed, a loss of 3240 yuan could occur within 3 days.

[0134] Step S52, the steps for prioritizing early warnings:

[0135] All fault warnings are prioritized using a warning priority sorting unit, and high-risk fault warnings are marked as Level 1 warnings.

[0136] Step S53, the steps for outputting operation and maintenance decisions:

[0137] Based on fault warnings and revenue loss assessments, corresponding operation and maintenance plans are generated and pushed to operation and maintenance personnel. These personnel complete repairs within 12 hours, thus preventing further power generation losses.

[0138] Example 3:

[0139] As shown in Figure 2, this embodiment provides a system for photovoltaic equipment cluster benchmarking analysis and latent fault early warning based on multi-dimensional data. The system includes multiple functional modules, as follows:

[0140] The data acquisition module 1 includes a basic parameter acquisition unit 11, an implicit correlation data acquisition unit 12, and a data transmission unit 13. The basic parameter acquisition unit 11 is responsible for real-time acquisition of electrical information data, environmental data, photovoltaic equipment factory parameters, and historical operation and maintenance data of the photovoltaic equipment, ensuring comprehensive real-time status monitoring through high-frequency data acquisition. The implicit correlation data acquisition unit 12 monitors photovoltaic equipment aging data and partial shading data to capture potential fault risks and provide a basis for fault prediction. The data transmission unit 13 is responsible for transmitting the collected data to the data processing layer in real-time through different communication protocols, ensuring the real-time performance and accuracy of data transmission.

[0141] This module provides real-time and accurate monitoring of photovoltaic equipment operation status by efficiently collecting multi-dimensional data, providing complete data support for subsequent fault early warning and health assessment.

[0142] Data fusion module 2 processes and fuses the collected data, ensuring data integrity and consistency, and generating standardized data sequences. This module includes a data cleaning unit 21, a data fusion unit 22, and a data standardization unit 23. The data cleaning unit 21 processes the collected raw data, removing noise and outliers to ensure the accuracy of subsequent analysis. The data fusion unit 22 generates a unified, standardized dataset by weighted fusion of data from different sources, providing reliable input for other modules in the system. The data standardization unit 23 is responsible for transforming multi-dimensional, multi-format data into a consistent standard format, ensuring data compatibility and consistency across different modules within the system.

[0143] The data fusion module ensures the accuracy and consistency of the system's input data, improves the precision and reliability of the system's analysis, and provides more accurate basic data for early warning and decision-making.

[0144] The cluster benchmarking module 3 is used to divide photovoltaic (PV) devices into different clusters, construct a group health benchmark, and improve model accuracy through individual difference calibration. This module includes a PV device clustering unit 31, a group health benchmark generation unit 32, and an individual difference calibration unit 33. The PV device clustering unit 31 groups PV devices with similar characteristics into the same cluster. The group health benchmark generation unit 32 generates a health benchmark for each cluster based on the behavioral patterns of the PV device group, serving as a standard for assessing the health status of the PV devices. The individual difference calibration unit 33 corrects for individual differences in the PV devices, ensuring more accurate health assessments and improving the accuracy of fault prediction.

[0145] The cluster benchmarking module can group photovoltaic devices according to their characteristics and environmental conditions, thereby providing a more realistic health benchmark assessment and effectively improving the accuracy and adaptability of the fault prediction model.

[0146] The dynamic early warning module 4 uses a deep learning model to predict the failure risk of photovoltaic (PV) equipment and generate early warning information, providing a basis for subsequent fault investigation. This module includes a working condition adaptation unit 41, a life cycle stage identification unit 42, and a fault risk prediction unit 43. The working condition adaptation unit 41 uses a working condition classification model built based on the random forest algorithm, combining light intensity and ambient temperature to divide the working environment of PV equipment into nine working condition combinations, and dynamically adjusts the early warning threshold to improve early warning accuracy. The life cycle stage identification unit 42 uses a gradient boosting tree model to automatically identify the life cycle stage of PV equipment based on historical operation and maintenance data, electrical information data, and aging data, and calls different early warning strategies according to the stage results. The fault risk prediction unit 43 combines an LSTM-XGBoost fusion model to predict the future failure risk probability of PV equipment based on the fused multi-dimensional data and cluster benchmarking results. When the risk probability reaches or exceeds 60%, a latent fault warning is generated, and the specific fault type is output. Through the dynamic early warning module, photovoltaic equipment can accurately identify potential failure risks under different operating conditions and life cycle stages, trigger early warnings in a timely manner, reduce the risk of photovoltaic equipment downtime, and extend the service life of photovoltaic equipment.

[0147] The intelligent operation and maintenance module 5, based on early warning information and economic loss assessment, prioritizes early warnings, generates operation and maintenance plans, pushes them to operation and maintenance personnel, and provides optimization guidance. This module includes a revenue loss quantification unit 51, an early warning priority ranking unit 52, and an operation and maintenance decision output unit 53. The revenue loss quantification unit 51 quantifies the potential economic loss of photovoltaic equipment failures using revenue loss values, helping to assess the impact of photovoltaic equipment failures on power plant revenue. The early warning priority ranking unit 52, based on revenue loss values ​​and the urgency of the failure, uses the analytic hierarchy process (AHP) to rank early warning information, ensuring that high-risk photovoltaic equipment is addressed first. The operation and maintenance decision output unit 53 automatically generates operation and maintenance plans based on early warning priority and the location of photovoltaic equipment, pushes them to operation and maintenance personnel, and provides optimization guidance to ensure timely execution of operation and maintenance tasks.

[0148] The intelligent operation and maintenance module effectively improves the accuracy and timeliness of operation and maintenance decisions by accurately quantifying economic losses and prioritizing them, ensuring that operation and maintenance personnel can prioritize the most critical faults and reduce downtime and economic losses of photovoltaic equipment.

[0149] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The methods disclosed in the embodiments are described simply because they correspond to the systems disclosed in the embodiments; relevant details can be found in the method section.

[0150] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0151] In the embodiments provided by this invention, it should be understood that the disclosed systems, methods, and approaches can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0152] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0153] In addition, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit.

[0154] Similarly, in the various embodiments of the present invention, each processing unit can be integrated into a functional module, or each processing unit can exist physically, or two or more processing units can be integrated into a functional module.

[0155] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0156] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or photovoltaic device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or photovoltaic device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or photovoltaic device that includes said element.

[0157] The above-disclosed embodiments are merely preferred embodiments of the present invention, but the present invention is not limited thereto. Any non-creative variations that can be conceived by those skilled in the art, as well as any improvements and modifications made without departing from the principles of the present invention, should fall within the protection scope of the present invention.

Claims

1. A method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters, characterized in that, Includes the following steps: Step S1: Data acquisition step, real-time acquisition of basic parameter data and implicit correlation data of photovoltaic equipment through sensors and uploading the data; Step S2: Data fusion step, preprocessing data from different sources using interpolation and outlier removal methods, and unifying them into a standard format; Step S3: Cluster benchmarking analysis step, based on the operating parameters and historical data of photovoltaic equipment, using clustering algorithms to divide similar photovoltaic equipment into clusters, constructing a group health benchmark, and calibrating according to the differences of photovoltaic equipment; Step S4: Dynamic early warning step, using the LSTM-XGBoost fusion model, combining the operating data of photovoltaic equipment and the cluster benchmark, predicting the future failure risk of photovoltaic equipment and generating implicit fault early warnings; Step S5: The intelligent operation and maintenance steps involve prioritizing the early warning information based on fault warnings and revenue loss assessments, generating an operation and maintenance decision plan, and pushing it to the operation and maintenance personnel. Step S3 includes photovoltaic equipment clustering, group health benchmark construction, and individual difference calibration. Step S31, photovoltaic equipment clustering: Based on the K-means clustering algorithm, photovoltaic equipment is clustered according to its model, batch, and installation area. By calculating similarity, the clustering algorithm divides photovoltaic equipment with a similarity ≥ 90% into the same cluster. Step S32, group health benchmark construction: For each photovoltaic equipment cluster, the most recent 3... Photovoltaic equipment with stable performance and no faults within a month is used as a health sample. The multi-dimensional parameters of these health samples are statistically analyzed, the standard values ​​of each parameter are calculated, and a group health benchmark model is constructed based on this. Step S33, the individual difference calibration step: For the individual differences of photovoltaic equipment in actual operation, an individual calibration coefficient is constructed. The deviation rate between the operating parameters of photovoltaic equipment in the historical health period and the group health benchmark is used as input. The individual calibration coefficient is calculated by the formula: individual calibration coefficient = average historical parameters of equipment / average group health benchmark. During real-time monitoring, the real-time parameters of the photovoltaic equipment to be evaluated are multiplied by the individual calibration coefficient and compared with the group health benchmark.

2. The method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters according to claim 1, characterized in that, Step S1 includes steps for acquiring basic parameters, acquiring implicit correlation data, and transmitting data. Step S11 involves acquiring basic parameter data, which includes electrical information data, environmental data, factory parameters, and historical operation and maintenance data of the photovoltaic equipment. This step involves: setting string-level current and voltage sensors at both ends of the photovoltaic string, setting a power acquisition module inside the inverter, and setting an environmental monitoring station at the power station site to collect electrical information data and environmental data of the photovoltaic equipment in real time; and collecting historical operation and maintenance records and data of the photovoltaic equipment. Prepare factory parameters; among them, the sampling frequency for collecting photovoltaic string voltage and inverter power is 10 seconds / time; the temperature of combiner box branch and photovoltaic string backplane is collected using a patch-type temperature sensor, and the collected photovoltaic equipment operating parameters and environmental parameters are transmitted in real time through RS485, LoRa, and 4G / 5G communication protocols, with a sampling frequency of 30 seconds / time; step S12, the step of collecting implicit correlation data: collection of photovoltaic equipment aging data and partial shading data; the equivalent series resistance value of the capacitor in the inverter is collected by the capacitor sensor, and the equivalent series resistance value is converted into an IEC standard. The 61215 standard IV curve is used to test and obtain the degradation rate data of the photovoltaic equipment, which is the aging data of the photovoltaic equipment. The partial shading data is obtained by collecting images from high-definition cameras installed in the power station, and the shading area and degree of shading are identified by image recognition algorithm. At the same time, shading data is collected in combination with dust thickness sensor, with sampling frequencies of every 15 minutes / time and every hour / time, respectively. Step S13, data transmission steps: All the collected data of the photovoltaic equipment are transmitted to the data transmission layer in real time through RS485 or LoRa or 4G / 5G communication protocol, with a sampling frequency of usually 1 minute to 5 minutes / time.

3. A method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters according to claim 1 or 2, characterized in that, The collected data from the photovoltaic equipment includes basic parameter data and implicit correlation data. The basic parameter data includes electrical information data, environmental data, factory parameters, and historical operation and maintenance data of the photovoltaic equipment. The implicit correlation data includes aging data and partial shading data of the photovoltaic equipment. Among the basic parameter data, the electrical information data of the photovoltaic equipment includes the output voltage, output current, and power factor of the photovoltaic string; the input power, output power, conversion efficiency, module temperature, and backplane temperature of the photovoltaic string of the inverter; and the DC bus voltage, branch current, and branch temperature of the combiner box. The environmental data of the photovoltaic equipment includes irradiance, ambient temperature, wind speed, and precipitation.

4. The method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters according to claim 3, characterized in that, Step S2 includes a data cleaning step and a data fusion step; Step S21, Data Cleaning Step: For outliers and missing values ​​in the collected data, an improved preprocessing algorithm is used for cleaning. This improved preprocessing algorithm includes: using an outlier removal method, combining the 3σ criterion with the physical characteristics of photovoltaic equipment to construct a dual verification mechanism of rules and statistics to avoid erroneously removing valid data; for missing values, different frequency data interpolation methods are used: linear interpolation is used for high-frequency data, while interpolation based on the similarity of adjacent time periods is used for low-frequency data. Step S22, Data Fusion Step: A weighted fusion algorithm is used to uniformly convert data from different sources and frequencies into a standardized 1-minute / time data sequence.

5. The method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters according to claim 4, characterized in that, Step S4 includes the following steps: operating condition adaptation, life cycle stage identification, and fault risk prediction. Step S41, operating condition adaptation: A working condition classification model based on the random forest algorithm is used to divide the working environment of the photovoltaic equipment into nine combinations of operating conditions according to light intensity and ambient temperature. For each operating condition, a corresponding early warning threshold is trained based on the photovoltaic equipment data. Step S42, life cycle stage identification: Using a gradient boosting tree model, the current life cycle stage of the photovoltaic equipment is automatically identified based on historical operation and maintenance data, electrical information data, and aging data. The life cycle of the photovoltaic equipment is divided into three stages: break-in period, stable period, and aging period. The performance characteristics of the photovoltaic equipment will differ in each stage, therefore different early warning strategies are invoked based on the stage identification results. Step S43, fault risk prediction: An LSTM-XGBoost fusion model is used, combining the multi-dimensional data after fusion with cluster benchmarking results, to predict the future fault risk probability of the photovoltaic equipment. When the risk probability reaches or exceeds 60%, it is determined as a latent fault warning, and the specific fault type is output.

6. The method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters according to claim 5, characterized in that, The light intensity is divided into: weak light ≤200W / ㎡, medium light 200-800W / ㎡, and strong light ≥800W / ㎡; the ambient temperature is divided into ≤15℃, 15-35℃, and ≥35℃.

7. The method for benchmarking analysis and latent fault early warning of photovoltaic equipment clusters according to claim 6, characterized in that, Step S5 includes steps for quantifying revenue loss, prioritizing early warnings, and outputting operation and maintenance decisions; Step S51, the step for quantifying revenue loss: establishing an anomaly-revenue loss correlation formula to calculate the revenue loss value, quantifying the potential economic loss caused by photovoltaic equipment failure, the formula for the revenue loss value s is: s=p×t×m×(1 α), where p is the rated power of the equipment, t is the duration of the abnormality, m is the electricity price, and α is the power attenuation rate under the abnormal state; Step S52, the steps for prioritizing early warnings: Based on the revenue loss value and the urgency of the fault, the analytic hierarchy process is used to comprehensively sort the various early warning information; the weight of the revenue loss value is set at 60%, and the weight of the urgency of the fault is set at 40%, and the early warning level is determined by calculating the comprehensive score; When the comprehensive score is ≥80, it is judged as a Level 1 warning; when the score is 60≤score<80, it is judged as a Level 2 warning; when the score is <60, it is judged as a Level 3 warning; Step S53, the steps of operation and maintenance decision output: based on the warning type and the actual location of the photovoltaic equipment, a specific operation and maintenance plan is automatically generated; at the same time, the operation and maintenance decision is displayed through the Web terminal and mobile terminal platform, including the warning priority list and operation and maintenance route planning, and is simultaneously pushed to the terminal photovoltaic equipment of the operation and maintenance personnel.

8. A photovoltaic equipment cluster benchmarking analysis and latent fault early warning system, characterized in that, include: The system comprises a data acquisition module (1), a data fusion module (2), a cluster benchmarking module (3), a dynamic early warning module (4), and an intelligent operation and maintenance module (5). The data acquisition module (1) is used to collect basic parameters and implicit correlation data of photovoltaic equipment. The data fusion module (2) is used to process and fuse the collected data to generate standardized data sequences. The cluster benchmarking module (3) is used to divide photovoltaic equipment into different clusters, construct a group health benchmark, and improve the model accuracy through individual difference calibration. The dynamic early warning module (4) predicts the failure risk of photovoltaic equipment through a deep learning model and generates early warning information. The intelligent operation and maintenance module (5) prioritizes early warnings and generates operation and maintenance plans based on early warning information and economic loss assessment, pushes them to operation and maintenance personnel, and provides optimization guidance. The cluster benchmarking module (3) includes a photovoltaic equipment clustering unit (31), a group health benchmark generation unit (32), and an individual difference calibration unit. Unit (33); Photovoltaic equipment clustering unit (31) divides photovoltaic equipment with similar characteristics into the same cluster; Group health benchmark generation unit (32) generates health benchmarks for each cluster based on the behavior patterns of the photovoltaic equipment group, which serve as the standard for assessing the health status of photovoltaic equipment; Individual difference calibration unit (33) corrects the individual differences of photovoltaic equipment; In photovoltaic equipment clustering unit (31): Based on the K-means clustering algorithm, photovoltaic equipment is clustered according to the model, batch and installation area of ​​photovoltaic equipment; By calculating similarity, the clustering algorithm divides photovoltaic equipment with similarity ≥90% into the same cluster; In group health benchmark generation unit (32): For each photovoltaic equipment cluster, photovoltaic equipment with stable performance and no faults in the last 3 months is selected as health samples, and the multi-dimensional parameters of these health samples are statistically analyzed to calculate the standard values ​​of each parameter, and a group health benchmark model is constructed accordingly; In the individual difference calibration unit (33): for individual differences in actual operation of photovoltaic equipment, an individual calibration coefficient is constructed. The deviation rate between the operating parameters of photovoltaic equipment in the historical health period and the group health benchmark is used as input. The individual calibration coefficient is calculated by the formula individual calibration coefficient = average historical parameters of equipment / average group health benchmark. During real-time monitoring, the real-time parameters of the photovoltaic equipment to be evaluated are multiplied by the individual calibration coefficient and compared with the group health benchmark.

9. A photovoltaic equipment cluster benchmarking analysis and latent fault early warning system according to claim 8, characterized in that, The data acquisition module (1) includes a basic parameter acquisition unit (11), an implicit correlation data acquisition unit (12), and a data transmission unit (13). The basic parameter acquisition unit (11) is responsible for collecting real-time electrical information data, environmental data, photovoltaic equipment factory parameters, and historical operation and maintenance data of the photovoltaic equipment, and conducting comprehensive monitoring through high-frequency data acquisition. The implicit correlation data acquisition unit (12) captures potential fault risks by monitoring photovoltaic equipment aging data and partial shading data. The data transmission unit (13) is responsible for transmitting the collected data to the data processing layer in real time through different communication protocols; the data fusion module (2) includes a data cleaning unit (21), a data fusion unit (22), and a data standardization unit (23); the data cleaning unit (21) processes the collected raw data to remove noise and outliers; the data fusion unit (22) generates a unified standardized dataset by weighted fusion of data from different sources; the data standardization unit (23) is responsible for converting multi-dimensional and multi-format data into a consistent standard format; the dynamic early warning module (4) includes a working condition adaptation unit (41), a life cycle stage identification unit (42), and a fault risk prediction unit (43); the working condition adaptation unit (41) uses a working condition classification model based on the random forest algorithm, combined with light intensity and ambient temperature, to divide the working environment of photovoltaic equipment into 9 working condition combinations, and dynamically adjusts the early warning threshold; the life cycle stage identification unit (42) uses a gradient boosting tree model, based on the historical operation and maintenance data of photovoltaic equipment, electricity Information data and photovoltaic equipment aging data automatically identify the life cycle stage of photovoltaic equipment and call different early warning strategies according to the stage results; the fault risk prediction unit (43) combines the LSTM-XGBoost fusion model and predicts the future fault risk probability of photovoltaic equipment based on the multi-dimensional data after fusion and the cluster benchmarking results; when the risk probability reaches or exceeds 60%, a hidden fault early warning is generated and the specific fault type is output; the intelligent operation and maintenance module (5) includes a revenue loss quantification unit (51), an early warning priority ranking unit (52), and an operation and maintenance decision output unit (53); the revenue loss quantification unit (51) quantifies the potential economic loss of photovoltaic equipment failure through revenue loss value and assesses the impact of photovoltaic equipment failure on power plant revenue; the early warning priority ranking unit (52) ranks the early warning information based on the revenue loss value and the fault urgency using the hierarchical analysis method; the operation and maintenance decision output unit (53) automatically generates operation and maintenance plan according to the early warning priority and the location of photovoltaic equipment and pushes it to the operation and maintenance personnel, while providing optimization guidance.

Citation Information

Patent Citations

  • Remote fault diagnosis method for photovoltaic power generation equipment

    CN119834735A

  • Distributed photovoltaic equipment detection method and system based on big data analysis

    CN120263102A

  • Photovoltaic power station intelligent analysis and fault intelligent diagnosis method and system

    CN120655260A

  • Intelligent fault diagnosis method integrating state monitoring and multi-mode large model

    CN120995768A

  • 5G base station micro-photovoltaic multi-source data dynamic charging method and system

    CN121012157A