Photovoltaic power generation fault diagnosis method based on AI analysis
By using an AI-based multi-source data fusion model, the problem of low efficiency in fault diagnosis of photovoltaic power plants has been solved, enabling accurate fault judgment and automated diagnosis, thereby improving the operational stability and safety of photovoltaic power plants.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SANAS ZHIWEI (QINGDAO) ELECTRIC POWER CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-21
AI Technical Summary
Existing fault diagnosis technologies for photovoltaic power plants are inefficient, unable to fully locate the root cause of faults, prone to misdiagnosis and omission, and pose safety risks, and cannot adapt to complex and multi-dimensional operating environments.
A photovoltaic power generation fault diagnosis method based on AI analysis is adopted. Through a multi-source data fusion model, fault feature extraction and in-depth analysis are achieved. Combined with data from drone inspections, environmental monitoring, video surveillance, etc., a multi-dimensional and multi-level data fusion model is constructed to perform fault judgment and predictive maintenance.
It significantly improves the accuracy and efficiency of fault diagnosis in photovoltaic power plants, reduces the false positive and false negative rates, automates fault type identification, location positioning, and early warning, provides reliable data support, and ensures the stable operation of power plants.
Smart Images

Figure CN122432929A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power fault diagnosis technology, and in particular to a photovoltaic power generation fault diagnosis method based on AI analysis. Background Technology
[0002] As an important form of new energy power generation, photovoltaic power plants are generally characterized by large land areas, a large number of components and equipment, and dispersed deployment. Furthermore, some power plants are built in remote areas such as deserts and mountains, operating in complex and harsh environments. Currently, the operation monitoring and fault diagnosis of photovoltaic power plants mainly rely on existing technologies such as manual on-site inspections, local threshold alarms for equipment, and monitoring of single electrical parameters. These technologies can only achieve simple status monitoring and explicit fault alarms for key equipment, and cannot perform collaborative analysis and comprehensive judgment of multi-type and multi-dimensional operational data.
[0003] However, the aforementioned existing technologies have significant shortcomings. On the one hand, manual inspection and traditional monitoring methods have low diagnostic efficiency and limited coverage, delayed fault response, and pose operational safety risks. On the other hand, single-dimensional monitoring methods have low fault identification accuracy, cannot distinguish between environmental interference and actual equipment faults, and are prone to misjudgment and missed judgment, making it difficult to support the efficient and stable operation of photovoltaic power plants. Summary of the Invention
[0004] To address the aforementioned issues, this invention provides a photovoltaic power generation fault diagnosis method based on AI analysis.
[0005] The photovoltaic power generation fault diagnosis method based on AI analysis provided by this invention includes the following steps: S1: Collect raw operating data of the photovoltaic power station, filter the raw operating data, and perform format standardization processing to generate standardized collected data; S2: Encrypt the standardized collected data to generate secure transmission data; S3: Decrypt the secure transmission data, extract the fault characteristics of the decrypted original transmission data, make a fault judgment on the fault characteristics, and generate a preliminary diagnosis result; S4: Based on the decrypted original transmission data and the preliminary diagnostic results, perform in-depth fault analysis to generate predictive maintenance suggestions, fault location information and fault level information.
[0006] In summary, the present invention has at least the following beneficial effects: 1. This invention, by comprehensively incorporating equipment operating status, environmental monitoring conditions, maintenance history information, and a multi-source data cross-validation mechanism, can fully adapt to the multi-source, coupled, and complex characteristics of photovoltaic power plant fault causes. It effectively overcomes the inherent limitation of single-dimensional monitoring data, which can only reflect local anomalies and cannot comprehensively pinpoint the root cause of faults. Through the mutual supplementation and synergistic verification of multi-source information, it achieves a comprehensive analysis and accurate judgment of the fault generation mechanism, influencing factors, and propagation path. Simultaneously, by constructing a multi-dimensional, multi-level data fusion model, this invention can fully explore the inherent correlations and logic between various types of data, effectively improving the completeness, systematicness, and objectivity of fault diagnosis, providing reliable data support for subsequent fault location and early warning.
[0007] 2. This invention, based on an artificial intelligence multi-source data fusion model, achieves intelligent reasoning throughout the entire process from data anomaly extraction to fault root cause localization. It can effectively distinguish between equipment malfunctions and state changes caused by environmental fluctuations, significantly reducing the false positive and false negative rates, and effectively improving the reliability, stability, and intelligence level of photovoltaic power plant fault diagnosis. Simultaneously, this invention can automatically complete a series of intelligent diagnostic processes, including fault type identification, fault location, and fault warning output, achieving automation and intelligence in fault diagnosis without manual intervention. This greatly improves the efficiency and accuracy of photovoltaic power plant fault diagnosis, providing strong technical support for the long-term stable, safe, and reliable operation of photovoltaic power plants. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a system architecture diagram of the photovoltaic power station multi-source data acquisition, secure transmission and multi-terminal application of the present invention; Figure 2 This is a schematic diagram of the visual monitoring interface for the operating status of the transformer substation of the present invention; Figure 3 This is a schematic diagram of the parameter monitoring interface of the photovoltaic power station environmental monitoring instrument of the present invention; Figure 4 These are thermal infrared images captured by the drone of this invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] The following is in conjunction with the appendix Figures 1 to 4 The present invention will be described in further detail below.
[0012] The AI-based photovoltaic power generation fault diagnosis method includes the following steps: Step 1: Collection of raw operating data.
[0013] The data acquisition module is used to collect basic data from photovoltaic power plants. The types of data mainly include branch current data, combiner box data, inverter data, transformer data, environmental monitoring data, dust monitoring data, drone inspection data, and video surveillance data.
[0014] A branch circuit is a basic power generation unit formed by connecting photovoltaic modules in series. Its current data directly reflects the health status of the module, and the data collected includes: Real-time branch current: The output current value (unit: A) of each photovoltaic module, monitoring for abnormalities such as low current (e.g., module microcracks, shading) or zero current (e.g., line disconnection).
[0015] Branch voltage: The output voltage value of the corresponding branch (unit: V). Combined with the current, the branch power is calculated to determine whether there is an abnormal voltage drop in the cascade.
[0016] Branch current consistency: The difference in current among branches in the same array. A large difference usually indicates that some branches are faulty or blocked.
[0017] Branch circuit temperature: Some data acquisition devices integrate temperature sensors to monitor the temperature of branch circuit junction boxes and prevent overheating faults caused by loose wiring.
[0018] The combiner box is responsible for combining the current from multiple photovoltaic branches and sending it to the inverter, and for collecting and focusing the combiner status and electrical safety data. Input current / voltage for each branch: Monitor the current and voltage connected to the combiner box for each branch to locate the faulty branch.
[0019] Total current / voltage after merging: Electrical parameters on the output side of the combiner box, used to determine whether the combiner function is normal.
[0020] Surge protector status: The on / off state of the surge protection module, records the number of lightning strikes, and evaluates the effectiveness of the surge protection.
[0021] Fuse status: The on / off status of each input fuse; an alarm is triggered when a fuse blows.
[0022] Enclosure environmental data: internal temperature and humidity, to prevent short circuits or component aging caused by moisture and high temperature.
[0023] Inverters are the core conversion equipment in photovoltaic power plants, and the collected data needs to cover power conversion, operating status, and fault alarms. DC side data: input voltage, input current, input power, to determine whether the DC transmission from the combiner box to the inverter is normal.
[0024] AC side data: output voltage, output current, output power, power factor, frequency, to monitor whether the grid-connected power quality meets the standards.
[0025] Conversion efficiency: Calculates the inverter's DC-to-AC conversion efficiency in real time. A decrease in efficiency indicates a fault in the inverter's internal components.
[0026] Operating status parameters: inverter module temperature, cooling fan speed, and IGBT (Insulated Gate Bipolar Transistor) operating status.
[0027] Fault alarm information: Fault codes detected by the inverter itself (such as overvoltage, overcurrent, islanding protection, module failure, etc.).
[0028] Power generation data: daily power generation, monthly power generation, and cumulative power generation, used for energy efficiency assessment and fault tracing.
[0029] like Figure 2 As shown, the transformer substation is responsible for boosting the low-voltage AC power output from the inverter to the grid connection voltage, and collecting data on boosting and grid connection safety: Electrical parameters of high and low voltage sides: low voltage side input voltage / current, high voltage side output voltage / current, active power, reactive power.
[0030] Tap changer position: The position of the voltage regulating tap changer is used to monitor whether the voltage regulation is normal.
[0031] Circuit breaker / disconnector status: Opening and closing status of the high-voltage side circuit breaker, and record the number of operations.
[0032] Grounding resistance: The resistance value of the grounding system of the transformer substation, ensuring the safety of personnel and equipment.
[0033] like Figure 3 As shown, environmental data is used to distinguish between power drops caused by faults and power fluctuations caused by natural environmental factors. The collected data includes: Irradiance: Solar irradiance intensity on the surface of a photovoltaic module (unit: W / m²) 2 ), is the core parameter for calculating theoretical power generation.
[0034] Ambient temperature: The air temperature at the power plant site affects the power generation efficiency of the components and the operating temperature of the equipment.
[0035] Module backsheet temperature: Directly monitor the operating temperature of the photovoltaic module. Excessive temperature will lead to a decrease in power generation efficiency.
[0036] Wind speed and direction: Affect the rate of heat dissipation and dust accumulation on the component surface, and are used to assess the impact of wind and sand on power plants.
[0037] Humidity and rainfall: Humidity affects the insulation performance of equipment, while rainfall can be used to determine changes in the cleanliness of component surfaces.
[0038] Dust accumulation is a significant factor contributing to the power degradation of photovoltaic modules. This study focuses on the degree of dust accumulation and cleaning requirements. Dust accumulation thickness: The thickness of dust on the component surface is monitored by an optical sensor (unit: μm).
[0039] Light transmittance degradation rate: The percentage decrease in light transmittance of the module glass due to dust, which is directly related to the loss of power generation efficiency.
[0040] Dust type identification: Some high-precision detectors can distinguish different types of pollutants such as dust, sand, and bird droppings.
[0041] Dust accumulation rate: The increase in dust accumulation thickness per unit time, used to develop differentiated cleaning strategies (such as shortening the cleaning cycle for areas with rapid dust accumulation).
[0042] Drone inspections primarily collect images of physical defects and thermal imaging data of photovoltaic modules for fault identification using models. The types of data collected include: Visible light image data: visible light features on the component surface such as cracks, breakage, delamination of the encapsulation, damage to the junction box, and obstruction by weeds.
[0043] like Figure 4 As shown, infrared thermal imaging data: temperature distribution map of the component, identifying hidden defects such as hot spot effect (abnormal local temperature rise), microcracks, and diode failure.
[0044] Inspection path and location data: The drone's flight trajectory and GPS positioning information accurately mark the location coordinates of the faulty components.
[0045] Multispectral image data: Some drones are equipped with multispectral cameras to collect spectral information in different bands and identify early faults such as component PID effect and aging.
[0046] Video surveillance is used for power plant security and visual monitoring of equipment operation status. The collected content includes: Equipment status video stream: Real-time video monitoring of key equipment areas such as inverter room, transformer substation, and combiner box to identify abnormalities such as smoke, unusual noises, and foreign object intrusion.
[0047] Security monitoring data: Video footage of the power station perimeter and entrances / exits to prevent malfunctions caused by human factors such as theft and vandalism.
[0048] Image recognition auxiliary data: By analyzing video footage through algorithms, problems such as component obstruction and personnel violations can be automatically identified.
[0049] The system uses the sensing and acquisition layer as the data source, and its core function is to directly connect to various monitoring devices to complete the collection, preliminary filtering and format standardization of raw data.
[0050] Branch current monitoring uses series current sensors and voltage sensors to collect current, voltage, and temperature data for each component string.
[0051] The combiner box / inverter / substation collects data such as electrical parameters, switch status, and fault codes through its built-in RS485 / Modbus / Profinet interface.
[0052] Environmental / dust monitoring uses irradiance sensors, temperature and humidity sensors, and dust transmittance meters to collect data such as irradiance, ambient temperature, component backplane temperature, and dust accumulation thickness.
[0053] The drone inspection uses visible light cameras, infrared thermal imaging cameras, and multispectral cameras to collect surface images, temperature maps, and spectral data of components, and simultaneously records GPS positioning information.
[0054] Video surveillance uses high-definition cameras and intelligent PTZ cameras to collect real-time video streams of the equipment area and the perimeter of the power station.
[0055] Step 2: Standardization processing of raw operating data.
[0056] Industrial-grade data acquisition units (DTUs / RTUs) are used as the core sensing terminals. They have the ability to parse multiple protocols and convert them into a unified format. They implement standardized processing for the complex multi-source data collected by photovoltaic power plants to ensure the data quality and transmission efficiency for subsequent fault diagnosis.
[0057] Industrial-grade data acquisition units (DTUs / RTUs) are compatible with various industrial communication protocols such as Modbus, IEC 61850, and MQTT via serial ports, Ethernet, and wireless communication. They connect to current and voltage sensors in photovoltaic modules, operating parameter acquisition modules in inverters / combiner boxes, environmental monitoring instruments (irradiance, temperature, wind speed), and UAV image acquisition terminals. After protocol parsing, the industrial-grade data acquisition unit converts all heterogeneous data into structured JSON or CSV formats. JSON format stores multi-dimensional correlated data including device number, acquisition time, and parameter type, while CSV format stores purely numerical time-series electrical and environmental data. This achieves unified management of cross-protocol and cross-type data, eliminating data format conflicts caused by protocol differences.
[0058] In the edge-side preprocessing stage, the industrial-grade data acquisition device performs preliminary filtering and cleaning operations based on local computing power, focusing on removing the following types of obviously invalid data: First, there are abnormal values due to sensor disconnection. When the current or voltage sensor of a certain channel fails to return data for several consecutive sampling cycles due to communication failure or power outage, the industrial-grade data acquisition device automatically marks the data for that period as invalid and discards it.
[0059] Second, invalid data that exceeds the limits, such as sudden changes in photovoltaic module current far exceeding the rated value, abnormal values of irradiance data that are positive at night, and extreme jumps in ambient temperature that exceed common geographical and climatic knowledge, are determined by using preset equipment rated parameter ranges and environmental knowledge thresholds.
[0060] Third, duplicate and redundant invalid data: For the same device and the same parameter, if the same duplicate value appears within the collection period, or the collection frequency is higher than the effective frequency required for subsequent system processing (e.g., 10 collections per second but only 1 collection per second is required), the industrial-grade data acquisition device will perform deduplication and downsampling processing.
[0061] Fourthly, invalid data due to local interference, such as blurred, overexposed, or completely distorted image data caused by cloud cover, birds flying by, or equipment reflection during drone aerial image acquisition, as well as abnormal hot spot data in fixed areas in infrared thermal images caused by lens dirt.
[0062] Through the above multi-dimensional filtering, the industrial-grade data acquisition device effectively removes noise and invalid data fragments, retaining only valid data fragments, significantly reducing the bandwidth pressure of uplink transmission and reducing the resource consumption of invalid data on cloud and edge computing nodes.
[0063] Step 3: Encrypting standardized data.
[0064] The system uses the edge transmission layer as a data communication hub, and its core function is to securely and stably transmit the data from the sensing and acquisition layer to the edge computing nodes or cloud platform.
[0065] The edge transmission layer uses an edge gateway and a security chip to work together to build a complete data encryption transmission mechanism. The input data for the encryption process is the standardized raw data output by the sensing and acquisition layer. This data includes structured JSON / CSV data such as branch current and voltage, combiner box, inverter, transformer, environment, and dust monitoring, as well as unstructured binary data such as UAV visible light images, infrared thermal imaging images, and video surveillance streams. This data is the only source for encryption processing.
[0066] First, the data to be transmitted is uniformly encapsulated into data frames. A fixed-format data header is constructed for each group of data. The data header includes, in sequence, the total length of the data header, the unique device code, the millisecond-level timestamp, the data type identifier, the data length, and the CRC32 checksum. The unique device code is used to locate the device to which the data belongs, the timestamp is used to prevent replay attacks, the data type identifier distinguishes between structured and unstructured data, and the CRC32 checksum is used to verify the integrity of the data after decryption. After encapsulation, a plaintext data packet to be encrypted is formed, consisting of the data header and the original data body.
[0067] Subsequently, key and security parameter generation operations are performed. The system pre-programs a 32-byte AES-256 master key into the security chip of the edge gateway and edge computing node. Each time a transmission task is started, a 12-byte non-repeating initialization vector IV and a 32-byte one-time session key SK are randomly generated. The session key SK is then encrypted using the master key to obtain the encrypted session key ESK. This key will be transmitted along with the ciphertext data to ensure the secure transmission of the session key.
[0068] Next, differentiated encryption operations are performed according to data type. For structured data in JSON and CSV formats, the data is divided into blocks of 1024 bytes each, with padding of zeros for any shortfall. AES-256-GCM is used to encrypt each block, generating a 16-byte authentication tag to verify data integrity. For unstructured data in image and video stream formats, a streaming, segment-by-segment encryption method is used, with each segment being 4096 bytes long. The original file format and encoding rules are not modified; only the data content is encrypted using AES-256-GCM, maintaining the data transmission order.
[0069] After encryption, the transmission frames are finally assembled. The Initialization Vector (IV), Encryption Session Key (ESK), ciphertext header, ciphertext data body, and authentication tag (TAG) are combined in a fixed order to form a complete encrypted transmission frame. A unique frame sequence number is added to each frame to mark the transmission progress when resuming interrupted transmission. For data frames from critical equipment such as inverters and transformer substations, multi-path redundancy transmission configuration is enabled to improve data transmission stability in extreme environments.
[0070] Finally, the encrypted final data is output, namely the encrypted secure transmission data. This data is in ciphertext form and can be transmitted through communication links such as fiber optic, Ethernet, 4G / 5G, LoRa, and NB-IoT. It effectively prevents eavesdropping, tampering, and replay attacks during transmission. Moreover, this ciphertext data cannot be directly used for subsequent calculations and processing; it must be decrypted before entering the edge computing process, thereby ensuring the security, integrity, and reliability of the entire data transmission process.
[0071] Step 4: Decrypting the securely transmitted data.
[0072] The system achieves localized decryption and real-time analysis through the edge computing layer. Its core functions are: to complete data decryption and rapid fault early warning locally at the power plant, reduce dependence on the cloud, and reduce response latency.
[0073] The input data for the decryption process is encrypted and secure transmission data from the edge transmission layer. This data completely preserves the transmission frame structure of the encryption stage and is the only data source for decryption processing. The decryption operation is completed before the edge computing node starts, ensuring the continuity of subsequent data processing.
[0074] First, the encrypted secure transmission data is decomposed into a frame structure, and a 12-byte initialization vector (IV), an encryption session key (ESK), a ciphertext data body, and a 16-byte authentication tag (TAG) are extracted from fixed positions. All parameters are extracted strictly in the order of assembly during encryption to ensure the accuracy of parameter acquisition and provide a foundation for subsequent decryption and verification.
[0075] The session key decryption operation is then performed. The AES-256 master key pre-stored in the edge computing node's security chip is used to decrypt the encrypted session key ESK and restore the session key SK used for this transmission. If key mismatch or decryption failure occurs during the decryption process, the current data frame is immediately discarded, a key abnormality alarm is reported to the cloud, and the decryption process is terminated.
[0076] Next, AES-256-GCM decryption and integrity verification are performed. The restored session key SK, the extracted initialization vector IV, and the authentication tag TAG are input into the decryption algorithm to decrypt the ciphertext data. At the same time, the integrity of the data is verified through the authentication tag TAG. If the tag verification fails, it is determined that the data has been tampered with or the transmission is corrupted. The data is discarded and a tampering alarm is reported. If the verification passes, the subsequent parsing steps continue.
[0077] After decryption, the data header is parsed to extract fixed-format data header information from the decrypted plaintext, obtain the device unique code, collection timestamp, data type, data length and CRC32 checksum, determine whether the data is replay data by the timestamp, determine the subsequent data restoration method by the data type identifier, and complete the data ownership location by the device code.
[0078] Then, plaintext restoration is performed according to data type. For structured data identified as JSON or CSV, it is directly restored to standard text format, which can be directly used for data cleaning, normalization, and feature extraction. For unstructured data identified as images or video streams, it is restored to the original format of JPG, PNG, infrared images, or H.264 / H.265 video streams, which can be directly fed into lightweight AI models for image recognition and fault detection.
[0079] Finally, a final data integrity check is performed. The CRC32 checksum is recalculated using the restored plaintext data and compared with the checksum carried in the data header. If the comparison matches, the data is confirmed to be complete and valid. If the comparison does not match, it is marked as bad data, triggering a retransmission request. After the check passes, standardized raw collected plaintext data is output. This data can be directly used as input data for the edge computing layer for subsequent processing processes such as real-time threshold alarms, fault feature extraction, and local fault diagnosis, realizing the end-to-end connection of encryption, decryption, and data processing.
[0080] Step 5: Identification of explicit faults.
[0081] In the real-time monitoring and threshold alarm stage, the edge computing layer monitors real-time electrical parameters such as branch current, inverter output power, and combiner box circuit status, and directly compares them with preset thresholds. When parameters exceed limits, drop suddenly, or are interrupted, audible and visual alarms and maintenance terminal reminders are immediately triggered. At the same time, the video monitoring stream is analyzed in real time to identify obvious problems such as component obstruction and personnel violations, achieving a response time of seconds.
[0082] In the data preprocessing and feature extraction stage, the edge computing layer cleans, removes noise, fills in missing values and normalizes the collected raw data, and extracts fault feature parameters such as branch current consistency deviation, inverter conversion efficiency change rate and component infrared hot spot temperature difference to form standardized feature data, providing support for in-depth analysis in the cloud.
[0083] In the localized fault diagnosis stage, lightweight AI models, such as lightweight CNN models and lightweight decision tree models, are deployed at the edge computing layer. Based on thresholds and key operational data, they directly identify explicit faults, enabling rapid diagnosis of common explicit faults locally without uploading massive amounts of raw data. The specific diagnostic steps are as follows: By monitoring that the output circuit current of the combiner box is zero and the voltage of the upstream and downstream branches is normal, it can be directly determined that the fuse of the corresponding circuit has blown.
[0084] By monitoring whether the DC-side input current or AC-side output current of the inverter exceeds the rated threshold, or whether the DC-side input voltage or AC-side output voltage exceeds the preset range, the inverter can be directly identified as having overcurrent or overvoltage faults.
[0085] If the monitoring device continuously loses communication messages and does not update data, and the power supply and link of the corresponding device are normal, it can be directly determined that the device communication is interrupted.
[0086] By monitoring the photovoltaic branch current to be zero or drop sharply, and the branch voltage to deviate abnormally, the photovoltaic branch disconnection fault can be directly determined.
[0087] By monitoring the temperature of the component backplane and the inverter module, if they exceed the safety threshold, the system can directly determine if the component is overheating or the equipment is overheating.
[0088] Step 6: In-depth fault diagnosis.
[0089] The system uses a cloud-based analytics layer as its core fault diagnosis engine. Its functions include: leveraging the massive computing power and storage in the cloud to perform large-scale in-depth data analysis, complex fault diagnosis, and predictive maintenance.
[0090] Build time-series databases (such as InfluxDB and TimescaleDB) to store electrical operation data and environmental data; build image databases to store drone inspection images and video frame data.
[0091] Establish a fault knowledge base, integrating historical fault cases, equipment manuals, and diagnostic rules to provide training basis for AI models.
[0092] Deploy multi-model fusion diagnostic algorithms in the cloud analytics layer: Based on the gradient boosting tree model: analyze the correlation of electrical parameters and diagnose equipment faults such as abnormal branch current, inverter faults, and transformer substation faults.
[0093] Based on deep learning models, feature extraction and target detection are performed on infrared and visible light images of UAVs to identify physical defects such as component microcracks, hot spots, and encapsulation delamination.
[0094] Based on the time-series prediction model: Time-series prediction and deviation analysis are performed on operating data such as current, power, and efficiency to identify trend faults such as year-on-year / month-on-month anomalies, efficiency degradation, and abnormal power loss.
[0095] Based on the association rule mining model: analyze the correlation between equipment replacement frequency, failure rate, runtime and environmental conditions, and identify maintenance-related faults such as equipment batch quality problems and abnormal maintenance matching.
[0096] The steps of using the gradient boosting tree model to diagnose electrical operation data are as follows: The gradient boosting tree model takes residual iterative correction as its core idea. First, it constructs an initial base model to provide a basic prediction benchmark for subsequent iterations.
[0097] Construct a gradient boosting tree model with electrical and environmental parameters as inputs and equipment operating status as output, and perform initialization calculations: In the formula, denoted as the initial prediction value of the gradient boosting tree model, i.e., the prediction function in the 0th iteration; x represents the input sample vector, which includes electrical operation data such as photovoltaic branch current / voltage, combiner box electrical parameters, inverter AC / DC parameters, and transformer high and low voltage side parameters, as well as environmental monitoring data such as irradiance, ambient temperature, and module backsheet temperature; c represents the constant term to be optimized, i.e., the optimal prediction value of the initial gradient boosting tree model. This represents the label value of the i-th sample. In regression tasks, it is a continuous electrical parameter value, and in classification tasks, it is a discrete label of normal / fault. This represents the loss function, used to measure the difference between the predicted value c and the true label. The error between them; 'a' represents the total number of training samples; This means finding the optimal value among all possible values of c that minimizes the sum of the loss function.
[0098] Based on the different task types, the initial gradient boosting tree model is divided into regression tasks and classification tasks. Regression tasks are used for continuous electrical parameter prediction and deviation analysis, while classification tasks are used for normal / fault binary classification identification.
[0099] The regression task uses the squared loss function to measure the prediction error, which takes the form: In the formula, y represents the label value of the sample, including continuous electrical parameter values such as branch current and inverter output power; This indicates the squared loss function used in the regression task, and F represents the predicted value of the gradient boosting tree model.
[0100] Under the squared loss function, the optimal solution for the initial gradient boosting tree model is the arithmetic mean of the training set labels: In the formula, The arithmetic mean of the sample labels is used as the initial prediction baseline for the initial gradient boosting tree and is used for subsequent fitting of parameter bias.
[0101] The classification task uses a log loss function to measure prediction error. The optimal solution of the initial gradient boosting tree model is the log odds of the positive class samples in the training set. The classification task is set as follows: In the formula, p represents the initial prior probability of faulty samples in the training set, that is, the proportion of faulty samples to the total number of samples; 1-p represents the initial prior probability of normal samples in the training set. This represents the logarithmic probability function, which maps the probability to the real number space, facilitating subsequent iterative optimization.
[0102] With the initial base model Starting with the previous iteration, the decision tree is trained round by round to fit the residuals of the previous model, gradually correcting the prediction bias, and finally obtaining the complete model after T iterations: In the formula, This represents the model for the final gradient boost, i.e., the complete model used for electrical fault diagnosis; The step size represents the learning rate; T represents the total number of iterations, i.e., the number of decision trees; t represents the index of the current iteration round, with values of 1, 2, 3, ..., T. This represents the decision tree model trained in the t-th iteration, used to fit the residuals of the previous iteration model and correct prediction bias.
[0103] During training, the gradient boosting tree model aims to minimize the overall loss and updates the prediction function round by round, ultimately obtaining the optimal diagnostic model that balances bias and variance.
[0104] The feature importance score is calculated solely based on the decrease in loss function caused by the split of the base decision tree node. It is accumulated layer by layer along the path of single split gain → single tree feature score → total feature score of the whole model, which is completely consistent with the model iteration structure and strongly correlated with each iteration.
[0105] In single-base decision trees During the construction process, each node split involves traversing all input features and selecting the feature that maximizes the decrease in the loss function along with the split point to complete the partition. The split gain of feature f at node m is calculated. : In the formula, Representation of features The decrease in loss brought about by this split directly reflects the contribution of this feature to improving prediction accuracy; This represents the total loss before node splitting, corresponding to the squared loss function in regression tasks and the correspondence loss function in classification tasks. These represent the total loss of the left and right child nodes after a node splits, respectively.
[0106] The same feature may be used for splitting in multiple nodes of a tree. Therefore, the importance score of feature f within a single tree is the sum of the gains of all nodes split by feature f, calculated as follows: In the formula, Let f represent the cumulative importance score of feature f in the t-th decision tree; This represents the set of all nodes in the tree that undergo splitting using feature f.
[0107] This step summarizes the contributions of a single split into feature importance at the level of a single decision tree.
[0108] Final diagnostic model Depend on Weighted ensemble of decision trees, learning rate To ensure consistent weighting coefficients, the feature importance scores across the entire model must be weighted and accumulated over the entire tree, maintaining consistency with the model iteration formula. In the formula, This represents the global importance score of feature f in the entire gradient boosting tree model.
[0109] By sorting all input features in descending order according to the global importance score I(f), the contribution weight of each feature in fault diagnosis can be obtained, providing a quantitative basis for subsequent fault root cause localization.
[0110] After the gradient boosting tree model is trained, the root cause of the fault can be directly located based on the combination of high-contribution features. A typical mapping relationship is as follows: When the high-contribution feature combination results in current feature importance being much greater than voltage feature importance, the output photovoltaic module shading fault and module microcrack fault will be detected.
[0111] When the high-contribution feature combination results in increased importance of both current and voltage features, it indicates an open circuit fault in the output branch or a poor contact fault in the connector.
[0112] When the high-contribution feature combination results in overcurrent / overvoltage features being far more important than other parameter features, the output inverter will experience overcurrent protection faults and system short-circuit faults.
[0113] When the combination of high-contribution features increases the importance of both irradiance and current features, the output environmental fluctuation judgment result is obtained.
[0114] The steps for deep learning models to diagnose drone image data are as follows: Acquire infrared images of photovoltaic modules captured by a drone, and define the dimensions for both the single-channel grayscale image and the three-channel color image: the input dimensions for the single-channel grayscale image are H×W, and the input dimensions for the three-channel color image are H×W×C. in C in Indicates the number of input channels, C for RGB images in =3.
[0115] Perform convolution calculations on the input infrared image, and calculate the output feature map size based on the input size, padding value, kernel size, and stride. The calculation formula is as follows: middle, Indicates the height of the input image. Indicates the width of the input image. Indicates the fill value. Indicates the height of the convolution kernel. Indicates the kernel width. Indicates the convolution stride. Indicates the height of the output feature map. This indicates the width of the output feature map.
[0116] The convolution output feature map is normalized using the following formula: In the formula, Represents the coordinates of the convolutional feature map eigenvalues at that location This represents the minimum value of the feature map. This represents the maximum value of the feature map. This represents the normalized eigenvalues.
[0117] Max pooling is performed on the normalized feature map, with a pooling region size of [size missing]. The pooling formula is: In the formula, Represents the eigenvalues after pooling. This represents the feature value at coordinates (x, y) of the convolutional feature map after normalization.
[0118] The pooling results are subjected to nonlinear activation processing, and the activation formula is: In the formula, This indicates that the output feature is activated, generating a high-temperature region location feature map.
[0119] Extract the hotspot region from the localization feature map and calculate the hotspot area: In the formula, Represents the set of hotspot pixels. This indicates the area of the hot spot.
[0120] Calculate the temperature difference between the hot spot region and the normal region: In the formula: This indicates the average temperature of the hot spot region. This represents the average temperature of the normal area. This represents the temperature difference.
[0121] The hot spot identification results are matched with the corresponding branch current data to calculate the current attenuation magnitude: In the formula, This represents the average current of the normal branches in the same array. This indicates the real-time current of the branch corresponding to the hot spot. This indicates the magnitude of current attenuation.
[0122] The degree of power generation attenuation is determined based on the temperature difference and current attenuation amplitude, and the output includes the hot spot location coordinates, hot spot area, temperature difference, current attenuation amplitude, and power generation attenuation determination result.
[0123] The steps for time series forecasting models to diagnose statistical comparison data are as follows: First, an autoregressive moving average model ARMA(p,q) is constructed to predict normal operating conditions based on historical time-series data. The model expression is as follows: In the formula, p represents the autoregression order, which is determined by the PACF plot; Here, c represents the autoregressive coefficient, and c represents the constant term. This represents white noise with a mean of 0 and a constant variance. q represents the historical white noise over the first m time steps, and q represents the moving average order, which is determined by the ACF plot. Represents the moving average coefficient. This represents the predicted value at time b. This represents the true value at the nth historical moment.
[0124] The absolute deviation is calculated based on the predicted value and the real-time collected value. This is used to measure the deviation of actual values from theoretical normal values. In the formula, This represents the actual collected value at time b.
[0125] Calculate the relative deviation rate based on the absolute deviation. This is used to normalize deviations, enabling comparison of the degree of anomalies between different devices and parameters: Obtain the irradiance at time b Component temperature and dust transmittance Normalization is then performed to unify the magnitude of environmental parameters and eliminate the influence of dimensions. In the formula, Indicates standard irradiance. Indicates the standard test temperature. Indicates standard transmittance; Environmental correction coefficients are constructed based on the normalized environmental parameters. This is used to quantify the impact of environmental factors on power generation parameters. In the formula, Indicates the weighting coefficient. ; Perform a correction operation on the relative deviation rate to obtain the corrected deviation. This is used to eliminate environmental interference and retain the true deviations caused by the deterioration of the equipment's own condition. Extract historical data from the time series database and calculate the year-on-year deviation rate. Used to identify long-term performance degradation trends over an annual cycle: In the formula, This indicates a historical moment 52 weeks prior to moment b; Extract data from the same period last week from the time series database and calculate the month-on-month deviation rate. Used to identify short-term operational fluctuations and sudden anomalies: In the formula, This represents a historical moment one week prior to moment b. The corrected deviation The year-on-year deviation rate The aforementioned month-on-month deviation rate Each value is compared with a corresponding preset threshold, and a fault determination is performed based on the comparison result, outputting the corresponding fault type.
[0126] The steps for the association rule mining model to analyze historical operation and maintenance data are as follows: Collect equipment replacement records, fault types, runtime, environmental conditions, and component batch information to construct a transaction dataset D for association analysis, providing a data foundation for subsequent frequent itemset mining and rule extraction.
[0127] Frequent itemset mining was performed using the Apriori algorithm, and support was calculated: In the formula, This indicates equipment replacement events, such as the replacement of photovoltaic modules, inverters, combiner boxes, and other equipment. This indicates the associated attributes related to equipment replacement, specifically including fault type, such as PID effect, hot spot fault, component damage, etc., equipment running time, environmental conditions, such as dust accumulation thickness, high temperature environment, etc., component batch information, etc. This represents the number of times that device replacement event A and attribute B occur simultaneously in transaction dataset D, i.e., the number of transactions where device replacement occurs simultaneously and is accompanied by the corresponding attribute B. Represents the dataset The total number of transactions in the dataset; support is used to measure the probability of association rules appearing in the global data and to filter high-frequency association patterns.
[0128] Calculate the confidence level: In the formula, This indicates the confidence level of the association rule if an equipment replacement event A occurs, which is accompanied by attribute B. It is mainly used to measure the certainty of the rule and reflects the probability that the corresponding attribute B will occur simultaneously when the equipment replacement occurs. This represents the total number of times that device replacement event A occurs alone, i.e., the total number of transactions related to all device replacements.
[0129] Confidence level is used to measure the degree of certainty of a rule, reflecting the probability that a corresponding fault or attribute will occur when equipment replacement takes place.
[0130] Calculate lift: In the formula, The association rule indicates that if device replacement event A occurs, the promotion of attribute B will follow. It is mainly used to determine whether there is a real association between A and B and to exclude false associations that occur by chance. support(A) represents the support of device replacement event A, that is, the probability of device replacement event A occurring in the global transaction data. support(B) represents the support of attribute B, that is, the probability of attribute B occurring in the global transaction data.
[0131] Lift is used to determine whether a rule has a real correlation and to exclude invalid rules that are only occasionally related.
[0132] It should be noted that when lift(A) When B)>1, it indicates a positive correlation between A and B, meaning there is a real correlation between equipment replacement and attribute B; when lift(A)>1, it indicates a positive correlation between A and B. When lift(A)=1, it means that A and B are independent, that is, they are unrelated; when lift(A)=1, it means that A and B are independent, that is, they are unrelated. When B) < 1, it indicates that there is a negative correlation between A and B, that is, the equipment replacement is inversely correlated with attribute B.
[0133] We select strong association rules that meet the minimum support, minimum confidence, and lift greater than 1, eliminate weak associations and spurious associations, and retain effective patterns with diagnostic value.
[0134] The strong association rules are sorted by support and confidence, and the strong associations between equipment replacement frequency and fault type, runtime, environmental conditions, and component batch are extracted.
[0135] For example, when a batch of components' replacement frequency itemset forms a strong association rule with PID effect, abnormal heating, and performance degradation itemset, it is determined that the batch of components has a quality defect problem, realizing the association reasoning from abnormal replacement frequency to root cause location.
[0136] The final output includes association rules, support, confidence, lift, and anomaly detection results for equipment replacement, providing a basis for component quality assessment, operation and maintenance strategy optimization, and preventive replacement.
[0137] Step 7: Probabilistic analysis of multidimensional data diagnostic results.
[0138] Multi-dimensional cross-validation employs a Bayesian model averaging method to fuse the outputs of multiple models, including electrical diagnostics, image recognition, time series analysis, and association mining, eliminating single-model bias and improving the accuracy and robustness of fault diagnosis. By aggregating diagnostic conclusions from different data sources and algorithm models, a joint decision on the final fault type and root cause is achieved through probability weighting.
[0139] Bayesian model average core fusion formula: In the formula, represents the global failure probability after fusion, and represents the comprehensive probability of finally determining the failure type y under input multi-source data X; M represents the total number of models participating in the fusion. This represents the prediction distribution of the m-th model, i.e., the probability of the fault type output by a single model; This represents the posterior weight of the m-th model, which is dynamically allocated based on the model's historical accuracy and data matching degree, and is used to reflect the decision credibility of different models.
[0140] The output probabilities of each independent model are uniformly input into the fusion layer to achieve the aggregation of multi-dimensional diagnostic information. The input content includes the electrical fault probability output by the gradient boosting tree model, the hot spot and hardware defect probability output by the CNN model, the efficiency anomaly and decay probability output by the ARMA time series model, and the equipment replacement and batch quality anomaly probability output by the Apriori correlation model, realizing the unified access of all-dimensional diagnostic results.
[0141] Based on the accuracy and recall of each model in historical samples, posterior weights are assigned to strengthen the decision-making proportion of high-confidence models and reduce the impact of weak or mismatched models. When electrical data is stable, the weight of the gradient boosting tree model is increased; when infrared images are clear, the weight of the CNN hotspot recognition model is increased; when performing long-term trend analysis, the weight of the time series prediction model is increased; and when conducting operation and maintenance data analysis, the weight of the Apriori association mining model is increased.
[0142] The comprehensive fault probability is obtained by weighted summation according to the Bayesian fusion formula. The fault type corresponding to the maximum probability is taken as the final diagnosis result. By using probability weighted fusion, a unique, stable and highly reliable fault conclusion is obtained, avoiding diagnostic bias caused by misjudgment by a single model.
[0143] Based on the fusion results, fault determination is completed, forming standardized decision logic for multiple scenarios. When the CNN hotspot recognition probability is 90%, the gradient boosting tree branch current deviation probability is 80%, and the timing efficiency decay probability is 75%, the fusion diagnosis is a hotspot fault in the component. When the gradient boosting tree voltage abnormality probability is 85%, the CNN infrared weak high temperature probability is 70%, the timing power continuous decay probability is 90%, and the association rule high replacement rate in the same batch is 80%, the fusion diagnosis is a PID effect fault in the component.
[0144] The final output includes the fault type, fault location, overall confidence level, contribution weights of each model, and processing suggestions. This completes the entire decision-making loop from independent diagnosis by multiple models to Bayesian weighted fusion, achieving high accuracy, low misjudgment rate, and strong robustness in AI fault diagnosis of photovoltaic systems, providing a reliable basis for operation and maintenance.
[0145] Step 8: Presentation of the entire decision-making process.
[0146] The system uses applications as the interaction entry point between the system and users. Its core function is to visualize the diagnostic results, predictive information and operation and maintenance suggestions output by the cloud analysis layer. It provides an intuitive, unified and operable interactive interface for operation and maintenance personnel and managers, enabling rapid response to fault alarms, closed-loop execution of operation and maintenance tasks, and comprehensive visibility of operation data, thereby improving the efficiency of power plant operation and maintenance and the level of management standardization.
[0147] The application presentation layer is built on a B / S architecture, supporting access from multiple terminals such as computers, large screens, mobile phones, and tablets. No dedicated client installation is required, facilitating remote management and on-site maintenance. The mobile terminal combines an app and a WeChat mini-program, balancing performance and convenience. It is suitable for outdoor maintenance scenarios in photovoltaic power plants, including those in the field, at high altitudes, and in remote areas, meeting the needs for viewing, handling, and providing feedback anytime, anywhere.
[0148] The real-time monitoring screen is used to display the overall operating status of the power plant, including key indicators such as power generation and equipment health. It also pushes fault alarm information in real time and displays the fault location intuitively on a map, making it easy for managers to quickly grasp the overall operating status of the plant.
[0149] The fault diagnosis report module can automatically generate structured reports that cover fault type, fault location, root cause analysis and handling suggestions. It supports historical query, export and archiving, providing data support for fault tracing and operation and maintenance summary.
[0150] The operation and maintenance management system mainly realizes functions such as work order dispatch, maintenance route planning, and maintenance record management. It tracks and manages the entire process of fault handling, forming a closed-loop operation and maintenance mechanism from diagnosis, dispatch, handling to archiving.
[0151] Mobile apps and mini-programs enable on-site maintenance personnel to remotely view fault information, receive alarm notifications, and upload maintenance results. They support offline operations and online synchronization, meeting the needs of efficient handling in remote areas.
[0152] like Figure 1 As shown, the system acquires raw electrical, environmental, image, and equipment status data from the perception and acquisition layer, and then securely and stably uploads it to the edge computing layer via the edge transmission layer. The edge computing layer performs real-time data preprocessing, threshold judgment, local rapid alarm, and feature extraction, and then uploads the effective data to the cloud analysis layer. The cloud analysis layer uses multi-model fusion AI algorithms to complete in-depth fault diagnosis, root cause location, fault classification, and predictive analysis. Finally, the diagnostic results, fault information, and operation and maintenance instructions are pushed to the application display layer, presented to users through a visual interface, and support operation and maintenance scheduling execution, forming a closed loop of the entire process of acquisition, transmission, calculation, diagnosis, presentation, operation and maintenance, and feedback.
[0153] The above are merely preferred embodiments of the invention and are not intended to limit the invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the protection scope of the invention.
Claims
1. A photovoltaic power generation fault diagnosis method based on AI analysis, characterized in that, Includes the following steps: S1: Collect raw operating data of the photovoltaic power station, filter the raw operating data, and perform format standardization processing to generate standardized collected data; S2: Encrypt the standardized collected data to generate secure transmission data; S3: Decrypt the secure transmission data, extract the fault characteristics of the decrypted original transmission data, make a fault judgment on the fault characteristics, and generate a preliminary diagnosis result; S4: Based on the decrypted original transmission data and the preliminary diagnostic results, perform in-depth fault analysis to generate predictive maintenance suggestions, fault location information and fault level information.
2. The photovoltaic power generation fault diagnosis method based on AI analysis according to claim 1, characterized in that, The original operational data includes structured data and unstructured data; The structured data includes branch current and voltage data of photovoltaic power plants, electrical parameters of combiner boxes, inverter operation data, transformer operation parameters, data collected by environmental monitoring instruments, and data from dust monitoring instruments. The unstructured data includes drone inspection image data and video surveillance stream data.
3. The photovoltaic power generation fault diagnosis method based on AI analysis according to claim 2, characterized in that, The specific steps for filtering the raw operational data and performing format standardization to generate standardized collected data are as follows: Outlier removal is performed on the structured data, and the data types are uniformly converted to JSON or CSV format; The original data format of the unstructured data is preserved, and device code, timestamp, and location information are added.
4. The photovoltaic power generation fault diagnosis method based on AI analysis according to claim 2, characterized in that, The specific steps for encrypting the standardized collected data to generate secure transmission data are as follows: The structured data is divided into blocks of 1024 bytes each. If the length of the structured data is less than 1024 bytes, the missing length is padded with "0" to generate block data. The segmented data is encrypted block by block, and a 16-byte authentication tag is generated simultaneously. The unstructured data is divided into blocks of 4096 bytes each, and each block is encrypted, while the authentication tag is generated. The encrypted structured and unstructured data are assembled into transmission frames, and the initialization vector, encryption session key, data header ciphertext, data body ciphertext, and authentication tag are combined in sequence to generate secure transmission data.
5. The photovoltaic power generation fault diagnosis method based on AI analysis according to claim 4, characterized in that, The specific steps for decrypting the securely transmitted data are as follows: The secure transmission data is decomposed into frames and extracted sequentially according to the order in which the transmission frames were assembled, to generate decomposed secure transmission data. The encrypted session key is decrypted using the master key. If decryption fails, a key error is reported and decryption is terminated. After a successful key match, the decrypted securely transmitted data is decrypted using a decryption algorithm, and the data integrity is verified using the authentication tag.
6. The photovoltaic power generation fault diagnosis method based on AI analysis according to claim 2, characterized in that, The specific steps for extracting fault characteristics from the decrypted original transmission data, judging the fault characteristics, and generating preliminary diagnostic results are as follows: Extract explicit faults from the decrypted original transmission data, and generate preliminary diagnostic results based on the information of the explicit faults; wherein, the explicit faults include electrical parameter over-limit faults, protection action faults, equipment status abnormality faults, and safety monitoring faults.
7. The photovoltaic power generation fault diagnosis method based on AI analysis according to claim 6, characterized in that, The specific steps for performing in-depth fault analysis based on the decrypted original transmission data and the preliminary diagnostic results to generate predictive maintenance suggestions, fault location information, and fault level information are as follows: Based on the unique code of the photovoltaic equipment, an equipment index is established for the decrypted original transmission data and the preliminary diagnostic results. The electrical operation data, UAV image data, statistical comparison data, and operation and maintenance history data are multi-dimensionally correlated and fused to generate multi-dimensional fused data. Based on the equipment health assessment model, the aging trend, failure probability and remaining service life of photovoltaic equipment are assessed using the multi-dimensional fused data, and predictive maintenance suggestions are generated. The multi-dimensional fused data is diagnosed through a multi-model collaborative mechanism to generate intermediate diagnostic results; By combining the intermediate diagnostic results with the device index, fault location information is generated; The intermediate diagnostic results are probabilistically weighted and cross-validated using a Bayesian fusion model to generate the final fault diagnosis results. Based on the scope of the fault's impact, the degree of harm, and the priority of its handling, the final fault diagnosis results are classified into emergency faults, general faults, and potential faults.
8. The photovoltaic power generation fault diagnosis method based on AI analysis according to claim 7, characterized in that, The multi-model collaborative mechanism includes a gradient boosting tree model, a deep learning model, a time series prediction model, and an association rule mining model; The steps for the gradient boosting tree model to diagnose the electrical operation data are as follows: Construct a gradient boosting tree model with electrical and environmental parameters as inputs and equipment operating status as output, and perform initialization calculations: In the formula, Let represent the initial predictions of the gradient boosting tree model, x represent the input sample vector, and c represent the constant term to be optimized. Let represent the label value of the i-th sample, and 'a' represent the total number of training samples. Based on the different task types, the initial gradient boosting tree model is divided into regression tasks and classification tasks; The regression task is set as follows: In the formula, This indicates the squared loss function used in the regression task, and F represents the predicted value of the gradient boosting tree model; Under the squared loss function, the optimal solution for the initial gradient boosting tree model is the arithmetic mean of the training set labels: In the formula, Represents the arithmetic mean of the sample labels; The classification task is set as follows: In the formula, p represents the initial prior probability of faulty samples in the training set; With the initial base model Starting with the previous iteration, the decision tree is trained round by round to fit the residuals of the previous model, gradually correcting the prediction bias, and finally obtaining the complete model after T iterations: In the formula, This represents the model for the final gradient boosting. Let T represent the learning rate, T represent the total number of iterations, and t represent the index of the current iteration round. Let represent the decision tree model trained in the t-th iteration; In a single decision tree Perform node splitting, traverse all input features, determine the feature and split point that maximizes the decrease in loss function, and calculate the splitting gain of feature f at node m. : In the formula, This represents the total loss before the node splits. These represent the total losses of the left and right child nodes after a node splits, respectively. For all nodes in the single decision tree that are split using feature f, perform split gain accumulation to obtain the importance score of feature f in that tree. : In the formula, This represents the set of all nodes in the tree that undergo splitting using feature f; The importance scores of feature f within all T decision trees are weighted and summed to obtain the global importance score of feature f. : Sort all input features in descending order according to their global importance score I(f), and output a sequence of feature contribution weights. Extract the highest-scoring features from the feature contribution weight sequence to form a high-contribution feature combination. Match the high-contribution feature combination with a preset feature combination and output the corresponding fault type based on the matching result.
9. The photovoltaic power generation fault diagnosis method based on AI analysis according to claim 8, characterized in that, The steps for the time-series prediction model to diagnose the statistical comparison data are as follows: Construct an autoregressive moving average model The expression is: In the formula, p represents the autoregressive order. Here, c represents the autoregressive coefficient, and c represents the constant term. Represents white noise. Let q represent the historical white noise over the first m time steps, and let q represent the moving average order. Represents the moving average coefficient. This represents the predicted value at time b. Represents the true value at the nth historical moment; The absolute deviation is calculated based on the predicted value and the real-time collected value. : In the formula, This represents the actual value collected at time b; Calculate the relative deviation rate based on the absolute deviation. : Obtain the irradiance at time b Component temperature and dust transmittance And perform normalization: In the formula, Indicates standard irradiance. Indicates the standard test temperature. Indicates standard transmittance; Environmental correction coefficients are constructed based on the normalized environmental parameters. : In the formula, Indicates the weighting coefficient. ; Perform a correction operation on the relative deviation rate to obtain the corrected deviation. : Extract historical data from the time series database and calculate the year-on-year deviation rate. : In the formula, This indicates a historical moment 52 weeks prior to moment b; Extract data from the same period last week from the time series database and calculate the month-on-month deviation rate. : In the formula, This represents a historical moment one week prior to moment b. The corrected deviation The year-on-year deviation rate The aforementioned month-on-month deviation rate Each value is compared with a corresponding preset threshold, and a fault determination is performed based on the comparison result, outputting the corresponding fault type.