Unmanned aerial vehicle inspection and AI fault diagnosis intelligent operation and maintenance system of photovoltaic power station
The intelligent operation and maintenance system, which combines drone inspection and AI diagnosis, utilizes high-resolution spectral sensors and current and voltage probes, along with similar component screening and adaptive feature selection, to improve the accuracy and efficiency of photovoltaic power plant fault diagnosis. It is highly adaptable and reduces operation and maintenance costs.
Patent Information
- Application Number
- CN202511867877.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-12-11
AI Technical Summary
In existing photovoltaic power plant operation and maintenance systems, drone inspection data is not fully utilized, and multispectral information is not deeply mined; feature selection methods are static and singular, and do not consider individual differences of components; fault diagnosis models have poor adaptability and are greatly affected by noise.
By using a drone equipped with a high-resolution spectral sensor and current and voltage probes, combined with similar component screening and adaptive feature selection, and training a fault diagnosis model through a deep neural network, fault type predictions and maintenance suggestions can be generated in real time.
It improves the accuracy and efficiency of photovoltaic power plant operation and maintenance diagnosis, adapts to changes in the power plant environment and the aging trend of components, reduces false alarms, optimizes maintenance resources, and extends the service life of components.
Smart Images

Figure CN121304142A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent operation and maintenance of photovoltaic power stations, in particular to an unmanned aerial vehicle inspection and AI fault diagnosis intelligent operation and maintenance system for photovoltaic power stations. BACKGROUND
[0002] As an important form of renewable energy, photovoltaic power generation plays a key role in global energy transformation. With the continuous expansion of photovoltaic power stations, the number of components has increased dramatically, and the operation and maintenance work is facing great challenges. Traditional photovoltaic power station operation and maintenance mainly relies on manual inspection. Inspectors detect component surface abnormalities such as hot spots, cracks, dirt, or electrical connection faults through visual observation or simple instruments such as infrared thermal imagers. This method is not only inefficient and has limited coverage, but is also affected by subjective factors, making it easy to miss subtle faults. Especially in large power stations, manual inspection is costly and time-consuming, and cannot meet the real-time monitoring needs. In recent years, unmanned aerial vehicle technology has been introduced into photovoltaic power station inspection. By carrying visible light or infrared cameras for aerial photography, component image data is obtained, and image processing algorithms are used to identify abnormal areas. This way of inspection significantly improves the speed, but existing unmanned aerial vehicle systems are mostly based on single type sensor data, such as visible light or thermal infrared data, with limited diagnostic capabilities. Visible light images mainly identify physical damage, and thermal infrared images detect temperature abnormalities, but neither can fully reflect the internal state of the components or early faults. Multispectral imaging technology can capture the spectral reflectance characteristics of components in different wavebands, providing more rich fault information, such as dirt accumulation, cell aging, or arc signs that may exhibit unique patterns in multispectral data. However, multispectral data has high dimensionality and large information volume, and how to effectively extract features related to faults becomes a difficulty. Existing processing methods often use conventional dimensionality reduction techniques such as principal component analysis or empirical band selection, but these methods often ignore the individual differences of photovoltaic components and environmental influences, resulting in inaccurate feature selection and weak generalization ability of diagnostic models.
[0003] Fault types of photovoltaic modules are diverse, including electrical faults (such as bypass diode failure, increased series resistance) and physical faults (such as cell micro-cracks, encapsulant degradation). Different faults may exhibit subtle differences in spectral reflectance, requiring high-precision sensors and advanced algorithms for analysis. In addition, photovoltaic modules are affected by manufacturing tolerances, installation angles, local environments (such as shadows, dust, temperature), etc., so even modules from the same batch may have different aging characteristics and failure modes. Existing fault diagnosis methods usually apply a uniform model to all modules without considering these individual differences, which can easily produce false positives or negatives. For example, a module may exhibit similar spectral characteristics to a fault due to local shading, but it is not actually a fault, and a uniform model may not be able to distinguish it. Artificial intelligence technologies, especially deep learning, have made progress in image recognition and fault diagnosis, and some studies have attempted to use convolutional neural networks for photovoltaic module defect detection. However, these models often require a large amount of labeled data, which is difficult to collect on-site at photovoltaic power stations and is costly to label. Moreover, on-site data is disturbed by weather, light, seasonal changes, and has a lot of noise, so directly training the model can easily overfit. Feature selection is a key step to improve model performance, and existing feature selection methods (such as filtering, wrapping, or embedding) are mostly based on global statistics to calculate the correlation between features and targets, but do not incorporate the concept of module similarity, resulting in poor adaptability in the diagnosis of specific modules. For example, global feature selection may highlight bands that are relevant to multiple modules, but ignore key features for a small number of modules.
[0004] On the other hand, historical fault information accumulated by photovoltaic power station operation and maintenance contains valuable knowledge, such as fault occurrence time, type, repair records, etc. Existing systems often simply use historical data as model input without further mining its relevance to current detection. The concept of similar component screening has been applied in reliability engineering, such as fault prediction based on similar equipment, but in photovoltaic power stations, the number of components is large, and there is still a lack of mature solutions for efficiently screening similar components and using them for feature selection. In addition, multi-spectral data contains a large number of redundant bands, which not only do not help in diagnosis, but also may introduce noise. Existing technologies lack dynamic noise suppression mechanisms and cannot adjust feature weights according to environmental changes. In summary, existing photovoltaic power station operation and maintenance systems have the following shortcomings: insufficient use of unmanned aerial vehicle inspection data, lack of deep mining of multi-spectral information; feature selection methods are static and single, without considering individual differences between components; fault diagnosis models have poor adaptability and are greatly affected by noise. The present invention proposes an intelligent system that integrates unmanned aerial vehicle inspection and AI diagnosis, improves diagnosis accuracy and efficiency through similar component screening and adaptive feature selection. SUMMARY
[0005] The present invention aims to provide an intelligent unmanned aerial vehicle inspection and AI fault diagnosis system for photovoltaic power stations to solve the problems raised in the background art.
[0006] To achieve the above object, the present application provides an unmanned aerial vehicle inspection and AI fault diagnosis intelligent operation and maintenance system for photovoltaic power stations, which comprises: A data acquisition component acquires multispectral reflectance data and electrical characteristic data of photovoltaic components through high-resolution spectral sensors and current-voltage probes carried by an unmanned aerial vehicle, and synchronously records fault history information of each photovoltaic component; A similar component screening component selects a set of similar photovoltaic components from photovoltaic components of the same installation batch according to the fault type classification of the current target photovoltaic component; A feature influence evaluation component calculates a fault response index of the target photovoltaic component at a specific spectral band based on the multispectral reflectance value distribution and fault history information distribution of the set of similar photovoltaic components of the target photovoltaic component; A key feature selection engine identifies a set of core spectral bands from all spectral bands according to the fault response index; An intelligent model construction component trains a fault diagnosis model based on a deep neural network using the multispectral reflectance data, electrical characteristic data and fault history information corresponding to the set of core spectral bands; An online diagnosis component inputs data acquired by real-time inspection of the unmanned aerial vehicle into the trained fault diagnosis model to generate fault type prediction and maintenance suggestions for photovoltaic components.
[0007] Preferably, the key feature selection engine further comprises: A correlation measurement component calculates a baseline correlation score of a specific spectral band according to the multispectral reflectance value distribution and fault history information distribution of all photovoltaic components, and optimizes the baseline correlation score in combination with the unique code of the photovoltaic component and the fault response index to obtain an enhanced correlation score; A noise influence estimation component calculates an environmental interference factor of a specific spectral band based on the fault history information distribution of the set of similar photovoltaic components of the target photovoltaic component and the fault history information distribution of all photovoltaic components; A key feature selection component identifies a set of core spectral bands from all spectral bands according to the environmental interference factor and the enhanced correlation score.
[0008] Preferably, the specific operation of the feature influence evaluation component comprises: For the target photovoltaic component, the reflectance value sequence of the set of similar photovoltaic components at a specific spectral band is extracted, and the entropy value of the sequence is calculated as a feature uncertainty indicator; The variance of the fault frequency sequence of the set of similar photovoltaic components is obtained as a fault fluctuation reference; calculating a correlation weight between the sequence of feature uncertainty indicators and the reference sequence of fault fluctuations, the weight being calculated by means of the mutual information quantity in information theory; multiplying the correlation weight by the feature uncertainty indicators to obtain the fault response index of the target photovoltaic module in the specific spectral band.
[0009] Preferably, the correlation measurement component performs the following when calculating the baseline correlation score: constructing a cumulative distribution sequence of the multispectral reflectance values and a cumulative distribution sequence of the fault history information of all photovoltaic modules; calculating the Hellinger distance between the cumulative distribution sequence of the multispectral reflectance values and the cumulative distribution sequence of the fault history information; linearly transforming and mapping the Hellinger distance to obtain the baseline correlation score in the specific spectral band.
[0010] Preferably, the correlation measurement component performs the following when optimizing the baseline correlation score: arranging all photovoltaic modules in descending order according to the fault severity level to generate a fault level sequence; arranging all photovoltaic modules in ascending order according to the multispectral reflectance value to generate a feature value sequence; calculating the Kendall concordance coefficient of the fault level sequence and the feature value sequence as a sequence consistency measure; eliminating the unique codes of the photovoltaic modules with matching positions from the fault level sequence and the feature value sequence to obtain a remaining fault level sequence and a remaining feature value sequence; replacing each unique code in the remaining fault level sequence with the fault response index of the corresponding photovoltaic module to form a fault response sequence; replacing each unique code in the remaining feature value sequence with the fault response index of the corresponding photovoltaic module to form a feature response sequence; calculating the Pearson correlation coefficient of the fault response sequence and the feature response sequence as a content consistency measure; calculating an optimization weight based on the sequence consistency measure and the content consistency measure, multiplying the optimization weight by the baseline correlation score to obtain an enhanced correlation score.
[0011] Preferably, the noise influence estimation component calculates the environmental interference factor by the following way: for the target photovoltaic module, obtaining a fault history information distribution density sequence of the set of similar photovoltaic modules; obtaining a fault history information distribution density sequence of all photovoltaic modules; calculating the Jensen-Shannon divergence between the fault history information distribution density sequence of the set of similar photovoltaic modules and the fault history information distribution density sequence of all photovoltaic modules; The average Jensen-Shannon divergence of all photovoltaic modules is inversely scaled to obtain the environmental interference factor of a specific spectral band.
[0012] Preferably, the key feature selection component operation includes: For each spectral band, the quotient value of the enhanced correlation score and the environmental interference factor is calculated. The quotient value is subjected to minimum-maximum normalization to obtain the selection priority score of the spectral band. The spectral bands with selection priority scores higher than the dynamic threshold are screened to form the core spectral band set.
[0013] Preferably, the intelligent model construction component uses a long short-term memory network architecture for model training. The multispectral reflectance data and time series electrical characteristic data corresponding to the core spectral band set are used as input feature vectors. The failure class code of the photovoltaic module is used as the output label. The network parameters are iteratively updated by the adaptive moment estimation algorithm to make the loss function converge, and the construction of the fault diagnosis model is completed.
[0014] Preferably, the online diagnosis component includes the following when processing real-time data: The reflectance values of the core spectral band set are extracted from the real-time streaming data of the unmanned aerial vehicle. The real-time electrical characteristic data are fused to construct a multi-dimensional input tensor. The multi-dimensional input tensor is fed into the fault diagnosis model, and the model output includes the fault probability distribution and the confidence evaluation result.
[0015] Preferably, the similar component screening component defines the fault similarity range for each target photovoltaic module, specifically: A preset tolerance window is expanded based on the fault type code of the target photovoltaic module to form the fault similarity range. Other photovoltaic modules with fault type codes falling within the range are retrieved from the database and included in the similar photovoltaic module set.
[0016] Compared with the prior art, the present application has the following advantages: The photovoltaic power station unmanned aerial vehicle inspection and AI fault diagnosis intelligent operation and maintenance system of the application utilizes an unmanned aerial vehicle to load high-resolution spectral sensors and current-voltage probes, synchronously collects multispectral reflectance data and electrical characteristic data, and records fault history information. The multispectral reflectance data provides the optical characteristics of the component surface and interior, the electrical characteristic data reflects the electrical performance, and the fault history information contains the time evolution law, and the three are complementary to each other, enhancing the depth and breadth of fault detection. Similar component screening components are targeted at target photovoltaic components, classified according to fault types, and similar component sets are selected from the same installation batch. This approach takes into account the manufacturing and installation background of the components, reduces the deviation caused by individual differences, and makes the diagnosis more in line with the actual conditions. For example, components in the same batch may experience similar environmental stresses, and comparing their performance can improve the rationality of diagnosis. Feature influence evaluation components calculate the fault response index of the target component in a specific spectral band based on the data distribution of the similar component set. The index dynamically reflects the sensitivity of the wave band to the fault, avoiding the rigid problem of global feature weight. The key feature selection engine further optimizes this process, calculates the baseline correlation score through the correlation measurement component, and combines the component unique code and the fault response index for enhancement, so that the correlation between the feature and the fault is more accurate. The noise influence estimation component calculates the environmental interference factor based on the similar component set and the global fault history distribution, quantifies the influence of external factors (such as shadows or dust) on spectral data, and balances fault correlation and noise suppression in feature selection. The key feature selection component integrates the environmental interference factor and the enhanced correlation score to identify the core spectral band set, effectively reduces the data dimension, highlights the key information, and reduces redundancy and computational burden.
[0017] The intelligent model construction component trains a deep neural network model using the multispectral reflectance data, electrical characteristic data, and failure history information corresponding to the core spectral band set. The deep neural network can learn complex nonlinear relationships in multi-source data and adapt to the diversity of failure modes. Due to the denoising and optimization in the feature selection stage, the model training is more efficient, the generalization ability is enhanced, and overfitting is avoided. The online diagnosis component inputs the real-time inspection data of the unmanned aerial vehicle into the trained model to quickly generate failure type prediction and maintenance recommendations. The entire process has high automation, seamlessly connects from data acquisition to diagnosis output, reduces manual intervention, and improves response speed. The system diagnosis accuracy is improved because the feature selection is targeted at specific component similar groups, the model focuses on relevant features, and false positives are reduced. For example, in the diagnosis of dirt failure, the system may preferentially select spectral bands sensitive to pollution and ignore interference bands caused by shadows. The operation and maintenance efficiency is significantly improved, the unmanned aerial vehicle has wide coverage and fast speed, and combined with real-time AI analysis, the detection period is shortened, which is suitable for regular inspection of large power stations. The system has strong adaptability and can adapt to changes in the power station environment and component aging trends through historical information learning and comparison of similar components. For example, as data accumulates, the model can be updated regularly to incorporate new failure modes and maintain diagnostic capabilities. In terms of cost-effectiveness, due to accurate diagnosis, unnecessary component replacement is reduced, maintenance resources are optimally allocated, component service life is extended, and power generation efficiency is ensured. The modular design of the system allows flexible upgrades, such as replacing sensors or algorithms, to adapt to technological development. Overall, the present application realizes the intelligent transformation of photovoltaic power station operation and maintenance through innovative integration of unmanned aerial vehicles, multispectral sensing, and artificial intelligence, providing a feasible solution for the industry. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings, which are incorporated in and form a part of the specification, illustrate one embodiment consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. It is appreciated that the accompanying drawings are only some embodiments of the present disclosure, and other drawings can be obtained from the accompanying drawings without creative labor for those skilled in the art. In the drawings: Figure 1 Performance evaluation chart for failure diagnosis model; Figure 2 Flowchart of feature influence evaluation component operation; Figure 3 Flowchart of noise influence estimation component calculating environmental disturbance factor; Figure 4 Accuracy and loss trend chart during model training process. DETAILED DESCRIPTION
[0019] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the following further describes the present disclosure in conjunction with accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, and not all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort belong to the scope of the present disclosure.
[0020] The terms used in the embodiments of the present disclosure are merely for the purpose of describing particular embodiments and are not intended to limit the present disclosure. The singular forms "a," "an," and "the" used in the embodiments of the present disclosure and the appended claims are intended to include plural forms as well, unless the context clearly indicates otherwise. "Multiple" generally includes at least two, and other quantifiers are similar.
[0021] It should be understood that, although the terms first, second, third, etc. can be used in the embodiments of the present disclosure to describe, these descriptions should not be limited to these terms. These terms are only used to distinguish the described objects. For example, without departing from the scope of the embodiments of the present disclosure, first can also be referred to as second, and similarly, second can also be referred to as first. In addition, the terms "first", "second", "third", etc. are only for the purpose of description, and cannot be understood as indicating or implying relative importance.
[0022] In the description of the present disclosure, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected", "connected" should be understood broadly, for example, can be fixedly connected, can be detachably connected, or integrally connected; can be mechanically connected, or electrically connected; can be directly connected, or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0023] It should also be noted that the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the goods or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such goods or devices. Without more limitations, the element defined by the sentence "including a" does not exclude the existence of other identical elements in the goods or devices including the element.
[0024] The following describes the optional embodiments of the present disclosure in conjunction with the accompanying drawings.
[0025] Please refer to Figures 1 to 4The application provides an unmanned aerial vehicle inspection and AI fault diagnosis intelligent operation and maintenance system for photovoltaic power stations, which comprises a data acquisition component, a similar component screening component, a feature influence evaluation component, a key feature selection engine, an intelligent model construction component and an online diagnosis component.
[0026] The data acquisition component uses high-resolution spectral sensors and current-voltage probes carried by the unmanned aerial vehicle to collect multispectral reflectance data and electrical characteristic data of photovoltaic components during the inspection process, and simultaneously retrieves fault history information of each component from the system database, including fault type, occurrence time and repair record, etc. The similar component screening component retrieves and selects a set of components with similar fault characteristics from the photovoltaic components of the same installation batch according to the fault type classification of the target photovoltaic component that needs to be diagnosed. The feature influence evaluation component calculates the fault response index of the target component on the specific spectral band based on the multispectral reflectance value distribution and fault history information distribution of the similar component set, which quantifies the sensitivity of spectral features to faults. The key feature selection engine selects a core spectral band set from all available spectral bands according to the fault response index, reducing the data dimension and highlighting the key features. The intelligent model construction component trains a fault diagnosis model based on deep neural network using the multispectral reflectance data, electrical characteristic data and fault history information corresponding to the core band, and the model can learn the complex mapping relationship between fault patterns and features. The online diagnosis component inputs the extracted core band reflectance values and electrical characteristic data into the trained model when processing real-time inspection data from the unmanned aerial vehicle, and the model outputs include fault type prediction probability and specific repair suggestions, thereby supporting operation and maintenance decisions.
[0027] Embodiment 1: The implementation process of the feature influence evaluation component involves multispectral data analysis and fault information processing of the similar component set of the target photovoltaic component to calculate the fault response index of the specific spectral band, which is the basis for subsequent key feature selection. For the selected target photovoltaic component, the feature influence evaluation component first calls the similar photovoltaic component set from the system database, which is determined by the similar component screening component, and this set contains a list of photovoltaic components that belong to the same installation batch as the target component and have similar fault history. The component then extracts the reflectance measurement values of each component in the similar set on the specific spectral band to be analyzed from the stored multispectral reflectance data set, and these measurement values are sorted according to the unique code of the component and form a numerical sequence. The extraction process of the reflectance value sequence includes data cleaning steps, such as removing abnormal values caused by instantaneous sensor failure, and using linear interpolation method to complete individual missing data points to ensure the integrity of the sequence.
[0028] After obtaining the reflectance value sequence, the feature influence evaluation component calculates the entropy value of the sequence as the feature uncertainty indicator. The entropy value is calculated based on the Shannon entropy concept in information theory, which measures the degree of confusion or uncertainty of a data sequence. The calculation process is to divide the value range of the reflectance value into several continuous intervals, and count the frequency of the value sequence falling into each interval, so as to approximately obtain the probability distribution of the reflectance value. According to the probability distribution, the feature influence evaluation component applies the Shannon entropy formula to calculate, and the higher the entropy value, the more dispersed and unpredictable the reflectance value distribution of the spectral band in the similar component set, and vice versa. This feature uncertainty indicator reflects the information content and stability of the band data itself. The feature influence evaluation component needs to process the fault history information to obtain the fault fluctuation reference. The component retrieves the historical fault records of the similar photovoltaic component set from the database, which contains the timestamp and fault type code of each fault occurrence. The component counts the total number of faults occurred in the similar set within a fixed time window, thereby generating a fault occurrence frequency sequence arranged in chronological order. The calculation of the fault fluctuation reference is realized by calculating the variance of the fault occurrence frequency sequence. The variance calculation reveals the fluctuation amplitude of the fault frequency in the time dimension, and a large variance value indicates that the time distribution of the fault occurrence is uneven, and there is obvious volatility; a small variance value indicates that the fault occurrence is relatively stable. This fault fluctuation reference quantifies the instability of the similar component set in the historical performance. The feature influence evaluation component needs to evaluate the correlation strength between the feature uncertainty indicator and the fault fluctuation reference, that is, to calculate the correlation weight. This step is completed by the mutual information quantity calculation in information theory. Mutual information measures the mutual dependence between two random variables, that is, how much uncertainty of the other variable can be reduced after knowing one variable. The component takes the feature uncertainty indicator sequence (a series of entropy values calculated by different spectral bands) and the fault fluctuation reference sequence (also corresponding to the variance values of different bands or different time scales) as input. The calculation of mutual information requires joint probability distribution and marginal probability distribution, and the component uses histogram method or kernel density estimation method to approximate these distributions, and then solves according to the definition formula of mutual information. The calculated mutual information value is the correlation weight, and the larger the weight value, the stronger the statistical correlation between the data uncertainty of the spectral band and the volatility of the fault occurrence.
[0029] Finally, the feature influence assessment component synthesizes the results from the above steps. It multiplies the calculated correlation weight with the previously obtained feature uncertainty indicator (i.e. the entropy value of the reflectance numerical sequence). The significance of this multiplication operation is to combine the data uncertainty of the waveband itself with its correlation degree with the fault fluctuation. The product is the fault response index of the target photovoltaic component at this specific spectral waveband. A higher fault response index means that not only the data itself of this waveband has a higher uncertainty (high entropy value), but also this uncertainty is closely related to the fluctuation of fault occurrence (high mutual information), therefore this waveband may have a higher indicative value for diagnosing the fault of the target component. The feature influence assessment component repeats the entire process above for all spectral wavebands that need to be analyzed for the target component, thus generating a corresponding fault response index for each waveband. These indices form an index list, providing a direct quantitative basis for the subsequent engine to select the core spectral wavebands. The entire calculation process is designed to be parallelizable, because the analysis of different spectral wavebands is independent of each other, which is conducive to using distributed computing frameworks such as Apache Spark to accelerate the processing of large-scale photovoltaic power station data and improve system response speed. Standardized data formats are used when data is transmitted between various links to ensure interface compatibility and calculation accuracy. The output of the feature influence assessment component is a structured data file that explicitly records the target component code, the list of analyzed wavebands, and the fault response index value corresponding to each waveband.
[0030] Example 2: The implementation of the key feature selection engine relies on the coordinated work of its two core components: the correlation measurement component is responsible for quantifying the association strength between spectral features and faults, and the noise influence estimation component is responsible for assessing the disturbance level of environmental factors on measurement data. After the correlation measurement component is started, its first task is to calculate the baseline correlation score of a specific spectral waveband, which begins with statistical analysis of the global data set. The component obtains the multispectral reflectance values of all photovoltaic components at a specific spectral waveband from the system database, and these values have been preprocessed to eliminate obvious acquisition errors. At the same time, the component calls the fault history information of all photovoltaic components, which is stored in a structured format and contains fault types, occurrence times, and severity levels. The correlation measurement component sorts the reflectance value sequence and calculates its empirical cumulative distribution function, generating a smooth cumulative distribution sequence that depicts the change in cumulative probability of reflectance values from the minimum to the maximum. Similarly, the component quantitatively processes the fault history information, such as mapping different fault types to numerical scores and calculating their cumulative distribution sequences.
[0031] After obtaining the two cumulative distribution sequences, the correlation measure component calculates the Hellinger distance between them. The Hellinger distance is a statistical tool for measuring the difference between two probability distributions, which is based on the square root integral of the difference between the distribution functions. In the discrete data environment, the component uses a numerical approximation method to subdivide the value interval of the cumulative distribution sequences, calculate the sum of the squares of the difference between the values of the two sequences in each small interval, and then take the square root. The result of the Hellinger distance calculation is a non-negative value, and the smaller the value, the more similar the two distributions, that is, the higher the consistency of the reflectance distribution of the spectral band with the fault information distribution. Subsequently, the component performs a linear transformation on the calculated Hellinger distance to convert it into a standardized score between zero and one. The parameters of the linear transformation, such as the slope and intercept, are pre-set according to historical data or domain knowledge to ensure the comparability of the score. The result of the transformation is the baseline correlation score of the spectral band, which preliminarily reflects the overall statistical correlation between the band reflectance and the component failure, and the higher the score, the stronger the potential correlation.
[0032] The execution of the noise influence estimation component is parallel or sequential to the correlation measure component, and its goal is to evaluate the degree of influence of environmental noise on a specific spectral band, and the calculation result is an environmental interference factor. The noise influence estimation component obtains the unique code list of the similar photovoltaic component set from the similar component filtering result of the current target photovoltaic component. According to this list, the component extracts the failure history information of all members of the similar set from the database, including detailed failure record time series. The component uses kernel density estimation method to construct the failure history information distribution density sequence of the similar photovoltaic component set, taking the failure occurrence time or failure type value as the variable. The kernel density estimation process needs to select an appropriate kernel function, such as the Gaussian kernel function, and determine the bandwidth parameter, in order to generate a smooth and continuous probability density curve to represent the distribution characteristics of the failure in the similar component. The noise influence estimation component obtains the failure history information of all photovoltaic components and uses the same kernel density estimation method and parameters to construct the global failure history information distribution density sequence. The global distribution density sequence reflects the background distribution of the failure mode of the entire photovoltaic power station. Next, the component calculates the Jensen-Shannon divergence between the failure history information distribution density sequence of the similar component set and the failure history information distribution density sequence of all photovoltaic components. The Jensen-Shannon divergence is a measure of the difference between two probability distributions, which is based on the concept of information entropy, and calculates the difference between the average relative entropy of the two distributions and the entropy of their mixed distribution. The calculation process involves operating on the values of the two density sequences at discrete points to obtain a scalar value representing the difference between the distributions. The larger the Jensen-Shannon divergence value, the more significant the difference between the failure mode of the similar component set and the global background, which may be due to local environmental noise factors.
[0033] To quantify the degree of noise influence, the noise influence estimation component further processes the Jensen-Shannon divergence values. The component calculates the average Jensen-Shannon divergence of all photovoltaic components over multiple spectral bands or historical time periods as a reference baseline. Then, the component inversely scales the calculated Jensen-Shannon divergence values for a particular spectral band relative to this average baseline. The inverse scaling is typically done by taking the inverse or using a formula of the form "baseline value / current value" to make the resulting environmental disturbance factor inversely proportional to the noise influence, i.e., the smaller the environmental disturbance factor value, the greater the disturbance of the band by environmental noise. This environmental disturbance factor provides an important de-noising basis for subsequent feature selection. After deriving the baseline correlation score, the correlation measure component does not immediately output, but enters an optimization phase. The optimization process requires the introduction of the unique codes of the photovoltaic components and the pre-computed failure response indices by the feature influence assessment component. The component ranks all photovoltaic components in descending order according to the failure severity level, which is assigned according to a predefined rule base, generating an ordered failure level sequence. At the same time, the component ranks all photovoltaic components in ascending order according to the multispectral reflectance values of the particular spectral band currently being analyzed, generating a feature value sequence. The elements in both sequences correspond to the unique codes of the photovoltaic components. Then the Kendall's tau coefficient of the failure level sequence and the feature value sequence is calculated. Kendall's tau coefficient is a non-parametric statistical measure used to assess the consistency of the rank order of two sequences. The calculation involves checking all possible pairs of photovoltaic components, counting the number of pairs that are in consistent order and the number of pairs that are in inconsistent order in the two sequences, and finally deriving a coefficient value between -1 and 1, which serves as a measure of sequence consistency. A coefficient value close to 1 indicates a high consistency of the ordering of the two sequences. Subsequently, the component performs a removal operation, removing those unique codes of photovoltaic components from the failure level sequence and the feature value sequence that have exactly matching rank positions in both sequences. This operation aims to filter out cases where the association is obvious, thereby focusing on a subset of components with inconsistent ordering that are more valuable for analysis, forming a remaining failure level sequence and a remaining feature value sequence. Each unique code in the remaining failure level sequence is replaced by the failure response index corresponding to that component, thereby generating a new failure response sequence. Similarly, each unique code in the remaining feature value sequence is replaced by the failure response index corresponding to that component, generating a feature response sequence. The failure response index is an indicator of the sensitivity of the component to failure, output by the feature influence assessment component. Then, the component calculates the Pearson correlation coefficient between the two new sequences, the failure response sequence and the feature response sequence. The Pearson correlation coefficient measures the degree of linear correlation between two sequences, and the calculation result is a value between -1 and 1, which serves as a measure of content consistency. The higher the coefficient value, the more highly correlated the change pattern of the failure response index even in the subset of components with inconsistent ordering.
[0034] The correlation measure component calculates a comprehensive optimization weight based on the sequence consistency measure (Kendall's tau) and the content consistency measure (Pearson's correlation coefficient). The calculation of the optimization weight can be a weighted average of the two coefficients, or other combinations, with the purpose of fusing the information of both sequence ranking and content correlation. The calculated optimization weight is then multiplied with the originally obtained baseline correlation score to produce the final enhanced correlation score. This enhanced correlation score not only contains the global statistical correlation information between spectral reflectance and faults, but also incorporates the local consistency check results based on individual component fault response index, making the evaluation of feature importance more comprehensive and robust. The key feature selection engine delivers the enhanced correlation score and the environmental disturbance factor as key inputs to the internal key feature selection component for the final decision.
[0035] Embodiment 3: The operation flow of the key feature selection component starts with receiving the enhanced correlation score from the correlation measure component and the environmental disturbance factor from the noise influence estimation component, which provide the quantitative evaluation of feature importance and noise influence for each spectral band. The key feature selection component performs a calculation for each candidate spectral band, dividing the enhanced correlation score by the environmental disturbance factor to obtain an initial priority quotient value for each band. The significance of the division operation is to measure the net correlation strength of the spectral band after excluding the influence of environmental noise. The higher the enhanced correlation score and the larger the environmental disturbance factor, the higher the quotient value of the band, indicating that the band has more significant feature importance. The calculation of the initial priority quotient value needs to handle the case where the divisor is zero, and a small positive number is preset as the lower limit value of the environmental disturbance factor to ensure numerical stability.
[0036] After obtaining the initial priority quotient of all spectral bands, the key feature selection component performs min-max normalization to map the quotient into a closed interval of zero to one, forming the normalized selection priority score. The normalization is based on the maximum and minimum of all initial priority quotients in the current computation cycle, and the conversion formula is: the selection priority score equals to the initial priority quotient of a certain band minus the minimum initial priority quotient among all bands, and the resulting difference is divided by the difference between the maximum and minimum initial priority quotients among all bands. The selection priority score intuitively reflects the relative importance ranking of each band relative to other bands, and the closer the score is to one, the more likely the band is to be selected into the core set. The determination of the dynamic threshold is the core decision-making link of the key feature selection component, and the threshold is used to screen the core spectral band set from all bands. The dynamic threshold is not a fixed value, but is dynamically generated according to the selection priority score distribution of all photovoltaic component data in the current batch. The component calculates the statistical characteristics of all band selection priority scores, such as the median, mean or specific quantile, and sets one of them as the threshold. A typical strategy is to select the seventy-fifth quantile of the selection priority score as the dynamic threshold, which means that only the bands with priority scores higher than the overall level of seventy-five percent will be retained. The dynamic threshold mechanism enables the feature selection to adapt to the data characteristics of different photovoltaic power stations or different inspection batches, avoiding selection bias due to differences in data distribution.
[0037] The screening process is to compare the selection priority score of each spectral band with the dynamic threshold, and the spectral band with a score higher than the threshold is determined as a key feature, and its unique identifier is added to the core spectral band set. The core spectral band set is an ordered list, usually sorted in descending order of selection priority score, to facilitate the subsequent model building component to use the most important features first. The key feature selection component finally outputs this core spectral band set as the input of the intelligent model building component.
[0038] After receiving the core spectral band set, the intelligent model building component starts the long short-term memory network-based fault diagnosis model training process. The long short-term memory network is a special recurrent neural network structure that can effectively learn and capture long-term dependencies in time series data, and is very suitable for processing photovoltaic component electrical characteristic data and multi-spectral reflectance data with time series characteristics. The model building component first constructs the input features, extracts the multi-spectral reflectance data corresponding to the core spectral band set from the original data set, and these data are usually collected in time order. At the same time, the component extracts the synchronous recorded electrical characteristic data, such as the time series measurement values of current and voltage. The reflectance data and electrical characteristic data are aligned and spliced to form a multi-dimensional input tensor, and the dimensions of the tensor include the number of samples, the time step and the feature dimension.
[0039] The preparation of fault labels is the basis of model training, and the intelligent model construction component encodes the fault categories of photovoltaic components. The fault categories come from historical maintenance records and expert diagnosis results. The component uses one-hot encoding to convert discrete fault types into binary vectors, for example, normal state is encoded as [1, 0, 0], hot spot fault is encoded as [0, 1, 0], and crack fault is encoded as [0, 0, 1]. The one-hot encoding vector is used as the target output of model training, so that the neural network can learn the mapping relationship from the input features to the fault probability. The architecture design of the long short-term memory network model includes an input layer, one or more long short-term memory network hidden layers, and an output layer. The number of input layer neurons matches the feature dimension, and receives a multi-dimensional input tensor. Each long short-term memory network hidden layer contains complex gating mechanisms, including input gate, forget gate and output gate. These gating structures work together through sigmoid function and tanh function to control the flow, memory and forgetting of information. The number of hidden layers and time steps are configured according to the complexity of the data and the computing resources. Deep network can learn more abstract feature representation. The output layer usually uses softmax activation function to convert the output of the hidden layer into the predicted probability of each fault category, and the sum of the probabilities of all categories is one. The model training process uses the adaptive moment estimation algorithm for optimization, which dynamically adjusts the learning rate of each parameter by calculating the first and second moment estimates of the parameters. The training goal is to minimize the difference between the predicted probability and the true label, and the loss function selects the classification cross-entropy loss commonly used in classification tasks. The training process is performed in multiple rounds, each round includes four steps: forward propagation to calculate the predicted value, calculation of the loss function, backward propagation to calculate the gradient, and update of network parameters using the adaptive moment estimation algorithm. In order to prevent the model from overfitting on the training data, the intelligent model construction component introduces regularization techniques such as dropout layer, and monitors the performance on the validation set during training. When the validation set loss no longer decreases, the training is terminated in advance. The trained fault diagnosis model is finally serialized and saved for the online diagnosis component to load. The entire model construction process is executed in a computing environment with graphics processor acceleration to meet the computing needs of large-scale data training.
[0040] Referring to Figure 4The figure clearly presents the training dynamic process of the photovoltaic power station AI fault diagnosis model, including the two core dimensions of accuracy change and loss value change. The upper subgraph shows the accuracy trend of the training set and the validation set: the training set accuracy gradually rises with the increase of training rounds, and stabilizes above 0.9 in the later period, reflecting the continuous strengthening of the model's fitting ability to the training data; the accuracy of the validation set fluctuates slightly, but overall shows an upward trend, indicating that the model can better generalize to unseen data and avoid overfitting. The lower subgraph shows the loss value trend of the training set and the validation set: the training set loss decreases rapidly in the early stage and maintains at a very low level in the later stage; the validation set loss also gradually decreases and stabilizes. Both of them finally converge to a low level, indicating that the model has effectively optimized the network parameters through the self-adaptive matrix estimation algorithm, and the loss function has successfully converged, proving that the model has fully learned the complex mapping relationship between "core spectral band reflectance + electrical characteristic data" and "photovoltaic module fault category", providing reliable model support for subsequent online diagnosis of photovoltaic module faults.
[0041] Example 4: The process of the online diagnosis component processing real-time data starts with receiving real-time streaming data packets sent by the UAV through a wireless data transmission link. The data packets contain raw reflectance data collected by the high-resolution spectral sensor and electrical parameters measured by the current-voltage probe. The data parsing module first unpacks and checks the data packets, verifying data integrity and timeliness, and discarding damaged data frames caused by transmission errors. The parsed data enters the core spectral band filtering link, and the component loads the core spectral band set configuration file generated by the key feature selection engine in advance. This file stores the selected feature band numbers in a list format. The data filtering algorithm accurately extracts the corresponding band values from the complete reflectance spectral data based on the band numbers, forming a core spectral reflectance subset. At the same time, the current-voltage measurement values are processed through signal conditioning and analog-to-digital conversion, and are strictly aligned with the reflectance data through timestamps, ensuring that the optical and electrical parameters collected at the same time can be correctly associated. The construction of multi-dimensional input tensors requires the fusion of feature data of different dimensions. The core spectral reflectance values form the static part of the feature vector, representing the optical characteristics of the component at the current time. The time series electrical characteristic data need to retain their dynamic change information, and the component maintains a fixed-length sliding window to continuously store the current and historical electrical parameters for several sampling periods. For each real-time data point, the current and historical electrical parameters are extracted from the sliding window, and statistical features such as average current and voltage fluctuation rate are calculated. The reflectance subset and electrical feature vector are concatenated into a complete multi-dimensional input tensor, with dimensions designed as (batch size, time step, feature number). The batch size is set to 1 when processing a single real-time sample, the time step is determined by the size of the electrical characteristic sliding window, and the feature number is the sum of the core spectral band number and the electrical feature dimension. After the tensor is constructed, it is standardized using the mean and standard deviation parameters saved during the training phase to normalize each feature dimension, ensuring that the input data distribution is consistent with the model training data. The inference execution of the fault diagnosis model feeds the preprocessed multi-dimensional input tensor into the trained deep neural network model loaded into memory. The forward propagation process of the model passes through the input layer, multiple hidden layers, and the output layer in sequence. The long short-term memory network units in the hidden layer encode the time series electrical characteristics and capture their long-term dependencies; the fully connected layer processes the static spectral reflectance features. The network finally generates a fault probability distribution vector through the softmax activation function in the output layer, with each element in the vector corresponding to the prediction probability of a predefined fault class, and the sum of all elements being 1. The model also outputs a confidence assessment result, which quantifies the credibility of this diagnosis by calculating the entropy value or the highest probability value of the prediction probability distribution. The probability distribution vector and the confidence value are packaged into a structured diagnosis result. The post-processing module of the diagnosis result converts the numerical result into an operable maintenance recommendation.The fault type with the highest probability in the probability distribution of the fault is determined as the final diagnosis result, and the system maps the fault code to a specific maintenance operation description according to the preset rule base, such as "suggest cleaning the component surface" or "there may be internal cracks, suggest detailed inspection". The confidence evaluation result is recorded together with the diagnosis result, and a low confidence diagnosis will trigger a system flag to prompt the operation and maintenance personnel to manually review. The final generated diagnosis report contains component number, detection time, fault type, maintenance suggestion and confidence level, which is displayed in real time through the human-computer interaction interface and stored in the historical database. Similar component screening component takes the fault type code of the target photovoltaic component as input, which comes from the standardized fault classification system in the power station asset management system. A tolerance window parameter is preset in the component, which defines the similarity judgment range of the fault type. The setting of the tolerance window is based on the physical characteristics of the fault type and historical data analysis, and the window size can be differentiated for fault types of different severity or different nature. The following table shows the correspondence between fault type code and tolerance window setting:
[0042] Similar component screening component queries the corresponding tolerance window value according to the fault type code of the target component from the above table. Then, the component calculates the fault similarity range, the lower bound of which is the target code minus the tolerance window value, and the upper bound of which is the target code plus the tolerance window value. The component constructs a database query statement, and sets the retrieval condition as: the component installation batch is consistent with the target component, and the fault type code in its historical fault record falls within the calculated similarity range. The database query is executed in the asset performance management system of the power station, and returns a list of unique identifiers of all photovoltaic components that meet the conditions, which constitutes the similar photovoltaic component set of the target component. The query process considers the time context of fault occurrence, and prioritizes components that have recently failed to ensure the timeliness of the set. The returned similar component set information is cached and associated with the target component identifier for subsequent use by the feature influence evaluation component. The entire screening process ensures that the component group used for comparative analysis is comparable in terms of fault nature, providing a reliable data foundation for subsequent feature analysis and model diagnosis. The online diagnosis component and the similar component screening component work together to form a closed loop from real-time data acquisition to fault diagnosis decision-making, realizing the intelligentization and automation of photovoltaic power station operation and maintenance.
[0043] The above examples are only used to illustrate the technical solutions of the present disclosure, rather than limit them; although the present disclosure has been described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. A drone inspection and AI fault diagnosis intelligent operation and maintenance system for photovoltaic power plants, characterized in that, The system includes: The data acquisition component uses a high-resolution spectral sensor and current and voltage probe mounted on a drone to collect multispectral reflectance data and electrical characteristic data of photovoltaic modules, and simultaneously records the fault history information of each photovoltaic module. Similar component screening: For the current target photovoltaic module, a set of similar photovoltaic modules is selected from the photovoltaic modules in the same installation batch based on their failure type. The feature impact assessment component calculates the fault response index of the target photovoltaic module in a specific spectral band based on the multispectral reflectance value distribution and fault history information distribution of a similar photovoltaic module set to the target photovoltaic module. The key feature selection engine identifies the core spectral band set from all spectral bands based on the fault response index. The intelligent model building component uses multispectral reflectance data, electrical characteristic data, and fault history information corresponding to the core spectral band set to train a fault diagnosis model based on a deep neural network. The online diagnostic component inputs data acquired in real-time by drone inspections into a trained fault diagnosis model to generate fault type predictions and maintenance suggestions for photovoltaic modules.
2. The intelligent operation and maintenance system for photovoltaic power plants using drone inspection and AI fault diagnosis as described in claim 1, characterized in that, The key feature selection engine further includes: The correlation measurement component calculates the baseline correlation score for a specific spectral band based on the distribution of multispectral reflectance values and fault history information of all photovoltaic modules. It then optimizes the baseline correlation score by combining the unique code of the photovoltaic module and the fault response index to obtain the enhanced correlation score. The noise impact estimation component calculates the environmental interference factor for a specific spectral band based on the distribution of fault history information of a set of similar photovoltaic modules of the target photovoltaic module and the distribution of fault history information of all photovoltaic modules. The key feature selection component identifies the core spectral band set from all spectral bands based on environmental interference factors and enhanced correlation scores.
3. The intelligent operation and maintenance system for photovoltaic power plants using drone inspection and AI fault diagnosis as described in claim 2, characterized in that, The specific operations of the feature influence assessment component include: For a target photovoltaic module, extract the reflectance numerical sequence of its similar photovoltaic module set in a specific spectral band, and calculate the entropy value of the numerical sequence as a characteristic uncertainty index; The variance of the fault occurrence frequency sequence of similar photovoltaic module sets is obtained as a reference for fault fluctuation; The correlation weight between the characteristic uncertainty index sequence and the fault fluctuation reference sequence is calculated, and this weight is obtained by mutual information in information theory. By multiplying the correlation weight by the characteristic uncertainty index, the fault response index of the target photovoltaic module in a specific spectral band is obtained.
4. The intelligent operation and maintenance system for photovoltaic power plants using drone inspection and AI fault diagnosis as described in claim 3, characterized in that, The correlation measurement component performs the following when calculating the baseline correlation score: Construct the cumulative distribution sequence of multispectral reflectance values and the cumulative distribution sequence of fault history information for all photovoltaic modules; Calculate the Hellinger distance between the cumulative distribution sequence of multispectral reflectance values and the cumulative distribution sequence of fault history information; A linear transformation mapping of the Hellinger distance yields the baseline correlation score for a specific spectral band.
5. The intelligent operation and maintenance system for photovoltaic power plants using drone inspection and AI fault diagnosis as described in claim 4, characterized in that, The correlation metric component performs the following when optimizing the baseline correlation score: All photovoltaic modules are sorted in descending order according to the severity of the fault, generating a fault level sequence; All photovoltaic modules are sorted in ascending order according to their multispectral reflectance values to generate a sequence of characteristic values. Calculate the Kendall's harmony coefficient between the fault level sequence and the eigenvalue sequence as a measure of sequence consistency; The unique codes of photovoltaic modules with matching locations are removed from the fault level sequence and the feature value sequence to obtain the remaining fault level sequence and the remaining feature value sequence. Each unique code in the remaining fault level sequence is replaced with the fault response index of the corresponding photovoltaic module to form a fault response sequence; Each unique code in the remaining feature value sequence is replaced with the corresponding fault response index of the photovoltaic module to form a feature response sequence; Calculate the Pearson correlation coefficient between the fault response sequence and the characteristic response sequence as a measure of content consistency; The optimized weights are calculated based on sequence consistency and content consistency metrics. The optimized weights are then multiplied by the baseline relevance score to obtain the enhanced relevance score.
6. The intelligent operation and maintenance system for photovoltaic power plants using drone inspection and AI fault diagnosis as described in claim 5, characterized in that, The noise impact estimation component calculates the environmental disturbance factor in the following manner: For a target photovoltaic module, obtain the distribution density sequence of its fault history information for a set of similar photovoltaic modules; Obtain the distribution density sequence of all photovoltaic modules' fault history information; Calculate the Jensen-Shannon divergence between the fault history information distribution density sequence of a similar photovoltaic module set and the fault history information distribution density sequence of all photovoltaic modules; By inversely scaling the average Jensen-Shannon divergence of all photovoltaic modules, an environmental interference factor for a specific spectral band is obtained.
7. The intelligent operation and maintenance system for photovoltaic power plants using drone inspection and AI fault diagnosis as described in claim 6, characterized in that, The key feature selection component operation includes: For each spectral band, calculate the quotient of the enhanced correlation score to the environmental interference factor; The quotient is subjected to minimum-maximum normalization to obtain the selection priority score of the spectral band. Spectral bands with priority scores higher than the dynamic threshold are selected to form a core spectral band set.
8. The intelligent operation and maintenance system for photovoltaic power plants using drone inspection and AI fault diagnosis as described in claim 7, characterized in that, The intelligent model building component uses a long short-term memory network architecture for model training. The multispectral reflectance data and time-series electrical property data corresponding to the core spectral band set are used as input feature vectors. The fault category code of the photovoltaic module is used as the output label; By iteratively updating the network parameters using an adaptive moment estimation algorithm, the loss function converges, thus completing the construction of the fault diagnosis model.
9. The intelligent operation and maintenance system for photovoltaic power plants using drone inspection and AI fault diagnosis as described in claim 8, characterized in that, The online diagnostic component processes real-time data including: Extract reflectance values of the core spectral band set from real-time streaming data from drones; By integrating real-time electrical characteristic data, a multidimensional input tensor is constructed. The multidimensional input tensor is fed into the fault diagnosis model, and the model output includes the fault probability distribution and confidence assessment results.
10. The intelligent operation and maintenance system for photovoltaic power plants using drone inspection and AI fault diagnosis as described in claim 9, characterized in that, The similar component screening component defines a fault similarity range for each target photovoltaic module, specifically: Based on the fault type coding of the target photovoltaic module, a preset tolerance window is extended to form a fault similarity range; Retrieve other photovoltaic modules whose fault type codes fall within this range from the database and include them in a set of similar photovoltaic modules.
Citation Information
Patent Citations
Distributed photovoltaic power station unmanned aerial vehicle inspection method and system
CN119597003A
Photovoltaic module damage intelligent detection method and system based on unmanned aerial vehicle image
CN120235872A
Photovoltaic power station unmanned aerial vehicle automatic inspection system integrated with machine vision
CN120908187A
Distributed photovoltaic diagnosis method and system based on multi-modal time sequence data
CN121093019A
Monitoring system of a photovoltaic plant, monitoring method of a photovoltaic plant and photovoltaic plant
EP4482025A2