A photovoltaic module fault recognition method, device and equipment
Patent Information
- Application Number
- CN202611140491.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-30
AI Technical Summary
[0005]本说明书实施例的目的是提供一种光伏组件故障识别方法、装置及设备,以克服现有方法中存在的故障识别精度差强人意的问题
[0022]由以上本说明书实施例提供的技术方案可见,本说明书实施例可以将无人机采集的光伏组件热红外图像输入视觉特征提取器,得到全局类别特征和局部块特征矩阵;对局部块特征矩阵中的多个图像块特征进行空间上下文编码,得到全局故障震源特征;所述全局故障震源特征包括热红外图像多个空间位置的全局故障指示信息;对各图像块特征进行门控映射,得到局部响应矩阵;所述局部响应矩阵包括用于表征热红外图像中各空间位置对全局故障指示信息吸收强度的多个局部响应系数;耦合全局故障震源特征与局部响应矩阵,得到传播增强特征;根据传播增强特征和阻尼系数,确定残差增量特征;所述阻尼系数用于表征对热红外图像背景成分的抑制程度;根据残差增量特征,修正局部块特征矩阵;根据全局类别特征和修正后的局部块特征矩阵,使用预设分类器确定光伏组件的故障类别。通过对局部块特征进行空间上下文编码得到全局故障震源特征,并结合局部响应矩阵进行耦合传播,使全局故障语义能够定向传递至各局部空间位置,有效解决了故障信号淹没于大面积背景的问题。进一步地,通过阻尼系数对背景成分进行抑制,进一步突出故障相关特征,提升了对微弱热异常的感知灵敏度。以及基于残差增量特征修正局部块特征矩阵,在不破坏原始特征结构的前提下实现了故障语义的增量式增强,保留了基础模型的原始判别能力。最终通过全局与局部的融合分类,实现了对光伏组件故障类别的准确识别。
Smart Images

Figure CN122634410B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the fields of image processing and fault identification technology, specifically to a method, apparatus, and device for identifying faults in photovoltaic modules. Background Technology
[0002] With the continued growth in global demand for clean energy and the expanding installed capacity of photovoltaic power plants, the need for operation and maintenance (O&M) testing is becoming increasingly urgent. During long-term outdoor operation, photovoltaic modules are susceptible to faults such as hot spots, bypass diode failure, sub-string open circuits, and battery overheating due to manufacturing defects, environmental stress, and partial shading. Failure to detect these faults in a timely manner can lead to decreased power generation efficiency and even safety accidents. Traditional manual inspection methods are inefficient and costly, making them unsuitable for the O&M needs of large-scale power plants. Automated inspection methods using drones equipped with thermal infrared cameras have become the mainstream technology for fault detection in photovoltaic power plants due to their advantages of low cost, wide coverage, and fast response.
[0003] At the fault identification algorithm level, existing methods mainly fall into two categories. Early methods relied on manual feature extraction and rule-based threshold judgment, which lacked robustness to environmental changes. In recent years, deep learning methods based on convolutional neural networks and object detection frameworks have gradually become mainstream, achieving high recognition accuracy under specific site and imaging conditions. However, these methods are usually supervised training on limited data collected from a single site, resulting in limited model generalization ability. Their performance degrades significantly in real-world deployment scenarios such as cross-site, cross-climate, and cross-flight altitude scenarios, making it difficult to support the operation and maintenance needs of large-scale heterogeneous photovoltaic power plants.
[0004] Furthermore, the design goal of the visual foundation model is to extract global semantic representations, aggregating information from the entire image. In photovoltaic thermal infrared images, fault areas often occupy only a small portion of the total image area, and the temperature anomalies caused by the fault are easily masked at the global feature level, resulting in unsatisfactory fault identification accuracy. Summary of the Invention
[0005] The purpose of the embodiments in this specification is to provide a method, apparatus, and device for identifying photovoltaic module faults, so as to overcome the problem of unsatisfactory fault identification accuracy in existing methods.
[0006] To address the aforementioned technical problems, this specification provides a photovoltaic module fault identification method, comprising: The thermal infrared images of photovoltaic modules collected by the drone are input into the visual feature extractor to obtain global category features and local block feature matrices; Spatial context encoding is performed on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image; Gated mapping is performed on the features of each image block to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image; By coupling the global fault source characteristics with the local response matrix, propagation enhancement characteristics are obtained; The residual increment characteristics are determined based on the propagation enhancement characteristics and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of background components in thermal infrared images. Based on the residual increment characteristics, the local block feature matrix is corrected; Based on the global category features and the corrected local block feature matrix, a preset classifier is used to determine the fault category of the photovoltaic module.
[0007] Furthermore, the step of inputting the thermal infrared image of the photovoltaic module collected by the UAV into the visual feature extractor to obtain a global category feature and a local block feature matrix includes: The thermal infrared image is divided into multiple image blocks; Each image patch is mapped to a feature vector, and the position encoding of each feature vector is superimposed to obtain the image patch feature sequence; A learnable category feature vector is inserted at the beginning of the image patch feature sequence to obtain the input sequence; The input sequence is fed into the multi-layer transform encoder of the visual feature extractor. Each layer transform encoder is used to transform the input sequence layer by layer through a multi-head self-attention and feedforward network. The head vector and the tail vector output by the last encoder are used as the global category feature and the local block feature matrix, respectively.
[0008] Furthermore, the step of performing spatial context encoding on multiple image block features in the local block feature matrix to obtain global fault source features includes: Global mean pooling is performed on multiple image block features along the dimension of the local block feature matrix to obtain fault aggregation features; The first linear projection transformation is performed on the fault aggregation features to obtain the global fault source features.
[0009] Furthermore, the gating mapping of features for each image patch to obtain a local response matrix includes: A second linear projection transformation is performed on the features of each image patch to obtain the initial gated response matrix; The initial gated response matrix is activated to obtain the local response matrix.
[0010] Furthermore, the method also includes: Normalize the feature matrix of the local block; The process of spatial context encoding multiple image block features in the local block feature matrix to obtain global fault source features includes: Spatial context encoding is performed on the normalized features of multiple image patches to obtain global fault source features; The process of gating and mapping the features of each image patch to obtain the local response matrix includes: Gated mapping is performed on the normalized features of each image patch to obtain the local response matrix.
[0011] Furthermore, the coupling of global fault source features and local response matrix yields propagation enhancement features, including: The global fault source features are extended to the same spatial dimension as the features of each image patch to obtain broadcast source features; The propagation increment features of each image patch location are obtained by multiplying the broadcast source features element-wise with the local response matrix. The propagation increment features are superimposed on the features of each image patch to obtain the propagation enhancement features.
[0012] Furthermore, the step of extending the global fault source features to the same spatial dimension as the features of each image patch to obtain broadcast source features includes: Based on the semantic correlation between the features of each image patch and the global fault source features, the correlation coefficient of each image patch location is obtained; the semantic correlation is used to characterize the degree of matching between the local fault indication information carried by each image patch feature and the global fault indication information. Based on the correlation coefficient, the propagation modulation coefficient of each image block location is determined; the propagation modulation coefficient is positively correlated with the correlation coefficient; the propagation modulation coefficient is used to characterize the intensity scaling ratio when the global fault source features are transmitted to each image block location. The broadcast source characteristics are obtained by multiplying the global fault source characteristics element by element with the propagation modulation coefficients of each image block location.
[0013] Further, determining the residual increment characteristics based on the propagation enhancement characteristics and damping coefficient includes: The background suppression features are obtained by multiplying the normalized features of each image patch by the damping coefficient. Subtracting the background suppression feature from the propagation enhancement feature yields the residual increment feature.
[0014] Furthermore, the step of multiplying the normalized features of each image patch by the damping coefficient to obtain the background suppression features includes: Context-aware projection is performed on the normalized features of each image patch to obtain the feature activity of each image patch feature. Based on feature activity, the spatial modulation coefficients of each image block feature are determined; the spatial modulation coefficients are positively correlated with feature activity; the spatial modulation coefficients are used to characterize the local adjustment intensity of the damping coefficient. By fusing the spatial modulation coefficients and damping coefficients of each image patch feature, the modulation damping coefficients of each image patch feature are obtained. The background suppression features are obtained by multiplying the normalized features of each image patch by the corresponding modulation damping coefficient.
[0015] Furthermore, the step of performing context-aware projection on the normalized image patch features to obtain the feature activity of each image patch feature includes: For each image patch location, extract features from multiple neighboring image patches within a preset neighborhood centered on that location to obtain a local context feature matrix; The local context feature matrix is subjected to a third linear projection transformation to obtain the local context features of the image patch location; A fourth linear projection transformation is performed on the image patch features corresponding to the image patch location to obtain the self-projection features of the image patch location. By fusing local context features and self-projection features, the neighborhood aggregation features of the image patch location are obtained; A nonlinear projection transformation is performed on the neighborhood aggregation features of each image patch location to obtain the feature activity of each image patch feature.
[0016] Further, the step of correcting the local block feature matrix based on the residual increment features includes: The residual increment feature is added element by element to the local block feature matrix to obtain the corrected local block feature matrix.
[0017] Further, the step of determining the fault category of the photovoltaic module using a preset classifier based on global category features and the corrected local block feature matrix includes: The global category features are input into the first classifier to obtain the global fault identification result; The corrected local block feature matrix is input into the second classifier to obtain the local fault identification result; The fault category of the photovoltaic module is determined by fusing the global fault identification results and the local fault identification results.
[0018] Furthermore, embodiments of this specification provide a photovoltaic module fault identification device, comprising: The extraction module is used to input the thermal infrared images of photovoltaic modules collected by the UAV into the visual feature extractor to obtain global category features and local block feature matrices; The encoding module is used to perform spatial context encoding on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image; The mapping module is used to perform gated mapping on the features of each image block to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image; The coupling module is used to couple the global fault source characteristics with the local response matrix to obtain propagation enhancement characteristics; The first determining module is used to determine the residual increment characteristics based on the propagation enhancement characteristics and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of the background components of the thermal infrared image. The correction module is used to correct the local block feature matrix based on the residual increment characteristics; The second determination module is used to determine the fault category of the photovoltaic module using a preset classifier based on the global category features and the corrected local block feature matrix.
[0019] Furthermore, embodiments of this specification provide a computer device, including: Memory, used to store computer programs; A processor is used to execute the computer program to implement the above-described photovoltaic module fault identification method.
[0020] Furthermore, embodiments of this specification provide a computer storage medium storing computer program instructions, which, when executed, implement the aforementioned photovoltaic module fault identification method.
[0021] Furthermore, embodiments of this specification provide a computer program product, including a computer program that, when executed by a processor, implements the aforementioned photovoltaic module fault identification method.
[0022] As can be seen from the technical solutions provided in the embodiments of this specification above, the embodiments of this specification can input the thermal infrared images of photovoltaic modules collected by UAVs into a visual feature extractor to obtain global category features and local block feature matrices; perform spatial context encoding on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image; perform gating mapping on each image block feature to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image; couple the global fault source features and the local response matrix to obtain propagation enhancement features; determine residual increment features based on the propagation enhancement features and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of background components in the thermal infrared image; correct the local block feature matrix based on the residual increment features; and determine the fault category of the photovoltaic module using a preset classifier based on the global category features and the corrected local block feature matrix. Global fault source features are obtained by spatial context encoding local block features and coupled with local response matrices for propagation. This enables the global fault semantics to be directionally transmitted to various local spatial locations, effectively solving the problem of fault signals being submerged in a large background. Furthermore, background components are suppressed by using damping coefficients to further highlight fault-related features and improve the sensitivity to weak thermal anomalies. Additionally, the local block feature matrix is modified based on residual incremental features, achieving incremental enhancement of fault semantics without destroying the original feature structure and preserving the original discriminative ability of the basic model. Finally, accurate identification of photovoltaic module fault categories is achieved through global and local fusion classification. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the accompanying drawings used in the description of the embodiments or prior art will be briefly introduced below.
[0024] Figure 1 This is a flowchart of a photovoltaic module fault identification method provided in the embodiments of this specification; Figure 2 This is a schematic diagram of the framework of a photovoltaic module fault identification model provided in the embodiments of this specification; Figure 3 This is a schematic diagram of the structural composition of a photovoltaic module fault identification device provided in the embodiments of this specification; Figure 4 This is a schematic diagram of the structural composition of a computer device provided in the embodiments of this specification. Detailed Implementation
[0025] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0026] It should be noted that the terms "first," "second," etc., used in this specification, claims, and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0027] In some embodiments, a photovoltaic module is the core power generation unit of a photovoltaic power generation system. It consists of multiple solar cells connected in series and parallel and encapsulated between materials such as glass, EVA film, and a backsheet, used to convert solar energy into electrical energy. During long-term outdoor operation, photovoltaic modules are susceptible to various types of failures due to factors such as changes in ambient temperature, humidity corrosion, wind and sand abrasion, bird droppings, local shading, and manufacturing defects. These failures mainly include hot spot failures, bypass diode failures, sub-string open circuit failures, and cell overheating failures. Hot spot failures manifest as one or more solar cells in the module being in a power-consuming state due to shading or current mismatch, with a local temperature significantly higher than that of the surrounding normal solar cells. Bypass diode failures manifest as short circuits or open circuits in the bypass diodes, causing the corresponding cell string to malfunction. Sub-string open circuit failures manifest as a break in a sub-string inside the module due to poor welding or corrosion, resulting in a decrease in the module's output power. Cell overheating failures manifest as abnormal temperature rises in the solar cells due to internal defects or excessive contact resistance. If the above-mentioned faults are not detected in time, they will not only lead to a decrease in power generation efficiency, but may also cause safety accidents such as electric arcs and fires in serious cases, threatening the safe operation of photovoltaic power stations.
[0028] Traditional manual ground inspection methods rely on maintenance personnel carrying handheld thermal imagers or visually inspecting each module, which has limitations such as low efficiency, high cost, and limited coverage, making it difficult to meet the maintenance needs of large-scale photovoltaic power plants. Unmanned aerial vehicles (UAVs) equipped with thermal infrared cameras can achieve rapid, large-area coverage at a lower cost and have become the mainstream technology for fault detection in photovoltaic power plants. The UAV flies over the photovoltaic array along a preset route, collecting temperature distribution images of each module's surface using thermal infrared cameras. The thermal infrared images are then analyzed by a ground station or onboard device to identify whether there are temperature anomalies and the type of anomaly. However, due to limitations in onboard computing resources and the low resolution and signal-to-noise ratio of thermal infrared images themselves, existing UAV thermal infrared fault identification methods still have significant shortcomings in identification accuracy and cross-scenario generalization ability.
[0029] At the fault identification algorithm level, existing methods mainly fall into two categories. One type relies on manually designed feature extraction operators to extract temperature statistics, texture features, or shape features from thermal infrared images, then combines these with rule-based thresholds or traditional machine learning classifiers for fault determination. This type of method is highly sensitive to changes in environmental temperature, lighting conditions, and shooting angles; when the inspection scenario changes, the thresholds need to be readjusted or the features redesigned, resulting in poor robustness. The other type of method is based on deep learning models such as convolutional neural networks or visual Transformers. It performs end-to-end supervised training on thermal infrared image datasets collected from specific sites, allowing the network to automatically learn fault-related discriminative features. This achieves high recognition accuracy under the scenario conditions covered by the training set. However, this type of method is typically trained on data collected from a single or a few power plants. The model parameters are highly coupled with the scenario distribution of the training data. When deployed to new power plants with different geographical locations, climate conditions, flight altitudes, or component models, the model's recognition performance significantly degrades due to the distribution shift between training and test data, making it difficult to meet the actual operation and maintenance needs of large-scale heterogeneous photovoltaic power plants. In addition, both of the above methods extract global features of the entire image for classification decisions. However, the fault area in photovoltaic thermal infrared images often only occupies a very small part of the total image area. The local temperature anomaly signal caused by the fault is easily submerged by the large area of normal background information during the global feature aggregation process, resulting in a low recall rate for small-area faults.
[0030] This specification provides a method for identifying photovoltaic module faults, as illustrated in the embodiments below. Figure 1 As shown, the specific implementation may include the following steps: S101: Input the thermal infrared image of the photovoltaic module collected by the UAV into the visual feature extractor to obtain the global category feature and local block feature matrix.
[0031] A visual feature extractor performs multi-layer semantic abstraction on the input image, extracting global scene-level information and local spatial detail information. Global category features characterize the global semantic attributes of the scene to which the entire thermal infrared image belongs, serving as the global basis for subsequent fault classification. The local block feature matrix contains feature vectors for multiple image blocks, characterizing the local temperature distribution patterns and texture details at various spatial locations in the thermal infrared image, serving as the spatial carrier for subsequent local fault semantic enhancement. By outputting both global category features and the local block feature matrix simultaneously through the visual feature extractor, a complementary feature foundation is provided for subsequent dual-branch fusion classification, solving the problem that a single global feature cannot adequately capture local detail perception.
[0032] In some embodiments, a thermal infrared image can be divided into multiple image blocks.
[0033] An image patch consists of several non-overlapping rectangular region units obtained by spatially segmenting a thermal infrared image according to a preset grid division method. Each image patch is of uniform size and collectively covers the entire spatial range of the thermal infrared image. By discretizing the entire image into a regularly arranged sequence of image patches, subsequent processing can independently encode and selectively enhance local spatial information using image patches as basic units, providing a spatial granularity basis for distinguishing faulty areas from normal areas.
[0034] In some embodiments, each image patch can be mapped to a feature vector, and the superposition position of each feature vector can be encoded to obtain an image patch feature sequence.
[0035] Feature vectors are obtained by linear projection transformation of pixel values within an image patch, used to map the image patch from pixel space to a high-dimensional feature space. Positional encoding represents the spatial arrangement order of each image patch in the original image. By superimposing positional encodings onto the feature vectors of the corresponding image patches, the subsequent transform encoder can perceive the relative spatial positional relationships of each feature vector. The image patch feature sequence consists of the feature vectors of all image patches arranged in spatial order, serving as the serialization input for the visual feature extractor. Through the superposition of positional encodings, the transform-based visual feature extractor can utilize the positional information of the feature vectors in the sequence to perceive the spatial distribution of each image patch, providing a positional basis for subsequent spatial context modeling.
[0036] In some embodiments, a learnable class feature vector can be inserted at the beginning of the image patch feature sequence to obtain the input sequence.
[0037] The learnable class feature vector can be a trainable parameter vector that is continuously updated and optimized through backpropagation during training. It does not depend on the local information of any specific image patch, but rather learns how to aggregate global discriminative information from the entire image during training. Placing the class feature vector at the beginning of the image patch feature sequence allows it to interact with all image patch features through a multi-head self-attention mechanism during the layer-by-layer processing of the transform encoder, ultimately converging the global semantic information of the entire image. The input sequence is formed by sequentially concatenating the class feature vector with all image patch feature vectors, serving as the complete input to the multi-layer transform encoder. By inserting the class feature vector, the visual feature extractor can directly extract the feature at that location as the global class feature at output, eliminating the need for additional global pooling operations and avoiding the information loss caused by the averaging of local fault information in global pooling.
[0038] In some embodiments, the input sequence can be input into a multilayer transform encoder of a visual feature extractor, where each transform encoder is used to transform the input sequence layer by layer through a multi-head self-attention and feedforward network.
[0039] The transform encoder consists of a multi-head self-attention module and a feedforward network module connected in series. The multi-head self-attention module calculates the semantic association strength between any two positions in the input sequence, enabling features at each position to aggregate relevant information from other global positions during updates. The feedforward network module performs independent nonlinear transformations on the features output by the self-attention module, enhancing the expressive power and nonlinear fitting ability of the features. Each layer of the transform encoder includes residual connections and layer normalization. Residual connections maintain the stability of feature propagation, while layer normalization eliminates numerical scale differences between different samples. Through the progressive abstraction of the multi-layer transform encoder, features at each position in the input sequence gradually fuse with global contextual information, outputting a deep feature sequence rich in semantic information. The multi-layer transform encoder, through the global receptive field of the self-attention mechanism, enables the features of each image patch to perceive the global contextual information of the entire image, providing a semantically rich feature foundation for subsequent global fault source extraction and local feature enhancement.
[0040] In some embodiments, the head vector and the tail vector output by the last layer encoder can be used as the global category feature and the local block feature matrix, respectively.
[0041] The output of the final encoder layer is a feature sequence, with a length equal to the input sequence. Each position in the sequence corresponds to a feature vector. The first position of the sequence corresponds to the final state of the learnable class feature vector inserted during input, after multiple transformations. This feature vector is optimized during training to aggregate global discriminative information of the entire image, serving as the global class feature output. The remaining parts of the sequence, excluding the first position, correspond to the final state of the feature vectors at each image patch position after multiple transformations. These feature vectors are arranged according to their spatial positions, serving as the local block feature matrix output. By separating the first and last vectors, the goal of simultaneously extracting global and local representations from a single transform encoder is achieved. The two outputs are aligned in the feature space but semantically complementary, providing a unified feature source for downstream dual-branch processing.
[0042] S102: Perform spatial context encoding on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image.
[0043] Spatial context coding is used to aggregate the scattered semantic information carried by multiple local block features distributed at different spatial locations in an image into a compact vector representation that can represent the global fault tendency of the entire image. Each image block feature in the local block feature matrix represents the temperature distribution pattern and texture details of its corresponding spatial location in the thermal infrared image. The global fault source feature is used to represent the collective fault indication information at multiple spatial locations in the thermal infrared image. As the source signal in the subsequent wave propagation process, it is responsible for providing scene-level fault semantic guidance to each local location. By aggregating the weak fault signals scattered in each image block into a global fault source feature through spatial context coding, the originally independent local features obtain a unified fault semantic expression from a global perspective. This provides an information source for the subsequent directional propagation of fault signals to each local location, solving the problem that a single image block feature cannot easily determine whether it is a fault area.
[0044] In some embodiments, the local block feature matrix may be normalized before step S102.
[0045] Normalization transformation is used to eliminate the inconsistency in numerical scale between image patch features in the local block feature matrix caused by sample differences. The numerical ranges of features in each image patch within the local block feature matrix may vary significantly. These differences may originate from actual fault temperature anomalies or from interference from non-fault factors such as ambient temperature, shooting distance, and camera response. Normalization transformation, by performing layer-by-layer normalization on the local block feature matrix, adjusts the numerical distribution of features in each image patch to a uniform scale range. This ensures that subsequent spatial context coding operations are not biased by differences in numerical scale, guaranteeing that the encoded global fault source features accurately reflect the relative fault indication intensity at each spatial location, rather than its absolute numerical value. By eliminating numerical scale interference introduced by differences in imaging conditions between different samples, normalization transformation provides numerically stable feature inputs for spatial context coding.
[0046] In some embodiments, spatial context encoding can be performed on the normalized features of multiple image blocks to obtain global fault source features.
[0047] The normalized features of multiple image patches maintain a consistent numerical scale. The numerical differences between these features primarily reflect the deviation of each spatial location from the average temperature level of the entire image, rather than scale differences caused by varying imaging conditions. Spatial context encoding of the normalized features ensures that the generation of global fault source features is based on the relative comparisons between the features, rather than being dominated by excessively large absolute values of individual features. Among the normalized features, the features corresponding to fault regions exhibit statistically outlier characteristics compared to normal regions. Spatial context encoding, by aggregating outlier information from all feature patches, enables the obtained global fault source features to more accurately reflect the presence and overall distribution of temperature anomalies in the entire image, preventing extreme temperature values in one region from masking anomalous information in other regions.
[0048] In some embodiments, global mean pooling can be performed on multiple image block features along the dimension of the local block feature matrix to obtain fault aggregation features.
[0049] The dimension of the local block feature matrix can be the direction of the number of image block features in the local block feature matrix, with each image block feature occupying an independent position in the matrix dimension. Global mean pooling is used to calculate the arithmetic mean of each image block feature in each feature channel along the sequence dimension, compressing multiple image block features in the sequence dimension into a single feature vector. This operation does not introduce additional learnable parameters and extracts the central tendency of each image block feature in a fixed calculation method. Fault aggregation features are used to characterize the average feature level of the entire thermal infrared image at all spatial locations, serving as the input basis for subsequent linear projection transformations. By compressing all image block features into a single vector through global mean pooling, the subsequent transformation process can generate fault source features based on global information rather than local information.
[0050] In some embodiments, a first linear projection transformation can be performed on the fault aggregation features to obtain global fault source features.
[0051] The first linear projection transformation applies a learnable linear mapping to the fault aggregation features, transforming them from the intermediate space after mean pooling to a global fault source feature space aligned with the original feature space. After global mean pooling, the fault aggregation features retain the average response level of each channel feature in the entire image. However, the distribution of this average response level is not directly aligned with the original feature space. Therefore, a learnable linear projection matrix is used to map it into global fault source features with clear fault semantic indications. This linear projection matrix contains learnable weight parameters, which are continuously updated and optimized through backpropagation during training, learning how to extract the most discriminative fault indication signals from the average response. The introduction of the learnable projection layer enables the global fault source features to have task-adaptive expressive capabilities, automatically adjusting the semantic emphasis of the source features based on the supervision signals of the training data, resulting in stronger discriminative power compared to fixed-calculation aggregation methods.
[0052] S103: Perform gated mapping on the features of each image block to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image.
[0053] Gating mapping transforms the features of each image patch from the original feature space to a gating space. In this gating space, each image patch location is assigned a scalar value that controls the received intensity of the global fault source features at that location. Each image patch feature represents the local temperature distribution pattern and texture detail information at each spatial location in the thermal infrared image. The local response matrix represents the absorption intensity of the global fault indication information at each spatial location in the thermal infrared image, with each spatial location corresponding to a local response coefficient. Image patch locations with larger response coefficients in the local response matrix indicate a higher correlation with the global fault semantics and are assigned a higher fault signal absorption intensity; locations with smaller response coefficients indicate a lower correlation with the global fault semantics and are assigned a lower fault signal absorption intensity. By adaptively converting the features of each image patch into absorption intensity coefficients through gating mapping, the subsequent propagation process can determine how much global fault signal to absorb based on the characteristic attributes of each location, rather than uniformly distributing it across all locations. This solves the problem of differentiated enhancement required for different spatial locations due to varying correlations with fault semantics.
[0054] In some embodiments, gating mapping can be performed on the normalized features of each image patch to obtain a local response matrix.
[0055] After normalization, the features of each image patch maintain a consistent numerical scale. The numerical differences among the features of each image patch mainly reflect the deviation of each spatial location from the average feature level of the entire image. Gating mapping is applied to the normalized features of each image patch, ensuring that the generation of local response coefficients is based on the relative comparison between the features of each image patch, rather than being dominated by excessively large absolute values of individual image patch features. In the normalized feature space, the image patch features corresponding to faulty regions and those corresponding to normal regions are distinguishable in numerical distribution. Gating mapping automatically discovers this distinguishability through learnable parameters and maps it to different response coefficient values. By performing gating mapping after normalization, the interference of differences in imaging conditions between different samples on the gating coefficient generation process is eliminated, ensuring that the values of the gating coefficients accurately reflect the deviation of each spatial location from the global feature distribution of the entire image, rather than being dominated by differences in numerical scale between samples.
[0056] In some embodiments, a second linear projection transformation can be performed on the features of each image patch to obtain an initial gating response matrix.
[0057] The second linear projection transformation applies a learnable linear mapping to the features of each image patch, independently projecting the features of each patch from the original feature space to the gated score space. This linear projection transformation does not involve interactive computation between image patches; the gated score for each patch location is obtained solely from the linear projection of its own feature vector. The initial gated response matrix characterizes the original gated score of each image patch location before the nonlinear mapping by the activation function; this score is not limited in numerical range. A positive gated score for each image patch location indicates a positive correlation between the feature pattern at that location and the fault semantics, while a negative value indicates a negative correlation. Through the second linear projection transformation, each image patch location obtains an initial gated score independent of other locations, providing an original response strength benchmark for subsequent nonlinear mapping.
[0058] In some embodiments, an activation transformation can be performed on the initial gating response matrix to obtain a local response matrix.
[0059] The activation transformation maps the original gating scores of each image patch in the initial gating response matrix to a preset numerical range, giving each gating score a clear physical meaning within this range. The lower limit of this preset numerical range corresponds to the "complete non-absorption" state, the upper limit to the "complete absorption" state, and the middle value to the "partial absorption" state. The activation transformation is implemented using a nonlinear activation function that monotonically increases within the numerical range, maintaining the relative magnitude of the initial gating scores—locations with high initial scores still obtain high response coefficients after the activation transformation, and locations with low initial scores still obtain low response coefficients, but all response coefficients are constrained to the preset range. By compressing the unbounded gating scores into a physically meaningful numerical range, the activation transformation allows local response coefficients to directly participate in subsequent element-wise multiplication operations as absorption intensity. Simultaneously, the nonlinear mapping enhances the sensitivity to subtle differences in gating scores, improving the distinction between fault candidate regions and normal backgrounds in terms of absorption intensity.
[0060] S104: Coupling the global fault source characteristics with the local response matrix yields propagation enhancement characteristics.
[0061] Global fault source features are used to characterize the collective fault indication information of the entire thermal infrared image, while the local response matrix is used to characterize the absorption intensity of the global fault indication information at each spatial location. Coupling combines the global fault source features and the local response matrix according to their spatial correspondence, allowing the scene-level fault semantics carried in the global fault source features to be allocated to each spatial location, with the allocation ratio determined by the absorption intensity at each location. Propagation enhancement features characterize the enhanced feature representation formed by gating and selecting each image patch location after receiving the global fault semantic injection. In this feature representation, fault-related locations receive additional fault semantic information injection, while background-related locations do not receive the same level of injection. By coupling the global fault source features and the local response matrix, selective propagation of global fault semantics from the entire image level to each local location is achieved, resulting in stronger semantic enhancement for locations with high correlation to fault semantics, and weaker or no semantic enhancement for locations with low correlation to fault semantics.
[0062] In some embodiments, the global fault source features can be extended to the same spatial dimension as the features of each image patch to obtain broadcast source features.
[0063] The global fault source feature is a single vector in the spatial dimension, lacking spatial distribution information and thus unable to be directly computed positionally with the image patch features at each local location. The extension is used to spatially replicate the global fault source feature, ensuring an identical copy appears at every image patch location. Broadcasting the source feature characterizes the uniform distribution of the global fault source feature across image patch locations, giving it the same spatial dimension as the local response matrix, providing a prerequisite for spatial dimension matching in subsequent positionally computations. By broadcasting the global fault source feature to all image patch locations, global fault semantic information is available at every location, providing an information basis for subsequent selective absorption based on the response coefficients at each location. During broadcasting, the semantic content itself remains consistent across locations, preserving the uniformity of the global fault semantics and avoiding deviations introduced at different locations.
[0064] The correlation coefficient of each image patch location can be obtained based on the semantic correlation between the features of each image patch and the global fault source features. Semantic correlation is used to characterize the degree of matching between the local fault indication information carried by the features of each image patch and the global fault indication information.
[0065] Before broadcasting, the semantic correlation between the features of each image patch and the global fault source features can be calculated. This correlation coefficient is used to characterize the degree of matching between the local indication information carried by each image patch feature and the global fault indication information. One calculation process is as follows: after normalizing the features of each image patch and the global fault source features, the dot product similarity between the two is calculated, and the scalar value is used as the correlation coefficient of the image patch location.
[0066] The propagation modulation coefficients for each image patch can be determined based on the correlation coefficient. The propagation modulation coefficients are positively correlated with the correlation coefficients. The propagation modulation coefficients characterize the intensity scaling ratio as global fault source features are transmitted to each image patch location.
[0067] Based on the correlation coefficient of each image patch location, the propagation modulation coefficient of each image patch location is determined. The propagation modulation coefficient is positively correlated with the correlation coefficient; locations with high correlation receive larger propagation modulation coefficients, while locations with low correlation receive smaller propagation modulation coefficients. The propagation modulation coefficient is used to characterize the intensity scaling ratio when global fault source features are transmitted to each image patch location, serving as a differential adjustment factor for uniform broadcasting operations.
[0068] The global fault source characteristics can be multiplied element-wise with the propagation modulation coefficients of each image block location to obtain the broadcast source characteristics.
[0069] The global fault source features are multiplied element-wise with the propagation modulation coefficients of each image block location to obtain the broadcast source features. This results in the global fault source features received at each image block location not being completely consistent, but rather being differentiated and scaled according to the degree of matching between the location and the fault semantics. Locations that match the fault semantics better receive stronger source signals, while locations that match the fault semantics less receive weaker source signals.
[0070] By introducing semantic correlation calculation and propagation modulation, the broadcast source features retain global semantic uniformity while possessing spatial differentiation characteristics, thus avoiding the background false enhancement problem that may be introduced by uniform broadcasting, which causes fault semantics to be transmitted to all locations to the same extent.
[0071] In some embodiments, the broadcast source features can be multiplied element-wise with the local response matrix to obtain the propagation increment features of each image block location.
[0072] Each location in the broadcast source features carries a complete global fault source feature vector, and each location in the local response matrix corresponds to a scalar response coefficient. Element-wise multiplication is used to multiply the response coefficient at each location by each element of the global fault source feature vector at that location, scaling the global fault source feature vector proportionally at each location according to the response coefficient corresponding to that location. The propagation increment feature is used to characterize the portion of fault semantic feature increment allocated from the global fault source after absorption intensity modulation at each image patch location. Locations with response coefficients close to 1 receive nearly complete fault source features as increments, while locations with response coefficients close to 0 receive fault source feature increments approaching zero. Through element-wise multiplication, global fault semantic information is differentially allocated to each spatial location according to the absorption intensity at each location, achieving selective injection of fault semantics.
[0073] In some embodiments, the propagation increment features can be superimposed on the features of each image patch to obtain propagation enhancement features.
[0074] In this embodiment, propagation increment features are superimposed on the features of each image patch to obtain propagation enhancement features. Superposition is used to fuse the propagation increment features with the original normalized image patch features in an element-wise addition manner. The propagation increment features contain the fault semantic increment allocated to each location from the global fault source, while the normalized image patch features contain the original local temperature distribution and texture information of each location. Adding them element-wise allows the original local features at each location to receive incremental supplementation from the global fault semantics. The propagation enhancement features characterize the enhanced feature expression of each image patch location after fusing the global fault semantic increment. In this feature expression, the fault semantic response of fault-related locations is significantly enhanced, while the fault semantic response of background-related locations is not enhanced equally. Through the superposition operation, scene-level fault indication information from the global fault source features is incrementally injected into each local location without changing the basic structure of the original features, enabling subsequent processing to make more accurate fault classification decisions based on the enhanced features.
[0075] The feature vectors at different locations in the propagation increment features may differ significantly in their numerical norms. Direct superposition may cause some original features to be submerged. Gated compression is used to perform a nonlinear transformation on the propagation increment features, compressing the propagation increment feature vectors at each image patch location into scalar weights. The numerical range of these scalar weights is constrained within a preset interval. The upper limit of this preset interval corresponds to the "fully superimposed" state, and the lower limit corresponds to the "slightly superimposed" state. Based on the adaptive fusion weights at each image patch location, the normalized image patch features and the propagation increment features are weighted and summed to obtain the propagation enhancement features at each image patch location. During weighted summation, a larger adaptive fusion weight results in a higher contribution ratio of the propagation increment features and a correspondingly lower contribution ratio of the original features; conversely, a smaller adaptive fusion weight results in a higher contribution ratio of the original features and a correspondingly lower contribution ratio of the propagation increment features. Through gated compression and weighted summation, the superposition process has input-dependent adjustment capabilities, avoiding excessive interference of the numerical amplitude of the propagation increment features on the original features and maintaining the stability of feature updates.
[0076] S105: Determine the residual increment characteristics based on the propagation enhancement characteristics and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of the background components of the thermal infrared image.
[0077] The propagation enhancement feature characterizes the enhanced feature representation of each image patch location after receiving global fault semantic injection. This representation includes both the enhanced response of fault-related locations and the residual response of normal background locations. The damping coefficient characterizes the degree of suppression of background components in the thermal infrared image. The damping coefficient is a learnable scalar parameter that adaptively adjusts during training. The residual increment feature characterizes the pure fault increment signal obtained after removing background components from the propagation enhancement feature. This signal contains only the differential information introduced during propagation that is distinguishable from the background. By introducing the damping coefficient to remove background components from the propagation enhancement feature, the feature responses of normal regions in the residual increment feature are actively suppressed while the feature responses of fault regions are preserved. This solves the problem of simultaneous amplification of fault signals and background noise in the propagation enhancement feature, ensuring that subsequent superposition onto the original feature only introduces beneficial fault increments without introducing background noise increments.
[0078] In some embodiments, the normalized features of each image patch can be multiplied by the damping coefficient to obtain background suppression features.
[0079] The normalized image patch features contain the original local temperature distribution and texture pattern information at each spatial location, including both the feature response at the fault location and the intrinsic response at the background location. The damping coefficient, as a scalar parameter, acts on all channels of each image patch feature, scaling the features proportionally. Background suppression features characterize the portion of the original local features labeled as "background components" and prepared for subtraction from the propagation enhancement features. By multiplying the normalized image patch features by the damping coefficient, the subsequent subtraction operation removes a certain proportion of the original background baseline from the propagation enhancement features. The retained incremental signal only contains the portion exceeding the original baseline introduced during propagation; this excess is more significant at fault locations and approaches zero at background locations, providing a quantitative basis for suppressing excessive responses in the background region.
[0080] Context-aware projection can be performed on the normalized features of each image patch to obtain the feature activity of each image patch feature.
[0081] Context-aware projection applies learnable linear and nonlinear transformations and mappings to the normalized features of each image patch, transforming the features from the original feature space to a feature activity space. During context-aware projection, each image patch location not only projects using its own feature vector but also incorporates information from other image patch features in its local neighborhood. This ensures that the final output feature activity reflects not only the response intensity of that location but also the degree of difference between that location and its local neighborhood. Feature activity characterizes the salience of each image patch feature relative to its local neighborhood; locations with high feature activity indicate that their feature patterns differ significantly from their surrounding locations, while locations with low feature activity indicate that their feature patterns tend to be consistent with their surrounding locations.
[0082] The spatial modulation coefficients of each image patch feature can be determined based on feature activity. The spatial modulation coefficients are positively correlated with feature activity. They characterize the local modulation intensity of the damping coefficient.
[0083] The spatial modulation coefficient is positively correlated with the characteristic activity; locations with high characteristic activity have larger spatial modulation coefficients, while locations with low characteristic activity have smaller spatial modulation coefficients. The spatial modulation coefficient is used to characterize the local adjustment strength of the damping coefficient.
[0084] The spatial modulation coefficients and damping coefficients of each image patch feature can be fused to obtain the modulation damping coefficient of each image patch feature. The normalized image patch feature can be multiplied by the corresponding modulation damping coefficient to obtain the background suppression feature.
[0085] The modulation damping coefficient takes independent values at each location, achieving spatial differentiation of damping intensity. The background suppression features are obtained by multiplying the normalized features of each image patch by the corresponding modulation damping coefficient.
[0086] By introducing context-aware projection and spatial modulation, the background suppression intensity is no longer a globally uniform value, but is adaptively adjusted according to the salience of each location relative to its local neighborhood. In normal regions, features tend to be consistent with their neighborhood, resulting in a larger modulation damping coefficient, and background components are fully subtracted. In fault regions, features are significantly different from their neighborhood, resulting in a smaller modulation damping coefficient, thus avoiding excessive attenuation of the fault signal.
[0087] The spatial modulation coefficient can be a scalar value positively correlated with feature activity, taking independent values at each spatial location; the damping coefficient is a globally learnable scalar parameter, taking a uniform value across the entire feature map. The fusion operation uses the damping coefficient as the base value and locally adjusts it using the spatial modulation coefficient. Specifically, the damping coefficient is multiplied by the spatial modulation coefficient to obtain a modulation damping coefficient that takes independent values at each image patch location. This modulation damping coefficient varies at different locations. Locations with a spatial modulation coefficient greater than 1 receive a higher modulation damping coefficient than the global damping coefficient, indicating stronger background suppression at these locations; locations with a spatial modulation coefficient less than 1 receive a lower modulation damping coefficient than the global damping coefficient, indicating weaker background suppression at these locations. By fusing the global damping coefficient and the local spatial modulation coefficient, the final background suppression strength possesses both global learnability and spatial adaptability. The global damping coefficient ensures a controllable baseline for the overall suppression level, while the local spatial modulation coefficient ensures that the suppression strength adapts to changes in spatial location.
[0088] In some embodiments, the residual increment feature can be obtained by subtracting the background suppression feature from the propagation enhancement feature.
[0089] The subtraction operation is an element-wise subtraction, that is, subtracting the value of the corresponding channel in the background suppression feature from each feature channel at each position in the propagation enhancement feature. The propagation enhancement feature is a superposition of the original feature and the propagation increment, while the background suppression feature contains the background components extracted from the original feature that need to be subtracted. After subtracting the background suppression feature, the background components in the original feature are actively removed, and the parts in the propagation increment feature that are considered to belong to the background are also removed simultaneously, leaving the pure fault increment signal. The residual increment feature is used to characterize the beneficial increment information remaining after subtracting the background components from the propagation enhancement feature, which can be superimposed on the original local block feature. This feature has a significant non-zero response at the fault location and a response close to zero at the normal background location. By subtracting the background suppression feature from the propagation enhancement feature, the residual increment feature focuses on the differential signal introduced by propagation rather than the original feature itself, ensuring that when it is subsequently superimposed on the original local block feature in residual form, it only brings beneficial fault semantic enhancement without introducing additional amplification of background noise.
[0090] S106: Correct the local block feature matrix based on the residual increment characteristics.
[0091] The residual increment feature is used to characterize the pure fault increment signal remaining after subtracting the background component from the propagation enhancement feature. This signal has a significantly non-zero response at the fault location and a response approaching zero at the normal background location. The local block feature matrix is used to characterize the original spatial local information of each image block location output by the visual feature extractor, including the feature responses of both the fault region and the normal background region. The correction is used to introduce the fault increment signal from the residual increment feature into the local block feature matrix, so that the original local features receive incremental fault semantic supplementation while maintaining the basic structure of the original features. By correcting the local block feature matrix with the residual increment feature, additional fault semantic responses are added to the fault-related locations without overwriting the original features. This enhances the separability of the fault region and the normal background in the feature space of the corrected feature matrix, while retaining other effective information unrelated to the fault in the original features extracted by the backbone network.
[0092] In some embodiments, the residual increment feature can be added element by element to the local block feature matrix to obtain the corrected local block feature matrix.
[0093] Element-wise addition is used to add the value of each feature channel at each spatial location in the residual increment feature to the value of the corresponding channel at the corresponding location in the local block feature matrix. The residual increment feature and the local block feature matrix have the same size in both spatial and feature dimensions, ensuring a one-to-one correspondence between each location and each channel during element-wise addition. Through element-wise addition, the fault increment signal in the residual increment feature is superimposed onto the original feature at the corresponding location. The feature response at fault-related locations is enhanced, while the normal background locations remain largely unchanged because the residual increment feature response approaches zero. The corrected local block feature matrix is used to represent the enhanced local feature expression after fusing fault increment information. In this feature expression, the discriminative information of the fault region is amplified, the background information of the normal region does not receive additional enhancement, and the feature distance between the two types of regions is increased. By incrementally superimposing residuals, the fault semantic information in the original features is supplemented by addition rather than covered by replacement. This ensures that all effective information related to fault discrimination in the original features output by the visual feature extractor is completely preserved. Only beneficial fault incremental signals are added on this basis, avoiding the problem of information loss caused by feature replacement.
[0094] S107: Based on the global category features and the corrected local block feature matrix, use a preset classifier to determine the fault category of the photovoltaic module.
[0095] Global category features are used to characterize the global semantic attributes of the scene to which the entire thermal infrared image belongs, serving as the basis for scene-level fault identification. The corrected local block feature matrix is used to characterize the enhanced local feature representation of each image block location after fusing fault increment information, where the separability between fault and normal regions in the feature space is significantly enhanced. A pre-defined classifier maps the input features from the feature space to the fault category space, outputting the probability distribution of each fault category. By simultaneously utilizing global category features and the corrected local block feature matrix for classification decisions, the final fault category determination references both the robust global judgment provided by the backbone network and the enhanced local fault details through the wave residual adapter. These two types of information complement each other, solving the problems of single global features being insensitive to small-area faults or single local features lacking global contextual constraints.
[0096] In some embodiments, global category features can be input into a first classifier to obtain global fault identification results.
[0097] The first classifier is a global classifier. The global classifier can be a learnable linear classification layer containing a weight matrix and a bias vector. This linear classification layer applies a linear mapping to the global category features, projecting them from the feature space to the category space, and outputting a raw score vector with a dimension equal to the number of fault categories. The global fault identification result is used to characterize the preliminary fault category judgment made solely based on the global semantic information of the entire image. This result reflects the visual feature extractor's tendency to discriminate scene-level attributes of the entire image. The global category features, through layer-by-layer abstraction by a multi-layer transform encoder, aggregate comprehensive information from various spatial locations within the entire image. Therefore, the output of the global classifier tends to reflect the dominant fault type in the entire image or an overall judgment that the entire image is fault-free. By outputting the global fault identification result separately, the robust judgment capability of the visual base model for scene-level attributes of the entire image is preserved, providing a global perspective reference for the final fusion decision.
[0098] In some embodiments, the corrected local block feature matrix can be input into a second classifier to obtain local fault identification results.
[0099] The second classifier is a local classifier. The local classifier comprises a global average pooling layer and a learnable linear classification layer. The global average pooling layer calculates the arithmetic mean of each feature channel along the sequential dimension of the corrected local block feature matrix, compressing the features of multiple image block locations into a single feature vector. This feature vector characterizes the average level of the corrected local block feature matrix across each feature channel. The learnable linear classification layer applies a linear mapping to this single feature vector, projecting it from the feature space to the class space, outputting a raw score vector with a dimension equal to the number of fault categories. The local fault identification result characterizes the preliminary fault category judgment based on the enhanced local features. This result reflects the discriminative tendency of the aggregated fault semantic information of each image block location after wave residual adapter enhancement. The response of fault-related locations in the corrected local block feature matrix is significantly enhanced; therefore, the output of the local classifier tends to reflect the presence and specific type of local thermal anomalies in the image, exhibiting higher sensitivity to small-area faults. By outputting the local fault identification result separately, quantitative evaluation and independent utilization of the wave residual adapter enhancement effect are achieved.
[0100] In some embodiments, the fault category of a photovoltaic module can be determined based on the fusion of global fault identification results and local fault identification results.
[0101] Fusion is used to combine the scene-level discrimination information provided by the global fault identification results and the local anomaly discrimination information provided by the local fault identification results into a unified decision output. The fusion method involves combining the original score vectors corresponding to the global and local fault identification results according to a preset fusion strategy to obtain a fused score vector. Each dimension of this vector corresponds to the comprehensive score of a fault category. Preset fusion strategies include, but are not limited to, weighted summation, linear transformation after concatenation, and confidence-gated selection. Based on the comprehensive scores of each fault category in the fused score vector, the fault category with the highest comprehensive score is determined as the fault category of the photovoltaic module. By fusing the identification results of the global and local branches, the final decision benefits from both the robust scene judgment capability of the global branch and the sensitive local anomaly perception capability of the local branch, achieving complementary advantages of the two types of information. This effectively improves the identification accuracy of small-area faults while retaining the strong generalization capability of the backbone network.
[0102] In some embodiments, step S102 above, which involves performing global mean pooling on multiple image patch features along the dimension of the local block feature matrix to obtain fault aggregated features, may further include: performing a fifth linear projection transformation on each image patch feature to obtain the embedded features corresponding to each image patch feature; constructing a similarity matrix based on the similarity between the embedded features; determining the outlier degree of each image patch location relative to the global feature distribution based on the similarity matrix; the outlier degree is used to characterize the deviation between the feature pattern of each image patch location and the global mainstream feature distribution; determining the aggregation weight of each image patch location based on the outlier degree of each image patch location; the aggregation weight is negatively correlated with the outlier degree; and performing weighted global mean pooling on multiple image patch features in the local block feature matrix based on the aggregation weight to obtain aggregated features.
[0103] A fifth linear projection transformation is applied to the features of each image patch to obtain the corresponding embedded features. This fifth linear projection transformation applies a learnable linear mapping to each image patch feature, projecting it from the original feature space to the embedding space. The embedded features characterize the low-dimensional representation of each image patch feature in the embedding space, where the similarity between image patch features can be measured more accurately. By projecting the image patch features into the embedding space, subsequent similarity calculations are based on task-adaptive feature representations rather than direct distances in the original feature space, thus improving the discriminative power of the similarity measurement.
[0104] A similarity matrix is constructed based on the similarity between each embedded feature. Similarity characterizes the closeness of the embedded features of any two image patch locations in the embedding space, obtained by calculating the cosine similarity or the reciprocal of the Euclidean distance between each pair of embedded features. The similarity matrix represents the complete set of pairwise similarities between all image patch locations; its dimension is equal to the total number of image patch locations. The element in the i-th row and j-th column of the matrix represents the similarity between the i-th and j-th image patch locations. By constructing the complete similarity matrix, the outlier assessment of each image patch location can be based on an overall comparison with all other locations, rather than just a comparison with individual locations.
[0105] Based on the similarity matrix, the outlier degree of each image patch location relative to the global feature distribution is determined. Outlier degree characterizes the deviation of the feature pattern of each image patch location from the global mainstream feature distribution. The outlier degree is determined as follows: for each image patch location, a statistical measure of the similarity between that location and all other locations is calculated. This statistical measure can be the average or median of the similarity between that location and other locations. The lower the average similarity, the worse the consistency between the feature pattern of that location and the global mainstream feature distribution, and the higher the outlier degree; conversely, the higher the average similarity, the more consistent the feature pattern of that location is with the global mainstream feature distribution, and the lower the outlier degree. By assessing outlier degree based on the similarity matrix, the determination of outlier degree is based on a comprehensive comparison with all other locations, avoiding misjudgments caused by individual extreme values.
[0106] The aggregation weights for each image patch location are determined based on the outlier level. These aggregation weights characterize the contribution of each patch location to global mean pooling, and are negatively correlated with the outlier level. Locations with higher outlier levels indicate that their feature patterns deviate more from the global mainstream distribution, making them more likely to be faulty or noisy regions; therefore, they are assigned lower aggregation weights to reduce their interference with the aggregation results. Conversely, locations with lower outlier levels indicate that their feature patterns are closer to the global mainstream distribution; therefore, they are assigned higher aggregation weights to enhance the dominant role of mainstream features in the aggregation results. The aggregation weights are determined by taking the inverse ratio to the outlier level and then normalizing the result so that the sum of the aggregation weights for each location equals 1.
[0107] Based on the aggregation weights, weighted global mean pooling is performed on the features of multiple image blocks in the local block feature matrix to obtain aggregated features. Weighted global mean pooling is used to sum the features of each image block according to their corresponding aggregation weights, resulting in a single aggregated feature vector. Compared to equal-weighted global mean pooling, weighted global mean pooling reduces the contribution weight of outlier positions and increases the contribution weight of mainstream positions, making the aggregated features more accurately reflect the central trend of the global mainstream feature distribution. This avoids the problem of the aggregation result deviating from the true global distribution due to interference from a few extreme feature values.
[0108] Based on the above method, the aggregated global fault source features can be more robust and better represent the collective fault indication information that is truly common in the whole image.
[0109] In some embodiments, the activation transformation of the initial gating response matrix in step S103 to obtain the local response matrix may further include: for each image patch location, extracting the initial gating scores of multiple neighboring locations within a preset neighborhood centered on that location to construct a neighborhood gating score set for that location. Based on the neighborhood gating score set, calculating the neighborhood gating statistical features for that location. The neighborhood gating statistical features include at least one of the mean, median, or maximum value of the neighborhood gating scores. Based on the comparison between the initial gating score and the neighborhood gating statistical features for that location, determining the neighborhood salience coefficient for that location. The neighborhood salience coefficient characterizes the prominence of the initial gating score relative to its local neighborhood. Based on the initial gating scores and corresponding neighborhood salience coefficients of each image patch location, determining the local response coefficients for each image patch location to form a local response coefficient matrix.
[0110] For each image patch location, initial gating scores are extracted from multiple neighboring locations within a preset neighborhood centered on that location, constructing a neighborhood gating score set for that location. The preset neighborhood centered on the target image patch location and covering its surrounding neighboring image patches; the size of this window determines the spatial range of the neighborhood evaluation. The neighborhood gating score set is used to characterize the distribution of the original response intensity of each neighboring location within the neighborhood centered on the target image patch location without activation mapping.
[0111] Based on the neighborhood gating score set, the neighborhood gating statistical characteristics for that location are calculated. These characteristics characterize the central tendency or extreme value level of the initial gating scores within the neighborhood of the target location. The statistical characteristics include at least one of the mean, median, or maximum value of the neighborhood gating scores. The mean reflects the average baseline level of the neighborhood response, the median reflects the central location of the neighborhood response, and the maximum value reflects the peak level of the neighborhood response.
[0112] The neighborhood significance coefficient is determined by comparing the initial gating score at a given location with the neighborhood gating statistics. The comparison is obtained by calculating the difference or ratio between the initial gating score and the neighborhood gating statistics. The neighborhood significance coefficient characterizes the prominence of the initial gating score relative to its local neighborhood. When the initial gating score is higher than the neighborhood statistics, the neighborhood significance coefficient is greater than 1, indicating that the location has a significantly prominent response strength in its local neighborhood. When the initial gating score is not higher than the neighborhood statistics, the neighborhood significance coefficient is not greater than 1, indicating that the response strength at that location is not prominent relative to its neighborhood.
[0113] The local response coefficients for each image patch location are determined based on the initial gating score and the corresponding neighborhood saliency coefficient. This is achieved by multiplying the initial gating score by the neighborhood saliency coefficient, amplifying the initial gating score at salient locations and suppressing it at insignificant locations. The local response coefficients for all image patch locations are then arranged spatially to form a local response coefficient matrix. By introducing neighborhood saliency evaluation, the gating coefficient matrix effectively suppresses erroneous responses from isolated noise points while protecting the normal responses of continuous fault regions, thus resolving the problem of effectively distinguishing between isolated noise points and continuous faults in the initial gating response.
[0114] In some embodiments, the activation transformation of the initial gated response matrix in step S103 to obtain the local response matrix may further include: determining the spatial proximity and series-parallel attribution relationships between the image block locations based on the spatial coordinate information of each image block location, to construct a graph structure of the photovoltaic module. The graph structure includes multiple graph nodes, each corresponding to a different image block location, and the edge connections between graph nodes are determined based on the spatial proximity and series-parallel attribution relationships. Gating mapping is performed on the image block features corresponding to each graph node to obtain the initial gated response coefficients of each graph node. Using the initial gated response coefficients of each graph node as node features, graph convolution propagation is performed on the graph structure to obtain the propagated gated response coefficients of each graph node. A local response coefficient matrix is constructed based on the propagated gated response coefficients of each graph node.
[0115] Based on the spatial coordinates of each image patch location, the spatial proximity and series-parallel attribution relationships between image patch locations are determined to construct the graph structure of the photovoltaic module. Spatial coordinates characterize the two-dimensional spatial location of each image patch in the original thermal infrared image. Spatial proximity is determined based on the Euclidean distance between image patch locations; two locations with a distance less than a preset threshold are considered spatially adjacent. Series-parallel attribution relationships are determined based on the electrical connection topology of the solar cells in the photovoltaic module; image patch locations belonging to the same cell string or the same parallel branch have a series-parallel attribution relationship. The graph structure characterizes the association relationships between image patch locations. This graph structure consists of multiple graph nodes and edge connections between graph nodes. Each graph node corresponds to a specific image patch location, and the edge connections between graph nodes are determined based on spatial proximity and series-parallel attribution relationships—when two image patch locations are spatially adjacent or belong to the same series-parallel branch in electrical connection, an edge connection is established between the corresponding two graph nodes. Gating mapping is performed on the image patch features corresponding to each graph node to obtain the initial gating response coefficients of each graph node.
[0116] Using the initial gating response coefficients of each graph node as node features, graph convolutional propagation is performed on the graph structure to obtain the propagated gating response coefficients of each graph node. Graph convolutional propagation is used to propagate and aggregate the gating coefficients of each graph node along the edge connections in the graph structure through multiple iterations. In each iteration, each graph node passes its current gating coefficient to its neighboring nodes, while simultaneously receiving gating coefficients from neighboring nodes and weighting and aggregating them. After multiple rounds of graph convolutional propagation, the gating response coefficients of each graph node incorporate the influence of other nodes in its local neighborhood graph structure, thus smoothing the differences in gating coefficients between physically connected adjacent battery cell positions.
[0117] Based on the propagated gating response coefficients of each graph node, a local response coefficient matrix is constructed. By introducing the constraints of the graph structure, the spatial distribution of the gating coefficients better matches the physical structure of the photovoltaic module—cells located in the same string or spatially adjacent cells obtain similar gating coefficients, while the changes in gating coefficients between different strings better reflect the actual situation of physical isolation. This solves the problem that purely data-driven gating learning cannot utilize prior knowledge of the inherent physical structure of the photovoltaic module.
[0118] In some embodiments, step S105 above, which involves performing context-aware projection on the normalized image patch features to obtain the feature activity of each image patch feature, may further include: for each image patch location, extracting multiple neighboring image patch features within a preset neighborhood centered on that location to obtain a local context feature matrix; performing a third linear projection transformation on the local context feature matrix to obtain the local context features of that image patch location; performing a fourth linear projection transformation on the image patch features corresponding to that image patch location to obtain the self-projection features of that image patch location; fusing the local context features and the self-projection features to obtain the neighborhood aggregation features of that image patch location; and performing a nonlinear projection transformation on the neighborhood aggregation features of each image patch location to obtain the feature activity of each image patch feature.
[0119] For each image patch location, features from multiple neighboring image patches within a preset neighborhood are extracted, centered on that location, to obtain a local context feature matrix. The preset neighborhood is a spatial window centered on the target image patch location, covering its surrounding neighboring image patches; the size of this window determines the spatial scope of context awareness. The local context feature matrix represents the set of feature vectors for each neighboring image patch within the neighborhood surrounding the target image patch location. The rows of this matrix correspond to the locations of each neighboring image patch, and the columns correspond to each feature channel.
[0120] A third linear projection transformation is applied to the local context feature matrix to obtain the local context features of the image patch location. This third linear projection transformation applies a learnable linear mapping to the local context feature matrix, aggregating features from multiple adjacent image patches into a single vector. This aggregation is achieved by weighted summation or mean pooling along the row directions of the local context feature matrix followed by the linear projection transformation. The local context features characterize the overall feature patterns of the neighborhood surrounding the target image patch location, serving as a contextual reference benchmark for evaluating the salience of that location relative to its neighborhood.
[0121] A fourth linear projection transformation is applied to the image patch features corresponding to the image patch location to obtain the self-projection features of that image patch location. The fourth linear projection transformation applies a learnable linear mapping to the feature vector of the image patch location itself, transforming it from the original feature space to the same space as the local context features. The self-projection features characterize the feature information of the image patch location itself, serving as a self-comparison benchmark for evaluating the saliency of that location relative to its neighborhood. The local context features and self-projection features are fused to obtain the neighborhood aggregation features of the image patch location. The fusion method involves concatenating the local context features and self-projection features along the channel dimension and then compressing them to the original feature dimension through a learnable linear transformation, or performing element-wise addition or weighted summation of the two. The neighborhood aggregation features characterize the fused representation of the image patch location after integrating its own feature information and neighborhood context feature information, in which the differences between the own features and neighborhood features are explicitly encoded.
[0122] A nonlinear projection transformation is applied to the neighborhood aggregation features of each image patch location to obtain the feature activity of each patch location. The nonlinear projection transformation applies a learnable transformation to the neighborhood aggregation features, including at least one layer of linear transformation and activation function mapping, mapping the neighborhood aggregation features from the feature space to a scalar activity space. Feature activity characterizes the significance of each image patch location relative to its local neighborhood. A higher feature activity value is obtained when the feature pattern of the location itself differs significantly from the overall feature pattern of its neighborhood; a lower feature activity value is obtained when the feature pattern of the location itself tends to be consistent with the overall feature pattern of its neighborhood. By combining the difference between its own features and the features of its neighborhood context to evaluate feature activity, the determination of activity has local contrast perception capability—isolated noise points, although having a strong response, have suppressed feature activity because their neighborhood lacks similar responses; within continuous fault regions, all locations have strong responses and their neighborhoods have similar responses, so feature activity is preserved normally, solving the problem that simply relying on the absolute response intensity cannot distinguish between isolated noise points and continuous fault regions.
[0123] In the embodiments of this specification, the terms "first linear projection transformation" to "fifth linear projection transformation" are used to distinguish linear projection transformation operations located in different processing stages, performing different functions, or receiving different input sources. The linear projection transformations defined by each number are mathematically identical; they all apply a learnable linear mapping to the input features, that is, transforming the features from the input space to the output space through matrix multiplication. The first linear projection transformation is configured in the spatial context encoding stage, applying a linear mapping to the aggregated feature vector obtained after global mean pooling, transforming it into global fault source features. Its input is the aggregated feature vector, and its output is the global fault source features. The second linear projection transformation is configured in the gating mapping stage, applying a linear mapping to the features of each image patch, projecting the features from the original feature space to the gating score space, and outputting an initial gating response matrix. Its input is the features of each image patch, and its output is the initial gating score. The third linear projection transformation is configured in the context-aware projection stage, applying a linear mapping to the local context feature matrix, aggregating multiple adjacent image patch features into a single local context feature vector. Its input is the local context feature matrix, and its output is the local context features. The fourth linear projection transformation is also configured in the context-aware projection stage. It applies a linear mapping to the feature vector of a single image patch location, transforming it into its own projected features. Its input is the feature vector of the single image patch, and its output is the projected features. The fifth linear projection transformation is configured in the weighted global mean pooling stage. It applies a linear mapping to the features of each image patch, projecting them from the original feature space to the embedding space for subsequent similarity calculation. Its input is the feature vector of each image patch, and its output is the embedded features. There is no sequential or hierarchical relationship between the above numbering; the numbering is only used to distinguish the different configuration positions and functional directions of the transformations.
[0124] Corresponding to the above identification model, the embodiments of this specification provide an identification model for photovoltaic module fault identification, such as... Figure 2 As shown.
[0125] In some embodiments, the recognition model may include an image patch embedding module, a visual feature extractor, a wave residual adapter, a first classifier, and a second classifier.
[0126] The image patch embedding module is used to divide the input thermal infrared image into several non-overlapping image patches, map each image patch into a feature vector, and output the image patch feature sequence.
[0127] The visual feature extractor receives the image patch feature sequence output by the image patch embedding module. After being transformed layer by layer by the multi-layer transformer encoder, it outputs global category features and local patch feature matrices.
[0128] The input of the wave residual adapter is connected to the output of the local block feature matrix of the visual feature extractor. It is used to perform spatial context encoding and gated propagation processing on the local block feature matrix and output the enhanced local block feature matrix.
[0129] The input of the first classifier is connected to the global category feature output of the visual feature extractor, and is used to classify the global category features.
[0130] The input of the second classifier is connected to the output of the enhanced local block feature matrix of the wave residual adapter, and is used to classify the enhanced local block feature matrix.
[0131] The final fault category prediction result is obtained by weighted fusion of the outputs of the first classifier and the second classifier.
[0132] Through the collaborative work of the above modules, the recognition model retains the global discrimination capability of the visual feature extractor while using the wave residual adapter to perform targeted enhancement of local fault semantics, thereby achieving accurate identification of photovoltaic module faults.
[0133] The visual feature extractor can employ a self-supervised pre-trained model based on a transformer architecture. The image patch embedding module divides the input thermal infrared image into blocks and performs linear projection, with position encoding superimposed on the features of each image patch to preserve spatial location information. During training, the parameters of the visual feature extractor remain frozen and do not participate in gradient updates, thus preserving its general visual representation capabilities learned from hundreds of millions of general image datasets and avoiding the destruction of pre-trained knowledge by downstream task training. The wave residual adapter includes a normalization unit, a spatial context encoding unit, and a gating mapping unit, connected sequentially. It extracts the global fault source from local block features and propagates it directionally to various local locations through a gating mechanism, while suppressing the background response with a damping coefficient. The first and second classifiers each contain a linear classification layer. The recognition model uses a dual-branch structure to simultaneously utilize global category features and local features enhanced by the wave residual adapter for classification decisions, taking into account both scene-level semantic judgment and local fault detail perception. This solves the problem of insufficient response to local small-area fault features after freezing the base model, and achieves sensitive identification of local thermal faults while retaining the strong generalization ability of the base model. At the same time, it improves the stability of cross-scenario deployment and the recall rate of small-area faults.
[0134] In some embodiments, the training method for the wave residual adapter, the first classifier, and the second classifier includes: acquiring thermal infrared image samples of photovoltaic modules and their corresponding fault category labels; inputting the thermal infrared image samples into a visual feature extractor, whose parameters are frozen during training; inputting the local block feature matrix output by the visual feature extractor into the wave residual adapter; processing the global category features output by the visual feature extractor into the first classifier; processing the enhanced local block feature matrix output by the second classifier; updating the gradient only for the learnable parameters in the wave residual adapter, the learnable parameters of the first classifier, the learnable parameters of the second classifier, and the fusion coefficients of the outputs of the two classifiers; the learnable parameters in the wave residual adapter include linear projection layer parameters of spatial context encoding, linear projection layer parameters of gated mapping, and damping coefficients; calculating the cross-entropy loss based on the fusion result of the outputs of the first classifier and the second classifier, using the fault category label as a supervision signal; and performing end-to-end training on the wave residual adapter, the first classifier, and the second classifier with the objective of minimizing the cross-entropy loss.
[0135] The frozen state of the visual feature extractor ensures that the general visual representation learned during the large-scale pre-training phase is not destroyed during downstream task training, thus maintaining cross-scene transferability. The learnable parameters of the wave residual adapter serve to adapt the general visual representation to the photovoltaic thermal infrared fault recognition task, learning the ability to extract and propagate fault semantics from local block features through end-to-end supervised training. The first classifier learns the scene-level fault category discrimination boundary from global category features, while the second classifier learns the mapping relationship from local fault semantics to fault categories from enhanced local block features. The learnable design of the fusion coefficients allows the contribution of the two branches to be adaptively determined by the training data, without the need for manual setting of weight ratios. The cross-entropy loss uses the fault category label as a supervision signal to drive the parameters of the wave residual adapter, the two classification branches, and the fusion coefficients to optimize in the direction of maximizing classification accuracy. By freezing the backbone and training only the lightweight adaptation module, this training method can avoid overfitting problems under the condition of limited sample size of the thermal infrared photovoltaic dataset, achieving the dual goals of training stability and performance reliability.
[0136] The learnable parameters of the wave residual adapter can include the parameters of the spatial context-encoded linear projection layer, the parameters of the gated mapping linear projection layer, and the damping coefficient. The learnable parameters of the first classifier include the weight matrix and bias vector of its linear classification layer. The learnable parameters of the second classifier include the weight matrix and bias vector of its linear classification layer. All parameters of the visual feature extractor are not updated during training. The gradient backpropagation path during training is as follows: the cross-entropy loss passes sequentially through the learnable parameters of the second classifier, the learnable parameters of the wave residual adapter, and the learnable parameters and fusion coefficients of the first classifier. The forward computation of the visual feature extractor does not participate in gradient backpropagation. The parameters of the spatial context-encoded linear projection layer map the aggregated feature vector after global mean pooling to global fault source features. This projection layer learns during training how to extract the most discriminative collective fault indication information from all image patch features. The parameters of the gated mapping linear projection layer map each image patch feature to an initial gating score. This projection layer learns during training the absorption intensity of global fault source features at each local location. The damping coefficient, as a learnable scalar parameter, adaptively adjusts the suppression ratio of the original local features during training. The linear classification layer parameters of the first and second classifiers learn mappings from the global semantic space and the local augmented semantic space to the fault category space, respectively, during training. The fusion coefficient learns the optimal weight ratio of the global branch and the local branch in the final decision during training.
[0137] As can be seen from the technical solutions provided in the embodiments of this specification above, the embodiments of this specification can input the thermal infrared images of photovoltaic modules collected by UAVs into a visual feature extractor to obtain global category features and local block feature matrices; perform spatial context encoding on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image; perform gating mapping on each image block feature to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image; couple the global fault source features and the local response matrix to obtain propagation enhancement features; determine residual increment features based on the propagation enhancement features and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of background components in the thermal infrared image; correct the local block feature matrix based on the residual increment features; and determine the fault category of the photovoltaic module using a preset classifier based on the global category features and the corrected local block feature matrix. Global fault source features are obtained by spatial context encoding local block features and coupled with local response matrices for propagation. This enables the global fault semantics to be directionally transmitted to various local spatial locations, effectively solving the problem of fault signals being submerged in a large background. Furthermore, background components are suppressed by using damping coefficients to further highlight fault-related features and improve the sensitivity to weak thermal anomalies. Additionally, the local block feature matrix is modified based on residual incremental features, achieving incremental enhancement of fault semantics without destroying the original feature structure and preserving the original discriminative ability of the basic model. Finally, accurate identification of photovoltaic module fault categories is achieved through global and local fusion classification.
[0138] The following are two specific embodiments of this specification: Example 1: Implementation of fault identification based on large-scale photovoltaic fault dataset.
[0139] This embodiment uses a large-scale UAV thermal infrared photovoltaic fault dataset as the experimental object.
[0140] (1) Data set and preprocessing.
[0141] A large-scale thermal infrared photovoltaic (PV) fault dataset was used, comprising eight PV power plants and covering four application scenarios: agriculture, rooftop, fisheries, and centralized systems. The dataset contains 5579 cropped thermal infrared images of PV modules, covering ten fault categories including battery overheating, junction box overheating, sub-string open circuits, shading of modules, and dust accumulation strips. The data preprocessing steps were as follows: all thermal infrared images were uniformly scaled to 224×224 pixels; single-channel grayscale thermal infrared images were copied three times along the channel dimension to construct three-channel pseudo-color images; the dataset was divided into training, validation, and test sets in an 8:1:1 ratio; training set images were augmented using random horizontal flipping, random vertical flipping, and random rotation (angle range 0 to 90 degrees); all images were normalized according to channel means of 0.485, 0.456, and 0.406 and standard deviations of 0.229, 0.224, and 0.225.
[0142] (2) Network configuration.
[0143] The backbone network employs a large-scale self-supervised pre-trained visual model, using a basic-size visual transformer structure with an image patch size of 14×14 pixels and a feature dimension of 768. It consists of 12 transformer encoder layers, with all backbone network parameters frozen and not participating in gradient updates. In the wave residual adapter, both linear projection matrices have 768-dimensional input and output dimensions, and the damping coefficient is initialized to 0.1. The dual-branch classification heads are both single-layer linear layers with an input dimension of 768 and an output dimension of 10, corresponding to ten fault categories. The fusion coefficients of the two branches are initialized to 1.0.
[0144] (3) Training configuration.
[0145] The optimizer employs an adaptive moment estimation algorithm, with an initial learning rate of 0.0001 and a weight decay coefficient of 0.01. Cosine annealing is used for learning rate scheduling, with a minimum learning rate of 0.000001. The batch size is 32, and training is conducted for 100 epochs. Cross-entropy loss is used as the loss function. All experiments were trained on a single GPU.
[0146] (4) Evaluation indicators.
[0147] The recognition performance is comprehensively evaluated using four metrics: overall accuracy, macro-average precision, macro-average recall, and macro-average composite score. The macro-average metric assigns equal weight to each category to fully account for the impact of class imbalance.
[0148] (5) Implementation results.
[0149] The method of this invention achieves an overall precision of 90.97%, a macro average precision of 87.49%, a macro average recall of 86.31%, and a macro average comprehensive score of 86.71% on this dataset. It outperforms all the comparison methods in terms of overall precision, recall, and comprehensive score, and the number of trainable parameters is only 4.16 million.
[0150] Example 2: Fault identification implementation based on low-resolution multi-category thermal infrared dataset.
[0151] This embodiment uses a low-resolution, multi-category UAV and manned aircraft thermal infrared photovoltaic fault dataset as the experimental object.
[0152] (1) Data set and preprocessing.
[0153] A dataset of photovoltaic module faults was collected using mid-wave and long-wave thermal infrared sensors mounted on manned and unmanned aircraft. The dataset contains 20,000 thermal infrared images with an original resolution of 24×40 pixels, covering eleven categories of anomalies (single-cell hotspots, multi-cell hotspots, thin-film module hotspots, bypass diode activation, shading, contamination, etc.) and one normal category, for a total of twelve categories. The preprocessing steps are as follows: all images were bilinearly interpolated and upsampled to 224×224 pixels; single-channel images were duplicated into three channels; the training, validation, and test sets were divided in an 8:1:1 ratio; the training set underwent random flipping and rotation for data augmentation; and normalization was performed using the standard mean and standard deviation.
[0154] (2) Network configuration.
[0155] The backbone network configuration is the same as in Example 1, and the backbone parameters are completely frozen. The output dimension of the dual-branch classification head is adjusted to 12, corresponding to twelve categories. The wave residual adapter structure is completely consistent with that in Example 1, requiring no structural adjustments for the dataset, demonstrating the adaptability of the method in this embodiment to different datasets.
[0156] (3) Training configuration.
[0157] The training configuration is consistent with that of Example 1, with an initial learning rate of 0.0001, a batch size of 32, 100 training rounds, and a cosine annealing learning rate scheduling strategy.
[0158] (4) Implementation results.
[0159] The method of this invention achieves an overall precision of 79.97%, a macro average recall of 61.95%, and a macro average comprehensive score of 63.11% on this dataset, outperforming all comparable methods in both overall precision and recall. Comparative experiments show that the pure attention mechanism structure achieves an overall precision of only 49.95% on this low-resolution dataset, approaching random levels, while the method of this invention maintains stable performance due to the strong generalization ability of the frozen backbone, fully demonstrating the superiority of the method of this invention in challenging scenarios with low resolution and small sample sizes.
[0160] Based on the above-described photovoltaic module fault identification method, this specification also proposes embodiments of a photovoltaic module fault identification device. For example... Figure 3 As shown, the photovoltaic module fault identification device 300 may specifically include the following modules: Extraction module 301 is used to input the thermal infrared image of photovoltaic module collected by UAV into the visual feature extractor to obtain global category features and local block feature matrix; Encoding module 302 is used to perform spatial context encoding on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image; The mapping module 303 is used to perform gated mapping on the features of each image block to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image; Coupling module 304 is used to couple global fault source characteristics with local response matrix to obtain propagation enhancement characteristics; The first determining module 305 is used to determine the residual increment characteristics based on the propagation enhancement characteristics and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of the background components of the thermal infrared image. The correction module 306 is used to correct the local block feature matrix based on the residual increment characteristics; The second determination module 307 is used to determine the fault category of the photovoltaic module using a preset classifier based on the global category features and the corrected local block feature matrix.
[0161] In some embodiments, the above-mentioned input of the thermal infrared image of the photovoltaic module collected by the UAV into the visual feature extractor to obtain a global category feature and a local block feature matrix includes: The thermal infrared image is divided into multiple image blocks; Each image patch is mapped to a feature vector, and the position encoding of each feature vector is superimposed to obtain the image patch feature sequence; A learnable category feature vector is inserted at the beginning of the image patch feature sequence to obtain the input sequence; The input sequence is fed into the multi-layer transform encoder of the visual feature extractor. Each layer transform encoder is used to transform the input sequence layer by layer through a multi-head self-attention and feedforward network. The head vector and the tail vector output by the last encoder are used as the global category feature and the local block feature matrix, respectively.
[0162] In some embodiments, the above-described spatial context encoding of multiple image block features in the local block feature matrix to obtain global fault source features includes: Global mean pooling is performed on multiple image block features along the dimension of the local block feature matrix to obtain fault aggregation features; The first linear projection transformation is performed on the fault aggregation features to obtain the global fault source features.
[0163] In some embodiments, the above-described gating mapping of features of each image patch to obtain a local response matrix includes: A second linear projection transformation is performed on the features of each image patch to obtain the initial gated response matrix; The initial gated response matrix is activated to obtain the local response matrix.
[0164] In some embodiments, the above method further includes: Normalize the feature matrix of the local block; The process of spatial context encoding multiple image block features in the local block feature matrix to obtain global fault source features includes: Spatial context encoding is performed on the normalized features of multiple image patches to obtain global fault source features; The process of gating and mapping the features of each image patch to obtain the local response matrix includes: Gated mapping is performed on the normalized features of each image patch to obtain the local response matrix.
[0165] In some embodiments, the above-mentioned coupling of global fault source features and local response matrix yields propagation enhancement features, including: The global fault source features are extended to the same spatial dimension as the features of each image patch to obtain broadcast source features; The propagation increment features of each image patch location are obtained by multiplying the broadcast source features element-wise with the local response matrix. The propagation increment features are superimposed on the features of each image patch to obtain the propagation enhancement features.
[0166] In some embodiments, the above-described expansion of the global fault source features to the same spatial dimension as the features of each image patch to obtain broadcast source features includes: Based on the semantic correlation between the features of each image patch and the global fault source features, the correlation coefficient of each image patch location is obtained; the semantic correlation is used to characterize the degree of matching between the local fault indication information carried by each image patch feature and the global fault indication information. Based on the correlation coefficient, the propagation modulation coefficient of each image block location is determined; the propagation modulation coefficient is positively correlated with the correlation coefficient; the propagation modulation coefficient is used to characterize the intensity scaling ratio when the global fault source features are transmitted to each image block location. The broadcast source characteristics are obtained by multiplying the global fault source characteristics element by element with the propagation modulation coefficients of each image block location.
[0167] In some embodiments, determining the residual increment characteristics based on the propagation enhancement characteristics and the damping coefficient includes: The background suppression features are obtained by multiplying the normalized features of each image patch by the damping coefficient. Subtracting the background suppression feature from the propagation enhancement feature yields the residual increment feature.
[0168] In some embodiments, the above-mentioned multiplication of the normalized image patch features with the damping coefficient to obtain background suppression features includes: Context-aware projection is performed on the normalized features of each image patch to obtain the feature activity of each image patch feature. Based on feature activity, the spatial modulation coefficients of each image block feature are determined; the spatial modulation coefficients are positively correlated with feature activity; the spatial modulation coefficients are used to characterize the local adjustment intensity of the damping coefficient. By fusing the spatial modulation coefficients and damping coefficients of each image patch feature, the modulation damping coefficients of each image patch feature are obtained. The background suppression features are obtained by multiplying the normalized features of each image patch by the corresponding modulation damping coefficient.
[0169] In some embodiments, the above-described context-aware projection of the normalized image patch features to obtain the feature activity of each image patch feature includes: For each image patch location, extract features from multiple neighboring image patches within a preset neighborhood centered on that location to obtain a local context feature matrix; The local context feature matrix is subjected to a third linear projection transformation to obtain the local context features of the image patch location; A fourth linear projection transformation is performed on the image patch features corresponding to the image patch location to obtain the self-projection features of the image patch location. By fusing local context features and self-projection features, the neighborhood aggregation features of the image patch location are obtained; A nonlinear projection transformation is performed on the neighborhood aggregation features of each image patch location to obtain the feature activity of each image patch feature.
[0170] In some embodiments, the above-mentioned correction of the local block feature matrix based on the residual increment features includes: The residual increment feature is added element by element to the local block feature matrix to obtain the corrected local block feature matrix.
[0171] In some embodiments, the above-mentioned method of determining the fault category of a photovoltaic module using a preset classifier based on global category features and a corrected local block feature matrix includes: The global category features are input into the first classifier to obtain the global fault identification result; The corrected local block feature matrix is input into the second classifier to obtain the local fault identification result; The fault category of the photovoltaic module is determined by fusing the global fault identification results and the local fault identification results.
[0172] As can be seen from the technical solutions provided in the embodiments of this specification above, the embodiments of this specification can input the thermal infrared images of photovoltaic modules collected by UAVs into a visual feature extractor to obtain global category features and local block feature matrices; perform spatial context encoding on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image; perform gating mapping on each image block feature to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image; couple the global fault source features and the local response matrix to obtain propagation enhancement features; determine residual increment features based on the propagation enhancement features and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of background components in the thermal infrared image; correct the local block feature matrix based on the residual increment features; and determine the fault category of the photovoltaic module using a preset classifier based on the global category features and the corrected local block feature matrix. Global fault source features are obtained by spatial context encoding local block features and coupled with local response matrices for propagation. This enables the global fault semantics to be directionally transmitted to various local spatial locations, effectively solving the problem of fault signals being submerged in a large background. Furthermore, background components are suppressed by using damping coefficients to further highlight fault-related features and improve the sensitivity to weak thermal anomalies. Additionally, the local block feature matrix is modified based on residual incremental features, achieving incremental enhancement of fault semantics without destroying the original feature structure and preserving the original discriminative ability of the basic model. Finally, accurate identification of photovoltaic module fault categories is achieved through global and local fusion classification.
[0173] This specification also provides a computer device for a photovoltaic module fault identification method, including a processor and a memory for storing processor-executable instructions. Specifically, the processor can perform the following tasks according to the instructions: inputting a photovoltaic module thermal infrared image collected by a UAV into a visual feature extractor to obtain global category features and a local block feature matrix; performing spatial context encoding on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image; performing gating mapping on each image block feature to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image; coupling the global fault source features and the local response matrix to obtain propagation enhancement features; determining residual increment features based on the propagation enhancement features and a damping coefficient; the damping coefficient characterizes the degree of suppression of background components in the thermal infrared image; correcting the local block feature matrix based on the residual increment features; and determining the fault category of the photovoltaic module using a preset classifier based on the global category features and the corrected local block feature matrix.
[0174] To execute the above instructions more accurately, please refer to... Figure 4 As shown in the embodiments of this specification, another specific computer device 400 is also provided, wherein the computer device 400 includes a network communication port 401, a processor 402 and a memory 403, and the above structures are connected by internal cables so that the various structures can perform specific data interaction.
[0175] The processor 402 can be specifically used to: input the thermal infrared image of the photovoltaic module collected by the UAV into a visual feature extractor to obtain a global category feature and a local block feature matrix; perform spatial context encoding on multiple image block features in the local block feature matrix to obtain a global fault source feature; the global fault source feature includes global fault indication information at multiple spatial locations in the thermal infrared image; perform gating mapping on each image block feature to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of the global fault indication information at each spatial location in the thermal infrared image; couple the global fault source feature with the local response matrix to obtain a propagation enhancement feature; determine the residual increment feature based on the propagation enhancement feature and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of the background components of the thermal infrared image; correct the local block feature matrix based on the residual increment feature; and determine the fault category of the photovoltaic module using a preset classifier based on the global category feature and the corrected local block feature matrix.
[0176] The memory 403 can be used to store the corresponding instruction program.
[0177] In this embodiment, the network communication port 401 can be a virtual port bound to different communication protocols, thereby enabling the sending or receiving of different data. For example, the network communication port can be a port responsible for web data communication, a port responsible for FTP data communication, or a port responsible for email data communication. Furthermore, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM or CDMA; it can also be a Wi-Fi chip; or it can be a Bluetooth chip.
[0178] In this embodiment, the processor 402 can be implemented in any suitable manner. For example, the processor can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc. This specification is not limiting.
[0179] In this embodiment, the memory 403 includes volatile memory and non-volatile memory. The memory 403 can include multiple layers. In digital systems, anything that can store binary data can be a memory; in integrated circuits, a circuit with storage function but no physical form is also called a memory, such as RAM, FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory stick, TF card, etc.
[0180] Furthermore, embodiments of this specification also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described... Figure 1 The instructions for the photovoltaic module fault identification method are shown.
[0181] Furthermore, embodiments of this specification provide a computer program product comprising a computer program that, when executed by a processor, implements the above-described... Figure 1 The photovoltaic module fault identification method shown.
[0182] It should be understood that in the various embodiments of this specification, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this specification.
[0183] It should also be understood that, in the embodiments of this specification, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this specification generally indicates that the preceding and following related objects have an "or" relationship.
[0184] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0185] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0186] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0187] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational tasks to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The task is a function specified in one or more boxes.
[0188] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for identifying faults in photovoltaic modules, characterized in that, include: The thermal infrared images of photovoltaic modules collected by a drone are input into a visual feature extractor to obtain global category features and local block feature matrices. This process includes: dividing the thermal infrared image into multiple image blocks; mapping each image block to a feature vector and superimposing positional encodings on each feature vector to obtain an image block feature sequence; inserting a learnable category feature vector at the beginning of the image block feature sequence to obtain an input sequence; inputting the input sequence into a multi-layer transform encoder of the visual feature extractor, where each transform encoder layer transforms the input sequence layer by layer using a multi-head self-attention and feedforward network; and using the beginning and end vectors output by the last encoder layer as the global category feature and local block feature matrix, respectively. Spatial context encoding is performed on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image; Gated mapping is performed on the features of each image block to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image; The propagation enhancement features are obtained by coupling global fault source features with local response matrices, including: obtaining the correlation coefficient of each image patch location based on the semantic correlation between each image patch feature and the global fault source features; the semantic correlation is used to characterize the degree of matching between the local fault indication information carried by each image patch feature and the global fault indication information; determining the propagation modulation coefficient of each image patch location based on the correlation coefficient; the propagation modulation coefficient is positively correlated with the correlation coefficient; the propagation modulation coefficient is used to characterize the intensity scaling ratio when the global fault source features are transmitted to each image patch location; multiplying the global fault source features and the propagation modulation coefficients of each image patch location element by element to obtain broadcast source features; multiplying the broadcast source features and the local response matrix element by element to obtain the propagation increment features of each image patch location; and superimposing the propagation increment features onto the image patch features to obtain the propagation enhancement features. The residual increment characteristics are determined based on the propagation enhancement characteristics and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of background components in thermal infrared images. The local block feature matrix is corrected based on the residual increment characteristics; Based on the global category features and the corrected local block feature matrix, a preset classifier is used to determine the fault category of the photovoltaic module.
2. The method according to claim 1, characterized in that, The process of spatial context encoding multiple image block features in the local block feature matrix to obtain global fault source features includes: Global mean pooling is performed on multiple image block features along the dimension of the local block feature matrix to obtain fault aggregation features; The first linear projection transformation is performed on the fault aggregation features to obtain the global fault source features.
3. The method according to claim 1, characterized in that, The process of gating and mapping the features of each image patch to obtain the local response matrix includes: A second linear projection transformation is performed on the features of each image patch to obtain the initial gated response matrix; The initial gated response matrix is activated to obtain the local response matrix.
4. The method according to claim 1, characterized in that, The method further includes: Normalize the feature matrix of the local block; The process of spatial context encoding multiple image block features in the local block feature matrix to obtain global fault source features includes: Spatial context encoding is performed on the normalized features of multiple image patches to obtain global fault source features; The process of gating and mapping the features of each image patch to obtain the local response matrix includes: Gated mapping is performed on the normalized features of each image patch to obtain the local response matrix.
5. The method according to claim 4, characterized in that, The determination of residual increment characteristics based on propagation enhancement characteristics and damping coefficient includes: The background suppression features are obtained by multiplying the normalized features of each image patch by the damping coefficient. Subtracting the background suppression feature from the propagation enhancement feature yields the residual increment feature.
6. The method according to claim 5, characterized in that, The step of multiplying the normalized features of each image patch by the damping coefficient to obtain background suppression features includes: Context-aware projection is performed on the normalized features of each image patch to obtain the feature activity of each image patch feature. Based on feature activity, the spatial modulation coefficients of each image block feature are determined; the spatial modulation coefficients are positively correlated with feature activity; the spatial modulation coefficients are used to characterize the local adjustment intensity of the damping coefficient. By fusing the spatial modulation coefficients and damping coefficients of each image patch feature, the modulation damping coefficients of each image patch feature are obtained. The background suppression features are obtained by multiplying the normalized features of each image patch by the corresponding modulation damping coefficient.
7. The method according to claim 6, characterized in that, The step of performing context-aware projection on the normalized image patch features to obtain the feature activity of each image patch feature includes: For each image patch location, extract features from multiple neighboring image patches within a preset neighborhood centered on that location to obtain a local context feature matrix; The local context feature matrix is subjected to a third linear projection transformation to obtain the local context features of the image patch location; A fourth linear projection transformation is performed on the image patch features corresponding to the image patch location to obtain the self-projection features of the image patch location. By fusing local context features and self-projection features, the neighborhood aggregation features of the image patch location are obtained; A nonlinear projection transformation is performed on the neighborhood aggregation features of each image patch location to obtain the feature activity of each image patch feature.
8. The method according to claim 1, characterized in that, The step of correcting the local block feature matrix based on the residual increment features includes: The residual increment feature is added element by element to the local block feature matrix to obtain the corrected local block feature matrix.
9. The method according to claim 1, characterized in that, The step of determining the fault category of the photovoltaic module using a preset classifier based on global category features and the corrected local block feature matrix includes: The global category features are input into the first classifier to obtain the global fault identification result; The corrected local block feature matrix is input into the second classifier to obtain the local fault identification result; The fault category of the photovoltaic module is determined by fusing the global fault identification results and the local fault identification results.
10. A photovoltaic module fault identification device, characterized in that, include: The extraction module is used to input the thermal infrared image of photovoltaic modules collected by the UAV into the visual feature extractor to obtain global category features and local block feature matrices. This includes: dividing the thermal infrared image into multiple image blocks; mapping each image block to a feature vector and superimposing positional encodings on each feature vector to obtain an image block feature sequence; inserting a learnable category feature vector at the beginning of the image block feature sequence to obtain an input sequence; inputting the input sequence into the multi-layer transform encoder of the visual feature extractor, where each transform encoder layer transforms the input sequence layer by layer using a multi-head self-attention and feedforward network; and using the beginning and end vectors output by the last encoder layer as the global category feature and local block feature matrix, respectively. The encoding module is used to perform spatial context encoding on multiple image block features in the local block feature matrix to obtain global fault source features; the global fault source features include global fault indication information at multiple spatial locations in the thermal infrared image; The mapping module is used to perform gated mapping on the features of each image block to obtain a local response matrix; the local response matrix includes multiple local response coefficients used to characterize the absorption intensity of global fault indication information at each spatial location in the thermal infrared image; A coupling module is used to couple global fault source features with local response matrices to obtain propagation enhancement features. This includes: obtaining a correlation coefficient for each image patch location based on the semantic correlation between each image patch feature and the global fault source features; the semantic correlation characterizes the degree of matching between the local fault indication information carried by each image patch feature and the global fault indication information; determining the propagation modulation coefficient for each image patch location based on the correlation coefficient; the propagation modulation coefficient is positively correlated with the correlation coefficient; the propagation modulation coefficient characterizes the intensity scaling ratio when the global fault source features are transmitted to each image patch location; multiplying the global fault source features element-wise with the propagation modulation coefficients of each image patch location to obtain broadcast source features; multiplying the broadcast source features element-wise with the local response matrix to obtain propagation increment features for each image patch location; and superimposing the propagation increment features onto the image patch features to obtain propagation enhancement features. The first determining module is used to determine the residual increment characteristics based on the propagation enhancement characteristics and the damping coefficient; the damping coefficient is used to characterize the degree of suppression of the background components of the thermal infrared image. The correction module is used to correct the local block feature matrix based on the residual increment characteristics; The second determination module is used to determine the fault category of the photovoltaic module using a preset classifier based on the global category features and the corrected local block feature matrix.
11. A computer device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the method of any one of claims 1-9.
12. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed, implement the steps of the method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Fault prediction method for multi-modal cross-attention enhancement graph neural network
CN120871803A
Photovoltaic fault identification method based on unmanned aerial vehicle inspection
CN121258964A