Intelligent photovoltaic hot spot fault detection method, system, medium, device and terminal
By employing a knowledge distillation module and feature fusion technology through collaborative training between teachers and students, the problem of balancing accuracy and lightweight design in photovoltaic hot spot fault detection has been solved, improving the detection performance of small-scale targets and enabling efficient and accurate photovoltaic power station inspections.
Patent Information
- Application Number
- CN202310028596.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2043-01-09
AI Technical Summary
Existing photovoltaic hot spot fault detection methods struggle to balance detection accuracy with lightweight models. They lack the ability to represent infrared hot spot target features, have poor performance in detecting small-scale targets, and suffer from high false alarm and missed alarm rates, failing to meet the needs of large-scale, high-efficiency inspections of photovoltaic power plants.
A knowledge distillation module trained collaboratively by teachers and students is adopted, which is combined with the CSPHN backbone network and BiMAF model to enhance feature representation ability through local and global feature fusion. Furthermore, a CgT module is constructed before the decoupled prediction head to improve the performance of small-scale target detection.
It enables efficient and accurate photovoltaic hot spot fault detection in complex environments, reduces network computing costs, improves detection efficiency and accuracy, and ensures the safe and stable operation of photovoltaic power generation systems.
Smart Images

Figure CN117095311B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent photovoltaic inspection technology, and in particular relates to an intelligent photovoltaic hot spot fault detection method, system, medium, equipment and terminal. Background Technology
[0002] Currently, with the rapid development of the economy and society, my country's energy shortage is becoming increasingly severe, and traditional fossil fuels face numerous problems such as non-renewable resources and serious pollution. Therefore, my country's energy supply pattern is shifting from fossil fuels to renewable and clean energy, vigorously developing new clean energy industries. Among these, solar photovoltaic power generation, as a key technology in the new energy field, has received widespread attention due to its advantages such as being green and low-carbon, flexible in its use, and widely distributed.
[0003] Photovoltaic (PV) power generation systems convert solar radiation into electrical energy directly using the photovoltaic effect of solar cell semiconductor materials. However, over long-term use, PV modules inevitably accumulate dust, fallen leaves, and other obstructions, causing localized shading and fluctuations in the current and voltage of some individual cells. This leads to increased localized power consumption within the PV module, resulting in a hotspot effect. The hotspot effect exhibits localized temperature rise; in severe cases, it can melt solder joints and damage the grid lines, shortening the lifespan of the PV module. Simultaneously, the hotspot effect can obscure some normally functioning solar panels, preventing them from operating properly and reducing the efficiency of large-scale PV power plants. Therefore, to ensure the long-term stable operation of PV power generation systems, regular hotspot fault detection of PV modules is essential.
[0004] Supported by increasingly mature power technologies, photovoltaic (PV) power plants are now widely distributed in vast, sunny, and complex areas, such as mountain power plants and hydroelectric power plants. Their vast coverage areas and scattered distribution due to terrain limitations present significant challenges to power system maintenance. Existing traditional manual inspection methods are limited by maintenance costs, working conditions, and labor efficiency, making it difficult to meet the high-precision, high-efficiency requirements of PV module maintenance. With the development of the drone industry, advancements in computer vision technology, and the increasing maturity of intelligent PV inspection systems, data can be collected using visible light and thermal infrared sensors, and machine vision technology can be used to detect hot spot faults in PV modules. This allows for automated, rapid, precise, and intelligent operation and maintenance of PV power plants across a wide area. Compared to traditional manual inspection methods, drone inspection effectively improves the accuracy and efficiency of hot spot detection.
[0005] Photovoltaic hot spot fault detection algorithms, as a key technology in intelligent photovoltaic inspection systems, can be mainly divided into two categories: fault diagnosis methods based on electrical characteristics and image diagnosis methods based on computer vision technology. Among them, fault diagnosis methods based on electrical characteristics utilize mathematical models or machine learning to monitor and analyze electrical data such as the output voltage, output current, and output power of photovoltaic modules, which can accurately diagnose hot spot faults in photovoltaic modules. However, these algorithms require the addition of external circuitry for auxiliary detection, resulting in high costs and low efficiency, making it difficult to meet the needs of actual large-scale photovoltaic power plant operation and maintenance.
[0006] Image diagnostic methods based on computer vision technology can be divided into two types: traditional image processing algorithms and deep learning algorithms. Traditional hot spot fault detection algorithms use sliding window technology to extract features manually and combine them with a classifier to complete the fault detection task. Although such algorithms can accurately detect targets in certain specific scenarios, they are difficult to capture the high-level semantic information of hot spot targets and have poor generalization ability in complex environments. Compared with traditional hot spot fault detection algorithms based on manual feature extraction, hot spot image diagnostic algorithms based on deep learning can automatically learn target features by utilizing the excellent feature extraction and nonlinear fitting capabilities of convolutional neural networks, thereby accurately determining whether there are hot spot targets in images or video sequences and providing precise location. They show excellent detection results in terms of detection accuracy, speed, and generalization ability. However, at present, various photovoltaic hot spot fault detection methods generally believe that algorithm detection accuracy and model lightweighting are mutually exclusive. Therefore, in order to ensure inspection performance, large-scale detection networks with high computational costs are usually selected to perform fault location and identification on photovoltaic inspection data. Although this approach improves the accuracy of photovoltaic power plant operation and maintenance, the inspection efficiency is generally low and cannot meet the needs of real-time detection. Meanwhile, due to limitations in the drone's shooting angle, altitude, and performance, the collected photovoltaic data is relatively blurry, and most hot spot faults are small-scale distorted targets, making it difficult for the fault detection algorithm to effectively extract the features of the target to be detected. This limits the detection performance of infrared small-scale hot spot faults, and consequently, the network reliability cannot be guaranteed.
[0007] Therefore, how to balance the real-time efficiency and algorithm accuracy of the detection network, improve the ability to express infrared hot spot target features, improve the detection performance of small-scale targets, and reduce the false alarm rate and missed alarm rate of the network are the technical problems that urgently need to be solved at this stage.
[0008] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:
[0009] (1) Photovoltaic power stations are widely distributed in complex areas with vast areas and abundant sunshine, such as mountain power stations and floating power stations. Their coverage area is huge, and they are generally scattered and disorderly due to terrain limitations, which brings great challenges to the maintenance of power systems.
[0010] (2) Traditional manual inspection methods are limited by operation and maintenance costs, working conditions and labor efficiency, making it difficult to meet the high-precision and high-efficiency operation and maintenance tasks of photovoltaic modules.
[0011] (3) Fault diagnosis methods based on electrical characteristics require the addition of external circuit auxiliary detection, which is costly and inefficient, making it difficult to meet the needs of actual large-scale photovoltaic power plant operation and maintenance.
[0012] (4) Traditional hot spot image detection algorithms have difficulty capturing the high-level semantic information of hot spot targets. Therefore, they can only accurately detect targets in certain specific scenarios and have poor generalization ability in complex environments.
[0013] (5) The detection performance of the algorithm is usually positively correlated with the number of detection model parameters and the network depth. However, the larger the number of model parameters and the deeper the network, the greater the computational cost of the algorithm and the longer the running time. Therefore, balancing detection accuracy and model lightweighting is a major challenge for current intelligent photovoltaic inspection algorithms.
[0014] (6) Due to the varied shooting angles of drones, the captured images are prone to feature distortion and severe noise, which makes it difficult to effectively express the hot spot target features and thus affects the final detection accuracy.
[0015] (7) Due to objective factors such as the drone's shooting altitude and the geographical location of the data, hot spot targets are generally small in size, randomly distributed, and lack appearance information such as texture, shape, and color. As a result, the prior anchor box is less adaptable to small targets, leading to missed alarms and false alarms.
[0016] (8) In dense multi-target scenarios, current intelligent photovoltaic inspection algorithms are unable to accurately detect all targets in images or videos, resulting in varying degrees of missed and false detections. This leads to low reliability of the inspection system and directly limits the performance of intelligent photovoltaic inspection. Summary of the Invention
[0017] To address the problems existing in the prior art, this invention provides an intelligent photovoltaic hot spot fault detection method, system, medium, equipment, and terminal.
[0018] This invention is implemented as follows: an intelligent photovoltaic hot spot fault detection method, the intelligent photovoltaic hot spot fault detection method comprising:
[0019] From the perspective of model lightweighting, a knowledge distillation module for collaborative training of "teacher + student" was constructed, which improved the detection accuracy of the algorithm while ensuring the model's inference efficiency. Considering the positive correlation between target feature extraction and detection performance, a backbone network based on CSPHN was designed, which combines the advantages of CSPNet and HorNet structures to significantly enhance the algorithm's feature representation ability without consuming too much network memory. Inspired by the hierarchical cognitive principle of visual neurons, a BiMAF model was constructed, which strengthens the neck network's ability to aggregate target features by fusing input features from global and local perspectives using a parallel method. At the same time, in order to further enhance the network's detection performance for dense small targets, a CgT module was proposed before decoupling the prediction head. By mining static and dynamic context information to generate an attention matrix, the detection system can still promptly detect hot spot faults in photovoltaic power generation systems in dense small target scenarios and eliminate and repair them, thereby ensuring the safe and stable operation of photovoltaic power generation systems.
[0020] Furthermore, the intelligent photovoltaic hot spot fault detection method includes the following steps:
[0021] Step 1: Based on the concept of knowledge distillation, a new collaborative training model combining teacher and student networks is designed. This leverages the advantages of the teacher's deep detection network to improve algorithm detection accuracy, while combining the small parameter count of the student network to enhance inference efficiency, thus balancing detection accuracy and model lightweighting requirements.
[0022] Step 2: For the UAV intelligent inspection dataset provided by a photovoltaic power station operation and maintenance company, a student network based on the YOLOX algorithm is built. By constructing an anchorless model, the number of network parameters is greatly reduced. Combined with a decoupled prediction head, the network convergence speed is accelerated while improving the detection accuracy of the algorithm.
[0023] Step 3: Considering the positive correlation between target feature extraction and detection performance, a CSPHN (HorNet-based cross-stage partial network) module is constructed in the teacher backbone network. While retaining the beneficial inductive bias of the convolutional neural network, it uses gated convolution and recursive design to obtain high-order spatial interaction information similar to Transformer to enhance the algorithm's feature representation ability for hot spot targets.
[0024] Step four: To extract more discriminative target features from the high-level semantic information and low-level localization information output by the backbone network, a BiMAF (Bi-branch multi-level feature adaptive fusion) module was designed in the teacher's neck network. This module aggregates multi-level features from both global and local perspectives in a parallel fusion manner, enabling the network to selectively focus on regions containing target saliency information, thereby enhancing the feature aggregation capability of the neck network.
[0025] Step 5: To further improve the detection performance of small-scale targets, CgT(g) is constructed before the teacher decoupling prediction head. n The conv-based contextual transformer module improves the expressive power of the Transformer architecture by performing self-attention learning operations and mining contextual information between input keys of the two-dimensional feature map, thereby enhancing the accuracy of hot spot small-scale target detection in various dense scenes.
[0026] Furthermore, the construction of the novel knowledge distillation collaborative training model includes:
[0027] (1.1) To address the imbalance between foreground and background features in teacher and student networks, a local distillation function is designed to separate image foreground and background information and guide student networks to focus on important pixel and channel features.
[0028] (1.2) Local distillation severs the correlation between the foreground and background of an image, making it difficult for the detection network to capture global information and limiting the network's detection performance. To address this, a global distillation function is proposed to reconstruct the relationships between different pixels and pass them back from the teacher network to the student network, thereby compensating for the global information lost during local distillation.
[0029] (1.3) By combining local and global distillation methods, information from a large teacher network is inherited into a compact student network, enabling it to achieve powerful performance without incurring additional costs during inference.
[0030] Furthermore, the specific process of step (1.1) includes:
[0031] (1.1.1) For the feature map F with horizontal and vertical coordinates i and j respectively, construct a binary mask M. i,j Separate the foreground and background of the image as shown in the following formula:
[0032]
[0033] In the formula, r is the target ground truth bounding box;
[0034] (1.1.2) To balance the loss of the detection network on targets of different scales and foreground and background regions, a scale mask S is set. i,j :
[0035]
[0036] In the formula, H r and W r These represent the height and width of the target bounding box, respectively.
[0037] (1.1.3) Constructing a spatial attention mask and channel attention mask To improve the performance of model distillation:
[0038]
[0039] In the formula, H, W, and C represent the feature height, width, and channel, respectively; T is the temperature hyperparameter; F c and F i,j These represent the feature information of the c-th channel and the feature information of size i×j, respectively;
[0040] (1.1.4) During training, a binary mask M is used. i,j Scale mask S i,j And attention mask A S (F) and A C (F) Guide students to learn about the teacher's key network space and channel information online, thereby constructing the following feature loss function L. fea and attention loss function L at :
[0041]
[0042] In the formula, and These represent the feature graphs of the teacher and student networks, respectively; f(·) is the result of F S Adjust to F T Reconstruction operators of the same dimension; α, β, and γ are hyperparameters that balance various losses; l represents the l1 norm operator; and These represent the spatial attention mask and channel attention mask for the teacher and student networks, respectively.
[0043] (1.1.5) The final local distillation function is obtained by calculating the feature loss and attention loss:
[0044] L focal =L fea +L at
[0045] Furthermore, the specific process of step (1.2) includes:
[0046] The global loss L is obtained by capturing global information of the image using the GcBlock module. global :
[0047] L global =λ·Σ(R(F) T )-R(F S )) 2
[0048] In the formula, λ represents the balancing loss hyperparameter; R(F) represents the global feature information, which can be expressed as:
[0049]
[0050] In the formula, W k W v1 and W v2 LN represents a convolutional layer; LN represents layer normalization; N p Represents the number of feature pixels; ReLU represents the linear activation function.
[0051] Furthermore, the design of the student network includes:
[0052] The YOLOX network includes a backbone network, a neck network, and a decoupled prediction head;
[0053] (1) The backbone mainly consists of three parts: Focus, CSPNet (Cross Stage Partial Network), and SPP (Spatial Pyramid Pooling). Among them, the Focus slicing module not only expands the receptive field of the network, but also effectively suppresses the loss of image feature information, thereby accelerating the training speed. The CSPNet structure aims to solve the problem of excessive computational cost caused by the repetition of gradient information during network optimization. The SPP structure uses multi-level pooling operations to expand the receptive field of the backbone network and improve the multi-scale feature fusion capability.
[0054] (2) The Neck network borrows from CSPNet to construct the CSP2_X structure to enhance the network's feature fusion capability, and uses the FPN+PAN feature pyramid structure to aggregate features at different scales. Among them, FPN transmits strong semantic information from top to bottom, and PAN transmits strong localization information from bottom to top.
[0055] (3) Finally, a decoupled prediction head is constructed. By decoupling the classification branch that focuses on texture information and the localization branch that focuses on edge information, the spatial misalignment problem is effectively solved and the convergence speed is improved.
[0056] Furthermore, the construction of the CSPHN module includes:
[0057] (1) From the perspective of network structure design, the feature mapping of the input layer is divided into two parts and the two are merged by using a cross-stage hierarchical structure, thereby reducing the network memory usage and enhancing the learning ability of the convolutional neural network.
[0058] (2) To enable the backbone network to extract both local and global features, the residual units in the original CSPNet model were replaced with HorNet modules. This module inherits the meta-architecture of the Transformer model's spatial hybrid layer and feedforward network cascade, and utilizes gated recurrent convolution (g... n Conv) captures higher-order spatial interactions in feature maps, thereby avoiding the secondary computational complexity caused by multiple dot products during the execution of multi-head attention mechanisms, and improving the backbone network's ability to extract features of the detected target.
[0059] The specific process includes:
[0060] First, assume g n Conv's input features are Then a set of projection features p0 and It can be obtained through the following formula:
[0061]
[0062] In the formula, φ(·) represents the projection operator, and
[0063] Then, perform gated recursive convolution, with the following formula:
[0064] p k+1 =DW k (q k )⊙g k (p k ) / α,k=0,1,...,n-1
[0065] In the formula, α is the scaling factor, {DW k} represents a set of depthwise convolutions, ⊙ represents a dot product operation, and {g k} is used to match dimensions of different sequences, as shown in the following formula:
[0066]
[0067] Finally, feature projection is performed after the top-level recursive operation to obtain g. n The output of Conv.
[0068] Furthermore, the design of the BiMAF model includes:
[0069] (2.1) Global Feature Fusion Branch: Global pooling is typically used for global encoding of spatial information, but it compresses global spatial information into channel descriptors, making it difficult to retain feature localization information. Therefore, to enable the feature fusion module to capture remote spatial interaction features with precise location information, the global pooling operation is transformed into a one-to-one feature operator. Next, to fully utilize the feature information obtained after the coordinate information embedding operation, concatenation and convolution operations are performed, and the global feature weights are output using the sigmoid normalization function. Finally, the global feature fusion result, Output, is obtained through weight allocation. g .
[0070] (2.2) Construct local feature weights using 1×1 convolution and sigmoid normalization function, and then use weight allocation operation to obtain the local feature fusion result Output. l .
[0071] (2.3) The BiMAF model aggregates global and local feature outputs through parallel fusion to obtain the final feature fusion result.
[0072] Output = [Output] g Output l ]
[0073] Furthermore, the calculation process of the global feature fusion result in step (2.1) includes:
[0074] (2.1.1) Assume the multi-level feature inputs are as follows: The element-wise AND operation is performed as follows:
[0075] M = Input1 + Input2
[0076] Using two pooling kernels of size (H,1) and (1,W) to encode each channel along the horizontal and vertical directions respectively, the output of the c-th channel with height h is obtained. and the output of the c-th channel with width w
[0077]
[0078] (2.1.2) Perform the following concatenation and convolution operations:
[0079] f=δ(F1([z h ,z w ]))
[0080] In the formula, [z h ,z w] represents a connection operation along the spatial dimension; δ is the activation function; f is the intermediate feature map obtained by encoding in the horizontal and vertical directions; F1 represents a 1×1 convolution operation;
[0081] By decomposing f along the spatial dimensions, we obtain two independent tensors: f h ∈R C / r×H and f h ∈R C / r×H Where r is the downsampling ratio; simultaneously, two 1×1 convolutions F are performed. h and F w The operation ensures that the two independent tensors have the same number of channels, thus obtaining the global feature weights W. g :
[0082] W g =sigmoid(F h ([f h ]))×sigmoid([f w (2.1.3) Output of the global feature fusion branch g for:
[0083] Output g =Input1×W g +Input2×(1-W g )
[0084] The calculation process of the local feature fusion result in step (2.2) includes:
[0085] Local feature fusion is performed on the input features, and the local feature weights W l Represented as:
[0086] W l =sigmoid(F1(δ(BN(F1(M)))))
[0087] Furthermore, the output of the local feature fusion branch is represented as:
[0088] Output l =I1×W l +I2×(1-W l )
[0089] Furthermore, the design of the CgT model includes:
[0090] Assume the two-dimensional input feature map is The key (key, K), query (query, Q), and value (value, V) are defined as K = XW. k Q = XW q and V=XW q Among them, Wq and W v W is a linear transformation matrix consisting of 1×1 convolutions; k It is a linear transformation matrix composed of groups of convolutions with a kernel size of 3×3. By performing context encoding on the input feature map, it can reflect the static context features f. s Simultaneously, auxiliary self-attention learning is used to explore the interaction between key contextual features and query features, thereby constructing a dynamic attention weight matrix A as follows:
[0091] A = g n conv([K,Q])
[0092] In the formula, g n Conv represents the gated recursive convolution operator.
[0093] Then, the dynamic context feature f is calculated. d :
[0094] f d =A⊙V
[0095] By capturing static and dynamic context features, the following output is obtained:
[0096] output = f s +f d .
[0097] Another object of the present invention is to provide a photovoltaic hot spot fault detection system applying the aforementioned photovoltaic hot spot fault detection method, the photovoltaic hot spot fault detection system comprising:
[0098] The data acquisition module is used to automatically plan the drone's patrol route by analyzing the geographical information and patrol range of the photovoltaic power station, and at the same time, it uses the thermal radiation imaging characteristics of the infrared sensor to collect inspection images and videos.
[0099] The data transmission module is used to transmit inspection data back to the ground control station and store it accordingly, taking advantage of the high speed and low latency of the 5G wireless network, so that the high-performance computer can process the data later.
[0100] The fault diagnosis module utilizes a dual-branch collaborative training algorithm based on a designed knowledge distillation mechanism to perform feature extraction, information aggregation, and fault location tasks on photovoltaic data. Finally, by combining image information and GPS (Global Positioning System) positioning data, the hot spot fault diagnosis result is obtained.
[0101] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program, which, when executed by the processor, causes the processor to perform the steps of the intelligent photovoltaic hot spot fault detection method.
[0102] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the intelligent photovoltaic hot spot fault detection method.
[0103] Another objective of this invention is to provide an information data processing terminal for implementing the intelligent photovoltaic hot spot fault detection system.
[0104] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:
[0105] First, addressing the technical problems existing in the prior art and the difficulty in solving them, and closely combining the technical solution to be protected by this invention with the results and data from the research and development process, this paper provides a detailed and in-depth analysis of how the technical solution of this invention solves the technical problems, and the inventive technical effects brought about after solving the problems. The specific description is as follows:
[0106] This invention aims to provide a new approach for photovoltaic hot spot fault detection, which will have a profound impact on accelerating the intelligent transformation of the photovoltaic inspection technology field. Specifically, addressing the four major technical problems in the current field of photovoltaic hot spot fault detection—the difficulty in balancing algorithm accuracy and model lightweighting, the inability to effectively represent infrared hot spot target features, the difficulty in guaranteeing small-scale target detection performance, and the difficulty in reducing network false alarm and missed alarm rates—this invention provides a dual-branch collaborative training photovoltaic hot spot fault detection system based on a knowledge distillation mechanism. Based on the concept of knowledge distillation, a new teacher + student network collaborative training model is designed. This model leverages the advantages of the teacher's deep detection network to improve algorithm detection accuracy and combines the small parameter count of the student network to improve algorithm inference efficiency, thus balancing detection accuracy and model lightweighting requirements. Specifically, the YOLOX-s algorithm is selected as the student network, which significantly reduces the number of network parameters by constructing an anchor-free model and, combined with a decoupled prediction head, improves algorithm detection accuracy while accelerating network convergence. To enhance the backbone network's ability to represent hot spot targets, a CSPHN module was constructed within the teacher network; a BiMAF module was designed to improve the neck network's feature aggregation capability; and a CgT module was proposed to further improve the detection performance of small-scale hot spots in various dense scenarios. Finally, image information was integrated with GPS positioning data to obtain the final hot spot fault diagnosis result. This system, combining UAV aerial photography and computer vision technology, can autonomously complete large-scale, rapid, and precise intelligent hot spot fault detection tasks, ensuring the safe and stable operation of photovoltaic power generation systems.
[0107] Second, considering the technical solution as a whole or from a product perspective, the technical effects and advantages of the technical solution to be protected by this invention are specifically described as follows:
[0108] This invention integrates UAV inspection technology, sensor technology, and target detection algorithms to propose a photovoltaic hot spot fault detection network suitable for complex environments, named KDBiDet (Abi-branch collaborative training algorithm based on knowledge distillation for photovoltaic hot spot detection system). From the perspective of model lightweighting, this detection network constructs a knowledge distillation module for teacher-student collaborative training to improve algorithm inference efficiency. Considering the positive correlation between target feature extraction and detection performance, a CSPHN-based backbone network is designed, combining the advantages of CSPNet and HorNet structures to significantly enhance the algorithm's feature representation capability without consuming excessive network memory. Inspired by the hierarchical cognitive principle of visual neurons, a BiMAF model is constructed, using a parallel method to fuse input features from both global and local perspectives, strengthening the neck network's ability to aggregate target features. Furthermore, to further enhance the network's detection performance for dense small targets, a CgT module is proposed before decoupling the prediction head. This module generates an attention weight matrix by mining static and dynamic contextual information, enabling the detection system to promptly detect and eliminate hot spot faults in photovoltaic power generation systems even in dense small target scenarios. This network will enable "equipment to speak and the power grid to think," allowing for the timely detection, elimination, and repair of hot spot faults in photovoltaic power generation systems, thereby ensuring the safe and stable operation of these systems. With the continuous promotion and application of this product technology, this invention will accelerate the intelligent transformation of the photovoltaic inspection industry, thereby promoting the transformation and upgrading of the power grid.
[0109] Third, as supporting evidence of the inventiveness of this invention, it is also reflected in the following important aspects:
[0110] (1) This invention relies on deep learning algorithms and utilizes a cloud-edge-device integrated architecture to propose a dual-branch collaborative training detection algorithm based on a knowledge distillation mechanism for photovoltaic hot spot fault datasets. This algorithm can be extended to various sub-sectors of the smart photovoltaic inspection market, enabling timely detection, elimination, and repair of various faults in photovoltaic power generation systems, thereby ensuring the safe and stable operation of photovoltaic power generation systems. This has significant research value and commercial potential. In the future, with the continuous promotion and application of product technologies, intelligent photovoltaic inspection systems will realize "equipment that can speak, and the power grid that can think," thereby accelerating the intelligent transformation of the photovoltaic inspection industry and promoting the transformation and upgrading of the power grid.
[0111] (2) This invention addresses four major technical challenges in the current field of photovoltaic hot spot fault detection: the difficulty in balancing algorithm accuracy and model lightweighting, the inability to effectively represent infrared hot spot target features, the difficulty in guaranteeing small-scale target detection performance, and the difficulty in reducing network false alarm and missed alarm rates. It proposes a dual-branch collaborative training photovoltaic hot spot fault detection algorithm based on a knowledge distillation mechanism. From the perspective of model lightweighting, a teacher + student collaborative training knowledge distillation module is constructed to improve algorithm inference efficiency. Considering the positive correlation between target feature extraction and detection performance, a CSPHN-based backbone network is designed to enhance the algorithm's ability to represent the features of the target to be detected. Inspired by the hierarchical cognitive principle of visual neural networks, a BiMAF module is constructed to strengthen the neck network's ability to aggregate features of multi-scale targets. Finally, to further suppress the adverse effects of missed and false detections on detection accuracy, a CgT-structured prediction head is proposed. By performing self-attention learning operations and mining the contextual information between the input keys of the two-dimensional feature map, the model's detection accuracy is improved, thereby enabling the automatic completion of large-scale, rapid, precise, and intelligent operation and maintenance tasks of photovoltaic power plants. In summary, this invention can solve the technical problems that urgently need to be addressed at present.
[0112] (3) Currently, various intelligent photovoltaic inspection systems generally believe that algorithm detection accuracy and model lightweighting are mutually exclusive. Therefore, to ensure inspection performance, large-scale detection networks with higher computational costs are usually selected to locate and identify faults in the collected photovoltaic inspection data. Although this approach improves the accuracy of photovoltaic power station operation and maintenance, the inspection efficiency is generally low and cannot meet the real-time detection requirements. To address this, inspired by the concept of knowledge distillation, this paper utilizes both local and global distillation methods to inherit information from a large teacher network into a compact small student network. This leverages the advantages of the teacher deep detection network to improve algorithm detection accuracy and combines the small parameter characteristics of the student network to accelerate algorithm operation. Consequently, the student network achieves powerful performance without incurring additional costs during inference, ensuring large-scale, rapid, precise, and intelligent operation and maintenance tasks for photovoltaic power stations. Attached Figure Description
[0113] Figure 1 This is a flowchart of the intelligent photovoltaic hot spot fault detection system provided in an embodiment of the present invention;
[0114] Figure 2 This is a flowchart of the intelligent photovoltaic hot spot fault detection algorithm provided in an embodiment of the present invention;
[0115] Figure 3 This is a schematic diagram of the intelligent photovoltaic hot spot fault detection method provided in an embodiment of the present invention;
[0116] Figure 4 This is a schematic diagram of the student detection network provided in an embodiment of the present invention;
[0117] Figure 5 This is a schematic diagram of the global distillation principle provided in an embodiment of the present invention;
[0118] Figure 6 This is a schematic diagram of the CSPHN structure provided in an embodiment of the present invention;
[0119] Figure 7 This is a schematic diagram of the BiMAF structure provided in an embodiment of the present invention;
[0120] Figure 8 This is a framework diagram of the intelligent photovoltaic hot spot fault detection system provided in an embodiment of the present invention;
[0121] Figure 9 These are schematic diagrams of three typical scenarios provided in the embodiments of the present invention; wherein, Figure (a) is a schematic diagram of a cluttered small target scenario, Figure (b) is a schematic diagram of a noise interference scenario, and Figure (c) is a schematic diagram of a motion blur scenario;
[0122] Figure 10 Figure 1 is a schematic diagram comparing the detection results of different algorithms provided in the embodiments of the present invention; wherein, Figure (a) is a schematic diagram of the detection results of each algorithm in a cluttered small target scene, Figure (b) is a schematic diagram of the detection results of each algorithm in a noisy interference scene, and Figure (c) is a schematic diagram of the detection results of each algorithm in a motion blur scene. Detailed Implementation
[0123] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0124] To enable those skilled in the art to fully understand how the present invention is specifically implemented, this section provides an explanatory description of the embodiments that expand upon the technical solutions of the claims.
[0125] like Figure 1 As shown, the intelligent photovoltaic hot spot fault detection system provided in this embodiment of the invention includes the following steps:
[0126] S101 Data Acquisition Module: This system can automatically plan the drone's patrol route by analyzing the geographical information and patrol range of the photovoltaic power station, and at the same time, it can collect inspection images and videos by utilizing the thermal radiation imaging characteristics of the infrared sensor.
[0127] S102 Data Transmission Module: Leveraging the high speed and low latency advantages of 5G wireless network, the inspection data is transmitted back to the ground control station and stored accordingly, so that the high-performance computer can process the data later.
[0128] The S103 fault diagnosis module utilizes the designed KDBiDet detection algorithm to perform feature extraction, information aggregation, and fault location tasks on photovoltaic data. Finally, by combining image information and GPS positioning data, it obtains the fault diagnosis results for hot spots.
[0129] like Figure 2 As shown, the intelligent photovoltaic hot spot fault detection method provided in this embodiment of the invention includes the following steps:
[0130] S101 constructs a new collaborative training model for knowledge distillation, which leverages the advantages of the teacher's deep network to improve the algorithm's detection accuracy and combines the small parameter characteristics of the student network to improve the algorithm's inference speed, thereby balancing the requirements of detection accuracy and model lightweighting.
[0131] S102, the student network uses the YOLOX detection network. By constructing an anchorless model, the number of network parameters is greatly reduced. Combined with the decoupled prediction head, the network convergence speed is accelerated while improving the detection accuracy of the algorithm. It consists of three parts: the backbone network that completes feature extraction, the neck network that performs feature extraction, and the decoupled prediction head that performs classification and regression tasks.
[0132] S103, the teacher network, builds a backbone network based on the YOLOX detection network and the CSPHN module. While retaining the beneficial inductive bias of the convolutional neural network, it uses gated convolution and recursive design to obtain high-order spatial interaction information similar to the Transformer, thus taking into account both local and global feature representation capabilities.
[0133] S104. A BiMAF module is designed in the teacher network to aggregate multi-level features from both global and local perspectives in a parallel fusion manner, enabling the network to selectively focus on regions containing target saliency information, thereby enhancing the feature aggregation capability of the neck network.
[0134] S105 constructs a CgT module before the teacher decouples the prediction head. By performing self-attention learning operations and mining the contextual information between the input keys of the two-dimensional feature map, it improves the expressive power of the Transformer architecture, thereby enhancing the accuracy of hot spot target detection in various dense scenes.
[0135] As a preferred embodiment, such as Figure 3 As shown, the intelligent photovoltaic hot spot fault detection method provided in this embodiment of the invention specifically includes:
[0136] 1. Knowledge Distillation Collaborative Training Module
[0137] Knowledge distillation aims to inherit information from a large teacher network into a compact student network, enabling it to achieve robust performance without incurring additional costs during inference. However, during distillation, imbalances exist between the spatial and channel feature maps of the teacher and student networks, which can negatively impact distillation. To address this issue, this invention proposes a local-global knowledge distillation structure, consisting of local and global distillation components.
[0138] 1.1 Local distillation
[0139] To address the imbalance between foreground and background features in teacher and student networks, a local distillation function is designed to separate foreground and background information in images and guide student networks to focus on important pixel and channel features.
[0140] (1) Constructing a binary mask M i,j Separate the image foreground (real bounding box r) and background as shown in the following formula.
[0141]
[0142] (2) Large-scale targets have a large number of pixels, resulting in a correspondingly large loss value, which makes it difficult for distillation to improve the detection performance of small-scale targets. Therefore, in order to balance the loss of the detection network for targets of different scales and foreground and background regions, a scale mask S is set. i,j :
[0143]
[0144] In the formula, H r and W r These represent the height and width of the target's true bounding box, respectively.
[0145] (3) To enable students to focus their attention on important spatial and channel information of the teacher's network, a spatial attention mask is constructed. and channel attention mask To improve the performance of the model distillation.
[0146]
[0147]
[0148] In the formula, H, W, and C represent the feature height, width, and channel, respectively; T is the temperature hyperparameter; F c and F i,j These represent the feature information of the c-th channel and the feature information of size i×j, respectively.
[0149] During training, a binary mask M is used. i,j Scale mask S i,j And attention mask A S(F) and A C (F) Guide students to learn about the teacher's key network space and channel information online, thereby constructing the following feature loss function and attention loss function:
[0150]
[0151] In the formula, and These represent the feature graphs of the teacher and student networks, respectively; f(·) is the result of F S Adjust to F T Reconstruction operators of the same dimension; α, β, and γ are hyperparameters that balance various losses; l represents the l1 norm operator; and Let represent the spatial attention mask and channel attention mask for the teacher and student networks, respectively. The final local distillation function is obtained by calculating the feature loss and attention loss:
[0152] L focal =L fea +L at (7)
[0153] 1.2 Global Distillation
[0154] Local distillation functions improve distillation performance by separating foreground and background information and forcing the student network to learn key spatial and channel features from the teacher network. However, this distillation method severs the correlation between the foreground and background, making it difficult for the detection network to capture global image information and limiting its detection performance. Therefore, this invention proposes a global distillation function to reconstruct the relationships between different pixels and pass them back from the teacher network to the student network, thereby compensating for the global information lost during local distillation. Figure 4 As shown.
[0155] This invention utilizes the GcBlock module to capture global image information, thereby setting the following global loss function:
[0156] L global =λ·Σ(R(F) T )-R(F S )) 2
[0157]
[0158] In the formula, W k W v1 and W v2 LN represents a convolutional layer; LN represents layer normalization; N p λ represents the number of feature pixels; λ represents the balancing loss hyperparameter; R(F) represents the global feature information.
[0159] In summary, this invention constructs a novel knowledge distillation collaborative training model, leveraging the advantages of the teacher deep detection network to improve algorithm detection accuracy, and combining the small parameter count of the student network to enhance algorithm inference efficiency. Specifically, YOLOX-s is used as the student network, and an improved YOLOX-l teacher network is proposed to achieve a balance between detection accuracy and model lightweighting.
[0160] 2. Student Assessment Network Module
[0161] This invention selects YOLOX as the student detection network, which utilizes an anchorless structure to significantly reduce manually designed hyperparameters. Simultaneously, the combination with a decoupled prediction head effectively improves detection accuracy and accelerates network convergence. Figure 5 As shown, it consists of three parts: the backbone network, the neck network, and the decoupled prediction head.
[0162] (1) The backbone mainly consists of three parts: Focus, CSPNet, and SPP. Among them, the Focus slicing module not only expands the receptive field of the network, but also effectively suppresses the loss of image feature information, thereby accelerating the training speed. The CSPNet structure aims to solve the problem of excessive computational cost caused by the repetition of gradient information during network optimization. The SPP structure uses multi-level pooling operations to expand the receptive field of the backbone network and improve the multi-scale feature fusion capability.
[0163] (2) The Neck network borrows from CSPNet to construct the CSP2_X structure to enhance the network's feature fusion capability, and uses the FPN+PAN feature pyramid structure to aggregate features at different scales. Among them, FPN conveys strong semantic information from top to bottom, and PAN conveys strong localization information from bottom to top.
[0164] (3) Finally, a decoupled prediction head is constructed. By decoupling the classification branch that focuses on texture information and the localization branch that focuses on edge information, the spatial misalignment problem is effectively solved and the convergence speed is improved.
[0165] 3. Teacher Assessment Network Module
[0166] 3.1 Backbone Network Based on CSPHN Structure
[0167] Convolutional neural networks (CNNs) possess two main characteristics: local equivariance and translational equivariance. By focusing on neighboring information in local feature maps and applying the same processing rules to different regions, they obtain inductive biases, achieving high detection performance even on small datasets. However, CNNs have limitations in representing global features of the target during multiple convolution and pooling operations. To address this, the Transformer model was developed. By constructing cascaded self-attention modules, it can reflect complex high-order spatial interactions and long-distance feature dependencies, thus possessing excellent global modeling capabilities. However, this model performs poorly on small datasets and incurs secondary computational complexity due to multiple dot product operations.
[0168] Therefore, considering the positive correlation between the target feature extraction capability and detection performance of the backbone network, this invention proposes a cross-stage local network based on HorNet. While retaining the beneficial inductive bias of convolutional neural networks, it utilizes gated convolution and recursive design to obtain high-order spatial interaction information similar to that of the Transformer, thus taking into account global modeling characteristics and enhancing the algorithm's ability to represent the features of the target to be detected. The principle diagram is shown below. Figure 6 As shown.
[0169] like Figure 6 As shown, the CSPHN model, from a network architecture design perspective, reduces network memory usage and enhances the learning ability of convolutional neural networks by dividing the feature mapping of the input layer into two parts and merging them using a cross-stage hierarchical structure. Simultaneously, to ensure the backbone network can handle both local and global feature extraction, the residual units in the original CSPNet model are replaced with HorNet modules. This module inherits the meta-architecture of the Transformer model's spatial mixing layer and cascaded feedforward network, and utilizes gated recurrent convolution (g... n Conv) captures higher-order spatial interactions in feature maps, thereby avoiding the secondary computational complexity caused by multiple dot products during the execution of multi-head attention mechanisms, and improving the backbone network's ability to extract features of the detected target.
[0170] Assume g n Conv's input features are Then a set of projection features p0 and It can be obtained through the following formula:
[0171]
[0172] In the formula, φ(·) represents the projection operator, and
[0173] Next, gated recursive convolution is performed using formula (10).
[0174] pk+1 =DW k (q k )⊙g k (p k ) / α,k=0,1,...,n-1 (10)
[0175] In the formula, α is the scaling factor, {DW k} represents a set of depthwise convolutions, ⊙ represents a dot product operation, and {g k} is used to match dimensions of different sequences, as shown in the following formula.
[0176]
[0177] Finally, after performing the feature projection following the top-level recursive operation, g can be obtained. n The output of Conv. This invention combines the advantages of CSPNet and HorNet, enabling the backbone network to handle both local and global feature extraction capabilities. To evaluate the superiority of the CSPHN model constructed in this invention, the last layer of the original YOLOX backbone network was replaced with the CSPHN model, and compared with HorNet under the same conditions. The experimental results are shown in Table 1.
[0178] Table 1 Comparison of backbone networks based on HorNet and CSPHN architectures
[0179]
[0180] As shown in Table 1, the backbone network based on HorNet has a high number of network parameters and floating-point operations due to the introduction of gated recurrent convolution. In contrast, the CSPHN module constructed in this invention can effectively reduce the network memory usage and improve the backbone network's ability to capture local and global features by building a cross-stage hierarchical structure, thereby improving the algorithm's detection performance.
[0181] 3.2 Neck Network Based on Dual-Branch Multi-Level Feature Adaptive Fusion Model
[0182] Due to the limitations of long-distance drone imaging, the captured infrared images generally have high background complexity, and hotspot faults are mostly small- to medium-scale distorted targets. Therefore, to suppress feature differences between multiple scales, the original YOLOX network uses a "FPN+PAN" structure to transmit strong semantic information from the bottom up while transmitting strong localization information from the top down. However, this structure directly adds and aggregates feature maps of different scales after resizing, failing to fully utilize the cross-scale information at the input, thus affecting the final detection accuracy.
[0183] To address the aforementioned shortcomings, this invention draws inspiration from the AFF model and designs a dual-branch, multi-level feature adaptive fusion module. This module employs a parallel approach to fuse input features from both global and local perspectives, enabling the neck network to selectively focus on target regions containing salient information while suppressing other irrelevant background information. This improves the model's detection performance for small-scale targets. The principle diagram is shown below. Figure 7 As shown.
[0184] (1) Global Feature Fusion Branch
[0185] Assume the multi-level feature inputs are respectively The element-wise AND operation is performed as follows.
[0186] M = Input1 + Input2 (12)
[0187] Global pooling is typically used for global encoding of spatial information, but it compresses this information into channel descriptors, making it difficult to retain feature localization information. Therefore, to enable the feature fusion module to capture remote spatial interaction features with precise location information, the global pooling operation is transformed into a one-to-one one-dimensional feature operator. Specifically, this invention uses two pooling kernels of size (H,1) and (1,W) to encode each channel along the horizontal and vertical directions respectively, thus obtaining the output of the c-th channel with height h. and the output of the c-th channel with width w
[0188]
[0189] Next, in order to make full use of the feature information obtained after the coordinate information embedding operation, the following concatenation and convolution operations are performed on it.
[0190] f=δ(F1([z h ,z w (15)
[0191] In the formula, [z h ,z w ] represents the connection operation along the spatial dimension; δ is the activation function; f is the intermediate feature map obtained by encoding in the horizontal and vertical directions; F represents the 1×1 convolution operation.
[0192] By decomposing f along the spatial dimensions, we obtain two independent tensors: f h ∈R C / r×H and f h ∈R C / r×H Where r is the downsampling ratio. To ensure that the two independent tensors have the same number of channels, the following operation is performed:
[0193] g h=F h ([f h (16)
[0194] g w =F w ([f w (17)
[0195] Furthermore, the global feature weights W g It can be calculated using the following formula.
[0196] W g =sigmoid(g h )×sigmoid(g w (18)
[0197] Finally, the output of the global feature fusion branch is... g It can be represented as:
[0198] Output g =Input1×W g +Input2×(1-W g (19)
[0199] (2) Local feature fusion branch
[0200] Leveraging the excellent local feature extraction capabilities of convolutional layers, this invention performs local feature fusion on the input features, with local feature weights W. l It can be represented as:
[0201] W l =sigmoid(F1(δ(BN(F1(M))))) (20)
[0202] Next, the output of the local feature fusion branch can be expressed as:
[0203] Output l =Input1×W l +Input2×(1-W l ) (twenty one)
[0204] In summary, the BiMAF model aggregates global and local feature outputs in a parallel mode to obtain the final feature fusion result.
[0205] Output = [Output] g Output l ] (twenty two)
[0206] To verify the superiority of the designed BiMAF model, Table 2 presents the comparison results between the proposed BiMAF model and models that only construct the global feature fusion branch (GFF), only construct the local feature fusion branch (LFF), and global-local serial feature fusion (GLFF).
[0207] Table 2 Comparison results of multiple feature fusion modules
[0208]
[0209] 3.3 Prediction Head Based on CgT Model
[0210] The original YOLOX detection network employs an anchor-free framework, effectively reducing the number of network parameters and avoiding the problem of poor adaptability of prior anchor boxes to small targets. However, such anchor-free algorithms require the construction of auxiliary methods to obtain the final target detection box, which is prone to angle deviation and even semantic ambiguity of the target, thus affecting detection accuracy. To overcome these limitations, this invention draws inspiration from the CoT module and constructs a g-based... n The CoT model (CgT) of conv improves the expressive power of the Transformer architecture by performing self-attention learning and mining the contextual information between the input keys of the two-dimensional feature map, thereby improving the accuracy of hot spot target detection in various dense scenes.
[0211] Specifically, assuming the two-dimensional input feature map is The key (key, K), query (query, Q), and value (value, V) are defined as K = XW. k Q = XW q and V=XW q Among them, W q and W v W is a linear transformation matrix consisting of 1×1 convolutions; k It is a linear transformation matrix composed of groups of convolutions with a kernel size of 3×3. By performing context encoding on the input feature map, it can reflect the static context features f. s Simultaneously, auxiliary self-attention learning is used to explore the interaction between key contextual features and query features, thereby constructing a dynamic attention weight matrix A as follows:
[0212] A = g n conv([K,Q]) (23)
[0213] In the formula, gn Conv represents the gated recursive convolution operator.
[0214] Compared to traditional self-attention weight matrices that rely on isolated key-query pairs, dynamic attention weight matrices can enhance the feature learning ability of the self-attention mechanism at each spatial location by leveraging static context features. Then, the dynamic context features f can be computed. d .
[0215] f d =A⊙V (24)
[0216] By capturing static and dynamic context features, the following output can be obtained:
[0217] output = f s +f d (25)
[0218] It is worth noting that although the dot product operation of the CgT module has a quadratic complexity problem, its computational complexity is within a manageable range due to the relatively small size of the input feature map of the prediction layer (20×20, 40×40 and 80×80), and the detection performance is significantly improved.
[0219] In summary, a backbone network based on the CSPHN structure was constructed within the teacher network, a neck network with adaptive feature fusion was designed, and a prediction head based on the CgT module was proposed, achieving high-performance detection in high-density, small-target scenarios. Furthermore, by leveraging the concept of knowledge distillation, important information from the large teacher network was transferred to the compact student network, ultimately achieving a fast, precise, and intelligent photovoltaic hotspot fault detection task.
[0220] The intelligent photovoltaic hot spot fault detection system provided in this embodiment of the invention includes:
[0221] Supported by increasingly mature power technologies, photovoltaic power plants are now widely distributed in vast, sunny, and complex areas, such as mountain power plants and floating hydroelectric power plants. Their coverage area is enormous, and due to terrain limitations, they are generally scattered and disorganized, posing a significant challenge to power system maintenance. To address this issue, this invention proposes an intelligent automatic hot spot detection system based on unmanned aerial vehicles (UAVs). The system framework diagram is shown below. Figure 8 As shown.
[0222] First, this system analyzes the geographical information and patrol range of the photovoltaic power station to automatically plan the drone's patrol route. Simultaneously, it utilizes the thermal radiation imaging characteristics of infrared sensors to collect inspection images and videos. Second, leveraging the high speed and low latency advantages of 5G wireless networks, the inspection data is transmitted back to the ground control station and stored accordingly for subsequent processing by a high-performance computer. Finally, the designed deep learning algorithm is used to perform feature extraction, information aggregation, and fault location tasks on the photovoltaic data. Ultimately, by combining image information and GPS positioning data, the final hot spot fault diagnosis result is obtained.
[0223] This system is designed to autonomously perform hot spot fault detection tasks in large-scale, complex photovoltaic power plants. Compared to traditional manual inspection methods, which are characterized by high operation and maintenance costs, poor working conditions, and low labor efficiency, this system can perform high-precision and high-efficiency photovoltaic module operation and maintenance, promptly identify and eliminate potential safety hazards in photovoltaic modules, and is of great significance for ensuring the safe and stable operation of photovoltaic power plants.
[0224] To demonstrate the inventiveness and technical value of the technical solution of this invention, this section provides specific product or related technology application examples of the technical solution claimed.
[0225] This invention is based on deep learning algorithms and focuses on the compression, storage, transmission, and processing of images and videos. It addresses the complexity and danger of the geographical locations of photovoltaic equipment, the severity of photovoltaic panel operation and maintenance issues, and the thermal radiation imaging characteristics of infrared sensors. By equipping a drone with a thermal infrared sensor to collect photovoltaic inspection data and using machine vision technology to detect hot spot faults in photovoltaic modules, this invention proposes a data-driven photovoltaic hot spot fault detection system.
[0226] Once implemented, this system can automatically complete large-scale, rapid, precise, and intelligent operation and maintenance tasks for photovoltaic power plants. It has a profound impact on improving the power generation efficiency of photovoltaic power plants, reducing their operation and maintenance costs, and realizing intelligent operation and maintenance. Furthermore, it can be extended to various sub-fields of smart construction sites, such as photovoltaic multi-type defect detection, substation equipment inspection, and multi-target fault detection of transmission lines, and has significant application value.
[0227] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.
[0228] The embodiments of the present invention have achieved some positive results during the research and development or use process, and have indeed great advantages compared with the prior art. The following content describes them in conjunction with the data, charts and other information of the experimental process.
[0229] The model training and result analysis provided in this embodiment of the invention are as follows:
[0230] To evaluate the effectiveness of the proposed algorithm, this invention selected 500 fault images (640*512 pixels) from UAV inspection data. To ensure sample completeness, motion-blurred images were generated by moving the image 8 pixels counter-clockwise at a 30-degree angle, and noise interference images were generated using a Gaussian filter with a variance of 0.2, ultimately producing 1500 experimental images. This invention verified that the intelligent photovoltaic detection system can complete large-scale, rapid, and precise operation and maintenance tasks of photovoltaic power plants under various harsh conditions through ablation experiments on each improved module of the algorithm and qualitative and quantitative comparative analysis of seven classic algorithms.
[0231] 1. Ablation test
[0232] To verify the superiority of each improved module in the teacher network, ablation experiments were conducted on the dataset based on the YOLOX-1 detection network by adding different improvement strategies. All experiments used the same data samples and parameter settings. The comparison results are shown in Table 3.
[0233] Table 3 Comparison results of ablation experiments in teacher networks
[0234]
[0235] Table 3 shows that, compared to the original YOLOX-1 detection network, the constructed CSPHN module enhances the backbone network's ability to capture local and global features. The designed BiMAF module effectively improves the feature aggregation capability of the neck network by weighting multi-scale features. The proposed CgT module, combining static and dynamic contextual information, can improve detection performance in dense scenes. Furthermore, to verify the superiority of collaborative training of different modules, ablation experiments were conducted on CSPHN+BiMAF, CSPHN+CgT, and BiMAF+CgT networks. The results show that, compared to the original detection network, the AP index improved by 0.8%, 1.3%, and 1.0%, respectively. Finally, by integrating multiple improved modules, the detection accuracy of this study reaches 0.845, and the detection performance of small and medium-sized hot spot targets is significantly improved, enabling high-precision completion of photovoltaic power plant operation and maintenance tasks even under various harsh conditions.
[0236] Meanwhile, in order to evaluate the effectiveness of the knowledge distillation mechanism, this invention uses the YOLOX-s algorithm as the student network and the improved YOLOX-l algorithm as the teacher network.
[0237] Table 4 presents the comparison results of ablation experiments before and after the introduction of the knowledge distillation mechanism. The experimental results show that the constructed "teacher + student" collaborative training model can improve the detection accuracy of the algorithm while ensuring inference efficiency, thus balancing the requirements of algorithm detection accuracy and model lightweightness.
[0238] Table 4 Comparison of ablation experiments before and after the introduction of the knowledge distillation mechanism
[0239]
[0240]
[0241] 1. Comparative Experiment
[0242] To objectively evaluate the detection performance of the proposed detection algorithm, seven detection algorithms, namely SSD, Faster-RCNN, Retinanet, FCOS, ATSS, Dynamic-RCNN, and YOLOX, were selected for comparative experiments. The detection results are shown in Table 5.
[0243] Table 5 Comparison results of different detection algorithms
[0244]
[0245] As shown in Table 5, by constructing a knowledge distillation model to inherit important information from the teacher network to the student network, the proposed algorithm achieves high performance on AP and AP. 50 AP 75In terms of metrics, it significantly outperforms the other seven compared algorithms, and substantially improves the detection performance of weak hotspot targets without increasing additional cost. Although compared to the original YOLOX algorithm, AP... M The value decreased slightly, but is still close to the optimal value.
[0246] To further demonstrate the superiority of the proposed detection system in various complex environments, the following three typical scenarios were selected for testing and verification: cluttered small target scenarios, noisy interference scenarios, and motion-blurred scenarios, such as... Figure 9 As shown. Meanwhile, Figure 10 The results of various algorithms in different scenarios are shown. For ease of observation and subsequent analysis, the areas of missed detection and false detection for each algorithm have been marked with solid white lines.
[0247] observe Figure 10 (a) It can be seen that all seven comparison algorithms suffer from varying degrees of missed detection when dealing with chaotic multi-objective scenarios. In contrast, the KDBiDet system, by designing various improvement strategies in the teacher network and utilizing the idea of knowledge distillation to transfer important information to the student network, is able to complete photovoltaic module operation and maintenance tasks with high precision and high efficiency.
[0248] observe Figure 10 (b) It can be seen that due to interference from complex noise, the SSD, Retinanet, and YOLOX algorithms suffer from serious false negatives, while the other four contrastive detection networks lack the ability to aggregate feature information across different scales, making it difficult to achieve accurate multi-scale hotspot target detection. The constructed KDBiDet network aggregates multi-scale features from both global and local perspectives in a parallel fusion manner, enabling the network to selectively focus on regions containing target saliency information, thereby enhancing the algorithm's cross-scale feature aggregation capability.
[0249] observe Figure 10 (c) It can be seen that in motion-blurred scenes, infrared images cannot accurately represent the contour features of hot spot faults, thus limiting the detection performance of the algorithm in dense multi-target situations. The designed KDBiDet improves the expressive power of the Transformer architecture by performing self-attention learning operations and mining the contextual information between the input keys of the two-dimensional feature map, thereby improving the detection accuracy of small-scale hot spot targets in various dense scenes.
[0250] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for intelligent photovoltaic hot spot fault detection, the method comprising: The intelligent photovoltaic hot spot fault detection method comprises: The knowledge distillation module of the "teacher + student" collaborative training is constructed; the feature expression ability of the to-be-detected target is enhanced by designing a CSPHN-based backbone network enhancement algorithm; the BiMAF model is constructed, the input features are fused from the global and local angles by the parallel method, and the aggregation ability of the neck network to the target features is strengthened; at the same time, the CgT module is proposed to find the hot spot fault existing in the photovoltaic power generation system under various harsh environments and to exclude and repair the hot spot fault; The intelligent photovoltaic hot spot fault detection method comprises the following steps: Step one, based on the knowledge distillation idea, a new mode of teacher + student network collaborative training is constructed, the algorithm detection precision is improved by the advantage of teacher deep detection network, and the algorithm reasoning efficiency is improved by combining the small parameter quantity characteristics of student network, so as to balance the requirements of detection precision and model lightweight; Step two, based on the unmanned aerial vehicle intelligent inspection data set, the student network based on YOLOX algorithm is built, the anchor frame is not set, the position coordinates of the target are directly predicted through the key points, and the decoupling prediction head is combined, so that the algorithm detection precision is improved and the network convergence speed is accelerated; Step three, the CSPHN model based backbone network is constructed in the teacher network, the high-order space interaction information similar to the Transformer is obtained by using the gate convolution and recursion design on the basis of retaining the beneficial inductive bias of the convolutional neural network, so as to enhance the feature expression ability of the algorithm to the hot spot target; Step four, the BiMAF module is applied in the neck network, and the multi-level features are aggregated from the global and local angles in a parallel fusion manner, so that the network can selectively focus on the area containing the target saliency information, and the feature aggregation ability of the neck network is enhanced; Step five, the decoupling prediction head based on CgT module is constructed in the teacher network, the attention matrix is generated by mining static and dynamic context information, so that the detection system can still find the hot spot fault existing in the photovoltaic power generation system in the dense small target scene and exclude and repair the hot spot fault; The construction of the collaborative training new mode comprises: Step 1.1, a local distillation function is designed to separate image foreground and background information, and guide the student network to pay attention to important pixels and channel features; Step 1.2, a global distillation function is proposed to reconstruct the relationship between different pixels, and the relationship between different pixels is transmitted from the teacher network to the student network to compensate for the global information lost in the local distillation process; Step 1.3, the information in the teacher network is inherited into the student network by fusing the local distillation and the global distillation; The construction of the CSPHN module comprises: The feature mapping of the input layer is divided into two parts, and the two parts are combined by using a cross-stage hierarchy; a residual unit in the original CSPNet model is replaced by a HorNet module which inherits the meta-architecture of the spatial mixing layer and the cascade of the feedforward network of the Transformer model, and g n Conv captures high-order spatial interactions in the feature map; The specific process of the BiMAF module comprises: Step 2.1 Global feature fusion branch: convert the global pooling operation into a one-to-one dimensional feature operator, output the global feature weight through the cascade and convolution operation and utilize the sigmoid normalization function, and obtain the global feature fusion result Output through the weight distribution operation g ; Step 2.2 Local feature fusion branch, using 1x1 convolution and sigmoid normalization function to construct local feature weight, and then using weight distribution operation to get local feature fusion result Output l ; Step 2.3, the BiMAF model aggregates the global feature fusion result and the local feature fusion result by parallel fusion to obtain the final feature fusion result: Output = [Output g , Output l ]; The specific process of the CgT model comprises: Assume that a two-dimensional input feature map is The keys K, queries Q and values V are defined as K=XW k , Q=XW q and V=XW q respectively; wherein W q and W v are linear transformation matrices composed of 1*1 convolution; W k is a linear transformation matrix composed of group convolution with a convolution kernel size of 3*3, which can reflect static context features f s by encoding the context of the input feature map, at the same time, auxiliary self-attention learning is used to mine the interaction between the context key features and the query features, and then a dynamic attention weight matrix A is constructed as follows: A = g n conv([K, Q]) wherein g n Conv denotes a gated recurrent convolution operator; Then, the dynamic context feature f is calculated d : f d = A O V By capturing static context features and dynamic context features, the following output is obtained: output = f s + f d .
2. The method of claim 1, wherein the method further comprises: The specific process of step 1.1 comprises: Step 1.1.1 Construct a binary mask M for feature map F with horizontal and vertical coordinates i, j, respectively i,j Separate image foreground and background as follows: In the formula, r is the target real box; Step 1.1.2 is to set the scale mask S for the balance detection network to loss of different scale targets and the front background area i,j : In the formula, H r and W r respectively represent the height and width of the target real box. Step 1.1.3 Constructing spatial attention mask and channel attention mask to improve model distillation performance: where H, W and C represent the characteristic height, width and channel, respectively; T is a temperature hyper-parameter; F c and F i,j represent the characteristic information of the c-th channel and the characteristic information of size i x j, respectively. Step 1.1.4 During training, a binary mask M is used. i,j Scale mask S i,j and attention mask and To guide students in online learning, the teacher's key network space and channel information are analyzed, thereby constructing the following feature loss function L. fea and attention loss function L at : where, and denote the feature maps of the teacher and student networks, respectively; f(·) is the reconstruction operator that adjusts the feature maps of the student network to the same dimension as the teacher network; a, b, and g are hyperparameters that balance the losses; and l denotes the l1 norm operator. S and T denote the feature maps of the teacher and student networks, respectively; f(·) is the reconstruction operator that adjusts the feature maps of the student network to the same dimension as the teacher network; a, b, and g are hyperparameters that balance the losses; and l denotes the l1 norm operator. and denote the spatial and channel attention masks of the teacher and student networks, respectively. Step 1.1.5, the final local distillation function is obtained by calculating the feature loss and attention loss: L focal = L fea + L at .
3. The method of claim 1, wherein the method further comprises: The specific process of step 1.2 includes: The global information of the captured image is obtained by using the GcBlock module, so as to obtain a global loss L global : L global = λ · Σ(R(F T )- R(F S )) 2 Wherein, λ represents the balance loss hyperparameter; R(F) represents the feature global information, which can be represented as: where W k , W v1 , and W v2 represent convolution layers; LN represents layer normalization; N p represents the number of feature pixels; and ReLU represents a linear activation function.
4. The method of claim 1, wherein the method further comprises: determining a temperature of the photovoltaic module; and determining whether the temperature is within a predetermined temperature range. The design of the student network includes: The YOLOX network includes a backbone network Backbone, a neck network Neck and a decoupling prediction head; The Backbone includes a Focus slice module, a CSPNet and an SPP; wherein, the Focus slice module is used to expand the network receptive field and suppress the loss of image feature information; the CSPNet structure is used to solve the problem of high computational cost caused by repeated gradient information in the network optimization process; the SPP structure uses multi-level pooling operation to expand the receptive field of the backbone network; The Neck is composed of a CSP2_X structure, and uses an FPN+PAN feature pyramid structure to aggregate features of different scales, wherein, the FPN transmits strong semantic information from top to bottom, and the PAN transmits strong positioning information from bottom to top; The decoupling prediction head is used to decouple the classification branch of focused texture information and the positioning branch of focused edge information.
5. The intelligent photovoltaic hot spot fault detection method of claim 1, wherein, The specific process includes: First, assume g n The input features of Conv are A set of projection features p0 and can be obtained by the following formula: where φ(·) denotes the projection operator, and Then, the gated recurrent convolution is performed, and the formula is: p k+1 = DW k (q k )⊙g k (p k ) / α,k=0,1,…,n-1 where a is a scaling factor, {DW k} denotes a set of depthwise convolutions, and denotes a dot product operation. k} is used to match the dimensions of the sequences, as shown in the following equation: Finally, feature projection is performed after the top-level recursive operation to obtain g n The output result of Conv.
6. The method of claim 1, wherein the method further comprises: The calculation process of the global feature fusion result in step 2.1 includes: Step 2.1.1 assumes that the multi-level feature input is respectively The element-wise operation is performed as follows: M = Input1 + Input2 The output of the cth channel with height h and width w is obtained by using two pooling kernels with size (H, 1) and (1, W) to encode each channel along the horizontal and vertical directions, respectively and the output of the cth channel with height h and width w Step 2.1.2 performs the following cascade and convolution operations: f = δ(F1([z h ,z w ])) wherein [z h ,z w ] represents a connection operation along the spatial dimension; δ is an activation function; f is an intermediate feature map obtained by encoding in the horizontal and vertical directions; and F1 represents a 1x1 convolution operation. By decomposing f along the spatial dimension, we get two independent tensors: f h ∈R C / r×H and f h ∈R C / r×H where r is the down-sampling ratio; meanwhile, two 1x1 convolutions F h and F w operations are performed to ensure that the two independent tensors have the same number of channels, resulting in the global feature weight W g : W g = sigmoid(F h ([f h ])) x sigmoid([f w ]) Step 2.1.3 Output of the global feature fusion branch g is: Output g = I1 x W g + I2 x (1 - W g ) The calculation process of the local feature fusion result in step 2.2 includes: The local feature fusion is performed on the input features, and the local feature weight W l is represented as: W l = sigmoid(F1(δ(BN(F1(M))))) Further, the output of the local feature fusion branch is represented as: Output l = Input1 x W l + Input2 x (1 - W l ).
7. Photovoltaic hot spot fault detection implementing the method of any one of claims 1-6, characterized in that, The photovoltaic hot spot fault detection system includes: A data acquisition module is configured to automatically plan a UAV cruising route by analyzing geographic information and a cruising range of a photovoltaic power station, and to collect inspection images and videos by using thermal radiation imaging characteristics of an infrared sensor; A data transmission module is configured to transmit inspection data back to a ground control station and store the inspection data correspondingly by using high-speed and low-latency advantages of a 5G wireless network, so that a subsequent high-performance computer can process the data; A fault diagnosis module is configured to use a double-branch collaborative training photovoltaic hot spot fault detection algorithm based on a knowledge distillation mechanism to perform feature extraction, information aggregation and fault positioning tasks on photovoltaic data, and finally obtain a hot spot fault diagnosis result in combination with image information and GPS positioning data.
Citation Information
Patent Citations
Photovoltaic module hot spot detection method and system based on fused image
CN115409814A