A Video-based Precise Detection and Tracking Method and System for Photovoltaic Modules
By shooting photovoltaic module videos by drones and using eYOLO-PV model and SORT algorithm, accurate detection and tracking of photovoltaic modules are achieved, solving the problems of inaccurate component detection and in real-time tracking in the existing technology, and improving detection accuracy and tracking accuracy.
Patent Information
- Application Number
- CN202410719639.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-06-05
AI Technical Summary
The existing photovoltaic power plant component detection technology is difficult to achieve accurate component detection and tracking, especially when the correspondence between different video frames is not considered, resulting in the same component being recognized multiple times, and the end-side real-time tracking and completely accurate tracking cannot be achieved.
The video-based photovoltaic module accurate detection and tracking method is adopted, and the photovoltaic module video is captured by a drone, the video frame is extracted as an image object detection data set, the component object frame is identified using the eYOLO-PV target detection model, and real-time tracking is performed based on the lightweight component tracking algorithm SORT.
It realizes accurate detection and tracking of photovoltaic modules, improves detection accuracy, ensures unique tracking of components, and realizes real-time tracking on the end side to achieve completely accurate tracking effect.
Smart Images

Figure CN118711084B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method and system for precise detection and tracking of photovoltaic modules based on video. Background Art
[0002] The existing detection technologies for components in photovoltaic power stations are divided into three categories: traditional image threshold processing algorithms, semantic segmentation algorithms, and object detection algorithms. Both traditional threshold segmentation and semantic segmentation achieve string-level detection. The object detection algorithm can distinguish different objects of the same category and can help quickly detect different components in the photovoltaic scenario. However, in the current research on using object detection algorithms to identify components, most of them only process component pictures taken offline and do not consider the correspondence between components in different frames of the video. In actual inspections, most of the collected component images overlap, so the object detection algorithm will cause the same component to be recognized multiple times. Using an object tracking algorithm to detect components can ensure that the components are uniquely tracked and a unique identifier is assigned to the components. Moreover, the few existing algorithms for tracking photovoltaic modules cannot achieve real-time tracking at the edge side and cannot accurately track the components completely. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for precise detection and tracking of photovoltaic modules based on video, which can accurately detect and track the photovoltaic modules in a photovoltaic power station.
[0004] To achieve the above purpose, the present invention provides the following solutions:
[0005] In the first aspect, the present invention provides a method for precise detection and tracking of photovoltaic modules based on video, including:
[0006] Obtain a video of photovoltaic modules taken by a drone.
[0007] Extract images from video frames in the photovoltaic module video to obtain an image object detection data set.
[0008] Input the image object detection data set into a trained object detection model to identify the component target boxes of each photovoltaic module in the photovoltaic module video; the object detection model is a detection model constructed by the eYOLO-PV object detection algorithm.
[0009] Based on the photovoltaic module video and the component target boxes of the photovoltaic modules, perform real-time tracking of the photovoltaic modules based on the lightweight component tracking algorithm SORT.
[0010] In the second aspect, the present invention provides a system for precise detection and tracking of photovoltaic modules based on video, including:
[0011] A video acquisition module, configured to acquire a video of a photovoltaic module captured by a drone.
[0012] An image extraction module, configured to extract video frames in the video of the photovoltaic module as images, thereby obtaining an image target detection data set.
[0013] A detection module, configured to input the image target detection data set into a trained target detection model, and identify the component target boxes of each photovoltaic module in the video of the photovoltaic module.
[0014] A tracking module, configured to perform real-time tracking on the photovoltaic modules according to the video of the photovoltaic module and the component target boxes of the photovoltaic modules, based on the lightweight component tracking algorithm SORT.
[0015] Optionally, it further includes: a training module, configured to train the target detection model.
[0016] Optionally, the training module specifically includes:
[0017] An image acquisition sub-module, configured to acquire sample infrared image data of photovoltaic modules in a photovoltaic power station; the sample infrared image data is obtained from historical videos of photovoltaic modules.
[0018] A labeling sub-module, configured to perform component box labeling on the sample infrared image data, thereby obtaining labeled sample infrared image data.
[0019] A training sub-module, configured to train the eYOLO-PV target detection algorithm based on the labeled sample infrared image data, thereby obtaining a trained target detection model; wherein, the eYOLO-PV target detection algorithm is an algorithm obtained by replacing the backbone network with an EfficientViT network on the basis of the improved YOLOv8 target detection algorithm.
[0020] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0021] The present invention provides a method and system for precise detection and tracking of photovoltaic modules based on video. The method includes: obtaining a video of photovoltaic modules captured by a drone; extracting video frames from the video of photovoltaic modules as images to obtain an image target detection data set; inputting the image target detection data set into a trained target detection model to identify the component target boxes of each photovoltaic module in the video of photovoltaic modules; and based on the lightweight component tracking algorithm SORT, performing real-time tracking on the photovoltaic modules according to the video of photovoltaic modules and the component target boxes of the photovoltaic modules. The eYOLO-PV model proposed by the present invention is specially designed for video monitoring of photovoltaic power stations, improving the detection accuracy. At the same time, the lightweight SORT in the present invention can be used for efficient tracking of photovoltaic modules in consecutive multi-frame videos, ensuring accurate positioning ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a schematic flowchart of a method for precise detection and tracking of photovoltaic modules based on video provided in Embodiment 1 of the present invention.
[0024] FIG. 2 is an infrared video screenshot of photovoltaic modules in a photovoltaic power station provided in Embodiment 1 of the present invention; FIG. 2(a) is an infrared video screenshot of photovoltaic modules in Power Station 1; FIG. 2(b) is an infrared video screenshot of photovoltaic modules in Power Station 2; FIG. 2(c) is an infrared video screenshot of photovoltaic modules in Power Station 3; FIG. 2(d) is an infrared video screenshot of photovoltaic modules in Power Station 4.
[0025] Figure 3 It is a schematic diagram of the YOLOv8 network structure provided in Embodiment 1 of the present invention.
[0026] Figure 4 It is a schematic diagram of the EfficientViT model structure provided in Embodiment 1 of the present invention.
[0027] Figure 5 It is a schematic diagram of the eYOLO-PV network structure provided in Embodiment 1 of the present invention.
[0028] FIG. 6 is a loss curve diagram of three models trained in Embodiment 1 of the present invention; FIG. 6(a) is a loss curve diagram of the eYOLO-PV target detection model; FIG. 6(b) is a loss curve diagram of the YOLOv8 target detection model; FIG. 6(c) is a loss curve diagram of the YOLOv7 target detection model.
[0029] Figure 7 It is a schematic diagram of the curve of the average accuracy mean index for different model training in the first embodiment of the present invention.
[0030] Figure 8 It is a flowchart of the SORT tracking algorithm provided in the first embodiment of the present invention.
[0031] Figure 9 is a schematic diagram of the DarkLabel interface provided in the first embodiment of the present invention.
[0032] Figure 10 is a schematic diagram of the component target tracking results of four power stations provided in the first embodiment of the present invention; Figure 10(a) is a schematic diagram of the component target tracking results of power station 1; Figure 10(b) is a schematic diagram of the component target tracking results of power station 2; Figure 10(c) is a schematic diagram of the component target tracking results of power station 3; Figure 10(d) is a schematic diagram of the component target tracking results of power station 4.
[0033] Figure 11 It is a schematic diagram of the structure of a video-based precise detection and tracking system for photovoltaic modules provided in the second embodiment of the present invention. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] The purpose of the present invention is to provide a video-based precise detection and tracking method and system for photovoltaic modules, which can accurately detect and track the photovoltaic modules of a photovoltaic power station.
[0036] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0037] Embodiment 1
[0038] As Figure 1 shown, this embodiment provides a video-based precise detection and tracking method for photovoltaic modules, including:
[0039] Step 101: Obtain a video of photovoltaic modules taken by a drone.
[0040] Step 102: Extract images from the video frames in the photovoltaic module video to obtain an image target detection data set.
[0041] Step 103: Input the image target detection dataset into the trained target detection model to identify the component target boxes of each photovoltaic component in the photovoltaic component video; the target detection model is a detection model constructed by the eYOLO-PV target detection algorithm.
[0042] Step 104: Based on the photovoltaic component video and the component target boxes of the photovoltaic components, perform real-time tracking of the photovoltaic components based on the lightweight component tracking algorithm SORT.
[0043] Wherein, before performing step 101, it further includes:
[0044] Judge whether the climate at the shooting location meets the shooting standards.
[0045] When it meets the shooting standards, set the flight speed, flight altitude, camera angle and shooting route of the drone, and control the drone to shoot the photovoltaic components.
[0046] Specifically, the requirements for drone inspection are as follows:
[0047] The infrared video contains the correlation information of the components in the front and back frames compared with the infrared image. Therefore, in order to detect and track the components more efficiently in the follow-up. All videos are shot under clear weather conditions that meet the IEC61215-2:2021 standard. And the following requirements need to be met:
[0048] (1) The resolution is 640×512 pixels and the frame rate is 29.97 frames per second.
[0049] (2) The camera moves monotonically along the row or column without obvious backward movement.
[0050] (3) The component rows must be approximately horizontal or vertical in each frame.
[0051] (4) The flight speed, altitude and camera angle of the drone are basically stable and unchanged.
[0052] As Figure 2(a)-Figure 2(d) shown, this embodiment uses the thermal infrared videos of two photovoltaic power stations in each of Place A and Place B, including 2 centralized photovoltaic power stations and 2 rooftop distributed photovoltaic power stations. The total video length of the four power stations is about 1.514 hours, with a total of 163,555 frames, and each frame contains about 70 photovoltaic components on average.
[0053] Wherein, when performing step 102, it can be specifically as follows:
[0054] According to the infrared video of the photovoltaic power station components, extract the infrared video frames as images and divide the infrared image target detection dataset.
[0055] Specifically, in the training and evaluation stages of supervised learning, the optimization and evaluation of the model rely on the true labels of the dataset. In this embodiment, out of the 163,555 video frames of photovoltaic power plants 1, 2, 3, and 4, a total of 284 frames (71 frames for each power plant) were selected, and the components in each frame were labeled using the LableImg tool. The dataset was divided in a ratio of 7:2:1, where 200 frames were used to train the photovoltaic component target detection model, 56 frames were used to validate the model, and 28 frames were used to test the model's performance. The specific division is shown in Table 1 below:
[0056] Table 1 Division of Photovoltaic Component Dataset
[0057]
[0058]
[0059] Among them, before performing step 103, it also includes:
[0060] Train the target detection model, which can be specifically as follows:
[0061] Obtain the sample infrared image data of photovoltaic components in the photovoltaic power plant; the sample infrared image data is obtained from historical photovoltaic component videos.
[0062] Perform component box annotation on the sample infrared image data to obtain the annotated sample infrared image data.
[0063] Based on the annotated sample infrared image data, train the eYOLO-PV target detection algorithm to obtain a trained target detection model; among them, the eYOLO-PV target detection algorithm is an algorithm obtained by replacing the backbone network with the EfficientViT network on the basis of the improved YOLOv8 target detection algorithm.
[0064] Specifically, the efficient component target detection model eYOLO-PV incorporates the EfficientViT structure as the backbone network into the YOLOv8 model. The designed network not only utilizes the advanced feature extraction ability of EfficientViT but also inherits the advantages of YOLOv8 in real-time target detection, which is suitable for this experimental scenario.
[0065] Among them, the specific structure of the YOLOv8 model is as follows:
[0066] As Figure 3 shown, the YOLOv8 network structure mainly consists of a backbone, a neck, and a head structure.
[0067] (1) Backbone structure. YOLOv8 uses the improved CSPDarket53 as the backbone network, which downsamples the input features five times to obtain five features of different scales in sequence. The structure of the backbone network is shown in the figure. The Cross Stage Partial (CSP) module in the backbone network is replaced by the C2f module (n represents the quantity). The full name of C2f is "CSPDarknet53 to 2-Stage FPN". It realizes the gradient shunt connection based on the CSPDarknet53 backbone network and the two-stage Feature Pyramid Network (2-Stage FPN), enriching the information flow of the feature extraction network while maintaining lightweight. The CBS (Convolution Conv, Batch Normalization BN, Activation function SiLU) module performs convolution operations on the input information, then conducts batch normalization, and finally uses the SiLU activation function to obtain the output result. The backbone network finally uses the Spatial Pyramid Pooling-Fast (SPPF) module to pool the input feature map into a fixed-size map for adaptive size output. SPPF connects three max pooling layers in the order of gradients, reducing the computational workload and having lower latency.
[0068] (2) Neck structure. Inspired by the Path Aggregation Network (PANet), YOLOv8 designs the Path Aggregation Network + Feature Pyramid Network (PANet+Feature PyramidNetwork, PAN+FPN) in the neck part. PAN+FPN constructs a top-down and bottom-up network structure, realizing the complementarity of shallow position information and enabling full fusion between multi-scale information. Except that the C2f module is used in FPN+PAN, the rest is basically the same as the PAN+FPN structure of YOLOv5.
[0069] (3) Head Structure. The detection part of YOLOv8 adopts a Decoupled Head structure. This decoupled head structure sets two independent branches for object classification and prediction bounding box regression respectively, and uses different loss functions for these two tasks. For the classification task, Binary Cross-Entropy Loss (BCE Loss for short) is used. For the prediction bounding box regression task, Distribution Focal Loss (DFL for short) and CIoU are adopted. This detection structure helps to improve the detection accuracy and accelerate the model convergence. YOLOv8 is an Anchor-Free detection model that can simply specify positive and negative samples. In addition, it also adopts a Task-Aligned Assigner to dynamically allocate samples, which helps to improve the detection accuracy and robustness of the model.
[0070] The YOLOv8 model has strong capabilities in real-time detection. However, the infrared images of photovoltaic modules are greatly affected by environmental factors and show diversity under different lighting and background conditions, such as temperature, humidity, and solar radiation. These factors will affect the quality of infrared images. For example, when the solar radiation is weak, the infrared temperature difference between the module and the background is small, resulting in unclear boundaries in the infrared image, making it more complex to identify and analyze the defect problems of photovoltaic modules, which poses higher requirements for the target detection algorithm. Due to the high model capabilities and excellent performance of Vision Transformer (ViT), it has caused a great upsurge in the field of computer vision. ViT effectively solves the problem that traditional convolutional networks cannot extract global feature layer order and content information, and significantly improves the accuracy of the detection algorithm. However, the continuously improving accuracy comes at the cost of an increase in model size and computational overhead. To improve memory efficiency and enhance communication between channels, Microsoft proposed an efficient memory transformer model called EfficientViT.
[0071] Therefore, in this embodiment, the EfficientViT structure is added to the YOLOv8 model, as Figure 4 shown, where part (a) in the figure is the EfficientViT network architecture, part (b) is the EfficientViT module of the sandwich structure, and part (c) is the cascaded group attention module.
[0072] The EfficientViT module mainly improves the model's efficiency in terms of memory occupancy and computational redundancy through the Sandwich Layout and the Cascaded Group Attention (CGA) mechanism respectively.
[0073] The Sandwich Layout module. This structure reduces the use of the memory-consuming Self-Attention mechanism and adds memory-efficient Feed Forward Network (FFN) layers for channel communication. Specifically, this layout sandwiches a Self-Attention layer between FFN layers with a Self-Attention layer in between The calculation can be expressed as shown in (1):
[0074]
[0075] where X i is the complete input feature of the i-th block. Before and after a single Self-Attention, this block transforms X i into X i+1 . This design reduces the memory and time consumption caused by the self-attention layer in the model and applies more FFN layers to effectively achieve communication between different feature channels. In addition, a Token Interaction layer is added before each FFN. Token Interaction uses Depthwise Separable Convolution (DWConv) to enhance the model's capabilities.
[0076] The Cascaded Group Attention module. Under the Multi-Headed Self-Attention (MHSA) mechanism, each head performs the same task, resulting in low computational efficiency. First, inspired by group convolution in convolutional neural networks, CGA inputs each head with different partitions of the complete feature, thus explicitly decomposing the attention calculation into different heads. Second, by calculating the attention maps of each head in a cascaded manner, the output of each head is added to the subsequent heads to increase the diversity of the attention maps. The CGA expression is (2):
[0077]
[0078] where the j-th head calculates the self-attention on the input feature partition X ij , that is, X i = [X i1 , Xi2 ,..., X ij , where 1 ≤ j ≤ h and h is the total number of heads, maps the input feature segmentation to different subspaces, is a linear layer that projects the concatenated output features back to the same dimension as the input.
[0079] Calculate the attention map for each head in a cascaded manner, as shown in Figure 4 (c), add the output of each head to the subsequent head to gradually improve the feature representation. The expression is shown in (3), where X' ij is the j-th input segmentation X ij and the output of the (j - 1)-th head calculated by formula (2) are added together.
[0080]
[0081] This cascaded design has two advantages. First, feeding different feature segmentations to each head can improve the diversity of the attention map. Similar to group convolution, each head only receives a slice of the full feature, thus saving computational cost and parameters. Second, cascaded attention heads allow increasing the network depth to further improve the model capacity without introducing any additional parameters. It only incurs a slight latency overhead because the attention map calculation in each head uses a smaller QK channel dimension.
[0082] The overall architecture of the EfficientViT network is presented in Figure 4 (a). Specifically, first, the model introduces the OverlapPatchEmbedding method to downsample the image by 16 times, which enhances the model's ability to capture more local details and context information. The overall architecture of EfficientViT contains three stages, and each stage stacks the proposed EfficientViT modules. Each subsampling layer (downsampling by a factor of 2 in resolution) reduces the number of tokens by 4 times. To achieve efficient subsampling, the author further proposes the SubSample block of EfficientViT, which also has a sandwich layout, except that the CGA layer of the EfficientViT module is replaced with an inverted residual to reduce information loss during subsampling.
[0083] Specifically, the improvement ideas and specific structure of the eYOLO-PV model are as follows:
[0084] Facing the photovoltaic scenario, although YOLOv8 has improved its performance by improving the model architecture and optimizing the loss function, etc., it still faces challenges when dealing with this task. Compared with ordinary object detection tasks, the main differences in this embodiment are reflected in the following aspects:
[0085] First, only one detection target, i.e., the photovoltaic module, is involved in this embodiment. This single detection target means that the target detection model does not need to overly focus on the extraction of deep information for multi-classification, nor does it need to handle possible overlaps or occlusions between targets.
[0086] Secondly, the infrared images of photovoltaic modules are greatly affected by the environment. When the infrared radiation of the photovoltaic module is relatively weak, the boundaries and features of the photovoltaic module in the infrared image will not be obvious enough; and in this embodiment, infrared videos are processed, and the resolution of infrared videos is usually low, which results in the loss of detailed information of the module in the video frames; moreover, the stability of the video may be affected by wind and operation control. Unstable shooting will not only cause the picture to be blurred, but also interfere with the performance of the target detection algorithm.
[0087] Therefore, although the overall structure of this embodiment needs to be relatively simplified, higher requirements are put forward for the accurate recognition of component edges and the extraction of component features.
[0088] In this context, EfficientViT combines the deep representation learning ability of the vision Transformer structure and the computational efficiency advantages of traditional convolutional neural networks, and can significantly improve the overall accuracy and robustness of the photovoltaic detection model while maintaining real-time processing speed. The eYOLO-PV efficient component detection model dedicated to infrared images of photovoltaic power stations designed in this embodiment replaces the backbone network of YOLOv8 with the EfficientViT network to utilize the advanced feature extraction ability of EfficientViT and inherit the advantages of YOLOv8 in real-time object detection at the same time. This combination not only reduces the computational complexity, but also is effective for interfering factors such as low resolution, light changes, and unstable shooting in the photovoltaic infrared scenario.
[0089] The structure diagram of the efficient eYOLO-PV model is shown in Figure 5 , where the resolution of the input infrared image is 640x640 pixels. The backbone network part consists of three main stages. The size of the image patch is 16, and the dimensions of the patch embeddings corresponding to the three stages are 64, 128, and 192 respectively. The number of CGA attention heads in the three stages is 4. For efficiency considerations, this architecture uses BatchNorm (BN) instead of LayerNorm (LN) because BN can be integrated into the previous convolutional or linear layer, thus providing runtime advantages. ReLU is used as the activation function, mainly for speed considerations and support for various inference deployment platforms.
[0090] Therefore, the training process of the target detection model is specifically as follows:
[0091] The hardware condition for training is an NVIDIA A800 GPU with 80GB of video memory, providing sufficient computing resources to process large-scale infrared image data of photovoltaic modules. Under the Ubuntu 20.04 operating system, the PyTorch 2.0 framework is used for training.
[0092] Preprocess 200 frames of videos in the training set, including cropping, scaling to a unified size (640x640 pixels), and data augmentation to improve the model's ability to recognize the boundaries of photovoltaic modules.
[0093] During the training process, the settings of key parameters include: using the Adam optimizer, setting the initial learning rate to 0.001, and adopting a learning rate decay strategy to gradually reduce the learning rate during training to promote model convergence. According to the available computing resources, the batch size is set to 16, ensuring the stability and efficiency of training. A composite loss function that combines the regression loss of CIoU Loss and DFLLoss and the classification loss of BCE Loss is adopted, aiming to optimize the localization accuracy and classification accuracy of the model simultaneously. The model was trained for a total of 300 epochs, which is sufficient to ensure that the model can fully learn the features in the dataset.
[0094] To improve the generalization ability and accuracy of the model, the following strategies are adopted during training: One is the Early Stopping method. Monitor the loss on the validation set and stop training when the loss does not decrease significantly for several consecutive epochs to avoid overfitting. The other is weight decay. Introduce weight decay (L2 regularization) to reduce the model complexity and improve the generalization performance of the model.
[0095] Through the implementation of the above training process and strategies, the photovoltaic module defect detection algorithm of eYOLO-PV proposed in this embodiment can effectively learn and identify photovoltaic modules, achieving the goals of high accuracy and fast detection. The design of the experimental training process takes into account the maximization of model performance and the requirements of actual application scenarios, ensuring that the algorithm has high practical value and reliability in the application of real-world photovoltaic power stations.
[0096] The evaluation process after the eYOLO-PV model training is completed is as follows:
[0097] To verify the effectiveness of the model, the performances of three different object detection models, eYOLO-PV, YOLOv8, and YOLOv7, during the training process were compared. Figures 6(a), (b), and (c) respectively show the changes in three different losses (Box Loss, Class Loss, and DFL Loss) with the training cycle under different network architectures.
[0098] After training, the Box Losses of the three networks in Figure 6 are: 0.31785, 0.58495, and 0.76211 respectively. This loss value measures the difference between the bounding boxes predicted by the network and the ground truth bounding boxes. For photovoltaic module detection, a lower Box Loss means that the network can more accurately locate the position of the photovoltaic panel. The performance of eYOLO-PV in this item indicates that it can effectively identify the accurate position of the photovoltaic module and can be used for subsequent positioning.
[0099] After training, the Class Losses of the three networks in Figure 6 are: 0.23494, 0.38907, and 0.46316 respectively. Although there is only one class, the Class Loss is still important because it measures the recognition confidence of the model for this class. In a single-class detection task, a low Class Loss indicates that the model is very confident in the photovoltaic modules it detects and rarely misdetects the background or other non-target objects as photovoltaic modules.
[0100] After training, the DFL Losses of the three networks in Figure 6 are: 0.63127, 0.87673, and 1.23281 respectively. DFL Loss represents the fine-grained feature learning loss, which involves learning different subtle features of photovoltaic modules in a single-class detection task. Due to the differences in the component features of the four power stations in this experiment and the different details of different component defects, the DFL Loss is higher than the Box Loss and the Class Loss. The good performance of eYOLO-PV in this metric means that the network can not only detect the presence of photovoltaic modules but also capture the important details of the modules.
[0101] Overall, the eYOLO-PV network provides better feature representation, enabling the network to converge to better results at a faster speed and being more accurate than YOLOv8 and YOLOv7 in component positioning, classification, and feature learning.
[0102] As Figure 7 shown, the curves of the mean average precision (mAP) metric of the three models for the intersection over union (IoU) from 0.5 to 0.95. After 300 rounds of iterative training, the mAP@0.5:0.95 of the training set for eYOLO-PV, the original YOLOv8, and YOLOv7 are: 0.98154, 0.94861, and 0.92175 respectively. It can be observed that eYOLO-PV has a higher precision from the beginning and converges quickly, showing fast learning ability and strong performance, and maintaining the leading precision throughout the training process, demonstrating the efficiency of the eYOLO-PV model.
[0103] To verify the robustness of the model, after training is completed, the model is evaluated on the test set. Table 2 compares the mean average precision (mAP) of the eYOLO-PV, YOLOv8, and YOLOv7 models. Among them, the eYOLO-PV model performs the best, reaching 99.73%. Although the model size of eYOLO-PV is 8.35 MiB, slightly larger than YOLOv8. However, in terms of inference speed, eYOLO-PV also shows high efficiency. For an infrared video stream containing multiple photovoltaic modules, it only takes 3.12 milliseconds to detect multiple targets simultaneously in each frame, which is 22.6% and 59.2% faster than YOLOv8 and YOLOv7 respectively. Finally, the floating point operations (FLOPs) of eYOLO-PV are 910 million, indicating high computational efficiency while maintaining high accuracy. In summary, the eYOLO-PV model demonstrates remarkable efficiency in terms of accuracy, inference speed, model and parameter size, and computational efficiency.
[0104] Table 2 Metrics on the Test Sets of Different Models
[0105]
[0106] Among them, when performing step 104, it can be specifically as follows:
[0107] Based on the Kalman filter, predict the motion state of the tracking target in the current video frame; the tracking target is the component target box of the photovoltaic module.
[0108] Based on the tracking target in the next video frame and the tracking target in the current video frame, use the Hungarian algorithm for target matching to obtain the matching result of the tracking target.
[0109] When the matching result is a matching failure, determine that the tracking target is lost, and delete the component target box of the photovoltaic module corresponding to the tracking target.
[0110] When the matching result is a matching success, based on the tracking target in the current video frame, update the state of the component target box of the photovoltaic module; the state is the position and speed of the component target box.
[0111] Specifically, the SORT tracking algorithm is an object tracking algorithm based on the Kalman Filter and Hungarian algorithms. Its core idea is to first detect objects in the video through an object detection algorithm and represent them as detection boxes. SORT can be used in conjunction with various object detectors (such as YOLO, Faster R-CNN, etc.) to extract the object position information in each frame. Then, the Kalman filter algorithm is used to estimate and predict the state of the object. Finally, the Hungarian algorithm is used to match the object detection boxes in the current frame with the objects tracked in the previous frame to achieve object tracking. Figure 8 Shows the process of the SORT tracking algorithm. The basic steps are summarized as: eYOLO-PV component detection, Kalman filter component prediction, Hungarian IoU object matching, object ID creation and destruction.
[0112] The Kalman filter can predict the position of the object in the next frame according to the linear constant velocity model by considering the dynamic model of the object and measurement noise. When new detections match existing objects, these detections are used to update the state of the object. Expressions (4) and (5) are the state vector x and observation vector z of the original Kalman filter respectively:
[0113]
[0114] where u and v represent the horizontal and vertical pixel positions of the object center, s and r represent the scale (area) and aspect ratio of the object bounding box respectively, is the speed of the component box movement.
[0115] The Hungarian algorithm performs associative matching on the detected objects Detections and predicted objects Tracks based on the IoU intersection over union distance metric function. There are three cases when using the Hungarian algorithm for matching: the case of unmatched tracks (UnmatchedTracks). If the mismatch persists for T times, the ID is destroyed; the case of unmatched detections (Unmatched Detections), that is, a new object appears in the detection box, but this new object does not exist in the tracking box. None of the Tracks can match the Detections. Assign a Track to this Detection and create a new ID; the case of matched tracks (Matched Tracks). If the Detection and Track match successfully, then update the ID state with this detection result as the observation value.
[0116] The lightweight SORT tracking algorithm is specifically as follows:
[0117] In this embodiment, more than 70 components are detected and tracked in each frame of video on average. The SORT algorithm has a simple logic and only makes matches based on the position information of the detection boxes without involving complex training processes. Therefore, the tracking speed is very fast and it is very suitable for scenarios of real-time processing of large-scale data. The SORT algorithm has two limitations: First, the SORT algorithm highly depends on the detection results during the tracking process. Even if the actual target still exists but has not been detected for a long time, the SORT algorithm cannot continue to track. However, the component detection algorithm in this embodiment adopts the high-precision detection model of eYOLO-PV, and the component recognition accuracy rate within the same frame can reach more than 99.73%. It can ensure that the same component can only be not correctly detected in very few frames, which does not affect the actual tracking effect. Second, the SORT algorithm only uses the position information of the detection boxes to make matches and does not utilize the image texture features of the components. In scenarios where the target is occluded, a large number of ID switching phenomena may occur, affecting the stability of tracking. However, in the photovoltaic scenario, all component targets are tiled and arranged regularly. And when there is no debris on the component surface and the battery is healthy, the image features of different components in the same picture are basically similar. Therefore, it is suitable to use the simple and efficient SORT multi-object tracking algorithm.
[0118] The purpose of this embodiment is to achieve real-time panoramic stitching of a photovoltaic power station, which puts forward high requirements for the processing speed of each link. In this section, the SORT tracking algorithm is lightweighted. During the inspection of a photovoltaic power station by an unmanned aerial vehicle, the unmanned aerial vehicle flies at a fixed height at a constant speed. Therefore, in the same power station scenario, the sizes of the photovoltaic components in the video are almost unchanged, and the moving speeds of the components in the picture are basically uniform. In the original SORT model, the Kalman filtering algorithm is used to predict the position of the component in the next frame. In expression (4), s represents the area of the component target bounding box. In this system, the variables in the state vector of the Kalman filtering are removed because the physical meaning of is the change rate of the component area, and under this scenario limitation, the change rate of the area is basically 0. After removing the variable , when tracking the components in the same frame, not only is the speed doubled compared to the original, and the tracking speed per frame is only 2 ms, but the accuracy is not affected either.
[0119] Wherein, after the real-time tracking of the photovoltaic components, it further includes:
[0120] Using MOTA to evaluate the lightweight SORT target tracking algorithm to obtain the tracking accuracy rate of the lightweight SORT target tracking algorithm.
[0121] Specifically, this embodiment uses the Multiple Object Tracking Accuracy (MOTA) to evaluate the overall performance of the lightweight SORT target tracking algorithm.
[0122] MOTA calculates by considering three types of errors: False Positives (FP), False Negatives (FN), and ID Switches (IDS), thus providing a quantitative way to measure the performance of tracking algorithms in different aspects. The definition of MOTA is shown in expression (6):
[0123]
[0124] Among them, FN t represents the number of targets missed in time frame t. A missed detection refers to a target that actually exists but is not detected by the tracking algorithm. FP t represents the number of false positives in time frame t. A false positive is a situation where the algorithm incorrectly identifies the background or other non-target objects as targets. IDSW t represents the number of ID switches that occur in time frame t. An ID switch refers to a situation where the algorithm incorrectly assigns the ID of a tracked target to another target. GT t represents the number of ground truth targets in time frame t.
[0125] The value of MOTA ranges from -∞ to 1, where 1 represents perfect tracking with no false positives, false negatives, or ID switches. The closer the value is to 1, the better the performance of the algorithm. On the contrary, if the sum of false positives, false negatives, and ID switches exceeds the total number of ground truth targets, the value of MOTA may even be negative, indicating that the performance of the algorithm is very poor and needs further improvement and optimization.
[0126] Table 3 statistics the component accuracy of Power Station 1 based on eYOLO-PV component detection and lightweight SORT algorithm tracking. Power Station 1 is divided into 4 segments of videos for convenient evaluation. The real component boxes are manually labeled using the DarkLabel video annotation software to prepare for the evaluation. Figure 9(a) shows the visualization result of the tracking targets under DarkLabel. It can be seen that the component box with ID 313 is misaligned due to being close to the lower boundary of the screen. Figure 9(b) corrects the position of the component with ID 313 as the original annotation ground truth.
[0127] Table 3 Tracking performance indicators of Power Station 1
[0128]
[0129] In Table 3, taking the first video segment of Power Station 1 as an example, the actual number of components in the video is 357, and each component will appear for more than 100 frames. A total of 37,389 component bounding boxes need to be detected and tracked in this video. The algorithm of this system can ensure that all components are detected and tracked. However, in the video frames, there are individual errors in the 37,389-scale component bounding boxes. There are 9 ID switching situations and 6 false detection situations in the first video. A total of 37,394 component bounding boxes are actually tracked and detected. Through formula calculation, MOTA is approximately 1.
[0130] Among the five video segments of Power Station 1, the fourth segment has the worst effect, with MOTA being 0.995. After viewing and analyzing, there is a frame break situation in the fourth segment of the video. The drone has a video transmission interruption of more than 30 frames, and the position of the components in the video frame suddenly changes from uniform motion. To solve this problem, in this section, we try to adopt: video frame interpolation algorithm; tracking algorithm with an image feature extraction network (advanced ByteTrac tracking algorithm). Among them, the frame interpolation algorithm can improve MOTA to a certain extent. Due to the complex feature matching network of the ByteTrack tracking algorithm, it greatly affects the tracking speed instead. Finally, in this section, the fourth video segment is split again from the frame break point, and the two small segments are separately detected and tracked, and then spliced together. This can not only effectively avoid false detection, missed detection, and ID switching situations, but also ensure high-speed video tracking. This method is used for subsequent power stations in this system when frame breaks occur. After correction, a small number of false detection, missed detection, and ID switching situations in the four power station videos do not affect the final component extraction effect. Finally, the component tracking MOTA of the five video segments of Power Station 1 is all 1.
[0131] As shown in Figures 10(a)-(d), the processing effects of Photovoltaic Power Stations 1, 2, 3, and 4 based on the eYOLO-PV object detection algorithm and the lightweight SORT tracking algorithm are shown. Each power station shows two subgraphs on the left and right. They respectively capture the position dynamics of the components in different frames, and the movement of the components over time can be clearly observed through comparison.
[0132] In addition, Figure 10 also shows the adaptability of the algorithm of this system to different types of power stations under different lighting conditions and environmental changes. In the case of low-resolution infrared videos, it can still accurately identify and track photovoltaic components, with high robustness.
[0133] Embodiment 2
[0134] As Figure 11 shown, this embodiment provides a video-based precise detection and tracking system for photovoltaic components, including:
[0135] A video acquisition module 1101, which is used to acquire the video of photovoltaic components taken by the drone.
[0136] An image extraction module 1102, configured to extract video frames in the photovoltaic module video as images, and obtain an image target detection dataset.
[0137] A detection module 1103, configured to input the image target detection dataset into a trained target detection model, and identify the component target boxes of each photovoltaic module in the photovoltaic module video.
[0138] A tracking module 1104, configured to perform real-time tracking on the photovoltaic modules based on the photovoltaic module video and the component target boxes of the photovoltaic modules, based on the lightweight component tracking algorithm SORT.
[0139] The system further includes: a training module, configured to train the target detection model.
[0140] Wherein, the training module specifically includes:
[0141] An image acquisition sub-module, configured to acquire sample infrared image data of photovoltaic modules in a photovoltaic power station; the sample infrared image data is obtained through historical photovoltaic module videos.
[0142] A labeling sub-module, configured to perform component box labeling on the sample infrared image data to obtain labeled sample infrared image data.
[0143] A training sub-module, configured to train the eYOLO-PV target detection algorithm based on the labeled sample infrared image data to obtain a trained target detection model; wherein, the eYOLO-PV target detection algorithm is an algorithm obtained by replacing the backbone network with an EfficientViT network on the basis of the improved YOLOv8 target detection algorithm.
[0144] In summary, the beneficial effects of the present invention are as follows:
[0145] The present invention provides a method for precise detection and tracking of photovoltaic modules based on video, which uses a specific unmanned aerial vehicle flight method to collect qualified infrared video data of a photovoltaic power station. Through the specially designed eYOLO-PV model for photovoltaic power station videos, photovoltaic modules are detected. The average precision on the test set reaches 99.73%, and for an infrared video stream containing multiple photovoltaic modules, the time required to detect multiple targets simultaneously in each frame is only 3.12 milliseconds. On the basis of component target detection, a lightweight SORT photovoltaic module tracking algorithm based on video is developed to achieve efficient tracking of the same photovoltaic module in multiple consecutive different video frames. It can ensure that the MOTA of the tracking is 1, achieving 100% accuracy, and the tracking speed per frame is only 2 ms.
[0146] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0147] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A method for accurate detection and tracking of photovoltaic modules based on video, characterized in that: include: Get the video of PV panels taken by drone; Performing image extraction on video frames in the photovoltaic module video to obtain an image target detection data set; Input the image target detection data set into a trained target detection model to identify the component target frame of each photovoltaic component in the photovoltaic component video; the target detection model is a detection model constructed by the eYOLO-PV target detection algorithm; According to the photovoltaic component video and the component target frame of the photovoltaic component, based on the lightweight component tracking algorithm SORT, the photovoltaic component is tracked in real time; The training process of the target detection model is: Acquire sample infrared image data of photovoltaic modules in a photovoltaic power station; the sample infrared image data is obtained through historical photovoltaic module videos; Performing component box annotation on the sample infrared image data to obtain annotated sample infrared image data; Based on the labeled sample infrared image data, the eYOLO-PV target detection algorithm is trained to obtain a trained target detection model; wherein the eYOLO-PV target detection algorithm is an algorithm obtained by replacing the backbone network with the EfficientViT network on the basis of the improved YOLO target detection algorithm; the YOLO target detection algorithm is an algorithm obtained by using the improved CSPDarket53 as the backbone network, replacing the CSP module in the backbone network with the C2f module, using the PAN+FPN structure in the neck part to construct a feature pyramid, and using the decoupling head structure as the detection part on the basis of the original YOLO series target detection algorithm; According to the photovoltaic component video and the component target frame of the photovoltaic component, based on the lightweight component tracking algorithm SORT, the photovoltaic component is tracked in real time, specifically including: Based on the lightweight component tracking algorithm SORT, the motion state of the tracking target in the current video frame is predicted; the tracking target is the component target frame of the photovoltaic component; the lightweight component tracking algorithm SORT is used to use the Kalman filter algorithm to predict the position of the component in the next frame. The lightweight component tracking algorithm SORT is an algorithm obtained by removing the change rate of the component area in the state vector of the Kalman filter on the basis of the original SORT model; Based on the tracking target in the next video frame and the tracking target in the current video frame, the Hungarian algorithm is used to perform target matching to obtain a matching result of the tracking target; When the matching result is a matching failure, it is determined that the tracking target is lost, and the component target frame of the photovoltaic component corresponding to the tracking target is deleted; When the matching result is a successful match, the state of the component target frame of the photovoltaic component is updated based on the tracking target in the current video frame; the state is the position and speed of the component target frame.
2. A method for accurate detection and tracking of photovoltaic components based on video according to claim 1, characterized in that: Before obtaining the PV module video taken by the drone, it also includes: Determine whether the climate of the shooting location meets the shooting standards; When the shooting standards are met, the flight speed, flight altitude, camera angle and shooting route of the drone are set, and the drone is controlled to shoot the photovoltaic components.
3. The method for accurate detection and tracking of photovoltaic components based on video according to claim 1, characterized in that: After real-time tracking of the photovoltaic components, the method further includes: MOTA is used to evaluate the lightweight SORT target tracking algorithm, and the tracking accuracy of the lightweight SORT target tracking algorithm is obtained.
4. The method for accurate detection and tracking of photovoltaic components based on video according to claim 3, characterized in that: The expression of the MOTA is as follows: Among them, FN t Indicates the number of missed targets in time frame t; missed targets refer to targets that actually exist but are not detected by the tracking algorithm; FP t It represents the number of misdetected targets in time frame t; misdetection refers to the situation where the algorithm mistakenly identifies the background or other non-target objects as targets; IDSW t represents the number of ID switches that occurred in time frame t; ID switching refers to the situation where the algorithm incorrectly assigns the ID of a tracked target to another target; GT t represents the true number of targets in time frame t.
5. A photovoltaic component precision detection and tracking system based on the video-based photovoltaic component precision detection and tracking method according to any one of claims 1 to 4, characterized in that: include: Video acquisition module, used to acquire the video of photovoltaic modules taken by drone; An image extraction module is used to extract video frames in the photovoltaic module video as images to obtain an image target detection data set; A detection module, used to input the image target detection data set into a trained target detection model to identify a component target frame of each photovoltaic component in the photovoltaic component video; A tracking module, used for tracking the photovoltaic component in real time based on the photovoltaic component video and the component target frame of the photovoltaic component and based on a lightweight component tracking algorithm SORT; Also included: a training module, used for training the target detection model; The training module specifically includes: An image acquisition submodule is used to acquire sample infrared image data of photovoltaic modules in a photovoltaic power station; the sample infrared image data is obtained through historical photovoltaic module videos; A labeling submodule, used for labeling the sample infrared image data with component frames to obtain labeled sample infrared image data; The training submodule is used to train the eYOLO-PV target detection algorithm based on the labeled sample infrared image data to obtain a trained target detection model; wherein the eYOLO-PV target detection algorithm is an algorithm obtained by replacing the backbone network with the EfficientViT network on the basis of the improved YOLOv8 target detection algorithm.
Citation Information
Patent Citations
Forward driving method and driving device in photovoltaic panel matrix
CN116880475A
Moving target detection and tracking method for robot navigation
CN117333833A