A method and system for panoramic stitching of infrared videos of insulators taken by drones
Through deep learning and feature point optimization algorithms, the noise and feature matching problems in the stitching of infrared video images of drones aerial photography are solved, and efficient and accurate insulator panoramic image stitching is achieved, which improves the efficiency and accuracy of power grid patrol.
Patent Information
- Application Number
- CN202211166423.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2042-09-23
AI Technical Summary
There are problems such as low resolution, high noise, strong background characteristics than foreground characteristics, and high repeatability of insulator films in the aerial infrared video stitching, resulting in image stitching errors and missing.
Deep learning method is adopted to build infrared image noise removal networks and insulator image segmentation networks, combine feature point detection and optimization algorithms, infrared video image preprocessing and feature point registration, and optimize the stitching process using global transformation matrix.
It realizes efficient removal of infrared image noise, eliminates background interference, and accurately splices the panoramic images of insulators, improving the efficiency and accuracy of power grid patrols.
Smart Images

Figure CN115578256B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image stitching, and in particular to a method and system for stitching a panoramic infrared video of an insulator photographed by an unmanned aerial vehicle (UAV). Background Art
[0002] To ensure the normal operation of power transmission lines, grid inspectors are required to regularly inspect them. However, the complex geography of power lines often creates significant inconvenience for grid inspectors. In recent years, with the advancement of drone and infrared technology, grid inspectors can now use drones equipped with infrared cameras to take aerial photos of power lines, and then conduct inspections based on image or video analysis.
[0003] During drone aerial photography, the closer the drone is to the target, the more detail the image can capture, but the smaller the field of view. Conversely, the farther the drone is from the target, the larger the field of view, but the smaller the detail captured. Therefore, drone aerial photography requires a trade-off between image detail and field of view, making it difficult to obtain detailed, wide-field images. To obtain rich, wide-field images, it's necessary to capture and stitch them together to create panoramic images with a large field of view, providing more information for inspection and analysis.
[0004] Infrared imaging has extensive applications in military, industrial, and medical fields. Compared to visible light detectors, infrared detectors can collect temperature information of objects in a non-contact manner. They can also capture scenes that are not visible light, even in low-light environments such as at night, in heavy rain, fog, and snow. Due to limitations in infrared thermal imaging detector manufacturing technology, each sensor element is larger than that of a visible light detector. For sensors of the same mechanical size, the resolution of infrared detectors is significantly lower than that of visible light sensors. Furthermore, due to material and industrial imperfections in infrared detectors, each pixel has different response characteristics. These numerical variations in response characteristics also lead to non-uniformity in infrared images. During infrared detector imaging, uneven brightness around the target, as well as noise from the circuit components themselves and their interactions within the imaging system, can all cause point-like Gaussian noise. Consequently, the resolution of infrared detectors commonly used in inspections is 640×480, and the resulting infrared images suffer from low resolution, reduced detail, and significant non-uniform noise and point-like Gaussian noise.
[0005] During power line inspections, power grid inspectors need to locate faulty insulator segments within the insulator string. Using aerial infrared panoramic images of the insulators allows for rapid identification of these locations. The production of infrared panoramic images taken by drones for power line inspections primarily involves video frame preprocessing, feature point detection, feature point registration, transformation matrix calculation, and image stitching. Video frame preprocessing primarily involves image denoising and insulator region segmentation.
[0006] To obtain richer insulator information, panoramic stitching of drone-photographed infrared videos of insulators requires the drone to be in close spatial proximity to the insulators, a type of close-range image stitching. Unlike conventional long-range image stitching, in aerial infrared insulator videos, the horizontal movement and rotation of the drone result in significant angular differences between the foreground and background, and the background features are much stronger than those of the foreground. Therefore, the homography matrix calculated using background feature points and registration information is not suitable for the projective transformation of the foreground insulators. Furthermore, the insulator segments within an insulator string are highly repetitive. When registering the feature points of the insulator segments, using the interval sampling method used in traditional image stitching to obtain video frames can easily misalign adjacent insulator feature points, resulting in missing or redundant insulator segments in the panoramic image. Summary of the Invention
[0007] In view of this, in order to solve the problems in the prior art, the present invention proposes a method and system for panoramic stitching of infrared videos of insulators taken by drones. Based on deep learning, the infrared videos of insulators taken by drones are input to obtain corresponding infrared panoramic images, and panoramic stitching of infrared videos of insulators taken by drones is realized, thereby realizing preprocessing such as infrared image denoising and infrared insulator image segmentation, as well as specific stitching solutions for the characteristics of insulators such as few details, weak features, high noise, and high repeatability.
[0008] The present invention solves the above problems through the following technical means:
[0009] On the one hand, the present invention provides a method for stitching a panoramic infrared video of an insulator taken by a drone, comprising the following steps:
[0010] Input the infrared video of the insulator taken by the drone to obtain multiple infrared images of the insulator;
[0011] Build and train an infrared image noise removal network model, and use the trained infrared image noise removal network model to denoise the insulator infrared image to obtain a noise-free insulator infrared image;
[0012] Build and train an infrared insulator image segmentation network model, and use the trained infrared insulator image segmentation network model to perform background removal on the noise-free insulator infrared image to obtain a noise-free insulator infrared image containing only the insulator string part;
[0013] Select feature point detection algorithm to detect feature points of the pre-processed insulator infrared image;
[0014] After extracting the feature points, the feature point matching algorithm is used to align the extracted feature points, and the optimization algorithm is used to screen the alignment;
[0015] Calculate a homography matrix for adjacent video frames, and use the homography matrix to project the key frame of the next frame onto the plane where the key frame of the previous frame is located;
[0016] When projecting video frames using a transformation matrix, key frames that are farther from the reference frame require multiple projection transformations before being projected onto the plane where the reference frame is located. By solving the global transformation matrix, the farther reference frame can be transformed onto the plane where the reference frame is located with only one projection. Key frames that require three or more single homography matrix projection transformations to be projected onto the reference frame are considered farther key frames.
[0017] When the length of the insulator string does not exceed the length of the spliced together of two images, the first frame is selected as the reference frame, the subsequent video frames are used as key frames, and the last frame is used as the splicing frame. When the length of the insulator string exceeds the length of the spliced together of two images, the middle frame of the full video is used as the reference frame, the forward and backward middle frames are used as key frames, and the first and last frames are used as splicing frames. They are projected onto the reference frame through the transformation matrix respectively, and the spliced frames are projected onto the plane where the reference frame is located. The insulator splicing panoramic image is output, thereby realizing the splicing of the insulator panoramic image.
[0018] As a preferred method, building and training an infrared image noise removal network model specifically includes:
[0019] Use a drone equipped with infrared thermal imaging equipment to shoot infrared videos of multiple insulators;
[0020] According to the mathematical model of non-uniform noise and the mathematical model of Gaussian noise, the corresponding parameter values in the noise mathematical model are determined by simulation software;
[0021] In an infrared video of an insulator captured by an infrared thermal imaging device, a plurality of infrared images of the insulator that are nearly noise-free are selected, and non-uniform noise and Gaussian noise that obey the noise mathematical model are added to the images to form input images and reference images for training an infrared image noise removal network model;
[0022] Most of the insulator infrared image pairs are used as training sets, and the remaining insulator infrared image pairs are used as test sets. Finally, multiple sets of training pairs are obtained for the two noise mathematical models respectively.
[0023] Build and train an infrared image noise removal network model based on convolutional neural network.
[0024] As a preferred method, a network model for infrared insulator image segmentation is constructed and trained, specifically:
[0025] Use a drone equipped with infrared thermal imaging equipment to shoot infrared videos of multiple insulators;
[0026] In the infrared video of insulators taken by drones, multiple groups of insulator infrared images are evenly sampled according to different insulator categories;
[0027] An image segmentation training dataset was constructed by enclosing the entire insulator region with a polygonal box. The sampled insulator infrared images were used to create an image segmentation training dataset using a data annotation tool, forming the input images and reference images for training the infrared insulator image segmentation network model.
[0028] Most of the insulator infrared image pairs are used as training sets, and the remaining insulator infrared image pairs are used as test sets, and finally multiple sets of training pairs are obtained for the infrared insulator image segmentation network model;
[0029] Build and train an infrared insulator image segmentation network model based on the FCN network.
[0030] Preferably, the infrared image noise removal network model includes feature extraction, nonlinear mapping and reconstruction; feature extraction is to extract features from the noise in the input infrared image through convolution to obtain multiple high-dimensional spatial matrices containing infrared noise features; nonlinear mapping maps the high-dimensional spatial matrix containing infrared noise features to another high-dimensional spatial matrix, and at the same time introduces a pooling layer, an activation layer and a deconvolution layer, and adopts a maximum pooling method to highlight the characteristics of the noise and perform a downsampling operation on the image; the activation layer introduces a nonlinear activation function so that the increase in the number of network layers is not offset by linear simplification; the deconvolution layer performs an upsampling operation on the image; the reconstruction process adopts a long residual method to dimensionally superimpose the feature matrix after the deconvolution layer and the feature matrix before the pooling layer to obtain a fused feature matrix, and finally reconstruct it into a residual image of striped noise through a convolution layer.
[0031] Preferably, the infrared image noise removal network model includes 9 convolution layers, a pooling layer, a sub-pixel convolution layer and a dimensional overlay layer; the input image is only the Y channel of an image, the layer parameters of the first convolution layer are set to Conv(1,32,3,1,1), the subsequent convolution layers are set to Conv(32,32,3,1,1), the convolution layer parameters before the sub-pixel convolution are set to (32,128,3,1,1), the sub-pixel convolution layer parameters are 2, and the convolution layer after the overlay layer is set to Conv(64,1,3,1,1); the pooling layer convolution kernel is 2, and the step size is 2.
[0032] Preferably, the infrared insulator image segmentation network model includes feature extraction, feature fusion and pixel classification; the feature extraction part adopts the VGGNet16 network structure, and the VGGNet16 network adopts a method of replacing a large convolution kernel with multiple small convolution kernels. The receptive field obtained by stacking two 3x3 convolution kernels is equivalent to the receptive field of a 5x5 convolution kernel, and the receptive field obtained by stacking three 3x3 convolution kernels is equivalent to the receptive field of a 7x7 convolution kernel. Therefore, the use of a small convolution kernel reduces the parameters under the same receptive field. In addition, the use of a small convolution kernel is equivalent to performing more feature mapping; in the feature fusion part, the FCN network is based on the VG After the last three pooling layers of the GNet16 network, 8x, 16x, and 32x upsampling layers are established respectively. After the 5th pooling layer, 32x upsampling is performed directly to obtain the output of FCN-32s. At the same time, the 5th pooling layer is upsampled by 2 times and superimposed with the 4th pooling layer. The superimposed result is upsampled by 16 times to obtain the output of FCN-16s. Finally, the result of superimposing the 5th pooling layer and the 4th pooling layer is upsampled by two times and superimposed with the 3rd pooling layer. The superimposed result is upsampled by 8 times to obtain the final FCN-8s output. For pixel classification, a convolutional layer is used as the classifier with 32 input channels and the number of output channels being the number of categories.
[0033] Preferably, according to the mathematical model of non-uniform noise and the mathematical model of Gaussian noise, the corresponding parameter values in the noise mathematical model are determined by simulation software, specifically including:
[0034] Assuming that the actual pixel value of a point on the infrared image is V(i,j), where i is the horizontal coordinate and j is the vertical coordinate; the strip noise S(i,j) is expressed as:
[0035] S(i,j)=G(V(i,j))
[0036] Where G represents the nonlinear mapping function of the strip noise. A polynomial is usually used to simulate the noise. The noise in the jth column is expressed as:
[0037]
[0038] in is the coefficient of the j-th column polynomial, which is set to a random number in the range [-0.1, 0.1], and the polynomial order M is set to 3; V m (i,j),V m-1 (i,j)…V 0(i, j) is an instantiation of V(i, j), which refers to the Mth power of the pixel value at the coordinate (i, j), where m, m-1…0 are all instantiations of M. Assuming that the actual pixel value of a point on the infrared image is V(i, j), the point noise F(i, j) is expressed as:
[0039] F(i,j)=H(V(i,j))
[0040] Where H represents the nonlinear mapping function of point noise, which is simulated in the form of Gaussian function;
[0041]
[0042] where σ 2 is the variance of Gaussian noise, and u is the mean of Gaussian noise. According to the simulation results of the simulation software, the variance of Gaussian noise is set to a random number in the range of [0.003, 0.004], and the expectation is set to 0.
[0043] Preferably, the following restrictions are added to the parameters of the homography matrix:
[0044] For the homography matrix h:
[0045]
[0046] where h ij Represents the parameters of the i-th row and j-th column in the homography matrix h, with the following restrictions:
[0047] 5<|h 13 |<50
[0048] |h 21 |<0.01
[0049] |h 31 |<0.01
[0050] When the homography matrix parameters of adjacent frames do not meet the above requirements, the frame is skipped and the homography matrix is solved using the next frame and the previous frame until the above requirements are met.
[0051] Preferably, the global transformation matrix is solved so that the distant reference frame can be transformed to the plane where the reference frame is located by only one projection, specifically:
[0052] In the projection process, the transformation matrix and matrix multiplication form a group, which is denoted as:
[0053] G=(A,·),
[0054] Where A is the transformation matrix, · is the matrix multiplication, and G is the group consisting of the transformation matrix and the matrix multiplication. Due to the closed property of the group, that is:
[0055]
[0056] Among them, A1 and A2 are specific elements of the transformation matrix set A. In the continuous projection process, it is equivalent to multiplying adjacent homography matrices. The homography matrix from the nth frame image to the first frame image is:
[0057]
[0058] Where H is the global homography matrix, h i is the homography matrix from the i+1th frame to the ith frame. The transformation matrix from the subsequent farther video frames to the reference frame is obtained by directly multiplying multiple transformation matrices through matrix multiplication. The subsequent farther video frames can be transformed to the plane where the reference frame is located through only one projection transformation.
[0059] On the other hand, the present invention provides a panoramic stitching system for infrared video of insulators taken by drones, comprising:
[0060] The video input module is used to input infrared videos of insulators taken by drones to obtain multiple infrared images of insulators;
[0061] Image denoising module, used to build and train an infrared image noise removal network model, and use the trained infrared image noise removal network model to denoise the insulator infrared image to obtain a noise-free insulator infrared image;
[0062] The background removal module is used to build and train an infrared insulator image segmentation network model. The trained infrared insulator image segmentation network model is used to perform background removal on the noise-free insulator infrared image to obtain a noise-free insulator infrared image containing only the insulator string part.
[0063] A feature point detection module is used to select a feature point detection algorithm to perform feature point detection on the pre-processed infrared image of the insulator;
[0064] The feature point registration module is used to extract feature points, register them using a feature point matching algorithm, and filter the registration using an optimization algorithm;
[0065] The homography matrix calculation module is used to calculate a homography matrix for adjacent video frames. The homography matrix can be used to project the key frame of the next frame onto the plane where the key frame of the previous frame is located.
[0066] The global homography matrix calculation module is used to project the video frames using the transformation matrix. When a key frame that is farther away from the reference frame needs to undergo multiple projection transformations before it can be projected onto the plane where the reference frame is located, the global transformation matrix is solved so that the farther reference frame can be transformed onto the plane where the reference frame is located with only one projection. The key frame that needs to undergo three or more single homography matrix projection transformations to be projected onto the reference frame is the farther key frame.
[0067] The image stitching module is used to select the first frame as the reference frame, the subsequent video frames as the key frames, and the last frame as the stitching frame when the length of the insulator string does not exceed the length of the stitching of the two images; when the length of the insulator string exceeds the length of the stitching of the two images, the middle frame of the full video is used as the reference frame, the forward and backward middle frames are used as the key frames, the first frame and the last frame are used as the stitching frames, and they are projected onto the reference frame through the transformation matrix respectively, and the stitching frame is projected onto the plane where the reference frame is located, and the insulator stitching panoramic image is output, thereby realizing the stitching of the insulator panoramic image.
[0068] Compared with the prior art, the beneficial effects of the present invention include at least:
[0069] This paper, focusing on stitching panoramic infrared images of insulators taken from drones, proposes a panoramic image stitching scheme that includes infrared image preprocessing and insulator-specific feature-based stitching. Training sets for infrared image denoising and infrared insulator image segmentation were created, addressing the lack of specific infrared datasets in deep learning, which typically requires training with visible training sets and then transferring them to infrared data. An infrared image denoising network was designed and trained, achieving good denoising results while meeting real-time requirements. A specific insulator image segmentation network was trained based on the FCN image segmentation network, eliminating interference from background features with strong features on foreground insulators during feature point detection and registration. This paper proposes a specific stitching scheme for insulators with high reproducibility, low detail, high noise, and pixel values susceptible to temperature. This allows power grid inspectors to simply use drones equipped with infrared equipment to capture infrared video and record insulator information, which can then be automatically output as a panoramic image. This improves the efficiency of insulator fault detection in the power grid and provides a solution for insulator string centerline detection and further fault analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0071] Figure 1Some images in the non-uniform noise denoising training set produced by the present invention;
[0072] Figure 2 Some images in the Gaussian noise denoising training set produced for the present invention;
[0073] Figure 3 Some images from the insulator image segmentation training set of two different categories produced for the present invention;
[0074] Figure 4 This is the network model diagram for infrared image noise removal of the present invention;
[0075] Figure 5 This is a flow chart of a method for stitching panoramic infrared images of insulators taken by drones of the present invention;
[0076] Figure 6 This is a mosaic result diagram of panoramic images of two different types of insulator strings stitched together in the present invention;
[0077] Figure 7 This is a structural diagram of the panoramic image stitching system of the infrared video of insulators taken by UAVs of the present invention. DETAILED DESCRIPTION
[0078] To make the above-mentioned objectives, features, and advantages of the present invention more clearly understood, the technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are also within the scope of protection of the present invention.
[0079] Example 1
[0080] The present invention proposes a panoramic stitching method for infrared videos of insulators taken by drones, which mainly includes the following three steps.
[0081] Step 1: Build an infrared dataset
[0082] A drone equipped with infrared thermal imaging equipment captured 300 videos of insulators. The drone was kept as parallel as possible to the insulator string, slowly moving from one end of the string to the other, capturing the entire string. "Insulator" in this context refers to a complete insulator substructure, consisting of one to three "insulator strings," each containing multiple "insulator segments."
[0083] Based on the mathematical model of non-uniform noise and the mathematical model of Gaussian noise, the corresponding parameter values in the noise mathematical model were determined by simulation using MATLAB simulation software. Assuming that the actual pixel value of a point on the infrared image is V(i,j), where i is the horizontal coordinate and j is the vertical coordinate; the stripe noise S(i,j) can be expressed as:
[0084] S(i,j)=G(V(i,j))
[0085] Where G represents the nonlinear mapping function of the strip noise. A polynomial is usually used to simulate the noise. The noise in the jth column can be expressed as:
[0086]
[0087] in is the coefficient of the j-th column polynomial, which is set to a random number in the range [-0.1, 0.1], and the polynomial order M is set to 3. m (i,j),V m-1 (i,j)…V 0 (i, j) is an instantiation of V(j, j), which refers to the Mth power of the pixel value at the coordinate (i, j), where m, m-1…0 are all instantiations of M. Assuming that the actual pixel value of a point on the infrared image is V(i, j), the point noise F(i, j) can be expressed as:
[0088] F(i,j)=H(V(i,j))
[0089] Where H represents the nonlinear mapping function of point noise, which is simulated in the form of Gaussian function
[0090]
[0091] where σ 2 is the variance of Gaussian noise, and u is the mean of Gaussian noise. According to the MATLAB simulation results, the variance of Gaussian noise is set to a random number in the range [0.003, 0.004], and the expectation is set to 0.
[0092] 200 nearly noiseless infrared images were selected from videos captured by infrared equipment. Non-uniform noise and Gaussian noise conforming to the above model were added to them to form the input and reference images for neural network training. 180 image pairs were used as the training set, and the remaining 20 image pairs were used as the test set. To increase the amount of data and the number of iterations of the neural network, the images were randomly cropped into 54x54 sub-images. Data augmentation was performed through rotation and horizontal flipping. Ultimately, 100,000 training pairs were obtained for each of the two noise models. Figure 1These are some of the non-uniform noise model denoising training pairs, where the first row is the infrared image without non-uniform noise, which serves as the reference data for neural network training, and the second row is the infrared image with non-uniform noise added, which serves as the input data for neural network training. Figure 2 These are some of the Gaussian noise model denoising training pairs, where the first row is the infrared image without Gaussian noise, which is used as the supervision data for neural network training, and the second row is the infrared image with Gaussian noise added, which is used as the input data for neural network training.
[0093] When constructing the image segmentation network, we discovered that the traditional method of creating an image segmentation dataset along the object wheel resulted in missing corner points along the insulator's edge during registration. This also resulted in the addition of some image segmentation masks, which introduced erroneous corner points and reduced registration accuracy. In insulator videos, the insulators comprise a significant portion of the foreground area of the insulator string. Therefore, we constructed the image segmentation dataset by enclosing the entire insulator region with a polygonal box. We sampled 1,000 sets of raw images from drone aerial video footage of insulators, balanced according to the insulator type. These images were then used to create an image segmentation training dataset using the labelme data annotation tool. 900 of these image pairs served as the training set, and the remaining 100 served as the test set. Figure 3 For some of the insulator image segmentation model training pairs, the first line is the denoised video frame, which is used as the input data for image segmentation network training, and the second line is the segmentation mask annotated using the labelme data annotation tool, which is used as the supervision data for image segmentation network training.
[0094] Step 2: Preprocessing the neural network and training it
[0095] The infrared noise removal model based on convolutional neural network mainly includes three parts: feature extraction, nonlinear mapping and reconstruction. Feature extraction mainly extracts features from the noise in the input infrared image through convolution to obtain multiple high-dimensional space matrices containing infrared noise features. Nonlinear mapping maps the high-dimensional space matrix containing infrared noise features to another high-dimensional space matrix. In this process, pooling layer, activation layer and deconvolution layer are introduced at the same time. The maximum value pooling method is adopted to highlight the characteristics of the noise more, and at the same time perform an operation similar to downsampling on the image, reducing computational overhead and increasing the receptive field; the activation layer introduces a nonlinear activation function so that the increase in the number of network layers is not offset by linear simplification; the deconvolution layer is equivalent to an upsampling operation. The reconstruction process adopts a long residual method to superimpose the feature matrix after the deconvolution layer and the feature matrix before the pooling layer in terms of dimension to obtain the fused feature matrix. Finally, it is reconstructed into a residual image of striped noise through a convolution layer. Its network structure model is as follows: Figure 4 shown.
[0096] Figure 4 The convolutional neural network shown in the figure consists of nine convolutional layers, one pooling layer, one sub-pixel convolution layer, and a dimensional stacking layer. The input image is only the Y channel of an image. The layer parameters of the first convolutional layer are set to Conv(1,32,3,1,1), and the subsequent convolutional layers are set to Conv(32,32,3,1,1). The convolutional layer parameters before the sub-pixel convolution are set to (32,128,3,1,1), the sub-pixel convolution layer parameters are 2, and the convolutional layer after the stacking layer is set to Conv(64,1,3,1,1). The pooling layer has a convolution kernel of 2 and a stride of 2. The Adam optimizer is used as the optimizer, with a bitch_size of 64 and an initial learning rate of 0.0003. The learning rate is decayed to one-tenth of its original value after every 40 cycles. The entire network was trained on a PC running Windows 10, using the PyTorch deep learning platform, CUDA version 10.1, Cudnn version 7.0, an NVIDIA GTX1060 graphics card (6GB of video memory), and an Intel Core i5 6300HQ CPU. Processing a 640x480 image took approximately 0.438 seconds on the CPU and 0.015 seconds on the GPU.
[0097] The infrared insulator image segmentation network based on the FCN network mainly includes three parts: feature extraction, feature fusion, and pixel classification. The feature extraction part adopts the VGGNet16 network structure. The VGGNet16 network uses multiple small convolution kernels instead of a large convolution kernel. The receptive field obtained by stacking two 3x3 convolution kernels is equivalent to the receptive field of a 5x5 convolution kernel, and the receptive field of three 3x3 convolution kernels is equivalent to that of a 7x7 convolution kernel. Therefore, using small convolution kernels can reduce parameters while maintaining the same receptive field. In addition, using small convolution kernels is equivalent to performing more feature mapping, which can further enhance the network's fitting ability. In the feature fusion part, the FCN network establishes 8x, 16x, and 32x upsampling layers after the last three pooling layers of the VGGNet16 network, respectively. After the fifth pooling layer, the output of FCN-32s is directly upsampled by a factor of 32. The fifth pooling layer is then upsampled by a factor of 2 and superimposed with the fourth pooling layer. The superimposed result is then upsampled by a factor of 16 to obtain the output of FCN-16s. Finally, the superimposed result of the fifth and fourth pooling layers is upsampled by a factor of two and superimposed with the third pooling layer. The superimposed result is then upsampled by a factor of 8 to obtain the final FCN-8s output. Because FCN-8s combines the features extracted from the VGGNet16 network with the feature maps of the upsampled output, it achieves better semantic segmentation results. For pixel classification, a convolutional layer is used as the classifier, with 32 input channels and the number of output channels equal to the number of categories.
[0098] Step 3: Aerial infrared insulator video stitching
[0099] Image denoising and segmentation preprocessing address the noise issues in the infrared video data of insulators captured by drones, as well as the problem of background features dominating foreground features in close-up stitching. This results in a noise-free image containing only the insulator strings. To address the lack of detail, changes in video angle caused by drone rotation, and temperature-induced changes in image pixel values, the SIFT feature point detection algorithm is used to detect feature points in the preprocessed image. SIFT feature points fully account for changes in illumination, scale, and rotation during image transformation, extracting precise image features and improving the accuracy of subsequent feature point registration. After feature point extraction, the extracted SIFT feature points are registered using a brute force matching algorithm, and the K-nearest neighbor algorithm is used to filter the alignment. Due to the high repetitiveness of insulators, full sampling is used for panoramic stitching to address the issue of insulator feature point registration errors caused by intermittent sampling, which can result in accurate or redundant insulators in the final panoramic image. In the input video, when the length of the insulator string does not exceed the length of the stitched two images, the first frame is selected as the reference frame, the subsequent frames as key frames, and the last frame as the stitching frame. When the input length exceeds two images, the middle frame of the full video is used as the reference frame, the forward and backward middle frames are used as key frames, and the first frame and the last frame are used as splicing frames, which are projected to the reference frame through the transformation matrix. The feature points of adjacent video frames are aligned, and a homography matrix is calculated for the adjacent video frames. The key frame of the latter frame can be projected to the plane where the key frame of the previous frame is located through the homography matrix, thereby realizing the splicing of panoramic images. When the video frame is projected and transformed by the transformation matrix, the key frame that is farther away from the reference frame needs to undergo multiple projection transformations before it can be projected into the plane where the reference frame is located, resulting in a reduction in the resolution of the video frame. The key frame corresponding to the projection transformation to the reference frame after three or more single homography matrices is the farther key frame; this embodiment solves the global transformation matrix so that the farther reference frame can be transformed to the plane where the reference frame is located through only one projection. During the projection process, the transformation matrix and the matrix multiplication can form a group, which is recorded as:
[0100] G=(A,·),
[0101] Where A is the transformation matrix, · is the matrix multiplication, and G is the group consisting of the transformation matrix and the matrix multiplication. Due to the closed property of the group, that is:
[0102]
[0103] Among them, A1 and A2 are specific elements of the transformation matrix set A. In the continuous projection process, it is equivalent to multiplying adjacent homography matrices. The homography matrix from the nth frame image to the first frame image is:
[0104]
[0105] Where H is the global homography matrix, h i The homography matrix from the i+1th frame to the ith frame can be directly multiplied by matrix multiplication to obtain the transformation matrix from the subsequent farther video frame to the reference frame. The subsequent farther video frame can be transformed to the plane where the reference frame is located through only one projection transformation, which greatly reduces the reduction in resolution.
[0106] However, the method of solving the homography matrix using the entire video sequence requires a high quality of drone aerial video. If there is a problem with one frame of video, it will cause an error in the global homography matrix, reducing the accuracy of the stitching. Therefore, the following restrictions are added to the parameters of the homography matrix:
[0107] For the homography matrix h:
[0108]
[0109] where h ij Represents the parameters of the i-th row and j-th column in the homography matrix h. Since the main motion component in the drone aerial photography process is the horizontal component and the rotation component is small, the restriction is:
[0110] 5<|h 13 |<50
[0111] |h 21 |<0.01
[0112] |h 31 |<0.01
[0113] If the homography matrix parameters of adjacent frames do not meet the above requirements, the frame is skipped and the homography matrix is solved using the next and previous frames until the requirements are met. By adding these restrictions, errors caused by video frame quality issues can be filtered out, greatly improving the robustness of the stitching.
[0114] The process of panoramic stitching method of infrared video of insulators taken by UAV is as follows: Figure 5 As shown, Figure 6 The panoramic image is automatically output after inputting drone aerial video and processing it with a panoramic stitching algorithm. The main insulators in the panoramic image are well stitched together, and the number of insulators is consistent with the actual number of insulators.
[0115] Example 2
[0116] like Figure 7 As shown, the present invention proposes a panoramic stitching system for infrared videos of insulators taken by drones, including a video input module, an image denoising module, a background removal module, a feature point detection module, a feature point registration module, a homography matrix calculation module, a global homography matrix calculation module and an image stitching module;
[0117] The video input module is used to input infrared videos of insulators taken by drones to obtain multiple infrared images of insulators;
[0118] The image denoising module is used to build and train an infrared image noise removal network model, and use the trained infrared image noise removal network model to denoise the insulator infrared image to obtain a noise-free insulator infrared image;
[0119] The background removal module is used to build and train an infrared insulator image segmentation network model, and uses the trained infrared insulator image segmentation network model to perform background removal processing on the noise-free insulator infrared image to obtain a noise-free insulator infrared image containing only the insulator string part;
[0120] The feature point detection module is used to select a feature point detection algorithm to perform feature point detection on the pre-processed infrared image of the insulator;
[0121] The feature point registration module is used to extract feature points, register the extracted feature points using a feature point matching algorithm, and filter the registration using an optimization algorithm;
[0122] The homography matrix calculation module is used to calculate a homography matrix for adjacent video frames, and the homography matrix can be used to project the key frame of the subsequent frame onto the plane where the key frame of the previous frame is located;
[0123] The global homography matrix calculation module is used to perform a projection transformation on the video frame through the transformation matrix. When the key frame farther from the reference frame needs to undergo multiple projection transformations before it can be projected onto the plane where the reference frame is located, the global transformation matrix is solved so that the farther reference frame can be transformed onto the plane where the reference frame is located through only one projection. The key frame corresponding to the key frame that needs to undergo three or more single homography matrix projection transformations to be projected onto the reference frame is the farther key frame.
[0124] The image stitching module is used to select the first frame as the reference frame, the subsequent video frames as the key frames, and the last frame as the stitching frame when the length of the insulator string does not exceed the length of the stitching of the two images; when the length of the insulator string exceeds the length of the stitching of the two images, the intermediate frame of the full video is used as the reference frame, the forward and backward intermediate frames are used as the key frames, the first frame and the last frame are used as the stitching frames, and they are respectively projected onto the reference frame through a transformation matrix, the stitching frame is projected and transformed onto the plane where the reference frame is located, and an insulator stitching panoramic image is output, thereby realizing the stitching of the insulator panoramic image.
[0125] The other features of this embodiment are the same as those of embodiment 1, so they will not be repeated here.
[0126] This paper, focusing on stitching panoramic infrared images of insulators taken from drones, proposes a panoramic image stitching scheme that includes infrared image preprocessing and insulator-specific feature-based stitching. Training sets for infrared image denoising and infrared insulator image segmentation were created, addressing the lack of specific infrared datasets in deep learning, which typically requires training with visible training sets and then transferring them to infrared data. An infrared image denoising network was designed and trained, achieving good denoising results while meeting real-time requirements. A specific insulator image segmentation network was trained based on the FCN image segmentation network, eliminating interference from background features with strong features on foreground insulators during feature point detection and registration. This paper proposes a specific stitching scheme for insulators with high reproducibility, low detail, high noise, and pixel values susceptible to temperature. This allows power grid inspectors to simply use drones equipped with infrared equipment to capture infrared video and record insulator information, which can then be automatically output as a panoramic image. This improves the efficiency of insulator fault detection in the power grid and provides a solution for insulator string centerline detection and further fault analysis.
[0127] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A panoramic stitching method for infrared video of insulators taken by drones, characterized in that: The steps include: Input the infrared video of the insulator taken by the drone to obtain multiple infrared images of the insulator; Build and train an infrared image noise removal network model, and use the trained infrared image noise removal network model to denoise the insulator infrared image to obtain a noise-free insulator infrared image; Build and train an infrared insulator image segmentation network model, and use the trained infrared insulator image segmentation network model to perform background removal on the noise-free insulator infrared image to obtain a noise-free insulator infrared image containing only the insulator string part; Select feature point detection algorithm to detect feature points of the pre-processed insulator infrared image; After extracting the feature points, the feature point matching algorithm is used to align the extracted feature points, and the optimization algorithm is used to screen the alignment; Calculate a homography matrix for adjacent video frames, and use the homography matrix to project the key frame of the next frame onto the plane where the key frame of the previous frame is located; When projecting video frames using a transformation matrix, key frames that are farther from the reference frame require multiple projection transformations before being projected onto the plane where the reference frame is located. By solving the global transformation matrix, the farther reference frame can be transformed onto the plane where the reference frame is located with only one projection. Key frames that require three or more single homography matrix projection transformations to be projected onto the reference frame are considered farther key frames. When the length of the insulator string does not exceed the length of the spliced together of two images, the first frame is selected as the reference frame, the subsequent video frames are used as key frames, and the last frame is used as the splicing frame. When the length of the insulator string exceeds the length of the spliced together of two images, the middle frame of the full video is used as the reference frame, the forward and backward middle frames are used as key frames, and the first and last frames are used as splicing frames. They are projected onto the reference frame through the transformation matrix respectively, and the spliced frames are projected onto the plane where the reference frame is located. The insulator splicing panoramic image is output, thereby realizing the splicing of the insulator panoramic image.
2. The panoramic stitching method of infrared video of insulators taken by drones according to claim 1 is characterized in that: Build and train a network model for infrared image noise removal, including: Use a drone equipped with infrared thermal imaging equipment to shoot infrared videos of multiple insulators; According to the mathematical model of non-uniform noise and the mathematical model of Gaussian noise, the corresponding parameter values in the noise mathematical model are determined by simulation software; In an infrared video of an insulator captured by an infrared thermal imaging device, a plurality of infrared images of the insulator that are nearly noise-free are selected, and non-uniform noise and Gaussian noise that obey the noise mathematical model are added to the images to form input images and reference images for training an infrared image noise removal network model; Most of the insulator infrared image pairs are used as training sets, and the remaining insulator infrared image pairs are used as test sets. Finally, multiple sets of training pairs are obtained for the two noise mathematical models respectively. Build and train an infrared image noise removal network model based on convolutional neural network.
3. The panoramic stitching method of infrared video of insulators taken by drones according to claim 1 is characterized in that: Build and train the infrared insulator image segmentation network model, specifically: Use a drone equipped with infrared thermal imaging equipment to shoot infrared videos of multiple insulators; In the infrared video of insulators taken by drones, multiple groups of insulator infrared images are evenly sampled according to different insulator categories; An image segmentation training dataset was constructed by enclosing the entire insulator region with a polygonal box. The sampled insulator infrared images were used to create an image segmentation training dataset using a data annotation tool, forming the input images and reference images for training the infrared insulator image segmentation network model. Most of the insulator infrared image pairs are used as training sets, and the remaining insulator infrared image pairs are used as test sets, and finally multiple sets of training pairs are obtained for the infrared insulator image segmentation network model; Build and train an infrared insulator image segmentation network model based on the FCN network.
4. The panoramic stitching method of infrared video of insulators taken by drones according to claim 1 is characterized in that: The infrared image noise removal network model includes feature extraction, nonlinear mapping and reconstruction; feature extraction is to extract features from the noise in the input infrared image through convolution to obtain multiple high-dimensional spatial matrices containing infrared noise features; The nonlinear mapping maps the high-dimensional spatial matrix containing infrared noise features to another high-dimensional spatial matrix. At the same time, the pooling layer, activation layer and deconvolution layer are introduced. The maximum pooling method is used to highlight the characteristics of the noise and perform a downsampling operation on the image. The activation layer introduces a nonlinear activation function so that the increase in the number of network layers is not offset by linear simplification. The deconvolution layer performs an upsampling operation on the image. The reconstruction process uses the long residual method to superimpose the feature matrix after the deconvolution layer and the feature matrix before the pooling layer in terms of dimension to obtain a fused feature matrix. Finally, a convolution layer is used to reconstruct it into a residual image of striped noise.
5. The panoramic stitching method of infrared video of insulators taken by drones according to claim 1 is characterized in that: The infrared image noise removal network model includes 9 convolutional layers, a pooling layer, a sub-pixel convolution layer and a dimensional overlay layer; the input image is only the Y channel of an image, the layer parameters of the first convolutional layer are set to Conv(1,32,3,1,1), the subsequent convolutional layers are set to Conv(32,32,3,1,1), the convolutional layer parameters before the sub-pixel convolution are set to (32,128,3,1,1), the sub-pixel convolution layer parameters are 2, and the convolution layer after the overlay layer is set to Conv(64,1,3,1,1); the pooling layer convolution kernel is 2, and the stride is 2.
6. The panoramic stitching method of infrared video of insulators taken by drones according to claim 1 is characterized in that: The infrared insulator image segmentation network model includes feature extraction, feature fusion and pixel classification; the feature extraction part adopts the VGGNet16 network structure, and the VGGNet16 network adopts multiple small convolution kernels instead of a large convolution kernel. The receptive field obtained by stacking two 3x3 convolution kernels is equivalent to the receptive field of a 5x5 convolution kernel, and the receptive field obtained by stacking three 3x3 convolution kernels is equivalent to the receptive field of a 7x7 convolution kernel. Therefore, using a small convolution kernel reduces parameters under the same receptive field. In addition, using a small convolution kernel is equivalent to performing more feature mapping; in the feature fusion part, the FCN network is based on the VGGNe After the last three pooling layers of the t16 network, 8x, 16x, and 32x upsampling layers are established respectively. After the 5th pooling layer, 32x upsampling is performed directly to obtain the output of FCN-32s. At the same time, the 5th pooling layer is upsampled by 2 times and superimposed with the 4th pooling layer. The superimposed result is upsampled by 16 times to obtain the output of FCN-16s. Finally, the result of superimposing the 5th pooling layer and the 4th pooling layer is upsampled by two times and superimposed with the 3rd pooling layer. The superimposed result is upsampled by 8 times to obtain the final FCN-8s output. For pixel classification, a convolutional layer is used as the classifier with 32 input channels and the number of output channels being the number of categories.
7. The method for panoramic stitching of infrared videos of insulators taken by drones according to claim 2, characterized in that: According to the mathematical model of non-uniform noise and the mathematical model of Gaussian noise, the corresponding parameter values in the noise mathematical model are determined by simulation software, including: Assume that the actual pixel value of a point on the infrared image is V(i, j), where i is the horizontal coordinate and j is the vertical coordinate; then the strip noise S(i, j) is expressed as: S(i, j) = G(V(i, j)) Where G represents the nonlinear mapping function of the strip noise. A polynomial is usually used to simulate the noise. The noise in the jth column is expressed as: in is the coefficient of the j-th column polynomial, which is set to a random number in the range [-0.1, 0.1], and the polynomial order M is set to 3; V m (i, j), V m-1 (i, j)...V 0 (i, j) is an instantiation of V(i, j), which refers to the Mth power of the pixel value at the coordinate (i, j), where m, m-1…0 are all instantiations of M. Assuming that the actual pixel value of a point on the infrared image is V(i, j), the point noise F(i, j) is expressed as: F(i, j) = H(V(i, j)) Where H represents the nonlinear mapping function of point noise, which is simulated in the form of Gaussian function; where σ 2 is the variance of Gaussian noise, and u is the mean of Gaussian noise. According to the simulation results of the simulation software, the variance of Gaussian noise is set to a random number in the range of [0.003, 0.004], and the expectation is set to 0.
8. The panoramic stitching method of infrared video of insulators taken by drones according to claim 1 is characterized in that: Add the following restrictions to the parameters of the homography matrix: For the homography matrix h: where h ij Represents the parameters of the i-th row and j-th column in the homography matrix h, with the following restrictions: 5<|h 13 |<50 |h 21 |<0.01 |h 31 |<0.01 When the homography matrix parameters of adjacent frames do not meet the above requirements, the frame is skipped and the homography matrix is solved using the next frame and the previous frame until the above requirements are met.
9. The panoramic stitching method of infrared video of insulators taken by drones according to claim 1 is characterized in that: By solving the global transformation matrix, the distant reference frame can be transformed to the plane where the reference frame is located with only one projection, specifically: In the projection process, the transformation matrix and matrix multiplication form a group, which is denoted as: G=(A,·), Where A is the transformation matrix, · is the matrix multiplication, and G is the group consisting of the transformation matrix and the matrix multiplication. Due to the closed property of the group, that is: Among them, A1 and A2 are specific elements of the transformation matrix set A. In the continuous projection process, it is equivalent to multiplying adjacent homography matrices. The homography matrix from the nth frame image to the first frame image is: Where H is the global homography matrix, h i is the homography matrix from the i+1th frame to the ith frame. The transformation matrix from the subsequent farther video frames to the reference frame is obtained by directly multiplying multiple transformation matrices through matrix multiplication. The subsequent farther video frames can be transformed to the plane where the reference frame is located through only one projection transformation.
10. A panoramic mosaic system for infrared video of insulators taken by drones, characterized in that: include: The video input module is used to input infrared videos of insulators taken by drones to obtain multiple infrared images of insulators; Image denoising module, used to build and train an infrared image noise removal network model, and use the trained infrared image noise removal network model to denoise the insulator infrared image to obtain a noise-free insulator infrared image; The background removal module is used to build and train an infrared insulator image segmentation network model. The trained infrared insulator image segmentation network model is used to perform background removal on the noise-free insulator infrared image to obtain a noise-free insulator infrared image containing only the insulator string part. A feature point detection module is used to select a feature point detection algorithm to perform feature point detection on the pre-processed infrared image of the insulator; The feature point registration module is used to extract feature points, register them using a feature point matching algorithm, and filter the registration using an optimization algorithm; The homography matrix calculation module is used to calculate a homography matrix for adjacent video frames. The homography matrix can be used to project the key frame of the next frame onto the plane where the key frame of the previous frame is located. The global homography matrix calculation module is used to project the video frames using the transformation matrix. When a key frame that is farther away from the reference frame needs to undergo multiple projection transformations before it can be projected onto the plane where the reference frame is located, the global transformation matrix is solved so that the farther reference frame can be transformed onto the plane where the reference frame is located with only one projection. The key frame that needs to undergo three or more single homography matrix projection transformations to be projected onto the reference frame is the farther key frame. The image stitching module is used to select the first frame as the reference frame, the subsequent video frames as the key frames, and the last frame as the stitching frame when the length of the insulator string does not exceed the length of the stitching of the two images; when the length of the insulator string exceeds the length of the stitching of the two images, the middle frame of the full video is used as the reference frame, the forward and backward middle frames are used as the key frames, the first frame and the last frame are used as the stitching frames, and they are projected onto the reference frame through the transformation matrix respectively, and the stitching frame is projected onto the plane where the reference frame is located, and the insulator stitching panoramic image is output, thereby realizing the stitching of the insulator panoramic image.
Citation Information
Patent Citations
Electric power inspection robot positioning method based on multi-sensor fusion
CN111739063A
Image processing method, image processing device and terminal equipment
CN111833285A