Super-Resolution Reconstruction Method and System for Motion Blurred Images in Transmission Line Inspection
By applying the super-resolution reconstruction method of transmission line patrol motion blur image in insulator detection, the problems of insufficient image quality and motion blur are solved, and the reconstruction of high-resolution images and the accuracy of target detection are improved.
Patent Information
- Application Number
- CN202510214372.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In the insulator detection, traditional technology has difficulty in accurately identifying the insulator state due to insufficient image quality and motion blur in the insulator detection.
A super-resolution reconstruction method for motion blur image inspection of transmission line is adopted to reconstruct high-resolution images through technical means such as timing processing, time attention mechanism, multi-layer convolution processing and 3D convolution processing.
It effectively improves the resolution of the image and enhances the clarity of image details, allowing the object detection technology to more accurately identify and analyze the state of insulator discharge and heating, and improves the accuracy and reliability of detection.
Smart Images

Figure CN119722461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method and system for super-resolution reconstruction of motion-blurred images in transmission line inspection. Background Art
[0002] Insulators are indispensable and important components in the power system, mainly used to support and isolate high-voltage wires to prevent current leakage through air or other media, thus ensuring the stability and safety of power transmission. With the continuous growth of power demand and the increasing complexity of the power system, higher requirements are put forward for the maintenance and detection of insulators.
[0003] Traditional insulator detection methods mainly include manual inspection and regular maintenance. Manual inspection relies on experienced technicians who check by visual observation and manual tools. This method is not only inefficient but also easily affected by human factors, making it difficult to detect potential faults in a timely manner. Although regular maintenance can prevent faults to a certain extent, due to the lack of real-time nature, there is often a large lag and it cannot meet the high standards of modern power systems.
[0004] With the rapid development of computer vision and deep learning technologies, image-based object detection technology has gradually become a new trend in insulator detection. Object detection technology can automatically identify and locate insulators in images, and monitor their discharge and heating states in real time through ultraviolet imagers, greatly improving the accuracy and efficiency of detection. Especially in the daily inspection of transmission lines, object detection technology can quickly detect defects in insulators, such as cracks, breakages, contaminations, etc., and take measures in a timely manner to avoid power outages and safety accidents caused by insulator failures.
[0005] However, although object detection technology shows great potential in insulator detection, it still faces some challenges. For example, the transmission line environment is complex, there are many background interferences, the shapes and colors of insulators are diverse, and the quality of infrared and ultraviolet images is often affected. These factors all increase the difficulty of detection. Especially when using drones for shooting, there may also be motion blur or insufficient image resolution, which makes it difficult for object detection technology to accurately identify the insulator state. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and system for super-resolution reconstruction of motion-blurred images in transmission line inspection, aiming to solve the problem that traditional technologies are difficult to accurately identify the insulator state due to motion blur and insufficient image resolution during the shooting of insulators.
[0007] In a first aspect, the present invention provides a method for super-resolution reconstruction of motion-blurred images in transmission line inspection, and the method includes:
[0008] Obtain a blurred image of an insulator, perform temporal processing on the blurred image to obtain a temporal image, and sequentially perform temporal attention, temporal aggregation, and temporal reconstruction operations on the temporal image to obtain a reconstructed image;
[0009] Perform multi-layer convolution processing, dilated convolution processing, residual processing, attention mechanism processing, dimensionality reduction processing, and dimensional transformation processing on the reconstructed image in sequence to obtain a first feature map;
[0010] Perform 3D convolution processing, spatial attention mechanism processing, tensor reshaping, permutation operation, 3D multi-head self-attention mechanism processing, shape restoration, high-dimensional convolution processing, residual connection, dynamic pooling, flattening operation, fully connected operation, tensor expansion, and dimensionality repetition processing on the reconstructed image in sequence to obtain a second feature map;
[0011] Increase the width dimension of the second feature map, splice the first feature map and the increased-width second feature map in the channel dimension to obtain a spliced feature map, perform channel dimensionality reduction and feature fusion on the spliced feature map through a 1x1x1 convolution to obtain a fused feature map, and weight the fused feature map to obtain a super-resolution image.
[0012] In a second aspect, the present invention provides a super-resolution reconstruction system for motion-blurred images of transmission line inspections, and the system includes:
[0013] A blurred image acquisition module, configured to obtain a blurred image of an insulator, perform temporal processing on the blurred image to obtain a temporal image, and sequentially perform temporal attention, temporal aggregation, and temporal reconstruction operations on the temporal image to obtain a reconstructed image;
[0014] A first convolution processing module, configured to perform multi-layer convolution processing, dilated convolution processing, residual processing, attention mechanism processing, dimensionality reduction processing, and dimensional transformation processing on the reconstructed image in sequence to obtain a first feature map;
[0015] A second convolution processing module, configured to perform 3D convolution processing, spatial attention mechanism processing, tensor reshaping, permutation operation, 3D multi-head self-attention mechanism processing, shape restoration, high-dimensional convolution processing, residual connection, dynamic pooling, flattening operation, fully connected operation, tensor expansion, and dimensionality repetition processing on the reconstructed image in sequence to obtain a second feature map;
[0016] A feature fusion module, configured to increase the width dimension of the second feature map, splice the first feature map and the increased-width second feature map in the channel dimension to obtain a spliced feature map, perform channel dimensionality reduction and feature fusion on the spliced feature map through a 1x1x1 convolution to obtain a fused feature map, and weight the fused feature map to obtain a super-resolution image.
[0017] In a third aspect, the present invention provides a storage medium that stores one or more programs, which, when executed by a processor, implement the above-mentioned super-resolution reconstruction method for motion-blurred images in transmission line inspections.
[0018] In a fourth aspect, the present invention provides an electronic device, which includes a memory and a processor, wherein:
[0019] The memory is used to store a computer program;
[0020] When the processor executes the computer program stored on the memory, it implements the above-mentioned super-resolution reconstruction method for motion-blurred images in transmission line inspections.
[0021] Compared with the prior art, the embodiments of the present invention have the following advantages:
[0022] 1. Through the super-resolution reconstruction technology proposed by the present invention, low-quality and blurred images can be enhanced to high-resolution images, thereby enhancing the clarity of image details, enabling the target detection technology to more accurately identify and analyze the state of insulator discharge and heating. This method effectively solves the problem of insufficient traditional image quality, improves the detection accuracy and reliability, and ensures the real-time and accuracy of insulator maintenance in the power system. Specifically, for the continuous frame images captured by the unmanned aerial vehicle, first, the blurred images are arranged in sequence to prevent the problem of temporal disorder in the images entering the backbone network, and then the temporal images with temporal characteristics are processed to combine temporal convolution with the temporal attention mechanism to analyze the regions in the image that are blurred due to motion. By extracting the information in the non-blurred frames, the information in the blurred frames is repaired, thereby improving the motion-blurred images and obtaining reconstructed images; then the improved reconstructed images are processed in two parallel branches. One branch processes the extraction of comprehensive spatial features from the input feature map (extracting spatial information features from low-resolution and spatially detailed-lacking pictures); the other branch processes the scenarios that require fine-grained feature enhancement and the processing of complex high-dimensional features (tasks with significant blurring and the need to enhance feature details in the input data). The first feature map and the second feature map after the two processes are fused. The fusion method is first to splice the two feature maps along the channel dimension, and then perform channel reduction and feature fusion through a 1x1x1 convolution (recombining channel information at each pixel position by learning weights), which can enhance the information exchange between the two streams while reducing the number of channels, and then weighted by the attention mechanism. The important channels are amplified, while the unimportant channels are suppressed, thereby enhancing the model's ability to focus on key features, and then through upsampling, the output image is made to have the same size as the input image to obtain a high-resolution image. Description of the Drawings
[0023] Figure 1 It is a flowchart of a super - resolution reconstruction method for motion - blurred images in transmission line inspection proposed in an embodiment of the present invention;
[0024] Figure 2 It is a schematic structural diagram of a super - resolution reconstruction system for motion - blurred images in transmission line inspection proposed in an embodiment of the present invention.
[0025] The following specific embodiments will further illustrate the present invention in conjunction with the above - mentioned drawings. Specific Embodiments
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention pertains. The words such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.
[0027] As Figure 1 shown, an embodiment of the present invention provides a super - resolution reconstruction method for motion - blurred images in transmission line inspection. The method includes steps S101 to S104, where:
[0028] Step S101: Obtain a blurred image of an insulator, perform temporal processing on the blurred image to obtain a temporal image, and sequentially perform temporal attention, temporal aggregation, and temporal reconstruction operations on the temporal image to obtain a reconstructed image;
[0029] It should be noted that the blurred image is obtained by the drone equipped with an imager. Then, the low-resolution image (blurred image) captured by the drone is input into the temporal convolutional module. The temporal convolutional module is a convolutional layer for processing temporal data. Its main function is to keep the order of the image unchanged in the temporal dimension (i.e., the dimension of the sequence) through convolutional operations, and at the same time perform convolutional operations on the spatial dimension. This module uses 3D convolution to process the temporal image data, maintaining temporal consistency without disturbing the temporal order of the image. The shape of the input image is (batch_size, 5, 416, 416, 3). batch_size represents the batch. Each sample has 5 frames of images, the size of each frame is 416x416, and each frame has 3 channels. After permute, the shape of X is converted to (2, 3, 5, 416, 416). After convolutional operations, the output shape is (batch_size, 64, 5, 416, 416), and then after permute, the output image is (batch_size, 5, 416, 416, 64). In this way, each sample still has 5 frames of images, the size of each frame is 416x416, but the number of channels has been converted from 3 to 64.
[0030] Specifically, the goal of the temporal convolutional module is to process multi-dimensional data with a temporal dimension. No convolutional operations are performed in the temporal dimension, and only convolutional operations are performed in the spatial dimensions (height and width). In this way, the model can capture temporal data and spatial features while maintaining the continuity of temporal information. The input of this module is (batch_size, 5, 416, 416, 3) (batch, time step, height, width, number of channels); the output is (batch_size, 5, 416, 416, 64) (batch, time step, height, width, number of channels).
[0031] By inputting (batch_size, 5, 416, 416, 3), where 5 is the number of image frames and can be adjusted according to specific circumstances, 416, 416 corresponds to the image size, and 3 is the number of channels. After permute for dimension transformation, (batch_size, 3, 5, 416, 416) is generated; after Conv3D convolution, (batch_size, 64, 5, 416, 416) is generated; after permute for dimension transformation, (batch_size, 5, 416, 416, 64) is generated.
[0032] Then it is input into the time feature network module, which extracts features, reconstructs in the time dimension, and restores time information from the input reconstructed image. By processing the features in the time and space positions, it finally outputs a feature map with rich time information, that is, learning the features of several adjacent steps and then reconstructing the best feature map. The specific working method is as follows: Input (batch_size, 5, 416, 416, 64) into the Conv3D, pooling layer, and activation function (this process does not change the image dimension), and output (batch_size, 5, 416, 416, 64) and input it into the time attention mechanism module.
[0033] The time attention mechanism module assigns weights to each time step in the time dimension, thereby enhancing the model's attention to important time steps. Its input is (batch_size, 5, 416, 416, 64) (batch, time step, height, width, number of channels); its output is (batch_size, 64, 1, 416, 416) (batch, number of channels, time step, height, width); the working method is as follows: First, pass through permute (dimension rearrangement) to output (batch_size, 64, 5, 416, 416); then flatten the spatial dimension to generate .
[0034] Calculate query, key, and value:
[0035]
[0036]
[0037]
[0038] Among them: Q (Query) is used to query the weight vectors of other time steps, represents the combination of all samples and channels, 5 represents the time dimension, represents the spatial features of the time step; K (Key) represents the features of each time step, used to calculate the similarity with Q. Through the inner product operation with Q, K can help determine which time steps should be paid more attention to, and its shape is the same as Q; V (Value) is the actual value calculated at each time step, which carries the feature information of each time step. By weighting with the attention weights (the similarity between Q and K), the weighted features are finally output, mainly providing the information of "needs to be aggregated". Then calculate the attention weights and output (batch_size, 64, 5, 416, 416), and then input (batch_size, 64, 5, 416, 416) into the time aggregation module.
[0039] The time aggregation module refers to aggregating the input features in the time dimension. It aggregates the time features at each position through a fully connected layer to obtain a compressed representation of time information. Its input is: (batch_size, 64, 5, 416, 416) (batch, number of channels, number of time steps, height, width); its output is: (batch_size, 64, 1, 416, 416) (batch, number of channels, number of time steps, height, width); and its working method is as follows:
[0040] (1) permute (dimension rearrangement, moving the time dimension to the last for easy operation on the time series of each spatial position) outputs (batch_size, 64, 416, 416, 5).
[0041] (2) Flatten the dimensions through x.view, flatten the height, width of the spatial dimension and the channel dimension, and flatten the first four quantities into one quantity , extract the time step feature vectors of each spatial position to form a two-dimensional matrix. Each row represents the features of a spatial position at different time steps, and each column corresponds to the features of all spatial positions at the same time step.
[0042] Specifically, form a two-dimensional matrix according to the following formula:
[0043] ;
[0044] Among them, represents the flattened matrix, represents the feature value at the H-th row and W-th column at time step T, B represents the number of samples of the Figure 1 -th input, C represents the number of channels of the feature map, H represents the height of the input feature map, W represents the width of the input feature map, and T represents the number of time steps.
[0045] (3) Through the fully connected layer, aggregate into output.
[0046] Perform time aggregation according to the following formula:
[0047] ;
[0048] Among them, represents the aggregated image, is a weight matrix, are all weight values learned by the fully connected layer, represents the transpose of the weight matrix.
[0049] (4) For By rearranging the tensor dimension operation, its function is to change the shape (i.e., dimensions) of the tensor, but without changing the data therein. The output (batch_size, 64, 1, 416, 416) is input into the temporal feature reconstruction module.
[0050] The temporal feature reconstruction module mainly performs "transposed convolution" operations on the input feature map to help the model better learn and represent features in the temporal dimension, enhance the expression ability of temporal information, and optimize the detail recovery ability of features through transposed convolution operations. Its input is: (batch_size, 64, 1, 416, 416) (batch, number of channels, number of time steps, height, width); the output is: (batch_size, 64, 1, 416, 416) (batch, number of channels, number of time steps, height, width); and its working method is:
[0051] The calculation formula for the transposed convolution operation is as follows:
[0052]
[0053] Among them, represents the size of the output feature map; represents the size of the input feature map; represents the offset of the convolution operation, which is determined by the size and stride of the convolution kernel; represents the stride, the moving stride of the convolution kernel on the input feature map each time; 2 represents the influence of padding, because padding is performed on both ends of the input simultaneously, so it needs to be multiplied by 2; represents the padding, the number of zeros added on each dimension (depth, height, width) of the input data; represents the size of the convolution kernel, the size of the convolution kernel on each dimension. represents the output padding, which is used to add additional zero padding in the output feature map. Here, input , substitute it into the above formula for calculation to obtain the output depth and width. So the final output is (batch_size, 64, 1, 416, 416), that is, the reconstructed image.
[0054] Step S102: Perform multi-layer convolution processing, dilated convolution processing, residual processing, attention mechanism processing, dimensionality reduction processing, and dimension transformation processing on the reconstructed image in sequence to obtain the first feature map;
[0055] It should be noted that in this step, the first feature map is obtained by inputting the reconstructed image into the multi-scale spatial extraction network module. The multi-scale spatial extraction network module mainly performs deep feature extraction and optimization on the output feature map of the temporal feature reconstruction module to extract deeper spatial and temporal features; the receptive field is increased and feature learning is optimized through the dilated convolutional kernel residual block; the important features are weighted through the attention mechanism to enhance the model's attention to key features; dimensionality reduction is performed and the optimized feature map is output. Its input is: (batch_size, 64, 1, 416, 416) (batch, number of channels, number of time steps, height, width); the output is: (batch_size, 1, 416, 416, 64) (batch, number of time steps, height, width, number of channels). Its working method is as follows:
[0056] First, the number of channels of the input feature map (batch_size, 64, 1, 416, 416) is increased to (batch_size, 256, 1, 416, 416) through two consecutive basic 3D convolution operations; then the receptive field is expanded through dilated convolution and input into the dilated convolution DilatedConv.3D, and the calculation formula is:
[0057] ;
[0058] Among them, represents the output size, including the depth, width and height of the output, represents the input size, represents the padding, represents the depth of the convolutional kernel, represents the size of the convolutional kernel, represents the stride.
[0059] Exemplarily, calculate the depth dimension:
[0060] ;
[0061] Among them, ;
[0062] Calculate the height and width dimensions:
[0063]
[0064] Among them, ;
[0065] Similarly, the feature size output by the pooling layer and the activation function after the dilated convolution is: (batch_size, 512, 1, 416, 416).
[0066] Then, the result of the dilated convolution is processed through a residual block (consisting of two convolutional layers, with each convolutional layer keeping the number of channels and the spatial dimensions unchanged), and then the output of the residual block is added to the result of the dilated convolution through a residual connection to output a feature map, with the output feature being (batch_size, 512, 1, 416, 416).
[0067] Then, the attention mechanism is processed:
[0068] (1) Average pooling: The feature map is subjected to average pooling operation to generate a weight for each channel, so as to enhance important features and suppress unimportant features. The formula for global pooling is as follows:
[0069]
[0070] Among them, and are the height and width of the input feature map respectively; is the value of the b-th sample, c-th channel, the -th time step, the h-th height position, and the -th width position of the input tensor; after global pooling, a weighted eigenvalue of a feature map (batch_size, channels, 1, 1) is generated. Here, assuming the input b = 1 and the first channel c = 1, the calculation process of average pooling is as follows:
[0071] ;
[0072] Among them, is the weighted eigenvalue of channel 1 in sample 1; 416 and 416 are the height and width of the input feature map; is the value of the 1st sample, 1st channel, 1st time step, 1st height position, and 1st width position of the input tensor.
[0073] (2) Convolution to generate attention weights: First, through a 1x1 convolutional layer, the feature map after global pooling can be mapped to a new tensor, and this mapping will be restricted by an activation function. The convolution operation is as follows:
[0074] ;
[0075] Among them: is the attention weight of the -th sample and the -th channel; represents mapping through a convolutional layer, and a tensor of (batch_size, channels, 1, 1) will be generated; Denote the activation function (Sigmoid), which is used to ensure that the reasonable range of weights outputs values between 0 and 1.
[0076] (3) Expand the attention weights to each spatial position (d, h, w, i.e., depth, height, width) through the broadcasting mechanism.
[0077]
[0078] Among them, are the attention weights expanded through the broadcasting mechanism.
[0079] (4) Apply the broadcasted attention weights:
[0080] ;
[0081] Among them, is the output tensor, and the features at each spatial position are adjusted by the attention weights of the corresponding channels. is the value of the b-th sample, c-th channel, d-th time step, h-th height position, and -th width position of the input feature map.
[0082] Then perform the final convolution to reduce the dimension back to the number of input channels: the Conv.3D convolution kernel is (1, 1, 1), padding = 0, reduce the number of image channels to be consistent with the input image dimension, and output the feature map (batch_size, 64, 1, 416, 416); perform a permute again for dimension transformation; make the feature map transform from (batch_size, 64, 1, 416, 416) to (batch_size, 1, 416, 416, 64); the output here (batch_size, 1, 416, 416, 64) is the final output of the multi-scale spatial extraction network module, that is, the first feature map is obtained.
[0083] Step S103: Perform 3D convolution processing, spatial attention mechanism processing, tensor reshaping, permutation operation, 3D multi-head self-attention mechanism processing, shape recovery, high-dimensional convolution processing, residual connection, dynamic pooling, flattening operation, fully connected operation, tensor expansion, and dimension repetition processing on the reconstructed image in sequence to obtain the second feature map;
[0084] It should be noted that in this step, the reconstructed image is input into the multi-dimensional detail enhancement module to process complex data containing spatio-temporal information. Through multi-dimensional feature extraction, feature fusion, and time series analysis, the accuracy and efficiency of data processing are improved. The input of the multi-dimensional detail enhancement module is: (batch_size, 64, 1, 416, 416) (batch, number of channels, time step, height, width); the output is: (batch_size, 1000, 416, 416, 1) (batch, number of channels, time step, height, width); and its working method is as follows:
[0085] First, 3D convolution is used to process the time series data. Specifically, three consecutive 3D convolutions are used, and the parameters of the three convolutions are all convolution kernel size (3, 3, 3), and padding = (1, 1, 1). The number of channels of (batch_size, 64, 1, 416, 416) is increased to (batch_size, 1024, 1, 416, 416), and then it is input into the spatial attention mechanism. The spatial attention mechanism module dynamically adjusts the position of each weight in the feature map for different spatial positions of the input feature map. Specifically, the spatial attention mechanism module generates a spatial attention map and applies it to the input data, so that the model pays more attention to important spatial positions and ignores unimportant regions. The working method of the spatial attention mechanism module is as follows:
[0086] (1) Convolution layer: The number of channels of the input feature map (batch_size, 1024, 1, 416, 416) is reduced to (batch_size, 1, 1, 416, 416) through three 3D convolutions (convolution kernel is 1, padding is 0), and then non-linearity is introduced through the application of an activation function. Then the sigmoid function compresses the values output by the convolution to the range [0, 1], generates attention weights and applies them.
[0087] (2) Weighted input feature map: The generated spatial attention map is multiplied element-wise with the input feature map to achieve weighting. The features at each spatial position will be amplified or reduced according to their corresponding attention weights. This process can enhance the features in the attention region and suppress the features in unimportant regions.
[0088] (3) The output is the weighted feature map, with the same shape as the input feature map. Through the spatial attention mechanism, the model can dynamically adjust the features at each position according to the attention values of the spatial positions.
[0089] That is: ;
[0090] Here the output shape is: (batch_size, 1024, 1, 416, 416).
[0091] Then reshape (batch_size, 1024, 1, 416, 416) into ; Then perform a permute operation, and then process it through a 3D multi-head attention mechanism. The input and output shapes of the 3D multi-head attention mechanism are both (L, N, E), where L is the sequence length; N is the batch size; E is the embedding dimension. In this embodiment, L = 416 x 416 = 173056; ; E = 1024.
[0092] In addition, it should be noted that the feature map processed by the 3D multi-head attention mechanism generates query, key, and value matrices through three linear layers respectively;
[0093] Divide the query, key, and value matrices into multiple heads, with each head corresponding to an embedding dimension. The embedding dimension of each head is: , represents the total embedding dimension in the 3D multi-head attention mechanism, represents the total number of heads in the 3D multi-head attention mechanism. Specifically, the parameters of the first multi-head attention layer (attn1) are: embedding dimension (embed_dim) = 1024; number of heads (num_heads) = 8; the embedding dimension of each head is 128.
[0094] In addition, generate query (Query), key (Key), and value (Value) matrices for the input image through three linear layers respectively.
[0095]
[0096] Among them, is the input image; is the weight matrix obtained through learning, corresponding to the mappings of query, key, and value respectively; the calculated represents the embedding dimension of each position.
[0097] Specifically, for the input :
[0098]
[0099] Among them, has a shape of (173056, Batch_size, 1024), indicating that the embedding dimension of each position is 1024.
[0100] Then perform head splitting: Split the query, key, and value of the attention mechanism into multiple heads to process information independently, and each head has a lower dimension. Suppose there are h heads, and the embedding dimension of each head is:
[0101]
[0102] Among them, E is the embedding dimension of the input vector; h is the number of heads, which is used in multi-head attention to increase the parallelism of the model and capture information from different aspects; represents the dimension of the feature subset processed by each head, allowing each head to capture different aspects of the features in a smaller dimensional space. For each head i:
[0103]
[0104] Among them, are all the linear transformation weights of each head. Divide Q, K, and V into 8 heads, with the dimension of each head being 128. For each head i:
[0105]
[0106] Then perform scaled dot-product attention: For each head, calculate the attention scores and apply them to the values.
[0107]
[0108] Among them, is the query matrix; is the key matrix; value matrix; is the scaling factor to prevent the dot product from being too large and causing the gradient to vanish. Since there are 173056 positions and a 2x2 matrix is generated for each position, the overall shape of the dot product result is (173056, 2, 2).
[0109] Then perform concatenation and linear transformation: Concatenate the outputs of all heads:
[0110]
[0111] Among them, represents concatenating the output vectors of each head on the feature of the last dimension; assuming there are 8 attention heads, is the output of each head; represents the result after concatenation. Here, the dimension of the output of each head is set to After concatenation, a complete 1024-dimensional vector is obtained, and the shape of the concatenated output is .
[0112] Then perform a linear transformation:
[0113]
[0114] Among them, represents a weight matrix of dimension , and the output after the weight transformation, whose shape is determined by the number of columns of the weight matrix, in order to match the input dimension of the next layer of the model. The calculation method of the parameters of the second multi-head attention layer is similar to that of the first one, and will not be repeated in this embodiment.
[0115] Then, reshape the output processed by the 3D multi-head attention mechanism:
[0116] The first step: permute, convert (173056, Batch_size, 1024) to (Batch_size, 1024, 173056).
[0117] The second step: reshape the result of the first-step permutation, that is, convert (Batch_size, 1024, 173056) to (Batch_size, 1, 1024, 416, 416).
[0118] The third step: permute, convert (Batch_size, 1, 1024, 416, 416) reshaped in the second step to (Batch_size, 1024, 1, 416, 416).
[0119] Then, generate (Batch_size, 2048, 1, 416, 416) by performing a convolution on (Batch_size, 1024, 1, 416, 416), and then generate (Batch_size, 4096, 1, 416, 416) by performing another convolution on (Batch_size, 2048, 1, 416, 416).
[0120] Then perform the residual connection:
[0121]
[0122] Among them, is the result after adding the residual term, is the feature map input to the residual block, is the residual term.
[0123] Take as the input of the main path; is the first residual term; is the second residual term.
[0124] Next, average pooling is performed. The purpose of average pooling is to reduce the spatial dimension of the feature map by averaging within the local neighborhood, which helps to reduce overfitting, ensure the main information of the feature map, and at the same time reduce the consumption of computing resources. Its input is (B, 4096, 1, 416, 416), and the output is (B, 4096, 1, 1, 1).
[0125] Formula for adaptive average pooling layer:
[0126] ;
[0127] Among them, represents the value of the feature map obtained after pooling at batch , channel c, depth , height , width position, represents the value of the original input feature map at batch b, channel c, depth d, height h, width position, represents the average pooling factor, and respectively represent the scaling ratios of the pooling window in the height and width directions, represents the floor operation, i and j respectively represent the row and column indices within the pooling window, , respectively represent the height and width of the output feature map.
[0128] Exemplarily, it is determined that in the height direction of the pooling window: input height , output height , height of the pooling window ; in the width direction: input width , output width , height of the pooling window ; then the entire 416x416 feature map is divided into a large window for pooling:
[0129]
[0130] Depth (D') = 1, so the only depth index is 0; height (H') = 1, so the only height index is 0; width (W') = 1, so the only width index is 0.
[0131] For each batch and each channel , calculate the average value of the elements at all spatial positions . The final output Y retains only one numerical value for each batch and channel, representing the average activation value of the channel over the entire space.
[0132] In addition, the flatten operation converts a multi-dimensional tensor into a two-dimensional tensor, which is usually used to connect the output of a convolutional layer to a fully connected layer. Flattening the pooled [B, 4096, 1, 1, 1] to [B, 4096] removes the extra dimensions.
[0133]
[0134] Among them, represents the flattened output matrix, which is a two-dimensional matrix where each row represents the flattened feature vector of a sample and is input to the fully connected layer of the subsequent network; represents the feature map output from the average pooling layer, which is a multi-dimensional feature tensor and is [B, 4096, 1, 1, 1] here.
[0135] The input to the fully connected layer is: [B, 4096], and a linear transformation is performed. The output is: [B, 1000]. The fully connected layer can be regarded as a linear transformation, expressed as:
[0136]
[0137] Among them, is the input feature map of the fully connected layer, whose dimension is [B, 4096]. Here, B represents the batch, and each sample has 4096 features; W is the weight matrix of the fully connected layer, with a dimension of [4096, 1000]. This matrix converts 4096 input features into 1000 output nodes, and each node corresponds to a part of the final output layer of the network; b is the bias vector, with a dimension of 1000, which is used to add a bias value to each output node to help the model better fit the data; Y is the output feature map of the fully connected layer, with a dimension of [B, 1000]. For each sample in the batch, the fully connected layer outputs a 1000-dimensional vector, and these vectors are usually used for subsequent classification and regression tasks.
[0138] The specific calculation steps of the linear transformation are as follows:
[0139] For each input sample , if , the corresponding Batch_size = 2.
[0140] The first step of matrix multiplication:
[0141]
[0142] Among them, W is the weight matrix of the fully connected layer, is the output feature map corresponding to the i-th sample.
[0143] Specifically, for the i-th sample, the j-th feature of the output is calculated as:
[0144]
[0145] where is the value of the input tensor X at the i-th sample and the k-th feature position, is the value of the weight matrix W at the k-th row and the j-th column, represents the output feature map corresponding to the j-th feature of the i-th sample.
[0146] The second step is to add the bias:
[0147]
[0148] That is:
[0149]
[0150] where is the value of the bias vector at the j-th position.
[0151] Then, expand and repeat the tensor, specifically:
[0152] (1) Expand:
[0153] Input: [B, 1000], expand the dimension.
[0154] Output: [B, 1000, 1, 1, 1].
[0155]
[0156] where corresponds to the batch size B, corresponds to the number of output channels , (because the size of the dimension after expansion is 1), represents the feature map after expanding the input feature map.
[0157] (2) Repeat:
[0158] Input: [B, 1000, 1, 1, 1], repeat.
[0159] Output: [B, 1000, 416, 416, 1].
[0160]
[0161] where means repeating 416 times in the depth dimension, Indicates that the height dimension is repeated 416 times, Indicates that the width dimension remains unchanged because the target dimension size is 1. [B, 1000, 416, 416, 1] will be used as the output feature map of the multi-dimensional detail enhancement module, that is, the second feature map. Indicates the feature map obtained by repeating the input feature map.
[0162] Step S104: Increase the width dimension of the second feature map, and splice the first feature map and the enhanced second feature map in the channel dimension to obtain a spliced feature map, and perform channel reduction and feature fusion on the spliced feature map through a 1x1x1 convolution to obtain a fused feature map, and weight the fused feature map to obtain a super-resolution image.
[0163] It should be noted that in this step, the first feature map and the second feature map are input into the dual-stream feature fusion module. The first function of this module is to integrate the output features from two streams of the multi-scale spatial extraction network module and the multi-dimensional detail enhancement module, effectively combining two different perspectives and feature expressions, and increasing the diversity of features; the second function is to apply the attention mechanism to dynamically adjust the feature responses of each channel through global pooling and fully connected layers, not only compressing the refined features, but also emphasizing the more informative parts. The specific working method is as follows:
[0164] First, increase the width dimension of the output of the multi-dimensional detail enhancement module:
[0165]
[0166] Among them, Indicates the function used to adjust the size of the input feature map; Indicates the input feature map, that is, the feature map input into the multi-dimensional detail enhancement module; Defines the target size of the interpolated size; Determines the alignment strategy of the corner pixels in the interpolation method. When set to True, the corner pixels of the input and output tensors will be aligned, thus keeping the position of the image edge unchanged, which helps to reduce geometric deformation. Is the adjusted feature map.
[0167] Then perform the splicing feature:
[0168] Splice the MESNet and the adjusted SFGNN output in the channel dimension:
[0169]
[0170] Among them, It is an operation of concatenating along dimensions; M represents the feature map output by the multi-scale spatial extraction network module; C is the new feature map after concatenating the feature maps output by the multi-scale spatial extraction network module and the multi-dimensional detail enhancement module; It means assuming values are brought in to obtain the concatenated graph.
[0171] Then perform feature fusion:
[0172] Use 1X1X1 convolution for channel reduction and fusion:
[0173]
[0174] Among them, represents the three-dimensional convolution operation of 1X1X1; represents the fused feature map, with the shape of {batch_size, 512, 416, 416, 64}; B represents the batch of the fused feature map; represents the number of channels of the fused feature map; represents the depth of the fused feature map; represents the height of the fused feature map; represents the width of the fused feature map. For each batch_size, output channel , depth , height , width , the formula for the fused feature is:
[0175]
[0176] Among them, represents the batch of the fused image; represents the number of channels of the fused feature map; is the depth of the fused feature map; is the height of the fused feature map; is the width of the fused feature map; is the original number of channels, that is, the number of channels of the output feature map after the previous concatenation operation; is the additional number of channels; the feature map output by the concatenation operation, that is, the input feature map of the feature fusion operation; is a weight tensor used to transform the input feature from channel c to the new ; is a point convolution, that is, 1x1x1 convolution, used to change the channel dimension without affecting the spatial dimension; is the bias term, providing additional numerical adjustment for the new channel .
[0177] Subsequently, the CBAM attention mechanism is applied, introducing the Convolutional Block Attention Module (CBAM), which combines channel attention and spatial attention to enhance the fused features. Specifically:
[0178] Calculate the channel attention weights through the channel attention module:
[0179]
[0180] Among them, represents the output attention feature map, which has the same shape as the input feature map, but each channel has been weighted and adjusted by the attention mechanism; represents global average pooling, which helps to extract global spatial information, reduce spatial dimension information, and retain information in the channel dimension for calculating channel attention; (assuming r = 16 is the dimensionality reduction ratio) and both represent convolutional weights; The activation function is used to introduce non-linearity and perform non-linear mapping on the result.
[0181] Perform global average pooling:
[0182]
[0183] Among them, represents the output after global average pooling, which is a new tensor containing the global information obtained through the pooling operation. In the channel dimension, the output is a feature vector representing the average feature of each channel; represents the 3D average pooling operation, which is applied to the fused feature map to calculate the average value of one channel from each spatial position, thereby extracting global information.
[0184]
[0185] Among them, is the output feature map after the max pooling operation, representing the maximum value extracted from each channel. The output of each channel is compressed into a scalar (the maximum value); is the 3D max pooling operation, which is applied to the input feature map. This operation applies max pooling to each channel in the spatial dimension (depth, height, width), selects the maximum value in each spatial region. The pooling operation can extract the most significant features while reducing the spatial dimension and retaining the global significant information; represents the output feature map after pooling.
[0186] Add the results of the global average pooling:
[0187]
[0188] Among them, M is the result of the final global pooling obtained by adding the output feature maps after global average pooling and max pooling operations; represents the output after global average pooling, which is a brand-new tensor containing the global information obtained through the pooling operation. In the channel dimension, the output is a feature vector representing the average feature of each channel.
[0189] Perform two-layer convolutional kernel ReLU activation:
[0190]
[0191] Among them, represents the output tensor processed by the convolutional kernel (ReLU), containing the final weighted kernel adjustment result of the fused features; represents the convolutional operation; is the calculation of the convolutional operation plus the bias term; is the convolutional layer weight applied to the pooled tensor for processing the input spatial information; M is the feature map after the previous pooling operation, containing the global information after pooling; is the bias term, which is added to the feature map together with the convolutional operation for further adjusting the output; is the output tensor shape.
[0192] Perform the first weighted operation:
[0193]
[0194] Among them, represents the weighted feature map with the shape of [Batch_size, Channels, Depth, Height, Width], meaning that the spatial dimensions (depth, height, width) and the number of channels of the feature map have been adjusted through the channel attention mechanism.
[0195] Calculate the spatial attention weight:
[0196]
[0197] Among them, represents concatenating and on the channel dimension, indicating that the concatenation operation occurs on the channel dimension.
[0198] Perform spatial pooling:
[0199]
[0200] Among them, represents the spatially averaged pooling result of the output, which represents the average value of all channels in the feature map; is the normalization factor; is to accumulate each channel c in the weighted feature map, which is to sum all channel values at a given spatial position ; represents performing the above operations for all batches b and all spatial positions .
[0201]
[0202] Among them, represents the spatially maximum pooling result of the output, which represents the maximum value of all channels in the feature map; represents selecting the maximum value for each channel c in the weighted feature map, which is to extract the maximum value of all channel values at a given spatial position .
[0203] The concatenated pooling result is:
[0204]
[0205] Among them, represents the result of concatenating the spatial average pooling and the spatial maximum pooling.
[0206] Perform convolution and Sigmoid activation:
[0207]
[0208] Among them, is the convolutional kernel weight, which is used to extract useful spatial features from the concatenated feature map to generate attention weights; represents the result of concatenating the spatial average pooling and the spatial maximum pooling; is the bias vector, which is used for adjustment after convolution operation to ensure an appropriate range of offset of the output, represents the output spatial attention map;
[0209] Perform the second weighting operation:
[0210]
[0211] Among them, represents the output weighted eigenvalue map.
[0212] Perform the activation function (finally apply the ReLU activation function to introduce non-linearity):
[0213]
[0214] Wherein: is the final output feature map, that is, the final output feature map (super-resolution image) of the entire network.
[0215] In summary, since the traditional super-resolution reconstruction network can have a good reconstruction effect on a single blurred image, but because the objects it faces are mostly the restoration of old pictures or single pictures generated by low-resolution visible light devices, it does not consider that the shooting device will generate motion blur due to movement in actual applications, making it difficult to promote the traditional super-resolution reconstruction network in the actual application of the power field. Therefore, to solve the above problems, the embodiments of the present invention propose a combination of spatio-temporal feature extraction and dual attention mechanisms (time + space) so that the model can aggregate information from both time and space dimensions when processing time-series data, thereby enhancing the expression ability of the overall features, and being able to reconstruct and restore the motion-blurred images during the shooting process of the drone, providing a reference method for insulator image reconstruction.
[0216] As Figure 2 shown, an embodiment of the present invention also proposes a super-resolution reconstruction system for motion-blurred images of transmission line inspections, and the system includes:
[0217] A blurred image acquisition module 10, configured to acquire a blurred image of an insulator, perform temporal processing on the blurred image to obtain a temporal image, and sequentially perform time attention, time aggregation, and time reconstruction operations on the temporal image to obtain a reconstructed image;
[0218] A first convolution processing module 20, configured to sequentially perform multi-layer convolution processing, dilated convolution processing, residual processing, attention mechanism processing, dimensionality reduction processing, and dimensional transformation processing on the reconstructed image to obtain a first feature map;
[0219] A second convolution processing module 30, configured to sequentially perform 3D convolution processing, spatial attention mechanism processing, tensor reshaping, permutation operation, 3D multi-head self-attention mechanism processing, shape restoration, high-dimensional convolution processing, residual connection, dynamic pooling, flattening operation, fully connected operation, tensor expansion, and dimensionality repetition processing on the reconstructed image to obtain a second feature map;
[0220] A feature fusion module 40, configured to increase the width dimension of the second feature map, splice the first feature map and the enhanced second feature map in the channel dimension to obtain a spliced feature map, perform channel dimensionality reduction and feature fusion on the spliced feature map through a 1x1x1 convolution to obtain a fused feature map, and weight the fused feature map to obtain a super-resolution image.
[0221] On the other hand, the present invention also provides a storage medium, on which one or more programs are stored, and when the program is executed by a processor, the above-mentioned super-resolution reconstruction method for motion-blurred images of transmission line inspections is implemented.
[0222] On the other hand, the present invention also provides an electronic device, including a memory and a processor, where the memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the above-mentioned super-resolution reconstruction method for motion-blurred images of transmission line inspections.
[0223] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0224] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0225] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0226] Although the embodiments of the present invention have been described in detail above, it will be obvious to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are all within the scope and spirit of the present invention as described in the claims. Moreover, the present invention as described herein can have other embodiments and can be implemented or realized in various ways.
Claims
1. A method for super-resolution reconstruction of motion blurred images of power transmission line inspection, characterized in that: The method comprises: Acquire a blurred image of the insulator, perform time series processing on the blurred image to obtain a time series image, and perform time attention, time aggregation and time reconstruction operations on the time series image in sequence to obtain a reconstructed image; The reconstructed image is sequentially subjected to multi-layer convolution processing, hole convolution processing, residual processing, attention mechanism processing, dimensionality reduction processing and dimensionality conversion processing to obtain a first feature map; The channels of the reconstructed image are enlarged by two consecutive 3D convolution operations, and the enlarged reconstructed image is subjected to dilated convolution processing according to the following formula: Among them, O represents the output size, including the depth, width and height of the output, I represents the input size, P represents padding, D1 represents the depth of the convolution kernel, K represents the size of the convolution kernel, and S1 represents the step size; The result of the dilated convolution is processed by a residual block, and then the output of the residual block is added to the result of the dilated convolution through a residual connection to output a feature map. The residual block consists of two convolutional layers, and each convolutional layer keeps the number of channels and spatial dimensions unchanged. Perform average pooling on the feature map: Among them, y b,c is the weighted feature value of channel c on sample b, H and W are the height and width of the input feature map, respectively, x b,c,1,h,w is the value of the input tensor at the bth sample, cth channel, 1st time step, hth height position, and wth width position; Map the average pooled feature map to a new tensor: A b,c =σ(Conv1×1(y b,c )); Among them, A b,c is the attention weight of the bth sample and the cth channel, Conv1×1 means mapping through a 1×1 convolution layer, and σ represents the activation function, which is used to make the weight range output between 0 and 1; The attention weights are extended to each spatial position via a broadcasting mechanism: A b,c,d,h,w =A b,c ; Among them, A b,c,d,h,w is the attention weight after expansion through the broadcast mechanism; Attention weights after applying broadcasting: AND b,c,d,h,w =X b,c,d,h,w ×A b,c,d,h,w ; Among them, Y b,c,d,h,w is the output tensor, where the features of each spatial position are adjusted by the attention weight of the corresponding channel, X b,c,d,h,w is the value of the bth sample, cth channel, dth time step, hth height position, and wth width position of the input feature map; Perform dimension reduction and dimension conversion on the output tensor to obtain the first feature map; The reconstructed image is sequentially subjected to 3D convolution processing, spatial attention mechanism processing, tensor reshaping, permutation operation, 3D multi-head self-attention mechanism processing, shape recovery, high-dimensional convolution processing, residual connection, dynamic pooling, flattening operation, full connection operation, tensor expansion and dimension repetition processing to obtain a second feature map; The width dimension of the second feature map is increased, and the first feature map and the increased second feature map are spliced in the channel dimension to obtain a spliced feature map, and the spliced feature map is subjected to channel dimension reduction and feature fusion through a 1x1x1 convolution to obtain a fused feature map, and the fused feature map is weighted to obtain a super-resolution image; Get the concatenated feature map according to the following formula: Among them, concat(·) represents the concatenation operation along the dimension, M represents the first feature map of the output, and C represents the concatenated feature map. It means that we assume that we bring in a value to get the concatenated graph; Feature fusion is performed according to the following formula: Among them, F fused represents the fused feature map, b′ represents the fused image batch, c′ represents the number of fused feature map channels, d′ represents the depth of the fused feature map, h′ represents the height of the fused feature map, w′ represents the width of the fused feature map, and C in Indicates the original number of channels, C s represents the number of additional channels, C[b,c,d,h,w] represents the feature map output by the concatenation operation, Represents a weight tensor used to transform the input feature from channel c to the new c′, 1, 1, 1 represents a 1x1x1 convolution, represents the bias term, C f Indicates the number of channels of the fused feature map; The weighting is based on the following formula: F final =Re LU(F cbam ) F cbam =F ca ⊙S2 F ca =F fused ⊙A2; Among them, F cbam Represents the weighted eigenvalue map of the output, F ca represents the first weighted operation result obtained in the global average pooling operation, ⊙ represents element-by-element multiplication, S2 represents the output spatial attention map, and F fused represents the feature map after the fusion operation, A2 represents the attention feature map, F final represents the output super-resolution image, and ReLU represents the activation function; Improve the width dimension of the output of the Multi-Dimensional Detail Enhancement module: Among them, interpolate(·) represents the function used to resize the input feature map; S3 represents the input feature map, that is, the feature map input by the multi-dimensional detail enhancement module; size = (416, 416, 64) defines the target size after interpolation; align_corners = True determines the alignment strategy of corner pixels in the interpolation method. When set to True, the corner pixels of the input and output tensors will be aligned.
2. The method for super-resolution reconstruction of motion blurred images of power transmission line inspection according to claim 1, characterized in that: The steps of acquiring a blurred image of the insulator, performing time series processing on the blurred image to obtain a time series image, and sequentially performing time attention, time aggregation and time reconstruction operations on the time series image to obtain a reconstructed image include: Input blurred image, and then go through permute function dimension transformation, Conv3D convolution, and permute function dimension transformation to get time series image. Input the time series image into the Conv3D convolution layer, the pooling layer and the activation function layer to obtain a time feature map; The temporal feature graph is sequentially subjected to dimension rearrangement and spatial dimension expansion by a permute function to obtain a spatial dimension graph, and the query, key, and value of the spatial dimension graph are calculated to weight the query, key, and value of the spatial dimension graph to output a weighted feature image; The weighted feature image is rearranged in dimension by a permute function and expanded in spatial dimension to obtain a feature vector, and the feature vector of each spatial position is extracted to form a two-dimensional matrix; Performing temporal aggregation according to the two-dimensional matrix to obtain an aggregated image, and rearranging the tensor dimension of the aggregated image; The rearranged aggregated image is subjected to a deconvolution operation to obtain a reconstructed image.
3. The method for super-resolution reconstruction of motion blurred images of power transmission line inspection according to claim 2 is characterized in that: The step of rearranging the weighted feature image through the permute function dimension and expanding the spatial dimension to obtain a feature vector, and extracting the feature vector of each spatial position to form a two-dimensional matrix includes: The two-dimensional matrix is formed according to the following formula: Where X′∈R (B×C×H×W,T) represents the flattened matrix, X (H,W,T) represents the feature value of the Hth row and the Wth column at time step T, B represents the number of samples input to the feature map at one time, C represents the number of channels of the feature map, H represents the height of the input feature map, W represents the width of the input feature map, and T represents the number of time steps; In the two-dimensional matrix, each row represents a feature vector of a spatial position at different time steps, and each column corresponds to the feature vectors of all spatial positions at the same time step; The step of performing time aggregation according to the two-dimensional matrix to obtain an aggregated image comprises: Time aggregation is performed according to the following formula: Where aggregated represents the aggregated image, E0=[E1 E2 … E T ] is a weight matrix, E1, E2, …, E T are all weight values learned by the fully connected layer, E0 T represents the transpose of the weight matrix; The step of performing a deconvolution operation on the rearranged aggregated image to obtain a reconstructed image comprises: The deconvolution operation is performed according to the following formula: Outputsize=(Inputsize-1)×Stride-2×padding+kernel_size+outpadding Among them, Outputsize represents the size of the output feature map, Inputsize represents the size of the input feature map, stride represents the step size, the step size of the convolution kernel on the input feature map each time, padding represents the padding, and the number of zeros added to each dimension of the input data, kernel_size represents the size of the convolution kernel, the size of the convolution kernel in each dimension, outpadding represents the output padding, which is used to add additional zero padding to the output feature map.
4. The method for super-resolution reconstruction of motion blurred images of power transmission line inspection according to claim 3 is characterized in that: The step of sequentially performing 3D convolution processing, spatial attention mechanism processing, tensor reshaping, permutation operation, 3D multi-head self-attention mechanism processing, shape recovery, high-dimensional convolution processing, residual connection, dynamic pooling, flattening operation, full connection operation, tensor expansion and dimension repetition processing on the reconstructed image to obtain the second feature map includes: The number of channels of the reconstructed image is increased by using 3D convolution, and then input into the spatial attention mechanism, the number of channels of the input feature map is reduced by three 3D convolutions, and nonlinearity is introduced by applying an activation function, the generated spatial attention map is element-wise multiplied with the input feature map, and a weighted feature map is output; The weighted feature map is passed through three linear layers to generate query, key, and value matrices respectively; Divide the query, key, and value matrices into multiple heads, each head corresponds to an embedding dimension, and the embedding dimension of each head is: E k represents the total embedding dimension in the 3D multi-head attention mechanism, N u Represents the total number of heads in the 3D multi-head attention mechanism; For each head, the attention score is calculated and applied to the value: in, is the query matrix; is the key matrix; Value matrix, L is the sequence length; N is the batch size; The outputs of all heads are stitched together, and the shape of the stitched image is restored; The residual connection is performed according to the following formula: x′=x+res_x; Among them, x′ is the result after adding the residual term, x is the feature input to the residual module, and res_x is the residual term; Dynamic pooling is performed according to the following formula: Among them, Y[b,c,d′,h′,w′] represents the value of the feature map obtained after the pooling operation at the position of batch b, channel c, depth d′, height h′, and width w′, and X[b,c,d,h,w] represents the value of the original input feature map at the position of batch b, channel c, depth d, height h, and width w. represents the average pooling factor, k H and k W Respectively represent the scaling ratio of the pooling window in height and width, represents the floor operation, i and j represent the row and column indices in the pooling window, respectively, and H′ and W′ represent the height and width of the output feature map, respectively.
5. A transmission line inspection motion blurred image super-resolution reconstruction system, characterized in that: The system comprises: A fuzzy image acquisition module is used to acquire a fuzzy image of the insulator, perform time series processing on the fuzzy image to obtain a time series image, and perform time attention, time aggregation and time reconstruction operations on the time series image in sequence to obtain a reconstructed image; A first convolution processing module, used for sequentially performing multi-layer convolution processing, hole convolution processing, residual processing, attention mechanism processing, dimensionality reduction processing and dimensionality conversion processing on the reconstructed image to obtain a first feature map; The channels of the reconstructed image are enlarged by two consecutive 3D convolution operations, and the enlarged reconstructed image is subjected to dilated convolution processing according to the following formula: Among them, O represents the output size, including the depth, width and height of the output, I represents the input size, P represents padding, D1 represents the depth of the convolution kernel, K represents the size of the convolution kernel, and S1 represents the step size; The result of the dilated convolution is processed by a residual block, and then the output of the residual block is added to the result of the dilated convolution through a residual connection to output a feature map. The residual block consists of two convolutional layers, and each convolutional layer keeps the number of channels and spatial dimensions unchanged. Perform average pooling on the feature map: Among them, y b,c is the weighted feature value of channel c on sample b, H and W are the height and width of the input feature map, respectively, x b,c,1,h,w is the value of the input tensor at the bth sample, cth channel, 1st time step, hth height position, and wth width position; Map the average pooled feature map to a new tensor: A b,c =σ(Conv1×1(y b,c )); Among them, A b,c is the attention weight of the bth sample and the cth channel, Conv1×1 means mapping through a 1×1 convolution layer, and σ represents the activation function, which is used to make the weight range output between 0 and 1; The attention weights are extended to each spatial position via a broadcasting mechanism: A b,c,d,h,w =A b,c ; Among them, A b,c,d,h,w is the attention weight after expansion through the broadcast mechanism; Attention weights after applying broadcasting: AND b,c,d,h,w =X b,c,d,h,w ×A b,c,d,h,w ; Among them, Y b,c,d,h,w is the output tensor, where the features of each spatial position are adjusted by the attention weight of the corresponding channel, X b,c,d,h,w is the value of the bth sample, cth channel, dth time step, hth height position, and wth width position of the input feature map; Perform dimension reduction and dimension conversion on the output tensor to obtain the first feature map; A second convolution processing module is used to sequentially perform 3D convolution processing, spatial attention mechanism processing, tensor reshaping, permutation operation, 3D multi-head self-attention mechanism processing, shape recovery, high-dimensional convolution processing, residual connection, dynamic pooling, flattening operation, full connection operation, tensor expansion and dimension repetition processing on the reconstructed image to obtain a second feature map; A feature fusion module is used to increase the width dimension of the second feature map, and to splice the first feature map and the improved second feature map in the channel dimension to obtain a spliced feature map, and to perform channel dimension reduction and feature fusion on the spliced feature map through a 1x1x1 convolution to obtain a fused feature map, and to weight the fused feature map to obtain a super-resolution image; Get the concatenated feature map according to the following formula: Among them, concat(·) represents the concatenation operation along the dimension, M represents the first feature map of the output, and C represents the concatenated feature map. It means that we assume that we bring in a value to get the concatenated graph; Feature fusion is performed according to the following formula: Among them, F fused represents the fused feature map, b′ represents the fused image batch, c′ represents the number of fused feature map channels, d′ represents the depth of the fused feature map, h′ represents the height of the fused feature map, w′ represents the width of the fused feature map, and C in Indicates the original number of channels, C s represents the number of additional channels, C[b, c, d, h, w] represents the feature map output by the concatenation operation, Represents a weight tensor used to transform the input feature from channel c to the new c′, 1, 1, 1 represents a 1x1x1 convolution, represents the bias term, C f Indicates the number of channels of the fused feature map; The weighting is based on the following formula: F final =ReLU(F cbam ) F cbam =F ca ⊙S2 F ca =F fused ⊙A2; Among them, F cbam Represents the weighted eigenvalue map of the output, F ca represents the first weighted operation result obtained in the global average pooling operation, ⊙ represents element-by-element multiplication, S2 represents the output spatial attention map, and F fused represents the feature map after the fusion operation, A2 represents the attention feature map, F final represents the output super-resolution image, and ReLU represents the activation function; Improve the width dimension of the output of the Multi-Dimensional Detail Enhancement module: Among them, interpolate(·) represents the function used to resize the input feature map; S3 represents the input feature map, that is, the feature map input by the multi-dimensional detail enhancement module; size = (416, 416, 64) defines the target size after interpolation; align_corners = True determines the alignment strategy of corner pixels in the interpolation method. When set to True, the corner pixels of the input and output tensors will be aligned.
6. A storage medium, characterized in that: The storage medium stores one or more programs, which, when executed by the processor, implement the method for super-resolution reconstruction of motion blurred images of power transmission line inspection as described in any one of claims 1 to 4.
7. An electronic device, comprising a memory and a processor, wherein: The memory is used to store computer programs; When the processor is used to execute the computer program stored in the memory, it implements the method for super-resolution reconstruction of motion blurred images of power transmission line inspection as described in any one of claims 1-4.
Citation Information
Patent Citations
Tongue picture semantic segmentation method and device, equipment and medium
CN117576405A