Pipe wall fisheye distortion correction method based on multi-scale discriminator and center self-attention

Through the multi-scale discriminator and the central self-attention method of fisheye distortion correction method, the correction problem of double distortion of fisheye lenses and pipe inner walls in the monitoring of the inner wall of the pipe is solved, and a high-precision image correction effect is achieved.

CN120339141APending Publication Date: 2025-07-18CHINA WEST NORMAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510267659.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to correct the double distortion of fisheye lens and pipe inner wall simultaneously in pipeline inner wall monitoring, and the existing deep learning methods are difficult to meet the requirements in pipeline exploration with high accuracy requirements.

Method used

The fish-eye distortion correction method of tube wall based on multi-scale discriminator and central self-attention is adopted. Through a distortion network composed of generator and discriminator, combined with the flow estimation module and distortion correction module, the central self-attention module and multi-scale convolution are introduced, and multiple loss functions are used for training to build rich image features.

Benefits of technology

Improve the accuracy and quality of image correction, generate higher quality correction images, and adapt to the high-precision requirements of pipeline inner wall monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339141A_ABST
    Figure CN120339141A_ABST
Patent Text Reader

Abstract

The invention discloses a pipe wall fisheye distortion correction method based on a multi-scale discriminator and center self-attention, the method is used for pipe wall fisheye distortion network correction, a distortion network is composed of a generator and a discriminator, the generator comprises a flow estimation module and a distortion correction module, and the flow estimation module and the distortion correction module both adopt UNet structures. After down-sampling and up-sampling are carried out on an input image for multiple times, feature reconstruction is carried out to obtain a corrected image, and a discriminator is used for discriminating; the discriminator adopts multi-scale convolution to extract features, and introduces a center self-attention module to highlight center area features. Aiming at the characteristic that fisheye image information is concentrated in a central area, a central self-attention module is introduced into a generator and a discriminator, so that the quality and the structural similarity of a generated image and the performance of network extraction features are improved; a multi-scale discriminator network is constructed, rich image features are obtained, the performance of the discriminator network is improved, and the whole network is guided to generate images with higher quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fisheye distortion correction for the inner wall of pipelines, and particularly relates to a method for correcting fisheye distortion of the pipeline wall based on a multi-scale discriminator and central self-attention. Background Art

[0002] A fisheye lens is an optical imaging system with a super-large field of view and a large aperture, which is commonly used in fields such as intelligent driving, surveillance and security, underwater photography, aerospace, etc. While providing a wide field of view, it inevitably introduces significant optical distortion, which makes the image prone to distortion and distortion at the edge part and difficult to identify; therefore, correcting the distortion of the image taken by the fisheye lens is a key step before higher-level image operations such as target tracking, motion prediction, and image segmentation. Due to the characteristics of short focal length and large viewing angle of the fisheye lens, it is applied to the special occasion of pipeline inner wall monitoring. The images collected for pipeline inner wall monitoring have not only the distortion of the pipeline inner wall caused by the fisheye lens, but also the distortion caused by the concave surface of the pipeline inner wall. Therefore, directly using the fisheye lens correction method cannot obtain the correct corrected image, and a separate correction method needs to be studied.

[0003] Traditional fisheye lens image correction methods, such as the checkerboard calibration method, rely on extracting feature points (such as corner points, straight lines, specific points) from the image to estimate the distortion parameters of the lens, and these parameters are then used to correct the image to reduce or eliminate the distortion. However, the effects of these traditional methods are limited by the detectability and quantity of the feature points, which may lead to inconsistency in the correction results. With the progress of deep learning technology, correction methods based on neural networks have become an effective alternative. These methods train the network to learn the mapping relationship between the distorted image and the undistorted image, so as to achieve correction. The deep learning model can handle complex distortion patterns, and through training on a large amount of data sets, it can learn a generalized correction function. Although the deep learning method has advantages in dealing with complex distortions, it requires a large amount of labeled data during the training process, and it is difficult to obtain real distorted image annotation data.

[0004] Fisheye image correction based on deep learning is usually divided into methods based on convolutional neural networks (CNNs) and methods based on generative adversarial networks (GANs). The method based on convolutional neural networks extracts features through multiple convolutional layers and pooling layers, etc., and learns the mapping model between the distorted image and the undistorted image through training on a large amount of fisheye distortion image data sets; the generative adversarial network consists of a generator and a discriminator, and is trained in an adversarial manner. The generator generates an undistorted image, and the discriminator compares the image generated by the generator with the undistorted image to carry out an adversarial process, so as to improve the quality of the image generated by the generator and make it closer to the real undistorted image.

[0005] Existing fish-eye correction methods usually focus on their versatility. To adapt to various environments and lens parameters, they may sacrifice some precision in certain cases. However, during pipeline exploration, the fish-eye lens parameters are known, and a high level of precision is required. Existing methods are difficult to meet the precision requirements. At the same time, traditional methods are difficult to correct both the distortion of the pipeline inner wall and the fish-eye lens distortion, and there is no corresponding research on the correction of fish-eye images on the pipe wall in existing deep learning methods. Summary of the Invention

[0006] To solve the technical problems existing in the prior art, the present invention proposes a method for correcting fish-eye distortion on the pipe wall based on a multi-scale discriminator and central self-attention.

[0007] To solve the above technical problems, the present invention is implemented as follows:

[0008] A method for correcting fish-eye distortion on the pipe wall based on a multi-scale discriminator and central self-attention is used for correcting the fish-eye distortion network. The distortion network consists of two parts: a generator and a discriminator. The generator includes a flow estimation module and a distortion correction module. Both the flow estimation module and the distortion correction module adopt the UNet structure. After the input image is downsampled and upsampled multiple times, the features are reconstructed to obtain the corrected image, and the discriminator is used for discrimination. The discriminator uses multi-scale convolution to extract features and introduces a central self-attention module to highlight the features in the central region.

[0009] Furthermore, the image input to the flow estimation module has a size of 256×256. In the encoder of the flow estimation module, a convolutional kernel with a size of 4×4, a stride of 2, and a padding of 1 is used for downsampling. In the decoder, a multi-level upsampling network is adopted. The resolution of the feature map is gradually enlarged through convolutional modules, residual blocks, and transposed convolutional modules. The downsampled feature map is skip-connected with the upsampled feature map of the corresponding size to generate a series of appearance flows. Then, a 3×3 convolution is performed on the output appearance flows to obtain five pre-corrected feature maps of dual channels with sizes of 128, 64, 32, 16, and 8 respectively, which are used for the decoder part of the distortion correction module.

[0010] The input image of the distortion correction module has a size of 256×256. In the encoder of the distortion correction module, a convolutional kernel with a size of 3×3, a stride of 2, and a padding of 1 is used for downsampling to obtain distorted feature maps with sizes of 128, 64, 32, 16, and 8 respectively. After the third downsampling, a central self-attention module is added to improve the feature extraction ability. The downsampled distorted feature maps of the distortion correction module are fused with the corresponding-sized dual-channel appearance flow pre-corrected features output by the flow estimation module, and after fusion, they are skip-connected with the corresponding-sized upsampled feature maps. In the decoder of the distortion correction module, a convolution with a kernel size of 3, a stride of 1, and a padding of 1 and the interpolate function provided by pytorch are used for bilinear interpolation upsampling to gradually reconstruct the corrected image.

[0011] Further, the discriminator is a discriminator network using multi-scale convolution. It extracts image features by convolving the input image with a size of 256×256. In the backbone feature network of the discriminator, a convolutional kernel with a size of 5×5 is used, and two auxiliary feature extraction networks are introduced in the discriminator, using convolutional kernels with sizes of 3×3 to extract more detailed features of the original image and 7×7 to extract local texture features of the original image. The features at the three scales are fused with each other, and a central self-attention module is introduced after the third layer of convolution to improve the performance of the discriminator network.

[0012] Further, the central self-attention module performs V_conv on the input feature map to obtain a value matrix, multiplies the input feature map with a Gaussian weight matrix to obtain a Gaussian feature, then uses Q_conv and K_conv to convolve the Gaussian feature respectively to obtain a query matrix and a key matrix, and further obtains the attention. The expression of the central self-attention module is as follows:

[0013]

[0014] Among them, Q_conv represents the convolution of the query matrix, V_conv represents the convolution of the value matrix, K_conv represents the convolution of the key matrix, GCW represents the Gaussian weight matrix, and d k represents the vector dimension in the key matrix;

[0015] The calculation of the Gaussian weight matrix uses the following formula:

[0016]

[0017] Among them, the standard deviation parameter σ of the Gaussian matrix is 1.5.

[0018] Preferably, in order to improve the image quality, the present application also adopts a total loss function training strategy including multiple loss functions. The total loss function is obtained from a reconstruction loss, an adversarial loss, a multi-layer loss, and an enhancement loss. The total loss function is expressed as follows:

[0019] L = λ l1 L1 + λ adv L adv + λ m L m + λ e L e

[0020] Among them, λ l1 , λ adv , λ m and λ e all represent hyperparameters, L1 represents the reconstruction loss, L adv represents the adversarial loss, L m represents the multi-layer loss, L e represents the enhancement loss;

[0021] Among them, the expression of the adversarial loss L adv is as follows:

[0022] L adv = min G max D (E[logD(Img gt )] + E[log(1 - D(G(Img out )))])

[0023] Among them, G represents the generator, D represents the discriminator, E represents the expectation, Img gt represents the true value image, Img out represents the corrected image generated by the generator;

[0024] The expression of the multi-layer loss L m is as follows:

[0025]

[0026] Among them, N represents the number of convolutional layers, S represents the downsampling operation, and the function S(x, n) represents downsampling the input x to times the original size, and respectively represent the original feature on the decoder of the distortion correction module and the feature after correction by the feature correction layer. cat(·) represents feature concatenation, and Conv represents a 3×3 convolution that decodes the output feature into a 3-channel RGB image;

[0027] The expression of the enhancement loss L e is as follows:

[0028] L e = L c + λ s L s

[0029] Among them, L c represents the content loss, L s represents the style loss, and λ s represents the style loss weight. The expression of L c is as follows:

[0030]

[0031] Among them, and respectively represent the feature maps obtained by convolving the corrected image and the ground-truth image using the j-th layer in the pre-trained VGG19 network. C j H j W j represent the size of the feature map of the j-th layer;

[0032] L s The expression of is as follows:

[0033]

[0034] represents the square Frobenius norm of the Gram matrix, which is a matrix with a shape of C j × C j The expression of its elements is as follows:

[0035]

[0036] Among them, c and c' represent the specified channels, and are the weights at the specified channels when the size is h*w.

[0037] Compared with the prior art, the beneficial effects of the present invention are:

[0038] This application constructs a double-distorted image dataset of pipe wall and fish-eye. In view of the characteristic that fish-eye image information is concentrated in the central area, a central self-attention module is introduced into the generator and the discriminator to improve the quality and structural similarity of the generated image and the performance of the network in extracting features; a multi-scale discriminator network is constructed to obtain rich image features to improve the performance of the discriminator network, guide the overall network to generate higher-quality images, and a large number of comparative experiments are carried out on the synthesized double-distorted dataset of pipe wall and fish-eye to verify the excellent performance of the method in this paper. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 Schematic diagram of the tube wall fish-eye distortion correction network structure of the present invention;

[0040] Figure 2 Schematic diagram of the multi-scale discriminator network structure of the present invention;

[0041] Figure 3 Schematic diagram of the central self-attention structure of the present invention;

[0042] Figure 4 Schematic diagram of the design of various multi-scale discriminator network structures of the present invention. Detailed implementation manners

[0043] The following further elaborates on the detailed implementation manners of the present invention in conjunction with the accompanying drawings and specific embodiments.

[0044] As Figure 1 shown, a tube wall fish-eye distortion correction method based on a multi-scale discriminator and central self-attention is used for tube wall fish-eye distortion network correction. The distortion network consists of two parts: a generator and a discriminator. The generator includes a flow estimation module and a distortion correction module. Both the flow estimation module and the distortion correction module adopt the UNet structure. After multiple downsamplings and upsamplings of the input image, feature reconstruction is performed to obtain the corrected image, and the discriminator is used for discrimination; the discriminator uses multi-scale convolution to extract features and introduces a central self-attention module to highlight the features of the central region.

[0045] The image input to the flow estimation module has a size of 256×256. In the encoder of the flow estimation module, a convolutional kernel with a size of 4×4, a stride of 2, and a padding of 1 is used for downsampling. In the decoder, a multi-level upsampling network is adopted. The resolution of the feature map is gradually enlarged through a convolutional module, a residual block, and a transposed convolutional module. The downsampled feature map is jump-connected with the upsampled feature map of its corresponding size to generate a series of appearance flows. Then, a 3×3 convolution is performed on the output appearance flows to obtain five two-channel appearance flow pre-correction feature maps with sizes of 128, 64, 32, 16, and 8 respectively, which are used for the decoder part of the distortion correction module;

[0046] The input image of the distortion correction module has a size of 256×256. A convolutional kernel with a large stride is used for downsampling convolution to reduce the size of the feature map. In the encoder of the distortion correction module, a convolutional kernel with a size of 3×3, a stride of 2, and a padding of 1 is used for downsampling to obtain distorted feature maps with sizes of 128, 64, 32, 16, and 8 respectively. After the third downsampling, a central self-attention module is added to improve the feature extraction ability. The downsampled distorted feature maps of the distortion correction module are fused with the corresponding-sized dual-channel appearance flow pre-correction features output by the flow estimation module. After fusion, they are skip-connected with the corresponding-sized upsampled feature maps. In the decoder of the distortion correction module, a convolution with a kernel size of 3, a stride of 1, and a padding of 1 and the interpolate function provided by PyTorch are used for bilinear interpolation upsampling to gradually reconstruct the corrected image. Fusing the appearance flow pre-correction features of the flow estimation module enriches the correction features, thus generating a better corrected image.

[0047] As Figure 2 shown, in the generative adversarial network, moderately strengthening the discriminator can improve the overall performance and training stability of the network to a certain extent. Moderately strengthening the discriminator enriches the features extracted by the discriminator and improves the quality and diversity of the generated samples. The discriminator is a discriminator network using multi-scale convolution. It extracts image features by convolving an input image with a size of 256×256. In the main feature network of the discriminator, a convolutional kernel with a size of 5×5 is used, and two auxiliary feature extraction networks are introduced into the discriminator. Convolutional kernels with sizes of 3×3 and 7×7 are used to extract more detailed features of the original image and local texture features of the original image respectively. The features at the three scales are fused with each other, and a central self-attention module is introduced after the third layer of convolution to improve the performance of the discriminator network.

[0048] Due to the particularity of fisheye lens imaging, the information contained in the fisheye lens distorted image gradually decreases from the center to the edge, and the distortion becomes more and more serious, which is similar to the Gaussian matrix distribution. Therefore, the self-attention is improved to make it more focused on extracting the features of the central region. A two-dimensional Gaussian matrix that approximately reflects the features of the fisheye lens distorted image is constructed. The central self-attention module performs V_conv on the input feature map to obtain a value matrix, multiplies the input feature map by the Gaussian weight matrix to obtain Gaussian features, then uses Q_conv and K_conv to convolve the Gaussian features respectively to obtain a query matrix and a key matrix, and further obtains the attention; as Figure 3 shown; the expression of the central self-attention module is as follows:

[0049]

[0050] Among them, Q_conv represents the convolution of the query matrix, V_conv represents the convolution of the value matrix, K_conv represents the convolution of the key matrix, GCW represents the Gaussian weight matrix, and d k represents the vector dimension in the key matrix;

[0051] The calculation of the Gaussian weight matrix adopts the following formula:

[0052]

[0053] Among them, the standard deviation parameter σ of the Gaussian matrix is 1.5.

[0054] To improve the image quality, the present application also adopts a training strategy of a total loss function including multiple loss functions. The total loss function is obtained through a reconstruction loss, an adversarial loss, a multi-layer loss, and an enhancement loss. The total loss function is expressed as follows:

[0055] L = λ l1 L1 + λ adv L adv + λ m L m + λ e L e

[0056] Among them, λ l1 , λ adv , λ m and λ e all represent hyperparameters, L1 represents the reconstruction loss, L adv represents the adversarial loss, L m represents the multi-layer loss, L e represents the enhancement loss; The multiple loss constraint training process adopted by the present application, where the L1 loss performs pixel-level feature constraint on the training, the adversarial loss improves the local texture quality of the image, the enhancement loss is obtained by weighting the style loss and the content loss, and further improves the image texture quality; The multi-layer loss can supervise the feature maps of each layer on the decoder during training and better guide the training

[0057] Among them, the expression of the adversarial loss L adv is as follows:

[0058] L adv = min G max D (E[logD(Img gt )] + E[log(1 - D(G(Img out )))]))

[0059] Among them, G represents the generator, D represents the discriminator, E represents the expectation, Img gt represents the true value image, Img outDenote the corrected image generated by the generator;

[0060] During the generation process, the distortion of each layer of features should be minimized. A multi-layer distortion loss function is introduced to improve the quality of the finally obtained feature map. The multi-layer loss L m has the following expression:

[0061]

[0062] where N represents the number of convolutional layers, S represents the downsampling operation, and the function S(x, n) represents downsampling the input x to times its original size, and represent the original feature on the decoder of the distortion correction module and the feature after correction by the feature correction layer respectively. cat(·) represents feature concatenation, and Conv represents a 3×3 convolution that decodes the output feature into a 3-channel RGB image. In this way, each feature map on the decoder can be effectively supervised.

[0063] The enhancement loss L e further enriches the texture details and has the following expression:

[0064] L e = L c + λ s L s

[0065] where L c represents the content loss, L s represents the style loss, λ s represents the style loss weight, and the expression of L c is as follows:

[0066]

[0067] where, and represent the feature maps obtained by convolving the corrected image and the ground truth image using the j-th layer in the pre-trained VGG19 network respectively. C j H j W j represents the size of the j-th layer feature map;

[0068] The expression of L s is as follows:

[0069]

[0070] represents the squared Frobenius norm of the Gram matrix, which is a matrix of shape C j × C jThe matrix, whose element expressions are as follows:

[0071]

[0072] where c and c' represent the specified channels, which are set as hyperparameters, and is the weight at the specified channel when the size is h*w. Specific embodiments

[0074] 1) Dataset synthesis

[0075] There are few real images taken in the pipeline using a fisheye lens. The dataset is obtained by synthesis; the Cocotrain2017 dataset is selected as the original data. This dataset has a total of 118,287 images. 5,000 of them are selected as the test dataset, and the others are used as the training dataset. Before synthesis, all images in the dataset are first transformed into images with a size of 256*256, and then each image is distorted and synthesized with the same parameters to obtain pairs of original images and distorted images, which constitute the dataset for the experiments in this paper. The lens viewing angle is set to 120 degrees in the part of the pipe wall distortion, and the lens focal length is 600 pixel distances. In the part of the fisheye lens distortion, the fisheye radial distortion parameters are set as k1 = 1e -5 and k2 = 1e -10 and k3 = 1e -15 and k4 = 1e -20 .

[0076] 2) Experimental settings

[0077] Loss function weight settings: The weight of the L1 loss λ l1 = 12, the weight of the multi-layer loss (pyramid loss) λ m = 0.5, the weight of the adversarial loss λ adv = 0.1, the weight of the style loss λ s = 250, the weight of the content loss λ e = 0.08; The Adam optimizer is used, and the optimizer parameters are set as β1 = 0.5 and β2 = 0.999. The learning rate is set to 1×10 e-4 , and a total of 1,000,000 batches are trained. The size of each batch is 8, and the network is trained using an NVIDIA GeForce GTX 3090 GPU.

[0078] 3) Comparative experiments

[0079] Traditional fish-eye correction methods use the method of establishing a mathematical model by extracting the changes of straight lines and curves from distorted images for image correction. Therefore, they cannot be applied to the synthetic dataset with two types of deformations. The data labels in the synthetic dataset only include the corresponding relationship between the original image and the distorted image, and do not include more complex labels such as distortion annotation and pixel displacement. Therefore, several network models suitable for processing the synthetic dataset in this paper are selected for comparative experiments. The comparison methods include Blind, DR-GAN, and PCN, and all three methods are based on the generation method. In the experiment, ssim (structural similarity), psnr (peak signal-to-noise ratio), and fid (feature vector distance) are used as the criteria for evaluating the quality of the generated images. The results of the comparative experiment are shown in Table 1. The model has better performance than other comparative models in the three index values. Compared with the second-place PCN network, the model has improved by 0.036 in ssim, 2.492 in psnr, and 6.943 in fid.

[0080] Table 1 Comparison of the effects of existing methods

[0081] ssim↑ psnr↑ fid↓ blind 0.817 25.226 36.64 DR-GAN 0.786 26.569 53.363 PCN (Original Network) 0.846 26.612 11.873 Ours 0.882 29.104 4.930

[0082] Three representative images are selected for qualitative comparison. The three images are a person, a distant view, and a close view. It can be clearly seen that the image obtained by Blind has more bad points than that of this application, the image obtained by DR-GAN is more blurred than that of this application, and PCN is generally not much different from that of this application, but there is a blurred phenomenon on the right side of its portrait image, and the details of the railway tracks in the distant view image are not as clear as those of this application. Generally, the images obtained by the model of this application perform better in terms of the quality of the generated images and the restoration of object edges.

[0083] 4) Ablation experiment

[0084] To verify the effectiveness of the two improved modules proposed in this paper, the two modules are replaced or added correspondingly on the original network, and each change is trained and tested. Each training round is 100000 batches, and the size of each batch is 8. The ablation experiments are designed as follows: 1. Only add a multi-scale discriminator on the basis of the original network to verify the effectiveness of the multi-scale discriminator; 2. Only add central self-attention to the generator of the original network to judge its influence on the generator; 3. Only add central self-attention to the discriminator of the original network to judge its influence on the overall network; 4. Add original self-attention to both the generator and discriminator of the original network; 5. Add central self-attention to both the generator and discriminator of the original network, and compare it with the 4th ablation experiment to prove the effectiveness of central self-attention. The results of the ablation experiment are shown in Table 2.

[0085] Table 2 Performance of networks with different structures

[0086] ssim↑ psnr↑ fid↓ Original Network 0.854 27.667 10.010 Only add multi-scale discriminator E 0.866 28.311 9.683 Only add central self-attention in the generator 0.864 28.320 9.706 Only add central self-attention in the original discriminator 0.868 27.438 7.458 Add original self-attention to both the generator and discriminator of the original network 0.865 27.818 10.378 Add central self-attention to both the generator and discriminator of the original network 0.873 28.414 7.904 Add central self-attention to both the generator and multi-scale discriminator 0.874 28.430 8.439

[0087] According to the ablation experiment, it can be seen that introducing a multi-scale discriminator into the network can effectively improve the comprehensive quality of the generated images. Adding a central self-attention module to the generator can effectively improve the quality of the generated images, while adding a central self-attention module to the discriminator can effectively improve the correction similarity of the images, and the central self-attention performs better than the original self-attention on the synthetic dataset. Therefore, adding a central self-attention module to both the generator and the discriminator can effectively improve the overall quality and accuracy of the generated images. Finally, the network that comprehensively uses a multi-scale discriminator and introduces a central self-attention module into both the generator and the multi-scale discriminator has a significant improvement in the quality of the generated images and the correction similarity compared to the original network.

[0088] Central self-attention:

[0089] When using the central self-attention module, it is necessary to make the two-dimensional Gaussian matrix distribution fit the distortion change trend of the fish-eye lens as much as possible. In order to obtain the standard deviation of the Gaussian matrix that fits the current dataset, a two-dimensional Gaussian matrix with a standard deviation of 1.5 is selected. The central self-attention module is introduced into the generator and discriminator of the original network, and the experimental results obtained by experimenting with the Gaussian weight matrix in the central self-attention module are shown in Table 3.

[0090] Table 3 Experimental results obtained using two-dimensional Gaussian matrices with different standard deviations

[0091] Standard Deviation ssim↑ Psnr↑ Fid↓ 0.5 0.869 28.012 9.435 1 0.873 28.345 9.019 1.5 0.873 28.414 7.904 2 0.867 28.240 8.278 2.5 0.873 28.182 11.487 5 0.871 28.370 9.919

[0092] Multi-scale discriminator network:

[0093] After determining the central self-attention module to be used, in order to obtain a better discriminator network, when designing the multi-scale, it is considered that the convolutional layers of each path are in parallel or serial combination, as Figure 4 shown. A total of 6 multi-scale discriminator networks are designed and used to replace the discriminator in the original network for experimental comparison. Among them, (a), (c), and (e) are parallel multi-scale convolutions, and (b), (d), and (f) are serial multi-scale convolutions.

[0094] In discriminator A (as Figure 4 (a) shown), 3×3 and 5×5 convolutions are used to extract the features of the input image. The auxiliary convolution 3×3 extracts features independently and splices the obtained feature maps onto the corresponding feature maps of the main convolution 5×5. The convolution operation before the CSA module can be expressed as follows:

[0095]

[0096] Among them, represents the n*n convolutional feature map of the i-th layer.

[0097] In discriminator B (as shown in Figure 4 (b)), the feature maps obtained by 3×3 and 5×5 convolutions are concatenated each time, and the concatenated feature maps are then convolved using 3×3 and 5×5 convolutions, which can be expressed as follows:

[0098]

[0099] In discriminator C (as shown in Figure 4 (c)), the same structure as discriminator A is used, and only the 3×3 convolution is replaced by a 7×7 convolution, which can be expressed as follows:

[0100]

[0101] In discriminator D (as shown in Figure 4 (d)), the same structure as discriminator B is used, and only the 3×3 convolution is replaced by a 7×7 convolution, which can be expressed as follows:

[0102]

[0103] In discriminator E (as shown in Figure 4 (e)), a total of three different convolution kernel sizes are used to extract features. Among them, the 3×3 convolution and the 7×7 convolution extract features in parallel independently, and the obtained feature maps are concatenated with the feature maps of the corresponding size of the 5×5 convolution, which can be expressed as follows:

[0104]

[0105] In discriminator F (as shown in Figure 4 (f)), three different convolution kernel sizes are also used to extract features. Each time, the feature maps obtained by the three convolutions are concatenated and then the next convolution is performed, which can be expressed as follows:

[0106]

[0107] The comparison experiment results of the 6 multi-scale discriminators are shown in Table 4. Three indicators in the obtained experiment results are comprehensively calculated. The multi-scale discriminator network can provide richer features, effectively improve the ability of the discriminator to judge real images and generated images, and effectively improve the quality and similarity of the generated images.

[0108] Table 4 Comparison Results of 6 Multi-scale Discriminators

[0109] Original Network 0.854 27.667 10.010 Central Self-attention + Discriminator A 0.874 28.325 9.292 Central Self-attention + Discriminator B 0.867 27.901 9.037 Central Self-attention + Discriminator C 0.872 28.345 10.849 Central Self-attention + Discriminator D 0.874 28.421 8.885 Central Self-attention + Discriminator E 0.874 28.430 8.439 Central Self-attention + Discriminator F (Rerun) 0.858 25.844 8.930

[0110] This patent application is for the correction of fisheye distortion images of oil and gas pipe walls. The synthetic dataset uses fixed specific radial distortion parameters and fixed pipe wall distortion parameters, and the distortion center is located at the center of the image. The proposed model is applied to other fisheye distortion images. The standard deviation parameter of the two-dimensional Gaussian matrix used in the central self-attention module is only applicable under the current fisheye lens distortion parameters. When it needs to be applied to images synthesized using other fisheye lens distortion parameters, the standard deviation parameter may need to be adjusted accordingly.

[0111] The above are only the embodiments of the present invention. Once again, it is stated that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements can be made to the present invention, and these improvements are also included in the protection scope of the claims of the present invention.

Claims

1. A method for correcting the fisheye distortion of the pipe wall based on a multi-scale discriminator and central self-attention, which is used for correcting the fisheye distortion network of the pipe wall, and is characterized in that: The distortion network consists of two parts: a generator and a discriminator. The generator includes a flow estimation module and a distortion correction module. Both the flow estimation module and the distortion correction module adopt the UNet structure. After the input image is downsampled and upsampled multiple times, the corrected image is reconstructed by feature extraction, and the discriminator is used for discrimination. The discriminator uses multi-scale convolution to extract features and introduces a central self-attention module to highlight the features of the central region.

2. The method for correcting the fisheye distortion of the pipe wall based on the multi-scale discriminator and the central self-attention according to claim 1, wherein: The input image of the flow estimation module has a size of 256×256. In the encoder of the flow estimation module, a convolutional kernel with a size of 4×4, a stride of 2, and a padding of 1 is used for downsampling. In the decoder, a multi-level upsampling network is adopted. The resolution of the feature map is gradually enlarged through convolutional modules, residual blocks, and transposed convolutional modules. The downsampled feature map is skip-connected with the upsampled feature map of the corresponding size to generate a series of appearance flows. Then, a 3×3 convolution is performed on the output appearance flow to obtain five dual-channel appearance flow pre-correction feature maps with sizes of 128, 64, 32, 16, and 8 respectively, which are used for the decoder part of the distortion correction module. The input image of the distortion correction module has a size of 256×256. In the encoder of the distortion correction module, a convolutional kernel with a size of 3×3, a stride of 2, and a padding of 1 is used for downsampling to obtain distortion feature maps with sizes of 128, 64, 32, 16, and 8 respectively. After the third downsampling, a central self-attention module is added to improve the feature extraction ability. The downsampled distortion feature map of the distortion correction module is fused with the corresponding-size dual-channel appearance flow pre-correction output by the flow estimation module, and then skip-connected with the upsampled feature map of the corresponding size. In the decoder of the distortion correction module, a convolution with a kernel size of 3, a stride of 1, and a padding of 1 and the interpolate function provided by pytorch are used for upsampling by bilinear interpolation to gradually reconstruct the corrected image by feature extraction.

3. The method for correcting the fisheye distortion of the pipe wall based on the multi-scale discriminator and the central self-attention according to claim 2, wherein: The discriminator is a discriminator network using multi-scale convolution. It extracts image features by convolving the input image with a size of 256×256. A convolutional kernel with a size of 5×5 is used in the main feature network of the discriminator. Two auxiliary feature extraction networks are introduced in the discriminator, using convolutional kernels with sizes of 3×3 and 7×7 respectively to extract more detailed features and local texture features of the original image. The features of the three scales are fused with each other, and a central self-attention module is introduced after the third layer of convolution to improve the performance of the discriminator network.

4. The method for correcting the fisheye distortion of the pipe wall based on the multi-scale discriminator and the central self-attention according to claim 1, wherein: The central self-attention module performs V_conv on the input feature map to obtain a value matrix, multiplies the input feature map by a Gaussian weight matrix to obtain Gaussian features, then uses Q_conv and K_conv to convolve the Gaussian features respectively to obtain a query matrix and a key matrix, and further obtains the attention. The expression of the central self-attention module is as follows: Among them, Q_conv represents the convolution of the query matrix, V_conv represents the convolution of the value matrix, K_conv represents the convolution of the key matrix, GCW represents the Gaussian weight matrix, and d k represents the vector dimension in the key matrix; The calculation of the Gaussian weight matrix uses the following formula: Among them, the standard deviation parameter σ of the Gaussian matrix is 1.

5.

5. The method for correcting the fish-eye distortion of the pipe wall based on the multi-scale discriminator and the central self-attention according to claim 1, wherein: A total loss function training strategy including multiple loss functions is also adopted. The total loss function is obtained through reconstruction loss, adversarial loss, multi-layer loss and enhancement loss. The expression of the total loss function is as follows: L = λl1L1 + λ adv L adv + λ m L m + λ e L e Among them, λ l1 , λ adv , λ m and λ e all represent hyperparameters, L1 represents the reconstruction loss, L adv represents the adversarial loss, L m represents the multi-layer loss, L e represents the enhancement loss; Among them, the adversarial loss L adv has the following expression: L adv = min G max D (E[logD(Img gt )] + E[log(1 - D(G(Img out )))]) Among them, G represents the generator, D represents the discriminator, E represents the expectation, Img gt represents the true value image, Img out represents the corrected image generated by the generator; Multi-layer loss L m The expression is as follows: Among them, N represents the number of convolutional layers, S represents the downsampling operation, and the function S(x, n) represents downsampling the input x to times the original size, and respectively represent the original features on the decoder of the distortion correction module and the features after correction by the feature correction layer. cat(·) represents feature concatenation, and Conv represents a 3×3 convolution that decodes the output features into a 3-channel RGB image; Enhanced loss L e has the following expression: L e = L c + λ s L s Among them, L c represents the content loss, and L s represents the style loss. λ s represents the style loss weight. The expression of L c is as follows: Among them, and respectively represent the feature maps obtained by convolving the corrected image and the ground truth image using the j-th layer in the pre-trained VGG19 network. C j H j W j represent the size of the j-th layer feature map; L s The expression is as follows: represents the squared Frobenius norm of the Gram matrix, which is a matrix of shape C j ×C j and its element expression is as follows: Among them, c and c' represent the specified channels, and is the weight of the specified channel when the size is h*w.

Citation Information

Cited By

  • Fisheye image correction method and system based on synthetic distortion enhancement

    CN120807371A