Tunnel inspection image enhancement method and system based on conditional generative adversarial network
Through the image enhancement method of condition-generating adversarial networks, using the pix2pix model and spatial attention mechanism, the problem of poor image quality in tunnel patrol is solved, and efficient and robust image quality improvement is achieved to meet real-time requirements.
Patent Information
- Application Number
- CN202510740549.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The complex lighting conditions and environmental interference inside the tunnel lead to poor image acquisition quality. The existing image enhancement methods cannot effectively improve the image quality of tunnel patrols, and the calculation complexity and cost are high, making it difficult to meet real-time requirements.
The image enhancement method based on conditional generation adversarial network is adopted, and the pix2pix model is used as the core architecture, combined with the spatial attention mechanism, adaptively focus on key areas and improve image quality.
Significantly improve the image quality of tunnel patrols, improve model robustness, reduce computing complexity, meet real-time requirements, and provide reliable data support.
Smart Images

Figure CN120259090A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of tunnel image processing, and in particular to a tunnel inspection image enhancement method and system based on a conditional generative adversarial network. Background Art
[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] As a key component of transportation infrastructure, the safe and stable operation of tunnels is directly related to the stability of transportation and the safety of people's lives and property. To ensure the safe operation of tunnels, regular inspections have become an essential part. Among various inspection methods, the inspection method based on image acquisition is widely used in tunnel inspection work. By regularly inspecting and obtaining tunnel images, it can intuitively reflect the disease conditions such as cracks and water seepage on the surface of the tunnel lining, provide key basis for tunnel maintenance decisions, and play a key role in the safe operation of the tunnel structure. However, the unique environmental conditions inside the tunnel bring many problems to image acquisition.
[0004] Due to the extremely complex and harsh lighting conditions in the tunnel, the lighting intensity is high in the tunnel entrance area because of direct contact with external light, and the collected images are prone to over-bright or even over-exposed phenomena, resulting in the loss of key detail information; while deep in the tunnel, the lighting is severely insufficient, and the images often appear dim, and many disease characteristics are difficult to clearly show. This lighting non-uniformity greatly increases the difficulty of comprehensively obtaining effective information. Secondly, the humid air, diffuse dust and electromagnetic interference generated by the operation of various devices in the tunnel make the collected images inevitably mixed with noise, seriously reducing the image quality and creating many obstacles for subsequent disease identification and analysis work. In addition, the existing method of obtaining high-quality images by strengthening the acquisition equipment has a high acquisition cost, a high later maintenance cost, and low detection efficiency. Therefore, a method for efficiently enhancing the quality of tunnel inspection images is needed.
[0005] Currently, traditional image enhancement methods have obvious limitations when dealing with the quality optimization of tunnel inspection images. For example, histogram equalization can improve the overall contrast of the image, but in the scene where the brightness difference of the tunnel image is extremely large, it is easy to cause over-enhancement or loss of local details, and at the same time, it will also amplify the noise in the image, further deteriorating the image quality. Another homomorphic filtering method can separate the illumination and reflection components of the image, but it has extremely high requirements for the setting of filter parameters, is difficult to adapt to the complex and changeable environmental conditions in the tunnel, and the algorithm has a high computational complexity and cannot meet the strict requirements for real-time performance of the inspection work. Summary of the Invention
[0006] To solve the above problems, the present disclosure proposes a tunnel inspection image enhancement method and system based on a conditional generative adversarial network. The pix2pix model is used as the core architecture of the conditional generative adversarial network, and the tunnel inspection image is used as a conditional constraint to ensure that the quality-enhanced image output by the generator is completely aligned with the tunnel inspection image in terms of spatial position. At the same time, a spatial attention mechanism is introduced to adaptively focus on the key regions most relevant to the enhancement task in the image, better capture the local structure and texture information of the image, ignore the irrelevant information and noise in the image, and improve the robustness of the model.
[0007] According to some embodiments, the present disclosure adopts the following technical solutions: A tunnel inspection image enhancement method based on a conditional generative adversarial network, comprising: Obtain a tunnel image to be enhanced; input the tunnel image to be enhanced into a trained image enhancement model, and output an enhanced optimized image; Wherein, the image enhancement model is constructed based on a conditional generative adversarial network, and the specific training process is as follows: Obtain a high-quality tunnel image under ideal conditions, and perform pre-processing on the high-quality image to obtain a low-quality tunnel image; Extract feature maps of different channel numbers and different sizes of the low-quality tunnel image through the generator of the conditional generative adversarial network, and output the corresponding quality-enhanced tunnel inspection image after weighted processing by the spatial attention mechanism, denoted as a fake high-quality image; Splice the low-quality tunnel image and the high-quality tunnel image to form a real image pair, splice the low-quality tunnel image pair and the fake high-quality image to form a fake image pair, input the fake image pair and the real image pair into the discriminator of the conditional generative adversarial network respectively, and obtain a discrimination result; perform cyclic training until the requirements are met.
[0008] According to some embodiments, the present disclosure adopts the following technical solutions: A tunnel inspection image enhancement system based on a conditional generative adversarial network, comprising: An image acquisition module, configured to obtain a tunnel image to be enhanced; An image enhancement module, configured to input the tunnel image to be enhanced into a trained image enhancement model, and output an enhanced optimized image; Wherein, the image enhancement model is constructed based on a conditional generative adversarial network, and the specific training process is as follows: Obtain a high-quality tunnel image under ideal conditions, and perform pre-processing on the high-quality image to obtain a low-quality tunnel image; The generator of the conditional generative adversarial network extracts feature maps of different numbers of channels and different sizes from low-quality tunnel images, and outputs corresponding quality-enhanced tunnel inspection images through weighted processing by a spatial attention mechanism, which are recorded as fake high-quality images; The low-quality tunnel images and high-quality tunnel images are spliced to form real image pairs, and the low-quality tunnel image pairs and fake high-quality images are spliced to form fake image pairs. The fake image pairs and real image pairs are respectively input into the discriminator of the conditional generative adversarial network to obtain discrimination results; The training is cycled until the requirements are met.
[0009] According to some embodiments, the present disclosure adopts the following technical solutions: A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the tunnel inspection image enhancement method based on a conditional generative adversarial network.
[0010] According to some embodiments, the present disclosure adopts the following technical solutions: A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the tunnel inspection image enhancement method based on a conditional generative adversarial network is implemented.
[0011] According to some embodiments, the present disclosure adopts the following technical solutions: An electronic device includes: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the tunnel inspection image enhancement method based on a conditional generative adversarial network.
[0012] Compared with the prior art, the beneficial effects of the present disclosure are: The tunnel inspection image enhancement method based on a conditional generative adversarial network of the present disclosure designs an image enhancement model based on a conditional generative adversarial network, adopts the pix2pix model as the core architecture of the network, and uses tunnel inspection images as conditional constraints to ensure that the quality-enhanced images output by the generator are completely aligned with the tunnel inspection images in terms of spatial position.
[0013] The tunnel inspection image enhancement method based on a conditional generative adversarial network of the present disclosure introduces a spatial attention module to further improve the feature extraction ability of the model, enabling the pix2pix model to adaptively focus on the key regions most relevant to the enhancement task in the image, better capture the local structure and texture information of the image, ignore the irrelevant information and noise in the image, and improve the robustness of the model.
[0014] The tunnel inspection image enhancement method based on conditional generative adversarial network of the present disclosure effectively makes up for the deficiencies of the above traditional image enhancement methods, significantly improves the quality of tunnel inspection images, and provides more reliable and efficient data and technical support for tunnel safety operation and maintenance work. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the disclosure. The illustrative embodiments and descriptions thereof of the disclosure are used to explain the disclosure and do not constitute an improper limitation of the disclosure.
[0016] Figure 1 It is a schematic diagram of the generator network structure of an embodiment of the present disclosure; Figure 2 It is a schematic diagram of the discriminator network structure of an embodiment of the present disclosure; Figure 3 It is a schematic diagram of the training process of the image enhancement model based on conditional generative adversarial network of an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0018] It should be noted that the following detailed descriptions are all illustrative and are intended to provide a further description of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0019] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0020] TERMINOLOGY EXPLANATION pix2pix model: As an image derivative model based on the GAN architecture, it can learn the mapping relationship between the input image and the target image and performs well in tasks such as image restoration and enhancement.
[0021] GAN adversarial loss: It is a binary cross-entropy loss function. The input of the adversarial loss is the image generated by the generator, and the output is the degree to which the generated image is judged as a real image by the discriminator. The purpose is to make the image generated by the generator able to deceive the discriminator.
[0022] L1 Loss: To ensure the similarity between the generated image and the real image at the pixel level. The input of the L1 loss is the image generated by the generator, and the output is a scalar value representing the pixel-level difference degree between the generated image and the real image. The smaller the output value, the closer the generated image and the real image are at the pixel level.
[0023] Spatial Attention-guided Reconstruction Loss Function: Aims to ensure that the image generated by the generator using spatial attention is more similar to the target image in key regions. By combining the spatial attention map with the reconstruction loss, it prompts the generator to pay more attention to the image reconstruction quality in important spatial regions.
[0024] tf.image.resize function: A function in TensorFlow used to resize images, supporting various interpolation methods and configuration options.
[0025] High-quality Tunnel Image: A tunnel image obtained under ideal conditions. Specifically, an industrial camera with a resolution greater than 300 DPI is used to obtain an image with a reasonable gray value distribution range of 50 - 255 under set brightness, temperature, humidity, and non-reflective conditions, a resolution greater than 1920×1080, and noise (measured by standard deviation) less than 5%.
[0026] Low-quality Tunnel Image: An image obtained by processing a high-quality tunnel image with noise addition, brightness reduction, and blurring, etc. More specifically, the gray value of the low-quality tunnel image is lower than 50, the resolution is lower than 1920×1080, and the noise (measured by standard deviation) is higher than 5%.
[0027] Example 1 In one embodiment of the present disclosure, a tunnel inspection image enhancement method based on a conditional generative adversarial network is provided, and the steps include: Step 1: Obtain the tunnel image to be enhanced; Step 2: Input the tunnel image to be enhanced into the trained image enhancement model, and output the optimized enhanced image; Among them, the image enhancement model is constructed based on a conditional generative adversarial network, and the specific training process is as follows: Obtain high-quality tunnel images under ideal conditions, and perform pre-processing on the high-quality images to obtain low-quality tunnel images; Extract feature maps of different channel numbers and different sizes of the low-quality tunnel images through the generator of the conditional generative adversarial network, and output the corresponding quality-enhanced tunnel inspection images after weighted processing by the spatial attention mechanism, denoted as fake high-quality images; The low-quality tunnel images and high-quality tunnel images are stitched together to form real image pairs, and the low-quality tunnel image pairs and fake high-quality images are stitched together to form fake image pairs. The fake image pairs and real image pairs are respectively input into the discriminator of the conditional generative adversarial network to obtain discrimination results; training is performed in a loop until the requirements are met.
[0028] As an embodiment, a tunnel inspection image enhancement method based on a conditional generative adversarial network of the present disclosure uses deep learning methods, especially generative adversarial networks (GANs) and their derivative models, which exhibit powerful image generation and processing capabilities. The pix2pix model, as an image derivative model based on the GAN architecture, can learn the mapping relationship between the input image and the target image and performs excellently in tasks such as image restoration and enhancement. The present disclosure proposes an image enhancement model based on a conditional generative adversarial network, using the pix2pix model as the core architecture of this learning network, with tunnel inspection images as conditional constraints to ensure that the quality-enhanced images output by the generator are completely aligned with the tunnel inspection images in terms of spatial position. To further improve the feature extraction ability of the model, a spatial attention module is introduced, enabling the pix2pix model to adaptively focus on the key regions most relevant to the enhancement task in the image, better capture the local structure and texture information of the image, ignore the irrelevant information and noise in the image, and improve the robustness of the model. The specific training implementation process of the image enhancement model based on the conditional generative adversarial network of the present disclosure is as follows: Step 1: Image data acquisition. Tunnel inspection images are collected through various methods and preprocessed.
[0029] Specifically, high-quality tunnel images under ideal conditions are obtained through various methods. The process of obtaining high-quality tunnel images under ideal conditions includes: deploying industrial cameras with high resolution and low noise, equipped with optical lenses with automatically adjustable parameters; placing the cameras in specially made waterproof, dustproof and anti-reflection covers to isolate interference such as humid water vapor, dust and light reflection in the tunnel; installing an LED matrix fill light and adjusting it in real time according to the tunnel lighting to obtain high-quality tunnel images with high brightness, high resolution and low noise.
[0030] The obtained high-quality tunnel images are processed by means such as adding noise, reducing brightness and blurring to obtain low-quality tunnel images. The low-quality tunnel images are images obtained by processing high-quality tunnel images with noise addition, brightness reduction and blurring, and have the characteristics of low brightness, low resolution and high noise.
[0031] Among them, in practical applications, due to the dark light in the tunnel, the actually obtained images are low-quality tunnel images to be enhanced.
[0032] Furthermore, the obtained high-quality tunnel images under ideal conditions are used as the target reference that the generator expects to output during the subsequent training process of the image enhancement model.
[0033] Step 2: Data preprocessing. Perform operations such as size adjustment and normalization on the obtained low-quality and high-quality tunnel images in sequence to provide standardized data samples for the construction of the subsequent model training dataset.
[0034] 1) Size adjustment: Use bilinear interpolation to scale the tunnel inspection images. By calling the tf.image.resize function in TensorFlow, the image can be adjusted to the specified height and width, such as 256×256.
[0035] 2) Normalization: Normalize the pixel values of the tunnel inspection images to the range [-1, 1]. Specifically, it can be performed through the following formula:
[0036] where, is the original pixel value, x is the pixel value after normalization.
[0037] Step 3: Dataset division. The preprocessed dataset is usually divided into a training set, a validation set, and a test set in the ratio of 70%, 15%, and 15%. The training set is used to learn the training parameters of the model, the validation set is used to evaluate the performance of the model during training and adjust the hyperparameters, and the test set is used to evaluate the generalization ability of the model.
[0038] Step 4: Construct an image enhancement model based on a conditional generative adversarial network with the tunnel inspection images as conditional constraints. The image enhancement model based on the conditional generative adversarial network adopts the Pix2Pix architecture, including a generator network and a discriminator network. The specific structure is as follows: 1) The generator network is the NetG network, including a downsampling module, an upsampling module, and a skip connection module. A spatial attention mechanism is added to the downsampling module and the upsampling module. And the generator loss function includes a GAN adversarial loss function, an L1 loss function, and a reconstruction loss function guided by spatial attention.
[0039] The downsampling module consists of 1 convolutional layer, 1 activation function layer (LeakyReLU function), and 1 normalization layer. The spatial dimension of the input image is gradually reduced through the convolutional layer, while the number of feature channels is increased to extract the high-level features of the tunnel inspection images, helping the generator network capture the semantic information and context information in the tunnel inspection images.
[0040] The upsampling module consists of 1 transposed convolution layer, 1 activation function layer (ReLU function), and 1 normalization layer, and a dropout layer is added to improve the robustness of the generator network, reduce the dependence on specific neurons, and avoid overfitting. The upsampling module gradually restores the spatial dimension of the feature map to the original image size through transposed convolution while reducing the number of feature channels.
[0041] During the downsampling process of the skip connection module, after operations through multiple convolution layers and normalization layers, feature maps with different numbers of channels and different sizes will be obtained. When starting upsampling, the feature map obtained by upsampling is concatenated with the feature map with the same spatial size during the downsampling process to achieve skip connection, which helps to retain feature information with different numbers of channels, avoid information loss, and enables the generator to utilize low-level and high-level feature information to generate more accurate tunnel inspection images.
[0042] Among them, to solve the problems of insufficient model feature representation ability and insufficient utilization of spatial information by the model, a spatial attention mechanism (Spatial Attention Mechanism, SAM) is added to both the downsampling module and the upsampling module of the generator network to improve the generalization ability of the model, reduce the risk of overfitting, and enhance the interpretability of the model.
[0043] As an embodiment, the generator extracts feature maps with different numbers of channels and different sizes from the conditional input image (low-quality tunnel image), and outputs the corresponding quality-enhanced tunnel inspection image, denoted as the fake high-quality image, after weighted processing by the spatial attention mechanism. The processing procedures of different modules in the generator are as follows: ① Downsampling module The input is the preprocessed low-quality tunnel inspection image, which is used as the input feature map of the downsampling module; Perform convolution operations on the input feature map (convolution kernel size is 4×4, stride is 2). After each convolution, the size of the feature map becomes half of the original, and the number of channels becomes twice the original, so as to extract more advanced and abstract features in the image; Perform max-pooling and average-pooling operations on the convolved feature map respectively, and output two single-channel feature maps to screen out important information in the input feature map; Concatenate the two single-channel feature maps output after max-pooling and average-pooling, and output a more comprehensive concatenated feature map; Perform convolution operations on the concatenated feature map (convolution kernel size is 3×3). The convolution kernel will learn the importance of each element in the concatenated feature map in the spatial position, and then output the spatial attention map; To enable the model to pay more attention to key feature regions such as abnormal lighting, blurriness, and noise interference, the generated spatial attention map is weighted element-wise with the original input feature map at different spatial positions to assign different weights to different regions in the feature map, resulting in a weighted feature map; The finally obtained weighted feature map will be used as the new input feature map for reprocessing. The downsampling module will repeat the above operations several times (until the number of channels of the feature map is 512 and the size is 4×4), and each operation will reduce the size of the feature map and increase the number of channels to extract more advanced feature information.
[0044] ② Upsampling module Except that the input for the initial upsampling is the final feature map generated by the downsampling module (with 512 channels and a size of 4×4), the input for the remaining upsampling modules is the feature map output by the previous upsampling layer; The input for upsampling is the fused feature map concatenated by the skip connection module. Max-pooling and average-pooling operations are performed on the fused feature map to obtain two single-channel feature maps to screen important information; The two single-channel feature maps output after max-pooling and average-pooling are concatenated to form a concatenated feature map; Convolution operation (with a kernel size of 3×3) is performed on the concatenated feature map to generate a spatial attention map.
[0045] The generated spatial attention map is weighted element-wise with the original input feature map at different spatial positions to adjust the weights of different regions in the feature map, resulting in a weighted feature map; Transposed convolution method (with a kernel size of 4×4 and a stride of 2) is used for upsampling the weighted feature map. After each transposed convolution, the size of the feature map becomes twice the original, and the number of channels becomes half of the original, thereby gradually restoring the resolution of the original image; The upsampling operation will expand the size of the feature map and reduce the number of channels through transposed convolution to gradually restore the resolution of the image. At the same time, the upsampling process will perform skip connections with the corresponding feature maps in the downsampling process (see the skip connection module for details).
[0046] ③ Skip connection module The input feature map of the skip connection module is the feature map output by the downsampling. The feature map of the corresponding layer in the downsampling is directly concatenated with the feature map of the corresponding layer in the upsampling in the channel dimension, that is, the downsampling feature map and the upsampling feature map with the same spatial size are stacked in the channel direction. After concatenation, a fused feature map is obtained, which contains both high-level semantic information and low-level detail information, thus obtaining a more comprehensive and rich feature representation; The fused feature map, as the output of the skip connection module, will be passed to the subsequent upsampling module and continue to participate in the model's calculation and feature extraction process; The role of the skip connection module is to fuse the low-level feature information extracted in the downsampling module with the feature information with restored resolution in the upsampling module, so that the generated image can retain both detailed information and have a high resolution.
[0047] 2) The discriminator network is the NetD network. The NetD network adopts the discriminator of PatchGAN. The discriminator divides the input image into multiple small blocks, discriminates each small block, and finally takes the average of the sum of the evaluations of all small blocks as the final evaluation result. The discriminator network consists of 3 convolutional layers, 1 normalization layer, 1 fully connected layer and 1 spatial attention module. It extracts image features through convolution and normalization processing, converts the input image into a single-channel feature map, and then focuses on the spatial region of the key discriminant information in the image through the spatial attention module. The specific structure is as follows: The discriminator adopts the discriminator of PatchGAN. Instead of discriminating the entire tunnel inspection image, the discriminator divides the tunnel inspection image into multiple small blocks, discriminates each small block, and finally takes the average of the sum of the evaluations of all small blocks as the final evaluation result.
[0048] The discriminator network consists of 3 convolutional layers (convolution kernel size 4×4), 1 normalization layer, 1 fully connected layer and 1 spatial attention module. It extracts the features of the tunnel inspection image through convolution and normalization processing, and can convert the input image into a single-channel feature map; then through the spatial attention module, the discriminator can more accurately focus on the spatial region with key discriminant information in the image, thereby improving its ability to distinguish between generated images and real images; through the fully connected layer, the feature map is mapped to a low-dimensional vector space, and finally through the Sigmoid activation function, the probability value is output. The structure of this discriminator network can arbitrarily adjust the number of convolutional layers, and can also choose whether to set the bias term (vector).
[0049] As an embodiment, different from the discriminator in the traditional adversarial generation network, the input of the discriminator here is no longer a single input, but a pair of images. The low-quality tunnel image, high-quality tunnel image and fake high-quality image are paired in different forms to obtain real image pairs and fake image pairs. The real image pair is composed of a low-quality tunnel image and a high-quality tunnel image spliced together; the fake image pair is composed of a low-quality tunnel image pair and a fake high-quality image output by the generator spliced together. Specifically as follows: The first type of real image pair (low-quality tunnel image, high-quality tunnel image), where the low-quality image is the input image of the generator (preprocessed low-brightness / blurry / noisy image), and the high-quality image is the image corresponding to the low-quality image (preprocessed normal-brightness / clear image). The spliced real image pair is input into the discriminator to determine whether this pair of images is a truly matching pair.
[0050] The second type of fake image pair (low-quality tunnel image, fake high-quality image), where the low-quality tunnel image is the input image of the generator (preprocessed low-brightness / blurry / noisy image), and the fake high-quality image is the image generated by the generator based on the low-quality image. The spliced fake image pair is input into the discriminator to determine whether this pair of images is a truly matching pair.
[0051] As an embodiment, a spatial attention module is also introduced into the discriminator. The spatial attention module is added after the second convolutional layer in the discriminator network because at this time, the feature map has undergone feature extraction and has semantic information, which can enable the discriminator to better focus on the key areas before making a final judgment and enhance the feature expression ability.
[0052] Specifically, the processing flow in the discriminator includes: 1) Input image pair splicing There are two types of input image pairs for the discriminator: The first type of real image pair is the real image pair spliced from a low-quality tunnel image (preprocessed low-brightness / blurry / noisy image) and a high-quality tunnel image (preprocessed normal-brightness / clear image).
[0053] The second type of fake image pair is the fake image pair spliced from a low-quality tunnel image (preprocessed low-brightness / blurry / noisy image) and a fake high-quality image (image generated by the generator based on the low-quality tunnel image).
[0054] 2) Two-layer convolution to extract initial features Perform a convolution operation on the spliced image pair using a convolutional layer with a kernel size of 4×4. Through two convolution operations, preliminary feature information is extracted from the image pair, while gradually reducing the image resolution and increasing the number of feature channels, converting the image pair into a feature map representation.
[0055] After two-layer convolution processing, the feature map contains more semantic information. At this time, adding a spatial attention module can enable the discriminator to better focus on the key areas before making a final judgment and enhance the feature expression ability.
[0056] 3) Spatial attention mechanism: Perform max pooling and average pooling operations on the feature maps output by the second convolutional layer, respectively generating two single-channel feature maps. Max pooling highlights the significant features in the feature map, while average pooling pays more attention to the overall feature distribution, thereby screening out the important information in the feature map.
[0057] Concatenate the two single-channel feature maps output by max pooling and average pooling to form a concatenated feature map containing more comprehensive information.
[0058] Perform a convolution operation on the concatenated feature map (with a convolution kernel size of 3×3). The convolution kernel will learn the importance of each element in the new feature map in terms of spatial position, generating a spatial attention map, which represents the importance degree of different spatial positions in the original feature map.
[0059] Perform element-wise weighted processing on the generated spatial attention map and the feature map output by the second convolutional layer according to different spatial positions to obtain a weighted feature map. Through weighted adjustment, different weights are assigned to different regions in the feature map, enabling the discriminator to more specifically focus on the key regions and features in the image pair during subsequent processing, such as the regions where important information such as brightness differences and structural consistency in the image pair are located.
[0060] 4) Feature extraction by the third convolutional layer Perform a convolution operation on the weighted feature map processed by the spatial attention module using a 4×4 convolution kernel to further extract and refine features, enhancing the abstraction degree and discriminative ability of the features.
[0061] 5) Normalization processing Perform normalization processing on the feature map output by the third convolutional layer. Normalization can stabilize the training process of the network, accelerate convergence, and help alleviate the problems of gradient disappearance or gradient explosion, making the data distribution of the feature map more stable and reasonable.
[0062] 6) Output the judgment result After the above series of processes, the discriminator makes a final judgment on the feature map output by the normalization layer. Map the feature map to a low-dimensional vector space through a fully connected layer, and then use the Sigmoid activation function to output a scalar value. This scalar value represents the discriminator's judgment result on whether the input image pair is a true match. The closer the value is to 1, the higher the probability that the discriminator believes the image pair is a true match, and the closer it is to 0, the higher the probability that it is a false match.
[0063] After adding a spatial attention module to the discriminator, the discriminator can automatically focus on the spatial regions in the image that have a greater impact on the discrimination result. For example, key feature regions such as abnormal lighting, blurriness, and noise interference that may exist in tunnel inspection images, which helps the discriminator more accurately identify the differences between the generated image and the real image, thus prompting the generator to continuously improve the quality of the generated image and enhancing the performance of the entire generative adversarial network.
[0064] Step 5: Construct the loss function of the image enhancement model based on the conditional generative adversarial network. The loss function of the image enhancement model based on the conditional generative adversarial network includes the generator loss function and the discriminator loss function. The generator loss function includes the GAN adversarial loss function (Generative Adversarial Network Adversarial Loss), the L1 loss function (Mean Absolute Error Loss), and the spatial attention-guided reconstruction loss function (Spatial Attention -Guided Reconstruction Loss).
[0065] (1) The generator loss function is as follows: ①GAN adversarial loss function The GAN adversarial loss is a binary cross-entropy loss function. The input of the adversarial loss is the image generated by the generator, and the output is the degree to which the generated image is judged as a real image by the discriminator. The purpose is to make the image generated by the generator able to deceive the discriminator. The adversarial loss is used to measure the discrimination difference between the image generated by the generator and the real image in the discriminator, prompting the image generated by the generator to be judged as a real image in the discriminator.
[0066]
[0067] Where: x is the input image; G(x) is the generator G Based on the input image x The generated image; D(G(x)) is the discriminator D's discrimination result on the generated image G(x) is a probability value, representing the probability that the discriminator believes the image is a real image. E x Represents the mathematical expectation of the distribution that the variable x follows (i.e., over the entire distribution of x , perform an average calculation on (D(G(x)) - 1)² to measure the average situation of the discriminator's discrimination result on the samples generated by the generator).
[0068] ②L1 loss function The L1 loss is to ensure the similarity between the generated image and the real image at the pixel level. The input of the L1 loss is the image generated by the generator, and the output is a scalar value representing the degree of pixel-level difference between the generated image and the real image. The smaller the output value, the closer the generated image and the real image are at the pixel level. It can directly calculate the average of the absolute differences between the generated image and the real image at the pixel level, ensuring that the generated image is closer to the real image in terms of details and content.
[0069]
[0070] Where: x is the input image; y is the real image corresponding to the input image x; G(x) is the generator G Based on the input image x The generated image. E xy Represents the mathematical expectation of the joint distribution with respect to the variables x and y (i.e., on the joint distribution of x and y Perform an average calculation on y - G (x 1 ) To measure the average difference between the generator output and the real target).
[0071] ③ Reconstruction loss function guided by spatial attention The reconstruction loss function guided by spatial attention aims to ensure that the image generated by the generator using spatial attention is more similar to the target image in key regions. By combining the spatial attention map with the reconstruction loss, it prompts the generator to pay more attention to the image reconstruction quality in important spatial regions.
[0072] Let the input image of the generator be x, the generated image be G(x) , the target image be y , and the spatial attention map be A G . First, normalize the attention map A G So that its element values are in the range of [0,1], and then calculate the per-pixel mean squared error (MSE) loss and weight it according to the attention map:
[0073] Where: H, W, C Are the height, width, and number of channels of the image respectively, A ij Is the element value at the position ( i, j ) in the attention map, G(x)ijk and y ijk are the pixel values of the generated image and the target image at the position ( i, j, k ), respectively.
[0074] Combining GAN the adversarial loss, L1 the loss, and the spatial attention-guided reconstruction loss function together can take into account the fidelity of the generated image (adversarial loss), detail accuracy ( L1 the loss), and the quality of key regions (spatial attention-guided reconstruction loss function). By introducing the hyperparameter to adjust the weights of the adversarial loss and L1 the loss, the flexibility and adaptability of the model are increased.
[0075]
[0076] Among them: is a hyperparameter used to adjust GAN the weights of the adversarial loss, L1 the loss, and the spatial attention-guided reconstruction loss.
[0077] (2) The discriminator loss function is as follows: The discriminator loss function uses the least squares loss function (Least Squares GAN, LSGAN) and the attention region feature difference loss function (Attention Region Feature Difference Loss).
[0078] ① Least squares loss function When the discriminator input image is a pair of fake images, it is expected that the output D of the discriminator D for the generated image pair G (low-quality tunnel image,
[0079]
[0080] (low-quality tunnel image)) is close to 0. By minimizing this part of the loss, the discriminator is trained to recognize the generated image pair as fake, that is, to output a lower probability value (close to 0). D When the discriminator input image is a pair of real images (low-quality tunnel image, high-quality tunnel image), it is expected that the output
[0081]
[0082] ② Attention Region Feature Difference Loss Function This loss function focuses on the feature differences in the regions of interest of spatial attention, prompting the discriminator to more keenly capture the feature differences between the generated image and the real image in the key regions, and enhancing the discriminative ability for the generated image.
[0083] Let the real image be y , and the generated image be G(x) . The spatial attention map of the discriminator is A D . First, input the real image and the generated image into the discriminator to obtain the corresponding feature maps F y and F G(x) . Binarize the attention map A D to obtain the mask M D , only retaining the regions of interest of attention, and calculate the mean square error (MSE) of the corresponding elements in the feature maps at the corresponding positions within the attention region:
[0084] where: N M is the total number of elements in the mask M D , F yi and F G(x)i are the feature element values of the real image and the generated image at the corresponding positions within the mask region, respectively.
[0085] Introduce the hyperparameter to adjust the weights of the least squares loss and the attention region feature difference loss, increasing the flexibility and adaptability of the model. The total loss of the discriminator is as follows:
[0086] Step 6: Use Adam optimizer to update the parameters of the generator and the discriminator to minimize their loss functions. Set the hyperparameter learning rate ( learning rate ) to 0.0002, betas to be (0.5, 0.999).
[0087] Step 7: Training loop. The training loop process includes: (1) Forward propagation: In each training iteration, the input image and the real target image are used as a real image pair and input into the discriminator to obtain the output of the discriminator for the real image pair. Meanwhile, the input image is passed through the generator to obtain a generated image, and the input image and the generated image are used as a generated image pair and input into the discriminator to obtain the output of the discriminator for the generated image pair.
[0088] (2)Calculate the loss: For the discriminator, according to its outputs for the real image pair and the generated image pair, use the loss function of the discriminator to calculate the loss of the discriminator. For the generator, according to the generated image and the real target image, use the loss function of the generator to calculate the loss of the generator.
[0089] (3)Backpropagation: Calculate the gradient of the discriminator loss with respect to the discriminator parameters, and the gradient of the generator loss with respect to the generator parameters. Update the parameters of the model through the backpropagation algorithm, continuously adjust the weights of the model, so that the generated image gradually approaches the real high-quality enhanced image.
[0090] (4)Update the parameters: Use the optimizer of the discriminator to update the parameters of the discriminator according to the calculated gradient. Use the optimizer of the generator to update the parameters of the generator according to the calculated gradient.
[0091] Step 8: Adjust the learning rate. Adjust the learning rate according to a predefined strategy. After a specific number of training steps or training epochs (such as every 100 epochs), multiply the learning rate by a decay factor (such as 0.1) to achieve the adjustment of the learning rate.
[0092] Step 9: Regularly save the checkpoint of the model during training. Set the save frequency epoch. After each set epoch, save the model parameters of the generator and the discriminator for subsequent resumed training, evaluating the model performance, or performing model inference.
[0093] Step 10: Model application. Set the total number of training epochs epoch. When the training reaches the set number of epochs of the predetermined epoch, stop the training, test and evaluate the quality of the generated pictures, and perform specific applications after passing the test.
[0094] Embodiment 2 In an embodiment of the present disclosure, a tunnel inspection image enhancement system based on a conditional generative adversarial network is provided, including: An image acquisition module for acquiring a tunnel image to be enhanced; An image enhancement module for inputting the tunnel image to be enhanced into a trained image enhancement model and outputting an enhanced optimized image; Wherein, the image enhancement model is constructed based on a conditional generative adversarial network, and the specific training process is as follows: Obtain high-quality tunnel images under ideal conditions, and perform preprocessing on the high-quality images to obtain low-quality tunnel images; Extract feature maps of different numbers of channels and different sizes of the low-quality tunnel images through the generator of the conditional generative adversarial network, and output corresponding quality-enhanced tunnel inspection images after weighted processing by the spatial attention mechanism, denoted as fake high-quality images; Stitch the low-quality tunnel images and high-quality tunnel images to form real image pairs, and stitch the low-quality tunnel image pairs and fake high-quality images to form fake image pairs. The fake image pairs and real image pairs are respectively input into the discriminator of the conditional generative adversarial network to obtain discrimination results; perform cyclic training until the requirements are met.
[0095] Embodiment 3 In an embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the tunnel inspection image enhancement method based on the conditional generative adversarial network.
[0096] Embodiment 4 In an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, it implements the tunnel inspection image enhancement method based on the conditional generative adversarial network.
[0097] Embodiment 5 In an embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the tunnel inspection image enhancement method based on the conditional generative adversarial network.
[0098] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or boxes Figure 1 one process or a plurality of processes and / or boxes Figure 1 steps for realizing the functions specified in one box or a plurality of boxes
[0100] Although the specific embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.
Claims
1. A tunnel inspection image enhancement method based on conditional generative adversarial network, characterized in that Including: Obtain the tunnel image to be enhanced; Input the tunnel image to be enhanced into the trained image enhancement model, and output the enhanced optimized image; Among them, the image enhancement model is constructed based on the conditional generative adversarial network, and the specific training process is as follows: Obtain high-quality tunnel images under ideal conditions, and perform pre-processing on the high-quality images to obtain low-quality tunnel images; Extract feature maps of different channel numbers and different sizes of the low-quality tunnel image through the generator of the conditional generative adversarial network, and output the corresponding quality-enhanced tunnel inspection image after weighted processing by the spatial attention mechanism, denoted as the fake high-quality image; Stitch the low-quality tunnel image and the high-quality tunnel image to form a real image pair, and stitch the low-quality tunnel image pair and the fake high-quality image to form a fake image pair. The fake image pair and the real image pair are respectively input into the discriminator of the conditional generative adversarial network to obtain the discrimination result; perform cyclic training until the requirements are met.
2. The tunnel inspection image enhancement method based on conditional generative adversarial network according to claim 1, wherein The image enhancement model based on the conditional generative adversarial network includes a generator network and a discriminator network. The generator network is the NetG network, which includes a downsampling module, an upsampling module, and a skip connection module. A spatial attention mechanism is added to the downsampling module and the upsampling module, and the generator loss function includes a GAN adversarial loss function, an L1 loss function, and a reconstruction loss function guided by spatial attention.
3. The tunnel inspection image enhancement method based on conditional generative adversarial network according to claim 1, characterized in that, The discriminator network is the NetD network. The NetD network uses a PatchGAN discriminator. The discriminator divides the input image into multiple small blocks, discriminates each small block, and finally takes the average of the sum of the evaluations of all small blocks as the final evaluation result. The discriminator network consists of 3 convolutional layers, 1 normalization layer, and 1 spatial attention module. It extracts image features through convolution and normalization processing, converts the input image into a single-channel feature map, and then focuses on the spatial region of the image with key discriminative information through the spatial attention module.
4. The tunnel inspection image enhancement method based on conditional generative adversarial network according to claim 1, characterized in that, The downsampling module of the generator network consists of 1 convolutional layer, 1 activation function layer, and 1 normalization layer. The convolutional layer gradually reduces the spatial dimension of the input low-quality tunnel image while increasing the number of feature channels to extract the high-level features of the image and capture the semantic information and context information in the image; the upsampling module consists of 1 transposed convolutional layer, 1 activation function layer, 1 normalization layer, and a dropout layer. The upsampling module gradually restores the spatial dimension of the feature map to the size of the original image through transposed convolution while reducing the number of feature channels.
5. The tunnel inspection image enhancement method based on conditional generative adversarial network according to claim 4, characterized in that, The input of the downsampling module is a low-quality tunnel image, and the input of the upsampling module is the feature map output by the previous upsampling layer; the input of the skip connection module is the concatenated feature map output by the downsampling module. Two feature maps output after max pooling and average pooling are concatenated again to form a new feature map; a convolutional operation is performed on the new feature map to generate a spatial attention map; after the generated spatial attention map and the original input are weighted element by element according to different spatial positions, different weights are assigned to different regions, and feature fusion is performed. The fused feature map is used to output a quality-enhanced tunnel inspection image, that is, a false high-quality image.
6. The tunnel inspection image enhancement method based on conditional generative adversarial network according to claim 1, characterized in that, In the discriminator network, the spatial attention module is added after the second convolutional layer in the discriminator network, focusing on the spatial regions in the image that have a great impact on the discrimination result, including key feature regions with existing lighting anomalies, blurriness, and noise interference; the discriminator loss function uses the least squares loss function and the attention region feature difference loss function.
7. Tunnel inspection image enhancement system based on conditional generative adversarial network, characterized in that, Including: An image acquisition module for acquiring a tunnel image to be enhanced; An image enhancement module for inputting the tunnel image to be enhanced into a trained image enhancement model and outputting an enhanced optimized image; Among them, the image enhancement model is constructed based on a conditional generative adversarial network, and the specific training process is as follows: Obtain high-quality tunnel images under ideal conditions, and perform preprocessing on the high-quality images to obtain low-quality tunnel images; Extract feature maps of different channel numbers and different sizes of the low-quality tunnel image through the generator of the conditional generative adversarial network, and output corresponding quality-enhanced tunnel inspection images, denoted as false high-quality images, after weighted processing by the spatial attention mechanism; Concatenate the low-quality tunnel image and the high-quality tunnel image to form a real image pair, and concatenate the low-quality tunnel image pair and the false high-quality image to form a false image pair. The false image pair and the real image pair are respectively input into the discriminator of the conditional generative adversarial network to obtain discrimination results; perform cyclic training until the requirements are met.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the tunnel inspection image enhancement method based on a conditional generative adversarial network according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, it implements the tunnel inspection image enhancement method based on a conditional generative adversarial network according to any one of claims 1-6.
10. An electronic device, characterized in that, Including: A processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the tunnel inspection image enhancement method based on a conditional generative adversarial network according to any one of claims 1-6.
Citation Information
Patent Citations
Rain and fog image enhancement method under low illumination based on TF-GAN
CN117197016A
Low-illumination image enhancement method and system based on improved generative adversarial network
CN117593238A
Retinal vessel segmentation method based on feature enhancement and multi-scale perceptual feature fusion
CN119131047A
CT image denoising system and method
WO2022000183A1