Tunnel disease detection method based on illumination compensation and dynamic feature fusion
Through the method of illumination compensation and dynamic feature fusion, the problems of uneven illumination and feature aliasing in tunnel disease detection are solved, high-precision tunnel disease detection is achieved, and the robustness and versatility of the detection system are improved.
Patent Information
- Application Number
- CN202510816423.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies for tunnel defect detection suffer from loss of defect details and feature aliasing due to uneven illumination and fixed fusion strategies, making it difficult to achieve high-precision, low-miss detection detection.
A method based on illumination compensation and dynamic feature fusion is adopted. Illumination compensation is performed through Retinex decomposition and adaptive gamma correction modules. Combined with the hierarchical attention fusion module and the adaptive spatial pyramid fast pooling module, multi-scale feature fusion weights are dynamically scheduled to improve the accuracy and robustness of disease detection.
It effectively alleviates the problem of feature aliasing caused by uneven lighting and background interference, improves the detection accuracy of minor tunnel defects and the robustness of the detection system, and realizes the accurate identification and positioning of defect areas of multiple types and sizes.
Smart Images

Figure CN120673042A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and image detection methods, and particularly relates to a tunnel disease detection method based on illumination compensation and dynamic feature fusion. Background Art
[0002] Tunnels are a crucial component of transportation infrastructure, and their structural health is directly linked to operational safety and service life. Currently, tunnel defect detection primarily relies on manual inspections or deep learning-based automatic identification technologies. While manual inspections can provide relatively intuitive structural status information, they suffer from low efficiency and high risks to personnel safety. Conventional deep learning methods, while improving the ability to perceive objects of varying scales through multi-layer convolutional networks and feature pyramids, remain insufficient for tunnels, with their complex lighting and changing environments.
[0003] Inside tunnels, due to interference from factors such as interlaced light intensity, humidity reflections, and obstructions, minute defects like crack edges and spalling patches often appear lost or overly dull in both dark and bright areas. Furthermore, traditional fixed splicing and fusion strategies easily blend defect features with the complex background, leading to high rates of missed detections and false positives. Therefore, a new detection framework is urgently needed that can adaptively compensate for illumination variations at the pixel level and dynamically adjust multi-scale feature fusion weights in both channel and spatial dimensions. This framework can address the challenges of uneven illumination and feature blending, achieving high-precision, low-miss detection detection of minute tunnel defects. Summary of the Invention
[0004] The purpose of the present invention is to provide a tunnel defect detection method based on illumination compensation and dynamic feature fusion, which solves the problems of defect detail loss and feature aliasing caused by uneven illumination and fixed fusion strategies in the prior art.
[0005] The technical solution adopted by the present invention is a tunnel disease detection method based on illumination compensation and dynamic feature fusion, which is specifically implemented in the following steps:
[0006] Step 1: Data preprocessing;
[0007] Step 2: Construct a lighting adaptive compensation module;
[0008] Step 3: Construct a hierarchical attention fusion module;
[0009] Step 4: Construct an adaptive spatial pyramid fast pooling module.
[0010] The present invention is also characterized in that:
[0011] Step 1 is implemented as follows:
[0012] Step 1.1: Collect images of the tunnel interior surface from the image capture device and perform pixel-level annotation of cracks, spalling, and water damage areas to construct a labeled tunnel disease training set;
[0013] Step 1.2: perform data augmentation on the labeled images to expand sample diversity. The transformed image set is denoted as {I i};
[0014] Step 1.3, each image I i Normalization is performed by channel, and the calculation formula is shown as follows:
[0015]
[0016] Where μ c and σ c are the mean and standard deviation of channel c in the training set respectively;
[0017] Step 1.4: Input the normalized image into the lightweight backbone network and extract the feature map by level. Where s represents the sth layer, H s , W s , C s are the height, width, and number of channels of the layer respectively.
[0018] Step 2 is implemented as follows:
[0019] In step 2.1, perform Retinex decomposition and input the normalized image into a lightweight neural network to generate an illumination map. Then, perform pixel-by-pixel division on the original image to obtain a reflectance map, thus separating the illumination and reflectance components.
[0020] Step 2.2, perform adaptive gamma correction, predict the gamma parameter map through a small network, and perform pixel-by-pixel power operation on the illumination map;
[0021] Step 2.3, enhancing the contrast of the reflection image and performing a stretching operation on the reflection image;
[0022] In step 2.4, the image is reconstructed and the channel structure is unified. The enhanced brightness map is fused with the reflectance map to output a feature map in a standard format.
[0023] Step 2.1 is implemented as follows:
[0024] Step 2.1.1, normalize I(x,y,z)∈R 3×W×3The input is fed into the Retinex encoder network, which first passes through a lightweight encoder network consisting of three layers of convolution. The first two layers have a 3×3 kernel size and 16 and 32 channels, respectively. The last layer uses a 1×1 convolution to generate a single-channel grayscale illumination map L(x,y). Each convolution layer is equipped with Batch Normalization and ReLU activation. The calculation formula is shown below:
[0025] L(x,y)=f θ (I)(x,y) (2)
[0026] In step 2.1.2, the illumination map generated in step 2.1.1 is divided pixel by pixel by the original image to obtain the reflected image. The reflected image retains the original image information that is independent of the illumination. The calculation formula is as follows:
[0027]
[0028] Here, ε is a very small constant to prevent division by zero.
[0029] Step 2.2 is implemented as follows:
[0030] In step 2.2.1, extract the global statistical features of the illumination map and input them into a lightweight neural network to predict the gamma value of each pixel. The illumination map L(x, y) is first subjected to global average pooling to obtain the brightness vector g, which is then input into a multi-layer perceptron to obtain the gamma map. The calculation formula is as follows:
[0031] γ(x,y)=MLP(GAP(L)) (4)
[0032] In step 2.2.2, perform gamma transformation on each pixel of the illumination map to adjust the brightness response and generate a corrected illumination map. The calculation formula is as follows:
[0033] L′(x,y)=[L(x,y)] γ(x,y) (5).
[0034] Step 2.3 is implemented as follows:
[0035] In step 2.3.1, perform minimum and maximum calculations on each channel of the reflectance image to determine the stretching interval. For each channel c, the calculation formula is as follows:
[0036] Min c =Min x,y R(x,y,c) (6)
[0037] Max c =Max x,y R(x,y,c) (7)
[0038] In step 2.3.2, each pixel of the reflectance image is mapped to the standard contrast interval to generate the enhanced reflectance image. The calculation formula is as follows:
[0039]
[0040] The step 2.4 is specifically implemented according to the following steps:
[0041] In step 2.4.1, the gamma-corrected luminance map L′(x, y) is multiplied pixel by pixel with the enhanced reflectance map R′(x, y, c) to reconstruct the illumination-compensated image. The calculation formula is as follows:
[0042] I enh (x,y,c)=R′(x,y,c)·L′(x,y) (9)
[0043] Step 2.4.2, enhance the image I enh Input 1×1 convolution layer for channel mapping output, unify the number of output channels, follow convolution with BN and ReLU, and output standard feature map F 0 , the calculation formula is as follows:
[0044] F 0 (x,y,:)=ReLU(W·I enh (x,y,:)+b) (10)
[0045] Among them, W and b are the convolution weight and bias respectively, and the output is the input of the subsequent feature extraction module.
[0046] Step 3 is implemented as follows:
[0047] Step 3.1: Apply channel attention, calculate channel weights for each layer of feature maps, and guide the network to focus on important semantic channels;
[0048] Step 3.2: Apply spatial attention to generate a response map for the spatial dimension, highlighting the location distribution of the target area;
[0049] In step 3.3, multiple layers of feature maps are fused, upsampled and weighted merged in sequence to achieve cross-scale semantic alignment.
[0050] Step 3.1 is implemented as follows:
[0051] Step 3.1.1, input feature map F s After global average pooling, in the spatial dimension H s ×W s Aggregate on it to generate the channel description vector g s , the vector is then fed into a one-dimensional convolution module with a kernel size of 1, a stride of 1, and an output channel number of C s, and normalized using the Sigmoid function to obtain the channel attention weight vector, input feature map After pooling, the channel is generated, and the calculation formula is as follows:
[0052]
[0053] W s =σ(Conv1D(g s )) (12)
[0054] Step 3.1.2, the attention weight W s Broadcasting is extended to the spatial dimension, and we get the same as F s The attention map of the same shape is then multiplied by channel by channel, multiplied with the input feature map, and the output channel weighted feature map The calculation formula is as follows:
[0055] F′ s (x,y,c)=W s (c) F s (x,y,c) (13)
[0056] The step 3.2 is specifically implemented according to the following steps:
[0057] Step 3.2.1, input channel weighted feature map Perform maximum pooling and average pooling operations along the channel dimension to generate two single-channel spatial maps F max ,F avg , stitch the two images in the channel dimension to get The concatenated feature map is fed into a two-dimensional convolutional layer with a kernel size of 7×7, a stride of 1, and a padding of 3. The size is kept constant, and the convolution output is activated with Sigmoid to generate a spatial attention map. The calculation formula is as follows:
[0058] M s =σ(Conv 7×7 (Maxpool(F′ s ),Avgpool(F′ s ))) (14)
[0059] Step 3.2.2: The spatial attention map is copied to a channel with C through dimension expansion. s , and F′ s Perform pixel-by-pixel and channel-by-channel multiplication to generate the spatial enhancement feature map F″ s The feature map is then fed into a convolutional layer with a kernel size of 3×3, a stride of 1, and a padding of 1. The output channel is C s, convolution is followed by Batch Normalization and ReLU activation function, and the output size remains unchanged for subsequent multi-layer fusion operations. The calculation formula is as follows:
[0060] F" s (x,y,c)=M s (x,y)·F′ s (x,y,c) (15)
[0061] The step 3.3 is specifically implemented according to the following steps:
[0062] Step 3.3.1: The previous level attention feature map F″ s+1 Upsample to the current scale through bilinear interpolation and then compare it with the current layer feature map F″ s At the same resolution, element-by-element addition is performed to generate a fused feature map. The fused result is input into a convolution layer with a kernel size of 3×3, a stride of 1, and a padding of 1 for feature extraction. The convolution is followed by BN and ReLU, and the output size is The calculation formula is as follows:
[0063]
[0064] Step 3.3.2: Repeat the upsampling, addition fusion and convolution operations in step 3.3.1 to gradually align the high-level semantic features to the shallow layers, and finally obtain the fused feature map F at the bottom layer. HA ∈R H×W×C Used for subsequent spatial pyramid pooling module processing, the calculation formula is as follows:
[0065]
[0066] Step 4 is implemented as follows:
[0067] Step 4.1: predict the pooling kernel size, perform global modeling on the fused feature map, and generate multi-scale pooling kernel parameters;
[0068] Step 4.2: Perform multi-scale pooling, performing pooling operations at the original and downsampled resolutions to obtain multi-receptive field feature maps;
[0069] Step 4.3: Apply channel attention, calculate the channel weight for each pooled feature, and then perform weighted fusion;
[0070] In step 4.4, multi-scale features are fused and the final semantic feature map is output for use by the detection head.
[0071] Step 4.1 is as follows: Input fusion feature map F HA ∈R H×W×CFirst, the global channel vector g∈R is extracted by GlobalAverage Pooling 1×1×C , the vector is then input into MLP to predict the multi-scale kernel size. MLP contains two fully connected layers. The output dimension of the first layer is 64, and the second layer outputs 3 pooling scale parameters, corresponding to the kernel sizes k1, k2, and k3 used subsequently. The calculation formula is as follows:
[0072] [k1,k2,k3]=MLP(GAP(F HA )) (18)
[0073] The step 4.2 is specifically as follows: fusion feature map F HA The first channel is processed directly on the original size using the maximum pooling layer with kernel size k1 and k2, with a stride of 1 and padding of Keep the spatial size unchanged; the second feature map is first downsampled through a 2×2 convolution layer, with the channel unchanged, and then the maximum pooling operation with a kernel size of k3 is used for spatial compression to obtain three pooled feature maps. The calculation formula is as follows:
[0074]
[0075] Among them, the Downsample operation is a convolution with a kernel size of 2×2 and a stride of 2;
[0076] The step 4.3 is specifically implemented according to the following steps:
[0077] Step 4.3.1, for each pooled feature map P i Perform global average pooling operation to obtain the channel description vector g i ∈R 1×1×C ,Then input the shared MLP module, which contains a 1×1,convolution and a Sigmoid activation function, and outputs the channel attention coefficient w i , the calculation formula is as follows:
[0078] w i =σ(MLP(GAP(P i )) (twenty two)
[0079] Step 4.3.2, the attention coefficient w i Multiply the corresponding pooling feature map P channel by channel i , generate weighted feature map P′ i , the calculation formula is as follows:
[0080] P′ i (x,y,c)=w i (c)·P i (x,y,c) (23)
[0081] The step 4.4 is specifically implemented according to the following steps:
[0082] Step 4.4.1: Perform spatial size alignment on the three pooled feature maps after channel weighting. If there is a P′3 with inconsistent scale, use bilinear interpolation to upsample it to H×W size, and then use F HA , P′1, P′2, Upsample(P′3) are spliced along the channel dimension to generate a multi-scale fusion feature map. The calculation formula is as follows:
[0083] F concat =Concat(F HA ,P′1,P′2,Upsample(P′3)) (24)
[0084] In step 4.4.2, the fused feature map is fed into a 3×3 convolutional layer with the number of output channels set to C. The convolution is followed by Batch Normalization and ReLU activation functions to unify the channel dimensions and obtain the final output feature map as the input of the detection head module. The calculation formula is as follows:
[0085] F out =Conv 3×3 (F concat ) (25).
[0086] The beneficial effects of the present invention are:
[0087] This tunnel defect detection method, based on illumination compensation and dynamic feature fusion, introduces an adaptive illumination compensation mechanism and a multi-dimensional dynamic fusion structure during model training, effectively enhancing the feature recognition capability of defect targets in complex illumination and multi-scale backgrounds. By introducing Retinex decomposition and an adaptive gamma correction module, unified processing of strong and weak light areas is achieved, reducing the interference of illumination differences on image structural features. Furthermore, a hierarchical attention fusion module is constructed to dynamically model the contextual relationships between semantics at different scales, improving the fusion quality of multi-layer features. Furthermore, the proposed adaptive spatial pyramid fast pooling module, combined with dynamic kernel scale prediction and channel attention scheduling mechanisms, enhances the model's responsiveness to structural defects such as large spans and fine cracks while maintaining the compactness of the feature map. This method can effectively alleviate the feature aliasing problem caused by uneven illumination, background interference, and scale mismatch in tunnel defect detection, enabling accurate identification and location of defect areas of multiple types and sizes, and improving the robustness and versatility of the detection system in actual engineering scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 This is a network architecture diagram of the tunnel disease detection method based on illumination compensation and dynamic feature fusion of the present invention;
[0089] Figure 2 This is a network architecture diagram of the lighting adaptive compensation module in step 2 of the tunnel disease detection method based on illumination compensation and dynamic feature fusion of the present invention;
[0090] Figure 3 This is a network architecture diagram of the hierarchical attention fusion module in step 3 of the tunnel disease detection method based on illumination compensation and dynamic feature fusion of the present invention;
[0091] Figure 4 This is a network architecture diagram of the adaptive spatial pyramid fast pooling module in step 4 of the tunnel disease detection method based on illumination compensation and dynamic feature fusion of the present invention. DETAILED DESCRIPTION
[0092] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0093] The present invention is based on the tunnel disease detection method of illumination compensation and dynamic feature fusion, and its network architecture is shown in the figure below: Figure 1 As shown, the specific implementation steps are as follows:
[0094] Step 1: Data preprocessing;
[0095] Step 2: Construct a lighting adaptive compensation module;
[0096] Step 3: Construct a hierarchical attention fusion module;
[0097] Step 4: Construct an adaptive spatial pyramid fast pooling module.
[0098] Example 1
[0099] The present invention is based on a tunnel disease detection method that integrates illumination compensation and dynamic features, wherein step 1 is specifically implemented according to the following steps:
[0100] Step 1.1: Collect images of the tunnel interior surface from the image capture device and perform pixel-level annotation of cracks, spalling, and water damage areas to construct a labeled tunnel disease training set;
[0101] Step 1.2: Perform random horizontal / vertical flipping, rotation, scaling, and brightness perturbation data enhancement operations on the labeled images to expand sample diversity. The transformed image set is denoted as {I i};
[0102] Step 1.3, each image I i Normalization is performed by channel, and the calculation formula is shown as follows:
[0103]
[0104] Where μ c and σ c are the mean and standard deviation of channel c in the training set respectively;
[0105] Step 1.4: Input the normalized image into the lightweight backbone network and extract the feature map by level. Where s represents the sth layer, H s , W s , C s are the height, width, and number of channels of the layer respectively.
[0106] Example 2
[0107] The present invention is based on the tunnel disease detection method of the fusion of illumination compensation and dynamic features. The network architecture of the illumination adaptive compensation module is shown in the figure below: Figure 2 As shown, step 2 is specifically implemented according to the following steps:
[0108] Step 2 is implemented as follows:
[0109] In step 2.1, perform Retinex decomposition and input the normalized image into a lightweight neural network to generate an illumination map. Then, perform pixel-by-pixel division on the original image to obtain a reflectance map, thus separating the illumination and reflectance components.
[0110] Step 2.2, perform adaptive gamma correction, predict the gamma parameter map through a small network, and perform pixel-by-pixel power operation on the illumination map;
[0111] Step 2.3, enhancing the contrast of the reflection image and performing a stretching operation on the reflection image;
[0112] In step 2.4, the image is reconstructed and the channel structure is unified. The enhanced brightness map is fused with the reflectance map to output a feature map in a standard format.
[0113] Example 3
[0114] The present invention is based on a tunnel disease detection method that integrates illumination compensation and dynamic features, wherein step 2.1 is specifically implemented according to the following steps:
[0115] Step 2.1.1, normalize I(x,y,z)∈R H×W×3 The input is fed into the Retinex encoder network, which first passes through a lightweight encoder network consisting of three layers of convolution. The first two layers have a 3×3 kernel size and 16 and 32 channels, respectively. The last layer uses a 1×1 convolution to generate a single-channel grayscale illumination map L(x,y). Each convolution layer is equipped with Batch Normalization and ReLU activation. The calculation formula is shown below:
[0116] L(x,y)=fθ (I)(x,y) (2)
[0117] In step 2.1.2, the illumination map generated in step 2.1.1 is divided pixel by pixel by the original image to obtain the reflected image. The reflected image retains the original image information that is independent of the illumination. The calculation formula is as follows:
[0118]
[0119] Here, ε is a very small constant to prevent division by zero.
[0120] Step 2.2 is implemented as follows:
[0121] In step 2.2.1, extract the global statistical features of the illumination map and input them into a lightweight neural network to predict the gamma value of each pixel. The illumination map L(x, y) is first subjected to global average pooling to obtain the brightness vector g, which is then input into a multi-layer perceptron to obtain the gamma map. The calculation formula is as follows:
[0122] γ(x,y)=MLP(GAP(L)) (4)
[0123] In step 2.2.2, perform gamma transformation on each pixel of the illumination map to adjust the brightness response and generate a corrected illumination map. The calculation formula is as follows:
[0124] L′(x,y)=[L(x,y)] γ(x,y) (5).
[0125] Step 2.3 is implemented as follows:
[0126] In step 2.3.1, perform minimum and maximum calculations on each channel of the reflectance image to determine the stretching interval. For each channel c, the calculation formula is as follows:
[0127] Min c =Min x,y R(x,y,c) (6)
[0128] Max c =Max x,y R(x,y,c) (7)
[0129] In step 2.3.2, each pixel of the reflectance image is mapped to the standard contrast interval to generate the enhanced reflectance image. The calculation formula is as follows:
[0130]
[0131] Step 2.4 is implemented as follows:
[0132] In step 2.4.1, the gamma-corrected luminance map L′(x, y) is multiplied pixel by pixel with the enhanced reflectance map R′(x, y, c) to reconstruct the illumination-compensated image. The calculation formula is as follows:
[0133] I enh (x,y,c)=R′(x,y,c)·L′(x,y) (9)
[0134] Step 2.4.2, enhance the image I enh Input 1×1 convolution layer for channel mapping output, unify the number of output channels, follow convolution with BN and ReLU, and output standard feature map F 0 , the calculation formula is as follows:
[0135] F 0 (x,y,:)=ReLU(W·I enh (x,y,:)+b) (10)W and b are the convolution weight and bias respectively, and the output is the input of the subsequent feature extraction module.
[0136] Example 4
[0137] The present invention is based on the tunnel disease detection method of illumination compensation and dynamic feature fusion. The network architecture of the hierarchical attention fusion module is shown in the figure below: Figure 3 As shown, step 3 is specifically implemented as follows:
[0138] Step 3.1: Apply channel attention, calculate channel weights for each layer of feature maps, and guide the network to focus on important semantic channels;
[0139] Step 3.1 is implemented as follows:
[0140] Step 3.1.1, input feature map F s After global average pooling, in the spatial dimension H s ×W s Aggregate on it to generate the channel description vector g s , the vector is then fed into a one-dimensional convolution module with a kernel size of 1, a stride of 1, and an output channel number of C s , and normalized using the Sigmoid function to obtain the channel attention weight vector, input feature map After pooling, the channel is generated, and the calculation formula is as follows:
[0141]
[0142] W s =σ(Conv1D(g s )) (12)
[0143] Step 3.1.2, the attention weight W sBroadcasting is extended to the spatial dimension, and we get the same as F s The attention map of the same shape is then multiplied by channel by channel, multiplied with the input feature map, and the output channel weighted feature map The calculation formula is as follows:
[0144] F′ s (x,y,c)=W s (c) F s (x,y,c) (13);
[0145] Step 3.2: Apply spatial attention to generate a response map for the spatial dimension, highlighting the location distribution of the target area;
[0146] Step 3.2 is implemented as follows:
[0147] Step 3.2.1, input channel weighted feature map Perform maximum pooling and average pooling operations along the channel dimension to generate two single-channel spatial maps F max ,F avg , stitch the two images in the channel dimension to get The concatenated feature map is fed into a two-dimensional convolutional layer with a kernel size of 7×7, a stride of 1, and a padding of 3. The size is kept constant, and the convolution output is activated with Sigmoid to generate a spatial attention map. The calculation formula is as follows:
[0148] M s =σ(Conv 7×7 (Maxpool(F′ s ),Avgpool(F′ s ))) (14)
[0149] Step 3.2.2: The spatial attention map is copied to a channel with C through dimension expansion. s , and F′ s Perform pixel-by-pixel and channel-by-channel multiplication to generate the spatial enhancement feature map F″ s The feature map is then fed into a convolutional layer with a kernel size of 3×3, a stride of 1, and a padding of 1. The output channel is C s , convolution is followed by Batch Normalization and ReLU activation function, and the output size remains unchanged for subsequent multi-layer fusion operations. The calculation formula is as follows:
[0150] F″ s (x,y,c)=M s (x,y)·F′ s (x,y,c) (15);
[0151] Step 3.3: fuse multiple layers of feature maps, upsample and weighted merge them sequentially to achieve cross-scale semantic alignment;
[0152] Step 3.3 is implemented as follows:
[0153] Step 3.3.1: The previous level attention feature map F″ s+1 Upsample to the current scale through bilinear interpolation and then compare it with the current layer feature map F″ s At the same resolution, element-by-element addition is performed to generate a fused feature map. The fusion result is input into a convolution layer with a convolution kernel size of 3×3, a stride of 1, and a padding of 1 for feature extraction. The convolution is followed by BN and ReLU, and the output size is The calculation formula is as follows:
[0154]
[0155] Step 3.3.2: Repeat the upsampling, addition fusion and convolution operations in step 3.3.1 to gradually align the high-level semantic features to the shallow layers, and finally obtain the fused feature map F at the bottom layer. HA ∈R H×W×c Used for subsequent spatial pyramid pooling module processing, the calculation formula is as follows:
[0156]
[0157] Example 5
[0158] The present invention is based on the tunnel disease detection method of illumination compensation and dynamic feature fusion, and the network architecture of the adaptive spatial pyramid fast pooling module is shown in the figure below. Figure 4 As shown, step 4 is specifically implemented according to the following steps:
[0159] Step 4.1: predict the pooling kernel size, perform global modeling on the fused feature map, and generate multi-scale pooling kernel parameters;
[0160] Step 4.1 is as follows: Input fusion feature map F HA ∈R H×W×C First, the global channel vector g∈R is extracted by GlobalAverage Pooling 1×1×C , the vector is then input into MLP to predict the multi-scale kernel size. MLP contains two fully connected layers. The output dimension of the first layer is 64, and the second layer outputs 3 pooling scale parameters, corresponding to the kernel sizes k1, k2, and k3 used subsequently. The calculation formula is as follows:
[0161] [k1,k2,k3]=MLP(GAP(F HA )) (18);
[0162] Step 4.2: Perform multi-scale pooling, performing pooling operations at the original and downsampled resolutions to obtain multi-receptive field feature maps;
[0163] Step 4.2 is as follows: fusion feature map F HA The first channel is processed directly on the original size using the maximum pooling layer with kernel size k1 and k2, with a stride of 1 and padding of Keep the spatial size unchanged; the second feature map is first downsampled through a 2×2 convolution layer, with the channel unchanged, and then the maximum pooling operation with a kernel size of k3 is used for spatial compression to obtain three pooled feature maps. The calculation formula is as follows:
[0164]
[0165] Among them, the Downsample operation is a convolution with a kernel size of 2×2 and a stride of 2;
[0166] Step 4.3: Apply channel attention, calculate the channel weight for each pooled feature, and then perform weighted fusion;
[0167] Step 4.3 is implemented as follows:
[0168] Step 4.3.1, for each pooled feature map P i Perform global average pooling operation to obtain the channel description vector g i ∈R 1×1×C ,Then input the shared MLP module, which contains a 1×1,convolution and a Sigmoid activation function, and outputs the channel attention coefficient w i , the calculation formula is as follows:
[0169] w i =σ(MLP(GAP(P i )) (twenty two)
[0170] Step 4.3.2, the attention coefficient w i Multiply the corresponding pooling feature map P channel by channel i , generate weighted feature map P′ i , the calculation formula is as follows:
[0171] P′ i (x,y,c)=w i (c)·P i (x,y,c) (23);
[0172] Step 4.4: fuse multi-scale features and output the final semantic feature map for use by the detection head;
[0173] Step 4.4 is implemented as follows:
[0174] Step 4.4.1: Perform spatial size alignment on the three pooled feature maps after channel weighting. If there is a P′3 with inconsistent scale, use bilinear interpolation to upsample it to H×W size, and then use F HA , P′1, P′2, Upsample(P′3) are spliced along the channel dimension to generate a multi-scale fusion feature map. The calculation formula is as follows:
[0175] F concat =Concat(F HA ,P′1,P′2,Upsample(P′3)) (24)
[0176] In step 4.4.2, the fused feature map is fed into a 3×3 convolutional layer with the number of output channels set to C. The convolution is followed by Batch Normalization and ReLU activation functions to unify the channel dimensions and obtain the final output feature map as the input of the detection head module. The calculation formula is as follows:
[0177] F out =Conv 3×3 (F concat ) (25).
[0178] Example 6
[0179] The tunnel disease detection method based on illumination compensation and dynamic feature fusion of the present invention is used, and the results are compared with those of the traditional detection method as shown in Table 1 below:
[0180] Table 1
[0181] Model AP50 F1 score Param Retinanet 0.563 0.484 29.5M Deformable_detr 0.650 0.587 31.7M CE-FPN 0.571 0.477 26.2M DEYO 0.727 0.680 25.2M Yolov11 <![CDATA[ 0.798 ]]> <![CDATA[ 0.709 ]]> 18.7M Ours 0.872 0.786 <![CDATA[ 19.1M ]]>
[0182] AP50: Average precision. F1 score: F1 score. Param: Parameter count. The best and second-best scores are bold and underlined, respectively.
[0183] The network model trained by the method proposed in the present invention achieved high AP50 and F1 score values on the tunnel disease dataset using the detected disease results as input. In addition, while achieving high performance, the number of parameters was also effectively controlled.
Claims
1. A tunnel disease detection method based on illumination compensation and dynamic feature fusion, characterized in that: Please follow the steps below to implement: Step 1: Data preprocessing; Step 2: Construct an adaptive lighting compensation module; Step 3: Construct a hierarchical attention fusion module; Step 4: Construct an adaptive spatial pyramid fast pooling module.
2. The tunnel disease detection method based on illumination compensation and dynamic feature fusion according to claim 1 is characterized in that: The step 1 is specifically implemented according to the following steps: Step 1.1: Collect images of the tunnel interior surface from the image capture device and perform pixel-level annotation of cracks, spalling, and water damage areas to construct a labeled tunnel disease training set; Step 1.2: perform data augmentation on the labeled images to expand sample diversity. The transformed image set is denoted as {I i }; Step 1.3, each image I i Normalization is performed by channel, and the calculation formula is shown as follows: Where μ c and σ c are the mean and standard deviation of channel c in the training set respectively; Step 1.4: Input the normalized image into the lightweight backbone network and extract the feature map by level. Where s represents the sth layer, H s , W s , C s are the height, width, and number of channels of the layer respectively.
3. The tunnel disease detection method based on illumination compensation and dynamic feature fusion according to claim 1 is characterized in that: The step 2 is specifically implemented according to the following steps: In step 2.1, perform Retinex decomposition and input the normalized image into a lightweight neural network to generate an illumination map. Then, perform pixel-by-pixel division on the original image to obtain a reflectance map, thus separating the illumination and reflectance components. Step 2.2, perform adaptive gamma correction, predict the gamma parameter map through a small network, and perform pixel-by-pixel power operation on the illumination map; Step 2.3, enhancing the contrast of the reflection image and performing a stretching operation on the reflection image; In step 2.4, the image is reconstructed and the channel structure is unified. The enhanced brightness map is fused with the reflectance map to output a feature map in a standard format.
4. The tunnel disease detection method based on illumination compensation and dynamic feature fusion according to claim 3 is characterized in that: The step 2.1 is specifically implemented according to the following steps: Step 2.1.1, normalize I(x,y,z)∈R H×W×3 The input is fed into the Retinex encoder network, which first passes through a lightweight encoder network consisting of three layers of convolution. The first two layers have a 3×3 kernel size and 16 and 32 channels, respectively. The last layer uses a 1×1 convolution to generate a single-channel grayscale illumination map L(x,y). Each convolution layer is equipped with Batch Normalization and ReLU activation. The calculation formula is shown below: L(x,y)=f θ (I)(x,y) (2) In step 2.1.2, the illumination map generated in step 2.1.1 is divided pixel by pixel by the original image to obtain the reflected image. The reflected image retains the original image information that is independent of the illumination. The calculation formula is as follows: Here, ε is a very small constant to prevent division by zero.
5. The tunnel disease detection method based on illumination compensation and dynamic feature fusion according to claim 3 is characterized in that: The step 2.2 is specifically implemented according to the following steps: In step 2.2.1, extract the global statistical features of the illumination map and input them into a lightweight neural network to predict the gamma value of each pixel. The illumination map L(x, y) is first subjected to global average pooling to obtain the brightness vector g, which is then input into a multi-layer perceptron to obtain the gamma map. The calculation formula is as follows: γ(x,y)=MLP(GAP(L)) (4) In step 2.2.2, perform gamma transformation on each pixel of the illumination map to adjust the brightness response and generate a corrected illumination map. The calculation formula is as follows: L(x,y)=[L(x,y)] γ(x,y) (5) 6. The tunnel disease detection method based on illumination compensation and dynamic feature fusion according to claim 3 is characterized in that: The step 2.3 is specifically implemented according to the following steps: In step 2.3.1, perform minimum and maximum calculations on each channel of the reflectance image to determine the stretching interval. For each channel c, the calculation formula is as follows: My c =My x,y R(x,y,c) (6) Max c =Max x,y R(x,y,c) (7) In step 2.3.2, each pixel of the reflectance image is mapped to the standard contrast interval to generate the enhanced reflectance image. The calculation formula is as follows: The step 2.4 is specifically implemented according to the following steps: Step 2.4.1, the gamma-corrected luminance map L ′ (x,y) and enhanced reflection map R ′ Multiply (x, y, c) pixel by pixel to reconstruct the illumination compensation image. The calculation formula is as follows: I enh (x,y,c)=R′(x,y,c)·L′(x,y) (9) Step 2.4.2, enhance the image I enh Input 1×1 convolution layer for channel mapping output, unify the number of output channels, follow convolution with BN and ReLU, and output standard feature map F 0 , the calculation formula is as follows: F 0 (x,y,:)=ReLU(W·I enh (x,y,:)+b) (10) Among them, W and b are the convolution weight and bias respectively, and the output is the input of the subsequent feature extraction module.
7. The tunnel defect detection method based on illumination compensation and dynamic feature fusion according to claim 1 is characterized in that: The step 3 is specifically implemented as follows: Step 3.1: Apply channel attention, calculate channel weights for each layer of feature maps, and guide the network to focus on important semantic channels; Step 3.2: Apply spatial attention to generate a response map for the spatial dimension, highlighting the location distribution of the target area; In step 3.3, multiple layers of feature maps are fused, upsampled and weighted merged in sequence to achieve cross-scale semantic alignment.
8. The tunnel disease detection method based on illumination compensation and dynamic feature fusion according to claim 7 is characterized in that: The step 3.1 is specifically implemented as follows: Step 3.1.1, input feature map F s After global average pooling, in the spatial dimension H s ×W s Aggregate on it to generate the channel description vector g s , the vector is then fed into a one-dimensional convolution module with a kernel size of 1, a stride of 1, and an output channel number of C s , and normalized using the Sigmoid function to obtain the channel attention weight vector, input feature map After pooling, the channel is generated, and the calculation formula is as follows: IN s =σ(Conv1D(g s )) (12) Step 3.1.2, the attention weight W s Broadcasting is extended to the spatial dimension, and we get the same as F s The attention map of the same shape is then multiplied by channel by channel, multiplied with the input feature map, and the output channel weighted feature map The calculation formula is as follows: F s ′(x,y,c)=W s (c)·F s (x,y,c) (13) The step 3.2 is specifically implemented according to the following steps: Step 3.2.1, input channel weighted feature map Perform maximum pooling and average pooling operations along the channel dimension to generate two single-channel spatial maps F max ,F avg , stitch the two images in the channel dimension to get The concatenated feature map is fed into a two-dimensional convolutional layer with a kernel size of 7×7, a stride of 1, and a padding of 3. The size is kept constant, and the convolution output is activated with Sigmoid to generate a spatial attention map. The calculation formula is as follows: M s =σ(Conv 7×7 (Maxpool(F′ s ),AvgPool(F′ s ))) (14) Step 3.2.2: The spatial attention map is copied to a channel with C through dimension expansion. s , and F′ s Perform pixel-by-pixel and channel-by-channel multiplication to generate the spatial enhancement feature map F″ s The feature map is then fed into a convolutional layer with a kernel size of 3×3, a stride of 1, and a padding of 1. The output channel is C S ,, the convolution is followed by Batch Normalization and ReLU activation function, and the output size remains unchanged for subsequent multi-layer fusion operations. The calculation formula is as follows: F″ s (x,y,c)=M s (x,y)·F′ s (x,y,c) (15) The step 3.3 is specifically implemented according to the following steps: Step 3.3.1: The previous level attention feature map F″ s+1 Upsample to the current scale through bilinear interpolation and then compare it with the current layer feature map F″ s At the same resolution, element-by-element addition is performed to generate a fused feature map. The fusion result is input into a convolution layer with a convolution kernel size of 3×3, a stride of 1, and a padding of 1 for feature extraction. The convolution is followed by BN and ReLU, and the output size is The calculation formula is as follows: Step 3.3.2: Repeat the upsampling, addition fusion and convolution operations in step 3.3.1 to gradually align the high-level semantic features to the shallow layers, and finally obtain the fused feature map F at the bottom layer. HA ∈R H×W×C Used for subsequent spatial pyramid pooling module processing, the calculation formula is as follows:
9. The tunnel disease detection method based on illumination compensation and dynamic feature fusion according to claim 1 is characterized in that: The step 4 is specifically implemented according to the following steps: Step 4.1: predict the pooling kernel size, perform global modeling on the fused feature map, and generate multi-scale pooling kernel parameters; Step 4.2: Perform multi-scale pooling, performing pooling operations at the original and downsampled resolutions to obtain multi-receptive field feature maps; Step 4.3: Apply channel attention, calculate the channel weight for each pooled feature, and then perform weighted fusion; In step 4.4, multi-scale features are fused and the final semantic feature map is output for use by the detection head.
10. The tunnel disease detection method based on illumination compensation and dynamic feature fusion according to claim 9 is characterized in that: The step 4.1 is specifically as follows: input the fusion feature map F HA ∈R H×W×C First, the global channel vector g∈R is extracted by Global Average Pooling 1×1×c , the vector is then input into MLP to predict the multi-scale kernel size. MLP contains two fully connected layers. The output dimension of the first layer is 64, and the second layer outputs 3 pooling scale parameters, corresponding to the kernel sizes k1, k2, and k3 used subsequently. The calculation formula is as follows: [k1,k2,k3]=MLP(GAP(F HA )) (18) The step 4.2 is specifically as follows: fusion feature map F HA The first channel is processed directly on the original size using the maximum pooling layer with kernel size k1 and k2, with a stride of 1 and padding of Keep the spatial size unchanged; the second feature map is first downsampled through a 2×2 convolution layer, with the channel unchanged, and then the maximum pooling operation with a kernel size of k3 is used for spatial compression to obtain three pooled feature maps. The calculation formula is as follows: Among them, the Downsample operation is a convolution with a kernel size of 2×2 and a stride of 2; The step 4.3 is specifically implemented according to the following steps: Step 4.3.1, for each pooled feature map P i Perform global average pooling operation to obtain the channel description vector g i ∈R 1 ×1×C ,Then input the shared MLP module, which contains a 1×1,convolution and a Sigmoid activation function, and outputs the channel attention coefficient w i , the calculation formula is as follows: w i σ(MLP(GAP(P i )) (22) Step 4.3.2, the attention coefficient w i Multiply the corresponding pooling feature map P channel by channel i , generate weighted feature map P′ i , the calculation formula is as follows: P′ i (x,y,c)=w i (c)·P i (x,y,c) (23) The step 4.4 is specifically implemented according to the following steps: Step 4.4.1: Perform spatial size alignment on the three pooled feature maps after channel weighting. If there are P3 with inconsistent scales, ′ , then use bilinear interpolation to upsample it to H×W size, and then F HA , P′1, P′2, Upsample(P′3) are spliced along the channel dimension to generate a multi-scale fusion feature map. The calculation formula is as follows: F concat =Concat(F HA ,P′1,P′2,Upsample(P′3)) (24) In step 4.4.2, the fused feature map is fed into a 3×3 convolutional layer with the number of output channels set to C. The convolution is followed by Batch Normalization and ReLU activation functions to unify the channel dimensions and obtain the final output feature map as the input of the detection head module. The calculation formula is as follows: F out =Conv 3×3 (F concat ) (25)。
Citation Information
Cited By
Hyperspectral image classification method and device and storage medium
CN122223461A