Tunnel-based multi-scale adaptive low-illumination image enhancement algorithm

Through the image enhancement algorithm of U-shaped network and multi-dimensional attention mechanism, the problem of poor low-light image quality in tunnel construction is solved, the image brightness and details are improved, and the performance of the intelligent monitoring system is improved.

CN120634877APending Publication Date: 2025-09-12XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510740578.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

During tunnel construction, images taken in low-light environments have poor quality, with problems such as insufficient brightness, noise interference, and blurred details, which affect the accuracy and efficiency of image processing and intelligent monitoring systems.

Method used

A multi-scale adaptive low-light image enhancement algorithm based on a U-shaped network is adopted. Image features are extracted through a multi-scale fusion module and a multi-dimensional attention mechanism. Image enhancement is performed in combination with an illumination estimation module to generate high-quality brightness and detail texture information.

Benefits of technology

It improves the brightness and detailed texture information of the image, enhances the image quality, improves the image processing effect, and enhances the accuracy and efficiency of the intelligent monitoring system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634877A_ABST
    Figure CN120634877A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale adaptive low-light image enhancement algorithm based on a tunnel, and the algorithm comprises the following steps: S1, collecting a low-light image of a tunnel face constructed in the tunnel, and making a data set to facilitate the enhancement of the low-light image; s2, performing feature extraction on the low-illumination image, and obtaining texture detail information of the tunnel face image by fusing feature information of multiple scales; s3, using a multi-dimensional attention mechanism to further extract space and channel information of different dimensions, and generating an attention feature map including feature information, including color and brightness information of an original map; and S4, combining the generated attention feature map with an adaptive illumination estimation module to obtain the illumination weight of the image, fusing the extracted feature information and the illumination factor to predict an intermediate control map layer, and combining the original image to obtain an enhanced image result. According to the method, the enhancement effect of the low-illumination image in the tunnel is remarkably improved, and the high-quality enhanced image with the original resolution can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and in particular is a tunnel-based multi-scale adaptive low-light image enhancement algorithm. Background Art

[0002] The continuous advancement of intelligent construction technology has placed higher demands on image processing for underground tunnel projects. During tunnel construction, low light levels, strong noise, and uneven brightness at tunnel faces pose challenges to image processing. Due to the complete lack of natural light deep within the tunnel, artificial lighting is required. However, lighting equipment often has limited power, resulting in insufficient and uneven lighting, which in turn leads to poor image quality. The captured images are often dark overall, with blurred details and difficulty distinguishing cracks and textures in the rock mass. Furthermore, there is significant noise, resulting in a grainy image that affects subsequent analysis. Furthermore, uneven light distribution, with some areas being too bright or too dark, affects the judgment of computer vision algorithms. These issues not only increase the difficulty of manual identification but also affect the accuracy of intelligent monitoring systems, posing challenges to construction safety and efficiency.

[0003] Image enhancement in low-light environments has always been a challenging problem in the field of computer vision. Due to insufficient lighting conditions, captured images often suffer from severe brightness attenuation, noise interference, and loss of detail. To address these issues, researchers have proposed a variety of effective solutions. Traditional methods are primarily based on image processing techniques. For example, improved histogram equalization algorithms can effectively expand the dynamic range of images, while homomorphic filtering can simultaneously enhance contrast and suppress noise. In recent years, the introduction of deep learning technology has brought breakthroughs in this field. By learning the mapping relationship between low-light images and normal-light images using large amounts of data, it has achieved remarkable results in detail restoration and noise suppression. These technologies further provide reliable technical support for practical application scenarios such as security monitoring and medical image analysis. Therefore, in-depth research on low-light image enhancement is of great significance.

[0004] Existing methods have achieved certain results in improving brightness, but the enhanced results are prone to problems such as insufficient brightness, low contrast, blurring, and artifacts. In addition, noise processing is not adequately considered, resulting in certain limitations in practical applications of the enhanced images and affecting the robustness of downstream visual tasks. Summary of the Invention

[0005] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a tunnel-based multi-scale adaptive low-light image enhancement algorithm, which can not only improve the brightness of the image, but also make the enhanced image detail texture information richer.

[0006] In order to achieve the above object, the solution of the present invention is:

[0007] Step S1: collect low-light images of the tunnel face under construction inside the tunnel and create a data set to enhance the low-light images.

[0008] Step S2: extract features from the low-light image and obtain texture detail information of the tunnel face image by fusing feature information at multiple scales.

[0009] In step S3, a multi-dimensional attention mechanism is used to further extract spatial and channel information of different dimensions, and generate an attention feature map containing feature information, including the color and brightness information of the original image.

[0010] In step S4, the generated attention feature map is combined with the adaptive illumination estimation module to obtain the illumination weight of the image, the extracted feature information and the illumination factor are fused to predict the intermediate control layer, and the enhanced image result is obtained by combining the original image.

[0011] Compared with the prior art, the present invention has the following beneficial effects:

[0012] The first improvement of the present invention is to utilize the feature extraction capability of the U-shaped network, and through jump connections, it can better extract detailed texture information in low-light images, design a multi-scale fusion module to extract feature maps at different levels, so that the network can capture feature information at different scales and perform multi-scale feature information learning.

[0013] The second improvement of this invention lies in the design of a multi-dimensional attention mechanism module, which considers both spatial and channel dimensions and fuses them to produce a fused feature. It can weight both texture details and channels, cascading mean pooling operations in different directions to obtain local information of the image. Through the full connection of the four branches, channel weighting is achieved, and the final output feature map contains both local and channel information.

[0014] The third improvement of the present invention lies in the design of an illumination estimation module and a predicted intermediate control layer. The generated attention feature map is combined with the adaptive illumination estimation module to obtain the illumination weight of the image, and the intermediate control layer is predicted, thereby realizing the joint control of brightness and structure, and guiding the network to generate high-quality brightness-enhanced images. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 Schematic diagram of the structure of the low-light image enhancement algorithm used in the embodiment of the present invention.

[0016] Figure 2 Schematic diagram of a multi-scale fusion module used in an embodiment of the present invention.

[0017] Figure 3Schematic diagram of the multi-dimensional attention mechanism module used in the embodiment of the present invention.

[0018] Figure 4 Schematic diagram of an illumination estimation module used in an embodiment of the present invention.

[0019] Figure 5 Schematic diagram of the prediction intermediate control layer used in the embodiment of the present invention.

[0020] Figure 6 FIG. 2 is a partial schematic diagram of a data set used in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] To explain the technical solution of the present invention, the present invention is described in detail below with reference to specific embodiments and the accompanying drawings.

[0022] like Figure 1 As shown in the figure, the present invention designs a tunnel-based multi-scale adaptive low-light image enhancement algorithm. It includes the following steps:

[0023] Step S1: collect low-light images of the tunnel face under construction inside the tunnel and create a data set to enhance the low-light images.

[0024] On-site construction personnel use technical means to implement quality control of tunnel excavation, and take pictures and archive them every 10 meters along the excavation face. Due to different image acquisition hardware, the image resolution quality of the collected images is inconsistent. Images with various pixel specifications are used as the low-light image dataset of the tunnel face, such as Figure 6 shown.

[0025] In step S2, the training set is imported into the image enhancement network, and features are extracted from the low-light image. By fusing feature information at multiple scales, texture detail information of the tunnel face image is obtained.

[0026] The original image is transferred to the encoder of the U-shaped network, and feature information is extracted by reducing the spatial resolution of the image layer by layer. Each level in the encoder consists of multiple convolution blocks and downsampling operations to reduce the size of the feature map.

[0027] Among them, the input of the feature extraction module is input, which includes the number of feature channels and the height and width of the input feature map. First, the input feature input undergoes convolution and pooling operations to obtain the feature map F0.

[0028] F0=Max_Pool(Conv(input))

[0029] F1=Max_Pool(Conv(F0))

[0030] F2=Max_Pool(Conv(F1))

[0031] F3=Max_Pool(Conv(F2))

[0032] Where Conv is the convolution module, Max_Pool is the pooling module, F1, F2, and F3 represent the first, second, and third levels of the encoder stage, respectively, which reduce the spatial resolution of the feature map layer by layer to obtain higher-level feature information.

[0033] The downsampled feature map is upsampled, and the low-resolution feature map is gradually restored to the size of the input image, and jump-connected and feature-fused with the encoder features at the same level.

[0034] F4=deConv(F3)

[0035] Concat1=Cat(F4,F2)

[0036] F5=deConv(F4)

[0037] Concat2=Cat(F5,F1)

[0038] F6=deConv(F5)

[0039] Concat3=Cat(F6,F0)

[0040] F out =sigmoid(deConv(F6))

[0041] Where: Cat represents the cascade module, deConv represents the upsampling module, sigmoid represents the activation function, F4, F5, F6 and F out They represent the upsampling feature map results of different levels respectively, and Concat1, Concat2 and Concat3 represent the feature maps corresponding to the skip connections respectively.

[0042] Skip connections better combine shallow features with deep features, and better help the network learn global and local features.

[0043] A multi-scale fusion module is designed to extract feature maps at different levels, enabling the network to capture feature information at different scales and perform multi-scale feature information learning. The process is as follows:

[0044] It is divided into two channels, and three scales, Concat2, Concat3 and F6, are extracted for processing. A 3×3 convolutional layer is used to extract local spatial features and capture texture details and edge information in the tunnel face image.

[0045] The 1×1 convolution layer changes the number of channels in the feature map to achieve feature reorganization and dimensionality reduction in the channel dimension.

[0046] Y=ω 1×1 (ReLU(BN(ω 3×3 (X))))

[0047] Where: Y is the output result, ω 1×1 and ω 3×3 Represent the weight matrices of 3×3 and 1×1 convolution kernels respectively.

[0048] In step S3, a multi-dimensional attention mechanism is used to further extract spatial and channel information of different dimensions, and generate an attention feature map containing feature information, including the color and brightness information of the original image.

[0049] Design a multi-dimensional attention mechanism module, considering the spatial dimension and channel dimension respectively.

[0050] First, the horizontal and vertical feature vectors of the input features are extracted, a cascade operation is performed in the spatial dimension, and a 1×1 convolution is used to complete the channel compression. The spatial attention map of the image is obtained through the activation function.

[0051] Here, two pooling kernels (H, 1) or (1, W) with two spatial ranges are used to represent the horizontal channel and the vertical channel respectively. The two pooling kernel formulas under the z-th channel are:

[0052]

[0053]

[0054] Where: D z (i, j) represents the pixel value at the i-th row and j-th column position on the z-th channel of the input feature map, and W and H represent the width and height of the feature map respectively.

[0055] A multi-branch structure is used to process the feature vector obtained by the coordinate attention mechanism, and the feature vector is input into multiple fully connected layers arranged in parallel. It is combined with the activation function to generate channel weights, and finally an attention feature map is generated.

[0056] In step S4, the generated attention feature map is combined with the adaptive illumination estimation module to obtain the illumination weight of the image, the extracted feature information and the illumination factor are fused to predict the intermediate control layer, and the enhanced image result is obtained by combining the original image.

[0057] The prediction intermediate control layer is designed, and the extracted feature map is first processed by two 3×3 convolution kernels and a sigmoid activation function.

[0058] The channel attention mechanism is introduced to obtain the channel weight map through global average pooling, fully connected layer and sigmoid activation function to predict the intermediate control layer.

[0059] The lighting estimation module is designed, which consists of two lighting adjustment modules. It calculates the brightness of each pixel in the low-light image and adjusts it in combination with the constraint function.

[0060]

[0061]

[0062] Where: L is the input low-light image, F is the intermediate control layer predicted by the network, and C B and C D It is the enhanced image result obtained after fusion.

[0063] Construct the total loss function by the illumination loss function L G , image structure loss L ssim , color consistency loss L c , texture detail loss L p constitute.

[0064] L total =λL G ++λ1L ssim +λ2L c +λ3L p

[0065]

[0066] Where: represents the gradient operator, c is the balance coefficient that controls the consistency of the structural strength, and constrains the illumination component by reducing excessive changes between gradients.

[0067]

[0068]

[0069] Where: M represents the texture feature image generated by the model, and A represents the texture feature of the original image.

Claims

1. A tunnel-based multi-scale adaptive low-light image enhancement algorithm, characterized by: The steps include: Step S1: Collect low-light images of the tunnel face under construction inside the tunnel and create a data set to enhance the low-light images. Step S2: Design a multi-scale fusion module to extract feature maps at different levels, so that the network can capture feature information at different scales and perform multi-scale feature information learning. Step S3: A multi-dimensional attention mechanism is proposed to further extract spatial and channel information of different dimensions, guiding the network to generate an attention feature map containing feature information, which includes the color and brightness information of the original image. Step S4: Combine the generated attention feature map with the adaptive illumination estimation module to obtain the illumination weight of the image, fuse the extracted feature information and the illumination factor prediction intermediate control layer, and combine it with the original image to obtain the enhanced image result.

2. The tunnel-based multi-scale adaptive low-light image enhancement algorithm according to claim 1, characterized in that: The specific experimental steps of step S1 are as follows: Workers at different construction stages photographed images of the tunnel face to obtain on-site data of the construction environment. The obtained tunnel construction images were integrated to obtain a dataset.

3. The tunnel-based multi-scale adaptive low-light image enhancement algorithm according to claim 1, characterized in that: The specific experimental steps of step S2 are as follows: (1) The original image is transferred to the encoder of the U-shaped network. Feature information is extracted by reducing the spatial resolution of the image layer by layer. Each layer in the encoder consists of multiple convolution blocks and downsampling operations to reduce the size of the feature map. The input of the feature extraction module is input, which contains the number of feature channels and the height and width of the input feature map. The input feature is first subjected to convolution and pooling operations to obtain the feature map F0. F0=Max_Pool(Conv(input)) F1=Max_Pool(Conv(F0)) F2=Max_Pool(Conv(F1)) F3=Max_Pool(Conv(F2)) Where Conv is the convolution module, Max_Pool is the pooling module, F1, F2, and F3 represent the first, second, and third levels of the encoder stage, respectively, which reduce the spatial resolution of the feature map layer by layer to obtain higher-level feature information. (2) The downsampled feature map is upsampled, and the low-resolution feature map is gradually restored to the size of the input image, and jump-connected and feature-fused with the encoder features at the same level. F4=deConv(F3) Concat1=Cat(F4,F2) F5=deConv(F4) Concat2=Cat(F5,F1) F6=deConv(F5) Concat3=Cat(F6,F0) F out =sigmoid(deConv(F6)) Where: Cat represents the cascade module, deConv represents the upsampling module, sigmoid represents the activation function, F4, F5, F6 and F out They represent the upsampling feature map results of different levels respectively. Concat1, Concat2 and Concat3 represent the feature maps corresponding to the skip connection respectively. The skip connection better combines the shallow features and the deep features, and better helps the network learn global and local features. (3) A multi-scale fusion module is designed to extract feature maps at different levels, enabling the network to capture feature information at different scales and perform multi-scale feature information learning. The process is as follows: The image is divided into two channels, extracting three scales: Concat2, Concat3, and F6. A 3×3 convolutional layer is used to extract local spatial features, capturing texture details and edge information in the tunnel face image. A 1×1 convolutional layer changes the number of channels in the feature map to achieve feature reorganization and dimensionality reduction in the channel dimension. Y=ω 1×1 (ReLU(BN(ω 3×3 (X)))) Where: Y is the output result, ω 1×1 and ω 3×3 Represent the weight matrices of 3×3 and 1×1 convolution kernels respectively.

4. The tunnel-based multi-scale adaptive low-light image enhancement algorithm according to claim 1, characterized in that: The specific experimental steps of step S3 are as follows: (1) Design a multi-dimensional attention mechanism module, considering the spatial dimension and channel dimension respectively. First, extract the horizontal and vertical feature vectors of the input features, implement cascade operations in the spatial dimension, and use 1×1 convolution to complete channel compression. The spatial attention map of the image is obtained through the activation function. Here, two spatial range pooling kernels (H, 1) or (1, W) are used to represent the horizontal channel and the vertical channel respectively. The two pooling kernel formulas under the zth channel are: Where: D z (i, j) represents the pixel value at the i-th row and j-th column position on the z-th channel of the input feature map, and W and H represent the width and height of the feature map respectively. (2) A multi-branch structure is used to process the feature vector obtained by the coordinate attention mechanism, and the feature vector is input into multiple fully connected layers arranged in parallel. It is combined with the activation function to generate channel weights, and finally an attention feature map is generated.

5. The tunnel-based multi-scale adaptive low-light image enhancement algorithm according to claim 1, characterized in that: The specific experimental steps of step S4 are as follows: (1) Design and predict the intermediate control layer. The extracted feature map is first processed by two 3×3 convolution kernels and a sigmoid activation function. Subsequently, the channel attention mechanism is introduced to obtain the channel weight map through global average pooling, a fully connected layer and a sigmoid activation function to predict the intermediate control layer. (2) Design an illumination estimation module, which consists of two illumination adjustment modules. The brightness of each pixel in the low-light image is calculated and adjusted in combination with the constraint function. Where: L is the input low-light image, F is the intermediate control layer predicted by the network, and C B and C D It is the enhanced image result obtained after fusion. (3) Construct the total loss function by the illumination loss function L G , image structure loss L ssim , color consistency loss L c , texture detail loss L p constitute. L total =λL G ++λ1L ssim +λ2L c +λ3L p Where: represents the gradient operator, c is the balance coefficient that controls the consistency of the structural strength, and constrains the illumination component by reducing excessive changes between gradients. Where: M represents the texture feature image generated by the model, A represents the texture feature of the original image, N represents the number of mappings of different features, H and W represent the length and width of the channel, μ and θ represent the gradient loss coefficients of H and W, represents the gradient operator, and Indicates H and W directions respectively.

Citation Information

Cited By

  • Method and system for identifying abnormal behaviors of personnel in tunnel construction environment

    CN121838256A