A multi-scale enhancement and clarification method for visible light images in underground coal mines
By introducing octave convolution and multi-scale enhancement modules into GCANet and combining them with the Mine-dehaze dataset, the problem of sharpening visible light images in coal mines under different dust and fog environments was solved, improving image clarity and information extraction capabilities.
Patent Information
- Application Number
- CN202310524162.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-10
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-05-10
AI Technical Summary
Existing technologies are insufficient to effectively improve the clarity of visible light images in coal mines under varying degrees of dust and fog, especially in their ability to extract information at different scales.
Based on the Gated Context Aggregation Network (GCANet), an octave convolution and multi-scale enhancement processing module are introduced. The parameters are optimized using the Mine-dehaze dataset to construct a sharpening processing model. Image sharpening is achieved through multi-scale fusion of feature maps and decoding processing.
It improves the clarity of visible light images in coal mines under different dust and fog environments, enhances the ability to extract information at different scales, and reduces model complexity.
Smart Images

Figure CN116703751B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to visible light image sharpening, in particular to a multi-scale enhanced sharpening method of a visible light image, and belongs to the field of image preprocessing. BACKGROUND
[0002] With the continuous development of machine vision technology, visible light image processing has been applied in many fields, including coal mine underground scenes. In this scene, there are a large amount of water mist and coal dust, which reduces the visibility and affects the sharpness and contrast of the captured visible light image. Therefore, the sharpening processing of the visible light image is crucial.
[0003] Currently, the visible light image sharpening methods aiming at dehazing are mainly concentrated in image enhancement, image restoration and convolutional neural network-based methods. Among them, the image enhancement technology aims to highlight the image details and improve the image contrast. This kind of method is basically driven by theory, and the algorithm design parameters need to be adjusted according to different images, which is difficult to obtain wide adaptability. The image restoration method is to solve the transmission rate and global atmosphere light value of the fog image according to the physical model of atmospheric degradation, and then perform inverse operation on the formation process of the fog image to realize fog image restoration and obtain a clear fog-free image. This kind of method needs rich prior knowledge as a guide to ensure the sharpening effect. Finally, the convolutional neural network-based method mainly refers to the end-to-end sharpening model. Under this research background, taking the visible light image in the coal mine underground scene as the processing object, how to improve the sharpening effect of the model on different degrees of dust and fog is a problem worth studying. Among them, improving the extraction ability of different scale information is the key to solving the above problems. SUMMARY
[0004] The purpose of the present application is to realize the sharpening method of the visible light image affected by dust and fog in the coal mine underground scene. First, a sharpening dataset Mine-dehaze is constructed for the measured visible light image in the coal mine underground scene. Then, an octuple convolution and multi-scale enhancement processing module are introduced based on the gated context aggregation network (GCANet) to realize the establishment of a sharpening processing model. Finally, the Mine-dehaze dataset image is constructed based on the sharpening processing model to optimize the parameters and complete the sharpening processing of the visible light image.
[0005] Specifically, the present application provides a visible light image sharpening method, and the steps thereof include:
[0006] 1) Constructing the coal mine underground scene visible light image defogging dataset Mine-dehaze, collecting the actual scene visible light image by Flir T660 infrared camera, according to the guidance of the visible light imaging physical model, selecting multiple cases of atmospheric light value A [0.6, 0.8, 1.0] and atmospheric dispersion coefficient β [0.4, 0.6, 0.8, 1.0, 1.2, 1.4, 1.6], to synthesize different degrees of blurred images;
[0007] 2) On the basis of step 1), using GCANet as the baseline model, using multiple step lengths of 1 and step lengths of 2 octaves convolution processing, the low frequency and high frequency features of the image to be processed are respectively processed by the low frequency and high frequency regions of the convolution kernel, and the processed different frequency information is interacted to form a multi-scale fusion feature map, compared with the classical convolution processing, the model complexity is effectively reduced while the feature extraction effect is guaranteed as much as possible;
[0008] 3) On the basis of step 2), the multi-scale fusion feature map is decoded, the feature map to be decoded is deconvolved and superimposed with the convolution processing result of the same level coding feature map, so as to repair the feature map, thereby forming the decoding feature map output by the multi-scale enhancement layer, which effectively fuses the coding feature map of the same level, enhances the expression of the feature map information, and realizes the establishment of the defogging model;
[0009] 4) On the basis of step 3), the parameter optimization of the defogging model is completed through the constructed Mine-dehaze dataset, and the defogging processing of the visible light image is realized.
[0010] The advantage of the present application is to provide a feasible scheme for visible light image defogging processing in the field of coal mine underground scene reconstruction. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 : Visible light image defogging method flow chart in specific embodiments of the present application;
[0012] Figure 2 : Octave convolution processing flow chart in specific embodiments of the present application;
[0013] Figure 3 : Multi-scale enhancement module processing flow chart in specific embodiments of the present application;
[0014] Figure 4 : Image defogging result comparison using Mine-dehaze dataset in specific embodiments of the present application. DETAILED DESCRIPTION
[0015] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application.Figure 1 As shown, the multi-scale enhancement-based visible light image sharpening method disclosed in the embodiment includes the following steps:
[0016] S101: Collect the visible light image of the actual scene in the coal mine by the Flir T660 infrared camera. Guided by the visible light imaging physical model, select multiple cases of atmospheric light value A ∈ [0.6, 0.8, 1.0] and atmospheric dispersion coefficient β ∈ [0.4, 0.6, 0.8, 1.0, 1.2, 1.4, 1.6], and then synthesize images with different degrees of blur. The constructed Mine-dehaze dataset can be used for subsequent sharpening model establishment.
[0017] S102: Use multiple eight-degree convolution processing with step size 1 and step size 2 based on the GCANet encoding module to realize information interaction of different frequencies. Compared with the classic convolution processing, the model complexity is effectively reduced while the feature extraction effect is guaranteed as much as possible.
[0018] S103: Decode the multi-scale fusion feature map, that is, add multi-scale enhancement processing to the GCANet decoding module, perform deconvolution processing on the feature map to be decoded, and superimpose the convolution processing result of the same level encoding feature map, so as to repair the feature map, thereby forming the decoding feature map output by the multi-scale enhancement layer, fusing the information of the same level encoding feature map and the current feature map, and realizing the establishment of the sharpening model.
[0019] S104: Complete the parameter optimization of the sharpening model through the Mine-dehaze dataset, and realize the sharpening processing of the visible light image.
[0020] In S101, the visible light image of the actual scene in the coal mine is collected by the Flir T660 infrared camera. On this basis, different degrees of blurred images are simulated according to the atmospheric scattering model. The working principle of the model is as follows:
[0021] I(x) = J(x) t(x) + A (1-t(x)) (1)
[0022] Where I(x) is the blurred image; J(x) is the clear image of the target object without fog radiation; A is the global atmospheric light; t(x) is the atmospheric transmittance, and its definition is
[0023] t(x) = e -βd(x) (2)
[0024] Wherein, β is the atmospheric scattering coefficient, the larger the value corresponds to the worse atmospheric projection effect; d(x) is the distance between the target and the picture acquisition system. Based on the above atmospheric scattering model, on the basis of the distance between the target and the picture acquisition system, the global atmospheric light A and the atmospheric scattering coefficient β are respectively limited according to A∈[0.6, 0.8, 1.0], β∈[0.4, 0.6, 0.8, 1.0, 1.2, 1.4, 1.6], and the collected actual image is fogged, so as to obtain different levels of blurred images, and constitute the Mine-dehaze data set. In the data set, 21 blurred images of different levels corresponding to each collected visible light image are generated. Accordingly, the data set contains a total of 3885 blurred images, of which 3108 images are selected as the training set, 389 images are selected as the verification set, and 387 images not overlapping with the training set are selected as the test set.
[0025] In S102, multiple eight-degree convolutions with a step of 1 and a step of 2 are used on the basis of the GCANet coding module to realize information interaction of different frequencies, and the process is as shown in Figure 2 . Specifically, the input image is divided into a high-frequency region X H and a low-frequency region X L , and the corresponding convolution kernel W and the output Y are also divided into high-frequency and low-frequency parts. In order to realize information updating within the same frequency and information exchange between different frequencies, the convolution kernel of the eight-degree convolution is divided into four parts, a high-frequency to high-frequency convolution kernel W H→H , a high-frequency to low-frequency convolution kernel W H→L , a low-frequency to high-frequency convolution kernel W L→H , and a low-frequency to low-frequency convolution kernel W L→L . The high-frequency feature map and the low-frequency feature map obtained after the eight-degree convolution can be defined as:
[0026] Y H =f(X H ;W H→H )+upsample(f(X L ;W L→H ),2) (3)
[0027] Y L =f(X L ;W L→L )+f(pool(X L ,2);W L→H )) (4)
[0028] where f(X; W) represents the convolution of image X with the convolution kernel W, pool(X, k) represents the average pooling operation with the kernel size of k x k, and upsample(X, k) represents the k times up-sampling operation. Then, the feature map extracted by the octave convolution can be obtained through the reconstruction of the low-frequency and high-frequency images.
[0029] In S103, a multi-scale enhancement process is added to the GCANet decoding module to fuse the information of the same level coding feature map and the current feature map. The process is shown in Figure 3 It is worth noting that in the decoding processing module, the solution of the current layer decoding feature map is completed according to the interaction of the next layer decoding feature map and the current layer coding feature map. This process can be described by the following formula:
[0030]
[0031] where j n is the deconvolution processing result of the n-th layer decoding feature map, i n is the convolution processing result of the coding feature map in the same layer as the n-th layer coding feature map, is the repair processing operation of the n-th layer decoding processing, which includes the classic residual group processing, and the parameter θ n is adjustable through training. Specifically,
[0032] S301: The deconvolution processing result j n+1 of the next layer decoding feature map is obtained through 2 times up-sampling processing to obtain a feature map with the same size as the current layer decoding feature map;
[0033] S302: The coding feature map i n in the same layer is subjected to convolution processing and pixel-by-pixel summation with the result of S301;
[0034] S303: The summed feature map is subjected to repair processing including classic residual group processing
[0035] S304: The repaired feature map is subtracted from the result of S301 pixel by pixel to obtain the current layer decoding feature map j n .
[0036] In S104, the training set, validation set and test set images included in the Mine-dehaze data set are used to assist in the parameter optimization of the clarification model.
[0037] In addition, in order to prove that the method of the application has a better effect on the clarification of the visible light image, the Mine-dehaze data set data is used for comparison of the clarification effect, and the comparison result is shown in Table 1. In Table 1, DCP is a dark channel prior method, CAP is a color attenuation model, GRM is a gradient regulation model, AOD-Net is an integrated defogging network, DehazeNet is an end-to-end single image defogging network, and GCANet is a gating context information aggregation network. In addition, Figure 4 Further, the qualitative results of the clarification of the images in the Mine-dehaze data set are shown, which proves that the method of the application has good clarification ability in the complex environment of the coal mine underground.
[0038] Table 1 Comparison of clarification effect for Reside data
[0039]
[0040]
[0041] Finally, it should be noted that the purpose of the disclosed embodiments is to help further understand the application, but those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the application and the appended claims. Therefore, the application should not be limited to the disclosed embodiments, and the scope of the application claimed is defined by the scope of the claims.
Claims
1. A method for dehazing of visible light images, comprising the steps of: 1) constructing a dehazing dataset Mine-dehaze for visible light images of underground coal mine scenes, collecting visible light images of actual scenes by a Flir T660 infrared camera, selecting an atmospheric light value A ∈ [0.6, 0.8, 1.0] and an atmospheric dispersion coefficient β ∈ [0.4, 0.6, 0.8, 1.0, 1.2, 1.4, 1.6] according to a physical model of visible light imaging to synthesize images with different degrees of blurring; 2) using a GCANet as a baseline model, using eight convolution processing with a step size of 1 and a step size of 2, performing convolution operations on low-frequency and high-frequency features of the image to be processed through low-frequency and high-frequency regions of the convolution kernel, and interacting different frequency information processed to form a multi-scale fusion feature map; Wherein the octave convolution processing specifically includes: dividing the input image into high-frequency region X H and low-frequency region X L , and the corresponding convolution kernel W and output Y are also divided into high-frequency and low-frequency two parts, the convolution kernel of octave convolution is divided into four parts, high-frequency to high-frequency convolution kernel W H→H , high-frequency to low-frequency convolution kernel W H→L , low-frequency to high-frequency convolution kernel W L→H , low-frequency to low-frequency convolution kernel W L→L , and the high-frequency feature map and low-frequency feature map obtained after octave convolution are defined as: Y H = f(X H ; W H→H )+ upsample(f(X L ; W L→H ), 2) (1) Y L = f(X L ; W L→L ) + f(pool(X L , 2); W L→H )) (2) wherein f(X; W) represents convolution of the image X with the convolution kernel W, pool(X, k) represents an average pooling operation with a convolution kernel size of k x k, and upsample(X, k) represents a k times upsampling operation; 3) performing decoding processing on the multi-scale fusion feature map, performing deconvolution processing on the feature map to be decoded, and superimposing the result of convolution processing with the same level of encoding feature map to form a decoding feature map output by the multi-scale enhancement layer, and realizing the establishment of the dehazing model; wherein the solution of the current layer decoding feature map is completed according to the interaction of the decoding feature map of the next layer and the encoding feature map of the current layer, which is described by the following formula: wherein j n is the result of the deconvolution processing of the decoded feature map of the nth layer, i n is the result of the convolution processing of the encoded feature map of the same layer as the encoded feature map of the nth layer, is the repair processing operation of the decoding processing of the nth layer, which contains the classical residual group processing, and the parameter θ n is adjusted through training. 4) completing parameter tuning of the dehazing model through the Mine-dehaze dataset to realize dehazing processing of the visible light image.
2. The method of sharpening a visible light image according to claim 1, wherein, The physical model of visible light imaging in step 1) is: I(x) = J(x) t(x) + A (1-t(x)) (4) wherein I(x) is a blurred image; J(x) is a clear image of the target object without fog radiation; A is a global atmospheric light; t(x) is an atmospheric transmittance, and its definition is t(x) = e -βd(x) (5) wherein β is an atmospheric scattering coefficient, and the larger the value, the worse the atmospheric projection effect; d(x) is the distance between the target and the image acquisition system.
3. The method of sharpening a visible light image according to claim 1, wherein, The solution process of the decoding feature map specifically includes: 1) the latter layer decoding feature map is processed by deconvolution to obtain result j n+1 Through 2 times upsampling processing, a feature map with the same size as the current layer decoding feature map is obtained. 2) the encoded feature map of the same level i n convolution processing and pixel-by-pixel summation with the result of step 1); 3) performing a repair process on the summed feature maps including a classical residual group process 4) subtract the result of step 1) pixel-wise from the repaired feature map, i.e. obtain the current layer decoded feature map j n .