A method for low-quality image reconstruction in underground coal mine monitoring scenarios
Patent Information
- Application Number
- CN202611002542.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-07-07
AI Technical Summary
[0007]本发明的目的是提供一种面向煤矿井下监控场景的低质图像重建方法,以解决井下图像空间非均匀退化难以建模、传统跳跃连接易传递干扰特征、无清晰真值时网络训练不稳定,以及重建质量与模型计算复杂度难以兼顾的问题
本发明通过一组重建参数向量得到多组增强参数向量,进而得到多个候选图像,使同一低质图像在不同组的增强参数向量控制下生成多个候选图像;随后,利用无参考图像质量评价函数对多个候选图像进行打分和排序,计算质量评分最高的候选图像和最低的候选图像之间的对比损失来训练学生网络,因此,本发明在缺少清晰真值图像的井下场景中,将无参考质量评价由外部评价指标转化为可参与网络训练的排序约束信号,从而为学生网络提供稳定的自监督优化方向;
Smart Images

Figure CN122510116B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing, and in particular relates to a method for reconstructing low-quality images for underground coal mine monitoring scenarios. Background Technology
[0002] Coal mine underground monitoring images exhibit significant spatial non-uniform degradation, with the same frame simultaneously containing areas of clarity, dust and fog, low illumination, and strong light interference. Existing image reconstruction methods suffer from the following drawbacks: 1. Lack of regional adaptive reconstruction: The degradation intensity field is not explicitly constructed, and the enhancement intensity cannot be dynamically adjusted according to the local degradation degree, which can easily lead to over-enhancement of clear areas and insufficient recovery of severely degraded areas.
[0003] 2. Traditional skip connections introduce noise artifacts: The shallow high-frequency features directly transmitted contain interference information such as coal dust scattering edges and strong light halos, which are difficult to distinguish from real details, resulting in distortion of the reconstructed image.
[0004] 3. Training is unstable when there is no clear ground truth: Relying on paired clear images or teacher network supervision, it is difficult to obtain the corresponding ground truth for the underground scene, resulting in unclear training objectives.
[0005] 4. Redundant and inefficient model: While improving detail restoration through dense connections and multi-scale fusion, the number of parameters and computational load is large, and the effective improvement for severely degraded regions is limited.
[0006] 5. Attention modulation lacks degenerate prior guidance: Weights are generated solely based on feature statistics, which can easily misclassify noise and pseudo-textures as valid details in low signal-to-noise ratio regions, affecting reconstruction stability. Summary of the Invention
[0007] The purpose of this invention is to provide a low-quality image reconstruction method for underground coal mine monitoring scenarios, in order to solve the problems of difficulty in modeling the spatial non-uniform degradation of underground images, easy transmission of interference features by traditional skip connections, unstable network training when there is no clear ground truth, and difficulty in balancing reconstruction quality and model computational complexity.
[0008] The present invention adopts the following technical solution: a low-quality image reconstruction method for underground coal mine monitoring scenarios, which acquires low-quality images of actual underground coal mine monitoring and inputs them into a trained student network to obtain reconstructed images; The method for training student networks is as follows: The student network encoder extracts and fuses multi-scale features from low-quality images to obtain reconstructed features. These reconstructed features are then mapped through fully connected layers or one-dimensional convolutional layers to obtain a set of reconstructed parameter vectors. Next, multiple sets of enhancement parameter vectors are sampled within the neighborhood of these reconstructed parameter vectors. These enhancement parameter vectors are then fused with the reconstructed features and decoded to obtain multiple candidate images. A no-reference quality score is applied to these candidate images, and the contrast loss between the candidate image with the highest and lowest quality score is calculated. The alignment loss between the multi-scale features output by the student network encoder and the multi-scale distilled reference features output by the teacher network encoder is also calculated to train the student network. The distilled reference features are obtained by the teacher network encoder through multi-scale feature extraction and fusion of low-quality images. In the decoding stages, the student network adaptively modulates the jump features of the student network encoder at each scale, and then fuses the adaptively modulated jump features at each scale with the upsampled features of the student network decoder layer by layer, and outputs the reconstructed image after multi-level decoding.
[0009] The beneficial effects of this invention are: This invention obtains multiple sets of enhancement parameter vectors from a set of reconstruction parameter vectors, and then obtains multiple candidate images, so that the same low-quality image generates multiple candidate images under the control of different sets of enhancement parameter vectors. Subsequently, the multiple candidate images are scored and ranked using a no-reference image quality evaluation function, and the contrast loss between the candidate image with the highest quality score and the candidate image with the lowest quality score is calculated to train the student network. Therefore, in the underground scene where there is a lack of clear ground truth images, this invention transforms the no-reference quality evaluation from an external evaluation index into a ranking constraint signal that can participate in network training, thereby providing a stable self-supervised optimization direction for the student network. This invention accurately characterizes the severity of mixed degradation such as dust, strong light, low illumination, and noise in various parts of downhole images by constructing a dedicated degradation intensity map. It can guide the network to adjust the image enhancement intensity pixel by pixel, avoiding the problems of local overexposure, loss of detail, or insufficient repair caused by traditional global uniform processing. This invention adaptively modulates the skip features of the student network, and then fuses and decodes the adaptively modulated skip features with the upsampled features of the student network decoder layer by layer. In this way, high-frequency information is selectively retained or suppressed according to the degree of local degradation, instead of directly transmitting shallow high-frequency features. This can suppress noise, halo edges and pseudo textures in heavily degraded areas, and can also preserve real textures and edge details in clear areas. This invention sets up a channel-space joint attention branch to jointly recalibrate the features in the student network in both channel and spatial dimensions. Without introducing large-scale self-attention or complex transformer structures, it performs lightweight modulation on image reconstruction features, balancing reconstruction quality and inference efficiency. Attached Figure Description
[0010] Figure 1 This is an architecture diagram of the student network of the present invention; Figure 2 for Figure 1 A schematic diagram of the architecture of the adaptive modulation module (DSM) in the image; Figure 3 A visual comparison of the method used in this embodiment and existing reconstruction methods in scenario 1; Figure 4 A visual comparison of the method used in this embodiment and existing reconstruction methods in scenario 2; Figure 5 The image shows a visual comparison of the method used in this embodiment with existing reconstruction methods in scenario 3. Detailed Implementation
[0011] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0012] This invention discloses a method for reconstructing low-quality images for underground coal mine monitoring scenarios. The method acquires low-quality images of actual underground coal mine monitoring and inputs them into a trained student network to obtain reconstructed images.
[0013] like Figure 1 As shown, the student network can adopt a lightweight four-layer pyramid encoder-decoder structure. Preferably, the number of channels in the student network increases progressively in the encoding stage according to 8, 16, 32, 64, and 128, forming a multi-scale downsampling representation. In the decoding stage, the image resolution is restored by progressively upsampling in the opposite direction, and the final output image has the same size as the input image. The teacher network is only used for distillation supervision during the training stage.
[0014] Therefore, the student network includes the following components: (1) Initial feature extraction layer: After inputting the low-quality image from the well, shallow feature extraction is first performed on the image. A convolutional layer with a kernel size of 11 and a stride of 1, as well as residual blocks, are used to extract structural features such as the edge and texture of the image and generate a shallow feature map with 8 channels.
[0015] (2) Multi-scale encoder: The feature maps output from the initial feature extraction layer are sequentially input into the multi-scale encoder. The multi-scale encoder consists of multiple levels of downsampling convolutional units, which reduce the spatial resolution of the feature maps by downsampling at each level, while increasing the number of feature channels at each level. The number of feature channels in the encoding stage is 8, 16, 32, 64 and 128 respectively. Shallow encoded features mainly preserve the local texture and edge details of the image, while high-level encoded features have a larger receptive field and stronger global semantic expression ability.
[0016] (3) Bottleneck layer: A bottleneck layer is set at the end of the encoder to further enhance and reconstruct the deep features with the lowest resolution and the highest number of channels, strengthen the network’s global modeling ability for complex degraded regions, and provide a deep feature basis for subsequent decoding and reconstruction.
[0017] (4) Adaptive modulation module, i.e., DSM: Multiple adaptive modulation modules are set between the encoder and the decoder to replace the traditional skip connections. Each skip connection transmits the features of the corresponding scale at the encoder end to the decoder end, and the modulated skip features are fused with the upsampled features of the corresponding scale at the decoder end.
[0018] (5) Multi-scale decoder: The deep features output from the bottleneck layer are input to the multi-scale decoder. The multi-scale decoder consists of multiple upsampling convolutional units. It restores the spatial resolution of the feature map by upsampling step by step and gradually restores the number of channels from 128 to 64, 32, 16 and 8. Each decoding stage receives the corresponding modulated jump features and fuses them with the current decoded features by adding them element by element.
[0019] (6) Output reconstruction layer: After multi-scale decoding and feature fusion, the decoded features are finally mapped to an image with the same spatial size as the input image.
[0020] The teacher network employs a multi-scale image reconstruction structure with strong representational capabilities; therefore, the teacher network may include the following components: (1) Teacher Encoder: The teacher encoder is used to extract multi-scale hierarchical features from low-quality downhole images. Res2Net (multi-scale residual network) is used as the backbone to extract hierarchical features. The multi-scale residual network extracts image features at different resolutions step by step, enabling the network to simultaneously obtain local texture information and global structural information. Preferably, the teacher encoder outputs encoded features at four main scales. Since the encoded features at different scales differ in channel dimension and semantic level, channel compression and alignment are performed on the encoded features at each scale. This is used to compress or map the teacher encoded features at different scales to a channel dimension that matches the subsequent teacher decoder and distillation reference features, facilitating subsequent decoding and reconstruction as well as multi-scale feature distillation.
[0021] (2) Bottleneck layer: This layer consists of 18 stacked residual blocks, which are used to enhance deep features. The residual blocks can maintain the flow of original information while improving the network's ability to represent complex degenerate regions.
[0022] (3) Teacher decoder: The spatial resolution of the image is gradually restored through multi-level upsampling. The corresponding scale coding features are fused in each level of decoding to restore structural information such as image edges and textures.
[0023] (4) Multiple auxiliary feature heads: output multi-scale distilled reference features, which are used by the student network to learn intermediate representations of the teacher network at different scales. During training, the student network can learn multi-scale priors through feature distillation loss, thereby achieving good image reconstruction capabilities with a lightweight structure. The teacher network is only used during the training phase to provide high-quality reconstructed images and multi-scale distilled reference features; the teacher network is not used during the actual inference phase, and only the trained student network is retained.
[0024] The method for training student networks is as follows: The student network encoder extracts and fuses multi-scale features from low-quality images to obtain reconstructed features. These reconstructed features are then mapped through fully connected layers or one-dimensional convolutional layers to obtain a set of reconstructed parameter vectors. Next, multiple sets of enhancement parameter vectors are sampled within the neighborhood of these reconstructed parameter vectors. These enhancement parameter vectors are then fused with the reconstructed features and decoded to obtain multiple candidate images. A no-reference quality score is applied to these candidate images, and the contrast loss between the candidate image with the highest and lowest quality score is calculated. The alignment loss between the multi-scale features output by the student network encoder and the multi-scale distilled reference features output by the teacher network encoder is also calculated to train the student network. The distilled reference features are obtained by the teacher network encoder through multi-scale feature extraction and fusion of low-quality images. In the decoding stages, the student network adaptively modulates the jump features of the student network encoder at each scale, and then fuses the adaptively modulated jump features at each scale with the upsampled features of the student network decoder layer by layer, and outputs the reconstructed image after multi-level decoding.
[0025] Finally, the reconstruction distillation loss, contrast loss, degradation intensity loss, and alignment loss are calculated separately to train the student network. The reconstruction distillation loss is the loss between the student network's reconstructed image and the teacher network's reconstructed image; the contrast loss is the loss between the candidate image with the highest quality score and the candidate image with the lowest quality score; the degradation intensity loss is the loss between the dedicated degradation intensity map at each scale and the degradation intensity prediction field; and the alignment loss is the loss between the multi-scale features output by the student network's encoder and the multi-scale distillation reference features output by the teacher network. The four types of losses are weighted and fused to obtain the total loss. The student network parameters are updated based on the total loss through backpropagation to obtain the trained student network.
[0026] The purpose of training the student network in this way is to address the problem of insufficient constraints in model training under conditions where clear ground truth images are unavailable. This is because it is typically difficult to obtain strictly paired low-quality images and clear ground truth images in underground scenes, making it impossible to directly construct a supervised mapping from degraded images to clear images. Therefore, by comparing the quality of candidate images obtained from the same input image under multiple sets of enhancement parameter vectors, the candidate image with the highest quality score is selected as the positive sample, and the candidate image with the lowest quality score is selected as the negative sample. This constructs a contrastive loss to constrain the student network to update in a more optimal direction. The contrastive loss is only executed during the training phase; during the inference phase, there is no need for repeated sampling or referenceless quality ranking comparisons. The specific steps are as follows: Suppose the input low-quality downhole image is Then we get a set of reconstruction parameter vectors as follows: ; in, To reconstruct the parameter vector; This represents the dimension of the reconstruction parameter vector. The reconstruction parameter vector consists of several factor magnitudes with image enhancement implications, used to describe the overall recovery tendency of the student network on the current input image.
[0027] The reconstruction parameter vector consists of any one or more of the following: brightness adjustment intensity, contrast adjustment intensity, detail restoration intensity, dust and fog suppression intensity, glare suppression intensity, and color balance intensity. The reconstruction parameter vector is not a fixed set of artificial parameters, but rather a learnable or adjustable variable during network training to describe the enhancement strategy employed for the current image.
[0028] This invention samples within the neighborhood of the reconstructed parameter vector to obtain multiple sets of enhancement parameter vectors. To determine which set of enhancement parameter vectors is superior in the absence of a clear ground truth image, multiple candidate images are generated using these multiple sets of enhancement parameter vectors. These multiple sets of enhancement parameter vectors are obtained by sampling within the neighborhood of the reconstructed parameter vector. During the sampling phase, a random perturbation is generated within the neighborhood of the reconstructed parameter vector according to a Gaussian distribution. This perturbation is then truncated and limited to add perturbation to the reconstructed parameter vector. The range of values for the random perturbation can be limited according to the physical meaning of each parameter component to avoid generating candidate results that are too dark, overexposed, or severely distorted. Specifically: in, For the first Group of random perturbations; To enhance the parameter vector.
[0029] A no-reference quality evaluation function is used to rank candidate images by quality. This function does not rely on sharp ground truth images, but rather comprehensively evaluates candidate image quality based on the imaging characteristics of downhole images, considering factors such as brightness and contrast balance, local sharpness, overexposure in strong light, and dust / fog degradation. Specifically: ; in, A no-reference quality score is given to the candidate images; Basic quality item; Penalty for overexposure due to strong light; This is a penalty for dust and fog degradation. These are all weighting coefficients, with values between 0.2 and 0.8, used to balance the impact of strong light overexposure penalties and dust and fog degradation penalties on the overall score; For any candidate image .
[0030] Low-quality underground images often contain reflections from miners' lamps, vehicle lights, equipment, and localized strong light sources. Over-enhancing these strong light areas can lead to diffusion, resulting in halos, overexposure, and edge distortion. Therefore, this invention introduces a strong light overexposure penalty term. First, the image is moderately smoothed or blurred to reduce the impact of random noise on the statistics of bright areas. Then, the proportion of pixels close to the saturation range is statistically analyzed. For example, pixels with brightness values greater than a preset brightness threshold are considered overexposure candidate areas. The larger the proportion of overexposure areas, the larger the strong light overexposure penalty term, and the lower the final quality score. This prevents the network from excessively amplifying localized strong light areas in pursuit of increased brightness.
[0031] Coal dust, water mist, and suspended particles in underground coal mine images can cause low image contrast, low color saturation, and poor visibility. This invention introduces a dust and fog degradation penalty term into the image quality score calculation. The higher the proportion of suspected dust and fog areas in the image, the worse the image visibility, the larger the corresponding dust and fog degradation penalty term value, and the lower the final image quality score.
[0032] For underground coal mine images, if the overall image is too dark, the visibility of structures is poor; if local contrast is insufficient, areas such as coal walls, conveyor belts, and equipment edges are difficult to identify; if the gradient energy is too low, it indicates insufficient recovery of texture and edge details. Therefore, this invention uses brightness, contrast, and sharpness together as basic quality parameters, specifically expressed as follows: ; in, This is the brightness quality item, used to determine whether the overall brightness of the image is suitable for observation and subsequent recognition; This is a contrast quality item used to evaluate whether the grayscale changes in local areas of an image are sufficient; The sharpness quality term can be obtained from gradient energy, gradient variance, edge strength, or Laplacian response statistics. , , These are all weighting coefficients in the basic quality item.
[0033] When scoring quality, For anchor point samples, As a positive sample, The positive samples represent the reconstruction directions with better quality in the neighborhood of the current parameters, while the negative samples represent the reconstruction directions with poorer quality. This transforms the unreferenced quality evaluation results into a trainable contrastive loss, enabling the student network to approach better reconstruction results and move away from poorer reconstruction results in the feature space.
[0034] The formula for calculating the contrast loss between the candidate image with the highest quality score and the candidate image with the lowest quality score is as follows: ; in, To compare the losses; The baseline image obtained by fusing and decoding the reconstructed parameter vector and reconstructed features; The candidate image with the highest quality score; The candidate image with the lowest quality score; For image feature extraction functions, For similarity measurement function, This is the temperature coefficient.
[0035] Furthermore, to improve training stability, this invention can set a quality difference sensitive weight based on the quality difference between positive and negative samples. Let the quality difference between positive and negative samples be: ; Then weights can be constructed: ; in, The weights are those used for construction. It is a monotonically increasing function; A sigmoid type weight function is used.
[0036] when When the value is large, it indicates a significant difference in the quality of positive and negative samples, making the ranking results more reliable, and the weight of the contrast loss increases; when When the value is small, it indicates that the quality of positive and negative samples is similar, and the ranking may be affected by noise without reference evaluation, thus reducing the weight of the contrast loss.
[0037] The final weighted comparison loss can be expressed as: ; in, This represents the weighted comparison loss.
[0038] In this process, during each decoding stage, the student network adaptively modulates the jump features at each scale of the student network encoder. These adaptively modulated jump features are then fused layer by layer with the upsampled features of the student network decoder, and the reconstructed image is output after multi-stage decoding. The method for adaptively modulating the jump features at each scale of the student network encoder is as follows: Step 1: Low-pass filter the jump features at each scale to obtain the low-frequency structure components. Subtract the jump features at each scale from the corresponding low-frequency structure components to obtain the high-frequency residual components of the jump features at each scale. Step 2: Generate high-frequency modulation weights based on the dedicated degradation intensity maps for each scale; Step 3: Multiply the high-frequency modulation weights at each scale element-wise with the high-frequency residual components of the corresponding scale jump features to obtain the modulated high-frequency residual components; fuse the modulated high-frequency residual components with the low-frequency structural components at the corresponding scale to obtain the modulated jump features, such as... Figure 2 As shown.
[0039] The method for obtaining the dedicated degradation intensity map in step 2 is as follows: Step 201: Register the reference frame to the coordinate system of the low-quality image through coarse registration transformation to obtain the registered reference image, and divide both the low-quality image and the reference image into multiple local image blocks according to the same grid.
[0040] This is because the degradation of low-quality underground images often exhibits spatial non-uniformity. Within the same frame, some areas suffer reduced visibility due to dust, fog, and strong scattering, while other areas remain relatively clear. Relying solely on enhancement parameter vectors often introduces over-enhancement and noise amplification in relatively clear areas, while failing to provide sufficient reconstruction capabilities for heavily degraded areas. Furthermore, the same monitoring scene underground is affected by dust concentration, ventilation status, and equipment operation at different times, often resulting in frames with significant differences in clarity. Therefore, a relatively clear reference frame can be selected from the same time window to provide weak supervision information for the degradation localization of the current low-quality image. To reduce the impact of viewpoint shift and slight camera movement, a coarse registration transformation is used to register the reference frame to the coordinate system of the low-quality image, resulting in a registered reference image. Considering that viewpoint changes, equipment vibration, and local motion are common in underground videos, pixel-by-pixel alignment is difficult to guarantee stability; therefore, local structural changes are statistically analyzed at the block scale. Both the low-quality image and the reference image are divided into the same grid. Local image patch .
[0041] Step 202: Calculate the high-frequency attenuation term and structural attenuation term for the corresponding local image patches in the low-quality image and the reference image, and then weight and fuse the high-frequency attenuation term and structural attenuation term to obtain the degradation feature map; specifically: For each corresponding local image patch, the structural similarity between the low-quality image and the reference image is calculated. A lower structural similarity indicates more severe structural information attenuation in that region. Therefore, the formula for calculating the structural attenuation term is: ; in, For structural attenuation terms; This represents the low-pass filter operator. Low-quality image; For reference image; This refers to a local image patch.
[0042] For each corresponding local image patch, calculate the high-frequency energy change between the low-quality image and the reference image. If the high-frequency energy of the current low-quality image is significantly lower than that of the reference image, it indicates that the region is blurred, obscured by dust or fog, or has lost details. Therefore, the formula for calculating the high-frequency attenuation term is: ; in, This is a high-frequency attenuation term; For local image patches Average gradient energy within; It is a very small constant.
[0043] Step 203: Perform bilinear interpolation on the degraded feature map to enlarge the resolution to match the size of the low-quality image, thereby obtaining the pseudo-degraded intensity field; Step 204: Scale the pseudo-degradation intensity field to the size of the corresponding scale jump feature to obtain the dedicated degradation intensity map for each scale.
[0044] During student network training, the degradation intensity loss between the dedicated degradation intensity map and the degradation intensity prediction field at each scale is calculated, constraining the pixel distribution of the degradation intensity prediction field and the dedicated degradation intensity map to tend to be consistent. The degradation intensity prediction field is output by a lightweight convolutional branch located at the end of the student network encoder.
[0045] The lightweight convolution branch includes a first convolutional layer, a first normalization layer, a first nonlinear activation function, a second convolutional layer, a second normalization layer, a second nonlinear activation function, an output convolutional layer, and an output normalization function, which are connected in sequence. The first and second convolutional layers are both 3×3 convolutional layers. The output convolutional layer is a 1×1 convolutional layer. The output normalization function is the Sigmoid function.
[0046] The formula for calculating the degradation intensity loss between the dedicated degradation intensity map and the degradation intensity prediction field at each scale is as follows: in, For degradation intensity loss; This is the degradation intensity prediction field output by the lightweight convolutional branch at the end of the student network encoder. This is a dedicated degradation intensity map.
[0047] To enhance the feature representation capability of complex degraded regions, a lightweight, fine-grained texture-oriented channel-space joint attention branch is set up within the student network. This branch integrates all enhancement parameter vectors and local degradation intensity information to perform channel and spatial recalibration of features. The branch receives the fused features obtained by fusing adaptively modulated skip features at the current scale with upsampled features from the student network decoder, the global enhancement parameter vector, and the dedicated degradation intensity map at the current scale as input. Based on the input, modulation information is generated, and channel recalibration is performed using this modulation information to obtain a calibration feature map, enhancing the channel responses related to dehazing, structural compensation, and texture restoration. The dedicated degradation intensity map is then concatenated and mapped to generate a spatial gating map. Multiplying the spatial gating map and the calibration feature map yields the spatial modulation result. Adding the spatial modulation result to the fused features results in lightweight, fine-grained features, concentrating the enhancement effect on highly degraded regions such as dust, fog, blur, and strong light, while avoiding over-enhancement of clear regions.
[0048] During inference, lightweight, fine-grained features are embedded into the key feature layer of the student network, specifically at the following locations: a. After the bottleneck layer and before the decoder; b. The position following the skip fusion feature with 64 channels in the decoding stage; c. The position after the skip fusion feature with 32 channels in the decoding stage.
[0049] This branch is implemented through global average pooling, lightweight mapping functions, one-dimensional local channel interactions, 1×1 convolution, and sigmoid activation functions. It does not require the construction of a full-channel correlation matrix or the introduction of a large-scale Transformer structure, thus achieving joint feature recalibration of the channel and spatial dimensions with low computational overhead.
[0050] Example: This example uses a self-built dataset covering key operational scenarios in a coal mine in Shaanxi Province, including tunneling faces, coal mining faces, belt conveyor systems, and return airways. All raw data was acquired from fixed mine monitoring cameras, with the monitoring system having a maximum acquisition resolution of 3 megapixels. During data preprocessing, frames were extracted and cropped from the underground monitoring videos, and severely blurred, black, and duplicate frames were removed. Simultaneously, nearly duplicate images were deduplicated to reduce the bias in statistical results caused by redundancy between adjacent frames. The final retained data sample covers real underground images from different operational areas, under different working conditions, and at different operational stages. The image resolutions primarily include 1920×1080 and 2560×1440.
[0051] The CUMID dataset was acquired using a KBA12B intrinsically safe mining camera captured images in multiple underground coal mines. This camera has a maximum resolution of 2560×1920 (5 megapixels) and a lighting distance of 30 meters. Video images were captured in various scenes within multiple coal mines using this intrinsically safe mining camera, including underground roadways, underground workshops, and coal mine conveyor belts.
[0052] Ablation experiment: The experiments were conducted using the PyTorch framework, and all training was performed on a single NVIDIA GeForce RTX 4060 (8GB) GPU. The operating system was Windows 11. The algorithm employed the Adam optimizer with first-order moment smoothing coefficients set. and second-order moment smoothing coefficient The values are 0.9 and 0.999 respectively, the batch size is set to 4, and the initial learning rate is set to 1×10. -4 And a cosine annealing strategy is used to gradually decay the temperature to 1×10⁻⁶ during training. -6 The input is an RGB three-channel image. To enhance the model's generalization ability, various data augmentation methods are used on the training samples during the training phase, such as random 90°, 180°, and 270° rotation, horizontal flipping, and random 512×512 cropping.
[0053] To verify the method of this invention, ablation experiments were set up on a self-built dataset and a CUMID dataset, using a student network as the basic framework.
[0054] Specifically, the following comparison models are set up in Tables 1 and 2: Model 1 is a teacher network; Model 2 is a student baseline that only incorporates knowledge distillation supervision, used to evaluate the fundamental role of teacher-student distillation in lightweight reconstruction; Model 3 adds contrast loss to Model 2 to analyze the impact of enhanced parameter vectors on the restoration effect; Model 4 builds upon Model 3 by introducing adaptive modulation to evaluate the ability to reconstruct local differences in degradation perception. Model 5 is based on Model 4 and adds a channel space joint attention branch.
[0055] The relevant quantitative results and comparative analysis of models 1-5 are shown in Tables 1 and 2.
[0056] Table 1 Ablation experiments on the self-built dataset Table 2 Ablation experiments on the CUMID dataset In the table: NIQE is the quality assessment metric for no-reference natural images; BRISQUE is the quality assessment metric for no-reference spatial domain images; FPS is the number of frames processed per second; Params are the model parameters; M is millions, ↑ indicates that the larger the value of the metric, the better the performance, ↓ indicates that the smaller the value of the metric, the better the performance.
[0057] The results in Tables 1 and 2 show that Model 1, with its standalone teacher network, improves overall restoration stability, indicating that the prior constraints provided by the teacher network help alleviate the uncertainty of training without ground truth. While Model 2, relying solely on basic teacher-student distillation, significantly reduces the number of parameters and improves inference speed, it also results in a noticeable decrease in restoration quality. Model 3, by introducing contrast loss, improves the no-reference quality index, demonstrating that quality-driven contrast constraints can learn more reasonable global enhancement directions. However, its effect is mainly at the global policy level; when sharp and heavily degraded regions coexist in the same frame, local overprocessing due to globally consistent enhancement may still occur. Model 4, by introducing adaptive modulation, significantly reduces artifacts and noise amplification in sharp regions while improving structural detail recovery in heavily degraded regions, verifying that pixel-level degradation scheduling plays a crucial role in balancing sharp region detail preservation and heavily degraded region recovery in lightweight models. Model 5, by setting a channel-space joint attention branch, combines global policy learning with local scheduling. This improves overall no-reference quality and regional restoration while maintaining inference efficiency, demonstrating their good complementarity. The enhancement parameter vector provides stable global enhancement behavior, while the degradation intensity prediction field allocates limited reconstruction capabilities in the spatial dimension. Together, they suppress unnecessary enhancement and noise amplification in clear regions and implement more targeted detail restoration and structural compensation in heavily degraded regions, thereby improving overall restoration quality, reducing model parameters, and maintaining low computational cost. Compared to Models 3 and 4, Model 5 outperforms them, indicating that the conditional fusion of degradation intensity and global policy helps achieve more accurate feature recalibration under low signal-to-noise ratio conditions.
[0058] Comparative experiments and results analysis: The dedicated degradation intensity map shows a higher response near dense dust and fog diffusion areas and in distant areas with weak textures, while the response is lower in relatively clear areas such as the ground and tunnel edges. This indicates that the constructed dedicated degradation intensity map can better reflect the spatial non-uniformity of downhole image degradation and provide more targeted prior information for subsequent skip-connect modulation and regional differential reconstruction.
[0059] To further verify the effectiveness of the reconstruction method in this embodiment for dust and fog recovery and detail restoration tasks in coal mines, three typical underground scenarios were selected and compared with several mainstream methods.
[0060] Table 3. Quantitative comparison between the algorithm in this embodiment and existing image dehazing algorithms. In the table: H represents information entropy; FADE represents fog concentration assessment index; NIQE represents no-reference image quality assessment; BI represents no-reference sharpness / fog comprehensive index, where ↑ indicates that the larger the value of the index, the better the performance, and ↓ indicates that the smaller the value of the index, the better the performance.
[0061] IDE (Source: Ju M, Ding C, Ren W, et al. IDE: Image dehazing and exposureusing an enhanced atmospheric scattering model[J]. IEEE Transactions on ImageProcessing, 2021, 30:2180-2192.); ROP + (Source: Liu J, Liu RW, Sun J, et al. Rank-one prior: Real-timescene recovery[J]. IEEE Transactions on Pattern Analysis and MachineIntelligence, 2023, 45(7):8845-8860.); ALSP (Source: He L, Yi Z, Liu J, et al. ALSP+: Fast scene recovery via ambient light similarity prior[J]. IEEE Transactions on Image Processing, 2025,34:4470-4484.) IHDCP (Source: Liu Y, Li T, Tan C, et al. IHDCP: Single Image Dehazing Using Inverted Haze Density Correction Prior[J]. IEEE Transactions on Image Processing, 2026, 35:1448-1461.); SLP (Source: Ling P, Chen H, Tan X, et al. Single Image Dehazing Using Saturation Line Prior[J]. IEEE Transactions on Image Processing, 2023, 32:3238-3253.); FFA-Net (Source: Qin X, Wang Z, Bai Y, Xie X, Jia H. FFA-Net: Feature Fusion Attention Network for Single Image Dehazing[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2020, 34(7): 11908-11915.); DeHamer (Source: Song Y, He Z, Qian H, et al. Vision transformers for single image dehazing[J]. IEEE Transactions on Image Processing, 2023, 32:1927-1941.); SFSNiD (Source: Cong X, Gui J, Zhang J, et al. A semi-supervised nighttime dehazing baseline with spatial-frequency aware and realistic brightness constraint[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024:2631-2640.); DehazeSB (Source: Lan Y, Cui Z, Luo X, et al. When Schrodinger bridgemeets real-world image dehazing with unpaired training[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2025:8756-8765.); KA_Net (Feng Y, Ma L, Meng IPC (Source: Fu J, Liu S, Liu Z, et al. Iterative predictor-critic codedecoding for real-world image dehazing[C] / / Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition. Nashville, TN, USA: Computer Vision Foundation / IEEE, 2025:12700-12709.); CoA (Source: Ma L, Feng Y, Zhang Y, et al. CoA: Towards real imagedehazing via compression-and-adaptation[C] / / Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition. 2025:11197-11206.).
[0062] Table 3 shows that, on the self-built dataset, the algorithm in this embodiment achieves the best result with FADE=0.39. The next best method is ROP. +The performance decreased by 7.14%, with H and NIQE also being optimal, indicating that it can effectively balance maintaining naturalness and overall visual quality while improving image sharpness. On the CUMID dataset, the algorithm in this embodiment achieved optimal results in H, FADE, and NIQE, with scores of 7.83, 0.43, and 4.00, respectively. Compared with the best-performing algorithm DehazeSB, H is improved by 7.85%; FADE / NIQE is improved by 41.89% / 3.15%; and in terms of BI, it achieved 15.71, only slightly higher than the optimal value of 15.07, with a difference of only 4.07%, which is considered a suboptimal result. This embodiment demonstrates superior quantitative performance under different data distributions, indicating that it has good scene adaptability and stable enhancement capabilities.
[0063] from Figure 3 , 4 As can be seen from point 5: 1. IDE and ROP + Although the algorithm can improve the overall brightness of the image to some extent, there is still a significant halo diffusion phenomenon near the bright areas, and the dust and fog residue in the distant areas is quite obvious. 2. The outputs of FFA-Net and DeHamer are generally dark, with insufficient recovery of details in dark areas and distant working surfaces; 3. IPC introduces noticeable artifacts and noise during the enhancement process; 4. CoA enhances the overall brightness of the image, but there is still strong halo diffusion and color imbalance around the light source; 5. The KA_Net algorithm can effectively suppress fog, especially in the foreground area of Scene 1, but the fog in the background is still quite noticeable.
[0064] Therefore, compared to other methods, the algorithm in this embodiment, while suppressing dust and fog, better preserves the outlines of people, the edges of alleyways, and the texture of the ground, resulting in a more natural overall brightness and contrast. + While SFSNiD still exhibits residual fog or blurred edges in some areas, KA_Net's dehazing effect is relatively stable, but cable edges and distant coal wall textures remain insufficiently clear. Therefore, the algorithm in this embodiment effectively suppresses dust, fog, and strong light interference in various scenarios, while preserving details such as coal wall textures, cable edges, and distant working surfaces. Especially in strong light neighborhoods and weak texture areas, the brightness transition is more natural, avoiding over-enhancement and halo diffusion in local areas, and mitigating the loss of detail and structural blurring in dark areas. Overall, the algorithm in this embodiment achieves a better balance between dust and fog suppression, detail restoration, and brightness equalization, resulting in reconstructed images with superior structural integrity and visual clarity.
Claims
1. A method for low-quality image reconstruction for a coal mine underground monitoring scene, characterized in that, Low-quality images of actual underground coal mine monitoring are acquired and input into a trained student network to obtain reconstructed images; The method for training the student network is as follows: The student network encoder extracts and fuses multi-scale features from low-quality images to obtain reconstructed features. These reconstructed features are then mapped through a fully connected layer or a one-dimensional convolutional layer to obtain a set of reconstructed parameter vectors. Next, multiple sets of enhancement parameter vectors are sampled within the neighborhood of these reconstructed parameter vectors. These enhancement parameter vectors are then fused with the reconstructed features and decoded to obtain multiple candidate images. A no-reference quality score is applied to these candidate images, and the contrast loss between the candidate image with the highest and lowest quality score is calculated. Additionally, the alignment loss between the multi-scale features output by the student network encoder and the multi-scale distillation reference features output by the teacher network encoder is calculated to train the student network. The distillation reference features are obtained by the teacher network encoder through multi-scale feature extraction and fusion of low-quality images. In the decoding stages at each level, the student network adaptively modulates the jump features of the student network encoder at each scale, and then fuses the adaptively modulated jump features at each scale with the upsampled features of the student network decoder layer by layer, and outputs the reconstructed image after multi-level decoding. The method for adaptively modulating the scale jump features of the student network encoder is as follows: Step 1: Low-pass filter the jump features at each scale to obtain low-frequency structural components, and subtract the jump features at each scale from the corresponding low-frequency structural components to obtain the high-frequency residual components of the jump features at each scale. Step 2: Generate high-frequency modulation weights based on the dedicated degradation intensity maps for each scale; Step 3: Multiply the high-frequency modulation weights of each scale element-wise with the high-frequency residual components of the corresponding scale jump features to obtain the modulated high-frequency residual components; fuse the modulated high-frequency residual components with the low-frequency structural components of the corresponding scale to obtain the modulated jump features. The method for obtaining the dedicated degradation intensity map in step 2 is as follows: Step 201: Register the reference frame to the coordinate system of the low-quality image through coarse registration transformation to obtain the registered reference image, and divide both the low-quality image and the reference image into multiple local image blocks according to the same grid. Step 202: Calculate the high-frequency attenuation term and structural attenuation term of the corresponding local image blocks in the low-quality image and the reference image, and then weight and fuse the high-frequency attenuation term and structural attenuation term to obtain the degradation feature map; Step 203: Perform bilinear interpolation on the degraded feature map to enlarge the resolution to match the size of the low-quality image, thereby obtaining the pseudo-degraded intensity field; Step 204: Scale the size of the pseudo-degradation intensity field to the size of the corresponding scale jump feature to obtain a dedicated degradation intensity map for each scale.
2. The low-quality image reconstruction method for a coal mine underground monitoring scene according to claim 1, characterized in that, in, The reconstruction parameter vector is any one or more of the following: brightness adjustment intensity, contrast adjustment intensity, detail recovery intensity, dust and fog suppression intensity, strong light suppression intensity, and color balance intensity.
3. The low-quality image reconstruction method for coal mine underground monitoring scene according to claim 1, characterized in that, in, The formula for calculating the contrast loss between the candidate image with the highest quality score and the candidate image with the lowest quality score is: ; wherein, is a contrast loss; is a reference image decoded from the reconstructed parameter vector and the reconstructed feature fusion; is a candidate image with the highest quality score; is a candidate image with the lowest quality score; is an image feature extraction function, is a similarity measure function, is a temperature coefficient.
4. The method for reconstructing low-quality images for underground coal mine monitoring scenarios according to claim 1, characterized in that, A stacked residual block is provided between the encoder and decoder of the teacher network. The stacked residual block is used to extract and refine the structural features and texture details of low-quality images. The stacked residual block consists of 18 residual blocks.
Citation Information
Patent Citations
Lightweight image super-resolution reconstruction method based on multi-dimensional knowledge distillation
CN113240580A
Image moire removing method and device suitable for intelligent terminal and storage medium
CN114596479A