A Dehazing Method for UAV Aerial Images Based on Multi-Scale Hybrid Model

Through the defogging method based on the multi-scale hybrid model, semantic features of different scales are extracted and fused, and the problems of low contrast, color distortion and detail loss in harsh foggy days are solved, achieving efficient defogging effect and robustness.

CN115760632BActive Publication Date: 2025-07-01FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211497560.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-27
Publication Date
2025-07-01
Estimated Expiration
2042-11-27

AI Technical Summary

Technical Problem

Drone aerial photos are prone to low contrast, color distortion and loss of background target details in harsh foggy conditions.

Method used

The defogging method based on multi-scale hybrid model is adopted, and semantic features of different scales are extracted and fused through compression excitation module and cavity convolution module, network is built with multi-scale learning mechanism, background features are gradually learned, and object detection performance verification is used using YOLOv4.

Benefits of technology

Based on real-time operation speed and high target detection accuracy, it improves the robustness of drone aerial photography in different outdoor foggy environments, and has a wide range of application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760632B_ABST
    Figure CN115760632B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for dehazing UAV aerial images based on a multi-scale hybrid model. First, preprocess the UAV aerial images, and use the squeeze-and-excitation module SE to compress the image features to obtain a global receptive field and capture long-range information. Secondly, design the context expansion module CDB by using the dilated convolutional modules DCLs with different dilation factors to extract and fuse multi-scale semantic features and obtain rich context information. Finally, compare the result predicted by the trained model with the original UAV aerial image to judge the completion of the dehazing task, and use YOLOv4 to visualize the object detection performance of the dehazing result on the real UAV aerial image dataset. Based on the above steps, deploy the single-image dehazing method based on the multi-scale hybrid model in the UAV preprocessing stage. Compared with other single-image dehazing algorithms, the method of the present invention shows the best performance in terms of visual performance and object detection accuracy, and has a relatively fast processing speed, with obvious advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision denoising, and particularly to a method for dehazing unmanned aerial vehicle (UAV) aerial images based on a multi-scale hybrid model. Background Art

[0002] Currently, with the continuous development of single-chip microcomputer technology and aerial photography sensor technology, the demand for UAV aerial photography technology has been enhanced in multiple fields such as water conservancy construction and film shooting. However, the degradation of UAV aerial images in outdoor complex scenes, especially due to severe foggy weather, will inevitably lead to problems such as low contrast, color distortion, and loss of background target details in UAV aerial images. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method for dehazing UAV aerial images based on a multi-scale hybrid model, which achieves a real-time running speed, a high target detection accuracy, and robustness in different outdoor foggy weather environments compared with other methods.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions: A method for dehazing UAV aerial images based on a multi-scale hybrid model, comprising the following steps:

[0005] Step S1: Preprocess the UAV aerial image, and use a squeeze-and-excitation module (SE) to compress the image features, so as to obtain a global receptive field, capture long-range information, and adjust the problems of low contrast and color distortion caused by haze weather;

[0006] Step S2: Use dilated convolutional layers (DCLs) with different dilation factors to extract and fuse semantic features of different scales, obtain rich context information, and prevent over-smoothing and detail loss in UAV aerial images;

[0007] Step S3: Use a multi-scale learning mechanism to construct a network to progressively learn background features; compare the predicted dehazing result with the original UAV aerial image, and use YOLOv4 to visualize the target detection performance of the dehazing result image, and execute the running time of the model to verify the reliability of the algorithm.

[0008] In a preferred embodiment, the specific implementation of step S1 is as follows:

[0009] Step S11: Obtain the UAV aerial image I, preprocess it, and use a convolutional module to extract texture features;

[0010] Step S12: Based on the background texture features extracted in S11, design a squeeze-and-excitation residual block (SRB) by combining a squeeze-and-excitation operation (SE) with a skip connection;

[0011] Step S13: In step S12, the compression excitation residual block (SRB) adaptively recalibrates the feature responses λ of each feature map using the SE operation, and then multiplies them with the SRB input features for channel-wise weighting to obtain λX, completing the recalibration of the original features in the channel dimension; based on the feature recalibration of the SRB module, a preliminary defogged image is generated.

[0012] In a preferred embodiment, step S2 is specifically implemented as follows:

[0013] Step S21: The features obtained in step S1 are further input into the context expansion block (CDB) for learning using a multi-scale mechanism. A CDB module mainly consists of two convolutional layers and two dilated convolutional layers (DCLs).

[0014] Step S22: Based on the multi-scale mechanism of S21, the DCLs aggregate three convolutional paths with different dilation factors dilated = 1, 3, 5 and the same receptive field receptive field = 3×3. Each path obtains different scales of context information d1, d2, d3 according to the change of the dilation factor.

[0015] Step S23: The features d1, d2, d3 obtained in step S22 are respectively input into the dilated convolutional layers corresponding to dilated = 1, 3, 5 to obtain features r1, r2, r3. Finally, the features r1, r2, r3 extracted through multi-scale are fused and output as feature r, and are restored through convolution to obtain the defogging result map I. dehazed 。

[0016] In a preferred embodiment, step S3 is specifically implemented as follows:

[0017] Step S31: Introduce the mean square error (MSE) and perceptual loss to optimize network learning, evaluate the model training process in real time, and save the training model and data in real time.

[0018] Step S32: Introduce the reference-based performance evaluation metrics PSNR, SSIM, and LPIPS, and the reference-free performance evaluation metrics DHQI and FRFSIM to evaluate the completion of the defogging task.

[0019] Step S33: Introduce the running time to evaluate the processing speed of the network model; and use the popular pre-trained object detection model YOLOv4 to detect the targets of interest from the defogged images on real foggy day UAV aerial images.

[0020] Step S34: Based on the detected target bounding boxes in S33, sort all the defogging results by the average confidence coefficient, and verify that the defogging method using the multi-scale hybrid model significantly improves the target detection performance, which is helpful for outdoor UAV aerial photography applications.

[0021] Compared with the prior art, the present invention has the following beneficial effects: The method of the present invention provides a defogging method for UAV aerial photography images based on a multi-scale hybrid model, which effectively improves the robustness and target detection accuracy of UAV aerial photography technology in dealing with complex foggy days, and has a very broad application prospect. Description of the Drawings

[0022] Figure 1 is the overall work flow chart in the preferred embodiment of the present invention.

[0023] Figure 2 is the diagram of the squeeze-and-excitation residual block SRB and the context expansion block CDB in the preferred embodiment of the present invention.

[0024] Figure 3 is the network framework diagram in the preferred embodiment of the present invention.

[0025] Figure 4 is the comparison diagram before and after defogging in the preferred embodiment of the present invention.

[0026] Figure 5 is the visualization comparison diagram of the detection accuracy of YOLOv4 for the images before and after defogging in the preferred embodiment of the present invention. Detailed Embodiment

[0027] The present invention will be further described below with reference to the drawings and embodiments.

[0028] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0029] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0030] This embodiment provides a defogging method for UAV aerial photography images based on a multi-scale hybrid model, referring to Figures 1 to 5, including the following steps: Step S1: Preprocess the drone aerial image. Use the Squeeze-and-Excitation (SE) module to compress the image features, thereby obtaining a global receptive field, capturing long-range information, and adjusting problems such as low contrast and color distortion caused by haze weather; Step S2: Use the Dilated Convolution Layers (DCLs) with different dilation factors to extract and fuse semantic features at different scales, obtain rich context information, and prevent over-smoothing and detail loss in the drone aerial image; Step S3: Use a multi-scale learning mechanism to construct a network to progressively learn background features. Compare the predicted haze-removal result with the original drone aerial image, and use YOLOv4 to visualize the object detection performance of the haze-removal result image, execute the model running time, and verify the reliability of the algorithm.

[0031] In this embodiment, in step S1, preprocess the drone aerial image. Use the Squeeze-and-Excitation (SE) module to compress the image features, thereby obtaining a global receptive field, capturing long-range information, and adjusting problems such as low contrast and color distortion caused by haze weather. Specifically, it includes the following steps:

[0032] Step S11: Obtain the drone aerial image I, preprocess it, and use the convolutional module to extract texture features.

[0033] Step S12: Based on the background texture features extracted in S11, design a Squeeze-and-Excitation Residual Block (SRB) by combining the squeeze and excitation operation SE with skip connections.

[0034] In step S13, the Squeeze-and-Excitation Residual Block (SRB) adaptively recalibrates the feature response λ of each feature map using the SE operation, and then multiplies it with the SRB input feature to obtain λX by channel-wise weighting, completing the recalibration of the original features in the channel dimension. Generate a preliminary haze-removal image based on the feature recalibration of the SRB module

[0035] The diagram of the SRB module is as Figure 2 (a) described: In this embodiment, the Squeeze-and-Excitation Residual Block (SRB) is used to provide long-range information supplementation. In this embodiment, step S2: Use the multi-scale dilated convolution layers (DCLs) to extract and fuse semantic features at different scales, obtain rich context information, and prevent over-smoothing and detail loss in the drone aerial image. Specifically, it includes the following steps:

[0036] Step S21: Input the features obtained in step S1 into the Context Expansion Block (CDB) for learning using a multi-scale mechanism. A CDB module is mainly composed of two convolutional layers and two dilated convolution layers (DCLs).

[0037] Step S22: Based on the multi-scale mechanism of S21, DCLs combines three convolutional paths with different dilation factors (dilated = 1, 3, 5) and the same receptive field (receptive field = 3×3). Each path obtains context information d1, d2, d3 at different scales according to the change of the dilation factor.

[0038] Step S23: Input the features d1, d2, d3 obtained in Step S22 into the dilated convolutional layers corresponding to dilated = 1, 3, 5 respectively to obtain features r1, r2, r3. Finally, fuse the features r1, r2, r3 extracted through multi-scale and output them as feature r, and restore it through convolution to the defogging result image I. dehazed 。

[0039] The diagram of the CDB module is as Figure 2 (b) shown: In this embodiment, the context expansion block CDB is used for multi-scale information representation and enriching image details. Further, Step 3: Use the multi-scale learning mechanism to construct a network to progressively learn background features. Compare the predicted defogging result with the original UAV aerial photo, and execute the running time of the model to verify the reliability of the algorithm. The specific steps are as follows:

[0040] Step S31: Introduce the root mean square loss MSE and perceptual loss Perceptual Loss to optimize network learning, evaluate the model training process in real time, and save the training model and data in real time.

[0041] Step S32: Introduce the reference-based performance evaluation metrics PSNR, SSIM, and LPIPS, and the reference-free performance evaluation metrics DHQI, FRFSIM to evaluate the completion of the defogging task.

[0042] Finally, Figure 3 is the network framework diagram in the embodiment of the present invention. The entire multi-scale hybrid defogging model contains 12 SRB modules and 12 CDB modules, which progressively remove haze and supplement image texture details in multiple scales.

[0043] Figure 4 is the comparison diagram before and after defogging in the embodiment of the present invention. Figure 5 is the visualization comparison diagram of the detection accuracy of YOLOv4 for the images before and after defogging in the embodiment of the present invention. The specific steps are as follows:

[0044] Step S33: Introduce the running time to evaluate the processing speed of the network model. And use the popular pre-trained object detection model YOLOv4 to detect the target of interest from the defogged image on the real foggy day UAV aerial photo.

[0045] Step S34: Based on the detected target bounding boxes in S33, sort all the defogging results by the average confidence coefficient, verify that the defogging method using the multi-scale hybrid model significantly improves the target detection performance, which is helpful for outdoor UAV aerial photography applications.

[0046] The above are the preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention in terms of the functions and effects produced shall fall within the protection scope of the present invention.

Claims

1. A method for removing haze from UAV aerial images based on a multi-scale hybrid model, characterized in that, It includes the following steps: Step S1: Preprocess the aerial drone images. Use the Squeeze-and-Excitation (SE) module to compress the image features, thereby obtaining the global receptive field, capturing long-range information, and adjusting the problems of low contrast and color distortion caused by haze weather; Step S2: Use the Dilated Convolution Layers (DCLs) with different dilation factors to extract and fuse semantic features at different scales, obtain rich context information, and prevent over-smoothing and detail loss in the aerial drone images; Step S3: Use the multi-scale learning mechanism to construct a network for progressive learning of background features; Compare the predicted dehazed results with the original aerial drone images, and use YOLOv4 to visualize the object detection performance of the dehazed result images, execute the model running time, and verify the reliability of the algorithm; The specific implementation of step S1 is as follows: Step S11: Obtain the aerial drone image I, preprocess it, and use the convolutional module to extract texture features; Step S12: Based on the background texture features extracted in S11, design a Squeeze-and-Excitation Residual Block (SRB) by combining the squeeze and excitation operation SE with skip connections; Step S13: In step S12, the compression excitation residual block (SRB) adaptively recalibrates the feature responses λ of each feature map using the SE operation, and then multiplies them with the SRB input features for per-channel weighting to obtain λX, completing the recalibration of the original features in the channel dimension; a preliminary dehazed image is generated based on the feature recalibration of the SRB module The specific implementation of step S2 is as follows: Step S21: Further input the features obtained in Step S1 into the Context Expansion Module CDB for learning using a multi-scale mechanism. A CDB module mainly consists of two convolutional layers and two dilated convolutional layers DCLs; ​ Step S22: Based on the multi-scale mechanism of S21, the DCLs combines three convolutional paths with different dilation factors dilated = 1, 3, 5 and the same receptive field receptive field = 3×3. Each path obtains different-scale context information d1, d2, d3 according to the change of the dilation factor; Step S23: Input the features d1, d2, and d3 obtained in step S22 into the corresponding dilated convolutional layers with dilated = 1, 3, and 5 respectively to obtain features r1, r2, and r3; finally, fuse the features r1, r2, and r3 extracted at multiple scales and output them as feature r, and restore it through convolution to obtain the dehazed result image I dehazed ; The specific implementation of step S3 is as follows: Step S31: Introduce the Mean Squared Error (MSE) and Perceptual Loss to optimize network learning, evaluate the model training process in real time, and save the training model and data in real time; Step S32: Introduce the reference-based performance evaluation metrics Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS), and the reference-free performance evaluation metrics Dehazing Quality Index (DHQI), Fast Robust Foggy-Scene Image Metric (FRFSIM) to evaluate the completion of the dehazing task; Step S33: Introduce the running time to evaluate the processing speed of the network model; and use the popular pre-trained object detection model YOLOv4 to detect the objects of interest from the dehazed images on the real foggy aerial drone images; Step S34: Based on the object bounding boxes detected in S33, sort all the dehazed results by the average confidence coefficient, verify that the dehazing method using the multi-scale hybrid model has a significant improvement in object detection performance, which is helpful for outdoor aerial drone applications.