Aerial image real-time defogging method and system based on dynamic negative correlation coupling and bidirectional degradation modeling
The aerial image dehazing method based on dynamic negative correlation coupling and bidirectional degradation modeling solves the problems of color distortion and insufficient detail retention in complex environments in existing aerial image dehazing methods, and achieves efficient and accurate real-time dehazing processing, which is suitable for applications such as drones and satellite remote sensing.
Patent Information
- Application Number
- CN202510870850.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing aerial image dehazing methods suffer from color distortion, insufficient detail retention, limited model generalization, difficulty in meeting real-time requirements, and inaccurate modeling of degradation mechanisms in complex aerial environments. These problems make it difficult to meet the high-quality, real-time dehazing processing requirements of applications such as drones, satellite remote sensing, and aerial photography.
A method based on dynamic negative correlation coupling and bidirectional degradation modeling is adopted. By constructing an encoder module, a dynamic negative correlation coupling module, a decoder module and an inverse haze module, and combining multi-objective loss functions for joint optimization training, multi-scale feature adaptive calibration and physical model constraints are achieved, and a bidirectional degradation path is constructed to ensure the reversibility and accuracy of the dehazing process.
The robustness and accuracy of the dehazing effect have been significantly improved, ensuring accurate restoration of image details in complex foggy environments, more accurate color reproduction, and more complete detail retention. This meets the real-time processing needs of edge devices such as drones and satellite remote sensing, reduces power consumption, and significantly reduces computational complexity.
Smart Images

Figure CN120725919A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aerial image processing, and in particular to a real-time aerial image defogging method and system based on dynamic negative correlation coupling and bidirectional degradation modeling. Background Art
[0002] Aerial image dehazing, a key technology in aerial image processing, is crucial for improving image quality in applications such as drones, satellite remote sensing, and aerial photography. However, existing aerial image dehazing methods suffer from numerous limitations, severely hindering both the effectiveness of dehazing and their performance in practical applications.
[0003] Traditional physics-based methods, such as the Dark Channel Prior (DCP), rely on the core assumption that atmospheric light is uniformly distributed and that the scene reflectivity satisfies specific conditions. However, in aerial scenes, atmospheric conditions are complex and variable, with high-altitude turbulence and significant variations in water vapor distribution at different altitudes. This assumption of uniform atmospheric light can easily fail. This failure directly leads to color distortion and insufficient detail preservation in dehazed images, severely impacting the dehazing effect.
[0004] On the other hand, learning-based methods, such as deep learning, use large amounts of data to learn the characteristics of foggy images, attempting to achieve more efficient dehazing. However, the high cost of acquiring aerial imagery data and the large variations in sceneries limit the generalization of models. Furthermore, aerial missions require extremely high real-time processing, and complex deep learning models, with limited onboard computing resources, take too long to infer, making it difficult to meet real-time requirements.
[0005] Crucially, existing methods often focus on single-directional degradation effects, considering only image degradation caused by fog scattering. They ignore bidirectional degradation effects in aerial imaging, such as changes in sensor noise characteristics and nonlinear changes in imaging system response caused by fog. For example, the image restoration method based on dynamic decomposition and fusion described in CN119540100B improves image restoration performance through dynamic decomposition and fusion strategies, but does not deeply model the bidirectional degradation effects of aerial foggy images.
[0006] Furthermore, existing methods fail to accurately capture the negatively correlated coupling relationship between the dynamic distribution of fog and image features, such as the dynamic correlation that the higher the fog concentration, the lower the proportion of effective image detail features. For example, the large multimodal model-guided adaptive image defogging method described in CN119693272A utilizes a large multimodal model, a Mixture of Experts (MoE) model, and a Mamba architecture to dynamically select appropriate experts for processing, improving defogging effectiveness. However, this method also fails to accurately model the negatively correlated coupling relationship between fog and image features.
[0007] This neglect results in an inability to accurately model the degradation mechanisms of aerial foggy images, making it difficult to achieve efficient, accurate, and real-time dehazing. For example, in complex aerial environments, existing methods may fail to accurately capture the dynamic relationship between fog and image features, leading to color distortion and insufficient detail preservation in dehazed images, severely impacting dehazing effectiveness and real-time performance.
[0008] In summary, existing aerial image dehazing methods suffer from color distortion, insufficient detail preservation, limited model generalization, difficulty meeting real-time requirements, and inaccurate modeling of degradation mechanisms when dealing with complex aviation environments. Therefore, there is an urgent need for an aerial image dehazing method that can adapt to complex aviation environments, accurately characterize the bidirectional degradation and dynamic negative correlation between fog and image, and possess real-time processing capabilities to meet the high-quality, real-time dehazing requirements of aviation missions. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide a real-time defogging method for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling, so as to solve the technical defects of existing defogging methods in the field of aerial image processing technology, especially in applications such as drones, satellite remote sensing and aerial photography. Specifically, when dealing with complex aerial environments, existing methods have problems such as color distortion, insufficient detail retention, limited model generalization, difficulty in meeting real-time requirements, and inaccurate degradation mechanism modeling. The present invention overcomes the above-mentioned limitations of the prior art by providing a real-time defogging method and system for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling, thereby achieving efficient and accurate real-time defogging processing.
[0010] In order to achieve the above objectives, the technical solution adopted by the present invention is: The present invention proposes a real-time defogging method for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling, comprising the following steps: Step 1: Build a bidirectional degradation modeling network framework: The framework consists of an encoder module, a dynamic negative correlation coupling module, a decoder module, an inverse atomization module and a joint optimization module.
[0011] The encoder module is responsible for extracting multi-scale features, the dynamic negative correlation coupling module is used to generate multi-scale calibration features, the decoder module is responsible for reconstructing the dehazed image, the inverse haze module constructs a parameter estimation subnetwork, and the joint optimization module is used to construct a multi-objective loss function and implement joint optimization training.
[0012] Step 2: Implement the dynamic negative correlation coupling mechanism: A negative correlation weight matrix is generated from atmospheric light features to adaptively calibrate the multi-scale transmission map features. This mechanism explicitly encodes the physical negative correlation between atmospheric light and the transmission map, enabling adaptive calibration of multi-scale features and ensuring accurate recovery of image details in areas with varying fog concentrations.
[0013] Step 3: Construct a bidirectional degradation path: A bidirectional degradation modeling framework, comprising a forward dehazing path and a reverse dehazing path, models degradation mechanisms through physical reversibility constraints. The forward path is responsible for dehazing, while the reverse path reconstructs the hazy image through a parameter estimation subnetwork, enforcing physical model constraints and ensuring the reversibility and accuracy of the dehazing process.
[0014] Step 4: Implement joint optimization training: We define a multi-objective loss function and implement joint optimization training to enhance model generalization through physical consistency constraints. The loss function includes dehazing loss, physical consistency loss, perceptual loss, and multi-scale structural similarity loss, ensuring that the model achieves optimal results in terms of dehazing, physical constraints, detail recovery, and structure preservation.
[0015] To complement the above method, the present invention also designs a real-time defogging system for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling. The system includes the following key components: 1. Data input interface: Configuration parameters support simultaneous input of multiple image data, and adopts ping-pong buffer mechanism to eliminate transmission delay and ensure the real-time and stability of data.
[0016] 2. Three-stage processing pipeline: This includes a feature extraction engine, a negative correlation calculation unit, and a reconstruction acceleration core. The feature extraction engine, composed of multiple parallel computing units, is responsible for extracting multi-scale features from the image. The negative correlation calculation unit generates a negative correlation weight matrix based on DSP (Digital Signal Processing) cluster computing and calibrates the multi-scale features. The reconstruction acceleration core rapidly reconstructs the dehazed image through operations such as bilinear interpolation, channel attention, and pixel-level fusion.
[0017] 3. Hardware Optimization: 8-bit fixed-point quantization and feature map block processing are used to reduce storage space and computational complexity, while improving processing speed and efficiency. These optimizations enable the system to achieve real-time 1080p@45fps processing on the Xilinx ZU9EG platform, while reducing power consumption to 4.8W, meeting the fast processing requirements of edge devices such as drones and satellite remote sensing.
[0018] The real-time defogging method and system for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling provided by the present invention have the following beneficial effects: 1. The present invention proposes a real-time defogging method and system for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling, which has shown significant beneficial effects in the field of aerial image processing technology, especially in applications such as drones, satellite remote sensing and aerial photography.
[0019] 2. This invention successfully addresses the technical shortcomings of existing dehazing methods in complex aerial environments, including color distortion, insufficient detail preservation, limited model generalization, difficulty meeting real-time requirements, and inaccurate degradation mechanism modeling. By leveraging a dynamic negative correlation coupling mechanism and a bidirectional degradation modeling framework, it achieves efficient and accurate real-time dehazing of aerial images, significantly improving the robustness and accuracy of the dehazing effect, ensuring accurate restoration of image details even in complex foggy environments, with more accurate color reproduction and more complete detail preservation.
[0020] 3. The dynamic negative correlation coupling mechanism of the present invention explicitly models the physical negative correlation between atmospheric light and transmission maps, realizing adaptive calibration of multi-scale features. It combines forward dehazing and reverse haze paths to construct closed-loop constraints, significantly improving the matching accuracy of physical models. Experiments show that the color error in the sky area is greatly reduced, and the edge distortion rate is significantly compressed. On the NH-HAZE (Non-Homogeneous Hazy, non-uniform haze image dataset) dataset, the PSNR (Peak Signal-to-Noise Ratio) reaches 27.06dB. The PSNR in dense fog areas is significantly improved, and the PSNR fluctuation range across fog concentration scenes is small, which has obvious advantages over existing technologies.
[0021] 4. The bidirectional degradation modeling framework and joint optimization strategy of the present invention enable the model to maintain good defogging effect in different scenes and fog concentrations, overcoming the technical limitations of the existing defogging method model with limited generalization and enhancing the generalization ability of the model.
[0022] 5. This invention utilizes a three-stage pipelined hardware acceleration system based on an FPGA (Field-Programmable Gate Array). Through a pipelined architecture of feature extraction, negative correlation calculation, and reconstruction output, combined with 8-bit fixed-point quantization and block processing, it overcomes real-time bottlenecks. It achieves 1080p @ 45fps processing on the Xilinx ZU9EG platform, with low latency, significantly reduced computational effort, and significantly lower power consumption, meeting the fast processing requirements of edge devices such as drones and satellite remote sensing.
[0023] 6. The hardware acceleration system of the present invention reduces power consumption and computational complexity while maintaining good defogging effect through optimized design and fixed-point quantization techniques.
[0024] 7. This invention constructs a lightweight, cross-level adaptive feature fusion mechanism that achieves adaptive weighted fusion of multi-scale calibrated features in the decoder through the collaboration of channel attention and pixel attention modules. This significantly reduces the number of parameters and computational complexity in the parameter estimation subnetwork. On the RS-Haze (Remote Sensing-Hazy) satellite remote sensing dataset, a large-scale, non-uniform remote sensing image dehazing dataset, the algorithm achieves excellent PSNR performance in both light and dense fog scenes, with significantly improved detail recovery.
[0025] 8. The hardware acceleration system of the present invention adopts a three-stage pipeline architecture and a ping-pong cache mechanism, which achieves efficient data processing and transmission, reduces power consumption and cost, and has good scalability and configurability, and can adapt to the needs of different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a diagram of the architecture of a multi-stage image dehazing and feature fusion model according to an embodiment of the present invention; Figure 2 This is a diagram of the architecture of a deep learning image dehazing model for estimating atmospheric light characteristics and transmission map characteristics according to Example 3 of the present invention. DETAILED DESCRIPTION
[0027] The technical solutions of the present invention are further described below with reference to the embodiments and accompanying drawings: Example 1 This embodiment describes a method for real-time dehazing of aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling. The specific implementation steps are as follows: Step 1. Data preparation: Dataset Collection: Collect aerial image datasets containing different fog concentrations and scenes, such as the NH-HAZE dataset and the RS-Haze dataset. These datasets cover a variety of weather conditions and scene types, which helps train defogging models with strong generalization capabilities. Data preprocessing: The collected image data is preprocessed, including image normalization and resizing. Image normalization scales pixel values to a specific range (such as 0 to 1) to facilitate model training; resizing ensures that all images have the same size for batch processing.
[0028] Step 2: Model construction: Encoder module: This module uses a convolutional neural network (CNN) structure, consisting of multiple convolutional layers, pooling layers, and activation functions. Convolutional layers are used to extract image features, pooling layers are used to reduce feature dimensions, and activation functions (such as ReLU, or Rectified Linear Unit) introduce nonlinear factors. Dynamic Negative Correlation Coupling Module: This module generates a negative correlation weight matrix based on atmospheric light characteristics to adaptively calibrate multi-scale transmission map features. Specifically, this module estimates the atmospheric light value first, then generates a weight matrix based on the negative correlation between the atmospheric light value and the transmission map to perform weighted calibration of the transmission map features. Decoder module: Use the calibrated features to reconstruct the dehazed image. The decoder module can use structures such as deconvolution layers and upsampling layers to gradually restore the spatial resolution of the image. Reverse fogging module: This module builds a parameter estimation subnetwork to implement the reverse fogging path. This module reconstructs the fogged image by estimating fogging parameters (such as atmospheric light value and transmission map), thereby strengthening the physical model constraints. Joint Optimization Module: Defines a multi-objective loss function, including dehazing loss (such as mean squared error loss), physical consistency loss (such as ensuring that the reconstructed hazy image is similar to the original hazy image), perceptual loss [such as using a pre-trained VGG (Visual Geometry Group) network to extract features and calculate differences] and multi-scale structural similarity loss [such as MS-SSIM (Multi-Scale Structural Similarity Index Measure) loss] to implement joint optimization training, update model parameters through the back-propagation algorithm, and minimize the multi-objective loss function.
[0029] Step 3: Model training: Training parameter settings: Set appropriate parameters such as learning rate, batch size, and number of training rounds. The learning rate controls the step size of parameter updates, the batch size determines the number of images processed in each iteration, and the number of training rounds determines the total number of times the model is trained. Training process: Use the prepared dataset to train the model. During the training process, the model parameters are updated through the backpropagation algorithm to minimize the multi-objective loss function. The model is verified using the validation set, and the training parameters are adjusted to obtain the best performance. For example, the learning rate or batch size can be adjusted based on the dehazing effect on the validation set (such as indicators such as PSNR).
[0030] Step 4: Model testing and application: Model testing: Use the test set to test the trained model and evaluate the dehazing effect. The test set should contain image data different from the training set to verify the generalization ability of the model. Model Application: The model is applied to practical aerial image dehazing tasks, such as those used by drones and satellite remote sensing. The model takes the aerial image to be dehazed as input, and outputs a clear, dehazed image. In practical applications, the model can be optimized and accelerated as needed to meet real-time processing requirements.
[0031] Example 2 In another preferred embodiment, based on the above embodiment 1, refer to Figure 1 and Figure 2 This embodiment describes in detail the specific implementation steps of the real-time defogging method for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling as follows: Step 1: Construct a bidirectional degradation modeling network framework, which consists of an encoder module, a dynamic negative correlation coupling module, a decoder module, an inverse fogging module, and a joint optimization module. The encoder module consists of four levels of convolutional layers, each of which contains Conv (Convolutional Layer), IN (Instance Normalization), and ReLU (Rectified Linear Unit); the dynamic negative correlation coupling module is used to generate multi-scale calibration features; the decoder module includes three levels of deconvolutional layers: Deconv (Deconvolutional Layer), CA (Channel Attention Module), and PA (Pixel Attention Module); the inverse fogging module is a lightweight parameter estimation subnetwork; and the joint optimization module is used to construct a multi-objective loss function. Step 2: Implement a dynamic negative correlation coupling mechanism, generate a negative correlation weight matrix based on atmospheric light characteristics, and perform adaptive calibration on multi-scale transmission map features; Step 3: Construct a bidirectional degradation path, including a forward defogging path and a reverse defogging path. A bidirectional degradation modeling framework is constructed, and degradation mechanism modeling is implemented through physical reversibility constraints. Step 4: Implement joint optimization training, define a multi-objective loss function and implement joint optimization training, and enhance the model generalization ability through physical consistency constraints. Through the above steps, an embedded U-shaped remote sensing dehazing network with cross-level feature adaptive fusion based on the U-shaped network is constructed. Step 5: Deploy the hardware acceleration system to implement hardware acceleration of the algorithm on the FPGA platform, and meet real-time processing requirements through customized architecture design; Step 6: Real-time dehazing inference.
[0032] In step 2, if Figure 1 As shown, the dynamic negative correlation coupling mechanism in the present invention generates a negative correlation weight matrix through atmospheric light characteristics, and performs adaptive calibration on the multi-scale transmission map characteristics. Specifically, the formula and middle, and Represent the input feature maps respectively; and represents the estimated transmission map parameters; and represents atmospheric light parameters; and is the characteristic map after calibration under different interactions; Indicates in The second characteristic of the step; Indicates in These parameters are calculated by the atmospheric light-guided physical negative correlation perception module to achieve adaptive calibration of the features.
[0033] The specific steps to implement the dynamic negative correlation coupling mechanism are as follows: 1. Extract the fourth layer output features of the encoder (1) Where, Represents the feature map output by the fourth layer of the encoder. This feature map contains the image features after multiple layers of convolution and nonlinear transformation, which is used for subsequent processing and analysis; It is the spatial dimension of the feature map. Due to four convolution and downsampling operations (such as convolution with a stride of 2), the height H and width W of the original image are reduced by 16 times; Indicates the number of channels of the feature map, each channel contains different feature information; It means "belonging to"; represents the set of real numbers; 2. Extract atmospheric light features: (2) is the input data. GAP (Global Average Pooling) refers to global average pooling of feature maps, compressing spatial dimensions to retain channel semantics. The global average pooling operation is as follows: (3) Where, is the spatial dimension of the feature map (height times width), and its inverse indicates averaging the summation results; Indicates the sum of the spatial dimensions of each channel of the feature map; Represents the index of the feature map in the height (H) dimension; Represents the index of the feature map in the width (W) dimension.
[0034] 3. Generate negative correlation weights: (4) Dimensionality reduction through the first fully connected layer: (5) Generate weights through the second fully connected layer: (6) Generate bias through the third fully connected layer: (7) Where, 、 、 Represent the weight matrices of the 1st, 2nd, and 3rd layers respectively; 、 、 Represents the bias items of the 1st, 2nd and 3rd layers respectively, which are vectors with the same dimensions as 、 、 The same number of columns; represents the output after dimensionality reduction through the first fully connected layer; represents the weights generated by the second fully connected layer; represents the bias generated by the third fully connected layer; represents the activation function; Represents the intermediate features of dimensionality reduction; Represents the channel negative correlation weight (range [0, 1]); Represents spatial adaptive bias (value range [-1, 1]); Adopting dual-path activation design, Sigmoid (S-type function) is used to constrain to [0, 1] to simulate negative correlation. Tanh (hyperbolic tangent function) constrained to [-1, 1] provides adaptive bias.
[0035] 4. Perform multi-scale feature calibration: Adopt fog concentration adaptive strategy; dense fog area (low Value): Enable 1 / 8 scale deep feature calibration; mist area (high Value): Enable 1 / 2 scale shallow feature calibration; 1) 1 / 2 scale processing: (8) Represents a feature map; Represents an upsampling operation; Indicates bilinear interpolation upsampling by 4 times; Represents the calibrated feature map, whose spatial dimension is 1 / 2 of the height H and width W of the original image, and the number of channels is C; features of this scale are suitable for processing hazy areas, because hazy areas have relatively more details and require higher spatial resolution to capture the details; represents convolution; The number of channels representing the feature map is reduced to a quarter of the original; Will 、 Bilinear interpolation upsampling by 4 times , performing channel-by-channel multiplication and accumulation operations.
[0036] 2) 1 / 4 scale processing: (9) 、 Broadcast directly to , perform dimension matching; Where, represents the feature map before calibration, Represents a calibrated feature map with a spatial dimension of 1 / 4 the height (H) and width (W) of the original image and a channel number of C. Features of this scale are suitable for processing areas with general fog concentration.
[0037] 3) 1 / 8 scale processing: (10) Where, represents the downsampling operation, Indicates average pooling downsampling by 2 times; represents the feature map before calibration, Represents the calibrated feature map, whose spatial dimension is 1 / 8 of the height H and width W of the original image, and the number of channels is C; this scale feature is suitable for dense fog areas, because dense fog areas have fewer details, and lower spatial resolution is sufficient to capture the main structural information while reducing computational complexity; the downsampling operation is 、 Perform 2×2 average pooling, output size matching .
[0038] 5. Generate the calibrated three-scale feature group: (11) The dynamic negative correlation coupling mechanism generates a dynamic negative correlation weight matrix and bias , explicitly encoding atmospheric light With transmission diagram The physical negative correlation ( ), achieving adaptive calibration of multi-scale features.
[0039] In step 3, if Figure 1 As shown in the figure, the architecture of a bidirectional defogging network includes a forward process and a reverse process; and Represent the convolution operations of the first and second layers respectively; , ,…, Represents the feature map or feature layer in the forward process; , ,…, Represents the feature map or feature layer in the reverse process. Figure 2 As shown, the architecture also includes an atmospheric light feature estimation branch and a transmission map feature estimation branch. In the atmospheric light feature estimation branch, the input feature map After processing through the GAP (Generic Access Profile) layer and multiple MlpBlock (Multilayer Perceptron Block) modules, atmospheric light signatures are generated. ; In the transmission map feature estimation branch, input feature map Generate transmission map features through multi-scale processing and NCGM (Negative Control Mechanism) .
[0040] Among them, the Residual Block is used to increase the depth of the network without adding additional parameters while keeping the input and output dimensions consistent; It represents a 1×1 convolutional layer, which is used to adjust the number of channels of the feature map without changing its spatial dimension. Element-wise Addition represents element-level addition, that is, the addition of elements at corresponding positions, which is used to fuse feature maps from different sources. Element-wise Multiplication represents element-level multiplication, that is, the multiplication of elements at corresponding positions, which is used for weighting feature maps. Concatenation represents the concatenation operation of feature maps, which is used to merge feature maps of different scales or sources together to retain more information. Chunk here means dividing the feature map into multiple blocks for parallel processing or specific operations.
[0041] The detailed implementation process of the bidirectional degradation path is as follows: The steps of the forward path (dehazing process) are as follows: 1. Input foggy image , Represents an image tensor, which is a three-dimensional array containing the pixel values of the image.
[0042] 2. Extract multi-scale features through a four-stage encoder: (12) Where, 、 、 、 Refers to the output feature maps of the 1st, 2nd, 3rd, and 4th level encoders respectively; Refers to a 3×3 convolution operation with a stride of 2 and a padding of 1; the number of instance normalization (IN) channel groups is 32, and the activation function uses ReLU.
[0043] 3. Input encoder fourth layer features , the calibration signature is generated by the dynamic negative correlation coupling module (step S2): (13).
[0044] 4. Reconstruction of multi-level decoding: (14) Where, Represents the calibrated deepest scale features, which contain lower spatial resolution but important semantic information; Represents the intermediate feature map generated by deconvolution and attention mechanisms, with spatial resolution increased to H / 4 × W / 4. Concat is an operation used to connect two or more tensors along a specified dimension to form a new tensor. This operation is used to merge feature information from different sources so that it can be processed together in subsequent network layers. Refers to a 3×3 deconvolution (also called transposed convolution) operation; represents calibrated mid-scale features; Represents the intermediate feature map generated by deconvolution and attention mechanism, with the spatial resolution increased to H / 2×W / 2; represents calibrated shallow-scale features; Represents the final dehazed image, generated by deconvolution and attention mechanism, and restored to the original resolution H×W.
[0045] The steps of the reverse path (atomization process) are as follows: 1. Construct parameter estimation subnetwork: (15) Where, Represents the true value of the clear image; represents the estimated parameters of the transmission graph; Represents the estimated parameters of atmospheric light; GAP (Global Average Pooling) means that by calculating the average value of each channel of the feature map, the spatial dimension is compressed to retain the channel semantics; GMP (Global Max Pooling) means that by calculating the maximum value of each channel of the feature map, the global maximum feature is extracted; Represents a 1×1 convolution operation, which is used to perform a linear transformation on the pooled features and generate parameter estimates.
[0046] 2. Generate fog image: (16) represents the generated image; the physical constraints include the spatially varying transmission map ( ) and the global atmospheric light vector ( ).
[0047] The detailed process of step 4 is as follows: Step 4.1. Calculate the loss function Step 4.1.1. Define the dehazing loss function Measure the pixel-level difference between the dehazed image and the real clear image, and use the smooth L1 function to reduce the influence of outliers: (17) Where, Represents the loss value of the dehazed image, which is used to measure the quality of the dehazed effect; Represents the first pixel values; Indicates the first pixel values; The total number of representative samples in, The loss function is defined as follows: (18) Indicates the value of the input variable, for the high error region ( ) uses linear penalty to avoid gradient explosion; for low error areas ( ) uses quadratic penalty to improve convergence accuracy.
[0048] Step 4.1.2: Physical consistency loss Ensure that the fog image reconstructed by the reverse fogging path is consistent with the input fog Figure 1 To strengthen the physical model constraints: (19) Where, Represents the physical consistency loss value, which is used to measure the difference between the fog image reconstructed by the reverse fogging path and the input fog image; represents the first pixel values; The first pixel values; Indicates the input fog image and the reconstructed fog image in The absolute difference in pixels.
[0049] Calculation process: 1) Clear truth value (20) 2) Foggy input (twenty one) 3) Calculate the absolute difference (twenty two) The atmospheric scattering model is embedded in the loss function; a closed-loop verification mechanism of "clear → foggy image → clear" is established; the weight coefficient 0.1 balances physical constraints and visual quality.
[0050] Step 4.1.3. Perceptual Loss Constraining the similarity between the output image and the true value in the feature space enhances texture detail recovery: (twenty three) Where, Represents the perceptual loss, which is used to measure the dehazed image and true clear images Similarity in feature space,enhances the recovery of texture details; Dehazed image In the VGG network Feature maps of layers; Indicates a true and clear image In the VGG network Feature maps of layers; 、 、 Respectively represent The number of channels, height, and width of the layer feature map; is a normalization factor used to balance the scale differences of feature maps in different layers; Indicates the dehazed image and the real clear image in the The squared Euclidean distance in the layer feature space measures the feature difference between the two; Represents the VGG network 、 and The feature differences of the three layers are summed to obtain the total perceptual loss; Freeze VGG16 (Visual Geometry Group 16-layer) parameters, only used as feature extractor, normalization factor Eliminate scale effects while multi-level features complement each other to enhance detail recovery.
[0051] The feature extraction layer configuration is shown in Table 1: Table 1 Feature extraction layer
[0052] Step 4.1.4, multi-scale structural similarity loss Multi-scale modeling adapts to areas with different fog concentrations. The separate evaluation of brightness and contrast is more consistent with human perception. The product form strengthens the inter-scale dependency and maintains the similarity of image structure. (twenty four) Where, It stands for multi-scale structural similarity loss, which is used to measure the structural similarity between the dehazed image and the real clear image at multiple scales; Indicates the number of multiple scales, usually multiple scales (such as 5 scales); Indicates the index of the sample; Dehazed image In the The mean on each scale; Indicates a true and clear image In the The mean on each scale; Dehazed image In the variance on the scale; Indicates a true and clear image In the variance on the scale; Dehazed image and true clear images In the Covariance on scales; 、 They represent two different small constant terms, which are used to stabilize the denominator of the formula and prevent division by zero errors. Their values are usually very small. 、 : Represents two different weight coefficients at each scale, used to balance the contribution of different scales to the total loss; these weights can be adjusted as needed, and different values are usually set according to the importance of the scale.
[0053] Parameter settings: Brightness constant
[0054] Contrast constant
[0055] Luminance Weight
[0056] Contrast Weight
[0057] Scale number
[0058] Step 4.2: Joint optimization objectives Step 4.2.1. Total loss function Integrate four types of loss functions to balance pixel accuracy, physical constraints, detail recovery, and structure preservation: (25) Where, is the total loss function, Represents the dehazing loss, which measures the pixel-level difference between the dehazed image and the true clear image. A smooth loss function is used to reduce the impact of outliers and ensure that the dehazed image is close to the true image at the pixel level. It is the physical consistency loss, ensuring that the fog image reconstructed by the reverse fogging path is consistent with the input fog Figure 1 By minimizing the difference between the reconstructed fog image and the input fog image, the constraints of the physical model are strengthened, and the physical consistency and generalization ability of the model are improved; It is a perceptual loss that constrains the similarity between the output image and the real image in the feature space, and enhances the texture detail recovery of the dehazed image through multi-layer feature extraction of the VGG network; It is a multi-scale structural similarity loss that adapts to different fog density areas. It maintains the structural similarity of the image and improves the visual quality by separate evaluation of brightness and contrast. 、 、 、 They are 、 、 and The weight parameter is used to balance the contribution of different loss items in the total loss.
[0059] Step 4.2.2, configure weights The optimized strategy is shown in Table 2: Table 2 Optimized strategies
[0060] Step 4.2.3, gradient back propagation, the negative correlation module uses a higher learning rate (1.2 , is the learning rate), accelerating the convergence of key parameters. Using a hierarchical learning rate mechanism, deep parameters , shallow parameter 0.8 The gradient clipping threshold is set to 1.0 to prevent gradient explosion.
[0061] Step 4.3: Training strategy implementation Step 4.3.1. Optimizer configuration Adaptive learning rate adapts to different parameter characteristics, and momentum mechanism accelerates convergence in flat areas: (26) Step 4.3.2, learning rate scheduling (cosine annealing) The mathematical expression is as follows: (27) The parameter configuration is shown in Table 3: Table 3 Optimizer parameter configuration
[0062] Step 4.3.3, batch training Parameter configuration is shown in Table 4: Table 4 Batch training parameter settings
[0063] Step 4.4: Training Monitoring and Stabilization Step 4.4.1. Convergence Monitoring Strategy This step includes the training phase, the validation phase, and the early stopping mechanism. The training process starts from the first round and ends at the 300th round. Each round of training includes processing the training data and periodic validation. The validation phase is used to evaluate the model performance and decide whether to save the model or trigger the early stopping mechanism based on performance indicators (such as PSNR). Set a loop from 1 to 300, indicating that the model training process is carried out for a total of 300 rounds; each round represents a complete traversal of the entire training dataset; In each round of training, traverse each batch in the training dataset and set the data loader train_loader to provide the training data to the model in batches; For each batch, the following four steps are performed: 1) Forward propagation: The input data is passed through the model for forward calculation to obtain the model output; 2) Calculate loss: Based on the model output and the target value, calculate the value of the loss function, which measures the difference between the model output and the true target; 3) Backpropagation: Backpropagation is performed by calculating the gradient of the loss function with respect to the model parameters. This step is to determine how to adjust the model parameters to minimize the loss. 4) Parameter update: Update the model parameters based on the gradients obtained through backpropagation. This is done through an optimizer such as Adam (Adaptive Moment Estimation). After every 10 rounds of training, the validation phase begins. This is to regularly evaluate the performance of the model during training and avoid overly frequent validation, thereby saving computing resources. Calculate the validation set PSNR / SSIM. Calculate the model's performance indicators, such as PSNR and SSIM (Structural Similarity Index Measure), on the validation dataset. These indicators are used to measure the performance of the model on the validation set.
[0064] If the PSNR value of the current round is greater than the best PSNR value (best_PSNR) recorded previously, the following operations are performed: 1) Update the best PSNR value: best_PSNR = PSNR; 2) Save the parameters of the current model for subsequent use; 3) Reset the early stop counter: early_stop_counter = 0; If the PSNR value of the current round does not exceed the optimal PSNR value, the early stopping counter is incremented by 1; If the value of the early stopping counter exceeds 20, it means that the model performance has not improved further (that is, the PSNR has not exceeded the optimal value) in the past 200 rounds (because it is verified every 10 rounds), and the early stopping operation is performed; The early stopping operation stops the training process to prevent the model from overfitting and returns the best model parameters saved previously. These parameters correspond to the moment when the model performed best on the validation set.
[0065] Step 4.4.2. Gradient stabilization First perform gradient clipping: (28) Where, The gradient vector represents the model parameters. The gradient indicates the direction and magnitude of the model parameters that need to be updated in the current training step. : represents the L2 norm of the gradient vector, that is, the Euclidean length of the gradient vector, which measures the magnitude of the gradient; is the threshold for gradient clipping, set to 1.0, which is used to control the maximum allowed size of the gradient; This part is used to determine whether the gradient needs to be clipped. If the L2 norm of the gradient is greater than the threshold , then scale the gradient to the threshold Otherwise, the gradient remains unchanged.
[0066] Then normalize the weights: (29) Where, The weight vector of the model is The value at the iteration; Represents the weight vector The L2 norm of , that is, the Euclidean length of the weight vector, is calculated as: (30) In the formula is the dimension of the weight vector; finally, the gradient is accumulated and the parameters are updated every 4 mini-batches, and the equivalent batch size is increased to 32.
[0067] The implementation results are shown in Table 5: Table 5 Implementation effect
[0068] The training efficiency breakthrough is shown in Table 6: Table 5 Training efficiency breakthrough
[0069] In step 5, the detailed implementation process of deploying the hardware acceleration system is as follows: Step 5.1: System Architecture Design Step 5.1.1. Data input interface The configuration parameters are shown in Table 7: Table 7 System architecture configuration parameters
[0070] It supports simultaneous input of 4 channels of image data (128-bit per channel) and uses a ping-pong buffer mechanism to eliminate transmission delays.
[0071] Step 5.1.2, three-stage processing pipeline First level: Feature extraction engine The input data is the starting point of the network and is fed into a processing pipeline consisting of multiple parallel computing units. Each parallel computing unit performs the same sequence of operations to extract features at different scales. In each parallel computing unit, the input data first undergoes a 3×3 convolution operation to extract local features. After the convolution operation, the data undergoes instance normalization. This process normalizes each feature map, which helps stabilize the training process and improve the generalization ability of the model. Finally, the data passes through the ReLU activation function to introduce nonlinearity, enabling the network to learn complex feature representations.
[0072] These operations are performed sequentially in each parallel computing unit as follows: 1) Parallel computing unit 1: Input data passes through Convolution, then instance normalization, and finally ReLU activation function; 2) Parallel computing unit 2: The output of unit 1 goes through the same →IN→ReLU process; 3) Parallel computing unit 3: The output of unit 2 continues through →IN→ReLU process; 4) Parallel computing unit 4: The output of unit 3 passes through again →IN→ReLU process.
[0073] After processing by these four parallel computing units, the network outputs multi-scale features, including: : Indicates features with a scale of 1 / 2; : Indicates features with a scale of 1 / 4; : Indicates features with a scale of 1 / 8; These multi-scale features can be used for subsequent tasks such as image segmentation, target detection, or super-resolution reconstruction because they can capture detailed information at different scales. The number of computing units in this part is 4, and the single-unit throughput is 1 pixel / clock cycle, supporting convolution, normalization, and activation operations.
[0074] Second level: negative correlation calculation unit The input features are first sent to a processing flow, the core of which is DSP calculation. In the DSP cluster calculation stage, the input features are first Perform a global average pooling operation to calculate a global feature representation Specifically, It is through All spatial positions (height and width ) is averaged, that is ; Next, based on Generate two vectors and ,vector is through two weight matrices and Linear transformation and nonlinear activation function and The specific calculation is ; The vector is passed through the weight matrix The linear transformation and hyperbolic tangent activation function tanh are obtained, and the calculation formula is ; After the DSP cluster calculation is completed, it enters the feature calibration stage. In this stage, the generated and With multi-scale features Perform a calibration operation; specifically, the calibrated features is by and Multiply element-wise, then add Obtained, that is ; This process can be regarded as a dynamic adjustment of the input features, where It plays a role in adjusting the feature weight, and A bias term is added to the feature; finally, after this series of processing, the calibrated multi-scale feature is output These features can be used in subsequent deep learning tasks such as classification, segmentation, or generation to improve the model's representation and generalization performance of input data. The number of DSP slices (Digital Signal Processing Slices) is 48, the parallelism is 16 channels / cycle, and the calculation accuracy is 8-bit fixed point.
[0075] Level 3: Rebuilding the Accelerator Core The input calibration features are the starting point of this part. After preliminary processing, these features are sent to the bilinear interpolation module. Here, the resolution of the feature map is increased by 2 or 4 times through upsampling. This process aims to improve the spatial resolution of the feature map so that subsequent operations can process image details more finely. After passing through the bilinear interpolation module, the features enter the channel attention hardware implementation stage. In this stage, the features are first globally average pooled to compress the spatial dimensions of the feature map into a global feature vector. This global feature vector is then linearly transformed through a fully connected layer (FC), and finally the output value is mapped to a range between 0 and 1 using the sigmoid activation function. The purpose of this process is to generate attention weights for each channel, thereby enhancing the features of important channels and suppressing unimportant channels. After completing the channel attention processing, the features enter the pixel-level fusion stage, where the calibration features of each scale are All the deconvolution operations are performed to further improve their spatial resolution; then, all the deconvolution features are fused through the summation operation to finally generate the output image. This fusion process ensures that features at different scales can work together to produce high-quality output at the pixel level. The entire process starts with input calibration features and ultimately generates a high-quality output image through gradual upsampling, channel attention mechanism, and pixel-level fusion. The synergy of multi-scale features and the guidance of the attention mechanism improves the model's representation and processing capabilities of input data. The pipelined bilinear interpolator in this part has a latency of <5ns, hardwareizing channel attention and reducing memory access by 90%.
[0076] Step 5.2: Hardware Optimization Step 5.2.1, 8-bit fixed-point quantization To reduce the model's storage space and computational complexity, the data in the model is quantized. Quantization means converting floating-point data into integer representations with a fixed number of bits while retaining as much valid information as possible. In this process, different data types (such as feature maps, weights, and biases) are assigned different numbers of bits based on their characteristics to ensure that the model's performance does not degrade significantly after quantization. Specifically, the quantization of the feature map is allocated as 1 sign bit + 3 integer bits + 4 decimal bits, which means that each quantized value of the feature map is represented by 8 bits, of which 1 bit is used to represent the sign (positive or negative), 3 bits are used to represent the integer part, and 4 bits are used to represent the decimal part. This allocation method can effectively reduce the storage space and computational complexity of the feature map while ensuring a certain accuracy. The quantization of weights is allocated as 1 sign bit + 2 integer bits + 5 decimal bits. Weights are parameters in neural networks and usually require higher precision to ensure model performance. Therefore, the quantization of weights allocates more decimal bits (5 bits) to retain more detailed information. At the same time, the integer bit of the weight is 2 bits, the sign bit is 1 bit, and the total is 8 bits. The quantization of the bias is allocated as 1 sign bit + 4 integer bits + 3 decimal bits. The bias is usually used to adjust the offset of the activation function in neural networks, and its value range may be large; therefore, the quantization of the bias is allocated more integer bits (4 bits) to ensure that a larger range of values can be represented; at the same time, the decimal bits of the bias are 3 bits and the sign bit is 1 bit, for a total of 8 bits; This bit allocation method can balance accuracy and storage efficiency during the quantization process. Although the quantization bit allocation of feature maps, weights and biases is different, they are all designed to ensure that the model can still maintain good performance after quantization while significantly reducing the storage and computing requirements of the model.
[0077] Step 5.2.2: Feature map block processing The input image size is 1920×1080 pixels. First, the input image is divided into 8×8 small blocks to form multiple image tiles. These image tiles are then assigned to different parallel processing units for processing; Specifically, the segmented image blocks are fed into four parallel processing units, labeled as parallel processing unit 1, parallel processing unit 2, parallel processing unit 3, and parallel processing unit 4; each processing unit independently processes the image blocks assigned to it, which may include various image processing operations such as filtering, feature extraction, or other computationally intensive tasks; The independent operation of the parallel processing units allows multiple image blocks to be processed simultaneously, significantly improving the overall processing speed and efficiency. Once all the parallel processing units have completed their respective tasks, the processed image blocks need to be reassembled or stitched together to form the complete output image. Finally, through the result stitching step, all processed image blocks are merged and restored to the original image size; this method not only speeds up image processing, but also allows the system to effectively utilize parallel computing resources such as multi-core processors or GPUs (Graphics Processing Units) to achieve high-performance image processing tasks.
[0078] Step 5.2.3, three-stage pipeline scheduling The clock frequency is 150HZ, and the scheduling mode is shown in Table 8: Table 8 Three-line assembly line scheduling mode
[0079] The detailed steps of step S6, real-time defogging reasoning, are as follows: Step 6.1. Input preprocessing 1) Standardize the input data format to adapt to hardware processing requirements The image input source is Gigabit Ethernet or camera interface, with a resolution of 1920×1080 (compatible with 720p / 1080p / 2K input) and a format of YUV420→RGB conversion (hardware accelerated). Normalization processing: (31) Where, Represents the pixel value of the original image, usually in the range of [0, 255]; Represents the normalized image pixel value, and the value range is mapped to [-1, 1].
[0080] Step 6.2: Multi-scale feature extraction This module gradually extracts image features through a series of convolutional layers, instance normalization, activation functions, and downsampling operations, and finally outputs multi-scale feature maps; The input image first passes through a processing unit consisting of a 3×3 convolution, instance normalization, and a ReLU activation function. This processing unit performs preliminary feature extraction on the image. The convolutional layer captures local features, instance normalization helps accelerate training and improve the model's generalization ability, and the ReLU activation function introduces nonlinearity, enabling the network to learn more complex feature representations. After completing the preliminary feature extraction, the image features pass through a downsampling layer, usually a maximum pooling or a convolution with a stride of 2, to reduce the spatial dimension of the feature map by a factor of 2. This downsampling operation helps reduce computational complexity and makes the feature map more abstract while retaining important feature information. The downsampled feature map goes through a similar processing unit again, which is another layer containing 3×3 convolution, instance normalization, and ReLU activation function. This process is repeated, and each processing is followed by another 2x downsampling operation to further extract higher-level features and reduce the spatial dimension of the feature map. After multiple convolution and downsampling operations, the network finally outputs a set of multi-scale feature maps, marked as , and These feature maps contain different levels of feature information from low-level to high-level, which can be used for subsequent image recognition, classification or other visual tasks.
[0081] Step 6.3: Dynamic Negative Correlation Feature Calibration First, input features It is converted into a global feature representation through the global average pooling operation ,Global average pooling is a downsampling technique that generates a single output value by calculating the average of all elements of the feature map, thereby reducing the dimension of the data and extracting the most important feature information; Next, based on the global features , generating two weights and ; Finally, multi-scale calibration is performed, which involves calibrating feature maps of different scales. , and Make adjustments, specifically: For the coarsest scale features , using an upsampling operation ( and ) for weight and To enlarge the space size, and then Perform element-by-element multiplication and addition operations to obtain the calibrated features ; For intermediate-scale features , use directly and and Perform element-by-element multiplication and addition operations to obtain the calibrated features ; For the finest scale features , using a downsampling operation ( and ) for weight and Reduce the space size and then Perform element-by-element multiplication and addition operations to obtain the calibrated features ; In this way, the model is able to fine-tune features at different scales, thereby improving the accuracy of feature expression and model performance.
[0082] Step 6.4: Image reconstruction and output Decoding the calibration features produces the final dehazing result: First, the calibrated feature map from the third layer Initially, it is processed through a 3×3 deconvolution layer; deconvolution is usually used for upsampling, that is, increasing the spatial dimension of the feature map, which helps to restore the details of the image; after deconvolution, the feature map then passes through the channel attention (CA) and pixel attention (PA) modules; the channel attention module can enhance important feature channels while suppressing unimportant channels, while the pixel attention module focuses on improving the importance of each pixel in the feature map. After these processes, the resulting feature map is labeled ; Next, Feature map after calibration with the second layer Perform element-by-element addition, which realizes the fusion of feature maps and combines feature information at different levels; the fused feature map passes through a 3×3 deconvolution layer again, and then passes through the channel attention and pixel attention modules to obtain a new feature map ; at last, Feature map after calibration with the first layer Perform element-by-element addition to further fuse features; this fused feature map is also processed by the 3×3 deconvolution layer and the channel attention and pixel attention modules to finally obtain the output feature map ; The entire process aims to gradually restore the spatial resolution of the image while improving the expressiveness of the feature map through layer-by-layer feature fusion and the introduction of the attention mechanism. This method not only helps to retain and enhance important information in the image, but also effectively combines features of different scales, thereby improving the performance of the model.
[0083] Example 4 In another preferred embodiment, based on the above-mentioned embodiment 3, this embodiment provides a system for implementing the real-time defogging method for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling described in embodiments 1 and 2, namely, a real-time defogging system for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling. The specific implementation method is as follows: 1. System architecture design: Data input interface: Design a data input interface for receiving input aerial images. This interface should support the simultaneous input of multiple image data channels and adopt a ping-pong buffer mechanism to eliminate transmission delays. The ping-pong buffer mechanism ensures the continuity of data during transmission and processing by alternating between two buffer areas. Three-stage pipeline processing module: including feature extraction engine, negative correlation calculation unit and reconstruction acceleration core. These three modules process the input image in sequence to form a pipeline operation and improve processing efficiency. Output interface: Design an output interface to output the dehazed image. The output interface can transmit the processed image to a display device or storage device.
[0084] 2. Feature extraction engine implementation: Parallel computing unit: The feature extraction engine consists of multiple parallel computing units, each of which uses a convolutional neural network (CNN) structure. These computing units can process different regions or different feature channels of the input image in parallel, improving the feature extraction speed; Feature extraction process: The input aerial image passes through the feature extraction engine to extract multi-scale features; these features include edge, texture, color and other information, providing a basis for subsequent processing.
[0085] 3. Implementation of negative correlation calculation unit: DSP cluster computing: The negative correlation calculation unit generates a negative correlation weight matrix based on DSP cluster computing. DSP has high-speed computing capabilities and is suitable for processing complex mathematical operations. Weight Matrix Generation: This unit receives the multi-scale features output by the feature extraction engine and generates a negative correlation weight matrix using the atmospheric light features. The atmospheric light features can be obtained by estimating the atmospheric light values. The negative correlation weight matrix is generated based on the negative correlation between the atmospheric light values and the transmission map. Feature calibration: Adaptively calibrate the multi-scale transmission map features and use the generated negative correlation weight matrix to weight the features, which helps eliminate the impact of fog on the image and improve the dehazing effect.
[0086] 4. Rebuild the acceleration kernel to achieve: Bilinear interpolation: The reconstruction acceleration kernel quickly restores the spatial resolution of an image through operations such as bilinear interpolation. Bilinear interpolation is a commonly used image scaling method that estimates the value of a new pixel based on the values of surrounding pixels. Channel attention and pixel-level fusion: The channel attention mechanism is introduced to perform weighted fusion of features from different channels to improve feature representation capabilities. At the same time, a pixel-level fusion strategy is adopted to fuse features of different scales to generate more refined dehazed images. Image reconstruction: After a series of processing, the acceleration kernel is reconstructed to generate a clear image after defogging. The image should have high visual quality and detail retention; 5. System integration and testing: System Integration: The feature extraction engine, negative correlation calculation unit, and reconstruction acceleration core are integrated into the FPGA platform to form a complete hardware acceleration system. FPGAs are highly flexible and configurable, making them suitable for implementing complex digital signal processing algorithms.
[0087] System testing: The system is tested to verify its defogging effectiveness and real-time processing capabilities. During testing, aerial images of varying fog concentrations and scenes are input to observe the system's defogging effectiveness and real-time performance. On the Xilinx ZU9EG platform, the system can achieve real-time processing at 1080p@45fps, with power consumption reduced to 4.8W, meeting the fast processing requirements of edge devices such as drones and satellite remote sensing.
[0088] Example 3 In another preferred embodiment, based on the above-mentioned embodiment 2, the physical model derivation of the dynamic negative correlation coupling of the real-time defogging method for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling of the present invention is described as follows: 1. Atmospheric scattering model and negative correlation theory The traditional atmospheric scattering model describes the fog formation process as follows: (32) Where, is the observed fog map; Radiate for real scenes; For the transmission map, and scene depth Negatively correlated with the scattering coefficient β; Global atmospheric light.
[0089] Key physical constraints: Transmission diagram With atmospheric light There is an explicit negative correlation: when → 0:00 , the higher the fog concentration, the stronger the dominance of atmospheric light; when →1 o'clock: , the influence of atmospheric light can be ignored.
[0090] Traditional independent estimation and , ignoring the correlation between the two, resulting in: 1) The sky area is white: Overrated, Failure to make corresponding adjustments is a violation physical constraints; 2) Edge halo artifacts: When the object boundary suddenly changes, No dynamic adaptation occurs, resulting in distortion.
[0091] The present invention explicitly encodes the weight generator and Negative correlation: 1. Atmospheric light feature extraction Output features to the encoder Compress spatial dimensions while preserving channel semantics: (33) Where, represents the feature vector after global average pooling processing, is the input feature map.
[0092] 2. Negative correlation weight generation Generate dynamic weights through fully connected layers and dual-path activation functions: (34) Where, represents the weights in the neural network, Represents the bias parameter in the neural network.
[0093] 3. Transmission map feature calibration Apply weights to multi-scale transfer graph features: (35) (36) Where, Dynamic Suppression Zhongyu Positively correlated channels (weights in dense fog areas approach 0); Provides spatially adaptive compensation to alleviate over-suppression. Indicates that at time step and scale The following feature map; Indicates that after processing at time step and scale The following feature map.
[0094] Comparison and verification with traditional methods: The error difference between independent estimation and dynamic coupling is quantified by simulation experiments as shown in Table 9: Table 9 Comparison of the error difference between independent estimation and dynamic coupling in simulation experiments
[0095] The dynamic coupling mechanism reduces the physical model error by 67%, verifying its necessity.
[0096] Quantitative comparison experimental settings with mainstream methods: The dataset uses the standard aerial image dehazing dataset NH-HAZE, which includes light fog, dense fog, and mixed fog scenes.
[0097] Comparison method: FFA-Net (Feature Fusion Attention Network): an end-to-end dehazing network based on channel attention; DehazeFormer (Dehazing Transformer Network): a method that combines Transformer with physical models; C2PNet (Curricular Contrastive Regularization for Physics-aware Single Image Dehazing Network): a method that explicitly models the atmospheric scattering model; Evaluation Metrics: Restoration quality: PSNR, SSIM; Hardware efficiency: FLOPs (Floating Point Operations), memory usage, and latency (at 1080p resolution); Platform: NVIDIA RTX 3090 GPU; The results are shown in Table 10: Table 10 Comparison of evaluation results
[0098] Algorithm performance comparison results: The PSNR of the proposed method is improved by 1.74dB and the SSIM is improved by 0.027 compared with FFA-Net, especially in dense fog areas (PSNR difference > 2.1dB); The computational effort (FLOPs) is only 7.7% of DehazeFormer, and the memory usage is reduced by 69%.
[0099] The results are shown in Table 11 when tested on the satellite remote sensing dataset RS-Haze, which was not used in the training. Table 11 Test results
[0100] Bidirectional degradation modeling compresses cross-scenario PSNR fluctuations to ±0.8dB, significantly outperforming traditional methods.
[0101] In the preferred solution, the dynamic negative correlation coupling module in step 1 explicitly encodes the physical negative correlation between atmospheric light and the transmission map, enabling adaptive calibration of multi-scale features. This setup not only enhances the image dehazing algorithm's adaptability to complex scenes but also significantly improves the naturalness and clarity of the dehazing effect. Furthermore, this module, through deep learning optimization, further accelerates processing speed, ensuring the feasibility of real-time applications.
[0102] In the preferred solution, the bidirectional degradation modeling network framework in step 1 includes a forward defogging path and a reverse defogging path, and the degradation mechanism modeling is realized through physical reversible constraints; the above settings ensure that the image can be effectively defogged during the forward defogging process, while the reverse defogging path can verify the authenticity of the defogging effect. The two-way paths work together to improve the accuracy and efficiency of image degradation and restoration.
[0103] In the preferred solution, the multi-objective loss function in step 6 includes dehazing loss, physical consistency loss, perceptual loss, and multi-scale structural similarity loss. The above setting can comprehensively evaluate the dehazing effect, ensuring that the generated image conforms to physical laws and is close to a real clear image in visual perception. At the same time, the multi-scale structural similarity loss further improves the quality of image detail restoration.
[0104] In the preferred solution, the feature extraction engine is composed of multiple parallel computing units for extracting multi-scale features of the input image. The above setting can significantly improve the efficiency and accuracy of feature extraction. Each computing unit focuses on feature analysis at different scales, thereby comprehensively capturing image details and laying a solid foundation for subsequent image recognition or classification tasks.
[0105] In the preferred solution, the negative correlation calculation unit generates a negative correlation weight matrix based on the DSP cluster calculation of atmospheric light characteristics, and adaptively calibrates the multi-scale transmission map features; the above settings can effectively reduce the impact of illumination changes on image quality and enhance the image detail performance; at the same time, by introducing a deep learning algorithm, the negative correlation weight distribution is further optimized to achieve accurate extraction and efficient processing of image transmission map features, which not only effectively improves the accuracy and efficiency of image defogging, but also ensures the clear restoration of image details under different lighting conditions, further optimizes the visual experience, and brings innovative breakthroughs to the field of image processing.
[0106] In the preferred solution, the reconstruction acceleration kernel quickly reconstructs the dehazed image through operations such as bilinear interpolation, channel attention, and pixel-level fusion. The above settings effectively improve the speed and quality of image reconstruction, making the dehazed image clearer and more natural. At the same time, the solution further reduces computing resource consumption by optimizing the algorithm structure, achieving high-efficiency image processing.
[0107] In a preferred solution, the three-stage pipeline processing module is implemented using an FPGA, supporting 8-bit fixed-point quantization and block processing. This significantly improves data processing efficiency and reduces resource consumption. Furthermore, the module integrates an efficient memory management mechanism, ensuring high-speed data access and smooth processing.
[0108] In a preferred solution, the system also includes a ping-pong cache mechanism to eliminate data transmission delays and ensure real-time processing; this configuration effectively improves data transmission efficiency and stability. Furthermore, the mechanism incorporates a built-in intelligent scheduling algorithm that automatically optimizes data flow paths, further shortening response times and providing users with a seamless and smooth experience.
[0109] In the preferred solution, the system achieves real-time processing of 1080p@45fps on the Xilinx ZU9EG platform, with power consumption reduced to 4.8W. This setup fully utilizes the platform's hardware acceleration capabilities while further improving processing efficiency through algorithm optimization. Furthermore, the solution is highly scalable and flexible, easily adapting to the needs of different application scenarios.
[0110] In the preferred solution, the method or system is suitable for real-time dehazing of aerial images of edge devices such as drones and satellite remote sensing. The above settings can effectively improve image clarity and enhance target recognition accuracy. Under complex meteorological conditions, the solution can still work stably, ensuring the high quality and real-time nature of aerial monitoring data, and providing strong support for environmental protection, disaster warning and other fields.
[0111] In summary, the proposed real-time dehazing method and system for aerial imagery based on dynamic negative correlation coupling and bidirectional degradation modeling provides an effective dehazing solution for the field of aerial image processing technology, particularly for applications such as drones, satellite remote sensing, and aerial photography. This invention addresses key issues inherent in existing dehazing methods when processing complex aerial environments, including color distortion, insufficient detail preservation, limited model generalization, difficulty meeting real-time requirements, and inaccurate degradation mechanism modeling. It offers comprehensive and in-depth improvements.
[0112] In terms of technical implementation, the present invention innovatively introduces a dynamic negative correlation coupling mechanism, which achieves adaptive calibration of multi-scale features by explicitly encoding the physical negative correlation between atmospheric light and the transmission map. This mechanism can automatically adjust feature weights based on different fog concentration areas, thereby ensuring accurate restoration of image details even in complex fog environments, significantly improving the robustness and accuracy of the defogging effect. At the same time, the present invention constructs a bidirectional degradation modeling framework that includes a forward defogging path and a reverse defogging path. By strengthening the degradation mechanism modeling through physically reversible constraints, it not only improves the defogging effect but also enhances the model's generalization ability.
[0113] To meet real-time requirements, the present invention designs a three-stage pipeline hardware acceleration system based on FPGA, which supports 8-bit fixed-point quantization and block processing. Through customized architecture design, the system significantly improves processing speed and efficiency, meeting the fast processing requirements of edge devices such as drones and satellite remote sensing. In addition, the present invention also combines physical models with deep learning technology, and through dynamic negative correlation coupling and bidirectional degradation modeling, realizes the organic combination of physical constraints and data-driven, which not only utilizes the prior knowledge of the physical model, but also gives play to the powerful learning ability of deep learning, thereby further improving the defogging effect and model generalization.
[0114] In terms of optimization strategy, the present invention defines a multi-objective loss function, including defogging loss, physical consistency loss, perception loss, and multi-scale structural similarity loss, and implements joint optimization training. This multi-objective joint optimization strategy can comprehensively consider multiple aspects such as defogging effect, physical constraints, detail recovery, and structure preservation, ensuring that the model is optimized in all aspects. At the same time, the hardware acceleration system designed by the present invention adopts a three-level pipeline architecture and ping-pong cache mechanism to achieve efficient data processing and transmission, reduce power consumption and cost, and provide a feasible hardware solution for real-time defogging processing. The system also has good scalability and configurability, can adapt to the needs of different application scenarios, and shows a wide range of application prospects.
Claims
1. A real-time dehazing method for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling, characterized by: The following steps are involved: Step 1: Construct a bidirectional degradation modeling network framework, which includes an encoder module, a dynamic negative correlation coupling module, a decoder module, an inverse atomization module, and a joint optimization module; Step 2: Use the encoder module to extract multi-scale features of the input aerial image; Step 3: Generate a negative correlation weight matrix based on atmospheric light characteristics through the dynamic negative correlation coupling module to adaptively calibrate the multi-scale transmission map features; Step 4: Use the decoder module to reconstruct the dehazed image based on the calibrated features; Step 5: Through the reverse atomization module, build a parameter estimation subnetwork to implement the reverse atomization path and strengthen the physical model constraints; Step 6: Use the joint optimization module to define a multi-objective loss function and implement joint optimization training to improve the dehazing effect and model generalization ability.
2. The real-time dehazing method for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling according to claim 1 is characterized by: The dynamic negative correlation coupling module of step 1 explicitly encodes the physical negative correlation between atmospheric light and transmission map, and realizes adaptive calibration of multi-scale features.
3. According to claim 1, the method for real-time dehazing of aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling is characterized by: The bidirectional degradation modeling network framework in step 1 includes a forward defogging path and a reverse fogging path, and the degradation mechanism modeling is achieved through physical reversible constraints.
4. The real-time dehazing method for aerial images based on dynamic negative correlation coupling and bidirectional degradation modeling according to claim 3 is characterized by: The multi-objective loss function in step 6 includes dehazing loss, physical consistency loss, perception loss and multi-scale structural similarity loss.
5. A real-time aerial image defogging system based on dynamic negative correlation coupling and bidirectional degradation modeling is a system that implements the real-time aerial image defogging method based on dynamic negative correlation coupling and bidirectional degradation modeling as described in claim 4, characterized in that: include: A data input interface for receiving input aerial images; A three-stage pipeline processing module, including a feature extraction engine, a negative correlation calculation unit, and a reconstruction acceleration core; Output interface, used to output the dehazed image.
6. The real-time aerial image defogging system based on dynamic negative correlation coupling and bidirectional degradation modeling according to claim 5 is characterized by: The feature extraction engine is composed of multiple parallel computing units and is used to extract multi-scale features of the input image.
7. The real-time aerial image defogging system based on dynamic negative correlation coupling and bidirectional degradation modeling according to claim 6 is characterized by: The negative correlation calculation unit generates a negative correlation weight matrix based on the atmospheric light feature DSP cluster calculation, and performs adaptive calibration on the multi-scale transmission map features.
8. The real-time aerial image defogging system based on dynamic negative correlation coupling and bidirectional degradation modeling according to claim 7 is characterized by: The reconstruction acceleration kernel quickly reconstructs the dehazed image through operations such as bilinear interpolation, channel attention, and pixel-level fusion.
9. The real-time aerial image defogging system based on dynamic negative correlation coupling and bidirectional degradation modeling according to claim 8 is characterized by: The three-stage pipeline processing module is implemented using FPGA and supports fixed-point quantization and block processing.
10. The real-time aerial image defogging system based on dynamic negative correlation coupling and bidirectional degradation modeling according to claim 9 is characterized by: The system also includes a ping-pong buffer mechanism for eliminating data transmission delays and ensuring real-time processing; the system implements 1080p@45fps real-time processing on the Xilinx ZU9EG platform to reduce power consumption.