A Video Dehazing Network Training Method Based on Bidirectional Optical Flow Loss
Through the video defog network training method based on bidirectional optical flow loss, the pre-trained video defog network is fine-tuned, which solves the problems of high computing complexity and inability to take into account both the defog effect, real-time and low power consumption in the prior art, and achieves an efficient and low-power video defog effect.
Patent Information
- Application Number
- CN202510466660.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing video defog removal technology has high computational complexity when processing high frame rate and high resolution videos, making it difficult to achieve efficient defog removal under limited computing resources, and cannot take into account both the defog removal effect and real-time performance, while meeting the needs of portable devices for low power consumption.
The video defogging network training method based on bidirectional optical flow loss is adopted. Through multi-stage convolutional encoding and upsampling decoding, combined with smooth L1 loss, perceptual loss and bidirectional optical flow loss, the pre-trained video defogging network is fine-tuned and end-to-end training, and a fully trained video defogging network is obtained.
It effectively improves the accuracy and time consistency of the video defog network in detail, reduces the amount of calculation, and achieves the goal of taking into account the defog effect, low power consumption and real-time.
Smart Images

Figure CN119992259B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent vision technology, and in particular, to a method for training a video dehazing network based on bidirectional optical flow loss. Background Art
[0002] In modern intelligent vision systems, the quality of video images is affected by complex weather conditions, which seriously reduces the recognition and detection capabilities of the vision system for scenes. In this regard, video dehazing technology, as a very important technology, is widely used in fields such as autonomous driving, drone navigation, and intelligent monitoring. Currently, in existing video dehazing research, the proportion of deep learning methods in video dehazing is increasing continuously. Among them, the convolutional neural network CNN is one of the typical methods. By extracting features from the input image, frame-by-frame dehazing processing is realized, and finally the processed images are stitched into a complete video. This method performs well in single-frame images. However, in the video dehazing scenario, frame-by-frame processing ignores the dynamic information between frames, resulting in insufficient temporal consistency, uneven transitions between video frames, and easy occurrence of problems such as jumps or flickers. To address this problem, some research introduces a generative adversarial network, and through the adversarial learning of the generator and the discriminator, a clearer dehazed image is generated. Or an attention mechanism. By combining channel attention and spatial attention, the model can better focus on the key information in the image, thereby improving the dehazing effect. Some research also introduces temporal processing modules such as long short-term memory network LSTM. Such modules can capture the temporal information between video frames and ensure smooth transitions between frames. In addition, the temporal processing method of 3D convolutional network is also widely used in video dehazing tasks. The 3D convolutional network enhances temporal consistency by adding convolutional operations in the time dimension.
[0003] However, in existing methods for improving the dehazing effect and temporal consistency, whether it is a generative adversarial network, an attention mechanism, a long short-term memory network, or a 3D convolutional network, they all have a high computational complexity. Especially when processing high-frame-rate and high-resolution videos, the consumption of computing resources increases significantly. In the application scenarios of portable devices such as embedded devices, it is impossible to meet the low-power performance requirements while ensuring the dehazing effect and real-time performance.
[0004] Therefore, there are technical problems existing that the computational complexity is relatively high, it is difficult to achieve efficient video dehazing under limited computing resources when processing high-frame-rate and high-resolution videos, and it is impossible to balance the dehazing effect and real-time performance while meeting the low-power performance requirements of portable devices such as embedded devices, which need to be improved. Summary of the Invention
[0005] In view of this, it is necessary to provide a method for training a video dehazing network based on bidirectional optical flow loss to solve the technical problems existing in the prior art, such as high computational complexity, difficulty in achieving efficient video dehazing under limited computational resources when processing high-frame-rate and high-resolution videos, and inability to balance the dehazing effect and real-time performance while meeting the low-power performance requirements of portable devices such as embedded devices.
[0006] On the one hand, to solve the above technical problems, the present invention provides a method for training a video dehazing network based on bidirectional optical flow loss, including:
[0007] Input the obtained hazy video data and corresponding real haze-free video data into a pre-trained video dehazing network, perform multi-level convolutional encoding and multi-level upsampling decoding on the hazy video data to obtain dehazing output data;
[0008] Determine the smooth L1 loss according to the dehazing output data and the real haze-free video data, perform multi-level convolutional encoding on the dehazing output data and the real haze-free video data to obtain dehazing perception features and real perception features, determine the perception loss according to the dehazing perception features and the real perception features, and perform bidirectional optical flow estimation based on an adaptive pyramid and cost volume improvement on the dehazing output data to obtain the bidirectional optical flow loss;
[0009] Iteratively perform parameter fine-tuning and end-to-end training on the pre-trained video dehazing network according to the smooth L1 loss, the perception loss, and the bidirectional optical flow loss to obtain a trained video dehazing network.
[0010] In a possible implementation, the pre-trained video dehazing network is obtained by training an initial dehazing network. Training the initial dehazing network includes:
[0011] Input the obtained hazy image data and corresponding real haze-free image data into the initial dehazing network, perform multi-level convolutional encoding and multi-level upsampling decoding on the hazy image data to obtain dehazing output image data;
[0012] Determine the pre-training smooth L1 loss according to the dehazing output image data and the real haze-free image data, perform multi-level convolutional encoding on the dehazing output image data and the real haze-free image data to obtain dehazing image perception features and real haze-free image perception features, determine the pre-training perception loss according to the dehazing image perception features and the real haze-free image perception features, and iteratively train the initial dehazing network according to the pre-training smooth L1 loss and the pre-training perception loss to obtain the pre-trained video dehazing network.
[0013] In a possible implementation, the multi-level convolutional encoding includes an initial convolutional layer and several inverted residual blocks, and each inverted residual block includes an expansion convolution and a depthwise separable convolution.
[0014] In a possible implementation, the multi-level upsampling decoding includes a deconvolution upsampling operation layer by layer and a batch normalization operation, and the output features of the batch normalization operation at each level are fused with the output features of the corresponding inverted residual block through skip connection feature fusion as the input features of the next deconvolution upsampling operation.
[0015] In a possible implementation, a pre-training smooth L1 loss is determined based on the defogged output image data and the real fog-free image data. The defogged output image data and the real fog-free image data are subjected to multi-level convolutional encoding to obtain defogged image perception features and real fog-free image perception features. The pre-training perception loss is determined based on the defogged image perception features and the real fog-free image perception features, including:
[0016] Performing pixel difference calculation on the defogged output image data and the real fog-free image data based on the mean square error method and the L1 loss to obtain the pre-training smooth L1 loss;
[0017] Performing multi-level convolutional encoding on the defogged output image data and the real fog-free image data to obtain defogged image perception features and real fog-free image perception features;
[0018] Performing feature difference calculation on the defogged image perception features and the real fog-free image perception features based on the Euclidean distance method to obtain the pre-training perception loss.
[0019] In a possible implementation, performing bidirectional optical flow estimation based on an adaptive pyramid and cost volume improvement on the defogged output data to obtain a bidirectional optical flow loss, including:
[0020] Arranging the defogged output data in the order of video frames, and taking every three consecutive frames as a group of optical flow input data, where each group of optical flow input data includes a forward frame image, a current frame image, and a backward frame image;
[0021] Inputting the optical flow input data into a preset adaptive feature pyramid for multi-scale feature extraction, deconvolution upsampling, and feature fusion to obtain optical flow features;
[0022] Performing cost volume method optical flow estimation and frame warping operation on the optical flow features to obtain a forward prediction frame image and a backward prediction frame image;
[0023] Determining the forward optical flow loss based on the forward prediction frame image and the current frame image, determining the backward optical flow loss based on the backward prediction frame image and the current frame image, and combining the forward optical flow loss and the backward optical flow loss to obtain the bidirectional optical flow loss.
[0024] In a possible implementation, iteratively performing parameter fine-tuning and end-to-end training on the pre-training video defogging network according to the smooth L1 loss, the perception loss, and the bidirectional optical flow loss to obtain a trained complete video defogging network, including:
[0025] Determine the total defogging loss based on the smooth L1 loss, perceptual loss, and bidirectional optical flow loss;
[0026] Keep the network parameters of other parts of the pre-trained defogging network unchanged, and adjust the network parameters of each preset part to be adjusted in the pre-trained defogging network in turn according to the total defogging loss;
[0027] Perform end-to-end training on the pre-trained video defogging network according to the total defogging loss, using the cosine annealing strategy, learning rate warm-up strategy, and gradient clipping strategy, and obtain a trained video defogging network through the iterative training process.
[0028] On the other hand, the present invention also provides a method for applying a video defogging network based on bidirectional optical flow loss, including:
[0029] Obtain the video to be defogged;
[0030] Input the video to be defogged into the trained video defogging network to obtain a defogged output video;
[0031] Wherein, the trained video defogging network is determined according to the above-mentioned video defogging network training method based on bidirectional optical flow loss.
[0032] On the other hand, the present invention also provides an embedded device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned video defogging network training method based on bidirectional optical flow loss and / or the above-mentioned method for applying a video defogging network based on bidirectional optical flow loss.
[0033] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned video defogging network training method based on bidirectional optical flow loss and / or the above-mentioned method for applying a video defogging network based on bidirectional optical flow loss.
[0034] The beneficial effects of the present invention are as follows: The video haze removal network training method based on bidirectional optical flow loss provided by the embodiments of the present invention first inputs the obtained hazy video data and corresponding real haze-free video data into a pre-trained video haze removal network, performs multi-level convolutional encoding and multi-level upsampling decoding on the hazy video data to obtain haze removal output data; then determines the smooth L1 loss according to the haze removal output data and the real haze-free video data, performs multi-level convolutional encoding on the haze removal output data and the real haze-free video data to obtain haze-aware features and real-aware features, determines the perceptual loss according to the haze-aware features and the real-aware features, and performs bidirectional optical flow estimation based on an adaptive pyramid and cost volume improvement on the haze removal output data to obtain the bidirectional optical flow loss; finally, iteratively performs parameter fine-tuning and end-to-end training on the pre-trained video haze removal network according to the smooth L1 loss, the perceptual loss, and the bidirectional optical flow loss to obtain a trained video haze removal network. During the training process of the present invention, the smooth L1 loss is used to calculate the loss at the pixel level, the perceptual loss is used to calculate the loss at the feature level to improve the accuracy of the network in details, the bidirectional optical flow loss is obtained by improving the bidirectional optical flow estimation to capture the dynamic information between video frames and improve the temporal consistency of the network in the video haze removal task, so that the video haze removal network can effectively improve the haze removal effect with a relatively low computational amount, and obtain a video haze removal network that takes into account the haze removal effect, low power consumption, and real-time performance. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a schematic flowchart of an embodiment of the video haze removal network training method based on bidirectional optical flow loss provided by the present invention;
[0037] Figure 2 It is a schematic flowchart of training the initial haze removal network in the embodiments of the present invention;
[0038] Figure 3 It is a schematic flowchart of calculating the smooth L1 loss and the perceptual loss in the embodiments of the present invention;
[0039] Figure 4 It is a schematic flowchart of the iterative training of the pre-trained video haze removal network in the embodiments of the present invention;
[0040] Figure 5 It is a schematic flowchart of an embodiment of the video haze removal network application method based on bidirectional optical flow loss provided by the present invention;
[0041] Figure 6Schematic structural diagram of an embodiment of the embedded device provided by the present invention. Detailed implementation manners
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0043] In the description of the embodiments of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example: A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0044] The descriptions such as "first" and "second" involved in the embodiments of the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one such feature.
[0045] Referring to "embodiment" in this article means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0046] The present invention provides a video defogging network training method based on bidirectional optical flow loss. The video defogging network training method based on bidirectional optical flow loss will be described separately below.
[0047] Figure 1 It is a schematic flowchart of an embodiment of the video defogging network training method based on bidirectional optical flow loss provided by the present invention. As Figure 1 shown, the video defogging network training method based on bidirectional optical flow loss includes:
[0048] S101. Input the obtained foggy video data and corresponding real fog-free video data into the pre-trained video defogging network, and perform multi-level convolutional encoding and multi-level upsampling decoding on the foggy video data to obtain defogged output data;
[0049] S102. Determine the smooth L1 loss based on the defogging output data and the real fog-free video data, perform multi-level convolutional encoding on the defogging output data and the real fog-free video data to obtain defogging perception features and real perception features, determine the perception loss based on the defogging perception features and the real perception features, and perform bidirectional optical flow estimation on the defogging output data based on the improved adaptive pyramid and cost volume to obtain the bidirectional optical flow loss;
[0050] S103. Iteratively perform parameter fine-tuning and end-to-end training on the pre-trained video defogging network according to the smooth L1 loss, the perception loss, and the bidirectional optical flow loss to obtain a trained complete video defogging network.
[0051] Compared with the prior art, the video defogging network training method based on the bidirectional optical flow loss provided by the embodiment of the present invention first inputs the obtained foggy video data and the corresponding real fog-free video data into the pre-trained video defogging network, performs multi-level convolutional encoding and multi-level upsampling decoding on the foggy video data to obtain the defogging output data; then determines the smooth L1 loss according to the defogging output data and the real fog-free video data, performs multi-level convolutional encoding on the defogging output data and the real fog-free video data to obtain defogging perception features and real perception features, determines the perception loss according to the defogging perception features and the real perception features, and performs bidirectional optical flow estimation on the defogging output data based on the improved adaptive pyramid and cost volume to obtain the bidirectional optical flow loss; finally, iteratively perform parameter fine-tuning and end-to-end training on the pre-trained video defogging network according to the smooth L1 loss, the perception loss, and the bidirectional optical flow loss to obtain a trained complete video defogging network. In the training process of the present invention, the smooth L1 loss is used to calculate the loss at the pixel level, the perception loss is used to calculate the loss at the feature level to improve the accuracy of the network in details, the improved bidirectional optical flow estimation is used to obtain the bidirectional optical flow loss to capture the dynamic information between video frames and improve the temporal consistency of the network in the video defogging task, so that the video defogging network can effectively improve the defogging effect while having a lower computational amount, and obtain a video defogging network that takes into account the defogging effect, low power consumption, and real-time performance.
[0052] In some embodiments of the present invention, Figure 2 is a schematic flowchart of training an initial defogging network according to an embodiment of the present invention. As Figure 2 shown, the pre-trained video defogging network is obtained by training the initial defogging network. The initial defogging network includes:
[0053] S201. Input the obtained foggy image data and the corresponding real fog-free image data into the initial defogging network, perform multi-level convolutional encoding and multi-level upsampling decoding on the foggy image data to obtain the defogging output image data;
[0054] S202. Determine the pre-trained smooth L1 loss based on the dehazed output image data and the ground-truth haze-free image data. Perform multi-level convolutional encoding on the dehazed output image data and the ground-truth haze-free image data to obtain the dehazed image perceptual features and the ground-truth haze-free image perceptual features. Determine the pre-trained perceptual loss based on the dehazed image perceptual features and the ground-truth haze-free image perceptual features. Iteratively train the initial dehazing network according to the pre-trained smooth L1 loss and the pre-trained perceptual loss to obtain the pre-trained video dehazing network.
[0055] Specifically, considering that in the video dehazing task, there are relatively few high-quality public video datasets, which are difficult to provide sufficient training samples for deep learning models. Therefore, in this embodiment, through the pre-training method, the complete encoder layer and decoder layer are pre-trained on a large image dataset first, so that the model can effectively capture and learn the feature representations under different haze conditions. Subsequently, the pre-trained video dehazing network is transferred to the video dehazing task through transfer learning and fine-tuned on a small-scale video dataset, solving the problem of the lack of high-quality public video datasets.
[0056] During the pre-training process, for the image dataset, the embodiment first performs format unification processing on the image dataset. All images are uniformly converted into the standard RGB color space, and the image resolution is adjusted to 512×512 to meet the input requirements of the model. Then, for different foggy conditions in the image dataset, the dataset is classified based on the fog degree of the images, generating three different levels of sub-datasets: light fog, medium fog, and heavy fog, which is convenient for the model to specifically learn the dehazing effects of different foggy scenarios.
[0057] In addition, to enhance the robustness of the model, the image dataset also undergoes image enhancement processing, including image rotation, translation, mirror flipping, and brightness adjustment, to improve the generalization ability of the model to different perspectives and lighting conditions. After image enhancement, the embodiment divides the image dataset into a training set, a validation set, and a test set. The training set is used for the parameter learning of the model, the validation set is used for the evaluation of the model performance during the training process, and the test set is used for the evaluation of the final model effect. Through these preprocessing steps, it is ensured that the image data can effectively support the training and performance verification of the dehazing model.
[0058] In the pre-training loss calculation, the embodiments use the smooth L1 loss and the perceptual loss. The smooth L1 loss calculates the pixel differences between the defogged output image data and the real fog-free image output data at the pixel level to determine the pre-training smooth L1 loss. The perceptual loss first performs multi-level convolutional encoding on the defogged output image data and the real fog-free image data to obtain the defogged image perceptual features and the real fog-free image perceptual features, and then calculates the feature differences between the defogged image perceptual features and the real fog-free image perceptual features at the feature level to determine the pre-training perceptual loss. After iteratively training the initial defogging network according to the pre-training smooth L1 loss and the pre-training perceptual loss, a pre-trained video defogging network is obtained.
[0059] In some embodiments of the present invention, the multi-level convolutional encoding includes an initial convolutional layer and several inverted residual blocks, and each inverted residual block includes an expansion convolution and a depthwise separable convolution.
[0060] Specifically, considering that the traditional convolutional neural network has a large computational load and it is usually difficult to meet the real-time requirements in actual application scenarios. In embedded devices or low-power devices, due to limited computing resources, a large computational load is difficult to operate efficiently on the hardware platform. Therefore, the embodiments design a lightweight encoder based on MobileNetV2 to efficiently extract multi-scale features of the input image.
[0061] In the encoder, for the input image after preprocessing and normalization , first, low-level feature extraction is performed through the initial convolutional layer. The initial convolutional operation can be expressed by the formula:
[0062]
[0063] where represents the initial feature map, and is the weight matrix of the initial convolutional layer.
[0064] Then, for the initial feature map after the initial convolutional processing, the embodiments use six inverted residual blocks to gradually extract multi-scale high-level features , and the obtained series of multi-scale feature maps are used as the encoded features, and the design of the inverted residual blocks makes the size of the feature maps gradually decrease. Each inverted residual block is composed of an expansion convolution and a depthwise separable convolution.
[0065] Among them, the main function of the expansion convolution is to expand the number of input channels through a 1×1 convolutional operation, allowing subsequent convolutional operations to perform feature extraction in a higher-dimensional space. This convolutional operation can not only increase the feature expression ability of the model but also does not change the spatial dimension of the input feature map. The expansion convolutional operation can be expressed by the formula:
[0066]
[0067] Among them, is the weight matrix of the channel expansion convolution, which is used to expand the number of channels of the input feature map from the original lower number of channels to a higher-dimensional feature space, so as to enhance the model's ability to extract features on different channels.
[0068] Then, perform depth convolution on the expanded feature map Depth convolution performs convolution operations independently on each channel, focusing on capturing the spatial information in the image without interacting the information between channels. This convolution method significantly reduces the computational complexity by avoiding cross-channel calculations. The depth convolution formula is expressed as:
[0069]
[0070] Among them, is the convolution kernel of the depth convolution. The convolution kernel independently performs 3×3 convolution operations on each input channel to extract spatial features. This operation can retain the spatial dimension of the feature map while greatly reducing the number of parameters and the amount of computation.
[0071] After the depth convolution is completed, the feature map undergoes channel fusion and adjustment through pointwise convolution. The role of pointwise convolution is to map the expanded number of channels to the specified number of output channels and fuse the features in different channels through cross-channel convolution operations. This step of operation can further reduce the computational overhead while keeping the spatial dimension of the feature map unchanged. The pointwise convolution formula is expressed as:
[0072]
[0073] Among them, is the weight matrix of the pointwise convolution, which is responsible for mapping the information in each channel to the specified number of output channels and realizing the interaction between the features of different channels. In this way, the model can integrate the features extracted on different channels, thereby enhancing the model's comprehensive understanding ability of the input image.
[0074] Finally, in order to prevent information loss during the convolution process, the inverted residual block also adopts a skip connection, directly adding the input feature map to the feature map after the convolution operation. Through the skip connection, the original information of the input features can be retained, and at the same time, the backpropagation of the gradient can be accelerated through the direct connection method, which helps the information transmission and optimization in the deep network. This step can be expressed by the formula:
[0075]
[0076] Among them, The activation function restricts the value range of the output features, preventing numerical overflow, and at the same time introduces non-linearity, enhancing the expressive ability of the model. The skip connection ensures that the original input information can be directly transmitted to the output, enabling the model to retain more semantic information during the feature extraction process and alleviating the problem of gradient disappearance.
[0077] Up to this point, after multiple layers of stacking and processing of the inverted residual blocks, the obtained multi-scale feature maps can capture different levels of information in the image, from low-level edge features to high-level semantic features, ensuring multi-dimensional processing of the defogging process. These feature maps will be passed to the decoder part in the subsequent steps for image reconstruction and detail restoration.
[0078] In some embodiments of the present invention, the multi-level upsampling decoding includes layer-by-layer transposed convolution upsampling operations and batch normalization operations, and the output features of the batch normalization operation at each level are fused with the output features of the corresponding level inverted residual block through skip connection feature fusion as the input features for the next transposed convolution upsampling operation.
[0079] Specifically, for the decoder part, the embodiment constructs a decoder structure that combines skip connection and transposed convolution upsampling to gradually restore the image resolution, and uses the multi-scale feature maps generated by the encoder for feature reconstruction.
[0080] First, the decoder restores each layer of low-resolution feature maps to a higher resolution through layer-by-layer transposed convolution operations, where the size of the transposed convolution kernel is set to 3×3, the stride is set to 2, and the padding is set to 1. The size of the feature map output by the transposed convolution can be expressed as:
[0081]
[0082] Among them, is the size of the input feature map, is the stride, is the size of the convolution kernel, is the padding value.
[0083] The th layer of the feature map after transposed convolution is expressed as:
[0084]
[0085] Among them, is the target size after transposed convolution, is the first transposed convolution layer of the decoder (the layers of the decoder are arranged from bottom to top).
[0086] After the deconvolution operation, to ensure the stability of training, the embodiment introduces a batch normalization operation. By normalizing the input distribution of the feature map after deconvolution, it reduces the problems of gradient vanishing and gradient explosion, and accelerates the training convergence. The batch normalization operation is as follows:
[0087]
[0088] After batch normalization, to maintain the detailed features of the image during the decoding process, the decoder also adopts skip connections. In the embodiment, the specific skip connection method is: the encoder feature map is connected to the decoder feature map and is connected to and is connected to and is connected to and is connected to . The fused feature map after skip connection is expressed by the formula:
[0089]
[0090] where, represents the feature map after skip connection, represents the operation of concatenating the encoder feature map and the decoder feature map in the channel dimension. Through this fusion operation, the decoder can simultaneously utilize the high-level semantic information in the encoder and the detailed features in the deconvolution process, ensuring the reconstruction quality of the image while accelerating the information transmission speed.
[0091] For each fused feature map after skip connection fusion, the embodiment gradually restores the number of channels through a convolutional layer, thereby retaining more high-resolution information. The entire decoding process ensures that both the high-level semantic features of the encoder are retained and the image details are effectively enhanced through repeated deconvolution and skip connection operations. Finally, a high-resolution image after defogging is generated. This process can be expressed by the formula:
[0092]
[0093] where, represents the deconvolution operation of the decoder, which is responsible for extracting key information from the fused features and restoring the image details. The deconvolution of the decoder and the skip connection work together to ensure that the features of the input image at different scales are fully utilized, achieving efficient defogging processing and image reconstruction.
[0094] In some embodiments of the present invention, Figure 3 is a schematic flowchart of calculating the smooth L1 loss and the perceptual loss in the embodiments of the present invention. As Figure 3 shown, the pre-trained smooth L1 loss is determined according to the defogging output image data and the real fog-free image data. The defogging output image data and the real fog-free image data are subjected to multi-level convolutional encoding to obtain the defogging image perceptual feature and the real fog-free image perceptual feature. The pre-trained perceptual loss is determined according to the defogging image perceptual feature and the real fog-free image perceptual feature, including:
[0095] S301. Calculate the pixel difference between the defogging output image data and the real fog-free image data based on the mean square error method and the L1 loss to obtain the pre-trained smooth L1 loss;
[0096] S302. Perform multi-level convolutional encoding on the defogging output image data and the real fog-free image data to obtain the defogging image perceptual feature and the real fog-free image perceptual feature;
[0097] S303. Calculate the feature difference between the defogging image perceptual feature and the real fog-free image perceptual feature based on the Euclidean distance method to obtain the pre-trained perceptual loss.
[0098] Specifically, considering that the traditional convolutional neural network is trained only based on the pixel difference at the image level for the defogging network, it is difficult to be consistent with the original image in terms of details such as structure and texture. In the pre-training task and the transfer learning task, the embodiment designs a composite loss function, which combines the smooth L1 loss and the perceptual loss based on MobileNetV2 to ensure that the defogged image can be close to the real fog-free image in both the pixel space and the feature space.
[0099] First, the smooth L1 loss optimizes the defogging effect by measuring the difference between the defogging output image and the real fog-free image at the pixel level. Different from the standard L1 loss, the smooth L1 loss is more robust. It can adopt the mean square error form when the error is small to ensure accuracy, and turn into the L1 loss when the error is large to reduce the influence of outliers on the training. The formula is defined as:
[0100]
[0101] Therefore, the calculation formula of the smooth L1 loss can be expressed as:
[0102]
[0103] Among them, represents the total number of pixels in the image, and respectively represent the a pixel value. The role of this loss function is to ensure that the output image is consistent with the real haze-free image in the pixel space, especially important for the restoration of fine details.
[0104] However, since relying solely on the loss function in the pixel space may lead to insufficient reconstruction of the image in terms of high-level semantic features. To address this, the embodiment designs a perceptual loss based on MobileNetV2. By using the encoder of the pre-trained network to extract multi-level features of the image, the dehazed output image and the real haze-free image are calculated for the difference in the feature space, so as to ensure that the reconstructed image is not only close to the real scene in the pixel space but also consistent in the feature space. Especially when dealing with video dehazing tasks, the perceptual loss can help the model better understand the temporal consistency of video frames and avoid the "flickering" phenomenon of the video caused by discontinuities during the haze removal process. First, the embodiment inputs the dehazed output image and the real haze-free image into the pre-trained encoder based on MobileNetV2 respectively to obtain the perceptual features of the dehazed image and the perceptual features of the real haze-free image , and calculates the perceptual loss based on this, where represents the feature extracted by MobileNetV2 at the th layer. The calculation formula of the perceptual loss is expressed as:
[0105]
[0106] Among them, represents the Euclidean distance, which is used to measure the difference between the dehazed image and the real haze-free image in the feature space. The feature map of each layer corresponds to image information at different scales, from low-level edge features to high-level semantic features. Through this multi-scale feature comparison, the perceptual loss can effectively improve the global structure and semantic consistency of the dehazed image.
[0107] Finally, the composite loss function combines the smooth L1 loss and the perceptual loss, and balances the contributions of the pixel loss and the feature loss through weights. The formula is expressed as:
[0108]
[0109] Among them, and are hyperparameters used to control the relative weights of the smooth L1 loss and the perceptual loss. By adjusting these two hyperparameters, the trade-off between the pixel-level reconstruction and the high-level semantic consistency of the model can be flexibly controlled, making the model show stronger adaptability in different scenarios.
[0110] The design of the composite loss function not only improves the quality of the dehazed image, but also reduces the problems of detail loss or texture distortion that may occur under complex weather conditions through reasonable feature space alignment. Through this multi-level loss optimization, the model can generate images with high visual quality and coherent structure in the dehazing task.
[0111] In some embodiments of the present invention, bidirectional optical flow loss is obtained by performing bidirectional optical flow estimation on the dehazing output data based on an adaptive pyramid and cost volume improvement, including:
[0112] Arrange the dehazing output data in the order of video frames, and take every three consecutive frames as a set of optical flow input data, where each set of optical flow input data includes a forward frame image, a current frame image, and a backward frame image;
[0113] Input the optical flow input data into a preset adaptive feature pyramid for multi-scale feature extraction, deconvolution upsampling, and feature fusion to obtain optical flow features;
[0114] Perform cost volume method optical flow estimation and frame warping operations on the optical flow features to obtain a forward prediction frame image and a backward prediction frame image;
[0115] Determine the forward optical flow loss according to the forward prediction frame image and the current frame image, determine the backward optical flow loss according to the backward prediction frame image and the current frame image, and merge the forward optical flow loss and the backward optical flow loss to obtain the bidirectional optical flow loss.
[0116] Specifically, in order to improve the temporal consistency of the video dehazing network in processing video dehazing tasks, the embodiment also introduces bidirectional optical flow loss when training the video dehazing network. To calculate the bidirectional optical flow loss, the embodiment designs a lightweight preset bidirectional optical flow network model based on an adaptive feature pyramid. The preset bidirectional optical flow network model takes every three adjacent video image sequences in the dehazing input video as a set of optical flow input data, where represents the forward frame image of the previous frame, represents the current frame image, represents the backward frame image of the next frame. Through the adaptive feature pyramid, the network generates multi-scale feature maps at different resolutions , where represents the number of layers of the feature pyramid. The design of the feature pyramid enables the network to simultaneously capture global motion information at low resolution and detailed motion information at high resolution, ensuring that the optical flow estimation can handle complex motion scenarios. This multi-scale feature extraction helps to improve the accuracy and robustness of the optical flow estimation. For each frame image , extract the multi-scale feature map , which can be expressed by the formula:
[0117]
[0118] Among them, represents the convolutional network in the feature pyramid. The feature pyramid structure is constructed bottom-up by decreasing the resolution, with a total of 4 layers. Each layer uses the original resolution as the benchmark, and the scaling ratio is , where represents the level. The low-level pyramid features (high-resolution features) mainly capture local motion details, while the high-level pyramid features (low-resolution features) are used to capture global motion information.
[0119] In addition, the multi-layer features included in the feature pyramid not only contain high-resolution detailed features but also extract low-resolution global features through successive downsampling. In order to combine global and local feature information in subsequent optical flow estimation, a deconvolution upsampling module is also designed to transmit the global information of the low-resolution level upward to the high-resolution level and perform feature fusion. The formula for feature fusion is expressed as:
[0120]
[0121] Among them, represents the deconvolution upsampling operation, is the optical flow feature after fusion. This operation can not only retain the local motion information of high-resolution features but also transmit the global motion trend contained in low-resolution features to the high-resolution level, effectively enhancing the robustness of optical flow estimation. Through this successive convolutional operation, the feature map not only covers the local detailed information of motion but also can reflect large-scale motion changes. The multi-scale processing of the pyramid can capture useful features in the case of motion blur or thick fog, improving the generalization ability of the model.
[0122] After obtaining the optical flow features, a cost volume structure is introduced in the embodiment when calculating the optical flow. The cost volume can capture the displacement relationship of pixels with relatively low computational complexity, avoiding the high computational cost brought by directly performing brute-force matching on all pixels. At the same time, the use of multi-scale feature maps also makes the optical flow estimation more accurate and can adapt to motion patterns in different scenarios.
[0123] During the process of calculating the optical flow, the specific approach is to calculate the forward optical flow and the backward optical flow , which respectively represent the motion estimation from the previous frame to the current frame and from the next frame to the current frame . The formula is expressed as:
[0124]
[0125]
[0126] Among them, represents the pixel position, and represent the pixel displacement, represents the th layer. The cost volume measures the similarity between two feature maps at different pixel displacements and provides a basis for optical flow estimation. By comparing the feature maps of adjacent frames, the network gradually regresses the forward optical flow and the backward optical flow . These two optical flow fields describe the pixel motion information between video frames.
[0127] Then, optical flow estimation is performed based on the constructed cost volume. This process is to gradually regress the forward optical flow field and the backward optical flow field through a deep convolutional network layer by layer, and the formula is expressed as:
[0128]
[0129]
[0130] At this time, the obtained optical flow field represents the motion information of the frames from frame to frame to frame in the th layer, represents the motion information of the frames from frame to frame in the
[0131] Then, in order to accurately estimate the forward optical flow and the backward optical flow, based on the forward optical flow and the backward optical flow , a frame warping operation is performed to warp the pixel positions of the previous frame and the next frame into the pixel positions of the current frame , and the forward prediction frame image and the backward prediction frame image are respectively generated. This operation adopts the bilinear interpolation method to enable accurate transformation of pixel positions, and the formula is expressed as
[0132]
[0133]
[0134] Among them, represents the value of the predicted frame image obtained by warping through the forward optical flow at the pixel, Denote the value of the predicted frame image obtained by backward optical flow deformation at the pixel. Through the action of forward and backward optical flows, the frame deformation operation can align the pixel information of adjacent frames to the current frame, which enables the predicted frame to have a high degree of consistency with the real haze-free image of the current frame. In this way, the embodiments can not only effectively fuse the information of adjacent frames, but also help the network obtain the motion trajectory between frames during the haze removal process, thereby improving the inter-frame consistency. In the video haze removal task, pixel-level alignment deformation can significantly reduce the inter-frame discontinuity and improve the smoothness of the video.
[0135] Finally, the embodiments define a bidirectional optical flow loss function to measure the pixel difference between the predicted frame and the real current frame. The forward optical flow loss is calculated by computing the L1 loss between the forward predicted frame image and the current frame image to measure the error between the two, and the formula is expressed as:
[0136]
[0137] where, is the pixel coordinate, is the total number of pixels.
[0138] The backward frame sampling similarity method calculates the L1 loss between the backward predicted frame image and the current frame image :
[0139]
[0140] The bidirectional optical flow loss is obtained by combining the losses of the forward optical flow and the backward optical flow, and the formula is expressed as:
[0141]
[0142] where, and represent the weight hyperparameters of the forward and backward optical flow losses.
[0143] In some embodiments of the present invention, Figure 4 is a schematic diagram of the iterative training process of the pre-trained video haze removal network according to the embodiments of the present invention. As shown in Figure 4 , according to the smooth L1 loss, the perceptual loss, and the bidirectional optical flow loss, the parameters of the pre-trained video haze removal network are iteratively fine-tuned and end-to-end trained to obtain a trained video haze removal network, including:
[0144] S401. Determine the total haze removal loss according to the smooth L1 loss, the perceptual loss, and the bidirectional optical flow loss;
[0145] S402. Keep the network parameters of other parts of the pre-trained dehazing network unchanged, and adjust the network parameters of each preset part to be adjusted in the pre-trained dehazing network according to the total dehazing loss in turn;
[0146] S403. Perform end-to-end training on the pre-trained video dehazing network according to the total dehazing loss, using the cosine annealing strategy, learning rate warm-up strategy, and gradient clipping strategy, and obtain a well-trained video dehazing network through the iterative training process.
[0147] Specifically, to enable the model to adapt to the video dehazing task, based on the pre-trained model, the embodiment performs parameter fine-tuning and small-sample transfer learning, and combines the bidirectional optical flow loss for optimization. The input is a collected hazy video dataset, represented as a frame sequence , where is the total number of frames of the video.
[0148] In transfer learning, considering that the pre-trained model has obtained excellent image dehazing ability on a large hazy image dataset. To maintain this ability, the embodiment performs transfer learning on the model by fine-tuning the parameters of the pre-trained model.
[0149] For the encoder part, the embodiment fine-tunes the last two high-level feature layers of the encoder to adapt to the dynamic change characteristics between video frames. The high-level feature layers are used to capture the temporal information in the video, helping the model achieve inter-frame consistency and dehazing effect in the video dehazing task. During the transfer learning process, the parameters of the high-level feature layers of the encoder are gradually updated through backpropagation to adapt to the temporal dynamic changes of the video. The update formula of the high-level feature layers is as follows:
[0150]
[0151] where, and represent the new parameters and initial parameters of the high-level feature layers respectively, is the learning rate, which is used to control the step size of parameter update, is the total loss function, and by optimizing the error of video dehazing is reduced.
[0152] For the decoder part, the embodiment fine-tunes the parameters to adapt to the dynamic changes and temporal consistency between video frames. For the decoder, its role is to gradually restore the spatial resolution of the input image, and fuse the feature maps generated by the encoder with the corresponding layers in the decoder through upsampling and skip connections. To ensure the temporal consistency in the video dehazing task, selectively unfreeze the key layers near the output end of the decoder, so that it can be adjusted and updated according to the video frame characteristics.
[0153] In the embodiment, first, the last convolutional layer is unfrozen. The last convolutional layer is responsible for mapping the upsampled feature map to the final dehazed image. This layer directly affects the quality of the video frame output. Therefore, during the transfer learning process, the parameters of this layer need to be unfrozen so that it can adapt to the temporal variation characteristics in video dehazing. By fine-tuning this layer, the model can learn how to handle the dynamic changes in the video, thereby generating dehazed results with temporal consistency. Its convolutional formula is expressed as:
[0154]
[0155] where, is the activation function, and are the weights and biases of this convolutional layer respectively. Through fine-tuning, and are updated according to the video frame dynamics.
[0156] Then, unfreeze the second-to-last upsampling module. The upsampling module is responsible for restoring the low-resolution feature map to a higher resolution and performing skip connections by combining the high-level features in the encoder. For the video dehazing task, the purpose of unfreezing this layer is to enable the model to better handle the changes between video frames. The upsampling operation is implemented by transposed convolution, and the formula is expressed as:
[0157]
[0158] where, is the upsampled feature map generated after the transposed convolution operation, is the input feature map, is the adjustable parameter matrix. By fine-tuning , this layer can be adjusted for different dynamic scenarios in the video dehazing task, enabling the model to gradually restore the resolution of the video layer by layer and ensuring smooth transitions between frames.
[0159] In addition to the fine-tuning of the above layers, the intermediate layers of the decoder in the embodiment remain frozen. These layers are mainly responsible for gradually promoting the low-level features in the encoder to high-level semantic information. Since these features are common in image and video dehazing tasks, keeping their parameters unchanged can reduce the risk of overfitting while maintaining the original dehazing ability of the model. The frozen intermediate layer parameters are expressed as:
[0160]
[0161] where, and represent the intermediate parameters of the decoder after and before parameter adjustment respectively.
[0162] Meanwhile, considering that during video dehazing, the differences between frames may cause single-frame features to be insufficient to reflect the continuous changes in the video. By fine-tuning the skip connection module in the decoder, the model can more effectively fuse features at different levels in the encoder, ensuring enhanced feature transfer between video frames while maintaining semantic consistency. The fusion formula for the skip connection is:
[0163]
[0164] In the embodiment, by unfreezing relevant parts of the decoder, the skip connection module can adjust the feature fusion method according to temporal continuity, thereby generating a more coherent video frame sequence.
[0165] In addition to the fine-tuning of the above-mentioned layers, the intermediate layers of the decoder in the embodiment remain frozen. These layers are mainly responsible for gradually promoting the low-level features in the encoder to high-level semantic information. Since these features are common in image and video dehazing tasks, keeping their parameters unchanged can reduce the risk of overfitting while maintaining the original dehazing ability of the model. The frozen intermediate layer parameters are represented as:
[0166]
[0167] Among them, and respectively represent the intermediate parameters of the decoder after and before parameter adjustment.
[0168] In addition, during this process, the embodiment continues to sample the smooth L1 loss function and the perceptual loss as the main task loss functions, and on this basis, introduces the bidirectional optical flow loss to construct an optimization strategy with multiple loss functions. The total loss function formula is expressed as:
[0169]
[0170] Among them , is the bidirectional optical flow loss function, and the main task loss The formula is expressed as:
[0171]
[0172] Among them, is the L1 smooth loss, is the perceptual loss, and are hyperparameters for controlling the loss weights.
[0173] Then, after completing the parameter fine-tuning, the embodiment introduces a variety of optimization strategies in combination with the total loss function to perform end-to-end training on the pre-trained video dehazing network, including the cosine annealing strategy, the learning rate warm-up strategy, the gradient clipping strategy, and the learning rate decay strategy.
[0174] During the training process, the embodiments introduce the Adam optimizer to effectively improve the speed of the training process and enhance the generalization ability of the model. Adam maintains the first-order moment estimate and second-order moment estimate of each parameter and uses this information to update the weights of the model. For each input sequence , the model processes the input sequence and updates the weights according to the total loss function . Specifically, the update rule formula of the Adam optimizer is expressed as:
[0175]
[0176]
[0177]
[0178]
[0179] where and represent the first-order momentum and second-order momentum of the gradient respectively, and are hyperparameters that control the momentum decay rate, is the learning rate, is a small constant to prevent division by zero operations.
[0180] To avoid the learning rate dropping too fast during the training process and causing the model to fall into a local optimum prematurely, the embodiments adopt a cosine annealing learning rate adjustment strategy. This strategy dynamically adjusts the learning rate so that the model can update parameters at a higher learning rate in the early stage of training, accelerating the convergence speed.
[0181] As the training progresses, the learning rate gradually decays to ensure that the model can converge smoothly in the later stage and improve the final performance. The formula is expressed as follows:
[0182]
[0183] where represents the learning rate at the -th step, and are the set minimum learning rate and maximum learning rate respectively, ensuring that the upper and lower limits of the learning rate are controlled, represents the total number of training steps, is the current number of training steps.
[0184] Considering the different roles of different layers of the model, among which, the low-level network structure of the encoder is mainly used to learn low-level features of images (such as edges and textures), while the high-level network structure focuses more on capturing more abstract deep features. Therefore, in the embodiments, a strategy of hierarchical learning rate is adopted here, and different learning rates are assigned to different layers according to the frozen state of the layers and the task characteristics to achieve better fine-tuning effects.
[0185] More specifically, for the frozen low-level part of the encoder, these layers are responsible for extracting general low-level features such as edges, textures, and illumination changes. When migrating to the video dehazing task, these low-level features usually have strong adaptability and can maintain consistent recognition effects under different concentrations of haze conditions, and have achieved strong feature extraction capabilities based on a large-scale haze dataset in the pre-trained model. Therefore, in order to avoid over-updating these general features in the new task, these layers are frozen and the learning rate is set to 0 to keep their weights unchanged, ensuring that the model can effectively retain the low-level dehazing features in the pre-trained model during the transfer learning stage, so as to focus on high-level temporal consistency and detail reconstruction.
[0186] Similarly, for the frozen decoder part, these layers are mainly used to restore the basic structure and detail information of video frames in the dehazing task and have learned generally applicable dehazing features in the pre-training of single-frame image dehazing. These basic layers of the decoder mainly focus on low-level reconstruction operations such as removing large fog patches in the image and enhancing contours, etc. These features do not change much in different dehazing tasks and are also applicable to the video dehazing task. Therefore, these layers are frozen and the learning rate is set to 0 to retain their original dehazing capabilities, ensure stable performance in the new task, and reduce the computational overhead.
[0187] For the last two high-level parts of the unfrozen encoder, the features extracted by these layers are high-level features related to specific tasks. In transfer learning, these layers need to be fine-tuned according to the characteristics of the new data to adapt to the dynamic features in video data and the requirements of the dehazing task. Therefore, the learning rate range is set to , This can ensure smooth fine-tuning of high-level features in transfer learning to prevent excessive parameter updates from destroying the pre-training effect.
[0188] For the two layers of the unfrozen decoder, only the last two layers are unfrozen in the decoder to restore high-resolution features and ensure output consistency and dehazing effects. The learning rate range is set to , ,Only for the two unfrozen layers, ensure that the decoder part can quickly adapt to the reconstruction requirements of the new task.
[0189] To prevent the learning rate from being too low at the initial stage of training and affecting the effective training of the model, the embodiment adopts the Warm-up learning warm-up strategy. During the stage of using the Warm-up learning warm-up strategy at the initial stage of training, the learning rate will gradually increase from a small initial value to the set , helping the model quickly enter the effective training state. The formula is as follows:
[0190]
[0191] where represents the initial learning rate in the Warm-up stage, and its setting varies with different levels. For the unfrozen high-level part of the encoder, the initial learning rate is set to 0.00005, and the corresponding maximum learning rate . For the two unfrozen layers of the decoder, the initial learning rate is set to 0.0005, and the corresponding maximum learning rate . represents the number of steps in the Warm-up stage, that is, during this period, the learning rate gradually increases from to the corresponding maximum learning rate to ensure that each layer enters a stable training state.
[0192] At the same time, to prevent the gradient from being too large during the backpropagation process and causing unstable training, the embodiment also adopts a gradient clipping strategy. After calculating the gradient in each backpropagation, the gradient clipping operation is performed to ensure that the norm of the gradient does not exceed the preset threshold , and this value is initially set to . This operation can effectively avoid the gradient explosion problem caused by too large a gradient, thereby ensuring the smoothness of the training process and the stable update of network parameters. The formula for gradient clipping is as follows:
[0193]
[0194] Before the optimizer performs parameter update, the gradient clipping function is called after each gradient calculation to limit the magnitude of the gradient and ensure the stability of the network. By comparing the gradient of the current step with the preset gradient threshold , if the norm of the gradient exceeds this threshold, it will be scaled down by the preset ratio and kept within a reasonable range. According to the depth of the network and the characteristics of multi-task training, this preset gradient threshold ensures that the model can prevent unstable phenomena caused by too large a gradient during the training process.
[0195] The embodiments can significantly enhance the robustness of training by introducing a gradient clipping strategy. Especially in multi-task learning, it can prevent the problem of excessive gradients caused by a certain loss term, ensuring that each task can be reasonably optimized. At the same time, by restricting the maximum norm of the gradients, the model can avoid unstable phenomena during training. Especially in deep networks or complex tasks, it can better control the amplitude of gradient updates.
[0196] In addition, the embodiments dynamically adjust the learning rate according to the training loss of the model in each iteration. The learning rate decay strategy uses an exponential decay formula, which is expressed as:
[0197]
[0198] where, is the decay factor. By gradually reducing the learning rate during the training process, the embodiments can effectively prevent the network from repeatedly deviating from the optimal solution due to an overly large learning rate in the later stage of training, ensuring that the network converges stably when approaching the optimal solution.
[0199] Finally, the weight update is performed by weighting with historical gradient information, and the update formula is:
[0200]
[0201]
[0202] where, is the momentum term, is the momentum coefficient, is the gradient of the loss function with respect to the weights.
[0203] In end-to-end training, the embodiments adopt the above methods. The total loss function that combines the smooth L1 loss, perceptual loss, and bidirectional optical flow loss is optimized through multiple rounds by the cosine annealing strategy, learning rate warm-up strategy, and gradient clipping strategy, and finally a well-trained video dehazing network is obtained, enabling the pre-trained video dehazing network to effectively adapt to the video dehazing task after training and obtaining a well-trained video dehazing network.
[0204] In summary, in order to improve the accuracy and temporal consistency of video dehazing in details, the present invention first inputs the obtained hazy video data and corresponding real haze-free video data into a pre-trained video dehazing network, performs multi-level convolutional encoding and multi-level upsampling decoding on the hazy video data to obtain dehazing output data; then determines the smooth L1 loss according to the dehazing output data and the real haze-free video data, performs multi-level convolutional encoding on the dehazing output data and the real haze-free video data to obtain dehazing perception features and real perception features, determines the perception loss according to the dehazing perception features and the real perception features, and performs bidirectional optical flow estimation based on an adaptive pyramid and cost volume improvement on the dehazing output data to obtain a bidirectional optical flow loss; finally, iteratively performs parameter fine-tuning and end-to-end training on the pre-trained video dehazing network according to the smooth L1 loss, the perception loss, and the bidirectional optical flow loss to obtain a trained video dehazing network. During the training process of the present invention, the smooth L1 loss is used to calculate the loss at the pixel level, the perception loss is used to calculate the loss at the feature level to improve the accuracy of the network in details, and the bidirectional optical flow loss is obtained by improving the bidirectional optical flow estimation to capture the dynamic information between video frames and improve the temporal consistency of the network in the video dehazing task, so that the video dehazing network can effectively improve the dehazing effect with a relatively low computational amount, and obtain a video dehazing network that takes into account the dehazing effect, low power consumption, and real-time performance.
[0205] The present invention also provides a method for applying a video dehazing network based on a bidirectional optical flow loss. Considering Figure 5 as follows, Figure 5 is a schematic flowchart of an embodiment of the method for applying a video dehazing network based on a bidirectional optical flow loss provided by the present invention. As Figure 5 shown, the method for applying a video dehazing network based on a bidirectional optical flow loss includes:
[0206] S501. Obtain a video to be dehazed;
[0207] S502. Input the video to be dehazed into the trained video dehazing network to obtain a dehazed output video.
[0208] Specifically, during the application process of the video dehazing network, it is first necessary to effectively obtain the video to be dehazed, and then input the video to be dehazed into the trained video dehazing network. After the video dehazing network completes the dehazing task, all dehazed frames are merged in chronological order to obtain a frame sequence, and combined with the frame rate setting, it is converted into a video file in a standard format to obtain a dehazed output video, and the frame rate of the finally output video is the same as that of the input video.
[0209] As Figure 6 shown, the present invention also correspondingly provides an embedded device 600. The embedded device 600 includes a processor 601, a memory 602, and a display 603. Figure 6Only some components of the embedded device 600 are shown, but it should be understood that it is not necessary to implement all the shown components, and more or fewer components can be alternatively implemented.
[0210] In some embodiments, the processor 601 may be a central processing unit (CPU), a microprocessor, or other data processing chips, and is used to run the program code stored in the memory 602 or process data, such as the video defogging network training method based on bidirectional optical flow loss and / or the video defogging network application method in the present invention.
[0211] In some embodiments, the memory 602 may be an internal storage unit of the embedded device 600, such as a hard disk or memory of the embedded device 600. In some other embodiments, the memory 602 may also be an external storage device of the embedded device 600, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the embedded device 600.
[0212] Furthermore, the memory 602 may also include both the internal storage unit of the embedded device 600 and the external storage device. The memory 602 is used to store the application software installed on the embedded device 600 and various types of data.
[0213] In some embodiments, the display 603 may be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (organic light-emitting diode) toucher, etc. The display 603 is used to display the information of the embedded device 600 and to display a visual user interface. The components 601-603 of the embedded device 600 communicate with each other through a system bus.
[0214] In one embodiment, when the processor 601 executes the video defogging network training program in the memory 602, the following steps may be implemented:
[0215] Input the obtained foggy video data and the corresponding real fog-free video data into a pre-trained video defogging network, and perform multi-level convolutional encoding and multi-level upsampling decoding on the foggy video data to obtain defogged output data;
[0216] Determine the smooth L1 loss based on the defogging output data and the real fog-free video data, perform multi-level convolutional encoding on the defogging output data and the real fog-free video data to obtain the defogging perception feature and the real perception feature, determine the perception loss based on the defogging perception feature and the real perception feature, and perform bidirectional optical flow estimation on the defogging output data based on the improved adaptive pyramid and cost volume to obtain the bidirectional optical flow loss;
[0217] Iteratively perform parameter fine-tuning and end-to-end training on the pre-trained video defogging network according to the smooth L1 loss, the perception loss, and the bidirectional optical flow loss to obtain a trained complete video defogging network.
[0218] In one embodiment, when the processor 601 executes the video defogging network application program in the memory 602, the following steps can be implemented:
[0219] Obtain the video to be defogged;
[0220] Input the video to be defogged into the trained complete video defogging network to obtain the defogged output video.
[0221] It should be understood that when the processor 601 executes the video defogging network training program and / or the video defogging network application program in the memory 602, in addition to the above functions, other functions can also be implemented. For specific details, reference can be made to the description of the corresponding method embodiments above.
[0222] Furthermore, the type of the embedded device 600 mentioned in the embodiments of the present invention is not specifically limited. The embedded device 600 can be a portable device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, etc. Exemplary embodiments of the portable embedded device include, but are not limited to, portable embedded devices equipped with IOS, android, microsoft, or other operating systems. The above portable embedded devices can also be other portable embedded devices, such as a laptop computer with a touch-sensitive surface (such as a touch panel). It should also be understood that in some other embodiments of the present invention, the embedded device 600 may not be a portable embedded device, but a desktop computer with a touch-sensitive surface (such as a touch panel).
[0223] Correspondingly, the embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium is used to store computer-readable programs or instructions. When the programs or instructions are executed by a processor, the steps or functions in the video defogging network training method based on the bidirectional optical flow loss and / or the video defogging network application method based on the bidirectional optical flow loss provided by the above method embodiments can be implemented.
[0224] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as processors, controllers, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0225] The above has introduced in detail the video dehazing network training method, application method, embedded device and medium based on bidirectional optical flow loss provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A video dehazing network training method based on bidirectional optical flow loss, characterized in that: include: The obtained foggy video data and the corresponding real fog-free video data are input into the pre-trained video defogging network, and the foggy video data is multi-level convolutionally encoded and multi-level upsampling decoded to obtain defogging output data; Determine a smooth L1 loss according to the defogging output data and the real fog-free video data, perform multi-level convolution encoding on the defogging output data and the real fog-free video data to obtain defogging perception features and real perception features, determine the perception loss according to the defogging perception features and the real perception features, arrange the defogging output data in the order of video frames, and use every three consecutive frames as a group of optical flow input data, wherein each group of the optical flow input data includes a forward frame image, a current frame image and a backward frame image; input the optical flow input data into a preset adaptive feature pyramid to perform multi-scale feature extraction, deconvolution upsampling and feature fusion to obtain optical flow features; Performing cost volume method optical flow estimation and frame deformation operations on the optical flow features to obtain a forward prediction frame image and a backward prediction frame image; determining a forward optical flow loss according to the forward prediction frame image and the current frame image, determining a backward optical flow loss according to the backward prediction frame image and the current frame image, and combining the forward optical flow loss and the backward optical flow loss to obtain a bidirectional optical flow loss; The pre-trained video dehazing network is iteratively parameter-fine-tuned and end-to-end trained according to the smooth L1 loss, the perceptual loss, and the bidirectional optical flow loss to obtain a fully trained video dehazing network.
2. The video dehazing network training method based on bidirectional optical flow loss according to claim 1 is characterized in that: The pre-trained video defogging network is obtained by training an initial defogging network, and the training initial defogging network includes: The acquired foggy image data and the corresponding real fog-free image data are input into the initial defogging network, and the foggy image data is multi-level convolutionally encoded and multi-level upsampling decoded to obtain defogging output image data; A pre-trained smooth L1 loss is determined according to the defogging output image data and the real haze-free image data, multi-level convolution encoding is performed on the defogging output image data and the real haze-free image data to obtain defogging image perception features and real haze-free image perception features, a pre-trained perceptual loss is determined according to the defogging image perception features and the real haze-free image perception features, and the initial defogging network is iteratively trained according to the pre-trained smooth L1 loss and the pre-trained perceptual loss to obtain the pre-trained video defogging network.
3. The video dehazing network training method based on bidirectional optical flow loss according to claim 1 is characterized in that: The multi-stage convolutional coding includes an initial convolutional layer and a plurality of inverted residual blocks, each of which includes an extended convolution and a depth-separable convolution.
4. The video defogging network training method based on bidirectional optical flow loss according to claim 3 is characterized in that: The multi-level upsampling decoding includes layer-by-layer deconvolution upsampling operations and batch normalization operations, and the output features of the batch normalization operation at each level are skip-connected and fused with the output features of the inverted residual block at the corresponding level to serve as the input features of the deconvolution upsampling operation of the next level.
5. The video dehazing network training method based on bidirectional optical flow loss according to claim 2 is characterized in that: The method of determining a pre-trained smoothed L1 loss according to the defogging output image data and the real fog-free image data, performing multi-level convolution encoding on the defogging output image data and the real fog-free image data to obtain defogging image perception features and real fog-free image perception features, and determining a pre-trained perception loss according to the defogging image perception features and the real fog-free image perception features, comprises: Calculating the pixel difference between the defogging output image data and the real non-fogging image data based on the mean square error method and the L1 loss to obtain a pre-trained smoothing L1 loss; Performing multi-level convolution coding on the defogging output image data and the real fog-free image data to obtain defogging image perception features and real fog-free image perception features; The pre-trained perceptual loss is obtained by calculating the feature difference between the dehazed image perceptual features and the real haze-free image perceptual features based on the Euclidean distance method.
6. The video dehazing network training method based on bidirectional optical flow loss according to claim 1 is characterized in that: The iteratively performing parameter fine-tuning and end-to-end training on the pre-trained video defogging network according to the smooth L1 loss, the perceptual loss, and the bidirectional optical flow loss to obtain a fully trained video defogging network includes: Determine a total dehazing loss according to the smoothed L1 loss, the perceptual loss, and the bidirectional optical flow loss; Keeping the network parameters of other parts of the pre-trained video defogging network unchanged, adjusting the network parameters of each preset part to be adjusted of the pre-trained video defogging network in turn according to the total defogging loss; The pre-trained video defogging network is end-to-end trained using a cosine annealing strategy, a learning rate warm-up strategy, and a gradient clipping strategy according to the total defogging loss, and a fully trained video defogging network is obtained through an iterative training process.
7. A video defogging network application method based on bidirectional optical flow loss, characterized in that: include: Get the video to be dehazed; Inputting the video to be defogged into a well-trained video defogging network to obtain a defogged output video; The fully trained video dehazing network is determined according to the video dehazing network training method based on bidirectional optical flow loss according to any one of claims 1 to 6.
8. An embedded device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the video defogging network training method based on bidirectional optical flow loss according to any one of claims 1 to 6 and / or the video defogging network application method based on bidirectional optical flow loss according to claim 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the video defogging network training method based on bidirectional optical flow loss according to any one of claims 1 to 6 and / or the video defogging network application method based on bidirectional optical flow loss according to claim 7 are implemented.
Citation Information
Patent Citations
Super-resolution reconstruction method and device based on deep learning
CN116645270A
Image defogging method based on multi-receptive-field feature fusion and mixed attention
CN116912130A