Video defogging network training method based on bidirectional optical flow loss
Through the video defogging network training method based on bidirectional optical flow loss, the parameters of the pre-trained video defogging network are fine-tuned, which solves the problem of high computing complexity in the existing technology, and achieves high-efficiency, low-power consumption and real-time video defogging effect.
Patent Information
- Application Number
- CN202510466660.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing video defog removal technology has high computational complexity when processing high frame rate and high resolution videos, making it difficult to achieve efficient defog removal under limited computing resources, and cannot take into account the performance requirements of defog removal, real-time and low power consumption.
The video defogging network training method based on bidirectional optical flow loss is adopted. Through multi-stage convolutional encoding and multi-stage upsampling and decoding, combined with smooth L1 loss, perceptual loss and bidirectional optical flow loss, the pre-trained video defogging network is fine-tuned and end-to-end training to obtain a fully trained video defogging network.
It effectively improves the accuracy and time consistency of the video defog network in detail, so that the video defog network can achieve efficient defog removal with low computing volume, taking into account the defog removal effect, low power consumption and real-time.
Smart Images

Figure CN119992259A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent vision technology, and in particular to a video defogging network training method based on bidirectional optical flow loss. Background Art
[0002] In modern intelligent vision systems, the quality of video images is affected by complex weather conditions, which seriously reduces the ability of the vision system to recognize and detect scenes. In this regard, video defogging technology, as a very important technology, is widely used in fields such as autonomous driving, drone navigation and intelligent monitoring. At present, in the existing research on video defogging, the proportion of deep learning methods in video defogging continues to increase. Among them, convolutional neural network CNN is one of the typical methods. It extracts features from the input image to achieve frame-by-frame defogging, and finally splices the processed images into a complete video. This method performs well in single-frame images, but in video defogging scenarios, frame-by-frame processing ignores the dynamic information between frames, resulting in insufficient temporal consistency, uneven transitions between video frames, and prone to jumps or flickers. To address this problem, some studies have introduced generative adversarial networks to generate clearer defogging images through adversarial learning between generators and discriminators. Or attention mechanism, through the combination of channel attention and spatial attention, the model can better focus on the key information in the image, thereby improving the defogging effect. Some studies have also introduced time processing modules such as long short-term memory networks LSTM. This type of module can capture the timing information between video frames and ensure smooth transition between frames. In addition, the timing processing method of 3D convolutional networks is also widely used in video dehazing tasks. 3D convolutional networks improve temporal consistency by adding convolution operations in the time dimension.
[0003] However, the existing methods for improving dehazing effect and temporal consistency, whether it is generative adversarial networks, attention mechanisms, long short-term memory networks or 3D convolutional networks, all have high computational complexity, especially when processing high frame rate and high-resolution videos. The consumption of computing resources increases significantly. In the application scenarios of portable devices such as embedded devices, it is impossible to ensure the dehazing effect and real-time performance while meeting its low power consumption performance requirements.
[0004] Therefore, there are existing technical problems such as high computational complexity, difficulty in achieving efficient video dehazing with limited computing resources when processing high frame rate and high resolution videos, and inability to balance dehazing effect and real-time performance while meeting the low power consumption performance requirements of portable devices such as embedded devices, which need to be improved. Summary of the invention
[0005] In view of this, it is necessary to provide a video dehazing network training method based on bidirectional optical flow loss to solve the technical problems existing in the prior art, such as high computational complexity, difficulty in achieving efficient video dehazing under limited computing resources when processing high frame rate and high resolution videos, and inability to take into account both dehazing effect and real-time performance while meeting the performance requirements of portable devices such as embedded devices for low power consumption.
[0006] In order to solve the above technical problems, on the one hand, the present invention provides a video defogging network training method based on bidirectional optical flow loss, comprising: The acquired foggy video data and the corresponding real fog-free video data are input into the pre-trained video defogging network, and the foggy video data is subjected to multi-level convolution encoding and multi-level upsampling decoding to obtain the defogging output data; Determine the smooth L1 loss according to the defogging output data and the real haze-free video data, perform multi-level convolution encoding on the defogging output data and the real haze-free video data to obtain the defogging perception features and the real perception features, determine the perception loss according to the defogging perception features and the real perception features, and perform the bidirectional optical flow estimation based on the adaptive pyramid and cost volume improvement on the defogging output data to obtain the bidirectional optical flow loss; The pre-trained video dehazing network is iteratively parameter-fine-tuned and end-to-end trained according to the smooth L1 loss, perceptual loss, and bidirectional optical flow loss to obtain a fully trained video dehazing network.
[0007] In a possible implementation, the pre-trained video defogging network is obtained by training an initial defogging network, and the training of the initial defogging network includes: The acquired foggy image data and the corresponding real fog-free image data are input into the initial defogging network, and the foggy image data is subjected to multi-level convolution encoding and multi-level upsampling decoding to obtain the defogging output image data; A pre-trained smooth L1 loss is determined according to the dehazed output image data and the real haze-free image data, and multi-level convolutional encoding is performed on the dehazed output image data and the real haze-free image data to obtain the dehazed image perception features and the real haze-free image perception features. A pre-trained perceptual loss is determined according to the dehazed image perception features and the real haze-free image perception features, and an initial dehazing network is iteratively trained according to the pre-trained smooth L1 loss and the pre-trained perceptual loss to obtain a pre-trained video dehazing network.
[0008] In one possible implementation, the multi-stage convolutional coding includes an initial convolutional layer and a plurality of inverted residual blocks, each of which includes a dilated convolution and a depth-wise separable convolution.
[0009] In a possible implementation, multi-level upsampling decoding includes layer-by-layer deconvolution upsampling operations and batch normalization operations, and the output features of the batch normalization operation at each level are jump-connected and fused with the output features of the inverted residual block at the corresponding level to serve as the input features of the deconvolution upsampling operation of the next level.
[0010] In a possible implementation, a pre-trained smooth L1 loss is determined according to the defogging output image data and the real haze-free image data, multi-level convolution encoding is performed on the defogging output image data and the real haze-free image data to obtain defogging image perception features and real haze-free image perception features, and a pre-trained perception loss is determined according to the defogging image perception features and the real haze-free image perception features, including: Based on the mean square error method and L1 loss, the pixel difference between the dehazed output image data and the real haze-free image data is calculated to obtain the pre-trained smooth L1 loss; Perform multi-level convolution coding on the defogging output image data and the real fog-free image data to obtain the defogging image perception features and the real fog-free image perception features; Based on the Euclidean distance method, the feature difference between the dehazed image perception features and the real haze-free image perception features is calculated to obtain the pre-trained perception loss.
[0011] In one possible implementation, the dehazing output data is subjected to bidirectional optical flow estimation based on adaptive pyramid and cost volume improvement to obtain a bidirectional optical flow loss, including: Arrange the defogging output data in the order of video frames, and use every three consecutive frames as a set of optical flow input data, where each set of optical flow input data includes a forward frame image, a current frame image, and a backward frame image; The optical flow input data is input into the preset adaptive feature pyramid for multi-scale feature extraction, deconvolution upsampling and feature fusion to obtain the optical flow features; Performing cost volume method optical flow estimation and frame deformation operations on the optical flow features to obtain a forward prediction frame image and a backward prediction frame image; The forward optical flow loss is determined according to the forward predicted frame image and the current frame image, the backward optical flow loss is determined according to the backward predicted frame image and the current frame image, and the forward optical flow loss and the backward optical flow loss are combined to obtain the bidirectional optical flow loss.
[0012] In one possible implementation, the pre-trained video dehazing network is iteratively parameter-tuned and end-to-end trained according to the smooth L1 loss, perceptual loss, and bidirectional optical flow loss to obtain a fully trained video dehazing network, including: Determine the total loss of dehazing based on smooth L1 loss, perceptual loss and bidirectional optical flow loss; Keep the network parameters of other parts of the pre-trained defogging network unchanged, and adjust the network parameters of each preset part to be adjusted of the pre-trained defogging network in turn according to the total defogging loss; According to the total dehazing loss, the pre-trained video dehazing network is end-to-end trained with cosine annealing strategy, learning rate warm-up strategy and gradient clipping strategy, and the fully trained video dehazing network is obtained through the iterative training process.
[0013] On the other hand, the present invention also provides a video defogging network application method based on bidirectional optical flow loss, comprising: Get the video to be dehazed; Input the video to be dehazed into the well-trained video dehazing network to obtain the dehazed output video; The fully trained video dehazing network is determined according to the above-mentioned video dehazing network training method based on bidirectional optical flow loss.
[0014] On the other hand, the present invention also provides an embedded device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned video defogging network training method based on bidirectional optical flow loss and / or the above-mentioned video defogging network application method based on bidirectional optical flow loss are implemented.
[0015] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned video defogging network training method based on bidirectional optical flow loss and / or the above-mentioned video defogging network application method based on bidirectional optical flow loss are implemented.
[0016] The beneficial effects of the present invention are as follows: the video defogging network training method based on bidirectional optical flow loss provided by the embodiment of the present invention first inputs the acquired foggy video data and the corresponding real fog-free video data into the pre-trained video defogging network, performs multi-level convolution encoding and multi-level upsampling decoding on the foggy video data to obtain defogging output data; then determines the smooth L1 loss according to the defogging output data and the real fog-free video data, performs multi-level convolution encoding on the defogging output data and the real fog-free video data to obtain defogging perception features and real perception features, determines the perception loss according to the defogging perception features and the real perception features, performs bidirectional optical flow estimation based on adaptive pyramid and cost volume improvement on the defogging output data to obtain bidirectional optical flow loss; finally, fine-tunes the parameters of the pre-trained video defogging network according to the smooth L1 loss, the perception loss and the bidirectional optical flow loss, and obtains a fully trained video defogging network. During the training process, the present invention calculates the loss from the pixel level through smoothing L1 loss, calculates the loss at the feature level through perceptual loss, improves the accuracy of the network in details, obtains the bidirectional optical flow loss through improving the bidirectional optical flow estimation, captures the dynamic information between video frames, improves the time consistency of the network in the video defogging task, enables the video defogging network to effectively improve the defogging effect with lower computational complexity, and obtains a video defogging network that takes into account the defogging effect, low power consumption and real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A schematic diagram of a flow chart of an embodiment of a video defogging network training method based on bidirectional optical flow loss provided by the present invention; Figure 2 A schematic diagram of the process of training an initial defogging network according to an embodiment of the present invention; Figure 3 A schematic diagram of a process for calculating smoothed L1 loss and perceptual loss according to an embodiment of the present invention; Figure 4 A schematic diagram of the process of iterative training of a pre-trained video dehazing network according to an embodiment of the present invention; Figure 5 A schematic flow chart of an embodiment of a video defogging network application method based on bidirectional optical flow loss provided by the present invention; Figure 6 A schematic structural diagram of an embodiment of an embedded device provided by the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0020] In the description of the embodiments of the present invention, unless otherwise specified, "multiple" means two or more than two. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone.
[0021] The descriptions of "first" and "second" in the embodiments of the present invention are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the technical features defined as "first" and "second" may explicitly or implicitly include at least one of the features.
[0022] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0023] The present invention provides a video defogging network training method based on bidirectional optical flow loss and a video defogging network training method based on bidirectional optical flow loss, which are described below respectively.
[0024] Figure 1 A flow chart of an embodiment of a video defogging network training method based on bidirectional optical flow loss provided by the present invention is shown in FIG. Figure 1 As shown, the video dehazing network training method based on bidirectional optical flow loss includes: S101, inputting the acquired foggy video data and the corresponding real fog-free video data into a pre-trained video defogging network, performing multi-level convolution encoding and multi-level upsampling decoding on the foggy video data to obtain defogging output data; S102, determining a smooth L1 loss according to the defogging output data and the real fog-free video data, performing multi-level convolution encoding on the defogging output data and the real fog-free video data to obtain defogging perception features and real perception features, determining a perception loss according to the defogging perception features and the real perception features, and performing bidirectional optical flow estimation based on adaptive pyramid and cost volume improvement on the defogging output data to obtain a bidirectional optical flow loss; S103, iteratively fine-tune parameters and perform end-to-end training on the pre-trained video dehazing network according to the smooth L1 loss, the perceptual loss, and the bidirectional optical flow loss to obtain a fully trained video dehazing network.
[0025] Compared with the prior art, the video defogging network training method based on bidirectional optical flow loss provided by the embodiment of the present invention first inputs the acquired foggy video data and the corresponding real fog-free video data into the pre-trained video defogging network, performs multi-level convolution encoding and multi-level upsampling decoding on the foggy video data to obtain defogging output data; then determines the smooth L1 loss according to the defogging output data and the real fog-free video data, performs multi-level convolution encoding on the defogging output data and the real fog-free video data to obtain defogging perception features and real perception features, determines the perception loss according to the defogging perception features and the real perception features, performs bidirectional optical flow estimation based on adaptive pyramid and cost volume improvement on the defogging output data to obtain bidirectional optical flow loss; finally, fine-tunes the parameters of the pre-trained video defogging network according to the smooth L1 loss, the perception loss and the bidirectional optical flow loss, and obtains a fully trained video defogging network. During the training process, the present invention calculates the loss from the pixel level through smoothing L1 loss, calculates the loss at the feature level through perceptual loss, improves the accuracy of the network in details, obtains the bidirectional optical flow loss through improving the bidirectional optical flow estimation, captures the dynamic information between video frames, improves the time consistency of the network in the video defogging task, enables the video defogging network to effectively improve the defogging effect with lower computational complexity, and obtains a video defogging network that takes into account the defogging effect, low power consumption and real-time performance.
[0026] In some embodiments of the present invention, Figure 2 FIG. 1 is a flow chart of training an initial defogging network according to an embodiment of the present invention. Figure 2 As shown, the pre-trained video defogging network is obtained by training the initial defogging network, and the training of the initial defogging network includes: S201, inputting the acquired foggy image data and the corresponding real fog-free image data into the initial defogging network, performing multi-level convolution encoding and multi-level upsampling decoding on the foggy image data to obtain defogging output image data; S202. Determine a pre-trained smooth L1 loss based on the defogging output image data and the real haze-free image data, perform multi-level convolution encoding on the defogging output image data and the real haze-free image data to obtain defogging image perception features and real haze-free image perception features, determine a pre-trained perceptual loss based on the defogging image perception features and the real haze-free image perception features, and iteratively train an initial defogging network based on the pre-trained smooth L1 loss and the pre-trained perceptual loss to obtain a pre-trained video defogging network.
[0027] Specifically, considering that there are relatively few high-quality public video datasets in the video defogging task, it is difficult to provide enough training samples for the deep learning model. To this end, the embodiment uses a pre-training method to first pre-train the complete encoder layer and decoder layer on a large image dataset, so that the model can effectively capture and learn the feature representation under different degrees of haze conditions. Subsequently, the pre-trained video defogging network is migrated to the video defogging task through transfer learning, and fine-tuned on a small-scale video dataset, solving the problem of fewer high-quality public video datasets.
[0028] During the pre-training process, for the image dataset, the embodiment first unifies the format of the image dataset, converts all images into the standard RGB color space, and adjusts the image resolution to 512×512 to meet the input requirements of the model. Then, for different foggy conditions in the image dataset, the dataset is classified based on the fog degree of the image, and three sub-datasets of different levels of light fog, moderate fog, and heavy fog are generated, so that the model can conduct special learning on the defogging effect of different foggy scenes.
[0029] In addition, in order to enhance the robustness of the model, the image dataset has also undergone image enhancement processing, including image rotation, translation, mirror flipping and brightness adjustment, to improve the model's generalization ability for different viewing angles and lighting conditions. After image enhancement, the embodiment divides the image dataset into a training set, a validation set and a test set, where the training set is used for model parameter learning, the validation set is used to evaluate the model performance during training, and the test set is used to evaluate the final model effect. Through these preprocessing steps, it is ensured that the image data can effectively support the training and performance verification of the dehazing model.
[0030] In the calculation of pre-trained loss, the embodiment uses smooth L1 loss and perceptual loss. Smooth L1 loss calculates the pixel difference between the defogging output image data and the real fog-free image output data at the pixel level to determine the pre-trained smooth L1 loss. Perceptual loss first performs multi-level convolution encoding on the defogging output image data and the real fog-free image data to obtain the defogging image perceptual features and the real fog-free image perceptual features, calculates the feature difference between the defogging image perceptual features and the real fog-free image perceptual features at the feature level, and determines the pre-trained perceptual loss. The pre-trained video defogging network is obtained by iteratively training the initial defogging network according to the pre-trained smooth L1 loss and the pre-trained perceptual loss.
[0031] In some embodiments of the present invention, the multi-stage convolutional coding includes an initial convolutional layer and several inverted residual blocks, each of which includes a dilated convolution and a depth-wise separable convolution.
[0032] Specifically, considering that traditional convolutional neural networks have a large computational load, they are usually difficult to meet real-time requirements in actual application scenarios. In embedded devices or low-power devices, due to limited computing resources, large computational loads are difficult to run efficiently on the hardware platform. For this embodiment, a lightweight encoder based on MobileNetV2 is designed to efficiently extract multi-scale features of the input image.
[0033] In the encoder, for the preprocessed normalized input image , firstly, low-level feature extraction is performed through the initial convolution layer. The initial convolution operation can be expressed as:
[0034] in, represents the initial feature map, is the weight matrix of the initial convolutional layer.
[0035] Then the initial feature map after the initial convolution , the embodiment gradually extracts multi-scale high-level features through six inverted residual blocks , and the obtained series of multi-scale feature maps As the encoding feature, the size of the feature map is gradually reduced through the design of the inverted residual block, where each inverted residual block consists of a dilated convolution and a depth-wise separable convolution.
[0036] Among them, the main function of the extended convolution is to expand the number of input channels through the 1×1 convolution operation, allowing subsequent convolution operations to extract features in a higher-dimensional space. This convolution operation can not only increase the feature expression ability of the model, but also does not change the spatial dimension of the input feature map. The extended convolution operation can be expressed by the formula:
[0037] in, It is the weight matrix of the channel expansion convolution, which is used to expand the number of channels of the input feature map from the original lower number of channels to a higher dimensional feature space, so as to enhance the model's ability to extract features on different channels.
[0038] Then the expanded feature map Perform deep convolution. Deep convolution performs convolution operations independently on each channel, focusing on capturing spatial information in the image without interacting with information between channels. This convolution method significantly reduces computational complexity by avoiding cross-channel calculations. The deep convolution formula is expressed as:
[0039] in, It is the convolution kernel of the deep convolution, which performs a 3×3 convolution operation on each input channel independently to extract spatial features. This operation can preserve the spatial dimension of the feature map while greatly reducing the number of parameters and calculations.
[0040] After the deep convolution process is completed, the feature map Channel fusion and adjustment are performed through point-by-point convolution. The function of point-by-point convolution is to map the expanded number of channels to the specified number of output channels, and to fuse the features in different channels through cross-channel convolution operations. This step can further reduce the computational overhead while keeping the spatial dimension of the feature map unchanged. The point-by-point convolution formula is expressed as:
[0041] in, is the weight matrix of the point-by-point convolution, which is responsible for mapping the information in each channel to the specified number of output channels and realizing the interaction between different channel features. In this way, the model can integrate the features extracted on different channels, thereby enhancing the model's comprehensive understanding of the input image.
[0042] Finally, in order to prevent information loss during the convolution process, the inverted residual block also uses a skip connection to convert the input feature map Add directly to the feature map after the convolution operation In the process, the original information of the input features can be retained through the jump connection, and the back propagation of the gradient can be accelerated through the direct connection method, which is helpful for information transmission and optimization in the deep network. This step can be expressed as:
[0043] in, The activation function limits the value range of the output feature to prevent numerical overflow, while introducing nonlinearity and enhancing the expressiveness of the model. The jump connection ensures that the original input information can be directly transmitted to the output, allowing the model to retain more semantic information during feature extraction and alleviate the problem of gradient disappearance.
[0044] At this point, after multiple layers of stacking and processing of multiple inverted residual blocks, the multi-scale feature map obtained It can capture different levels of information of the image, from low-level edge features to high-level semantic features, ensuring multi-dimensional processing of the dehazing process. These feature maps will be passed to the decoder part of the subsequent steps for image reconstruction and detail recovery.
[0045] In some embodiments of the present invention, multi-level upsampling decoding includes layer-by-layer deconvolution upsampling operations and batch normalization operations, and the output features of the batch normalization operations at each level are jump-connected and fused with the output features of the inverted residual blocks at the corresponding level to serve as the input features of the deconvolution upsampling operations at the next level.
[0046] Specifically, for the decoder part, the embodiment constructs a decoder structure combining skip connection and deconvolution upsampling to gradually restore the image resolution and utilize the multi-scale feature map generated by the encoder. Perform feature reconstruction.
[0047] First, the decoder performs layer-by-layer deconvolution operations to transform the low-resolution feature maps of each layer into Restore to a higher resolution, where the size of the deconvolution kernel is set to 3×3, the stride is set to 2, and the padding is set to 1. The feature map size of the deconvolution output can be expressed as:
[0048] in, is the size of the input feature map, is the stride, is the convolution kernel size, is the fill value.
[0049] No. Feature map after deconvolution of the layer It is expressed as:
[0050] in, is the target size after deconvolution, is the first deconvolution layer of the decoder (the layers of the decoder are arranged from bottom to top).
[0051] After the deconvolution operation, in order to ensure the stability of training, the embodiment introduces a batch normalization operation, which standardizes the input distribution of the feature map after deconvolution, reduces the gradient vanishing and gradient exploding problems, and speeds up the training convergence. The batch normalization operation is as follows:
[0052] After batch normalization, in order to maintain the detailed features of the image during the decoding process, the decoder also uses a jump connection. In the embodiment, the specific jump connection method is: the encoder feature map With decoder feature map Connected, and Connected, and Connected, and Connected, and The fused feature map after the jump connection The formula is:
[0053] in, represents the feature map after skip connection, Represents the feature map of the encoder Feature map with decoder The concatenation operation is performed on the channel dimension. Through this fusion operation, the decoder can simultaneously utilize the high-level semantic information in the encoder and the detailed features in the deconvolution process, speeding up information transmission while ensuring the quality of image reconstruction.
[0054] For each fusion feature map after skip connection fusion , the embodiment gradually restores the number of channels through the convolution layer, thereby retaining more high-resolution information. The entire decoding process ensures that the high-level semantic features of the encoder are retained and effective enhancement is achieved in image detail restoration through repeated deconvolution and jump connection operations. Finally, a high-resolution image after defogging is generated This process can be expressed as:
[0055] in, Denotes the deconvolution operation of the decoder, which is responsible for extracting key information from the fused features and restoring image details. The decoder's deconvolution works in conjunction with the jump connection to ensure that the features of the input image at different scales are fully utilized, achieving efficient dehazing and image reconstruction.
[0056] In some embodiments of the present invention, Figure 3 FIG. 1 is a flow chart of calculating the smoothed L1 loss and the perceptual loss according to an embodiment of the present invention. Figure 3 As shown, the pre-trained smooth L1 loss is determined according to the defogging output image data and the real fog-free image data, the defogging output image data and the real fog-free image data are multi-level convolutionally encoded to obtain the defogging image perception features and the real fog-free image perception features, and the pre-trained perception loss is determined according to the defogging image perception features and the real fog-free image perception features, including: S301, calculating the pixel difference between the defogging output image data and the real fog-free image data based on the mean square error method and L1 loss to obtain a pre-trained smooth L1 loss; S302, performing multi-level convolution coding on the defogging output image data and the real fog-free image data to obtain the defogging image perception features and the real fog-free image perception features; S303, calculating the feature difference between the dehazed image perception features and the real haze-free image perception features based on the Euclidean distance method to obtain the pre-trained perception loss.
[0057] Specifically, considering that the traditional convolutional neural network only trains the defogging network based on pixel differences at the image level, it is difficult to keep the structure and texture details consistent with the original image. In the pre-training task and in the transfer learning task, the embodiment designs a composite loss function that combines the smooth L1 loss and the perceptual loss based on MobileNetV2 to ensure that the defogging image can be close to the real fog-free image in both the pixel space and the feature space.
[0058] First, the smooth L1 loss optimizes the dehazing effect by measuring the difference between the dehazing output image and the real haze-free image at the pixel level. Different from the standard L1 loss, the smooth L1 loss is more robust and can use the mean square error form when the error is small to ensure accuracy. When the error is large, it can be converted to L1 loss to reduce the impact of outliers on training. The formula is defined as:
[0059] Therefore, the calculation formula of smooth L1 loss can be expressed as:
[0060] in, represents the total number of pixels in the image, and represent the true haze-free image and the dehazed output image. The role of this loss function is to ensure that the output image is consistent with the real haze-free image in the pixel space, which is especially important for the recovery of fine details.
[0061] However, since the loss function that simply relies on the pixel space may result in insufficient reconstruction of the image at the high-level semantic features, the embodiment designs a perceptual loss based on MobileNetV2, which extracts the multi-level features of the image by using the encoder of the pre-trained network and calculates the dehazed output image. and real fog-free image In the feature space, the reconstructed image is not only close to the real scene in the pixel space, but also consistent in the feature space. Especially when dealing with video defogging tasks, the perceptual loss can help the model better understand the temporal consistency of video frames and avoid the video "flickering" phenomenon caused by the discontinuity of the fog removal process. First, the embodiment will defog the output image and real fog-free image Input the pre-trained encoder based on MobileNetV2 respectively to obtain the dehazed image perception features and real fog-free image perception features , and use this to calculate the perceptual loss, where Indicates that MobileNetV2 is in The calculation formula of the perceptual loss for the features extracted by the layer is expressed as:
[0062] in, Represents the Euclidean distance, which is used to measure the difference between the dehazed image and the real haze-free image in the feature space. Corresponding to image information of different scales, from low-level edge features to high-level semantic features, through this multi-scale feature comparison, perceptual loss can effectively improve the global structure and semantic consistency of the dehazed image.
[0063] Finally, the composite loss function It combines the smooth L1 loss and the perceptual loss, and balances the contribution of pixel loss and feature loss through weights. The formula is expressed as:
[0064] in, and It is a hyperparameter used to control the relative weight of smooth L1 loss and perceptual loss. By adjusting these two hyperparameters, we can flexibly control the trade-off between the model's pixel-level reconstruction and high-level semantic consistency, making the model more adaptable in different scenarios.
[0065] The design of the composite loss function not only improves the quality of the dehazed image, but also reduces the problem of detail loss or texture distortion that may occur under complex weather conditions through reasonable feature space alignment. Through this multi-level loss optimization, the model can generate images with high visual quality and structural coherence in the dehazed task.
[0066] In some embodiments of the present invention, the dehazing output data is subjected to bidirectional optical flow estimation based on adaptive pyramid and cost volume improvement to obtain bidirectional optical flow loss, including: Arrange the defogging output data in the order of video frames, and use every three consecutive frames as a set of optical flow input data, where each set of optical flow input data includes a forward frame image, a current frame image, and a backward frame image; The optical flow input data is input into the preset adaptive feature pyramid for multi-scale feature extraction, deconvolution upsampling and feature fusion to obtain the optical flow features; Performing cost volume method optical flow estimation and frame deformation operations on the optical flow features to obtain a forward prediction frame image and a backward prediction frame image; The forward optical flow loss is determined according to the forward predicted frame image and the current frame image, the backward optical flow loss is determined according to the backward predicted frame image and the current frame image, and the forward optical flow loss and the backward optical flow loss are combined to obtain the bidirectional optical flow loss.
[0067] Specifically, in order to improve the temporal consistency of the video defogging network in processing the video defogging task, the embodiment also introduces a bidirectional optical flow loss when training the video defogging network. To calculate the bidirectional optical flow loss, the embodiment designs a lightweight preset bidirectional optical flow network model based on an adaptive feature pyramid. The preset bidirectional optical flow network model converts each adjacent three-frame video image sequence in the defogging input video into a single image. As a set of optical flow input data, where Indicates the previous frame image a while ago, Indicates the current frame image. Represents the next frame image of the next frame. Through the adaptive feature pyramid, the network generates multi-scale feature maps at different resolutions ,in Indicates the number of layers of the feature pyramid. The design of the feature pyramid enables the network to capture both low-resolution global motion information and high-resolution detail motion information, ensuring that optical flow estimation can cope with complex motion scenes. This multi-scale feature extraction helps improve the accuracy and robustness of optical flow estimation. For each frame of image , extract multi-scale feature maps , which can be expressed as:
[0068] in, Represents the convolutional network in the feature pyramid. The feature pyramid structure is built from bottom to top in a decreasing resolution manner. There are 4 layers in total. Each layer is based on the original resolution and the reduction ratio is ,in Representation hierarchy. Low-level pyramid features (high-resolution features) mainly capture local motion details, while high-level pyramid features (low-resolution features) are used to capture global motion information.
[0069] In addition, the multi-layer features contained in the feature pyramid not only contain high-resolution detail features, but also extract low-resolution global features through layer-by-layer downsampling. In order to combine global and local feature information in subsequent optical flow estimation, a deconvolution upsampling module is also designed to pass the global information of the low-resolution layer upward to the high-resolution layer and perform feature fusion. The formula for feature fusion is expressed as:
[0070] in, represents the deconvolution upsampling operation, is the fused optical flow feature. This operation not only retains the local motion information of the high-resolution features, but also transfers the global motion trend contained in the low-resolution features to the high-resolution layer, effectively enhancing the robustness of the optical flow estimation. Through this layer-by-layer convolution operation, the feature map not only covers the local details of the motion, but also reflects the large-scale motion changes. The multi-scale processing of the pyramid can capture useful features in the case of motion blur or heavy fog, and improve the generalization ability of the model.
[0071] After obtaining the optical flow features, the embodiment introduces a cost volume structure when calculating the optical flow. The cost volume can capture the displacement relationship of pixels with lower computational complexity, avoiding the high computational overhead caused by brute force matching of all pixels. At the same time, the use of multi-scale feature maps also makes the optical flow estimation more accurate and can adapt to motion modes in different scenarios.
[0072] In the process of calculating optical flow, the specific method is to calculate the forward optical flow and backward optical flow , respectively representing the number of frames from the previous frame To current frame and the next frame To current frame The motion estimation formula is expressed as:
[0073]
[0074] in, represents the pixel position, and represents pixel displacement, Indicates The cost volume measures the similarity of two feature maps under different pixel displacements, providing a basis for optical flow estimation. By comparing the feature maps of adjacent frames, the network gradually regresses the forward optical flow. and backward optical flow , these two optical flow fields describe the pixel motion information between video frames.
[0075] Then, the optical flow is estimated based on the constructed cost volume. This process is to obtain the forward optical flow field through layer-by-layer regression of the deep convolutional network. and backward optical flow field , the formula is:
[0076]
[0077] At this time, the optical flow field obtained Indicates Layer midframe To Frame sports information, Indicates Layer midframe To Frame sports information.
[0078] Then, in order to accurately estimate the forward optical flow and the backward optical flow, the embodiment uses the forward optical flow as the basis. and backward optical flow , perform frame deformation operation, and transform the previous frame and the next frame The pixel position of the current frame is transformed into The pixel positions of and backward prediction frame image This operation uses bilinear interpolation method, so that the pixel position can be accurately transformed. The formula is expressed as
[0079]
[0080] in, Represents the value of the predicted frame image at the pixel obtained by forward optical flow deformation, Represents the value of the predicted frame image at the pixel obtained by backward optical flow deformation. Through the action of forward and backward optical flow, the frame deformation operation can align the pixel information of adjacent frames to the current frame, which makes the predicted frame highly consistent with the real fog-free image of the current frame. In this way, not only can the information of adjacent frames be effectively fused, but also the network can be helped to obtain the motion trajectory between frames during the defogging process, thereby improving the consistency between frames. In the video defogging task, pixel-level alignment deformation can significantly reduce the discontinuity between frames and improve the smoothness of the video.
[0081] The last embodiment defines a bidirectional optical flow loss function to measure the pixel difference between the predicted frame and the real current frame. The forward optical flow loss is calculated by and the current frame image The L1 loss is used to measure the error between the two, and the formula is expressed as:
[0082] in, is the pixel coordinate, is the total number of pixels.
[0083] Backward frame sampling similarity method to calculate backward prediction frame image and the current frame image L1 loss:
[0084] The bidirectional optical flow loss is obtained by combining the losses of forward optical flow and backward optical flow, and the formula is expressed as:
[0085] in, and Represents the weight hyperparameters of the forward and backward optical flow losses.
[0086] In some embodiments of the present invention, Figure 4 FIG. 1 is a flow chart of iterative training of a pre-trained video defogging network according to an embodiment of the present invention. Figure 4 As shown in the figure, the pre-trained video dehazing network is iteratively parameter-tuned and end-to-end trained according to the smooth L1 loss, perceptual loss, and bidirectional optical flow loss to obtain a fully trained video dehazing network, including: S401, determining the total defogging loss according to the smooth L1 loss, the perceptual loss and the bidirectional optical flow loss; S402, keeping the network parameters of other parts of the pre-trained defogging network unchanged, and adjusting the network parameters of each preset part to be adjusted of the pre-trained defogging network in turn according to the total defogging loss; S403, performing end-to-end training of the pre-trained video defogging network using a cosine annealing strategy, a learning rate warm-up strategy, and a gradient clipping strategy according to the total defogging loss, and iterating the training process to obtain a fully trained video defogging network.
[0087] Specifically, in order to adapt the model to the video defogging task, based on the pre-trained model, the embodiment optimizes the model by fine-tuning parameters and small sample transfer learning, combined with bidirectional optical flow loss. The input is the collected foggy video dataset, represented as a frame sequence ,in The total number of frames in the video.
[0088] In transfer learning, considering that the pre-trained model has obtained excellent image defogging capabilities on a large haze picture dataset, in order to maintain this capability, the embodiment performs transfer learning on the model by fine-tuning parameters of the pre-trained model.
[0089] For the encoder part, the embodiment fine-tunes the last two high-level feature layers of the encoder to adapt to the dynamic change characteristics between video frames. The high-level feature layer is used to capture the timing information in the video, helping the model to achieve inter-frame consistency and defogging effect in the video defogging task. During the transfer learning process, the high-level feature layer of the encoder gradually updates the parameters through back propagation to adapt to the dynamic changes in the timing of the video. The update formula of the high-level feature layer is as follows:
[0090] in, and Represent the new parameters and initial parameters of the high-level feature layer respectively, is the learning rate, which is used to control the step size of parameter update. is the total loss function, which is optimized To reduce the error of video defogging.
[0091] For the decoder part, the embodiment fine-tunes parameters to adapt to the dynamic changes and temporal consistency between video frames. For the decoder, its role is to restore the spatial resolution of the input image layer by layer, and fuse the feature map generated by the encoder with the corresponding layer in the decoder through upsampling and skip connections. In order to ensure temporal consistency in the video defogging task, the key layers near the output end of the decoder are selectively unfrozen so that they can be adjusted and updated according to the characteristics of the video frame.
[0092] The embodiment first unfreezes the last convolution layer, which is responsible for mapping the upsampled feature map to the final defogging image. This layer directly affects the quality of the video frame output, so the parameters of this layer need to be unfrozen during the transfer learning process so that it can adapt to the temporal variation characteristics in video defogging. By fine-tuning this layer, the model can learn how to deal with dynamic changes in the video, thereby generating a defogging result with temporal consistency. Its convolution formula is expressed as:
[0093] in, is the activation function, and are the weights and biases of the convolutional layer respectively. Through fine-tuning, and Updated dynamically based on the video frame.
[0094] Then unfreeze the penultimate upsampling module, which is responsible for restoring the low-resolution feature map to a higher resolution and performing jump connections with the high-level features in the encoder. For the video dehazing task, the purpose of unfreezing this layer is to enable the model to better handle changes between video frames. The upsampling operation is implemented through deconvolution, and the formula is expressed as:
[0095] in, is the upsampled feature map generated after the deconvolution operation, is the input feature map, is an adjustable parameter matrix. This layer can be adjusted for different dynamic scenes in the video dehazing task, so that the model can restore the resolution of the video layer by layer and ensure smooth transitions between frames.
[0096] In addition to the fine-tuning of the above layers, the intermediate layers of the decoder in the embodiment remain frozen. These layers are mainly responsible for gradually upgrading the low-level features in the encoder to high-level semantic information. Since these features are common in image and video defogging tasks, keeping their parameters unchanged can reduce the risk of overfitting while maintaining the original defogging ability of the model. The parameters of the frozen intermediate layers are expressed as:
[0097] in, and They represent the intermediate parameters of the decoder after and before parameter adjustment respectively.
[0098] At the same time, considering that in the process of video dehazing, the difference between frames may cause the single-frame features to be insufficient to reflect the continuous changes in the video. By fine-tuning the jump connection module in the decoder, the model can more effectively fuse the features of different levels in the encoder, ensuring that the feature transfer effect between video frames is enhanced while maintaining semantic consistency. The fusion formula of the jump connection is:
[0099] By unfreezing the relevant parts of the decoder, the skip connection module can adjust the fusion mode of features according to the temporal continuity, thereby generating a more coherent video frame sequence.
[0100] In addition to the fine-tuning of the above layers, the intermediate layers of the decoder in the embodiment remain frozen. These layers are mainly responsible for gradually upgrading the low-level features in the encoder to high-level semantic information. Since these features are common in image and video defogging tasks, keeping their parameters unchanged can reduce the risk of overfitting while maintaining the original defogging ability of the model. The parameters of the frozen intermediate layers are expressed as:
[0101] in, and They represent the intermediate parameters of the decoder after and before parameter adjustment respectively.
[0102] In addition, in this process, the embodiment continues to sample the smooth L1 loss function and the perceptual loss as the main task loss function, and introduces the bidirectional optical flow loss on this basis to construct an optimization strategy of multiple loss functions. The total loss function formula is expressed as:
[0103] in , is the bidirectional optical flow loss function, the main task loss The formula is:
[0104] in, is the L1 smoothing loss, is the perceived loss, and is a hyperparameter that controls the loss weight.
[0105] Then, after completing parameter fine-tuning, the embodiment introduces a variety of optimization strategies in combination with the total loss function to perform end-to-end training on the pre-trained video dehazing network, including a cosine annealing strategy, a learning rate warm-up strategy, a gradient clipping strategy, and a learning rate decay strategy.
[0106] During the training process, the embodiment introduces the Adam optimizer to effectively speed up the training process and improve the generalization ability of the model. Adam maintains the first-order moment estimate and the second-order moment estimate of each parameter and uses this information to update the model weights. For each input sequence , the model processes the input sequence and calculates the total loss function To update the weights, the Adam optimizer’s update rule formula is:
[0107]
[0108]
[0109]
[0110] in, and denote the first-order momentum and second-order momentum of the gradient, respectively. and is a hyperparameter that controls the momentum decay rate, is the learning rate, A tiny constant to prevent division by zero.
[0111] In order to prevent the learning rate from dropping too fast during training and causing the model to fall into a local optimum early on, the embodiment adopts a cosine annealing learning rate adjustment strategy. This strategy dynamically adjusts the learning rate , which enables the model to update parameters at a higher learning rate in the early stage of training and accelerate the convergence speed.
[0112] As training progresses, the learning rate gradually decays to ensure that the model can converge smoothly in the later stages and improve the final performance. The formula is as follows:
[0113] in, Indicates The learning rate of the step, and are the minimum and maximum learning rates set respectively to ensure that the upper and lower limits of the learning rate are controlled. represents the total number of training steps, is the number of steps in the current training.
[0114] Considering the different functions of different layers of the model, the low-level network structure of the encoder is mainly used to learn low-level features of the image (such as edges and textures), while the high-level network structure focuses more on capturing more abstract deep-level features. Therefore, in the embodiment, a layered learning rate strategy is adopted here, and different learning rates are assigned to different layers according to the frozen state of the layer and the task characteristics to achieve better fine-tuning effect.
[0115] More specifically, for the frozen low-level part of the encoder, these layers are responsible for extracting common low-level features, such as edges, textures, and illumination changes. When migrating to the video dehazing task, these low-level features are usually highly adaptable and can maintain consistent recognition effects under different concentrations of haze conditions. In addition, the pre-trained model has achieved strong feature extraction capabilities based on a large-scale haze dataset. Therefore, in order to avoid over-updating these common features in the new task, these layers are frozen and the learning rate is set to 0 to keep their weights unchanged, ensuring that the model can effectively retain the low-level dehazing features in the pre-trained model during the transfer learning phase, thereby focusing on high-level temporal consistency and detail reconstruction.
[0116] Similarly, for the frozen decoder part, these layers are mainly used to restore the basic structure and detail information of the video frame in the defogging task. The generally applicable defogging features have been learned in the pre-training of single-frame image defogging. These basic layers of the decoder mainly focus on low-level reconstruction operations, such as removing foggy patches and strengthening contours in the image. These features do not change much in different defogging tasks and are also applicable to video defogging tasks. Therefore, these layers are frozen and the learning rate is set to 0 to retain their original defogging capabilities, ensure stable performance in new tasks, and reduce computational overhead.
[0117] For the last two high-level layers of the unfrozen encoder, these layers extract high-level features related to specific tasks. In transfer learning, these layers need to be fine-tuned according to the characteristics of the new data to adapt to the dynamic features in the video data and the requirements of the defogging task. Therefore, the learning rate range is set to , This ensures that high-level features can be smoothly fine-tuned in transfer learning to prevent excessive parameter updates from destroying the effect of pre-training.
[0118] For the two layers of the unfrozen decoder, only the last two layers are unfrozen in the decoder to restore high-resolution features and ensure the consistency and dehazing effect of the output. The learning rate range is set to , , only for the two unfrozen layers, ensuring that the decoder part can quickly adapt to the reconstruction requirements of new tasks.
[0119] To prevent the learning rate from being too low in the early stage of training, which would affect the effective training of the model, the embodiment adopts the warm-up learning strategy. It will start from a smaller initial value Gradually increase to the set , which helps the model quickly enter an effective training state. The formula is as follows:
[0120] in, Represents the initial learning rate in the warm-up phase. Its setting varies with different layers. For the high-level part of the unfrozen encoder, the initial learning rate Set to 0.00005, corresponding to the maximum learning rate For the two layers of the unfrozen decoder, the initial learning rate is Set to 0.0005, corresponding to the maximum learning rate , Represents the duration of the Warm-up phase, i.e., during this period, the learning rate gradually changes from Increase to the corresponding maximum learning rate to ensure that each layer enters a stable training state.
[0121] At the same time, in order to prevent the gradient from being too large during the back propagation process, which leads to unstable training, the embodiment also adopts a gradient clipping strategy. After each back propagation, a gradient clipping operation is performed to ensure that the norm of the gradient does not exceed a preset threshold. , which is initially set to This operation can effectively avoid the gradient explosion problem caused by excessive gradients, thereby ensuring the smoothness of the training process and the stable update of network parameters. The formula for gradient clipping is as follows:
[0122] Before the optimizer performs parameter updates, the gradient clipping function is called after each gradient calculation to limit the size of the gradient and ensure the stability of the network. With preset gradient threshold For comparison, if the norm of the gradient exceeds this threshold, it will be reduced by a preset ratio to keep it within a reasonable range. According to the depth of the network and the characteristics of multi-task training, this preset gradient threshold ensures that the model can prevent instability caused by excessive gradients during training.
[0123] The embodiment can significantly enhance the robustness of training by introducing a gradient clipping strategy, especially in multi-task learning, to prevent the problem of excessive gradients caused by a certain loss term, and ensure that each task can be reasonably optimized. At the same time, by limiting the maximum norm of the gradient, the model can avoid instability during training, especially in deep networks or complex tasks, and can better control the amplitude of gradient updates.
[0124] In addition, the embodiment dynamically adjusts the learning rate according to the training loss of the model in each iteration, and the learning rate decay strategy uses an exponential decay formula, which is expressed as:
[0125] in, The embodiment can effectively prevent the network from repeatedly deviating from the optimal solution due to an excessively large learning rate in the later stage of training by gradually reducing the learning rate during the training process, thereby ensuring that the network converges stably when approaching the optimal solution.
[0126] Finally, the weight is updated by combining the historical gradient information weights. The update formula is:
[0127]
[0128] in, is the momentum term, is the momentum coefficient, is the gradient of the loss function with respect to the weight.
[0129] In the end-to-end training, the embodiment adopts the above method, combines the total loss function of smooth L1 loss, perceptual loss and bidirectional optical flow loss through the cosine annealing strategy, the learning rate warm-up strategy and the gradient clipping strategy, and finally obtains a fully trained video defogging network after multiple rounds of optimization, so that the pre-trained video defogging network can effectively adapt to the video defogging task after training, and obtains a fully trained video defogging network.
[0130] In summary, in order to improve the accuracy of video defogging in details and time consistency, the present invention first inputs the acquired foggy video data and the corresponding real fog-free video data into a pre-trained video defogging network, performs multi-level convolution encoding and multi-level upsampling decoding on the foggy video data to obtain defogging output data; then, a smooth L1 loss is determined according to the defogging output data and the real fog-free video data, multi-level convolution encoding is performed on the defogging output data and the real fog-free video data to obtain defogging perception features and real perception features, and perceptual loss is determined according to the defogging perception features and the real perception features, and bidirectional optical flow estimation based on adaptive pyramid and cost volume improvement is performed on the defogging output data to obtain bidirectional optical flow loss; finally, the pre-trained video defogging network is iteratively parameterized and end-to-end trained according to the smooth L1 loss, perceptual loss and bidirectional optical flow loss to obtain a fully trained video defogging network. During the training process, the present invention calculates the loss from the pixel level through smoothing L1 loss, calculates the loss at the feature level through perceptual loss, improves the accuracy of the network in details, obtains the bidirectional optical flow loss through improving the bidirectional optical flow estimation, captures the dynamic information between video frames, improves the time consistency of the network in the video defogging task, enables the video defogging network to effectively improve the defogging effect with lower computational complexity, and obtains a video defogging network that takes into account the defogging effect, low power consumption and real-time performance.
[0131] The present invention also provides a video defogging network application method based on bidirectional optical flow loss, combined with Figure 5 Come and see, Figure 5 A schematic diagram of a flow chart of an embodiment of a video defogging network application method based on bidirectional optical flow loss provided by the present invention, as shown in FIG. Figure 5 As shown, the video defogging network application method based on bidirectional optical flow loss includes: S501, obtaining a video to be defogged; S502: Input the video to be defogged into a well-trained video defogging network to obtain a defogged output video.
[0132] Specifically, in the application process of the video defogging network, it is first necessary to effectively obtain the video to be defogged, and then input the video to be defogged into a well-trained video defogging network. After the video defogging network completes the defogging task, all defogged frames are merged in time order to obtain a frame sequence, and combined with the frame rate setting, it is converted into a video file in a standard format to obtain the defogged output video, and the frame rate of the final output video remains the same as the frame rate of the input video.
[0133] like Figure 6 The present invention also provides an embedded device 600. The embedded device 600 includes a processor 601, a memory 602 and a display 603. Figure 6Only some components of the embedded device 600 are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0134] In some embodiments, the processor 601 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 602, such as the video defogging network training method based on bidirectional optical flow loss and / or the video defogging network application method based on bidirectional optical flow loss in the present invention.
[0135] In some embodiments, the memory 602 may be an internal storage unit of the embedded device 600, such as a hard disk or memory of the embedded device 600. In other embodiments, the memory 602 may also be an external storage device of the embedded device 600, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the embedded device 600.
[0136] Furthermore, the memory 602 may include both an internal storage unit of the embedded device 600 and an external storage device. The memory 602 is used to store application software installed in the embedded device 600 and various data.
[0137] In some embodiments, the display 603 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 603 is used to display information of the embedded device 600 and to display a visual user interface. The components 601-603 of the embedded device 600 communicate with each other through a system bus.
[0138] In one embodiment, when the processor 601 executes the video defogging network training program in the memory 602, the following steps may be implemented: The acquired foggy video data and the corresponding real fog-free video data are input into the pre-trained video defogging network, and the foggy video data is subjected to multi-level convolution encoding and multi-level upsampling decoding to obtain the defogging output data; Determine the smooth L1 loss according to the defogging output data and the real haze-free video data, perform multi-level convolution encoding on the defogging output data and the real haze-free video data to obtain the defogging perception features and the real perception features, determine the perception loss according to the defogging perception features and the real perception features, and perform the bidirectional optical flow estimation based on the adaptive pyramid and cost volume improvement on the defogging output data to obtain the bidirectional optical flow loss; The pre-trained video dehazing network is iteratively parameter-fine-tuned and end-to-end trained according to the smooth L1 loss, perceptual loss, and bidirectional optical flow loss to obtain a fully trained video dehazing network.
[0139] In one embodiment, when the processor 601 executes the video defogging network application in the memory 602, the following steps may be implemented: Get the video to be dehazed; The video to be dehazed is input into the well-trained video dehazing network to obtain the dehazed output video.
[0140] It should be understood that: when the processor 601 executes the video defogging network training program and / or the video defogging network application program in the memory 602, in addition to the above functions, other functions can also be implemented. For details, please refer to the description of the corresponding method embodiments above.
[0141] Furthermore, the embodiment of the present invention does not specifically limit the type of the embedded device 600 mentioned, and the embedded device 600 may be a portable device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, etc. Exemplary embodiments of portable embedded devices include but are not limited to portable embedded devices equipped with IOS, Android, Microsoft or other operating systems. The above-mentioned portable embedded device may also be other portable embedded devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the present invention, the embedded device 600 may not be a portable embedded device, but a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0142] Accordingly, an embodiment of the present application also provides a computer-readable storage medium, which is used to store computer-readable programs or instructions. When the program or instructions are executed by a processor, it can implement the steps or functions of the video defogging network training method based on bidirectional optical flow loss and / or the video defogging network application method based on bidirectional optical flow loss provided in the above-mentioned method embodiments.
[0143] Those skilled in the art will appreciate that all or part of the processes of the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium, wherein the computer-readable storage medium is a disk, an optical disk, a read-only storage memory, or a random access memory, etc.
[0144] The above is a detailed introduction to the video dehazing network training method, application method, embedded device and medium based on bidirectional optical flow loss provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A video dehazing network training method based on bidirectional optical flow loss, characterized in that: include: The obtained foggy video data and the corresponding real fog-free video data are input into the pre-trained video defogging network, and the foggy video data is multi-level convolutionally encoded and multi-level upsampling decoded to obtain defogging output data; Determine a smooth L1 loss according to the defogging output data and the real fog-free video data, perform multi-level convolution encoding on the defogging output data and the real fog-free video data to obtain defogging perception features and real perception features, determine a perception loss according to the defogging perception features and the real perception features, and perform bidirectional optical flow estimation based on adaptive pyramid and cost volume improvement on the defogging output data to obtain a bidirectional optical flow loss; The pre-trained video dehazing network is iteratively parameter-fine-tuned and end-to-end trained according to the smooth L1 loss, the perceptual loss, and the bidirectional optical flow loss to obtain a fully trained video dehazing network.
2. The video dehazing network training method based on bidirectional optical flow loss according to claim 1 is characterized in that: The pre-trained video defogging network is obtained by training an initial defogging network, and the training initial defogging network includes: The acquired foggy image data and the corresponding real fog-free image data are input into the initial defogging network, and the foggy image data is multi-level convolutionally encoded and multi-level upsampling decoded to obtain defogging output image data; A pre-trained smooth L1 loss is determined according to the defogging output image data and the real haze-free image data, multi-level convolution encoding is performed on the defogging output image data and the real haze-free image data to obtain defogging image perception features and real haze-free image perception features, a pre-trained perceptual loss is determined according to the defogging image perception features and the real haze-free image perception features, and the initial defogging network is iteratively trained according to the pre-trained smooth L1 loss and the pre-trained perceptual loss to obtain the pre-trained video defogging network.
3. The video dehazing network training method based on bidirectional optical flow loss according to claim 1 is characterized in that: The multi-stage convolutional coding includes an initial convolutional layer and a plurality of inverted residual blocks, each of which includes an extended convolution and a depth-separable convolution.
4. The video defogging network training method based on bidirectional optical flow loss according to claim 3 is characterized in that: The multi-level upsampling decoding includes layer-by-layer deconvolution upsampling operations and batch normalization operations, and the output features of the batch normalization operation at each level are skip-connected and fused with the output features of the inverted residual block at the corresponding level to serve as the input features of the deconvolution upsampling operation of the next level.
5. The video dehazing network training method based on bidirectional optical flow loss according to claim 2 is characterized in that: The method of determining a pre-trained smoothed L1 loss according to the defogging output image data and the real fog-free image data, performing multi-level convolution encoding on the defogging output image data and the real fog-free image data to obtain defogging image perception features and real fog-free image perception features, and determining a pre-trained perception loss according to the defogging image perception features and the real fog-free image perception features, comprises: Calculating the pixel difference between the defogging output image data and the real non-fogging image data based on the mean square error method and the L1 loss to obtain a pre-trained smoothing L1 loss; Performing multi-level convolution coding on the defogging output image data and the real fog-free image data to obtain defogging image perception features and real fog-free image perception features; The pre-trained perceptual loss is obtained by calculating the feature difference between the dehazed image perceptual features and the real haze-free image perceptual features based on the Euclidean distance method.
6. The video dehazing network training method based on bidirectional optical flow loss according to claim 1 is characterized in that: The step of performing an adaptive pyramid and cost volume-based improved bidirectional optical flow estimation on the defogging output data to obtain a bidirectional optical flow loss includes: Arranging the defogging output data in the order of video frames, and taking every three consecutive frames as a group of optical flow input data, wherein each group of the optical flow input data includes a forward frame image, a current frame image and a backward frame image; Inputting the optical flow input data into a preset adaptive feature pyramid to perform multi-scale feature extraction, deconvolution upsampling and feature fusion to obtain optical flow features; Performing cost volume method optical flow estimation and frame deformation operation on the optical flow features to obtain a forward prediction frame image and a backward prediction frame image; A forward optical flow loss is determined according to the forward predicted frame image and the current frame image, a backward optical flow loss is determined according to the backward predicted frame image and the current frame image, and the forward optical flow loss and the backward optical flow loss are combined to obtain a bidirectional optical flow loss.
7. The video dehazing network training method based on bidirectional optical flow loss according to claim 1 is characterized in that: The iteratively performing parameter fine-tuning and end-to-end training on the pre-trained video defogging network according to the smooth L1 loss, the perceptual loss, and the bidirectional optical flow loss to obtain a fully trained video defogging network includes: Determine a total dehazing loss according to the smoothed L1 loss, the perceptual loss, and the bidirectional optical flow loss; Keeping the network parameters of other parts of the pre-trained defogging network unchanged, adjusting the network parameters of each preset part to be adjusted of the pre-trained defogging network in turn according to the total defogging loss; The pre-trained video defogging network is end-to-end trained using a cosine annealing strategy, a learning rate warm-up strategy, and a gradient clipping strategy according to the total defogging loss, and a fully trained video defogging network is obtained through an iterative training process.
8. A video defogging network application method based on bidirectional optical flow loss, characterized in that: include: Get the video to be dehazed; Inputting the video to be defogged into a well-trained video defogging network to obtain a defogged output video; The fully trained video dehazing network is determined according to the video dehazing network training method based on bidirectional optical flow loss according to any one of claims 1 to 7.
9. An embedded device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the video defogging network training method based on bidirectional optical flow loss according to any one of claims 1 to 7 and / or the video defogging network application method based on bidirectional optical flow loss according to claim 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the video defogging network training method based on bidirectional optical flow loss according to any one of claims 1 to 7 and / or the video defogging network application method based on bidirectional optical flow loss according to claim 8 are implemented.
Citation Information
Patent Citations
Image defogging method of convolutional neural network based on fusion Transform
CN116012253A
Super-resolution reconstruction method and device based on deep learning
CN116645270A
Image defogging method based on multi-receptive-field feature fusion and mixed attention
CN116912130A
Self-supervised smoke removal method for operation video
CN117876262A
Multi-frame video interpolation using optical flow
US20190138889A1
Cited By
Deep learning image defogging method based on image alignment driving
CN120725921A
A deep learning image defogging method based on image alignment driving
CN120725921B