Video defogging method and system based on multi-scale spatio-temporal fusion network
Patent Information
- Application Number
- CN202411213355.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-08-30
AI Technical Summary
然而,现有的视频除雾方法存在一些局限性,从基于物理模型的分量估计中获得无雾帧,中间预测不准确,导致最终结果误差积累
[0056]本发明提供的基于多尺度时空融合网络的视频去雾方法及系统,其通过分帧处理生成视频序列图像,应用自动色彩均衡技术进行预处理,以优化视觉效果;采用基于通道注意力机制的编码器,分别从五个维度实现对图像的色彩恢复和逐层特征提取,有效保留图像中的低频信息;通过低频信息传递模块进一步传递浅层特征中的低频信息,动态调整特征响应以提高浅层特征的利用效率;通过解码器进行时空特征对齐和融合,增强了模型对时序和空间信息的捕捉能力;通过重建模块将融合后的高维特征转换为目标RGB格式的图像,确保输出图像的空间分辨率和色彩精度。整体上,本发明及系统提供了一种高效、精确的视频去雾技术,适用于智慧交通、智能驾驶等下游应用场景。
Smart Images

Figure CN119107268B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a video dehazing method and system based on a multi-scale spatiotemporal fusion network. Background Technology
[0002] Dehazing algorithms are an important research direction in computer vision, aiming to restore and enhance hazy images, making them clearer and more realistic, and effectively dealing with interference from severe weather and environments. With the development of intelligent driving and smart transportation, video data containing haze can significantly impact downstream tasks, including object detection and image segmentation. To reduce the resulting video degradation, dehazing of video data is essential. Haze in the atmosphere is caused by the scattering of light by tiny particles suspended in the atmosphere; this optical effect leads to visual degradation of images and videos.
[0003] The perceptual sensitivity of visual content is impaired due to the complex interactions of scattering and absorption phenomena caused by tiny atmospheric particles. Consequently, the resulting images suffer from reduced contrast and a white blurring effect. Such images, characterized by blurriness and substantial degradation, lead to the loss of critical information. Strategic deployment of defogging technology offers a viable approach to correcting the widespread fog-induced degradation in images and frames.
[0004] Video dehazing benefits from temporal cues such as highly correlated haze thickness and lighting conditions, as well as moving foreground objects and background. However, existing video dehazing methods have some limitations. Obtaining haze-free frames from physics-based component estimations leads to inaccurate intermediate predictions, resulting in accumulated errors in the final result. Secondly, these methods aggregate temporal information using input feature overlay or frame-to-frame alignment within local sliding windows, making it difficult to obtain global and long-term temporal information. Summary of the Invention
[0005] This invention provides a video dehazing method and system based on a multi-scale spatiotemporal fusion network, which solves the technical problem of how to perform high-quality dehazing on video images.
[0006] To address the above technical problems, this invention provides a video dehazing method based on a multi-scale spatiotemporal fusion network, comprising the following steps:
[0007] S1. Perform frame segmentation on the haze video to be processed to generate a sequence of original image frames composed of multiple original image frames in time order.
[0008] S2. The original image frame sequence is preprocessed using automatic color equalization technology to obtain the corresponding preprocessed image frame sequence.
[0009] S3. Input the original image frame sequence and the preprocessed image frame sequence into the constructed multi-scale spatiotemporal fusion network for dehazing to obtain a dehazed image frame sequence;
[0010] The multi-scale spatiotemporal fusion network includes an encoder, a low-frequency information transmission module, a decoder, and a reconstruction module. The encoder module takes the current frame, the three frames preceding the current frame (a total of four preprocessed image frames), and the original image frame as input, and extracts features from five different dimensions to obtain five-dimensional feature-enhanced image frames. The low-frequency information transmission module extracts low-frequency information from the five-dimensional feature-enhanced image frames to obtain corresponding five-dimensional low-frequency information, which is then input to the decoder. The decoder takes the fifth-dimensional feature-enhanced image frame and the five-dimensional low-frequency information as input, performs spatiotemporal feature shifting and spatiotemporal feature fusion, and obtains feature-fused image frames. The reconstruction module reconstructs the feature-fused image frames to obtain four corresponding fog-free image frames.
[0011] Furthermore, the encoder is sequentially equipped with five feature extraction units, and the encoder's processing flow includes the following steps:
[0012] B1. Input the four preprocessed image frames and the original image frames into the first feature extraction unit to extract features in the first dimension, and obtain four feature-enhanced image frames in the first dimension.
[0013] B2. Input the four feature-enhanced image frames of the first dimension and the four original image frames into the second feature extraction unit to extract features in the second dimension, and obtain four feature-enhanced image frames in the second dimension.
[0014] B3. Input the four feature-enhanced image frames of the second dimension and the four original image frames into the third feature extraction unit to extract features in the third dimension, and obtain the four feature-enhanced image frames of the third dimension.
[0015] B4. Input the four feature-enhanced image frames of the third dimension and the four original image frames into the fourth feature extraction unit to extract features in the fourth dimension, and obtain the four feature-enhanced image frames of the fourth dimension.
[0016] B5. Input the four feature-enhanced image frames of the fourth dimension and the four original image frames into the fifth feature extraction unit to extract features in the fifth dimension, and obtain the four feature-enhanced image frames of the fifth dimension.
[0017] B6. Input the four-frame feature-enhanced image frames of the five dimensions into the low-frequency information transmission module, and input the four-frame feature-enhanced image frames of the fifth dimension into the encoder.
[0018] Furthermore, the feature extraction process for each feature extraction unit includes the following steps:
[0019] G1, convolve the input feature map and then merge it with the input feature map F. G0 Perform channel connections to obtain feature map F G1 ;
[0020] G2, Transfer feature map F G1 Convolution, PReLU activation function, and convolution are performed sequentially to generate feature map F. G2 ;
[0021] G3, transfer feature map F G2 Global average pooling, convolution, PReLU activation function, convolution, and sigmoid function are performed sequentially to generate weight information; the weight information is then compared with the feature map F. G2 Multiplication yields the feature map F G3 ;
[0022] G4, Input feature map F G0 With feature map F G3 Adding them together yields the feature map F. G4 ;
[0023] G5, Transfer feature map F G4 By performing convolution, convolution, and convolution again, the feature map F is obtained. G5 .
[0024] Furthermore, the low-frequency information transmission module includes four low-frequency information transmission units, which are respectively used to extract the low-frequency information of the feature-enhanced image frames output by the second to fifth feature extraction units;
[0025] The operation of each of the low-frequency information transmission units includes the following steps:
[0026] L1. Extract shallow features: Extract initial low-frequency features from the feature extraction unit;
[0027] L2, Transmitting Shallow Features: Using long-distance hop connections, the initial low-frequency features are directly transmitted to the decoder.
[0028] Furthermore, during the propagation process in step L2, a multi-scale mapping unit is introduced to process the shallow features. The operations performed by the multi-scale mapping unit include the following steps:
[0029] L21. Four sets of feature maps of different scales are generated from the input feature map by four convolution operations with different kernel sizes.
[0030] L22. By performing element-wise summation, the four sets of feature maps at different scales are fused to obtain a fused feature map.
[0031] L23. Extract global information from the fused feature map using global average pooling;
[0032] L24. Compact features are obtained by passing global information through a fully connected layer;
[0033] L25. The compact features are passed through four fully connected layers to obtain the weights corresponding to each scale.
[0034] L26. The weights are summed with the input feature maps at the corresponding scales to obtain the final output feature map.
[0035] Following the multi-scale mapping unit, there is also a step:
[0036] L27. Reconstruct the output feature map of the multi-scale mapping unit to generate a high-quality reconstructed feature map, which is then input into the decoder.
[0037] Furthermore, the decoder includes four spatiotemporal feature shifting and fusion units. The feature-enhanced image frame from the fifth feature extraction unit and the reconstructed feature map output from the fourth low-frequency information transmission unit are input into the first spatiotemporal feature shifting and fusion unit for the first spatiotemporal feature shifting and fusion to obtain a first fused feature map. The first fused feature map and the reconstructed feature map output from the third low-frequency information transmission unit are input into the second spatiotemporal feature shifting and fusion unit for the second spatiotemporal feature shifting and fusion to obtain a second fused feature map. The second fused feature map and the reconstructed feature map output from the second low-frequency information transmission unit are input into the third spatiotemporal feature shifting and fusion unit for the third spatiotemporal feature shifting and fusion to obtain a third fused feature map. The third fused feature map and the reconstructed feature map output from the first low-frequency information transmission unit are input into the fourth spatiotemporal feature shifting and fusion unit for the fourth spatiotemporal feature shifting and fusion to obtain a fourth fused feature map, which is then input into the reconstruction module.
[0038] Furthermore, the operation of each spatiotemporal feature shifting and fusion unit specifically includes the following steps:
[0039] T1, Aggregate the current frame f i The previous frame f adjacent to it i-1 Aggregate the current frame f i The two preceding frames f adjacent to it i-1 f i-2 Aggregate the current frame f i The three adjacent frames f i-1 f i-2 f i-3 Three sets of aggregation features were obtained;
[0040] T2. Three spatiotemporal shifting modules are used to spatiotemporally shift the three sets of aggregated features respectively to obtain three sets of shifted aggregated features;
[0041] T3. The three sets of shift aggregation features are fused to obtain the fused aggregation features;
[0042] T4. Upsample the fused and aggregated features to obtain high-dimensional features, which are then input into the reconstruction module.
[0043] Furthermore, the operations performed by the spatiotemporal shift module include the following steps:
[0044] T21. Spatially shift a set of aggregated features from the input to obtain spatially shifted features;
[0045] T22, Applying a multi-head self-attention mechanism to spatial shift features and the current frame f i A portion of the process generates multiple local windows, and for each local window representing a spatial shift feature, it generates corresponding K-value matrices and V-value matrices. For each current frame f... i A local window is used to generate the Q-value matrix;
[0046] T23. Fuse the K-value matrix, V-value matrix, and Q-value matrix of each local window to obtain the local window fusion feature; fuse all local window fusion features to obtain the shift aggregation feature corresponding to the set of aggregation features.
[0047] Furthermore, for the current frame f i The previous frame f adjacent to it i-1 The first set of aggregation features, step T21 specifically includes the following steps:
[0048] T211, Set the current frame f i Divided into f i a and f i b Two parts, the previous frame f i-1 Divided into and Two parts;
[0049] T212, Feature group Slicing along the channel dimension yields M feature slices, where the m-th feature slice is denoted as... Where m = 1, ..., M are slice indices;
[0050] T213, For each feature piece Spatially shift it in the x and y directions to obtain the translated feature patch. Group all features along the channel dimension Series connection yields spatial displacement characteristics.
[0051] T214, f i b and Perform series connection, and analyze the series connection results and spatial displacement characteristics. After performing convolution, the result is added to the concatenated result. Then, the added result is convolved again and added to the concatenated result once more to obtain the first feature map.
[0052] T215, to After convolution and The sum is then convolved and added again to obtain the second feature map.
[0053] T216. Connect the first feature map and the second feature map in series to obtain the spatial shift feature of the first set of aggregated features;
[0054] T217, f i a As the current frame f i Part of the input window has a multi-head self-attention mechanism.
[0055] The present invention also provides a video dehazing system based on a multi-scale spatiotemporal fusion network, the key feature of which is that it has an intelligent agent, which is used to execute the above-mentioned video dehazing method based on a multi-scale spatiotemporal fusion network.
[0056] This invention provides a video dehazing method and system based on a multi-scale spatiotemporal fusion network. It generates video image sequences through frame-by-frame processing, applies automatic color equalization for preprocessing to optimize visual effects, employs an encoder based on a channel attention mechanism to achieve color restoration and layer-by-layer feature extraction from five dimensions, effectively preserving low-frequency information in the image. A low-frequency information transmission module further transmits low-frequency information from shallow features, dynamically adjusting feature responses to improve the utilization efficiency of shallow features. A decoder performs spatiotemporal feature alignment and fusion, enhancing the model's ability to capture temporal and spatial information. A reconstruction module converts the fused high-dimensional features into a target RGB format image, ensuring the spatial resolution and color accuracy of the output image. Overall, this invention and system provide an efficient and accurate video dehazing technology suitable for downstream applications such as intelligent transportation and autonomous driving. Attached Figure Description
[0057] Figure 1 This is an overall structural diagram of a multi-scale spatiotemporal fusion network provided in an embodiment of the present invention;
[0058] Figure 2 This is a structural diagram of the first feature extraction unit (GFEB1) provided in this embodiment of the invention;
[0059] Figure 3 This is a structural diagram of the Spatiotemporal Feature Shifting and Fusion Unit (STAF) provided in an embodiment of the present invention;
[0060] Figure 4This is a structural diagram of the spacetime shift module (STFS) provided in an embodiment of the present invention. Detailed Implementation
[0061] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. The embodiments are given for illustrative purposes only and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0062] The video dehazing method based on a multi-scale spatiotemporal fusion network provided in this invention includes the following steps:
[0063] S1. Perform frame segmentation on the haze video to be processed to generate a sequence of original image frames composed of multiple original image frames in time order.
[0064] S2. The original image frame sequence is preprocessed using automatic color equalization technology to obtain the corresponding preprocessed image frame sequence;
[0065] S3. Input the original image frame sequence and the preprocessed image frame sequence into the constructed multi-scale spatiotemporal fusion network for dehazing to obtain a dehazed image frame sequence.
[0066] (1) Step S1
[0067] Each frame of the smog video was extracted independently to provide basic data for subsequent processing steps.
[0068] (2) Step S2
[0069] Step S2 applies Automatic Color Equalization (ACE) to each original image frame, improving the overall visual effect by adjusting the image's brightness and contrast, and performing preliminary color correction. Then, the color-equalized image is saved as one of the inputs to the multi-scale spatiotemporal fusion network, providing optimized image data for subsequent processing steps.
[0070] Specifically, for each original image frame, the automatic color equalization technique includes the following steps:
[0071] S21. Color Spatial Adjustment: Adjusts the brightness, contrast, and color balance of the image to ensure that the image is more visually balanced and natural, completes the color difference correction of the image, and can significantly improve the visibility of the image, making it more suitable for subsequent processing.
[0072] S22. Dynamic Extension: Dynamically extend the corrected image to improve its contrast and brightness. The dynamic extension formula is as follows:
[0073]
[0074] Where R(x) represents the pixel value of pixel x in the original image frame obtained after step S21, and L(x) represents the pixel value of pixel x obtained after expansion. min R represents the minimum value among all pixel values in the entire original image frame after step S21. max This represents the maximum value among all pixel values in the entire original image frame after step S21. By subtracting the minimum value and dividing by the range, R(x) is transformed into the interval [0,1], thereby enhancing the contrast of the image.
[0075] Through the steps described above, ACE generates images with high contrast and enhanced brightness. ACE can be viewed as a simplified model of the human visual system; its enhancement process aligns with human perception, helping to generate higher-quality images that better match visual perception during video processing.
[0076] (3) Step S3
[0077] 1) Overall structure of multi-scale spatiotemporal fusion network
[0078] The structure of multi-scale spatiotemporal fusion networks is as follows: Figure 1 As shown, the system includes an encoder, a low-frequency information transmission module, a decoder, and a reconstruction module. The encoder module takes the current frame, the three preceding frames (a total of four preprocessed image frames), and the original image frame as input, and extracts features from five different dimensions to obtain five-dimensional feature-enhanced image frames. The low-frequency information transmission module extracts low-frequency information from the five-dimensional feature-enhanced image frames, inputting the corresponding five-dimensional low-frequency information into the decoder. The decoder takes the fifth-dimensional feature-enhanced image frame and the five-dimensional low-frequency information as input, performing spatiotemporal feature shifting and spatiotemporal feature fusion to obtain feature-fused image frames. The reconstruction module reconstructs the feature-fused image frames to obtain the corresponding four fog-free image frames.
[0079] The encoder, based on a channel attention mechanism, performs color restoration and layer-by-layer feature extraction from five dimensions, effectively preserving low-frequency information in the image. Subsequently, a low-frequency information transfer module further transfers low-frequency information from shallow features through long-distance skip connections and multi-scale mapping units, dynamically adjusting feature responses to improve the utilization efficiency of shallow features. The decoder enhances the model's ability to capture temporal and spatial information through grouped spatial displacement and window multi-head self-attention mechanisms. Finally, the reconstruction module converts the fused high-dimensional features into the target RGB format image, ensuring the spatial resolution and color accuracy of the output image.
[0080] 2) Encoder
[0081] The structure of the encoder is as follows Figure 1As shown, there are five feature extraction units (CFEB1 to CFEB5) arranged sequentially. The encoder's processing flow includes the following steps:
[0082] B1. Input the four preprocessed image frames and the original image frames into the first feature extraction unit (GFEB1) to extract features in the first dimension, and obtain four feature-enhanced image frames in the first dimension.
[0083] B2. Input the four feature-enhanced image frames of the first dimension and the four original image frames into the second feature extraction unit (GFEB2) to extract features in the second dimension, and obtain four feature-enhanced image frames in the second dimension.
[0084] B3. Input the four feature-enhanced image frames of the second dimension and the four original image frames into the third feature extraction unit (GFEB3) to extract features in the third dimension, and obtain four feature-enhanced image frames in the third dimension.
[0085] B4. Input the four feature-enhanced image frames of the third dimension and the four original image frames into the fourth feature extraction unit (GFEB4) to extract features in the fourth dimension, and obtain four feature-enhanced image frames in the fourth dimension.
[0086] B5. Input the four feature-enhanced image frames of the fourth dimension and the four original image frames into the fifth feature extraction unit (GFEB5) to extract features in the fifth dimension, and obtain four feature-enhanced image frames in the fifth dimension.
[0087] B6. Input the four-frame feature-enhanced image frames of the five dimensions into the low-frequency information transmission module, and input the four-frame feature-enhanced image frames of the fifth dimension into the encoder.
[0088] As an example, the input and output dimensions of the five feature extraction units are as follows:
[0089] GFEB1: Initial input dimension is 3, final output dimension is 64;
[0090] GFEB2: Initial input dimension is 64+12, final output dimension is 96;
[0091] GFEB3: The initial input dimension is 96+24, and the final output dimension is 192;
[0092] GFEB4: The initial input dimension is 192+48, and the final output dimension is 384;
[0093] GFEB5: The initial input dimension is 384+96, and the final output dimension is 768.
[0094] The feature extraction process is similar for each GFEB unit. Taking GFEB1 as an example, refer to... Figure 2The structure diagram of GFEB1 shown below illustrates the feature extraction process of GFEB1, which includes the following steps:
[0095] G1, convolve the input feature map and then merge it with the input feature map F. G0 Perform channel connection (Cat operation) to obtain feature map F. G1 ;
[0096] G2, Transfer feature map F G1 The model performs convolution (conv, input dimension: 3, output dimension: 3, kernel: 3×3, stride: 1), PReLU activation function (applies non-linear transformation to the convolutional feature map to improve the model's expressive power), and convolution (conv, input dimension: 3, output dimension: 3, kernel: 3×3, stride: 1) sequentially to generate feature map F. G2 ;
[0097] G3, transfer feature map F G2 The algorithm sequentially performs global average pooling (extracting global feature information to provide contextual information for the attention mechanism), convolution (input dimension: 3, output dimension: 1, kernel: 1×1, stride: 1), PReLU activation function (processing the channel feature maps while maintaining non-linearity), convolution (input dimension: 1, output dimension: 3, kernel: 1×1, stride: 1, providing finer feature adjustment), and Sigmoid function (generating an attention weight map for adaptively adjusting the feature weights of each channel), generating weight information. This weight information is then compared with the feature map F. G2 Multiplication yields the feature map F G3 ;
[0098] G4, Input feature map F G0 With feature map F G3 Adding the features (reducing the attenuation of low-frequency information and preserving more image details) yields the feature map F. G4 ;
[0099] G5, Transfer feature map F G4 The following convolutions are performed sequentially: (input dimension: 3, output dimension: 64, kernel: 3×3, stride: 1), (convolution operation on the input feature map, input dimension: 64, output dimension: 64, kernel: 3×3, stride: 2), and (input dimension: 64, output dimension: 64, kernel: 3×3, stride: 1), to obtain the feature map F. G5 .
[0100] 3) Low-frequency information transmission module
[0101] After obtaining the high-quality feature map processed by the encoder, in order to retain the low-frequency information in the shallow features extracted by the encoder and ensure that this information can be used efficiently, this embodiment directly transmits the low-frequency information to the decoder through long-distance skip connections, and introduces a multi-scale mapping unit in this process.
[0102] like Figure 1 As shown, the low-frequency information transmission module includes four low-frequency information transmission units LFITM1 to LFITM4, which are used to extract low-frequency information from the feature-enhanced image frames output by GFEB2 to GFEB5, respectively.
[0103] The structure of each low-frequency information transmission unit is identical. The operation of any low-frequency information transmission unit includes the following steps:
[0104] L1. Extract shallow features: Extract initial low-frequency features from the feature extraction unit;
[0105] L2, Propagating Shallow Features: Using long-distance skip connections, the initial low-frequency features are directly passed to the decoder. This prevents these shallow features from being lost in the deep network, thus preserving more image details.
[0106] During the transmission process in step L2, a multi-scale mapping unit is introduced to process the shallow features so as to better integrate them with the deep features extracted by the decoder.
[0107] Multi-scale mapping unit: Starting from multiple scales, several mapping layers with different receptive field sizes are designed, and global information representations are obtained based on them. Furthermore, softmax attention guided by information from different scales is used to obtain weighted representations for each scale, and finally, these weighted representations are summed. Specific steps include:
[0108] L21, Multi-scale Convolution Operation: Four sets of feature maps of different scales are generated from the input feature map by using four convolution operations with different kernel sizes (1×1, 3×3, 5×5, and 7×7 respectively).
[0109] L22, Multi-scale feature fusion: By performing element-wise summation, four sets of feature maps at different scales are fused to obtain a fused feature map;
[0110] L23, Global Average Pooling: Extracts global information from the fused feature map;
[0111] L24, Fully Connected Layer: Global information is processed through a fully connected layer to obtain compact features, thereby achieving more accurate and adaptive guidance;
[0112] L25, Four fully connected layers: The compact features are passed through four fully connected layers to obtain the weights corresponding to each scale;
[0113] L26, Feature Weighting: The weights are summed with the input feature maps of the corresponding scales to obtain the final output feature map.
[0114] Following the multi-scale mapping unit, there is a further step:
[0115] L27, Feature Reconstruction: The output feature map of the multi-scale mapping unit is reconstructed to generate a high-quality reconstructed feature map that is input into the decoder.
[0116] By employing a low-frequency information transmission module, we can effectively retain and utilize low-frequency information in shallow features, fully integrate all extracted high- and low-dimensional features, improve the overall feature expression capability, and ensure the stability of the training process.
[0117] 4) Decoder
[0118] like Figure 2 As shown, the decoder includes four spatiotemporal feature shifting and fusion units, STAF1 and STAF4. The feature-enhanced image frame from the fifth feature extraction unit and the reconstructed feature map output from the fourth low-frequency information transmission unit are input into the first spatiotemporal feature shifting and fusion unit STAF1 for the first spatiotemporal feature shifting and fusion to obtain the first fused feature map. The first fused feature map and the reconstructed feature map output from the third low-frequency information transmission unit are input into the second spatiotemporal feature shifting and fusion unit STAF2 for the second spatiotemporal feature shifting and fusion to obtain the second fused feature map. The second fused feature map and the reconstructed feature map output from the second low-frequency information transmission unit are input into the third spatiotemporal feature shifting and fusion unit STAF3 for the third spatiotemporal feature shifting and fusion to obtain the third fused feature map. The third fused feature map and the reconstructed feature map output from the first low-frequency information transmission unit are input into the fourth spatiotemporal feature shifting and fusion unit STAF4 for the fourth spatiotemporal feature shifting and fusion to obtain the fourth fused feature map, which is then input into the reconstruction module.
[0119] The operation of each spatiotemporal feature shifting and fusion unit is the same, refer to... Figure 3 The diagram shown illustrates the structure of the spatiotemporal feature shifting and fusion unit. The specific steps involved in operating each spatiotemporal feature shifting and fusion unit are as follows:
[0120] T1, Aggregate the current frame f i The previous frame f adjacent to it i-1 Aggregate the current frame f i The two preceding frames f adjacent to it i-1 f i-2 Aggregate the current frame f i The three adjacent frames f i-1 f i-2 f i-3 Three sets of aggregation features were obtained;
[0121] T2. Three spatiotemporal shift modules (STFS) are used to spatiotemporally shift the three sets of aggregated features to obtain three sets of shifted aggregated features (so that the model can better capture spatial information and align features);
[0122] T3. The three sets of shift aggregation features are fused to obtain the fused aggregation features;
[0123] T4. Upsample the fused and aggregated features to obtain the high-dimensional feature input reconstruction module.
[0124] Among them, reference Figure 4 The diagram shown illustrates the structure of the spacetime shift module (STFS). The operations performed by the STFS module include the following steps:
[0125] T21. Spatially shift a set of aggregated features from the input to obtain spatially shifted features;
[0126] T22, Applying Window Multi-Head Self-Attention (W-MSA) mechanism to spatial shift characteristics and the current frame f i A portion of the process generates multiple local windows, and for each local window representing a spatial shift feature, it generates a corresponding K(Key) value matrix and V(Value) value matrix. For each current frame f... i A local window generates a Q(Query) value matrix;
[0127] T23. Fuse the K-value matrix, V-value matrix, and Q-value matrix of each local window to obtain the local window fusion feature; fuse all local window fusion features to obtain the shift aggregation feature corresponding to the set of aggregation features (further improving the expressive power of the features, enabling the model to better capture and utilize spatial information).
[0128] like Figure 4 As shown, for the current frame f i The previous frame f adjacent to it i-1 The first set of aggregation features, step T21 specifically includes the following steps:
[0129] T211, Set the current frame f i Divided into f i a and f i b Two parts, the previous frame f i-1 Divided into and Two parts;
[0130] T212, Feature slice: Feature group M feature slices are obtained along the channel dimension, where the m-th feature slice is denoted as... Where m = 1, ..., M are slice indices;
[0131] T213, Translation Feature Patch: For each feature patch Apply Δx in the x and y directions m ,Δy m Spatial shift of pixels ∈{-9,-5,0,5,9} yields the translated feature patch.
[0132]
[0133] |Δx m |=k x *(s-1)+1,|Δy m |=k y *(s-1)+1
[0134] Where Shift() represents a shift operation, k x k y It is an integer, and s is defined as the base length of the spatial shift. We set s to zero when the spatial shift results in blank pixels in the boundary. For a Δx m Pixel displacement, the corresponding feature set in space by Δx m -1 pixel displacement, followed by a 3×3 convolution in the depth direction, which spans two displacement processing objects and achieves smooth translation between two adjacent displacement feature slices. Then all features are grouped along the channel dimension. By concatenating these components, we obtain the spatial displacement characteristics.
[0135]
[0136] T214, f i b and Perform concatenation to analyze the concatenation results and spatial displacement characteristics. After performing convolution, the result is added to the concatenated result. Then, the added result is convolved again and added to the concatenated result once more to obtain the first feature map.
[0137] T215, to After convolution and The sum is then convolved and added again to obtain the second feature map.
[0138] T216. Connect the first feature map and the second feature map in series to obtain the spatial shift feature of the first set of aggregated features;
[0139] T217, f i a As the current frame f i Part of the input window multi-head self-attention mechanism (W-MSA).
[0140] For including the current frame f i The two preceding frames f adjacent to it i-1 f i-2 The second set of aggregated features, step T21 is similar to steps T211 to T216 above, except that the generated second feature map incorporates f i-2 Spatial shift characteristics. For the current frame f... i The three adjacent frames f i-1 f i-2 f i-3 The second set of aggregated features, step T21 is similar to steps T211 to T216 above, except that the generated second feature map incorporates f i-2 The spatial shift characteristics were then used to generate the addition of f. i-3 The third feature map of spatial displacement features is obtained, and finally the first feature map, the second feature map and the third feature map are concatenated.
[0141] The window multi-head self-attention mechanism (W-MSA) specifically includes the following steps:
[0142] T221. Window Division: The feature map is divided into multiple local windows of the same size, each containing a specific number of pixel blocks.
[0143] T222, Multi-head Self-Attention Mechanism: Within each local window, the query, key, and value matrices are calculated separately. A multi-head self-attention mechanism is used to calculate the attention weight at each position, and the features are then summed using a weighted average. This mechanism allows features at each position to interact with information from other positions within the window, thereby enhancing the feature representation capability.
[0144] T223. Feature Fusion: This method fuses the features within each window to generate a new feature map. This further enhances the expressive power of the features, enabling the model to better capture and utilize spatial information.
[0145] Step T3 specifically involves stacking the three sets of spatiotemporally shifted features. An attention mechanism is used to model temporal dependencies, improving the understanding and processing capabilities of temporal features. The result of the multiplication is then subjected to Softmax normalization. These processed features are then aggregated with the features of the current frame. By fusing information from different time points, a more comprehensive feature representation is generated. An upsampling structure is used between adjacent STFS modules, employing PixelShufflePack to perform pixel shuffling of the channels, reducing dimensionality while increasing resolution.
[0146] More specifically, for the first STFS module, step T3 includes the following steps:
[0147] T31. Spatiotemporal Feature Stacking: Three sets of features processed by STFS are stacked together to form a feature set containing information across multiple temporal dimensions. These feature sets represent image information at different points in time, used to capture subtle changes over time. Assuming the three feature sets are F1, F2, and F3, the stacked feature can be represented as:
[0148] F stack =Stack(F1,F2,F3)
[0149] Stacked features F stack The dimensions are (N,3,C,H,W), where N is the batch size, C is the number of channels, and H and W are the height and width, respectively.
[0150] T32. Query Feature Generation: Extracting query features F from the current frame. query And flatten it into a two-dimensional matrix Q:
[0151] Q = Flatten(F) query )
[0152] T33. Relevance Calculation: Combine the two-dimensional matrix Q of the query features with the stacked feature matrix F. stack Perform matrix multiplication to calculate the correlation between the current frame and features from different time points. Stacked features F stack Converted into two-dimensional matrices K and V (through a flattening operation):
[0153] K = Flatten(F) stack )
[0154] V = Flatten(F) stack )
[0155] Then, the correlation between query feature Q and key K is calculated:
[0156]
[0157] Here, A has a dimension of (N×H×W,3), representing the correlation between features at different time points, and the superscript T indicates matrix transpose.
[0158] This step models the dependencies in the time dimension through an attention mechanism, enabling the model to better understand and process temporal features.
[0159] T34. Attention Weighting and Normalization: The above correlation result A is normalized using the Softmax function.
[0160] α = Softmax(A)
[0161] Here, α represents the weights assigned to features at different time points. The Softmax function can assign weights to features at different time points, allowing highly relevant features to receive greater attention, thereby enhancing the expressive power of the features.
[0162] T35. Feature Aggregation: Finally, the values are weighted and summed using the attention weight α to obtain the aggregated features.
[0163] F out =α·V
[0164] Finally, the aggregated features F out Features F of the current frame query To merge:
[0165] F final =Concat(F out ,F query )
[0166] Finally, it goes through a convolutional layer for further processing and optimization of the final feature representation.
[0167] T36. Upsampling: The feature map generated in step T35 is upsampled through convolution and pixel shuffling, reducing dimensionality while increasing resolution. The specific process is as follows:
[0168] T361, Convolutional Layer: Performs convolution operations on the input feature map to expand the number of channels, preparing for subsequent pixel shuffling.
[0169] T362, Pixel shuffling: Performs pixel shuffling on the convolutional feature map. Pixel shuffling improves the spatial resolution of the feature map and rearranges the increased number of channels into spatial dimensions.
[0170] T37. Return to the upsampled feature map and perform the above operation for the next STFS module.
[0171] 5) Reconstruction Module
[0172] The reconstruction module first performs a convolution operation on the input feature map, then applies a non-linear activation function to extract and enhance its salient features. Next, through further convolution operations, the number of channels in the feature map is gradually reduced from a high dimension to the three channels of the target output, ensuring that the final result matches the RGB format of the target image. Throughout this process, the spatial resolution of the feature map is preserved, thus ensuring that the features after non-linear transformation can be effectively extracted and integrated, providing a solid foundation for accurate image reconstruction.
[0173] The specific process of rebuilding the module is as follows:
[0174] R1, Convolutional Layer: Performs convolution operation on the input feature map, with a kernel size of 3×3 and a stride of 1, keeping the spatial dimensions of the output image unchanged.
[0175] R2, LeakyReLU activation function: The output of the first convolutional layer is activated using the LeakyReLU function with a negative slope parameter of 0.1, so that the output is 0.1 times the input value when the input is negative.
[0176] R3, Convolutional Layer: After receiving the LeakyReLU activation, it is processed using a 1×1 convolutional kernel to reduce the number of channels to the target output of 3 channels. The stride is 1, and the padding is 0, maintaining the spatial dimensions of the output image.
[0177] To apply the above method, this embodiment of the invention also provides a video dehazing system based on a multi-scale spatiotemporal fusion network, which includes an intelligent agent for executing the above-described video dehazing method based on a multi-scale spatiotemporal fusion network.
[0178] In summary, this invention provides a video dehazing method and system based on a multi-scale spatiotemporal fusion network. It generates video image sequences through frame-by-frame processing and applies automatic color equalization for preprocessing to optimize visual effects. An encoder based on a channel attention mechanism is employed to achieve color restoration and layer-by-layer feature extraction from five dimensions, effectively preserving low-frequency information in the image. A low-frequency information transmission module further transmits low-frequency information from shallow features, dynamically adjusting feature responses to improve the utilization efficiency of shallow features. A decoder performs spatiotemporal feature alignment and fusion, enhancing the model's ability to capture temporal and spatial information. A reconstruction module converts the fused high-dimensional features into a target RGB format image, ensuring the spatial resolution and color accuracy of the output image. Overall, this invention and system provide an efficient and accurate video dehazing technology suitable for downstream applications such as intelligent transportation and autonomous driving.
[0179] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A video dehazing method based on multi-scale spatiotemporal fusion networks, characterized in that, Including the following steps: S1. Perform frame segmentation on the haze video to be processed to generate a sequence of original image frames composed of multiple original image frames in time order. S2. The original image frame sequence is preprocessed using automatic color equalization technology to obtain the corresponding preprocessed image frame sequence. S3. Input the original image frame sequence and the preprocessed image frame sequence into the constructed multi-scale spatiotemporal fusion network for dehazing to obtain a dehazed image frame sequence; The multi-scale spatiotemporal fusion network includes an encoder, a low-frequency information transmission module, a decoder, and a reconstruction module. The encoder module takes the current frame, the three frames preceding it (a total of four preprocessed image frames), and the original image frame as input, and extracts features from five different dimensions to obtain five-dimensional feature-enhanced image frames. The low-frequency information transmission module extracts low-frequency information from the five-dimensional feature-enhanced image frames, inputting the corresponding five-dimensional low-frequency information into the decoder. The decoder takes the fifth-dimensional feature-enhanced image frame and the five-dimensional low-frequency information as input, performs spatiotemporal feature shifting and spatiotemporal feature fusion, and obtains feature-fused image frames. The reconstruction module reconstructs the feature-fused image frames to obtain four corresponding fog-free image frames. The low-frequency information transmission module introduces a multi-scale mapping unit to process shallow features. The operations performed by the multi-scale mapping unit include the following steps: L21. Four sets of feature maps of different scales are generated from the input feature map by four convolution operations with different kernel sizes. L22. By performing element-wise summation, the four sets of feature maps at different scales are fused to obtain a fused feature map. L23. Extract global information from the fused feature map using global average pooling; L24. Compact features are obtained by passing global information through a fully connected layer; L25. The compact features are passed through four fully connected layers to obtain the weights corresponding to each scale. L26. The weights are summed with the input feature maps at the corresponding scales to obtain the final output feature map. Following the multi-scale mapping unit, the following steps are further provided: L27. Reconstruct the output feature map of the multi-scale mapping unit to generate a high-quality reconstructed feature map, which is then input into the decoder. The spatiotemporal feature shift includes the following steps: T21. Spatially shift a set of aggregated features from the input to obtain spatially shifted features; T22, Applying a multi-head self-attention mechanism to spatial shift features and the current frame. A portion of the process generates multiple local windows, and for each local window representing a spatial shift feature, it generates corresponding K-value matrices and V-value matrices. For each current frame... A local window is used to generate the Q-value matrix; T23. Fuse the K-value matrix, V-value matrix, and Q-value matrix of each local window to obtain the local window fusion feature; All local window fusion features are fused to obtain the shifted aggregated features corresponding to this set of aggregated features.
2. The video dehazing method based on a multi-scale spatiotemporal fusion network according to claim 1, characterized in that: The encoder has five feature extraction units arranged sequentially, and the encoder's processing flow includes the following steps: B1. Input the four preprocessed image frames and the original image frames into the first feature extraction unit to extract features in the first dimension, and obtain four feature-enhanced image frames in the first dimension. B2. Input the four feature-enhanced image frames of the first dimension and the four original image frames into the second feature extraction unit to extract features in the second dimension, and obtain four feature-enhanced image frames in the second dimension. B3. Input the four feature-enhanced image frames of the second dimension and the four original image frames into the third feature extraction unit to extract features in the third dimension, and obtain the four feature-enhanced image frames of the third dimension. B4. Input the four feature-enhanced image frames of the third dimension and the four original image frames into the fourth feature extraction unit to extract features in the fourth dimension, and obtain the four feature-enhanced image frames of the fourth dimension. B5. Input the four feature-enhanced image frames of the fourth dimension and the four original image frames into the fifth feature extraction unit to extract features in the fifth dimension, and obtain the four feature-enhanced image frames of the fifth dimension. B6. Input the four-frame feature-enhanced image frames of the five dimensions into the low-frequency information transmission module, and input the four-frame feature-enhanced image frames of the fifth dimension into the encoder.
3. The video dehazing method based on a multi-scale spatiotemporal fusion network according to claim 2, characterized in that, The feature extraction process for each feature extraction unit includes the following steps: G1, convolve the input feature map and then merge it with the input feature map F. G0 Perform channel connections to obtain feature map F G1 ; G2, Transfer feature map F G1 Convolution, PReLU activation function, and convolution are performed sequentially to generate feature map F. G2 ; G3, transfer feature map F G2 Global average pooling, convolution, PReLU activation function, convolution, and sigmoid function are performed sequentially to generate weight information; the weight information is then compared with the feature map F. G2 Multiplication yields the feature map F G3 ; G4, Input feature map F G0 With feature map F G3 Adding them together yields the feature map F. G4 ; G5, Transfer feature map F G4 By performing convolution, convolution, and convolution again, the feature map F is obtained. G5 .
4. The video dehazing method based on a multi-scale spatiotemporal fusion network according to claim 3, characterized in that: The low-frequency information transmission module includes four low-frequency information transmission units, which are used to extract the low-frequency information of the feature-enhanced image frames output by the second to fifth feature extraction units, respectively. The operation of each of the low-frequency information transmission units includes the following steps: L1, Extracting shallow features: Extract initial low-frequency features from the feature extraction unit; L2, Transmitting Shallow Features: Using long-distance hop connections, the initial low-frequency features are directly transmitted to the decoder.
5. The video dehazing method based on a multi-scale spatiotemporal fusion network according to claim 1, characterized in that: The decoder includes four spatiotemporal feature shifting and fusion units. The feature-enhanced image frame from the fifth feature extraction unit and the reconstructed feature map output from the fourth low-frequency information transmission unit are input into the first spatiotemporal feature shifting and fusion unit for the first spatiotemporal feature shifting and fusion to obtain a first fused feature map. The first fused feature map and the reconstructed feature map output from the third low-frequency information transmission unit are input into the second spatiotemporal feature shifting and fusion unit for the second spatiotemporal feature shifting and fusion to obtain a second fused feature map. The second fused feature map and the reconstructed feature map output from the second low-frequency information transmission unit are input into the third spatiotemporal feature shifting and fusion unit for the third spatiotemporal feature shifting and fusion to obtain a third fused feature map. The third fused feature map and the reconstructed feature map output from the first low-frequency information transmission unit are input into the fourth spatiotemporal feature shifting and fusion unit for the fourth spatiotemporal feature shifting and fusion to obtain a fourth fused feature map, which is then input into the reconstruction module.
6. The video dehazing method based on a multi-scale spatiotemporal fusion network according to claim 5, characterized in that, The specific steps involved in the operation of each spatiotemporal feature shifting and fusion unit are as follows: T1, Aggregate the current frame The previous frame adjacent to it Aggregate the current frame The two adjacent frames , Aggregate the current frame The three adjacent frames , , Three sets of aggregation features were obtained; T2. Three spatiotemporal shifting modules are used to spatiotemporally shift the three sets of aggregated features respectively to obtain three sets of shifted aggregated features; T3. The three sets of shift aggregation features are fused to obtain the fused aggregation features; T4. Upsample the fused and aggregated features to obtain high-dimensional features, which are then input into the reconstruction module.
7. The video dehazing method based on a multi-scale spatiotemporal fusion network according to claim 1, characterized in that, For including the current frame The previous frame adjacent to it The first set of aggregation features, step T21 specifically includes the following steps: T211, Move the current frame Divided into and Two parts, the previous frame Divided into and Two parts; T212, Feature group Slicing along the channel dimension yields M feature slices, where the m-th feature slice is denoted as... ,in It is a slice index; T213, For each feature piece Spatially shift it in the x and y directions to obtain the translated feature patch. Group all features along the channel dimension Series connection yields spatial displacement characteristics. ; T214, will and Perform series connection, and analyze the series connection results and spatial displacement characteristics. After performing convolution, the result is added to the concatenated result. Then, the added result is convolved again and added to the concatenated result once more to obtain the first feature map. T215, to After convolution and The sum is then convolved and added again to obtain the second feature map. T216. Connect the first feature map and the second feature map in series to obtain the spatial shift feature of the first set of aggregated features; T217, will As the current frame A portion of the input window uses a multi-head self-attention mechanism.
8. A video dehazing system based on a multi-scale spatiotemporal fusion network, characterized in that: It includes an intelligent agent that performs the video dehazing method based on a multi-scale spatiotemporal fusion network as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-scale fusion defogging method based on stacked hourglass network
CN115330631A
Non-uniform remote sensing video image defogging method fusing space-time frequency information
CN117474801A