Video enhancement method and device in severe weather and computer equipment
Through the dual-path Mamba architecture and frequency domain adaptive enhancement video enhancement method, the problem of coupling of multiple degradation factors in bad weather is solved, and efficient and real-time video recovery effect is achieved. It is suitable for video surveillance, intelligent transportation and unmanned driving.
Patent Information
- Application Number
- CN202510442203.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively deal with the coupling problem of multiple degradation factors under severe weather conditions. The calculation complexity is high and the optical flow estimation depends on optical flow estimation, resulting in unsatisfactory video recovery effect and insufficient real-time performance.
The dual-path Mamba architecture is used to directly model spatiotemporal features, combined with frequency domain adaptive enhancement and contrast learning mechanisms, and enhance the frequency domain through adaptive mask generation and cross attention mechanisms to avoid optical flow estimation, and build positive and negative sample pairs for training.
It realizes efficient and high-quality video recovery, reduces computing complexity, improves robustness, can be processed in real time under severe weather conditions, and has good versatility.
Smart Images

Figure CN120339132A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video enhancement technologies, and in particular, to a method, apparatus, and computer device for enhancing videos in adverse weather conditions. Background Art
[0002] With the rapid development of technologies such as video surveillance, intelligent transportation, and autonomous driving, video data plays an increasingly important role in practical applications. However, under adverse weather conditions (such as rain, fog, snow, etc.), videos are usually simultaneously interfered by multiple degradation factors, resulting in blurred images, increased noise, and missing details, thus seriously affecting the accuracy of subsequent analysis and processing.
[0003] Video enhancement technology aims to improve video quality in order to provide clear and usable visual information under various environmental conditions. Existing video enhancement methods mainly focus on the following two aspects: (1) Methods based on image restoration technology. In the early stage, this method was mostly aimed at single-frame image restoration and was mainly implemented through technologies such as convolutional neural networks (CNNs) and generative adversarial networks (GANs). These methods usually rely on a large amount of labeled data for training and have poor performance when dealing with dynamic scenes. Recent studies such as DASR (Dynamic Adaptive SuperResolution), DuRRNet, etc. have demonstrated the feasibility of unsupervised degradation representation learning and can perform image restoration in the absence of labeled data. Image restoration methods based on All-In-One (AIO) (such as ADAIR, HAIR, etc.) strive to handle multiple degradations with a single model. However, when dealing with multiple degradations under complex weather conditions, these methods often have difficulty fully utilizing the temporal information in the video, resulting in unsatisfactory restoration effects. Works such as TransWeather have pioneered Transformer-based multi-weather restoration methods. Although they perform well in some scenarios, they still face problems such as high computational complexity and insufficient real-time performance.
[0004] (2) Methods for spatio-temporal information modeling of videos. For example, EDVR (Enhanced Deformable Video Restoration), BasicVSR, etc. mainly perform feature fusion through deformable alignment, which can capture temporal information to a certain extent, but have limited effectiveness in dealing with complex degradation interactions; RealBasicVSR processes unknown degradations through a pre-cleaning strategy, improving the restoration effect, but still relies on prior knowledge and is difficult to adapt to variable weather conditions; Transformer-based methods (such as ViWS-Net and Diff-TTA) further expand the unified video restoration model. Although they perform well in dealing with various weather conditions, their computational complexity is too high, restricting their use in real-time applications; the latest selective state space model (Mamba) shows great potential in sequence modeling and can effectively capture long-range spatio-temporal dependencies, but still has deficiencies in dealing with complex degradation interactions.
[0005] When existing technologies handle video restoration under adverse weather conditions, the following main drawbacks exist: 1) The problem of multiple degradation couplings Under adverse weather conditions, there are often multiple degradation factors acting simultaneously. The characteristics and influence mechanisms of different degradation factors (such as rain, fog, and snow) are different. Existing methods are difficult to effectively handle multiple degradation problems simultaneously.
[0006] 2) Insufficient research on frequency domain characteristics Different types of weather interferences show obvious differences in the frequency domain. Raindrops mainly manifest as high-frequency noise, while haze causes blurring in the low-frequency part. Current methods utilize less video frequency domain information.
[0007] 3) Balancing computational efficiency and effect Global modeling methods (such as Transformer) have high computational complexity and are difficult to meet the requirements of real-time processing. Lightweight methods often struggle to ensure the restoration effect. It is necessary to find a balance between computational efficiency and processing effect.
[0008] 4) Limitations of optical flow estimation Traditional optical flow-based spatio-temporal alignment methods are prone to failure under adverse weather conditions: weather such as heavy rain and thick fog will seriously affect the accuracy of optical flow estimation. The optical flow calculation itself brings additional computational overhead. Incorrect optical flow estimation will lead to feature misalignment and artifacts. There is an urgent need for an efficient spatio-temporal feature modeling method that does not rely on optical flow estimation. Summary of the Invention
[0009] Based on this, it is necessary to provide a method, device, and computer device for enhancing videos in bad weather to address the above technical problems. This method aims at multiple degradation problems existing simultaneously in videos under bad weather conditions and realizes the efficient restoration of video degradation by adaptively learning features in different frequency bands and constructing positive and negative sample pairs.
[0010] A method for enhancing videos in bad weather, the method comprising: Obtain a degraded video frame of bad weather.
[0011] Construct a video enhancement model; the video enhancement model includes a feature extraction module, a first dual-path Mamba network, a second dual-path Mamba network, a frequency domain enhancement module, a third dual-path Mamba network, and a restoration and reconstruction module connected in sequence; the feature extraction module is used to extract multi-scale features of the degraded video frame; the dual-path Mamba network is used to capture the short-term temporal dependencies of the input features and maintain the fine spatial structure by using two complementary spatio-temporal sequence Mamba blocks and Hilbert sequence Mamba blocks respectively to obtain potential features; the frequency domain enhancement module is used to perform adaptive enhancement and fusion in the frequency domain according to the degraded video frame and the potential features output by the second dual-path Mamba network by using an adaptive mask generation and cross-attention mechanism to obtain video enhancement features; the restoration and reconstruction module is used to process the potential features output by the third dual-path Mamba network by using upsampling convolution and residual connection to obtain a high-quality restored video frame.
[0012] Construct positive and negative sample pairs of the degraded video frame and the high-quality restored video frame, and train the video enhancement model through multi-degradation contrast learning to obtain a trained video enhancement model.
[0013] Process the degraded video frame of bad weather to be enhanced by using the trained video enhancement model to obtain a high-quality restored video frame of the degraded video frame of bad weather to be enhanced.
[0014] A device for enhancing videos in bad weather, the device comprising: A degraded video frame acquisition module, configured to obtain a degraded video frame of bad weather; Video enhancement model construction module, used to construct a video enhancement model; the video enhancement model includes a feature extraction module, a first dual-path Mamba network, a second dual-path Mamba network, a frequency domain enhancement module, a third dual-path Mamba network, and a restoration and reconstruction module connected in sequence; the feature extraction module is used to extract multi-scale features of low-quality video frames; the dual-path Mamba network is used to capture short-term temporal dependencies of input features and maintain fine spatial structures respectively by using two complementary spatio-temporal Mamba blocks and Hilbert sequence Mamba blocks, and obtain latent features; the frequency domain enhancement module is used to perform adaptive enhancement and fusion in the frequency domain according to the low-quality video frames and the latent features output by the second dual-path Mamba network by using an adaptive mask generation and cross-attention mechanism, and obtain video enhancement features; the restoration and reconstruction module is used to process the latent features output by the third dual-path Mamba network by using upsampling convolution and residual connection to obtain high-quality restored video frames; Video enhancement model training module, used to construct positive and negative sample pairs of low-quality video frames and high-quality restored video frames, and train the video enhancement model through multi-degradation contrast learning to obtain a trained video enhancement model; Severe weather video enhancement module, used to process the low-quality video frames of severe weather to be enhanced by using the trained video enhancement model to obtain high-quality restored video frames of the low-quality video frames of severe weather to be enhanced.
[0015] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the severe weather video enhancement method described in any one of the above are implemented.
[0016] The above severe weather video enhancement method, device and computer device. The method innovatively uses a dual-path Mamba architecture to directly model spatio-temporal features, completely avoiding the dependence on optical flow estimation; combined with frequency domain adaptive enhancement and contrast learning mechanisms, it realizes efficient and high-quality video restoration; this method not only reduces the computational complexity, but also improves the robustness under various severe weather conditions; this method has good generality, does not need to rely on prior weather type annotations, and can be widely applied to multiple fields such as video surveillance, intelligent transportation, and unmanned driving; due to its efficient computational performance, this method can meet the requirements of real-time processing. Description of the Drawings
[0017] Figure 1 It is a schematic flowchart of the severe weather video enhancement method in an embodiment; Figure 2 It is a schematic diagram of the overall architecture of the video enhancement model in another embodiment; Figure 3 It is a schematic diagram of the architecture of the dual-path Mamba modeling module in another embodiment; Figure 4 Schematic diagram of Hilbert curve scanning in another embodiment, where Figure 4 (a) is 3D Hilbert curve, Figure 4 and (b) is 3D Hilbert curve of recursive structure; Figure 5 Schematic diagram of the architecture of the frequency domain enhancement module in another embodiment; Figure 6 Color gamut comparison chart and frequency domain feature comparison chart of real image and degraded image on a rainy day in another embodiment, where Figure 6 (a) is the color gamut comparison chart of the degraded image, Figure 6 (b) is the color gamut comparison chart of the real image, Figure 6 (c) is the color gamut comparison chart of the difference between the degraded image and the real image, Figure 6 (d) is the frequency domain comparison chart of the degraded image, Figure 6 (e) is the frequency domain comparison chart of the real image, Figure 6 (f) is the frequency domain comparison chart of the difference between the degraded image and the real image; Figure 7 Color gamut comparison chart and frequency domain feature comparison chart of real image and degraded image on a foggy day in another embodiment, where Figure 7 (a) is the color gamut comparison chart of the degraded image, Figure 7 (b) is the color gamut comparison chart of the real image, Figure 7 (c) is the color gamut comparison chart of the difference between the degraded image and the real image, Figure 7 (d) is the frequency domain comparison chart of the degraded image, Figure 7 (e) is the frequency domain comparison chart of the real image, Figure 7 (f) is the frequency domain comparison chart of the difference between the degraded image and the real image; Figure 8 Color gamut comparison chart and frequency domain feature comparison chart of real image and degraded image on a snowy day in another embodiment, where Figure 8 (a) is the color gamut comparison chart of the degraded image, Figure 8 (b) is the color gamut comparison chart of the real image, Figure 8 (c) is the color gamut comparison chart of the difference between the degraded image and the real image, Figure 8 (d) is the frequency domain comparison chart of the degraded image, Figure 8 (e) is the frequency domain comparison chart of the real image, Figure 8 (f) is the frequency domain comparison chart of the difference between the degraded image and the real image; Figure 9 Comparison effect of the restored image of this method with existing algorithms and real images in one embodiment; Figure 10Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0018] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0019] In one embodiment, as Figure 1 shown, a method for enhancing videos in bad weather is provided, and the method includes the following steps: Step 100: Obtain a low-quality video frame of bad weather.
[0020] Specifically, the bad weather may be, but is not limited to, weather such as rain, heavy fog, snow, etc.
[0021] Step 102: Construct a video enhancement model; the video enhancement model includes a feature extraction module, a first dual-path Mamba network, a second dual-path Mamba network, a frequency domain enhancement module, a third dual-path Mamba network, and a restoration and reconstruction module connected in sequence; the feature extraction module is used to extract multi-scale features of the low-quality video frame; the dual-path Mamba network is used to capture short-term temporal dependencies and maintain fine spatial structures of the input features by using two complementary spatio-temporal sequence Mamba blocks and Hilbert sequence Mamba blocks respectively to obtain potential features; the frequency domain enhancement module is used to perform adaptive enhancement and fusion in the frequency domain according to the low-quality video frame and the potential features output by the second dual-path Mamba network by using an adaptive mask generation and cross-attention mechanism to obtain video enhancement features; the restoration and reconstruction module is used to process the potential features output by the third dual-path Mamba network by using upsampling convolution and residual connection to obtain a high-quality restored video frame.
[0022] Specifically, the video enhancement model is a bad weather video enhancement model based on the frequency domain enhancement of the dual-path Mamba architecture. The overall architecture of the video enhancement model is as Figure 2As shown in the figure. Among them, the feature extraction module: uses ConvNeXt as the backbone network to extract multi-scale features. The DPMM network backbone processes latent features through a three-stage hierarchical architecture, using a dual-path Mamba network with resolutions of 1 / 4, 1 / 8, and 1 / 4. Each level is a dual-path Mamba network, and the dual-path Mamba network includes a number of complementary spatio-temporal sequence Mamba blocks and Hilbert sequence Mamba blocks; among them, the dual-path Mamba modeling module is used to capture short-range temporal dependencies through SSMB, maintain fine spatial structure through HSMB, and parallelly model global dynamics and local details. The frequency domain enhancement module: adaptively enhances and fuses features in the frequency domain through adaptive mask generation and cross-attention mechanism. The restoration and reconstruction module: outputs high-quality restored video frames through upsampling convolution and residual connection.
[0023] Aiming at the problem that existing methods are difficult to effectively handle the problem of multiple degradation couplings simultaneously, the present invention adopts dual-path Mamba modeling, effectively captures short-range temporal dependencies through spatio-temporal sequence Mamba blocks (SSMB), and maintains fine spatial structure by using Hilbert curve scanning through Hilbert sequence Mamba blocks (HSMB). Dual-path parallel processing, combined with the linear time complexity characteristics of the Mamba model, realizes efficient spatio-temporal modeling and overcomes the limitations of existing methods.
[0024] Aiming at the different differential performances of different weather interferences in the frequency domain and the problem of insufficient research on frequency domain characteristics of existing methods, this application proposes a frequency adaptive fusion module (FAFB). Through adaptive mask generation, the model can focus on specific frequency bands; the cross-channel attention mechanism enhances the frequency feature representation; the frequency fusion module realizes the adaptive fusion of high and low frequency features. Thus, targeted degradation processing is achieved and the restoration quality is improved.
[0025] Through the innovative combination of the dual-path Mamba architecture, frequency adaptive enhancement, and general multi-degradation contrast learning, this video enhancement model has achieved remarkable technical effects in terms of restoration quality, computational efficiency, generality, practicality, etc., overcomes the deficiencies of the prior art, provides an efficient and reliable solution for video enhancement tasks in bad weather, and has broad application prospects.
[0026] Aiming at the problem of high computational complexity of existing global modeling methods, the dual-path Mamba architecture of this method has a linear time complexity. Under the 192-dimensional feature configuration, the number of parameters is 54.11M, and the FLOPs is 38.08G, which is significantly lower than 68.72G of ViWS-Net, and the average processing time is reduced by about 44.6%. By flexibly adjusting the feature dimension, the best balance can be achieved between computational resources and performance according to actual application requirements.
[0027] Users can flexibly select different model configurations according to the actual application scenarios, and make a trade-off between the restoration quality and computing resources.
[0028] Step 104: Construct positive and negative sample pairs of the degraded video frames and the high-quality restored video frames, and train the video enhancement model through multi-degradation contrast learning to obtain a trained video enhancement model.
[0029] Specifically, aiming at the problem that it is difficult to balance the computing efficiency and effect in the existing methods, this application introduces Universal Multi-Degradation Contrast Learning (UMCL). The sampling strategy based on the difference map ensures the effectiveness of the training samples; the selection of positive and negative samples with spatio-temporal constraints promotes feature learning; the hybrid sampling strategy enhances the generalization ability of the model. UMCL provides an unsupervised degradation mode discovery mechanism, which realizes the adaptive processing of multiple degradation modes without relying on prior annotations.
[0030] The positive and negative sample pairs are essentially triplets, which contain a positive sample pair and a negative sample pair. p a The patch on the restored video frame image is called the anchor patch. p p The block on the GT true clean image is the positive sample patch. p n The block on the degraded video frame is the negative sample patch. It is hoped that the feature map distance between the anchor patch and the positive sample patch is as close as possible, and the distance between the anchor patch and the negative sample patch is as far as possible.
[0031] Step 106: Process the degraded video frames of bad weather to be enhanced by using the trained video enhancement model to obtain high-quality restored video frames of the degraded video frames of bad weather to be enhanced.
[0032] In the above bad weather video enhancement method, the method directly models the spatio-temporal features by innovatively adopting the dual-path Mamba architecture, completely avoiding the dependence on optical flow estimation; combining frequency-domain adaptive enhancement and contrast learning mechanism, realizing efficient and high-quality video restoration; this method not only reduces the computational complexity, but also improves the robustness under various bad weather conditions; this method has good generality, does not need to rely on prior weather type annotations, and can be widely applied to multiple fields such as video surveillance, intelligent transportation, and driverless; due to the efficient computing performance, this method can meet the requirements of real-time processing.
[0033] In one embodiment, in the video enhancement model in step 102: the low-quality video frame is input into the feature extraction module to obtain multi-scale features; the multi-scale features are input into the first dual-path Mamba network to obtain the first latent feature; the first latent feature is downsampled and then input into the second dual-path Mamba network to obtain the second latent feature; the second latent feature and the low-quality video frame are input into the frequency domain enhancement module to obtain video enhancement features; the first latent feature and the upsampled video enhancement features are fused and then input into the third dual-path Mamba network to obtain the third latent feature, and the third latent feature is input into the restoration and reconstruction module to obtain the high-quality restored video frame.
[0034] In one embodiment, the first dual-path Mamba network in step 102 includes a number of dual-path Mamba modules; as Figure 3 shown, the dual-path Mamba modeling module includes two complementary spatio-temporal sequence Mamba blocks and a Hilbert sequence Mamba block; in the i-th dual-path Mamba modeling module: the module input feature is obtained; the spatio-temporal sequence Mamba block captures the short-range temporal dependencies of the module input feature through spatial-first rearrangement to obtain the output feature of the spatio-temporal sequence Mamba block; the expression of the output feature of the spatio-temporal sequence Mamba block is;
[0035]
[0036] where, is the output feature of the spatio-temporal sequence Mamba block, BiMamba represents the bidirectional Mamba operation, LN represents the layer normalization layer, and MLP is a multi-layer perceptron with depth convolution, is the feature input into the dual-path Mamba modeling module, is the matrix vector after dimensional deformation in the spatial-temporal first order, are respectively the i batch size, feature dimension, input frame length, height and width of the feature matrix of the i-th dual-path Mamba module; is the dimensional deformation of the feature matrix in the spatial-temporal first order.
[0037] The output feature of the spatio-temporal sequence Mamba block is input into the Hilbert sequence Mamba block to obtain the output feature of the i-th dual-path Mamba modeling module; the expression of the output feature of the i-th dual-path Mamba modeling module is: i th i th
[0038]
[0039]
[0040] Among them, is the output feature of the i th double-path Mamba modeling module, represents the Hilbert curve mapping determined by the module input feature dimension, represents the reconstruction operation of the Hilbert curve order (used to restore to the conventional feature matrix dimension according to the inverse Hilbert curve mapping transformation), is the output feature of the Hilbert sequence Mamba block (this feature is a feature matrix arranged in the order of Hilbert curve mapping), is the matrix vector after dimensional deformation by Hilbert curve mapping.
[0041] Specifically, the Hilbert sequence Mamba block uses Hilbert curve scanning to maintain fine spatial structure and local spatial relationships. Figure 4 shows a schematic diagram of Hilbert curve scanning, where Figure 4 (a) is a 3D Hilbert curve, Figure 4 (b) is a 3D Hilbert curve with a recursive structure.
[0042] In one embodiment, in the frequency domain enhancement module in step 102: perform instance normalization on the latent features output by the second double-path Mamba network to obtain normalized latent features; perform instance normalization on the low-quality video frame, perform a fast Fourier transform on the obtained normalized image to obtain frequency domain features; use the adaptive mask generation module to analyze the spectrum through adaptive average pooling and learnable thresholds for the frequency domain features to generate complementary binary masks centered on DC classification; multiply the complementary binary masks with the frequency domain features element by element and then perform inverse fast Fourier transforms respectively to obtain separated high-frequency components and low-frequency components; enhance the high-frequency components and the normalized latent features through a cross-channel attention model with learnable temperature scaling to obtain high-frequency enhanced features; enhance the low-frequency components and the normalized latent features through a cross-channel attention module to obtain low-frequency enhanced features; fuse the high-frequency enhanced features and the low-frequency enhanced features through a frequency fusion enhancement module to obtain fused features; enhance the fused features and the normalized latent features through a cross-channel attention module and then perform weighted element-by-element addition with the normalized latent features to obtain video enhancement features.
[0043] Specifically, the framework structure of the frequency domain enhancement module is as Figure 5 shown.
[0044] The Frequency-domain Adaptive Enhancement Module (FAFB) utilizes frequency-domain information to enhance the restoration quality. This module processes the degraded image I and the corresponding latent feature x from the DPMM backbone in sequence, and outputs enhanced features through frequency decomposition, feature enhancement, and fusion. Figure 6 For another embodiment, it is the color gamut comparison chart and frequency-domain feature comparison chart of a real image and a degraded image on a rainy day, where Figure 6 (a) is the color gamut comparison chart of the degraded image, Figure 6 (b) is the color gamut comparison chart of the real image, Figure 6 (c) is the color gamut comparison chart of the difference between the degraded image and the real image, Figure 6 (d) is the frequency-domain comparison chart of the degraded image, Figure 6 (e) is the frequency-domain comparison chart of the real image, Figure 6 (f) is the frequency-domain comparison chart of the difference between the degraded image and the real image. Figure 6 In (c), the left and right frames are respectively the color gamut comparison charts of the differences of Patch 1 and Patch 2 in the degraded image and the real image; Figure 6 (d) and Figure 6 In (e), the brighter points have higher frequencies, Figure 6 In (f), the brighter areas have greater differences.
[0045] Figure 7 For another embodiment, it is the color gamut comparison chart and frequency-domain feature comparison chart of a real image and a degraded image on a foggy day, where Figure 7 (a) is the color gamut comparison chart of the degraded image, Figure 7 (b) is the color gamut comparison chart of the real image, Figure 7 (c) is the color gamut comparison chart of the difference between the degraded image and the real image, Figure 7 (d) is the frequency-domain comparison chart of the degraded image, Figure 7 (e) is the frequency-domain comparison chart of the real image, Figure 7 (f) is the frequency-domain comparison chart of the difference between the degraded image and the real image. Figure 7 In (c), the left and right frames are respectively the color gamut comparison charts of the differences of Patch 1 and Patch 2 in the degraded image and the real image; Figure 7 (d) and Figure 7 In (e), the brighter points have higher frequencies, Figure 7 In (f), the brighter areas have greater differences.
[0046] Figure 8 For another embodiment, it is the color gamut comparison chart and frequency-domain feature comparison chart of a real image and a degraded image on a snowy day, where Figure 8 (a) is the color gamut comparison chart of the degraded image, Figure 8 (b) is the color gamut comparison chart of the real image, Figure 8 (c) is the color gamut comparison chart of the difference between the degraded image and the real image, Figure 8(d) is the frequency-domain comparison diagram of the degraded image, Figure 8 (e) is the frequency-domain comparison diagram of the real image, Figure 8 (f) is the frequency-domain comparison diagram of the difference between the degraded image and the real image. Figure 8 The left and right frames in (c) are the color gamut comparison diagrams of the differences between Patch 1 and Patch 2 in the degraded image and the real image respectively; Figure 8 (d) and Figure 8 The brighter points in (e) have higher frequencies, Figure 8 and the brighter areas in (f) have greater differences.
[0047] Figure 6 From (c) to Figure 6 (f), Figure 7 From (c) to Figure 7 (f), and Figure 8 From (c) to Figure 8 (f), the left frame corresponds to Patch 1 and the right corresponds to Patch 2.
[0048] First, the normalized image is converted to a frequency-domain representation through instance normalization and the fast Fourier transform (FFT): where F :
[0049] Here, represents the instance normalization layer, and
[0050] represents a small constant to prevent numerical instability.
[0051] Then, the Adaptive Mask Generation (AMG) module analyzes the spectrum through adaptive average pooling and learnable thresholding to generate a complementary binary mask centered on the DC component: and where clip ensures an appropriate frequency segmentation range.
[0052] The masked frequency components are transformed back to the spatial domain through the inverse fast Fourier transform (IFFT) to obtain the separated high-frequency and low-frequency components and :
[0053]
[0054] The separated frequency components are enhanced through cross-channel attention (CCA) with learnable temperature scaling and fused through the Frequency Fusion Enhancement (FFE) module. The FFE adopts spatial gating with max-average pooling Modulate high-frequency features using channel gating with an MLP Enhance low-frequency features for adaptive fusion:
[0055]
[0056]
[0057] Finally, output the features Obtained through a 3D upsampling convolutional layer and ReLU activation , and then concatenated with the 1 / 4 resolution features from the feature modeling path for final reconstruction through 3D convolution :
[0058]
[0059] Wherein, is a 3D convolution, is a feature concatenation operation.
[0060] In one embodiment, a frequency fusion enhancement module is used to modulate high-frequency features using spatial gating with max-average pooling, enhance low-frequency features using channel gating with an MLP, and then perform high-frequency and low-frequency adaptive fusion.
[0061] In one embodiment, step 104 includes: extracting an anchor patch set from a high-quality restored video frame, and the expression of the anchor patch set is:
[0062] Wherein, is the anchor patch set, represents the difference value of the anchor patch of, and respectively represent the mean and variance of the difference map; For each anchor patch, sample positive samples from the ground truth frames that satisfy spatial and temporal constraints to obtain a positive sample set; the expression of the positive sample set is:
[0063] Wherein, is the positive sample set of the anchor patch , is the positive sample patch, is the maximum sampling range of the positive sample, represents the normalized training progress, Control the exponential decay rate, is the time difference between frames; For each anchor patch, a hybrid sampling strategy is adopted to determine negative samples, and the negative sample set is obtained as:
[0064] where, is the anchor patch of the negative sample set, is the sampling probability, 、 are the basic interval distance and the adaptive interval distance of the negative samples respectively, is the negative sample patch of the difference value; Construct a contrastive loss; the expression of this contrastive loss is:
[0065] where, is the contrastive loss, 、 and respectively represent the anchor patch, positive sample and negative sample features extracted from the corresponding patches 、 and respectively, represents a small constant to prevent numerical instability, is the L1 distance between the anchor patch feature and the positive sample feature.
[0066] Specifically, the Universal Multi-Degradation Contrastive Learning module (UMCL) utilizes the inherent consistency of degradation patterns and the regional dynamics across spatial positions to effectively handle multiple weather degradations without explicit supervision. Its core is to construct informative triplets , and guide the sampling process through the difference map .
[0067] In one embodiment, the expression of the total loss function adopted in the video enhancement model training process is:
[0068] where, is the loss function, is the pixel-level reconstruction loss, is the perceptual loss is the contrastive loss, and are two hyperparameters.
[0069] In one embodiment, the feature extraction module is ConvNeXt.
[0070] In one embodiment, an existing algorithm and the algorithm of the present application are used to enhance degraded weather video frames. For the comparison effect of real images, the comparison effect between the restored images of the present method and the existing algorithm and real images is as Figure 9 shown.
[0071] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,
[0072] In one embodiment, a degraded weather video enhancement device is provided, including: a degraded video frame acquisition module, a video enhancement model construction module, a video enhancement model training module, and a degraded weather video enhancement module, where: The degraded video frame acquisition module is used to acquire degraded video frames of bad weather.
[0073] The video enhancement model construction module is used to construct a video enhancement model; the video enhancement model includes a feature extraction module, a first dual-path Mamba network, a second dual-path Mamba network, a frequency domain enhancement module, a third dual-path Mamba network, and a restoration and reconstruction module connected in sequence; the feature extraction module is used to extract multi-scale features of the degraded video frames; the dual-path Mamba network is used to capture short-term temporal dependencies and maintain fine spatial structures of the input features by using two complementary spatio-temporal sequence Mamba blocks and Hilbert sequence Mamba blocks respectively to obtain latent features; the frequency domain enhancement module is used to perform adaptive enhancement and fusion in the frequency domain according to the degraded video frames and the latent features output by the second dual-path Mamba network by using an adaptive mask generation and cross-attention mechanism to obtain video enhancement features; the restoration and reconstruction module is used to process the latent features output by the third dual-path Mamba network by using upsampling convolution and residual connection to obtain high-quality restored video frames.
[0074] The video enhancement model training module is used to construct positive and negative sample pairs of degraded video frames and high-quality restored video frames, and train the video enhancement model through multi-degradation contrast learning to obtain a trained video enhancement model.
[0075] A bad weather video enhancement module is used to process the degraded video frames of bad weather to be enhanced by using a trained video enhancement model, and obtain high-quality restored video frames of the degraded video frames of bad weather to be enhanced.
[0076] In one embodiment, in the video enhancement model in the video enhancement model construction module: the degraded video frames are input into the feature extraction module to obtain multi-scale features; the multi-scale features are input into the first dual-path Mamba network to obtain the first latent feature; the first latent feature is downsampled and then input into the second dual-path Mamba network to obtain the second latent feature; the second latent feature and the degraded video frames are input into the frequency domain enhancement module to obtain video enhancement features; the first latent feature and the upsampled video enhancement features are fused and then input into the third dual-path Mamba network to obtain the third latent feature, and the third latent feature is input into the restoration and reconstruction module to obtain high-quality restored video frames.
[0077] In one embodiment, the first dual-path Mamba network in the video enhancement model construction module includes a plurality of dual-path Mamba modeling modules; the dual-path Mamba modeling module includes two complementary spatio-temporal sequence Mamba blocks and Hilbert sequence Mamba blocks; in the i th dual-path Mamba modeling module: obtain the module input features; use the spatio-temporal sequence Mamba block to capture the short-term temporal dependencies of the module input features through spatial priority rearrangement, and obtain the output features of the spatio-temporal sequence Mamba block as shown in the expression of the output features of the spatio-temporal sequence Mamba block; input the output features of the spatio-temporal sequence Mamba block into the Hilbert sequence Mamba block to obtain the output features of the i th dual-path Mamba modeling module as shown in the expression of the output features of the i th dual-path Mamba modeling module.
[0078] In one embodiment, in the frequency domain enhancement module of the video enhancement model construction module: instance normalization is performed on the latent features output by the second dual-path Mamba network to obtain normalized latent features; instance normalization is performed on the low-quality video frame, and the obtained normalized image is subjected to fast Fourier transform to obtain frequency domain features; the frequency domain features are analyzed by the adaptive mask generation module through adaptive average pooling and learnable threshold to generate complementary binary masks centered on DC classification; the complementary binary masks are multiplied element-wise with the frequency domain features respectively and then subjected to inverse fast Fourier transform respectively to obtain separated high-frequency components and low-frequency components; the high-frequency components and the normalized latent features are enhanced through a cross-channel attention model with learnable temperature scaling to obtain high-frequency enhanced features; the low-frequency components and the normalized latent features are enhanced through a cross-channel attention module to obtain low-frequency enhanced features; the high-frequency enhanced features and the low-frequency enhanced features are fused through a frequency fusion enhancement module to obtain fused features; the fused features and the normalized latent features are enhanced through a cross-channel attention module and then weighted element-wise added to the normalized latent features to obtain video enhancement features.
[0079] In one embodiment, the frequency fusion enhancement module is used to spatially gate and modulate high-frequency features with max-average pooling, channel-gate enhance low-frequency features with MLP, and then perform high-frequency and low-frequency adaptive fusion.
[0080] In one embodiment, the video enhancement model training module is further used to extract an anchor patch set as shown in the expression of the anchor patch set above from the high-quality restored video frames; for each anchor patch, positive samples are sampled from the ground truth frames that satisfy spatial and temporal constraints to obtain a positive sample set as shown in the expression of the positive sample set above; for each anchor patch, a negative sample is determined by using a hybrid strategy to obtain a negative sample set as described in the expression of the negative sample set above; a contrast loss as shown in the representation of the contrast loss above is constructed.
[0081] In one embodiment, the total loss function as shown in the expression of the total loss function above is adopted during the video enhancement model training process in the video enhancement model training module.
[0082] In one embodiment, the feature extraction module in the video enhancement model construction module is ConvNeXt.
[0083] For the specific limitations of the adverse weather video enhancement device, reference may be made to the limitations of the adverse weather video enhancement method in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned adverse weather video enhancement device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0084] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an adverse weather video enhancement method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0085] Those skilled in the art can understand that Figure 10 the structure shown in
[0086] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0087] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0088] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for enhancing videos in bad weather, characterized in that, The method includes: Obtaining low-quality video frames of bad weather; Constructing a video enhancement model; the video enhancement model includes a feature extraction module, a first dual-path Mamba network, a second dual-path Mamba network, a frequency domain enhancement module, a third dual-path Mamba network, and a restoration and reconstruction module connected in sequence; the feature extraction module is used to extract multi-scale features of the low-quality video frames; the dual-path Mamba network is used to capture short-term temporal dependencies of the input features and maintain fine spatial structures by using two complementary spatio-temporal sequence Mamba blocks and Hilbert sequence Mamba blocks respectively, to obtain latent features; the frequency domain enhancement module is used to perform adaptive enhancement and fusion in the frequency domain according to the low-quality video frames and the latent features output by the second dual-path Mamba network by using an adaptive mask generation and cross-attention mechanism, to obtain video enhancement features; the restoration and reconstruction module is used to process the latent features output by the third dual-path Mamba network by using upsampling convolution and residual connection, to obtain high-quality restored video frames; Constructing positive and negative sample pairs of low-quality video frames and high-quality restored video frames, and training the video enhancement model through multi-degradation contrast learning to obtain a trained video enhancement model; Using the trained video enhancement model to process the low-quality video frames of bad weather to be enhanced, to obtain high-quality restored video frames of the low-quality video frames of bad weather to be enhanced.
2. The bad weather video enhancement method according to claim 1, characterized in that, In the video enhancement model: Inputting the low-quality video frames into the feature extraction module to obtain multi-scale features; Inputting the multi-scale features into the first dual-path Mamba network to obtain first latent features; Downsampling the first latent features and inputting them into the second dual-path Mamba network to obtain second latent features; Inputting the second latent features and the low-quality video frames into the frequency domain enhancement module to obtain video enhancement features; Fusing the first latent features and the upsampled video enhancement features and inputting them into the third dual-path Mamba network to obtain third latent features, Inputting the third latent features into the restoration and reconstruction module to obtain high-quality restored video frames.
3. The bad weather video enhancement method according to claim 1, wherein The first dual-path Mamba network includes a number of dual-path Mamba modeling modules; the dual-path Mamba modeling module includes two complementary spatio-temporal sequence Mamba blocks and Hilbert sequence Mamba blocks; in the i-th dual-path Mamba modeling module: Obtaining the module input features; Using the spatio-temporal sequence Mamba block to capture the short-term temporal dependencies of the module input features through space-first rearrangement, and obtaining the output features of the spatio-temporal sequence Mamba block as; Among them, is the output feature of the spatio-temporal sequence Mamba block, BiMamba represents the bidirectional Mamba operation, LN represents the layer normalization layer, and MLP is a multi-layer perceptron with depth convolution. is the feature input to the dual-path Mamba modeling module. is the matrix vector after dimensional transformation in the spatio-temporal priority order. are respectively the batch size, feature dimension, input frame length, height, and width of the feature matrix of the i-th dual-path Mamba module. is the dimensional transformation of the feature matrix in the spatio-temporal priority order. The output features of the spatio-temporal series Mamba block are input into the Hilbert series Mamba block, and the output features of the i th double-path Mamba modeling module are: Among them, is the output feature of the i th dual-path Mamba modeling module, represents the Hilbert curve mapping determined by the module input feature dimension, represents the reconstruction operation of the Hilbert curve order, is the output feature of the Hilbert sequence Mamba block, is the matrix vector after dimensional transformation by the Hilbert curve mapping.
4. The bad weather video enhancement method according to claim 1, characterized in that In the frequency domain enhancement module: Performing instance normalization on the latent features output by the second dual-path Mamba network to obtain normalized latent features; Performing instance normalization on the low-quality video frames, and performing fast Fourier transform on the obtained normalized images to obtain frequency domain features; Using the frequency domain features to generate complementary binary masks centered on DC classification through adaptive average pooling and learnable threshold analysis by using an adaptive mask generation module; The complementary binary masks are multiplied element-wise with the frequency-domain features respectively, and then inverse fast Fourier transforms are performed respectively to obtain separated high-frequency components and low-frequency components; The high-frequency components and the normalized latent features are enhanced through a cross-channel attention model with learnable temperature scaling to obtain high-frequency enhanced features; The low-frequency components and the normalized latent features are enhanced through the cross-channel attention module to obtain low-frequency enhanced features; The high-frequency enhanced features and the low-frequency enhanced features are fused through a frequency fusion enhancement module to obtain fused features; The fused features and the normalized latent features are enhanced using the cross-channel attention module and then weighted element-wise added to the normalized latent features to obtain video enhanced features.
5. The bad weather video enhancement method according to claim 4, characterized in that The frequency fusion enhancement module is used to spatially gate and modulate high-frequency features with max-average pooling, channel gate and enhance low-frequency features with MLP, and then perform high-frequency and low-frequency adaptive fusion.
6. The bad weather video enhancement method according to claim 1, wherein Positive and negative sample pairs of the low-quality video frames and the high-quality restored video frames are constructed, and the video enhancement model is trained through multi-degradation contrast learning to obtain a trained video enhancement model, including: Extract an anchor patch set from the high-quality restored video frames; The anchor patch set is: Among them, is the anchor patch set, represents the anchor patch difference value, and represent the mean and variance of the difference map respectively; For each anchor patch, positive samples are sampled from the ground truth frames that meet the spatial and temporal constraints to obtain a positive sample set as: Among them, is the positive sample set of the anchor patch , is the positive sample patch is the maximum sampling range of the positive sample represents the normalized training progress controls the exponential decay rate is the time difference between frames; For each anchor patch, a negative sample is determined using a hybrid sampling strategy to obtain a negative sample set as: Among them, is the negative sample set of the anchor patch , is the sampling probability, , are the basic interval distance and the adaptive interval distance of the negative sample respectively, is the negative sample patch is the difference value; Construct the contrast loss as: Among them, is the contrast loss, , and respectively represent the anchor patch, positive sample, and negative sample features extracted from the corresponding patches , and ; represents a constant to prevent numerical instability, is the L1 distance between the anchor patch feature and the positive sample feature.
7. The bad weather video enhancement method according to claim 1, wherein The total loss function used during the training process of the video enhancement model is: Among them, is the loss function, is the pixel-level reconstruction loss, is the perceptual loss is the contrastive loss, and are two hyperparameters.
8. The bad weather video enhancement method according to claim 1, characterized in that The feature extraction module is ConvNeXt.
9. A video enhancement device for bad weather, characterized in that, The device includes: A low-quality video frame acquisition module, configured to acquire low-quality video frames of bad weather; A video enhancement model construction module, configured to construct a video enhancement model; The video enhancement model includes a feature extraction module, a first dual-path Mamba network, a second dual-path Mamba network, a frequency-domain enhancement module, a third dual-path Mamba network, and a restoration and reconstruction module connected in sequence; The feature extraction module is used to extract multi-scale features of the low-quality video frames; The dual-path Mamba network is used to capture short-range temporal dependencies and maintain fine spatial structures of the input features respectively using two complementary spatio-temporal sequence Mamba blocks and Hilbert sequence Mamba blocks to obtain latent features; The frequency-domain enhancement module is used to perform adaptive enhancement and fusion in the frequency domain according to the low-quality video frames and the latent features output by the second dual-path Mamba network using an adaptive mask generation and cross-attention mechanism to obtain video enhanced features; The restoration and reconstruction module is used to process the latent features output by the third dual-path Mamba network using upsampling convolution and residual connections to obtain high-quality restored video frames; A video enhancement model training module, configured to construct positive and negative sample pairs of low-quality video frames and high-quality restored video frames, and train the video enhancement model through multi-degradation contrast learning to obtain a trained video enhancement model; A bad weather video enhancement module is used to process a bad weather degraded video frame to be enhanced by using a trained video enhancement model, so as to obtain a high-quality restored video frame of the bad weather degraded video frame to be enhanced.
10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the bad weather video enhancement method according to any one of claims 1 to 8 are implemented.