Real-time Video Super-Resolution Reconstruction Method and System with Multimodal Fusion
Through adaptive Kalman filter prediction and multimodal fusion method of extracting visual and text features of CLIP model, the existing video super-resolution methods in complex texture recovery, computational complexity and inter-frame consistency problems are solved, and high-quality and real-time video super-resolution reconstruction is achieved.
Patent Information
- Application Number
- CN202510383211.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing video super-resolution methods have limitations in dealing with complex textures and details recovery, and ignore the auxiliary role of other modal information such as text, and have high computational complexity, making it difficult to achieve high-quality real-time processing. Especially for high-resolution videos, the inter-frame consistency problem has not been fully solved.
The lightweight three-stream structure based on adaptive Kalman filter prediction is adopted, combined with the CLIP model to extract visual and text features, and through multimodal fusion, lightweight residual network and adaptive motion compensation technology, high-quality and real-time video super-resolution reconstruction is achieved.
It improves the quality and detailed performance of super-resolution reconstruction, significantly reduces the computing complexity, realizes the real-time processing capability of 4K video, and has a frame rate of up to 30fps, solving the problem of inter-frame consistency and reducing jitter and artifacts in the time dimension.
Smart Images

Figure CN119887530B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and deep learning, and particularly relates to a real-time video super-resolution reconstruction method and system based on multimodal fusion. Background Art
[0002] With the rapid development of high-definition display technology, the resolution requirements for video content are increasing day by day. However, due to transmission bandwidth limitations, storage capacity limitations, and physical constraints of acquisition devices, a large amount of video content is still stored and transmitted at a low resolution. Video super-resolution technology aims to reconstruct high-resolution videos from low-resolution videos to improve visual quality and detail performance.
[0003] Traditional video super-resolution methods mainly rely on interpolation algorithms or sample-based learning methods, which have significant limitations in dealing with complex textures and detail recovery. In recent years, deep learning methods have made remarkable progress in the field of video super-resolution, but the existing methods still face the following problems: First, most methods only rely on visual information for super-resolution reconstruction, ignoring the auxiliary role of other modal information such as text; Second, achieving real-time processing while ensuring super-resolution quality is still a challenge, especially for high-resolution videos such as 4K; Finally, the problem of inter-frame consistency has not been fully solved in the existing methods, resulting in jitter and incoherence in the reconstructed video in the time dimension.
[0004] In addition, existing deep learning methods usually have high computational complexity and are difficult to achieve real-time processing on ordinary hardware, especially for high-resolution videos. Although some lightweight network designs have been proposed to reduce computational complexity, they often lead to a decrease in reconstruction quality. Therefore, how to achieve computationally efficient real-time processing while ensuring high-quality super-resolution reconstruction is still a key problem to be solved urgently. Summary of the Invention
[0005] The purpose of the present invention is to provide a real-time video super-resolution reconstruction method and system based on multimodal fusion. The method adopts a lightweight three-stream structure based on adaptive Kalman filter prediction, combines visual and text features extracted by the CLIP model, and realizes high-quality and real-time video super-resolution reconstruction through technologies such as multimodal fusion, lightweight residual network, and adaptive motion compensation.
[0006] The present invention proposes a real-time video super-resolution reconstruction method based on multimodal fusion, including the following steps:
[0007] Obtain a low-resolution video sequence;
[0008] Extract visual features and text features using the CLIP model, and perform bimodal feature fusion to generate guidance features;
[0009] Align the guidance features and the low-resolution video through a multi-modal fusion module;
[0010] Extract high-quality features using a lightweight residual module;
[0011] Fuse multi-frame features through inter-frame information flow propagation and perform motion compensation using adaptive Kalman filtering;
[0012] Combine the high-quality features with the information after feature fusion to reconstruct a high-definition image.
[0013] Preferably, the steps for the CLIP model to extract visual features and text features specifically include:
[0014] Input the low-resolution video sequence into the ViT encoder and the ResNet network to extract visual features;
[0015] Input the context description in the low-resolution video frame into the Text Encoder to extract text features;
[0016] Calculate the similarity score between the visual features and the text features through a cross-modal similarity encoder;
[0017] Generate guidance features from the visual features and the text features through a gated fusion module.
[0018] Preferably, the multi-modal fusion module dynamically integrates the guidance features and the low-resolution video features through a gating mechanism, and the specific steps are as follows:
[0019] Input the guidance features and the low-resolution video features into two parallel convolutional layers and ReLU layers;
[0020] Add the outputs of the two parallel branches and then input them into another convolutional layer and ReLU layer;
[0021] Multiply the output result by the guidance gate in the guidance module to control the fusion degree and generate aligned features.
[0022] Preferably, the residual path of the lightweight residual module adopts a depthwise separable convolutional layer to reduce the computational complexity, and the specific steps are as follows:
[0023] Convert the video frame from a size of to a tensor with 4 channels ;
[0024] Convert the tensor to a tensor with a shape of , where represents the number of channels; reduce the number of channels to 2 through a convolutional layer , and duplicate two copies;
[0025] Add three tensors with the same shape and pass them through the ReLU activation function;
[0026] Reshape the tensor group from to and rearrange it into
[0027] through a convolutional layer, and double the number of tensor channels to obtain a feature representation with 2 channels.
[0028] Preferably, the specific steps of the adaptive Kalman filter for motion compensation include:
[0029] Based on the deep features extracted by the feature learning network, calculate the feature correlation using a 3×3 convolutional neural network;
[0030] Perform motion estimation on the low-resolution video sequence through the adaptive model parameters to obtain the motion estimation value at time t+1;
[0031] Send the motion estimation value into the fusion module to obtain the Kalman motion estimation value;
[0032] Correct the super-resolved low-resolution blocks through the motion compensation module to improve the inter-frame consistency.
[0033] Preferably, the adaptive Kalman filter includes a multi-branch multi-layer perceptron and a fusion module, and its Kalman motion estimation value is obtained by combining the feature maps of different branches; the feature correlation module and the fusion module use layer normalization to process the output of each layer of the network.
[0034] Preferably, the lightweight residual module includes a plurality of densely connected residual blocks, each residual block includes layer normalization, a convolutional layer, and a Leaky-ReLU activation layer; deformable convolutions are set in the residual blocks to extract local motion information of the video frame; dense connections are added between adjacent residual blocks to connect the input of each residual block to the outputs of all previous residual blocks.
[0035] Preferably, the method is adversarially trained through a discriminator, where the discriminator loss function consists of an adversarial loss function, a multi-scale loss function, a cycle consistency loss function, a guidance loss function, and a total variation regularization loss function; the multi-scale loss function is used to evaluate the reconstruction quality at different scales; the cycle consistency loss function is used to ensure feature consistency.
[0036] Preferably, the discriminator is initialized with a pre-trained model extracted from the VGG network; the perceptual difference loss is calculated by extracting multi-layer features of the input image; the feature learning network uses the output features of the last convolutional layer of the ImageNet pre-trained model.
[0037] A real-time video super-resolution reconstruction system for multimodal fusion implementing the method, comprising:
[0038] A feature extraction module, configured to extract visual features and text features of a low-resolution video using a CLIP model, and perform bimodal feature fusion to generate guiding features;
[0039] A feature alignment module, configured to perform feature alignment on the guiding features and the low-resolution video through a multimodal fusion module;
[0040] A feature optimization module, configured to extract high-quality features using a lightweight residual module;
[0041] A motion compensation module, configured to fuse multi-frame features through inter-frame information flow propagation and perform motion compensation using adaptive Kalman filtering;
[0042] An image reconstruction module, configured to combine the high-quality features with the information after feature fusion to reconstruct a high-definition image;
[0043] A network training module, configured to train and optimize the system through an adversarial generative network.
[0044] The present invention has the following beneficial effects:
[0045] 1. By introducing a CLIP model to extract visual and text features, multimodal information fusion is achieved, enhancing semantic understanding ability and improving the quality and detail performance of super-resolution reconstruction. Compared with the method of single visual features, the PSNR and SSIM metrics are improved by about 2dB and 0.05 respectively.
[0046] 2. By adopting a lightweight three-stream structure and channel-separated convolution technology, the computational complexity is significantly reduced, the computational amount is reduced by more than 70%, the number of parameters is reduced by more than 50%, and the real-time processing ability of 4K video is achieved, with a frame rate of up to 30fps.
[0047] 3. Innovatively introducing adaptive Kalman filtering for motion compensation, combining deep learning and classical signal processing technologies, effectively solves the inter-frame consistency problem, reduces jitter and artifacts in the time dimension, and improves video smoothness.
[0048] 4. A multi-component loss function system is designed to guide network learning from multiple perspectives, including adversarial loss, multi-scale loss, cycle consistency loss, etc., comprehensively improving the visual quality of video super-resolution.
[0049] 5. The overall solution has high adaptability, can dynamically adjust processing parameters for different scenarios, and has good performance in various application scenarios such as medical imaging, surveillance videos, and live content. Description of the Drawings
[0050] Figure 1 It is the overall flowchart of the multi-modal fusion real-time video super-resolution reconstruction method of the present invention.
[0051] Figure 2 It is the structural schematic diagram of the CLIP model extracting visual features and text features in the present invention.
[0052] Figure 3 It is the structural schematic diagram of the multi-modal fusion module in the present invention.
[0053] Figure 4 It is the structural schematic diagram of the channel separation convolutional layer of the lightweight residual module in the present invention.
[0054] Figure 5 It is the process schematic diagram of motion compensation by adaptive Kalman filtering in the present invention.
[0055] Figure 6 It is the structural schematic diagram of the dense connection of the lightweight residual module in the present invention.
[0056] Figure 7 It is the structural block diagram of the multi-modal fusion real-time video super-resolution reconstruction system of the present invention.
[0057] Figure 8 It is the detailed structural schematic diagram of the cross-modal attention mechanism of the CLIP model in the present invention.
[0058] Figure 9 It is the computational flowchart of the state update and prediction of adaptive Kalman filtering in the present invention. Specific Embodiments
[0059] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0060] Refer to Figure 1 , the present invention provides a multi-modal fusion real-time video super-resolution reconstruction method, which adopts a lightweight three-stream structure based on adaptive Kalman filtering prediction, and includes the following steps:
[0061] Step 1: Obtain a low-resolution video sequence. Preferably, this step can directly obtain a low-resolution video through a video acquisition device, or extract an existing low-resolution video sequence from a video database. In this embodiment, the resolution of the low-resolution video sequence can be 720p (1280×720 pixels) or lower.
[0062] The lightweight three-stream structure based on adaptive Kalman filter prediction in the present invention refers to three main feature paths that cooperate with each other but are processed in parallel in the system. The first path is the visual feature stream, which contains two complementary branches: the ViT encoder branch is responsible for capturing the global semantic features and long-range dependencies of video frames, while the ResNet network branch focuses on extracting local texture, edges and other detailed features. These two branches work together, taking low-resolution video frames as input and outputting a rich visual feature representation after fusion.
[0063] The second path is the text feature stream, which is composed of the Text Encoder of the CLIP model. It specifically processes the text descriptions related to the video content and extracts deep semantic features from them. This process takes the text descriptions related to the video as input and outputs a high-dimensional text feature representation to provide semantic guidance for video super-resolution.
[0064] The third path is the temporal consistency stream, which integrates an adaptive Kalman filter and a motion compensation module, specifically processes the temporal relationships and motion information between frames, and ensures the consistency of the super-resolution results in the temporal dimension. This process receives the feature representations of consecutive video frames and outputs temporally consistent features after motion compensation.
[0065] The cooperation mode of these three paths is extremely ingenious: the visual feature stream and the text feature stream first process the input data independently, and then perform deep feature fusion through a cross-modal similarity encoder and a gated fusion module to generate guiding features rich in semantic information. These guiding features are precisely integrated with the original low-resolution video features in the feature alignment module. At the same time, the temporal consistency stream receives the information after feature alignment, combines the historical frame information, and optimizes the consistency in the temporal dimension through the adaptive Kalman filter algorithm. Finally, the processing results of the three paths are unified and fused in the image reconstruction module to generate a super-resolution video with both high spatial details and temporal smoothness.
[0066] This three-stream architecture design enables the system to utilize spatial information, semantic information and temporal information simultaneously, realizing multi-dimensional and multi-modal comprehensive feature extraction and fusion. It is the core technical framework for the present invention to achieve high-quality real-time video super-resolution, significantly improving the processing efficiency while ensuring the video quality.
[0067] Step 2: Use the CLIP (Contrastive Language-Image Pre-training) model to extract visual features and text features, and perform bimodal feature fusion to generate guiding features.
[0068] In the present invention, various flexible methods are adopted for text information acquisition. Firstly, the system can directly utilize the metadata information attached to video files, such as video titles, description tags, etc., as text inputs. These metadata are usually provided by the creator during video generation or upload and can accurately reflect the semantic information of the video content. Secondly, in specific application scenarios, pre-manually annotated scene description texts can be used. These descriptions are divided according to the timeline and precisely correspond to the video content. In addition, the system also supports automatically generating text descriptions through a pre-trained scene description model to achieve real-time applications without manual intervention.
[0069] To solve the problem of time alignment between text and video frames, the present invention designs a multi-level matching strategy. For text descriptions with timestamps (such as subtitles or scene annotations), the system precisely matches the text with the video frames within the corresponding time range according to the timestamps. When processing continuous video streams, a sliding time window (such as 5 seconds) mechanism is adopted to assign the same text description to all frames within the window, and the window updates the description as the video progresses. The system can also intelligently detect scene switching points or key frames in the video and only update the text description at these key points. The ordinary frames between two key frames share the same text description, improving the processing efficiency.
[0070] To meet the real-time processing requirements, the present invention also implements a text caching mechanism and predictive text generation technology to prepare the next text description that may be needed in advance, ensuring that text features can be fused with video frame features in a timely manner and minimizing the processing delay to the greatest extent.
[0071] Referring to Figure 2 and Figure 8 , this step specifically includes: inputting the low-resolution video sequence into the ViT (Vision Transformer) encoder and ResNet network to extract visual features; inputting the context description in the low-resolution video frames into the Text Encoder to extract text features; calculating the similarity score between the visual features and the text features through a cross-modal similarity encoder; generating guiding features by the visual features and the text features through a gated fusion module.
[0072] In this embodiment, the dimension of the visual features is , where is the image batch size, is the image pixel size, is the number of feature channels. Preferably, can be set to 256 to balance the feature expression ability and computational efficiency. The dimension of the text features is , and after dimensionality reduction processing, it matches the number of channels of the visual features.
[0073] The tensor dimension transformation in the lightweight residual module is one of the technical essences of the present invention, and the detailed dimension change process is further elaborated below. The input dimension of the initial video frame is the batch size B multiplied by the height H, width W, and RGB three channels, that is, in the [B, H, W, 3] format. In the channel expansion stage, the system expands the number of channels from 3 to 4c (c = 64) through a carefully designed convolution operation, generating a feature tensor with the dimension of [B, H, W, 4c].
[0074] The subsequent shape reorganization stage is a key link. The system rearranges the above tensor into the shape of [B, H, W, 4c, Choi], where Choi = 4 is the channel selection factor. This transformation is achieved through the reshape operation, which does not change the total amount of data but only adjusts the organizational structure, preparing for the subsequent channel-separated convolution. In the channel reduction stage, through an innovative grouped convolution operation, the number of channels is effectively reduced from 4c×Choi to 2c, and two identical copies are generated, forming three feature tensors with the dimension of [B, H, W, 2c].
[0075] In the feature fusion stage, these three tensors are element-wise added and then processed through the ReLU activation function, and the dimension remains unchanged. Subsequently, in the downsampling and shape transformation link, the system performs spatial downsampling on the activated feature map, changing the dimension from [B, H, W, 2c] to [B, H / 2, W / 2, 2c], and at the same time completing the transformation of the spatial resolution. In the channel rearrangement stage, the features are rearranged into the form of [B, H, W, c] through precise convolution operations, restoring the original spatial resolution but reducing the number of channels to c. The final channel expansion step increases the number of channels from c to 2c, generating the final output tensor [B, H, W, 2c].
[0076] This precise dimension transformation and channel operation design significantly reduce the computational complexity while maintaining a sufficiently strong feature expression ability, which is the core technology for realizing lightweight and efficient processing.
[0077] The present invention innovatively designs a cross-modal feature extraction and fusion algorithm for the CLIP model, and the detailed process is as follows:
[0078] 1. Visual feature extraction: For the input low-resolution video frame , first extract visual features through the ViT encoder and ResNet dual-path network:
[0079] ,
[0080] Among them, Represents a visual encoder combination. For the ResNet path, a five-layer structure with 3×3 convolutional kernels is adopted, and batch normalization and ReLU activation functions are added between each layer; for the ViT path, the image is segmented into patches, where is the optimal segmentation size. Each patch is input into the Transformer encoder after linear projection. The Transformer contains 12 self-attention layers, and each layer contains 8 attention heads.
[0081] 2. Text feature extraction: For the text description of the video content , semantic features are extracted through the Text Encoder
[0082] ,
[0083] where represents the text encoder, is the text feature, is the text feature dimension. The TextEncoder is based on the Transformer architecture and contains 12 self-attention layers.
[0084] 3. Cross-modal attention calculation:
[0085] To effectively fuse visual and text features, the present invention designs a cross-modal attention mechanism, as shown in Figure 8 . First, calculate the similarity score between the visual feature and the text feature:
[0086] ,
[0087] where is the feature dimension (usually set to , is the transposed matrix of the text feature, is the extracted visual feature, and are the height and width of the feature map respectively, is the number of channels, is the transposed matrix of the text feature, softmax is the normalization function to ensure that the similarity value range is between [0,1]. Then, construct the cross-modal attention map :
[0088] ,
[0089] where and are learnable parameter matrices for feature mapping, is the Sigmoid activation function. The attention map The value range is [0, 1], indicating the response degree of each spatial position in the visual feature to the text feature.
[0090] 4. Gated feature fusion: Based on the calculated attention map , fuse visual and text features through a gating mechanism:
[0091] ,
[0092] Among them, is the fused guidance feature, and are learnable projection matrices for feature channel alignment and fusion, is the text feature is the feature map after spatial expansion of represents element-wise multiplication (Hadamard product).
[0093] Preferably, and are both convolutional layers for feature channel alignment and fusion. In practical applications, when the attention value is close to 1, the model relies more on visual features; when is close to 0, the model relies more on text features. Experiments show that in regions containing complex textures, values are usually higher, while in regions containing clear semantic information, values are lower.
[0094] Step 3: Align the guidance feature and the low-resolution video through a multimodal fusion module.
[0095] Referring to Figure 3 , the multimodal fusion module dynamically integrates the guidance feature and the low-resolution video feature through a gating mechanism. The specific steps are as follows: Input the guidance feature and the low-resolution video feature into two parallel convolutional layers and ReLU layers; Add the outputs of the two parallel branches and then input them into another convolutional layer and ReLU layer; Multiply the output result by the guidance gate in the guidance module to control the fusion degree and generate the aligned feature.
[0096] The mathematical expression of the gating mechanism is:
[0097] ,
[0098] Among them, and represent the guidance feature and the low-resolution video feature respectively, is the attention weight matrix, Denotes element-wise multiplication. This mechanism allows the system to dynamically adjust the importance of different modality features according to the content. Preferably, when the video content is complex, the system tends to assign higher weights to visual features; when the video content is simple but contains rich semantic information, the system increases the weight of text features. In a specific implementation, the guidance gate is calculated through the following steps:
[0099] ,
[0100] wherein, denotes the feature concatenation operation, denotes a 3×3 convolutional layer, is the Sigmoid activation function. In this way, the guidance gate can adaptively learn the importance weights between the guidance features and the low-resolution video features.
[0101] Step 4: Use a lightweight residual module to extract high-quality features.
[0102] Referring to Figure 4 , the residual path of the lightweight residual module uses a channel-separated convolutional layer to reduce the computational complexity. The specific steps are as follows: Convert the video frame from a size of to a tensor with 4 channels ; Convert the said tensor to a tensor with a shape of , where represents the number of channels; Reduce the number of channels to 2 through a convolutional layer, and duplicate two copies; Add three tensors with the same shape, and pass through the ReLU activation function; Change the tensor group shape from to , and rearrange it to through a convolutional layer; Expand the tensor channels twice to obtain a feature representation with 2 channels .
[0103] Preferably, is set to 64 and Choi is set to 4. This configuration can achieve the best computational efficiency on most hardware platforms. The channel-separated convolutional design reduces the computational complexity of the module from of the traditional convolution to , achieving a significant improvement in computational efficiency. The detailed calculation process of the channel-separated convolution can be expressed as:
[0104] ,
[0105] wherein, is the input feature, which is divided into groups ( is the optimal value), is the Group features is a 3×3 convolution operation for the group of features, and is the output feature. This grouped convolution strategy significantly reduces the number of parameters and computational complexity while maintaining the feature expression ability.
[0106] Step 5: Propagate and fuse multi-frame features through inter-frame information flow, and perform motion compensation using adaptive Kalman filtering.
[0107] Referring to Figure 5 and Figure 9 , the specific steps of performing motion compensation using adaptive Kalman filtering include: calculating feature correlations using a 3×3 convolutional neural network based on the deep features extracted by the feature learning network; performing motion estimation on the low-resolution video sequence through adaptive model parameters to obtain the motion estimation value at time t+1; sending the motion estimation value into the fusion module to obtain the Kalman motion estimation value; and correcting the super-resolved low-resolution blocks through the motion compensation module to improve inter-frame consistency.
[0108] The present invention innovatively develops a deep adaptive Kalman filtering algorithm, which combines traditional Kalman filtering with deep learning to achieve accurate motion prediction. The detailed algorithm is as follows:
[0109] 1. System state definition: In the video super-resolution task, define the state vector containing pixel position and motion speed information:
[0110] ,
[0111] where represents the pixel position at time, and represents the motion speed in the and directions. 2. State prediction stage: Predict the state at the current time based on the state at the previous time:
[0112] ,
[0113] where is the state transition matrix, is the control input matrix, is the control vector, is the process noise, which follows a Gaussian distribution with covariance . In the present invention, is obtained through learning instead of the fixed matrix in traditional Kalman filtering:
[0114] ,
[0115] Among them, is a small network composed of three 3×3 convolutional layers. The input is the state transition matrix and state vector at the previous moment, and the output is the state transition matrix at the current moment. represents the state transition matrix at the previous moment t-1. This adaptive design enables the model to dynamically adjust the state prediction strategy according to the content.
[0116] 3. Prediction covariance update: The covariance matrix of the predicted state is updated in the following way:
[0117] ,
[0118] Among them, is the state covariance matrix at the previous moment, is the process noise covariance matrix. In the present invention, is not a fixed value, but is obtained through network learning:
[0119] ,
[0120] Among them, is the noise estimation network, and the softplus activation function ensures that the covariance matrix is positive definite.
[0121] 4. State update stage: Based on the observation value obtained from the deep features , calculate the Kalman gain and update the state:
[0122] ,
[0123] ,
[0124] ,
[0125] Among them, is the observation matrix, is the observation noise covariance matrix, is the identity matrix, represents the updated state covariance matrix, reflecting the uncertainty of the state estimation. In the present invention, the observation value is extracted from the deep features through the feature correlation module:
[0126] ,
[0127] Among them, represents the feature correlation module, which is implemented by a 3×3 convolutional neural network, is The characteristics of the moment are the parameters of the feature correlation module. The observation noise covariance matrix is also obtained in an adaptive manner:
[0128] ,
[0129] wherein, is the observation noise estimation network.
[0130] 5. Motion compensation: Based on the updated state , calculate the motion offset, and achieve pixel-level motion compensation through bilinear interpolation:
[0131] ,
[0132] wherein, is extracted from the state vector , and represent the pixel coordinates in the target image, represents the input image.
[0133] Experiments show that compared with the traditional Kalman filter, the adaptive Kalman filter of the present invention shows significant advantages in dealing with complex motion patterns, and the motion prediction accuracy is improved by about 35%. Especially in dealing with scenarios such as non-rigid object motion and camera jitter, the effect is more obvious.
[0134] The present invention adopts an innovative composite strategy of first super-resolution and then compensation fine-tuning. Specifically, the system first performs preliminary super-resolution processing on the input low-resolution video frame sequence, and generates preliminary super-resolution results by extracting high-quality features through lightweight residual modules. Although these results have been significantly enhanced in the spatial dimension, there may be inconsistencies in the time dimension.
[0135] Subsequently, the system calculates the feature correlation between adjacent frames based on the deep feature extraction network to obtain the motion estimation value. It should be noted that the features of the original low-resolution video frames are used in this step, rather than the super-resolved features. This design takes into account the balance between computational efficiency and accuracy. Then, the motion estimation value is sent to the Kalman filter fusion module to be integrated with the historical motion information to generate a more accurate Kalman motion estimation value.
[0136] In the final key link, the motion compensation module uses the Kalman motion estimation value to precisely fine-tune and time-align the preliminary super-resolution results, and corrects the incoherence in the time dimension. The low-resolution blocks after super-resolution actually refer to the image blocks of preliminary super-resolution. Through motion compensation, adjacent frames are made consistent in time, eliminating jitter and flicker phenomena.
[0137] This strategy of first performing super-resolution and then compensation has multiple advantages compared to the traditional method of first compensating and then performing super-resolution: it allows each frame to first obtain the best spatial super-resolution reconstruction quality, and then specifically address the temporal domain consistency issue; effectively avoids the accuracy limitation of motion compensation in the low-resolution domain; better preserves high-frequency detail information; and simultaneously achieves the best balance between spatial quality and temporal smoothness.
[0138] Step 6: Combine the high-quality features with the information after feature fusion to reconstruct the high-definition image. In this step, by constructing a generator, the high-quality features extracted in Step 4 are combined with the information after feature fusion in Step 5, and the high-definition image is reconstructed through an upsampling and linear interpolation network. In this embodiment, sub-pixel convolution technology is used for upsampling. Compared with traditional bicubic interpolation, sub-pixel convolution can learn more complex upsampling patterns and produce higher-quality reconstruction results.
[0139] Preferably, the upsampling ratio can be 2x, 3x, or 4x, which is selected according to actual application requirements.
[0140] The calculation process of sub-pixel convolution is as follows:
[0141] ,
[0142] Among them, is the fused feature, Conv is the convolution operation, is the pixel shuffle operation, is the generated super-resolution image. The pixel shuffle operation is defined as:
[0143] ,
[0144] Among them, is the upsampling ratio, represents the floor operation, mod represents the modulo operation, represents the channel index in the output feature map.
[0145] In addition, the present invention also includes adversarial training of the method through a discriminator. The discriminator discriminates the authenticity of the generated super-resolution video and the ground-truth ultra-high-resolution video through an adversarial generation network. The loss function of the adversarial generation network is:
[0146] ,
[0147] Among them, is the adversarial loss function, is the multi-scale loss function, is the cycle consistency loss function, is the guidance loss function, is the total variation regularization loss function, and are the weight coefficients of the corresponding loss functions.
[0148] Preferably, . These weight values are determined through a large number of experiments and can achieve a good balance among various loss functions.
[0149] The adversarial loss function is calculated as follows:
[0150] ,
[0151] where, represents the discriminator, represents the generator, represents the real high-resolution image, represents the low-resolution input represents the generated super-resolution image, represents the expectation of the distribution of the real high-resolution image , represents the expectation of the distribution of the low-resolution input image .
[0152] The multi-scale loss function comprehensively evaluates the reconstruction quality by calculating the difference between the generated image and the real image at different scales:
[0153] ,
[0154] where, represents the operation of downsampling the image by times, is the number of scales considered (usually set to 3), is the weight coefficient of the th scale. This multi-scale design enables the model to simultaneously focus on the global structure and local details, represents the generated super-resolution image.
[0155] The formula for calculating the perceptual feature loss function is:
[0156] ,
[0157] where, represents extracting the features of the th layer, is the number of feature layers extracted in the loss function, is the generated super-resolution image, is the real ultra-high-resolution image, represents the weight coefficient of the features of the sth layer. Preferably, Set to 5, considering features from shallow to deep. The discriminator is initialized with a pre-trained model extracted from the VGG network, and the perceptual difference loss is calculated by extracting multi-layer features of the input image. The feature learning network outputs features from the last convolutional layer of the ImageNet pre-trained model. This pre-training strategy enables the model to converge faster and obtain better generalization ability.
[0158] Refer to Figure 6 , the lightweight residual module of the present invention includes a plurality of densely connected residual blocks, each residual block includes layer normalization, a convolutional layer, and a Leaky-ReLU activation layer; a deformable convolution is provided in the residual block for extracting local motion information of video frames; dense connections are added between adjacent residual blocks, connecting the input of each residual block to the outputs of all previous residual blocks.
[0159] The mathematical definition of the recursive residual network is:
[0160] ,
[0161] Among them, represents the input feature, represents the recursive convolutional layer, The layer number of the identity module for the residual connection operation is defined as:
[0162] ,
[0163] Among them, represents the low-resolution video frame feature, represents the fused feature, represents the residual connection operation, represents the feature at the previous moment t - 1, represents the recursive convolutional layer.
[0164] Preferably, three layers of recursive residual networks are stacked for progressive feature extraction. The number of convolutional kernels in the first recursive convolutional layer is 160, the second layer is 256, and the third layer is 256. This configuration achieves a good balance between model capacity and computational efficiency. The Leaky-ReLU activation function is defined as
[0165] ,
[0166] Among them, is the leakage parameter, preferably set to 0.2. Compared with the standard ReLU, Leaky-ReLU avoids the dead ReLU problem and enhances the expressive ability of the model, represents the input value, represents the condition if the input value.
[0167] The deformable convolution operation is defined as:
[0168] ,
[0169] where, is the input feature map, is the current position, is the predefined offset of the -th sampling point, is the learned offset, is the corresponding weight, is the number of sampling points. In this embodiment, is set to 9, corresponding to a 3×3 convolution kernel.
[0170] The implementation method of dense connection is:
[0171] ,
[0172] where, is the output of the -th layer, represents the connection of the outputs of the previous layers, is the non-linear transformation function of the -th layer. This dense connection mechanism promotes feature reuse, improves gradient flow, and accelerates model training.
[0173] Referring to Figure 7 , the present invention also provides a real-time video super-resolution reconstruction system for multi-modal fusion, including:
[0174] Feature extraction module 1, which is used to extract visual features and text features of a low-resolution video by using a CLIP model, and perform dual-modal feature fusion to generate guiding features;
[0175] Feature alignment module 2, which is used to align the guiding features with the low-resolution video through a multi-modal fusion module;
[0176] Feature optimization module 3, which is used to extract high-quality features by using a lightweight residual module;
[0177] Motion compensation module 4, which is used to fuse multi-frame features through inter-frame information flow propagation and perform motion compensation by using adaptive Kalman filtering;
[0178] Image reconstruction module 5, which is used to combine the high-quality features with the information after feature fusion to reconstruct a high-definition image;
[0179] Network training module 6, which is used to train and optimize the system through an adversarial generation network.
[0180] Preferably, the feature extraction module 1 includes a visual feature extraction unit 11 and a text feature extraction unit 12. The visual feature extraction unit 11 processes low-resolution video frames using a ViT encoder and a ResNet network, and the text feature extraction unit 12 processes text description information using the Text Encoder of the CLIP model.
[0181] The feature alignment module 2 includes a parallel convolution unit 21 and a fusion control unit 22. The parallel convolution unit 21 performs convolution processing on the guidance features and the low-resolution video features, and the fusion control unit 22 controls the fusion degree through a guidance gate.
[0182] The feature optimization module 3 includes a channel transformation unit 31 and a residual processing unit 32. The channel transformation unit 31 is responsible for the conversion of the tensor shape and the adjustment of the number of channels, and the residual processing unit 32 realizes the residual connection and enhancement of the features.
[0183] The motion compensation module 4 includes a feature correlation unit 41 and a Kalman filter unit 42. The feature correlation unit 41 extracts inter-frame correlation information, and the Kalman filter unit 42 performs motion estimation and prediction.
[0184] The image reconstruction module 5 includes a feature fusion unit 51 and an upsampling unit 52. The feature fusion unit 51 combines high-quality features and the fused information, and the upsampling unit 52 realizes high-resolution reconstruction through sub-pixel convolution.
[0185] The network training module 6 includes a generator training unit 61 and a discriminator training unit 62. The generator training unit 61 optimizes the parameters of the reconstruction network, and the discriminator training unit 62 provides an adversarial training signal.
[0186] The system of the present invention realizes a modular design, and the cooperation between the functional modules is close, forming a complete processing flow. The system has high scalability and adaptability, and can adjust the parameters and configurations of each module according to actual application requirements.
[0187] In the actual application process, the system can dynamically adjust the processing strategy according to different video contents. For example, for scenes with a large amount of details (such as landscape videos), the system will rely more on visual features; for scenes with clear semantic information (such as human conversations), the system will increase the weight of text features. Experiments show that this dynamic adjustment strategy enables the system to achieve good performance on different types of video contents.
[0188] In addition, the system supports deployment on multiple hardware platforms, including desktop GPUs, mobile GPUs, and dedicated chips. On the NVIDIA RTX 2080Ti, the system can process 1080p video at 30fps; on the NVIDIA RTX 3090, the system can process 4K video at 25fps. Through optimization techniques such as model quantization and pruning, the system can also achieve real-time processing of 720p video on mid-range mobile devices.
[0189] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A real-time video super-resolution reconstruction method based on multi-modal fusion, characterized in that It includes the following steps: Obtain a low-resolution video sequence; Adopt the CLIP model to extract visual features and text features, and perform bimodal feature fusion to generate guidance features; Align the features of the guidance features and the low-resolution video through a multimodal fusion module; Adopt a lightweight residual module to extract high-quality features; Fuse multi-frame features through inter-frame information flow propagation and perform motion compensation using adaptive Kalman filtering; Combine the high-quality features with the multi-frame feature information after inter-frame information flow propagation fusion to reconstruct a high-definition image; The specific steps of performing motion compensation by the adaptive Kalman filter include: Based on the depth features extracted by the feature learning network, use a 3×3 convolutional neural network to calculate feature correlations; Perform motion estimation on the low-resolution video sequence through an adaptive model parameter to obtain the motion estimation value at time t+1; Send the motion estimation value at time t+1 into the fusion module to obtain the Kalman motion estimation value; Correct the super-resolved low-resolution blocks through a motion compensation module to improve inter-frame consistency.
2. The real-time video super-resolution reconstruction method with multimodal fusion according to claim 1, characterized in that: The steps of the CLIP model extracting visual features and text features specifically include: Input the low-resolution video sequence into the ViT encoder and the ResNet network to extract visual features; Input the context description in the low-resolution video frame into the Text Encoder to extract text features; Calculate the similarity score between the visual features and the text features through a cross-modal similarity encoder; Generate guidance features by passing the visual features and the text features through a gated fusion module.
3. The real-time video super-resolution reconstruction method with multi-modal fusion according to claim 1, wherein: The multimodal fusion module dynamically integrates the guidance features and the visual features extracted from the low-resolution video sequence through a gating mechanism, The specific steps thereof are: Input the guidance features and the visual features extracted from the low-resolution video sequence into two parallel convolutional layers and ReLU layers; Add the outputs of the two parallel branches and then input them into another convolutional layer and ReLU layer; Multiply the output result by the guidance gate in the guidance module to control the fusion degree and generate aligned features.
4. The real-time video super-resolution reconstruction method with multi-modal fusion according to claim 1, wherein: The residual path of the lightweight residual module adopts a channel-separated convolutional layer to reduce the computational complexity, and the specific steps are as follows: Convert the video frame from a size of to a tensor with 4 channels ; Convert the said tensor to a tensor with a shape of where represents the number of channels; reduce the number of channels to 2 through a convolutional layer and duplicate two copies; Add three tensors with the same shape and pass through the ReLU activation function; Transform the shape of the tensor group from to , and rearrange it into through the convolutional layer; Double the tensor channels to obtain a feature representation with 2 channels. of the feature representation.
5. The real-time video super-resolution reconstruction method with multi-modal fusion according to claim 1, characterized in that: The adaptive Kalman filter includes a multi-branch multi-layer perceptron, a fusion module, and a feature correlation module for calculating feature correlations, and its Kalman motion estimation value is obtained by combining feature maps of different branches; wherein the feature correlation module is implemented by a 3×3 convolutional neural network, and the feature correlation module and the fusion module perform layer normalization on the output of each layer of the network.
6. The real-time video super-resolution reconstruction method with multi-modal fusion according to claim 1, characterized in that: The lightweight residual module includes a plurality of densely connected residual blocks, each residual block includes layer normalization, a convolutional layer, and a Leaky-ReLU activation layer; deformable convolutions are provided in the residual blocks for extracting local motion information of video frames; dense connections are added between adjacent residual blocks to connect the input of each residual block with the outputs of all previous residual blocks.
7. The real-time video super-resolution reconstruction method with multi-modal fusion according to claim 1, characterized in that: The method is trained adversarially by a discriminator, where the discriminator loss function consists of an adversarial loss function, a multi-scale loss function, a cycle consistency loss function, a guidance loss function, and a total variation regularization loss function; the multi-scale loss function is used to evaluate the reconstruction quality at different scales; the cycle consistency loss function is used to ensure feature consistency.
8. The real-time video super-resolution reconstruction method with multi-modal fusion according to claim 7, characterized in that: The discriminator is initialized with a pre-trained model extracted from the VGG network; by extracting multi-layer features of the input image, the perceptual difference loss is calculated; the feature learning network outputs features from the last convolutional layer of the ImageNet pre-trained model.
9. A real-time video super-resolution reconstruction system for multi-modal fusion implementing the method according to any one of claims 1-8, characterized in that, It includes: A feature extraction module, which is used to extract visual features and text features of a low-resolution video by using a CLIP model, and perform bimodal feature fusion to generate guidance features; A feature alignment module, which is used to align the guidance features with the low-resolution video through a multi-modal fusion module; A feature optimization module, which is used to extract high-quality features by using a lightweight residual module; A motion compensation module, which is used to fuse multi-frame features through inter-frame information flow propagation and perform motion compensation by using adaptive Kalman filtering; An image reconstruction module, which is used to combine the high-quality features with the information after feature fusion to reconstruct a high-definition image; A network training module, which is used to train and optimize the system through an adversarial generation network.
Citation Information
Patent Citations
Video resolution improving system and method based on pre-training video generation model
CN119048356A