Video super-resolution method and device based on spatial-temporal feature aggregation

Through the fusion method of optical flow calculation and multi-scale spatiotemporal feature fusion, combined with diffusion model and color repair technology, the artifacts and inter-frame discontinuity problems in complex motion scenes in video super-resolution are solved, and high-quality super-resolution video is generated.

CN120355572APending Publication Date: 2025-07-22XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510256527.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing video super-resolution methods are prone to visual artifacts and inter-frame discontinuity problems when dealing with complex motion scenes or fast motion scenes.

Method used

By obtaining the target video frame and its video frames at adjacent moments for optical flow calculation, reverse distortion and multi-scale spatiotemporal feature fusion, super-resolution video frames are generated using the pre-trained diffusion model, and image color repair is performed by combining low-rank adaptive model and adaptive example normalization technology.

Benefits of technology

It effectively reduces artifacts in fast motion or complex scenes, improves the spatial and temporal consistency and visual consistency between video frames, and generates clearer and more continuous videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355572A_ABST
    Figure CN120355572A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video super-resolution method and device based on spatial-temporal feature aggregation. The method comprises the following steps: acquiring a target video frame, a previous moment video frame and a next moment video frame in a to-be-processed video; acquiring an optical flow between two adjacent video frames; according to the optical flow between the two adjacent video frames, the video frame at the previous moment, the target video frame and the video frame at the next moment, reverse distortion and multi-scale spatial-temporal feature fusion are carried out, and context image features of the target video frame are obtained; and inputting the context image features of the target video frame into a pre-trained diffusion model to obtain a super-resolution video frame of the target video frame. According to the method, through multi-scale optical flow calculation, the influence of space-time dislocation is reduced, so that the reconstructed video is more continuous; and multi-scale spatio-temporal feature fusion is carried out based on three adjacent video frames, so that the spatio-temporal consistency and the visual consistency among the video frames are improved, and the artifact phenomenon of the diffusion model in a rapid motion or complex scene is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image restoration, and particularly to a video super-resolution method and device based on spatio-temporal feature aggregation. Background Art

[0002] The purpose of video super-resolution technology is to improve the spatial resolution of video frames, making the video visually clearer and more delicate, thereby improving the user's viewing experience. Video super-resolution methods can be divided into two categories: traditional methods and deep learning-based methods.

[0003] Traditional video super-resolution methods can, to a certain extent, improve the resolution of videos, but they are insufficient in dealing with complex motion scenarios or detail restoration, and are prone to problems such as artifacts, edge blurring, and detail loss, resulting in a low visual quality of the reconstructed video; although deep learning methods can effectively improve the clarity and detail performance of videos, they are still prone to problems of inter-frame discontinuity when dealing with fast motion scenarios. Therefore, current video super-resolution methods have problems such as visual artifacts and inter-frame discontinuity in fast motion scenarios. Summary of the Invention

[0004] Based on this, it is necessary to address the above problems and propose a video super-resolution method and device based on spatio-temporal feature aggregation to solve the problems of visual artifacts and inter-frame discontinuity in complex scenarios or fast motion scenarios.

[0005] To achieve the above object, a first aspect of the present application provides a video super-resolution method based on spatio-temporal feature aggregation, the method comprising:

[0006] Obtain a target video frame in the video to be processed, as well as the previous moment video frame and the next moment video frame adjacent to the target video frame, wherein the target video frame is any frame in the video to be processed;

[0007] Perform optical flow calculation based on the previous moment video frame, the target video frame, and the next moment video frame to obtain the optical flow between two adjacent video frames;

[0008] Perform inverse warping and multi-scale spatio-temporal feature fusion according to the optical flow between two adjacent video frames, the previous moment video frame, the target video frame, and the next moment video frame to obtain the context image feature of the target video frame after feature aggregation;

[0009] Input the context image feature of the target video frame into a pre-trained diffusion model to obtain the super-resolution video frame of the target video frame.

[0010] Further, performing optical flow calculation based on the previous video frame, the target video frame, and the next video frame to obtain the optical flow between two adjacent video frames, specifically including:

[0011] Extracting the image features of the previous video frame, the target video frame, and the next video frame;

[0012] Performing optical flow calculation based on the image features of the previous video frame and the target video frame to obtain the first optical flow;

[0013] Performing optical flow calculation based on the image features of the target video frame and the next video frame to obtain the second optical flow.

[0014] Further, performing inverse warping and multi-scale spatio-temporal feature fusion based on the optical flow between two adjacent video frames, the previous video frame, the target video frame, and the next video frame to obtain the context image features of the target video frame after feature aggregation, specifically including:

[0015] Performing pixel-level inverse warping operation according to the image features of the previous video frame and the first optical flow to obtain the first sub-image feature;

[0016] Performing pixel-level inverse warping operation according to the image features of the next video frame and the second optical flow to obtain the second sub-image feature;

[0017] Performing multi-scale spatio-temporal feature fusion on the first sub-image feature, the second sub-image feature, and the image features of the target video frame to obtain the context image features of the target video frame after feature aggregation.

[0018] Further, performing multi-scale spatio-temporal feature fusion on the first sub-image feature, the second sub-image feature, and the image features of the target video frame to obtain the context image features of the target video frame after feature aggregation, specifically including:

[0019] Inputting the first sub-image feature and the second sub-image feature into a preset aggregation module to obtain the aggregated feature output by the aggregation module;

[0020] Using a preset channel attention fusion module to perform feature fusion on the aggregated feature and the image features of the target video frame to obtain the context image features of the target video frame.

[0021] Further, inputting the context image features of the target video frame into a pre-trained diffusion model to obtain the super-resolution video frame of the target video frame, specifically including:

[0022] Adjust the weights in the diffusion model using the low-rank adaptation model and the context image features to obtain an adjusted target diffusion model;

[0023] Input the context image features of the target video frame into the target diffusion model to obtain a super-resolution video frame output by the target diffusion model.

[0024] Furthermore, the low-rank adaptation model includes a degradation evaluation network, a KAN layer, and an independent ID assignment strategy;

[0025] The adjusting the weights in the diffusion model using the low-rank adaptation model and the context image features to obtain an adjusted target diffusion model specifically includes:

[0026] Input the context image features into the degradation evaluation network for two-dimensional vector conversion and Gaussian Fourier conversion to obtain degradation features;

[0027] Use the independent ID assignment strategy to set a unique number embedding layer for each target network block in the diffusion model, where the target network block is a network block whose preset weights need to be adjusted;

[0028] Input the number embedding layer and the degradation features into the KAN layer to generate a fine-tuning matrix for the target network block;

[0029] Use the fine-tuning matrix of the target network block to adjust the weights of the target network block in the diffusion model to obtain an adjusted target diffusion model.

[0030] Furthermore, the degradation features are obtained based on the following formula:

[0031] f d =concat[sin(2πdM T ),cossin(2πdM T )]

[0032] In the formula, f d is the degradation feature, concat represents concatenation along a specified dimension, d is a two-dimensional vector generated based on the context features, and M is a randomly initialized matrix.

[0033] Furthermore, the using the fine-tuning matrix of the target network block to adjust the weights of the target network block in the diffusion model to obtain an adjusted target diffusion model specifically includes:

[0034] Obtain the original weights of the target network block;

[0035] Calculate the adjusted weight of the target network block based on the original weight, the fine-tuning matrix of the target network block, and a preset low-rank matrix;

[0036] Substitute the adjusted weight into the diffusion model to obtain an adjusted target diffusion model.

[0037] Further, the method further includes:

[0038] Based on adaptive instance normalization, perform image color restoration on the super-resolution video of the target video frame and the target video frame to obtain a color-restored super-resolution video.

[0039] To achieve the above object, a second aspect of the present application provides a video super-resolution device based on spatio-temporal feature aggregation. The device includes: a feature extraction module, a multi-scale spatio-temporal feature aggregation module, and an image reconstruction module;

[0040] The feature extraction module is configured to obtain a target video frame in a video to be processed, as well as a previous video frame and a next video frame adjacent to the target video frame. Wherein, the target video frame is any frame in the video to be processed;

[0041] The multi-scale spatio-temporal feature aggregation module is configured to perform optical flow calculation based on the previous video frame, the target video frame, and the next video frame to obtain the optical flow between two adjacent video frames;

[0042] Perform inverse warping and multi-scale spatio-temporal feature fusion according to the optical flow between two adjacent video frames, the previous video frame, the target video frame, and the next video frame to obtain the context image features of the target video frame after feature aggregation;

[0043] The image reconstruction module inputs the context image features of the target video frame into a pre-trained diffusion model to obtain a super-resolution video frame of the target video frame.

[0044] Adopting the embodiments of the present invention has the following beneficial effects:

[0045] An embodiment of the present invention proposes a video super-resolution method based on spatio-temporal feature aggregation. The method includes: obtaining a target video frame in the video to be processed, as well as the previous video frame and the next video frame adjacent to the target video frame, where the target video frame is any frame in the video to be processed; performing optical flow calculation based on the previous video frame, the target video frame, and the next video frame to obtain the optical flow between two adjacent video frames; performing inverse warping and multi-scale spatio-temporal feature aggregation according to the optical flow between two adjacent video frames, the previous video frame, the target video frame, and the next video frame to obtain the context image features of the target video frame after feature aggregation; inputting the context image features of the target video frame into a pre-trained diffusion model to obtain the super-resolution video frame of the target video frame. Through multi-scale optical flow calculation, the present invention can accurately align video frames at different times, thereby reducing the influence of spatio-temporal misalignment during feature aggregation and making the reconstructed video more continuous. Secondly, multi-scale spatio-temporal feature aggregation is performed based on three adjacent video frames to fully fuse spatio-temporal information and generate the context image features of adjacent frames, so as to improve spatio-temporal consistency and visual consistency between video frames, and effectively reduce the artifact phenomenon of the diffusion model in fast-moving or complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0047] Among them:

[0048] Figure 1 is a schematic flowchart of the video super-resolution method based on spatio-temporal feature aggregation in the embodiment of the present invention;

[0049] Figure 2 is a diagram of the video super-resolution model based on diffusion prior and spatio-temporal feature aggregation in the embodiment of the present invention;

[0050] Figure 3 is a diagram of the multi-scale spatio-temporal feature aggregation module in the embodiment of the present invention;

[0051] Figure 4 is a schematic diagram of the low-rank adaptation module and the diffusion model in the embodiment of the present invention;

[0052] Figure 5 is a structural block diagram of the video super-resolution device based on spatio-temporal feature aggregation in the embodiment of the present invention;

[0053] Figure 6This is the internal structure diagram of the computer device in the embodiment of the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] To make the video in fast motion or complex scenes clearer and more delicate, the present invention proposes a video super-resolution method based on spatio-temporal feature aggregation, which can be referred to Figure 1 , Figure 1 This is the flowchart of the video super-resolution method based on spatio-temporal feature aggregation in the embodiment of the present invention. The method includes:

[0056] Step 110: Obtain the target video frame in the video to be processed, as well as the previous video frame and the next video frame adjacent to the target video frame at the previous moment, where the target video frame is any frame in the video to be processed.

[0057] In the embodiment of the present invention, the video that needs to improve clarity is selected as the video to be processed, and any frame is arbitrarily selected from the video to be processed as the target video frame, so as to optimize the clarity of the target video frame.

[0058] Specifically, if the video frame at time t is selected as the target video frame, then continue to obtain the video frames at (t - 1) and (t + 1) adjacent to time t as the previous video frame and the next video frame respectively, so as to perform image feature analysis and fusion of the target video frame based on the adjacent three frames, and improve the continuity of the reconstructed video by capturing the inter-frame dependence relationship.

[0059] Step 120: Calculate the optical flow based on the previous video frame, the target video frame, and the next video frame to obtain the optical flow between two adjacent video frames.

[0060] In the embodiment of the present invention, the optical flow is calculated based on the image content in the previous video frame and the target video frame to estimate the optical flow from the target video frame to the previous video frame; the optical flow is calculated based on the target video frame and the next video frame to estimate the optical flow from the target video frame to the next video frame. By estimating the two-way intermediate optical flow, the motion attributes between frames are reflected.

[0061] Step 130: Perform inverse warping and multi-scale spatio-temporal feature fusion based on the optical flow between two adjacent video frames, the previous video frame, the target video frame, and the next video frame to obtain the context image features of the target video frame after feature aggregation.

[0062] In the embodiment of the present invention, an inverse warping operation is performed based on the previous video, the next video, and the optical flow between two adjacent videos to respectively obtain two sub-image features between the target video frame and the adjacent video frames. Then, feature fusion is performed according to the two sub-image features and the target video frame to obtain the context image features after multi-frame feature aggregation related to the target video frame. By fusing the features of multiple video frames, the spatio-temporal consistency and visual consistency are improved.

[0063] Step 140: Input the context image features of the target video frame into a pre-trained diffusion model to obtain the super-resolution video frame of the target video frame.

[0064] In the embodiment of the present invention, a pre-trained diffusion model is used to reconstruct a high-resolution image. By inputting the context image features of the target video frame obtained in Step 130 into the diffusion model, the super-resolution video frame output by the diffusion model can be obtained.

[0065] In an embodiment of the present invention, a single-step diffusion model is selected for image reconstruction, such as SD-Turbo. Since traditional diffusion models usually require a large number of time steps to gradually denoise and generate high-quality samples, while SD-Turbo distills the key steps of the diffusion process and successfully reduces the sampling steps significantly, thereby improving the inference speed. Therefore, the single-step diffusion model can efficiently handle the video reconstruction task.

[0066] Through the optical flow calculation between multiple frames, the present invention can accurately align video frames at different times, thereby reducing the impact of spatio-temporal misalignment during feature aggregation and making the reconstructed video more continuous. Secondly, multi-scale spatio-temporal feature fusion is performed based on three adjacent video frames to fully fuse spatio-temporal information and generate the context image features of adjacent frames, so as to improve spatio-temporal consistency and visual consistency between video frames, and effectively reduce the artifact phenomenon of the diffusion model in fast-moving or complex scenes.

[0067] In an embodiment of the present invention, a video super-resolution model based on spatio-temporal feature aggregation with diffusion prior is proposed, which can refer to Figure 2 , Figure 2This is a diagram of a video super-resolution model based on diffusion prior for spatio-temporal feature aggregation according to an embodiment of the present invention. The model takes the original video frame LR as the low-resolution video frame as input, uses a shallow feature extraction module for feature extraction, aggregates the image features of the previous video frame and the next video frame through a multi-scale spatio-temporal feature aggregation module, and takes the output and the image features of the target video frame as the input of a channel attention fusion module for feature fusion to obtain context image features. The context image features are used as input and input to a low-rank adaptation module and a diffusion model for image reconstruction to generate a super-resolution video frame, and the super-resolution video frame is subjected to adaptive instance normalization processing to generate the final super-resolution video frame HR.

[0068] Specifically, based on the proposed video super-resolution model based on diffusion prior for spatio-temporal feature aggregation, step 120: Calculate the optical flow based on the previous video frame, the target video frame, and the next video frame to obtain the optical flow between two adjacent video frames, specifically including:

[0069] Step210: Extract the image features of the previous video frame, the target video frame, and the next video frame.

[0070] Specifically, perform shallow feature extraction on the previous video frame, the target video frame, and the next video frame to obtain the image features of each moment respectively.

[0071] In the embodiment of the present invention, ResNet-18 can be used to perform shallow feature extraction on the previous video frame, the target video frame, and the next video frame to obtain the image feature F(t - 1) of the previous video frame, the image feature F(t) of the target video frame, and the image feature F(t + 1) of the next video frame respectively. ResNet-18 has a small number of parameters and computational complexity, and is suitable as a shallow feature extractor for extracting low-level features (such as edges, textures, and simple shapes). In addition, the skip connections in ResNet-18 allow the network to directly learn the residuals between the input and the output, reducing unnecessary feature transformations, thereby improving the efficiency and quality of feature extraction.

[0072] Step220: Calculate the optical flow according to the image features of the previous video frame and the target video frame to obtain the first optical flow.

[0073] Step230: Calculate the optical flow according to the image features of the target video frame and the next video frame to obtain the second optical flow.

[0074] In the embodiment of the present invention, the multi-scale spatio-temporal feature aggregation module includes a multi-scale spatio-temporal feature estimation module and an intermediate frame aggregation module. Refer to Figure 3 , Figure 3This is a diagram of the multi-scale spatio-temporal feature aggregation module in an embodiment of the present invention. After obtaining the image features F(t-1) of the previous video frame, the image features F(t) of the target video frame, and the image features F(t+1) of the next video frame, first, the image features F(t-1) and F(t+1) are input into the multi-scale spatio-temporal feature estimation module. The multi-scale spatio-temporal feature estimation module randomly selects downsampling scales, such as the original size, twice-downsampled size, or four-times-downsampled size, etc., to scale the image features F(t-1) and F(t+1). After processing through convolutional blocks and residual blocks, bidirectional intermediate optical flows between the previous video frame and the next video frame and the target video frame are generated respectively. Then, the estimated optical flows are enlarged to the original feature size.

[0075] Specifically, step A: After scaling the image features F(t-1) of the previous video frame and the image features F(t+1) of the next video frame, estimate the first initial optical flow from the image features F(t) of the target video frame to the image features F(t-1) of the previous video frame, and the second initial optical flow from the image features F(t) of the target video frame to the image features F(t+1) of the next video frame.

[0076] Step B: Enlarge the first initial optical flow and the second initial optical flow to the original size to obtain the first optical flow f t→(t-1) and the second optical flow f t→(t+1) .

[0077] In an embodiment of the present invention, step 130: According to the optical flow between two adjacent video frames, the previous video frame, the target video frame, and the next video frame, perform inverse warping and multi-scale spatio-temporal feature fusion to obtain the context image features of the target video frame after feature aggregation, specifically including:

[0078] Step310: Perform a pixel-level inverse warping operation according to the image features of the previous video frame and the first optical flow to obtain the first sub-image features.

[0079] Step320: Perform a pixel-level inverse warping operation according to the image features of the next video frame and the second optical flow to obtain the second sub-image features.

[0080] Specifically, step C: The image features F(t-1) of the previous video frame and the first optical flow f t→(t-1) Use a pixel-level inverse warping operation to obtain the first sub-image features; the image features F(t+1) of the next video frame and the second optical flow f t→(t+1) Use a pixel-level inverse warping operation to obtain the second sub-image features. As shown in formulas (1) and (2):

[0081]

[0082] In the formula, f (t-1)→t is the first sub-image feature, and F (t+1)→t is the second sub-image feature, represents a pixel-level inverse warping operation.

[0083] Step D: Use the first sub-image feature F (t-1)→t as the image feature F(t - 1) of the previous video frame, and use the second sub-image feature F (t+1)→t as the image feature F(t + 1) of the next video frame, and repeat steps A - C for m iterations. The number of iterations can be determined according to the actual situation, for example, four iterations can be performed. After completing the preset number of iterations, generate the final first sub-image feature F (t-1)→t and the second sub-image feature F (t+1)→t .

[0084] Step330: Perform multi-scale spatio-temporal feature fusion on the first sub-image feature, the second sub-image feature, and the image feature of the target video frame to obtain the context image feature of the target video frame after feature aggregation.

[0085] In the embodiments of the present invention, feature fusion is performed based on two sub-image features and the target video frame to obtain the context image feature after multi-frame feature aggregation related to the target video frame. By fusing the features of multiple video frames, spatio-temporal consistency and visual consistency are improved.

[0086] In an embodiment of the present invention, Step330: Perform multi-scale spatio-temporal feature fusion on the first sub-image feature, the second sub-image feature, and the image feature of the target video frame to obtain the context image feature of the target video frame after feature aggregation, which specifically includes:

[0087] Step331: Input the first sub-image feature and the second sub-image feature into a preset aggregation module to obtain the aggregation feature output by the aggregation module.

[0088] In the embodiments of the present invention, the preset aggregation module is an intermediate frame aggregation module. The intermediate frame aggregation module includes a 3×3 convolutional block, six residual blocks, a 1×1 convolutional block, and a Tanh activation function. The result of adding the first sub-image feature and the second sub-image feature is used as the input of the intermediate frame aggregation module. After being processed by a 3×3 convolutional block, six residual blocks, a 1×1 convolutional block, and a Tanh activation function, the output result is obtained, and the result of adding the first sub-image feature and the second sub-image feature is skip-connected to obtain the aggregation feature h t .

[0089] Step332. Use the preset channel attention fusion module to perform feature fusion on the aggregated features and the image features of the target video frame, and obtain the context image features of the target video frame.

[0090] In the embodiment of the present invention, the aggregated feature h t and the image feature F(t) of the target video frame are input into the preset channel attention fusion module for feature fusion, and the context image features of the target video frame after multi-frame feature aggregation are obtained, as shown in formula (3):

[0091] FT(t) = CAF(h t , F(t)) (3)

[0092] In the formula, FT(t) is the context image feature of the target video frame, and CAF() represents the channel attention fusion operation.

[0093] In an embodiment of the present invention, step 140. Input the context image features of the target video frame into the pre-trained diffusion model to obtain the super-resolution video frame of the target video frame, which specifically includes:

[0094] Step410. Use the low-rank adaptation model and the context image features to adjust the weights in the diffusion model, and obtain the adjusted target diffusion model.

[0095] In order to utilize the rich prior information contained in the diffusion model, the embodiment of the present invention designs a low-rank adaptation module to fine-tune the weights of the VAE encoder, Unet network, and VAE decoder in the pre-trained diffusion model, so that the model can better adapt to the super-resolution reconstruction task.

[0096] In an embodiment of the present invention, the low-rank adaptation model includes three parts: a degradation evaluation network, a KAN layer, and an independent ID assignment strategy. Refer to Figure 4 , Figure 4 which is the schematic diagram of the low-rank adaptation module and the diffusion model in the embodiment of the present invention. Step410. Use the low-rank adaptation model and the context image features to adjust the weights in the diffusion model, and obtain the adjusted target diffusion model, which specifically includes:

[0097] Step411. Input the context image features into the degradation evaluation network for two-dimensional vector conversion and Gaussian Fourier conversion to obtain degradation features.

[0098] In the embodiment of the present invention, the degradation evaluation network adopts the pre-trained unsupervised degradation estimation network UDEM, and UDEM estimates the input context image features as a two-dimensional vector The two-dimensional vector characterizes the noise and blur degree. After obtaining the two-dimensional vector d, a degradation feature is obtained based on the two-dimensional vector d by using a Gaussian Fourier embedding layer, as shown in formula (4).

[0099] f d = concat[sin(2πdM T ), cossin(2πdM T )]

[0100] In the formula, f d is the degradation feature, concat represents concatenation along a specified dimension, d is the two-dimensional vector generated based on the context feature, and M is a randomly initialized matrix.

[0101] Step412: Use an independent ID assignment strategy to set a unique numbered embedding layer for each target network block in the diffusion model, where the target network block is a network block with preset weights to be adjusted.

[0102] In the embodiments of the present invention, the independent ID assignment strategy is responsible for setting different numbered embedding layers for the network blocks that need to be adjusted in the diffusion model. For example, 6 embedding layers with index numbers from 0 to 5 are set for the VAE encoder in the diffusion model, 10 embedding layers with index numbers from 0 to 9 are set for the U-net network in the diffusion model, and 6 embedding layers with index numbers from 0 to 5 are set for the VAE decoder in the diffusion model.

[0103] Step413: Input the numbered embedding layer and the degradation feature into the KAN layer to generate a fine-tuning matrix for the target network block.

[0104] The KAN layer is an attention mechanism introducing a kernel function, which enhances the ability of the traditional attention mechanism through the kernel method and can efficiently process high-dimensional data. In the embodiments of the present invention, the embedding layers of each network block generated based on the independent ID assignment strategy and the degradation feature f d are input into the KAN layer to generate a fine-tuning matrix for the weights of each network block. As shown in formula (5):

[0105] K i = KAN(concat(f d , E i )) (5)

[0106] In the formula, K is the fine-tuning matrix, i is the embedding layer number, KAN() is the KAN network layer, and E i is the index-numbered embedding layer.

[0107] Step414: Use the fine-tuning matrix of the target network block to adjust the weights of the target network block in the diffusion model to obtain an adjusted target diffusion model.

[0108] Specifically, the original weights of the target network block in the diffusion model are fine-tuned by the fine-tuning matrix of the target network block and applied to the diffusion model to generate a high-resolution reconstructed image based on the adjusted target diffusion model.

[0109] In an embodiment of the present invention, Step414: Adjust the weights of the target network block in the diffusion model by using the fine-tuning matrix of the target network block to obtain an adjusted target diffusion model, specifically including:

[0110] Step4141: Obtain the original weights of the target network block.

[0111] Step4142: Calculate the adjusted weights of the target network block based on the original weights, the fine-tuning matrix of the target network block, and a preset low-rank matrix.

[0112] Step4143: Substitute the adjusted weights into the diffusion model to obtain an adjusted target diffusion model.

[0113] Specifically, the adjusted weights of the target network block are calculated by formula (6):

[0114] W n = W + AK i B (6)

[0115] In the formula, W n is the adjusted weight, W is the original weight, both A and B are low-rank matrices, and represent low-rank matrices, r << d and r << n.

[0116] Step420: Input the context image features of the target video frame into the target diffusion model to obtain a super-resolution video frame output by the target diffusion model.

[0117] The weights of the diffusion model are fine-tuned by the context image features of the target video frame, so that the super-resolution video frame output by the fine-tuned diffusion model is clearer.

[0118] In an embodiment of the present invention, to reduce the color deviation in the reconstructed image, image color restoration is also performed based on adaptive instance normalization according to the super-resolution video of the target video frame and the target video frame to obtain a super-resolution video after color restoration. By adjusting the mean and standard deviation of the feature map, the style in the target video frame is embedded into the super-resolution video frame, thereby realizing style control and feature transformation to reduce the color deviation of the generated super-resolution video frame.

[0119] By performing the above-mentioned video super-resolution processing based on spatio-temporal feature aggregation on each frame of the video to be processed, the reconstructed high-definition video can be obtained.

[0120] The present invention also proposes a video super-resolution device based on spatio-temporal feature aggregation. Please refer to Figure 5 , Figure 5 which is the structural block diagram of the video super-resolution device based on spatio-temporal feature aggregation in the embodiments of the present invention, and includes: a feature extraction module 501, a multi-scale spatio-temporal feature aggregation module 502, and an image reconstruction module 503.

[0121] The feature extraction module 501 is used to obtain the target video frame in the video to be processed, as well as the previous video frame and the next video frame adjacent to the target video frame. Among them, the target video frame is any frame in the video to be processed.

[0122] The multi-scale spatio-temporal feature aggregation module 502 is used to perform optical flow calculation based on the previous video frame, the target video frame, and the next video frame to obtain the optical flow between two adjacent video frames.

[0123] According to the optical flow between two adjacent video frames, the previous video frame, the target video frame, and the next video frame, inverse warping and multi-scale spatio-temporal feature fusion are performed to obtain the context image features of the target video frame after feature aggregation.

[0124] The image reconstruction module 503 inputs the context image features of the target video frame into a pre-trained diffusion model to obtain the super-resolution video frame of the target video frame.

[0125] The video super-resolution device based on spatio-temporal feature aggregation proposed by the present invention can accurately align video frames at different times through optical flow calculation between multiple frames, thereby reducing the influence of spatio-temporal misalignment during feature aggregation and making the reconstructed video more continuous. Secondly, multi-scale spatio-temporal feature fusion is performed based on three adjacent video frames to fully fuse spatio-temporal information and generate the context image features of adjacent frames, so as to improve spatio-temporal consistency and visual consistency between video frames, and effectively reduce the artifact phenomenon of the diffusion model in fast-moving or complex scenes.

[0126] Figure 6 shows the internal structure diagram of a computer device in an embodiment of the present invention. The computer device can specifically be a terminal or a system. For example Figure 6As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement each step in the above method embodiments. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute each step in the above method embodiments. Those skilled in the art can understand, Figure 6 The structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0127] In one embodiment, a computer device is proposed, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes each step in the above method embodiments.

[0128] In one embodiment, a computer-readable storage medium is proposed, storing a computer program. When the computer program is executed by the processor, the processor executes each step in the above method embodiments.

[0129] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0131] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A video super-resolution method based on spatio-temporal feature aggregation, characterized in that, The method includes: Obtain a target video frame in the video to be processed, as well as the previous video frame and the next video frame adjacent to the target video frame, where the target video frame is any frame in the video to be processed; Perform optical flow calculation based on the previous video frame, the target video frame, and the next video frame to obtain the optical flow between two adjacent video frames; Perform inverse warping and multi-scale spatio-temporal feature fusion according to the optical flow between two adjacent video frames, the previous video frame, the target video frame, and the next video frame to obtain the context image features of the target video frame after feature aggregation; Input the context image features of the target video frame into a pre-trained diffusion model to obtain the super-resolution video frame of the target video frame.

2. The method according to claim 1, wherein The performing optical flow calculation based on the previous video frame, the target video frame, and the next video frame to obtain the optical flow between two adjacent video frames specifically includes: Extract the image features of the previous video frame, the target video frame, and the next video frame; Perform optical flow calculation according to the image features of the previous video frame and the target video frame to obtain the first optical flow; Perform optical flow calculation according to the image features of the target video frame and the next video frame to obtain the second optical flow.

3. The method according to claim 2, characterized in that, The performing inverse warping and multi-scale spatio-temporal feature fusion according to the optical flow between two adjacent video frames, the previous video frame, the target video frame, and the next video frame to obtain the context image features of the target video frame after feature aggregation specifically includes: Perform a pixel-level inverse warping operation according to the image features of the previous video frame and the first optical flow to obtain the first sub-image features; Perform a pixel-level inverse warping operation according to the image features of the next video frame and the second optical flow to obtain the second sub-image features; Perform multi-scale spatio-temporal feature fusion on the first sub-image features, the second sub-image features, and the image features of the target video frame to obtain the context image features of the target video frame after feature aggregation.

4. The method according to claim 3, characterized in that, The performing multi-scale spatio-temporal feature fusion on the first sub-image features, the second sub-image features, and the image features of the target video frame to obtain the context image features of the target video frame after feature aggregation specifically includes: Input the first sub-image features and the second sub-image features into a preset aggregation module to obtain the aggregated features output by the aggregation module; Use a preset channel attention fusion module to perform feature fusion on the aggregated features and the image features of the target video frame to obtain the context image features of the target video frame.

5. The method according to claim 1, wherein The inputting the context image features of the target video frame into a pre-trained diffusion model to obtain the super-resolution video frame of the target video frame specifically includes: Use a low-rank adaptation model and the context image features to adjust the weights in the diffusion model to obtain an adjusted target diffusion model; Input the context image features of the target video frame into the target diffusion model to obtain the super-resolution video frame output by the target diffusion model.

6. The method according to claim 5, wherein The low-rank adaptation model includes a degradation evaluation network, a KAN layer, and an independent ID assignment strategy. Adjusting the weights in the diffusion model using the low-rank adaptation model and the context image features to obtain an adjusted target diffusion model specifically includes: Input the context image features into the degradation evaluation network for two-dimensional vector conversion and Gaussian Fourier conversion to obtain degradation features. Use the independent ID assignment strategy to set a unique number embedding layer for each target network block in the diffusion model, where the target network block is a network block whose preset weights need to be adjusted. Input the number embedding layer and the degradation features into the KAN layer to generate a fine-tuning matrix for the target network block. Use the fine-tuning matrix of the target network block to adjust the weights of the target network block in the diffusion model to obtain an adjusted target diffusion model.

7. The method according to claim 6, wherein The degradation features are obtained based on the following formula: f d = concat[sin(2πdM T ), cossin(2πdM T )] where f d is a degenerate feature, concat represents concatenation along a specified dimension, d is a two-dimensional vector generated based on the context feature, and M is a randomly initialized matrix.

8. The method according to claim 6, wherein Using the fine-tuning matrix of the target network block to adjust the weights of the target network block in the diffusion model to obtain an adjusted target diffusion model specifically includes: Obtain the original weights of the target network block. Calculate the adjusted weights of the target network block based on the original weights, the fine-tuning matrix of the target network block, and a preset low-rank matrix. Substitute the adjusted weights into the diffusion model to obtain an adjusted target diffusion model.

9. The method according to claim 1, wherein The method further includes: Based on adaptive instance normalization, perform image color restoration on the super-resolution video of the target video frame and the target video frame to obtain a color-restored super-resolution video.

10. A video super-resolution device based on spatio-temporal feature aggregation, characterized in that, The device includes: a feature extraction module, a multi-scale spatio-temporal feature aggregation module, and an image reconstruction module. The feature extraction module is used to obtain the target video frame in the video to be processed, as well as the previous video frame and the next video frame adjacent to the target video frame, where the target video frame is any frame in the video to be processed. The multi-scale spatio-temporal feature aggregation module is used to perform optical flow calculation based on the previous video frame, the target video frame, and the next video frame to obtain the optical flow between two adjacent video frames. Perform inverse warping and multi-scale spatio-temporal feature fusion based on the optical flow between two adjacent video frames, the previous video frame, the target video frame, and the next video frame to obtain the context image features of the target video frame after feature aggregation. The image reconstruction module inputs the context image features of the target video frame into a pre-trained diffusion model to obtain the super-resolution video frame of the target video frame.