Random-scale video super-resolution method based on diffusion model
Through implicit neural networks based on diffusion model and spatiotemporal information fusion technology, the problems of insufficient amplification of arbitrary scale and poor high-frequency details recovery in the prior art are solved, and high-quality and consistent arbitrary scale video super-resolution is achieved.
Patent Information
- Application Number
- CN202510154703.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-12
AI Technical Summary
Existing video super-resolution technology is difficult to adapt to amplification at any scale, insufficient recovery of high-frequency details, and easy to have artifact problems, resulting in blurred picture or unreal details.
The arbitrary scale video super-resolution method based on the diffusion model is adopted, and the upsampling of any scale is realized through an implicit neural network, combined with the spatial and temporal information fusion mechanism, the consistency and consistency between frames are improved, and image details are gradually denoised through a multi-step reverse generation process.
It realizes super-resolution support for videos of any scale, improves the visual quality and visual continuity of the video, reduces artifacts, and produces high-resolution videos with rich and consistent details.
Smart Images

Figure CN120088133A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video super-resolution, and particularly relates to a method for video super-resolution at any scale based on a diffusion model. Background Art
[0002] In recent years, as an emerging generative model, the diffusion model has achieved remarkable breakthroughs in image generation and restoration tasks. The core idea of the diffusion model is to gradually transform data from a structured distribution into noise and then restore the original data through an inverse process. This approach has natural advantages in image super-resolution because it can finely model the detail restoration process while maintaining the global consistency of the image.
[0003] In the video super-resolution task, more and more people have begun to notice the potential of diffusion models. However, the randomness of diffusion models makes it impossible to directly control the continuity between video frames. Upscale-A-Video inserts a temporal layer into the pre-trained image diffusion model to enhance the modeling of the dependencies between video frames, optimize the local consistency within short segments, and introduce spatio-temporal 3D residual blocks and spatial feature transform layers (SFT) in the VAE decoder to reduce flicker artifacts and color offsets during the decoding process. In addition, it uses flow-guided recurrent latent code propagation, which requires no additional training. Through optical flow estimation (such as RAFT), it realizes the forward and backward propagation and fusion of latent space features, and only selects regions with smaller optical flow consistency errors for propagation, thereby enhancing the overall coherence of long videos. By fine-tuning the U-Net network and the VAE Decoder, an additional temporal layer is introduced in the U-Net network to achieve local consistency constraints within video segments, and an additional spatio-temporal 3D residual block is introduced in the VAE Decoder to enhance low-level consistency. In addition, a flow-guided recurrent propagation module that requires no training is introduced in the latent space to ensure the global spatio-temporal consistency of long-sequence videos. SATeCo proposes a novel framework to calibrate the denoising and reconstruction process of high-resolution videos by learning the spatio-temporal features of low-resolution videos. Its core modules include a spatial feature adaptation module (SFA) and a temporal feature alignment module (TFA). The SFA adjusts the frame features by adaptively estimating the affine parameters of each pixel to ensure pixel-level guidance for high-resolution frame synthesis. The TFA uses self-attention and cross-attention mechanisms to perform feature interactions within a 3D local window (tubelet) and align them with their low-resolution counterparts to maintain temporal consistency between frames. SATeCo freezes all the parameters of the pre-trained UNet and VAE during the training process and only optimizes the SFA and TFA modules, thereby reducing the computational cost and improving efficiency. Using the spatial adaptation module in the diffusion model, it learns the spatial adaptability between high-resolution video frames and their corresponding low-resolution video frames, prompting the generated high-resolution videos to be rich in details and consistent with the corresponding content of the input low-resolution videos. After the spatial adaptation module, a temporal alignment module is further used to learn the temporal coherence between high-resolution video frames, ensuring the consistency of image detail content between different frames of the generated high-resolution videos.
[0004] The above methods have achieved good performance, but there are still some deficiencies in arbitrary-scale super-resolution:
[0005] (1) Difficulty in adapting to arbitrary magnification: Current video super-resolution technologies are usually trained and optimized at fixed magnification factors (such as ×2, ×4). However, at any scale, it is difficult to maintain consistent generation results. Especially when the magnification factor does not match the training magnification factor, the generated video often appears blurred, with details lost or the picture distorted, making it difficult to achieve high-quality dynamic enhancement effects. This dependence on fixed magnification factors limits the flexibility of the super-resolution model, making it difficult for existing technologies to meet the video magnification requirements in different scenarios.
[0006] (2) High-frequency detail loss and artifact problems: In video super-resolution magnification, especially at large magnification factors, existing methods are insufficient in restoring high-frequency details and are prone to detail loss or artifact problems, resulting in a blurred or unrealistic detailed generated picture. In scenarios with multi-scale changes, it is even more difficult for the super-resolution model to accurately restore the tiny details and complex textures in the video, thus affecting the overall visual effect and authenticity of the video.
[0007] The above problems need to be solved urgently. Therefore, an arbitrary-scale video super-resolution method based on a diffusion model is proposed. Summary of the Invention
[0008] The technical problem to be solved by the present invention is: how to solve the problems in the existing technology, such as difficulty in adapting to arbitrary magnification, high-frequency detail loss, and artifacts, and provides an arbitrary-scale video super-resolution method based on a diffusion model.
[0009] The present invention solves the above technical problems through the following technical solutions. The present invention includes the following steps:
[0010] S1: Input of low-resolution video frames
[0011] Obtain the input low-resolution video sequence L and the preset magnification factor s or the target resolution.
[0012] S2: Interpolation and optical flow information calculation
[0013] Perform interpolation processing on the low-resolution video sequence L, interpolate it to the target resolution size, and obtain a rough high-resolution video frame sequence. Use a pre-trained optical flow network of to calculate the optical flow information between adjacent frames and generate an optical flow map O.
[0014] S3: Generation of latent feature representation
[0015] Encode the rough high-resolution video frames to generate a latent feature representation z.
[0016] S4: Generation of initial input
[0017] Generate a noise of the same size as the latent feature representation z obtained according to the encoding, and initialize an initial input x according to the latent feature representation z and the noise T , and send the latent feature representation z into the guidance conditional network for processing to generate a conditional input; then combine the conditional input and the initial input x T and send them into the U-Net network of the diffusion model together;
[0018] S5: Spatiotemporal information fusion
[0019] In the encoder part of the U-Net network, the features input in step S4 are processed according to the original structure to extract multi-scale noisy features; in the decoder part of the U-Net network, the multi-scale noisy features are sent into the implicit neural upsampling module after being processed by residual blocks and Transformer blocks, and spatiotemporal information fusion processing is performed in the implicit neural upsampling module;
[0020] S6: Implicit neural upsampling processing
[0021] Generate a coordinate grid coord of the corresponding size according to the size of the multi-scale noisy features in the encoder part of the U-Net network, scale both the features after spatiotemporal information fusion processing and the size of the coordinate grid coord to the interval from -1 to 1, find the corresponding position in the feature map f of the features after spatiotemporal information fusion processing for each point in the coordinate grid and query the feature values, then scale back to the original size, and send the obtained feature map f new after querying into the implicit neural function layer parameterized by a multi-layer perceptron to obtain the output f query , and then connect it with the multi-scale noisy features output by the encoder and send it into the next-level module in the decoder part of the U-Net network; out
[0022] S7: Optical flow alignment processing
[0023] After the U-Net network finishes processing, send the output features into the optical flow alignment module, and use the optical flow map calculated in step S2 to perform curling processing on the features to align the features between adjacent frames;
[0024] S8: Generate high-resolution video frames and output
[0025] Send the features after optical flow alignment processing through a T-time iterative denoising process to gradually eliminate noise and restore the details of high-resolution video frames. The features f denoise obtained after the denoising process is decoded by the decoder of the diffusion model into a high-resolution video sequence H.
[0026] Furthermore, in the step S2, the rough high-resolution video frame sequence The calculation formula is as follows:
[0027]
[0028] Among them, "↑" represents interpolation processing.
[0029] Furthermore, in the step S2, the calculation formula of the optical flow map is as follows:
[0030]
[0031] Furthermore, in the step S3, the specific processing process is as follows:
[0032] Map the features of each frame to the latent space through the pre-trained encoder ε to generate the latent feature representation z:
[0033]
[0034] Furthermore, in the step S4, the initial input x T The generation formula is as follows:
[0035]
[0036] Among them, are learnable parameters and are optimized during the training process.
[0037] Furthermore, in the step S5, the spatio-temporal information fusion processing process is as follows:
[0038] S51: The input features are first unfolded according to the frame channels, and each frame is processed cyclically;
[0039] S52: For the features f i of each frame, obtain the features of its adjacent frames and concatenate them together according to the channel dimension to get the feature f c(i) ;
[0040] S53: Through the channel compression operation, compress the concatenated features back to the original channel size, so as to obtain the new feature f new(i) fusing spatio-temporal information, and then obtain the feature map f new .
[0041] Furthermore, in the step S52, the generation formula of the feature f c(i) is as follows:
[0042] f c(i) = concate(f i-1 , f i , f i+1 );
[0043] Among them, fi-1 , f i , f i+1 respectively represent the target frame feature and its adjacent front and rear frame features, and concate represents concatenation along the channel dimension.
[0044] Furthermore, in the step S53, the feature f new(i) is generated by the following formula:
[0045] f new(i) = Conv(f c(i) );
[0046] where Conv represents the convolutional layer processing operation.
[0047] Furthermore, in the step S6, the feature map f query and the output f out are generated by the following formula:
[0048] f query = Grid_Sample(f new , coord);
[0049] f out = INF(f query );
[0050] where Coord represents the two-dimensional coordinate grid of the generated target size, Grid_Sample represents finding the feature value at the corresponding position in f new for each coordinate in the two-dimensional coordinate grid, and INF represents the implicit neural upsampling module formed by the implicit neural function layer parameterized by the multi-layer perceptron, which processes the features through the multi-layer perceptron.
[0051] Furthermore, in the step S8, the high-resolution video sequence H is generated by the following formula:
[0052] H = D(f denoise );
[0053] where D represents the decoder.
[0054] The present invention has the following advantages compared with the prior art:
[0055] 1. Adapt to upsampling of any scale
[0056] The present invention uses an implicit neural network to replace the traditional convolutional upsampling module, making the system have stronger scale adaptability; the implicit neural network can achieve upsampling of any scale without relying on specific-scale convolutions through continuous spatial representations; this continuous feature expression method ensures that high-resolution results can be smoothly generated under different resolution requirements.
[0057] The introduction of the implicit neural network improves the adaptability of the upsampling module, enabling this method to be applicable to video super-resolution tasks of any scale without the need to redesign or adjust the model structure. This not only enhances the flexibility of the model but also reduces the dependence on training data of specific scales. In practical applications, this advantage can meet the video processing tasks with different resolution requirements and is particularly suitable for application scenarios with different device display requirements.
[0058] 2. Enhancement of spatio-temporal consistency
[0059] The present invention introduces an implicit neural upsampling module based on spatio-temporal information fusion. While generating the target frame, it utilizes the information of adjacent frames. With the spatio-temporal fusion mechanism, the motion information of adjacent frames is deeply fused with the target frame; this spatio-temporal fusion mechanism effectively extracts the dynamic features between frames, ensuring the coherence and consistency between frames of the generated high-resolution video.
[0060] Through the deep fusion of spatio-temporal information, the present invention reduces the jitter and blurring phenomena between frames during the video generation process, achieving smooth transitions in the video. This method is particularly suitable for video processing with fast motion or complex scenes, significantly enhancing the visual quality and visual continuity of the video, and ensuring no obvious visual breaks during the playback of the super-resolution video. This improvement can enhance the user viewing experience in practical applications and has important application value especially in streaming media playback and virtual reality scenarios.
[0061] 3. Efficient diffusion model generation process
[0062] The present invention adopts a diffusion model for the generation of high-resolution frames. The diffusion model generates images through a multi-step reverse process, gradually denoising and restoring image details at each step. This approach is more stable and reliable compared to single-step generation methods. The multi-step iterative reverse generation process of the diffusion model can effectively improve the accuracy of detail restoration and reduce the appearance of artifacts.
[0063] The step-by-step generation process of the diffusion model enables the noise in the low-resolution input frame to be gradually removed, thereby generating high-resolution frames with high details and low artifacts. This method is more stable in terms of generation effect and has significant advantages when processing low-resolution videos with noise or compression artifacts. Compared with traditional methods, the diffusion model solution of the present invention can significantly improve the overall quality of videos in practical video applications, making the generated videos more refined in terms of detail performance and suitable for the requirements of high-quality video playback and data storage. Brief description of the drawings
[0064] Figure 1 It is a schematic flowchart of the method for arbitrary-scale video super-resolution based on the diffusion model in the second embodiment of the present invention;
[0065] Figure 2 It is a schematic diagram of the overall architecture of the diffusion model in the third embodiment of the present invention;
[0066] Figure 3 It is a graph of integer multiple results of applying the method of the present invention in the third embodiment of the present invention, where the 2x super-resolution magnification result, the 4x super-resolution magnification result, and the 8x super-resolution magnification result;
[0067] Figure 4 It is a graph of non-integer multiple results of applying the method of the present invention in the third embodiment of the present invention, where the 2.5x super-resolution magnification result, the 3.2x super-resolution magnification result, and the 4.1x super-resolution magnification result. Detailed implementation manners
[0068] The following details the embodiments of the present invention. The present embodiment is implemented on the premise of the technical solution of the present invention, and provides detailed implementation manners and specific operation processes. However, the protection scope of the present invention is not limited to the following embodiments.
[0069] Embodiment 1
[0070] This embodiment provides a technical solution: a method for arbitrary-scale video super-resolution based on a diffusion model, and the specific content is as follows:
[0071] Task objectives of arbitrary-scale video super-resolution
[0072] Video super-resolution aims to generate a corresponding high-resolution video sequence based on the input low-resolution video sequence. There are many existing video super-resolution methods, such as methods based on Transformer, methods based on convolution, methods based on diffusion models, etc. However, most of these methods are restricted by fixed multiple magnifications, which greatly limits their use in real scenarios. In the present invention, we propose a diffusion model for arbitrary-scale video super-resolution. Specifically, we use an implicit neural upsampling module to process the features of the input video sequence, generate features with different resolutions at different levels in the U-Net decoder, connect them with the corresponding features in the encoder, and send them to the next level. Spatiotemporal fusion sampling fuses the information of the current frame with the information of its adjacent frames, and can focus on the information outside the current frame during sampling, effectively solving the problem of inconsistency between adjacent frames in video super-resolution, achieving super-resolution of arbitrary scale multiples, and generating high-quality videos with arbitrary resolutions. In this embodiment, the above objectives are specifically achieved through the following steps:
[0073] Implicit Neural Upsampling: To enable the U-Net network to process features of arbitrary resolution sizes and ensure smooth connection between the corresponding hierarchical outputs of the encoder and decoder, we designed an implicit neural upsampling module. During the upsampling of each level in the U-Net decoder, we use a coordinate-based approach to generate a grid of the size of the output features of the corresponding level in the encoder. Both the grid and the features are scaled to the two-dimensional coordinate range from -1 to 1. Each point in the grid locates the corresponding position in the features according to the coordinates, and calculates the feature value of this point according to the feature value of the feature map at that position in the nearest or interpolation manner and returns it. After all grid points are queried, they are scaled back to the original size and fed into the implicit neural function layer parameterized by a multi-layer perceptron. The result is connected to the corresponding output of the encoder and fed into the next layer of the decoder.
[0074] Spatio-temporal Information Fusion of Adjacent Frames: During the above-mentioned grid sampling process, since the query range of each point is only within the target feature, the useful information contained in its adjacent frames is ignored, and this method is prone to problems of inconsistency before and after. To solve this problem, we adopt the method of spatio-temporal fusion of adjacent frames. When feeding the features into the implicit neural upsampling module, we process the features, expand them along the frame dimension, extract the adjacent frames before and after for each frame and connect them along the channel dimension, and then process the channel dimension to finally obtain the spatio-temporal information fusion features with the same channel size as the target frame. After spatio-temporal fusion processing, the above sampling process is performed using this feature map.
[0075] Embodiment 2
[0076] In this embodiment, the video super-resolution method in Embodiment 1 is further described. This video super-resolution method replaces the upsampling module in the traditional U-Net network with an implicit neural network and adds a spatio-temporal information fusion mechanism in this module, realizing video super-resolution at any scale. The method mainly includes the following steps: input a low-resolution video frame sequence, gradually restore the high-resolution frames through a diffusion model, and at the same time add the spatio-temporal fusion of the target frame and adjacent frames in the implicit neural upsampling module to improve the continuity and consistency of the generated video frames.
[0077] The diffusion model is a method of gradually generating data and is used in the present invention to gradually restore a high-resolution video from a low-resolution video. The diffusion model restores the initial low-resolution frame to the target high-resolution frame through a multi-step reverse process. In each step, the model generates an intermediate frame with a higher resolution according to the degree of noise reduction, and finally reaches the target clarity and resolution.
[0078] In the traditional U-Net network, the upsampling module is usually implemented through convolution operations and cannot handle resolution changes at arbitrary scales well. The present invention uses an implicit neural network as the upsampling module, enabling the generation network to handle resolution improvement at arbitrary scales. The implicit neural network uses a continuous feature representation method, making upsampling no longer dependent on convolution operations at specific scales. It has strong scale adaptability and flexibility and can generate more detailed high-resolution frames.
[0079] To improve the continuity and consistency of the generated video frames, the present invention introduces a spatio-temporal fusion method based on the target frame and adjacent frames in the implicit neural upsampling module. This spatio-temporal fusion mechanism deeply fuses the spatio-temporal features of the target frame and adjacent frames to capture the motion information and correlation between frames.
[0080] Specifically, when feeding the features into the implicit neural upsampling module, we process the features, expand them along the frame dimension, extract the adjacent frames before and after each frame and concatenate them along the channel dimension, and then process the channel dimension to finally obtain the spatio-temporal information fusion features with the same channel size as the target frame. The high-resolution frames generated in this way are not only clear but also maintain high continuity in dynamic videos.
[0081] As Figure 1 shown, the implementation process of the specific technical solution of the present invention includes the following main steps, in sequence: input of low-resolution video frames, processing by the diffusion model, spatio-temporal information fusion, processing by the implicit neural upsampling module, generation of high-resolution video frames, and output. The following is a detailed description of each step:
[0082] Step 1: First, obtain the input low-resolution video sequence L = {L 1 , L 2 , …, L n} and the preset magnification factor s or the target resolution (the resolution of the target output). This video sequence L contains each low-resolution frame, and the information of these frames will be used for feature extraction and magnification processing later. The preset magnification factor s or the target resolution can flexibly meet user needs to generate a super-resolution video output that meets specific display requirements.
[0083] Step 2: First, perform interpolation processing on the low-resolution video sequence L to interpolate it to the target resolution size, thereby obtaining a rough high-resolution video frame sequence The rough high-resolution frame sequence obtained in this step is not the final output, but provides a preliminary high-resolution reference for subsequent processing. Next, a pre-trained optical flow network of is used to calculate the optical flow information between adjacent frames, generating optical flow maps O. These optical flow maps can effectively capture the motion information between frames and provide important dynamic guidance for the subsequent diffusion process. The interpolation process and the calculation formula of the optical flow map O are as follows:
[0084]
[0085] Step 3: Encode the rough high-resolution video frames. Specifically, the features of each frame are mapped into the latent space through a pre-trained encoder ε to generate a latent feature representation z. The feature representation in the latent space can compress the information of the original video frames and provide a more efficient feature description for the subsequent generation process. Through this encoding, the redundant information of the original video data can be greatly reduced while retaining the key information required to generate high-resolution videos. The generation formula of the latent feature representation z is as follows:
[0086]
[0087] Step 4: According to the latent feature representation z obtained by encoding, generate a noise noise of the same size, initialize an initial input x based on z and noise T , and send the latent feature representation z into a guidance conditional network (the guidance conditional network is similar to the encoder structure of U-Net, which processes the input features step by step and performs downsampling operations, aiming to ensure that the connection process in the decoder part of the U-Net network will not be unable to connect due to size mismatch) for processing to generate a conditional input. This conditional input is feature auxiliary information used to guide the denoising process, helping the network gradually restore video details during the denoising process. Then, send the conditional input and the initial input x T into the U-Net network together. The generation formula of the initial input x T is as follows:
[0088]
[0089] where are learnable parameters and are optimized during the training process.
[0090] Step 5: In the encoder part of the U-Net network, the input features are processed according to the original structure to extract multi-scale noisy features. In the decoder part, the features are processed through residual blocks and Transformer blocks and then sent into an implicit neural upsampling module. In the implicit neural upsampling module, the input features are first unfolded according to the frame channels and each frame is processed in a loop. For each frame feature f i, obtain its features with adjacent frames and concatenate them together along the channel dimension to obtain the feature f c(i) , then, through the channel compression operation, compress the concatenated features back to the original channel size, so as to obtain the new feature f that fuses spatio-temporal information new(i) . Such a processing method can effectively fuse the inter-frame information in the decoding stage and improve the coherence between frames. The feature f c(i) and the feature f new(i) are generated as follows:
[0091] f c(i) = concate(f i-1 , f i , f i+1 );
[0092] f new(i) = Conv(f c(i) );
[0093] Among them, in the first formula above, the features in the brackets represent the target frame feature and its adjacent frame features before and after, concate represents concatenation along the channel dimension; Conv represents the convolutional layer processing operation.
[0094] Step 6: Based on the multi-scale noisy features generated in the encoder part of the U-Net network, the present invention further introduces the generation of the coordinate grid coord. Generate the coordinate grid coord of the corresponding size according to the size of the multi-scale noisy features in the corresponding encoder part of the U-Net network, scale the processed features (feature f new(i) ) and the coordinate grid size to the interval from -1 to 1, find the corresponding position in the feature map f new for each point position in the coordinate grid and query the feature value, then scale back to the original size, and send the obtained feature map f query into the implicit neural function layer parameterized by the multi-layer perceptron to obtain the output f out , and then connect it with the output of the encoder and send it into the next-level module of the decoder. The specific processing process is as follows:
[0095] f query = Grid_Sample(f new , coord);
[0096] f out = INF(f query );
[0097] Among them, Coord represents the two-dimensional coordinate grid of the generated target size, and Grid_Sample represents finding f at each coordinate in the two-dimensional coordinate grid newThe eigenvalue at the corresponding position, INF represents an implicit neural upsampling module formed by an implicit neural function layer parameterized by a multi-layer perceptron, and the features are processed by the multi-layer perceptron.
[0098] Step 7: After all the processing of the U-Net network is completed, the noisy feature f out will be fed into the optical flow alignment module (SpyNet). This module uses the optical flow map calculated in Step 2 to warp the features, aligning the features between adjacent frames, thereby further enhancing the spatio-temporal consistency between frames. The optical flow alignment operation ensures the coherence of features between frames, effectively reducing inconsistent phenomena such as jitter and blur in video frames, making the generated high-resolution video smoother and more natural.
[0099] Step 8: The features after optical flow alignment are passed through a denoising process with T iterations to gradually eliminate noise, thereby restoring the details of high-resolution video frames. The denoising process further eliminates noise and restores details in each iteration. After the final denoising process is completed, the resulting feature f denoise will be decoded by the decoder into a high-resolution video sequence H:
[0100] H = D(f denoise );
[0101] where D represents the decoder.
[0102] Embodiment 3
[0103] In recent years, some works have also been studying how to achieve super-resolution at arbitrary scales. Although these methods meet the requirements of achieving super-resolution at arbitrary scales in terms of results, it is usually difficult to achieve high-fidelity details at high magnification factors because these regression-based methods tend to calculate the average result of possible super-resolution predictions due to the regression loss. Since the emergence of the Diffusion model, it has shown convincing performance in generating high-quality images with high-fidelity details and has also achieved good results in the field of video super-resolution. Nevertheless, due to the fixed-factor upsampling and downsampling of the U-Net network, the Diffusion model-based methods are still limited by the fixed magnification factor. For tasks with different magnification factors, the model needs to be retrained or a complex cascade structure needs to be adopted, resulting in additional training costs.
[0104] To solve this problem, an innovative implicit expression super-resolution diffusion model is proposed in this embodiment to achieve arbitrary-scale video super-resolution with high-quality details. We introduce the implicit neural representation function (i.e., the implicit neural function layer in Embodiment 1) into the diffusion model to solve the fixed-scale limitation while retaining the advantages of the diffusion model in generating higher-quality details. Specifically, as Figure 2As shown, we first calculate the optical flow between adjacent frames and encode the video frames into a continuous latent feature space. To enable the representation of images at arbitrary size resolutions, we adopt an implicit neural representation function in the decoder structure of the U-Net network and use a coordinate-based multi-layer perceptron (MLP) to parameterize it. Before performing the upsampling step, we perform a spatio-temporal information fusion operation on the input features. Overall, the model alternately uses a diffusion model and an implicit neural representation function. To retain more information from the low-resolution video frames, we fuse the original features obtained by encoding the low-resolution video frames through an adjustment network with the features obtained at each level of the encoder part of the U-Net network during the diffusion process. At the same time, to enable the diffusion model to obtain outputs of different resolution sizes, we introduce a magnification factor as an input to the network. The magnification factor is encoded as a feature vector and used as the weight for fusing the original features and the noise processed by the U-Net network. After each denoising process is completed, the features are warped using the initially obtained optical flow to obtain new features as the input for the next denoising. Finally, the decoder decodes the features to obtain high-resolution video frames. The experimental results of our method are as Figure 3 , Figure 4 shown.
[0105] In summary, the arbitrary-scale video super-resolution method based on the diffusion model in the above embodiments can effectively solve the deficiencies of existing video super-resolution methods in terms of arbitrary scales through spatio-temporal feature fusion and arbitrary-scale implicit neural upsampling. It can not only solve the limitations of existing methods that can only perform integer multiples and single multiples, but also generate high-resolution video sequences of arbitrary sizes. At the same time, it can ensure the consistency between video sequences, enrich the application scenarios of video super-resolution, and expand the technical boundaries of this field.
[0106] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations of the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for super-resolution of arbitrary-scale videos based on a diffusion model, characterized in that: The following steps are involved: S1: Low-resolution video frame input Obtain an input low-resolution video sequence L and a preset magnification s or target resolution; S2: Interpolation and optical flow information calculation Interpolate the low-resolution video sequence L to the target resolution to obtain a rough high-resolution video frame sequence Use a pre-trained optical flow network of to calculate the optical flow information between adjacent frames and generate an optical flow map O; S3: Latent feature representation generation Encode the rough high-resolution video frame to generate a latent feature representation z; S4: Initial input generation According to the encoded latent feature representation z, a noise of the same size is generated, and an initial input x is initialized according to the latent feature representation z and the noise noise T , and feed the latent feature representation z into the guided conditional network to generate the conditional input; then the conditional input and the initial input x T They are sent together to the U-Net network of the diffusion model; S5: Spatiotemporal information fusion In the encoder part of the U-Net network, the features input in step S4 are processed according to the original structure to extract multi-scale noisy features; In the decoder part of the U-Net network, multi-scale noisy features are processed by residual blocks and Transformer blocks and then sent to the implicit neural upsampling module, where spatiotemporal information fusion is performed; S6: Implicit Neural Upsampling Processing According to the multi-scale noisy feature size of the corresponding encoder part of the U-Net network, a coordinate grid coord of the corresponding size is generated. The size of the feature after spatiotemporal information fusion processing and the coordinate grid coord are scaled to the range of -1 to 1. The feature map f after spatiotemporal information fusion processing is found for each point in the coordinate grid. new The corresponding position in the query is queried and the feature value is then scaled back to the original size. The feature map f query It is sent to the implicit neural function layer parameterized by the multi-layer perceptron to get the output f out , and then connected with the multi-scale noisy features output by the encoder and sent to the next level module of the decoder part of the U-Net network; S7: Optical flow alignment processing After the U-Net network is processed, the output features are sent to the optical flow alignment module, and the features are warped using the optical flow map calculated in step S2 to align the features between adjacent frames; S8: Generate high-resolution video frames and output The features processed by optical flow alignment are subjected to T-iteration denoising processes to gradually eliminate noise and restore high-resolution video frame details. After the denoising process, the feature f denoise Decoded by the diffusion model decoder into a high-resolution video sequence H.
2. The method for super-resolution of arbitrary-scale videos based on a diffusion model according to claim 1, characterized in that: In step S2, a rough high-resolution video frame sequence The calculation formula is as follows: Among them, "↑" indicates interpolation processing.
3. The method for super-resolution of arbitrary-scale videos based on a diffusion model according to claim 1, characterized in that: In step S2, the calculation formula of the optical flow map is as follows:
4. The method for super-resolution of arbitrary-scale videos based on a diffusion model according to claim 1, characterized in that: In step S3, the specific processing process is as follows: The features of each frame are mapped to the latent space through the pre-trained encoder ε to generate the latent feature representation z:
5. The method for super-resolution of arbitrary-scale videos based on a diffusion model according to claim 1, characterized in that: In step S4, the initial input x T The generation formula is as follows: in, is a learnable parameter that is optimized during training.
6. The method for super-resolution of arbitrary-scale videos based on a diffusion model according to claim 1, characterized in that: In step S5, the spatiotemporal information fusion processing process is as follows: S51: The input features are first expanded according to the frame channel, and each frame is processed in a loop; S52: Feature f for each frame i , obtain the features of the adjacent frames and connect them together according to the channel dimension to obtain the feature f c(i) ; S53: Through the channel compression operation, the spliced features are compressed back to the original channel size, thereby obtaining a new feature f that integrates spatiotemporal information new(i) , and then get the feature map f new .
7. The method for super-resolution of arbitrary-scale videos based on a diffusion model according to claim 6, characterized in that: In step S52, feature f c(i) The generation formula is as follows: f c(i) =concate(f i-1 ,f i ,f i+1 ); Among them, f i-1 、f i 、f i+1 They represent the target frame features and the features of the frames before and after them respectively, and concate means connecting by channel dimension.
8. The method for super-resolution of arbitrary-scale videos based on a diffusion model according to claim 7, characterized in that: In step S53, feature f new(i) The generation formula is as follows: f new(i) =Conv(f c(i) ); Among them, Conv represents the convolutional layer processing operation.
9. The method for super-resolution of arbitrary-scale videos based on a diffusion model according to claim 1, characterized in that: In step S6, the feature map f query And the output f out The generation formula is as follows: f query =Grid_Sample(f new ,coord); in out =INF(f query ); Among them, Coord represents the generated two-dimensional coordinate grid of the target size, and Grid_Sample represents the search for f at each coordinate in the two-dimensional coordinate grid. new INF represents the implicit neural upsampling module formed by the implicit neural function layer parameterized by the multi-layer perceptron, and the features are processed by the multi-layer perceptron.
10. The method for super-resolution of arbitrary-scale videos based on a diffusion model according to claim 1, characterized in that: In step S8, the high-resolution video sequence H is generated by the following formula: H=D(f denoise ); Wherein, D represents a decoder.
Citation Information
Patent Citations
Video super-resolution reconstruction method and system based on multi-scale local self-attention
CN115082308A
Multi-attention-based video super-resolution reconstruction network construction method and application thereof
CN116993585A
Double-branch image high-multiple super-resolution enhancement method based on diffusion model
CN118864245A
Remote sensing image super-resolution reconstruction method and system based on prior diffusion model
CN119251054A
High-resolution video generation using image diffusion models
US20240171788A1
Cited By
Video super-resolution method and device for complex operation scene based on space-time consistency and medium
CN121213358A
Historical document repairing method and system based on implicit interpolation network enhancement
CN121258846A
A historical document restoration method and system based on implicit interpolation network enhancement
CN121258846B