Real-world video super-resolution method, system, device and medium

Through the biaxial spatiotemporal attention mechanism module and two-dimensional convolution reconstruction technology, the problem of insufficient reconstruction quality and robustness in real-world video super-resolution tasks is solved, and efficient high-quality video output is achieved.

CN120339063AActive Publication Date: 2025-07-18SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510193801.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-18
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The prior art in real-world video super-resolution tasks has low video reconstruction quality, insufficient robustness, and high computational cost, especially when processing low-quality and degraded videos.

Method used

The biaxial spatiotemporal attention mechanism module is adopted, combining vertical-time and horizontal-time attention blocks, and feature capture capabilities are enhanced through rotating position coding, and space-time reconstruction is carried out using two-dimensional convolution, and combined with pre-training-fine-tuning strategies to adapt to different types of low-quality video inputs.

Benefits of technology

High-quality video reconstruction is achieved, improving robustness and reducing computing costs, and able to generate stable high-resolution video output in complex real-world scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339063A_ABST
    Figure CN120339063A_ABST
Patent Text Reader

Abstract

The invention relates to a real-world video super-resolution method, system and device and a medium. The method comprises the following steps: inputting an original video sequence; video embedding: extracting a spatial feature of each frame from the original video sequence to obtain a first feature, and inputting the first feature to a biaxial space-time attention mechanism module to obtain a second feature; wherein the double-axis space-time attention mechanism module comprises a vertical-time attention block and a horizontal-time attention block, a feature block generated after first feature fragmentation processing is converted into a token sequence through rotation position coding, the token sequence is rearranged and then is sent into the vertical-time attention block and the horizontal-time attention block, and the vertical-time attention block and the horizontal-time attention block are sent to the double-axis space-time attention mechanism module; simulating spatial texture and motion characteristics; and time-space reconstruction: integrating time information by adopting time attention, and reconstructing and generating video output with higher time-space quality. Compared with the prior art, the method has the advantages of good video reconstruction quality, high robustness and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video image processing, and in particular, to a real-world video super-resolution method, system, device and medium. Background Art

[0002] The real-world video super-resolution task aims to recover high-resolution results from degraded video inputs that are of low quality and may contain noise, blur, and compression artifacts. However, directly applying the ViViT architecture to the real-world VSR task does not yield good results. Different from high-level vision tasks, low-level vision tasks require pixel-level accuracy and detail consistency, and thus are more vulnerable to information loss. The compression patching, tokenization, and sequential spatio-temporal attention processes in ViViT may lead to significant detail loss, weakening the model's ability to generate high-quality reconstructions.

[0003] To address these issues, ViT-based image and video super-resolution models have introduced specific adjustments, such as smaller patch sizes, CNN-based pixel refinement, and local attention mechanisms, and combined with optical flow-based feature alignment and fusion modules. By using a cyclic structure design and auxiliary modules, CNN-based RealBasicVSR and transformer-based RealViformer have achieved significantly better results than ViViT-VSR. However, the above solutions still face the following limitations: local attention limits the receptive field and scalability, and relying on optical flow for alignment often leads to cumulative errors under long-term sequences or realistic degradation conditions, significantly affecting performance. Summary of the Invention

[0004] The purpose of the present invention is to provide a real-world video super-resolution method, system, device and medium with good video reconstruction quality, high robustness and low cost, in order to overcome the defects of the above-mentioned existing technologies.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] According to a first aspect of the present invention, there is provided a real-world video super-resolution method, including:

[0007] Input an original video sequence;

[0008] Video embedding: Extract spatial features of each frame from the original video sequence to obtain a first feature, and input the first feature into a bi-axial spatio-temporal attention mechanism module to obtain a second feature; wherein, the bi-axial spatio-temporal attention mechanism module includes a vertical-temporal attention block and a horizontal-temporal attention block, and the feature blocks generated by slicing the first feature through rotational position encoding are converted into token sequences, and after the token sequences are rearranged, they are respectively sent into the vertical-temporal attention block and the horizontal-temporal attention block to simulate spatial texture and motion characteristics;

[0009] Spatio-temporal reconstruction: Based on the second feature and the original video sequence, time information is integrated using temporal attention to reconstruct and generate a video output with higher spatio-temporal quality.

[0010] Preferably, extracting the spatial features of each frame from the original video sequence specifically includes: using two-dimensional convolution to extract the spatial features of each frame from the original video sequence, and using a temporal attention mechanism to perform temporal continuity enhancement operations.

[0011] Preferably, the feature blocks generated after the first feature is sliced are specifically: slicing the first feature into 1×1 along the time, height, and width dimensions.

[0012] Preferably, the rotational position encoding is set by multiplying a bias in the complex vector space. For the u-th query vector q u and the v-th key vector k v , the rotational position encoding expression is:

[0013] f q (q u , u) = e iuΘ q u ,

[0014] f k (k v , v) = e ivΘ k v ,

[0015] In the formula: Θ is a diagonal matrix containing elements θ d = b -2d / |D| , b is the rotation base, D is the embedding dimension of the attention block; where the score A v corresponding to the attention block is the inner product of the real parts of two complex vectors f q and f k .

[0016] Preferably, in the biaxial spatio-temporal attention mechanism module, after rearranging the token sequence, it is respectively fed into the vertical-temporal attention block and the horizontal-temporal attention block to simulate spatial texture and motion characteristics, specifically including:

[0017] First, perform token embedding along the height-time dimension. After rearranging the token dimensions from [B D n N n H n W to [(BD n W ) n H n N , B is the batch size, D is the embedding dimension of the attention block, n N , n H , nW Corresponding to the number of blocks in the time, height, and width directions respectively, it is fed into the vertical-time attention block VTAB to calculate self-attention on the vertical-time plane for spatial texture and vertical motion modeling;

[0018] Then, token embedding is performed along the width-time dimension, and the token dimension is rearranged from [(B D n W )n H n N to [(B D n H )n W n N . After that, it is fed into the horizontal-time attention block HTAB to calculate self-attention on the horizontal-time plane for spatial texture and horizontal motion modeling.

[0019] Preferably, based on the second feature and the original video sequence, temporal attention is used to integrate temporal information to reconstruct and generate a video output with higher spatio-temporal quality, which specifically includes:

[0020] Using temporal attention to integrate temporal information;

[0021] Using two-dimensional convolution and pixel shuffling operations for frame-by-frame reconstruction to generate a video output with higher spatio-temporal quality.

[0022] Preferably, the biaxial spatio-temporal attention mechanism module is trained using a pre-training - fine-tuning strategy.

[0023] According to the second aspect of the present invention, a system adopting the real-world video super-resolution method described above is provided, including:

[0024] An input module for inputting the original video sequence;

[0025] A video embedding module for extracting the spatial features of each frame from the original video sequence to obtain the first feature, and inputting the first feature into the biaxial spatio-temporal attention mechanism module to obtain the second feature; wherein, the biaxial spatio-temporal attention mechanism module includes a vertical-time attention block and a horizontal-time attention block, and the feature blocks generated by slicing the first feature through rotational position encoding are converted into a token sequence, and after rearranging the token sequence, it is respectively fed into the vertical-time attention block and the horizontal-time attention block to simulate spatial texture and motion characteristics;

[0026] A spatio-temporal reconstruction module for the user to integrate temporal information based on the second feature and the original video sequence to reconstruct and generate a video output with higher spatio-temporal quality.

[0027] According to a third aspect of the present invention, there is provided an electronic device, including a memory and a processor, wherein a computer program is stored on the memory, and when the processor executes the program, any one of the above methods is implemented.

[0028] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, any one of the above methods is implemented.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] (1) High video reconstruction quality: The dual-axis spatio-temporal attention mechanism module constructed by the present invention realizes the synchronous utilization of spatial features and motion characteristics by simultaneously performing video embedding along the height-time and width-time dimensions, and applying vertical-time attention blocks and horizontal-time attention blocks, enhancing the ability to capture details and dynamic changes, and enabling more accurate and high-quality video reconstruction effects.

[0031] (2) Higher robustness: The present invention further enhances the ability to capture long-range dependencies by performing rotational position encoding on the features extracted from the original video sequence, ensuring spatial and temporal consistency, and enabling high-quality video super-resolution task outputs in complex real-world degradation scenarios; at the same time, a pre-training-fine-tuning strategy is adopted to stably output high-quality results under different types of low-quality video inputs.

[0032] (3) Low-cost and high-quality video super-resolution output: In the spatio-temporal reconstruction process, the present invention improves the traditional frame-by-frame processing method, enabling the simultaneous fusion of spatial and temporal information during the reconstruction process. Compared with three-dimensional convolution, using two-dimensional convolution for reconstruction not only reduces the computational cost but also incurs almost no performance loss, that is, while maintaining a low computational overhead, it generates video outputs with higher spatio-temporal quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic flow diagram of the real-world video super-resolution method of the present invention;

[0034] Figure 2 It is a schematic diagram of the dual-axis spatio-temporal dual-axis attention mechanism;

[0035] Figure 3 It is a schematic diagram showing the structural differences between the model of the present invention and the existing models. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] Embodiment

[0038] As Figure 1 shown, in this embodiment, a two-axis spatio-temporal transformer model for real-world video super-resolution tasks is constructed, aiming to improve the resolution quality of videos. The input low-resolution video is converted into a high-resolution output through a series of processing steps. The specific implementation steps of the real-world video super-resolution method include:

[0039] S1. Input the original video sequence, and the original low-resolution video sequence is represented as where i represents the frame index, B is the batch size, C is the number of channels, N is the total number of frames, and H and W are the height and width of each frame respectively.

[0040] S2. Video embedding: Extract the spatial features of each frame from the original video sequence to obtain the first feature, and input the first feature into the two-axis spatio-temporal attention mechanism module to obtain the second feature.

[0041] In this embodiment, the two-axis spatio-temporal attention mechanism module includes a vertical-time attention block and a horizontal-time attention block. The feature blocks generated by slicing the first feature through rotated position encoding are converted into a token sequence. After rearranging the token sequence, it is respectively sent into the vertical-time attention block and the horizontal-time attention block to simulate spatial texture and motion characteristics. The specific process includes:

[0042] S21. Apply two-dimensional convolution to each frame to capture its spatial features and obtain shallow features d represents the dimension of the shallow features.

[0043] S22. Perform a 1×1 slicing operation on the shallow features, and enhance the temporal continuity through the temporal attention mechanism to divide the shallow features into blocks where n N , n H , n W correspond to the number of blocks in the time, height, and width directions respectively, and n, h, and w represent the block sizes in each dimension.

[0044] S23. Perform token embedding along the height-time dimension, and rearrange the token dimensions from [B D n N n H nW is changed to [(BD n W )n H n N . After that, B is the batch size, D is the embedding dimension of the attention block, and n N ,n H ,n W correspond to the number of blocks in the time, height, and width directions respectively, and are fed into the vertical-time attention block VTAB for calculating self-attention on the vertical-time plane and performing spatial texture and vertical motion modeling;

[0045] Specifically, the rotated position encoding RoPE is used to embed the sharded feature blocks into the tokens where D is the embedding dimension of the attention block. For the u-th query vector q u and the v-th key vector k v , RoPE multiplies the bias in the complex vector space, and the formula is f q (q u ,u) = e iuΘ q u ,f k (k v ,v) = e ivΘ k v , where Θ is a diagonal matrix containing the element θ d = b -2d / |D| , and the rotation base b = 10000. The attention score is obtained by calculating the inner product of the real parts of two complex vectors: A v = Re < f q (q u ,u), f k (k v ,v) >. The advantage of RoPE is that it can better handle long-range dependencies and maintain spatial and temporal consistency, which is very important for video super-resolution tasks.

[0046] S24. Perform token embedding along the width-time dimension, and rearrange the token dimension from [(B D n W )n H n N to [(B D n H )n W n N . After that, it is fed into the horizontal-time attention block HTAB for calculating self-attention on the horizontal-time plane and performing spatial texture and horizontal motion modeling.

[0047] In this embodiment, the specific settings of the vertical-time attention block VTAB and the horizontal-time attention block HTAB are as follows: The input token sequence A1 passes through a convolutional layer and a multi-head dot product attention block in sequence to generate a feature A2. After the features A1 and A2 are fused, a feature A3 is obtained. The feature A3 passes through layer normalization processing and a multi-layer perceptron in sequence to obtain a feature A4. The features A3 and A4 are fused to obtain a feature A5, which is used as the output of the vertical-time attention block VTAB / horizontal-time attention block HTAB.

[0048] Through the above settings, the model can effectively utilize spatial features and motion characteristics in both the vertical and horizontal directions, enhancing the ability to capture details and dynamic changes.

[0049] S3. Spatiotemporal reconstruction: Based on the second feature and the original video sequence, temporal attention is used to integrate temporal information to reconstruct and generate a video output with higher spatiotemporal quality. The output high-resolution video is denoted as where s represents the magnification ratio, usually set to 4 times.

[0050] Specifically, in the reconstruction process after decoding and de-sharding, temporal attention is first applied to integrate temporal information, and then operations such as two-dimensional convolution and pixel shuffling are used for frame-by-frame reconstruction. Experiments show that 3D convolution does not bring performance improvement compared to 2D convolution, but significantly increases the computational cost. Therefore, 2D convolution is selected for reconstruction. This approach enables the model to combine spatial and temporal information, thereby generating a video output with higher spatiotemporal quality.

[0051] To make the model adapt to the degradation situation in the real world, a pre-training and fine-tuning strategy is adopted. The entire architecture is designed based on ViViT, including an input video embedding and a spatiotemporal reconstruction module. A dual-axis spatiotemporal attention mechanism is particularly introduced to achieve enhanced fusion and modeling of spatiotemporal information.

[0052] One of the core points of the present invention is the first proposed dual-axis spatiotemporal attention mechanism (Dual Axial Spatial X Temporal Attention Mechanism), as Figure 2 shown:

[0053] When re - examining the attention mechanisms in video Transformers, it is noted that different types of attention mechanisms have their own advantages and disadvantages. Vision Transformer (ViT) captures static spatial information by dividing an image into non - overlapping patches and applying self - attention on them. However, in video tasks, single - spatial attention cannot handle the temporal dependencies between frames. ViViT introduces temporal attention to capture the dynamic relationships between frames, calculating self - attention by concatenating tokens across time at the same spatial location, which helps capture motion patterns but may lead to loss of spatial details within a frame. Spatiotemporal attention that combines spatial and temporal information enables the model to both focus on key regions within a frame and understand the inter - frame dynamics.

[0054] Existing methods usually use these two types of attention sequentially or alternately, but still process spatial and temporal information separately, limiting the ability to capture complex spatiotemporal patterns. Local window attention is widely used in video super - resolution models (such as RVRT, PSRT, IART), adopting the Swin Transformer architecture, calculating attention within local windows to improve low - level task performance. However, this mechanism is limited by a finite receptive field and poor scalability, making it difficult to capture long - range dependencies. To study the impact of different attention mechanisms on video super - resolution (VSR) performance, this embodiment extends the ViViT architecture to the VSR task and conducts experiments on the REDS dataset. The results show that using only spatial or temporal attention performs poorly, with temporal attention being slightly better because it can utilize inter - frame similarities to make up for the lack of spatial details.

[0055] The biaxial spatiotemporal attention mechanism of the present invention combines vertical - temporal and horizontal - temporal attention, effectively integrating spatial and temporal information. The input video is segmented into a series of tokens, each token representing a spatiotemporal region. These tokens are first embedded along the height - temporal dimension and then fed into the vertical - temporal attention block (VTAB) to model spatial texture and vertical motion. Subsequently, they are embedded along the width - temporal dimension, and the horizontal - temporal attention block (HTAB) is applied to capture spatial texture and horizontal motion. This method enables the model to simultaneously utilize spatial features and motion characteristics in both vertical and horizontal directions, enhancing the ability to capture details and dynamic changes. Experiments show that this method not only improves parameter efficiency and running speed but also makes significant progress in the performance of video super - resolution tasks.

[0056] Table 1

[0057]

[0058] This embodiment also provides a real - world video super - resolution system, which includes:

[0059] An input module for inputting an original video sequence;

[0060] A video embedding module for extracting spatial features of each frame from the original video sequence to obtain a first feature, and inputting the first feature into a dual-axis spatio-temporal attention mechanism module to obtain a second feature; wherein, the dual-axis spatio-temporal attention mechanism module includes a vertical-temporal attention block and a horizontal-temporal attention block, and converts the feature blocks generated by slicing the first feature through rotated position encoding into a token sequence, rearranges the token sequence and then sends it into the vertical-temporal attention block and the horizontal-temporal attention block respectively to simulate spatial texture and motion characteristics;

[0061] A spatio-temporal reconstruction module, which uses temporal attention to integrate temporal information based on the second feature and the original video sequence, and reconstructs and generates a video output with higher spatio-temporal quality.

[0062] In summary, through in-depth analysis of the core spatial and temporal attention mechanisms of video Transformers, it can be seen that existing spatio-temporal attention mechanisms usually process spatial and temporal information independently and sequentially, and this sequential method cannot fully integrate spatio-temporal information, restricting coherent video representation. Based on this, the present invention proposes a new dual-axial spatial and temporal Transformer for real-world video super-resolution (Dual Axial SpatialXTemporalTransformer for Real-World Video Super-Resolution, DualX-VSR).

[0063] Different from traditional spatio-temporal attention mechanisms, the present invention simultaneously processes vertical-temporal and horizontal-temporal attention, projects spatio-temporal information along orthogonal directions to achieve integrated modeling of spatial and temporal information, which avoids sequential stacking of spatial and temporal modules and provides a more unified and coherent representation. The effectiveness of this method has been verified in real-world video super-resolution tasks. As Figure 3 shown, DualX-VSR retains the simplicity of the ViViT architecture, does not add additional auxiliary modules, and exhibits superior performance compared to current video super-resolution methods. By eliminating explicit motion estimation and long-range feature propagation, DualX-VSR provides a robust solution to the unique challenges of real-world video super-resolution and sets a new standard for high-quality video restoration in complex real-world scenarios.

[0064] The electronic device of the present invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0065] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a magnetic disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0066] The processing unit executes the various methods and processes described above, such as method S1 - S3. For example, in some embodiments, method S1 - S3 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of method S1 - S3 described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute method S1 - S3 by any other suitable means (e.g., by means of firmware).

[0067] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip systems (SOC), complex programmable logic devices (CPLD), and so on.

[0068] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0069] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0070] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A real-world video super-resolution method, characterized in that, Including: Input the original video sequence; Video embedding: Extract the spatial features of each frame from the original video sequence to obtain the first feature, and input the first feature into the biaxial spatio-temporal attention mechanism module to obtain the second feature; wherein, the biaxial spatio-temporal attention mechanism module includes a vertical-temporal attention block and a horizontal-temporal attention block, and the feature blocks generated by slicing the first feature through rotational position encoding are converted into token sequences, and after rearranging the token sequences, they are respectively sent into the vertical-temporal attention block and the horizontal-temporal attention block to simulate spatial texture and motion characteristics; Spatio-temporal reconstruction: Based on the second feature and the original video sequence, use temporal attention to integrate temporal information and reconstruct to generate a video output with higher spatio-temporal quality.

2. The real-world video super-resolution method according to claim 1, wherein The extraction of the spatial features of each frame from the original video sequence is specifically: Use two-dimensional convolution to extract the spatial features of each frame from the original video sequence, and use a temporal attention mechanism to perform a temporal continuity enhancement operation.

3. A real-world video super-resolution method according to claim 1, characterized in that, The feature blocks generated after slicing the first feature are specifically: The first feature is sliced into 1×1 along the temporal, height, and width dimensions.

4. A real-world video super-resolution method according to claim 1, characterized in that The rotation position encoding is a bias setting multiplied in the complex vector space for the \(u\)-th query vector \(\mathbf{q}\) u and the \(v\)-th key vector \(\mathbf{k}\) v , and the rotation position encoding expression is as follows: f q (q u ,u) = e iuΘ q u , f k (k v , v) = e ivΘ k v , where: Θ is a diagonal matrix containing the elements θ d = b -2d / |D| , b is the rotation base, D is the embedding dimension of the attention block; wherein, the score A corresponding to the attention block v is the inner product of the real parts of two complex vectors f q and f k .

5. A real-world video super-resolution method according to claim 1, characterized in that In the biaxial spatio-temporal attention mechanism module, after rearranging the token sequences, they are respectively sent into the vertical-temporal attention block and the horizontal-temporal attention block to simulate spatial texture and motion characteristics, which specifically includes: First, perform token embedding along the height-time dimension, and rearrange the token dimensions from [B D n N n H n W to [(B Dn W )n H n n . Here, B is the batch size, D is the embedding dimension of the attention block, and n N ,n H ,n W correspond to the number of blocks in the time, height, and width directions respectively. Then, send them into the vertical-time attention block VTAB to calculate self-attention on the vertical-time plane for spatial texture and vertical motion modeling; Then, perform token embedding along the width-time dimension, and rearrange the token dimension from [(B D n W )n H n N to [(B Dn H )n W n N . After that, send it into the horizontal-time attention block HTAB for calculating self-attention on the horizontal-time plane to perform spatial texture and horizontal motion modeling.

6. A real-world video super-resolution method according to claim 1, characterized in that Based on the second feature and the original video sequence, using temporal attention to integrate temporal information and reconstruct to generate a video output with higher spatio-temporal quality, which specifically includes: Use temporal attention to integrate temporal information; Use two-dimensional convolution and pixel shuffling operations for frame-by-frame reconstruction to generate a video output with higher spatio-temporal quality.

7. A real-world video super-resolution method according to claim 1, characterized in that The biaxial spatio-temporal attention mechanism module is trained using a pre-training - fine-tuning strategy.

8. A system adopting the real-world video super-resolution method according to claim 1, characterized in that, Including: An input module for inputting the original video sequence; A video embedding module for extracting the spatial features of each frame from the original video sequence to obtain the first feature, and inputting the first feature into the biaxial spatio-temporal attention mechanism module to obtain the second feature; wherein, the biaxial spatio-temporal attention mechanism module includes a vertical-temporal attention block and a horizontal-temporal attention block, and the feature blocks generated by slicing the first feature through rotational position encoding are converted into token sequences, and after rearranging the token sequences, they are respectively sent into the vertical-temporal attention block and the horizontal-temporal attention block to simulate spatial texture and motion characteristics; A spatio-temporal reconstruction module for the user to, based on the second feature and the original video sequence, use temporal attention to integrate temporal information and reconstruct to generate a video output with higher spatio-temporal quality.

9. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image denoising method, system and device and storage medium

    CN116012266A

  • Image super-resolution method and device based on cross attention mechanism and Swin-Transform

    CN117237197A

  • Monitoring video real-time super-resolution reconstruction method and system fused with space-time attention mechanism

    CN119130805A

  • Video super-resolution network, and video super-resolution, encoding and decoding processing method and device

    WO2023000179A1