Video super-resolution method and system based on cyclic depth estimation guidance

By introducing depth perception and optical flow perception modules into the video super-resolution method and combining it with the inter-frame cyclic alignment structure, the shortcomings of existing methods in dealing with complex motion and stereoscopic objects are solved, and better depth of field information extraction and video reconstruction effects are achieved.

CN120672578APending Publication Date: 2025-09-19ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510759907.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing video super-resolution methods perform poorly when dealing with complex motion and stereoscopic objects, and are unable to effectively extract and utilize depth information in video frames, resulting in occlusion and boundary problems in the reconstructed video.

Method used

Using a method guided by recurrent depth estimation, we design depth-sensing and optical flow-sensing modules, providing additional depth information to guide the network. This information is combined with optical flow to extract spatiotemporal details. By propagating and fusing features and depth maps from previous and subsequent frames through an inter-frame recurrent alignment structure, we fully exploit the temporal information between video frames.

Benefits of technology

It effectively enhances the network's ability to model long-distance dependencies, improves the extraction and utilization of depth information in video frames, and improves the visual effect of super-resolution videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672578A_ABST
    Figure CN120672578A_ABST
Patent Text Reader

Abstract

The invention discloses a video super-resolution method and system based on cyclic depth estimation guidance, and belongs to the technical field of video super-resolution. The method comprises the following steps: performing image enhancement on an original low-resolution video, extracting an original spatial feature of each video frame after image enhancement, performing interpolation on the original spatial features, and extracting a low-level feature map from the interpolated video frame; respectively calculating a depth map and an optical flow map of each video frame according to the low-level feature map; the low-level feature map, the depth map and the optical flow graph are sent to an inter-frame cyclic alignment structure for feature alignment and fusion, and a high-resolution reconstruction result of each video frame is obtained; and traversing all video frames to obtain a super-divided high-resolution video. According to the method, extra depth information is provided as guidance information of the network and is complementary with optical flow information to extract space-time detail information, and time sequence information between video frames is fully mined through propagation and fusion of features and depth maps of front and back frames in a time dimension, so that a super-resolution visual effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video super-resolution, and in particular relates to a video super-resolution method and system based on cyclic depth estimation guidance. Background Art

[0002] Video, a common multimedia format, consists of a series of consecutive images. Super-resolution aims to reconstruct a low-resolution input image or video into a high-resolution output. This task is a classic and challenging one in computer science, with numerous practical applications, such as satellite imagery, remote sensing, digital high-definition imaging, and microscopic imaging.

[0003] Super-resolution input can be broadly categorized into two types: images and videos. Image super-resolution aims to reconstruct a single low-resolution image into a high-resolution result; video super-resolution, on the other hand, takes multiple consecutive frames as input and leverages the relationships between them to restore the high-resolution image. Traditional video super-resolution methods have evolved through methods based on interpolation and motion estimation, but these methods struggle to effectively estimate complex motion in videos. With the widespread application of deep learning in various computing fields, the field of video super-resolution has also seen extensive research on deep learning-based methods. Currently, deep learning-based video super-resolution methods primarily include three architectures: convolutional neural network (CNN)-based methods, which use a sliding window approach to extract effective detail features between adjacent video frames for computation; recurrent neural network (RNN)-based methods, which utilize hidden states to focus on historical information in previous video frames and future information in future video frames for computation; and Transformer-based methods, which utilize a self-attention mechanism to extract both long-range and short-range features for computation.

[0004] Existing methods perform poorly when focusing on objects at different depths in the same scene. They only focus on the "plane" of objects in the scene, but not on the three-dimensional sense of the objects, and do not have a good grasp of the spatial layout of the objects. This is insufficient when extracting spatiotemporal detail information and capturing object motion information, which can easily cause occlusion and boundary problems in the reconstructed video. In addition, existing methods perform well in local spatiotemporal modeling, but are unable to focus on information in distant frames during the two-way propagation of videos, failing to achieve progressive utilization of video information. This leads to deficiencies in the model's global information modeling, which may result in the loss of detailed information and affect the reconstruction effect of long-term time-series videos. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention discloses a video super-resolution method and system based on cyclic depth estimation guidance, proposes a depth perception module and an optical flow perception module, provides additional depth information as guidance information of the network, and complements the optical flow information to extract spatiotemporal detail information, thereby strengthening the long-distance dependence of the network and better extracting and utilizing the depth of field information in the video frames; proposes a cyclic structure, which fully mines the timing information between video frames by propagating and fusing the features of the previous and next frames and the depth map in the time dimension, so that the video after super-resolution has better visual effects.

[0006] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0007] In a first aspect, the present invention proposes a video super-resolution method based on cyclic depth estimation guidance, characterized in that it includes the following steps:

[0008] Perform image enhancement on the original low-resolution video, extract the original spatial features of each video frame after image enhancement, interpolate the original spatial features, and extract low-level feature maps from the interpolated video frames;

[0009] The depth map and optical flow map of each video frame are calculated based on the low-level feature map. When calculating the depth map, the low-level feature map is first downsampled to extract the high-level feature map, and then the high-level feature map is upsampled and connected with the residual of the low-level feature map to obtain the depth map. When calculating the optical flow map, the low-level feature map is sequentially subjected to optical flow prediction, smoothing, and optical flow field optimization to obtain the optical flow map of each video frame.

[0010] The low-level feature map, depth map and optical flow map are sent to the inter-frame cyclic alignment structure for feature alignment and fusion to obtain the high-resolution reconstruction result of each video frame; all video frames are traversed to obtain the high-resolution video after super-resolution.

[0011] Furthermore, image enhancement includes horizontal 90° flipping, vertical 90° flipping, and mirror symmetry.

[0012] Furthermore, the extraction process of low-level feature maps is as follows:

[0013] After image enhancement, each video frame is expanded from 3 to 64 channel dimensions using a convolutional layer while maintaining the same resolution, and then the original spatial features are extracted.

[0014] The original spatial features are interpolated with an upsampling rate of 0.5, and the obtained upsampled spatial features are then convolved to extract the low-level feature map of the image.

[0015] Furthermore, pixel binning and sub-pixel convolution operations are used for downsampling and upsampling, respectively.

[0016] Furthermore, the optical flow prediction is performed on the low-level feature map, which is expressed as:

[0017] f′=Conv(F,W flow )+b flow

[0018] Among them, f ′ Represents the optical flow prediction result, which consists of the horizontal component u(x,y) and the vertical component v(x,y), F represents the low-level feature map, W flow represents the convolution kernel weight of optical flow estimation, b flow represents the bias term.

[0019] Furthermore, the optical flow field is optimized based on the loss function of the photometric consistency loss hypothesis, and the loss function is as follows:

[0020]

[0021] Among them, I t-1 represents the t-1 frame image, I t (x,y) represents image I t where u(x,y) and v(x,y) represent the horizontal and vertical components of the optical flow, respectively, and L represents the loss function based on the photometric consistency loss assumption.

[0022] Furthermore, the calculation process of the inter-frame cyclic alignment structure is as follows:

[0023] For two adjacent video frames, the optical flow map of the latter frame and the low-level feature map F of the previous frame t-1 Perform feature alignment to obtain the alignment space feature F (t-1)→t ; and, the optical flow map of the next frame and the depth map of the previous frame Perform feature alignment to obtain the aligned depth feature D (t-1)→t Similarly, we get the alignment spatial feature F of the t+1th frame aligned to the tth frame (t+1)→t and aligned deep features D (t+1)→t ;

[0024] The low-level feature map F t , alignment space feature F (t-1)→t and F (t+1)→t , Depth Map Align deep features D (t-1)→t and D (t+1)→t Splicing, as the fusion feature of the t-th frame;

[0025] Convolution and sub-pixel convolution are performed on the fused features to obtain the final high-resolution output.

[0026] Furthermore, when processing the fusion features, the convolution operation is first performed to adjust the channel so that its dimension changes from becomes H×W×C; then perform sub-pixel convolution to make its dimension become Finally, the convolution operation is performed to adjust the channel to obtain the final high-resolution output X HR , the dimension of the output feature is 2H×2W×3.

[0027] In a second aspect, the present invention proposes a video super-resolution system based on cyclic depth estimation guidance, which is used to implement the above-mentioned video super-resolution method based on cyclic depth estimation guidance.

[0028] Beneficial effects of the present invention:

[0029] This paper designs a complete network structure for video super-resolution based on cyclic depth estimation guidance. It innovatively introduces a depth perception module and an optical flow perception module to provide the network with additional depth information as guidance, and combines optical flow information to extract spatiotemporal details. This design effectively enhances the network's ability to model long-range dependencies and further improves the extraction and utilization of depth information in video frames. By introducing a cyclic structure, the method can propagate and fuse the features of the previous and next frames and the depth map in the temporal dimension, thereby fully exploiting the temporal information between video frames, ultimately making the super-resolution video show a more outstanding visual effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 4 is a structural block diagram of a video super-resolution system based on cyclic depth estimation guidance adopted in an embodiment of the present invention;

[0031] Figure 2 is an overall flow chart of an embodiment of the present invention;

[0032] Figure 3 is a structural block diagram of a depth perception module in an embodiment of the present invention;

[0033] Figure 4 is a structural block diagram of an optical flow perception module in an embodiment of the present invention;

[0034] Figure 5 4 is a structural block diagram of an inter-frame cyclic alignment module in an embodiment of the present invention. DETAILED DESCRIPTION

[0035] The method of the present invention will be further described below with reference to the accompanying drawings.

[0036] The accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0037] The flowcharts shown in the accompanying drawings are merely exemplary and do not necessarily include all steps. For example, some steps may be decomposed, while some steps may be combined or partially combined, so the actual execution order may change according to actual circumstances.

[0038] The structural block diagram of the video super-resolution method and system based on cyclic depth estimation guidance of the present invention is as follows Figure 1 As shown in Figure 2, six modules are designed in this method: data preprocessing module, spatial feature extraction module, depth perception module, optical flow perception module, and inter-frame cyclic alignment module. The flow chart of the video super-resolution method implemented by each module is shown in Figure 2. Figure 2 shown.

[0039] The data preprocessing module is used to process the input original video data stream, obtain low-resolution video and perform image enhancement on each video frame; specifically, the method in the following step (1) is executed.

[0040] Step (1). Obtain a low-resolution video sequence and perform mirror symmetry, horizontal 90° flip, and vertical 90° flip on each frame to achieve image enhancement. The low-resolution video sequence after image enhancement is recorded as I t , where t represents the t-th frame image; it is then input into the spatial feature extraction module frame by frame.

[0041] The spatial feature extraction module is used to extract the spatial features of each frame image in the low-resolution video and interpolate the original spatial features to obtain the upsampled spatial features of each video frame; specifically, the method in the following step (2) is executed.

[0042] Step (2). Take the t(t∈[1,T])th frame image I of the low-resolution video obtained in step (1) t The convolutional layer is used to change the channel dimension of the input frame from 3 to 64, and then the original spatial feature extraction is performed.

[0043] Step (2.1) The spatial feature extraction process is expressed as:

[0044] X t =R(RELU(conv(I t )))

[0045] Among them, Xt Represents the t-th frame image I t The original spatial features of the network are obtained by multiplying the spatial features by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 6

[0046] Interpolate the spatial features of each frame of the low-resolution video with an upsampling rate of 0.5:

[0047] X t ′=Bilinear(X t ,0.5)

[0048] Among them, X t ′ Represents the t-th frame image X t The upsampled spatial features of the t-th frame image X t 0.5 times, Bilinear(.,.) represents the bilinear interpolation method. At this time, the feature expression after upsampling of each frame is: in And X t ′ The channel dimension is X t 4 times the channel dimension, that is C represents the number of channels of the original frame image.

[0049] Step (2.2). Upsample the spatial feature X t ′ Perform convolution operation to extract the low-level feature map F of the image:

[0050] F=Convs(X t ′)

[0051] Among them, Convs represents the stacked convolution operation layer.

[0052] A core design point of the present invention is to use the depth perception module to calculate the depth map of each frame in the low-resolution video, such as Figure 3 As shown, the depth perception module is used to perform convolution operations and up- and down-sampling operations on the upsampled spatial features to obtain a depth map of each video frame; specifically, the method in the following step (3) is executed.

[0053] Step (3). For the low-level feature map F obtained in step (2), calculate the depth map of each video frame separately, specifically:

[0054] Step (3.1). For the low-level feature map F obtained in step (2), use the Pixel Unshuffle operation to downsample it to obtain the downsampled feature map Reduce the resolution of the feature map and extract the main details:

[0055] F out =PixelUnshuffle(F)

[0056] Here, Pixel Unshuffle is a spatial downsampling operation that reduces the spatial resolution of the input feature map while increasing the number of channels to retain more local information.

[0057] Step (3.2). For the downsampled feature map F obtained in step (3.1) out , perform convolution operation on it to further extract detail features and obtain high-level feature map F ′ :

[0058] F′=Convs(F out )

[0059] Among them, Convs represents the stacked convolution operation layer

[0060] Step (3.3). For the high-level feature map F obtained in step (3.2) ′ , use Pixel Shuffle (sub-pixel convolution) operation to upsample the high-level feature map and perform residual connection with the low-level feature map F obtained in step (2.2) to obtain the depth map F out′ :

[0061] F iut′ =PixelShuffle(F ′ )+F

[0062] Here, Pixel Shuffle is an efficient spatial upsampling method that achieves resolution improvement through channel reorganization. Pixel Unshuffle and Pixel Shuffle are inverse operations of each other.

[0063] like Figure 4 As shown, the optical flow perception module is used to obtain the optical flow map of each video frame through optical flow prediction, smoothing and optical flow field optimization for the low-level feature map; specifically, the method in the following step (4) is executed.

[0064] Step (4). For the low-level feature map F obtained in step (2), calculate the optical flow map of each video frame separately, specifically:

[0065] Step (4.1). Perform optical flow prediction on the low-level feature map F obtained in step (2) to obtain the optical flow prediction result f for each pixel. ′ , including horizontal and vertical components:

[0066] f′=Conv(F,W flow )+b flow

[0067] Among them, W flow represents the convolution kernel weight of optical flow estimation, b flow represents the bias term. ′ Contains the horizontal and vertical components of each pixel:

[0068] f ′ =[u(x,y),v(x,y)]

[0069] Where u(x,y) represents the horizontal displacement at the pixel position (x,y), and v(x,y) represents the vertical displacement at the pixel position (x,y).

[0070] Step (4.2). For the optical flow prediction result f obtained in step (4.1) ′ , use the filter to perform smoothing and obtain the filtered optical flow field f smooth :

[0071] f smooth =Fliter(f′)

[0072] Where Fliter(·) represents a smoothing operation.

[0073] Step (4.3). For the optical flow field f obtained in step (4.2) smooth , use the loss function based on the photometric consistency loss assumption to optimize the optical flow field and obtain the optical flow map Measures the ability of optical flow estimation to reconstruct the input:

[0074]

[0075] Among them, I t-1 Indicates I t The previous frame image, I t (x,y) represents image I t where u(x,y) and v(x,y) represent the horizontal and vertical components of the optical flow, respectively, and L represents the loss function based on the photometric consistency loss assumption.

[0076] The inter-frame cyclic alignment module is used to obtain high-resolution reconstruction results for each video frame by propagating and fusing the features of the previous and next frames, as well as the depth map and optical flow map, in the time dimension; it traverses all video frames to obtain a high-resolution video after super-resolution. Figure 5 As shown, specifically perform the method in the following step (5).

[0077] Step (5). Input the low-level feature map obtained in step (2), the depth map obtained in step (3), and the optical flow map obtained in step (4) into the inter-frame cyclic alignment module to perform feature alignment, feature fusion, feature reconstruction and upsampling to obtain a high-resolution super-resolution result, such as Figure 4 As shown (for ease of understanding, Figure 4 The right figure only shows the t-1th frame and the tth frame). Specifically:

[0078] Step (5.1). For two adjacent video frames, the optical flow map of the latter frame obtained in step (4.3) And the low-level feature map F of the previous frame obtained in step (2) t-1 Perform feature alignment to obtain the alignment space feature F (t-1)→t :

[0079]

[0080] Among them, Warp(·,·) represents the feature alignment operation.

[0081] Similarly, we get the alignment spatial feature F of the t+1th frame aligned to the tth frame (t+1)→t :

[0082]

[0083] Step (5.2). For the two adjacent video frames described in step (5.1), the optical flow map of the latter frame obtained in step (4.3) is And the depth map of the previous frame obtained in step (3.3) Perform feature alignment to obtain the aligned depth feature D (t-1)→t :

[0084]

[0085] Similarly, we get the aligned depth feature D of the t+1th frame aligned to the tth frame (t+1)→t :

[0086]

[0087] Step (5.3). For the two adjacent video frames described in step (5.1), the low-level feature map F of the latter frame obtained in step (3.1) is t, the depth map of the next frame obtained in step (3.3) The alignment space feature F obtained in step (5.1) (t-1)→t and the alignment spatial feature F of the t+1th frame aligned to the tth frame (t+1)→t , the aligned depth feature D obtained in step (5.2) (t-1)→t and the aligned depth feature D of the t+1th frame aligned to the tth frame (t+1)→t Perform fusion to obtain the fusion features of the next frame

[0088]

[0089] Step (5.4). Use 3×3 convolution to adjust the channel of the fusion feature so that its dimension changes from It becomes H×W×C to facilitate subsequent high-resolution upsampling operations; then PixelShuffle is used to perform sub-pixel convolution, making its dimension become Finally, 3×3 convolution is used to adjust the channels to obtain the final high-resolution output X HR , the dimension of the output feature is 2H×2W×3:

[0090]

[0091] Among them, Conv 3×3 (·) represents a 3×3 convolution operation, PixelShuffle(·,4) represents a sub-pixel convolution operation, and 4 represents a 4x reduction in the channel size, expanding the channel dimension to the spatial resolution. This achieves a 4x increase in resolution. In this embodiment, 4x can also be replaced by other multiples of perfect square numbers, such as 9x.

[0092] Step (5.5). Traverse all video frames to obtain a high-resolution result after the entire video resolution is increased by 4 times.

[0093] The combination of the data preprocessing module, spatial feature extraction module, depth perception module, optical flow perception module, and inter-frame cyclic alignment module described above can construct a complete video super-resolution method and system based on cyclic depth estimation guidance. Each module may or may not be physically separate, may be located in one place, or may be distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention. Those skilled in the art can understand and implement it without any creative work.

[0094] The video super-resolution method and system based on cyclic depth estimation guidance can be applied to any device with data processing capabilities, such as a computer or other device. The system embodiment can be implemented through software, hardware, or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory and executing them.

[0095] The above description is merely a specific embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the concept of this application. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A video super-resolution method based on cyclic depth estimation guidance, characterized in that The following steps are involved: Perform image enhancement on the original low-resolution video, extract the original spatial features of each video frame after image enhancement, interpolate the original spatial features, and extract low-level feature maps from the interpolated video frames; The depth map and optical flow map of each video frame are calculated based on the low-level feature map. When calculating the depth map, the low-level feature map is first downsampled to extract the high-level feature map, and then the high-level feature map is upsampled and connected with the residual of the low-level feature map to obtain the depth map. When calculating the optical flow map, the low-level feature map is sequentially subjected to optical flow prediction, smoothing, and optical flow field optimization to obtain the optical flow map of each video frame. The low-level feature map, depth map and optical flow map are sent to the inter-frame cyclic alignment structure for feature alignment and fusion to obtain the high-resolution reconstruction result of each video frame; all video frames are traversed to obtain the high-resolution video after super-resolution.

2. The video super-resolution method based on cyclic depth estimation guidance according to claim 1, characterized in that Image enhancement includes horizontal 90° flip, vertical 90° flip, and mirror symmetry.

3. The video super-resolution method based on cyclic depth estimation guidance according to claim 1, characterized in that The extraction process of low-level feature maps is as follows: After image enhancement, each video frame is expanded from 3 to 64 channel dimensions using a convolutional layer while maintaining the same resolution, and then the original spatial features are extracted. The original spatial features are interpolated with an upsampling rate of 0.5, and the obtained upsampled spatial features are then convolved to extract the low-level feature map of the image.

4. The video super-resolution method based on cyclic depth estimation guidance according to claim 1, characterized in that Pixel rebinning and sub-pixel convolution are used for downsampling and upsampling respectively.

5. The video super-resolution method based on cyclic depth estimation guidance according to claim 1, characterized in that Optical flow prediction is performed on the low-level feature map, expressed as: f′=Conv(F,W flow )+b flow Among them, f ′ Represents the optical flow prediction result, which consists of the horizontal component u(x,y) and the vertical component v(x,y), F represents the low-level feature map, W flow represents the convolution kernel weight of optical flow estimation, b flow represents the bias term.

6. The video super-resolution method based on cyclic depth estimation guidance according to claim 1, characterized in that The optical flow field is optimized based on the loss function of the photometric consistency loss hypothesis. The loss function is as follows: Among them, I t-1 represents the t-1 frame image, I t (x,y) represents image I t where u(x,y) and v(x,y) represent the horizontal and vertical components of the optical flow, respectively, and L represents the loss function based on the photometric consistency loss assumption.

7. The video super-resolution method based on cyclic depth estimation guidance according to claim 1, characterized in that The calculation process of the inter-frame cyclic alignment structure is as follows: For two adjacent video frames, the optical flow map of the latter frame and the low-level feature map F of the previous frame t-1 Perform feature alignment to obtain the alignment space feature F (t-1)→t ; and, the optical flow map of the next frame and the depth map of the previous frame Perform feature alignment to obtain the aligned depth feature D (t-1)→t Similarly, we get the alignment spatial feature F of the t+1th frame aligned to the tth frame (t+1)→t and aligned deep features D (t+1)→t ; The low-level feature map F t , alignment space feature F (t-1)→t and F (t+1)→t , Depth Map Align deep features D (t-1)→t and D (t+1)→t Splicing, as the fusion feature of the t-th frame; Convolution and sub-pixel convolution are performed on the fused features to obtain the final high-resolution output.

8. The video super-resolution method based on cyclic depth estimation guidance according to claim 7, characterized in that: When processing the fusion features, first perform a convolution operation to adjust the channel so that its dimension changes from becomes H×W×C; then perform sub-pixel convolution to make its dimension become Finally, the convolution operation is performed to adjust the channel to obtain the final high-resolution output X HR , the dimension of the output feature is 2H×2W×3.

9. A video super-resolution system based on cyclic depth estimation guidance, used to implement the video super-resolution method according to claim 1, characterized in that: The system comprises: An image preprocessing module, which is used to perform image enhancement on the original low-resolution video; A spatial feature extraction module is used to extract the original spatial features of each video frame after image enhancement and interpolate the original spatial features, and extract low-level feature maps from the interpolated video frames; The depth perception module is used to calculate the depth map of each video frame based on the low-level feature map. When calculating the depth map, the low-level feature map is first downsampled to extract the high-level feature map, and then the high-level feature map is upsampled and concatenated with the low-level feature map residual to obtain the depth map. The optical flow perception module is used to calculate the optical flow map of each video frame based on the low-level feature map. When calculating the optical flow map, the low-level feature map is sequentially subjected to optical flow prediction, smoothing, and optical flow field optimization to obtain the optical flow map of each video frame. The inter-frame cyclic alignment module is used to obtain the high-resolution reconstruction result of each video frame by propagating and fusing the low-level feature maps, depth maps and optical flow maps of the previous and next frames in the temporal dimension; it traverses all video frames to obtain the high-resolution video after super-resolution.