Video super-resolution device based on multiple time differences and dynamic optimization framework

The multi-time difference and dynamic optimization framework addresses the challenge of complex motion patterns in video super-resolution by efficiently capturing and aligning motion information, enhancing video clarity and accuracy.

CN120318071APending Publication Date: 2025-07-15INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510408945.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing video super-resolution methods face challenges in effectively handling complex and diverse motion patterns in videos, particularly in satellite and low-light environments, due to high computational costs and limitations in time and spatial modeling, leading to noise interference and motion blur.

Method used

A video super-resolution method utilizing a multi-time difference and dynamic optimization framework, comprising a total model with modules for multi-level temporal difference analysis, time-difference-guided dynamic path optimization, multi-attention enhancement and correction, and reconstruction, to capture and align motion information across varying scales and speeds.

Benefits of technology

The method efficiently captures and aligns motion information across multiple scales and speeds, reducing computational costs while maintaining temporal and spatial consistency, resulting in clearer and more accurate video reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318071A_ABST
    Figure CN120318071A_ABST
Patent Text Reader

Abstract

The invention discloses a video super-resolution device based on multiple time differences and a dynamic optimization framework. Comprising a total model; the total model comprises a multi-stage time difference analysis mechanism module, a time difference guided dynamic path optimization module, a multi-attention enhancement and correction module and a reconstruction module; capturing motion features at different time intervals by calculating time differences among adjacent frames, jump frames and cross frames; a motion area is positioned according to the time difference extracted by the multi-stage time difference analysis mechanism module; dynamically selecting different processing paths according to different inter-frame intervals, namely motion scales, and obtaining multi-scale characteristics of different paths; information fusion and enhancement of time, space and channel dimensions are carried out on the output features of the dynamic path optimization module guided by the time difference, and finally, super-resolution reconstruction is carried out on the output features of the multi-attention enhancement and correction module to form a target frame. According to the method, target alignment of different motion scales is optimized, alignment errors are reduced, and spatial information performance is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video processing. The present invention relates to a video super-resolution device based on multi-temporal difference and dynamic optimization framework, which can be applied to fields such as satellite remote sensing video and natural scene video for super-resolution reconstruction and recognition of moving targets. Background Art

[0002] With the continuous development of video processing technology, video super-resolution has become an important research direction for improving video quality. Compared with single-frame image super-resolution, video super-resolution not only utilizes spatial information but also can improve spatio-temporal consistency and enhance the detail restoration effect through the temporal correlation between multiple frames, and has wide application value in multiple fields such as video enhancement, target detection, and intelligent monitoring.

[0003] However, complex shooting environments and hardware limitations may lead to problems such as noise, compression distortion, and motion blur in video data, affecting visual quality and restricting subsequent applications. Therefore, how to effectively improve the spatial resolution of videos, reduce noise interference, and enhance motion details has become a key problem that needs to be solved urgently.

[0004] The key to video super-resolution lies in how to utilize the redundant information between multiple frames for motion compensation and reconstruction. Traditional super-resolution methods rely on constructing prior information and introducing additional constraints to optimize the objective function. However, the selection and design of prior information are both complex and time-consuming. In recent years, convolutional neural networks (CNNs) have become the mainstream technology in the super-resolution field due to their powerful learning ability and non-linear feature representation ability. Currently, most methods rely on optical flow and deformable convolution techniques for motion estimation and compensation. Optical flow methods usually require additional components to generate inter-frame flow maps, or are combined with the overall network through an optical flow estimation sub-network, resulting in complex methods and high computational costs. Although the method based on deformable convolution reduces some computational complexity, it has limitations in temporal prior modeling and the expansion of spatio-temporal receptive fields. In addition, ordinary VSR algorithms are usually trained based on simple natural videos, and may be difficult to adapt to diverse motion patterns and content features when processing different types of videos (such as remote sensing videos, medical videos, or low-light environment videos).

[0005] In recent years, temporal difference has been proposed as a new method for modeling motion information. Research shows that temporal difference can effectively extract motion information, especially suitable for multi-scale targets in complex scenes. However, relying solely on short-term or long-term temporal differences is difficult to comprehensively model different types of motion patterns, and at the same time, a single receptive field may lead to the loss of detail information. In addition, a simple residual block structure is difficult to balance local and global motion information, which may cause the accumulation of alignment errors. Therefore, how to design an efficient video super-resolution method to accurately capture spatio-temporal information while reducing computational costs remains a problem worthy of in-depth exploration.

[0006] In the future, the research on video super-resolution will likely develop in the directions of adaptive spatio-temporal modeling, cross-modal information fusion, lightweight network architectures, etc., to further improve the video reconstruction quality and provide a clearer and smoother video experience for various application scenarios. Summary of the Invention

[0007] In order to overcome the problems of difficult capture of motion information and motion alignment of multi-scale moving targets in videos, the technical solutions adopted in the present invention are as follows:

[0008] A video super-resolution device based on multi-temporal difference and dynamic optimization framework, comprising: a total model; inputting a sequence of frame images into the total model, where is the target frame, and the rest are support frames; the total model includes a multi-level temporal difference analysis mechanism module, a temporal difference-guided dynamic path optimization module, a multi-attention enhancement and correction module, and a reconstruction module;

[0009] The multi-level temporal difference analysis mechanism module is used to calculate the temporal differences between adjacent frames, skipped frames, and cross frames, and capture the motion features at different time intervals;

[0010] The temporal difference-guided dynamic path optimization module locates the motion regions according to the temporal differences extracted by the multi-level temporal difference analysis mechanism module; dynamically selects different processing paths according to different inter-frame intervals, i.e., motion scales, and obtains multi-scale features of different paths;

[0011] The multi-attention enhancement and correction module performs information fusion and enhancement on the output features of the temporal difference-guided dynamic path optimization module in the temporal, spatial, and channel dimensions, and selectively corrects the alignment errors of the difference information of different paths;

[0012] The reconstruction module super-resolution reconstructs the output features of the multi-attention enhancement and correction module into the target frame.

[0013] The present invention has the following advantages compared with the prior art:

[0014] 1. A video super-resolution method based on multi-temporal difference and dynamic optimization framework disclosed in the present invention can simultaneously capture the motion information of targets with different scales and motion speeds using a multi-level temporal difference analysis mechanism, can comprehensively explore the frame sequence, and fully adapt to the diverse changes of moving targets.

[0015] 2. A video super-resolution method based on multi-temporal differences and a dynamic optimization framework disclosed by the present invention adopts a dynamic path optimization module guided by temporal differences, effectively solving the problem of information omission that may be caused by a single receptive field, thereby comprehensively improving the dynamic alignment performance for different moving targets.

[0016] 3. A video super-resolution method based on multi-temporal differences and a dynamic optimization framework disclosed by the present invention adopts a multi-attention enhancement and correction module to strengthen key features from multiple dimensions, uses a hybrid attention mechanism to model local and global information, selectively corrects and optimizes the aligned features, reduces the alignment error, and maintains the consistency of spatio-temporal information through residual connections and enhances the fidelity of the reconstructed image.

[0017] 4. A video super-resolution method based on multi-temporal differences and a dynamic optimization framework disclosed by the present invention is superior to other mainstream methods in performance and achieves a good balance between performance and inference time, having broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a model framework diagram of the present invention;

[0019] Figure 2 are RGB temporal difference maps and super-resolution reconstructed images of adjacent temporal differences and two-level temporal differences (adjacent temporal differences and jump temporal differences);

[0020] Figure 3 is a comparison diagram of experimental results for qualitative evaluation of the present invention;

[0021] Figure 4 is a comparison diagram of the relationship between the peak signal-to-noise ratio (PSNR) and the number of parallel loss operations per second (PLOPs) of the present invention and a comparative method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above objectives, the present invention adopts the following technical solutions.

[0023] In video super-resolution methods, current methods still cannot process complex and diverse moving objects in video scenes in an efficient manner while achieving spatio-temporal information consistency. Therefore, the present invention utilizes the characteristic that time difference information can be used for efficient motion alignment to perform training and learning of the model. The specific model process is as follows: First, the low-resolution frames are input into a multi-level time difference analysis mechanism to capture motion information at multiple time scales; the motion regions are locked using the differences between images, and then the alignment paths are dynamically selected according to the distance between frames; then important information in multiple dimensions is fused, redundant features are filtered, and the error of time difference alignment is reduced; finally, the high-resolution target frames are obtained.

[0024] More specifically, the present invention proposes a video super-resolution device based on a multi-time difference and dynamic optimization framework. As Figure 1 shown, the technical solution adopted by the present invention is as follows:

[0025] The low-resolution image sequence of frames is input into the overall model, where is the target frame to be super-resolved, and the rest are support frames. The overall model includes four main modules, namely:

[0026] Multi-level Temporal Difference Analysis Mechanism module (MTDM), which calculates the time differences between adjacent frames, skip frames, and cross frames, and captures motion feature information at different time intervals. By combining local and global motion features, it can accurately simulate the fast and slow motions of objects. MTDM inputs the low-resolution image sequence into three branches respectively to obtain time difference information ;

[0027] Temporal Differences-Guided Dynamic Routing Optimization Module (T-DROM), which locates the motion regions according to the time differences extracted by MTDM. It dynamically selects different processing paths (Route1, Route2, Route3) according to different inter-frame intervals (i.e., motion scales), prevents information omission caused by a single receptive field, and improves the alignment accuracy of multi-scale motion information to obtain multi-scale features of different paths ;

[0028] The Multi-Attention Enhancement and Correction Module (MAECM) fuses and enhances the information in the temporal, spatial, and channel dimensions of the output of T-DROM, and selectively corrects the alignment errors of the differential information in different paths to maintain spatio-temporal consistency and enhance the fidelity of the reconstructed image; The Reconstruction Module reconstructs the output features of MAECM into the target frame with super-resolution;

[0029] The Reconstruction Module (Reconstruction) super-resolution reconstructs the output features of MAECM into the target frame ; 。

[0030] The multi-level time difference analysis mechanism module MTDM contains three branch modules, namely:

[0031] The Adjacent Time Difference Module (ATDM) is used to effectively capture the fine details of fast-moving objects in the video. The low-resolution image sequence is input, and the local information about motion alignment is obtained by subtracting adjacent frames ; ;

[0032] The Skipped Temporal Difference Module (STDM) is used to supplement the deficiency that ATDM cannot capture slow-moving targets in the video, and further improve the comprehensiveness and accuracy of local information. The input low-resolution image sequence is subtracted at the position of skipping one frame to extract the second supplementary local information ; ;

[0033] The Crossed Temporal Difference Module (CTDM) is used to capture the global dependencies based on the entire frame sequence. The two low-resolution image sequences in opposite directions are input and subtracted to obtain the global-based temporal information. CTDM aims to explore the temporal differences between different frames not covered by ATDM and STDM, and further supplements the deficiencies of the first two modules. and ; The combination of the three modules realizes a comprehensive exploration of the entire frame sequence, including both local information and global information, so as to more comprehensively capture the diverse changes of moving targets.

[0034] For ease of description, it is set that

[0035] For ease of description, it is set that The low-resolution image sequence is respectively input into three branch modules of the MTDM module: Input frame by frame into the ATDM. Since the inter-frame distance is short, small-scale local motion information can be obtained An RGB difference map is obtained by subtracting adjacent frames , as Figure 2 shown in the RGB temporal difference map of ATDM in. Then, the difference maps are stacked along the temporal dimension, and feature extraction is performed using convolution. Subsequently, it is downsampled to the low-resolution space through average pooling operations. Processing the sparse difference information in the low-resolution space can significantly reduce the time and computational costs. Next, the difference maps are concatenated in the channel dimension, which not only improves the processing efficiency but also enhances the strength of the sparse signal. The above process can be expressed as:

[0036] (1)

[0037] where represents stacking, represents channel connection. Finally, a dual-branch structure is adopted to refine the difference information . In the first branch, the difference information is first upsampled to the original size and the restored difference information is added to the features of the target frame for preliminary compensation. Subsequently, the compensated features are refined through a series of convolution operations to obtain . In the second branch, without the support of the target frame , the difference information is directly refined, and then it is upsampled to the original size to obtain . The combination of the two branches realizes double information compensation for the target frame. The operation processes of the two branches are as follows:

[0038] (2)

[0039] (3)

[0040] (4)

[0041] where represents bilinear upsampling, and are two balance coefficients. Here we set α = β = 0.5.

[0042] As described by ATDM, ATDM can effectively capture the details of fast-moving targets. However, in practical applications, there are multi-scale moving targets with different moving speeds in video images, and the temporal difference information between adjacent frames is sparse, which is easily misidentified as interference information and deleted. ATDM can only capture partial local information, such as Figure 2 shown in the temporal difference map and super-resolution reconstruction image of ATDM in . Therefore, we designed STDM according to the motion characteristics of the target. By appropriately increasing the frame interval, more comprehensive motion information can be captured Figure 2 , as shown in the RGB temporal difference map and super-resolution reconstruction image of ATDM+STDM in . The experimental results show that the temporal difference effect of skipping one frame is the best. The specific operation is as follows: First, input the low-resolution image sequence into STDM, but the number of input frames is only half of that in ATDM, thus significantly reducing the computational amount. Then, calculate the temporal difference between the skipped frames to obtain the RGB difference map

[0043] (5)

[0044] Finally, the difference information also performs dual information compensation on the target frame through a dual-branch structure to enhance the representation ability of moving targets. The formula is expressed as follows:

[0045] (6)

[0046] (7)

[0047] (8)

[0048] where and are two weight coefficients used to balance the proportions of the two parts. Similarly, .

[0049] In the last branch CTDM, the low-resolution image sequence is rearranged in the reverse order to obtain the forward and backward frame sequences and . The two sequences are respectively subjected to convolution, channel connection, compression, and convolution operations to smooth the large variance. The specific process is as follows:

[0050] (9)

[0051] (10)

[0052] Indicates channel compression. Finally, calculate the forward sequence and the reverse sequence The difference between them is as follows:

[0053] (11)

[0054] (12)

[0055] Through the subtraction between these two bidirectional sequences and the supplementary information at different positions in the frame sequence can be obtained and collectively referred to as the cross-frame time difference feature .

[0056] The Temporal Differences-Guided Dynamic Routing Optimization Module (T-DROM) targets diverse temporal difference information , and and dynamically selects optimization paths with different receptive fields, including Route1, Route2, Route3. Since the difference between adjacent frames in ATDM is small, the alignment process is relatively simple. Therefore, in Route1, the target frame extracted by convolution is used to enhance and correct the difference output by ATDM, and then the number of channels is compressed to the original size through a convolution operation, and the formula is as follows:

[0057] (13)

[0058] The farther the support frame is from the target frame, the greater the alignment difficulty. In STDM, the captured moving targets usually have multi-scale characteristics. Route2 is designed to handle diverse motion information. First, the temporal difference Channel connections are made to enhance the spatial information of the target frame. Subsequently, a Multi-Scale Feature Enhancement Unit (MFEU) is designed to extract multi-scale features of moving targets using different receptive fields. Specifically, the first branch in the MFEU retains the original difference through direct connection; the second branch extracts features at the original size through 3×3 convolution; the third branch realizes feature propagation from small to large sizes through pooling, convolution, and bilinear upsampling operations. Finally, the features of the three branches are aggregated, and the number of channels is restored to the original size through convolution. The operations of the whole process are as follows:

[0059] (14)

[0060] Route3 is used to process the two difference sequence outputs of CTDM and . Since it is complex to explore the global information of the entire frame sequence and the time lengths of the differences are not uniform, the MFEU is first used to explore the multi-scale information of the two difference sequences, and the number of channels is expanded to multiples of the number of its frames through convolution. Then, the multi-scale enhanced difference values are activated to obtain two masks. Finally, the two masks are added together and the frame sequence after feature extraction using convolution is modulated to obtain a long-term dependence based on the frame sequence . The above process is expressed as follows:

[0061] (15)

[0062] Among them, represents the activation function. , and are collectively referred to as the output results of T-DROM .

[0063] In the Multi-Attention Enhancement and Correction Module (MAECM), all input frames are convolved to extract features, obtaining . Then, because the super-resolution of the sequence not only needs to retain the fine spatial structure but also pay attention to temporal continuity, a Temporal-Spatial Attention Fusion Module (TSA) is introduced to perform feature fusion in the temporal and spatial dimensions, obtaining the feature . In the time dimension, weights are assigned to the support frames in sequence according to the similarity between the support frames and the target frame, so as to capture the temporal dependencies of the entire sequence; in the spatial dimension, a pyramid structure is designed to effectively process multi-scale spatial feature information. Aiming at the misalignment error caused by temporal differences, a Motion Alignment and Correction Unit (MACU) is further proposed to optimize the features . It should be noted that due to the small temporal differences between adjacent frames and skip frames, the experimental results show that the pixel ratio of the residual information after MACU correction is extremely low. Therefore, MACU is dedicated to correcting the temporal differences between cross frames . Specifically, first, the temporally-differenced aligned features are subtracted from the features fused by spatio-temporal attention to obtain the residual information. Then, five residual blocks are used to refine the residual features. Among them, the skip connection mechanism enhances the detail restoration ability while maintaining the main structure. The outermost skip connection adds the original sequence to the difference features refined multiple times, thus ensuring the spatio-temporal consistency of the entire sequence. As follows:

[0064] (16)

[0065] Finally, the target frame features , the adjacent frame differences , the skip frame differences and the corrected cross frame differences are fused and further reconstructed into local and global features through a Hybrid Attention Transformer (HAT). HAT combines channel attention and window-based self-attention mechanisms, making full use of the complementary characteristics of both in local fitting and global modeling, so as to more effectively extract the key information in short-term, medium-term and long-term temporal differences. The implementation process of MAECM is as follows:

[0066] (17)

[0067] Among them, the balance coefficients and are both 0.5 respectively.

[0068] Finally, the temporal differences of frames at different distances, after being modulated and corrected by the attention of MAECM, are transmitted to the Reconstruction Module for final super-resolution reconstruction. The reconstruction module processes the input features Perform a convolution operation and implement upsampling through a PixelShuffle layer to generate a high-resolution reconstructed target frame :

[0069] (18)

[0070] To evaluate the method of the present invention, research is carried out in combination with embodiments in satellite videos. First, we collected a large number of satellite video clips from the Jilin-1 satellite. According to previous work, we composed 189 cropped scenes into our training set, 40 scenes as the validation set to verify the effect of the model during training, and randomly selected 5 scenes to form the test set Jilin-T, where the size of the cropped images is 640×640. In addition, to obtain more moving target data, we also selected scenes from a large-scale dataset VISO for moving target detection and tracking in satellite videos, cropped the scenes that met the size, and randomly selected 11 scenes to form the test set VISO-T.

[0071] The parameters designed in the experiment are as follows: the size of the input image I LR is H×W×C, and the size of the super-resolved image I SR t is Hr×Wr×C, where C is the number of channels, C = 3, H and W are the height and width of the input image, and r is the scale factor. In the present invention, we only focus on r = 4. The number of input frames is 5 frames. During training, we sample 4 patches of size 64×64 in each mini-batch. Data augmentation is achieved by performing operations such as random flipping and rotation on the data. The model of the present invention was trained for 35 epochs on an NVIDIA GeForce RTX 4090. The model training took nearly 40 hours.

[0072] The method proposed by the present invention is compared with several state-of-the-art techniques. To ensure the fairness of the comparison, these methods were all retrained on the Jilin-1 training set and used the same validation set and test set for qualitative and quantitative evaluations.

[0073] The visualization results in the qualitative evaluation are as Figure 3As shown, the present invention can restore clearer boundaries and more reliable details, and the error between the reconstructed super-resolution image and the real image is significantly reduced, and it performs outstandingly in restoring the simple and complex line information of buildings. For the vertical line area, obvious distortions have occurred in the RBPN and Zooming Slow-Mo methods, and the MSDTGP and LGTD methods have lost some information and the boundary lines are distorted. For complex boundaries, except for the present invention, boundary distortions have occurred in the remaining methods, and even unrealistic artifacts have been generated. These results show that the present invention can not only generate clear boundary information through motion information alignment, but also maintain the consistency of spatio-temporal information through an optimized compensation structure, thus providing excellent visual effects while ensuring the authenticity of the reconstruction target.

[0074] Two objective metrics, PSNR and SSIM, were used for quantitative evaluation, where the calculation of PSNR is based on the luminance (Y) channel. As shown in Table 1, the present invention has improved by 0.55 dB compared to RBPN based on optical flow alignment. This indicates that the multi-scale temporal difference information structure can not only capture the global motion information in a long time series, but also supplement fine motion details, thus achieving more comprehensive motion alignment. The present invention can fully exploit rich information through a simple and efficient model structure, achieve fast and effective image super-resolution reconstruction, and show higher practicality in video scenarios. The correlation between the FLOPs and PSNR of each method is as Figure 4 shown, where the PSNR value is the average of 16 scenes in two satellite videos, Jilin-T and VISO-T. It can be observed that the present invention is superior to all other methods in terms of performance and achieves a good balance between performance and computational complexity.

[0075] Table 1

[0076] Finally, ablation experiments were conducted to verify the effectiveness of each module and study and compare the model performance with different input frame numbers and skip frame numbers.

[0077] As can be seen from the above embodiments, the complex motion alignment problem in video super-resolution can be solved by the multi-temporal difference and dynamic optimization framework proposed by the present invention, and the inference time is fast, the system is simple, and the effect is remarkable.

[0078] Figure 1 The components included in it are respectively:

[0079] Multi-level Temporal Difference Analysis Mechanism (MTDM);

[0080] Adjacent Time Difference Module (ATDM);

[0081] Skipped Temporal Difference Module (STDM);

[0082] Crossed Temporal Difference Module (CTDM);

[0083] Temporal Differences-Guided Dynamic Routing Optimization Module (T-DROM);

[0084] Multi-Scale Feature Enhancement Unit (MFEU);

[0085] Multi-Attention Enhancement and Correction Module (MAECM);

[0086] Temporal-Spatial Attention Fusion Module (TSA);

[0087] Motion Alignment and Correction Unit (MACU);

[0088] Hybrid Attention Transformer (HAT);

[0089] Reconstruction Module (Reconstruction).

[0090] The content not described in detail in the specification of the present invention belongs to the prior art well-known to those skilled in the art. It is easy for those skilled in the art to understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A video super-resolution device based on multi-time difference and dynamic optimization framework, characterized in that Including: Total model; Input a sequence of frame images into the overall model, where is the target frame and the rest are support frames; the overall model includes a multi-level temporal difference analysis mechanism module, a temporal difference-guided dynamic path optimization module, a multi-attention enhancement and correction module, and a reconstruction module; Multi-level time difference analysis mechanism module, which is used to calculate the time differences between adjacent frames, jump frames and cross frames, and capture the motion features at different time intervals; Time difference-guided dynamic path optimization module, which locates the motion region according to the time differences extracted by the multi-level time difference analysis mechanism module; dynamically selects different processing paths according to different inter-frame intervals, i.e., motion scales, to obtain multi-scale features of different paths; Multi-attention enhancement and correction module, which performs information fusion and enhancement in the time, space and channel dimensions on the output features of the time difference-guided dynamic path optimization module, and selectively corrects the alignment errors of the difference information of different paths; Reconstruction module, which super-resolution reconstructs the output features of the multi-attention enhancement and correction module into the target frame.

2. The video super-resolution device based on a multi-time difference and dynamic optimization framework according to claim 1, wherein The multi-level time difference analysis mechanism module contains three branch modules, which are respectively: An adjacent time difference module, which is used to capture details of fast-moving objects in a video and input an image sequence to obtain local information about motion alignment by subtracting between adjacent frames ; A jump time difference module for capturing slow-moving targets in a video, the input image sequence Subtract at the position of jumping one frame to extract the locally supplemented information for the second time ; Cross-time difference module, which is used to capture global dependencies based on the entire frame sequence, and inputs two image sequences in opposite directions and subtract them to obtain global-based temporal information .

3. The video super-resolution device based on a multi-time difference and dynamic optimization framework according to claim 2, wherein In the adjacent time difference module, Frame by frame, the input enters the adjacent time difference module to obtain small-scale local motion information , and an RGB difference map is obtained by subtracting adjacent frames , then, the difference maps are stacked along the time dimension and convolutions are used for feature extraction, and then through an average pooling operation it is downsampled to a low-resolution space. Next, the difference maps are concatenated in the channel dimension to obtain difference information , and the above process is expressed as: (1) Among them, represents stacking, represents channel connection, and the difference information enters the double-branch structure respectively: In the first branch, the difference information is first upsampled to the original size, and the restored difference information is added to the features of the target frame for preliminary compensation. The compensated features are refined through a series of convolution operations to obtain , and the specific formula is as follows: (2) Among them, denotes bilinear upsampling, and are two balance coefficients; In the second branch, in the absence of a target frame supported, directly refine the difference information and then upsample it to the original size to obtain , and the specific formula is as follows: ​ (3) The combination of the two branches realizes the dual information compensation of the target frame. The operation processes of the two branches are as follows: (4) Among them, and are two balance coefficients.

4. The video super-resolution device based on a multi-time difference and dynamic optimization framework according to claim 3, wherein In the jump time difference module, first, the image sequence is input into the jump time difference module, and the number of input frames is half of that in the adjacent time difference module. Then, the time difference between the jump frames is calculated to obtain the RGB difference map . The difference maps are stacked along the time dimension and convolution is used for feature extraction. Subsequently, through the average pooling operation it is downsampled to the low-resolution space. Next, the difference maps are concatenated in the channel dimension to obtain the difference information . The specific processing method is as follows: (5) Final difference information Perform dual information compensation on the target frame through a dual-branch structure The formula is expressed as follows: (6) (7) (8) Among them, and are two weighting coefficients.

5. The video super-resolution device based on a multi-temporal difference and dynamic optimization framework according to claim 4, characterized in that, In the cross-time difference module, the image sequence is rearranged in the reverse order to obtain the forward and backward frame sequences and . Convolution, channel connection, compression, and convolution operations are respectively performed on the two sequences to smooth the large variance. The specific process is as follows: (9) (10) Indicates channel compression. Finally, calculate the forward sequence and the reverse sequence The difference between them is as follows: (11) (12) By subtracting between these two bidirectional sequences and a difference sequence of supplementary information at different positions in the frame sequence can be obtained and , collectively referred to as cross-frame time difference features .

6. The video super-resolution device based on the multi-time difference and dynamic optimization framework according to claim 5, wherein The time-difference-guided dynamic path optimization module targets diverse time-difference information , and , and dynamically selects optimization paths with different receptive fields, including Route1, Route2, and Route3; In Route1, the target frame extracted by convolution for the difference output by the adjacent time difference module is enhanced and corrected, and then the number of channels is compressed to the original size through a convolution operation. The formula is as follows: (13) In Route2, first, the temporal difference between the convolved target frame and the skip frame is concatenated in the channel dimension to enhance the spatial information of the target frame. Subsequently, through a multi-scale feature enhancement unit, multi-scale feature extraction of the moving target is performed using different receptive fields; In the first branch of the multi-scale feature enhancement unit, the original difference is retained through direct connection; in the second branch, features are extracted on the original size through 3×3 convolution; The third branch realizes the feature propagation from small to large sizes through pooling, convolution and bilinear upsampling operations. Finally, the features of the three branches are aggregated, and the number of channels is restored to the original size through convolution. The operations of the whole process are as follows: (14) Route3 is used to process two difference sequences of the cross-time difference module and . First, the multi-scale feature enhancement unit is used to explore the multi-scale information of the two difference sequences, and convolution is used to expand the number of channels to multiples of their number of frames. Then, the difference values after multi-scale enhancement are activated to obtain two masks. Finally, the two masks are added together and the frame sequence after feature extraction using convolution is modulated to obtain a long-term dependence based on the frame sequence . The above process is shown as follows: (15) Among them, denotes the activation function, , and are collectively referred to as the output results of T-DROM .

7. The video super-resolution device based on a multi-temporal difference and dynamic optimization framework according to claim 6, wherein In the multi-attention enhancement and correction module, for all input frames perform convolution to extract features and obtain , then, introduce a spatio-temporal attention fusion module to perform feature fusion in the temporal and spatial dimensions to obtain feature , in the temporal dimension, assign weights to the support frames in sequence according to the similarity between the support frames and the target frame, so as to capture the temporal dependencies of the entire sequence; In the spatial dimension, a pyramid structure is designed to effectively process multi-scale spatial feature information. To address the misalignment error caused by temporal differences, a motion alignment and correction unit is further proposed to optimize the features The motion alignment and correction unit is used to correct the temporal differences between cross frames .

8. The video super-resolution device based on multi-time difference and dynamic optimization framework according to claim 7, characterized in that In the motion alignment and correction unit, first, the time-difference alignment feature is subtracted from the feature fused by spatio-temporal attention to obtain the residual information. Then, five residual blocks are used to refine the residual features. The skip connection mechanism enhances the ability to restore details while maintaining the main structure. The outermost skip connection adds the original sequence to the difference features refined multiple times as follows: (16)。 9. The video super-resolution device based on multi-time difference and dynamic optimization framework according to claim 8, characterized in that Target frame features , Adjacent frame differences , Jump frame differences and the corrected cross-frame differences are fused and further reconstructed into local and global features through a hybrid attention transformer.

10. A video super-resolution device based on multi-time difference and dynamic optimization framework according to claim 9, characterized in that, The implementation process of the multi-attention enhancement and correction module is as follows: (17) Among them, the balance coefficient and are 0.5 respectively, and HAT is the attention transformer.

11. The video super-resolution device based on a multi-time difference and dynamic optimization framework according to claim 10, characterized in that, The reconstruction module processes the input features through a convolution operation and performs upsampling through a PixelShuffle layer to generate the reconstructed target frame : (18)。