Method for Video Restoration Using Spatiotemporal Transformer Network
Through the combination of encoder, related embedding modules, low-frequency and high-frequency feature converters of the space-time Transformer network, the problem of insufficient information integration in video repair is solved, and high-quality video repair effects are achieved, suitable for various video frames.
Patent Information
- Application Number
- CN202211247146.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-10-12
AI Technical Summary
Existing video repair technologies are difficult to effectively integrate time domain and airspace information, resulting in poor repair results, especially in complex video frames.
The spatiotemporal Transformer network is adopted to accurately extract and fuse the spatiotemporal information of video frames through the combination of encoder, related embedding modules, rough low-frequency feature converters, refined high-frequency feature transferrs and decoders.
It realizes high-quality repair of any video occlusion and damaged areas, improving the repair effect, especially in complex video frames, with high versatility and repair accuracy.
Smart Images

Figure CN115829857B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent vision restoration, and relates to a method for video restoration using a spatio-temporal Transformer network. Background Art
[0002] Video patching involves using a mask to smear moving or stationary objects in a video frame sequence, and then filling in the blurred parts based on the content information of the current frame and other frames of the video. The restored video should have the effect that the blurred positions "disappear". Typical applications are video restoration, watermark removal, object removal... The closer the restored blurred area is to the actual video, the better the restoration effect. Video restoration needs to combine temporal and spatial information to process video frames, search for spatial information in the current frame, then search for appropriate frames in other frames as reference frames to search for temporal information, and finally integrate these two parts of information and fill them back into the original frame to complete the restoration of the occluded position. Video restoration first needs to consider whether the missing information in the current frame is "exposed" in other frames. If the missing information of the current frame is found in other frames, the current frame should be used as a reference frame, match and extract valuable features, and then transmit them as information to the input frame to repair the occluded position. Although the recently emerging deep learning has made significant progress in image and video restoration, the ability of the model to capture useful information and reconstruct it in video frames is still fragile. To sum up, video restoration needs to integrate the information obtained in time and space and effectively transform and fill it back into the restored image. The more complex and challenging part in this process is: what information should be extracted from the reference frame (in time), and how to effectively extract and use the information of the reference frame and the current frame (in space). Summary of the Invention
[0003] The purpose of the present invention is to provide a method for video restoration using a spatio-temporal Transformer network, which solves the defect that accurate restoration cannot be achieved in previous restoration work, realizes the restoration of any video occlusion and damage, and has high restoration and integrity.
[0004] The technical solution adopted by the present invention is a method for video restoration using a spatio-temporal Transformer network, and this method is implemented according to the following steps:
[0005] Step 1: Construct a video restoration network STTTN, where the video restoration network STTTN includes an encoder, a correlation embedding module, a rough low-frequency feature transformer, a refined high-frequency feature transferrer, and a decoder; preprocess the video frames using a pre-restoration method to obtain pre-restored video frames, and the video frames include reference frames and input frames;
[0006] Step 2: Add the basic region normalization RN-B to the Encoder, and input the pre-repaired input frame, the pre-repaired reference frame, and the non-pre-repaired reference frame into the Encoder respectively to obtain the query Q, the key value K, and the content V respectively;
[0007] Step 3: Input the query Q and the key value K of each video frame into the relevant embedding module RE to obtain the feature block correlation W and the position index , where W records the correlation of the most relevant feature blocks between the current frame and all other frames, records the position index of the most relevant feature blocks between the current frame and all other frames;
[0008] Step 4: Input the position index of each video frame and the content V into the coarse low-frequency feature converter CLFT to obtain the texture feature map T;
[0009] Step 5: Input the preprocessed video frame repaired in Step 1 into the deep neural network DNN to obtain the feature value F, and then input the feature value F and the texture feature map T obtained by the coarse low-frequency feature converter into the refined high-frequency feature transfer RHFT for fusion;
[0010] Step 6: Finally, input the information of the video frame after fusion in Step 5 into the Decoder, and the fused video frame is decoded by the Decoder to obtain the repaired video frame.
[0011] The feature of the present invention also lies in that,
[0012] The process of pre-repairing the video frame in Step 1 is as follows: Using the Fast Marching Method, more weights are assigned to the pixels near the point, near the boundary normal, and on the boundary contour. Once a pixel is repaired, it will move to the next nearest pixel using the fast marching method, thereby repairing the video frame.
[0013] The process of adding the basic region normalization to the Encoder in Step 2 is as follows:
[0014] By introducing the basic region normalization RN-B, the spatial pixels are divided into different regions according to the occlusion, and then the mean and variance are calculated in different regions.
[0015] In the relevant embedding module in Step 3, first, expand the query Q and the key value K into small patches respectively, denoted as and , and calculate their similarity by dot product and :
[0016]
[0017] Among them denotes and the similarity between, and T represents the transpose operation.
[0018] The conversion process of the rough low-frequency feature converter in step 4 is as follows:
[0019] First, the low-frequency features in the values in different time domains are associated with the input frame through the rough low-frequency feature converter to calculate the rough low-frequency feature transfer map P. The i-th element in the low-frequency transfer feature map P is obtained according to the calculation formula: , where represents finding the maximum index value of the input value , represents the correlation.
[0020] Each value in the low-frequency feature transfer map P represents the position index of the current frame that is most relevant to the i-th position of the input frame among all reference frames. The specific calculation process returns the second item of the value of the torch.max() function to obtain the index corresponding to the maximum value. After obtaining the most relevant position index , only need to sequentially take the -th position index of the content V to obtain the texture feature map T, where each position of T contains the high-frequency texture features of the most similar positions in the reference frames. After obtaining the input frame texture feature map T, it is then used to refine the high-frequency feature transfer.
[0021] In step 5, the operation steps of the refined high-frequency feature transferrer are as follows:
[0022] Step 5.1: Calculate a feature block correlation W from to represent the confidence of the transmitted texture features at each position in the texture feature map T. The specific calculation process for obtaining the feature block correlation W is to obtain the maximum value of through the first item of the return value of the torch.max() function, where the feature block correlation W records the specific correlation of the most relevant feature block: ;
[0023] Step 5.2: Obtain the texture feature map T of multiple frames of images in the time domain through the rough low-frequency feature converter, then perform a Concat connection operation on the texture feature map T and the feature value F obtained from the deep neural network DNN and then perform a convolution Conv operation, and then multiply the connected result by the feature block correlation W;
[0024] Step 5.3: Then add the calculation result of Step 5.2 to the eigenvalue F obtained from the deep neural network DNN, where the deep neural network DNN is composed of multiple layers of convolution and residual connections, the convolution kernel is 3*3, and the stride and padding are 1; the above operation can be expressed by the following formula:
[0025]
[0026] where F represents the eigenvalue obtained by inputting the input frame to be pre-repaired into the deep neural network, and Conv, Concat, and ⊙ represent convolution, concatenation operation, and dot product respectively, is the feature output by the spatio-temporal texture of the input frame combined with the reference frame.
[0027] The learnable region normalization RN-L inserted in the decoder is used to automatically detect whether there is occlusion in the video frame.
[0028] The beneficial effects of the present invention are as follows: The method for video repair using the spatio-temporal Transformer network of the present invention has the following effects:
[0029] (1) Pre-repair of video frames with more accurate feature similarity measurement;
[0030] (2) By introducing region normalization (RN), the spatial pixels are divided into different regions according to an occlusion, solving the deviation problem of the mean and variance, and thus constructing an encoder with stronger information extraction ability;
[0031] (3) Embed information in the image, similar to the standard Transformer network structure, and introduce a related texture information embedding module (RE) to embed the reference and input frames;
[0032] (4) The coarse low-frequency feature transformer (CLFT) is used to transfer low-frequency information such as contours from the reference frame to the input frame;
[0033] (5) The precise refinement high-frequency feature transfer (RHFT) is used to transfer finer texture information such as image details to the input frame and improve the repair of occlusion;
[0034] (6) Similar to the encoder, learnable region normalization is added to the decoder to help fuse damaged and undamaged regions and modify the video frame more stably.
[0035] The beneficial effects of the present invention are as follows: The present invention is a method for video restoration using a spatio-temporal Transformer network, which can reconstruct (restore) occluded or damaged video regions. Throughout the invention, the Transformer is beneficial for the present invention to capture temporal and spatial information in different video frames. The basic region normalization RN-B and the learnable region normalization RN-L included in the invention are beneficial for stabilizing the restoration process of the present invention and improving the restoration effect. This video method has a good restoration effect on damaged videos both in terms of details and overall, and the restoration method is applicable to videos of any size and duration, having high versatility. Description of the Drawings
[0036] Figure 1 It is the overall flowchart of the method of the present invention. Detailed Implementation Manner
[0037] The overall architecture of the video restoration network STTTN of the present invention consists of five basic components, namely an encoder, a correlation embedding (RE), a coarse low-frequency feature transfer (CLFT), a refined high-frequency feature transfer (RHFT), and a decoder. Generally speaking, from the overall design, implicit restoration is used to restore the video at the feature level. In order to identify the defects of the previous architectures and achieve better restoration effects, after deleting the temporal search part of other previous models, it is found that the restoration ability of the model is significantly reduced, even lagging behind the effects of many non-spatio-temporal video restoration networks. This indicates that most previous works explored how to search for information in the time domain but ignored checking the deep representations of the obtained images. Therefore, the present invention establishes a new encoder and decoder architecture, which enables STTTN to have a stronger ability to capture image structures and temporal information. The overall architecture idea is that first, the encoder obtains the deep representation of the image, then, image restoration is performed, and finally the image is mapped back to the decoder to generate the restored image frame. The design of the overall structure is as Figure 1 shown.
[0038] Component 1: Encoder: Traditional video processing and image restoration methods use feature normalization (FN) to assist model training, but they usually perform it on the entire frame without considering the influence of pixels in the damaged region on the mean / variance. Here, the present invention introduces region normalization (RN), divides spatial pixels into different regions according to the occlusion, and then calculates the mean and variance in different regions. The basic region normalization (RN-B) is embedded in the encoder, which normalizes the damaged and undamaged regions according to the input mask, making the mean and variance offsets more accurate and more conducive to obtaining a deep image representation, and can extract useful information from video frames more comprehensively.
[0039] The input of the encoder network consists of three frames (pre-repaired input frame , pre-repaired reference frame and reference frame). Input frame (Inp), Input frame (Inp ), Reference frames (Ref) and Reference frames (Ref ) represent the input frame, pre-repaired input frame , reference frame and pre-repaired reference frame respectively. Inp and Ref consist of RGB images, hole occlusions and non-hole occlusions. The hole occlusion on the RGB image is a single-channel grayscale image, and the non-hole occlusion is the area outside the hole occlusion area. These inputs are concatenated along the channel axis to form a 5-channel image. Ref consists of a three-channel RGB image, given the input feature and a binary region mask indicating the damaged area, where C, H, and W represent the number of channels, width, and height of the image respectively, and R represents that the numbers in the set all belong to real numbers. For each channel, there are two sets of learnable parameters γ and β for the affine transformation of each region. Through the encoder, the obtained image representation consists of three parts: Q (query), K (key value), and V (content).
[0040] Component 2: Correlation Embedding Module (RE): Different from the previous operation of obtaining query Q, key value K, and content V through a linear transformer and then calculating the attention, the present invention obtains query, key value, and content with sufficient texture feature information through the encoder, which makes it easier to find the correlation between the input frame and the reference frame in the time domain. First, the present invention uses the correlation embedding module to estimate the similarity between the query and the key value, thereby establishing the correlation between the input frame and the reference frame. The present invention unfolds the query and the key value into small patches, denoted as and . The present invention calculates their similarity by taking the dot product of q and , where T represents the transpose operation, and represent the i-th small patch and the j-th small patch in the query and the key value respectively:
[0041]
[0042] The larger, the stronger the correlation between the two feature blocks, and the more texture information can be transferred, and vice versa. Using the correlation obtained by the correlation embedding module , the present invention can obtain two parts P and W, which are respectively used for rough low-frequency feature transfer and fine high-frequency feature transfer, and the specific calculation details will be discussed below.
[0043] Component 3: Coarse Low-Frequency Feature Transformer (CLFT): To better transmit the low-frequency information (such as contours) of an image, the present invention designs a coarse low-frequency feature transformer. The previous attention mechanism directly converts into a weight and then multiplies this weight by the content, which is actually the weighted average of the content. However, doing so may transfer a large amount of textures that are useless for the input frame to the target frame, resulting in blurred repaired areas. To improve the ability of the reference frame to transfer low-frequency texture features, the present invention will associate the low-frequency features in the content in different time domains with the input frame through CLET: The present invention first calculates the rough low-frequency feature transfer map P, where the i-th element is obtained by finding the maximum value according to the correlation :
[0044]
[0045] where represents finding the maximum index value of the input value , that is, each value in the map P represents the position index of a frame that is most relevant to the i-th position of the input frame among all reference frames. The specific calculation process obtains the index corresponding to the maximum value through the second item of the return value of the torch.max() function. After obtaining the most relevant position index, the present invention extracts the low-frequency texture features that should be transferred the most. Therefore, the present invention only needs to take the position of the frame to be transferred in the unfolded patches, and then the present invention can obtain the texture feature map T, where each position of T contains the high-frequency texture features of the most similar position in Ref. The present invention obtains the rough feature representation T of the input frame and then uses it for the refinement of the high-frequency feature transfer of the present invention.
[0046] Component 4: Refined High-Frequency Feature Transformer (RHFT): High-frequency detail information is also essential for video restoration, so the present invention designs a refined high-frequency feature transformer. To fuse the most suitable high-frequency textures in the time domain and the spatial domain with the input frame, a feature block correlation W is calculated from to represent the confidence of each position in T in transferring texture features. The specific calculation process for obtaining W is to obtain the maximum value of through the first item of the return value of the torch.max() function, where W records the specific correlation of the most relevant feature block.
[0047]
[0048] Among them is to find the maximum value of the input value. In order to make full use of the original image information of the input frame, the present invention divides the features of each layer into two steps: First, obtain the low-frequency texture features T of multiple frames of images in the time domain through a low-frequency feature converter, and fuse the features of the input frame, and then multiply by the feature block correlation W. At this time, W is equivalent to a weighted average of the features, which can more accurately transmit the texture features of the reference frame. Only two feature transfers cannot extract the information of the input frame well, so the present invention extracts the features of the input frame again and fuses the high-frequency and low-frequency features. The feature F extracted by the deep neural network DNN is a deep neural network composed of multiple layers of convolution and residual connections, the convolution kernel is 3*3, and the stride and padding are 1. The above operations can be expressed by the following formula:
[0049]
[0050] Conv, Concat, and ⊙ respectively represent convolution (the convolution operation used here is the same as the convolution operation used by the above DNN), concatenation operation, and dot product, is the feature output by the spatio-temporal texture of the input frame combined with the reference frame.
[0051] Component 5: Decoder
[0052] In the deep network, it is increasingly difficult to distinguish each damaged area and undamaged area, and the corresponding mask is difficult to obtain. To enhance the image reconstruction ability, the present invention inserts a learnable RN in the decoder to automatically detect occlusion and non-occlusion. The regions are individually normalized, and a global affine transformation is performed to enhance their fusion. Finally, the repaired video frame is output through the decoder, and finally the video frames are integrated to obtain the repaired video.
[0053] The method for video repair using a spatio-temporal Transformer network of the present invention constructs a powerful spatio-temporal transformer video repair network. The repair is divided into three steps: pre-repairing the input frame and the reference frame before the encoder, transmitting and repairing rough low-frequency texture features, and finally further refining and repairing high-frequency texture features related to details. The three parts complement each other from low to high and form a complete patching process layer by layer.
[0054] The method for video repair using a spatio-temporal Transformer network of the present invention is specifically implemented according to the following steps:
[0055] Step 1: Construct a video restoration network STTTN, where the video restoration network STTTN includes an encoder, a correlation embedding module, a rough low-frequency feature converter, a refined high-frequency feature transferrer, and a decoder; preprocess video frames using a pre-restoration method to obtain pre-restored video frames, where the video frames include reference frames and input frames;
[0056] The process of pre-restoring the video frames in Step 1 is as follows: Use the Fast Marching Method to assign more weights to the pixels near the points, near the boundary normals, and on the boundary contours. Once a pixel is restored, it will move to the next nearest pixel using the fast marching method, thereby restoring the video frames.
[0057] Step 2: Add a basic region normalization RN-B to the encoder Encoder, and input the pre-restored input frames, pre-restored reference frames, and un-pre-restored reference frames into the encoder respectively to obtain queries Q, key values K, and contents V;
[0058] The process of adding basic region normalization to the encoder is as follows:
[0059] By introducing the basic region normalization RN-B, the spatial pixels are divided into different regions according to the occlusion, and then the mean and variance are calculated in different regions.
[0060] Step 3: Input the queries Q and key values K of each video frame into the correlation embedding module RE to obtain the feature block correlation W and the position index , where W records the correlation of the most relevant feature blocks in the current frame with all other frames, and P records the position index of the most relevant feature blocks in the current frame with all other frames;
[0061] In the correlation embedding module in Step 3, first, expand the queries Q and key values K into small patches, denoted as and , and calculate their similarity by taking the dot product and :
[0062]
[0063] where represents the similarity between and , and T represents the transpose operation.
[0064] Step 4: Input the position index of each video frame and the content V into the rough low-frequency feature converter CLFT to obtain the texture feature map T;
[0065] In step 4, the conversion process of the rough low-frequency feature converter is as follows:
[0066] First, the low-frequency features in the values in different time domains are associated with the input frame through the rough low-frequency feature converter to calculate the rough low-frequency feature transfer map P. The i-th element in the low-frequency transfer feature map P is obtained according to the calculation formula: , where represents finding the maximum index value of the input value , and represents the correlation.
[0067] Each value in the low-frequency feature transfer map P represents the position index of the current frame that is most relevant to the i-th position of the input frame among all reference frames. The specific calculation process obtains the second item of the return value of the torch.max() function to get the index corresponding to the maximum value. After obtaining the most relevant position index , only need to sequentially take the -th position index of the content V to obtain the texture feature map T, where each position of T contains the high-frequency texture features of the most similar position in the reference frame. After obtaining the input frame texture feature map T, it is then used to refine the high-frequency feature transfer.
[0068] Step 5: Input the preprocessed video frame repaired in step 1 into the deep neural network DNN to obtain the feature value F, and then input the feature value F and the texture feature map T obtained by the rough low-frequency feature converter into the refined high-frequency feature transferer RHFT for fusion;
[0069] In step 5, the operation steps of the refined high-frequency feature transferer are as follows:
[0070] Step 5.1: Calculate a feature block correlation W from to represent the confidence of the transferred texture features at each position in the texture feature map T. The specific calculation process for obtaining the feature block correlation W is to obtain the maximum value of through the first item of the return value of the torch.max() function, where the feature block correlation W records the specific correlation of the most relevant feature block: ;
[0071] Step 5.2: Obtain the texture feature map T of multiple frames in the time domain through the rough low-frequency feature converter, then perform a Concat connection operation on the texture feature map T and the feature value F obtained by the deep neural network DNN and then perform a convolution Conv operation, and then multiply the connected result by the feature block correlation W;
[0072] Step 5.3: Then add the calculation result of Step 5.2 to the eigenvalue F obtained by the deep neural network DNN, where the deep neural network DNN is composed of multiple layers of convolution and residual connections, the convolution kernel is 3*3, and the stride and padding are 1; the above operation can be expressed by the following formula:
[0073]
[0074] where F represents the eigenvalue obtained by inputting the input frame to be repaired into the deep neural network, and Conv, Concat, and ⊙ represent convolution, concatenation operation, and dot product respectively, which is the eigenvalue output from the spatio-temporal texture of the input frame combined with the reference frame.
[0075] Step 6: Finally, input the information of the video frame after fusion in Step 5 into the decoder Decoder, and the video frame after fusion is decoded by the decoder to obtain the repaired video frame.
[0076] The learnable region normalization RN-L inserted in the decoder is used to automatically detect whether there is occlusion in the video frame.
[0077] To prove the effectiveness of the video restoration network of the present invention, experiments were conducted on the YouTube-VOS dataset and the DAVIS dataset respectively. We compared with other currently popular models: the research of Lee et al. (Lee, S., Oh, S. W., Won, D., and Kim, S. J. (2019). Copy-and-paste networks for deep video inpainting. In Proceedings of the IEEE / CVF International Conference on Computer Vision. 4413–4421), the research of Zeng et al. (Zeng, Y., Fu, J., and Chao, H. (2020). Learning joint spatial-temporal transformations for video inpainting. In European Conference on Computer Vision. 528–543). We compared on three metrics respectively: video Frechet Inception distance score, peak signal-to-noise ratio, and structural similarity. On the tests of these two datasets, all three metrics were improved, with the minimum improvements being 0.002, 0.34, and 0.0047 respectively. It was thus found that the more complex the video frames are, the better the model restoration effect is, because the video restoration network of the present invention captures temporal and spatial information well. In addition, it was observed that the overall restoration effect of the video restoration network of the present invention on videos with large scene changes is smoother and more uniform than other methods, and there is no excessive distortion of people and scenes, which proves that the transformer can better capture the global information of images.
Claims
1. A method for video restoration using a spatio-temporal Transformer network, characterized in that, The method is implemented according to the following steps: Step 1: Construct a video restoration network STTTN, where the video restoration network STTTN includes an encoder, a correlation embedding module, a rough low-frequency feature converter, a refined high-frequency feature transferrer, and a decoder; preprocess video frames using a pre-restoration method to obtain pre-restored video frames, where the video frames include reference frames and input frames; Step 2: Add basic region normalization RN-B to the encoder Encoder, and input the pre-restored input frame, the pre-restored reference frame, and the non-pre-restored reference frame into the encoder respectively to obtain query Q, key value K, and content V; Step 3: Input the query Q and key value K of each video frame into the relevant embedding module RE to obtain the feature block correlation W and the position index , where W records the correlation of the most relevant feature block in the current frame with all other frames, and records the position index of the most relevant feature block in the current frame with all other frames; In the relevant embedding module in step 3, first, the query Q and the key value K are respectively expanded into small patches, denoted as and , and their similarity is calculated by dot product and : Among them denotes and the similarity between, where T represents the transpose operation; and represent the i -th chunk in the query and the key value, respectively, and the j -th chunk, , ; Step 4: Input the position index of each video frame and the content V into the coarse low-frequency feature transformer CLFT to obtain the texture feature map T; the conversion process of the coarse low-frequency feature transformer is as follows: First, the low-frequency features in the values in different time domains are associated with the input frame through a rough low-frequency feature converter, and a rough low-frequency feature transfer map P is calculated, where the position index of the i-th position in the low-frequency transfer feature map P is obtained according to the calculation formula: , where represents finding the maximum index value of the input value , represents the correlation; Each value in the low-frequency feature transfer map P represents the position index of the current frame that is most relevant to the i-th position of the input frame among all reference frames. The specific calculation process returns the second item of the value through the torch.max() function to obtain the index corresponding to the maximum value; after obtaining the most relevant position index it is only necessary to sequentially take the position indices to obtain the texture feature map T, where each position in T contains the high-frequency texture features of the most similar position in the reference frame. After obtaining the input frame texture feature map T, it is then used to refine the high-frequency feature transfer Step 5: Input the preprocessed video frame restored in Step 1 into a deep neural network DNN to obtain a feature value F, and then input the feature value F and the texture feature map T obtained by the rough low-frequency feature converter into the refined high-frequency feature transferrer RHFT for fusion; Step 6: Finally, input the information of the video frame after fusion in Step 5 into the decoder Decoder, and the fused video frame is decoded by the decoder to obtain the restored video frame.
2. The method for video restoration using a spatio-temporal Transformer network according to claim 1, wherein, The process of pre-restoring the video frame in Step 1 is as follows: Use the Fast Marching Method to assign more weights to the pixels near the point, near the boundary normal, and on the boundary contour. Once a pixel is restored, it will move to the next nearest pixel using the fast marching method, thereby restoring the video frame.
3. The method for video restoration using a spatio-temporal Transformer network according to claim 2, wherein The process of adding basic region normalization to the encoder in Step 2 is as follows: By introducing basic region normalization RN-B, spatial pixels are divided into different regions according to occlusion, and then the mean and variance are calculated in different regions.
4. The method for video restoration using a spatio-temporal Transformer network according to claim 1, characterized in that, In Step 5, the operation steps of the refined high-frequency feature transferrer are as follows: Step 5.1: Calculate a feature block correlation W from to represent the confidence of the transmitted texture features at each position in the texture feature map T. The specific calculation process for obtaining the feature block correlation W is to obtain it through the first item of the return value of the torch.max() function of the maximum value, where the feature block correlation W records the specific correlation of the most relevant feature block: ; Step 5.2: Obtain the texture feature map T of multiple frames in the time domain through the rough low-frequency feature converter, then perform a Concat connection operation on the texture feature map T and the feature value F obtained by the deep neural network DNN, and then perform a convolution Conv operation, and then multiply the connected result by the feature block correlation W; Step 5.3: Then add the calculation result of Step 5.2 to the feature value F obtained by the deep neural network DNN, where the deep neural network DNN is composed of multiple layers of convolution and residual connections, the convolution kernel is 3*3, and the stride and padding are 1; the above operations can be expressed by the following formula: Among them, F represents the eigenvalue obtained by inputting the input frame to be pre-repaired into a deep neural network, and Conv, Concat, and ⊙ represent convolution, concatenation operation, and dot product respectively. is the feature output by the spatio-temporal texture of the input frame combined with the reference frame.
5. The method for video restoration using a spatio-temporal Transformer network according to claim 1, wherein, In Step 6, the learnable region normalization RN-L inserted in the decoder is used to automatically detect whether there is occlusion in the video frame.