A video inpainting method, system, device and medium based on optical flow guidance
Patent Information
- Application Number
- CN202311696259.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-12
AI Technical Summary
这些方法可以很好地从视频中聚合远程信息,但容易使恢复结果的局部结构与现实相矛盾
在本发明中,使用时间-空间维度的高效Transformer结构,在兼顾性能的同时从全局角度交互视频序列之间的信息,找到视频序列中和缺失区域相似度高的部分用于内容修复,充分利用了相邻帧信息保证修复结果连续性的同时,也利用了非相邻帧的内容信息,使修复结果更为合理;使用编解码结构从局部角度进行特征约束,使修复区域和周围特征纹理一致且结果合理,最终得到一段内容完整的视频内容,提高了内容修复的效率并提高了修复内容的合理性。
Smart Images

Figure CN117830158B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video processing technology, and in particular relates to a video restoration method, system, device and medium based on optical flow guidance. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Restoration work initially began with images, aiming to render missing areas using retrieved known content. Video inpainting evolved from image inpainting, but with the added temporal dimension. In video inpainting, due to the temporal consistency of the video, information about the damaged area can be obtained not only from the current frame but also from nearby frames. Traditional video rendering methods typically utilize the geometric relationship between the damaged area in the target frame and the valid area in the reference frame, such as homography or optical flow, using patch blocks to generate high-quality results. However, due to the cumbersome optimization process, these methods often have high computational complexity, limiting their practical application.
[0004] With the rapid development of deep learning, more efficient video filling schemes have been proposed based on deep learning, achieving significant improvements in filling quality and speed. These deep learning-based methods can be broadly categorized into three types based on their design principles: 3D convolution methods, optical flow methods, and attention methods.
[0005] 3D convolutional methods typically reconstruct missing content by directly aggregating temporal information from adjacent frames through 3D temporal convolution. For example, Wang et al. proposed the first deep learning-based video painting network, which consisted of a 3D CNN for temporal prediction and a 2D CNN for spatial detail recovery. Furthermore, Kim et al. employed a recurrent 3D-2D feedforward network to aggregate temporal information from adjacent frames into the missing regions of the target frame. While 3D convolutional methods can effectively aggregate temporal information, the relatively higher computational complexity of 3D CNNs compared to 2D CNNs limits their application.
[0006] To alleviate this problem, some researchers have transformed the aggregation of temporal information into a pixel propagation problem and used optical flow information to represent the changes between pixels. These methods first introduce a deep flow completion network to recover the flow sequence, and then use the recovered flow sequence to fill in the relevant pixels in missing regions of adjacent frames. For example, Xu et al. used a coarse-to-fine deep flow completion network to complete the flow field, guiding relevant pixels to the missing regions. Building on this, Gao et al. further improved the performance of video inpainting by explicitly completing the flow edges. Zou et al. corrected spatial misalignments in the temporal feature propagation stage using completed optical flow. Although these methods have shown encouraging results, due to the nature of optical flow itself, they are not very good at aggregating visible content from distant frames.
[0007] To effectively simulate long-range communication, state-of-the-art methods utilize attention mechanisms to capture long-term communications. For example, Zeng et al. proposed the first transformer model for video rendering by learning multi-layer multi-head transformers. Furthermore, Liu et al. improved edge details of missing content by using soft segmentation and soft synthesis operations in the transformer. Ren et al. developed a novel discrete latent transformer, DLFormer, which formulates the video rendering task as a discrete latent space. These methods can effectively aggregate long-range information from video, but they are prone to causing the local structure of the restored result to contradict reality. While Transformer achieves impressive performance, its self-attention module also introduces the same high computational complexity. When designing video reconstruction structures, it is crucial to consider not only the quality of the reconstruction result but also improving the efficiency of video reconstruction. Summary of the Invention
[0008] To overcome the shortcomings of the prior art, this invention provides a video restoration method, system, device, and medium based on optical flow guidance. For video features with missing content, optical flow information is combined with Transformer to restore the target video frame from both temporal and spatial dimensions under the guidance of optical flow information. This fully utilizes the information of adjacent frames to ensure the continuity of the restoration result, while also utilizing the content information of non-adjacent frames to make the restoration result more reasonable. Moreover, the restoration process uses Transformer to consider global feature relationships and uses the encoding and decoding structure to consider local feature relationships in the preliminary restoration result, which can effectively solve the computational complexity caused by the self-attention mechanism.
[0009] To achieve the above objectives, a first aspect of the present invention provides a video restoration method based on optical flow guidance, comprising: Obtain the video frame sequence to be repaired; An optical flow estimation network is used to estimate the video frame sequence to be repaired to obtain a missing optical flow information sequence. The missing optical flow information sequence is then filled in based on the local correlation between optical flows to obtain complete optical flow information. Efficient self-attention mechanisms are introduced into the temporal transformer structure and the spatial transformer structure, respectively. The video frame sequence to be repaired and the complete optical flow information are guided by the complete optical flow information from the global temporal and global spatial dimensions. Features with high similarity to the missing region are used for repair operation to obtain preliminary repair results. Based on the local feature correlation of the preliminary repair results, the preliminary repair results are adjusted using the encoding and decoding structure to obtain the final repaired video.
[0010] A second aspect of the present invention provides an optical flow-guided video restoration system, comprising: Acquisition module: Acquires the video frame sequence to be repaired; Optical flow information module: An optical flow estimation network is used to estimate the video frame sequence to be repaired to obtain a missing optical flow information sequence, and the missing optical flow information sequence is filled based on the local correlation between optical flows to obtain complete optical flow information; Preliminary Repair Module: Efficient self-attention mechanisms are introduced into the temporal transformer structure and the spatial transformer structure, respectively. Guided by the complete optical flow information, the video frame sequence to be repaired and the complete optical flow information are processed from the global temporal and global spatial dimensions, and features with high similarity to the missing region are used for repair operations to obtain preliminary repair results. Final Repair Module: Based on the local feature correlation of the preliminary repair results, the module adjusts the preliminary repair results using the encoding and decoding structure to obtain the final repaired video.
[0011] A third aspect of the present invention provides a computer device, comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, an optical flow-guided video restoration method is performed.
[0012] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs an optical flow-guided video restoration method.
[0013] The above one or more technical solutions have the following beneficial effects: In this invention, a highly efficient Transformer structure with a time-space dimension is used to interact with information between video sequences from a global perspective while taking performance into account. It finds parts of the video sequence with high similarity to the missing area for content repair, making full use of information from adjacent frames to ensure the continuity of the repair results, while also utilizing content information from non-adjacent frames to make the repair results more reasonable. An encoding and decoding structure is used to constrain features from a local perspective, making the repaired area consistent with the surrounding feature texture and the result reasonable. Finally, a complete video content is obtained, which improves the efficiency of content repair and the reasonableness of the repaired content.
[0014] In this invention, an efficient self-attention mechanism simulates the interaction between different heads by introducing DW convolution. The attention function of each head can depend on all keys and queries. The introduction of DW convolution weakens the ability of single-head attention to focus on information from different representative subsets from different positions. By introducing instance normalization, not only is the ability to diversity restored, but the time complexity of the self-attention mechanism can also be effectively reduced, and the amount of computation can be reduced.
[0015] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0016] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0017] Figure 1 This is a flowchart of the optical flow-guided video restoration method in Embodiment 1 of the present invention; Figure 2 This is a flowchart of the optical flow repair module in Embodiment 1 of the present invention; Figure 3 This is a flowchart of the efficient Transformer structure in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the window division for the self-attention mechanism in Embodiment 1 of the present invention; Figure 5 This is a flowchart illustrating the fusion of optical flow information and feature information in Embodiment 1 of the present invention; Figure 6 This is a flowchart of the structural adjustment module in Embodiment 1 of the present invention; Figure 7 This is a schematic diagram comparing the FGVC, STTN, FFM, DSTT, E2, FGT methods and the optical flow-guided video inpainting method in Embodiment 1 of the present invention in a scene region. Figure 8This is a schematic diagram showing the results of the video restoration method based on optical flow in Embodiment 1 of the present invention, displaying consecutive frames. Figure 9 This is a schematic diagram showing a visual comparison of the ablation experiment in Embodiment 1 of the present invention. Detailed Implementation
[0018] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0019] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0020] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0021] Example 1 This embodiment discloses a video restoration method based on optical flow guidance, including: Obtain the video frame sequence to be repaired; An optical flow estimation network is used to estimate the video frame sequence to be repaired to obtain a missing optical flow information sequence. The missing optical flow information sequence is then filled in based on the local correlation between optical flows to obtain complete optical flow information. Efficient self-attention mechanisms are introduced into the temporal transformer structure and the spatial transformer structure, respectively. The video frame sequence to be repaired and the complete optical flow information are guided by the complete optical flow information from the global temporal and global spatial dimensions. Features with high similarity to the missing region are used for repair operation to obtain preliminary repair results. Based on the local feature correlation of the preliminary repair results, the preliminary repair results are adjusted using the encoding and decoding structure to obtain the final repaired video.
[0022] This embodiment discloses a video restoration method based on optical flow guidance. The core idea of this method is to restore the missing areas of the video by using the content information of adjacent and non-adjacent frames under the guidance of video optical flow data. Even when the object is partially missing, the original structure of the object can be restored well.
[0023] The method in this embodiment combines the structures of CNN and Transformer and utilizes complete optical flow information to repair missing regions. It includes two parts: the first part is a Transformer network that fills the missing regions under the guidance of optical flow information, and the second part is a CNN network that performs overall fine-tuning of the filling structure.
[0024] Guided by complete optical flow, a transformer network is used to extract relevant and effective features from video frames in both temporal and spatial dimensions to repair missing regions in the target frame. In the temporal transformer, an efficient self-attention mechanism, EMSA, is used on the frame sequence, while in the spatial transformer, efficient self-attention EMSA is performed within tokens originating from the same frame. In the temporal and spatial transformers, due to the significant temporal offset between the target frame and non-adjacent frames, a normal-sized window cannot contain all the effective information. Therefore, EMSA is performed within a large window to ensure content integrity. Considering that optical flow information is generated and may contain some errors, a flow reweighting module is used between the two transformer structures to adaptively control the influence weight of optical flow information on the recovered content based on the interaction between tokens.
[0025] Because the Transformer focuses more on the interaction of long-range information, the repaired content, while having a reasonable texture, may have local structures that differ from the real structure. Therefore, a CNN network with stronger local constraints is used to perform local feature interactions on the repaired content and adjust its structure.
[0026] In this embodiment, optical flow estimation and image inpainting techniques are used to obtain incomplete optical flow information of the input video sequence at the beginning and repair the incomplete optical flow information, which is then used as guidance information for subsequent filling work to achieve stability of the entire video.
[0027] The framework proposed in this embodiment can be subdivided into three modules: optical flow restoration module, content repair module, and structure adjustment module.
[0028] like Figure 1 The framework shown is encapsulated as a loop scheme that, for a video with missing content, iterates from the first frame to the last frame to obtain a complete video. The input data is the sequence of video frames with missing content. In extracting features from the input sequence Simultaneously, the optical flow recovery module is used to obtain the complete optical flow information of the sequence. Patch Embedding is used to divide the two types of data into token blocks. and The data is then passed to the content restoration module. In this module, high similarity information from other frames is used to fill in the missing areas of the target frame. Next, based on the results from the content restoration module, the structure adjustment module fine-tunes its local structural features. After completing these steps, a complete video segment is finally obtained through decoding.
[0029] The following section details the process steps of the optical flow-guided video restoration method.
[0030] The optical flow restoration module is the first step in the entire framework. Since the missing areas in the video may contain portions of moving objects, relying solely on high similarity information within adjacent frames is insufficient to effectively distinguish the textures of moving objects from the background. However, the outlines of moving objects within the optical flow information are more prominent than within the image, enabling differentiation from the background. Therefore, complete optical flow information guides subsequent restoration tasks. The state-of-the-art optical flow estimation method RAFT is used to obtain the incomplete optical flow sequence from the input video frame sequence. Then, the P3D block is inserted into the LAFC encoder, and thanks to the powerful recovery capability of the U-net network, a complete optical flow sequence is obtained. The outlines of moving objects in the optical flow map are completely restored, which helps guide subsequent video restoration.
[0031] Figure 2 This is a detailed structural diagram of the optical flow recovery module in this embodiment. This embodiment first utilizes the state-of-the-art optical flow estimation network RAFT, estimating a size of H×W×3, where H represents height, W represents width, 3 represents the number of channels, and the length is T for the damaged video sequence. Incomplete optical flow information sequence The sequence includes Forward optical flow and backward optical flow , where i is the time interval between consecutive flows, and the length of the flow sequence is 2n + 1. Then, P3D blocks, or Pseudo-3D blocks, are used to aggregate features from local flows in both temporal and spatial dimensions.
[0032] This implementation inputs the optical flow sequence map with missing content into an improved U-net-type network designed in this embodiment to obtain an enhanced complete optical flow map, and uses this to utilize the local correlation between flows to analyze the incomplete optical flow information sequence. The filling process is performed, and the filling result is decoded to obtain a complete optical flow information sequence. .
[0033] Specifically, such as Figure 2As shown, the improvement in the U-net type network is as follows: In the existing U-net type network, P3D blocks are added to the coding convolutional layer through residual connections. For example, the input of the first coding convolution of the U-net type network is used as the input of the first P3D block, and the output of the first coding convolution is added to the output of the first P3D block and used as the input of the second coding convolution, and so on.
[0034] The input of the m-th coding layer is represented as The output is represented as The local feature aggregation process can be represented by formula (1):
[0035] (1)
[0036] Among them, TC() is a one-dimensional temporal convolution, and SC() is a two-dimensional spatial convolution. The two work together to form a P3D block. () denotes the encoding layer of the local flow feature aggregation network. The temporal resolution remains unchanged except for the last P3D block in the encoder and the P3D blocks inserted in the jump connections. Within these blocks, the aggregated flow features of the target flow are obtained by reducing the temporal resolution of the flow sequence.
[0037] This embodiment focuses on the local clustering features after dilated convolution. Decoding is performed to visualize the optical flow information and obtain complete flow information. The process of the optical flow recovery module is shown in formula (2):
[0038] (2)
[0039] Where DC() represents dilated convolution, Concat() represents feature concatenation, and Decoder() represents the decoding layer of the improved U-net network. This represents the stream features after n encoding layers, where n is the number of encoding layers and t is the number of video frames.
[0040] The content restoration module in this embodiment is the core component. Filling in missing areas involves establishing correspondences by comparing the content similarity between different frames. Since the receptive field of convolutional kernels is narrow, convolutional operations alone cannot effectively capture these correspondences. Therefore, this embodiment employs a hybrid local and global feature extraction method to enhance the long-distance dependence of features, thereby better capturing the spatial correspondence between two frames. Here, a 5-layer convolutional neural network is used to extract features from the video frame sequence to be restored, obtaining low-level features. , low-level features and complete optical flow information Image features are obtained after performing patch embedding. and optical flow characteristics Image features will be obtained. and optical flow characteristics As input, guided by optical flow information, a Transformer structure is used to identify features with high similarity to the missing regions in the video sequence from both temporal and spatial dimensions for the inpainting operation, ultimately yielding preliminary inpainting results. .
[0041] Specifically, after obtaining complete optical flow information from the previous stage, this embodiment will proceed according to the optical flow sequence. The missing regions in the video are repaired. The content repair module consists of both temporal and spatial Transformer structures. He et al. proposed that the similarity of content in adjacent frames leads to content redundancy when processing video information. Therefore, information from non-adjacent temporal domains can serve as a promising reference for missing regions in local neighborhoods. Here, this embodiment stacks multiple temporal-spatial Transformer structures to effectively combine information from adjacent and non-adjacent domains to achieve content filling. Considering the computational complexity of multiple Transformers, this embodiment adds channel-wise convolution (DW convolution) to the multi-head attention mechanism (MHSA) in the Transformer block, reducing the computational cost of the self-attention mechanism and making MHSA a more efficient EMSA.
[0042] A standard Transformer module consists of MHSA and FFN. The output of each Transformer block... As shown in formula (3):
[0043] (3)
[0044] Here, FFN() is the feedforward layer of the Transformer. For multi-head self-attention mechanisms, LN() is the normalization operation. This is the input for the standard Transformer module.
[0045] MHSA first obtains the query Q, key K, and value V by applying linear normalization to the input. Each convolutional group consists of K linear layers, which will... Dimension input mapping to Dimensional space, in which This refers to the head dimension. For ease of description, let's assume... Then, the multi-head attention mechanism MHSA can be simplified to single-head self-attention SA. The global relationship between token sequences can be shown in Equation (4):
[0046]
[0047]
[0048]
[0049] (4)
[0050] in, These represent linear layers with different parameters. This indicates a normalization operation. () indicates a regression operation. Then, the output values of each head are concatenated and linearly operated to form the final output. As shown above, for the standard Transformer, MSA has two disadvantages: (1) The calculation of the MSA scale is based on the fact that the input dimensions d and n are quadratic, resulting in huge training and inference overhead; (2) Each head in MSA is only responsible for a subset of the input, which may reduce the performance of the network, especially when the channel dimension d in each subset is too low, so that the dot product of Q and K cannot form an information matching function. In order to solve these problems, this embodiment proposes an efficient self-attention mechanism, as shown in formula (5):
[0051]
[0052]
[0053]
[0054] (5)
[0055] Here, Conv() is a standard 1×1 convolution operation that simulates the interaction between different heads. This is represented as a channel-wise convolution. Therefore, the attention function of each head can depend on all keys and queries. However, this weakens the ability of the self-attention mechanism (MSA) to collectively focus on information from different representative subsets from different locations. To restore this diversity, this embodiment adds an instance normalization (IN()) to the dot product matrix, i.e., after Softmax. This change effectively reduces the time complexity of the self-attention mechanism. This embodiment... Figure 3 The diagram shows the structure of an efficient Transformer.
[0056] The role of the temporal transformer is to enable the interaction of long-range information within a video. This allows information to be obtained from other video frames when there is no loss of region information within the target frame. The effectiveness of this module is particularly evident in the process of restoring moving objects within a video. In the temporal transformer, attention retrieval is performed on the tags of different frames. Since content moves along the time dimension, it is reasonable to use a large window to compensate for the reference offset. Therefore, in this embodiment, image tokens are divided into non-overlapping cubes with large window sizes, represented as "regions," along the height and width dimensions, and EMSA is performed within the cubes.
[0057] This embodiment is in Figure 4 The diagram illustrates the window partitioning of the self-attention mechanism, which uses token blocks with high similarity to other video frames in the input frame sequence to fill in the missing content of the target frame. The implementation of the time transformer is shown in formula (6).
[0058]
[0059] (6)
[0060] in, The token is the input image; LN() is the normalization operation; EMSA() is the efficient self-attention mechanism; and FFN() is the feedforward layer of the Transformer. The output of the temporal Transformer structure to introduce an efficient self-attention mechanism.
[0061] like Figure 5 As shown, in the spatial Transformer, complete optical flow information is used to guide the attention process. The simplest method is to directly concatenate the optical flow token and the image token. However, this operation has two problems. First, since the generated optical flow information may have errors compared to the real situation, this can mislead the judgment of related areas. Second, the internal texture information of the same motion region in the optical flow may differ, and operating it according to the same proportion may result in a difference between the result and the original content.
[0062] To mitigate these issues, this embodiment first uses a multilayer perceptron (MLP) to look up the results obtained from the temporal Transformer. and optical flow Similarity between information. Optical flow based on similarity. Perform weighting. The weighting process is shown in formula (7):
[0063] (7)
[0064] Here, MLP() represents the MLP operation. The weighted result is an enhanced stream. This enhances the filling capability of the spatial Transformer. The implementation of the spatial Transformer is shown in equation (8):
[0065]
[0066]
[0067] (8)
[0068] Where Rt represents the enhanced optical flow feature With the output of the time Transformer The cascaded result shows that Fusion() represents the fusion operation of formula (7), LN() is the normalization operation, EMSA() is the efficient self-attention mechanism, and FFN() is the feedforward layer of the Transformer. The output of the spatial Transformer structure, which introduces an efficient self-attention mechanism, is the initial repair result of our method.
[0069] The structural adjustment module is the final implementation part. Video restoration aims to ensure the obtained content is complete and displays correctly. This means that not only the completeness of the restoration result but also its rationality must be considered. The structural adjustment of the video content involves using a U-net-type network to integrate the initial restoration results. The process involves reshaping the content to enhance the correlation between local features, ensuring that the repaired content matches the surrounding area while also closely reflecting real-world patterns. This represents the video fill result after local structural adjustments. After Transformer decoding, the final video content is complete and reasonable.
[0070] Considering that the self-attention mechanism focuses more on the interaction of global information, this may lead to structural deformation of local information. The CNN structure, on the other hand, focuses more on the interaction of local information. Therefore, this embodiment adds a 3-layer encoder-decoder structure after the Transformer structure. This strengthens the constraint of the content around the missing region and fine-tunes the filling result. This makes the repair result of some damaged objects more realistic in terms of both structure and texture, and the effectiveness of this module is verified in the ablation experiment. The structural adjustment process is shown in formula (10): (10)
[0071] Where * denotes dot product operation. A mask representing a missing area in a video frame. This represents the result after the local structural adjustments have been completed. This indicates the initial repair of the damaged area in the video. ED() represents the codec's operation process; the decoding operation takes the features encoded by the structure adjustment module and the corresponding encoded features in the concatenation as input. The effectiveness of this module was verified in ablation experiments in this embodiment of the invention.
[0072] This implementation uses reconstruction loss in the damaged region and effective region and T-Patch GAN loss to supervise the training process, and uses hinge loss as adversarial loss.
[0073] This embodiment extensively compares representative stable methods from recent years in terms of both imaging effects and quantitative analysis to evaluate the method of this embodiment.
[0074] 1. Experimental setup This embodiment uses the YouTube-VOS and DAVIS datasets for evaluation. YouTube-VOS contains 541 videos, and DAVIS contains 90 videos, covering a variety of scenarios. This embodiment uses the YouTube-VOS training set to train the network. During training, irregular masks are randomly generated proportionally to simulate video content incompleteness. During testing, quantitative analysis is performed using moving rectangular masks and object-like masks, respectively. Based on previous work, this embodiment selects PSNR, SSIM, and LPIPS as evaluation metrics. This embodiment is compared with state-of-the-art methods, including VINet, STTN, FGVC, DSTT, FFM, E2, and FGT.
[0075] This embodiment uses three metrics—PSNR, SSIM, and LPIPS—to evaluate the performance of state-of-the-art video rendering methods. Specifically, PSNR and SSIM are commonly used metrics for evaluating distorted images and videos; higher values indicate better performance. LPIPS measures the perceptual similarity between two input videos; higher values indicate better performance and have been adopted in recent video inpainting works.
[40] .
[0076] 2. Comparison of Results Quantitative results tested on the YouTube-VOS and DAVIS datasets are presented, and the proposed method is compared with previous video inpainting methods, including VINet, STTN, FGVC, DSTT, FFM, E2, and FGT. During inference, the video resolution is set to 432×256. Square mask sets with continuous motion tracking are generated for the YouTube-VOS and DAVIS datasets. The average size of the masks in the square mask sets is 1 / 16 of the entire frame. The DAVIS object mask set is randomly shuffled, and frames are corrupted using these masks to evaluate video rendering performance on the object masks. As shown in Table 1, the proposed method significantly outperforms all previous state-of-the-art algorithms on three quantitative metrics. The excellent results demonstrate that the proposed method produces videos with less distortion in terms of peak signal-to-noise ratio and structural similarity, more visually plausible content in terms of image similarity, and better spatiotemporal consistency.
[0077] surface The bold text in 1 indicates the optimal method for this metric.
[0078] Table 1. Comparison results with various methods on the DAVIS and Youtube-VOS datasets.
[0079] This embodiment selects three representative methods—STTN, FGVC, DSTT, FFM, E2, and FGT—for a direct comparison. Figure 6 The results of video completion and object removal are shown in the example. In contrast, other methods struggle to recover reasonable details in occluded areas, while the method proposed in this embodiment generates faithful texture and structural information. This demonstrates the effectiveness of the proposed method. Figure 6 The results of restoring video damaged by a rectangular mask using various methods are shown. From the examples in the third row, it can be seen that the FGVC method produces a result with obvious distortion; the structural shape of the car in the example is distorted. From the examples in the second row, it can be seen that the STTN and DSTT methods result in missing object content; the cat's ears, which are hidden by the mask, are not filled in. From the examples in the fourth row, it can be seen that the E2 and FFM methods produce blurring; the dog's ears show obvious distortion. From the examples in the first row, it can be seen that the FGT method produces ghosting; there are extra artifacts in the dog's head area.
[0080] In addition, Figure 7 The results of object removal from video using the method of this embodiment are shown. It can be seen that the generated results are continuous in time, and there are no obvious ghosting or distortion. Figure 8 The results of the optical flow-guided video restoration method in this embodiment are shown in consecutive frames.
[0081] 3. Ablation test This embodiment uses the DAVIS dataset for ablation experiments, and the results are shown in Table 2. Each video in the DAVIS dataset contains a clear moving foreground, which is consistent with the problem considered in the design of this embodiment. The experiment performed three types of comparisons, and the results are shown in Table 2. The visual results are as follows: Figure 9 As shown.
[0082] As shown in Table 2, when the structural adjustment module is removed from the model, the values of the quantitative indicators drop sharply. Figure 9 As can be seen in line a, the video restoration effect deteriorates significantly after the module is removed. Figure 9 The position of the bird's wings in line a is not well depicted from the scene, and the reconstructed neck content does not match the original content. From Figure 9 As can be seen from line b, removing this module makes the outlines of objects within the generated area less distinct. Figure 9 In line b, the arm area of the person is masked, and after removing the module, the drawn arm outline is not clear.
[0083] Table 2. Ablation Experiment Results
[0084] As shown in Table 2, adding the probability map significantly improved the video's distortion. This is because the fine-tuning of the structure adjustment module strengthened the connections between local features, improving the accuracy of the object edge repair structure. This fully demonstrates the rationality of the stable frame design.
[0085] This embodiment proposes a method for video stabilization that considers both global and local information. Guided by optical flow information, a Transformer structure is first used to interact with long-range information of the video sequence from two dimensions to identify regions with high similarity to missing areas. Then, an encoding / decoding structure is used to constrain local information, strengthening the continuity between local features. Through the collaborative work of both approaches, video content restoration is achieved. The deep learning-based method proposed in this embodiment utilizes a computationally low-cost self-attention mechanism to obtain similarity between different information types, improving the efficiency and reasonableness of content restoration. This embodiment proposes a supervised deep learning-based algorithm that uses consecutive video frames and their missing region masks as input to achieve video stabilization. This embodiment effectively enhances moving objects within the video using optical flow information, and uses the contour information of these moving objects in the optical flow to guide subsequent content restoration. It employs an efficient temporal-spatial Transformer structure to interact with information between video sequences from a global perspective while maintaining performance, identifying regions in the video sequence with high similarity to the missing regions for content restoration. An encoding / decoding structure is used to constrain features from a local perspective, ensuring that the restored region matches the surrounding texture and that the result is reasonable, ultimately yielding a complete video content. This embodiment's method restores video frames while also addressing some low-quality videos.
[0086] Example 2 The purpose of this embodiment is to provide a video restoration system based on optical flow guidance, including: Acquisition module: Acquires the video frame sequence to be repaired; Optical flow information module: An optical flow estimation network is used to estimate the video frame sequence to be repaired to obtain a missing optical flow information sequence, and the missing optical flow information sequence is filled based on the local correlation between optical flows to obtain complete optical flow information; Preliminary Repair Module: Efficient self-attention mechanisms are introduced into the temporal transformer structure and the spatial transformer structure, respectively. Guided by the complete optical flow information, the video frame sequence to be repaired and the complete optical flow information are processed from the global temporal and global spatial dimensions, and features with high similarity to the missing region are used for repair operations to obtain preliminary repair results. Final Repair Module: Based on the local feature correlation of the preliminary repair results, the module adjusts the preliminary repair results using the encoding and decoding structure to obtain the final repaired video.
[0087] Example 3 The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method.
[0088] Example 4 The purpose of this embodiment is to provide a computer-readable storage medium.
[0089] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the above method.
[0090] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0091] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0092] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A video restoration method based on optical flow guidance, characterized in that, include: Obtain the video frame sequence to be repaired; An optical flow estimation network is used to estimate the video frame sequence to be repaired to obtain a missing optical flow information sequence. The missing optical flow information sequence is then filled in based on the local correlation between optical flows to obtain complete optical flow information. Efficient self-attention mechanisms are introduced into the temporal transformer structure and the spatial transformer structure, respectively. The video frame sequence to be repaired and the complete optical flow information are guided by the complete optical flow information from the global temporal and global spatial dimensions. Features with high similarity to the missing region are used for repair operation to obtain preliminary repair results. Based on the local feature correlation of the preliminary repair results, the preliminary repair results are adjusted using the encoding and decoding structure to obtain the final repaired video; The incomplete optical flow information sequence is input into an improved U-net network to obtain complete optical flow information. The improved U-net network is constructed by adding P3D blocks to the coding layer of the U-net network via residual connections. The result of the P3D block convolution operation on the input of the improved U-net network's coding layer is added to the corresponding coding layer's output to obtain local clustering features. The result of dilated convolution on these local clustering features is concatenated with the corresponding local clustering features, and the concatenated result is decoded through the decoding layer of the improved U-net network. The efficient self-attention mechanism is specifically as follows: DW convolution is introduced on the multi-head attention mechanism, which simulates the interaction between different heads, and an instance normalization layer is introduced after the normalized exponential function to enhance the diversity of single-head attention.
2. The video restoration method based on optical flow guidance as described in claim 1, characterized in that, A convolutional neural network is used to extract features from the video frame sequence to be repaired, obtaining low-level features. These low-level features and complete optical flow information are then patched and embedded to obtain image features and optical flow features. An efficient self-attention mechanism is introduced into the temporal transformer structure to operate on the image features, specifically: , , in, For image features, LN() is the normalization operation, EMSA() is the efficient self-attention mechanism, and FFN() is the feedforward layer of the Transformer. The output of the time transformer structure to introduce an efficient self-attention mechanism.
3. The video restoration method based on optical flow guidance as described in claim 2, characterized in that, A multilayer sensing mechanism is used to find the similarity between the output of the temporal transformer structure with an efficient self-attention mechanism and the optical flow features. The optical flow features are weighted according to the similarity to obtain enhanced optical flow information. Then, the enhanced optical flow information and the output of the temporal transformer structure with an efficient self-attention mechanism are concatenated. The result of the concatenation operation is input into the spatial transformer structure with an efficient self-attention mechanism for repair.
4. The video restoration method based on optical flow guidance as described in claim 3, characterized in that, The result of the cascaded operation is input into a spatial transformer structure with an efficient self-attention mechanism for repair, specifically: Where Rt is the result of cascading the enhanced optical flow information with the output of the time-varying transformer structure that incorporates an efficient self-attention mechanism, LN() is the normalization operation, EMSA() is the efficient self-attention mechanism, and FFN() is the feedforward layer of the Transformer. The output of a spatial Transformer structure that introduces an efficient self-attention mechanism.
5. The video restoration method based on optical flow guidance as described in claim 1, characterized in that, Based on the mask of the missing area of the video frame, the U-net-type network is used to reshape the preliminary repair result to obtain the video with local structural adjustment; After decoding the video filling result after local structural adjustments, the final repaired video is obtained.
6. A video restoration system based on optical flow guidance, characterized in that, include: Acquisition module: Acquires the video frame sequence to be repaired; Optical flow information module: An optical flow estimation network is used to estimate the video frame sequence to be repaired to obtain a fragmented optical flow information sequence. Based on the local correlation between optical flows, the fragmented optical flow information sequence is filled to obtain complete optical flow information. The fragmented optical flow information sequence is then input into an improved U-net network to obtain complete optical flow information. The improved U-net network is constructed by adding P3D blocks to the coding layer of a U-net network via residual connections. The result of the P3D block convolution operation on the input of the improved U-net network's coding layer is added to the corresponding coding layer's output to obtain local clustering features. The result of dilated convolution on these local clustering features is concatenated with the corresponding local clustering features. The concatenated result is then decoded by the decoding layer of the improved U-net network. Preliminary Repair Module: Efficient self-attention mechanisms are introduced into both the temporal and spatial transformer structures. Guided by the complete optical flow information, the video frame sequence to be repaired and the complete optical flow information are processed from both global temporal and global spatial dimensions, using features with high similarity to the missing regions for repair operations, resulting in preliminary repair results. Specifically, the efficient self-attention mechanism involves introducing DW convolution on top of the multi-head attention mechanism. This DW convolution simulates the interaction between different heads, and an instance normalization layer is introduced after the normalized exponential function to enhance the diversity of single-head attention. Final Repair Module: Based on the local feature correlation of the preliminary repair results, the module adjusts the preliminary repair results using the encoding and decoding structure to obtain the final repaired video.
7. A computer device, characterized in that, include: The computer device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform a video restoration method based on optical flow as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs a video restoration method based on optical flow as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Aerial video classification method based on space-time multi-scale Transform
CN115223082A
Temporal feature alignment network for video inpainting
US20220284552A1