Video repair method and device
By introducing deformation processing and spatial time attention operation in the video repair method, the problem that existing methods cannot take into account videos of different motion amplitudes is solved, and the adaptive repair effect of videos with different motion amplitudes is achieved.
Patent Information
- Application Number
- CN202111223609.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-10-20
AI Technical Summary
The existing video repair method based on Transformer cannot be applied to videos with small and large motion amplitudes at the same time, resulting in unsatisfactory repair results.
By obtaining the image characteristics of each image frame of the video to be repaired, deformation information is calculated and deformation processing is performed, combining spatial attention and temporal attention operations, repaired image characteristics are fused to generate the repaired video.
Adaptive repair of videos with different motion amplitudes is achieved, and the repair effect is improved. It can be applied to videos with small motion amplitudes and large motion amplitudes at the same time.
Smart Images

Figure CN113962916B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of video processing, and in particular to a video restoration method and device. Background Art
[0002] Video restoration is to repair unknown (damaged) areas based on information from known areas of the video to obtain a reconstructed area, so that the restored video is visually coherent and natural. That is, after deleting the pixels that need to be repaired, the time and space gaps are filled with reasonable content in the video to reconstruct a reasonable background. Video restoration can help many video editing and restoration tasks, such as removing unwanted objects, scratch or damage recovery, and repositioning. More importantly, video restoration can also be used in conjunction with augmented reality (AR) to provide a better visual experience.
[0003] At present, the video restoration method is basically based on the spatiotemporal joint Transformer. Transformer is a model proposed based on the self-attention mechanism to solve the Seq2Seq problem. However, the video restoration method based on Transformer treats all input videos equally, but different videos have different spatiotemporal information. For example, the background motion of some videos is not large, while the background motion of some videos is very intense. If they are treated equally, the two modes of video cannot be compatible, resulting in unsatisfactory restoration effect of one mode of video. Summary of the invention
[0004] The present disclosure provides a video restoration method and device to at least solve the problem that the video restoration method in the related art cannot be applied to both videos with small motion amplitude and videos with large motion amplitude.
[0005] According to a first aspect of an embodiment of the present disclosure, a video restoration method is provided, comprising: obtaining a video to be restored; obtaining deformation information between each image frame of the video to be restored based on image features of each image frame of the video to be restored; performing deformation processing on each image frame of the video to be restored based on the deformation information; for each image frame after the deformation processing, performing the following operations: obtaining a first restoration image feature of the current image frame by performing a spatial attention operation on the current image frame based on the image features of the current image frame, obtaining a second restoration image feature of the current image frame by performing a temporal attention operation on the current image frame based on the image features of each image frame, fusing the first restoration image feature and the second restoration image feature to obtain a fused restoration image feature of the current image frame; and obtaining a restored video based on the fused restoration image feature of each image frame after the deformation processing.
[0006] Optionally, fusing the first repaired image feature and the second repaired image feature to obtain the fused repaired image feature of the current image frame includes: determining the weights of the first repaired image feature and the second repaired image feature respectively based on the deformation information; and fusing the first repaired image feature and the second repaired image feature based on the weights of the first repaired image feature and the second repaired image feature to obtain the fused repaired image feature of the current image frame.
[0007] Optionally, obtaining the deformation information between each image frame of the video to be repaired based on the image features of each image frame of the video to be repaired includes: obtaining the similarity information based on the sub-image features of each image block in each image frame, where the image block is one of a predetermined number of image blocks obtained by respectively splitting each image frame; and obtaining the deformation information based on the similarity information.
[0008] Optionally, obtaining the similarity information based on the sub-image features of each image block in each image frame includes: obtaining the image features of each image frame of the video to be repaired; obtaining the Query matrix, the Key matrix, and the Value matrix based on the image features of each image frame; and obtaining the similarity matrix as the similarity information based on the similarity between the Query sub-matrix of each image block in the Query matrix and the Key sub-matrix of each image block in the Key matrix.
[0009] Optionally, obtaining the similarity matrix based on the similarity between the Query sub-matrix of each image block in the Query matrix and the Key sub-matrix of each image block in the Key matrix includes: obtaining the mask of each image block corresponding to the Query matrix and the mask of each image block corresponding to the Key matrix; when the masks of the image block in the Query matrix and the image block in the Key matrix are both 1, taking the product of the Query sub-matrix of the image block in the Query matrix and the Key sub-matrix of the image block in the Key matrix as the similarity between the two image blocks; when either the mask of the image block in the Query matrix or the mask of the image block in the Key matrix is 0, taking 0 as the similarity between the two image blocks; and merging the similarities after normalization processing to obtain the similarity matrix.
[0010] Optionally, obtaining the deformation information based on the similarity information includes: obtaining the deformation parameter matrix based on the similarity matrix as the deformation information.
[0011] Optionally, performing deformation processing on each image frame of the video to be repaired based on the deformation information includes: performing deformation processing on the Key matrix and the Value matrix based on the deformation parameter matrix to obtain the deformed Key matrix and the deformed Value matrix.
[0012] Optionally, based on the deformation parameter matrix, perform deformation processing on the Key matrix and the Value matrix to obtain the deformed Key matrix and the deformed Value matrix, including: for each image block, respectively deform the Key sub-matrix of the image block in the Key matrix and the Value sub-matrix of the image block in the Value matrix through the deformation parameter sub-matrix of the image block in the deformation parameter matrix, to obtain the deformed Key sub-matrix and the deformed Value sub-matrix of the image block; splice the deformed Key sub-matrices to obtain the deformed Key matrix; splice the deformed Value sub-matrices to obtain the deformed Value matrix.
[0013] Optionally, perform a spatial attention operation on the current image frame based on the image features of the current image frame to obtain the first restored image feature of the current image frame, including: obtaining the first restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix.
[0014] Optionally, perform a temporal attention operation on the current image frame based on the image features of each image frame to obtain the second restored image feature of the current image frame, including: obtaining the second restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix.
[0015] Optionally, the video restoration method is performed by a pre-trained video restoration model. The video restoration model includes an encoder, a deformation parameter prediction network, a first attention network, a second attention network, a fusion network, and a decoder. Among them, obtaining the image features of each image frame of the video to be restored includes: inputting the video to be restored into the encoder to obtain the feature matrix of each image frame of the video to be restored as the image features; among them, obtaining the deformation parameter matrix based on the similarity matrix includes: inputting the similarity matrix into the deformation parameter prediction network to obtain the deformation parameter matrix; among them, obtaining the first restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix includes: inputting the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix into the corresponding processing sub-network in the first attention network to obtain the first restored image feature of the current image frame, where the first attention network is used to perform spatial attention operation on the current image frame based on the image features of the current image frame; among them, obtaining the second restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix includes: inputting the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix into the second attention network to obtain the second restored image feature of the current image frame, where the second attention network is used to perform temporal attention operation on the current image frame based on the image features of each image frame; among them, fusing the first restored image feature and the second restored image feature to obtain the fused restored image feature of the current image frame includes: inputting the first restored image feature, the second restored image feature, and the deformation parameter matrix into the fusion network to obtain the fused restored image feature of the current image frame; among them, obtaining the restored video based on the fused restored image features of each image frame after deformation processing includes: inputting the fused restored image features of each image frame after deformation processing into the decoder to obtain the restored video.
[0016] Optionally, the video restoration model is trained through the following operations: obtaining a training sample set, where the training sample set includes a plurality of training videos and the clear video corresponding to each training video; inputting the training videos into an encoder to obtain the feature matrix of each image frame of the training videos; based on the feature matrix of each image frame, obtaining a Query matrix, a Key matrix, and a Value matrix; based on the similarity between the Query sub-matrix of each image patch in the Query matrix and the Key sub-matrix of each image patch in the Key matrix, obtaining a similarity matrix, where the image patch is one of a predetermined number of image patches obtained by splitting each image frame respectively; inputting the similarity matrix into a deformation parameter prediction network to obtain a deformation parameter matrix; based on the deformation parameter matrix, performing deformation processing on the Key matrix and the Value matrix to obtain a deformed Key matrix and a deformed Value matrix; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into a first attention network to obtain a restored first feature matrix, where the first feature matrix includes the first restored image features of each image frame; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into a second attention network to obtain a restored second feature matrix, where the second feature matrix includes the second restored image features of each image frame; inputting the first feature matrix, the second feature matrix, and the deformation parameter matrix into a fusion network to obtain an estimated restored feature matrix of each image frame, where the estimated restored feature matrix includes the fused restored image features of each image frame; inputting the estimated restored feature matrix into a decoder to obtain an estimated restored video; based on the estimated restored video and the clear video corresponding to the training video, adjusting the parameters of the encoder, the deformation parameter prediction network, the first attention network, the second attention network, the fusion network, and the decoder to train the video restoration model.
[0017] According to a second aspect of the embodiments of the present disclosure, a video restoration device is provided, including: a video acquisition unit configured to acquire a video to be restored; a deformation information acquisition unit configured to obtain deformation information between each image frame of the video to be restored based on the image features of each image frame of the video to be restored; a deformation processing unit configured to perform deformation processing on each image frame of the video to be restored based on the deformation information; a restoration unit configured to perform the following operations on each image frame after the deformation processing: obtaining a first restored image feature of the current image frame by performing a spatial attention operation on the current image frame based on the image features of the current image frame, obtaining a second restored image feature of the current image frame by performing a temporal attention operation on the current image frame based on the image features of each image frame, and fusing the first restored image feature and the second restored image feature to obtain a fused restored image feature of the current image frame; a restored video acquisition unit configured to obtain the restored video based on the fused restored image features of each image frame after the deformation processing.
[0018] Optionally, the restoration unit is further configured to determine the weight of the first restored image feature and the weight of the second restored image feature respectively based on the deformation information; and fuse the first restored image feature and the second restored image feature based on the weight of the first restored image feature and the weight of the second restored image feature to obtain the fused restored image feature of the current image frame.
[0019] Optionally, the deformation information acquisition unit is further configured to obtain similarity information based on the sub-image features of each image block in each image frame, where the image block is one of a predetermined number of image blocks obtained by respectively splitting each image frame; and obtain the deformation information based on the similarity information.
[0020] Optionally, the deformation information acquisition unit is further configured to obtain the image features of each image frame of the video to be restored; obtain a Query matrix, a Key matrix, and a Value matrix based on the image features of each image frame; and obtain a similarity matrix based on the similarity between the Query sub-matrix of each image block in the Query matrix and the Key sub-matrix of each image block in the Key matrix as the similarity information.
[0021] Optionally, the deformation information acquisition unit is further configured to obtain the mask of each image patch corresponding to the Query matrix and the mask of each image patch corresponding to the Key matrix; when the masks of the image patches in the Query matrix and the masks of the image patches in the Key matrix are both 1, take the product of the Query sub-matrix of the image patch in the Query matrix and the Key sub-matrix of the image patch in the Key matrix as the similarity between the two image patches; when any one of the masks of the image patches in the Query matrix and the masks of the image patches in the Key matrix is 0, take 0 as the similarity between the two image patches; perform normalization processing on the similarities and then merge them to obtain a similarity matrix.
[0022] Optionally, the deformation information acquisition unit is further configured to obtain a deformation parameter matrix based on the similarity matrix as the deformation information.
[0023] Optionally, the deformation processing unit is further configured to perform deformation processing on the Key matrix and the Value matrix based on the deformation parameter matrix to obtain a deformed Key matrix and a deformed Value matrix.
[0024] Optionally, for each image patch, the deformation processing unit is further configured to deform the Key sub-matrix of the image patch in the Key matrix and the Value sub-matrix of the image patch in the Value matrix respectively through the deformation parameter sub-matrix of the image patch in the deformation parameter matrix to obtain a deformed Key sub-matrix and a deformed Value sub-matrix of the image patch; splice the deformed Key sub-matrices to obtain a deformed Key matrix; splice the deformed Value sub-matrices to obtain a deformed Value matrix.
[0025] Optionally, the restoration unit is further configured to obtain a first restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix.
[0026] Optionally, the restoration unit is further configured to obtain a second restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix.
[0027] Optionally, the video restoration method is performed by a pre-trained video restoration model, which includes an encoder, a deformation parameter prediction network, a first attention network, a second attention network, a fusion network, and a decoder. Among them, the deformation information acquisition unit is further configured to input the video to be restored into the encoder to obtain the feature matrix of each image frame of the video to be restored as the image feature; the deformation information acquisition unit is further configured to input the similarity matrix into the deformation parameter prediction network to obtain the deformation parameter matrix; among them, the restoration unit is further configured to input the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix into the corresponding processing sub-network in the first attention network to obtain the first restored image feature of the current image frame, where the first attention network is used to perform spatial attention operation on the current image frame based on the image feature of the current image frame; the restoration unit is further configured to input the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix into the second attention network to obtain the second restored image feature of the current image frame, where the second attention network is used to perform temporal attention operation on the current image frame based on the image feature of each image frame; the restoration unit is further configured to input the first restored image feature, the second restored image feature, and the deformation parameter matrix into the fusion network to obtain the fused restored image feature of the current image frame; the restored video acquisition unit is further configured to input the fused restored image feature of each image frame after deformation processing into the decoder to obtain the restored video.
[0028] Optionally, the video restoration model is trained through the following operations: obtaining a training sample set, where the training sample set includes a plurality of training videos and the clear video corresponding to each training video; inputting the training video into an encoder to obtain the feature matrix of each image frame of the training video; based on the feature matrix of each image frame, obtaining a Query matrix, a Key matrix, and a Value matrix; based on the similarity between the Query sub-matrix of each image patch in the Query matrix and the Key sub-matrix of each image patch in the Key matrix, obtaining a similarity matrix, where the image patch is one of a predetermined number of image patches obtained by splitting each image frame; inputting the similarity matrix into a deformation parameter prediction network to obtain a deformation parameter matrix; based on the deformation parameter matrix, performing a deformation process on the Key matrix and the Value matrix to obtain a deformed Key matrix and a deformed Value matrix; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into a first attention network to obtain a restored first feature matrix, where the first feature matrix includes the first restored image feature of each image frame; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into a second attention network to obtain a restored second feature matrix, where the second feature matrix includes the second restored image feature of each image frame; inputting the first feature matrix, the second feature matrix, and the deformation parameter matrix into a fusion network to obtain the estimated restored feature matrix of each image frame, where the estimated restored feature matrix includes the fused restored image feature of each image frame; inputting the estimated restored feature matrix into a decoder to obtain the estimated restored video; based on the estimated restored video and the clear video corresponding to the training video, adjusting the parameters of the encoder, the deformation parameter prediction network, the first attention network, the second attention network, the fusion network, and the decoder to train the video restoration model.
[0029] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the instructions to implement the video restoration method according to the present disclosure.
[0030] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are run by at least one processor, causing at least one processor to execute the video restoration method as described above according to the present disclosure.
[0031] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including computer instructions, where the computer instructions, when executed by a processor, implement the video restoration method according to the present disclosure.
[0032] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0033] According to the video restoration method and device of the present disclosure, by performing spatial attention operation on the current image frame based on the image features of the current image frame, the first restored image feature of the current image frame is obtained, that is, key information is searched from the current image frame, and by performing temporal attention operation on the current image frame based on the image features of each image frame, the second restored image feature of the current image frame is obtained, that is, key information is searched from other image frames. Then, the first restored image feature and the second restored image feature are fused to obtain the fused restored image feature of the current image frame, which can control the flow of spatio-temporal information, and different spatio-temporal weights are assigned to different types of motion videos during fusion to obtain the optimal restored video. Therefore, the present disclosure solves the problem that the video restoration methods in the related art cannot be applied to videos with small motion amplitude and videos with large motion amplitude at the same time.
[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.
[0036] Figure 1 is a schematic diagram of an implementation scenario of a video restoration method according to an exemplary embodiment of the present disclosure;
[0037] Figure 2 is a flowchart of a video restoration method shown according to an exemplary embodiment;
[0038] Figure 3 is a schematic diagram of a structure applied to a video restoration method shown according to an exemplary embodiment;
[0039] Figure 4 is a schematic diagram of a structure of a homography estimator shown according to an exemplary embodiment;
[0040] Figure 5 is a schematic diagram of a structure of a deformation parameter prediction network shown according to an exemplary embodiment;
[0041] Figure 6 is a schematic diagram of a structure of a temporal-spatial domain selection module shown according to an exemplary embodiment;
[0042] Figure 7 is an attention map based on all image patches shown according to an exemplary embodiment;
[0043] Figure 8 It is a flowchart of a method for training a video restoration model shown according to an exemplary embodiment;
[0044] Figure 9 It is a schematic diagram of verification results shown according to an exemplary embodiment Figure 1 ;
[0045] Figure 10 It is a schematic diagram of verification results shown according to an exemplary embodiment Figure 2 ;
[0046] Figure 11 It is a schematic diagram of verification results shown according to an exemplary embodiment Figure 3 ;
[0047] Figure 12 It is a schematic diagram of verification results shown according to an exemplary embodiment Figure 4 ;
[0048] Figure 13 It is a block diagram of a video restoration device shown according to an exemplary embodiment;
[0049] Figure 14 It is a block diagram of a device for training a video restoration model shown according to an exemplary embodiment;
[0050] Figure 15 It is a block diagram of an electronic device 1500 according to an embodiment of the present disclosure. Detailed implementation manners
[0051] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0052] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0053] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations, namely, "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "executing at least one of Step 1 and Step 2" means the following three parallel situations: (1) executing Step 1; (2) executing Step 2; (3) executing Step 1 and Step 2.
[0054] Currently, video restoration methods based on spatio-temporal selection usually borrow valid content from adjacent image frames. There may be missing content in adjacent image frames. Moreover, when the image frames in a video are almost static, using the current methods often results in blurred results with few details. In this case, the video restoration problem can be degraded to a single-image restoration task, and the pixels in the known regions can provide structural or texture cues to make the result smooth at the boundaries and natural in details.
[0055] To address the above problems, the present disclosure provides a video restoration method that enables a video restoration model to adaptively handle two situations: videos where the foreground or background is almost static; videos where both the foreground and background are moving rapidly, that is, it can be applicable to videos with small motion amplitudes and videos with large motion amplitudes at the same time. Hereinafter, a video containing both an athlete running outdoors and an athlete running on a treadmill will be used as an example for illustration.
[0056] Figure 1 is a schematic diagram of an implementation scenario showing the video restoration method according to an exemplary embodiment of the present disclosure. As Figure 1 described, the implementation scenario includes a server 100, a user terminal 110, and a user terminal 120. Among them, the number of user terminals is not limited to 2, and includes but is not limited to devices such as mobile phones and personal computers. The user terminal can install an application program for video restoration. The server can be a single server, a server cluster composed of several servers, or a cloud computing platform or a virtualization center.
[0057] After the user terminal 110 or the user terminal 120 obtains the video to be repaired (a video that simultaneously includes an athlete running outdoors and an athlete running on a treadmill), it sends the video to be repaired to the server 100. After the server 100 inputs the video to be repaired into the encoder and obtains the feature matrix of each image frame of the video to be repaired, based on the feature matrix of each image frame, it obtains the Query matrix, the Key matrix, and the Value matrix. Then, based on the similarity between the Query sub-matrix of each image patch in the Query matrix and the Key sub-matrix of each image patch in the Key matrix, it obtains the similarity matrix, where an image patch is one of a predetermined number of image patches obtained by splitting each image frame; the server 100 inputs the similarity matrix into the deformation parameter prediction network, and after obtaining the deformation parameter matrix, based on the deformation parameter matrix, it performs a deformation process on the Key matrix and the Value matrix to obtain the deformed Key matrix and the deformed Value matrix; then, the server 100 inputs the Query matrix, the deformed Key matrix, and the deformed Value matrix into the first attention network to obtain the repaired first feature matrix, where the first feature matrix is obtained based on the reference information between the sub-matrices of the same image frame in the Query matrix, the deformed Key matrix, and the deformed Value matrix, and inputs the Query matrix, the deformed Key matrix, and the deformed Value matrix into the second attention network to obtain the repaired second feature matrix, where the second feature matrix is obtained based on the reference information between the sub-matrices of different image frames in the Query matrix, the deformed Key matrix, and the deformed Value matrix; finally, it inputs the first feature matrix, the second feature matrix, and the deformation parameter matrix into the fusion network, and after obtaining the repaired feature matrix of each image frame, it inputs the repaired feature matrix into the decoder to obtain the repaired video. Although the video to be repaired is a video that simultaneously includes an athlete running outdoors and an athlete running on a treadmill, the repaired video obtained by the video repair method of the present disclosure still achieves a good repair effect.
[0058] Next, a video repair method and apparatus according to an exemplary embodiment of the present disclosure will be described in detail with reference to Figures 2 to 14 A flowchart of a video repair method according to an exemplary embodiment is shown in
[0059] Figure 2 As shown in Figure 2 shown,
[0060] In step S201, a video to be repaired is obtained. The video to be repaired can be temporarily obtained through a terminal camera or a locally stored video, and the present disclosure does not limit this.
[0061] In step S202, based on the image features of each image frame of the video to be repaired, the deformation information between each image frame of the video to be repaired is obtained by using the deformation parameter prediction network. It should be noted that the deformation information between every two image frames needs to be obtained here.
[0062] According to an exemplary embodiment of the present disclosure, obtaining the deformation information between each image frame of the video to be repaired based on the image features of each image frame of the video to be repaired includes: obtaining similarity information based on the sub-image features of each image block in each image frame, where the image block is one of a predetermined number of image blocks obtained by respectively splitting each image frame; and obtaining deformation information based on the similarity information. According to this embodiment, the deformation information can be obtained conveniently and quickly.
[0063] According to an exemplary embodiment of the present disclosure, obtaining similarity information based on the sub-image features of each image block in each image frame includes: obtaining the image features of each image frame of the video to be repaired; obtaining a Query matrix, a Key matrix, and a Value matrix based on the image features of each image frame; and obtaining a similarity matrix based on the similarity between each Query sub-matrix of each image block in the Query matrix and each Key sub-matrix of each image block in the Key matrix as the similarity information. A conventional video encoder can be used for the encoder, and the present disclosure does not make any limitations.
[0064] According to an exemplary embodiment of the present disclosure, obtaining a Query matrix, a Key matrix, and a Value matrix based on the image features of each image frame may include: respectively splitting the feature matrix of each image frame into a predetermined number of matrices, where the predetermined number of matrices corresponds to the predetermined number of image blocks; splicing all the split matrices in the channel dimension to obtain a spliced matrix; and respectively inputting the spliced matrix into corresponding convolutions to obtain a Query matrix, a Key matrix, and a Value matrix. Through this embodiment, the three matrices can be obtained conveniently and quickly.
[0065] For example, the video to be repaired consists of T image frames, and each image frame is a picture with a width of W and a height of H. Figure 3 is a schematic structural diagram applied to a video repair method shown in an exemplary embodiment, as Figure 3 shown, Encoder is the encoder, which encodes each image frame of the video to be repaired. After the t-th image frame passes through the encoder Encoder, a feature matrix f T , f T is obtained, and the dimension of f is c×h×w, where c is the number of channels, h corresponds to H of the picture, and w corresponds to W of the picture. Then each image frame is split into a predetermined number of image blocks, such as each image frame is split into N pimage patches, that is, f T is divided into N p sub - feature matrices. Figure 3 where N p = 4, the width and height of each image patch are w / 2 and h / 2 respectively, then the total number of image patches in the entire video to be repaired is T * N p image patches. Then, the feature matrices of T * N p image patches are concatenated in the channel dimension to obtain the feature matrix F of the video to be repaired (i.e., the above - mentioned concatenated matrix). The concatenated matrix is respectively input into the corresponding convolutions to obtain the Query matrix, the Key matrix, and the Value matrix. The corresponding convolutions are not limited in this disclosure and can be 1×1 convolutions respectively.
[0066] According to an exemplary embodiment of the present disclosure, based on the similarity between the Query sub - matrix of each image patch in the Query matrix and the Key sub - matrix of each image patch in the Key matrix, a similarity matrix is obtained, including: obtaining the mask of each image patch corresponding to the Query matrix and the mask of each image patch corresponding to the Key matrix; when the masks of the image patches in the Query matrix and the Key matrix are both 1, taking the product of the Query sub - matrix of the image patch in the Query matrix and the Key sub - matrix of the image patch in the Key matrix as the similarity between the two image patches; when any one of the masks of the image patches in the Query matrix and the Key matrix is 0, taking 0 as the similarity between the two image patches; normalizing and then combining the similarities to obtain the similarity matrix. Through this embodiment, the similarity matrix can be obtained conveniently and quickly.
[0067] For example, in the above - mentioned embodiment, the similarity matrix can be implemented by Figure 3 the deformation homography estimator (DePtH) in, and the DePtH is a deformation homography estimator based on Transformer. Figure 4 is a schematic structural diagram of a homography estimator shown according to an exemplary embodiment. As Figure 4 shown, after F is input into DePtH, three matrices required for three self - attention operations are respectively generated through three 1×1 convolutions: the Query matrix, the Key matrix, and the Value matrix. For example, the patch encoder generates the Query matrix, the Key matrix, and the Value matrix after deforming F, where the Query matrix, the Key matrix, and the Value matrix are obtained by Q, K, V = M q (f i ), M k (f i ), M v (f i)It is determined that M is a 1×1 convolution. Then, these three matrices are input into Patch Matching (abbreviated as PM) to obtain a similarity matrix. The specific operation of PM is as follows: The similarity between each patch in the Query matrix and each patch in the Key matrix is calculated by multiplying the feature matrix of each patch in the Query matrix by the feature matrix of each patch in the Key matrix. For example, as shown in formula (1), C is the cosine similarity matrix calculated between f k (i) T and f q (j), where f k (i) T is the transpose of the feature matrix of each patch in the Key matrix after channel normalization, and f q (j) is the feature matrix of each patch in the Query matrix after channel normalization. Since there are a total of T·N p patches in the video to be modified, for the i-th patch in the Key matrix and the j-th patch in the Query matrix, only the parts containing valid information that are not covered in both patches can be calculated. For this purpose, a downsampled binary mask is used to filter the results, that is, only when m k (i) and m q (j) are both 1, the similarity matrix of the two patches is calculated.
[0068]
[0069] According to an exemplary embodiment of the present disclosure, based on the similarity information, deformation information is obtained, including: based on the similarity matrix, a deformation parameter matrix is obtained as the deformation information.
[0070] For example, as Figure 4 shown, the output of PM is input into a deformed parameter predictor (Deformed Transformer estimator), that is, the above-mentioned deformed parameter prediction network, to obtain a predicted bias coefficient matrix θ (i.e., the above-mentioned deformed parameter matrix). θ contains the deformation parameters between each patch in the Query matrix and each patch in the Key matrix. θ is a matrix with a dimension of T·Np×2×3, where for T·Np patches, each patch has a matrix θ' composed of 2×3 = 6 numbers as the transformation parameter sub-matrix of each patch.
[0071] According to an exemplary embodiment of the present disclosure, the deformation parameter prediction network includes a convolutional layer, a pooling layer, and a fully connected layer. Inputting the similarity information into the deformation parameter prediction network to obtain deformation information, including: inputting the similarity information into the convolutional layer to obtain the convolved information; inputting the convolved information into the pooling layer to obtain the pooled information; inputting the pooled information into the fully connected layer to obtain the deformation information. Through this embodiment, a simple structure of the deformation parameter prediction network is given, which can conveniently and quickly obtain the deformation parameter matrix.
[0072] For example, Figure 5 FIG. is a schematic structural diagram of a deformation parameter prediction network shown according to an exemplary embodiment, as Figure 5 shown, the left and right sides of the network respectively represent the input and output. The network can be composed of 3×3 convolutions with a stride of 2, where the number of channels in the middle layer is 128, and Average Pool represents average pooling over all positions h×w. Additionally, except for the last layer decoder, ReLU is used as the non-linear activation function after all convolutional layers.
[0073] In step S203, based on the deformation information, perform deformation processing on each image frame of the video to be repaired;
[0074] According to an exemplary embodiment of the present disclosure, based on the deformation information, performing deformation processing on each image frame of the video to be repaired includes: based on the deformation parameter matrix, performing deformation processing on the Key matrix and the Value matrix to obtain the deformed Key matrix and the deformed Value matrix. The above θ acts on the Key matrix and the Value matrix respectively, which can align the Key matrix and the Value matrix with reference to the Query matrix, and generate the deformed Key matrix (Deformed K) and the deformed Value matrix (Deformed V) respectively, while the Query matrix remains unchanged.
[0075] According to an exemplary embodiment of the present disclosure, based on the deformation parameter matrix, performing deformation processing on the Key matrix and the Value matrix to obtain the deformed Key matrix and the deformed Value matrix includes: for each image block, respectively deform the Key sub-matrix of the image block in the Key matrix and the Value sub-matrix of the image block in the Value matrix through the deformation parameter sub-matrix of the image block in the deformation parameter matrix to obtain the deformed Key sub-matrix and the deformed Value sub-matrix of the image block; splice the deformed Key sub-matrices to obtain the deformed Key matrix; splice the deformed Value sub-matrices to obtain the deformed Value matrix. Through this embodiment, rotation is performed in units of image blocks, and more accurate deformed results can be obtained.
[0076] For example, asFigure 4 The shown θ is a matrix with a dimension of T·Np×2×3. Among them, for T·Np image patches, each image patch has a matrix θ′ composed of 2×3 = 6 numbers as the transformation parameter sub-matrix of each image patch. The Key matrix and the Value matrix respectively perform rotation transformation on each image patch according to the corresponding θ′ to obtain the transformed Key matrix (Deformed K, that is, the deformed Key matrix mentioned above), and the transformed Value matrix (Deformed V, that is, the deformed Value matrix mentioned above).
[0077] In step S204, for each image frame after deformation processing, the following operations are performed: By performing spatial attention operation on the current image frame based on the image features of the current image frame, the first repaired image feature of the current image frame is obtained. By performing temporal attention operation on the current image frame based on the image features of each image frame, the second repaired image feature of the current image frame is obtained. The first repaired image feature and the second repaired image feature are fused to obtain the fused repaired image feature of the current image frame.
[0078] According to an exemplary embodiment of the present disclosure, obtaining the first repaired image feature of the current image frame by performing a spatial attention operation on the current image frame based on the image features of the current image frame includes: Based on the image features of the current image frame after deformation processing, using a first attention network to obtain the first repaired image feature of the current image frame, where the first attention network is used to perform a spatial attention operation on the current image frame based on the image features of the current image frame. For example, in some videos, the foreground motion amplitude is not large, such as the foreground moving on the same background. At this time, when repairing the current image frame, key information should be found from the current image frame.
[0079] According to an exemplary embodiment of the present disclosure, obtaining the first repaired image feature of the current image frame by performing a spatial attention operation on the current image frame based on the image features of the current image frame includes: Based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix, the first repaired image feature of the current image frame is obtained.
[0080] According to an exemplary embodiment of the present disclosure, the first attention network includes T processing sub-networks, where T represents the number of image frames of the video to be repaired. Among them, the Query matrix, the deformed Key matrix, and the deformed Value matrix are input into the first attention network to obtain the repaired first feature matrix, including: inputting the sub-matrix of the i-th image frame in the Query matrix, the sub-matrix of the i-th image frame in the deformed Key matrix, and the sub-matrix of the i-th image frame in the deformed Value matrix into the i-th processing sub-network to obtain the feature matrix of the i-th image frame, where i ∈ {1, 2... T}; splicing the feature matrices of all image frames to obtain the repaired first feature matrix. Through this embodiment, each image frame searches for key information from its own image frame, which helps to obtain the repair of videos with small background change amplitudes.
[0081] For example, Figure 6 is a schematic structural diagram of a time-domain and spatial-domain selection module shown according to an exemplary embodiment, as Figure 6 shown. The Spatial Transformer can be used as the above-mentioned first attention network. The dashed box inside represents the processing sub-network. After the Query matrix, the deformed Key matrix, and the deformed Value matrix are input into the Spatial Transformer, the first processing sub-network processes the sub-matrix of the first image frame in the Query matrix, the sub-matrix of the first image frame in the deformed Key matrix, and the sub-matrix of the first image frame in the deformed Value matrix to obtain the feature matrix of the first image frame, and so on, to obtain the feature matrices of all image frames. Splicing the feature matrices of all image frames gives T*N p .
[0082] According to an exemplary embodiment of the present disclosure, by performing a time attention operation on the current image frame based on the image features of each image frame, the second repaired image feature of the current image frame is obtained, including: based on the image features of each image frame after deformation processing, using the second attention network to obtain the second repaired image feature of the current image frame, where the second attention network is used to perform a time attention operation on the current image frame based on the image features of each image frame. For example, in some videos, the foreground moves on different backgrounds and the background movement amplitude is very large. At this time, when repairing the current image frame, more key information should be searched from the front and back frames of the current image frame, that is, search for information from different time dimensions.
[0083] According to an exemplary embodiment of the present disclosure, obtaining a second restored image feature of the current image frame by performing a temporal attention operation on the current image frame based on the image features of each image frame includes: obtaining the second restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix.
[0084] According to an exemplary embodiment of the present disclosure, the second attention network includes T processing sub-networks, where T represents the number of image frames of the video to be restored. Among them, inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into the second attention network to obtain a restored second feature matrix includes: inputting the sub-matrix of the j-th image frame in the Query matrix, the sub-matrices of image frames other than the j-th image frame in the deformed Key matrix, and the sub-matrices of image frames other than the j-th image frame in the deformed Value matrix into the j-th processing sub-network to obtain the feature matrix of the j-th image frame, where j ∈ {1, 2... T}; splicing the feature matrices of all image frames to obtain a restored second feature matrix. Through this embodiment, each image frame searches for key information from other image frames, which helps to obtain the restoration of videos with large background change amplitudes.
[0085] For example, as Figure 6 shown, a Temporal Transformer can be used as the above-mentioned second attention network. The dashed boxes inside represent processing sub-networks. After the Query matrix, the deformed Key matrix, and the deformed Value matrix are input into the Temporal Transformer, the first processing sub-network processes the sub-matrix of the first image frame in the Query matrix, the sub-matrices of image frames other than the first image frame in the deformed Key matrix, and the sub-matrices of image frames other than the first image frame in the deformed Value matrix to obtain the feature matrix of the first image frame, and so on, to obtain the feature matrices of all image frames. Splicing the feature matrices of all image frames gives T*N p .
[0086] According to an exemplary embodiment of the present disclosure, fusing the first restored image feature and the second restored image feature to obtain the fused restored image feature of the current image frame includes: determining the weight of the first restored image feature and the weight of the second restored image feature respectively based on the deformation information; fusing the first restored image feature and the second restored image feature based on the weight of the first restored image feature and the weight of the second restored image feature to obtain the fused restored image feature of the current image frame.
[0087] According to an exemplary embodiment of the present disclosure, the first restored image feature of each image frame, the second restored image feature of each image frame, and the deformation parameter matrix can be input into the fusion network to obtain the restored feature matrix of each image frame.
[0088] According to an exemplary embodiment of the present disclosure, the fusion network includes a first processing sub-network, a second processing sub-network, and a fusion sub-network. Among them, inputting the first restored image feature of each image frame, the second restored image feature of each image frame, and the deformation parameter matrix into the fusion network to obtain the restored feature matrix of each image frame includes: inputting the first restored image feature of each image frame and the deformation parameter matrix into the first processing sub-network to obtain the weighted first restored image feature; inputting the second restored image feature of each image frame and the deformation parameter matrix into the second processing sub-network to obtain the weighted second restored image feature; inputting the weighted first restored image feature and the weighted second restored image feature into the fusion sub-network to obtain the restored feature matrix of each image frame. Through this embodiment, the fusion network assigns weights to the two feature matrices based on the deformation parameter matrix, that is, assigns different weights to the two feature matrices for different types of motion videos to fuse and obtain the restored feature matrix.
[0089] For example, as Figure 6 shown, Spatial Temporal Gated can be used as the above fusion network. The horizontal 1×1 convolution above represents the first processing sub-network, the horizontal 1×1 convolution below represents the second processing sub-network, and the vertical 1×1 convolution represents the fusion sub-network. Input the first feature matrix (T*N p ) and the deformation parameter matrix θ into the first processing sub-network to obtain the weighted first feature matrix. At the same time, input the second feature matrix (T*N p ) and the deformation parameter matrix θ into the second processing sub-network to obtain the weighted second feature matrix. Then, splice the weighted first feature matrix and the weighted second feature matrix to obtain the splicing matrix 2T*N p , input the obtained splicing matrix into the fusion sub-network to obtain the restored feature matrix T×N of each image frame p .
[0090] In step S205, a restored video is obtained based on the fusion and restored image features of each image frame after deformation processing.
[0091] According to an exemplary embodiment of the present disclosure, the video restoration method is executed by a pre-trained video restoration model. The video restoration model includes an encoder, a deformation parameter prediction network, a first attention network, a second attention network, a fusion network, and a decoder. Among them, obtaining the image features of each image frame of the video to be restored includes: inputting the video to be restored into the encoder to obtain a feature matrix of each image frame of the video to be restored as the image features; among them, obtaining the deformation parameter matrix based on the similarity matrix includes: inputting the similarity matrix into the deformation parameter prediction network to obtain the deformation parameter matrix; among them, obtaining the first restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix includes: inputting the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix into the corresponding processing sub-network in the first attention network to obtain the first restored image feature of the current image frame, where the first attention network is used to perform spatial attention operations on the current image frame based on the image features of the current image frame; among them, obtaining the second restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix includes: inputting the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix into the second attention network to obtain the second restored image feature of the current image frame, where the second attention network is used to perform temporal attention operations on the current image frame based on the image features of each image frame; among them, fusing the first restored image feature and the second restored image feature to obtain the fused and restored image feature of the current image frame includes: inputting the first restored image feature, the second restored image feature, and the deformation parameter matrix into the fusion network to obtain the fused and restored image feature of the current image frame; among them, obtaining the restored video based on the fused and restored image features of each image frame after deformation processing includes: inputting the fused and restored image features of each image frame after deformation processing into the decoder to obtain the restored video.
[0092] For example, the fused and repaired image features of each image frame after deformation processing are input into a decoder to obtain a repaired video, as can be Figure 3 The output of the spatio-temporal selection module (STS) shown is T matrices of c×h×w (i.e., the above-mentioned fused and repaired image features). This matrix can be input into a decoder of an RNN structure for decoding. The video obtained after decoding is the video repaired by the network for each image frame.
[0093] According to an exemplary embodiment of the present disclosure, the video repair model is trained through the following operations: obtaining a training sample set, where the training sample set includes a plurality of training videos and a clear video corresponding to each training video; inputting the training videos into an encoder to obtain a feature matrix of each image frame of the training videos; based on the feature matrix of each image frame, obtaining a Query matrix, a Key matrix, and a Value matrix; based on the similarity between the Query sub-matrices of each image patch in the Query matrix and the Key sub-matrices of each image patch in the Key matrix, obtaining a similarity matrix, where an image patch is one of a predetermined number of image patches obtained by respectively slicing each image frame; inputting the similarity matrix into a deformation parameter prediction network to obtain a deformation parameter matrix; based on the deformation parameter matrix, performing deformation processing on the Key matrix and the Value matrix to obtain a deformed Key matrix and a deformed Value matrix; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into a first attention network to obtain a repaired first feature matrix, where the first feature matrix includes the first repaired image features of each image frame; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into a second attention network to obtain a repaired second feature matrix, where the second feature matrix includes the second repaired image features of each image frame; inputting the first feature matrix, the second feature matrix, and the deformation parameter matrix into a fusion network to obtain an estimated repaired feature matrix of each image frame, where the estimated repaired feature matrix includes the fused and repaired image features of each image frame; inputting the estimated repaired feature matrix into a decoder to obtain an estimated repaired video; based on the estimated repaired video and the clear video corresponding to the training video, adjusting the parameters of the encoder, the deformation parameter prediction network, the first attention network, the second attention network, the fusion network, and the decoder to train the video repair model.
[0094] It should be noted that the present disclosure can also divide each image frame in the video into N p1 、N p2 …N pnimage patches. Based on each division case, the above repair process is performed, and the repaired features corresponding to each division case are stitched together to obtain the total features. In practical applications, the spatial image patches with the shape of can be first extracted from the query features of each image frame, and for each image frame, N (N p1 , N p2 … or N pn ) query features of image patches are obtained. Then, the self-attention in both space and time is performed T times on T image frames respectively, that is, the repair process in the above embodiment is performed.
[0095] In summary, since the local image patches on the current image frame and other image frames in the video contribute differently to different types of motion videos, the present disclosure controls the flow of spatio-temporal information through the spatio-temporal selection module (STS), and assigns different spatio-temporal weight values to the features obtained based on the current image frame and other image frames for different types of motion videos. For example, when the motion amplitude of the target in the video is relatively large, more weights will be assigned to the extraction of temporal domain information. When the motion amplitude of the target in the video is relatively small, more weights will be assigned to the extraction of spatial domain information, so that the features of the targeted image patches can be extracted and matched. Figure 7 is an attention map based on all image patches shown in an exemplary embodiment. As Figure 7 shown, it shows the attention weight matrix between every two image patches of all T image frames. Among them, the right diagonal part represents the spatial domain attention, which is the feature attention of the image patches within the frame, and the left diagonal part represents the temporal domain attention, which is the feature attention of the image patches between frames.
[0096] Figure 8 is a flowchart of a training method for a video repair model shown in an exemplary embodiment. As Figure 8 shown, the video repair model includes an encoder, a deformation parameter prediction network, a first attention network, a second attention network, a fusion network, and a decoder. The training method includes the following main steps:
[0097] In step S801, a training sample set is obtained, where the training sample set includes multiple training videos and the clear video corresponding to each training video.
[0098] In step S802, the training video is input into the encoder to obtain the feature matrix of each image frame of the training video.
[0099] In step S803, based on the feature matrix of each image frame, a Query matrix, a Key matrix, and a Value matrix are obtained.
[0100] In step S804, a similarity matrix is obtained based on the similarity between the Query sub-matrix of each image patch in the Query matrix and the Key sub-matrix of each image patch in the Key matrix, where the image patch is one of a predetermined number of image patches obtained by splitting each image frame respectively.
[0101] In step S805, the similarity matrix is input into the deformation parameter prediction network to obtain a deformation parameter matrix.
[0102] In step S806, based on the deformation parameter matrix, deformation processing is performed on the Key matrix and the Value matrix to obtain a deformed Key matrix and a deformed Value matrix.
[0103] In step S807, the Query matrix, the deformed Key matrix, and the deformed Value matrix are input into the first attention network to obtain a repaired first feature matrix, where the first feature matrix is obtained based on the reference information between the sub-matrices of the same image frame in the Query matrix, the deformed Key matrix, and the deformed Value matrix.
[0104] In step S808, the Query matrix, the deformed Key matrix, and the deformed Value matrix are input into the second attention network to obtain a repaired second feature matrix, where the second feature matrix is obtained based on the reference information between the sub-matrices of different image frames in the Query matrix, the deformed Key matrix, and the deformed Value matrix.
[0105] In step S809, the first feature matrix, the second feature matrix, and the deformation parameter matrix are input into the fusion network to obtain an estimated repaired feature matrix for each image frame.
[0106] In step S810, the estimated repaired feature matrix is input into the decoder to obtain an estimated repaired video.
[0107] In step S811, based on the estimated repaired video and the clear video corresponding to the training video, the parameters of the encoder, the deformation parameter prediction network, the first attention network, the second attention network, the fusion network, and the decoder are adjusted to train the video repair model.
[0108] According to an exemplary embodiment of the present disclosure, the first attention network includes T processing sub-networks, where T represents the number of image frames of the training video. Among them, the Query matrix, the deformed Key matrix, and the deformed Value matrix are input into the first attention network to obtain the restored first feature matrix, including: inputting the sub-matrix of the i-th image frame in the Query matrix, the sub-matrix of the i-th image frame in the deformed Key matrix, and the sub-matrix of the i-th image frame in the deformed Value matrix into the i-th processing sub-network to obtain the feature matrix of the i-th image frame, where i ∈ {1, 2... T}; splicing the feature matrices of all image frames to obtain the restored first feature matrix.
[0109] According to an exemplary embodiment of the present disclosure, the second attention network includes T processing sub-networks, where T represents the number of image frames of the training video. Among them, the Query matrix, the deformed Key matrix, and the deformed Value matrix are input into the second attention network to obtain the restored second feature matrix, including: inputting the sub-matrix of the j-th image frame in the Query matrix, the sub-matrices of the image frames other than the j-th image frame in the deformed Key matrix, and the sub-matrices of the image frames other than the j-th image frame in the deformed Value matrix into the j-th processing sub-network to obtain the feature matrix of the j-th image frame, where j ∈ {1, 2... T}; splicing the feature matrices of all image frames to obtain the restored second feature matrix.
[0110] According to an exemplary embodiment of the present disclosure, the fusion network includes a first processing sub-network, a second processing sub-network, and a fusion sub-network. Among them, the first feature matrix, the second feature matrix, and the deformation parameter matrix are input into the fusion network to obtain the estimated restored feature matrix of each image frame, including: inputting the first feature matrix and the deformation parameter matrix into the first processing sub-network to obtain the weighted first feature matrix; inputting the second feature matrix and the deformation parameter matrix into the second processing sub-network to obtain the weighted second feature matrix; inputting the weighted first feature matrix and the weighted second feature matrix into the fusion sub-network to obtain the estimated restored feature matrix of each image frame.
[0111] To prove the feasibility of the method of the present disclosure, the corresponding verification results are provided below.
[0112] Figure 9 It is a schematic diagram of the verification results shown according to an exemplary embodiment Figure 1 as Figure 9 shown, the method of the present disclosure exceeds the existing video restoration methods, such as the VINet, Deep-Flow, CPT, and STTN methods, on the publicly available academic datasets YouTube VOS and DAVIS.
[0113] Figure 10 is a schematic diagram of verification results shown according to an exemplary embodiment Figure 2 , Figure 10 showing the quantitative contribution degree of each sub-module in the model to the entire model. The last line shows the PSNR and SSIM evaluation metrics of the entire model on the publicly available academic datasets YouTubeVOS and DAVIS. The higher these two evaluation metrics are, the better the performance. The second, third, and fourth lines show the performance after removing STS, removing the spatial branch of STS, and removing the temporal branch of STS, respectively. It can be seen that the STS module has brought positive performance improvement to the model on both the YouTube dataset and the DAVIS dataset.
[0114] Figure 11 is a schematic diagram of verification results shown according to an exemplary embodiment Figure 3 , such as Figure 11 shown. The first column is the input image, the second column is the result without adding STS, and the last column is the result after adding STS. It can be seen that when STS is not added, the repaired image is relatively blurred and the lines are relatively curved; while after adding the STS spatio-temporal selection module, the repair result is significantly improved. That is, the model can reallocate attention weights among many frames according to video features.
[0115] Figure 12 is a schematic diagram of verification results shown according to an exemplary embodiment Figure 4 , such as Figure 12 shown, which is the visualization result of the attention distribution of the missing area learned by the module of the present disclosure. The spatial branch (Spatial Branch) and the temporal branch (Temporal Branch) are respectively the spatial frame with the highest attention weight and the four image frames with the highest attention in the temporal order obtained by sorting the attention weights from high to low among all T image frames. In order to complete the repair of the object in the target image frame, the model can track the moving object in the video in both the spatial and temporal dimensions, and the attention area is highlighted. The original image (origin image) is the input image frame. For example, in the left example, the input image frame is the 16th frame. It can be seen that the temporal branch (Temporal Branch) allocates more attention to several image frames around the target image frame, so as to find information that can fill the covered area near the mask of the adjacent image frames to fill the occluded area.
[0116] Figure 13 is a block diagram of a video restoration device shown according to an exemplary embodiment. Refer to Figure 13, the video repair device includes: a video acquisition unit 1301, a deformation information acquisition unit 1302, a deformation processing unit 1303, a repair unit 1304, and a repaired video acquisition unit 1305.
[0117] The video acquisition unit 1301 is configured to acquire the video to be repaired; the deformation information acquisition unit 1302 is configured to obtain the deformation information between each image frame of the video to be repaired based on the image features of each image frame of the video to be repaired; the deformation processing unit 1303 is configured to perform deformation processing on each image frame of the video to be repaired based on the deformation information; the repair unit 1304 is configured to perform the following operations on each image frame after the deformation processing: perform a spatial attention operation on the current image frame based on the image features of the current image frame to obtain the first repaired image feature of the current image frame, perform a temporal attention operation on the current image frame based on the image features of each image frame to obtain the second repaired image feature of the current image frame, and fuse the first repaired image feature and the second repaired image feature to obtain the fused repaired image feature of the current image frame; the repaired video acquisition unit 1305 is configured to obtain the repaired video based on the fused repaired image features of each image frame after the deformation processing.
[0118] According to an exemplary embodiment of the present disclosure, the repair unit 1304 is further configured to determine the weight of the first repaired image feature and the weight of the second repaired image feature respectively based on the deformation information; fuse the first repaired image feature and the second repaired image feature based on the weight of the first repaired image feature and the weight of the second repaired image feature to obtain the fused repaired image feature of the current image frame.
[0119] According to an exemplary embodiment of the present disclosure, the deformation information acquisition unit 1302 is further configured to obtain the similarity information based on the sub-image features of each image block in each image frame, where the image block is one of a predetermined number of image blocks obtained by respectively splitting each image frame; obtain the deformation information based on the similarity information.
[0120] According to an exemplary embodiment of the present disclosure, the deformation information acquisition unit 1302 is further configured to obtain the image features of each image frame of the video to be repaired; obtain a Query matrix, a Key matrix, and a Value matrix based on the image features of each image frame; obtain a similarity matrix based on the similarity between each Query sub-matrix of each image block in the Query matrix and each Key sub-matrix of each image block in the Key matrix as the similarity information.
[0121] According to an exemplary embodiment of the present disclosure, the deformation information acquisition unit 1302 is further configured to acquire the mask of each image block corresponding to the Query matrix and the mask of each image block corresponding to the Key matrix; when the masks of the image blocks in the Query matrix and the masks of the image blocks in the Key matrix are both 1, take the product of the Query sub-matrix of the image block in the Query matrix and the Key sub-matrix of the image block in the Key matrix as the similarity between the two image blocks; when any one of the masks of the image blocks in the Query matrix and the masks of the image blocks in the Key matrix is 0, take 0 as the similarity between the two image blocks; perform normalization processing on the similarities and then merge them to obtain a similarity matrix.
[0122] According to an exemplary embodiment of the present disclosure, the deformation information acquisition unit 1302 is further configured to obtain a deformation parameter matrix based on the similarity matrix as the deformation information.
[0123] According to an exemplary embodiment of the present disclosure, the deformation processing unit 1303 is further configured to perform deformation processing on the Key matrix and the Value matrix based on the deformation parameter matrix to obtain a deformed Key matrix and a deformed Value matrix.
[0124] According to an exemplary embodiment of the present disclosure, the deformation processing unit 1303 is further configured to, for each image block, respectively deform the Key sub-matrix of the image block in the Key matrix and the Value sub-matrix of the image block in the Value matrix through the deformation parameter sub-matrix of the image block in the deformation parameter matrix to obtain a deformed Key sub-matrix and a deformed Value sub-matrix of the image block; splice the deformed Key sub-matrices to obtain a deformed Key matrix; splice the deformed Value sub-matrices to obtain a deformed Value matrix.
[0125] According to an exemplary embodiment of the present disclosure, the repair unit 1304 is further configured to obtain a first repaired image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix.
[0126] According to an exemplary embodiment of the present disclosure, the repair unit 1304 is further configured to obtain a second repaired image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix.
[0127] According to an exemplary embodiment of the present disclosure, the video restoration method is performed by a pre-trained video restoration model. The video restoration model includes an encoder, a deformation parameter prediction network, a first attention network, a second attention network, a fusion network, and a decoder. Among them, the deformation information acquisition unit is further configured to input the video to be restored into the encoder to obtain a feature matrix of each image frame of the video to be restored as image features; the deformation information acquisition unit is further configured to input the similarity matrix into the deformation parameter prediction network to obtain a deformation parameter matrix; among them, the restoration unit is further configured to input the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix into the corresponding processing sub-network in the first attention network to obtain the first restored image feature of the current image frame, where the first attention network is used to perform spatial attention operations on the current image frame based on the image features of the current image frame; the restoration unit is further configured to input the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix into the second attention network to obtain the second restored image feature of the current image frame, where the second attention network is used to perform temporal attention operations on the current image frame based on the image features of each image frame; the restoration unit is further configured to input the first restored image feature, the second restored image feature, and the deformation parameter matrix into the fusion network to obtain the fused restored image feature of the current image frame; the restored video acquisition unit is further configured to input the fused restored image features of each image frame after the deformation process into the decoder to obtain the restored video.
[0128] According to an exemplary embodiment of the present disclosure, the video restoration model is trained through the following operations: obtaining a training sample set, where the training sample set includes a plurality of training videos and a clear video corresponding to each training video; inputting the training videos into an encoder to obtain a feature matrix of each image frame of the training videos; based on the feature matrix of each image frame, obtaining a Query matrix, a Key matrix, and a Value matrix; based on the similarity between the Query sub-matrix of each image patch in the Query matrix and the Key sub-matrix of each image patch in the Key matrix, obtaining a similarity matrix, where an image patch is one of a predetermined number of image patches obtained by splitting each image frame; inputting the similarity matrix into a deformation parameter prediction network to obtain a deformation parameter matrix; based on the deformation parameter matrix, performing a deformation process on the Key matrix and the Value matrix to obtain a deformed Key matrix and a deformed Value matrix; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into a first attention network to obtain a restored first feature matrix, where the first feature matrix includes the first restored image feature of each image frame; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into a second attention network to obtain a restored second feature matrix, where the second feature matrix includes the second restored image feature of each image frame; inputting the first feature matrix, the second feature matrix, and the deformation parameter matrix into a fusion network to obtain an estimated restored feature matrix of each image frame, where the estimated restored feature matrix includes the fused restored image feature of each image frame; inputting the estimated restored feature matrix into a decoder to obtain an estimated restored video; based on the estimated restored video and the clear video corresponding to the training video, adjusting the parameters of the encoder, the deformation parameter prediction network, the first attention network, the second attention network, the fusion network, and the decoder to train the video restoration model.
[0129] Figure 14 FIG. 4 is a block diagram of a training apparatus for a video restoration model according to an exemplary embodiment. The video restoration model includes an encoder, a deformation parameter prediction network, a first attention network, a second attention network, a fusion network, and a decoder. The training apparatus includes: a training sample set acquisition unit 1401, an encoding unit 1402, a matrix acquisition unit 1403, a similarity acquisition unit 1404, a deformation parameter acquisition unit 1405, a deformation processing unit 1406, a first feature matrix acquisition unit 1407, a second feature matrix acquisition unit 1408, a fusion unit 1409, a decoding unit 1410, and a training unit 1411.
[0130] A training sample set acquisition unit 1401 is configured to acquire a training sample set, where the training sample set includes a plurality of training videos and a clear video corresponding to each training video; an encoding unit 1402 is configured to input a training video into an encoder to obtain a feature matrix of each image frame of the training video; a matrix acquisition unit 1403 is configured to obtain a Query matrix, a Key matrix, and a Value matrix based on the feature matrix of each image frame; a similarity acquisition unit 1404 is configured to obtain a similarity matrix based on the similarity between a Query sub-matrix of each image patch in the Query matrix and a Key sub-matrix of each image patch in the Key matrix, where an image patch is one of a predetermined number of image patches obtained by respectively splitting each image frame; a deformation parameter acquisition unit 1405 is configured to input the similarity matrix into a deformation parameter prediction network to obtain a deformation parameter matrix; a deformation processing unit 1406 is configured to perform a rotation process on the Key matrix and the Value matrix based on the deformation parameter matrix to obtain a deformed Key matrix and a deformed Value matrix; a first feature matrix acquisition unit 1407 is configured to input the Query matrix, the deformed Key matrix, and the deformed Value matrix into a first attention network to obtain a restored first feature matrix, where the first feature matrix is obtained based on the reference information between the sub-matrices of the same image frame in the Query matrix, the deformed Key matrix, and the deformed Value matrix; a second feature matrix acquisition unit 1408 is configured to input the Query matrix, the deformed Key matrix, and the deformed Value matrix into a second attention network to obtain a restored second feature matrix, where the second feature matrix is obtained based on the reference information between the sub-matrices of different image frames in the Query matrix, the deformed Key matrix, and the deformed Value matrix; a fusion unit 1409 is configured to input the first feature matrix, the second feature matrix, and the deformation parameter matrix into a fusion network to obtain an estimated restored feature matrix of each image frame; a decoding unit 1410 is configured to input the estimated restored feature matrix into a decoder to obtain an estimated restored video; a training unit 1411 is configured to adjust the parameters of the encoder, the deformation parameter prediction network, the first attention network, the second attention network, the fusion network, and the decoder based on the estimated restored video and the clear video corresponding to the training video, and train the video restoration model.
[0131] According to an exemplary embodiment of the present disclosure, the first attention network includes T processing sub-networks, where T represents the number of image frames of the training video. Among them, the first feature matrix acquisition unit 1407 is further configured to input the sub-matrix of the i-th image frame in the Query matrix, the sub-matrix of the i-th image frame in the deformed Key matrix, and the sub-matrix of the i-th image frame in the deformed Value matrix into the i-th processing sub-network to obtain the feature matrix of the i-th image frame, where i ∈ {1, 2... T}; and splice the feature matrices of all image frames to obtain the repaired first feature matrix.
[0132] According to an exemplary embodiment of the present disclosure, the second attention network includes T processing sub-networks, where T represents the number of image frames to be repaired during training. Among them, the second feature matrix acquisition unit 1408 is further configured to input the sub-matrix of the j-th image frame in the Query matrix, the sub-matrices of the image frames other than the j-th image frame in the deformed Key matrix, and the sub-matrices of the image frames other than the j-th image frame in the deformed Value matrix into the j-th processing sub-network to obtain the feature matrix of the j-th image frame, where j ∈ {1, 2... T}; and splice the feature matrices of all image frames to obtain the repaired second feature matrix.
[0133] According to an exemplary embodiment of the present disclosure, the fusion network includes a first processing sub-network, a second processing sub-network, and a fusion sub-network. Among them, the fusion unit 1409 is further configured to input the first feature matrix and the deformation parameter matrix into the first processing sub-network to obtain a weighted first feature matrix; input the second feature matrix and the deformation parameter matrix into the second processing sub-network to obtain a weighted second feature matrix; and input the weighted first feature matrix and the weighted second feature matrix into the fusion sub-network to obtain the estimated repaired feature matrix of each image frame.
[0134] According to an embodiment of the present disclosure, an electronic device can be provided. Figure 15 FIG. is a block diagram of an electronic device 1500 according to an embodiment of the present disclosure. The electronic device includes at least one memory 1501 and at least one processor 1502. A set of computer-executable instructions is stored in the at least one memory. When the set of computer-executable instructions is executed by the at least one processor, the video repair method and the training method of the video repair model according to the embodiments of the present disclosure are executed.
[0135] As an example, the electronic device 1500 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 1000 does not have to be a single electronic device, and can also be a collection of devices or circuits that can execute the above instructions (or instruction sets) alone or in combination. The electronic device 1500 can also be part of an integrated control system or system manager, or can be configured as a portable electronic device that interfaces with a local or remote (e.g., via wireless transmission).
[0136] In the electronic device 1500, the processor 1502 can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 1502 can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.
[0137] The processor 1502 can run instructions or code stored in the memory, where the memory 1501 can also store data. The instructions and data can also be sent and received over a network via a network interface device, where the network interface device can employ any known transmission protocol.
[0138] The memory 1501 can be integrated with the processor 1502, for example, by arranging RAM or flash memory within an integrated circuit microprocessor, etc. In addition, the memory 1502 can include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory 1501 and the processor 1502 can be operatively coupled, or can communicate with each other, for example, through an I / O port, a network connection, etc., such that the processor 1502 can read files stored in the memory 1501.
[0139] In addition, the electronic device 1500 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device can be connected to each other via a bus and / or a network.
[0140] According to an embodiment of the present disclosure, a computer-readable storage medium may also be provided, wherein when the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the video repair method and the training method of the video repair model according to the embodiments of the present disclosure. Examples of the computer-readable storage medium here include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0141] According to an embodiment of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, implement the video repair method and the training method of the video repair model according to the embodiments of the present disclosure.
[0142] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0143] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A video restoration method, characterized in that, Including: Obtain the video to be repaired; Based on the image features of each image frame of the video to be repaired, obtain the deformation information between each image frame of the video to be repaired; Based on the deformation information, perform deformation processing on each image frame of the video to be repaired; For each image frame after deformation processing, perform the following operations: obtain the first repaired image feature of the current image frame by performing spatial attention operation on the current image frame based on the image features of the current image frame, obtain the second repaired image feature of the current image frame by performing temporal attention operation on the current image frame based on the image features of each image frame, and fuse the first repaired image feature and the second repaired image feature to obtain the fused repaired image feature of the current image frame; Based on the fused repaired image features of each image frame after deformation processing, obtain the repaired video; Wherein, the fusing the first repaired image feature and the second repaired image feature to obtain the fused repaired image feature of the current image frame includes: Based on the deformation information, respectively determine the weight of the first repaired image feature and the weight of the second repaired image feature; Based on the weight of the first repaired image feature and the weight of the second repaired image feature, fuse the first repaired image feature and the second repaired image feature to obtain the fused repaired image feature of the current image frame.
2. The video repair method according to claim 1, wherein The obtaining the deformation information between each image frame of the video to be repaired based on the image features of each image frame of the video to be repaired includes: Based on the sub-image features of each image block in each image frame, obtain similarity information, where the image block is one of a predetermined number of image blocks respectively segmented from each image frame; Based on the similarity information, obtain the deformation information.
3. The video repair method according to claim 2, characterized in that The obtaining the similarity information based on the sub-image features of each image block in each image frame includes: Obtain the image features of each image frame of the video to be repaired; Based on the image features of each image frame, obtain a Query matrix, a Key matrix, and a Value matrix; Based on the similarity between the Query sub-matrix of each image block in the Query matrix and the Key sub-matrix of each image block in the Key matrix, obtain a similarity matrix as the similarity information.
4. The video repair method according to claim 3, characterized in that The obtaining the similarity matrix based on the similarity between the Query sub-matrix of each image block in the Query matrix and the Key sub-matrix of each image block in the Key matrix includes: Obtain the mask of each image block corresponding to the Query matrix and the mask of each image block corresponding to the Key matrix; When the masks of the image blocks in the Query matrix and the masks of the image blocks in the Key matrix are both 1, use the product of the Query sub-matrix of the image block in the Query matrix and the Key sub-matrix of the image block in the Key matrix as the similarity between the two image blocks; When any one of the masks of the image patches in the Query matrix and the masks of the image patches in the Key matrix is 0, 0 is taken as the similarity between the two image patches. The similarities are normalized and then combined to obtain the similarity matrix.
5. The video repair method according to claim 3, wherein The obtaining of the deformation information based on the similarity information includes: Based on the similarity matrix, a deformation parameter matrix is obtained as the deformation information.
6. The video restoration method according to claim 5, wherein The performing of the deformation processing on each image frame of the video to be repaired based on the deformation information includes: Based on the deformation parameter matrix, deformation processing is performed on the Key matrix and the Value matrix to obtain a deformed Key matrix and a deformed Value matrix.
7. The video restoration method according to claim 6, characterized in that The performing of the deformation processing on the Key matrix and the Value matrix based on the deformation parameter matrix to obtain a deformed Key matrix and a deformed Value matrix includes: For each image patch, through the deformation parameter sub-matrix of the image patch in the deformation parameter matrix, the Key sub-matrix of the image patch in the Key matrix and the Value sub-matrix of the image patch in the Value matrix are respectively deformed to obtain the deformed Key sub-matrix and the deformed Value sub-matrix of the image patch; The deformed Key sub-matrices are spliced to obtain the deformed Key matrix; The deformed Value sub-matrices are spliced to obtain the deformed Value matrix.
8. The video repair method according to claim 6, wherein The obtaining of the first repaired image feature of the current image frame by performing a spatial attention operation on the current image frame based on the image feature of the current image frame includes: Based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix, the first repaired image feature of the current image frame is obtained.
9. The video restoration method according to claim 8, wherein The obtaining of the second repaired image feature of the current image frame by performing a temporal attention operation on the current image frame based on the image feature of each image frame includes: Based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of the other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of the other image frames except the current image frame in the deformed Value matrix, the second repaired image feature of the current image frame is obtained.
10. The video repair method according to claim 9, wherein, The video repair method is executed by a pre-trained video repair model, and the video repair model includes an encoder, a deformation parameter prediction network, a first attention network, a second attention network, a fusion network, and a decoder. Among them, the obtaining of the image feature of each image frame of the video to be repaired includes: The video to be repaired is input into the encoder to obtain a feature matrix of each image frame of the video to be repaired as the image feature. Among them, the obtaining of the deformation parameter matrix based on the similarity matrix includes: Input the similarity matrix into the deformation parameter prediction network to obtain the deformation parameter matrix; Among them, obtaining the first restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix includes: Input the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix into the corresponding processing sub-network in the first attention network to obtain the first restored image feature of the current image frame, where the first attention network is used to perform spatial attention operations on the current image frame based on the image features of the current image frame; Among them, obtaining the second restored image feature of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix includes: Input the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and the Value sub-matrices of other image frames except the current image frame in the deformed Value matrix into the second attention network to obtain the second restored image feature of the current image frame, where the second attention network is used to perform temporal attention operations on the current image frame based on the image features of each image frame; Among them, fusing the first restored image feature and the second restored image feature to obtain the fused restored image feature of the current image frame includes: Input the first restored image feature, the second restored image feature, and the deformation parameter matrix into the fusion network to obtain the fused restored image feature of the current image frame; Among them, obtaining the restored video based on the fused restored image features of each image frame after deformation processing includes: Input the fused restored image features of each image frame after deformation processing into the decoder to obtain the restored video.
11. The video repair method according to claim 10, characterized in that, The video restoration model is trained through the following operations: Obtain a training sample set, where the training sample set includes multiple training videos and the clear video corresponding to each training video; Input the training video into the encoder to obtain the feature matrix of each image frame of the training video; Based on the feature matrix of each image frame, obtain the Query matrix, Key matrix, and Value matrix; A similarity matrix is obtained based on the similarity between the Query sub - matrices of each image patch in the Query matrix and the Key sub - matrices of each image patch in the Key matrix, where the image patch is one of a predetermined number of image patches obtained by splitting each of the image frames; The similarity matrix is input into the deformation parameter prediction network to obtain a deformation parameter matrix; Based on the deformation parameter matrix, deformation processing is performed on the Key matrix and the Value matrix to obtain a deformed Key matrix and a deformed Value matrix; The Query matrix, the deformed Key matrix, and the deformed Value matrix are input into the first attention network to obtain a restored first feature matrix, where the first feature matrix includes the first restored image features of each image frame; The Query matrix, the deformed Key matrix, and the deformed Value matrix are input into the second attention network to obtain a restored second feature matrix, where the second feature matrix includes the second restored image features of each image frame; The first feature matrix, the second feature matrix, and the deformation parameter matrix are input into the fusion network to obtain the estimated restored feature matrix of each image frame, where the estimated restored feature matrix includes the fused restored image features of each image frame; The estimated restored feature matrix is input into the decoder to obtain an estimated restored video; Based on the estimated restored video and the clear video corresponding to the training video, the parameters of the encoder, the deformation parameter prediction network, the first attention network, the second attention network, the fusion network, and the decoder are adjusted to train the video restoration model.
12. A video restoration device, characterized in that, Comprising: A video acquisition unit configured to acquire a video to be restored; A deformation information acquisition unit configured to obtain the deformation information between each image frame of the video to be restored based on the image features of each image frame of the video to be restored; A deformation processing unit configured to perform deformation processing on each image frame of the video to be restored based on the deformation information; A restoration unit configured to perform the following operations on each image frame after deformation processing: obtaining the first restored image features of the current image frame by performing spatial attention operation on the current image frame based on the image features of the current image frame, obtaining the second restored image features of the current image frame by performing temporal attention operation on the current image frame based on the image features of each image frame, and fusing the first restored image features and the second restored image features to obtain the fused restored image features of the current image frame; A restored video acquisition unit configured to obtain a restored video based on the fused restored image features of each image frame after deformation processing; Wherein, the restoration unit is further configured to respectively determine the weights of the first restored image features and the weights of the second restored image features based on the deformation information; Based on the weights of the first repaired image features and the weights of the second repaired image features, fuse the first repaired image features and the second repaired image features to obtain the fused repaired image features of the current image frame.
13. The video repair device according to claim 12, characterized in that, The deformation information acquisition unit is further configured to obtain similarity information based on the sub-image features of each image block in each image frame, where the image block is one of a predetermined number of image blocks obtained by respectively splitting each image frame; based on the similarity information, obtain the deformation information.
14. The video restoration device according to claim 13, wherein The deformation information acquisition unit is further configured to obtain the image features of each image frame of the video to be repaired; Based on the image features of each image frame, obtain a Query matrix, a Key matrix, and a Value matrix; based on the similarity between the Query sub-matrix of each image block in the Query matrix and the Key sub-matrix of each image block in the Key matrix, obtain a similarity matrix as the similarity information.
15. The video repair device according to claim 14, characterized in that, The deformation information acquisition unit is further configured to obtain the mask of each image block corresponding to the Query matrix and the mask of each image block corresponding to the Key matrix; when the masks of the image blocks in the Query matrix and the masks of the image blocks in the Key matrix are both 1, use the product of the Query sub-matrix of the image block in the Query matrix and the Key sub-matrix of the image block in the Key matrix as the similarity between the two image blocks; when any one of the masks of the image blocks in the Query matrix and the masks of the image blocks in the Key matrix is 0, use 0 as the similarity between the two image blocks; normalize and merge the similarities to obtain the similarity matrix.
16. The video repair device according to claim 14, characterized in that, The deformation information acquisition unit is further configured to obtain a deformation parameter matrix based on the similarity matrix as the deformation information.
17. The video repair device according to claim 16, wherein The deformation processing unit is further configured to perform deformation processing on the Key matrix and the Value matrix based on the deformation parameter matrix to obtain a deformed Key matrix and a deformed Value matrix.
18. The video repair device according to claim 17, wherein, The deformation processing unit is further configured to, for each image block, respectively deform the Key sub-matrix of the image block in the Key matrix and the Value sub-matrix of the image block in the Value matrix through the deformation parameter sub-matrix of the image block in the deformation parameter matrix to obtain the deformed Key sub-matrix and the deformed Value sub-matrix of the image block; Stitch the deformed Key sub-matrices to obtain the deformed Key matrix; stitch the deformed Value sub-matrices to obtain the deformed Value matrix.
19. The video repair device according to claim 17, characterized in that, The repair unit is further configured to obtain the first repaired image features of the current image frame based on the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix.
20. The video repair device according to claim 19, wherein, The repair unit is further configured to obtain a second repaired image feature of the current image frame based on a Query sub-matrix of the current image frame in the Query matrix, Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and Value sub-matrices of other image frames except the current image frame in the deformed Value matrix.
21. The video repair device according to claim 20, wherein, The video repair device operates based on a pre-trained video repair model, and the video repair model includes an encoder, a deformation parameter prediction network, a first attention network, a second attention network, a fusion network, and a decoder. Among them, the deformation information acquisition unit is further configured to input the video to be repaired into the encoder to obtain a feature matrix of each image frame of the video to be repaired as an image feature. The deformation information acquisition unit is further configured to input the similarity matrix into the deformation parameter prediction network to obtain the deformation parameter matrix. Among them, the repair unit is further configured to input the Query sub-matrix of the current image frame in the Query matrix, the Key sub-matrix of the current image frame in the deformed Key matrix, and the Value sub-matrix of the current image frame in the deformed Value matrix into corresponding processing sub-networks in the first attention network to obtain a first repaired image feature of the current image frame, where the first attention network is used to perform a spatial attention operation on the current image frame based on the image feature of the current image frame. The repair unit is further configured to input the Query sub-matrix of the current image frame in the Query matrix, Key sub-matrices of other image frames except the current image frame in the deformed Key matrix, and Value sub-matrices of other image frames except the current image frame in the deformed Value matrix into the second attention network to obtain a second repaired image feature of the current image frame, where the second attention network is used to perform a temporal attention operation on the current image frame based on the image feature of each image frame. The repair unit is further configured to input the first repaired image feature, the second repaired image feature, and the deformation parameter matrix into the fusion network to obtain a fused repaired image feature of the current image frame. The repaired video acquisition unit is further configured to input the fused repaired image feature of each image frame after the deformation process into the decoder to obtain the repaired video.
22. The video repair device according to claim 21, wherein, The video restoration model is trained through the following operations: obtaining a training sample set, where the training sample set includes multiple training videos and the clear video corresponding to each training video; inputting the training video into the encoder to obtain the feature matrix of each image frame of the training video; obtaining a Query matrix, a Key matrix, and a Value matrix based on the feature matrix of each image frame; obtaining a similarity matrix based on the similarity between the Query sub-matrix of each image patch in the Query matrix and the Key sub-matrix of each image patch in the Key matrix, where the image patch is one of a predetermined number of image patches obtained by respectively splitting each image frame; inputting the similarity matrix into the deformation parameter prediction network to obtain a deformation parameter matrix; performing a deformation process on the Key matrix and the Value matrix based on the deformation parameter matrix to obtain a deformed Key matrix and a deformed Value matrix; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into the first attention network to obtain a restored first feature matrix, where the first feature matrix includes the first restored image feature of each image frame; inputting the Query matrix, the deformed Key matrix, and the deformed Value matrix into the second attention network to obtain a restored second feature matrix, where the second feature matrix includes the second restored image feature of each image frame; inputting the first feature matrix, the second feature matrix, and the deformation parameter matrix into the fusion network to obtain the estimated restored feature matrix of each image frame, where the estimated restored feature matrix includes the fused restored image feature of each image frame; inputting the estimated restored feature matrix into the decoder to obtain an estimated restored video; and adjusting the parameters of the encoder, the deformation parameter prediction network, the first attention network, the second attention network, the fusion network, and the decoder based on the estimated restored video and the clear video corresponding to the training video to train the video restoration model.
23. An electronic device, characterized in that, Comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the video restoration method according to any one of claims 1 to 11.
24. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the video restoration method according to any one of claims 1 to 11.
25. A computer program product, comprising computer instructions, characterized in that, The computer instructions, when executed by a processor, implement the video restoration method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Video classification method and device, storage medium and electronic equipment
CN111259781A
Deblurred video recovery method and device, terminal equipment and storage medium
CN111932480A