Video frame deblurring method and device, computer equipment and storage medium
Through the method of fusion and feature extraction of adjacent video frames, cross-scale feature fusion is used for encoding and decoding, the accuracy and efficiency of video frame defuzzing are solved, and a more efficient video frame defuzzing effect is achieved.
Patent Information
- Application Number
- CN202311569978.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to effectively remove fuzzy video frames in videos, affecting the user experience.
By obtaining fuzzy video frames and their adjacent two video frames, fusion and feature extraction are performed, cross-scale feature fusion is used for encoding and decoding, reducing dependence on the time dimension, and making full use of the features of the spatial dimension.
It improves the accuracy and efficiency of defuzzing video frames, reduces the amount of computing, and enhances the processing ability of fuzzy video frames.
Smart Images

Figure CN120031746A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a video frame deblurring method, apparatus, computer equipment, and storage medium. Background Art
[0002] With the popularity of handheld photography devices such as mobile phones, video shooting is becoming more and more popular among users. However, when users shoot with photography devices, due to the shaking of the photography devices or the rapid movement of the objects being photographed, the captured videos may be blurred, affecting the user experience. Therefore, it is urgent to provide a solution for deblurring blurred video frames in a video. Summary of the invention
[0003] The embodiments of the present application provide a video frame deblurring method, apparatus, computer equipment and storage medium, which can reduce the dependence on the features in the time dimension, make full use of the important features in the spatial dimension, and improve the accuracy of video frame deblurring while reducing the amount of calculation. The technical solution is as follows:
[0004] In one aspect, a video frame deblurring method is provided, the method comprising:
[0005] Acquire a blurred video frame, a first video frame, and a second video frame, wherein the first video frame is a video frame before the blurred video frame, and the second video frame is a video frame after the blurred video frame;
[0006] The blurred video frame, the first video frame and the second video frame are fused to obtain a fused video frame, and features are extracted from the fused video frame to obtain first frame features of k different scales, wherein the scale of the mth first frame feature is greater than the scale of the m+1th first frame feature, where k is a positive integer greater than 1, and m is any positive integer from 1 to k-1;
[0007] For the mth first frame feature, the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature are fused to obtain the mth fused first frame feature, and the mth fused first frame feature is encoded to obtain the mth second frame feature; wherein the m+1th second frame feature is the encoding feature of the m+1th first frame feature, and the m+1th third frame feature is the decoding feature of the m+1th second frame feature;
[0008] For the first second frame feature, the first second frame feature is decoded to obtain a deblurred video frame.
[0009] Optionally, decoding the multiple second sub-features in the first second frame feature based on the weights of the multiple second sub-features in the first second frame feature to obtain the deblurred video frame includes:
[0010] For each second sub-feature in the first second frame feature, multiply the weight of the second sub-feature by the second sub-feature to obtain a decoded feature of the second sub-feature;
[0011] The decoded features of the plurality of second sub-features are combined to obtain the deblurred video frame.
[0012] Optionally, the training of the video frame processing model based on the difference between the predicted video frame and the sample clear video frame includes:
[0013] The video frame processing model is trained based on the difference between the predicted video frame and the sample clear video frame and the difference between the edge gradient of the predicted video frame and the edge gradient of the sample clear video frame.
[0014] Optionally, the training of the video frame processing model based on the difference between the predicted video frame and the sample clear video frame and the difference between the edge gradient of the predicted video frame and the edge gradient of the sample clear video frame includes:
[0015] Determining a first loss value based on a difference between the predicted video frame and the sample clear video frame;
[0016] Determine a second loss value based on a difference between an edge gradient of the predicted video frame and an edge gradient of the sample clear video frame;
[0017] Performing a weighted summation of the first loss value and the second loss value to obtain a third loss value;
[0018] Based on the third loss value, the video frame processing model is trained.
[0019] On the other hand, a video frame deblurring device is provided, the device comprising:
[0020] A video frame acquisition module, used to acquire a blurred video frame, a first video frame and a second video frame, wherein the first video frame is a video frame before the blurred video frame, and the second video frame is a video frame after the blurred video frame;
[0021] a feature extraction module, configured to fuse the blurred video frame, the first video frame and the second video frame to obtain a fused video frame, and perform feature extraction on the fused video frame to obtain first frame features of k different scales, wherein the scale of the mth first frame feature is greater than the scale of the m+1th first frame feature, where k is a positive integer greater than 1, and m is any positive integer from 1 to k-1;
[0022] The encoding module is used for fusing the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature to obtain the mth fused first frame feature, and encoding the mth fused first frame feature to obtain the mth second frame feature; wherein the m+1th second frame feature is the encoding feature of the m+1th first frame feature, and the m+1th third frame feature is the decoding feature of the m+1th second frame feature;
[0023] A decoding module is used to decode the first second frame feature to obtain a deblurred video frame.
[0024] Optionally, the encoding module is further configured to, when m is equal to k-1, encode the m+1th first frame feature to obtain the m+1th second frame feature for the m+1th first frame feature;
[0025] The decoding module is also used to, when m is not equal to 1, decode the mth second frame feature to obtain the mth third frame feature for the mth second frame feature; and when m is equal to k-1, decode the m+1th second frame feature to obtain the m+1th third frame feature for the m+1th second frame feature.
[0026] Optionally, the feature extraction module is used to:
[0027] Determine the blurred video frame, the first video frame and the second video frame as a first group of video frames, perform k-1 downsampling of different scales on the first group of video frames to obtain a second group of video frames to a kth group of video frames, wherein the scale of the mth group of video frames is greater than the scale of the m+1th group of video frames;
[0028] For each group of k groups of video frames, the blurred video frame, the first video frame and the second video frame in the current group of video frames are fused to obtain a fused video frame, and features are extracted from the current fused video frame to obtain a first frame feature.
[0029] Optionally, the feature extraction module is used to:
[0030] Determine first region information and second region information, wherein the first region information and the second region information indicate regions in a video frame;
[0031] The content in the area indicated by the first area information in the blurred video frame is replaced with the content in the area indicated by the first area information in the first video frame, and the content in the area indicated by the second area information in the blurred video frame is replaced with the content in the area indicated by the second area information in the second video frame to obtain the fused video frame.
[0032] Optionally, the feature extraction module is used to:
[0033] Fusing the blurred video frame, the first video frame and the second video frame to obtain a first fused video frame;
[0034] Downsampling the first fused video frame at different scales k-1 times to obtain second to k-th fused video frames, wherein the scale of the m-th fused video frame is greater than the scale of the m+1-th fused video frame;
[0035] For each of the k fused video frames, a feature of the current fused video frame is extracted to obtain a first frame feature.
[0036] Optionally, the encoding module is used to:
[0037] Based on the mth fused first frame feature, determine the weight of the first sub-feature on multiple regions in the mth fused first frame feature, where the weight of the first sub-feature indicates the importance of the first sub-feature;
[0038] Based on the weights of the multiple first sub-features in the m-th fused first frame feature, the multiple first sub-features in the m-th fused first frame feature are encoded to obtain the m-th second frame feature.
[0039] Optionally, the encoding module is used to:
[0040] Divide the mth fused first frame feature into multiple first sub-features according to the region;
[0041] For each first sub-feature, a first prediction convolution kernel is used to convolve the first sub-feature to obtain a weight of the first sub-feature.
[0042] Optionally, the encoding module is used to:
[0043] For each first sub-feature in the mth fused first frame feature, multiplying the weight of the first sub-feature by the first sub-feature to obtain an encoded feature of the first sub-feature;
[0044] The encoded features of the multiple first sub-features are combined to obtain the mth second frame feature.
[0045] Optionally, the decoding module is used to:
[0046] Based on the first second frame feature, determining weights of second sub-features on multiple regions in the first second frame feature, wherein the weights of the second sub-features represent the importance of the second sub-features;
[0047] Based on the weights of the multiple second sub-features in the first second frame feature, the multiple second sub-features in the first second frame feature are decoded to obtain the deblurred video frame.
[0048] Optionally, the decoding module is used to:
[0049] Dividing the first second frame feature into a plurality of second sub-features according to the region;
[0050] For each second sub-feature, a second prediction convolution kernel is used to perform convolution on the second sub-feature to obtain a weight of the second sub-feature.
[0051] Optionally, the decoding module is used to:
[0052] For each second sub-feature in the first second frame feature, multiply the weight of the second sub-feature by the second sub-feature to obtain a decoded feature of the second sub-feature;
[0053] The decoded features of the plurality of second sub-features are combined to obtain the deblurred video frame.
[0054] Optionally, the video frame deblurring device is implemented by a video frame processing model, and the device further includes a model training module, which is used to:
[0055] Acquire a sample blurred video frame, a sample first video frame, a sample second video frame and a sample clear video frame, wherein the sample first video frame is a video frame before the sample blurred video frame, the sample second video frame is a video frame after the sample blurred video frame, and the sample clear video frame is a video frame after the sample blurred video frame is deblurred;
[0056] Inputting the sample blurred video frame, the sample first video frame and the sample second video frame into the video frame processing model to obtain a predicted video frame output by the video frame processing model;
[0057] The video frame processing model is trained based on the difference between the predicted video frame and the sample clear video frame.
[0058] Optionally, the model training module is used to:
[0059] The video frame processing model is trained based on the difference between the predicted video frame and the sample clear video frame and the difference between the edge gradient of the predicted video frame and the edge gradient of the sample clear video frame.
[0060] Optionally, the model training module is used to:
[0061] Determining a first loss value based on a difference between the predicted video frame and the sample clear video frame;
[0062] Determine a second loss value based on a difference between an edge gradient of the predicted video frame and an edge gradient of the sample clear video frame;
[0063] Performing a weighted summation of the first loss value and the second loss value to obtain a third loss value;
[0064] Based on the third loss value, the video frame processing model is trained.
[0065] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the video frame deblurring method as described in the above aspects.
[0066] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed by the video frame deblurring method as described in the above aspects.
[0067] On the other hand, a computer program product is provided, comprising a computer program, wherein the computer program is loaded and executed by a processor to implement the operations performed by the video frame deblurring method as described in the above aspects.
[0068] The solution provided in the embodiment of the present application, when deblurring a blurred video frame, takes into account the influence of the two video frames before and after the blurred video frame on the blurred video frame because continuous video frames are correlated, and downsamples the video frame to obtain multiple video frames of different scales. During the encoding process, cross-scale feature fusion is performed, so that the encoding process can learn deeper and more fine-grained features, fully utilize important features in the spatial dimension, and reduce the amount of calculation while also improving the accuracy of video frame deblurring. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0070] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0071] Figure 2 is a flow chart of a video frame deblurring method provided by an embodiment of the present application;
[0072] Figure 3 is a flowchart of another video frame deblurring method provided by an embodiment of the present application;
[0073] Figure 4 is a schematic diagram of a video frame fusion method provided in an embodiment of the present application;
[0074] Figure 5 is a flowchart of another video frame deblurring method provided in an embodiment of the present application;
[0075] Figure 6 It is a schematic diagram of the structure of a coding network provided in an embodiment of the present application;
[0076] Figure 7 It is a structural diagram of a video frame processing model provided in an embodiment of the present application;
[0077] Figure 8 It is a flowchart of a training method of a video frame processing model provided in an embodiment of the present application;
[0078] Fig. 9 is a result comparison diagram of a video frame deblurring method provided by an embodiment of the present application;
[0079] Fig.10 is a result comparison diagram of another video frame deblurring method provided in an embodiment of the present application;
[0080] Fig.11 is a result comparison diagram of another video frame deblurring method provided in an embodiment of the present application;
[0081] Fig.12 This is a result comparison diagram of another video frame deblurring method provided in an embodiment of the present application;
[0082] Fig.13 This is a result comparison diagram of another video frame deblurring method provided in an embodiment of the present application;
[0083] Fig.14 It is a structural schematic diagram of a video frame deblurring device provided in an embodiment of the present application;
[0084] Fig.15 is a structural schematic diagram of another video frame deblurring device provided in an embodiment of the present application;
[0085] Fig.16 is a schematic diagram of the structure of a terminal provided in an embodiment of the present application;
[0086] Fig.17 It is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0087] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0088] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of this application, a first video frame may be referred to as a second video frame, and similarly, a second video frame may be referred to as a first video frame.
[0089] Among them, at least one refers to one or more than one, for example, at least one video frame can be one video frame, two video frames, three video frames, or any other integer greater than or equal to one. Multiple refers to two or more than two, for example, multiple video frames can be two video frames, three video frames, or any other integer greater than or equal to two. Each refers to each of at least one, for example, each video frame refers to each video frame in the multiple video frames. If the multiple video frames are 3 video frames, each video frame refers to each video frame in the 3 video frames.
[0090] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all fully authorized by users or relevant parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0091] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0092] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0093] Computer vision (CV) is a science that studies how to make machines "see". To be more specific, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further perform image processing so that the computer processes the images into images that are more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the field of vision such as Swin-Transformer, ViT (Vision Transformer), V-MoE (Vision MoE), and MAE (Masked Auto Encoder) can be quickly and widely applied to downstream specific tasks after fine-tuning (Fine Tune). Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (3 Dimensions) technology, virtual reality, augmented reality, simultaneous positioning and mapping, and also includes common biometric recognition technology.
[0094] The following will describe the video frame deblurring method provided in the embodiment of the present application based on artificial intelligence technology and computer vision technology.
[0095] The video frame deblurring method provided in the embodiment of the present application can be used in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network) and big data and artificial intelligence platforms. Server. Optionally, the terminal is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart terminal, etc., but is not limited to this.
[0096] In one possible implementation, the computer program involved in the embodiments of the present application may be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. Multiple computer devices distributed at multiple locations and interconnected through a communication network can constitute a blockchain system.
[0097] In one possible implementation, the computer device used to train the video frame processing model in the embodiment of the present application is a node in the blockchain system, and the node can store the trained video frame processing model in the blockchain. Thereafter, the node or the node corresponding to other devices in the blockchain can call the video frame processing model to deblur the video frame.
[0098] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application, see Figure 1 , the implementation environment includes: a terminal 101 and a server 102. The terminal 101 and the server 102 are connected via a wireless or wired network. Optionally, the server 102 is used to train a video frame processing model using the method provided in the embodiment of the present application, and send the video frame processing model to the terminal 101, and the video frame processing model is used to deblur the video frame. Subsequently, the terminal 101 calls the video frame processing model to deblur the video frame.
[0099] The video frame deblurring method provided in the embodiment of the present application can be applied to any scenario where the video frame needs to be deblurred.
[0100] For example, when a user is shooting a video with a handheld camera device, the shot video may be blurred due to the shaking of the camera device or the rapid movement of the shot object. In this case, the method provided in the embodiment of the present application can be used to deblur any video frame in the video by using the previous video frame and the next video frame of the video frame to obtain a clearer video frame.
[0101] Figure 2 is a flowchart of a video frame deblurring method provided in an embodiment of the present application. The embodiment of the present application is executed by a computer device. Figure 2 , the method comprising:
[0102] 201. A computer device obtains a blurred video frame, a first video frame, and a second video frame, where the first video frame is a video frame preceding the blurred video frame, and the second video frame is a video frame following the blurred video frame.
[0103] The first video frame, the blurred video frame and the second video frame are three consecutive video frames in the same video. Due to the shaking of the camera device or the movement of the object being shot when shooting the video, the video frames in the video may be blurred. The purpose of the embodiment of the present application is to deblur the blurred video frames. The first video frame and the second video frame are two video frames adjacent to the blurred video frame. The embodiment of the present application uses the first video frame and the second video frame to deblur the blurred video frame.
[0104] 202. The computer device fuses the blurred video frame, the first video frame and the second video frame to obtain a fused video frame, performs feature extraction on the fused video frame to obtain first frame features of k different scales, the scale of the mth first frame feature is greater than the scale of the m+1th first frame feature, k is a positive integer greater than 1, and m is any positive integer from 1 to k-1.
[0105] Since the images in the continuous video frames in the video are also continuous, when a part of the image in the blurred video frame is blurred, the part of the image in the first video frame and the second video frame may be clear, and the blurred video frame can be motion compensated using the first video frame and the second video frame. Therefore, the computer device fuses the blurred video frame, the first video frame and the second video frame to obtain a fused video frame, which fuses the image content in the blurred video frame, the first video frame and the second video frame.
[0106] After obtaining the fused video frame, the computer device extracts features from the fused video frame to obtain first frame features, which fuse features of the blurred video frame, the first video frame, and the second video frame.
[0107] It should be noted that, in the embodiment of the present application, by fusing the simulated video frame, the first video frame and the second video frame and then performing feature extraction, k first frame features of different scales are obtained, and the scales of the k first frame features decrease in sequence.
[0108] 203. The computer device fuses the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature to obtain the mth fused first frame feature, and encodes the mth fused first frame feature to obtain the mth second frame feature; wherein the m+1th second frame feature is the encoded feature of the m+1th first frame feature, and the m+1th third frame feature is the decoded feature of the m+1th second frame feature.
[0109] Wherein, m is any positive integer from 1 to k-1, that is, except for the kth first frame feature, for any mth first frame feature, the method of step 203 is used to determine the mth second frame feature corresponding to the mth first frame feature. The mth second frame feature can be understood as the second frame feature obtained by encoding the mth first frame feature after fusing the m+1th second frame feature and the m+1th third frame feature. That is, the encoding process of the frame feature of the mth scale fuses the encoding result of the m+1th scale and the decoding result of the m+1th scale.
[0110] 204. The computer device decodes the first second frame feature to obtain a deblurred video frame.
[0111] In the manner of step 203, the first first frame feature, the second second frame feature and the second third frame feature are fused to obtain the first fused first frame feature, and the first fused first frame feature is encoded to obtain the first second frame feature. In step 204, the second frame feature is decoded to obtain a deblurred video frame, which is the result of deblurring the blurred video frame.
[0112] The method provided in the embodiment of the present application, when deblurring a blurred video frame, takes into account the influence of the two preceding and succeeding video frames adjacent to the blurred video frame on the blurred video frame because continuous video frames are correlated, and downsamples the video frame to obtain multiple video frames of different scales. During the encoding process, cross-scale feature fusion is performed, so that the encoding process can learn deeper and more fine-grained features, fully utilize important features in the spatial dimension, and improve the accuracy of video frame deblurring while reducing the amount of calculation.
[0113] Above Figure 2The embodiment of the present invention is only a brief description of the video frame deblurring method. For a more detailed process, please refer to the following Figure 3 An embodiment of Figure 3 The embodiment of the present invention describes in detail the fusion process, encoding process and decoding process of the video frame in the video frame deblurring method. Figure 3 is a flowchart of another video frame deblurring method provided in an embodiment of the present application. The embodiment of the present application is executed by a computer device. Figure 3 , the method comprising:
[0114] 301. A computer device obtains a blurred video frame, a first video frame, and a second video frame, where the first video frame is a video frame before the blurred video frame, and the second video frame is a video frame after the blurred video frame.
[0115] The process of obtaining the video frame in step 301 is the same as the process of obtaining the video frame in step 201 above, and will not be repeated here.
[0116] 302. The computer device fuses the blurred video frame, the first video frame and the second video frame to obtain a fused video frame, performs feature extraction on the fused video frame to obtain first frame features of k different scales, the scale of the mth first frame feature is greater than the scale of the m+1th first frame feature, k is a positive integer greater than 1, and m is any positive integer from 1 to k-1.
[0117] Among them, the computer device obtains the first frame features of k different scales in two ways. The first way is to downsample first, then fuse and extract features. The second way is to fuse first, then downsample and extract features. The detailed process is shown in the following description.
[0118] The first method: the computer device determines the blurred video frame, the first video frame, and the second video frame as the first group of video frames, and downsamples the first group of video frames k-1 times at different scales to obtain the second group of video frames to the kth group of video frames, wherein the scale of the mth group of video frames is larger than the scale of the m+1th group of video frames; for each group of video frames in the k groups of video frames, the blurred video frame, the first video frame, and the second video frame in the current group of video frames are fused to obtain a fused video frame, and features are extracted for the current fused video frame to obtain a first frame feature.
[0119] That is, the first group of video frames are video frames of the original scale, and the computer device downsamples the first group of video frames to obtain the second group of video frames, and the scale of the first group of video frames is larger than the scale of the second group of video frames. The computer device downsamples the second group of video frames to obtain the third group of video frames, and the scale of the second group of video frames is larger than the scale of the third group of video frames. By analogy, the computer device downsamples k-1 times to obtain k-1 groups of video frames with different scales from the first group of video frames, thereby obtaining k groups of video frames with different scales, each group of video frames includes a blurred video frame, a first video frame, and a second video frame, and the scales of the three video frames in each group of video frames are the same.
[0120] For the mth group of video frames, the computer device fuses the blurred video frame, the first video frame, and the second video frame in the mth group of video frames to obtain the mth fused video frame, and extracts features from the mth fused video frame to obtain the mth first frame feature. For the kth group of video frames, the computer device fuses the blurred video frame, the first video frame, and the second video frame in the kth group of video frames to obtain the kth fused video frame, and extracts features from the kth fused video frame to obtain the kth first frame feature. That is, for each group of video frames in the k groups of video frames, the computer device fuses the blurred video frame, the first video frame, and the second video frame in the group of video frames to obtain a fused video frame, and extracts features from the fused video frame to obtain a first frame feature. Therefore, the computer device can obtain a total of k first frame features, and the scales of the k first frame features are different. Among them, the scale of the mth first frame feature is larger than the scale of the m+1th first frame feature, for example, the scale of the 1st first frame feature is larger than the scale of the 2nd first frame feature, and the scale of the 2nd frame feature is larger than the scale of the 3rd first frame feature.
[0121] Optionally, the computer device may use a convolutional neural network to extract features from the fused video frame to obtain features of the first frame.
[0122] Optionally, the computer device fuses the blurred video frame, the first video frame and the second video frame in the current group of video frames to obtain a fused video frame, including: determining first region information and second region information, the first region information and the second region information indicating regions in the video frame; replacing the content in the region indicated by the first region information in the blurred video frame with the content in the region indicated by the first region information in the first video frame, and replacing the content in the region indicated by the second region information in the blurred video frame with the content in the region indicated by the second region information in the second video frame, to obtain a fused video frame.
[0123] Since the first video frame, the blurred video frame and the second video frame are continuous video frames in the same video, the sizes of the first video frame, the blurred video frame and the second video frame are the same. The computer device determines the first area information and the second area information based on the sizes of the first video frame, the blurred video frame and the second video frame, and the first area information and the second area information indicate the areas in the first video frame, the blurred video frame and the second video frame. Among them, there is no overlapping area between the area indicated by the first area information and the area indicated by the second area information. Optionally, the first area information and the second area information can be randomly determined by the computer device, that is, the area to be fused is randomly determined. For example, taking the lower left corner of the video frame as the origin of the coordinate axis, the first coordinate information and the second coordinate information are randomly determined within the coordinate range where the video frame is located, and the first coordinate information is used as the first area information, and the second coordinate information is used as the second area information. The first coordinate information is the coordinate of a certain area in the video frame, and according to the first coordinate information, the computer device can determine an area in the video frame. The second coordinate information is the coordinate of a certain area in the video frame, and according to the second coordinate information, the computer device can determine another area in the video frame.
[0124] After determining the first region information and the second region information, the computer device determines the region indicated by the first region information in the first video frame and the blurred video frame, and replaces the content of the region in the blurred video frame with the content of the region in the first video frame, that is, replaces the image of the region in the blurred video frame with the image of the region in the first video frame, or replaces the voxels of the region in the blurred video frame with the voxels of the region in the first video frame, thereby realizing the fusion of the image of the first video frame to the image of the blurred video frame. The computer device determines the region indicated by the second region information in the second video frame and the blurred video frame, and replaces the content of the region in the blurred video frame with the content of the region in the second video frame, that is, replaces the image of the region in the blurred video frame with the image of the region in the second video frame, or replaces the voxels of the region in the blurred video frame with the voxels of the region in the second video frame, thereby realizing the fusion of the image of the second video frame to the image of the blurred video frame.
[0125] Among them, fusing the content of the first video frame into the blurred video frame can be understood as a forward inter-frame shift, and fusing the content of the second video frame into the blurred video frame can be understood as a backward inter-frame shift, thereby realizing a bidirectional inter-frame shift. Through the bidirectional inter-frame shift, the information of the forward frame and the backward frame is mixed with the information of the current frame, which not only avoids the use of additional methods for inter-frame comparison, but also expands the receptive field of the fused video frame.
[0126] Optionally, the video frame fusion process can be expressed by the following formula (1).
[0127]
[0128] in, represents the fused video frame, I t-1 Represents the first video frame, I t represents a blurred video frame, I t+1 represents the second video frame, I t-1 →I t Indicates that the content in the first video frame is merged into the blurred video frame, I t+1 →I t Indicates that the content in the second video frame is merged into the blurred video frame.
[0129] Figure 4 is a schematic diagram of a video frame fusion method provided in an embodiment of the present application, such as Figure 4 As shown, the first area information indicates areas marked with "1", "2", and "3", and the second area information indicates areas marked with "a", "b", and "c". The computer device replaces the images of the areas marked with "1", "2", and "3" in the blurred video frame with images of the areas marked with "1", "2", and "3" in the first video frame, and replaces the images of the areas marked with "a", "b", and "c" in the blurred video frame with images of the areas marked with "a", "b", and "c" in the second video frame, thereby obtaining a fused video frame.
[0130] The second method: the computer device fuses the blurred video frame, the first video frame and the second video frame to obtain the first fused video frame; downsamples the first fused video frame k-1 times at different scales to obtain the second fused video frame to the kth fused video frame, wherein the scale of the mth fused video frame is larger than the scale of the m+1th fused video frame; for each of the k fused video frames, extract features of the current fused video frame to obtain a first frame feature.
[0131] Among them, the blurred video frame, the first video frame and the second video frame are video frames of original scale, and the scale of the first fused video frame is the same as the scale of the blurred video frame, the first video frame and the second video frame, that is, the scale of the first fused video frame is a video frame of original scale. The computer device downsamples the first fused video frame to obtain the second fused video frame, and the scale of the first fused video frame is larger than the scale of the second fused video frame. The computer device downsamples the second fused video frame to obtain the third fused video frame, and the scale of the second fused video frame is larger than the scale of the third fused video frame. By analogy, the computer device downsamples k-1 times to obtain k-1 fused video frames with different scales from the first fused video frame, thereby obtaining k fused video frames with different scales.
[0132] For the mth fused video frame, the computer device extracts features from the mth fused video frame to obtain the mth first frame feature. For the kth fused video frame, the computer device extracts features from the kth fused video frame to obtain the kth first frame feature. That is, for each of the k fused video frames, the computer device extracts features from the fused video frame to obtain a first frame feature. Therefore, the computer device can obtain a total of k first frame features, and the scales of the k first frame features are different. Among them, the scale of the mth first frame feature is greater than the scale of the m+1th first frame feature, for example, the scale of the 1st first frame feature is greater than the scale of the 2nd first frame feature, and the scale of the 2nd frame feature is greater than the scale of the 3rd first frame feature.
[0133] Optionally, the computer device may use a convolutional neural network to extract features from the fused video frame to obtain features of the first frame.
[0134] Optionally, the fusion process of the blurred video frame, the first video frame and the second video frame in the second method is the same as the fusion process of the blurred video frame, the first video frame and the second video frame in the first method, and will not be repeated here.
[0135] 303. The computer device fuses the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature to obtain the mth fused first frame feature, and encodes the mth fused first frame feature to obtain the mth second frame feature; wherein the m+1th second frame feature is the encoded feature of the m+1th first frame feature, and the m+1th third frame feature is the decoded feature of the m+1th second frame feature.
[0136] Wherein, m is any positive integer from 1 to k-1, and the mth first frame feature is any first frame feature other than the k first frame features. For the mth first frame feature, the computer device first obtains the m+1th second frame feature and the m+1th third frame feature, and fuses the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature to obtain the mth fused first frame feature. That is, the computer device fuses the current first frame feature with the encoding result and decoding result of the previous first frame feature with a smaller scale. Wherein, the acquisition process of the m+1th second frame feature and the m+1th third frame feature can be referred to the following steps 304-305.
[0137] Taking the second first frame feature as an example, the computer device fuses the second first frame feature, the third second frame feature and the third third frame feature to obtain the second fused first frame feature, encodes the second fused first frame feature to obtain the second second frame feature.
[0138] In an embodiment of the present application, a computer device can downsample video frames to obtain multiple video frames of different scales. During the encoding process, cross-scale feature fusion is performed so that the encoding process can learn deeper and finer-grained features, which is beneficial to improving the accuracy and reliability of encoding and decoding.
[0139] Optionally, the m+1th second frame feature and the m+1th third frame feature are scaled so that the scales of the scaled m+1th second frame feature and the m+1th third frame feature are the same as the scale of the mth first frame feature, so that the scaled m+1th second frame feature and the m+1th third frame feature are fused with the mth first frame feature.
[0140] Optionally, taking the determination of the second second frame feature as an example, the feature encoding process can be expressed by the following formula (2).
[0141] E S2 =Encoder{SVFF(F S3 ), SVFF(D S3 ), F S2}; Formula (2)
[0142] Among them, the E S2 Indicates the second frame feature, E S3 Indicates the third second frame feature, D S3 Indicates the third frame feature, F S2 represents the second first frame feature, SVFF(·) represents the scale transformation, and Encoder{·} represents the encoding process.
[0143] In one possible implementation, a computer device encodes the mth fused first frame feature to obtain the mth second frame feature, including: determining weights of first sub-features on multiple regions in the mth fused first frame feature based on the mth fused first frame feature, the weight of the first sub-feature indicating the importance of the first sub-feature; encoding multiple first sub-features in the mth fused first frame feature based on the weights of multiple first sub-features in the mth fused first frame feature to obtain the mth second frame feature.
[0144] The first frame feature is a two-dimensional or three-dimensional feature, and the two-dimensional or three-dimensional first frame feature can be regarded as composed of first sub-features on multiple regions in the dimensional space. Taking the mth fused first frame feature as an example, the computer device determines the weight of the first sub-feature on each region in the fused first frame feature based on the fused first frame feature, and the weight of the first sub-feature indicates the importance of the first sub-feature. After obtaining the weight of the first sub-feature of each region, the computer device encodes any first sub-feature based on the weight of the first sub-feature, and then after encoding multiple first sub-features, the second frame feature can be obtained, and the second frame feature is composed of the features obtained by encoding the first sub-features of multiple regions. Among them, the second frame feature can be understood as the feature obtained by encoding the entire fused first frame feature.
[0145] Optionally, the computer device determines weights of first sub-features on multiple regions in the mth fused first frame feature based on the mth fused first frame feature, including: dividing the mth fused first frame feature into multiple first sub-features according to regions; for each first sub-feature, using a first predicted convolution kernel to convolve the first sub-feature to obtain the weight of the first sub-feature.
[0146] For any first sub-feature, the computer device uses the first prediction convolution kernel to convolve the first sub-feature to obtain the weight of the first sub-feature. The computer device can obtain the weight of each first sub-feature by using this method. The weight of the first sub-feature indicates the importance of the first sub-feature. The first prediction convolution kernel is a pre-stored convolution kernel, for example, the first prediction convolution kernel is a convolution kernel in a deep learning model, and the first prediction convolution kernel is obtained by training the deep learning model. The deep learning model can be as follows: Figure 5 Video frame processing model in an embodiment of the present invention.
[0147] Optionally, the computer device encodes multiple first sub-features in the mth fused first frame feature based on the weights of the multiple first sub-features in the mth fused first frame feature to obtain the mth second frame feature, including: for each first sub-feature in the mth fused first frame feature, multiplying the weight of the first sub-feature by the first sub-feature to obtain an encoded feature of the first sub-feature; and combining the encoded features of multiple first sub-features to obtain the mth second frame feature.
[0148] For any first sub-feature, the computer device multiplies the weight of the first sub-feature by the first sub-feature to obtain the encoded feature of the first sub-feature. The computer device can use this method to obtain the encoded feature of each first sub-feature. The encoded features of the multiple first sub-features can constitute the second frame feature, and the second frame feature is the feature obtained by encoding the entire first frame feature.
[0149] Optionally, the weight of the first sub-feature is a weight kernel, the weight kernel is data in matrix form, and the first sub-feature is also data in matrix form. The computer device cross-multiplies the weight kernel of the first sub-feature with the first sub-feature to obtain the encoded feature of the first sub-feature.
[0150] 304. When m is equal to k-1, the computer device encodes the m+1th first frame feature to obtain the m+1th second frame feature.
[0151] Wherein, when m is equal to k-1, the m+1th first frame feature is also the kth first frame feature, and the kth first frame feature is also the last frame feature among the k first frame features. For the m+1th first frame feature, there is no need to fuse the encoding features or decoding features of other scales, and the m+1 first frame features can be directly encoded.
[0152] In one possible implementation, the computer device determines weights of first sub-features on multiple regions in the m+1th first frame feature based on the m+1th first frame feature, and encodes multiple first sub-features in the m+1th first frame feature based on the weights of the multiple first sub-features in the m+1th first frame feature to obtain the m+1th second frame feature.
[0153] 305. When m is not equal to 1, the computer device decodes the mth second frame feature to obtain the mth third frame feature for the mth second frame feature; when m is equal to k-1, the computer device decodes the m+1th second frame feature to obtain the m+1th third frame feature for the m+1th second frame feature.
[0154] According to the above steps 303 and 304, the computer device can obtain k second frame features.
[0155] Wherein, when m is not equal to 1, the mth second frame feature is any other second frame feature except the 1st second frame feature and the kth second frame feature. When m is equal to k-1, the m+1th second frame feature is also the kth second frame feature. That is, for any other second frame feature except the 1st second frame feature, the computer device can obtain the corresponding third frame feature by decoding the second frame feature.
[0156] In one possible implementation, for the mth second frame feature, the computer device determines the weights of the second sub-features in multiple regions of the mth second frame feature based on the mth second frame feature, and decodes the multiple second sub-features in the mth second frame feature based on the weights of the multiple second sub-features in the mth second frame feature to obtain the mth third frame feature. For the m+1th second frame feature, the computer device determines the weights of the second sub-features in multiple regions of the m+1th second frame feature based on the m+1th second frame feature, and decodes the multiple second sub-features in the m+1th second frame feature based on the weights of the multiple second sub-features in the m+1th second frame feature to obtain the m+1th third frame feature.
[0157] Optionally, taking the determination of the second third frame feature as an example, the feature decoding process can be expressed by the following formula (3).
[0158] D S2 =Decoder(E S2 ); Formula (3)
[0159] Among them, D S2 Indicates the second and third frame features, E S2 represents the second frame feature, and Decoder(·) represents the decoding process.
[0160] 306. The computer device decodes the first second frame feature to obtain a deblurred video frame.
[0161] In one possible implementation, the first second frame feature is decoded to obtain a deblurred video frame, including: determining weights of second sub-features on multiple regions in the first second frame feature based on the first second frame feature, the weight of the second sub-feature indicating the importance of the second sub-feature; and decoding multiple second sub-features in the first second frame feature based on the weights of multiple second sub-features in the first second frame feature to obtain a deblurred video frame.
[0162] The second frame feature is a two-dimensional or three-dimensional feature, and the two-dimensional or three-dimensional second frame feature can be regarded as composed of second sub-features in multiple regions in the dimensional space. Based on the second frame feature, the computer device determines the weight of the second sub-feature in each region of the second frame feature, and the weight of the second sub-feature indicates the importance of the second sub-feature. After obtaining the weight of the second sub-feature in each region, the computer device decodes any second sub-feature based on the weight of the second sub-feature, and then after decoding multiple first sub-features, a deblurred video frame can be obtained, and the deblurred video frame is also a video frame obtained by deblurring the above-mentioned blurred video frame.
[0163] Optionally, the computer device determines weights of second sub-features on multiple regions in the first second frame feature based on the first second frame feature, including: dividing the first second frame feature into multiple second sub-features according to the region; for each second sub-feature, using a second prediction convolution kernel, convolving the second sub-feature to obtain the weight of the second sub-feature.
[0164] For any second sub-feature, the computer device uses the second prediction convolution kernel to convolve the second sub-feature to obtain the weight of the second sub-feature. The computer device can obtain the weight of each second sub-feature by using this method. The weight of the second sub-feature indicates the importance of the second sub-feature. The second prediction convolution kernel is a pre-stored convolution kernel, for example, the second prediction convolution kernel is a convolution kernel in a deep learning model, and the second prediction convolution kernel is obtained by training the deep learning model. The deep learning model can be as follows: Figure 5 Video frame processing model in an embodiment of the present invention.
[0165] Optionally, the computer device decodes multiple second sub-features in the first second frame feature based on the weights of the multiple second sub-features in the first second frame feature to obtain a deblurred video frame, including: for each second sub-feature in the first second frame feature, multiplying the weight of the second sub-feature by the second sub-feature to obtain a decoded feature of the second sub-feature; and combining the decoded features of multiple second sub-features to obtain a deblurred video frame.
[0166] For any second sub-feature, the computer device multiplies the weight of the second sub-feature by the second sub-feature to obtain a decoded feature of the second sub-feature. The computer device can obtain the decoded feature of each second sub-feature by using this method. The decoded features of the multiple second sub-features can constitute a deblurred video frame, and the deblurred video frame is a video frame obtained by decoding the entire second frame feature.
[0167] Optionally, the weight of the second sub-feature is a weight kernel, the weight kernel is data in matrix form, and the second sub-feature is also data in matrix form. The computer device cross-multiplies the weight kernel of the second sub-feature with the second sub-feature to obtain a decoded feature of the second sub-feature.
[0168] Since the features of different regions have their own weights, the features of the regions are encoded and decoded based on their respective weights, so that the encoding and decoding processes are spatially varying. Therefore, the encoding and decoding processes can focus more on important features in the spatial dimension, thereby improving the accuracy and reliability of encoding and decoding.
[0169] The method provided in the embodiment of the present application, when deblurring a blurred video frame, takes into account the influence of the two preceding and succeeding video frames adjacent to the blurred video frame on the blurred video frame because continuous video frames are correlated, and downsamples the video frame to obtain multiple video frames of different scales. During the encoding process, cross-scale feature fusion is performed, so that the encoding process can learn deeper and more fine-grained features, fully utilize important features in the spatial dimension, and improve the accuracy of video frame deblurring while reducing the amount of calculation.
[0170] Moreover, since only adjacent video frames are used to compensate for blurred video frames, the dependence on other video frames that are far apart in the time dimension is reduced. Then, in the encoding and decoding process, encoding and decoding are performed based on the weights of features in different regions, taking into account the importance of features in different regions, so that the encoding and decoding process is more focused on features that are more important in the spatial dimension. Therefore, the method of the embodiment of the present application reduces the dependence on features in the time dimension, makes full use of important features in the spatial dimension, and further improves the accuracy of video frame deblurring.
[0171] In some embodiments, based on the above embodiments, the video frame deblurring method is implemented by a video frame processing model. The detailed process is as follows: Figure 5 Embodiment of the invention. Figure 5 is a flowchart of another video frame deblurring method provided in an embodiment of the present application. The embodiment of the present application is executed by a computer device. Figure 5 , the method comprising:
[0172] 501. A computer device obtains a blurred video frame, a first video frame, and a second video frame, wherein the first video frame is a video frame preceding the blurred video frame, and the second video frame is a video frame following the blurred video frame.
[0173] The process of obtaining the video frame in step 501 is the same as the process of obtaining the video frame in step 201 above, and will not be repeated here.
[0174] 502. The computer device fuses the blurred video frame, the first video frame and the second video frame through a video frame processing model to obtain a fused video frame, performs feature extraction on the fused video frame to obtain first frame features of k different scales, the scale of the mth first frame feature is greater than the scale of the m+1th first frame feature, k is a positive integer greater than 1, and m is any positive integer from 1 to k-1.
[0175] In one possible implementation, the video frame processing model includes a fusion network, which is used to perform inter-frame shifting to fuse the video frames. The computer device inputs the blurred video frame, the first video frame, and the second video frame into the fusion network, and the fusion network fuses the blurred video frame, the first video frame, and the second video frame to output a fused video frame.
[0176] Optionally, the computer device determines the first region information and the second region information through a fusion network, the first region information and the second region information indicating regions in a video frame, replaces the content in the region indicated by the first region information in the blurred video frame with the content in the region indicated by the first region information in the first video frame, and replaces the content in the region indicated by the second region information in the blurred video frame with the content in the region indicated by the second region information in the second video frame, to obtain a fused video frame.
[0177] In one possible implementation, the video frame processing model includes a feature extraction network. The computer device extracts features from the fused video frame through the feature extraction network to obtain first frame features of k different scales. That is, the computer device inputs the fused video frame into the feature extraction network, and the feature extraction network outputs the first frame features. The feature extraction network may be a convolutional network.
[0178] In a possible implementation, the video frame processing model includes k-1 downsampling networks, k fusion networks and k feature extraction networks. The mth fusion network is connected to the mth feature extraction network, the mth fusion network is also connected to the m-1th downsampling network, and the first fusion network is not connected to any downsampling network. k is a positive integer greater than 1, and m is a positive integer less than k.
[0179] The computer device determines the blurred video frame, the first video frame and the second video frame as the first group of video frames, and performs k-1 downsampling of different scales on the first group of video frames through k-1 downsampling networks to obtain the second group of video frames to the kth group of video frames, wherein the scale of the mth group of video frames is greater than the scale of the m+1th group of video frames. That is, the first group of video frames is a video frame of the original scale, and the computer device inputs the first group of video frames into the first downsampling network and outputs the second group of video frames, and the scale of the first group of video frames is greater than the scale of the second group of video frames. The computer device inputs the second group of video frames into the second downsampling network and outputs the third group of video frames, and the scale of the second group of video frames is greater than the scale of the third group of video frames. By analogy, the k-1 downsampling networks can obtain k-1 groups of video frames with different scales from the first group of video frames, thereby obtaining k groups of video frames with different scales, each group of video frames includes the blurred video frame, the first video frame and the second video frame, and the scales of the three video frames in each group of video frames are the same.
[0180] The computer device fuses the blurred video frame, the first video frame, and the second video frame in the mth group of video frames through the mth fusion network to obtain the mth fused video frame. The computer device extracts features from the mth fused video frame through the mth feature extraction network to obtain the mth first frame feature. The blurred video frame, the first video frame, and the second video frame in the kth group of video frames are fused through the kth fusion network to obtain the kth fused video frame, and features are extracted from the kth fused video frame to obtain the kth first frame feature.
[0181] That is, the computer device inputs the mth group of video frames into the mth fusion network, the mth fusion network outputs the mth fused video frame, the mth fused video frame is input into the mth feature extraction network, and the mth feature extraction network outputs the mth first frame feature. Similarly, the computer device inputs the kth group of video frames into the kth fusion network, the kth fusion network outputs the kth fused video frame, the kth fused video frame is input into the kth feature extraction network, and the kth feature extraction network outputs the kth first frame feature.
[0182] 503. The computer device fuses the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature through a video frame processing model to obtain the mth fused first frame feature, and encodes the mth fused first frame feature to obtain the mth second frame feature; wherein the m+1th second frame feature is the encoded feature of the m+1th first frame feature, and the m+1th third frame feature is the decoded feature of the m+1th second frame feature.
[0183] The video frame processing model includes k encoding networks and k decoding networks, the mth encoding network is connected to the mth feature extraction network, and the mth encoding network is connected to the mth decoding network.
[0184] The computer device fuses the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature through the mth encoding network to obtain the mth fused first frame feature, and encodes the mth fused first frame feature to obtain the mth second frame feature. That is, the mth feature extraction network outputs the mth first frame feature, the mth first frame feature is input to the mth encoding network, the m+1th encoding network outputs the m+1th second frame feature, the m+1th second frame feature is input to the mth encoding network, the m+1th decoding network outputs the m+1th third frame feature, the m+1th third frame feature is input to the mth encoding network, and the mth encoding network processes the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature to output the mth second frame feature.
[0185] In one possible implementation, a computer device fuses the mth first frame feature, the m+1th second frame feature, and the m+1th third frame feature through an mth encoding network to obtain an mth fused first frame feature, and determines weights of first sub-features on multiple regions in the mth fused first frame feature based on the mth fused first frame feature, wherein the weight of the first sub-feature represents the importance of the first sub-feature; and encodes multiple first sub-features in the mth fused first frame feature based on the weights of multiple first sub-features in the mth fused first frame feature to obtain an mth second frame feature.
[0186] Optionally, the computer device divides the mth fused first frame feature into multiple first sub-features according to regions through the mth encoding network, and for each first sub-feature, convolves the first sub-feature with a first prediction convolution kernel to obtain a weight of the first sub-feature, and for each first sub-feature, multiplies the weight of the first sub-feature by the first sub-feature to obtain an encoded feature of the first sub-feature, and combines the encoded features of multiple first sub-features to obtain an mth second frame feature.
[0187] Figure 6 is a schematic diagram of a coding network structure provided in an embodiment of the present application, such as Figure 6As shown, the encoding network includes a convolution layer 601, a weight prediction layer 602, a reshaping layer 603 and a cross multiplication operator, and the weight prediction layer includes a first prediction convolution kernel. The first frame feature includes multiple first sub-features, the first sub-feature is input into the convolution layer 601, the convolution layer 601 outputs the convolution result, the convolution result is input into the weight prediction layer 602, the weight prediction layer 602 outputs the weight feature of the first sub-feature, the weight feature of the first sub-feature is normalized, the normalized weight feature of the first sub-feature is input into the reshaping layer 603, the reshaping layer 603 performs a scale transformation on the weight feature of the first sub-feature to obtain the weight of the first sub-feature, the cross multiplication operator cross-multiplies the first sub-feature with the weight of the first sub-feature to obtain the encoded feature of the first sub-feature, and the encoded features of multiple first sub-features can constitute the second frame feature.
[0188] Optionally, the scale of the first frame feature is The scale of the second frame feature is β represents the scaling factor of the encoding network.
[0189] Optionally, the process of predicting weights can be expressed by the following formula (4).
[0190] W l =Softmax(Klay(Conv(Fin))); Formula (4)
[0191] Among them, W l represents weight, Fin represents the first frame feature, Conv(·) represents convolutional layer, Klay(·) represents weight prediction layer, and Softmax(·) represents normalization.
[0192] Optionally, the process of encoding to obtain the second frame features can be expressed by the following formula (5).
[0193]
[0194] Among them, F out represents the second frame feature, (x, y) represents the center pixel, W l (x, y) represents the weight, F in(x+m,y+m) represents the first sub-feature, m refers to the distance from the edge of the region to the center of the region, and r represents the region range of the first frame feature.
[0195] In one possible implementation, the video frame processing model also includes a scale transformation network. Except for the k-th encoding network and the k-th decoding network, each other encoding network and each decoding network is connected to a scale transformation network. Taking the m+1-th encoding network and the m+1-th decoding network as an example, the m+1-th second frame feature output by the m+1-th encoding network is first input into the scale transformation network for scale transformation, so that the scale of the m+1-th second frame feature after scale transformation is the same as the scale of the m-th first frame feature, and then the m+1-th second frame feature after scale transformation is input into the m-th encoding network, and the m+1-th third frame feature output by the m+1-th decoding network is first input into the scale transformation network for scale transformation, so that the scale of the m+1-th third frame feature after scale transformation is the same as the scale of the m-th first frame feature, and then the m+1-th third frame feature after scale transformation is input into the m-th encoding network. Optionally, the network structure of the scale transformation network is the same as the above. Figure 6 The network structure of the encoding network is the same, only the network parameters are different, so they will not be repeated.
[0196] 504. When m is equal to k-1, the computer device encodes the m+1th first frame feature through the video frame processing model to obtain the m+1th second frame feature.
[0197] The video frame processing model includes k encoding networks and k decoding networks, the mth encoding network is connected to the mth feature extraction network, and the mth encoding network is connected to the mth decoding network.
[0198] For the m+1th first frame feature, the computer device encodes the m+1th first frame feature through the m+1th encoding network to obtain the m+1th second frame feature.
[0199] In a possible implementation, the computer device determines the weights of the first sub-features on multiple regions in the m+1th first frame feature based on the m+1th first frame feature through the m+1th encoding network, and encodes the multiple first sub-features in the m+1th first frame feature based on the weights of the multiple first sub-features in the m+1th first frame feature to obtain the m+1th second frame feature. That is, the m+1th feature extraction network outputs the m+1th first frame feature, the m+1th first frame feature is input to the m+1th encoding network, and the m+1th encoding network outputs the m+1th second frame feature.
[0200] 505. When m is not equal to 1, the computer device decodes the mth second frame feature through the video frame processing model to obtain the mth third frame feature; when m is equal to k-1, the computer device decodes the m+1th second frame feature through the video frame processing model to obtain the m+1th third frame feature.
[0201] The video frame processing model includes k encoding networks and k decoding networks, the mth encoding network is connected to the mth feature extraction network, and the mth encoding network is connected to the mth decoding network.
[0202] The computer device decodes the mth second frame feature through the mth decoding network to obtain the mth third frame feature. The computer device decodes the m+1th second frame feature through the m+1th decoding network to obtain the m+1th third frame feature.
[0203] In one possible implementation, taking the mth second frame feature as an example, for the mth second frame feature, the computer device determines the weights of the second sub-features on multiple regions in the mth second frame feature based on the mth second frame feature through the mth decoding network, and decodes the multiple second sub-features in the mth second frame feature based on the weights of the multiple second sub-features in the mth second frame feature to obtain the mth third frame feature. That is, the mth encoding network outputs the mth second frame feature, the mth second frame feature is input into the mth decoding network, and the mth decoding network outputs the mth third frame feature. That is, the mth encoding network outputs the mth second frame feature, the mth second frame feature is input into the mth decoding network, and the mth decoding network outputs the mth third frame feature.
[0204] Optionally, the computer device divides the mth second frame feature into multiple second sub-features according to regions through the mth decoding network; for each second sub-feature, the second prediction convolution kernel is used to convolve the second sub-feature to obtain a weight of the second sub-feature. For each second sub-feature, the weight of the second sub-feature is multiplied by the second sub-feature to obtain a decoded feature of the second sub-feature, and the decoded features of the multiple second sub-features are combined to obtain the mth third frame feature.
[0205] 506. The computer device decodes the first second frame feature through a video frame processing model to obtain a deblurred video frame.
[0206] The video frame processing model includes k encoding networks and k decoding networks, the mth encoding network is connected to the mth feature extraction network, and the mth encoding network is connected to the mth decoding network.
[0207] For the first second frame feature, the computer device decodes the first second frame feature through the first decoding network to obtain a deblurred video frame.
[0208] In a possible implementation, the computer device determines the weights of the second sub-features on multiple regions in the first second frame feature based on the first second frame feature through the first decoding network, and decodes the multiple second sub-features in the first second frame feature based on the weights of the multiple second sub-features in the first second frame feature to obtain a deblurred video frame. That is, the first encoding network outputs the first second frame feature, the first second frame feature is input into the first decoding network, and the first decoding network outputs the deblurred video frame.
[0209] Optionally, the computer device divides the first second frame feature into multiple second sub-features according to regions through the first decoding network; for each second sub-feature in the first second frame feature, the second sub-feature is convolved using a second prediction convolution kernel to obtain a weight of the second sub-feature. For each second sub-feature, the weight of the second sub-feature is multiplied by the second sub-feature to obtain a decoded feature of the second sub-feature, and the decoded features of the multiple second sub-features are combined to obtain a deblurred video frame.
[0210] In one possible implementation, the network structure of the decoding network is the same as that described above. Figure 6 The network structures of the encoding networks shown are the same, only the network parameters are different, so they will not be described one by one.
[0211] The method provided in the embodiment of the present application uses a video frame processing model to process blurred video frames, a first video frame, and a second video frame to achieve deblurring of blurred video frames, thereby simplifying the processing flow and helping to improve the efficiency of video frame deblurring. Since encoding and decoding are performed based on the weights of features in different regions, the importance of features in different regions is taken into account, making the encoding and decoding process more focused on features that are more important in the spatial dimension. Therefore, the method in the embodiment of the present application reduces the dependence on features in the time dimension, makes full use of important features in the spatial dimension, and while reducing the amount of computation, also improves the accuracy of video frame deblurring.
[0212] Above Figure 5 The embodiment illustrates the process of deblurring by fusing multi-scale features. For ease of understanding, the structure of the video frame processing model is introduced in detail by taking the fusion of features at three scales as an example. Figure 7 is a structural diagram of a video frame processing model provided in an embodiment of the present application, such as Figure 7As shown, the video frame processing model includes 2 downsampling networks, 3 fusion networks, 3 feature extraction networks, 3 encoding networks, 3 decoding networks and multiple scale conversion networks. The blurred video frame, the first video frame and the second video frame are recorded as video frame L1. The original scale of video frame L1 is the first scale. The blurred video frame, the first video frame and the second video frame are downsampled by the first downsampling network to obtain video frame L2. The scale of video frame L2 is the second scale. The video frame L2 is downsampled by the second downsampling network to obtain video frame L3. The scale of video frame L3 is the third scale. The first scale is larger than the second scale, and the second scale is larger than the third scale.
[0213] like Figure 7 As shown, the video frame L3 is input into the third fusion network, and the fused video frame CS3 is output; the fused video frame CS3 is input into the third feature extraction network, and the first frame feature FS3 is output; the first frame feature FS3 is input into the third encoding network, and the second frame feature ES3 is output; the second frame feature ES3 is input into the third decoding network, and the third frame feature DS3 is output; and the second frame feature ES3 and the third frame feature DS3 are input into the scale conversion network for scale transformation.
[0214] like Figure 7 As shown, the video frame L2 is input into the second fusion network, and the fused video frame CS2 is output; the fused video frame CS2 is input into the second feature extraction network, and the first frame feature FS2 is output; the first frame feature FS2, the second frame feature ES3 after scale transformation, and the third frame feature DS3 after scale transformation are input into the second encoding network, and the second frame feature ES2 is output; the second frame feature ES2 is input into the second decoding network, and the third frame feature DS2 is output; and the second frame feature ES2 and the third frame feature DS2 are input into the scale conversion network for scale transformation.
[0215] like Figure 7 As shown, the video frame L1 is input into the first fusion network to output the fused video frame CS1, the fused video frame CS1 is input into the first feature extraction network to output the first frame feature FS1, the first frame feature FS1, the scaled second frame feature ES2 and the scaled third frame feature DS2 are input into the first encoding network to output the second frame feature ES1, the second frame feature ES1 is input into the first decoding network to output the deblurred video frame DS1.
[0216] The training process of the video frame processing model in the above embodiment can be seen in the following Figure 8 Embodiment of the invention. Figure 8 is a flowchart of a method for training a video frame processing model provided in an embodiment of the present application. The embodiment of the present application is executed by a computer device. Figure 8 , the method comprising:
[0217] 801. A computer device obtains a sample blurred video frame, a sample first video frame, a sample second video frame and a sample clear video frame, wherein the sample first video frame is a video frame before the sample blurred video frame, the sample second video frame is a video frame after the sample blurred video frame, and the sample clear video frame is a video frame after the sample blurred video frame is deblurred.
[0218] Among them, the sample blurred video frame, the sample first video frame, and the sample second video frame are the same as the blurred video frame, the first video frame, and the second video frame in the embodiment of the above-mentioned video frame deprocessing method, and are not repeated here.
[0219] The sample clear video frame is a video frame after deblurring the sample blurry video frame, that is, a sufficiently clear video frame obtained by other means, such as a clear video frame obtained by manual adjustment, etc. The purpose of the embodiment of the present application is to use the above sample data to train a video frame processing model.
[0220] 802. The computer device inputs the sample blurred video frame, the sample first video frame and the sample second video frame into a video frame processing model to obtain a predicted video frame output by the video frame processing model.
[0221] The process of step 802 is the same as that of step Figure 5 The process of the embodiment of the present invention is similar, and will not be described in detail here.
[0222] 803. The computer device trains the video frame processing model based on the difference between the predicted video frame and the sample clear video frame.
[0223] The sample clear video frame is a video frame that is deblurred from the sample blurred video frame. The smaller the difference between the predicted video frame output by the video frame processing model and the sample clear video frame, the more accurate the video frame processing model is. Therefore, based on the difference between the predicted video frame and the sample clear video frame, the video frame processing model can be trained to reduce the difference between the predicted video frame obtained by the trained video frame processing model and the sample clear video frame.
[0224] In one possible implementation, the computer device trains the video frame processing model based on the difference between the predicted video frame and the sample clear video frame and the difference between the edge gradient of the predicted video frame and the edge gradient of the sample clear video frame.
[0225] Optionally, the computer device determines a first loss value based on the difference between the predicted video frame and the sample clear video frame; determines a second loss value based on the difference between the edge gradient of the predicted video frame and the edge gradient of the sample clear video frame; performs weighted summation of the first loss value and the second loss value to obtain a third loss value; and trains a video frame processing model based on the third loss value.
[0226] The computer device uses the following formula (6)-formula (8) to determine the third loss value.
[0227]
[0228]
[0229] L=L char +λL edge ; Formula (8)
[0230] Among them, L char Represents the first loss value, L edge represents the second loss value, L represents the third loss value, R represents the predicted video frame, G represents the sample clear video frame, ΔR represents the edge gradient of the predicted video frame, ΔG represents the edge gradient of the sample clear video frame. ε represents an adjustable parameter, and λ represents a weight adjustment parameter.
[0231] The method provided in the embodiment of the present application uses a sample blurred video frame, a sample first video frame, a sample second video frame and a sample clear video frame to train a video frame processing model, so that the video frame processing model learns how to deblur the blurred video frame based on the blurred video frame and two adjacent video frames, and reconstructs the deblurred video frame, thereby improving the deblurring ability of the video frame processing model. The video frame processing model is subsequently used to deblur the video frame, which simplifies the processing flow and is conducive to improving the efficiency of video frame deblurring.
[0232] Video frame deblurring is a challenging task. Since the blur in a video frame is usually spatially variable, the related art mainly performs deblurring based on the long-term temporal correspondence between video frames. The blurred video frame is deblurred by using multiple consecutive video frames around the blurred video frame. This has a strong dependence on information in the time dimension and a large amount of computation. In addition, the video frames far away from the blurred video frame have a weaker correlation with the blurred video frame. Therefore, deblurring with reference to these video frames will lead to misleading results and poor results.
[0233] In this regard, an embodiment of the present application provides a deblurring solution based on the short-term temporal correspondence between video frames, abandons the method of using video frames in the long time dimension as a reference, and only refers to the two adjacent video frames. During the encoding and decoding process, the weight of the features of each area is adaptively determined, and encoding and decoding are performed according to the weights, so that the encoding and decoding process focuses more on important features in the spatial dimension, better captures the spatial knowledge of video frame deblurring, and reduces the amount of calculation while improving the accuracy of video frame deblurring.
[0234] In order to verify the effectiveness of the embodiments of the present application, experiments were conducted on different data sets using the methods provided in the embodiments of the present application and the methods provided in Related Technologies 1 to 5. The experimental results are shown in Figure 2. Figure 9-12 As shown in Tables 1 to 3. Fig. 9 Table 1 and Table 1 are the results of experiments on the Go-Pro dataset. Fig.10 Table 2 and Table 2 are the results of experiments on the DVD dataset. Fig.11 Table 3 and Table 3 are the results of experiments on the RED dataset. Fig.12 These are the results of experiments on real fuzzy datasets.
[0235] Table 1
[0236]
[0237] Table 2
[0238]
[0239] Table 3
[0240]
[0241] Among them, PSNR stands for Peak Signal-to-Noise Ratio, and SSIM stands for Structural Similarity. Figure 9-12 From the comparison with Tables 1 to 3, it can be seen that the effect of the video frame deblurring method provided by the embodiment of the present application is better than the effect of the video frame deblurring method provided by the related art.
[0242] In order to further verify the effectiveness of the embodiment of the present application, 20 consecutive video frames, 10 consecutive video frames and 3 consecutive video frames are used as experimental data for deblurring. The experimental results are as follows: Fig.13 As shown, from Fig.13 It can be seen from the figure that the result of deblurring using 3 consecutive video frames in the embodiment of the present application is better than the result of deblurring using 10 consecutive video frames or 20 consecutive video frames.
[0243] Fig.14 is a schematic diagram of the structure of a video frame deblurring device provided in an embodiment of the present application. Fig.14 , the device comprises:
[0244] The video frame acquisition module 1401 is used to acquire a blurred video frame, a first video frame, and a second video frame, wherein the first video frame is a video frame before the blurred video frame, and the second video frame is a video frame after the blurred video frame;
[0245] The feature extraction module 1402 is used to fuse the blurred video frame, the first video frame and the second video frame to obtain a fused video frame, and perform feature extraction on the fused video frame to obtain first frame features of k different scales, wherein the scale of the mth first frame feature is greater than the scale of the m+1th first frame feature, where k is a positive integer greater than 1, and m is any positive integer from 1 to k-1;
[0246] The encoding module 1403 is used to fuse the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature to obtain the mth fused first frame feature, and encode the mth fused first frame feature to obtain the mth second frame feature; wherein the m+1th second frame feature is the encoding feature of the m+1th first frame feature, and the m+1th third frame feature is the decoding feature of the m+1th second frame feature;
[0247] The decoding module 1404 is used to decode the first second frame feature to obtain a deblurred video frame.
[0248] The video frame deblurring device provided in the embodiment of the present application, when deblurring a blurred video frame, takes into account the influence of the two video frames before and after the blurred video frame on the blurred video frame because continuous video frames are correlated, and downsamples the video frame to obtain multiple video frames of different scales. During the encoding process, cross-scale feature fusion is performed, so that the encoding process can learn deeper and finer-grained features, fully utilize important features in the spatial dimension, and reduce the amount of calculation while also improving the accuracy of video frame deblurring.
[0249] Optionally, the encoding module 1403 is further configured to, when m is equal to k-1, encode the m+1th first frame feature to obtain the m+1th second frame feature for the m+1th first frame feature;
[0250] The decoding module 1404 is also used to, when m is not equal to 1, decode the mth second frame feature to obtain the mth third frame feature for the mth second frame feature; and when m is equal to k-1, decode the m+1th second frame feature to obtain the m+1th third frame feature for the m+1th second frame feature.
[0251] Optionally, the feature extraction module 1402 is used to:
[0252] The blurred video frame, the first video frame and the second video frame are determined as the first group of video frames, and the first group of video frames is downsampled k-1 times at different scales to obtain the second group of video frames to the kth group of video frames, wherein the scale of the mth group of video frames is greater than the scale of the m+1th group of video frames;
[0253] For each group of k groups of video frames, the blurred video frame, the first video frame and the second video frame in the current group of video frames are fused to obtain a fused video frame, and features are extracted from the current fused video frame to obtain a first frame feature.
[0254] Optionally, the feature extraction module 1402 is used to:
[0255] Determine first region information and second region information, the first region information and the second region information indicating regions in the video frame;
[0256] The content in the area indicated by the first area information in the blurred video frame is replaced with the content in the area indicated by the first area information in the first video frame, and the content in the area indicated by the second area information in the blurred video frame is replaced with the content in the area indicated by the second area information in the second video frame to obtain a fused video frame.
[0257] Optionally, the feature extraction module 1402 is used to:
[0258] The blurred video frame, the first video frame and the second video frame are fused to obtain a first fused video frame;
[0259] Down-sampling the first fused video frame at different scales k-1 times to obtain the second fused video frame to the kth fused video frame, wherein the scale of the mth fused video frame is larger than the scale of the m+1th fused video frame;
[0260] For each of the k fused video frames, a feature of the current fused video frame is extracted to obtain a first frame feature.
[0261] Optionally, the encoding module 1403 is used to:
[0262] Based on the mth fused first frame feature, determine the weight of the first sub-feature on multiple regions in the mth fused first frame feature, where the weight of the first sub-feature indicates the importance of the first sub-feature;
[0263] Based on the weights of the multiple first sub-features in the m-th fused first frame feature, the multiple first sub-features in the m-th fused first frame feature are encoded to obtain the m-th second frame feature.
[0264] Optionally, the encoding module 1403 is used to:
[0265] Divide the mth fused first frame feature into multiple first sub-features according to the region;
[0266] For each first sub-feature, a first prediction convolution kernel is used to convolve the first sub-feature to obtain a weight of the first sub-feature.
[0267] Optionally, the encoding module 1403 is used to:
[0268] For each first sub-feature in the mth fused first frame feature, multiply the weight of the first sub-feature by the first sub-feature to obtain an encoded feature of the first sub-feature;
[0269] The encoded features of multiple first sub-features are combined to obtain the mth second frame feature.
[0270] Optionally, the decoding module 1404 is used to:
[0271] Based on the first second frame feature, determine the weight of the second sub-feature on multiple regions in the first second frame feature, the weight of the second sub-feature indicating the importance of the second sub-feature;
[0272] Based on the weights of the multiple second sub-features in the first second frame feature, the multiple second sub-features in the first second frame feature are decoded to obtain a deblurred video frame.
[0273] Optionally, the decoding module 1404 is used to:
[0274] Dividing the first second frame feature into a plurality of second sub-features according to the region;
[0275] For each second sub-feature, a second prediction convolution kernel is used to convolve the second sub-feature to obtain a weight of the second sub-feature.
[0276] Optionally, the decoding module 1404 is used to:
[0277] For each second sub-feature in the first second frame feature, multiply the weight of the second sub-feature by the second sub-feature to obtain a decoded feature of the second sub-feature;
[0278] The decoded features of the plurality of second sub-features are combined to obtain a deblurred video frame.
[0279] Alternatively, see Fig.15 The video frame deblurring device is implemented by a video frame processing model, and the device also includes a model training module 1405 for:
[0280] Obtain a sample blurred video frame, a sample first video frame, a sample second video frame and a sample clear video frame, wherein the sample first video frame is a video frame before the sample blurred video frame, the sample second video frame is a video frame after the sample blurred video frame, and the sample clear video frame is a video frame after the sample blurred video frame is deblurred;
[0281] Inputting the sample blurred video frame, the sample first video frame and the sample second video frame into the video frame processing model to obtain a predicted video frame output by the video frame processing model;
[0282] The video frame processing model is trained based on the difference between the predicted video frame and the sample clear video frame.
[0283] Optionally, the model training module 1405 is used to:
[0284] The video frame processing model is trained based on the difference between the predicted video frame and the sample clear video frame and the difference between the edge gradient of the predicted video frame and the edge gradient of the sample clear video frame.
[0285] Optionally, the model training module 1405 is used to:
[0286] Determining a first loss value based on a difference between the predicted video frame and the sample clear video frame;
[0287] Determine a second loss value based on a difference between an edge gradient of the predicted video frame and an edge gradient of the sample clear video frame;
[0288] Perform a weighted summation of the first loss value and the second loss value to obtain a third loss value;
[0289] Based on the third loss value, the video frame processing model is trained.
[0290] It should be noted that the video frame deblurring device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the video frame deblurring device provided in the above embodiment and the video frame deblurring method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0291] An embodiment of the present application also provides a computer device, which includes a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the video frame deblurring method of the above embodiment.
[0292] Optionally, the computer device is provided as a terminal. Fig.16 A schematic diagram of the structure of a terminal 1600 provided by an exemplary embodiment of the present application is shown.
[0293] The terminal 1600 includes a processor 1601 and a memory 1602 .
[0294] The processor 1601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1601 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1601 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0295] The memory 1602 may include one or more computer-readable storage media, which may be non-transitory. The memory 1602 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1602 is used to store at least one computer program, which is used to be possessed by the processor 1601 to implement the video frame deblurring method provided in the method embodiment of the present application.
[0296] In some embodiments, the terminal 1600 may further optionally include: a peripheral device interface 1603 and at least one peripheral device. The processor 1601, the memory 1602 and the peripheral device interface 1603 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 1603 via a bus, a signal line or a circuit board. Optionally, the peripheral device includes: at least one of a radio frequency circuit 1604, a display screen 1605, a camera assembly 1606, an audio circuit 1607 and a power supply 1608.
[0297] The peripheral device interface 1603 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1601 and the memory 1602. In some embodiments, the processor 1601, the memory 1602, and the peripheral device interface 1603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1601, the memory 1602, and the peripheral device interface 1603 may be implemented on a separate chip or circuit board, which is not limited in this embodiment.
[0298] The radio frequency circuit 1604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1604 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1604 converts the electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The radio frequency circuit 1604 can communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes, but is not limited to: a metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1604 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.
[0299] The display screen 1605 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1605 is a touch display screen, the display screen 1605 also has the ability to collect touch signals on the surface or above the surface of the display screen 1605. The touch signal can be input to the processor 1601 as a control signal for processing. At this time, the display screen 1605 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 1605 can be one, set on the front panel of the terminal 1600; in other embodiments, the display screen 1605 can be at least two, respectively set on different surfaces of the terminal 1600 or in a folding design; in other embodiments, the display screen 1605 can be a flexible display screen, set on a curved surface or a folding surface of the terminal 1600. Even, the display screen 1605 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1605 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0300] The camera assembly 1606 is used to capture images or videos. Optionally, the camera assembly 1606 includes a front camera and a rear camera. The front camera is arranged on the front panel of the terminal 1600, and the rear camera is arranged on the back of the terminal 1600. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1606 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0301] The audio circuit 1607 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 1601 for processing, or input them into the radio frequency circuit 1604 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1600. The microphone may also be an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 1601 or the radio frequency circuit 1604 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 1607 may also include a headphone jack.
[0302] The power supply 1608 is used to power various components in the terminal 1600. The power supply 1608 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 1608 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0303] Those skilled in the art will understand that Fig.16 The structure shown in the figure does not constitute a limitation on the terminal 1600, and the terminal 1600 may include more or less components than those shown in the figure, or combine some components, or adopt a different component arrangement.
[0304] Optionally, the computer device is provided as a server. Fig.17 It is a structural diagram of a server provided in an embodiment of the present application. The server 1700 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1701 and one or more memories 1702, wherein the memory 1702 stores at least one computer program, and the at least one computer program is loaded and executed by the processor 1701 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output, and the server may also include other components for implementing device functions, which will not be described in detail here.
[0305] An embodiment of the present application further provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed by the video frame deblurring method of the above embodiment.
[0306] The embodiment of the present application further provides a computer program product, including a computer program, wherein the computer program is loaded and executed by a processor to implement the operations performed by the video frame deblurring method of the above embodiment.
[0307] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0308] The above description is only an optional embodiment of the embodiments of the present application and is not intended to limit the embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the protection scope of the present application.
Claims
1. A video frame deblurring method, It is characterized in that The method comprises: Acquire a blurred video frame, a first video frame, and a second video frame, wherein the first video frame is a video frame before the blurred video frame, and the second video frame is a video frame after the blurred video frame; The blurred video frame, the first video frame and the second video frame are fused to obtain a fused video frame, and features are extracted from the fused video frame to obtain first frame features of k different scales, wherein the scale of the mth first frame feature is greater than the scale of the m+1th first frame feature, where k is a positive integer greater than 1, and m is any positive integer from 1 to k-1; For the mth first frame feature, the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature are fused to obtain the mth fused first frame feature, and the mth fused first frame feature is encoded to obtain the mth second frame feature; wherein the m+1th second frame feature is the encoding feature of the m+1th first frame feature, and the m+1th third frame feature is the decoding feature of the m+1th second frame feature; For the first second frame feature, the first second frame feature is decoded to obtain a deblurred video frame.
2. The method according to claim 1, It is characterized in that The method further comprises: When m is equal to k-1, for the m+1th first frame feature, encode the m+1th first frame feature to obtain the m+1th second frame feature; When m is not equal to 1, for the mth second frame feature, the mth second frame feature is decoded to obtain the mth third frame feature; when m is equal to k-1, for the m+1th second frame feature, the m+1th second frame feature is decoded to obtain the m+1th third frame feature.
3. The method according to claim 1, It is characterized in that The step of fusing the blurred video frame, the first video frame and the second video frame to obtain a fused video frame, and extracting features from the fused video frame to obtain k first frame features of different scales includes: Determine the blurred video frame, the first video frame and the second video frame as a first group of video frames, perform k-1 downsampling of different scales on the first group of video frames to obtain a second group of video frames to a kth group of video frames, wherein the scale of the mth group of video frames is greater than the scale of the m+1th group of video frames; For each group of k groups of video frames, the blurred video frame, the first video frame and the second video frame in the current group of video frames are fused to obtain a fused video frame, and features are extracted from the current fused video frame to obtain a first frame feature.
4. The method according to claim 3, It is characterized in that The step of fusing the blurred video frame, the first video frame and the second video frame in the current group of video frames to obtain a fused video frame includes: Determine first region information and second region information, wherein the first region information and the second region information indicate regions in a video frame; The content in the area indicated by the first area information in the blurred video frame is replaced with the content in the area indicated by the first area information in the first video frame, and the content in the area indicated by the second area information in the blurred video frame is replaced with the content in the area indicated by the second area information in the second video frame to obtain the fused video frame.
5. The method according to claim 1, It is characterized in that The step of fusing the blurred video frame, the first video frame and the second video frame to obtain a fused video frame, and extracting features from the fused video frame to obtain k first frame features of different scales includes: Fusing the blurred video frame, the first video frame and the second video frame to obtain a first fused video frame; Downsampling the first fused video frame at different scales k-1 times to obtain second to k-th fused video frames, wherein the scale of the m-th fused video frame is greater than the scale of the m+1-th fused video frame; For each of the k fused video frames, a feature of the current fused video frame is extracted to obtain a first frame feature.
6. The method according to claim 1, It is characterized in that The encoding of the mth fused first frame feature to obtain the mth second frame feature includes: Based on the mth fused first frame feature, determine the weight of the first sub-feature on multiple regions in the mth fused first frame feature, where the weight of the first sub-feature indicates the importance of the first sub-feature; Based on the weights of the multiple first sub-features in the m-th fused first frame feature, the multiple first sub-features in the m-th fused first frame feature are encoded to obtain the m-th second frame feature.
7. The method according to claim 6, It is characterized in that The step of determining weights of first sub-features on multiple regions in the m-th fused first frame feature based on the m-th fused first frame feature includes: Divide the mth fused first frame feature into multiple first sub-features according to the region; For each first sub-feature, a first prediction convolution kernel is used to convolve the first sub-feature to obtain a weight of the first sub-feature.
8. The method according to claim 6, It is characterized in that The method of encoding the multiple first sub-features in the m-th fused first frame feature based on the weights of the multiple first sub-features in the m-th fused first frame feature to obtain the m-th second frame feature includes: For each first sub-feature in the mth fused first frame feature, multiplying the weight of the first sub-feature by the first sub-feature to obtain an encoded feature of the first sub-feature; The encoded features of the multiple first sub-features are combined to obtain the mth second frame feature.
9. The method according to claim 1, It is characterized in that The decoding of the first second frame feature to obtain a deblurred video frame includes: Based on the first second frame feature, determining weights of second sub-features on multiple regions in the first second frame feature, wherein the weights of the second sub-features represent the importance of the second sub-features; Based on the weights of the multiple second sub-features in the first second frame feature, the multiple second sub-features in the first second frame feature are decoded to obtain the deblurred video frame.
10. The method according to claim 9, It is characterized in that The step of determining the weights of the second sub-features on the plurality of regions in the first second frame feature based on the first second frame feature includes: Dividing the first second frame feature into a plurality of second sub-features according to the region; For each second sub-feature, a second prediction convolution kernel is used to perform convolution on the second sub-feature to obtain a weight of the second sub-feature.
11. The method according to any one of claims 1 to 10, It is characterized in that The video frame deblurring method is implemented by a video frame processing model. The training process of the video frame processing model includes: Acquire a sample blurred video frame, a sample first video frame, a sample second video frame and a sample clear video frame, wherein the sample first video frame is a video frame before the sample blurred video frame, the sample second video frame is a video frame after the sample blurred video frame, and the sample clear video frame is a video frame after the sample blurred video frame is deblurred; Inputting the sample blurred video frame, the sample first video frame and the sample second video frame into the video frame processing model to obtain a predicted video frame output by the video frame processing model; The video frame processing model is trained based on the difference between the predicted video frame and the sample clear video frame.
12. A video frame deblurring device, It is characterized in that The device comprises: A video frame acquisition module, used to acquire a blurred video frame, a first video frame and a second video frame, wherein the first video frame is a video frame before the blurred video frame, and the second video frame is a video frame after the blurred video frame; a feature extraction module, configured to fuse the blurred video frame, the first video frame and the second video frame to obtain a fused video frame, and perform feature extraction on the fused video frame to obtain first frame features of k different scales, wherein the scale of the mth first frame feature is greater than the scale of the m+1th first frame feature, where k is a positive integer greater than 1, and m is any positive integer from 1 to k-1; The encoding module is used for fusing the mth first frame feature, the m+1th second frame feature and the m+1th third frame feature to obtain the mth fused first frame feature, and encoding the mth fused first frame feature to obtain the mth second frame feature; wherein the m+1th second frame feature is the encoding feature of the m+1th first frame feature, and the m+1th third frame feature is the decoding feature of the m+1th second frame feature; A decoding module is used to decode the first second frame feature to obtain a deblurred video frame.
13. A computer device, It is characterized in that The computer device includes a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the video frame deblurring method according to any one of claims 1 to 11.
14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the video frame deblurring method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, It is characterized in that The computer program is loaded and executed by a processor to implement the operations performed by the video frame deblurring method according to any one of claims 1 to 11.