Video deblurring method and device, electronic equipment, storage medium and product
By adjusting the number of attention blocks in the residual channels and implementing multi-stage processing in the deblurring network, the problem of insufficient adaptability to complex blurred scenes in existing technologies is solved, and better video deblurring effects are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-17
AI Technical Summary
Existing video deblurring methods are not adaptable enough to complex and blurred scenes, resulting in poor deblurring results.
By adjusting the number of attention blocks in the first residual channel of each encoder in the trained deblurring network, multi-stage deblurring is performed based on the blur level of the blurred video frame. The blurred video frame is then determined by the blur detector and adaptively processed.
It improves adaptability to complex and blurry scenes, enhances the deblurring effect of videos, and ensures the clarity and viewing experience of videos.
Smart Images

Figure CN121883304A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a video deblurring method, apparatus, electronic device, computer storage medium, and computer program product. Background Technology
[0002] With the development and popularization of portable imaging devices, more and more users are using portable imaging devices for video recording. Unlike professional video imaging devices, the image quality of portable imaging devices is affected by factors such as optical components or environment, and is very prone to blurring. Therefore, it is necessary to perform deblurring processing on the video.
[0003] In related technologies, most video deblurring methods are designed for a single level of blur, which makes these methods less adaptable to complex blurred scenes and affects the deblurring effect of the video. Summary of the Invention
[0004] This application provides a video deblurring method, apparatus, electronic device, computer storage medium, and computer program product that can improve the deblurring effect of videos.
[0005] The technical solution of this application is implemented as follows:
[0006] This application provides a video deblurring method, the method comprising:
[0007] Identify the individual blurred video frames in the video to be processed;
[0008] For each blurred video frame, the blur level of the blurred video frame is obtained. Based on the blur level, the number of first residual channel attention blocks (RCABs) included in each encoder of the trained deblurring network is adjusted to obtain the target deblurring network. The first residual channel attention block is a residual channel attention block with shared weights in the encoder.
[0009] The target deblurring network is used to perform multi-stage deblurring processing on the blurred video frame to obtain the target video frame;
[0010] The target video is obtained by combining the target video frames of each blurred video frame with the unprocessed clear video frames in the video to be processed.
[0011] This application provides a video deblurring apparatus, the apparatus comprising:
[0012] The determination module is used to identify each blurred video frame in the video to be processed;
[0013] An adjustment module is used to obtain the blur level of each blurred video frame, and adjust the number of first residual channel attention blocks included in each encoder of the trained deblurring network according to the blur level to obtain the target deblurring network; the first residual channel attention block is the residual channel attention block with shared weights in the encoder.
[0014] The deblurring module is used to perform multi-stage deblurring processing on the blurred video frame using the target deblurring network to obtain the target video frame;
[0015] The combination module is used to combine the target video frames of each blurred video frame with the unprocessed clear video frames in the video to be processed to obtain the target video.
[0016] This application provides an electronic device, the device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the video deblurring method provided by one or more of the foregoing technical solutions.
[0017] This application provides a computer storage medium storing a computer program; when the computer program is executed, it can implement the video deblurring method provided by one or more of the aforementioned technical solutions.
[0018] This application provides a computer program product, including a computer program that, when executed by a processor, implements the video deblurring method provided by one or more of the aforementioned technical solutions.
[0019] This application provides a video deblurring method, apparatus, electronic device, computer storage medium, and computer program product. The method includes: determining each blurred video frame in a video to be processed; for each blurred video frame, obtaining the blur level of the blurred video frame; adjusting the number of first residual channel attention blocks included in each encoder of a trained deblurring network according to the blur level to obtain a target deblurring network; the first residual channel attention blocks are residual channel attention blocks with shared weights in the encoder; performing multi-stage deblurring processing on the blurred video frames using the target deblurring network to obtain target video frames; and combining the target video frames of each blurred video frame with unprocessed clear video frames in the video to be processed to obtain a target video.
[0020] As can be seen, in this embodiment, the number of RCABs with shared weights in each encoder of the network can be adaptively adjusted according to the different blur levels of each blurred video frame, so as to achieve adaptive processing of images with different blur levels. Compared with the existing deblurring methods for a single blur level, this embodiment can adaptively process blurred video frames with different blur levels, thus improving the adaptability to complex blurred scenes and ensuring the deblurring effect of the video. Attached Figure Description
[0021] Figure 1 A flowchart of a video deblurring method provided in an embodiment of this application;
[0022] Figure 2 A schematic diagram illustrating a process for determining blurred video frames using a blur detector, provided as an embodiment of this application;
[0023] Figure 3 A schematic diagram of a network structure for an RCAB provided in an embodiment of this application;
[0024] Figure 4 A schematic diagram of the network structure of a Supervised Attention Module (SAM) provided for an embodiment of this application;
[0025] Figure 5 This is a schematic diagram of the network structure corresponding to the first stage network in the deblurring network provided in the embodiments of this application;
[0026] Figure 6 This is a schematic diagram of the structure of a deblurring network provided in an embodiment of this application;
[0027] Figure 7 A schematic diagram of the network structure of a Residual Channel Attention Module (RCAM) provided for an embodiment of this application;
[0028] Figure 8 This is a schematic diagram of the composition structure of the video deblurring device provided in the embodiments of this application;
[0029] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0030] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are merely illustrative of the present application and are not intended to limit the present application. Furthermore, the embodiments provided below are some embodiments for implementing the present application, and not all embodiments for implementing the present application. Unless otherwise specified, the technical solutions described in the embodiments of the present application can be implemented in any combination.
[0031] It should be noted that, in the embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method or system that includes a list of elements includes not only the elements expressly described, but also other elements not expressly listed, or elements inherent to implementing the method or system. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other related elements (e.g., steps in the method or units in the system; for example, a unit may be a portion of circuitry, a portion of a processor, a portion of a program or software, etc.) in the method or system that includes that element.
[0032] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, I and / or J can represent three cases: I alone, I and J simultaneously, and J alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of elements. For example, including at least one of I, J, and R can mean including any one or more elements selected from the set consisting of I, J, and R.
[0033] For example, the video deblurring method provided in this application includes a series of steps, but the video deblurring method provided in this application is not limited to the steps described. Similarly, the video deblurring apparatus provided in this application includes a series of modules, but the video deblurring apparatus provided in this application is not limited to the modules explicitly described, but may also include modules that need to be set up for obtaining relevant information or processing based on information.
[0034] The following are various embodiments.
[0035] In some embodiments of this application, the video deblurring method can be implemented using a processor in a video deblurring device. The processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), controller, microcontroller, and microprocessor.
[0036] In this embodiment, the video deblurring method can be applied to various scenarios that require video deblurring, including but not limited to medical imaging, security monitoring, and autonomous driving. Its main function is to eliminate or reduce the blurring problem in the video caused by various reasons, thereby improving the clarity and viewing experience of the video.
[0037] Figure 1 A flowchart of a video deblurring method provided in an embodiment of this application is shown below. Figure 1 As shown, the process may include:
[0038] Step 100: Identify the individual blurred video frames in the video to be processed.
[0039] In this embodiment of the application, the video to be processed refers to the video that needs to be deblurred; here, the source of the video to be processed is not specifically limited. For example, the video to be processed can be the original video directly captured by a camera, portable imaging device or other shooting device, or the video obtained after preprocessing the original video, or the video downloaded from the network, etc.
[0040] For example, the video to be processed may include multiple video frames, wherein in addition to clear video frames, the multiple video frames also include several blurry video frames, that is, the number of blurry video frames may be one or more.
[0041] Here, the factors that cause blurry video frames are not limited. For example, they may be caused by one or more factors such as camera shake, movement of the subject, out-of-focus, and low resolution.
[0042] In this embodiment, the video to be processed can be obtained first, and then each blurred video frame can be determined from the video to be processed. Then, the determined blurred video frames can be subjected to subsequent deblurring processing. The process of determining blurred video frames is described below by way of example.
[0043] In some embodiments, determining each blurred frame in the video to be processed may include: inputting the video to be processed into a blur detector in the form of consecutive frames to obtain the category results of each video frame in the video to be processed; and determining each blurred video frame in the video to be processed based on the category results of each video frame in the video to be processed.
[0044] In this embodiment of the application, a fuzz detector can be used to determine the category result of each video frame in the video to be processed; here, each video frame corresponds to a type result; the category result is also called the category label, which is used to characterize the specific category of each video frame, and the category result may include one of the fuzzy category and the clear category.
[0045] For example, the category results of each video frame output by the blur detector can be in numerical form; for example, a value of 0 represents the sharp category and a value of 1 represents the blurry category; that is, if the category result of a video frame is 0, it means that the video frame is a blurry video frame, and if the category result of a video frame is 1, it means that the video frame is a sharp video frame.
[0046] As can be seen, in this embodiment of the application, each blurred video frame in the video to be processed is determined by a blur detector. In this way, the blurring process is only performed on the determined blurred video frames, which effectively avoids unnecessary calculations and over-enhancement of clear video frames, and improves the processing speed while ensuring the deblurring effect.
[0047] In some embodiments, the video to be processed is input into a blur detector in the form of consecutive frames to obtain the category results of each video frame in the video to be processed. This may include: extracting features from each video frame in the video to be processed to obtain feature points of each video frame; clustering the feature points of all video frames in the video to be processed to obtain multiple cluster centers; transforming the feature points of each video frame using multiple cluster centers to obtain feature vectors of each video frame; and inputting the feature vectors of each video frame into a classifier to obtain the category results of each video frame.
[0048] In this embodiment, the fuzz detector may include a feature extraction module, a feature aggregation module, and a classifier. The feature extraction module uses a feature extraction algorithm to extract features from each video frame in the video to be processed, obtaining feature points for each video frame. The feature aggregation module uses a Bag-of-Features (BOF) cluster to group the feature points of all video frames in the video to be processed, obtaining multiple cluster centers, and then uses these cluster centers to transform the feature points of each video frame, obtaining feature vectors for each video frame. The classifier is used to determine the corresponding category based on the feature vectors of each video frame.
[0049] It should be noted that, in addition to obtaining each feature point of each video frame using the feature extraction algorithm, the feature extraction module also obtains descriptors of these feature points; these descriptors describe the texture, shape and other information around the feature points in the video frame, which plays an important role in distinguishing between blurry and sharp categories.
[0050] Here, the type of feature extraction algorithm is not limited. For example, it can be the Speed Up Robust Features (SURF) algorithm or the Scale-Invariant Feature Transform (SIFT) algorithm.
[0051] For example, cluster centers, also known as visual vocabulary, can be considered as the basic visual elements that constitute an image. After obtaining multiple cluster centers, the feature points of each video frame are transformed using these cluster centers to obtain the feature vectors of each video frame. This can include: for each feature point in each video frame, the frequency of its occurrence at each cluster center can be counted; based on the counted frequency of each feature point, a feature vector is constructed, and this feature vector is determined as the feature vector of the video frame; wherein, each video frame corresponds to one feature vector.
[0052] For example, the feature vector of a video frame reflects the distribution of each visual element in the video frame; therefore, the feature vector obtained by using BOF is essentially a statistical and encoded representation of the feature points in the video frame; it can transform the originally high-dimensional and sparse feature points into low-dimensional and dense feature vectors, so that subsequent classifiers can process and recognize them.
[0053] In this embodiment of the application, after obtaining the feature vectors of each video frame, the feature vectors of each video frame can be input into the classifier to obtain the category result of each video frame; here, the type of classifier is not limited, for example, it can be a Support Vector Machine (SVM) classifier, or other types of classifiers.
[0054] For illustrative purposes, the above process is illustrated using the SURF algorithm and SVM classifier as examples. (See [link to documentation]). Figure 2 When video frame I t After inputting the blur detector, the SURF algorithm can be used to extract features from the video frame to obtain the feature points F of the video frame. t The two satisfy the following relationship: F t =SURF(I t Next, the feature points F of the video frame are determined using BOF. t eigenvector V t The two satisfy the following relationship: V t =BOF(F t Finally, the feature vector V of this video frame is... t The data is input into an SVM classifier to obtain the category result Label for each video frame. Further, if the Label for a video frame is determined to be 0, it indicates that the video frame is blurry; if the Label for a video frame is not 0, it indicates that the video frame is clear. Therefore, by performing the above steps on each video frame in the video to be processed, the blurry video frames in the video can be identified.
[0055] Step 101: For each blurred video frame, obtain the blur level of the blurred video frame. Based on the blur level, adjust the number of attention blocks in the first residual channel of each encoder in the trained deblurring network to obtain the target deblurring network.
[0056] In this embodiment of the application, after obtaining each blurred video frame in the video to be processed according to the above steps, the blur level of each blurred video frame can be obtained, and then the number of first residual channel attention blocks included in each encoder in the trained deblurring network can be adjusted according to the blur level.
[0057] For example, the blur level is used to characterize the degree of blur in a blurred video frame, and it can be represented by a numerical value; here, the numerical value can be positively correlated with the blur level.
[0058] For example, if the numerical value is positively correlated with the blur level, it means that the larger the value, the higher the blur level, and the greater the blurriness of the blurred video frame; correspondingly, the smaller the value, the lower the blur level, and the less blurriness of the blurred video frame. For instance, assuming that the blur level of blurred video frame 1 is 7 and the blur level of blurred video frame 2 is 9, it means that the blurriness of blurred video frame 2 is higher than that of blurred video frame 1.
[0059] In this embodiment, the blur level of the blurred video frame can be obtained first, and then the number of first RCABs included in each encoder in the trained deblurring network can be adjusted according to the blur level.
[0060] The first RCAB is the RCAB in the encoder that has shared weights; it should be noted that when there are multiple RCABs with shared weights, these RCABs can share parameters.
[0061] In this embodiment, the deblurring network, also known as the Multi-Stage Adaptive Deblurring Network (MAD-Net), is a three-stage deblurring network obtained by improving the Multi-Stage Progressive Image Restoration Network (MPRNet). Its overall network structure includes a three-stage cascaded network, which can progressively deblurr video frames with different levels of blurriness.
[0062] The first-stage network and the second-stage network have encoder and decoder structures. That is, both the first-stage network and the second-stage network include encoders. In other words, after obtaining the blur level of the blurred video frame, the number of first RCABs included in the encoders of the first-stage network and the second-stage network will be adjusted according to the blur level to obtain the target deblurring network.
[0063] Here, the encoder includes a first RCAB with shared weights and a second RCAB with unique weights; wherein the second RCAB is connected in series with the first RCAB.
[0064] It should be noted that the network structures of the first and second stages, in addition to the encoder and decoder, also include a convolutional module (RCAB) before the encoder and a supervised attention module (SAM) after the decoder. The network structures of the RCAB and SAM are as follows: Figure 3 and Figure 4 As shown, since RCAB and SAM are both existing modules in the Multi-Stage Progressive Image Restoration Network (MPRNet), they will not be discussed in detail here.
[0065] For example, the initial number of the first RCAB and the second RCAB is related to the training dataset of the deblurring network, and no specific limitation is made here.
[0066] Here, the encoder is also called an Adaptive Deblurring Network (AD-Net) encoder, which differs from the ordinary encoder in MPRNet. The key to the AD-Net encoder is that it concatenates a series of RCABs (first RCABs) with shared weights and RCABs (second RCABs) with independent weights. In this way, when the encoder processes each blurred video frame, it can adaptively adjust the number of first RCABs according to the blur level of each blurred video frame, that is, it can enable an early exit mechanism for different blur levels.
[0067] In some embodiments, adjusting the number of first RCABs included in each encoder of the trained deblurring network according to the fuzziness level may include: adjusting the number of first RCABs included in each encoder of the trained deblurring network according to the positive correlation between the fuzziness level and the number of first RCABs.
[0068] Here, the positive correlation can be a functional expression (e.g., a monotonically increasing function) or a mapping table; for example, the positive correlation can be manually set or obtained by training a defuzzification network.
[0069] Understandably, at higher blur levels, more first RCABs are typically needed to process the image features of the blurred video frames to ensure effective deblurring. Conversely, at lower blur levels, fewer first RCABs can be used to process the image features, reducing computational complexity while maintaining deblurring effectiveness. Therefore, based on the positive correlation between blur level and the number of first RCABs, the number of first RCABs included in each encoder of the trained deblurring network can be adjusted.
[0070] For example, both the first-stage network and the second-stage network of the deblurring network include an encoder, and the adjustment methods for both are the same; the adjustment process will be further explained below using the first-stage network as an example.
[0071] Figure 5 This is a schematic diagram of the network structure corresponding to the first stage network in the deblurring network provided in the embodiments of this application, as shown below. Figure 5As shown, the AD-Net encoder includes M RCABs with unique weights and N RCABs with shared weights. For example, when a blurred video frame is input to the first stage of the deblurring network, it is first convolved to expand its dimensions for subsequent extraction of richer feature information. Here, the size of the convolution kernel is not specifically limited; for example, it can be 3x3. Next, features in each dimension are extracted using RCABs. Then, the AD-Net encoder adjusts the number of RCABs with shared weights (i.e., the number of the first RCABs) in the network according to the blur level of the blurred video frame to extract deeper semantic features. Figure 5 It can be seen that blurred video frames with lower blur levels can use an earlier encoder output (corresponding to the topmost path, the dashed line indicates skipping the subsequent first RCAB, i.e. using less first RCAB processing); conversely, blurred video frames with higher blur levels can use a later encoder output (corresponding to the middle or bottommost path, i.e. using more first RCAB processing); finally, the output of the AD-Net encoder is sent to the same decoder and SAM for subsequent processing.
[0072] As can be seen, in this embodiment of the application, the deblurring network can adaptively adjust the network architecture of the encoder according to the blur level of different blurred video frames, so as to achieve adaptive processing of complex blur problems and different blur levels. In addition, since the RCABs with shared weights in the encoder can share parameters, the encoder does not introduce additional parameters to the entire network architecture. Therefore, it is not necessary to store additional parameters for different blur levels, which reduces computational complexity and improves processing efficiency.
[0073] Step 102: Use a target deblurring network to perform multi-stage deblurring on the blurred video frames to obtain the target video frames.
[0074] In this embodiment of the application, after obtaining the target deblurring network according to the above steps, the target deblurring network can be used to perform multi-stage deblurring processing on the blurred video frame to obtain the target video frame.
[0075] In some embodiments, using a target deblurring network to perform multi-stage deblurring processing on a blurred video frame may include: segmenting the blurred video frame into multiple non-overlapping first image blocks, extracting features from the first image blocks using the encoder of a first-stage network, and restoring the extracted features using the decoder of a first-stage network; segmenting the blurred video frame into multiple non-overlapping second image blocks, extracting features from the second image blocks using the encoder of a second-stage network, and restoring the extracted features using the decoder of a second-stage network; and using the blurred video frame as input to a third-stage network to calculate the resolution of the blurred video frame to obtain the target video frame.
[0076] In this embodiment, the first-stage network and the second-stage network respectively segment the input blurred video frame, dividing it into multiple non-overlapping image blocks. Specifically, the first-stage network can divide the blurred video frame into four non-overlapping image blocks (top left, bottom left, top right, and bottom right); the second-stage network can divide the blurred video frame into two non-overlapping image blocks (left and right); that is, the number of first image blocks is 4, and the number of second image blocks is 2. The original blurred video frame is used in the third-stage network. A SAM (Segmentation Amplifier) is also provided between adjacent stages. The SAM is used to pass the features output by the decoder to the next stage, such as... Figure 6 As shown.
[0077] For example, both the first-stage network and the second-stage network use an encoder and decoder structure to extract deep semantic features of the blurred video frames, while the third-stage network calculates the resolution of the blurred video frames (without performing any downsampling operation), thereby preserving more detail and texture in the final output target video frame.
[0078] In this embodiment, a three-stage cascaded network can be used to progressively deblurr each blurred video frame to obtain the target video frame for each blurred video frame. For example, as shown... Figure 6 As shown, the first two stages are based on the AD-Net encoder, dynamically adjusting the number of RCABs with shared weights to adaptively deblur video frames at different blur levels step by step. The last stage utilizes stacked RCAMs to operate at the resolution of the original blurred video frames, extracting channel statistics and correspondingly enhancing the information signal, thereby preserving the fine texture required in the final output target video frame X3. A SAM is added between every two stages to effectively suppress less important features, ensuring that information features are passed to subsequent deblurring stages. In addition, a cross-stage feature fusion mechanism is introduced, utilizing the intermediate multi-scale contextual features of the previous stage sub-networks to consolidate the intermediate features of the subsequent stage sub-networks; here, the RCAM network structure is as follows. Figure 7 As shown, since RCAM is also an existing module in MPRNet, it will not be discussed in detail here.
[0079] Step 103: Combine the target video frames of each blurred video frame with the unprocessed clear video frames in the video to be processed to obtain the target video.
[0080] In this embodiment of the application, after the target video frame of each blurred video frame, the target video frame of each blurred video frame can be combined with the unprocessed clear video frame in the video to be processed to obtain the target video; the target video is the video after the video to be processed has been deblurred.
[0081] For example, the combination method can be as follows: delete each blurred video frame in the video to be processed, and then insert the target video frame of each blurred video frame into the corresponding position, thereby obtaining the target video. Understandably, compared to the video to be processed containing blurred video frames, the target video after deblurring has better coherence and smoothness, effectively improving the visual quality of the video content.
[0082] As can be seen, in this embodiment, the number of RCABs with shared weights in each encoder of the network can be adaptively adjusted according to the different blur levels of each blurred video frame, so as to achieve adaptive processing of images with different blur levels. Compared with the existing deblurring methods for a single blur level, this embodiment can adaptively process blurred video frames with different blur levels, thus improving the adaptability to complex blurred scenes and ensuring the deblurring effect of the video.
[0083] In some embodiments, the deblurring network is trained by the following steps: obtaining a training dataset; preprocessing the training dataset to obtain a preprocessed target training dataset; and iteratively training the pre-built initial deblurring network using the target training dataset to obtain a trained deblurring network.
[0084] In this embodiment of the application, the training dataset may include multiple training image pairs, and the training image pairs include blurred images and clear images corresponding to the blurred images. Here, the data source of the training dataset is not specifically limited. For example, the training dataset can be obtained from existing deblurred datasets (such as the GoPro deblurred dataset), or from video segments of different scenes and different motion states captured by the shooting device, or from both of the above data sources.
[0085] For example, the process of obtaining a training dataset from video segments is described below. Obtaining a training dataset may include: acquiring multiple sample video segments; for each sample video segment, continuously extracting frames from the sample video segment according to different preset frame numbers to obtain multiple sets of continuous frames; each set of continuous frames includes T sample video frames, where T is an odd number greater than 1; averaging each set of continuous frames to obtain the blurred sample video frames corresponding to each set of continuous frames; for each set of continuous frames, determining the middle sample video frame located in each set of continuous frames and the blurred sample video frame corresponding to each set of continuous frames as a set of training image pairs; and constructing a training dataset based on each set of training image pairs corresponding to each sample video segment in the multiple sample video segments.
[0086] Here, the preset number of frames is an odd number greater than 1, and there is no specific limitation; for example, it can be 7, 9, 11 or 13, etc.
[0087] For example, assuming the preset number of frames is {7, 9, 11, 13, 15}, if for each sample video segment, the sample video is continuously sampled according to the above five different preset number of frames, then five sets of continuous frames can be obtained. The first set of continuous frames includes 5 sample video frames, the second set of continuous frames includes 7 sample video frames, and so on.
[0088] Furthermore, by averaging the five sets of consecutive frames, we can obtain the blurred sample video frames corresponding to each set of consecutive frames; and the blur level of the blurred sample video frames can be set to the number of sample video frames T included in each set of consecutive frames.
[0089] For example, the intermediate sample video frame in each group of consecutive frames can be set as the ground truth of the blurred image. That is, the intermediate sample video frame and the blurred sample video frame in each group of consecutive frames can be determined as a training image pair. After obtaining multiple training image pairs corresponding to multiple sample video segments, a training dataset can be constructed based on these training image pairs.
[0090] Understandably, in order to reduce the dependence of deblurring networks on specific datasets, the existing deblurring dataset can be expanded using the obtained video segments to obtain a training dataset with a wider range of blur levels, which can ensure that the trained deblurring network has strong generalization ability.
[0091] In this embodiment of the application, after obtaining the training dataset, the training dataset can be preprocessed to obtain the preprocessed target training dataset; for example, the preprocessing may include cropping the image size in the training dataset to a uniform size.
[0092] For example, after obtaining the target training dataset, it can be divided into a training set and a test set, and appropriate training parameters can be set, such as batch size and learning rate; for example, the batch size can be set to 64; the deblurring network can be trained using the AdamW optimizer, and the maximum learning rate can be set to 2×10. -4 The minimum learning rate is 10. -6 .
[0093] For example, a 2000-epoch cosine annealing learning rate schedule can be adopted, with the first 50 epochs used for warm-up. The warm-up strategy helps the network to begin training stably in the early stages, preventing large gradient updates. The cosine annealing learning rate schedule helps the network smoothly adjust the learning rate during training, thereby improving convergence speed and stability. For example, for fuzzy levels {7, 9, 11, 13, 15}, the corresponding number of RCABs with shared weights in the encoder are {1, 2, 3, 4, 5}, respectively. For ease of understanding, the following will combine... Figure 6 The training process will be explained.
[0094] For example, as described above, a blurred image is input into a pre-constructed initial deblurring network and is segmented into non-overlapping image blocks. The first stage network has 4 blocks, the second stage network has 2 blocks, and the final stage does not segment the image, retaining the original blurred image. Figure 6 As shown, at any given stage S, the network does not directly predict the restored sharp image X. S Instead, it predicts a residual image R. S Add the input blurred image I to the residual image R S To obtain a clear image X S , that is, X S =I+R S The initial deblurring network can be trained end-to-end using the following loss function L:
[0095]
[0096] Where Y is the sharp image corresponding to the blurred image I (i.e., representing the ground-truth image), and L char The Charbonnier loss is defined as:
[0097]
[0098] Where ε is a constant, empirically set to 10. -3 L edge The edge loss is defined as:
[0099]
[0100] Where Δ is the Laplace operator. The parameter λ in formula (1) controls the relative importance of the two loss terms and can be set to 0.05.
[0101] For example, after obtaining the corresponding loss function L from the training set, the initial deblurring network can be iteratively trained using the loss function L to obtain the trained deblurring network.
[0102] For example, to improve the performance of the deblurring network, the trained network can be further optimized using a test set. During testing, a wider range of blur levels can be introduced to assess the network's generalization ability. For instance, the training image pairs in the test set are obtained by averaging frames {7, 9, 11, 13, 15, 17, 19, 23, 27, 31}. Here, continuously sampling frames 19, 23, 27, and 31 aims to significantly increase the blur levels, demonstrating whether the AD-Net encoder can help the network generalize to blur levels completely unseen during training. Finally, evaluation metrics such as Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) can be used to quantitatively analyze and evaluate the deblurred video frames. Based on the evaluation results, it can be determined whether to stop training the network. If training is stopped, the trained deblurring network is obtained.
[0103] As can be seen in this embodiment, when obtaining the training dataset, the existing deblurring dataset is expanded to cover a wider range of fuzziness, thereby improving the generalization ability of the deblurring network and reducing the dependence on a specific dataset.
[0104] In order to better reflect the purpose of this application, based on the above embodiments of this application, and in combination with Figure 6 Further explanation is needed.
[0105] For example, such as Figure 6 As shown, the process of deblurring blurred video frames using a target deblurring network includes the following steps:
[0106] Step A1: Divide the blurred video frame I into four non-overlapping image blocks: top left, bottom left, top right, and bottom right.
[0107] Step A2: Convolve each image patch to expand its dimensions for extracting richer feature information later;
[0108] Step A3: Each image block after convolution is processed by RCAB to obtain the features of each image block in each dimension, and then fed into the AD-Net encoder;
[0109] Step A4: The AD-Net encoder adjusts the number of RCABs with shared weights based on the blur level of the blurred video frame I to extract deeper semantic features;
[0110] Step A5: The AD-Net encoder merges deep features, combining the features of four image blocks into two image blocks distributed on the left and right, and then sends them to the decoder to extract the merged features;
[0111] Step A6: Divide the blurred video frame I into two image blocks, left and right;
[0112] Step A7: Input the left and right image patches and the large-scale feature maps output by the decoder of the first-stage network into SAM. During training, SAM can use real images to provide useful control signals for the current stage of deblurring process.
[0113] Step A8: The output of SAM is divided into two parts. One part is the left and right image patches of the second input (corresponding to step A6), which will continue the process of the second-stage network. The other part of the output is used to train the network.
[0114] Step B1: Divide the blurred video frame I into two image blocks, left and right;
[0115] Step B2: After the second-stage network performs convolutional dimension expansion and RCAB processing, the features before and after the decoder in the first-stage network are fed into the AD-Net encoder of the second-stage network.
[0116] Step B3: After passing through a decoder similar to that of the first-stage network, the SAM of the second-stage network also produces two parts of output. One part continues the process of the third-stage network, and the other part is used to train the network.
[0117] Step C1: The input to the third-stage network is an unsegmented, blurry video frame, with the aim of recovering image details using complete contextual information;
[0118] Step C2: Convolve the blurred video frame to expand the dimensions. The convolved blurred video frame is then processed by RCAB to obtain the features of the blurred video frame in each dimension.
[0119] Step C3: Feed the features before and after the decoder in the second-stage network into the RCAM stacked in the third-stage network without performing any downsampling operation to generate spatially rich high-resolution features.
[0120] Step C4: Finally, a convolution is performed to reduce the feature dimension to 3, outputting the target video frame including detailed textures.
[0121] It should be noted that steps A1 to A8 above are the processing flow for the first stage network; steps B1 to B3 are the processing flow for the second stage network; and steps C1 to C4 are the processing flow for the third stage network.
[0122] As can be seen, this application proposes a deblurring network, MAD-Net, which can decompose blurred video frames into multiple stages, with each stage progressively restoring image details, thereby better handling complex blur problems. Furthermore, the AD-Net encoder included in this network can dynamically adjust the network architecture to adapt to images with different blur levels without requiring retraining for specific blur levels, achieving flexible image processing.
[0123] Figure 8 This is a schematic diagram of the composition structure of the video deblurring device according to an embodiment of this application, as shown below. Figure 8 As shown, the device includes: a determining module 300, an adjusting module 301, a deblurring module 302, and a combining module 303, wherein:
[0124] The determination module 300 is used to determine each blurred video frame in the video to be processed;
[0125] The adjustment module 301 is used to obtain the blur level of each blurred video frame, and adjust the number of first residual channel attention blocks included in each encoder of the trained deblurring network according to the blur level to obtain the target deblurring network; the first residual channel attention block is the residual channel attention block with shared weights in the encoder.
[0126] The deblurring module 302 is used to perform multi-stage deblurring processing on the blurred video frame using the target deblurring network to obtain the target video frame;
[0127] The combination module 303 is used to combine the target video frames of each blurred video frame with the unprocessed clear video frames in the video to be processed to obtain the target video.
[0128] In some embodiments, the adjustment module 301 is further configured to:
[0129] Based on the positive correlation between the blur level and the number of the first residual channel attention blocks, the number of the first residual channel attention blocks included in each encoder of the trained deblurring network is adjusted.
[0130] In some embodiments, the deblurring network includes a three-stage cascaded network, the first stage network and the second stage network having an encoder and a decoder structure, the encoder further including a second residual channel attention block with exclusive weights, the second residual channel attention block being connected in series with the first residual channel attention block.
[0131] In some embodiments, the deblurring module 302 is further configured to:
[0132] The blurred video frame is segmented into multiple non-overlapping first image blocks. The encoder of the first stage network is used to extract features from the first image blocks, and the decoder of the first stage network is used to restore the extracted features.
[0133] The blurred video frame is segmented into multiple non-overlapping second image blocks. The encoder of the second-stage network is used to extract features from the second image blocks, and the decoder of the second-stage network is used to restore the extracted features.
[0134] The blurred video frame is used as input to the third-stage network, and the resolution of the blurred video frame is calculated using the third-stage network to obtain the target video frame.
[0135] In some embodiments, the apparatus further includes a training module, the training module being configured to:
[0136] Obtain a training dataset; the training dataset includes multiple pairs of training images, each pair of training images including a blurred image and a sharp image corresponding to the blurred image;
[0137] The training dataset is preprocessed to obtain the preprocessed target training dataset;
[0138] The pre-built initial deblurring network is iteratively trained using the target training dataset to obtain the trained deblurring network.
[0139] In some embodiments, the training module is further configured to:
[0140] Acquire multiple sample video segments;
[0141] For each of the sample video segments, the sample video segment is continuously frame-sampling according to different preset frame numbers to obtain multiple sets of continuous frames; each set of continuous frames includes T sample video frames, where T is an odd number greater than 1;
[0142] The average of each group of consecutive frames is taken to obtain the blurred sample video frame corresponding to each group of consecutive frames;
[0143] For each group of consecutive frames, the intermediate sample video frame located in each group of consecutive frames and the blurred sample video frame corresponding to each group of consecutive frames are determined as a training image pair;
[0144] The training dataset is constructed based on the training image pairs corresponding to each of the multiple sample video segments.
[0145] In some embodiments, the determining module 300 is further configured to:
[0146] The video to be processed is input into the blur detector in the form of consecutive frames to obtain the category results of each video frame in the video to be processed;
[0147] Based on the category results of each video frame in the video to be processed, each blurred video frame in the video to be processed is determined.
[0148] In some embodiments, the fuzz detector includes a classifier, and in some embodiments, the determining module 300 is further configured to:
[0149] Feature extraction is performed on each video frame in the video to be processed to obtain the feature points of each video frame;
[0150] Cluster the feature points of all video frames in the video to be processed to obtain multiple cluster centers;
[0151] The feature points of each video frame are transformed using the multiple cluster centers to obtain the feature vector of each video frame;
[0152] The feature vectors of each video frame are input into the classifier to obtain the category results of each video frame.
[0153] In practical applications, the aforementioned determining module 300, adjusting module 301, defuzzifying module 302, combining module 303, and training module can all be implemented by a processor located in an electronic device. The processor can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.
[0154] Furthermore, in this embodiment, the functional modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.
[0155] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0156] Specifically, the computer program instructions corresponding to a video deblurring method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the computer program instructions corresponding to a video deblurring method in the storage media are read or executed by an electronic device, any of the video deblurring methods in the aforementioned embodiments are implemented.
[0157] Based on the same technical concept as the foregoing embodiments, see Figure 9 It illustrates an electronic device 400 provided in an embodiment of this application, which may include: a memory 401 and a processor 402; wherein,
[0158] Memory 401 is used to store computer programs and data;
[0159] Processor 402 is configured to execute a computer program stored in memory to implement any of the video deblurring methods described in the foregoing embodiments.
[0160] In practical applications, the memory 401 mentioned above can be volatile memory, such as RAM; or non-volatile memory, such as ROM, flash memory, hard disk drive (HDD) or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 402.
[0161] The processor 402 described above can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor. It is understood that for different video deblurring devices, the electronic device used to implement the above processor function can also be other types, and this application embodiment does not specifically limit the specific implementation.
[0162] In some embodiments, the functions or modules of the apparatus provided in this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0163] This application also provides a computer storage medium storing a computer program; when the computer program is executed, it can implement any of the video deblurring methods described in the foregoing embodiments.
[0164] This application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the video deblurring methods described in the foregoing embodiments.
[0165] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0166] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict.
[0167] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0168] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0169] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0170] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0171] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0172] The above are merely preferred embodiments of this application and are not intended to limit the scope of protection of this application.
Claims
1. A method of video deblurring, characterized by, The method includes: Identify the individual blurred video frames in the video to be processed; For each blurred video frame, the blur level of the blurred video frame is obtained. Based on the blur level, the number of first residual channel attention blocks included in each encoder of the trained deblurring network is adjusted to obtain the target deblurring network. The first residual channel attention block is the residual channel attention block with shared weights in the encoder. The target deblurring network is used to perform multi-stage deblurring processing on the blurred video frame to obtain the target video frame; The target video is obtained by combining the target video frames of each blurred video frame with the unprocessed clear video frames in the video to be processed.
2. The method of claim 1, wherein, The step of adjusting the number of first residual channel attention blocks in each encoder of the trained deblurring network according to the blur level includes: Based on the positive correlation between the blur level and the number of the first residual channel attention blocks, the number of the first residual channel attention blocks included in each encoder of the trained deblurring network is adjusted.
3. The method of claim 1, wherein, The deblurring network includes a three-stage cascaded network. The first-stage network and the second-stage network have encoder and decoder structures. The encoder also includes a second residual channel attention block with unique weights. The second residual channel attention block is connected in series with the first residual channel attention block.
4. The method of claim 3, wherein, The multi-stage deblurring process of the blurred video frame using the target deblurring network includes: The blurred video frame is segmented into multiple non-overlapping first image blocks. The encoder of the first stage network is used to extract features from the first image blocks, and the decoder of the first stage network is used to restore the extracted features. The blurred video frame is segmented into multiple non-overlapping second image blocks. The encoder of the second-stage network is used to extract features from the second image blocks, and the decoder of the second-stage network is used to restore the extracted features. The blurred video frame is used as input to the third-stage network, and the resolution of the blurred video frame is calculated using the third-stage network to obtain the target video frame.
5. The method of claim 1, wherein, The deblurring network is trained through the following steps: Obtain a training dataset; the training dataset includes multiple pairs of training images, each pair of training images including a blurred image and a sharp image corresponding to the blurred image; The training dataset is preprocessed to obtain the preprocessed target training dataset; The pre-built initial deblurring network is iteratively trained using the target training dataset to obtain the trained deblurring network.
6. The method according to claim 5, characterized in that, The acquisition of the training dataset includes: Acquire multiple sample video segments; For each of the sample video segments, the sample video segment is continuously frame-sampling according to different preset frame numbers to obtain multiple sets of continuous frames; each set of continuous frames includes T sample video frames, where T is an odd number greater than 1; The average of each group of consecutive frames is taken to obtain the blurred sample video frame corresponding to each group of consecutive frames; For each group of consecutive frames, the intermediate sample video frame located in each group of consecutive frames and the blurred sample video frame corresponding to each group of consecutive frames are determined as a training image pair; The training dataset is constructed based on the training image pairs corresponding to each of the multiple sample video segments.
7. The method according to claim 1, characterized in that, The process of determining each blurred video frame in the video to be processed includes: The video to be processed is input into the blur detector in the form of consecutive frames to obtain the category results of each video frame in the video to be processed; Based on the category results of each video frame in the video to be processed, each blurred video frame in the video to be processed is determined.
8. The method according to claim 7, characterized in that, The fuzz detector includes a classifier. The step of inputting the video to be processed into the fuzz detector in the form of consecutive frames to obtain the category results of each video frame in the video to be processed includes: Feature extraction is performed on each video frame in the video to be processed to obtain the feature points of each video frame; Cluster the feature points of all video frames in the video to be processed to obtain multiple cluster centers; The feature points of each video frame are transformed using the multiple cluster centers to obtain the feature vector of each video frame; The feature vectors of each video frame are input into the classifier to obtain the category results of each video frame.
9. A video deblurring device, characterized in that, The device includes: The determination module is used to identify each blurred video frame in the video to be processed; An adjustment module is used to obtain the blur level of each blurred video frame, and adjust the number of first residual channel attention blocks included in each encoder of the trained deblurring network according to the blur level to obtain the target deblurring network; the first residual channel attention block is the residual channel attention block with shared weights in the encoder. The deblurring module is used to perform multi-stage deblurring processing on the blurred video frame using the target deblurring network to obtain the target video frame; The combination module is used to combine the target video frames of each blurred video frame with the unprocessed clear video frames in the video to be processed to obtain the target video.
10. An electronic device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method according to any one of claims 1 to 8.
11. A computer storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method described in any one of claims 1 to 8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 8.