Video deblurring method and device based on alignment and with occlusion area correction
By adopting the method of aligning and occlusion area correction in the video defuzzing technology, the shortcomings and occlusion effects of timing information processing in the prior art are solved, and a more efficient video defuzzing effect is achieved.
Patent Information
- Application Number
- CN202210424323.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-04-21
AI Technical Summary
The existing video defuzzing technology has forgotten problems and accumulated timing deviations when processing timing information, and the optical flow based on fuzzy frame estimation is inaccurate, which affects the defuzzing effect.
The video defuzzing method based on alignment and with occlusion area correction is adopted, and the alignment filter is used to align adjacent frames in the feature domain, and the defuzzing filter is calculated using correction parameters in the reconstruction and fusion module to achieve more accurate inter-frame information accumulation and defuzzing processing.
It effectively solves the impact of occlusion on video debuffing, improves the accuracy of inter-frame alignment and debuffing effect, and restores clearer video frames.
Smart Images

Figure CN114862703B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video deblurring, and in particular relates to a video deblurring method and device based on alignment and with occlusion area correction. Background Art
[0002] With the emergence of various handheld video shooting devices, people are increasingly fond of shooting various videos for recording. However, when shooting videos, the videos are often blurred due to camera shake and the rapid movement of moving objects in the scene. Video deblurring has therefore attracted extensive research. The main task of video deblurring is to restore clear frames based on existing blurred image frames. Depending on whether the blur kernel is known, this task can be divided into blind deblurring and non-blind deblurring. However, in real scenes, the blur kernel is usually unknown, and there are many related research technologies for restoring clear video frames when the blur kernel is unknown. At this stage, deblurring technology can be roughly divided into two categories. One method is to implement video deblurring based on traditional optimization models. The other method is to perform video deblurring based on deep convolutional networks.
[0003] The traditional optimization model-based approach usually requires the design of many different constraints to optimize the problem model. For image space constraints, sparse prior constraints, low-rank prior constraints, gradient prior constraints and other constraints are usually designed. The time series terms usually use optical flow to describe the motion relationship between multiple frames. Although the traditional optimization model-based approach can restore relatively clear video frames, the model solution process is very complicated, and the iterative solution process usually takes a lot of time.
[0004] The method based on deep convolutional network usually adopts an end-to-end network design method, inputting blurred continuous frames and outputting clear frames. For the video deblurring task, making full use of the temporal and spatial relationship between frames is the key to designing the network. In order to utilize the temporal relationship between frames, most methods choose to directly connect multiple frames in the temporal dimension at the beginning, and then input them into the network for end-to-end network training. However, simple splicing of the temporal dimension between frames does not fully explore the temporal relationship between multiple frames. How to fully and effectively utilize the temporal information between frames is a difficulty in the video deblurring task. For temporal information, one type of solution technology is to use recurrent neural network to transfer the information of the previous frame in the video sequence to the target frame to achieve the effect of temporal information accumulation. Another type mainly learns the motion optical flow between adjacent frames, and then uses the learned optical flow to align adjacent frames before deblurring. When using recurrent neural network to accumulate temporal information, there will be a problem of forgetting the temporal information of distant frames, and the non-alignment of multiple frames will also lead to an increasing deviation in the accumulation of temporal information. There is a difference between the accuracy of the optical flow learned based on blurred frames and the motion optical flow in real situations. Inaccurate optical flow alignment will also affect the subsequent deblurring effect.
[0005] In order to solve the problems of forgetting time series information and accumulating time series deviations in recurrent neural networks, as well as inaccurate estimation of optical flow based on blurred frames. STFAN proposes to learn a spatiotemporal adaptive alignment network, directly use the network to learn alignment filters, align adjacent frames in the feature domain, and then further learn filters for deblurring based on the alignment filters to achieve deblurring. This method does not require explicit estimation of optical flow, and the accumulation of information between frames after alignment is more accurate. However, the learned alignment filter must be able to achieve a good alignment effect. If it cannot be well aligned, it will have a great impact on the restoration effect of the next step. In real scenes, due to factors such as the rapid motion of moving objects and camera jitter, occlusion is very common in adjacent frames. The alignment filter learned by STFAN does not consider the impact of occlusion, and cannot achieve an ideal alignment effect in scenes with occlusion. Summary of the invention
[0006] In view of the above, the purpose of the present invention is to provide a video deblurring method and device based on alignment and with occlusion area correction, which can perform occlusion correction when mining the temporal relationship alignment to solve the occlusion problem encountered during alignment and achieve better video deblurring effect.
[0007] To achieve the above-mentioned object of the invention, an embodiment provides a video deblurring method based on alignment and with occlusion area correction, comprising the following steps:
[0008] Obtain a target frame, a previous frame of the target frame, and a deblurred image of the previous frame in the video frame; wherein the target frame and the previous frame are both blurred images;
[0009] The video deblurring model is used to calculate the input target frame, the previous frame and the deblurred image of the previous frame to obtain the deblurred result image of the target frame;
[0010] Among them, the video deblurring model includes a feature extraction module, an alignment and occlusion correction module and a reconstruction fusion module. The input target frame, the previous frame and the deblurred image of the previous frame are feature extracted by the feature extraction module to obtain three feature maps, and the combined feature map of the three feature maps is input into the alignment and occlusion correction module; in the alignment and occlusion correction module, the combined feature map is preliminarily aligned to obtain a first alignment filter; the three feature maps are subjected to occlusion pixel map calculation, correlation calculation and correction parameter calculation to obtain correction parameters, and the correction parameters are combined with the first alignment filter to construct a second alignment filter with a correction function, and the feature map of the previous frame is filtered by the second alignment filter to obtain an aligned feature map, and the second alignment filter and the aligned feature map are input into the reconstruction fusion module; in the reconstruction fusion module, a deblurring filter is calculated based on the second alignment filter, and the feature map of the target frame is deblurred by the deblurring filter, and the deblurring result is fused with the target frame to obtain a deblurred result map of the target frame.
[0011] In one embodiment, the alignment and occlusion correction module includes a preliminary alignment unit, an occlusion calculation unit, a correlation calculation unit, a correction parameter calculation unit, a filter calculation unit, and a final alignment unit;
[0012] The preliminary alignment unit performs preliminary alignment calculation on the input combined feature map, and uses the output three-dimensional feature matrix as the first alignment filter;
[0013] The occlusion calculation unit performs bidirectional optical flow calculation based on the target frame and the previous frame to obtain occluded pixels, and then generates an occluded pixel map;
[0014] The correlation calculation unit calculates the inter-frame correlation according to the feature graphs of the target frame and the previous frame to obtain a correlation coefficient matrix;
[0015] The correction parameter calculation unit calculates correction parameters of two branches according to the masked pixel map, the correlation coefficient matrix and the connected feature map of the feature map of the previous frame deblurred image to obtain two correction parameter matrices;
[0016] The filter calculation unit uses two correction parameter matrices as weights and biases to perform weighted fusion on the first alignment filter to obtain a second alignment filter with a correction function;
[0017] The final alignment unit uses the second alignment filter to filter the feature map of the previous frame to obtain an aligned feature map.
[0018] In one embodiment, in the preliminary alignment unit, a convolutional neural network is used to perform preliminary alignment calculation on the combined feature map;
[0019] In the correction parameter calculation unit, both branches use at least two convolutional layers to perform convolution operations on the connection feature map to calculate the correction parameters, and the convolutional layer parameters of the two branches are different.
[0020] In one embodiment, the alignment and occlusion correction module further includes an information enhancement unit, which performs feature pixel enhancement calculation based on the occlusion pixel map to obtain an information enhanced occlusion pixel map, and the information enhanced occlusion pixel map is used to calculate the correction parameters.
[0021] In one embodiment, the information enhancement unit includes three convolution components. The masked pixel map is calculated by the three convolution components to obtain three masked feature maps. Any two masked feature maps are dot-multiplied and activated to obtain an attention matrix. The attention matrix is dot-multiplied with the remaining masked feature matrix and then superimposed with the originally input masked pixel map to obtain an information-enhanced masked pixel map, wherein each convolution component includes at least one convolution layer, and the convolution layer parameters of each convolution component are different.
[0022] In one embodiment, the reconstruction fusion module includes a deblurring filter generating unit, a deblurring unit, and a fusion recovery unit;
[0023] The deblurring filter generating unit calculates a deblurring filter by convolving the second alignment filter using a convolution component; wherein the convolution component includes at least one convolution layer;
[0024] The deblurring unit utilizes a deblurring filter to multiply the feature map of the target frame to achieve deblurring processing to obtain a deblurring processing result;
[0025] The fusion recovery unit uses a convolution component to perform a convolution operation on the connection result of the deblurring processing result and the aligned feature map in the time dimension to achieve fusion recovery. The fusion recovery result is superimposed on the target frame to obtain the deblurring result map of the target frame.
[0026] In one embodiment, before the video deblurring model is applied, model parameter optimization is performed. When the model parameters are optimized, the loss function Loss used is:
[0027]
[0028] Among them, B represents the blurred target frame image, GT represents the clear image of the target frame object, C, H, W represent the number of channels, length and width of the image respectively, Ψ i () represents the feature map extracted by the pre-trained feature extraction network, C i ,H i ,W i , represents the feature map Ψ i () represents the number of channels, length and width, and γ represents the weights that control the two types of loss functions.
[0029] In one embodiment, when optimizing the model parameters of a video deblurring model, first, the first stage training is performed on the feature extraction module, the reconstruction fusion module, and the other parts of the alignment and occlusion correction module in the video deblurring model that do not include the correction parameter calculation part. During the training process, the second alignment filter is replaced with the first alignment filter, that is, the feature map of the previous frame is filtered based on the first alignment filter to obtain the aligned feature map, and the deblurring filter is calculated based on the first alignment filter, and the learning rate is reduced according to a fixed periodic interval;
[0030] After the first stage of training, the alignment and occlusion correction module is set to obtain the correction parameter part to participate in the training, that is, the entire video deblurring model participates in the second stage of training. During the second stage of training, the structure of the first stage of training still maintains the learning rate at the end of the first stage of training unchanged, and the learning rate of the correction parameter part is reduced at a fixed interval, and the initial learning rate is the same as the initial learning rate of the first stage of training.
[0031] To achieve the above-mentioned purpose of the invention, an embodiment further provides a video deblurring device based on alignment and with occlusion area correction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the memory stores a video deblurring model constructed by the above-mentioned video deblurring method, and the processor implements the following steps when executing the computer program:
[0032] Obtain a target frame, a previous frame of the target frame, and a deblurred image of the previous frame in the video frame; wherein the target frame and the previous frame are both blurred images;
[0033] The video deblurring model is used to calculate the input target frame, the previous frame and the deblurred image of the previous frame to obtain the deblurred result image of the target frame.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] An alignment method is used to mine the temporal relationship between frames. During the alignment process, the occlusion problem between frames is considered. An aligned occlusion correction method is proposed and utilized to better align adjacent frames, accumulate and combine the temporal relationship between frames, and restore clearer video frames. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0037] Figure 1 is a flow chart of a video deblurring method based on alignment and with occlusion area correction provided by an embodiment;
[0038] Figure 2 The structure of the video deblurring model provided in the embodiment and the flowchart of using the same to perform video deblurring;
[0039] Figure 3 is a structure and calculation flow chart of an alignment and occlusion correction module provided in an embodiment;
[0040] Figure 4 is a structure and calculation flow chart of the information enhancement unit provided in the embodiment;
[0041] Figure 5 This is an example diagram of video deblurring results after occlusion correction provided by an embodiment. DETAILED DESCRIPTION
[0042] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0043] To address the problem of deblurring multi-frame video images, especially the occlusion problem between frames, an embodiment provides a video deblurring method and device based on alignment and with occlusion area correction. Common occlusion scenarios are taken into account during alignment, achieving better image restoration effects compared to other video deblurring technologies.
[0044] Figure 1 1 is a flow chart of a video deblurring method based on alignment and with occlusion area correction provided by an embodiment. Figure 1 As shown, the video deblurring method based on alignment and with occlusion area correction provided by the embodiment includes the following steps:
[0045] Step 1, obtaining a target frame, a previous frame of the target frame, and a deblurred image of the previous frame in a video frame; wherein both the target frame and the previous frame are blurred images.
[0046] In an embodiment, video interval sampling can obtain a video frame sequence, wherein the video frame that needs to be deblurred at the current moment is the target frame, and if the target frame is the first frame in the video frame sequence, the target frame is copied as its previous frame and the deblurred image of the previous frame.
[0047] Step 2: Use the video deblurring model to calculate the input target frame, the previous frame and the deblurred image of the previous frame to obtain the deblurred result image of the target frame.
[0048] In the embodiment, Figure 2 As shown, the video deblurring model includes a feature extraction module, an alignment and occlusion correction module, and a reconstruction fusion module. The deblurring process using the video deblurring model includes:
[0049] The input target frame, the previous frame and the deblurred image of the previous frame are subjected to feature extraction by the feature extraction module to obtain three feature maps, and the combined feature map of the three feature maps is input into the alignment and occlusion correction module; in the alignment and occlusion correction module, the combined feature map is preliminarily aligned to obtain a first alignment filter; the three feature maps are subjected to occlusion pixel map calculation, correlation calculation and correction parameter calculation to obtain correction parameters, and the correction parameters are combined with the first alignment filter to construct a second alignment filter with a correction function, and the feature map of the previous frame is filtered by the second alignment filter to obtain an aligned feature map, and the second alignment filter and the aligned feature map are input into the reconstruction fusion module; in the reconstruction fusion module, a deblurring filter is calculated based on the second alignment filter, and the feature map of the target frame is deblurred by the deblurring filter, and the deblurring result is fused with the target frame to obtain a deblurred result map of the target frame.
[0050] In the embodiment, the feature extraction module uses CNN to extract the input target frame B t , previous frame B t-1 and the deblurred image S of the previous frame t-1 Perform feature extraction to output the target frame B t The feature map E t , previous frame B t-1 The feature map E t-1 And the deblurred graph S t-1 The characteristic map E_S t-1 , feature map E t 、E t-1 and S t-1 Linked together to form a combined feature map E all .
[0051] Figure 3 1 is a structure and calculation flow chart of the alignment and occlusion correction module provided in the embodiment. Figure 3 As shown, the alignment and occlusion correction module includes a preliminary alignment unit, an occlusion calculation unit, a correlation calculation unit, a correction parameter calculation unit, a filter calculation unit, and a final alignment unit.
[0052] Among them, the preliminary alignment unit inputs the combined feature map E all A preliminary alignment calculation is performed, and the output three-dimensional feature matrix is used as the first alignment filter Filter1, which does not include occlusion correction. In the embodiment, the preliminary alignment unit uses a convolutional neural network based on the combined feature map E all Perform preliminary alignment calculation, that is, the combined feature map E all Successive convolution operations are performed to obtain the first aligned filter Filter1.
[0053] The occlusion calculation unit calculates the target frame B t and the previous frame B t-1 Perform bidirectional optical flow calculation. If the current pixel appears in the forward optical flow but not in the reverse optical flow, it can be preliminarily determined that the current pixel is blocked, and the blocked pixel is obtained accordingly, thereby generating the blocked pixel map M.
[0054] The correlation calculation unit calculates the feature map E of the target frame t and the feature map E of the previous frame t-1 The correlation between frames is calculated to obtain the correlation coefficient matrix R.
[0055] The correction parameter calculation unit calculates the feature map E_S of the masked pixel map M, the correlation feature map R, and the deblurred image. t-1 The correction parameters of the two branches are learned and calculated based on the connection feature graph to obtain two correction parameter matrices α and β. In the embodiment, both branches use at least two convolutional layers to perform convolution operations on the connection feature graph, and the result matrix obtained by the convolution operation is used as the correction parameter matrix. It should be noted that the convolution layer parameters of the two branches are different.
[0056] The filter calculation unit uses two correction parameter matrices α and β as weights and biases to perform weighted fusion on the first alignment filter to obtain a second alignment filter Filter2 with a correction function; specifically, it can be expressed as Filter2=Filter1×α+β.
[0057] The final alignment unit uses the second alignment filter Filter2 to align the feature map E of the previous frame. t-1 Perform filtering to obtain the aligned feature map Specifically, the three-dimensional second alignment filter Filter2 is geometrically transformed, and the obtained two-dimensional transformation matrix is consistent with the feature map E t-1 Multiply to achieve the filtering process.
[0058] In the embodiment, in order to improve the quality of the occluded pixel map M, a self-attention mechanism is used to enhance the information of the occluded pixel map. Specifically, the alignment and occlusion correction module also includes an information enhancement unit, which performs feature pixel enhancement calculation based on the occluded pixel map to obtain an information-enhanced occluded pixel map, and the information-enhanced occluded pixel map is used to calculate the correction parameters.
[0059] Figure 4 : is a structure and calculation flow chart of the information enhancement unit provided in the embodiment. Figure 4 As shown in the figure, the information enhancement unit includes three convolution components. The masked pixel map is calculated by the three convolution components to obtain three masked feature maps f, g, and h. The dot product of any two masked feature matrices (such as f and g) (fg T ) is activated to obtain an attention matrix, which is dot-multiplied with the remaining masked feature matrix h and then superimposed with the input masked pixel map M to obtain an information-enhanced masked pixel map, wherein each convolution component contains at least one convolution layer, preferably one convolution layer with a convolution kernel of 1*1, and the convolution layer parameters of each convolution component are different.
[0060] In the embodiment, the reconstruction fusion module includes a deblurring filter generation unit, a deblurring unit, and a fusion recovery unit; wherein the deblurring filter generation unit uses a convolution component to perform convolution calculation on the second alignment filter Filter2, and uses the output feature matrix as the deblurring filter E deblur ; Wherein, the convolution component includes at least one convolution layer;
[0061] The deblurring unit uses a deblurring filter E deblur and the feature map E of the target frame t , multiplied to achieve deblurring and get the deblurring result
[0062] The fusion recovery unit uses the convolution component to deblur the result. After alignment, the feature map The convolution operation is performed on the result of the connection in the time series dimension to achieve fusion recovery, and the fusion recovery result is superimposed with the target frame to obtain the deblurred result image R of the target frame. t .
[0063] In the embodiment, before the video deblurring model is applied, the model parameters are optimized. When the model parameters are optimized, in order to enable the network to perform better, the sum of the MSE loss function and the perceptual loss function is mainly used as the total loss function when training the model. Specifically, the total loss function Loss used is:
[0064]
[0065] Among them, B represents the blurred target frame image, GT represents the clear image of the target frame object, C, H, W represent the number of channels, length and width of the image respectively, Ψ i () represents the feature map extracted by the pre-trained feature extraction network. Here, Ψ i () represents the feature map extracted from the 15th convolutional layer of the pre-trained VGG-19 network, C i ,H i ,W i , represents the feature map Ψ i () represents the number of channels, length and width, γ represents the weight controlling the two types of loss functions, and it is set to 0.01 in the experiment.
[0066] When training the model, in order to enhance the robustness of the model, the training data is also enhanced to a certain extent. The enhancement methods include color transformation, such as image brightness, contrast, saturation adjustment, etc. It also includes horizontal flipping and vertical flipping of the training data. In order to be closer to the real scene, Gaussian noise is also randomly added to the input image during training. The Adam optimizer is used during training. At the beginning of training, the parameters in the model are randomly initialized. In order to give full play to the role of each module in the model, a phased training method is adopted during training, including:
[0067] First, the first stage training is performed on the feature extraction module, reconstruction fusion module, and other parts of the alignment and occlusion correction module in the video deblurring model that do not contain the correction parameter acquisition part. During the training process, the second alignment filter is replaced by the first alignment filter, that is, the feature map of the previous frame is filtered based on the first alignment filter to obtain the aligned feature map, and the deblurring filter is calculated based on the first alignment filter. The learning rate is reduced at a fixed periodic interval. For example, the initial learning rate is 10 -4 The learning rate is reduced tenfold every 150 cycles. A total of 300 cycles are trained in the first stage. The learning rate of this part of the network is fixed at 10 -6 constant.
[0068] After the first stage of training, the correction parameter part obtained in the alignment and occlusion correction module is set to participate in the training, that is, the entire video deblurring model participates in the second stage of training. During the second stage of training, the structure of the first stage of training still maintains the learning rate adopted at the end of the first stage of training. The learning rate of the correction parameter part is reduced at a fixed interval, and the initial learning rate is the same as the initial learning rate of the first stage of training. For example, the initial learning rate of the correction parameter part is set to 10 -4 , the learning rate is reduced by ten times after every 100 cycles, and the learning rate of the first stage training part is 10 -6 constant. Figure 5 is an example diagram of video deblurring results after occlusion correction provided by the embodiment. Figure 5 As shown in the figure, for the input blurred frame, if deblurring is performed directly based on the deblurring method without occlusion correction, the result is shown in Figure (b). In comparison, the deblurring result after occlusion correction is shown in Figure (c). The deblurring method with alignment and occlusion area correction can achieve better image restoration results.
[0069] The embodiment further provides a video deblurring device based on alignment and with occlusion area correction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the memory stores a video deblurring model constructed by the above-mentioned video deblurring method, and the processor implements the following steps when executing the computer program:
[0070] Step 1, obtaining a target frame, a previous frame of the target frame, and a deblurred image of the previous frame in a video frame; wherein the target frame and the previous frame are both blurred images;
[0071] Step 2: Use the video deblurring model to calculate the input target frame, the previous frame and the deblurred image of the previous frame to obtain the deblurred result image of the target frame.
[0072] In practical applications, the computer memory can be a volatile memory at the near end, such as RAM, or a non-volatile memory, such as ROM, FLASH, floppy disk, mechanical hard disk, etc., or a remote storage cloud. The computer processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), that is, these processors can be used to implement the video deblurring step based on alignment and with occlusion area correction.
[0073] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A video deblurring method based on alignment and with occlusion area correction, characterized in that: The following steps are involved: Obtain a target frame, a previous frame of the target frame, and a deblurred image of the previous frame in the video frame; wherein the target frame and the previous frame are both blurred images; The video deblurring model is used to calculate the input target frame, the previous frame and the deblurred image of the previous frame to obtain the deblurred result image of the target frame; Among them, the video deblurring model includes a feature extraction module, an alignment and occlusion correction module and a reconstruction fusion module. The input target frame, the previous frame and the deblurred image of the previous frame are feature extracted by the feature extraction module to obtain three feature maps, and the combined feature map of the three feature maps is input to the alignment and occlusion correction module; in the alignment and occlusion correction module, the combined feature map is preliminarily aligned to obtain a first alignment filter; the three feature maps are subjected to occlusion pixel map calculation, correlation calculation and correction parameter calculation to obtain correction parameters, and the correction parameters are combined with the first alignment filter to construct a second alignment filter with a correction function, and the feature map of the previous frame is filtered by the second alignment filter to obtain an aligned feature map, and the second alignment filter and the aligned feature map are input to the reconstruction fusion module; in the reconstruction fusion module, a deblurring filter is calculated based on the second alignment filter, and the feature map of the target frame is deblurred by the deblurring filter, and the deblurring result is fused with the target frame to obtain a deblurred result map of the target frame; The alignment and occlusion correction module comprises a preliminary alignment unit, an occlusion calculation unit, a correlation calculation unit, a correction parameter calculation unit, a filter calculation unit, and a final alignment unit; The preliminary alignment unit performs preliminary alignment calculation on the input combined feature map, and uses the output three-dimensional feature matrix as the first alignment filter; The occlusion calculation unit performs bidirectional optical flow calculation based on the target frame and the previous frame to obtain occluded pixels, and then generates an occluded pixel map; The correlation calculation unit calculates the inter-frame correlation according to the feature graphs of the target frame and the previous frame to obtain a correlation coefficient matrix; The correction parameter calculation unit calculates correction parameters of two branches according to the masked pixel map, the correlation coefficient matrix and the connected feature map of the feature map of the previous frame deblurred image to obtain two correction parameter matrices; The filter calculation unit uses two correction parameter matrices as weights and biases to perform weighted fusion on the first alignment filter to obtain a second alignment filter with a correction function; The final alignment unit uses the second alignment filter to filter the feature map of the previous frame to obtain an aligned feature map.
2. The video deblurring method based on alignment and with occlusion area correction according to claim 1, characterized in that: In the preliminary alignment unit, a convolutional neural network is used to perform preliminary alignment calculations on the combined feature maps; In the correction parameter calculation unit, both branches use at least two convolutional layers to perform convolution operations on the connection feature map to calculate the correction parameters, and the convolutional layer parameters of the two branches are different.
3. The video deblurring method based on alignment and with occlusion area correction according to claim 1, characterized in that: The alignment and occlusion correction module also includes an information enhancement unit, which performs feature pixel enhancement calculation based on the occlusion pixel map to obtain an information-enhanced occlusion pixel map, and the information-enhanced occlusion pixel map is used to calculate the correction parameters.
4. The video deblurring method based on alignment and with occlusion area correction according to claim 3, characterized in that: The information enhancement unit includes three convolution components. The masked pixel map is calculated by the three convolution components to obtain three masked feature maps. Any two masked feature maps are dot-multiplied and activated to obtain an attention matrix. The attention matrix is dot-multiplied with the remaining masked feature matrix and then superimposed with the originally input masked pixel map to obtain an information-enhanced masked pixel map, wherein each convolution component includes at least one convolution layer, and the convolution layer parameters of each convolution component are different.
5. The video deblurring method based on alignment and with occlusion area correction according to claim 1, characterized in that: The reconstruction fusion module includes a deblurring filter generating unit, a deblurring unit, and a fusion recovery unit; The deblurring filter generating unit calculates a deblurring filter by convolving the second alignment filter using a convolution component; wherein the convolution component includes at least one convolution layer; The deblurring unit utilizes a deblurring filter to multiply the feature map of the target frame to achieve deblurring processing to obtain a deblurring processing result; The fusion recovery unit uses a convolution component to perform a convolution operation on the connection result of the deblurring processing result and the aligned feature map in the time dimension to achieve fusion recovery. The fusion recovery result is superimposed on the target frame to obtain the deblurring result map of the target frame.
6. The video deblurring method based on alignment and with occlusion area correction according to claim 1, characterized in that: Before the video deblurring model is applied, the model parameters are optimized. When the model parameters are optimized, the loss function Loss used is: Among them, B represents the blurred target frame image, GT represents the clear image of the target frame object, C, H, W represent the number of channels, length and width of the image respectively, Ψ i () represents the feature map extracted by the pre-trained feature extraction network, C i ,H i ,W i , represents the feature map Ψ i () represents the number of channels, length and width, and γ represents the weights that control the two types of loss functions.
7. The video deblurring method based on alignment and with occlusion area correction according to claim 1 or 6, characterized in that: When optimizing the model parameters of the video deblurring model, firstly, the first stage training is performed on the feature extraction module, the reconstruction fusion module, and the other parts of the alignment and occlusion correction module in the video deblurring model, which do not include the correction parameter calculation part. During the training process, the second alignment filter is replaced by the first alignment filter, that is, the feature map of the previous frame is filtered based on the first alignment filter to obtain the aligned feature map, and the deblurring filter is calculated based on the first alignment filter, and the learning rate is reduced according to a fixed periodic interval; After the first stage of training, the alignment and occlusion correction module is set to obtain the correction parameter part to participate in the training, that is, the entire video deblurring model participates in the second stage of training. During the second stage of training, the structure of the first stage of training still maintains the learning rate at the end of the first stage of training unchanged, and the learning rate of the correction parameter part is reduced at a fixed interval, and the initial learning rate is the same as the initial learning rate of the first stage of training.
8. A video deblurring device based on alignment and with occlusion area correction, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The memory stores a video deblurring model constructed by the video deblurring method according to any one of claims 1 to 7, and the processor implements the following steps when executing the computer program: Obtain a target frame, a previous frame of the target frame, and a deblurred image of the previous frame in the video frame; wherein the target frame and the previous frame are both blurred images; The video deblurring model is used to calculate the input target frame, the previous frame and the deblurred image of the previous frame to obtain the deblurred result image of the target frame.
Citation Information
Patent Citations
Deblurred video recovery method and device, terminal equipment and storage medium
CN111932480A
Video enhancement method and device, computer equipment and storage medium
CN113781312A