A deep image deblurring method based on spatiotemporal frequency perception
By introducing a spatiotemporal frequency perception method into video defuzzing technology, integrating frequency-time information, the problem of difficulty in effectively using time information in the prior art is solved, and a better image defuzzing effect is achieved.
Patent Information
- Application Number
- CN202311190096.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-09-14
AI Technical Summary
Existing video defuzzing techniques are difficult to effectively utilize time information, especially in fuzzy sequences, resulting in unsatisfactory defuzzing effects.
The depth image defuzzing method based on spatiotemporal frequency perception is adopted to integrate frequency-time information into the video defuzzing framework, and the image defuzzing is achieved through the space-frequency feature extraction module, the spectral prior guided alignment module and the timing energy attention module.
By effectively modeling time prior information, blurred images can be better restored and the debuffering effect can be improved, especially in unpredictable blur situations.
Smart Images

Figure CN116993623B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image deblurring, and in particular to a deep image deblurring method based on spatiotemporal frequency perception. Background Art
[0002] Video deblurring, as a fundamental vision task, aims to recover clear frames from blurry sequences by leveraging intrinsic temporal information. Therefore, many studies have been devoted to exploring the potential of temporal information hidden in blurry sequences, which can be divided into two categories: traditional optimization methods and deep learning-based methods.
[0003] Traditional optimization methods usually emphasize the assumptions about the blur degradation process and apply some hand-crafted temporal priors to alleviate video deblurring, such as temporal sharpness prior, motion blur prior, and temporal coherence prior. However, these priors are difficult to design and the corresponding methods are also difficult to optimize, which limits their practicality.
[0004] In recent years, deep learning-based video deblurring methods have made a lot of progress in addressing the above challenges. Most of them follow a common process: feature extraction, alignment, fusion, and optimization. For example, the pioneering EDVR implicitly aligns through redesigned deformable convolutions, and some works use progressive refinement schemes to perform motion compensation for more accurate temporal modeling. However, the above strategies only study temporal information from the perspective of the spatial domain and do not fully explore its potential.
[0005] The present invention provides a new solution to effectively model the temporal prior information and ultimately achieve image deblurring. Summary of the invention
[0006] To solve the above technical problems, the present invention first re-examines the differences between blurry-sharp pairs in the spatial domain and the frequency domain respectively, and then expands the energy information of blurry-clear image pairs in the frequency spectrum along the time dimension; through the effective exploration of this prior information, the present invention designs a deep image deblurring method based on spatiotemporal frequency perception.
[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0008] A deep image deblurring method based on spatiotemporal frequency perception integrates frequency-time information into the video deblurring framework and transforms (2N+1) consecutive blurred frames B t+i , i = 0, ± 1, ± 2, ..., ± N is input into the trained video deblurring framework. t Under the guidance of t ; t represents the current time;
[0009] The main elements of the video deblurring framework include the space-frequency feature extraction module, the spectral prior-guided alignment module, and the temporal energy attention module;
[0010] Blurred frame B t+i After several layers of convolutional layers, the initial features are obtained as follows:
[0011] The space-frequency feature extraction module includes a space branch and a frequency branch. The input of the space-frequency feature extraction module is the initial feature The output features are The spatial branch uses several convolutions to capture The frequency branch is responsible for blur degradation modeling and transforms the input features into Divided into amplitude spectrum and phase spectrum; the spatial branch and the frequency branch are fused to obtain the output characteristics
[0012] The characteristics Input spectral prior guided alignment module, output alignment features The spectral prior-guided alignment module first uses Fourier transform to obtain the spectrum of spatial characteristics. Since the spectrum contains information about the direction of motion, the input feature The alignment is better achieved; the alignment part uses a deformable convolution and image pyramid strategy to implicitly learn the deformation function of the features of the tth frame from the features of the t+ith frame;
[0013] The temporal energy attention module aligns the features The input is sent to the temporal energy attention module, which aggregates the clearer areas in adjacent frames into the restored image. Under the guidance of spatial and frequency information, multi-frame information is fused and the fusion feature is output.
[0014] Fusion Features After the space-frequency feature extraction module and several convolutional layers, the final restored image I is obtained. t ;
[0015] The restored image I is measured by using spatial loss and frequency loss t and clear image S t The difference between them is used to train the video deblurring framework; among them, the spatial loss L spa is the Charbonnier penalty function adopted in the spatial domain; the frequency loss is supervised by the ground truth amplitude L spec and supervision L from the energy spectrum ener Consists of.
[0016] Furthermore, the spatial-frequency feature extraction module obtains the output feature The process includes:
[0017] The spatial branch uses several convolutions to capture spatial content and details;
[0018] The frequency branch is responsible for fuzzy degradation modeling and transforms the input features into Divided into amplitude spectrum and phase spectrum;
[0019] The amplitude spectrum and phase spectrum are convolved, and the convolution results are inverted to the spatial domain through inverse discrete Fourier transform:
[0020]
[0021] in, represents the spectral features extracted from the spectrum, Amp(·) and Pha(·) represent the amplitude spectrum and phase spectrum respectively, and F -1 stands for Inverse Discrete Fourier Transform;
[0022] Combining the spatial branch and the frequency branch, and in order to ensure the convergence of the network, a residual connection is used, that is, the initial features are added to obtain the features
[0023]
[0024] where c 1 、c 2 and c 3 Represents different convolutional layers.
[0025] Furthermore, the spectral prior-guided alignment module is based on the feature The image pyramid strategy is used to achieve multi-level alignment and generate the alignment characteristics at the t+i moment. Specifically include:
[0026] right Perform Fourier transform to obtain spectral characteristics;
[0027] Deformable convolution is used to implicitly learn the deformation function of the features of the t-th frame from the features of the t+i-th frame to align adjacent frames. For deformable convolution, the learning offset of each position is usually obtained by every 2M channels, where M represents the size of the convolution kernel; the features of the t+i-th frame refer to the content of the t+i-th frame; the features of the t+i-th frame refer to the content of the t+i-th frame;
[0028] Perform global average pooling on the spectral features to adjust the offset of each position using global motion information: Given the extracted features and The corresponding spectral characteristics The deviation Δx at time t+i t+i for:
[0029]
[0030] where c 4 (·) and g(·) represent a function composed of several standard convolutions and a global average pooling function, respectively.
[0031] Furthermore, the alignment feature Input to the temporal energy attention module, output The process includes:
[0032] For the frequency domain, the energy spectrum of multiple frames is taken as a classifier to judge the importance of these frames. The frequency attention map is calculated according to the softmax function:
[0033]
[0034] in, represents the weight calculated by the attention module, f(·) represents a new function composed of Fourier transform, convolution layer and Lrelu activation function in sequence, |·| represents the L2 norm, Represents the characteristics after alignment at the (t+i)th moment.
[0035] Execute the spatial attention mechanism in parallel with the temporal energy attention module;
[0036] Under the guidance of spatial and frequency information, multiple frames of information are fused and output to obtain
[0037] Furthermore, when training the video deblurring framework, the spatial loss L spa for:
[0038]
[0039] Among them, ε is a constant, usually set to 0.001;
[0040] L spec and L ener They are:
[0041]
[0042] L ener =Norm[F(I i )-Norm[F(S i );
[0043] Where F(·) represents the Fourier transform operation, Norm[·] represents the L2 norm operation; the total loss function L of the training video deblurring framework is all for:
[0044] L all =L spa +λ(L spec +L ener );
[0045] Among them, λ is the weight factor.
[0046] Compared with the prior art, the beneficial technical effects of the present invention are:
[0047] 1. Feature extraction. It is found through experiments that blur degradation can be effectively modeled in the frequency domain. Therefore, the present invention designs a space-frequency feature extraction module using Fourier transform and pure convolution unit.
[0048] 2. Alignment. Temporal multi-frame alignment usually aims to reduce the contextual differences between adjacent frames. However, in the case of unpredictable blur, where high-frequency details are severely lost, alignment often becomes less effective. Based on experimental results, this paper proposes a spectral prior-guided alignment module to mitigate the negative impact caused by motion blur in alignment by mining the amplified blur degradation information in the frequency spectrum.
[0049] 3. Fusion. Temporal multi-frame fusion is responsible for aggregating clear segments from multiple frames. However, most existing methods only focus on local spatial features, which are limited by the previous alignment results. To solve this problem, the present invention proposes a temporal energy attention module, which uses temporal spectral energy as a global clarity guide, which has a complementary effect with the previous local spatial approach.
[0050] 4. Optimization. The present invention finds that the frequency spectrum can be regarded as an indicator of global fuzziness. Therefore, the present invention designs a frequency spectrum loss and an energy loss function to better optimize the proposed network algorithm in the frequency domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is an overall flow chart of the deblurring method in the present invention;
[0052] Figure 2 A comparison diagram of blurry-clear image pairs between frequency domain and spatial domain according to the present invention;
[0053] Figure 3 A schematic diagram of the invention for finding the prior of the spectrum in the fuzzy direction;
[0054] Figure 4 A schematic diagram of the invention for finding the prior of spectral energy on blur size;
[0055] Figure 5 Schematic diagram of a space-frequency feature extraction module of the present invention;
[0056] Figure 6 A schematic diagram of the alignment module guided by spectral priors of the present invention;
[0057] Figure 7 Schematic diagram of the temporal energy attention module of the present invention. DETAILED DESCRIPTION
[0058] A preferred embodiment of the present invention is described in detail below with reference to the accompanying drawings.
[0059] 1. Exploring the temporal priors in fuzzy sequences from a spectrum perspective
[0060] The present invention re-examines the blurred sequence in Fourier space and finds some intrinsic frequency-time priors, including prior one, prior two, and prior three, which mean that the temporal blur degradation can be easily decomposed in the latent frequency domain. Based on these priors, the present invention proposes a new Fourier-based frequency-time video deblurring solution, in which the core design is to adapt the temporal spectrum to a deep video deblurring process, including feature extraction, alignment, aggregation, and optimization. Finally, a large number of experiments show that the model of the present invention has excellent performance compared with other state-of-the-art methods, thus confirming the effectiveness of frequency-time prior modeling.
[0061] First of all, the present invention Figure 2 , Figure 3 , Figure 4 Some temporal prior information is extracted to guide video deblurring. In this part, the present invention gives some theoretical and visual analysis.
[0062] 1.1 Approximate motion blur modeling
[0063] For the sake of clarity and simplicity, the present invention ignores the influence of depth change. The camera jitter degradation model can be simplified into mathematical form:
[0064] B t =S t *K t +n t ;
[0065] Among them, K t is an unknown blur kernel, n t is additive white noise.
[0066] 1.2 Temporal Fuzzy Priors in Fourier Space
[0067] A priori, blur degradation can be better modeled in the frequency domain. It is well known that images can be represented by real grayscale values in the spatial domain and by complex frequency values in the frequency domain. From the perspective of image representation, since both motion and frequency are directional, while grayscale values are not, the motion of the camera is easier to observe in the frequency domain than in the spatial domain. Therefore, sharp and blurred images can be well distinguished in the frequency domain, such as Figure 2 shown.
[0068] A priori two, the frequency spectrum can amplify the unpredictable variations of temporal motion blur.
[0069] As mentioned above, Fourier transform is widely used to evaluate the frequency characteristics of an image. For images containing multiple color channels, Fourier transform is calculated separately for each channel.
[0070] Given an image S, the Fourier transform F converts it into the complex components of the Fourier space, expressed as:
[0071]
[0072] Assume that some motion blur kernels K(t) with θ direction at time t are known to simulate the motion of the camera. The frequency response of these kernels in the θ+π / 2 direction is relatively weak, while their frequency spectra have a large amplitude in the θ+π / 2 direction, such as Figure 3 shown.
[0073] In addition, according to the convolution theorem, the convolution operation in the spatial domain is equal to the product in the Fourier frequency domain, which can be expressed as:
[0074] F(B t )=F(S t ★K t )=F(S t )·F(K t );
[0075] Where F(·) represents the Fourier transform function.
[0076] According to the equation, the special stripe pattern in the motion kernel spectrum will be brought into the spectrum of the blurred frame.
[0077] Therefore, the frequency spectrum of the blurred frame exhibits a directional pattern that is roughly perpendicular to the direction of camera motion, e.g. Figure 3 As shown at the bottom. The frequency spectrum can amplify the unpredictable fluctuations in temporal motion blur. Inspired by this, the present invention designs a spectral prior-guided alignment module to mitigate the negative impact caused by motion blur in alignment.
[0078] Prior three, blur degradation may weaken the energy spectrum of sharp videos.
[0079] Without compromising generality, the blur kernel K t can be normalized. Since the integral of incoherent light is always non-negative, the blur kernel should be non-negative. In this case, the present invention has the following constraints on the kernel:
[0080] K t ≥0 and∫K t =1.
[0081] Based on this, the present invention infers that motion blur does not amplify the Fourier spectrum, as evidenced by the following formula:
[0082] |F(K t (x))|=|∫K t (x)e -jωx dx|≤∫|K_t(x)|dx=∫K_t(x)dx=1.
[0083] where (|.|) represents the modulus and x is the kernel K t The above formula shows that greater blur may mean more energy loss. The present invention further proves this conclusion through a toy experiment, such as Figure 4 Inspired by this finding, the present invention designs a temporal energy attention module to utilize frames whose energy is not degraded by blur in the time series.
[0084] 2. Design of video deblurring algorithm based on temporal spectrum prior
[0085] The overall framework of the method in the present invention is as follows Figure 1 As shown, its goal is to t Under the guidance of t , given (2N+1) consecutive blurred frames B t+i , where i = 0, ±1, ±2, ..., ±N. The present invention innovatively integrates frequency-time information into a widely used video deblurring framework, which includes feature extraction, alignment, fusion and optimization. The following subsections will describe the specific details in detail.
[0086] 2.1 Space-Frequency Feature Extraction Module (SFE Module)
[0087] In this work, we find that blur degradation can be effectively modeled in the frequency domain, as shown in the priori 1 in Section 1.2. Inspired by this, we design a space-frequency feature extraction module, which consists of a space branch and a frequency branch, as Figure 5 shown.
[0088] For the spatial branch, the present invention employs several convolutions to capture spatial content and details. Meanwhile, the frequency branch is responsible for fuzzy degradation modeling, which separates the input features into amplitude spectrum and phase spectrum through DFT. Then, in order to effectively condition the frequency information, these spectra are processed by (1×1) convolution kernels and inverted to the spatial domain through inverse discrete Fourier transform (IDFT), which can be expressed as:
[0089]
[0090] Among them, Amp and Pha represent the amplitude spectrum and phase spectrum respectively, and F -1 Stands for Inverse Discrete Fourier Transform.
[0091] Finally, the features extracted from the two branches of the SFE module are expressed as:
[0092]
[0093] where c 1 、c 2 、c 3 Represents different convolutional layers with (1×1) kernels.
[0094] 2.2 Spectral Prior-Guided Alignment Module
[0095] Existing temporal alignment methods usually only focus on spatial characteristics to reduce content differences. However, most of them only achieve suboptimal results due to the unpredictable motion blur in time. To address this problem, the present invention deepens its understanding by re-examining blur degradation in the Fourier domain and finds that the frequency spectrum can amplify the unpredictable changes of temporal motion blur, as shown in prior 2. Inspired by this prior, the present invention proposes an alignment module guided by spectral priors, which uses the amplified blur information in the frequency spectrum to mitigate the negative impact of unpredictable blur in alignment.
[0096] Specifically, if Figure 6 As shown, the present invention first applies Fourier transform to obtain the spectrum of spatial characteristics. In order to effectively align adjacent frames, the present invention adopts deformable convolution to implicitly learn the deformation function of the characteristics of the tth frame from the characteristics of the (t+i)th frame.
[0097] For deformable convolutions, the learned offset at each position is usually obtained for every 2M channels, where M represents the size of the convolution kernel.
[0098] To improve alignment in the case of large blur, we perform global average pooling on the spectral features to adjust the offset of each position using global motion information. and their spectral characteristics The deviation Δx at the (t+i)th moment t+i It is learned through:
[0099]
[0100] where c 4 (.) and g(.) represent several standard convolution and global average pooling functions. Through the learned offset, the present invention uses deformable convolution to warp adjacent frames to the reference frame. In addition, due to the depth change, the degree of motion is sensitive to the change. To solve this problem, the present invention uses a feature pyramid strategy to ensure adaptation to blurs of different scales.
[0101] 2.3 Temporal Energy Attention Module
[0102] In video deblurring, the key challenge is to exploit sharper regions in adjacent frames and aggregate them into the restored image. Most existing methods only focus on local spatial properties, which are limited by the previous alignment results. To alleviate this problem, the present invention explores global guidance in aggregation via Fourier transform and finds that sharper regions usually contain more spectral energy, as shown in the prior three in Section 1.2.
[0103] Therefore, the present invention adds time domain spectral energy to the aggregation to estimate the sharpness of the time domain frame, thereby performing time domain fusion more effectively. Specifically, the present invention designs the following Figure 7 The time-domain energy attention fusion shown, which combines the advantages of spatial and frequency domains.
[0104] For the frequency domain, the present invention uses the energy spectrum of multiple frames as a classifier to determine the importance of these frames. The frequency attention map can be calculated as:
[0105]
[0106] in Represents the characteristics after alignment at the (t+i)th moment.
[0107] Then, in order to aggregate more local textures and details, a spatial attention mechanism is performed in parallel with the temporal energy attention mechanism. Finally, under the guidance of spatial and frequency information, multi-frame information is effectively fused and output.
[0108] 2.4 Optimization
[0109] For video deblurring, the present invention considers two aspects, namely spatial domain and frequency domain, and uses two loss functions to measure the restored image I t+i and clear image S y+i The difference between.
[0110] In the spatial domain, the present invention adopts the Charbonnier penalty function as the spatial loss, focusing on the restored pixel-level details, which is defined as:
[0111]
[0112] Where ε is a constant, which is set to 1×10 -3 .
[0113] Received Figure 2 Inspired by , the present invention finds that the spectrum not only contains high-frequency and low-frequency signals, but also can separate them according to the difference in blur. Therefore, in order to further restore high-frequency details, in the training stage, the present invention designs a supervision from the ground truth amplitude and energy spectrum composed of L spec and: ener The frequency loss of the component is defined as:
[0114]
[0115] L ener =Norm[F(I i )]-Norm[F(S i )];
[0116] Among them, F and Norm represent Fourier transform and L2 norm operations.
[0117] L all =L spa +λ(L spec +L ener );
[0118] Here, λ is a weight factor, which is empirically set to 0.1.
[0119] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.
[0120] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A deep image deblurring method based on spatiotemporal frequency perception, which integrates frequency-time information into the video deblurring framework and converts (2N+1) consecutive blurred frames B t+i , i = 0, ± 1, ± 2, ..., ± N is input into the trained video deblurring framework. t Under the guidance of t ; t represents the current time; The video deblurring framework includes a space-frequency feature extraction module, a spectral prior-guided alignment module, and a temporal energy attention module; Blurred frame B t+i After several layers of convolutional layers, the initial features are obtained The space-frequency feature extraction module includes a space branch and a frequency branch. The input of the space-frequency feature extraction module is the initial feature The output features are The spatial branch uses several convolutions to capture The frequency branch is responsible for blur degradation modeling and transforms the input features into Divided into amplitude spectrum and phase spectrum; the spatial branch and the frequency branch are fused to obtain the output characteristics The characteristics Input spectral prior guided alignment module, output alignment features The spectral prior-guided alignment module first uses Fourier transform to obtain the spectrum of spatial characteristics. Since the spectrum contains information about the direction of motion, the feature Can be used to achieve alignment; The alignment part uses a deformable convolution and image pyramid strategy to implicitly learn the deformation function of the features of the tth frame from the features of the t+ith frame; The temporal energy attention module aligns the features The input is sent to the temporal energy attention module, which aggregates the clearer areas in adjacent frames into the restored image. Under the guidance of spatial and frequency information, multi-frame information is fused and the fusion feature is output. Fusion Features After the space-frequency feature extraction module and several convolutional layers, the final restored image I is obtained. t ; The restored image I is measured by using spatial loss and frequency loss t and clear image S t The difference between them is used to train the video deblurring framework; Among them, the space loss L spa is the Charbonnier penalty function adopted in the spatial domain; the frequency loss is supervised by the ground truth amplitude L spec and supervision L from the energy spectrum ener Consists of.
2. The deep image deblurring method based on spatiotemporal frequency perception according to claim 1, It is characterized in that The spatial-frequency feature extraction module obtains the output features The process includes: The spatial branch uses several convolutions to capture spatial content and details; The frequency branch is responsible for fuzzy degradation modeling and transforms the input features into Divided into amplitude spectrum and phase spectrum; The amplitude spectrum and phase spectrum are convolved, and the convolution result is inverted to the spatial domain through inverse discrete Fourier transform: in, represents the spectral features extracted from the spectrum, Amp(·) and Pha(·) represent the amplitude spectrum and phase spectrum respectively, and F -1 stands for Inverse Discrete Fourier Transform; Combining the spatial branch and the frequency branch, and in order to ensure the convergence of the network, a residual connection is used, that is, the initial features are added to obtain the features where c 1 、c 2 and c 3 Represents different convolutional layers.
3. The deep image deblurring method based on spatiotemporal frequency perception according to claim 1, Features: The spectral prior-guided alignment module is based on the features The image pyramid strategy is used to achieve multi-level alignment and generate the alignment characteristics at the t+i moment. Specifically include: right Perform Fourier transform to obtain spectral characteristics; Deformable convolution is used to implicitly learn the deformation function of the features of the tth frame from the features of the t+ith frame to align adjacent frames. For deformable convolution, the learning offset of each position is usually obtained by every 2M channels, where M represents the size of the convolution kernel; Perform global average pooling on the spectral features to adjust the offset of each position using global motion information: Given the extracted features and The corresponding spectral characteristics The deviation Δx at time t+i t+i for: where c 4 (·) and g(·) represent a function composed of several standard convolutions and a global average pooling function, respectively.
4. The deep image deblurring method based on spatiotemporal frequency perception according to claim 1, It is characterized in that Align Features Input to the temporal energy attention module, output The process includes: For the frequency domain, the energy spectrum of multiple frames is taken as a classifier to judge the importance of these frames. The frequency attention map is calculated according to the softmax function: in, represents the weight calculated by the attention module, f(·) represents a new function composed of Fourier transform, convolution layer and Lrelu activation function in sequence, |·| represents the L2 norm, represents the characteristics after alignment at the (t+i)th moment; Execute the spatial attention mechanism in parallel with the temporal energy attention module; Under the guidance of spatial and frequency information, multiple frames of information are fused and output to obtain 5. The deep image deblurring method based on spatiotemporal frequency perception according to claim 1, It is characterized in that When training the video deblurring framework, the spatial loss L spa for: Among them, ε is a constant; L spec and L ener They are: L ener =Norm[F(I i )]-Norm[F(S i )]; Where F(·) represents the Fourier transform operation, Norm[·] represents the L2 norm operation; the total loss function L of the training video deblurring framework is all for: L all =L spa +λ(L spec +L ener ); Among them, λ is the weight factor.
Citation Information
Patent Citations
Video saliency target detection method based on frequency domain prior
CN111178188A
Neural network video deblurring method based on multi-attention mechanism fusion
CN111539884A