Video image deblurring method

By constructing a multi-level image processing model, using multi-scale encoding and decoding and deep convolutional networks to screen frame images with high similarity and clarity, the problem of poor image restoration effect in the prior art is solved, and efficient image defuzzing is achieved.

CN120147185APending Publication Date: 2025-06-13DONGGUAN CITY COLLEGE
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510226014.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has poor image restoration effect due to blur problems in image recognition, and the calculation amount required for model parameter adjustment is huge.

Method used

By obtaining the target frame image and the upper and lower frame video images filtered by similarity and clarity, a multi-level image processing model is constructed, including an encoder, a long and short-term memory network LSTM and a decoder, and image defuzzing is performed using multi-scale encoding and decoding and deep convolutional networks.

Benefits of technology

The calculation amount required for model parameter adjustment is reduced, the efficiency and reusability of image deblurring are improved, and the accuracy of image restoration is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147185A_ABST
    Figure CN120147185A_ABST
Patent Text Reader

Abstract

The invention discloses a video image deblurring method, and relates to the technical field of image processing. Comprising the following steps: acquiring a to-be-processed blurred image set comprising a target frame image, and screening first K video images with the maximum sum of similarity S and image definition C in the target frame image and N frame images adjacent to the target frame image to obtain an extracted image set; dividing each video image into T blurred images of different levels; and inputting the target frame image and the blurred images of different levels into an image processing model in sequence to obtain a clear image after the target frame image is processed. According to the invention, the image processing model based on multi-scale coding and decoding and the deep convolutional network is designed by combining the advantages of SRN-DeburNet and separable convolution, and the reusability and efficiency are improved on the premise of ensuring the accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method for deblurring video images. Background Art

[0002] With the development of artificial intelligence and machine vision, the application scope of image recognition is becoming wider and wider. In image recognition, the objects to be recognized are often in motion, and even the camera itself may be in motion. Coupled with reasons such as weather and line of sight, motion blur, noise blur, and defocus blur will occur. In view of the characteristics of motion blur, an image blurring process can be constructed and then restored. In traditional methods, the method of estimating the blur kernel and directly performing inverse filtering transformation is used to restore the image. The accuracy of blur kernel estimation determines the quality of image restoration. Because this method often leads to poor image restoration effects due to the failure to accurately estimate the blur kernel.

[0003] Currently, the mainstream method is to use deep learning (such as using CNN to restore images). Chakrabarti proposed a neural method for blind motion deblurring, which uses CNN to predict the Fourier coefficients of the blur kernel and uses a non-blind deblurring method to restore the image. Nah S, Kim T H, and Lee K M et al. proposed a deep multi-scale convolutional neural network for dynamic scene deblurring, which improved the target detection effect by using a multi-scale neural network. TAO X, GAO H, and WANG Y et al. proposed a convolutional neural network for learning non-uniform motion blur removal, and proposed a scale recursive network SRN-DeblurNet, which uses a simpler network and fewer parameters and adopts a "coarse to fine" scheme to restore the image.

[0004] Existing methods need to make assumptions, simplifications, and modelings for blurs in sequence, and the computational complexity required for adjusting model parameters is huge, and the image restoration effect is poor. Summary of the Invention

[0005] Based on this, it is necessary to provide a method for deblurring video images in view of the above technical problems.

[0006] An embodiment of the present invention provides a method for deblurring video images, including:

[0007] Obtaining a to-be-processed blurred image set including a target frame image;

[0008] In the video image set, screening the top K video images with the largest sum of similarity S and image clarity C values in the N video images adjacent to the target frame image to obtain an extraction image set including upper and lower frame video images that are relevant to the target frame image; dividing each video image in the extraction image set into T blurred images at different levels;

[0009] Build an image processing model, which includes multiple processing layers; each processing layer includes an encoder, a long short-term memory network (LSTM), and a decoder; the input end of the encoder serves as the input end of the image processing model, the output end of the encoder is connected to the input end of the LSTM, the output end of the LSTM is connected to the input end of the decoder, and the output end of the decoder serves as the output end of the image processing model; wherein, the output end of each processing layer is connected to the input end of the next processing layer, and the output end of the LSTM in each layer is also connected to the input end of the LSTM in the next processing layer;

[0010] Input the target frame image and blurred images at different levels into the image processing model in sequence, encode them through the encoder, and then capture the long-term dependence relationship between the blurred images at different levels and the target frame image through the LSTM to obtain encoded features; decode the encoded features through the decoder to obtain feature points representing the similarity correlation between the target frame image and the upper and lower frame video images; process the feature points through the Softmax function to obtain the clear image after processing the target frame image.

[0011] Optionally, determining the similarity S specifically includes:

[0012] Convert the target frame image and N frames of images adjacent to the target frame image into feature vectors, and use the cosine similarity between the feature vectors as the similarity S;

[0013] Determine the image clarity C, and the calculation formula is:

[0014] Image clarity = (Spatial complexity - Noise complexity) / Spatial complexity;

[0015] Wherein, the spatial complexity refers to the number of pixel points, and the noise complexity refers to the number of noises.

[0016] Optionally, the encoder includes:

[0017] The input end of the encoder is the first encoding module Enblock, the output end of the first encoding module Enblock is sequentially connected to three separable convolutions SepConv, and the output end of the third separable convolution SepConv is connected to the second encoding module Enblock; the output end of the second encoding module Enblock is sequentially connected to seven separable convolutions SepConv, and the output end of the seventh separable convolution SepConv serves as the output end of the encoder;

[0018] The decoder includes:

[0019] The output end of the long short-term memory network (LSTM) is sequentially connected to seven separable convolutions (SepConv). The input end of the first separable convolution (SepConv) is used as the input end of the decoder. The output end of the seventh separable convolution (SepConv) is connected to the input end of the first decoding module (Deblock). The output end of the first decoding module (Deblock) is sequentially connected to three separable convolutions (SepConv). The output end of the third separable convolution (SepConv) is connected to the second decoding module (Deblock). The output end of the second decoding module (Deblock) is used as the output end of the decoder.

[0020] Optionally, the clear image obtained after processing the blurred image by the image processing model has the following calculation formula:

[0021] I i , h i = Net(B i , I i-1 , h i-1 ; θ);

[0022] f i = Net(B i , I i-1 );

[0023] h i , g i = LSTM(B i , I i-1 );

[0024] I i = Net(g i , θ D );

[0025] Among them, I represents the image intensity, h represents the image height, f represents the image color value, g represents the image processing result, B i represents the input and output of each module in the image processing model, and θ represents the parameter.

[0026] Optionally, it further includes that the loss function Losses of the image processing model is:

[0027]

[0028] Among them, N is the number of image frames, I i is the output of each layer, is the average value of all outputs I i , and κ i is the weight of each layer.

[0029] The above video image deblurring method provided by the embodiments of the present invention, compared with the prior art, has the following beneficial effects:

[0030] For the input data of the image processing model, the present invention reduces the computational complexity required for model parameter adjustment through a dual screening criterion of similarity S and clarity C.

[0031] More importantly, the present invention combines the advantages of SRN-DeblurNet and separable convolution, and designs an image processing model based on multi-scale encoding and decoding and deep convolutional network. The long short-term memory network LSTM is used to capture the long-term dependence relationship between blurred images at different levels and target frame images, and encoded features are obtained; the encoded features are decoded to obtain feature points representing the similarity correlation between the target frame image and the upper and lower frame video images; under the premise of ensuring the correct rate, the reusability and efficiency can be improved. Description of the Drawings

[0032] Figure 1 It is a schematic flow chart of a video image deblurring method provided in an embodiment;

[0033] Figure 2 It is a schematic diagram of a conventional convolution Conv of a video image deblurring method provided in an embodiment;

[0034] Figure 3 It is a schematic diagram of a separable convolution SepConv of a video image deblurring method provided in an embodiment;

[0035] Figure 4 It is a model flow chart of a video image deblurring method provided in an embodiment. Detailed Embodiments

[0036] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0037] In traditional methods, assumptions, simplifications, and modeling are performed on blurs in sequence, and the computational complexity required for adjusting the parameters of the model is huge, and the image restoration effect is poor. Real blurs often do not conform to the assumptions. Blurs are not uniform and linear, and blurs are not only generated by cameras. Therefore, it is very complicated to model the restoration process with real blurred images. Deep learning-based methods can be used to train the parameters. CNN has a qualitative improvement in the effect of training parameters compared with traditional methods. The scale cyclic network is a more effective deblurring network structure in CNN. It starts from coarse and gradually refines, which is a process from coarse to fine.

[0038] In a video, the upper and lower frames are similar and related. Image restoration of a blurred video is not simply treating each frame in the video as an individual blurred image for restoration. Instead, considering the correlation, as long as one of multiple similar pictures can be restored clearly, other images can "learn" restoration from the clear picture. Due to the large amount of work and low efficiency of existing video restoration methods, the present invention establishes a network structure with a "coarse to fine" scheme, combines the advantages of SRN-DeblurNet and separable convolution, and designs an image processing model based on a multi-scale encoding and decoding and separable convolution network. While ensuring the accuracy rate, it reduces the number of parameters, improves reusability and efficiency. By further optimizing the parameter settings, multi-scale encoding and decoding and separable convolution can be applied to other video image processing.

[0039] In one embodiment, a method for deblurring video images is provided, as Figure 1 shown, the method includes:

[0040] 1. Obtain a set of blurred images to be processed including the target frame image.

[0041] 2. In the video image set, screen the top K video images with the largest sum of similarity S and image clarity C values from the N video images adjacent to the target frame image to obtain an extraction image set containing upper and lower frame video images that are correlated with the target frame image. Divide each video image in the extraction image set into T different levels of blurred images.

[0042] 3. Construct an image processing model, where the image processing model includes multiple processing layers. Each processing layer includes an encoder, a long short-term memory network LSTM, and a decoder. The input end of the encoder is used as the input end of the image processing model. The output end of the encoder is connected to the input end of the long short-term memory network LSTM. The output end of the long short-term memory network LSTM is connected to the input end of the decoder. The output end of the decoder is used as the output end of the image processing model. Among them, the output end of each processing layer is connected to the input end of the next processing layer, and the output end of the long short-term memory network LSTM in each layer is also connected to the input end of the long short-term memory network LSTM in the next processing layer.

[0043] (1) Encoder

[0044] The input end of the encoder is the first encoding module Enblock. The output end of the first encoding module Enblock is sequentially connected to three separable convolutions SepConv. The output end of the third separable convolution SepConv is connected to the second encoding module Enblock. The output end of the second encoding module Enblock is sequentially connected to seven separable convolutions SepConv. The output end of the seventh separable convolution SepConv is used as the output end of the encoder.

[0045] (2) Decoder

[0046] The output end of the long short-term memory network LSTM is successively connected to seven separable convolutions SepConv. The input end of the first separable convolution SepConv is used as the input end of the decoder. The output end of the seventh separable convolution SepConv is connected to the input end of the first decoding module Deblock. The output end of the first decoding module Deblock is successively connected to three separable convolutions SepConv. The output end of the third separable convolution SepConv is connected to the second decoding module Deblock. The output end of the second decoding module Deblock is used as the output end of the decoder.

[0047] 4. Input the target frame image and blurred images at different levels into the image processing model in sequence. Encode them through the encoder, and then capture the long-term dependence relationship between the blurred images at different levels and the target frame image through the long short-term memory network LSTM to obtain encoded features. Decode the encoded features through the decoder to obtain feature points representing the similarity correlation between the target frame image and the upper and lower frame video images. Process the feature points through the Softmax function to obtain the clear image after processing the target frame image.

[0048] The specific implementation process is as follows:

[0049] As Figure 2 shown, the computational complexity of the conventional convolution: k*k*Cin*Wout*Hout*Cout. As Figure 3 shown, the computational complexity of the depthwise separable convolution: k*k*Cin*Wout*Hout.

[0050] The computational complexity of the conventional convolution is Cout (the number of output convolution channels) times that of the depthwise separable convolution. Among them, the input convolution: Win*Hin*Cin, Win: the width of the input convolution, Hin: the height of the input convolution, Cin: the number of input convolution channels, Wout: the width of the output convolution, Hout: the height of the output convolution, Cout: the number of output convolution channels, convolution kernel: k*k.

[0051] 1. Image Processing Model Based on Multi-Scale Encoding and Decoding and Depthwise Separable Convolution

[0052] (1) Algorithm Flow

[0053] A. For a certain frame of image (i.e., the target frame image), it is necessary to use a total of N images before and after in the context. Calculate the similarity S (0 <= S <= 1) between these N images and the original image, judge the image clarity C (0 <= C <= 1), and select the top K images with the largest S + C value.

[0054] The similarity calculation method is as follows: The image can be converted into feature vectors (such as the feature vectors extracted using a convolutional neural network), and then the cosine similarity between these feature vectors is calculated to measure the similarity of the images.

[0055] The image sharpness can be calculated using the following formula:

[0056] Image sharpness = (Spatial complexity - Noise complexity) / Spatial complexity;

[0057] Among them, the spatial complexity refers to the number of pixel points in the image, and the noise complexity refers to the number of noises in the image.

[0058] B. For a certain image, it can be divided into T different levels of blurred images. When the worst scale t = 1, it represents the worst scale, that is, the level with the smallest resolution. The deblurring of the picture is a process from coarse to fine. For the first layer, its input is an image, and the image is divided from coarse to fine.

[0059] The multi-scale image technology is also called the multi-resolution technology (MRA), which refers to the multi-scale expression of the image and separate processing at different scales. The reason for this is that in many cases, the features that are not easily seen or obtained at one scale are easily discovered or extracted at another scale. Therefore, the multi-scale technology is more commonly used in extracting image features. To process the image in the multi-scale case, first, the image needs to be expressed in the multi-scale case, and the mutual connection between each scale needs to be found. The pyramid structure is a form of multi-scale expression of the image. The multi-scale transformation technologies used to obtain the multi-scale expression can basically be divided into three categories: scale space technology, time-scale technology, and time-frequency technology.

[0060] The extraction process of the multi-scale image technology:

[0061] For an N*N image (N = 2^n), if one pixel is taken out every other pixel in both directions, the taken-out pixels will form an N / 2*N / 2 image. That is to say, by performing 1:2 sub-sampling in both directions, a relatively rough thumbnail of the original image can be obtained. This process is repeated until the original image becomes a 1*1 image. Through this process, a series of images can be obtained: N*N, N / 2*N / 2, N / 2^2*N / 2^2,..., N / 2^n*N / 2^n. The series of images obtained form the shape of a pyramid, and the original image corresponds to the 0th layer. Although there are many image sequences, the storage space of the pyramid structure is not very large. For a complete pyramid structure with N + 1 layers, the total number of units is N^2(1 + 1 / 4 + 1 / 4^2 +...... + 1 / 4^n) <= (4N^2) / 3.

[0062] C. For each layer, first, it passes through an encoding module Enblock, then through three separable convolutions SepConv, then through another encoding module Enblock and seven separable convolutions SepConv to complete the encoding stage. Then, it passes through a long short-term memory network LSTM to start the decoding stage. The decoding stage is symmetric to the encoding stage, passing through seven separable convolutions SepConv, a decoding module Deblock, then through three separable convolutions SepConv, and a decoding module Deblock. The model process is as Figure 4 shown, Figure 4 the dotted lines in i-1 ↑ represent skip connections, and ↑ also serves as the input to the current layer. The middle blue arrow indicates that the output of the upper long short-term memory network LSTM is used as the input to the current layer's long short-term memory network LSTM to more effectively capture spatio-temporal correlations.

[0063] (2) Model architecture

[0064] The image processing model architecture is shown in Table 1.

[0065] Table 1 Model architecture

[0066]

[0067] The first round of convolution operation is on a three-channel image with an input size of 55*55. After one convolution, the feature map has 96 elements, the convolution kernel size is 3*3, the activation function is ReLU, the stride is 2, which reduces the dimension of the features, and the output is 96 images with a size of 27*27. Then, it goes through three rounds of separable convolutions with a convolution stride of 1, without changing the dimension.

[0068] The second round of convolution operation passes through a convolution kernel of 192 with a size of 3*3 and a stride of 2, which reduces the dimension and doubles the number of channels at the same time. Then, it goes through 7 SepConv with a stride of 1 each, and the feature size is 13*13.

[0069] The third round of operation goes through an LSTM. The fourth and fifth rounds of convolution operations are both transposed convolution processes.

[0070] In the fourth round, first, it goes through a transposed convolution process to double the size and reduce the number of channels by half, and then goes through 7 SepConv.

[0071] In the fifth round, it also goes through a transposed convolution process to double the size and reduce the number of channels by half, and then goes through 3 SepConv. The result and the three-channel image of 55*55 are sent into the next layer loop and finally output. SepConv (Separable Convolutions) is a separable convolution.

[0072] The calculation formula is as follows:

[0073] I i , h i = Net(B i , I i-1 , h i-1 ; θ);

[0074] f i = Net(B i , I i-1 );

[0075] h i , g i = LSTM(B i , I i-1 );

[0076] I i = Net(g i , θ D );

[0077] Among them, I represents the image intensity, h represents the image height, f represents the image color value, g represents the image processing result, B i represents the input and output of each module in the image processing model, and θ represents the parameter.

[0078] Loss function Losses:

[0079]

[0080] Among them, N is the number of image frames, I i is the output of each layer, is the average value of all outputs I i , k i is the weight of each layer.

[0081] 2. Model training

[0082] The convergence of the deep neural network may be very slow and requires a faster optimizer. Adam is an adaptive learning rate algorithm. Use the Adam method for convergence. The set initial parameter values are: β 1 = 0.9, β 2 = 0.99, and η adopts the default value of 0.001.

[0083] Due to the performance of the machine, set ε = 10 -4 , and after more than 500 trainings, the model converges. Use Keras to create an optimizer:

[0084] Optimizer = keras.optimizers.Adam(lr = 0.001, beta_1 = 0.9, beta_2 = 0.999)

[0085] 150 videos were cropped and sampled, and 48 55*55 images in 3 seconds of each video were used as input.

[0086] 3. Conclusion

[0087] The mean squared error (MSE), MS-SSIM, and peak signal-to-noise ratio (PSNR) were used to judge the model recognition effect. MSE, that is, the mean squared error, the smaller this value, the more similar the two images. The range of the MS-SSIM index is [-1, 1], the smaller the value, the less similar. When SSIM = 1, it means the two images are the same. PSNR, that is, the peak signal-to-noise ratio. According to the results in Table 2, the accuracy of this method has been slightly improved, but the parameters of this method are fewer than the previous two methods, and the efficiency has been increased by at least 37.5%.

[0088] Table 2 Comparison of Model Recognition Effects

[0089]

[0090] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A video image deblurring method, characterized in that: include: Acquire a video image set including a target frame image; In the video image set, the first K video images with the largest sum of similarity S and image clarity C are selected from the N frames of video images adjacent to the target frame image, and an extracted image set containing upper and lower frame video images that are correlated with the target frame image is obtained; each video image in the extracted image set is divided into T fuzzy images of different levels; Construct an image processing model, wherein the image processing model includes multiple processing layers; each processing layer includes an encoder, a long short-term memory network LSTM and a decoder; the input end of the encoder serves as the input end of the image processing model, the output end of the encoder is connected to the input end of the long short-term memory network LSTM, the output end of the long short-term memory network LSTM is connected to the input end of the decoder, and the output end of the decoder serves as the output end of the image processing model; wherein the output end of each processing layer is connected to the input end of the next processing layer, and the output end of the long short-term memory network LSTM in each layer is also connected to the input end of the long short-term memory network LSTM in the next processing layer; The target frame image and blur images of different levels are sequentially input into the image processing model, encoded by the encoder, and then captured through the long short-term memory network (LSTM) to obtain the long-term dependency between blur images of different levels and the target frame image, and the encoded features are obtained; the encoded features are decoded by the decoder to obtain feature points that characterize the similarity correlation between the target frame image and the upper and lower frame video images; the feature points are processed by the Softmax function to obtain a clear image of the processed target frame image.

2. A video image deblurring method as claimed in claim 1, characterized in that: The determining of the similarity S specifically includes: The target frame image and N frame images adjacent to the target frame image are converted into feature vectors, and the cosine similarity between the feature vectors is used as the similarity S; The image definition C is determined by the following calculation formula: Image clarity = (spatial complexity - noise complexity) / spatial complexity; Among them, spatial complexity refers to the number of pixels, and noise complexity refers to the amount of noise.

3. A video image deblurring method as claimed in claim 1, characterized in that: The encoder comprises: The input end of the encoder is the first encoding module Enblock, the output end of the first encoding module Enblock is connected to three separable convolutions SepConv in sequence, and the output end of the third separable convolution SepConv is connected to the second encoding module Enblock; the output end of the second encoding module Enblock is connected to seven separable convolutions SepConv in sequence, and the output end of the seventh separable convolution SepConv is used as the output end of the encoder; The decoder comprises: The output end of the long short-term memory network LSTM is connected to seven separable convolutions SepConv in sequence, the input end of the first separable convolution SepConv is used as the input end of the decoder, the output end of the seventh separable convolution SepConv is connected to the input end of the first decoding module Deblock, the output end of the first decoding module Deblock is connected to three separable convolutions SepConv in sequence, the output end of the third separable convolution SepConv is connected to the second decoding module Deblock, and the output end of the second decoding module Deblock is used as the output end of the decoder.

4. A video image deblurring method as claimed in claim 1, characterized in that: The calculation formula for obtaining a clear image after the blurred image is processed by the image processing model is: I i ,h i =Net(B i ,I i-1 ,h i-1 ;θ); f i =Net(B i ,I i-1 ); h i ,g i =LSTM(B i ,I i-1 ); I i =Net(g i ,θ D ); Among them, I represents the image intensity, h represents the image height, f represents the image color value, g represents the image processing result, and B i Represents the input and output of each module in the image processing model, and θ represents the parameter.

5. A video image deblurring method as claimed in claim 1, characterized in that: The loss function Losses of the image processing model is also included: Where N is the number of image frames, I i For each layer output, For all outputs I i The average value of i is the weight of each layer.

Citation Information

Patent Citations

  • Image processing method and device, storage medium and electronic device

    CN108629743A

  • Depth image deblurring method based on scale recursive network

    CN108846820A

  • The invention discloses a video restoration model training method based on a deep network and a video restoration method

    CN109949234A

  • Dynamic scene blind deblurring method based on asymmetric U-Net network

    CN116188313A

  • Method for detecting and identifying target object by using video monitoring blurred image

    CN117237413A