Image deblurring method based on event data driving
By designing a motion adaptive Transformer defuzzing network, using an adaptive motion mask predictor, motion sparse event block and motion-aware image block, the problem of the spatial correlation between events and blurred areas in the prior art is not effectively mined and the motion-awareness ability is insufficient, and a more efficient image defuzzing effect is achieved.
Patent Information
- Application Number
- CN202510164884.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The existing event-driven defuzzing methods fail to effectively explore the spatial correlation between events and fuzzy areas, lack motion perception capabilities, and the traditional attention mechanism is susceptible to irrelevant tokens in high-density visual information, resulting in a decrease in feature aggregation quality.
A motion-adaptive Transformer is designed to defuzzy the image based on event data, and a motion-adaptive Transformer is used to defuzzy the network, including feature extraction layer, event-image encoder and image reconstruction decoder. Accurate motion area positioning and efficient cross-modal interaction are achieved through adaptive motion mask predictors, motion sparse event blocks and motion-aware image blocks.
The defuzzing performance in complex dynamic scenarios is improved, the number of parameters is reduced, and the robustness is better. The experimental results show that it is better than the state-of-the-art methods on multiple datasets.
Smart Images

Figure CN120070255A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image deblurring, and specifically to an event data-driven image deblurring method. Background Art
[0002] Motion blur is usually caused by the relative motion between the camera and the scene during exposure, and is a common degradation phenomenon in the imaging process of frame cameras. Image deblurring, as a typical ill-posed inverse problem, aims to recover a clear image from a blurred observation. Although deep learning-based solutions have significantly improved traditional deblurring performance, they are still limited by the inherent defect that traditional cameras cannot capture dynamic details during exposure. In contrast, event cameras record pixel intensity changes asynchronously with microsecond-level precision, and their high temporal resolution provides key motion clues for deblurring, showing enhanced potential for traditional methods in recent work.
[0003] Existing event-driven deblurring methods have the following key limitations: (1) The spatial correlation between events and blurred regions is not effectively exploited; (2) There is a lack of motion perception ability. Existing CNN architectures are difficult to model global motion patterns, and have limited ability to locate motion regions hidden in events. Traditional attention mechanisms are vulnerable to interference from irrelevant tokens in high-density visual information, reducing the quality of feature aggregation; (3) Existing solutions fuse event and image features through simple stitching or linear weighted fusion, ignoring the dynamic association between event intensity changes and image spatial consistency, resulting in insufficient detail recovery.
[0004] The above problems limit the application efficiency of event-driven deblurring methods in complex dynamic scenes. There is an urgent need to design a new attention mechanism with both motion perception and computational efficiency to accurately utilize event priors to guide cross-modal feature fusion. Summary of the Invention
[0005] In order to overcome the deficiencies of existing methods, the present invention provides an event data-driven image deblurring method, aiming to improve the deblurring performance in complex dynamic scenes through accurate motion region localization and efficient cross-modal interaction.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] The feature of an event data-driven image deblurring method of the present invention is that it is carried out according to the following steps:
[0008] Step 1: Obtain a set of blurred images, denoted as , where represents the th blurred image, , is the number of blurred images;
[0009] Get the blurred image set The corresponding event sequence is denoted as , Indicates Blurred image The corresponding sequence of events;
[0010] Get a clear image set, denoted as ,in, Indicates A clear image;
[0011] Step 2: Build an event-driven motion adaptive Transformer deblurring network, including: feature extraction layer, Layer events - image encoder, image reconstruction decoder;
[0012] Step 2.1, the feature extraction layer is composed of an image feature extraction layer and an event feature extraction layer, and the feature extraction layer is respectively Blurred image and Event sequence Perform feature extraction and obtain the corresponding Initial blurred image features and Initial event characteristics ;
[0013] Step 2.2, The layer event-image encoder consists of an adaptive motion mask predictor, a motion sparse event block, and a motion-aware image block;
[0014] when When all 1 matrices are used as the first Tier Initial motion mask and with Input into the adaptive motion mask predictor for processing and obtain the Tier Predicted motion mask ;
[0015] The said Tier Predicted motion mask and Input the motion sparse event block for processing to generate the first Tier Event characteristics ;
[0016] The said Tier A predicted motion mask 、 and are processed in the input motion perception image block to generate the th blurred image feature ;
[0017] When , the th predicted motion mask is used as the th initial motion mask and is input into the adaptive motion mask predictor together with the th event feature output by the event-image encoder of the th layer for processing, and the ;
[0018] The th predicted motion mask and the th event feature output by the event-image encoder of the th layer are input into the motion sparse event block for processing to generate the ;
[0019] The th predicted motion mask , the th event feature and the th blurred image feature output by the event-image encoder of the th layer are input into the motion perception image block for processing to generate the ; Thus, the th th blurred image feature is output by the event-image encoder of the
[0020] Step 2.3. After the image reconstruction decoder processes , the th predicted clear image , so as to obtain a predicted clear image set ;
[0021] Step 3. Use the formula to construct a loss function for backpropagation :
[0022] (4)
[0023] Step 4. Based on , and , train the motion adaptive Transformer deblurring network and calculate the loss function , and at the same time use the adaptive moment estimation optimization method with a learning rate to update the network weights. When the number of training iterations reaches the set number or the loss error is less than the set threshold, the training stops, so as to obtain an optimal deblurring model for processing blurred images to obtain corresponding clear images.
[0024] Another feature of the image deblurring method based on event data driving according to the present invention is that the adaptive motion mask predictor includes: a response calculation module, a response fusion unit, an adaptive adjustment unit, and a mask generation unit;
[0025] When , after linearly projecting by the response calculation module, then input it into the GELU activation function for processing to generate the th layer and the th local response . At the same time, based on , perform normalized weighted averaging on to generate the th layer and the th global response ;
[0026] The response fusion unit expands the dimension of to align it with in space, and then splices the expanded global response with to generate the th layer and the th first fusion feature . At the same time, multiply the expanded global response element by element with to generate the th layer and the th second fusion feature ;
[0027] The adaptive adjustment unit is based on , calculate the -th average sparsity rate , and then set the -th layer's -th learnable expansion factor , so as to calculate the -th layer's -th high-response token count :
[0028] (1)
[0029] In Equation (1), is the training stability coefficient, is the -th layer's -th total spatial token count;
[0030] The mask generation unit processes using a multi-layer perceptron to obtain spatial projection features. At the same time, after performing max-pooling and average-pooling operations on the channel dimension of , it is then input into a linear layer and a GELU activation function for processing to obtain channel projection features. Thus, the spatial projection features and the channel projection features are concatenated along the channel and then input into another linear layer and a GELU activation function for processing to obtain the -th layer's -th spatial motion score ; Finally, according to , select the top high-response positions to generate a predicted motion mask .
[0031] Furthermore, the motion sparse event block includes: a motion sparse attention unit and an extended control spatial gating unit;
[0032] When , the motion sparse attention unit normalizes layer by layer, and then successively processes it through a convolutional layer and a depthwise separable convolutional layer to obtain event local projection features, and performs a channel partitioning operation on the event local projection features to generate the -th layer's -th event query matrix , the -th event key matrix , the -th event value matrix ; Thus, use Equation (2) to calculate the -th layer's -th event motion sparse transposed attention map , and combine with After multiplication, the -th updated event feature of the -th layer is generated :
[0033] (2)
[0034] In formula (2), is the scaling factor to be learned, Softmax represents the activation function, represents element-wise multiplication, represents transpose;
[0035] The extended control spatial gating unit processes by inputting it into the GELU activation function to obtain the spatial modulation score. At the same time, the is converted into the hybrid transmission rate by using the Sigmoid function. Then, after element-wise multiplication of the spatial modulation score, the hybrid transmission rate, and , and then adding it to , the -th event feature of the -th layer is obtained ;
[0036] Furthermore, the motion-aware image patch includes: a motion-aware attention unit and a cross-modal intensity gating unit;
[0037] When , the motion-aware attention unit performs layer normalization on , and after being processed by a convolutional layer and a depthwise separable convolutional layer in sequence, the local image projection feature is obtained. The channel division operation is performed on the local image projection feature to generate the -th image query matrix , the -th image key matrix , the -th image value matrix . Then, the -th image motion sparse transposed attention map of the -th layer is calculated by using formula (3) . After multiplying with , the -th updated image feature of the -th layer is generated
[0038] (3)
[0039] In formula (3), is the motion factor to be learned;
[0040] The cross-modal intensity gating unit will process the input layer normalization layer to obtain normalized image features. At the same time, process the input layer normalization layer and the GELU activation function to obtain event-modulated features. Then, after element-wise multiplying the normalized image features and the event-modulated features, add them to to obtain the th cross-modal fusion image feature of the .
[0041] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the image deblurring method, and the processor is configured to execute the program stored in the memory.
[0042] A computer-readable storage medium according to the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the image deblurring method.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] 1. The present invention utilizes an event-driven video deblurring task. Through a motion-adaptive design, it can enable the Transformer to achieve a good end-to-end deblurring effect. And compared with existing deblurring methods, it reduces the number of parameters and has better robustness on different datasets. Experimental results show that the method proposed by the present invention is superior to the state-of-the-art methods on the GoPro dataset, REVD dataset, HS-ERGB dataset, and REBlur dataset.
[0045] 2. The present invention designs two specific attention mechanisms to extract event and image features respectively. First, through the rich high-temporal-resolution motion information in the event, the present invention designs an adaptive motion mask predictor to obtain spatial motion information. To cope with the sparse characteristic of the event in space, the motion sparse attention uses a motion mask to shield the influence of irrelevant regions in the event; to cope with the dense characteristic of the image in space, the motion perception attention uses a motion mask to highlight the influence of motion-related regions in the image, thereby solving the problem of the lack of motion region perception ability of the attention mechanism.
[0046] 3. The present invention designs two specific gating mechanisms. The extended control space gating enables the gradient of the adaptive motion mask predictor to be differentiable, and at the same time, this gating modulates the local dilation brought by the convolutional operations in the network; the cross-modal intensity gating realizes the efficient fusion of event features and image features. Through the modeling of the gating mechanism, the global information of the input features is deeply mined, thereby improving the defocusing performance of the image and increasing the interpretability of the model.
[0047] 4. The motion adaptive Transformer defocusing network designed by the present invention is constructed based on motion sparse event blocks and motion-aware image blocks. Among them, the motion sparse event block consists of motion sparse attention and extended control space gating; the motion-aware image block consists of motion-aware attention and cross-modal intensity gating. The network is trained in an end-to-end manner. This method breaks the limited receptive field of the traditional CNN network, and at the same time, the Transformer constructed for images and events makes more full use of the important motion information in events and images, thereby realizing more accurate motion blur modeling and achieving better defocusing effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flowchart of the inventive method;
[0049] Figure 2 is a structural diagram of the motion adaptive Transformer defocusing network method of the present invention;
[0050] Figure 3 is a structural diagram of the adaptive mask predictor in the present invention;
[0051] Figure 4 is a structural diagram of the motion sparse event block in the present invention;
[0052] Figure 5 is a structural diagram of the motion-aware image block in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In this embodiment, a video defocusing method based on event data driving, the specific process is shown in Figure 1 , which comprehensively considers the characteristics of event sparsity and image density, and realizes two specific feature extraction methods for the two types of data by designing a motion adaptive attention mechanism, and then realizes the fusion of the two features through the design of the gating mechanism to achieve the defocusing effect. The algorithm structure diagram of the whole method is shown in Figure 2 . Specifically, the method is carried out according to the following steps:
[0054] Step 1. Obtain a set of blurred images, denoted as , where represents the th blurred image, , is the number of blurred images;
[0055] Obtain a set of blurred images The corresponding event sequence, denoted as , indicating the th blurred image corresponding event sequence;
[0056] Obtain a set of clear images, denoted as , where indicates the th clear image;
[0057] In this embodiment, the GoPro dataset, REVD dataset, HS-ERGB dataset, and REBlur dataset are used to train and evaluate the model. Among them, the GoPro dataset and HS-ERGB dataset are synthetic datasets, and the REVD dataset and REBlur dataset are real-world datasets.
[0058] Step 2. In this embodiment, as Figure 2 shown, construct an event-driven motion adaptive Transformer deblurring network, including: a feature extraction layer, a layer of event-image encoder, and an image reconstruction decoder; in this embodiment, ;
[0059] Step 2.1. The feature extraction layer consists of an image feature extraction layer and an event feature extraction layer, and respectively extracts features from the th blurred image and the th event sequence to obtain the th initial blurred image feature and the th initial event feature .
[0060] Step 2.2. The layer of event-image encoder consists of 1 adaptive motion mask predictor, 1 motion sparse event block, and 1 motion-aware image block;
[0061] When , use a matrix of all 1s as the th initial motion mask of the th layer and process it in the input adaptive motion mask predictor to obtain the th predicted motion mask of the th layer th ;
[0062] The layer's th predicted motion mask is processed with the input motion sparse event block to generate the layer's th event feature ;
[0063] The layer's th predicted motion mask , and is processed with the layer's th input motion perception image block to generate the layer's
[0064] When it is, using the layer's th predicted motion mask as the layer's th initial motion mask and processing it with the th event feature output by the layer event-image encoder in the input adaptive motion mask predictor to obtain the layer's ;
[0065] The layer's th predicted motion mask and the layer's th event feature are processed with the layer's th input motion sparse event block to generate the layer's
[0066] The layer's th predicted motion mask , the layer's th event feature and the layer's th blurred image feature are processed with the The th layer of blurred image features ; thus, the th layer event-image encoder outputs the th layer of blurred image features .
[0067] Step 2.2.1. The adaptive motion mask predictor includes: a response calculation module, a response fusion unit, an adaptive adjustment unit, and a mask generation unit; the specific structure of the adaptive motion mask predictor is as Figure 3 shown;
[0068] When , the response calculation module performs a linear projection on , then inputs it into the GELU activation function for processing, generating the th layer of local responses , and at the same time, based on , performs a normalized weighted average on , generating the th layer of global responses ;
[0069] The response fusion unit expands the dimension of to align it with in space, then concatenates the expanded global response with , generating the th layer of first fusion features , and at the same time, multiplies the expanded global response element-wise with , generating the th layer of second fusion features ;
[0070] The adaptive adjustment unit calculates the th average sparsity rate , then sets the th layer of learnable expansion factor , thus using Equation (1) to calculate the th layer of high-response token count :
[0071] (1)
[0072] In Equation (1), is the training stability coefficient, is the total number of spatial tokens in the th layer and the th space; in this embodiment, ;
[0073] The mask generation unit processes using a multi-layer perceptron to obtain spatial projection features. At the same time, after performing max-pooling and average-pooling operations on the channel dimension of , it is then input into a linear layer and a GELU activation function for processing to obtain channel projection features. Thus, the spatial projection features and the channel projection features are concatenated along the channel and then input into another linear layer and a GELU activation function for processing to obtain the th layer and the th spatial motion score ; finally, according to , the top high-response positions are selected to generate a predicted motion mask .
[0074] Step 2.2.2. The motion sparse event block includes: a motion sparse attention unit and an extended control space gating unit; the specific structure of the motion sparse event block is as shown in Figure 4 ;
[0075] When , the motion sparse attention unit performs layer normalization on , and then successively processes it through a convolutional layer and a depthwise separable convolutional layer to obtain event local projection features. An operation of channel partitioning is performed on the event local projection features to generate the th layer and the th event query matrix , the th event key matrix , and the th event value matrix ; thus, the th layer and the th event motion sparse transposed attention map is calculated using Equation (2). is multiplied by to generate the th layer and the th updated event feature :
[0076] (2)
[0077] In Equation (2), is a learnable scaling factor, Softmax represents an activation function, represents element-wise multiplication, denotes transpose;
[0078] The extended control space gating unit will be processed in the input GELU activation function to obtain the spatial modulation fraction. At the same time, the is converted into the hybrid transmission rate, so that the spatial modulation fraction, the hybrid transmission rate, and are element-wise multiplied and then added to to obtain the th event feature .
[0079] Step 2.2.3. The motion-aware image block consists of a motion-aware attention unit and a cross-modal intensity gating unit; the specific structure of the motion-aware image block is as Figure 5 shown;
[0080] When , the motion-aware attention unit performs layer normalization on , and after being processed by a convolutional layer and a depthwise separable convolutional layer in sequence, the local image projection feature is obtained. The local image projection feature is subjected to a channel partitioning operation to generate the th image query matrix , the th image key matrix , and the th image value matrix . Thus, the th image motion sparse transposed attention map is calculated using Equation (3). After multiplying with , the th updated image feature is generated:
[0081] (3)
[0082] In Equation (2), is the learnable scaling factor, and is the learnable motion factor, which learns to dynamically adjust the contribution intensity of the motion region to the attention weight.
[0083] The cross-modal intensity gating unit processes in the input layer normalization layer to obtain the normalized image feature. At the same time, is processed in the input layer normalization layer and the GELU activation function to obtain the event modulation feature. Thus, after element-wise multiplying the normalized image feature and the event modulation feature, and then Add them up to obtain the image features after cross-modal fusion ;
[0084] Step 2.3. After the image reconstruction decoder processes it, the th predicted clear image is obtained, thus obtaining the set of predicted clear images .
[0085] Step 3. Use Equation to construct the loss function for backpropagation :
[0086] (4)
[0087] Step 4. Based on , and , train the motion adaptive Transformer deblurring network, calculate the loss function , and at the same time use the adaptive moment estimation optimization method with the learning rate to update the network weights. In this example, the learning rate is taken as 2e-4. When the number of training iterations reaches the set number or the loss error is less than the set threshold, the training stops, thereby obtaining the optimal deblurring model for processing blurred images to obtain the corresponding clear images.
[0088] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0089] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium. When the computer program is run by a processor, it executes the steps of the above method.
[0090] Embodiment
[0091] To verify the effectiveness of the method of the present invention, in this embodiment, the commonly used GoPro dataset, HS-ERGB dataset, REBlur dataset, and REVD dataset are selected for training and testing.
[0092] The method is trained based on the training set of the GoPro dataset, fine-tuned on the training sets of the other three datasets, and the metric evaluation is performed on the test sets of each dataset respectively.
[0093] In the present invention, the structural similarity (PSNR) and the peak signal-to-noise ratio (SSIM) are used as evaluation metrics.
[0094] In this embodiment, six methods and the method of the present invention are selected for effect comparison in the GoPro dataset, HS-ERGB dataset, and REBlur dataset. The selected methods are UFPNet, EFNet, EIFNet, MAENet, Restormer, FFTFormer, and MAT is the method of the present invention; in the REVD dataset, six methods and the method of the present invention are selected for effect comparison. The selected methods are REDNet, UEVD, FEVD, EFNet, EIFNet, MAENet, and MAT is the method of the present invention.
[0095] In Tables 1, 2, 3, and 4, EFNet refers to a cross-modal fusion event-driven image deblurring method, EIFNet refers to a modality-aware event-driven image deblurring method, and MAENet refers to a motion-aware event-driven image deblurring method.
[0096] In Tables 1, 3, and 4, UFPNet refers to an image deblurring method based on self-supervised kernel estimation, Restormer refers to an image deblurring method based on Transformer, and FFTFormer refers to an image deblurring method based on frequency-domain aware Transformer.
[0097] In Table 2, REDNet refers to an event-driven image deblurring method in a real environment, UEVD refers to an event-driven deblurring method with unknown image exposure duration, and FEVD refers to an event-driven deblurring method based on a frequency-domain aware network.
[0098] According to the experimental results, the results are shown in Tables 1, 2, 3, and 4:
[0099] Table Experimental results of deblurring of the method of the present invention and the six selected comparison methods on the GoPro dataset
[0100]
[0101] Table Experimental results of deblurring of the method of the present invention and the six selected comparison methods on the REVD dataset
[0102]
[0103] Table 3 Experimental results of deblurring of the method of the present invention and the six selected comparison methods on the HS-ERGB dataset
[0104]
[0105] Table 4 Experimental results of deblurring the proposed method of the present invention and six selected comparison methods on the REBlur dataset
[0106] The experimental results show that on four different datasets, the proposed method of the present invention has better effects than other methods, thus proving the feasibility of the proposed method of the present invention. The experiments show that the proposed method of the present invention can achieve more effective attention calculation according to the sparse characteristics of events and the dense characteristics of images, and the gating mechanism realizes efficient feature transmission, so as to achieve excellent performance in the task of deblurring blurred images.
Claims
1. An image deblurring method based on event data drive, characterized in that: The steps are as follows: Step 1: Get the fuzzy image set, denoted as ,in, Indicates A blurred image, , is the number of blurred images; Get the blurred image set The corresponding event sequence is denoted as , Indicates Blurred image The corresponding sequence of events; Get a clear image set, denoted as ,in, Indicates A clear image; Step 2: Build an event-driven motion adaptive Transformer deblurring network, including: feature extraction layer, Layer events - image encoder, image reconstruction decoder; Step 2.1, the feature extraction layer is composed of an image feature extraction layer and an event feature extraction layer, and the feature extraction layer is respectively Blurred image and Event sequence Perform feature extraction and obtain the corresponding Initial blurred image features and Initial event characteristics ; Step 2.2, The layer event-image encoder consists of an adaptive motion mask predictor, a motion sparse event block, and a motion-aware image block; when When all 1 matrices are used as the first Tier Initial motion mask and with Input into the adaptive motion mask predictor for processing and obtain the Tier Predicted motion mask ; The said Tier Predicted motion mask and Input the motion sparse event block for processing to generate the first Tier Event characteristics ; The said Tier Predicted motion mask , and The motion perception image block is input for processing to generate the first Tier Fuzzy image features ; when At the time, Tier Predicted motion mask As the Tier Initial motion mask And with the Layer event - the output of the image encoder Event characteristics Input into the adaptive motion mask predictor for processing and obtain the Tier Predicted motion mask ; The said Tier Predicted motion mask With Layer event - the output of the image encoder Event characteristics Input the motion sparse event block for processing to generate the first Tier Event characteristics ; The said Tier Predicted motion mask , No. Tier Event characteristics With Tier Fuzzy image features The motion perception image block is input for processing to generate the first Tier Fuzzy image features ; Thus, by Layer event - image encoder output Tier Fuzzy image features ; Step 2.3: The image reconstruction decoder After processing, we get A clear picture of the prediction , thus obtaining the predicted clear image set ; Step 3: Utilize Constructing the loss function for backpropagation : (4) Step 4: Based on , and , train the motion adaptive Transformer deblurring network and calculate the loss function , and use adaptive moment estimation optimization method with learning rate To update the network weights, when the number of training iterations reaches the set number or the loss error is less than the set threshold, the training stops, thereby obtaining the optimal deblurring model, which is used to process the blurred image to obtain the corresponding clear image.
2. The image deblurring method based on event data drive according to claim 1, characterized in that: The adaptive motion mask predictor comprises: a response calculation module, a response fusion unit, an adaptive adjustment unit and a mask generation unit; when When After linear projection, it is input into the GELU activation function for processing to generate the first Tier Local Response , and based on right Perform normalized weighted averaging to generate the Tier Global Response ; The response fusion unit is Expand the dimension to match In spatial alignment, the expanded global response is then compared with Splice to generate Tier First fusion feature , and the expanded global response is combined with Perform element-by-element multiplication to generate Tier Second fusion feature ; The adaptive adjustment unit is based on , calculate the Average sparsity rate , and then set the Tier Expansion factors to be learned , and then use formula (1) to calculate the Tier High response tokens : (1) In formula (1), is the training stability coefficient, For the Tier Total number of tokens in a space; The mask generation unit uses a multi-layer perceptron to Processing is performed to obtain spatial projection features. At the same time, After the maximum pooling and average pooling operations are performed on the channel dimension, the spatial projection features and the channel projection features are input into the linear layer and the GELU activation function for processing to obtain the channel projection features. Then, the spatial projection features and the channel projection features are concatenated along the channel and input into another linear layer and the GELU activation function for processing to obtain the first Tier Spatial motion score Finally, according to , before selecting The predicted motion mask is generated from the high response locations .
3. The image deblurring method based on event data drive according to claim 2, characterized in that: The motion sparse event block includes: a motion sparse attention unit and an extended control space gating unit; when When the motion sparse attention unit is After layer normalization, the event local projection features are obtained by sequentially processing through convolutional layers and depth-separable convolutional layers, and the event local projection features are divided into channels to generate the first Tier Event query matrix , No. Event Key Matrix , No. Event value matrix ; Then use formula (2) to calculate the Tier Event motion sparse transposed attention map ,Will and After multiplication, the first Tier Updated event features : (2) In formula (2), is the scaling factor to be learned, Softmax represents the activation function, represents element-wise multiplication, represents transpose; The extended control space gating unit will Input into GELU activation function for processing to obtain spatial modulation score, and use Sigmoid function to convert Convert to mixed transmission rate, thus converting spatial modulation fraction, mixed transmission rate and After element-by-element multiplication, Add together and get Tier Event characteristics .
4. The image deblurring method based on event data drive according to claim 2, characterized in that: The motion perception image block comprises: a motion perception attention unit and a cross-modal intensity gating unit; when When the motion perception attention unit The local projection features of the image are obtained by layer normalization, and the local projection features of the image are processed by the convolution layer and the depth-separable convolution layer in turn. The channel division operation is performed on the local projection features of the image to generate the first Tier Image query matrix , No. Image key matrix , No. Image value matrix , and then use formula (3) to calculate the Tier Image motion sparse transposed attention map , and After multiplication, the first Tier Updated image features : (3) In formula (3), is the motion factor to be studied; The cross-modal intensity gating unit will The input layer is processed by the normalization layer to obtain the standardized image features. The input layer normalization layer is processed with the GELU activation function to obtain the event modulation feature, so that the normalized image feature and the event modulation feature are multiplied element by element and then Add together and get Tier Image features after cross-modal fusion .
5. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the image deblurring method according to any one of claims 1 to 4, and the processor is configured to execute the program stored in the memory.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image deblurring method according to any one of claims 1 to 4 are executed.
Citation Information
Patent Citations
Image blind motion deblurring method based on CNN-Transform hybrid auto-encoder
CN113570516A
Video deblurring method based on event data driving
CN114463218A
Transform-based moving image deblurring method
CN115496676A
Image deblurring method based on event guidance
CN117726549A
Image deblurring method based on direction perception Transform
CN118333897A