Moving image deblurring method based on two-way imaging system and event double integration
By constructing a dual-path imaging system and an event-based dual-integration self-supervised learning framework, the motion blur problem is solved, achieving high-quality image deblurring in high-speed motion scenes. It adapts to training with real-world scene data, has high temporal and spatial registration accuracy, and is suitable for high-speed, high dynamic range scenes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies struggle to effectively address motion blur, especially in high-speed or nonlinear motion scenarios. Traditional algorithms are computationally complex and rely on prior assumptions about the blur kernel. Meanwhile, deep learning methods based on frame images struggle to recover texture details in real-world scenes, and the lack of effective datasets and clear image annotations limits the generalization ability of these algorithms.
A dual-path imaging system consisting of an event camera and a traditional camera is constructed. The beam path is split by a beam splitter prism, and high-precision time synchronization and sub-pixel-level spatial registration are achieved by combining FPGA. A two-stage self-supervised learning network based on event double integration is designed. Self-supervised learning is performed using blur-event loss, optical flow-event loss and blur-sharp loss to output sharp images.
It achieves high-quality image deblurring in high-speed motion scenes, adapts to real-world scene data training, requires no clear image annotation, has high temporal and spatial registration accuracy, and demonstrates superior deblurring performance.
Smart Images

Figure CN121865100A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing and sensor fusion technology, and in particular to a motion image deblurring method based on a dual-path imaging system and event dual integration. Background Technology
[0002] Motion blur is a long-standing and difficult-to-avoid problem in image acquisition. Traditional frame cameras generate images by integrating photons over an exposure period. If there is relative motion between the camera and the scene during the exposure, or if the subject moves rapidly, the imaging point on the sensor will shift, resulting in motion blur and trailing in the final output image. This blur not only reduces the visual quality of the image but also seriously affects the progress of subsequent visual tasks.
[0003] To address motion blur, existing technologies mainly fall into two categories: traditional algorithms based on blind deconvolution and deblurring networks based on deep learning. However, traditional algorithms often have high computational complexity and rely on prior assumptions about the blur kernel; while deep learning methods based on pure frame images have improved performance, they struggle to recover realistic texture details in high-speed or non-linear motion scenarios due to the loss of inter-frame information.
[0004] In recent years, event cameras, as a novel biomimetic asynchronous sensor, have attracted widespread attention. Unlike traditional cameras, event cameras do not output images at a fixed frame rate, but rather asynchronously detect brightness changes at each pixel. When the logarithmic change in light intensity of a pixel exceeds a set threshold, the event camera outputs an event signal containing coordinates, a timestamp, and polarity. Event cameras possess extremely high microsecond-level temporal resolution, high dynamic range, and low latency, enabling them to accurately record the trajectory of brightness changes during motion, providing a new approach to solving motion blur problems.
[0005] However, there are currently few publicly available event camera datasets, and most of these datasets use simulated event streams generated by event simulators, which greatly affects the effectiveness of network training. On the other hand, many supervised learning methods require pairs of "blurred-sharp" images as training sets, but obtaining perfectly sharp ground truth images in real-world scenes is extremely difficult, limiting the generalization ability of the algorithms.
[0006] Therefore, there is a need for a deblurring system and method that can accurately synchronize and register dual-channel camera data and utilize the physical characteristics of event streams for self-supervised learning. Summary of the Invention
[0007] The purpose of this invention is to provide a motion image deblurring method based on a dual-path imaging system and event-based dual integration. By constructing a hardware-level synchronous dual-path acquisition system, high-precision temporal and spatial registration is achieved. Furthermore, a self-supervised learning framework based on the event-based dual integration model is proposed, which improves image clarity in high-speed motion scenes.
[0008] The technical solution to achieve the purpose of this invention is: a motion image deblurring method based on a dual-path imaging system and event dual integration, comprising the following steps:
[0009] Step 1: Construct a dual-path imaging system consisting of an event camera and a traditional camera. Use a beam splitter to split the scene's light path into two paths, which are then input into the event camera and the traditional camera, respectively.
[0010] Step 2: Use FPGA to implement high-precision time synchronization trigger control of the two cameras, and complete sub-pixel level spatial registration based on homography transformation;
[0011] Step 3: Design a two-stage self-supervised learning network based on event double integration. In the first stage, event double integration features are learned from blurred images and event streams through a convolutional neural network. In the second stage, dual-path features are fused through an encoder-decoder structure to output a clear image.
[0012] Step 4: Introduce blur-event loss, optical flow-event loss, and blur-sharpness loss as self-supervised signals to constrain brightness, motion, and reconstruction consistency;
[0013] Step 5: Train the network in two stages on both the simulated and real datasets, and optimize the model using SGDR learning rate scheduling.
[0014] Step 6: Input the blurred image and event stream, process them through the network, and output the reconstructed clear image.
[0015] Furthermore, the construction of the dual-path imaging system consisting of an event camera and a traditional camera described in step 1 includes a dual-path imaging module, a temporal registration module, an integrated control module, a spatial registration module, an image deblurring network module, and a self-supervised training module.
[0016] The dual-path imaging module includes an event camera and a conventional camera coaxially configured via a beam splitter. The beam splitter is used to simultaneously acquire event streams and image frames of the same scene.
[0017] The time registration module generates a synchronization trigger signal based on the FPGA to control the time alignment between the event camera and the traditional camera during the exposure time.
[0018] The integrated control module includes an FPGA controller and a serial communication unit, which are used to receive instructions from the host computer and configure parameters for the traditional camera and the event camera, as well as send hardware trigger signals at the same time.
[0019] The spatial registration module achieves sub-pixel-level spatial alignment between the event stream and the blurred image based on homography transformation;
[0020] The image deblurring network module includes an event dual-integral network and an image fusion network, used to reconstruct a clear image from a blurred image and an event stream;
[0021] The self-supervised training module is designed with fuzzy-event loss, optical flow-event loss, and fuzzy-clear loss to achieve unsupervised learning.
[0022] Furthermore, the event camera is a dynamic vision sensor (DVS), the conventional camera is an industrial area array camera, and the beam splitter is a non-polarized cubic beam splitter.
[0023] Furthermore, the time registration module communicates with the host computer through the UART interface of the FPGA. After receiving the control command, the FPGA generates a multi-cycle pulse signal to trigger the event camera and the traditional camera respectively.
[0024] Furthermore, the integrated control module functions as follows: activating the event camera and the traditional camera respectively; enumerating available serial ports through the Serial Port Manager and establishing a connection with the FPGA; configuring the acquisition parameters and selecting the mode for the event camera and the traditional camera respectively; and sending control commands to the FPGA.
[0025] Furthermore, step 2 involves using an FPGA to implement high-precision time synchronization trigger control for the two cameras and completing sub-pixel-level spatial registration based on homography transformation, as detailed below:
[0026] Step 2.1: Drive the display device to display an alternating flashing black and white checkerboard pattern as a common reference plane;
[0027] Step 2.2: Control the conventional camera to capture an image of the pattern, and simultaneously control the event camera to record a synchronized event stream;
[0028] Step 2.3: Extract sub-pixel precision corner coordinates from traditional camera images using Harris corner detection. ;
[0029] Step 2.4: Accumulate the event stream of the event camera within a specific time window into accumulated frame images, and extract the corresponding corner coordinates from the accumulated frames. ;
[0030] Step 2.5: Select at least 4 sets of corresponding corner points, based on the homography transformation model. Construct an overdetermined system of equations;
[0031] Step 2.6: Solve the homography matrix H using the Direct Linear Transform (DLT) algorithm combined with Singular Value Decomposition (SVD) to complete spatial registration.
[0032] Furthermore, the homography transformation model described in step 2.5 is specifically expressed as follows:
[0033]
[0034] Will After expansion, the specific coordinate transformation equations are obtained:
[0035]
[0036] in, The coordinates of the event camera. For coordinates of a traditional camera.
[0037] Furthermore, the design described in step 3 is based on a two-stage self-supervised learning network with event dual integration. The first stage learns event dual integration features from the blurred image and event stream through a convolutional neural network. The second stage fuses the dual features through an encoder-decoder structure and outputs a clear image, as detailed below:
[0038] Step 3.1: Obtain the relationship between the blurred image, the potentially sharp image, and the event flow from the perspective of events, i.e., event double integration. Construct an event double integration network to approximate the event double integration process. The expression for event double integration is as follows:
[0039]
[0040] in, Indicates the start time point, corresponding to the clear image. The moment; The integral time length represents the time window for event acquisition; This is the threshold for triggering the event; It is the sequence of events captured by the event camera, which is a function that changes over time; Indicates the time corresponding to the clear image up to the current time The cumulative number of events; This represents an exponential function used to convert cumulative event data into an intensity or brightness representation.
[0041] Step 3.2: Construct an image fusion network containing a convolutional attention mechanism (CBAM). Input the blurred image and the corresponding double integral features, and output the clear image after deblurring.
[0042] Furthermore, step 4 introduces blur-event loss, optical flow-event loss, and blur-sharpness loss as self-supervised signals to constrain brightness, motion, and reconstruction consistency, as detailed below:
[0043] Step 4.1, Fuzzy-Event Loss Function The estimation of the event double integral is constrained by the brightness consistency between consecutive blurred images, and the calculation formula is as follows:
[0044]
[0045] in, and These are the left and right blurred images, respectively. and These represent the corresponding double integral approximation features of the events. It is a constant introduced to ensure numerical stability, which is less than a set threshold;
[0046] Step 4.2, Event-Optical Flow Loss Function The consistency between event data and image optical flow in describing physical motion is used to force the two motion fields to remain consistent. The calculation formula is as follows:
[0047]
[0048] in, and These are the length and width of the image, respectively. This represents the optical flow estimates of the two initially reconstructed left and right images under the RAFT-Sparse optical flow model. Indicates the image within the exposure time. gradient, This indicates the trigger threshold for the event camera. Represents an event flow;
[0049] Step 4.3, Blur-Clear Loss Function The final reconstructed clear frame is blurred and compared with the blurred images of the left and right sides to calculate the reconstruction consistency loss. The calculation formula is as follows:
[0050]
[0051] in, and These represent the exposure ranges for the left and right blurred images, respectively. For reconstructed clear frames;
[0052] Step 4.4, Total Loss Function The formula is as follows:
[0053]
[0054] in, , and These are the weights of the corresponding loss functions.
[0055] Furthermore, in the two-stage training described in step 5, the first stage trains the event dual-integral network, and the second stage trains the image fusion network. The loss weights are adjusted with each stage, and the learning rate adopts a cosine annealing and periodic restart strategy.
[0056] Compared with the prior art, the present invention has the following significant advantages: (1) It combines the dynamic information of the event camera with the static texture of the traditional camera to achieve high-quality image deblurring; (2) The self-supervised learning framework does not require clear image annotation and is suitable for training on real scene data; (3) The system has high temporal and spatial registration accuracy and is suitable for high-speed and high dynamic range scenes; (4) It exhibits superior deblurring performance in both public datasets and real scenes. Attached Figure Description
[0057] Figure 1 This is a flowchart of the motion image deblurring method based on a dual-path imaging system and event-based dual integration, as described in this invention.
[0058] Figure 2 This is a physical schematic diagram of the dual-path imaging system in an embodiment of the present invention.
[0059] Figure 3 This is a structural block diagram of the self-supervised learning network in an embodiment of the present invention.
[0060] Figure 4 This is a schematic diagram of the structure of the EDI network in an embodiment of the present invention.
[0061] Figure 5 This is a schematic diagram of the Fusion network structure in an embodiment of the present invention.
[0062] Figure 6 This is a qualitative result diagram based on a real dataset in an embodiment of the present invention. Detailed Implementation
[0063] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0064] like Figure 1 As shown, the present invention provides a motion image deblurring method based on a dual-path imaging system and event-based dual integration, comprising the following steps:
[0065] Step 1: Construct a dual-path imaging system consisting of an event camera and a traditional camera. Use a beam splitter to split the scene's light path into two paths, which are then input into the event camera and the traditional camera, respectively.
[0066] The dual-path imaging system includes a dual-path imaging module, a temporal registration module, an integrated control module, a spatial registration module, an image deblurring network module, and a self-supervised training module.
[0067] The dual-path imaging module includes an event camera and a conventional camera coaxially configured via a beam splitter. The beam splitter is used to simultaneously acquire event streams and image frames of the same scene.
[0068] The time registration module generates a synchronization trigger signal based on the FPGA to control the time alignment between the event camera and the traditional camera during the exposure time.
[0069] The integrated control module includes an FPGA controller and a serial communication unit, which are used to receive instructions from the host computer and configure parameters for the traditional camera and the event camera, as well as send hardware trigger signals at the same time.
[0070] The spatial registration module achieves sub-pixel-level spatial alignment between the event stream and the blurred image based on homography transformation;
[0071] The image deblurring network module includes an event dual-integral network and an image fusion network, used to reconstruct a clear image from a blurred image and an event stream;
[0072] The self-supervised training module is designed with fuzzy-event loss, optical flow-event loss, and fuzzy-clear loss to achieve unsupervised learning.
[0073] As a specific example, the event camera is a dynamic vision sensor (DVS), the conventional camera is an industrial area array camera, and the beam splitter is a non-polarized cubic beam splitter.
[0074] As a specific example, the time registration module communicates with the host computer through the UART interface of the FPGA. After receiving the control command, the FPGA generates multi-cycle pulse signals to trigger the event camera and the traditional camera respectively.
[0075] As a specific example, the main functions of the integrated control module are: to start the event camera and the traditional camera respectively; to enumerate the available serial ports through the Serial Port Manager and establish a connection with the FPGA; to configure the acquisition parameters and select the mode for the event camera and the traditional camera respectively; and to send control commands to the FPGA.
[0076] Step 2: Implement high-precision time synchronization trigger control for the two cameras using FPGA, and complete sub-pixel-level spatial registration based on homography transformation, as detailed below:
[0077] Step 2.1: Drive the display device to display an alternating flashing black and white checkerboard pattern as a common reference plane;
[0078] Step 2.2: Control the conventional camera to capture an image of the pattern, and simultaneously control the event camera to record a synchronized event stream;
[0079] Step 2.3: Extract sub-pixel precision corner coordinates from traditional camera images using Harris corner detection. ;
[0080] Step 2.4: Accumulate the event stream of the event camera within a specific time window into accumulated frame images, and extract the corresponding corner coordinates from the accumulated frames. ;
[0081] Step 2.5: Select at least 4 sets of corresponding corner points, based on the homography transformation model. Construct an overdetermined system of equations;
[0082] The homography transformation model is specifically expressed as follows:
[0083]
[0084] Will After unfolding, the specific coordinate transformation equations can be obtained:
[0085]
[0086] in, The coordinates of the event camera. For coordinates of a traditional camera.
[0087] Step 2.6: Solve the homography matrix H using the Direct Linear Transform (DLT) algorithm combined with Singular Value Decomposition (SVD) to complete spatial registration.
[0088] Step 3: Design a two-stage self-supervised learning network based on event double integration. The first stage learns event double integration features from the blurred image and event stream through a convolutional neural network. The second stage fuses the dual-path features through an encoder-decoder structure and outputs a clear image, as detailed below:
[0089] Step 3.1: Obtain the relationship between the blurred image, the potentially sharp image, and the event flow from the perspective of events, i.e., event double integration. Construct an event double integration network to approximate the event double integration process. The expression for event double integration is as follows:
[0090]
[0091] in, Indicates the start time point, corresponding to the clear image. The moment; The integral time length represents the time window for event acquisition; This is the threshold for triggering the event; It is the sequence of events captured by the event camera, which is a function that changes over time; Indicates the time corresponding to the clear image up to the current time The cumulative number of events; This represents an exponential function used to convert cumulative event data into an intensity or brightness representation.
[0092] Step 3.2: Construct an image fusion network containing a convolutional attention mechanism (CBAM). Input the blurred image and the corresponding double integral features, and output the clear image after deblurring.
[0093] Step 4: Introduce blur-event loss, optical flow-event loss, and blur-sharpness loss as self-supervised signals to constrain brightness, motion, and reconstruction consistency, as detailed below:
[0094] Step 4.1: The fuzzy-event loss function uses the brightness consistency between consecutively blurred images to constrain the estimation of the event double integral, thereby overcoming the problems caused by event noise and threshold instability. Its calculation formula is as follows:
[0095]
[0096] in, and These are the left and right blurred images, respectively. and These represent the corresponding double integral approximation features of the events. It is a minimal constant introduced to ensure numerical stability;
[0097] Step 4.2: The event-optical flow loss function uses the consistency between event data and image optical flow in describing physical motion to force the two motion fields to remain consistent, thereby improving the accuracy of motion estimation. Its calculation formula is as follows:
[0098]
[0099] in, and These are the length and width of the image, respectively. This represents the optical flow estimates of the two initially reconstructed left and right images under the RAFT-Sparse optical flow model. Indicates the image within the exposure time. gradient, This indicates the trigger threshold for the event camera. Represents an event flow;
[0100] Step 4.3: The blur-sharp loss function is calculated by blurring the final reconstructed sharp frame and comparing it with the blurred frames on the left and right sides. This loss is used to determine the reconstruction consistency loss. The calculation formula is as follows:
[0101]
[0102] in, and These represent the exposure ranges for the left and right blurred images, respectively. For reconstructed clear frames;
[0103] Step 4.4, the total loss function is as follows:
[0104]
[0105] in, , and These are the weights of the corresponding loss functions.
[0106] Step 5: Train the network in two stages on both the simulated and real datasets, and optimize the model using SGDR learning rate scheduling.
[0107] The first stage trains the event-based dual-integral network, and the second stage trains the image fusion network. The loss weights are adjusted with each stage, and the learning rate adopts a cosine annealing and periodic restart strategy.
[0108] Step 6: Input the blurred image and event stream, process them through the network, and output the reconstructed clear image.
[0109] Example
[0110] like Figure 1 As shown in the figure, this embodiment of a motion image deblurring method based on a dual-path imaging system and event dual integration includes the following steps:
[0111] Step 1: Construct a dual-path imaging system consisting of an event camera and a traditional camera. Use a beam splitter to split the scene's light path into two paths, which are then input into the event camera and the traditional camera, respectively.
[0112] The dual-path imaging system includes a dual-path imaging module, a temporal registration module, an integrated control module, a spatial registration module, an image deblurring network module, and a self-supervised training module.
[0113] The dual-path imaging module includes an event camera and a conventional camera coaxially configured via a beam splitter. The beam splitter is used to simultaneously acquire event streams and image frames of the same scene.
[0114] The time registration module generates a synchronization trigger signal based on the FPGA to control the time alignment between the event camera and the traditional camera during the exposure time.
[0115] The integrated control module includes an FPGA controller and a serial communication unit, which are used to receive instructions from the host computer and configure parameters for the traditional camera and the event camera, as well as send hardware trigger signals at the same time.
[0116] The spatial registration module achieves sub-pixel-level spatial alignment between the event stream and the blurred image based on homography transformation;
[0117] The image deblurring network module includes an event dual-integral network and an image fusion network, used to reconstruct a clear image from a blurred image and an event stream;
[0118] The self-supervised training module is designed with fuzzy-event loss, optical flow-event loss, and fuzzy-clear loss to achieve unsupervised learning.
[0119] As a specific example, such as Figure 2 As shown, the event camera is a Prophesee EVK4 HD dynamic vision sensor DVS, the conventional camera is a Daheng MER2-160-227U3M industrial area array camera, the beam splitter is a Soleb CM1-BP145B1 non-polarizing cubic beam splitter, and the synchronization circuit is an FPGA-based synchronization circuit.
[0120] As a specific example, the time registration module communicates with the host computer through the UART interface of the FPGA. After receiving the control command, the FPGA generates multi-cycle pulse signals to trigger the event camera and the traditional camera respectively.
[0121] As a specific example, the main functions of the integrated control module are: to start the event camera and the traditional camera respectively; to enumerate the available serial ports through the Serial Port Manager and establish a connection with the FPGA; to configure the acquisition parameters and select the mode for the event camera and the traditional camera respectively; and to send control commands to the FPGA.
[0122] Step 2: Implement high-precision time synchronization trigger control for the two cameras using FPGA, and complete sub-pixel-level spatial registration based on homography transformation, as detailed below:
[0123] Step 2.1: Drive the display device to display an alternating flashing black and white checkerboard pattern as a common reference plane;
[0124] Step 2.2: Control the conventional camera to capture an image of the pattern, and simultaneously control the event camera to record a synchronized event stream;
[0125] Step 2.3: Extract sub-pixel precision corner coordinates from traditional camera images using Harris corner detection. ;
[0126] Step 2.4: Accumulate the event stream of the event camera within a specific time window into accumulated frame images, and extract the corresponding corner coordinates from the accumulated frames. ;
[0127] Step 2.5: Select at least 4 sets of corresponding corner points, based on the homography transformation model. Construct an overdetermined system of equations;
[0128] The homography transformation model is specifically expressed as follows:
[0129]
[0130] Will After unfolding, the specific coordinate transformation equations can be obtained:
[0131]
[0132] Step 2.6: Solve the homography matrix H using the Direct Linear Transform (DLT) algorithm combined with Singular Value Decomposition (SVD) to complete spatial registration.
[0133] Step 3: Design a two-stage self-supervised learning network based on event double integration. The first stage learns event double integration features from the blurred image and event stream through a convolutional neural network. The second stage fuses the dual-path features through an encoder-decoder structure and outputs a clear image, as detailed below:
[0134] like Figure 3 As shown, the self-supervised learning network includes an event dual integral (EDI) module and an image fusion (Fusion) module. The EDI module takes a blurred image and an event voxel grid as input, extracts and fuses features through convolutional layers, and outputs an event dual integral approximation. The Fusion module adopts a U-Net structure, introduces residual connections and CBAM attention mechanism, deeply fuses the left and right preliminary reconstructed images, and outputs the final clear frame.
[0135] Step 3.1: Obtain the relationship between the blurred image, the potentially sharp image, and the event flow from the perspective of events, i.e., event double integration. Construct an event double integration network to approximate the event double integration process. The expression for event double integration is as follows:
[0136]
[0137] Combination Figure 4This is a structural diagram of the EDI network. The event stream undergoes preprocessing to generate event features with 2N channels, where N is the number of time bins; in this embodiment, N=8. Simultaneously, the blurred image undergoes feature transformation through two convolutional layers to obtain a feature map consistent with the number of event feature channels. Subsequently, the event features and image features are fused, and the fused features are then calculated through four convolutional layers. The first to third convolutional layers all incorporate normalization and ReLU activation functions, ultimately outputting a double-integral approximation of the event. The specific approximation is derived from the following formula:
[0138]
[0139]
[0140]
[0141]
[0142] in, and For blurred images Features of the blurred image after feature extraction For voxel pretreatment, This is an event information feature map. For ReLU activation function operations, and For the corresponding convolution operation, The combined convolutional operation after fusion includes the convolution, normalization, and ReLU activation functions in the first three layers, as well as the convolution in the last layer. This is a channel stitching operation. The final event double integral is obtained by comprehensively extracting features from the blurred image and the corresponding event information, playing a crucial role in image deblurring.
[0143] Step 3.2: Construct an image fusion network containing a convolutional attention mechanism (CBAM). Input the blurred image and the corresponding double integral features, and output the clear image after deblurring.
[0144] Combination Figure 5This diagram illustrates the Fusion network structure, which employs an encoder-decoder architecture. It consists of one input layer, two encoder layers, two residual blocks, two decoder layers, and one output layer. All layers include ReLU activation and batch normalization, while the final output layer uses the Sigmoid activation function. The network also utilizes skip connections; by connecting the output of the encoder layer to the input of the decoder layer, the network can effectively fuse multi-scale features and learn more texture details. An additional Convolutional Block Attention (CBAM) block is introduced in the middle of the network, combining channel attention and spatial attention to adaptively select important features. The specific calculation process for each feature is as follows:
[0145]
[0146]
[0147]
[0148]
[0149] in, The convolutional attention weights include channel attention and spatial attention mechanisms. Channel attention incorporates two adaptive computation modules: average pooling and max pooling in the spatial dimension, a multilayer perceptron (MLP), and a sigmoid function. This module first processes the input features... Perform average pooling and max pooling on the spatial dimension separately, pooling the features of each channel into a single value, resulting in two dimensions. The feature maps were then processed. Subsequently, a weight-shared MLP was used to learn channel-dimensional features and the importance of each channel, employing two convolutional kernels of size 3. A 3x3 convolution is used to reduce the dimensionality, and the activation function is ReLU. Finally, the two features output by the MLP are summed, and the Sigmoid activation function is used to map the features to the final channel attention weight matrix. The calculation process is shown in the following formula:
[0150]
[0151] in, Indicates channel attention weights. This represents the Sigmoid activation function. Average pooling in spatial dimension This represents max pooling in terms of spatial dimensions.
[0152] Spatial attention includes two adaptive pooling modules—average pooling and max pooling—along with the channel dimension, as well as a sigmoid activation function. This module sequentially performs max pooling and average pooling on the input features along the channel dimension, obtaining a single feature with the following dimensions. The features are then selected to facilitate subsequent feature learning. The pooled results are then concatenated using a convolution kernel of size 1. A convolution of 1 compresses the dimension of the concatenated features. Finally, the Sigmoid activation function maps the features to the final spatial attention weight matrix, the calculation process of which is shown in the following formula:
[0153]
[0154] in, Spatial attention weights, This represents average pooling along the channel dimension. This represents max pooling along the channel dimension, resulting in the final CBAM block output:
[0155]
[0156]
[0157] in, For channel attention, For spatial attention, which is the final output, the operator This is point-by-point multiplication.
[0158]
[0159] in, and These represent the first-level residual operation and the second-level residual operation, respectively.
[0160]
[0161]
[0162]
[0163] After the above process, the Fusion module outputs the final clear image. .
[0164] Step 4: Introduce blur-event loss, optical flow-event loss, and blur-sharpness loss as self-supervised signals to constrain brightness, motion, and reconstruction consistency, as detailed below:
[0165] Step 4.1: The fuzzy-event loss function uses the brightness consistency between consecutively blurred images to constrain the estimation of the event double integral, thereby overcoming the problems caused by event noise and threshold instability. Its calculation formula is as follows:
[0166]
[0167] in, and These are the left and right blurred images, respectively. and These represent the corresponding double integral approximation features of the events. It is a minimal constant introduced to ensure numerical stability;
[0168] Step 4.2: The event-optical flow loss function uses the consistency between event data and image optical flow in describing physical motion to force the two motion fields to remain consistent, thereby improving the accuracy of motion estimation. Its calculation formula is as follows:
[0169]
[0170] in, and These are the length and width of the image, respectively. This represents the optical flow estimates of the two initially reconstructed left and right images under the RAFT-Sparse optical flow model. Indicates the image within the exposure time. gradient, This indicates the trigger threshold for the event camera. Represents an event flow;
[0171] Step 4.3: The blur-sharp loss function is calculated by blurring the final reconstructed sharp frame and comparing it with the blurred frames on the left and right sides. This loss is used to determine the reconstruction consistency loss. The calculation formula is as follows:
[0172]
[0173] in, and These represent the exposure ranges for the left and right blurred images, respectively. For the reconstructed clear frame.
[0174] Step 4.4, the total loss function is as follows:
[0175]
[0176] in, , and These are the weights of the corresponding loss functions.
[0177] Step 5: Train the network in two stages on both the simulated and real datasets, and optimize the model using SGDR learning rate scheduling.
[0178] The first stage trains the event-based dual-integral network, and the second stage trains the image fusion network. The loss weights are adjusted with each stage, and the learning rate adopts a cosine annealing and periodic restart strategy.
[0179] Step 6: Input the blurred image and event stream, process them through the network, and output the reconstructed clear image.
[0180] The self-supervised learning network of this invention is actually divided into two stages: the first stage is the event double integral feature generation stage, and the second stage is the image fusion stage. Therefore, the network training strategy is also divided into two stages. The optimizer uses the Adam optimizer. and The learning rates were set to 0.9 and 0.999 respectively, and the learning strategy used SGDR scheduling combined with an optimization strategy of cosine annealing learning rate decay and periodic restart. For the first 100 epochs, the initial learning rate was 1e-3, focusing on event-based double integral deblurring training; therefore, the weights... Set to 256, Set to 1, Set to 1e-1; after the last 100 epochs, SGDR performs a "hot restart," resetting the initial learning rate to 5e-4 and simultaneously adjusting the weights. Reduced to 32, and This remains unchanged and is used for focused training of the fusion network.
[0181] Combination Figure 6 The results are qualitative findings of this invention on a real dataset. From left to right, they are the left blurred image, the right blurred image, and the finally reconstructed clear frame. It can be seen that the reconstructed clear frame has clear edges and no obvious artifacts, thus achieving the deblurring task well.
[0182] In summary, this invention proposes a motion image deblurring method based on a dual-path imaging system and event-based dual integration. The spatiotemporal consistency of multi-source data is ensured through FPGA synchronization and beam splitting path calibration using a homography transformation matrix. Furthermore, a self-supervised learning network based on physical model constraints is proposed, and an attention mechanism is introduced to optimize feature extraction. Experiments demonstrate that this method exhibits superior performance in image deblurring tasks and has significant application value.
[0183] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A motion image deblurring method based on a dual-path imaging system and event-based dual integration, characterized in that, Includes the following steps: Step 1: Construct a dual-path imaging system consisting of an event camera and a traditional camera. Use a beam splitter to split the scene's light path into two paths, which are then input into the event camera and the traditional camera, respectively. Step 2: Use FPGA to implement high-precision time synchronization trigger control of the two cameras, and complete sub-pixel level spatial registration based on homography transformation; Step 3: Design a two-stage self-supervised learning network based on event double integration. In the first stage, event double integration features are learned from blurred images and event streams through a convolutional neural network. In the second stage, dual-path features are fused through an encoder-decoder structure to output a clear image. Step 4: Introduce blur-event loss, optical flow-event loss, and blur-sharpness loss as self-supervised signals to constrain brightness, motion, and reconstruction consistency; Step 5: Train the network in two stages on both the simulated and real datasets, and optimize the model using SGDR learning rate scheduling. Step 6: Input the blurred image and event stream, process them through the network, and output the reconstructed clear image.
2. The motion image deblurring method based on a dual-path imaging system and event dual integration as described in claim 1, characterized in that, The construction of a dual-path imaging system consisting of an event camera and a traditional camera described in step 1 includes a dual-path imaging module, a temporal registration module, an integrated control module, a spatial registration module, an image deblurring network module, and a self-supervised training module. The dual-path imaging module includes an event camera and a conventional camera coaxially configured via a beam splitter. The beam splitter is used to simultaneously acquire event streams and image frames of the same scene. The time registration module generates a synchronization trigger signal based on the FPGA to control the time alignment between the event camera and the traditional camera during the exposure time. The integrated control module includes an FPGA controller and a serial communication unit, which are used to receive instructions from the host computer and configure parameters for the traditional camera and the event camera, as well as send hardware trigger signals at the same time. The spatial registration module achieves sub-pixel-level spatial alignment between the event stream and the blurred image based on homography transformation; The image deblurring network module includes an event dual-integral network and an image fusion network, used to reconstruct a clear image from a blurred image and an event stream; The self-supervised training module is designed with fuzzy-event loss, optical flow-event loss, and fuzzy-clear loss to achieve unsupervised learning.
3. The motion image deblurring method based on a dual-path imaging system and event dual integration according to claim 2, characterized in that, The event camera is a dynamic vision sensor (DVS), the conventional camera is an industrial area array camera, and the beam splitter is a non-polarized cubic beam splitter.
4. The motion image deblurring method based on a dual-path imaging system and event dual integration according to claim 2, characterized in that, The time registration module communicates with the host computer through the UART interface of the FPGA. After receiving the control command, the FPGA generates a multi-cycle pulse signal to trigger the event camera and the traditional camera respectively.
5. The motion image deblurring method based on a dual-path imaging system and event dual integration according to claim 2, characterized in that, The integrated control module functions as follows: it starts the event camera and the traditional camera respectively; it enumerates available serial ports through Serial PortManager and establishes a connection with the FPGA; it configures the acquisition parameters and selects the mode for the event camera and the traditional camera respectively; and it sends control commands to the FPGA.
6. The motion image deblurring method based on a dual-path imaging system and event dual integration according to claim 1, characterized in that, Step 2 describes the use of FPGA to implement high-precision time synchronization trigger control for the two cameras, and the completion of sub-pixel-level spatial registration based on homography transformation, as detailed below: Step 2.1: Drive the display device to display an alternating flashing black and white checkerboard pattern as a common reference plane; Step 2.2: Control the conventional camera to capture an image of the pattern, and simultaneously control the event camera to record a synchronized event stream; Step 2.3: Extract sub-pixel precision corner coordinates from traditional camera images using Harris corner detection. ; Step 2.4: Accumulate the event stream of the event camera within a specific time window into accumulated frame images, and extract the corresponding corner coordinates from the accumulated frames. ; Step 2.5: Select at least 4 sets of corresponding corner points, based on the homography transformation model. Construct an overdetermined system of equations; Step 2.6: Solve the homography matrix H using the Direct Linear Transform (DLT) algorithm combined with Singular Value Decomposition (SVD) to complete spatial registration.
7. The motion image deblurring method based on a dual-path imaging system and event dual integration according to claim 6, characterized in that, The homography transformation model described in step 2.5 is specifically expressed as follows: ; Will After expansion, the specific coordinate transformation equations are obtained: ; in, The coordinates of the event camera. For coordinates of a traditional camera.
8. The motion image deblurring method based on a dual-path imaging system and event dual integration according to claim 1, characterized in that, Step 3 describes a two-stage self-supervised learning network based on event-double integration. The first stage learns event-double integration features from the blurred image and event stream using a convolutional neural network. The second stage fuses the dual-path features through an encoder-decoder structure and outputs a clear image, as detailed below: Step 3.1: Obtain the relationship between the blurred image, the potentially sharp image, and the event flow from the perspective of events, i.e., event double integration. Construct an event double integration network to approximate the event double integration process. The expression for event double integration is as follows: ; in, Indicates the start time point, corresponding to the clear image. The moment; The integral time length represents the time window for event acquisition; This is the threshold for triggering the event; It is the sequence of events captured by the event camera, which is a function that changes over time; Indicates the time corresponding to the clear image up to the current time The cumulative number of events; This represents an exponential function used to convert cumulative event data into an intensity or brightness representation. Step 3.2: Construct an image fusion network containing a convolutional attention mechanism (CBAM). Input the blurred image and the corresponding double integral features, and output the clear image after deblurring.
9. The motion image deblurring method based on a dual-path imaging system and event dual integration according to claim 1, characterized in that, Step 4 introduces blur-event loss, optical flow-event loss, and blur-sharpness loss as self-supervised signals to constrain brightness, motion, and reconstruction consistency, as detailed below: Step 4.1, Fuzzy-Event Loss Function The estimation of the event double integral is constrained by the brightness consistency between consecutive blurred images, and the calculation formula is as follows: ; in, and These are the left and right blurred images, respectively. and These represent the corresponding double integral approximation features of the events. It is a constant introduced to ensure numerical stability, which is less than a set threshold; Step 4.2, Event-Optical Flow Loss Function The consistency between event data and image optical flow in describing physical motion is used to force the two motion fields to remain consistent. The calculation formula is as follows: ; in, and These are the length and width of the image, respectively. This represents the optical flow estimates of the two initially reconstructed left and right images under the RAFT-Sparse optical flow model. Indicates the image within the exposure time. gradient, This indicates the trigger threshold for the event camera. Represents an event flow; Step 4.3, Blur-Clear Loss Function The final reconstructed clear frame is blurred and compared with the blurred images of the left and right sides to calculate the reconstruction consistency loss. The calculation formula is as follows: ; in, and These represent the exposure ranges for the left and right blurred images, respectively. For reconstructed clear frames; Step 4.4, Total Loss Function The formula is as follows: ; in, , and These are the weights of the corresponding loss functions.
10. The motion image deblurring method based on a dual-path imaging system and event dual integration according to claim 1, characterized in that, In the two-stage training described in step 5, the first stage trains the event-based dual-integral network, and the second stage trains the image fusion network. The loss weights are adjusted with each stage, and the learning rate adopts a cosine annealing and periodic restart strategy.