Lightweight high-speed photographing system and method based on event camera

By combining hardware synchronization and data fusion technologies of visible light cameras and event cameras, the problems of data volume and power consumption of traditional high-speed cameras are solved, enabling lightweight high-speed photography with high frame rate and high dynamic range, and outputting clear video streams.

CN121645008APending Publication Date: 2026-03-10HUBEI SANJIANG AEROSPACE WANFENG TECH DEV
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Traditional high-speed cameras suffer from massive data volume, high power consumption, demanding storage and processing hardware requirements, short shooting time, and limited dynamic range due to full-frame-rate sampling, making it impossible to fully record scenes in extremely bright and dark areas; pure event cameras lack texture and color information and cannot capture static background details.

Method used

A hybrid imaging module is used to combine a visible light camera and an event camera. Time alignment is achieved through a hardware synchronization module. The high-frequency frame interpolation reconstruction algorithm of the data processing and fusion unit is used to fuse the event stream and visible light image frames to generate a high frame rate, high dynamic range, and motion-blur-free video stream.

Benefits of technology

It achieves high-performance, high-speed imaging with low cost and low power consumption, outputting high frame rate, rich texture and color video, suitable for portable devices, breaking through the data volume and time limitations of traditional cameras, and possessing high dynamic range and clear transient detail capture capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121645008A_ABST
    Figure CN121645008A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight high-speed photographing system and method based on an event camera, and the system comprises a hybrid imaging module which is used for synchronously collecting the image data and event data of a scene, and comprises a visible light camera and an event camera; the hardware synchronization module is electrically connected with the visible light camera and the event camera, and is used for receiving a frame synchronization signal of the visible light camera and sending a trigger signal to the event camera, so that the event camera inserts a time stamp corresponding to a visible light image frame exposure moment in an event stream; the data processing and fusion unit is in communication connection with the hybrid imaging module and is used for receiving the image data and the event data and executing the following steps: performing time alignment on an event stream and a visible light image frame stream based on a time stamp; operating a high-frequency frame insertion reconstruction algorithm, and reconstructing a multi-frame intermediate image by using event stream data between two frames; synthesizing the original visible light image frame and the reconstructed intermediate image into a high-frame-rate video stream; and the power supply and interface module is used for supplying power to the system and outputting a high-frame-rate video stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of high-speed imaging, and more specifically, to a lightweight high-speed photography system and method based on an event camera. Background Technology

[0002] Traditional high-speed cameras primarily rely on high-frame-rate image sensors to record high-speed dynamic processes by continuously capturing a series of complete image frames with extremely short exposure times. However, with the increasing demand for higher temporal resolution and longer recording times, traditional high-speed photography technology faces a series of insurmountable bottlenecks. Firstly, the data volume is enormous, reaching several gigabytes per second, placing extremely high demands on data transmission bandwidth, processing power, and storage capacity, resulting in high system costs and large size. Secondly, the dynamic range is limited; the dynamic range of traditional CMOS / CCD sensors is typically around 60-70 dB. In scenes with both extremely bright and extremely dark areas (such as welding, explosions, and shadows under sunlight), overexposure or underexposure easily occurs, making it impossible to fully record the scene's brightness information. Thirdly, the recording time is limited; due to the limitation of data storage speed, most traditional high-speed cameras can only record for a short period of a few seconds at the highest frame rate.

[0003] Unlike traditional cameras that capture the entire scene at a fixed frequency, event cameras operate with each pixel independently and asynchronously. An event is triggered and output only when a single pixel senses a change in light intensity exceeding a preset threshold. This event includes the pixel coordinates, a timestamp accurate to the microsecond level, and the polarity of the brightness change (increase or decrease). Event cameras possess microsecond-level temporal resolution, equivalent to hundreds of thousands or even millions of frame rates. Dynamic range easily exceeds 120dB, even reaching up to 140dB. When the scene is static or changing slowly, the camera generates almost no data, significantly reducing data bandwidth and power consumption. Events are recorded instantaneously with changes in brightness, therefore, motion blur is inherently absent.

[0004] However, event cameras also have limitations. They output a sparse stream of events rather than a complete image containing rich textures and colors. In static parts of the scene, the event camera produces no information, making it impossible to capture texture details of the static background. Furthermore, the spatial resolution of currently commercially available event cameras is generally lower than that of mainstream visible light cameras.

[0005] Related research considers fusing event cameras with traditional visible light cameras to combine the advantages of both, but mainly focuses on the algorithm level, such as using event data to improve the quality of video frame interpolation. Algorithms such as TimeLens and EFI-Net have demonstrated the potential to synthesize high-quality intermediate frames using high-temporal-precision motion information in event streams. However, these studies are mostly still in the theoretical and algorithm verification stage, and have not yet provided a complete, lightweight, low-power, and easy-to-deploy end-to-end hardware system solution.

[0006] In summary, there is an urgent need for a new high-speed photography solution that can overcome the many drawbacks of traditional high-speed cameras and make up for the shortcomings of pure event cameras, thereby achieving higher-performance high-speed, high dynamic range imaging at a lower cost, with a smaller size and power consumption. Summary of the Invention

[0007] To address at least one deficiency or improvement need in the prior art, this invention provides a lightweight high-speed photography system and method based on an event camera, which solves the problems in existing high-speed photography technology, such as huge data volume, high power consumption, demanding storage and processing hardware requirements, and short shooting time caused by full frame rate sampling. The final output is a video stream with high frame rate, high dynamic range, no motion blur, and rich texture and color.

[0008] To achieve the above objectives, according to a first aspect of the present invention, a lightweight high-speed photography system based on an event camera is provided. The system includes: a hybrid imaging module for synchronously acquiring image data and event data of a scene, the hybrid imaging module including a visible light camera and an event camera; a hardware synchronization module electrically connected to the visible light camera and the event camera, for receiving a frame synchronization signal from the visible light camera and sending a trigger signal to the event camera according to the frame synchronization signal, causing the event camera to insert a time stamp corresponding to the exposure time of a visible light image frame into the event stream; a data processing and fusion unit communicatively connected to the hybrid imaging module, for receiving the image data and event data, and performing the following operations: realigning the event stream with the visible light image frame stream based on the time stamp; running a high-frequency frame interpolation reconstruction algorithm to reconstruct multiple intermediate images using the event stream data between two frames; and combining the original visible light image frames and the reconstructed intermediate images into a high frame rate video stream; and a power supply and interface module for supplying power to the system and outputting the high frame rate video stream.

[0009] In an exemplary embodiment, the visible light camera and the event camera in the hybrid imaging module are coupled through a beam splitter optical path, and the rotation matrix and translation vector between them are obtained in advance through spatial calibration.

[0010] In one exemplary embodiment, the hardware synchronization module is configured to: detect the rising edge of the frame synchronization signal of the visible light camera; and in response to detecting the rising edge, generate a trigger pulse and send it to the event camera; wherein the event camera is configured to, upon receiving the trigger pulse, insert an external trigger event with a microsecond-level timestamp into the output asynchronous event stream to mark the start exposure time of each frame of the visible light camera.

[0011] In an exemplary embodiment, the data processing and fusion unit further includes a data input and alignment module, a shallow feature extraction module, a dual-channel attention module, a global feature fusion module, and an implicit temporal decoder. The data input and alignment module receives a visible light image frame sequence and an event stream, performs temporal alignment on the visible light image frame sequence and the event stream, and extracts the event set between two adjacent image frames. The shallow feature extraction module includes parallel image feature extraction paths and event feature extraction paths, used to encode the visible light image frames and event sets into spatiotemporal feature bodies, respectively. The dual-channel attention module performs cross-modal interaction, compensation, and fusion of image features and event features, including at least a dual attention transformer submodule and a channel attention submodule. The global feature fusion module fuses the multi-layer image features processed by the dual-channel attention module. The implicit temporal decoder decodes and generates a clear image at the corresponding time based on the input timestamp and the global features output by the global feature fusion module.

[0012] In one exemplary embodiment, the dual attention transformer submodule is configured to perform the following operations: divide the input image features and event features into local blocks and calculate the sum of their respective window-based multi-head self-attention weights; calibrate and update the self-attention weights of the event features using the self-attention weights of the image features, calculate the output event features using the calibrated self-attention weights of the event features; and calculate the output compensated image features using the self-attention weights of the image features and the weighted event features.

[0013] In an exemplary embodiment, the channel attention submodule is configured to: perform global pooling on the input image features and event features respectively to generate a first channel statistic and a second channel statistic; and use the first channel statistic and the second channel statistic as weights to adaptively rescale the image features and event features output in the processing stage respectively.

[0014] According to a second aspect of the present invention, a lightweight high-speed photography method based on an event camera is also provided, applied to the lightweight high-speed photography system based on an event camera as described above, comprising: acquiring a time-synchronized visible light image frame sequence and event stream through a hardware synchronization mechanism; performing spatiotemporal alignment on the visible light image frame sequence and event stream; inputting two adjacent visible light image frames and the event stream data between the two adjacent frames into a pre-trained high-frequency frame interpolation reconstruction model, wherein the model performs feature fusion through a network structure including a dual-channel attention mechanism to reconstruct multiple intermediate images; and synthesizing the original visible light image frames and the reconstructed intermediate images in chronological order into a high frame rate video stream and outputting it.

[0015] According to a third aspect of the invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to execute the above-described lightweight high-speed photography method based on an event camera when it is run.

[0016] According to a fourth aspect of the present invention, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the aforementioned lightweight high-speed photography method based on an event camera via the computer program.

[0017] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: (1) This invention provides a lightweight high-speed photography system based on an event camera, achieving an ultra-high equivalent frame rate and extremely low system overhead. By interpolating frames at high magnification between low-frame-rate visible light images, it is easy to achieve equivalent high-speed recording of thousands of FPS or even higher, while the system's data generation, transmission, and storage overhead is only slightly higher than that of an ordinary 60 FPS camera, and far lower than that of a traditional high-speed camera with the same frame rate. It also has high dynamic range and rich texture. This system combines the advantages of the ultra-high dynamic range of the event camera and the high resolution and rich colors of the visible light camera. The fusion algorithm can reconstruct the highlight or dark areas using event information while preserving complete texture details, and output high dynamic range video. The frame interpolation process is guided by an event stream accurate to microseconds, and motion estimation and compensation are completed at a microsecond temporal resolution, which fundamentally eliminates the motion blur caused by exposure time in traditional high-speed cameras and can capture clearer transient details; (2) The lightweight high-speed photography system based on an event camera provided in this application has the advantages of being lightweight, low-power, and applicable to a wide range of scenarios. Since the event camera, one of the core data sources, has extremely low power consumption (usually in the milliwatt to 1 watt range) and the overall data throughput is controllable, the entire system can be designed to be very compact and energy-efficient. This allows it to be used as a portable device or integrated into platforms such as drones and robots that have strict limitations on size, weight, and power consumption. In addition, since the pressure on raw data storage is greatly reduced, this system can achieve continuous high-speed recording for a long time (minutes or even longer), breaking through the second-level recording time limit of traditional high-speed cameras; (3) This invention provides a complete end-to-end solution from hardware to software, with high system integration and ease of use. Users can directly obtain high-quality high-speed video output without having to deal with complex dual-camera synchronization and data fusion issues themselves. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A schematic diagram of the overall hardware architecture of an optional lightweight high-speed photography system based on an event camera, provided for an embodiment of this application; Figure 2 A schematic block diagram of an optional spatiotemporal calibration provided for an embodiment of this application; Figure 3 A schematic diagram of an optional time synchronization hardware connection provided in an embodiment of this application; Figure 4 This application provides an optional spatial calibration process diagram. Figure 5 A general framework diagram of an optional high-frequency frame interpolation reconstruction algorithm provided for embodiments of this application; Figure 6 A schematic diagram of an optional dual-channel attention module provided in an embodiment of this application; Figure 7 A schematic diagram of an optional dual attention modulator provided in an embodiment of this application; Figure 8 A schematic diagram of an optional channel attention module provided in an embodiment of this application; Figure 9 The illustration shows the implementation effect of a lightweight high-speed photography system based on an event camera when selecting a combustion scene as the subject of the present application embodiment. Figure 10 An optional illustration of the comparison of single-frame reconstruction visualization results on the simulation dataset REDS, provided for an embodiment of this application; Figure 11 An optional visualization diagram of single-frame reconstruction results on the HQF simulation dataset, provided as an embodiment of this application; Figure 12 This is a schematic diagram of an optional electronic device provided in an embodiment of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0021] The terms "first," "second," "third," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0022] According to one aspect of an embodiment of this application, a lightweight high-speed photography system based on an event camera is provided. The following is in conjunction with... Figure 1 This application describes a lightweight high-speed photography system based on an event camera, as provided in the embodiments of this application.

[0023] Figure 1 A schematic diagram of an optional lightweight high-speed photography system based on an event camera is provided for embodiments of this application, as shown below. Figure 1 As shown, the system includes: A hybrid imaging module is used to simultaneously acquire image data and event data of a scene. The hybrid imaging module includes a visible light camera and an event camera. The hardware synchronization module is electrically connected to the visible light camera and the event camera. It is used to receive the frame synchronization signal from the visible light camera and send a trigger signal to the event camera according to the frame synchronization signal, so that the event camera inserts a time stamp corresponding to the exposure time of the visible light image frame into the event stream. The data processing and fusion unit, communicatively connected to the hybrid imaging module, is used to receive the image data and event data, and to perform the following operations: Based on the time stamp, the event stream is time-aligned with the visible light image frame stream; Run a high-frequency frame interpolation reconstruction algorithm to reconstruct the intermediate images of multiple frames using the event stream data between two frames; The original visible light image frames are combined with the reconstructed intermediate images to form a high frame rate video stream; The power supply and interface module is used to power the system and output the high frame rate video stream.

[0024] This application provides an optional lightweight high-speed photography system based on an event camera, applicable to fields such as scientific research, industrial inspection, sports analysis, and film and television special effects. At the hardware level, the system tightly integrates an event camera and a standard visible light (RGB) camera, achieving spatiotemporal synchronization. At the processing level, an efficient fusion algorithm utilizes the high temporal resolution information of the event stream to perform high-ratio frame interpolation reconstruction on the low frame rate images captured by the visible light camera. The final output is a high frame rate, high dynamic range, motion-blur-free video stream containing rich textures and colors. This solves the problems of existing high-speed photography technologies, such as massive data volume, high power consumption, stringent storage and processing hardware requirements, and excessively short shooting time due to full-frame-rate sampling.

[0025] like Figure 1 As shown, the lightweight high-speed photography system of the present invention mainly consists of a hybrid imaging module, a hardware synchronization module, a data processing and fusion unit, and a power supply and interface module.

[0026] Optionally, the hybrid imaging module includes a visible light camera and an event camera. In this embodiment, the visible light camera is an industrial-grade CMOS camera with a global shutter and a resolution of 1920x1080, configured to output 10-bit raw image data at a frame rate of 60 FPS via a MIPI CSI-2 interface. The event camera uses a sensor with a resolution of 1280x720, a dynamic range >120dB, and a timestamp accuracy of 1μs.

[0027] Optionally, the visible light camera and the event camera in the hybrid imaging module are optically coupled via a beam splitter prism, and their rotation matrix and translation vector are pre-calibrated spatially. The two cameras are precisely mounted via the beam splitter prism to reduce parallax.

[0028] In one exemplary embodiment, the hardware synchronization module is configured as follows: Detect the rising edge of the frame synchronization signal of the visible light camera; In response to the detection of the rising edge, a trigger pulse is generated and sent to the event camera; The event camera is configured to insert an external trigger event with a microsecond-level timestamp into the output asynchronous event stream when it receives the trigger pulse, so as to mark the start exposure time of each frame of the visible light camera.

[0029] Combination Figure 1 , Figure 2 and Figure 3 As shown, the hardware synchronization module directly utilizes the GPIO function of the data processing and fusion unit. The frame synchronization output signal (VSync) of the visible light camera is connected to the input of the hardware synchronization module. When the hardware synchronization module detects the rising edge of the VSync signal, it immediately generates a short pulse signal, which is sent to the external trigger input pin (EXT_TRIG) of the event camera. The firmware of the event camera is configured to insert a special external trigger event into its asynchronous event data stream upon receiving this pulse. This event contains a precise microsecond-level timestamp. In this way, the start exposure time of each frame of visible light image is precisely marked in the event stream.

[0030] Before system deployment, a one-time spatiotemporal calibration is required. (Refer to...) Figure 4 The calibration settings align the system with a checkerboard calibration board.

[0031] Spatial Calibration: The entire hybrid imaging module is moved, and a data segment is acquired. The visible light camera captures checkerboard image frames from different viewpoints, while the event camera records checkerboard black-and-white boundary events during the movement. These two sets of data are input into the Kalibr Camera Calibration Toolbox running on the PC. Kalibr can extract corner points from the visible light images and edges from the event images (reconstructed through short-term cumulative events). Through optimization, the rotation matrix R and translation vector T between the two cameras are accurately calculated.

[0032] Time calibration: While calibrating the runtime space, the Kalibr toolbox can also analyze the correlation between the timestamp of the hardware-triggered event and the timestamp of the visible light image frame, calculate a systematic small time offset tooffset and compensate for it.

[0033] After calibration, the parameters (R, T, toffset and their respective intrinsic parameters) are permanently stored in the non-volatile memory of the data processing and fusion unit.

[0034] In an exemplary embodiment, the data processing and fusion unit further includes a data input and alignment module, a shallow feature extraction module, a dual-channel attention module, a global feature fusion module, and an implicit temporal decoder; The data input and alignment module is used to receive visible light image frame sequences and event streams, perform time alignment on the visible light image frame sequences and event streams, and extract the event set between two adjacent image frames. The shallow feature extraction module includes parallel image feature extraction paths and event feature extraction paths, which are used to encode visible light image frames and event sets into spatiotemporal feature bodies, respectively. A dual-channel attention module is used for cross-modal interaction, compensation and fusion of image features and event features, and includes at least one dual attention transformer submodule and one channel attention submodule. A global feature fusion module is used to fuse multi-layer image features processed by the dual-channel attention module; An implicit time decoder is used to decode and generate a clear image at the corresponding time based on the input timestamp and the global features output by the global feature fusion module.

[0035] Preferably, the data processing and fusion unit employs an embedded information processing board based on a heterogeneous computing architecture. This processing board includes a general-purpose application processor (AP) and a dedicated hardware acceleration core. The AP runs a Linux operating system and upper-layer control software, and deploys a deep learning-based event camera frame interpolation and reconstruction network. The dedicated hardware acceleration core is responsible for receiving and buffering MIPI data streams from the two cameras, and performing data preprocessing and computationally intensive parts of the algorithm.

[0036] like Figures 5-8 As shown, the data input and alignment module in the data processing and fusion unit can receive image frame streams from a visible light camera. The event stream E from the event camera. Based on the trigger event markers inserted by the hardware synchronization module, the system can accurately segment any two frames of visible light images. and The set of all events that occurred between .

[0037] Shallow Feature Extraction (SFE) section: , and The data is fed into a convolutional neural network, which includes two parallel shallow feature extraction modules. The event stream shallow feature extraction employs a dual-path approach. , , The convolutional layers have 4 and 96 input channels and 96 output channels respectively; the shallow feature extraction of image frames uses a dual-path approach. , , The convolutional layer has 96 input and 96 output channels. This is used to handle sparse event sets. Encode into one or more dense event feature volumes, which contain rich motion information within that time period.

[0038] Optionally, in the Global Feature Fusion (GFF) section, after feature extraction from the DCAB module sequence, the global feature fusion module is used to fuse the image features output by each DCAB module to capture complete potential continuous motion features.

[0039] Implicit-time Decoder (ITD) differs from traditional fixed-output decoders. For example, the implicit-time decoder can combine the input timestamp with the global features of the GFF output to decode the potentially sharp image corresponding to the output time, and on this basis, recover high frame rate video and construct continuous dynamic scenes.

[0040] Furthermore, the video stream is output, that is, the original frames and all interpolated frames are combined in chronological order and output in real time through a gigabit Ethernet interface to form the final high-speed video.

[0041] In one exemplary embodiment, the dual attention transformer submodule is configured to perform the following operations: The input image features and event features are divided into local blocks, and the sum of their respective window-based multi-head self-attention weights is calculated. The self-attention weights of the event features are calibrated and updated using the self-attention weights of the image features, and the output event features are calculated using the calibrated self-attention weights of the event features. The self-attention weights of image features and the weighted event features are used to calculate the compensated image features.

[0042] Dual Channel Attention Block (DCAB): This attention-based dual channel attention block has two parallel inputs, including an image feature and an event feature. Details of DCAB are as follows... Figure 6As shown. Generally, image frames contain texture details and less noise, but lose motion information; events contain more complete motion and edge texture information, but have random noise. Therefore, the proposed DCAB uses image features to denoise events when calculating attention for image frames and event streams, and uses the corrected event features to feed back and enhance the image features, achieving the effect of de-motion blurring. In addition, DCAB also uses vertical channel attention to supplement the attention calculation in the vertical direction, achieving sufficient feature extraction from event streams and image frames. Specifically, the proposed DCAB module can be divided into two main steps: (1) Dual attention in the horizontal direction: In order to effectively extract cross-modal information from image frames and event streams and give full play to their respective advantages, such as Figure 7 As shown, a mutual compensation and fusion method based on Dual Attention Transformer (DAT) is proposed. In the... layer The input image features are stored in the DAT. Event characteristics The system is divided into several local blocks, and then the window-based multi-head self-attention of the two is calculated using the following formula:

[0043] Among them, self-attention weights Defined as:

[0044] Used to calculate weights , and These are encoded query vectors, key vectors, and value vectors. They utilize image features. Event characteristics Calculate their respective self-attention weights and Subsequently, to reduce the impact of event noise, the module uses self-attention weights based on nearly noise-free image features. Decalibrating self-attention weights of event features :

[0045] exist After being calibrated and updated, the module further analyzes the input event characteristics. Calculate the output event characteristics :

[0046]

[0047] in This corresponds to the value vector operator. This design helps reduce incorrect attention estimation caused by event noise while enhancing event features. On the other hand, image features are often affected by texture loss due to blurring caused by high-speed motion, which can be compensated for using events. This is achieved by inputting features of a blurred image. Features of weighted events Calculate the compensated image features :

[0048]

[0049] In one exemplary embodiment, the channel attention submodule is configured as follows: Global pooling is performed on the input image features and event features respectively to generate the first channel statistics and the second channel statistics; The first channel statistics and the second channel statistics are used as weights to adaptively rescale the image features and event features output during the processing stage.

[0050] Furthermore, regarding channel attention in the vertical direction, since the sliding window partitioning operation used in the DAT module to calculate attention divides the image plane into non-overlapping regions, the network uses a shift operation for DAT modules with even-numbered indices to maintain consistency between non-overlapping blocks. Although the DAT module helps to fully extract features in the horizontal direction, the interdependence information of feature channels in the vertical direction is not fully utilized. Therefore, for every two consecutive DAT modules, a Channel Attention Module (CAB) is used to connect the input of the first DAT module and the output of the second DAT module, forming residual structures in both the blurred image feature and event feature channels, thus forming a Dual Attention Module (DCAB) and supplementing channel attention in the vertical direction.

[0051] The structure of CAB is as follows: Figure 8 As shown. It consists of a simple structure of a global pooling layer, a convolutional layer, a ReLU activation layer, and a sigmoid activation function. In the first... layer In a DCAB, blurred image features and event characteristics The CABs entering their respective channels are used to calculate the statistics for the first channel. ), second channel statistics ( ):

[0052]

[0053] Using the generated first channel statistics Second channel statistics , respectively compared with the blurred image features output by the two DAT modules and event characteristics Multiply, adaptively rescale and update the two features:

[0054]

[0055] Therefore, the network can make full use of feature information in both horizontal and vertical directions, promote mutual compensation and fusion between blurred images and events, and further improve the effect of motion blur removal.

[0056] According to one aspect of the embodiments of this application, a lightweight high-speed photography method based on an event camera is provided, applied to the system as described above, including: The hardware synchronization mechanism is used to acquire time-synchronized visible light image frame sequences and event streams. Spatiotemporal alignment of the visible light image frame sequence and event stream; The visible light images of two adjacent frames and the event stream data between the two adjacent frames are input into a pre-trained high-frequency frame interpolation reconstruction model. The model performs feature fusion through a network structure containing a dual-channel attention mechanism to reconstruct multiple intermediate images. The original visible light image frames and the reconstructed intermediate images are combined in chronological order to form a high frame rate video stream and then output.

[0057] In this embodiment, after the system starts working, the hybrid imaging module synchronously acquires data. The visible light camera outputs a sequence of image frames at a preset frame rate (e.g., 60 FPS). The event camera asynchronously outputs an event stream. The hardware synchronization module ensures data synchronization by capturing the rising edge of the visible light camera's frame synchronization signal (VSync) and immediately sending a trigger pulse to the event camera. The event camera inserts an external trigger event containing a precise timestamp into the data stream. Thus, the initial exposure time of each frame is precisely marked in the event stream. Based on the trigger event markers in the event stream, all events corresponding to any two adjacent frames are precisely segmented to form an event set. For each event in the event set, its pixel coordinates are mapped to the image coordinate system of the visible light camera using a pre-calibrated rotation matrix and translation vector, resulting in aligned coordinates. This step ensures spatial consistency between the event and the image pixels, laying the foundation for subsequent pixel-level feature fusion. The aligned image pairs and the set of events between them are then input into a pre-trained model to reconstruct the intermediate images. The original frames and all reconstructed intermediate images are arranged according to their corresponding timestamps to form a continuous, smooth high frame rate video sequence. For example, if the original frame rate is 60 FPS, and N=99 intermediate frames are inserted between every two frames, the final output equivalent frame rate is 60×(99+1)=6000 FPS. This video stream is output or stored in real time via interfaces such as Gigabit Ethernet or USB.

[0058] Through this embodiment, a high-speed video with extremely high temporal resolution, no motion blur, and rich texture and color information can be output, perfectly reproducing the details of a high-speed dynamic process.

[0059] In an exemplary embodiment, the spatiotemporal alignment of the visible light image frame sequence and the event stream includes: Based on the timestamps of external triggering events inserted into the event stream by the hardware synchronization mechanism, the precise exposure time of each frame of visible light image is determined. The exposure time is compensated based on the time offset between the visible light camera and the event camera, which is obtained in advance through calibration. Based on the rotation matrix and translation vector obtained through pre-calibration, the event coordinates in the event stream are mapped to the image coordinate system of the visible light camera.

[0060] Reference Figure 9 The figure shows the effect of lightweight high-speed photography based on an event camera provided by this invention when a scene of intense combustion with gasoline being poured on a fire is selected as the subject of the photograph.

[0061] If a traditional high-speed camera is used, insufficient frame rate will result in noticeable motion blur and a jerky appearance. To capture images clearly, an extremely high frame rate is required, leading to an explosion in data volume.

[0062] The input sources for this invention are: a visible light camera that captures clear but time-spaced images at a low frame rate; and an event camera that generates a dense stream of events when gasoline is poured on a fire and detonates, accurately depicting the trajectory of the leaping flames.

[0063] The invention outputs the following: The system utilizes the trajectory information of an event stream to precisely synthesize multiple clear, blur-free intermediate images between two visible light images, ultimately forming smooth, clear high-speed video. Its effect rivals or even surpasses that of expensive traditional high-speed cameras. High-frequency frame interpolation reconstruction increases the frame rate of flame explosion video by 500 times.

[0064] In a typical embodiment, the visible light camera operates at 30 FPS, with an interpolation ratio of 499 (N=499). This means that 499 clear, continuous images are reconstructed by interpolating between two traditional images. The system can then output high-definition video at an equivalent frame rate of 30 * (499 + 1) = 15000 FPS. Its data output bandwidth and power consumption are primarily determined by the 30 FPS visible light camera and the data processing and fusion unit. Compared to a true 15000 FPS traditional high-speed camera, its data volume and power consumption are reduced by at least an order of magnitude. Figure 9 and Figure 10 As shown, simulation tests on the REDS and HQF datasets demonstrate that the image frames reconstructed by this system achieve PSNR and SSIM scores comparable to state-of-the-art pure software algorithms (eSL-Net and RED-Net, etc.) (PSNR > 30dB, SSIM > 0.90). Regarding dynamic range, by fusing event information, the system can correctly reconstruct scenes with brightness spanning multiple orders of magnitude, achieving an effective dynamic range > 120dB.

[0065] In summary, this invention, through innovative hardware and software co-design, successfully integrates the advantages of event cameras and visible light cameras, providing a novel lightweight, low-power, high-performance high-speed photography solution with extremely high practical value and broad application prospects.

[0066] According to another aspect of the embodiments of this application, a storage medium is also provided. Optionally, in this embodiment, the storage medium can be used to execute the program code of any of the lightweight high-speed photography methods based on event cameras described above in the embodiments of this application.

[0067] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: S1 acquires time-synchronized visible light image frame sequences and event streams through a hardware synchronization mechanism; S2, perform spatiotemporal alignment on the visible light image frame sequence and event stream; S3, input the visible light images of two adjacent frames and the event stream data between the two adjacent frames into the pre-trained high-frequency frame interpolation reconstruction model. The model performs feature fusion through a network structure containing a dual-channel attention mechanism to reconstruct the intermediate images of multiple frames. S4 combines the original visible light image frames and the reconstructed intermediate images in chronological order into a high frame rate video stream and outputs it.

[0068] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated in this embodiment.

[0069] The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, as well as magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0070] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described lightweight high-speed photography method based on an event camera is also provided. The electronic device may be a server, a terminal, or a combination thereof.

[0071] Figure 12 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application, such as... Figure 12 As shown, it includes a processor 1202, a communication interface 1204, a memory 1206, and a communication bus 208. The processor 1202, communication interface 1204, and memory 1206 communicate with each other via the communication bus 1208. Memory 1206 is used to store computer programs; When processor 1202 executes a computer program stored in memory 1206, it performs the following steps: S1 acquires time-synchronized visible light image frame sequences and event streams through a hardware synchronization mechanism; S2, perform spatiotemporal alignment on the visible light image frame sequence and event stream; S3, input the visible light images of two adjacent frames and the event stream data between the two adjacent frames into the pre-trained high-frequency frame interpolation reconstruction model. The model performs feature fusion through a network structure containing a dual-channel attention mechanism to reconstruct the intermediate images of multiple frames. S4 combines the original visible light image frames and the reconstructed intermediate images in chronological order into a high frame rate video stream and outputs it.

[0072] Optionally, the communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 The symbol is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned electronic device and other devices.

[0073] Memory may include RAM or non-volatile memory. Volatile memory, for example, at least one disk storage device. Alternatively, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0074] The processors mentioned above can be general-purpose processors, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; they can also be DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0075] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.

[0076] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0077] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0078] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0079] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of embodiments of this disclosure upon considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

[0080] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0081] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An event camera based lightweight high speed photography system, characterized in that, The application relates to a high-frame-rate hybrid imaging system. The application comprises: a hybrid imaging module for synchronously collecting image data and event data of a scene, the hybrid imaging module comprising a visible light camera and an event camera; a hardware synchronization module electrically connected to the visible light camera and the event camera, configured to receive a frame synchronization signal of the visible light camera, and send a trigger signal to the event camera according to the frame synchronization signal, so that the event camera inserts a time mark corresponding to an exposure time of a visible light image frame in an event stream; a data processing and fusion unit in communication connection with the hybrid imaging module, configured to receive the image data and the event data, and perform the following operations: time aligning the event stream and the visible light image frame stream based on the time mark; running a high-frequency frame insertion reconstruction algorithm to reconstruct multiple intermediate images by using event stream data between two frames; combining original visible light image frames and the reconstructed intermediate images into a high-frame-rate video stream; 2. The event camera based light-weighted high-speed photography system of claim 1, wherein, a power supply and interface module for supplying power to the system and outputting the high-frame-rate video stream.

3. The event camera based light-weighted high-speed photography system of claim 1, wherein, The visible light camera and the event camera in the hybrid imaging module are coupled through a light splitting prism optical path, and a rotation matrix and a translation vector between the two are obtained through spatial calibration in advance. The hardware synchronization module is configured to: detect a rising edge of the frame synchronization signal of the visible light camera; generate a trigger pulse and send it to the event camera in response to detecting the rising edge; 4. The event camera based light-weighted high-speed photography system of claim 1, wherein, wherein the event camera is configured to, upon receiving the trigger pulse, insert an external trigger event with a microsecond-level time stamp in an output asynchronous event stream to mark a starting exposure time of each image frame of the visible light camera. The data processing and fusion unit further comprises a data input and alignment module, a shallow feature extraction module, a dual-channel attention module, a global feature fusion module and an implicit time decoder. The data input and alignment module is configured to receive a visible light image frame sequence and an event stream, time align the visible light image frame sequence and the event stream, and extract an event set between two adjacent image frames. The shallow feature extraction module comprises a parallel image feature extraction path and an event feature extraction path, and is configured to encode the visible light image frame and the event set into a spatio-temporal feature body, respectively. The dual-channel attention module is configured to interact, compensate and fuse the image features and the event features across modalities, and at least comprises a dual attention transformer sub-module and a channel attention sub-module. The global feature fusion module is configured to fuse multi-layer image features processed by the dual-channel attention module.

5. The event camera based light-weighted high-speed photography system of claim 4, wherein, The implicit time decoder is configured to decode a clear image corresponding to a time instant according to an input time stamp and global features output by the global feature fusion module. The dual attention transformer sub-module is configured to perform the following operations: divide the input image features and event features into local blocks, respectively, and calculate window-based multi-head self-attention weights of the image features and the event features, respectively; update the self-attention weights of the event features by using the self-attention weights of the image features, and calculate output event features by using the updated self-attention weights of the event features. The output compensated image features are calculated using the self-attention weights of the image features and the weighted event features.

6. The event camera based light-weighted high-speed photography system of claim 4, wherein, The channel attention sub-module is configured to: perform global pooling on the input image features and event features respectively to generate first channel statistics and second channel statistics; use the first channel statistics and the second channel statistics as weights to respectively adaptively rescale the image features and the event features output by the processing stage.

7. An event camera based lightweight high speed photography method applied to the system of any one of claims 1-6, characterized in that, The method comprises: obtaining a time-synchronized sequence of visible light image frames and an event stream through a hardware synchronization mechanism; spatiotemporally aligning the sequence of visible light image frames and the event stream; inputting adjacent two frames of visible light images and event stream data between the adjacent two frames into a pre-trained high-frequency frame interpolation reconstruction model, the model performing feature fusion through a network structure containing a dual-channel attention mechanism to reconstruct multiple intermediate images; combining the original visible light image frames and the reconstructed intermediate images into a high-frame-rate video stream in chronological order and outputting the high-frame-rate video stream. 8.The event camera-based lightweight high-speed photography method of claim 7, wherein, The spatiotemporal alignment of the sequence of visible light image frames and the event stream comprises: determining the precise exposure time of each frame of visible light image according to the timestamp of the external trigger event inserted in the event stream by the hardware synchronization mechanism; compensating the exposure time based on the time offset between the visible light camera and the event camera obtained through calibration in advance; mapping the event coordinates in the event stream to the image coordinate system of the visible light camera based on the rotation matrix and the translation vector obtained through calibration in advance.

9. A computer readable storage medium, characterized in that, The computer-readable storage medium comprises a stored program, wherein the program, when executed, performs the method of any one of claims 7-8. 10.An electronic device comprising a memory and a processor, the electronic device characterized by, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 7-8 through the computer program.

Citation Information

Cited By

  • A defect detection and process localization method

    CN122196938A

  • A fusion event and rgb vision system and method

    CN122372816A