Event camera based high dynamic high speed imaging method and system
By using a field-of-view coded imaging system and a spatiotemporal grayscale completion neural network, the time limitations and artifact problems of traditional high dynamic and high-speed imaging are solved, achieving high-quality high dynamic and high-speed imaging.
Patent Information
- Application Number
- CN202410920890.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-10
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-07-10
AI Technical Summary
Traditional high dynamic range imaging requires longer exposure times, and traditional high-speed cameras cannot operate continuously for long periods. Event-guided high dynamic range high-speed imaging suffers from high artifacts and low realism.
A field-of-view coding imaging system is used, which controls the incident light through a spatial light modulator. An event camera records the timestamp and grayscale value of the conduction state, and combines the spatiotemporal grayscale completion neural network to recover high frame rate dense video.
It achieves high dynamic range and high-speed imaging, reduces the risk of image information aliasing, improves image quality and fidelity, and reduces system complexity and power consumption.
Smart Images

Figure CN118741327B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and computational imaging, specifically relating to a high dynamic range and high speed imaging method and system based on an event camera. Background Technology
[0002] Event cameras (or dynamic vision sensors) differ from traditional image sensors in that they can respond to changes in brightness illuminating the image surface with extremely high temporal resolution. The advantage of event cameras lies in their compression of spatiotemporal light field information in the brightness dimension, thus eliminating the static, textureless, and redundant information acquired by traditional image sensors, achieving frame rates far exceeding those of traditional image sensors. The disadvantage of event cameras is that, precisely because they compress the brightness dimension, they describe information at a point in spatiotemporal space using only three states of brightness change: "brighter," "darker," and "unchanged," thus lacking a reliable brightness reference. Furthermore, because they can only respond to changes in brightness, their output information is extremely sparse, often only appearing at the edges of moving objects, a stark contrast to the dense vision of humans.
[0003] High dynamic range (HDR) imaging and high-speed imaging are two classic challenges in the photography industry. HDR refers to an image sensor's ability to simultaneously capture both extremely dark and extremely bright texture details within a scene with a high dynamic range. Traditional methods for achieving HDR typically involve continuously exposing the same scene with different exposure parameters, followed by image registration and fusion to obtain a single HDR image with bracketed exposures. This method requires significantly more time than conventional image acquisition and suffers from motion blur and mismatched images. High-speed imaging, on the other hand, refers to an image sensor's ability to continuously expose a fast-moving scene at speeds of up to several thousand frames per second. Achieving high-speed photography places extremely high demands on both the image sensor's acquisition and readout speeds. To ensure that large amounts of image data are not lost during acquisition, commercially available high-speed cameras often have ultra-large capacity and ultra-fast write speed buffers. Even so, such products typically only support short continuous operation times of tens of seconds, unable to support extended high-speed imaging.
[0004] In recent years, some researchers have employed event cameras to guide traditional cameras in high dynamic range and high-speed imaging. The basic principle is to use a traditional camera to capture low-frame-rate, low-dynamic-range video images, and an event camera to capture events generated during video recording. This allows for the reconstruction of motion relationships between frames in the traditional image, enabling motion blur removal and frame interpolation of the original video stream. Simultaneously, the high dynamic range of the event camera helps compensate for bright and dark textures that traditional cameras cannot capture. However, such systems and methods require the fusion of traditional and event cameras, resulting in complex structures. Furthermore, due to inconsistencies in dynamic range and noise characteristics between the two cameras, the recovered "high dynamic range and high-speed" imaging results often contain numerous artifacts, and the bright and dark texture information supplemented by events lacks authenticity. Summary of the Invention
[0005] To overcome the problems of high dynamic range imaging requiring longer exposure times, the inability of traditional high-speed cameras to operate continuously for extended periods, and the high artifacts and low fidelity issues in event-guided high dynamic range high-speed imaging, this invention proposes a high dynamic range high-speed imaging method and system based on an event camera.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] In a first aspect, the present invention discloses a high dynamic range, high-speed imaging method based on an event camera, comprising the following steps:
[0008] 1) Construct a field-of-view coded imaging system for acquiring information on the brightness changes of incident light in high-speed motion scenes. The field-of-view coded imaging system includes an imaging objective lens, a front relay lens, a spatial light modulator, a rear relay lens, and an event camera.
[0009] 2) First, block the incident light rays corresponding to all pixels on the spatial light modulator. Then, adjust the spatial light modulator multiple times so that the incident light rays of all pixels can illuminate the image plane of the event camera. Each adjustment of the spatial light modulator will add a pixel in the conducting state, and the pixels that were originally in the conducting state will remain conducting. The event camera will record the timestamp of the first positive event triggered by each pixel that enters the conducting state. The positive event is the brightness change information recorded when the event camera detects that the photovoltage on a certain pixel increases and exceeds the change threshold. The recorded information includes spatial information and temporal information.
[0010] 3) Calculate the time difference between the timestamp of the first positive event triggered by the pixel entering the conduction state and the initial conduction time of the pixel, and calculate the gray value of each pixel at its conduction time to obtain a spatially sparse high dynamic range grayscale image. Arrange all spatially sparse high dynamic range grayscale images in chronological order to obtain a high frame rate spatially sparse grayscale image sequence.
[0011] 4) Based on the high frame rate spatially sparse grayscale image sequence, a dense high dynamic range high speed video is obtained through a spatiotemporal grayscale completion neural network.
[0012] Furthermore, the specific method for controlling the spatial light modulator is as follows: the modulation mask of the spatial light modulator is divided into multiple non-overlapping 3M*3M regions, each region containing 9 M*M blocks. Each time the spatial light modulator is controlled, an additional M*M block is added to each 3M*3M region to enter the conduction state.
[0013] Further, step 4) specifically includes:
[0014] 4.1) Downsample each spatially sparse high dynamic range grayscale image in the high frame rate spatially sparse grayscale image sequence to obtain a downsampled dense image sequence in which the length and width of the image are both 1 / 3 of the original image; then use the RAFT model to estimate the optical flow of the downsampled dense image sequence to obtain a downsampled dense optical flow in which the length and width of the image are both 1 / 3 of the original image; then upsample the downsampled dense optical flow to obtain a dense optical flow in which the length and width of the image are consistent with the original image.
[0015] 4.2) Based on the relative position vector of the M*M block with gray values within its own 3M*3M region, local quantization compensation is performed on the dense optical flow of each 3M*3M region to obtain the compensated dense optical flow.
[0016] 4.3) The compensated dense optical flow is used to propagate image domain information and feature domain information to the high frame rate spatially sparse grayscale image sequence to obtain a high frame rate spatially sparse grayscale image sequence after one completion.
[0017] 4.4) The high frame rate spatially sparse grayscale image sequence after one completion is further transformed and completed by a mask-guided sparse transformer to obtain dense high dynamic range high speed video.
[0018] Further, in step 4.1), the downsampling of each spatially sparse high dynamic range grayscale image in the high frame rate spatially sparse grayscale image sequence to obtain a downsampled dense image sequence whose length and width are both 1 / 3 of the original image; specifically:
[0019] For each spatially sparse high dynamic range grayscale image, create a new matrix whose length and width are both 1 / 3 of the original image. Then, select M*M blocks with grayscale values within each 3M*3M region of the spatially sparse high dynamic range grayscale image and place them into the new matrix of the spatially sparse high dynamic range grayscale image. The spatial position of the blocks is the same as the position of the 3M*3M region in the original spatially sparse high dynamic range grayscale image. This results in a new image whose length and width are both 1 / 3 of the original image. Arrange the new images in chronological order to obtain a downsampled dense image sequence whose length and width are both 1 / 3 of the original image.
[0020] In a second aspect, the present invention discloses a high dynamic range high speed imaging system based on an event camera, including a field-of-view coded imaging system, a spatially sparse grayscale image acquisition module, and an imaging video acquisition module.
[0021] The field-of-view coded imaging system is used to obtain the timestamp of the first positive event triggered by each pixel entering the conduction state;
[0022] The spatially sparse grayscale image acquisition module first calculates the time difference between the timestamp of the first positive event triggered by the pixel entering the conduction state and the initial conduction time of the pixel, and calculates the grayscale value of each pixel at its conduction time, and finally obtains a high frame rate spatially sparse grayscale image.
[0023] The imaging video acquisition module obtains dense, high-dynamic, high-speed video based on the high frame rate spatially sparse grayscale image through a spatiotemporal grayscale completion neural network.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0025] This invention utilizes the time difference between the timestamp of the first positive event triggered by the activated pixel and the activation time to calculate image grayscale, overcoming the problem that event cameras cannot acquire image grayscale information, thus enabling event cameras to be applied to high-speed, high-dynamic-range image acquisition. This invention combines a field-of-view coded imaging system with an event camera, overcoming the problem of existing field-of-view coded compressed sensing imaging systems outputting spatiotemporal image information superimposed on a single grayscale image, thereby obtaining spatiotemporally sparse grayscale information, eliminating the risk of image information superimposition, and enabling the recovery of higher-quality grayscale images with fewer artifacts. This invention uses only one event camera for image information acquisition, overcoming the deficiency of event camera-guided high-speed imaging of traditional cameras requiring multiple sensor fusion, improving system integration, solving recovery artifacts caused by multi-sensor fusion, and improving the fidelity of high-speed images. This invention employs a unique sparse image optical flow estimation method, overcoming the deficiency of traditional optical flow estimation methods that cannot be applied to spatiotemporally sparse grayscale images, thus realizing the conversion from sparse images to dense optical flow, providing an important reference for high-speed image recovery. Attached Figure Description
[0026] Figure 1 This invention uses a DMD (Digital Micromirror Device) modulated field-of-view coded imaging system.
[0027] Figure 2 This invention uses an LCoS (Liquid Crystal on Silicon) modulation field-of-view encoding imaging system.
[0028] Figure 3 This is the modulation mask pattern switched at the 1ms in this embodiment of the invention;
[0029] Figure 4 This is the process of updating the modulation mask at 2ms in this embodiment of the invention;
[0030] Figure 5 This is an example of the field-of-view coding timing used in the embodiments of the present invention;
[0031] Figure 6 This is an overall flowchart of the temporal grayscale completion network of the present invention;
[0032] Figure 7 This is a comparison of the dense optical flow before and after correction with the true optical flow obtained from the embodiments of the present invention;
[0033] Figure 8 This paper compares the actual performance of the spatiotemporal grayscale recovery neural network of this invention in processing sparsely acquired grayscale image sequences with that of traditional compressed sensing imaging methods. Detailed Implementation
[0034] The present invention will be further described and illustrated below with reference to specific embodiments. The embodiments described are merely examples of the content of this disclosure and do not limit the scope of the invention. The technical features of each embodiment in the present invention can be combined accordingly, provided that there is no mutual conflict.
[0035] The method of this invention is implemented based on a field-of-view coded imaging system, which includes an imaging objective lens, a front relay lens, a rear relay lens, a spatial light modulator (SLM), an event camera, and optional polarizers or polarization beam splitters (PBS). The imaging objective lens images the external scene onto a virtual image plane; the front relay lens images the image on the virtual image plane onto the surface of the spatial light modulator; the spatial light modulator can be high-speed modulated, selecting some of its pixels to conduct, thereby continuing to send part of the field-of-view light to the rear relay lens; the rear relay lens receives the field-of-view light modulated by the spatial light modulator and images it onto the event camera. The spatial light modulator can be implemented as a digital micromirror device (DMD) or a liquid crystal on silicon (LCoS). Due to their different characteristics, different spatial light modulators require slightly different optical system arrangements. Figure 1 and 2 The diagram shown is a schematic representation of two representative field-of-view coded imaging systems provided by this invention.
[0036] Figure 1 This invention utilizes a DMD modulation field-of-view coded imaging system, with a digital micromirror device (DMM) 3 as the spatial light modulator. The imaging objective 1 images the natural scene onto a virtual image plane; the front relay mirror 2 images the image from the virtual image plane onto the DMM 3; the DMM 3 changes the deflection angle of the micromirrors on different pixels, allowing the field-of-view rays from the white-indicated pixels to be reflected into the subsequent optical path, while the field-of-view rays from the black-indicated pixels are excluded from the optical system, thus generating a modulation mask; the rear relay mirror 4 images the modulated field-of-view rays onto the image plane of the event camera; and the event camera 5 acquires brightness change information from the image plane of the event camera.
[0037] Figure 2This invention utilizes an LCoS modulated field-of-view coded imaging system, with a spatial light modulator being a silicon-based liquid crystal 8 (SLC). The imaging objective 1 images the natural scene onto a virtual image plane. A front relay lens 2 sequentially passes the image from the virtual image plane through a first polarizer 6 and a polarizing prism 7 to further image it onto the surface of the SLC 8. It is necessary to control the polarization direction of the first polarizer 6 to be consistent with the polarization direction of the light transmitted through the polarizing prism 7. The SLC 8 changes the polarization direction of the incident light from each pixel, causing the polarization direction of white pixels to change by 90 degrees, while the polarization direction of black pixels remains unchanged. The polarization direction of the second polarizer 9 is perpendicular to that of the first polarizer 6, allowing field-of-view light whose polarization direction has changed by 90 degrees after modulation by the SLC to pass through the second polarizer 9, while field-of-view light whose polarization direction has not changed cannot pass through the second polarizer 9. A rear relay lens 4 images the surface of the SLC onto the image plane of an event camera. The event camera 5 is used to collect brightness change information.
[0038] Figure 1 and Figure 2 The spatial light modulator in the image is conjugate to both the virtual image plane and the image plane of the event camera. Its role in the optical path is to control how external light illuminates a sparse portion of the pixels, preventing it from illuminating the majority of the remaining pixels. Therefore, at each moment, the event camera only acquires spatially sparse brightness variation information. This sparse acquisition aims to reduce the bandwidth burden on the transmission link during data acquisition in high-speed imaging systems.
[0039] The selection of a spatial light modulator must meet the following requirements: 1. The switching frequency should not be lower than the target frame rate for high-speed imaging. 2. The contrast ratio should not be lower than 1:100. The switching frequency directly affects the operating frequency of the field-of-view coded imaging system; the contrast ratio refers to the ratio of the light intensity of the shaded part to the light intensity of the transmitted part in the modulation mask. A low contrast ratio will affect the signal-to-noise ratio of the acquired grayscale values.
[0040] The spatial light modulator is hard synchronized with the event camera, meaning the event camera can record the moment the spatial light modulator switches the modulation mask each time.
[0041] Combination Figure 1 and Figure 2 The high dynamic range and high speed imaging method based on an event camera provided by this invention can be implemented according to the following steps:
[0042] 1) Construction of the field-of-view coded imaging system. According to the present invention... Figure 1 Or example Figure 2 Establish a field-of-view coded imaging system.
[0043] 2) Adjust the spatial light modulator to block the incident light from all pixels, modulate the mask to be completely black, and enter the imaging preparation state.
[0044] The spatial light modulator is adjusted, and the modulation mask is switched so that incident light from some pixels can illuminate the sensor surface of the event camera, putting these pixels in a conductive state. The remaining pixels are not illuminated by incident light and are in a masked state. The event camera records the timestamp of the first positive event triggered by the pixels entering the conductive state. A positive event refers to the brightness change information recorded when the event camera detects an increase in photovoltage on a pixel that exceeds a change threshold; this information includes both spatial and temporal data.
[0045] The spatial light modulator is adjusted to update the modulation mask, bringing some previously blocked pixels to life while maintaining the functionality of previously active pixels. The event camera also records the timestamp of the first positive event triggered by the newly activated pixels. The time difference between two modulation mask switching operations is called the modulation period T (i.e., the time difference between two adjacent adjustments to the spatial light modulator). The minimum value of the modulation period T is determined by the highest switching frequency of the spatial light modulator.
[0046] The spatial light modulator is adjusted and the modulation mask is updated multiple times to ensure that every pixel of the event camera enters the overconducted state. During the modulation mask switching process, each pixel remains in the overconducted state for a time of mT. After each pixel has been in the overconducted state for mT, it will enter a masked state for a period of nT. This nT time is called the reset time of each pixel. m and n are preset parameters.
[0047] 3) Using the timestamp of the first positive event triggered by each activated pixel and the time difference between the timestamp and the activation time, the grayscale value of each pixel at its activation time is calculated. Thus, for each moment of switching modulation masks, this invention can obtain the grayscale values of newly activated pixels at that moment. These grayscale values are high frame rate but spatially sparse.
[0048] 4) The high frame rate spatially sparse grayscale data obtained in step 3) is fed into the spatiotemporal grayscale completion neural network to obtain dense high frame rate imaging video.
[0049] In one specific embodiment of the present invention, step 2) specifically includes:
[0050] 2.1) Switch the modulation mask from completely black to... Figure 3 The pattern shown. Figure 3 In this design, white pixels represent the on state, and black pixels represent the off state. The modulation mask is specifically designed as follows: it is divided into non-overlapping 3M*3M regions, each containing nine M*M blocks. For each mask modulation, exactly one M*M block within each region becomes on; that is, each time the spatial light modulator is adjusted, one new M*M block enters the on state within each 3M*3M region. In this embodiment, M is chosen as 3.
[0051] 2.2) The pixels of the event camera entering the conduction state collect incident photons, resulting in an increase in the photovoltage of the photodiode on that pixel:
[0052]
[0053] In the formula (reference: Bao Y, Sun L, Ma Y, et al. Temporal-Mapping Photography for Event Cameras[J]. arXiv preprint arXiv:2403.06443,2024.), C PD V is the capacitance of the photodiode in the event camera. ref ε is the reference voltage for the event camera, which is 0 because all pixels are occluded in the initial imaging preparation state; I(x,y) is the grayscale value (light intensity) of pixel (x,y) when it is turned on; TR(t) is the light intensity modulation function, which can be referenced to the response function of the spatial light modulator used. Generally, since the response speed of the spatial light modulator is extremely fast, this response function can be considered as a unit step function ε(t) that steps to 1 at t=0; t * (x,y) is the time difference between the timestamp of the first positive event after pixel (x,y) is turned on and the time of turn-on. In the formula, this invention defines the turn-on time as time 0; V thd The preset threshold voltage for the event camera.
[0054] The meaning of this formula is that when a certain pixel (x,y) is turned on, the event camera will, according to the different gray values (light intensities) I(x,y) of that pixel at different times t * (x, y) outputs the first positive event. The formula describes t. * There is an approximate inverse relationship between (x,y) and grayscale value (light intensity) I(x,y). Since there is a linear mapping relationship between grayscale value and light intensity, the concepts of light intensity and grayscale value will not be distinguished below.
[0055] 2.3) Maintain conduction state t * After (x,y), and after pixel (x,y) outputs the first positive event, an event tail removal filter is used to discard all subsequent positive events to reduce the data transmission burden.
[0056] 2.4) At the start of the next modulation period T, the modulation mask is updated, turning on the previously masked pixels, while the previously turned-on pixels remain turned on. In this embodiment, the proportion of newly added pixels is 10% of the total pixels, and the modulation period T is 1ms. Example Figure 4The process of updating the mask is shown: the mask updated in step four is based on the original mask modulated in step three, with the addition of 10% random conductive pixels.
[0057] 2.5) Referring to steps 2.2) and 2.3), for newly connected pixels, their grayscale values are also collected.
[0058] 2.6) The spatial light modulator is adjusted and the modulation mask is updated multiple times to ensure that every pixel of the event camera enters the overconducting state. This ensures that the grayscale of each pixel of the event camera can be captured at a certain moment.
[0059] During mask switching, each pixel remains on for a period of time, mT. The value of m needs to take into account the time required for the darkest pixel on the sensor to trigger the first positive event. In this embodiment, m is set to 5.
[0060] After each pixel goes through an mT conduction state, it enters a occlusion state for nT. This nT time is called the reset time for each pixel. The purpose of the reset time is to provide a reference voltage V for the event camera. ref Allow sufficient time for the voltage to return to zero. If n is too small, it will cause the reference voltage V to... ref If the value is not 0, the grayscale value collected when the pixel is turned on again will be distorted. In this embodiment, n is 2.
[0061] example Figure 5 This is an example of the field-of-view coding timing used in an embodiment of the present invention. Pixels (301, 28) and (51, 328) are turned on at different times and remain on for 5T time. During this period, due to the different light intensities received by the two pixels, they remain on for t seconds from their respective on-times. * The first positive event was triggered.
[0062] Step 3) contains the following details:
[0063] 3.1) For each conduction time of each pixel, find the timestamp of the first positive event within the subsequent mT time interval, and calculate the time difference t between the timestamp of this positive event and the conduction time. * (x,y).
[0064] 3.2) Referring to the formula mentioned in 2.2), since for each pixel of the camera at the same event, the photodiode capacitance C PD and V thd They are basically the same and can be considered as the same constant, therefore they can be determined by t. * The light intensity (grayscale value) I(x,y) is calculated from (x,y). The scaling factor between light intensity and grayscale is given by C. PD and V thd If it is necessary to measure the absolute physical quantity of light intensity, then C can be used.PD and V thd Perform actual physical calibration measurements.
[0065] After converting the timestamps of all positive events into light intensity (grayscale values), grayscale values with the same conduction time are stored in the same matrix to obtain spatially sparse high dynamic range grayscale images. All spatially sparse images are arranged in chronological order, resulting in a grayscale image sequence with a frame rate of 1 / T but spatially sparse, i.e., a high frame rate spatially sparse grayscale image sequence. In this embodiment, the frame rate of these grayscale images is 1000 FPS (frames per second), and the spatial sparsity is 10%, meaning that only 10% of the positions in each image have grayscale values, while the remaining positions are empty.
[0066] The grayscale image sequence obtained in step 3.2) exhibits high dynamic range and high grayscale quantization accuracy because: the grayscale measurement described in this invention is performed by measuring the time it takes for the photovoltage to reach a set threshold. The maximum light intensity that this invention can identify is determined by the minimum temporal resolution of the event camera, which is 1 μs in this embodiment. The minimum light intensity that this invention can identify is determined by the holding time (mT), which is 5 ms in this embodiment. Therefore, the equivalent dynamic range of the grayscale image is 74 dB. The grayscale resolution of the grayscale image in this invention is determined by the minimum temporal resolution of the event camera, which is 1 μs in this embodiment. Therefore, the quantization accuracy level of the grayscale image is 5000. It is worth noting that the imaging dynamic range of this invention can be flexibly varied by changing the holding time (mT), achieving adaptive dynamic range adjustment according to the minimum light intensity to be detected, thus improving the flexibility of the system.
[0067] The spatiotemporal grayscale completion neural network in step 4) is specifically as follows:
[0068] The spatiotemporal grayscale completion neural network comprises three core modules: a quantization optical flow estimation module, a dual-domain information propagation module, and a mask-guided sparse Transformer module. The quantization optical flow estimation module includes two sequential steps: downsampling optical flow estimation and optical flow quantization error compensation. The dual-domain information propagation module includes image domain information propagation and feature domain information propagation modules. Figure 6 The overall flowchart of the spatiotemporal grayscale completion neural network is shown.
[0069] 4.1) Downsampling Optical Flow Estimation. First, a new matrix with a width and height of 1 / 3 of the original size is created for each spatially sparse high dynamic range (HMR) grayscale image obtained in step 3.2). Then, M*M blocks with grayscale values are selected from each 3M*3M region of the spatially sparse HMR grayscale image and placed into the new matrix. The spatial position of the selected blocks is the same as the position of the 3M*3M region in the original spatially sparse HMR grayscale image. This yields a downsampled dense image sequence with a width and height of 1 / 3 of the original spatially sparse image. For this downsampled dense image sequence, the RAFT model (Reference: Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In ECCV, 2020.) is used for optical flow estimation to obtain the downsampled dense optical flow with a width and height of 1 / 3 of the original spatially sparse image. The obtained downsampled dense optical flow is then upsampled by 3 times, and the optical flow value is multiplied by 3 times the upsampling rate to obtain a dense optical flow with the same length and width as the original spatially sparse image.
[0070] 4.2) Optical Quantization Error Compensation. Based on the relative position vectors of the M*M blocks with grayscale values within a 3M*3M region, local quantization compensation is performed on the dense optical flow at each 3M*3M region location to correct quantization errors. The quantization error compensation process is as follows: Taking the forward optical flow of a certain 3M*3M region as an example, the local quantization error compensation amount is the relative position vector of the M*M blocks with grayscale values in the current frame within the 3M*3M region minus the relative position vector of the M*M blocks with grayscale values in the next frame within the 3M*3M region. This correction amount is then added to the dense optical flow of the 3M*3M region obtained in step 4.1). After this process, the compensated dense optical flow is obtained. Figure 7 The comparison between dense optical flow and true optical flow before and after optical flow error compensation is shown, and the compensated dense optical flow is more consistent with the true optical flow.
[0071] 4.3) Dual-domain information propagation. Dual-domain information propagation is performed using the compensated dense optical flow obtained in step 4.2), which includes image domain information propagation and feature domain information propagation. This invention performs global information propagation in the image domain and local information propagation in the feature domain. Both involve forward and backward propagation, which follow the same process.
[0072] Image domain information propagation. Global information propagation occurs in the image domain, but conditions C1, C2, and C3 must be satisfied simultaneously. Image domain information propagation employs an optical flow-based warp method, utilizing the following reliability evaluation strategy: the reliability of the completed optical flow is verified by checking the consistency of forward and backward optical flow. The evaluation method is as follows:
[0073]
[0074] Where x refers to the coordinates (x, y) of any pixel p. For backward optical flow, This represents the forward optical flow. ε t→t+1 (x) represents the forward and backward optical flow consistency error, which only occurs when condition C1:ε t→t+1 Image domain information propagation is performed on pixel p at time t only if (x) < ∈ . In this embodiment, ∈ is taken as 5. Image domain information propagation also requires satisfying C2: at time t, pixel p is masked. In addition, image domain information propagation also requires satisfying C3: image information propagated from adjacent frames cannot be masked.
[0075] After image domain information propagation, the grayscale value of pixel p that simultaneously satisfies conditions C1, C2, and C3 at any time t is updated using the optical flow warp method, while the grayscale values of the remaining pixels remain unchanged. After image domain information propagation, some missing parts in the spatially sparse image are filled in.
[0076] Feature domain information propagation. First, image sequence features are extracted from a local image sequence (in this embodiment, 10 adjacent spatially sparse images) using an image encoder (reference: Liu R, Deng H, Huang Y, et al. Fuseformer: Fusing fine-grained information in transformers for video inpainting [C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2021:14040-14049.). Then, the image sequence features are propagated using an optical flow-guided variable alignment module (reference: Chan K CK, Zhou S, Xu X, et al. Basicvsr++: Improving video super-resolution with enhanced propagation and alignment [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022:5972-5981.). After the information in the feature domain is propagated, the missing parts in the spatially sparse image are filled to a greater extent, but some areas are still incomplete.
[0077] 4.5) Mask-guided sparse Transformer. The sparse transformer guided by the mask is used to further transform and complete the regions that are still incomplete in step 4.4). First, the image sequence features obtained in step 4.4) are processed using a soft segmentation operation (reference: Liu R, Deng H, Huang Y, et al. Fuseformer: Fusing fine-grained information in transformers for video inpainting[C] / / Proceedings of the IEEE / CVF international conference on computer vision.2021:14040-14049.) to generate patch embeddings and segment them into non-overlapping 5×9 windows. Then, the query Q, key K, and value V are obtained from the patch embedding windows through a linear layer. For the query Q, spatiotemporal attention is applied only to the windows covered by the mask. For the key K and value V, odd-numbered windows participate in self-attention to reduce the computational and memory costs of the Transformer module. After filtering out unnecessary windows in the query Q, key K, and value V, self-attention is applied to the remaining windows to extract finer features. These features are integrated using a soft combination operation and reconstructed by an image decoder to obtain the final dense high-speed video.
[0078] Figure 8 The practical effect of the spatiotemporal grayscale recovery neural network of the present invention is demonstrated. Figure 8 The first image from top to bottom is the true value of a scene taken in a static state. Figure 8 The second image from top to bottom is a sparse grayscale image acquired at 1ms during the high-speed movement to the right by the system of the present invention. Figure 8 The third image from top to bottom is a dense grayscale image obtained after processing by the spatiotemporal grayscale recovery neural network in step 4). Figure 8 The fourth image from top to bottom is the exposure image from 0 to 9 ms obtained using the traditional compressed sensing imaging method. Figure 8The fifth image from top to bottom shows the result recovered at 1ms using the traditional compressed sensing imaging method (Reference: Liu Y, Yuan X, Suo J, et al. Rank minimization for snapshot compressive imaging[J]. IEEE transactions on pattern analysis and machine intelligence, 2018, 41(12):2990-3006.). The peak signal-to-noise ratio (PSNR) of the recovered result of this invention is 36.19dB, which is higher than the PSNR of 20.77dB based on the traditional compressed sensing imaging method. Compared with high-speed imaging technology guided by event cameras, this invention uses only one event camera sensor, reducing system complexity and eliminating reconstruction artifacts caused by inherent biases of different devices. Compared with conventional high-speed cameras, this invention greatly reduces the spatiotemporal redundancy of high-speed image data, enabling information to be read out in real time. Therefore, it can achieve high-speed image information acquisition at any time period without being limited by the high-speed cache space inside the high-speed camera. At the same time, the system power consumption is reduced due to the reduction of information acquisition redundancy.
[0079] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. Those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A high dynamic range, high-speed imaging method based on an event camera, characterized in that, Includes the following steps: 1) Construct a field-of-view coded imaging system for acquiring information on the brightness changes of incident light in high-speed motion scenes. The field-of-view coded imaging system includes an imaging objective lens, a front relay lens, a spatial light modulator, a rear relay lens, and an event camera. 2) First, block the incident light rays corresponding to all pixels on the spatial light modulator, and then adjust the spatial light modulator multiple times so that the incident light rays of all pixels can illuminate the image plane of the event camera. Each adjustment of the spatial light modulator will add pixels in the conducting state, and the pixels that were originally in the conducting state will remain conducting. The event camera records the timestamp of the first positive event triggered by each pixel entering the conduction state. The positive event is the brightness change information recorded when the event camera detects an increase in photovoltage on a certain pixel that exceeds the change threshold. The recorded information includes spatial information and temporal information. 3) Calculate the time difference between the timestamp of the first positive event triggered by the pixel entering the conduction state and the initial conduction time of the pixel, and calculate the gray value of each pixel at its conduction time to obtain a spatially sparse high dynamic range grayscale image. Arrange all spatially sparse high dynamic range grayscale images in chronological order to obtain a high frame rate spatially sparse grayscale image sequence. 4) Based on the high frame rate spatially sparse grayscale image sequence, a dense, high dynamic range, high-speed video is obtained through a spatiotemporal grayscale completion neural network; In step 2), the specific method for controlling the spatial light modulator is as follows: the modulation mask of the spatial light modulator is divided into multiple non-overlapping 3M*3M regions, each region containing 9 M*M blocks. Each time the spatial light modulator is controlled, an additional M*M block is added to each 3M*3M region to enter the conduction state. Step 4) specifically includes: 4.1) Downsample each spatially sparse high dynamic range grayscale image in the high frame rate spatially sparse grayscale image sequence to obtain a downsampled dense image sequence in which the length and width of the image are both 1 / 3 of the original image; then use the RAFT model to estimate the optical flow of the downsampled dense image sequence to obtain a downsampled dense optical flow in which the length and width of the image are both 1 / 3 of the original image; then upsample the downsampled dense optical flow to obtain a dense optical flow in which the length and width of the image are consistent with the original image. 4.2) Based on the relative position vector of the M*M block with gray values within its own 3M*3M region, local quantization compensation is performed on the dense optical flow of each 3M*3M region to obtain the compensated dense optical flow. 4.3) The compensated dense optical flow is used to propagate image domain information and feature domain information to the high frame rate spatially sparse grayscale image sequence to obtain a high frame rate spatially sparse grayscale image sequence after one completion. 4.4) The high frame rate spatially sparse grayscale image sequence after one completion is further transformed and completed by a mask-guided sparse transformer to obtain dense high dynamic range high speed video.
2. The method according to claim 1, characterized in that, The imaging objective lens images the external scene onto a virtual image plane. The front relay lens images the image on the virtual image plane onto the surface of the spatial light modulator. The spatial light modulator can select some pixels on its surface to conduct, thereby sending some field of view light to the rear relay lens. The rear relay lens receives the field of view light modulated by the spatial light modulator and images it onto the image plane of the event camera. The event camera collects the brightness change information on the image plane of the event camera.
3. The method according to claim 2, characterized in that, The spatial light modulator is a digital micromirror device.
4. The method according to claim 2, characterized in that, The spatial light modulator is a silicon-based liquid crystal, and the field-of-view coded imaging system further includes a first polarizer, a polarizing prism, and a second polarizer. The imaging objective lens images the external scene onto a virtual image plane. The front relay lens sequentially images the image on the virtual image plane onto the surface of the liquid crystal on silicon (LCD) through a first polarizer and a polarizing prism. The LCD changes the polarization direction of the incident light of each pixel, causing the polarization direction of the white pixels to change by 90 degrees, while the polarization direction of the black pixels remains unchanged. The polarization direction of the second polarizer is perpendicular to that of the first polarizer, causing the polarization direction of the field light modulated by the LCD to change by 90 degrees and allowing it to pass through the second polarizer. Field light whose polarization direction has not changed cannot pass through the second polarizer. The field light that passes through the second polarizer continues to the rear relay lens. The rear relay lens receives the field light modulated by the spatial light modulator and images it onto the image plane of the event camera. The event camera collects the brightness change information on the image plane of the event camera.
5. The method according to claim 1, characterized in that, In step 2), each pixel remains on for a duration of mT, and after a duration of mT, each pixel enters a duration of nT of a masking state; m and n are preset values; T is the modulation period, which is the time difference between two adjacent modulations of the spatial light modulator.
6. The method according to claim 1, characterized in that, In step 3), the grayscale value of each pixel at its on-time is calculated as follows: Where I(x,y) is the grayscale value of pixel (x,y) when pixel (x,y) is turned on; C PD The capacitance of the photodiode in the event camera; V ref ε is the reference voltage of the event camera, which is 0; TR(t) is the intensity modulation function, and ε(t) is the unit step function that jumps to 1 at t = 0; t * (x,y) is the time difference between the timestamp of the first positive event after pixel (x,y) is turned on and the time of turn-on, and the time of turn-on is defined as 0; V thd The preset threshold voltage for the event camera.
7. The method according to claim 1, characterized in that, In step 4.1), the process of downsampling each spatially sparse high dynamic range grayscale image in the high frame rate spatially sparse grayscale image sequence to obtain a downsampled dense image sequence whose length and width are both 1 / 3 of the original image is as follows: For each spatially sparse high dynamic range grayscale image, create a new matrix whose length and width are both 1 / 3 of the original image. Then, select M*M blocks with grayscale values within each 3M*3M region of the spatially sparse high dynamic range grayscale image and place them into the new matrix of the spatially sparse high dynamic range grayscale image. The spatial position of the blocks is the same as the position of the 3M*3M region in the original spatially sparse high dynamic range grayscale image. This results in a new image whose length and width are both 1 / 3 of the original image. Arrange the new images in chronological order to obtain a downsampled dense image sequence whose length and width are both 1 / 3 of the original image.
8. A high dynamic range, high-speed imaging system for implementing the method of claim 1, characterized in that, It includes a field-of-view coded imaging system, a spatially sparse grayscale image acquisition module, and an imaging video acquisition module; The field-of-view coded imaging system is used to obtain the timestamp of the first positive event triggered by each pixel entering the conduction state; The spatially sparse grayscale image acquisition module first calculates the time difference between the timestamp of the first positive event triggered by the pixel entering the conduction state and the initial conduction time of the pixel, and calculates the grayscale value of each pixel at its conduction time, and finally obtains a high frame rate spatially sparse grayscale image. The imaging video acquisition module obtains dense, high-dynamic, high-speed video based on the high frame rate spatially sparse grayscale image through a spatiotemporal grayscale completion neural network.