An event-driven optical sectioning microscopic image temporal super-resolution method and system

By employing an adaptive fusion method that integrates sparse event streams and frame-based image sequences, the problem of insufficient temporal resolution in light sheet microscopy is solved, achieving high-quality temporal super-resolution reconstruction and improving the reconstruction accuracy of rapid motion and brightness changes in dynamic scenes.

CN122434735APending Publication Date: 2026-07-21HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2026-05-07
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing light-sheet microscopy techniques struggle to improve temporal resolution under conditions of high spatial resolution and high signal-to-noise ratio, limiting their application in high-speed microscopy.

Method used

By employing an event-driven temporal super-resolution method for light sheet microscopy images, a sparse event stream and a frame-based image sequence are fused. The self-attention principle is used to adaptively fuse image-event data, reconstructing images at unsampled moments, improving the frame rate and maintaining structural continuity.

Benefits of technology

It achieves efficient recovery of temporally missing axial images without increasing the physical sampling rate, improving imaging speed and image quality, and solving the problems of motion artifacts and structural distortion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434735A_ABST
    Figure CN122434735A_ABST
Patent Text Reader

Abstract

The application discloses a kind of light sheet microscopic image time super-resolution method and system based on event driving.There is described as follows in the application: (1) the event data and image data of biological sample are registered and aligned on space-time, and combined as image-event data;(2) according to image event data, preliminary intermediate frame, light flow refinement intermediate frame and synthetic refinement intermediate frame reconstruction are carried out respectively;(3) preliminary intermediate frame, light flow refinement intermediate frame and synthetic refinement intermediate frame are adaptively fused based on self-attention principle.The method of the application can efficiently fuse the event stream containing high temporal resolution and the image frame containing high spatial resolution, realize high-quality time super-resolution reconstruction, and improve the reconstruction accuracy of fast motion and brightness change in dynamic scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of high-speed dynamic microscopy imaging, and more specifically, relates to an event-driven temporal super-resolution method and system for light sheet microscopy images. Background Technology

[0002] High-speed microscopy plays an irreplaceable role in revealing ultrafast dynamic life processes at the cellular and subcellular scales, and has become a core observational tool in cutting-edge biomedical fields such as neural activity detection, biomolecular dynamics research, organelle transport analysis, and microcirculation observation. Many life activities, such as the propagation of neuronal action potentials, vascular endothelial responses, or protein conformational changes, often occur on sub-millisecond or even microsecond timescales. This places stringent demands on the performance of imaging systems: they must maintain high spatial resolution to resolve fine biological structures while possessing temporal resolution above kilohertz, i.e., extremely high temporal sampling capability, in order to capture the complete dynamic trajectory of these transient biological events without distortion.

[0003] Light-sheet microscopy typically relies on axial scanning to acquire sample volume information layer by layer during three-dimensional volume imaging. Its imaging speed is limited by factors such as camera frame rate, scanning sampling rate, light-sheet thickness, and field of view. When maintaining high resolution with a large field of view is required, the system often faces a trade-off between speed, signal-to-noise ratio, and resolution. To improve volume imaging speed, some studies have used single-objective tilting light-sheet scanning structures, achieving rapid scanning of the light-sheet within the sample using a high-speed galvanometer, obtaining 3D images of hundreds of volumes per second without requiring mechanical movement of the sample or objective. Other studies have used controlled downsampling in the spatial domain to preserve the frequency domain signal in a recoverable aliasing form, achieving a 2 to 4-fold increase in imaging speed without compromising spatial resolution after frequency rearrangement. Furthermore, some methods employ optical shearing or multi-angle projection imaging, integrating information from different angles in a single exposure to achieve near real-time observation of dynamic processes.

[0004] Although these methods have achieved breakthroughs in the imaging rate of light-sheet scanning microscopy to varying degrees, they are still limited by the physical limitations of CMOS cameras. In particular, under the requirements of high spatial resolution and high signal-to-noise ratio imaging technology, the temporal resolution of light-sheet microscopy is difficult to improve further, which limits its promotion and application in the field of high-speed microscopy. Summary of the Invention

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides an event-driven temporal super-resolution method and system for light sheet microscopy images. Its purpose is to recover missing axial images in a three-dimensional scene by deeply fusing sparse event streams and frame-based image sequences, without increasing the physical sampling rate. This improves the frame rate while maintaining structural continuity, providing a feasible approach for observing high-speed dynamic processes. This solves the problem that traditional interpolation algorithms, due to the low signal-to-noise ratio of fluorescence images and the complexity of dynamic content, struggle to accurately reconstruct the true image at unsampled moments, easily introducing motion artifacts or structural distortion.

[0006] To achieve the above objectives, according to one aspect of the present invention, an event-driven temporal super-resolution method for light sheet microscopy images is provided, comprising the following steps: (1) Register and align event data and image data of biological samples in time and space. Discretize event data in the time dimension and accumulate it with weights to obtain voxel grid data, so that image data and voxel grid data correspond in the same temporal and spatial coordinate system, thus combining them into image-event data. (2) Based on the image event data obtained in step (1), perform preliminary intermediate frame reconstruction, optical flow thinning intermediate frame reconstruction, and synthetic thinning intermediate frame reconstruction respectively; (3) The preliminary intermediate frame, optical flow thinning intermediate frame and synthetic thinning intermediate frame obtained in step (2) are adaptively fused based on the self-attention principle to generate interpolated images and achieve temporal super-resolution.

[0007] Preferably, the event-driven temporal super-resolution method for light sheet microscopy images uses a voxel grid denoted as V(x, y, t), where the spatial dimensions x and y directly correspond to the location of the event, and the temporal dimension t is discretized by binning the time axis. Specifically: For the time interval [T] k , T k+1 ], use time node T to represent this time period. k, j Let j = 1, 2, 3, …, n-1 be uniformly divided into n time bins, i.e., n consecutive time sub-intervals {[T k , T k, 1 ],[T k, 1 , T k, 2 ], …, [T k, n-1 ,T k+1 ]}, where n is the expected time super-resolution factor; For any event (x, y, t, Its occurrence time t is located in [T k , T k+1 Within this scope, it is first mapped to consecutive positions on the time box axis: a=(tT k ) / (T k+1 - T k )*n Where 'a' is a continuous real number representing the relative position of the event on the time-box axis; Voxel value V(x, y, T) k, i The specific calculation method for i=1, 2, …, n is as follows:

[0008] in, This represents the polarity of the event.

[0009] Preferably, in the event-driven light sheet microscopy image temporal super-resolution method, the optical flow refinement narrows the time interval [T]... k , T k+1 The image-event data within the time frame is used to obtain event correlates by extracting the correlation of local distribution features between different time points, and image correlates by extracting the correlation of image features at the beginning and end of the time interval. These correlates are used to characterize the motion trend of pixels. The destination simulation pixel temporal trajectory is used to reconstruct n-1 time points T. k, j Images j=1, 2, 3, …, n-1 are used as intermediate frames for optical flow thinning.

[0010] Preferably, the optical flow refinement in the event-driven temporal super-resolution method for light sheet microscopy images is as follows: A1. For the time interval [T] k , T k+1 For the image-event data within the time frame, the entire voxel grid set is taken as the context grid ContextGrid. J sub-grids are extracted from the entire voxel grid as the relevant grid CorrGrid in the form of a sliding window on the time axis, as follows: ContextGrid= concat((V(x,y, t=T k, 1 V(x,y, t=T) k, 2 ), V(x,y, t=T) k+1 )),dim=t) CorrGrid = concat((V(x,y, t=T)) k,a V(x,y, t=T) k, a+1 ), V(x,y, t=T) k, a+J-1 )),dim=t) Where a = 1, 2, ..., n-J+1; A2, Image I k and I k+1 The context grid (ContextGrid) and related grid (CorrGrid) obtained in step A1 are respectively feature encoded, wherein the image I... k and I k+1 Convolutional encoders are used to extract static information features F, including texture and edge information. I,k and F I,k+1 A convolutional encoder is used to extract global spatiotemporal features F from the context grid ContextGrid. T To depict the overall motion trend; a convolutional encoder is used on the relevant CorrGrid to extract local distribution features F to characterize the distribution of local events at different time points. E ={F1,F2,…,F g}, where g is the number of related grids; A3. Measure the local distribution characteristics F obtained in step A2. E Autocorrelation in time series constructs event correlation C E Measure image features F I,k and F I,k+1 Constructing image correlation C based on correlation I The details are as follows:

[0011]

[0012] Where <·> represents the inner product operation to calculate correlation, which is used to measure the degree of matching between features and establish the correspondence between pixels in the time dimension.

[0013] A4. Obtain the event-related entity C based on step A3. E Image-related body C I According to the event-related body C E Image-related body C I Following the principle that stronger correlation in characterization requires fewer control points, the number of control points for the curve is determined, and the pixel temporal trajectory curve is simulated based on the contextual features F obtained in step A2. T Constraining the pixel grayscale update amplitude, iteratively updating the temporal trajectory curve to obtain the reconstructed n-1 time nodes T k, j The image of j=1, 2, 3, …, n-1 is used as its optical flow thinning image.

[0014] Preferably, in the event-driven temporal super-resolution method for light sheet microscopy images, step A4 uses Catmull-Rom curves to simulate the temporal trajectory of each pixel; specifically as follows: A temporal trajectory modeling method based on Catmull-Rom curves is introduced. For each pixel (x, y), its temporal trajectory is represented as:

[0015] in For curve control point parameters, Catmull-Rom basis functions; After initializing the curve control points, according to the event-related body C E Image-related body C I The increment of control points for the (x, y) temporal trajectory curve of a pixel is determined by the principle that the stronger the correlation of the representation, the fewer the number of control points.

[0016] in, For incremental mapping networks of time-series trajectories and control points of trajectory curves, a relatively simple implementation can be achieved using a two-layer convolutional network. This is the current time-series trajectory curve; Utilizing contextual features F T Generate confidence weights to adjust the update magnitude during each pixel iteration. :

[0017] in, The confidence mapping function, also implemented by a neural network, has an output value range of [0,1] and is used to reflect the update reliability at different spatial locations. This represents element-wise multiplication; According to the update range The optical flow is estimated by iteratively updating the pixels of each control point and reconstructing the time node image to obtain the optical flow-refined intermediate frame.

[0018] Preferably, in the event-driven light sheet microscopy image temporal super-resolution method, the preliminary intermediate frame reconstruction divides the time interval [T]... k , T k+1 The image-event data within the [database name] is input into the reconstruction network, and the output images are used as n-1 time nodes T. k, j The initial intermediate frames corresponding to the intermediate frame images j=1, 2, 3, …, n-1; and / or The synthesis refinement, from the time interval [T] k , T k+1 From the image-event data within the [database], extract image features at the start and end times, as well as the time node T. k, jThe voxel mesh data within a preset range of j=1, 2, 3, …, n-1 is used as the local event voxel mesh, encoded, and then input into the encoder-decoder to reconstruct the time node T. k, j The image is used as its composite thinned image.

[0019] Preferably, in the event-driven temporal super-resolution method for light sheet microscopy images, the synthesis and refinement are performed for any time node T. k, j The details are as follows: B1. For the time interval [T] k , T k+1 Image-event data within [the image / event data], taking time node T. k, j The voxel grid with a preset time window width nearby is a local event voxel grid, denoted as {V(x, y, ti), |t-ti|<δ}, where δ is the time window width; the local event voxel grid uses fragment event input to reduce event redundancy and polarity ambiguity caused by event accumulation; B2. The time interval [T] k , T k+1 Images of the start and end times I k and I k+1 The local event voxel grid {V(x, y, ti), |t-ti|<δ} extracted in step B1 is used for feature fusion and input into the encoder-decoder for image reconstruction to obtain the time node T. k, j The reconstructed image is used as the synthetic refinement intermediate frame at that time point.

[0020] According to another aspect of the present invention, an event-driven time-based super-resolution system for light sheet microscopy images is provided, comprising: The registration and alignment module is used to register and align event data and image data of biological samples in time and space. In the time dimension, the event data is binned, discretized, and weighted to obtain voxel grid data, so that the image data and voxel grid data correspond in the same temporal and spatial coordinate system, thereby combining them into image-event data. The image-event data is then submitted to the coarse frame generation module, the optical flow thinning module, and the synthesis thinning module, respectively. The coarse frame generation module is used to generate time intervals [T] k , T k+1 The image-event data within the [database name] is input into the reconstruction network, and the output images are used as n-1 time nodes T. k, j The preliminary intermediate frames corresponding to the intermediate frame images j=1, 2, 3, …, n-1 are submitted to the fusion module; The optical flow refinement module includes a control point incremental mapping network, a confidence mapping function network, and an optical flow estimation network; it is used to refine the optical flow over the time interval [T]. k , Tk+1 Image-event data within the time interval [T] is processed by extracting the correlation of local distribution features between different time points to obtain event correlates, and extracting the correlation of image features at the start and end of the time interval to obtain image correlates. These correlates are then mapped to the current time-series trajectory curve and input into a control point increment mapping network to obtain control point increments. The time interval [T] is then further processed. k , T k+1 Event data in the image-event data within the image is encoded as contextual features F. T The input confidence mapping function network obtains the update reliability at different spatial locations, and multiplies it with the control point increment to obtain the update magnitude at each iteration update; the optical flow estimation network iteratively updates the temporal trajectory based on the control point increment and the update magnitude at each location, reconstructs the optical flow refinement intermediate frames, and submits them to the fusion module. The synthesis refinement module is used to refine the data from the time interval [T]. k , T k+1 From the image-event data within the [database], extract image features at the start and end times, as well as the time node T. k, j The voxel mesh data within a preset range of j=1, 2, 3, …, n-1 is used as the local event voxel mesh, encoded, and then input into the encoder-decoder to reconstruct the time node T. k, j The image is submitted to the fusion module as its synthesized and refined image; The fusion module is used to adaptively fuse the initial intermediate frame, the optical flow-refined intermediate frame, and the synthetic refined intermediate frame based on the self-attention principle, thereby generating an interpolated image and achieving temporal super-resolution.

[0021] According to another aspect of the present invention, a training method for the event-driven light sheet microscopy image temporal super-resolution system is provided, comprising the following steps: S1. Constructing training data: A high temporal resolution continuous image sequence is acquired using an sCMOS camera, while an event camera synchronously records event stream data within the corresponding time period. The acquired sCMOS image sequence is downsampled along the time axis, and the downsampled image frames are used as input image frames. The unselected image frames in the downsampling are used as the ground truth for super-resolution interpolated images to construct training samples. S2. Iteratively update system parameters and use end-to-end training to minimize the loss function.

[0022] Preferably, the training method employs a weighted combination of multiple loss functions to assess end-to-end model performance, with the overall loss function as follows:

[0023] in, The brightness alignment loss is composed of both the basic pixel reconstruction loss and the brightness alignment loss. , and These are perceptual similarity loss, structural similarity loss, and edge gradient loss, used to enhance the perceptual fidelity of the model at the level of details, texture, and contour. , , Its weight.

[0024] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: The method of this invention can efficiently fuse event streams containing high temporal resolution and image frames containing high spatial resolution to achieve high-quality temporal super-resolution reconstruction and improve the reconstruction accuracy of fast motion and brightness changes in dynamic scenes. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the image-event data structure provided by the present invention; Figure 2 This is a flowchart of the time-driven light sheet microscopy image temporal super-resolution system provided in an embodiment of the present invention; Figure 3 This is a diagram illustrating the effect of temporal super-resolution on microscopic images in an embodiment of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0027] The event-driven temporal super-resolution method for light sheet microscopy provided by this invention is applied to an event-single-objective light sheet imaging device. The single-objective light sheet imaging device achieves scanning through a galvanometer and a rotating disk mirror. The scanning speed is the highest among light sheet imaging devices, far exceeding that of other light sheet modes that achieve scanning through a displacement stage. It is difficult to further improve the temporal resolution by increasing the frequency of mechanical movement.

[0028] The probe light of the event-single-objective light-sheet system is split into light-sheet probe light and event probe light by energy splitting, with a recommended splitting ratio of 50:50. The light-sheet probe light and event probe light are then sent to the sCMOS camera imaging and the event imaging camera, respectively, to achieve light-sheet-time synchronous imaging. Includes the following steps: (1) Register and align event data and image data of biological samples in space and time. Discretize and weight the event data in the time dimension to obtain voxel grid data, so that the image data and voxel grid data correspond in the same temporal and spatial coordinate system, thus combining them into image-event data, such as Figure 1 As shown.

[0029] In the time series {T k Image data {I} on the x, k=1, 2, 3……}. k Let I, k=1, 2, 3……} be the sequence of two-dimensional images captured sequentially by the light-sheet imaging device, where I k For T k The grayscale image captured by the sCMOS camera and captured by the time-lapse imaging device; the corresponding event data {E k :E k-k+1 k=1, 2, 3……}, that is, T k To T k+1 Optical flow data E captured by the time-event imaging camera t , t∈[T k , T k+1 ], where event data E t (x, y, t, p) includes the event coordinates (x, y), the event timestamp t, and the polarity p, which takes values ​​of 0 and 1.

[0030] Because the temporal resolution of event cameras is on the order of microseconds, their timestamp accuracy is far higher than that of image frames. Therefore, a large amount of event data can be observed between adjacent frames. This event data is asynchronous and unstructured, resembling a 3D point cloud, and cannot be directly used as input to deep learning networks. Therefore, it is necessary to convert the events into a regular, synchronous intermediate representation, namely a voxel grid.

[0031] A voxel grid V(x, y, t) is used, where the spatial dimensions x and y directly correspond to the location of the event, and the temporal dimension t is discretized by binning the time axis. Specifically: For the time interval [T] k , T k+1 ], use time node T to represent this time period. k, j Let j = 1, 2, 3, …, n-1 be uniformly divided into n time bins, i.e., n consecutive time sub-intervals {[T k , T k, 1 ],[T k, 1 , T k, 2 ], …, [T k, n-1 ,T k+1]}, where n is the desired time super-resolution factor; for example, when the super-resolution factor is 10, that is, [T1, T2] is divided into 10 consecutive time sub-intervals {[T1, T2]}. k, 1 ],[T k, 1 , T k, 2 ], …, [T 1, 9 , T2]}.

[0032] For any event (x, y, t, p), its occurrence time t lies within [T]. k , T k+1 Within this scope, it is first mapped to consecutive positions on the time box axis: a=(tT k ) / (T k+1 - T k )*n Here, 'a' is a continuous real number representing the relative position of the event on the timebox axis. For example, when a = 3.4, it means that the event occurs between the 3rd and 4th timeboxes.

[0033] To avoid rigidly assigning an event to a single timebox, a weighted cumulative approach is used to simultaneously allocate the event to its two adjacent timeboxes. Let n1 and n2 be the timebox indices to which the event will be assigned, respectively, and equal to the floor of a. Then, the weights of the event's contribution to these two timeboxes are as follows: W n1 =1-(a-n1) W n2 =1-(n²-a) The weights are inversely proportional to the distance; the greater the distance, the smaller the contribution. Therefore, the weight needs to be subtracted by 1. The event polarity p is accumulated into the voxel grid according to the above weights. During this process, the event's position (x, y) in the spatial dimension remains unchanged; only the time dimension is weighted according to its timestamp. The voxel value V(x, y, T) k, i The specific calculation method for i=1, 2, …, n is as follows:

[0034] Ultimately, the accumulated values ​​in each voxel grid cell reflect the events that occurred at the corresponding spatial location and time interval. Image-event data is represented as images at both ends of the time interval and event voxel grids at the middle time.

[0035] Every two frames of images I k and I k+1 Corresponding to a time interval [T] k , T k+1The event data generated within this interval is converted into an event voxel grid and compared with the image data I. k and I k+1 Both are used as inputs to the reconstruction network to reconstruct the time interval [T]. k , T k+1 Between n-1 time points T k, j The intermediate frame images, j=1, 2, 3, …, n-1.

[0036] Using the above method, the originally asynchronous, sparse and irregular event stream is transformed into a regular and dense three-dimensional voxel grid representation, making it similar to traditional images in terms of data form. This allows it to be integrated with image data, making it convenient to use image-event data as input for subsequent deep learning network processing.

[0037] (2) Based on the image event data obtained in step (1), perform preliminary intermediate frame reconstruction, optical flow thinning intermediate frame reconstruction, and synthetic thinning intermediate frame reconstruction respectively; The initial intermediate frame will be the time interval [T] k , T k+1 The image-event data within the [database name] is input into the reconstruction network, and the output images are used as n-1 time nodes T. k, j The initial intermediate frames are the intermediate frame images corresponding to j=1, 2, 3, …, n-1. The U-Net network is preferably used as the reconstruction network. U-Net is a network architecture that is widely used and effective in image reconstruction and generation tasks. It has a simple structure, stable performance, and good feature representation ability, and is suitable for the initial restoration of intermediate frames.

[0038] The optical flow refinement extends the time interval [T] k , T k+1 The image-event data within the time frame is used to obtain event correlates by extracting the correlation of local distribution features between different time points, and image correlates by extracting the correlation of image features at the beginning and end of the time interval. These correlates are used to characterize the motion trend of pixels. The destination simulation pixel temporal trajectory is used to reconstruct n-1 time points T. k, j Images j=1, 2, 3, …, n-1 are used as intermediate frames for optical flow thinning; specifically as follows: A1. For the time interval [T] k , T k+1 For the image-event data within the time frame, the entire voxel grid set is taken as the context grid ContextGrid. J sub-grids are extracted from the entire voxel grid as the relevant grid CorrGrid in the form of a sliding window on the time axis, as follows: ContextGrid= concat((V(x,y, t=T k, 1 V(x,y, t=T) k, 2 ), V(x,y, t=T) k+1 )),dim=t) CorrGrid = concat((V(x,y, t=T)) k,a V(x,y, t=T) k, a+1 ), V(x,y, t=T) k, a+J-1 )),dim=t) Where a = 1, 2, ..., n-J+1.

[0039] A2, Image I k and I k+1 The context grid (ContextGrid) and related grid (CorrGrid) obtained in step A1 are respectively feature encoded, wherein the image I... k and I k+1 Convolutional encoders are used to extract static information features F, including texture and edge information. I,k and F I,k+1 A convolutional encoder is used to extract global spatiotemporal features F from the context grid ContextGrid. T To depict the overall motion trend; a convolutional encoder is used on the relevant CorrGrid to extract local distribution features F to characterize the distribution of local events at different time points. E ={F1,F2,…,F g}, where g is the number of related grids.

[0040] A3. Measure the local distribution characteristics F obtained in step A2. E Autocorrelation in time series constructs event correlation C E Measure image features F I,k and F I,k+1 Constructing image correlation C based on correlation I The details are as follows:

[0041]

[0042] Where <·> represents the inner product operation to calculate correlation, which is used to measure the degree of matching between features and establish the correspondence between pixels in the time dimension.

[0043] A4. Obtain the event-related entity C based on step A3. E Image-related body C IAccording to the event-related body C E Image-related body C I Following the principle that stronger correlation in characterization requires fewer control points, the number of control points for the curve is determined, and the pixel temporal trajectory curve is simulated based on the contextual features F obtained in step A2. T Constraining the pixel grayscale update amplitude, iteratively updating the temporal trajectory curve to obtain the reconstructed n-1 time nodes T k, j Images of pixels j=1, 2, 3, …, n-1 are used as their optical flow thinning images; the preferred approach is to use Catmull-Rom curves to simulate the temporal trajectory of each pixel; specifically as follows: To characterize the continuous structure in the sheet data and effectively address the temporal inconsistency caused by the emergence of new structures and the disappearance of old structures, a temporal trajectory modeling method based on Catmull-Rom curves is introduced. For each pixel (x, y), its temporal trajectory is represented as:

[0044] in For curve control point parameters, These are Catmull-Rom basis functions.

[0045] After initializing the curve control points, according to the event-related body C E Image-related body C I The increment of control points for the (x, y) temporal trajectory curve of a pixel is determined by the principle that the stronger the correlation of the representation, the fewer the number of control points.

[0046] in, For incremental mapping networks of time-series trajectories and control points of trajectory curves, a relatively simple implementation can be achieved using a two-layer convolutional network. This is the current time-series trajectory curve.

[0047] Utilizing contextual features F T Generate confidence weights to adjust the update magnitude during each pixel iteration. :

[0048] in, The confidence mapping function, also implemented by a neural network, has an output value range of [0,1] and is used to reflect the update reliability at different spatial locations. This indicates element-wise multiplication.

[0049] According to the update range The optical flow is estimated by iteratively updating the pixels of each control point and reconstructing the time node image to obtain the optical flow-refined intermediate frame.

[0050] Because the Catmull-Rom curve possesses the characteristics of piecewise control, local adjustment, and global continuity, changes at local points do not affect the overall shape of the curve and can ensure smooth transitions between time periods. Therefore, it effectively avoids the trajectory discontinuities and matching errors that occur in traditional optical flow estimation in 3D axial scenes. This mechanism enables the model to achieve smooth and stable pixel trajectory modeling in complex nonlinear motion scenes, thereby improving the accuracy and temporal consistency of optical flow estimation.

[0051] The synthesis refinement, from the time interval [T] k , T k+1 From the image-event data within the [database], extract image features at the start and end times, as well as the time node T. k, j The voxel mesh data within a preset range of j=1, 2, 3, …, n-1 is used as the local event voxel mesh, encoded, and then input into the encoder-decoder to reconstruct the time node T. k, j The image is used as its synthesized thinned image; for any time node T k, j The details are as follows: B1. For the time interval [T] k , T k+1 Image-event data within [the image / event data], taking time node T. k, j The voxel grid with a preset time window width nearby is a local event voxel grid, denoted as {V(x, y, ti), |t-ti|<δ}, where δ is the time window width; the local event voxel grid uses fragment event input to reduce event redundancy and polarity ambiguity caused by event accumulation; B2. The time interval [T] k , T k+1 Images of the start and end times I k and I k+1 The local event voxel grid {V(x, y, ti), |t-ti|<δ} extracted in step B1 is used for feature fusion and input into the encoder-decoder for image reconstruction to obtain the time node T. k, j The reconstructed image is used as the intermediate frame for synthesis and refinement at that time point; specifically as follows: Image I k and I k+1 The local event voxel grid {V(x, y, ti) extracts multi-scale image features and event features through encoders with the same structure but different parameters. The image features and event features at the same scale are stitched together and fused. The decoder takes the fused features as input, upsamples step by step to restore the spatial resolution, and combines the fused features at different scales to gradually refine the reconstruction results, generating a synthesized intermediate frame image as the synthesized refined intermediate frame.

[0052] Existing methods for achieving image super-resolution based on event cameras and traditional CMOS cameras, such as the TimeLens model, often employ a network framework centered on optical flow estimation, or even rely entirely on optical flow information for temporal frame interpolation. Currently, the performance of super-resolution methods based on event cameras is highly dependent on the accuracy of optical flow estimation.

[0053] However, optical flow estimation is a complex and highly nonlinear process, and the traditional two-dimensional optical flow assumption is even more difficult to hold for light sheet microscopy imaging scenarios.

[0054] The traditional two-dimensional optical flow hypothesis assumes that the continuity of biological samples, like the continuity of object motion, is continuous throughout the entire process, possessing a completely continuous trajectory. While it can effectively reconstruct continuous temporal information in conventional scenarios...

[0055] However, the time series of light sheet microscopy data not only includes two-dimensional motion within the plane but also involves three-dimensional structural changes along the optical axis. Therefore, when expanding from two-dimensional motion interpolation to three-dimensional axial interpolation, the scene characteristics undergo a fundamental transformation. In three-dimensional axial imaging, the inter-frame correspondence is no longer a strict point-to-point mapping but may exhibit a complex relationship of one point corresponding to multiple points or multiple points corresponding to one point. Furthermore, situations such as structural disappearance and new structural generation frequently occur. For example, there are obvious interfaces between submicroscopic structures of organs, cells, and organelles, resulting in discontinuous image light intensity, exhibiting a "step" change.

[0056] Therefore, optical flow estimation based on the assumption of two-dimensional optical flow cannot accurately simulate the temporal optical flow variations in light sheet microscopy data involving three-dimensional structural changes along the optical axis, especially in regions with large dynamic range, high noise, or low texture, where the optical flow field often struggles to converge stably. Due to the inevitable biases in optical flow estimation, the overall frame interpolation quality of event-camera-based super-resolution models will significantly degrade, resulting in motion artifacts, blurred edges, and temporal inconsistencies. Existing frameworks that use optical flow modules as initial or core network components are highly susceptible to training instability or convergence difficulties on such data.

[0057] (3) The preliminary intermediate frames, optical flow-refined intermediate frames, and synthetic refined intermediate frames obtained in step (2) are adaptively fused based on the self-attention principle to generate T. k, i Interpolated frame image I at time k, i i=1, 2, ..., n-1, to achieve temporal super-resolution.

[0058] The event-driven temporal super-resolution system for light sheet microscopy provided by this invention includes: The registration and alignment module is used to register and align event data and image data of biological samples in time and space. In the time dimension, the event data is binned, discretized, and weighted to obtain voxel grid data, so that the image data and voxel grid data correspond in the same temporal and spatial coordinate system, thereby combining them into image-event data. The image-event data is then submitted to the coarse frame generation module, the optical flow thinning module, and the synthesis thinning module, respectively. The coarse frame generation module preferably uses a U-Net network as the reconstruction network to generate frames within the time interval [T]. k ,T k+1 The image-event data within the [database name] is input into the reconstruction network, and the output images are used as n-1 time nodes T. k, j The preliminary intermediate frames corresponding to the intermediate frame images j=1,2,3,…,n-1 are submitted to the fusion module; The optical flow refinement module includes a control point incremental mapping network, a confidence mapping function network, and an optical flow estimation network; it is used to refine the optical flow over the time interval [T]. k , T k+1 Image-event data within the time interval [T] is processed by extracting the correlation of local distribution features between different time points to obtain event correlates, and extracting the correlation of image features at the start and end of the time interval to obtain image correlates. These correlates are then mapped to the current time-series trajectory curve and input into a control point increment mapping network to obtain control point increments. The time interval [T] is then further processed. k , T k+1 Event data in the image-event data within the image is encoded as contextual features F. T The input confidence mapping function network obtains the update reliability at different spatial locations, and multiplies it with the control point increment to obtain the update magnitude at each iteration update; the optical flow estimation network iteratively updates the temporal trajectory based on the control point increment and the update magnitude at each location, reconstructs the optical flow refinement intermediate frames, and submits them to the fusion module.

[0059] The synthesis refinement module preferably employs a codec to refine data from the time interval [T]. k , T k+1 From the image-event data within the [database], extract image features at the start and end times, as well as the time node T. k, j The voxel mesh data within a preset range of j=1, 2, 3, …, n-1 is used as the local event voxel mesh, encoded, and then input into the encoder-decoder to reconstruct the time node T. k, j The image is submitted to the fusion module as its synthesized and refined image.

[0060] The fusion module is used to adaptively fuse the initial intermediate frame, the optical flow-refined intermediate frame, and the synthetic refined intermediate frame based on the self-attention principle, thereby generating an interpolated image and achieving temporal super-resolution.

[0061] The event-driven time-based super-resolution system for light sheet microscopy images is trained according to the following method: S1. Constructing training data: A high temporal resolution continuous image sequence is acquired using an sCMOS camera, while an event camera synchronously records event stream data within the corresponding time period. The acquired sCMOS image sequence is downsampled along the time axis. The downsampled image frames are used as input image frames, and the unselected image frames from the downsampled sequence are used as ground truth supervisory images for super-resolution interpolation, thus constructing training samples.

[0062] S2. Iteratively update system parameters and use end-to-end training to minimize the loss function; A weighted combination of multiple loss functions is preferred for end-to-end model performance. The overall loss function is as follows:

[0063] in, The brightness alignment loss is composed of both the basic pixel reconstruction loss and the brightness alignment loss. , and These are perceptual similarity loss, structural similarity loss, and edge gradient loss, used to enhance the perceptual fidelity of the model at the level of details, texture, and contour. , , Its weight; The specific calculation method for brightness alignment loss is as follows:

[0064] Among them, weight The parameters, which change dynamically during training, are obtained by calculating the Bach distance between the brightness distributions of the reconstructed and ground truth frames. This distance is used to adaptively balance the weights of the two types of loss terms during training. Since the event camera only records the polarity of brightness changes and not the magnitude, directly supervising pixel intensity makes stable matching difficult. Therefore, a reconstruction loss... Brightness distribution similarity is used as a constraint to guide the model to learn brightness consistency in a statistical manner.

[0065] When the brightness difference between the reconstructed frame and the real frame is large, the weights... The loss function is relatively large, prioritizing reconstruction accuracy, causing the model to learn the basic mapping between global brightness and structure; as training progresses, the brightness distribution gradually converges. As the brightness distribution gradually decreases, the model's training focus automatically shifts to detail optimization and contrast restoration. Through this dynamic balancing mechanism based on brightness distribution, the model can adaptively transition from the coarse reconstruction stage to the fine structure restoration stage, achieving a more stable convergence process and higher reconstruction fidelity.

[0066] The following is an example: The system employs a cylindrical lens group and a slit to form a thin film. A dichroic mirror separates the excitation and detection light paths. A tilted third objective lens is placed in the telecentric space to ensure aberration-free, large field-of-view detection. A galvanometer is positioned on the pupil plane between the primary and secondary objectives, enabling rapid scanning of the excitation film and descanning of the fluorescence signal. Through galvanometer scanning, the system can perform rapid three-dimensional volume scanning without mechanical movement of the primary objective lens and the sample, reducing interference with the sample and significantly accelerating the imaging speed.

[0067] The event-oblique light plate system employs a mature and efficient optical path design. During detection, a 50%-50% beam splitter is used to focus the light beam onto the photosensitive surfaces of the sCMOS and the event camera, enabling simultaneous imaging by the two cameras.

[0068] This embodiment performs super-resolution microscopic images with a super-resolution magnification of 10.

[0069] The event-driven time-series super-resolution system for light sheet microscopy provided in this embodiment includes: The registration and alignment module is used to register and align event data and image data of biological samples in time and space. In the time dimension, the event data is binned, discretized, and weighted to obtain voxel grid data, so that the image data and voxel grid data correspond in the same temporal and spatial coordinate system, thereby combining them into image-event data. The image-event data is then submitted to the coarse frame generation module, the optical flow thinning module, and the synthesis thinning module, respectively. The coarse frame generation module uses the U-Net network for reconstruction, dividing the time interval [T] into... k , T k+1 The image-event data within the [database name] is input into the reconstruction network, and the output images are used as n-1 time nodes T. k, j The preliminary intermediate frames corresponding to the intermediate frame images j=1, 2, 3, …, n-1 are submitted to the fusion module; The optical flow refinement module includes a control point incremental mapping network, a confidence mapping function network, and an optical flow estimation network; it is used to refine the optical flow over the time interval [T]. k , T k+1Image-event data within the time interval [T] is processed by extracting the correlation of local distribution features between different time points to obtain event correlates, and extracting the correlation of image features at the start and end of the time interval to obtain image correlates. These correlates are then mapped to the current time-series trajectory curve and input into a control point increment mapping network to obtain control point increments. The time interval [T] is then further processed. k , T k+1 Event data in the image-event data within the image is encoded as contextual features F. T The input confidence mapping function network obtains the update reliability at different spatial locations, and multiplies it with the control point increment to obtain the update magnitude at each iteration update; the optical flow estimation network iteratively updates the temporal trajectory based on the control point increment and the update magnitude at each location, reconstructs the optical flow refinement intermediate frames, and submits them to the fusion module.

[0070] The synthesis refinement module employs an encoder-decoder to refine data from the time interval [T]. k , T k+1 From the image-event data within the [database], extract image features at the start and end times, as well as the time node T. k, j The voxel mesh data within a preset range of j=1, 2, 3, …, n-1 is used as the local event voxel mesh, encoded, and then input into the encoder-decoder to reconstruct the time node T. k, j The image is submitted to the fusion module as its synthesized and refined image.

[0071] The event-driven temporal super-resolution system for light sheet microscopy images provided in this embodiment employs end-to-end training, as detailed below: S1. Constructing training data: A high-temporal-resolution continuous image sequence is acquired using an sCMOS camera, while an event camera synchronously records event stream data within the corresponding time period. The acquired sCMOS image sequence is downsampled along the time axis. Selected downsampled image frames are used as input image frames, while unselected downsampled image frames are used as ground truth supervisory images for super-resolution interpolation, thus constructing training samples. Specifically: Select two frames with a large time interval on the time axis as input frames, for example, select time t0 and t... 10 The image frames are used as input image frames, i.e., the preceding and following reference frames; while the intermediate frames (t1 to t9) within the time interval are used as supervision ground truths to supervise and constrain the network output results.

[0072] Using the above construction method, an input consisting of two sCMOS images with a large frame interval (t0 and t1) can be obtained. 10 ), and the corresponding event data (including t0 to t 10 The event), the output is a multi-frame truth image (t1 to t9) within the time interval.

[0073] S2. Iteratively update system parameters and use end-to-end training to minimize the loss function; This embodiment uses a weighted combination of multiple loss functions to assess end-to-end model performance. The overall loss function is as follows:

[0074] in, The brightness alignment loss is composed of both the basic pixel reconstruction loss and the brightness alignment loss. , and These are perceptual similarity loss, structural similarity loss, and edge gradient loss, used to enhance the perceptual fidelity of the model at the level of details, texture, and contour. , , Its weight; The specific calculation method for brightness alignment loss is as follows:

[0075] Among them, weight The parameters, which change dynamically during training, are obtained by calculating the Bach distance between the brightness distributions of the reconstructed and ground truth frames. This distance is used to adaptively balance the weights of the two types of loss terms during training. Since the event camera only records the polarity of brightness changes and not the magnitude, directly supervising pixel intensity makes stable matching difficult. Therefore, a reconstruction loss... Brightness distribution similarity is used as a constraint to guide the model to learn brightness consistency in a statistical manner. When the value is close to 1, priority is given to ensuring the model's basic training capability. When the brightness is close, With a value close to 0, the model focuses more on restoring non-luminance factors such as texture details.

[0076] When the brightness difference between the reconstructed frame and the real frame is large, the weights... The loss function is relatively large, prioritizing reconstruction accuracy, causing the model to learn the basic mapping between global brightness and structure; as training progresses, the brightness distribution gradually converges. As the brightness distribution gradually decreases, the model's training focus automatically shifts to detail optimization and contrast restoration. Through this dynamic balancing mechanism based on brightness distribution, the model can adaptively transition from the coarse reconstruction stage to the fine structure restoration stage, achieving a more stable convergence process and higher reconstruction fidelity.

[0077] The time-driven light sheet microscopy image temporal super-resolution system provided in this embodiment is used to perform temporal super-resolution of microscopic images, such as... Figure 2 As shown, the specific steps are as follows: (1) Register and align the event data and image data of biological samples in the spatiotemporal space. Discretize the event data in the time dimension and accumulate it with weights to obtain voxel grid data, so that the image data and voxel grid data correspond in the same temporal and spatial coordinate system, thus combining them into image-event data.

[0078] In the time series {T k Image data {I} on the x, k=1, 2, 3……}. k Let I, k=1, 2, 3……} be the sequence of two-dimensional images captured sequentially by the light-sheet imaging device, where I k For T k The grayscale image captured by the sCMOS camera and captured by the time-lapse imaging device; the corresponding event data {E k :E k-k+1 k=1, 2, 3……}, that is, T k To T k+1 Optical flow data E captured by the time-event imaging camera t , t∈[T k , T k+1 ], where event data E t (x, y, t, p) includes the event coordinates (x, y), the event timestamp t, and the polarity p, which takes values ​​of 0 and 1.

[0079] A voxel grid V(x, y, t) is used, where the spatial dimensions x and y directly correspond to the location of the event, and the temporal dimension t is discretized by binning the time axis. Specifically: For the time interval [T] k , T k+1 ], use time node T to represent this time period. k, j Let j = 1, 2, 3, …, n-1 be uniformly divided into n time bins, i.e., n consecutive time sub-intervals {[T k , T k, 1 ],[T k, 1 , T k, 2 ], …, [T k, n-1 ,T k+1 ]}, where n is the desired time super-resolution factor; in this embodiment, [T1, T2] is divided into 10 consecutive time sub-intervals {[T1, T} k, 1 ],[T k, 1 , T k, 2 ], …, [T 1, 9 , T2]}.

[0080] For any event (x, y, t, p), its occurrence time t lies within [T]. k , T k+1Within this scope, it is first mapped to consecutive positions on the time box axis: a=(tT k ) / (T k+1 - T k )*n Here, 'a' is a continuous real number representing the relative position of the event on the timebox axis. For example, when a = 3.4, it means that the event occurs between the 3rd and 4th timeboxes.

[0081] To avoid rigidly assigning an event to a single timebox, a weighted cumulative approach is used to simultaneously allocate the event to its two adjacent timeboxes. Let n1 and n2 be the timebox indices to which the event will be assigned, respectively, and equal to the floor of a. Then, the weights of the event's contribution to these two timeboxes are as follows: W n1 =1-(a-n1) W n2 =1-(n²-a) The weights are inversely proportional to the distance; the greater the distance, the smaller the contribution. Therefore, the weight needs to be subtracted by 1. The event polarity p is accumulated into the voxel grid according to the above weights. During this process, the event's position (x, y) in the spatial dimension remains unchanged; only the time dimension is weighted according to its timestamp. The voxel value V(x, y, T) k, i The specific calculation method for i=1, 2, …, n is as follows:

[0082] Ultimately, the accumulated values ​​in each voxel grid cell reflect the events that occurred at the corresponding spatial location and time interval. Image-event data is represented as images at both ends of the time interval and event voxel grids at the middle time.

[0083] Every two frames of images I k and I k+1 Corresponding to a time interval [T] k , T k+1 The event data generated within this interval is converted into an event voxel grid and compared with the image data I. k and I k+1 Both are used as inputs to the reconstruction network to reconstruct the time interval [T]. k , T k+1 Between n-1 time points T k, j The intermediate frame images, j=1, 2, 3, …, n-1.

[0084] (2) Based on the image event data obtained in step (1), perform preliminary intermediate frames, optical flow thinning intermediate frames, and composite thinning intermediate frames respectively; The initial intermediate frame will be the time interval [T] k , T k+1 The image-event data within the frame is input into the U-Net network of the coarse frame generation module, and the output image is used as n-1 time nodes T. k, j The initial intermediate frames corresponding to the intermediate frame images of j=1, 2, 3, …, n-1; The optical flow refinement extends the time interval [T] k , T k+1 The image-event data within the time frame is used to obtain event correlates by extracting the correlation of local distribution features between different time points, and image correlates by extracting the correlation of image features at the beginning and end of the time interval. These correlates are used to characterize the motion trend of pixels. The destination simulation pixel temporal trajectory is used to reconstruct n-1 time points T. k, j Images j=1, 2, 3, …, n-1 are used as intermediate frames for optical flow thinning; specifically as follows: A1. For the time interval [T] k , T k+1 For the image-event data within the time frame, the entire voxel grid set is taken as the context grid ContextGrid. J sub-grids are extracted from the entire voxel grid as the relevant grid CorrGrid in the form of a sliding window on the time axis, as follows: ContextGrid= concat((V(x,y, t=T k, 1 V(x,y, t=T) k, 2 ), V(x,y, t=T) k+1 )),dim=t) CorrGrid = concat((V(x,y, t=T)) k,a V(x,y, t=T) k, a+1 ), V(x,y, t=T) k, a+J-1 )),dim=t) Where a = 1, 2, ..., n-J+1.

[0085] A2, Image I k and I k+1 The context grid (ContextGrid) and related grid (CorrGrid) obtained in step A1 are respectively feature encoded, wherein the image I... k and I k+1 Convolutional encoders are used to extract static information features F, including texture and edge information. I,k and F I,k+1A convolutional encoder is used to extract global spatiotemporal features F from the context grid ContextGrid. T To depict the overall motion trend; a convolutional encoder is used on the relevant CorrGrid to extract local distribution features F to characterize the distribution of local events at different time points. E ={F1,F2,…,F g}, where g is the number of related grids, which is 25 in this embodiment.

[0086] A3. Measure the local distribution characteristics F obtained in step A2. E Autocorrelation in time series constructs event correlation C E Measure image features F I,k and F I,k+1 Constructing image correlation C based on correlation I The details are as follows:

[0087]

[0088] Where <·> represents the inner product operation to calculate correlation, which is used to measure the degree of matching between features and establish the correspondence between pixels in the time dimension.

[0089] A4. Obtain the event-related entity C based on step A3. E Image-related body C I According to the event-related body C E Image-related body C I Following the principle that stronger correlation in characterization requires fewer control points, the number of control points for the curve is determined, and the pixel temporal trajectory curve is simulated based on the contextual features F obtained in step A2. T Constraining the pixel grayscale update amplitude, iteratively updating the temporal trajectory curve to obtain the nine reconstructed time nodes T k, j Images with j=1, 2, 3, …, 9 are used as their optical flow thinned images; in this embodiment, Catmull-Rom curves are used to simulate the temporal trajectory of each pixel; specifically as follows: To characterize the continuous structure in the sheet data and effectively address the temporal inconsistency caused by the emergence of new structures and the disappearance of old structures, a temporal trajectory modeling method based on Catmull-Rom curves is introduced. For each pixel (x, y), its temporal trajectory is represented as:

[0090] in For curve control point parameters, These are Catmull-Rom basis functions.

[0091] After initializing the curve control points, according to the event-related body C E Image-related body C I The increment of control points for the (x, y) temporal trajectory curve of a pixel is determined by the principle that the stronger the correlation of the representation, the fewer the number of control points.

[0092] in, For incremental mapping networks of time-series trajectories and control points of trajectory curves, a relatively simple implementation can be achieved using a two-layer convolutional network. This is the current time-series trajectory curve.

[0093] Utilizing contextual features F T Generate confidence weights to adjust the update magnitude during each pixel iteration. :

[0094] in, The confidence mapping function, also implemented by a neural network, has an output value range of [0,1] and is used to reflect the update reliability at different spatial locations. This indicates element-wise multiplication.

[0095] By following the update magnitude The optical flow is estimated by iteratively updating the pixels of each control point and reconstructing the time node image to obtain the optical flow-refined intermediate frame.

[0096] Because the Catmull-Rom curve possesses the characteristics of piecewise control, local adjustment, and global continuity, changes at local points do not affect the overall shape of the curve and can ensure smooth transitions between time periods. Therefore, it effectively avoids the trajectory discontinuities and matching errors that occur in traditional optical flow estimation in 3D axial scenes. This mechanism enables the model to achieve smooth and stable pixel trajectory modeling in complex nonlinear motion scenes, thereby improving the accuracy and temporal consistency of optical flow estimation.

[0097] The synthesis refinement, from the time interval [T] k , T k+1 From the image-event data within the [database], extract image features at the start and end times, as well as the time node T. k, j The voxel mesh data within a preset range of j=1, 2, 3, …, 9 is used as the local event voxel mesh, encoded, and then input into the encoder-decoder to reconstruct the time node T. k, j The image is used as its synthesized thinned image; for any time node T k, j The details are as follows: B1. For the time interval [T] k , T k+1Image-event data within [the image / event data], taking time node T. k, j The voxel grid with a preset time window width is a local event voxel grid, denoted as {V(x, y, ti), |t-ti|<δ}, where δ is the time window width, which is taken as 2 in this embodiment; B2. The time interval [T] k , T k+1 Images of the start and end times I k and I k+1 The local event voxel grid {V(x, y, ti), |t-ti|<δ} extracted in step B1 is used for feature fusion and input into the encoder-decoder for image reconstruction to obtain the time node T. k, j The reconstructed image is used as the intermediate frame for synthesis and refinement at that time point; specifically as follows: Image I k and I k+1 The local event voxel grid {V(x, y, ti) extracts multi-scale image features and event features through encoders with the same structure but different parameters. The image features and event features at the same scale are stitched together and fused. The decoder takes the fused features as input, upsamples step by step to restore the spatial resolution, and combines the fused features at different scales to gradually refine the reconstruction results, generating a synthesized intermediate frame image as the synthesized refined intermediate frame.

[0098] (3) The preliminary intermediate frames, optical flow-refined intermediate frames, and synthetic refined intermediate frames obtained in step (2) are adaptively fused based on the self-attention principle to generate T. k, i Interpolated frame image I at time k, i i=1, 2, ..., n-1, to achieve temporal super-resolution.

[0099] For the captured microscopic images of microtubes, temporal super-resolution is performed. The first and eleventh layers are taken as two input sCMOS frames. Following the method provided in this embodiment, frames are added to fill in the missing nine layers, with the same events as those from the first to the eleventh layers. The effect is as follows: Figure 3 As shown, the first row contains the two input sCMOS frames, the second row contains the real images, and the third row contains the nine layers of images reconstructed by the network, excluding layers 1 and 11.

[0100] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An event-driven temporal super-resolution method for light sheet microscopy images, characterized in that, Includes the following steps: (1) Register and align event data and image data of biological samples in time and space. Discretize event data in the time dimension and accumulate it with weights to obtain voxel grid data, so that image data and voxel grid data correspond in the same temporal and spatial coordinate system, thus combining them into image-event data. (2) Based on the image event data obtained in step (1), perform preliminary intermediate frame reconstruction, optical flow thinning intermediate frame reconstruction, and synthetic thinning intermediate frame reconstruction respectively; (3) The preliminary intermediate frame, optical flow thinning intermediate frame and synthetic thinning intermediate frame obtained in step (2) are adaptively fused based on the self-attention principle to generate interpolated images and achieve temporal super-resolution.

2. The event-driven temporal super-resolution method for light sheet microscopy images as described in claim 1, characterized in that, The voxel grid is denoted as V(x, y, t), where the spatial dimensions x and y directly correspond to the location of the event, and the temporal dimension t is discretized by binning the time axis. Specifically: For the time interval [T] k , T k+1 ], use time node T to represent this time period. k, j Let j = 1, 2, 3, …, n-1 be uniformly divided into n time bins, i.e., n consecutive time sub-intervals {[T k , T k, 1 ],[T k, 1 , T k, 2 ], …, [T k, n-1 , T k+1 ]}, where n is the expected time super-resolution factor; For any event (x, y, t, Its occurrence time t is located in [T k , T k+1 Within this scope, it is first mapped to consecutive positions on the time box axis: a=(t-T k ) / (T k+1 - T k )*n Where 'a' is a continuous real number representing the relative position of the event on the time-box axis; Voxel value V(x, y, T) k, i The specific calculation method for i=1, 2, …, n is as follows: in, This represents the polarity of the event.

3. The event-driven temporal super-resolution method for light sheet microscopy images as described in claim 2, characterized in that, The optical flow refinement extends the time interval [T] k , T k+1 The image-event data within the time frame is used to obtain event correlates by extracting the correlation of local distribution features between different time points, and image correlates by extracting the correlation of image features at the beginning and end of the time interval. These correlates are used to characterize the motion trend of pixels. The destination simulation pixel temporal trajectory is used to reconstruct n-1 time points T. k, j Images j=1, 2, 3, …, n-1 are used as intermediate frames for optical flow thinning.

4. The event-driven temporal super-resolution method for light sheet microscopy images as described in claim 3, characterized in that, The optical flow refinement is as follows: A1. For the time interval [T] k , T k+1 For the image-event data within the time frame, the entire voxel grid set is taken as the context grid ContextGrid. J sub-grids are extracted from the entire voxel grid as the relevant grid CorrGrid in the form of a sliding window on the time axis, as follows: ContextGrid= concat((V(x,y, t=T k, 1 ),V(x,y, t=T k, 2 ), ,V(x,y, t=T k+1 )), dim=t) CorrGrid= concat((V(x,y, t=T k,a ),V(x,y, t=T k, a+1 ), ,V(x,y, t=T k, a+J-1 )), dim=t) Where a = 1, 2, ..., n-J+1; A2, Image I k and I k+1 The context grid (ContextGrid) and related grid (CorrGrid) obtained in step A1 are respectively feature encoded, wherein the image I... k and I k+1 Convolutional encoders are used to extract static information features F, including texture and edge information. I,k and F I,k+1 A convolutional encoder is used to extract global spatiotemporal features F from the context grid ContextGrid. T To depict the overall motion trend; a convolutional encoder is used on the relevant CorrGrid to extract local distribution features F to characterize the distribution of local events at different time points. E ={F1,F2,…,F g }, where g is the number of related grids; A3. Measure the local distribution characteristics F obtained in step A2. E Autocorrelation in time series constructs event correlation C E Measure image features F I,k and F I,k+1 Constructing image correlation C based on correlation I The details are as follows: Where <·> represents the inner product operation to calculate correlation, which is used to measure the degree of matching between features and establish the correspondence between pixels in the time dimension; A4. Obtain the event-related entity C based on step A3. E Image-related body C I According to the event-related body C E Image-related body C I Following the principle that stronger correlation in characterization requires fewer control points, the number of control points for the curve is determined, and the pixel temporal trajectory curve is simulated based on the contextual features F obtained in step A2. T Constraining the pixel grayscale update amplitude, iteratively updating the temporal trajectory curve to obtain the reconstructed n-1 time nodes T k, j The image of j=1, 2, 3, …, n-1 is used as its optical flow thinning image.

5. The event-driven temporal super-resolution method for light sheet microscopy images as described in claim 3, characterized in that, Step A4 uses Catmull-Rom curves to simulate the temporal trajectory of each pixel; the details are as follows: A temporal trajectory modeling method based on Catmull-Rom curves is introduced. For each pixel (x, y), its temporal trajectory is represented as: in For curve control point parameters, Catmull-Rom basis functions; After initializing the curve control points, according to the event-related body C E Image-related body C I The increment of control points for the (x, y) temporal trajectory curve of a pixel is determined by the principle that the stronger the correlation of the representation, the fewer the number of control points. in, For incremental mapping networks of time-series trajectories and control points of trajectory curves, a relatively simple implementation can be achieved using a two-layer convolutional network. This is the current time-series trajectory curve; Utilizing contextual features F T Generate confidence weights to adjust the update magnitude during each pixel iteration. : in, The confidence mapping function, also implemented by a neural network, has an output value range of [0,1] and is used to reflect the update reliability at different spatial locations. This represents element-wise multiplication; According to the update range The optical flow is estimated by iteratively updating the pixels of each control point and reconstructing the time node image to obtain the optical flow-refined intermediate frame.

6. The event-driven temporal super-resolution method for light sheet microscopy images as described in claim 2, characterized in that, The preliminary intermediate frame reconstruction will reconstruct the time interval [T] k , T k+1 The image-event data within the [database name] is input into the reconstruction network, and the output images are used as n-1 time nodes T. k, j The initial intermediate frames corresponding to the intermediate frame images j=1, 2, 3, …, n-1; and / or The synthesis refinement, from the time interval [T] k , T k+1 From the image-event data within the [database], extract image features at the start and end times, as well as the time node T. k, j The voxel mesh data within a preset range of j=1, 2, 3, …, n-1 is used as the local event voxel mesh, encoded, and then input into the encoder-decoder to reconstruct the time node T. k, j The image is used as its composite thinned image.

7. The event-driven temporal super-resolution method for light sheet microscopy images as described in claim 6, characterized in that, The synthesis refinement applies to any time node T. k, j The details are as follows: B1. For the time interval [T] k , T k+1 Image-event data within [the image / event data], taking time node T. k, j The voxel grid with a preset time window width nearby is a local event voxel grid, denoted as {V(x, y, ti), |t-ti|<δ}, where δ is the time window width; the local event voxel grid uses fragment event input to reduce event redundancy and polarity ambiguity caused by event accumulation; B2. The time interval [T] k , T k+1 Images of the start and end times I k and I k+1 The local event voxel grid {V(x, y, ti), |t-ti|<δ} extracted in step B1 is used for feature fusion and input into the encoder-decoder for image reconstruction to obtain the time node T. k, j The reconstructed image is used as the synthetic refinement intermediate frame at that time point.

8. An event-driven temporal super-resolution system for light sheet microscopy images, characterized in that, include: The registration and alignment module is used to register and align event data and image data of biological samples in time and space. In the time dimension, the event data is binned, discretized, and weighted to obtain voxel grid data, so that the image data and voxel grid data correspond in the same temporal and spatial coordinate system, thereby combining them into image-event data. The image-event data is then submitted to the coarse frame generation module, the optical flow thinning module, and the synthesis thinning module, respectively. The coarse frame generation module is used to generate time intervals [T] k , T k+1 The image-event data within the [database name] is input into the reconstruction network, and the output images are used as n-1 time nodes T. k, j The preliminary intermediate frames corresponding to the intermediate frame images j=1, 2, 3, …, n-1 are submitted to the fusion module; The optical flow refinement module includes a control point incremental mapping network, a confidence mapping function network, and an optical flow estimation network; it is used to refine the optical flow over the time interval [T]. k , T k+1 Image-event data within the time interval [T] is processed by extracting the correlation of local distribution features between different time points to obtain event correlates, and extracting the correlation of image features at the start and end of the time interval to obtain image correlates. These correlates are then mapped to the current time-series trajectory curve and input into a control point increment mapping network to obtain control point increments. The time interval [T] is then further processed. k , T k+1 Event data in the image-event data within the image is encoded as contextual features F. T The input confidence mapping function network obtains the update reliability at different spatial locations, and multiplies it with the control point increment to obtain the update magnitude at each iteration update; the optical flow estimation network iteratively updates the temporal trajectory based on the control point increment and the update magnitude at each location, reconstructs the optical flow refinement intermediate frames, and submits them to the fusion module. The synthesis refinement module is used to refine the data from the time interval [T]. k , T k+1 From the image-event data within the [database], extract image features at the start and end times, as well as the time node T. k, j The voxel mesh data within a preset range of j=1, 2, 3, …, n-1 is used as the local event voxel mesh, encoded, and then input into the encoder-decoder to reconstruct the time node T. k, j The image is submitted to the fusion module as its synthesized and refined image; The fusion module is used to adaptively fuse the initial intermediate frame, the optical flow-refined intermediate frame, and the synthetic refined intermediate frame based on the self-attention principle, thereby generating an interpolated image and achieving temporal super-resolution.

9. The training method for an event-driven temporal super-resolution system for light sheet microscopy images as described in claim 8, characterized in that, Includes the following steps: S1. Constructing training data: A high temporal resolution continuous image sequence is acquired using an sCMOS camera, while an event camera synchronously records event stream data within the corresponding time period. The acquired sCMOS image sequence is downsampled along the time axis, and the downsampled image frames are used as input image frames. The unselected image frames in the downsampling are used as the ground truth for super-resolution interpolated images to construct training samples. S2. Iteratively update system parameters and use end-to-end training to minimize the loss function.

10. The training method as described in claim 9, characterized in that, A weighted combination of multiple loss functions is used to assess end-to-end model performance. The overall loss function is as follows: in, The brightness alignment loss is composed of both the basic pixel reconstruction loss and the brightness alignment loss. , and These are perceptual similarity loss, structural similarity loss, and edge gradient loss, used to enhance the perceptual fidelity of the model at the level of details, texture, and contour. , , Its weight.