Single-molecule super-resolution microscopic imaging method based on depth spatio-temporal information integration
Through a single-molecule super-resolution microscopy imaging method integrating deep spatiotemporal information, deep neural networks integrate spatial and temporal features, the problem of difficult to take into account both high temporal resolution and high spatial resolution in the prior art is solved, and high-precision fluorescent molecular positioning and super-resolution image reconstruction are achieved.
Patent Information
- Application Number
- CN202510097678.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
While existing single-molecule positioning technologies achieve high temporal resolution, it is difficult to maintain high spatial resolution, especially in high-density three-dimensional data, where the positioning accuracy and detection rate are significantly reduced.
A single molecule super-resolution microscopy imaging method based on deep spatiotemporal information integration is adopted. By training deep neural network models, spatial and temporal characteristics are integrated, and the scintillation mechanism of fluorescent molecules and time window information are used to improve positioning accuracy and reconstruction accuracy.
The positioning accuracy and reconstruction accuracy are significantly improved in high-density data, the positioning error is reduced, the structural continuity and smoothness are enhanced, and the microscopic imaging with high spatiotemporal resolution is achieved.
Smart Images

Figure CN120147119A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a single-molecule super-resolution microscopy imaging method based on the integration of depth and spatio-temporal information, belonging to the field of biological super-resolution microscopy imaging technology. Background Art
[0002] Single-molecule localization technology is a highly promising imaging technique. It precisely locates individual fluorescent molecules from a sequence of diffraction-limited images, effectively breaking through the limitation of optical diffraction and significantly improving spatial resolution. This technology enables biological structures to be visualized at the molecular level, thus becoming a very valuable super-resolution imaging method in biomedical research. The core of single-molecule localization technology lies in the sparse emission, detection, and localization of fluorescent molecules to achieve super-resolution imaging with nanoscale spatial resolution. However, to achieve sufficient spatial resolution, this technology usually requires collecting tens of thousands of original image frames, which to a certain extent limits its temporal resolution. In addition, the limited temporal resolution and low labeling density also pose limitations on the selection of fluorescent dyes. Meanwhile, during the long-time image acquisition process, the use of high-excitation laser intensity will introduce the problem of phototoxicity.
[0003] To improve the temporal resolution of SMLM, the commonly adopted strategy is to increase the density of simultaneously emitting fluorescent molecules and use high-density analysis methods. By increasing the density of fluorescent molecules in a single-frame image, the number of original images required for super-resolution image reconstruction can be reduced, thereby improving the temporal resolution. However, a high density of fluorescent molecules will cause the overlap of multiple point spread functions in a single-frame image, making it difficult to distinguish, thus reducing spatial resolution. This poses a computational challenge for detecting and localizing adjacent fluorescent molecules. It is possible to identify the overlapping molecules in each frame through multi-emitter fitting or by using advanced analysis methods such as deconvolution, compressive sensing, and mean shift theory. To further improve the temporal resolution, the fluorescence fluctuations in the time domain of the image sequence are utilized to reduce the localization uncertainty of molecules in the sample. Localization techniques based on super-resolution optical fluctuation analysis and Bayesian inference have successfully achieved high-density reconstruction. However, these methods perform best only on two-dimensional samples with the same point spread function model, and their efficiency drops significantly when applied to three-dimensional imaging. Achieving high detection rates and high localization accuracies in high-density three-dimensional data remains a challenging issue to be solved.
[0004] As a key branch of artificial intelligence, deep learning has played a revolutionary role in the field of SMLM. Through deep neural networks, this technology can automatically extract features and patterns from complex biological images, thus processing multi-dimensional and multi-modal data more efficiently. These advancements include tasks such as extracting localization information from diffraction-limited images, reconstructing high-quality super-resolution images from sparse localizations, and optimizing the point spread function model for specific application scenarios. For example, DeepSTORM and Deeploco have shown superior performance to traditional single-molecule fitting algorithms when dealing with high molecular densities. These methods can extract features characterizing individual isolated molecules, such as coordinates, colors, molecular orientations, and distortions, from complex fluorescence images, thereby achieving precise localization of individual molecules. In addition, the recently developed deep context-dependent method (DECODE) has introduced an analysis tool for three-dimensional localization of high-density molecules under different imaging conditions. DECODE has improved the labeling density and imaging speed of SMLM through machine learning techniques, enabling dynamic live-cell SMLM reconstruction. However, although the DECODE technique fuses adjacent three-frame images in SMLM to improve the localization accuracy, it does not fully utilize the temporal information to further optimize the localization effect. Summary of the Invention
[0005] In view of the above problems, the objective of the present invention is to provide a single-molecule super-resolution microscopy imaging method based on the integration of deep spatio-temporal information, which makes full use of the blinking mechanism of fluorescent molecules and improves the reconstruction accuracy of the super-resolution spatio-temporal information integration microscopy imaging method SRST (super-resolution spatiotemporal information integration) based on single-molecule localization through spatio-temporal information integration. The present invention can determine the precise localization of fluorescent molecules used to label biological substances such as proteins in biological samples such as cells at the nanoscale, that is, achieve super-resolution imaging of biological samples.
[0006] To achieve the above objective, the present invention adopts the following technical solutions:
[0007] The single-molecule super-resolution microscopy imaging method based on the integration of deep spatio-temporal information disclosed by the present invention includes the following steps:
[0008] Step 1: Preparation of training data and construction of the model: Determine the true environmental parameters of the data to be measured, set the model hyperparameters, and construct the SRST neural network model structure. The entire model consists of three parts: a spatial feature extraction module, a temporal feature extraction module, and a spatio-temporal feature integration module;
[0009] Spatial Feature Extraction Module: Using a U-net-like architecture, it ensures that the data size of the input and output remains unchanged and is used to extract the spatial information of fluorescent molecules in the image, including two stages: downsampling and upsampling. In the downsampling stage, the image is processed through convolutional layers, and the image size is halved in each iteration while the number of filters is doubled to capture more detailed features in the image. In contrast, in the upsampling stage, the image size is doubled and the number of filters is halved to restore the spatial dimension of the image and gradually refine the feature map. Each stage consists of 3 fully convolutional layers using a 3×3 convolutional kernel. The initial stage of the spatial feature extraction module starts with 48 convolutional filters, providing a solid foundation for capturing the initial image features. To address the common checkerboard artifact problem in the upsampling process, the nearest neighbor interpolation method is selected.
[0010] Temporal Feature Extraction Module: The temporal feature extraction module combines the concept of Long Short-Term Memory (LSTM) and makes specific modifications. By integrating the powerful spatial feature extraction ability of Convolutional Neural Networks (CNN) and the advanced temporal analysis ability of LSTMs, this module effectively captures the temporal information in the image sequence. The module adopts a Bidirectional LSTM (Bi-LSTM) architecture, which not only considers past information but also predicts future information, thus achieving a more comprehensive understanding of temporal data. Different from traditional Bi-LSTM, this module directly operates on 2D feature maps instead of converting them into one-dimensional vectors, thereby retaining more spatial information.
[0011] Spatio-Temporal Feature Fusion Module: The spatio-temporal feature fusion module consists of multiple U-Net-like structures, Mix-Conv (Mixed Convolution) layers, and head convolutional layers, as Figure 1b , 1c shown, aiming to fuse spatio-temporal information and generate the final localization result. The spatial and temporal features are represented by three 48-channel feature maps and are combined and processed through Mix-Conv. This layer effectively integrates the features of different modules, laying the foundation for the final integration of spatio-temporal features. In the integration module, all features are divided into three groups in chronological order: the past set, the current set, and the future set, with the current set being the core. The model focuses on the current frame and realizes a more refined understanding of the temporal context by dividing the frames into past, current, and future sets. This strategy encapsulates the spatial details obtained by the spatial feature extraction module and the temporal dynamics captured by the temporal feature extraction module to increase the weight of the current set in the overall data features and ensure that the network can accurately understand the differences and importance between frames.
[0012] After the spatio-temporal feature fusion module, the SRST network deeply analyzes each pixel in the image through an additional convolutional layer. The network predicts the probability p of the presence of fluorescent molecules and the related coordinate information Δx, Δy, Δz, brightness ΔN, uncertainties σx, σy, σz, σN, background value B, and the probability of presence p. To ensure the rationality of the probability and uncertainties, the sigmoid function is used to limit p and the uncertainties within the interval [0, 1]. At the same time, the coordinate and brightness information is constrained within the interval [-1, 1] through the hyperbolic tangent non-linearity, which not only ensures the stability of the output values but also enables the effective utilization of information from adjacent pixels.
[0013] Step 2: Training of the SRST network model: Generate a sequence of training data images, train the SRST neural network model, and continuously monitor the model training effect metrics until the requirements are met to obtain a trained SRST neural network model;
[0014] Step 2.1: Simulate the luminescence pattern of fluorescent molecules, randomly determine the coordinates of fluorescent molecules, and generate a series of fluorescent images as the training data image sequence. In the present invention, by simulating the luminescence characteristics of fluorescent molecules, a corresponding data set is generated in each round of iteration during network training. The core advantages of this method are reflected in two key aspects: On the one hand, it accurately simulates the physical behavior of fluorescent molecules, especially the intermittent fluorescence emission phenomenon, i.e., the so-called "fluorescence blinking", presented under continuous excitation light. This blinking behavior occurs randomly and asynchronously among molecules, and each fluorescent molecule follows an independent Poisson distribution, through which the emission time and emission interval of the molecule are determined. Incorporating this mechanism into the simulation process can achieve a more accurate characterization of the behavior of fluorescent molecules, thereby improving the efficiency of model training. On the other hand, this method also involves an adjustable time window size, which defines the time context range referred to by the network when processing a single data frame and has a direct impact on the localization accuracy of fluorescent molecules. Effectively utilizing the time information carried by the previous and subsequent frames can provide strong support for the localization of the current frame.
[0015] To prevent the network from overfitting to specific biological structures, the coordinates of the fluorescent molecules in the present invention are randomly generated to ensure that the dataset does not exhibit the characteristics of a continuous structure. During each iteration, the training samples are randomly generated to ensure that each frame is used as a single target only once, thus avoiding the network from overfitting to specific frames. The performance of the present invention depends on an accurate generation model that contains information on the point spread function and camera parameters. If the simulated data does not match the experimental data, the performance of the present invention will be affected. During the preparation of the training data, the image size is standardized to 40×40 pixels. Initially, the total number of fluorescent molecules is estimated using the Poisson distribution according to the molecular density and the total number of image frames. To better conform to the data collected by SMLM, the present invention incorporates the fluorescence blinking mechanism into the simulated data and assigns independent blinking events to each fluorescent molecule. Specifically, the initial emission time point of each fluorescent molecule is determined by a uniform distribution, the number of blinks in the sequence is simulated by the Poisson distribution, and the duration of each emission is determined by the exponential distribution. After these steps, the training images are generated by convolving with the point spread function and adding the camera noise background. To prevent excessive blinking from affecting the robustness of the network, the number of blinks of each fluorescent molecule is limited to at most three times.
[0016] Step 2.2: After obtaining the training data, keep the data input each time consistent with the time window size set by the network, train the network, and simultaneously construct a loss function to evaluate the recognition effect of the model;
[0017] The loss function consists of three parts: the molecular counting loss L count 、the molecular localization loss L loc and the background loss L bg . These loss functions jointly guide the network to learn how to accurately predict whether there are fluorescent molecules at each pixel point on the image, as well as the specific positions and light intensity characteristics of these molecules. The probability p of the presence of fluorescent molecules predicted by the network follows a Bernoulli distribution, and the blinking event of each fluorescent molecule can be regarded as an independent binomial distribution event. Therefore, the counting distribution of fluorescent molecules can be modeled as a Poisson binomial distribution. In this case, the mean μ count and variance of the total number of predicted fluorescent molecules can be calculated as and respectively, where K is the total number of pixels in a single-frame image. When K is large enough, according to the central limit theorem, the Poisson binomial distribution can be approximated as a Gaussian distribution, which simplifies the statistical analysis and allows us to use the properties of the Gaussian distribution to further optimize the performance of the network:
[0018]
[0019] where E is the total number of fluorescent molecules. When μ countWhen approaching E, maximizing the logarithmic probability of E is equivalent to minimizing:
[0020]
[0021] Jointly design the positioning loss to optimize the output variables of the Gaussian mixture model (GMM) to approximate the true posterior regarding the molecular position and brightness. For each pixel k, approximate the true posterior probability with a Gaussian distribution weighted by the detection probability. Model the four-dimensional Gaussian P(u|μ k , Σ k ) as the distribution of the coordinates and brightness u = [x, y, z, I] of the molecule:
[0022]
[0023] where μ k = [x k + Δx k , y k + Δy k , z k + Δz k , I k + ΔI k , μ k represents the offset of the fluorescent molecule based on the center position in the current pixel (i.e., sub-pixel coordinates). Minimize the distance between the inferred posterior and the true posterior by optimizing the log-likelihood of the weighted Gaussian distribution on the GT (ground truth) (i.e., by minimizing the forward KL divergence).
[0024]
[0025] For the background loss, directly calculate the difference between the predicted background image and the true background image using MSE.
[0026]
[0027] Step 2.3: Repeat the above Steps 2.1 to 2.2 until the number of iteration rounds reaches the target number of iterations determined according to the environmental parameters in the model construction phase, and the model performance evaluation given by the above loss function is close to stable between training rounds and no longer shows obvious improvement. The obvious improvement refers to the improvement range not exceeding the preset threshold.
[0028] Step Three: Application of the SRST network model: Preprocess the data to be measured and input it into the trained SRST neural network to form the result of the super-resolution reconstruction image.
[0029] Furthermore, the specific method for the above Step One, training data preparation and model construction, is:
[0030] Step 1.1: Determine the real environment of the data to be measured, including the fluorescence intensity range, background noise, corresponding point spread function, and camera parameters used for data acquisition of the data to be measured;
[0031] Step 1.2: Set the training parameters involved in the model training process, including the required time window size, whether to introduce a blinking mechanism, the size of the training data image, the fluorescence molecule density, and the number of iterations;
[0032] Step 1.3: Set the model hyperparameters, including the time window size, the number of convolutional network feature channels, the depth of the convolutional network, etc.; then combine the spatial feature extraction module, the time feature extraction module, and the spatio-temporal feature integration module to construct the basic structure of the neural network model for the subsequent training process.
[0033] Furthermore, the specific method for applying the SRST network model in Step 3 is as follows:
[0034] Step 3.1: According to the set camera parameters, preprocess the data to be measured, remove the influence brought by the camera, and convert the original data into the number of photons;
[0035] Step 3.2: Input the preprocessed data into the trained network, and the network outputs the coordinate information and probability related to the fluorescence molecules;
[0036] Step 3.3: Set the probability threshold to obtain the final reconstructed super-resolution fluorescence molecule localization result, and draw the corresponding super-resolution image.
[0037] Furthermore, in the part of determining the real environment parameters of the data to be measured in the preparation of training data and model construction in Step 1, the following are satisfied:
[0038] ① The point spread function, fluorescence intensity, and background intensity are measured by the super-resolution microscopy analysis platform tool and used to generate training data;
[0039] ② The fluorescence intensity adopted by the model follows a normal distribution (μ, σ) and is used to generate training data;
[0040] ③ The camera parameters are extracted from the camera parameter database provided by the camera manufacturer according to the camera model and are used to add noise to the training data and simulate the real situation of sample acquisition.
[0041] Furthermore, in the part of constructing the basic structure of the neural network model in the preparation of training data and model construction in Step 1, the following are satisfied:
[0042] ① The spatial feature extraction module adopts a U-net-like architecture. In the downsampling stage, the convolutional layers reduce the image size and increase the number of filters. In the upsampling stage, the image size increases and the number of filters decreases. Each stage has 3 fully convolutional layers. There are 48 convolutional filters in the initial stage, and the nearest neighbor interpolation method is used to solve the grid-like artifact problem;
[0043] ② The temporal feature extraction module combines CNN and LSTM, adopts a bidirectional LSTM architecture, directly operates on the 2D feature map, retains spatial information, and captures the temporal information of the image sequence;
[0044] ③ The spatio-temporal feature fusion module consists of multiple U-nets, hybrid convolutional layers, and head convolutional layers. It fuses spatial and temporal features, processes them in groups according to time order, and takes the current frame as the center to achieve a fine understanding of the temporal context.
[0045] Furthermore, in the training data generation part of the SRST network model training in step 2, it satisfies:
[0046] ① According to the density of the molecules and the total number of frames in the image, use the Poisson distribution to estimate the total number of fluorescent molecules and the positions of individual fluorescent molecules;
[0047] ② Set the starting emission time frame of individual molecules using the uniform distribution, and then set the emission duration and emission interval of individual molecules through the Poisson distribution;
[0048] ③ Convolve the fluorescent molecules with the point spread function through the spline interpolation method, change the spot of the fluorescent molecule from a single pixel point to a spot containing multiple pixel points, and finally add the background intensity to obtain the training data.
[0049] Furthermore, in the model training execution part of the SRST network model training in step 2, it satisfies:
[0050] ① After obtaining the training data, make the single input data consistent with the time window size set by the network, and train the network through pytorch;
[0051] ② Use the Adam W optimizer and set the gradient norm clipping;
[0052] ③ Adopt algorithms such as the MSE root mean square error algorithm and the GMM Gaussian mixture model method to construct the loss function, and establish a loss function for evaluating the recognition effect of the network model.
[0053] Furthermore, in the application of the SRST network model in step 3, it satisfies:
[0054] ① When the preprocessed data is input into the trained network, set the single input data to be consistent with the time window size set by the network;
[0055] ②The network outputs the probability p of the existence of the coordinate information Δx, Δy, Δz, brightness ΔN, uncertainties σx, σy, σz, σN of the fluorescent molecule and the background value B;
[0056] ③Set a probability threshold, identify the pixel points above the threshold as having fluorescent molecules, record the position coordinate information of the fluorescent molecules on these pixel points, and the localization points above the probability threshold in all frames are the finally reconstructed super-resolution fluorescent molecule localization results, and the corresponding super-resolution images are drawn using drawing software.
[0057] Beneficial effects:
[0058] 1. In terms of localization performance, compared with the current state-of-the-art methods (such as DECODE), the single-molecule super-resolution microscopy imaging method based on the integration of depth and spatio-temporal information disclosed in the present invention, SRST, has a 10% increase in the Jaccard index (JI) in dense fluorescence data and a 12 nm reduction in the localization error. As the signal-to-noise ratio decreases and the density of fluorescent molecules increases, the super-resolution reconstruction of the present invention has better structural continuity, no structural loss, and high localization accuracy.
[0059] 2. The single-molecule super-resolution microscopy imaging method based on the integration of depth and spatio-temporal information disclosed in the present invention shows wide applicability in different imaging scenarios through training with different environmental parameters.
[0060] 3. The data generation method based on the fluorescence blinking mechanism disclosed in the present invention generates training data by ensuring that the fluorescent molecules have the characteristic of repeated luminescence and uses it for network training, realizing further optimization of the network performance.
[0061] 4. The single-molecule super-resolution microscopy imaging method based on the integration of depth and spatio-temporal information disclosed in the present invention is trained by setting network structures with different time window sizes, proving that spatio-temporal information can improve the accuracy of single-molecule localization.
[0062] 5. In the study of biological structures, the single-molecule super-resolution microscopy imaging method based on the integration of depth and spatio-temporal information disclosed in the present invention, SRST, can maintain the ability of accurate reconstruction even at ultra-high densities exceeding the original level. The present invention is good at capturing enhanced structural details in 3D imaging of mitochondrial and microtubule structures, while also reducing imaging artifacts and improving structural smoothness. Description of the Drawings
[0063] Figure 1 is a schematic diagram of the SRST process of the present invention. Figure (a) shows the training process and super-resolution reconstruction process of the method, and Figure (b) shows the network architecture of the deep spatio-temporal information integration microscopy method. The network consists of three main parts: a spatial feature extraction module, a temporal feature extraction module, and a spatio-temporal feature fusion module. Figure (c) shows the internal structural details of the sub-component. The MixConv component consists of a series of Conv2D layers and ELU activation functions, and the Head component consists of a series of Conv2D layers, ELU activation functions, and another Conv2D layer;
[0064] Figure 2 is the performance quantification of the present invention on simulated data. Figure (a) shows different density simulated frames with randomly generated representative actual coordinates (green crosses) and predicted coordinates (red dots) under high signal-to-noise ratio conditions. Figure (b) shows the influence of the fluorescence blinking mechanism on the detection performance, localization error, and efficiency. The detection accuracy, localization error, and efficiency of SRST trained with and without the blinking mechanism are quantified under low, medium, and high signal-to-noise ratio conditions within different fluorescence molecule density ranges. Figure (c) shows the influence of the time window size on the detection performance, localization error, and efficiency. Under high signal-to-noise ratio conditions, the detection accuracy, localization error, and efficiency of SRST are quantified with 5-frame, 7-frame, and 9-frame time windows. Figure (d) shows a comparison of SRST and DECODE over a wide density range under low, medium, and high signal-to-noise ratio conditions, with performance metrics including JI, RMSE, and efficiency. Scale bar, 1 μm (a);
[0065] Figure 3 is a schematic diagram of the reconstruction comparison between SRST and DECODE of the present invention on simulated data. Figure (a) shows the super-resolution reconstruction results of the SRST method and the DECODE method of the present invention for ultra-high density simulated data. Figures (b) and (c) are enlarged views of the area in Figure (a) for SRST and DECODE. Figures (d) and (e) show the normalized intensity measurements of the marked areas in Figures (b) and (c) through intensity fluctuations, showing the continuity of specific structures. Figure (f) shows the relationship between the efficiency of SRST and DECODE and the fluorescence molecule density under low signal-to-noise ratio conditions. Scale bar, 1 μm for Figure (a), 500 nm for Figure (b), and 100 nm for Figure (c);
[0066] Figure 4Schematic diagram for comparing the reconstruction results of the SRST method of the present invention on real sample microtubule data. Figure (a) shows the super-resolution reconstruction of microstructures based on SRST and DECODE. The first row shows the super-resolution reconstructions of SRST and DECODE at the original molecular density, while the second row shows the reconstructions with a density four times higher than the original density. Figure (b) is a view of the region enclosed by the orange dashed rectangle in Figure (a) enlarged to compare the structural details of SRST and DECODE in these regions. The red arrows indicate that the grid-like artifacts that appear in DECODE imaging increase with the increase in density, while SRST imaging remains smooth. The order of the images in Figure (b) corresponds to the order of the images in Figure (a). Figure (c) is a view of the region enclosed by the red dashed rectangle in Figure (a) enlarged to compare the structural details of SRST and DECODE in these regions. The red arrows highlight the characteristics of the reconstruction differences between the two methods. Figures (c(i)(ii)) are side views of the specified regions in Figure (c), and the red arrows indicate the differences in imaging between the two methods. Scale bars: 2 μm for Figures (a) and (b), 1 μm for Figure (c), and 500 nm for Figures (c(i)(ii)).;
[0067] Figure 5 Schematic diagram for comparing the reconstruction results of the SRST method of the present invention on real sample mitochondrial data. Figure (a) shows the super-resolution reconstruction image of mitochondria obtained using SRST; Figure (b) is a view of the specified region in Figure (a) enlarged to compare the structural details of SRST, DECODE, and the single emitter fitting algorithm in this region. The red arrows indicate the characteristics of the reconstruction differences among the three methods. Figure (c) is a side view of another specified region in Figure (a) enlarged to show the super-resolution reconstruction side views of SRST, DECODE, and the single radiation source fitting algorithm in this region, and the red arrows point out the differences in structural reconstruction among the three methods. Figure (d) is a side view of the third specified region in Figure (a) enlarged, and the red arrows indicate the differences in imaging details among the three methods. Scale bars: 2 μm for Figures (a) and (b), and 500 nm for Figures (c) and (d); Detailed implementation manners
[0068] The present invention will be described in detail below with reference to the accompanying drawings. However, it should be understood that the provision of the drawings is only for a better understanding of the present invention, and they should not be construed as a limitation to the present invention.
[0069] As Figure 1a shown, the depth spatio-temporal information integration microscopy imaging method (SRST) based on single molecule localization of the present embodiment can perform microscopy imaging with high spatio-temporal resolution under ultra-high density data, and the specific implementation steps are as follows:
[0070] Step 1: Preparation of training data and construction of the model.
[0071] Step 1.1: Measure the real environment of the data to be measured for generating simulated data image frames, including the fluorescence intensity range, background noise, corresponding point spread function, and camera parameters used for data acquisition. The point spread function, fluorescence intensity, and background intensity are obtained through a super-resolution microscopy analysis platform and are used for generating images of fluorescent point sources. The camera parameters are used to add noise to the images and convert the number of photons into data representations in real situations.
[0072] Step 1.2: Set the time window size for the network to process data at one time according to the experimental data, which is used to analyze the spatio-temporal information carried by the fluorescence image data. An appropriate time window size enables the network to better analyze the fluorescence image data and achieve more accurate single-molecule localization. The larger the time window, the more spatio-temporal information can be extracted, but it will bring a large consumption of computing resources and excessive noise introduction.
[0073] Step 1.3: Use deep learning to construct a neural network for realizing single-molecule localization super-resolution microscopy imaging under ultra-high-density fluorescence data. The whole model consists of three parts: a spatial feature extraction module, a temporal feature extraction module, and a spatio-temporal feature integration module.
[0074] Step Two: Training of the SRST network model.
[0075] Step 2.1: Each round of training data is randomly generated. According to the density of molecules and the total number of frames in the image, use the Poisson distribution to estimate the total number of fluorescent molecules and the positions of individual fluorescent molecules. Simulate the luminescence pattern of fluorescent molecules, set the starting luminescence time frame of individual molecules using a uniform distribution, and then set the luminescence duration and luminescence intervals of individual molecules through the Poisson distribution. It should be noted that the two luminescence durations of the same molecule are not equal. Finally, convolve the fluorescent molecules with the point spread function using the spline interpolation method, change the spot of the fluorescent molecule from a single pixel point to a spot containing multiple pixel points, and finally add the background intensity to generate a series of fluorescence images as the training data image sequence.
[0076] Step 2.2: Train the network, use the loss function to backpropagate the model parameters, and ensure that the model correctly understands the spatio-temporal information integration.
[0077] Step 2.3: Repeat the above steps until the number of iteration rounds reaches the target number of iterations determined according to the environmental parameters in the model construction stage, and the model performance evaluation given by the above loss function is close to stable between training rounds and no longer shows obvious improvement.
[0078] Step Three: Application of the SRST network model.
[0079] Step 3.1: Preprocess the data to be measured according to the set camera parameters, remove the influence brought by the camera, and convert the original data into the number of photons.
[0080] Step 3.2: Input the preprocessed data into the trained network, ensuring that the data input each time is consistent with the time window size set by the network. The network outputs the coordinate information of fluorescent molecules Δx, Δy, Δz, brightness ΔN, uncertainties σx, σy, σz, σN, and the probability p of the existence of the background value B.
[0081] Step 3.3: When outputting the information of localized molecules, set a probability threshold to identify the pixel points above the threshold as having fluorescent molecules and record the position coordinate information of the fluorescent molecules at these pixel points. The set of localization points for all frames is the reconstruction result, and the corresponding super-resolution image is drawn using drawing software. Through the module and feature integration strategy, multi-frame image data is analyzed efficiently. This implementation can capture the spatial details in the image and understand the temporal dynamics in the entire image sequence, thereby achieving high-precision fluorescent molecule localization and super-resolution image reconstruction.
[0082] As Figure 2 shown, when considering the fluorescence blinking mechanism during network training, the simulation results consistently show enhanced localization performance under all SNR conditions ( Figure 2 b). In the high signal-to-noise ratio scenario, the network trained with the fluorescence blinking mechanism shows a 1% increase in JI and a 3 nm decrease in RMSE. Under low signal-to-noise ratio conditions, the JI of the network increases by 3% and the RMSE decreases by 7 nm. These results indicate that incorporating the fluorescence blinking mechanism into network training can improve the localization performance, especially under low signal-to-noise ratio conditions. The relationship between network performance and molecular density under different time windows is as Figure 2 shown in c. When processing low-density data, the RMSE of the 9-frame model is 59 nm, and the RMSE of the 5-frame model is 64 nm, about 5 nm higher. For high-density data processing, the RMSE of the 9-frame model is 135 nm, while the RMSE of the 5-frame model is 141 nm, an improvement of about 6 nm. After introducing the fluorescence blinking mechanism and time window size, compared with DECODE, the spatio-temporal information integration analysis enables SRST to identify more localizable molecules. In terms of detection accuracy, localization error, and efficiency, SRST shows better performance than DECODE, especially in high-density scenarios, as Figure 2 shown in d. The analysis of high-density data shows that under high signal-to-noise ratio conditions, SRST has an 8% increase in JI, a 10 nm decrease in RMSE, and a 7% increase in efficiency. Similarly, under low signal-to-noise ratio conditions, SRST achieves a 10% increase in JI, a 14 nm decrease in RMSE, and an 11% increase in efficiency.
[0083] As Figure 3As shown, the super-resolution reconstruction results of SRST and DECODE for ultra-high density simulated data are compared. The reconstruction results of SRST and DECODE under ultra-high density and low signal-to-noise ratio conditions are as Figure 3 shown in Fig. Figure 3 2a. In specific structures in the ultra-high density region ( Figure 3 Figs. Figure 3 2b and 2c), SRST exhibits a smoother and more continuous structure, while DECODE produces obvious grid-like artifacts in these regions and there are obvious structural breaks in specific structures. To examine the structural continuity of SRST reconstruction, a quantitative analysis of specific structures was carried out. The normalized brightness differences of the two methods for specific structures are as Figure 3 shown in Figs. Figure 3 2d and 2e. It is worth noting that the regions reconstructed by SRST show a more uniform normalized intensity distribution, and the variance between adjacent peaks and valleys is on average about 0.1. In contrast, within the regions reconstructed by DECODE, the difference between adjacent peaks and valleys exceeds 0.6, resulting in serious discontinuities in the reconstruction ( Figure 3 Fig. Figure 3 2d). The overall normalized intensity fluctuation in another selected region reconstructed by SRST is smaller, and the slope between peaks and valleys is gentler. On the contrary, DECODE shows more significant fluctuations, presenting a multi-peak pattern, with a peak difference exceeding 0.8, resulting in structural mutations in the reconstructed regions ( Figure 3 Fig. Figure 3 2e). Through the quantitative comparison of the two methods, SRST has an 8% improvement in the efficiency index ( Figure 3 Fig.
[0084] As Figure 4 shown, the super-resolution reconstruction results of SRST and DECODE are compared. Within specific structures in the high density region ( Figure 4 Fig. Figure 4 2c), SRST shows smoother reconstruction results and accurate structural characterization under high density conditions, while DECODE produces grid-like artifacts in the corresponding regions. By magnifying the side view of a specific microstructure ( Figure 4 Fig. Figure 4 2c(i)), it can be seen that SRST shows a clear ring structure, while DECODE does not fully capture these features. In specific structures with four-fold density ( Figure 4 Fig. Figure 4 2c), SRST successfully reconstructs the continuous microtubule image at this increased density ( Figure 4 Fig. Figure 4 2c(ii)), while DECODE shows obvious structural breaks in these specific structures. The reconstructions of special regions by the two methods at the original density and quadruple density are as Figure 4 shown in Fig.
[0085] As Figure 5As shown, the present invention conducts a comparative analysis of SRST, DECODE, and single-emitter fitting algorithms under ultra-high density conditions. The mitochondrial super-resolution reconstruction processed by SRST under ultra-high density conditions is shown in Fig. (5a). In a specific structure in the ultra-high density region ( Figure 5 b), SRST shows a more complete subcellular structure, while the single-emitter fitting algorithm and DECODE show structural degradation at the positions marked by the red arrows. Specifically, the single-emitter fitting algorithm shows more obvious detail loss. The side view of part of the mitochondria is shown as Figure 5 c, with a slice thickness of 100 nm. SRST clearly captures the annular structure, while DECODE and the single-emitter fitting algorithm cannot fully represent this feature. In addition, in the scenario where the annular structure region is identified by all three methods ( Figure 5 d), SRST shows enhanced details and a more continuous structure.
[0086] The above embodiments are only used to illustrate the present invention. The implementation steps of the methods and the like can all be changed. Any equivalent transformation and improvement based on the technical solution of the present invention should not be excluded from the protection scope of the present invention.
Claims
1. A single-molecule super-resolution microscopy method based on deep spatiotemporal information integration, characterized by: The steps include: Step 1: Training data preparation and model construction: determine the real environment parameters of the data to be tested, set the model hyperparameters, and build the SRST neural network model structure; Step 2: SRST network model training: Generate training data image sequence, train the SRST neural network model, and continuously monitor the model training effect indicators until they meet the requirements, and obtain the trained SRST neural network model; Step 2.1: simulate the emission pattern of fluorescent molecules, randomly determine the coordinates of fluorescent molecules, and generate a series of fluorescent images as training data image sequences; Step 2.2: After obtaining the training data, keep the single input data consistent with the time window size set by the network, train the network, and construct a loss function to evaluate the model recognition effect; Step 2.3: Repeat steps 2.1 to 2.2 above until the number of iterations reaches the target number of iterations determined according to the environmental parameters in the model building phase, and the model effect evaluation given by the above loss function is close to stable between training rounds and no longer shows significant improvement. The significant improvement refers to an improvement range that is no greater than a preset threshold. Step 3: Application of SRST network model: Preprocess the data to be tested and input the trained SRST neural network to form a super-resolution reconstructed image result.
2. A single-molecule super-resolution microscopy method according to claim 1, characterized in that: The specific method of step 1 training data preparation and model construction is: Step 1.1: Determine the real environment of the data to be measured, including the fluorescence intensity range, background noise, corresponding point spread function and camera parameters used when collecting data; Step 1.2: Set the training parameters involved in the model training process, including the required time window size, whether to introduce a blinking mechanism, the training data image size, the fluorescent molecule density, and the number of iterations; Step 1.3: Set the model hyperparameters, including the time window size, convolutional network feature channels, convolutional network depth, etc.; then combine the spatial feature extraction module, the temporal feature extraction module and the spatiotemporal feature integration module to construct the basic structure of the neural network model for the subsequent training process.
3. A single-molecule super-resolution microscopy method according to claim 2, characterized in that: The specific method of applying the SRST network model in step 3 is: Step 3.1: Pre-process the data to be measured according to the set camera parameters, remove the influence of the camera, and convert the raw data into photon counts; Step 3.2: Input the preprocessed data into the trained network, and the network outputs the relevant coordinate information and probability of the fluorescent molecules; Step 3.3: Set the probability threshold, obtain the final reconstructed super-resolution fluorescence molecule localization result, and draw the corresponding super-resolution image.
4. A single-molecule super-resolution microscopy method according to claim 1, 2 or 3, characterized in that: In the step 1 of training data preparation and model building, the part of determining the real environment parameters of the data to be tested satisfies: ① The fluorescence intensity used in the training data is normally distributed (μ, σ); ② The point spread function, fluorescence intensity and background intensity are obtained through the super-resolution microscopy analysis platform tool to generate training data; ③ The camera parameters are extracted from the camera parameter database provided by the camera manufacturer according to the camera model and used to simulate the real sample sampling environment.
5. A single-molecule super-resolution microscopy method according to claim 1, 2 or 3, characterized in that: The basic structure of the neural network model in the step 1 of training data preparation and model construction satisfies: ① The spatial feature extraction module adopts a U-net-like architecture to ensure that the size of input and output data does not change, and is used to extract the spatial information of fluorescent molecules on the image; ② The temporal feature extraction module combines CNN and LSTM, adopts a bidirectional LSTM architecture, directly operates on 2D feature maps, retains spatial information, and captures the temporal information of image sequences; ③The spatiotemporal feature fusion module consists of multiple U-net-like, hybrid convolutional layers and head convolutional layers, which fuses spatial and temporal features, groups them in chronological order, and takes the current frame as the center to achieve a detailed understanding of the temporal context.
6. A single-molecule super-resolution microscopy method according to claim 1, 2 or 3, characterized in that: In the training data generation part of the step 2 SRST network model training, the following conditions are met: ① Based on the molecular density and the total number of frames in the image, the Poisson distribution is used to estimate the total number of fluorescent molecules and the position of individual fluorescent molecules; ②Use uniform distribution to set the starting luminescence time frame of a single molecule, and then use Poisson distribution to set the luminescence duration and luminescence interval of a single molecule; ③ Use the point spread function to convolve the fluorescent molecules through the spline interpolation method, change the light spot of the fluorescent molecule from one pixel to a spot containing multiple pixels, and finally add the background intensity to obtain the training data.
7. A single-molecule super-resolution microscopy method according to claim 1, 2 or 3, characterized in that: The model training execution part in the step 2 SRST network model training satisfies: ① After obtaining the training data, keep the single input data consistent with the time window size set by the network, and train the network through pytorch; ②Use the Adam W optimizer and set gradient norm clipping; ③The MSE root mean square error algorithm and GMM Gaussian mixture model method are used to construct the loss function, and a loss function is established to evaluate the recognition effect of the network model.
8. A single-molecule super-resolution microscopy method according to claim 1, 2 or 3, characterized in that: In the step 3, the SRST network model application satisfies: ① When the preprocessed data is input into the trained network, the single input data is set to be consistent with the time window size set by the network; ② The network outputs the relevant coordinate information of the fluorescent molecule Δx, Δy, Δz, brightness ΔN, uncertainty σx, σy, σz, σN and the probability p of the background value B; ③ Set a probability threshold, identify pixels above the threshold as containing fluorescent molecules, and record the position coordinate information of the fluorescent molecules at the pixel. The positioning points above the probability threshold in all frames are the final reconstructed super-resolution fluorescent molecule positioning results, and use drawing software to draw the corresponding super-resolution image.