Single-frame composite structure illumination obvious micro-imaging method based on ensemble learning
Through a single-frame composite structured illumination micro-imaging method based on integrated learning, the problems of complex parameter estimation and low temporal resolution in the prior art are solved, high-quality super-resolution imaging is achieved, and good reconstruction performance is maintained at low excitation power.
Patent Information
- Application Number
- CN202510282298.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-30
AI Technical Summary
The existing structured illumination microtechnology has shortcomings in parameter estimation and temporal resolution, and the existing frame reduction algorithm has poor reconstruction quality in actual imaging experiments.
Using a single-frame composite structured light illumination microimaging method based on integrated learning, three subnets and one adaptive integrated network are constructed, and the original illumination image of the sample is collected using an interferometric structured light illumination microsystem, and super-resolved results are generated through training data.
The original frame count of the traditional method is reduced by 9 times, complex parameter estimation steps are avoided, high-quality super-resolution results are obtained, and good reconstruction performance is maintained at low excitation power.
Smart Images

Figure CN120065496A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of super-resolution fluorescence microscopy, and specifically relates to a single-frame composite structured illumination microscopy method (eDL-cSIM) based on ensemble learning, which is used for super-resolution observation at the subcellular scale. Background Art
[0002] Fluorescence microscopes are of great significance to life sciences. However, limited by the diffraction limit, their spatial resolution can only reach 200 nm. In the past few decades, a variety of fluorescence super-resolution methods have been developed, such as single-molecule localization microscopy (PALM / STORM), structured illumination microscopy (SIM), stimulated emission depletion microscopy (STED). Among various fluorescence super-resolution techniques, SIM stands out due to its low phototoxicity, photobleaching, etc. The SIM technique uses periodic fringes with three phases in each of the three directions to modulate the high-frequency information of the sample within the cut-off frequency of the optical system, and reconstructs it in the frequency domain and then performs an inverse Fourier transform to obtain a super-resolution image. However, the principle of SIM determines the complexity of illumination parameter estimation and the low temporal resolution. To address this problem, a parameter estimation method based on principal component analysis (PCA-SIM) [Qian, Jiaming, et al. "Structured illumination microscopy based on principal component analysis." eLight 3.1 (2023): 4.] was proposed. This method can achieve fast and accurate non-iterative parameter estimation. However, this method does not improve the performance of SIM from the perspective of the number of original frames. Currently, there are already solutions to reduce the number of original frames required by SIM [Qian J, Cao Y, Xu K, et al. Robust frame-reduced structured illumination microscopy with accelerated correlation-enabled parameter estimation [J]. Applied Physics Letters, 2022, 121(15).]. However, due to relying on strict optical settings or complex noise models, the reconstruction quality of existing frame reduction algorithms in actual imaging experiments is inferior to that of conventional 9-frame reconstruction.
[0003] In recent years, the rise of deep learning technology has greatly promoted the development of SIM technology. Some studies have proved the effectiveness of deep learning for SIM [Qiao, Chang, et al. "Evaluation and development of deep neural networks for image super-resolution in optical microscopy." Nature methods 18.2 (2021): 194-202.]. However, the super-resolution of SIM based on deep learning often uses a model mainly based on convolutional neural networks for training and inference. Its reconstruction quality is limited by a single model, and due to the limited receptive field of convolutional neural networks, it is difficult to model global information. Summary of the Invention
[0004] The purpose of the present invention is to provide a single-frame composite structured illumination microscopy imaging method based on ensemble learning.
[0005] The technical solution for achieving the purpose of the present invention is: a single-frame composite structured illumination microscopy imaging method based on ensemble learning, and the specific steps are as follows:
[0006] Step 1: Use an interferometric structured illumination microscopy system to collect the original illumination image of the sample;
[0007] Step 2: Collect the original illumination image required for single-frame composite structured illumination super-resolution;
[0008] Step 3: Generate training data, train sub-networks and an ensemble network;
[0009] Step 4: Use the trained network to output the super-resolution result.
[0010] Preferably, the interferometric structured illumination microscopy system includes a laser, a mirror, a spatial filter, a spatial light modulator, a filtering system composed of four lenses and a mask, a microscope, and a camera. After the laser emits laser light, it is reflected by the mirror, collimated and expanded by the spatial filter, and then enters the spatial light modulator. The spatial light modulator loads nine fringe images and one composite fringe image. Among them, the nine fringe images are in three directions, and each direction has three-step phase-shifted fringe images. The composite fringe image is obtained by adding one frame from each of the nine fringe images in the three illumination directions. The light diffracts after passing through the spatial light modulator, then enters the microscope through the filtering system and interferes on the surface of a fixed bovine pulmonary artery endothelial sample. Then, the camera acquires nine original sample images obtained by interfering the nine original fringe images and a composite illumination image obtained by interfering the composite fringe image. The average value of the nine original sample images is taken to obtain a wide-field image, and then the nine original sample images are super-resolution reconstructed by structured illumination microscopy based on principal component analysis to obtain a ground truth image.
[0011] Preferably, an ensemble learning framework is constructed. The ensemble learning framework includes three sub-networks and an adaptive ensemble network based on a visual self-attention model. The input composite illumination image enters the three sub-networks respectively. Sub-network 1 extracts low-frequency information from the composite illumination image. Sub-network 2 extracts high-frequency information from the spatial domain of the composite illumination image. Sub-network 3 extracts high-frequency information from the frequency domain of the composite illumination image. The adaptive ensemble network based on the visual self-attention model fuses the output images obtained by the three sub-networks to obtain a super-resolution image.
[0012] Preferably, Sub-network 1 uses a U-shaped architecture, the output is set as a wide-field image, and the shape of the input composite illumination image is set as (B, C, H, W), where B represents the batch size, C is the number of channels, and H and W represent the height and width of the input respectively. Sub-network 1 is an encoder-decoder structure. The encoder stage has five layers, and each layer includes a convolutional block and a max-pooling layer;
[0013] Corresponding to the encoder, the decoder stage also contains five layers, and each layer includes a transposed convolutional layer and a convolutional block, doubling the size of the feature map and halving the number of channels; the decoder is residually connected to the same-layer branch of the encoder, and the transposed convolutional feature map is connected to the encoder's feature map along the channel dimension; the connected feature map is processed by a convolutional block to halve the number of channels. Finally, a set of convolutional layers is used to convert the number of channels to 1, and the features are refined to produce the final output without changing the spatial dimension.
[0014] Preferably, Sub-network 2 is a recursive residual network, the output is set as the ground truth image, local and global residual connections are used, the information flow is optimized through a set number of recursive iterations, and the output tensor is upsampled by pixel shuffling to finally generate an image.
[0015] Preferably, the sub-network 3 is constructed using a Fourier domain amplitude-phase channel attention mechanism, and the output is set to the ground truth image. The input composite illumination image is first processed by a head convolutional layer to generate an output tensor x of a certain dimension. The feature map generated by the head convolutional layer is then input into a series of residual groups. Each residual group contains 4 residual blocks, and each residual block consists of two partial convolutions and an amplitude-phase channel attention block. The amplitude-phase channel attention block is composed of an amplitude channel attention layer and a phase channel attention layer. After partial convolution within the residual block, x undergoes a Fourier transform:
[0016]
[0017] where represents the Fourier transform, F real represents the real part of the Fourier transform, F imag represents the imaginary part of the Fourier transform, j represents the imaginary unit, and PConv(x) represents performing partial convolution on x. The amplitude channel attention layer extracts the amplitude feature A from the Fourier-transformed feature map:
[0018]
[0019] Then, the amplitude feature A is processed through partial convolution, global average pooling, max pooling, and a one-dimensional convolution with both the input and output channels set to 1. The weight of the amplitude feature is obtained through a sigmoid function and compressed to the range [0, 1];
[0020] The weight is multiplied by the feature map extracted by partial convolution to implement the amplitude channel attention mechanism:
[0021] X ACALayer = PConv(x) × S(Conv 1×1 ((Pooling(Pconv(A))))) (10)
[0022] where Pooling represents a combination of average pooling and max pooling;
[0023] The phase channel attention layer takes the phase of the Fourier-transformed feature map:
[0024]
[0025] The feature map is upsampled through pixel shuffling and transformed into an image of (B, 1, 2×H, 2×W).
[0026] Preferably, the adaptive integration network includes three encoder stages and one decoder stage, and each encoder stage includes a number of sliding window self-attention modules;
[0027] The outputs of the three sub-networks first pass through a 3×3 convolutional layer for shallow feature extraction to obtain the feature map X c Entering the encoder stage, each sliding window self-attention module in the encoder and decoder will divide the input feature map into several windows using the method based on the context of the statistical language model. Subsequently, the multi-head self-attention mechanism is calculated and processed within each window, followed by layer normalization and a multi-layer perceptron to adjust the dimensions, and a residual connection is made with the input of the sliding window self-attention module to form the output feature map; the sliding window self-attention modules within each encoder stage are sequentially connected in order; the output of each encoder stage is downsampled by patch merging to reduce the size of the feature map while expanding the number of channels to achieve the purpose of compression coding. For an input of size (B, H, W, C), samples are taken every 1 element along the rows and columns to obtain four patches X 1 、X 2 、X 3 、X 4 , the dimension of each patch is (B, H / 2, W / 2, C). Subsequently, all patches are concatenated along the channel dimension to obtain a tensor of size (B, H / 2, W / 2, 4C), and then the dimensions are transformed to reshape the shape to (B, H / 2×W / 2, 4C); the encoder stage 2 and encoder stage 3 use the output of the previous stage's patch merging operation as the input to gradually compress the features to higher dimensions;
[0028] The output of each encoder stage is patch-expanded to ensure that the dimensions match during the residual connection;
[0029] The output of encoder stage i is concatenated with the patch-expanded output of encoder stage i + 1 along the channel dimension, i>1, and the result of the concatenated output is input into the decoder; the information fused by the encoder is decoded by the decoder.
[0030] Compared with the prior art, the significant advantages of the present invention are: (1) The present invention reduces the number of original frames of the traditional method by 9 times through the method of single-frame composite illumination; (2) The present invention can obtain high-quality super-resolution results without complex and time-consuming parameter estimation steps; (3) The integrated learning model based on convolutional neural network and Vision Transformer of the present invention has stronger global modeling ability and can maintain good reconstruction performance under low excitation power.
[0031] The following further describes the present invention in detail with reference to the accompanying drawings: Description of the Drawings
[0032] Figure 1 It is an architecture diagram of a single-frame composite structured illumination microscopy imaging method based on ensemble learning.
[0033] Figure 2 It is the schematic diagram of the sub-network 3 network structure.
[0034] Figure 3 It is the schematic diagram of the integrated network structure.
[0035] Figure 4 It is the schematic diagram of 19 frequency components generated by the composite structure light illumination.
[0036] Figure 5 It is the comparison between the super-resolution reconstruction result of the present invention and the reconstruction results of other different deep learning methods.
[0037] Figure 6 It is the super-resolution reconstruction result of the present invention under different signal-to-noise ratios. Specific implementation manners
[0038] A single-frame composite structure light illumination microscopy method based on ensemble learning. This method composites the illumination modes of structured illumination microscopy and constructs an ensemble learning network, and only one composite illumination image is required to achieve high-precision and high-robustness super-resolution imaging. The present invention first constructs three convolutional-based sub-networks and an adaptive ensemble learning network based on a visual self-attention model, composites the fringes in three directions of the SIM illumination mode by means of composite coding, uses the single-frame composite illumination data as the input of the network during training, and uses the reconstruction result of structured illumination microscopy based on principal component analysis as the ground truth for network training. After training, the test set data is fed into the neural network. The data is first output by the sub-networks, and the data of the sub-networks is input into the integrated network after being concatenated in the channel dimension to obtain the super-resolution image. This method includes the following three steps:
[0039] Step 1: Use an interferometric structured illumination microscopy system to collect the original illumination image of the sample.
[0040] In a further embodiment, the interferometric structured illumination microscopy system includes a laser, a mirror, a spatial filter, a spatial light modulator, a filtering system composed of four lenses and a mask, a microscope, and a camera. After the laser emits laser light, which is reflected by the mirror, it is collimated and expanded by the spatial filter and then enters the spatial light modulator. The spatial light modulator loads nine fringe images and one composite fringe image. Among them, the nine fringe images are in three directions, and each direction has three-step phase-shifted fringe images. The composite fringe image is obtained by adding one frame (one-step phase shift) from each of the nine fringe images in the three illumination directions. After the light diffracts through the spatial light modulator, it passes through the filtering system and then enters the microscope, where it interferes on the surface of a fixed bovine pulmonary artery endothelial sample. Then, the camera captures nine original sample images obtained by the interference of the nine original fringe images and a composite illumination image obtained by the interference of the composite fringe. The average value of the nine original sample images is taken to obtain a wide-field image, and the nine original sample images are super-resolved and reconstructed by structured illumination microscopy based on principal component analysis to obtain the ground truth.
[0041] The above process is repeated 200 times to obtain 200 sets of images. Each set contains one composite illumination image (resolution: 512×512), one wide-field image (resolution: 512×512), and one ground truth image (resolution: 1024×1024). Through random cropping, a total of 12,800 sets of composite illumination images and wide-field images with a resolution of 128×128, and ground truth images with a resolution of 256×256 are obtained, and random horizontal / vertical flipping and random scaling are used to further enhance the dataset. Subsequently, the dataset is divided in a ratio of 8:1:1, resulting in 10,240 sets of training data, 1,280 sets of validation data, and 1,280 sets of test data.
[0042] Step 2: As Figure 1 shown, construct an ensemble learning framework. The ensemble learning framework includes three sub-networks mainly based on convolution and an adaptive ensemble network based on a visual self-attention model. The input composite illumination images enter the three sub-networks respectively. Sub-network 1 extracts low-frequency information from the composite illumination image, sub-network 2 extracts high-frequency information from the spatial domain of the composite illumination image, and sub-network 3 extracts high-frequency information from the frequency domain of the composite illumination image. The adaptive ensemble network based on the visual self-attention model fuses the output images obtained by the three sub-networks to obtain a super-resolution image.
[0043] Sub-network 1 uses a U-shaped architecture. The output is set to a wide-field image, and the shape of the input composite illumination image is set to (B, C, H, W), where B represents the batch size, C is the number of channels (C = 1 for the input layer), and H and W represent the height and width of the input respectively. Sub-network 1 has an encoder-decoder structure. The encoder stage has five layers, each layer mainly consisting of a basic convolutional block and max pooling. The basic convolutional block is represented as follows:
[0044] X o = GELU(PConv 3×3 (Conv 3×3 (X i ))) (3)
[0045] Among them, X i and X o represent the input and output feature maps of the basic convolutional block. Conv 3×3 represents a regular convolution with a kernel size of 3. PConv 3×3 represents a partial convolution with a kernel size of 3. In the convolution process, the partial convolution divides the feature channels obtained by convolution according to a set ratio P. The first consecutive P channels are regarded as the representative of the entire feature map for calculation, and the remaining partial channels perform an identity mapping, thereby reducing the training and inference time. GELU is the Gaussian error linear unit activation function. Subsequently, max pooling with a kernel size of 2 and a stride of 2 is used to downsample the result, reducing the feature map size to half of the input. In the first basic convolutional block of the encoder, the number of channels is expanded from 1 to 64. In the remaining 4 layers, each layer consists of a basic convolutional block and max pooling, and the number of channels is successively expanded to 128, 256, 512, and 1024. The tensor is finally compressed to a size of (B, 1024, H / 16, W / 16). Corresponding to the encoder, the decoder stage also contains five layers, each layer mainly consisting of a transposed convolution and a basic convolutional block. The kernel size of the transposed convolution is 3 and the stride is 2, doubling the feature map size and halving the number of channels. Then, a residual connection is made with the same-layer branch of the encoder, and the feature map of the transposed convolution is connected to the feature map of the encoder along the channel dimension, thereby expanding the number of channels. The connected feature map is processed by a basic convolutional block (for example, for the first decoder layer, the input channels are set to 1024 and the output channels are set to 512) to halve the number of channels. Finally, a group of 1×1 convolutions is used to convert the number of channels to 1 and refine the features to produce the final output without changing the spatial dimension.
[0046] Sub-network 2 is a residual network with a recursive structure. The output is set to the ground truth image, and the input is (B, C, H, W). Local and global residual connections are used to optimize the information flow through 25 recursive iterations. The output tensor is upsampled by pixel shuffling and finally generates an image of (B, 1, 2×H, 2×W).
[0047] The sub-network 3 is constructed using the Fourier domain amplitude-phase channel attention mechanism, and the output is also set to the ground truth image. The input tensor size is (B, C, H, W). As Figure 2 shown, the input tensor is first processed by the head convolutional layer to produce an output tensor x with dimensions (B, 64, H, W). The feature maps generated by the head convolutional layer are then fed into a series of residual groups. The network contains 4 residual groups, and each group contains 4 residual blocks. Each residual block consists of two partial convolutions and an amplitude-phase channel attention block, and the amplitude-phase channel attention block consists of an amplitude channel attention layer and a phase channel attention layer. After the partial convolution within the residual block, x undergoes a Fourier transform:
[0048]
[0049] where represents the Fourier transform, F real represents the real part of the Fourier transform, F imag represents the imaginary part of the Fourier transform, and j represents the imaginary unit. The amplitude channel attention layer extracts the amplitude features from the Fourier-transformed feature maps:
[0050]
[0051] These amplitude features are then processed through partial convolution, global average pooling, max pooling, and a one-dimensional convolution with both the input and output channels set to 1. The weights of the amplitude features are obtained through a sigmoid function and compressed into the range [0, 1]. The sigmoid function is expressed as:
[0052]
[0053] These weights are multiplied by the feature maps extracted by the partial convolution to implement the amplitude channel attention mechanism, as shown in Equation (10), where Pooling represents the combination of average pooling and max pooling:
[0054] X ACALayer = PConv(x) × S(Conv 1×1 ((Pooling(Pconv(A))))) (10)
[0055] The phase channel attention layer takes the phase of the Fourier-transformed feature maps:
[0056]
[0057] The other steps are the same as those of the amplitude channel attention layer to implement the phase channel attention mechanism.
[0058] Finally, the feature maps are upsampled through pixel shuffling and converted into an image with dimensions (B, 1, 2×H, 2×W).
[0059] The outputs of the three sub-networks are stacked in the channel dimension to form an input adaptive integration network of (B, 3, H, W). The adaptive integration network is as Figure 2 shown. This network consists of three encoder stages and one decoder stage. Each encoder stage is composed of several sliding window self-attention modules. Within a stage, the odd-numbered sliding window self-attention modules perform the calculation of the multi-head self-attention mechanism within the window. The specific process is shown in equations (12 - 14): (12)
[0061] Head i = Attention(Q, K, V) (13)
[0062] MutiHead(Q, K, V) = Concat(head 1 , … head h )W O (14)
[0063] In the equations, Q, K, and V are matrices obtained by linear transformation of the input matrix. Attention, head, and MutiHead represent the operations of calculating the self-attention mechanism, single-head attention mechanism, and multi-head attention mechanism respectively. d k represents the dimension of Q, K, and V; softmax is the normalized exponential function; W O represents the weight matrix of the linear transformation of the multi-head attention mechanism; Concat represents the operation of concatenating head i (i = 1 … h) in the d k dimension.
[0064] The even-numbered block sliding window self-attention modules perform the calculation of the sliding window multi-head self-attention mechanism. The specific process is as follows:
[0065] The input matrix is slid one pixel to the right and down. At the same time, the first row and the first column of the matrix are filled with the third row and the third column before sliding. Then, the matrix is calculated according to the process of equations (12 - 14).
[0066] The output obtained from the sub-network first passes through a 3×3 convolutional layer for shallow feature extraction. The resulting feature map is X cEnter the encoder stage. Each sliding window self-attention module in the encoder and decoder will divide the input feature map into several windows using a method based on the context of the statistical language model. Subsequently, the multi-head self-attention mechanism is calculated and processed within each window. Then, layer normalization and a multi-layer perceptron are used to adjust the dimensions, and a residual connection is formed with the input of the sliding window self-attention module to form the output feature map. Encoder stage 1 consists of 6 sliding window self-attention modules, encoder stage 2 consists of 4 sliding window self-attention modules, and encoder stage 3 consists of 4 sliding window self-attention modules. The sliding window self-attention modules within each stage are sequentially connected in order. The output of each encoder stage will be subjected to patch merging to reduce the size of the feature map while expanding the number of channels to achieve the purpose of compression encoding. For an input of size (B, H, W, C), samples are taken every 1 element along the rows and columns to obtain four patches X 1 、X 2 、X 3 、X 4 , and the dimension of each patch is (B, H / 2, W / 2, C). Subsequently, all patches are concatenated along the channel dimension to obtain a tensor of size (B, H / 2, W / 2, 4C), and then the dimension is transformed to reshape the shape to (B, H / 2×W / 2, 4C). The second and third stages of the encoder use the output of the previous stage's patch merging operation as their input, gradually compressing the features to higher dimensions. In addition, to implement the residual connection, the output of each encoding stage needs to be patch-expanded to ensure that the dimensions match during the residual connection. This process is a combination of pixel shuffling, convolution, layer normalization, and dimension concatenation. For the i-th stage (i > 1) of the encoder, its output is concatenated with the patch-expanded output of the (i + 1)-th stage along the channel dimension: Let X i be the output feature of the i-th stage, X i+1 be the output feature of the (i + 1)-th stage, and X i+1 is first processed by pixel shuffling:
[0067] X up =PixelShuffle(X i+1 )(15)
[0068] Next, a 3×3 convolutional layer is used to adjust the number of channels to match the subsequent calculations and further enhance the feature information:
[0069] X d =Convolution(X up ,3×3)(16)
[0070] Then, the dimension is transformed and layer normalization is applied to process X d :
[0071] X p = LN(X d )(17)
[0072] Finally, the output of layer normalization is concatenated with X i along the channel dimension, and the resulting output is fed into the decoder. This step is expressed as:
[0073] X c = Concat(X i , X p )(18)
[0074] This concatenation effectively integrates multi-scale features and enhances the network's representation ability. The information fused by the encoder is decoded by a single decoder, and the number of sliding window self-attention modules in it is set to 6. Its output is connected with the shallow feature X c using a residual connection, and a 3×3 convolution is used to convert the feature channels from 64 to 1 to achieve image reconstruction.
[0075] Step 3: Train the sub-network and the integrated network.
[0076] The sub-network calculates the loss using the mean square error function (MSE), and the integrated network calculates the loss using a weighted combination of the structural similarity loss function (SSIM) and the perceptual loss (LP), and the weight of the perceptual loss is set to 0.2. The network parameters are updated using the adaptive momentum optimizer (adam), and a dynamic learning rate adjustment strategy is adopted to make finer adjustments to the model weights. The initial learning rate is set to 0.05. From 0 to the 60th training epoch, the learning rate decays by 0.5 times every 15 training epochs. From the 61st to the 150th training epoch, the learning rate decays by 0.5 times every 30 training epochs. Starting from the 150th training epoch, if the validation loss does not decrease for 20 consecutive training epochs, the learning rate decays by 0.5 times. The constructed neural network is calculated on a workstation equipped with an Intel Core i7-13700KF CPU, 32GB RAM, and an NVIDIA GeForce RTX4090 based on the PyTorch platform (version 2.0.1, using Python 3.10.0).
[0077] Step 4: Use the trained network to output the super-resolution results. Load the composite illumination images in the validation set, as well as the network model and its parameters. First, send the composite illumination image data into three sub-networks respectively for output, stack the output results of the sub-networks along the channel dimension, send them into the integrated network, and then output the final results.
[0078] Example
[0079] To test the feasibility of the present invention, the present invention (eDL-cSIM) was first compared with two other state-of-the-art neural networks for structured illumination super-resolution reconstruction, namely the amplitude-phase channel attention network (APCAN) and the Fourier-enhanced and shifted hierarchical network (FESTN). The two networks used for comparison were trained with composite structured illumination images and corresponding super-resolution images as input and output. The predicted results of bovine pulmonary artery endothelial cells obtained by these three methods are shown in the appendix Figure 5 as follows. Due to the severe spectral overlap of the input images posing a great challenge to the performance of deep learning, the reconstruction quality of FESTN was relatively affected, and the predicted results of mitochondria and actin did not show obvious super-resolution effects (in (b3) and (c3) of the appendix Figure 5 ). Although the amplitude-phase channel attention network obtained a resolution-enhanced reconstruction of mitochondria, its fidelity was affected by spectral overlap, and its actin results were still affected (in (b4) and (c4) of the appendix Figure 5 ). In contrast, eDL-cSIM can recover higher-quality super-resolution images. Especially in the actin results, the fine structures that could not be distinguished by other neural networks were well resolved (in (b5) and (c5) of the appendix Figure 5 ).
[0080] Secondly, the performance of eDL-cSIM in a complex imaging environment was tested. The excitation power was gradually reduced from the rated value (55 mW) to 25% of the original value, and 9 original images of mitochondria of bovine pulmonary artery endothelial cells were taken in a nine-frame fringe illumination mode at different powers, and 1 image was taken in a composite fringe illumination mode. And super-resolution images were obtained by applying principal component analysis-based structured illumination microscopy (PCA-SIM) and eDL-cSIM respectively. As shown in the appendix Figure 6 , the results of principal component analysis-based structured illumination microscopy (in (c1)-(c4) of the appendix Figure 6 ) showed increasingly serious reconstruction artifacts as the excitation power decreased, and the structural similarity between the results in different noise environments and the results at high excitation power decreased from 0.77 to 0.47, while eDL-cSIM (in (d1)-(d4) of the appendix Figure 6 ) always maintained a stable reconstruction quality, and the structural similarity was above 0.85. This example demonstrates the comprehensive advantages of the present invention in terms of accuracy and noise resistance, as well as its potential for long-term live cell imaging under weak excitation conditions.
Claims
1. A single-frame composite structure illumination microscopy imaging method based on ensemble learning, characterized in that: The specific steps are: Step 1: Collect the original illumination image of the sample using the interferometric structured light illumination microscopy system; Step 2: Collect the original illumination image required for single-frame composite structured light illumination super-resolution; Step 3: Generate training data, train sub-networks and integrate networks; Step 4: Use the trained network to output super-resolution results.
2. The single-frame composite structure illumination microscopy imaging technique based on ensemble learning according to claim 1, characterized in that: The interferometric structured light illumination microscopy system comprises a laser, a reflector, a spatial filter, a spatial light modulator, a filtering system consisting of four lenses and a mask, a microscope and a camera. The laser emitted by the laser is reflected by the reflector and then collimated and expanded by the spatial filter before entering the spatial light modulator. The spatial light modulator is loaded with nine fringe images and one composite fringe image, wherein the nine fringe images are fringe images in three directions with three-step phase shifts in each direction. The composite fringe image is obtained by adding one frame of the nine fringe images in each of the three illumination directions. The light is diffracted after passing through the spatial light modulator, and then enters the microscope after passing through the filtering system and interferes on the surface of a fixed bovine pulmonary artery endothelial sample. The camera then collects nine original sample images obtained by interfering the nine original fringe images and a composite illumination image obtained by interfering the composite fringe images. The nine original sample images are averaged to obtain a wide-field image. The nine original sample images are then super-resolved using structured light illumination microscopy based on principal component analysis to obtain a ground truth image.
3. The single-frame composite structure illumination microscopy imaging method based on ensemble learning according to claim 1, characterized in that: An integrated learning framework is constructed, which includes three sub-networks and an adaptive integrated network based on a visual self-attention model. The input composite illumination image enters the three sub-networks respectively. Sub-network 1 extracts low-frequency information of the composite illumination image, sub-network 2 extracts high-frequency information of the composite illumination image from the spatial domain, and sub-network 3 extracts high-frequency information of the composite illumination image from the frequency domain. The adaptive integrated network based on the visual self-attention model fuses the output images obtained by the three sub-networks to obtain a super-resolution image.
4. The single-frame composite structure illumination microscopy imaging method based on ensemble learning according to claim 3, characterized in that: The sub-network 1 uses a U-shaped architecture, the output is set to a wide-field image, and the input composite illumination image shape is set to (B, C, H, W), where B represents the batch size, C is the number of channels, H and W represent the height and width of the input, respectively. The sub-network 1 is an encoder-decoder structure, and the encoder stage has five layers, each of which includes a convolution block and a maximum pooling layer; Corresponding to the encoder, the decoder stage also contains five layers, each of which includes a deconvolution layer and a convolution block, which doubles the size of the feature map and halves the number of channels; the decoder is residually connected to the same-layer branch of the encoder, connecting the deconvolution feature map with the encoder feature map along the channel dimension; the connected feature map is processed by a convolution block to halve the number of channels. Finally, a set of convolution layers is used to convert the number of channels to 1, and the features are refined to produce the final output without changing the spatial dimension.
5. The single-frame composite structure illumination microscopy imaging method based on ensemble learning according to claim 3, characterized in that: Subnetwork 2 is a residual network with a recursive structure. The output is set to the ground truth image. Local and global residual connections are used to optimize the information flow through a set number of recursive iterations. The output tensor is processed by pixel shuffling and upsampling to finally generate an image.
6. The single-frame composite structure illumination microscopy imaging method based on ensemble learning according to claim 3, characterized in that: Subnetwork 3 is built using the Fourier domain amplitude-phase channel attention mechanism, and the output is set to the ground truth image. The input composite illumination image is first processed by the head convolution layer to produce an output tensor x of dimension x. The feature map generated by the head convolution layer is then input into a series of residual groups. Each residual group contains 4 residual blocks, each of which consists of two partial convolutions and an amplitude-phase channel attention block. The amplitude-phase channel attention block consists of an amplitude channel attention layer and a phase channel attention layer. x is Fourier transformed after the partial convolution in the residual block: in stands for Fourier transform, F real Represents the real part of Fourier transform, F imag Represents the imaginary part of Fourier transform, j represents the imaginary unit, PConv(x) represents the partial convolution performed on the input x, and the amplitude channel attention layer extracts the amplitude feature A from the feature map after Fourier transform: Then the amplitude feature A is processed by partial convolution, global average pooling, maximum pooling, and one-dimensional convolution with both input and output channels set to 1. The weight of the amplitude feature is obtained by the S-type function and the weight is compressed to the range of [0, 1]. The weights are multiplied with the feature maps extracted by the partial convolution to implement the amplitude channel attention mechanism: X ACALayer =PConv(x)×S(Conv 1×1 ((Pooling(Pconv(A))))) Among them, Pooling represents the combination of average pooling and maximum pooling; The phase channel attention layer takes the phase of the feature map after Fourier transformation: The feature map is upsampled by pixel shuffling and converted into a (B, 1, 2×H, 2×W) image.
7. The single-frame composite structure illumination microscopic imaging method based on ensemble learning according to claim 3, characterized in that: The adaptive ensemble network includes three encoder stages and one decoder stage, each encoder stage includes a number of sliding window self-attention modules; The outputs of the three sub-networks are first passed through a 3×3 convolutional layer for shallow feature extraction to obtain the feature map X c Entering the encoder stage, each sliding window self-attention module in the encoder and decoder will divide the input feature map into several windows based on the context of the statistical language model, and then perform multi-head self-attention mechanism calculation processing in each window, and then adjust the dimension through layer normalization and multi-layer perceptron, and perform residual connection with the input of the sliding window self-attention module to form the output feature map; each sliding window self-attention module in each encoder stage is connected sequentially; the output of each encoder stage is patch merged to reduce the size of the feature map and expand the number of channels to achieve For the purpose of compression coding, for an input of size (B, H, W, C), every other element along the rows and columns is sampled to obtain four patches X1, X2, X3, X4, each of which has a dimension of (B, H / 2, W / 2, C). All patches are then concatenated along the channel dimension to obtain a tensor of size (B, H / 2, W / 2, 4C), and then a dimensional transformation is performed to reshape it to (B, H / 2×W / 2, 4C). Encoder stage 2 and encoder stage 3 use the output of the patch merging operation in the previous stage as input to gradually compress the features to higher dimensions. The output of each encoder stage is patch-expanded to ensure matching sizes when performing residual connections; The output of encoder stage i is concatenated with the patch expansion output of encoder stage i+1 along the channel dimension, i>1, and the concatenated output is input to the decoder; The information fused by the encoder is decoded by the decoder.
8. The single-frame composite structure illumination microscopy imaging method based on ensemble learning according to claim 7, characterized in that: Encoder stage 1 consists of 6 sliding window self-attention modules, encoder stage 2 consists of 4 sliding window self-attention modules, and encoder stage 3 consists of 4 sliding window self-attention modules.
9. The single-frame composite structure illumination microscopic imaging method based on ensemble learning according to claim 7, characterized in that: The odd sliding window self-attention module in each encoder stage performs multi-head self-attention mechanism calculation within the window; The even-block sliding window self-attention module performs sliding window multi-head self-attention mechanism calculation.
10. The single-frame composite structure illumination microscopy imaging method based on ensemble learning according to claim 9, characterized in that: The specific process of the odd sliding window self-attention module to calculate the multi-head self-attention mechanism within the window is: head i =Attention(Q,K,V) MutiHead(Q,K,V)=Concat(head1,…head h )W O Where Q, K, and V are matrices obtained by linear transformation of the input matrix. Attention, head, and MultiHead represent the operations of calculating the self-attention mechanism, the single-head attention mechanism, and the multi-head attention mechanism, respectively. k represents the dimension of Q, K, V; softmax is the normalized exponential function; W O Represents the weight matrix of the linear transformation of the multi-head attention mechanism; Concat means that head i In d k The concatenation operation is performed on the dimension, i = 1…h, where h represents the number of heads of the multi-head attention mechanism. The even-block sliding window self-attention module performs sliding window multi-head self-attention mechanism calculation. The specific process is as follows: Slide the input matrix one pixel to the right and downward, and fill the first row and column of the matrix with the third row and column before sliding. Then calculate the matrix using the odd sliding window self-attention module to perform the multi-head self-attention mechanism calculation process within the window.
Citation Information
Patent Citations
Bimodal microimaging system and method
CN111610621A
Superspeed structured light illumination super-resolution microscopic imaging device based on compressed sensing
CN114967092A
Wide-field illumination fluorescence super-resolution microscopic imaging method based on deep learning
CN115841423A
High-speed single-frame label-free cell tomography
CN116263411A
Video-level real-time structured light illumination super-resolution microscopic imaging method and system based on deep learning
CN116503246A