A method for super-resolution reconstruction of microscopic three-dimensional images based on structured light illumination

CN122529974APending Publication Date: 2026-08-07CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV OF POSTS & TELECOMM
Filing Date
2026-05-25
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]1、多数方法针对单层(单张)重建或把多层视为独立帧逐层处理,忽略了3D-SIM层间的物理耦合和信息互补,导致无法充分利用跨层频谱与结构信息来提升轴向和层间分辨率与去噪能力;

Benefits of technology

[0013]本申请是基于跨层多尺度融合和频域-空间混合注意力的设计,对不同层之间互补的频谱与结构信息进行相互校正,相对在训练时同时约束空间域和频谱域需要更低的计算成本,通过联合利用层间信息提升横向与轴向分辨率,重建质量明显优于逐层处理的单层方法;在3D-SIM数据集上的实验表明,本申请设计的方法在PSNR、SSIM等定量指标,以及在视觉主观质量例如伪影/纹理恢复上,均优于DFCAN、APCAN等单层方法及HAT、MAT等自然图像超分方法(无论是做单层还是多层重建时),并优于维纳重建在低SNR下的表现;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122529974A_ABST
    Figure CN122529974A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of microscopic imaging and image reconstruction, in particular to a kind of super-resolution reconstruction method based on structured light illumination microscopic three-dimensional image, comprising: using improved DRCT network based on Transformer to process the multilayer observation stack of structured light illumination microscopic image, to obtain the corresponding reconstruction stack;The DRCT network constructs space-frequency hybrid representation through complementary fusion, carries out feature fusion and residual learning to space-frequency hybrid representation on different scales, generates the final reconstruction image based on layer-by-layer reconstruction mechanism or joint reconstruction mechanism.This application is directly recovered from original multilayer observation to high-quality three-dimensional reconstruction image by simultaneously using space domain and frequency domain information in a single end-to-end framework and carrying out information fusion between multiple layers, and the collaborative optimization of super-resolution and denoising is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of microscopic imaging and image reconstruction technology, specifically to a super-resolution reconstruction method based on structured illumination microscopic three-dimensional images. Background Technology

[0002] Structured illumination microscopy (SIM) is an important tool for cell / subcellular imaging, as it enhances spatial resolution through stripe illumination and frequency domain synthesis. Post-processing of SIM often employs classic methods such as Wiener deconvolution to obtain a reference reconstruction. However, in cases of low signal-to-noise ratio (SNR) or complex features (point, line, or ring organelles), Wiener reconstruction is prone to residual noise and artifacts, making it difficult to balance resolution and visual interpretability.

[0003] In existing technologies, deep learning-based super-resolution and denoising methods (such as HAT and MAT for natural images, and DFCAN and APCAN for SIM or microscopic images) have achieved certain results in single-layer or layer-by-layer reconstruction tasks, but they have the following shortcomings:

[0004] 1. Most methods target single-layer (single frame) reconstruction or treat multiple layers as independent frames for layer-by-layer processing, ignoring the physical coupling and information complementarity between 3D-SIM layers. This results in the inability to fully utilize cross-layer spectral and structural information to improve axial and inter-layer resolution and denoising capabilities.

[0005] 2. The Transformer structure, which mainly uses super-resolution of natural images, does not fully consider the frequency domain characteristics and interlayer correlations in the microscopic SIM scene, and therefore has limited effectiveness in dealing with the spectral aliasing and stripe artifacts unique to SIM imaging. In the existing technology, some papers have proposed the DDL-SIM model, which improves the reconstruction fidelity by simultaneously constraining the spatial domain and the spectral domain during training. However, it requires the network to learn information from both domains at the same time, resulting in high overall computational cost. In addition, the difference in the supervision signals of the two domains may cause the model to learn false features that perform well in a single domain but are inconsistent globally, thus affecting the physical accuracy of the reconstruction results.

[0006] 3. Under low SNR conditions, existing methods (including classical Wiener) do not adequately balance noise suppression and structure preservation, often resulting in loss of detail or noise residue. Summary of the Invention

[0007] In view of this, this application discloses a super-resolution reconstruction method based on structured illumination microscopic 3D images to solve the problems in the prior art, comprising: processing the multi-layer observation stack of the structured illumination microscopic image using an improved Transformer-based DRCT network to obtain the corresponding reconstruction stack; the improved Transformer-based DRCT network processes the data including:

[0008] S1. Obtain multi-layer raw observation input of structured illumination microscopic image, preprocess the multi-layer raw observation input, and perform preliminary feature extraction to obtain initial structured illumination features;

[0009] S2. Extract spatial and spectral features from the initial structured illumination features, and construct a spatial-frequency hybrid representation through complementary fusion.

[0010] S3. Perform feature fusion and residual learning on the spatial-frequency domain hybrid representation at different scales to obtain the fused representation;

[0011] S4. Based on the fusion representation, a layer-by-layer reconstruction mechanism or a joint reconstruction mechanism is used to generate the final reconstructed image.

[0012] The beneficial effects of this application include:

[0013] This application is based on a cross-layer multi-scale fusion and frequency-space hybrid attention design, which mutually corrects the complementary spectral and structural information between different layers. Compared with constraining the spatial and spectral domains simultaneously during training, it requires lower computational cost. By jointly utilizing inter-layer information, it improves lateral and axial resolution, and the reconstruction quality is significantly better than single-layer methods that process layer by layer. Experiments on the 3D-SIM dataset show that the method designed in this application outperforms single-layer methods such as DFCAN and APCAN, as well as natural image super-resolution methods such as HAT and MAT (whether performing single-layer or multi-layer reconstruction), in terms of quantitative indicators such as PSNR and SSIM, and in terms of visual subjective quality such as artifact / texture recovery (whether performing single-layer or multi-layer reconstruction). It also outperforms Wiener reconstruction at low SNR.

[0014] By introducing adaptive frequency domain attention and multiple loss constraints, the network can more reliably suppress background noise and preserve the real structure under low SNR. Experiments show that it has advantages over Wiener reconstruction in terms of noise suppression and structural fidelity.

[0015] By employing a multi-scale Transformer and local convolution refinement design, the high-frequency information of point-like microstructures can be recovered while maintaining the continuity of linear / ring-like structures. The method has shown good results on a variety of biological samples and is adaptable to a variety of organelle morphologies, indicating that the method designed in this application has good generalization ability. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the structure of the Transformer-based DRCT network in the embodiments of this application;

[0017] Figure 2 This is a schematic diagram comparing the reconstruction effects of the method in this embodiment and the Wiener reconstruction method;

[0018] Figure 3 This is a comparative diagram of layer-by-layer reconstruction and multi-layer joint reconstruction in the embodiments of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, features, and advantages of this application clearer and to enable those skilled in the art to better understand the technical solutions of this application, the following detailed description of this application is provided in conjunction with the accompanying drawings and embodiments.

[0020] This embodiment includes a super-resolution reconstruction method based on structured illumination microscopic 3D images, comprising: processing a multi-layer observation stack of the structured illumination microscopic image using an improved Transformer-based DRCT (Deep Residual Convolution-Transformer) network to obtain a corresponding reconstruction stack; the improved Transformer-based DRCT network, such as... Figure 1 As shown, the data processing includes:

[0021] S1. Obtain multi-layer raw observation input of structured illumination microscopic image, preprocess the multi-layer raw observation input, and perform preliminary feature extraction to obtain initial structured illumination features.

[0022] In this embodiment, multi-layer observation data acquired from the original 3D-SIM dataset is selected as multi-layer raw observation input in sample TIFF files; the TIFF file is saved in "multi-layer stacking" format, with each sample containing multiple focal planes / phases; the preprocessing includes: intensity normalization (e.g., percentile normalization) for each sample to eliminate differences in acquisition intensity; data augmentation, such as random flipping, rotation, and simulating different levels of noise to improve robustness; if memory or video memory is limited, the image can be cropped in blocks and stitched together and averaged during inference to restore the complete result.

[0023] The preliminary feature extraction includes: passing the preprocessed sample sequentially through a shallow feature extraction module and a projection layer to obtain network-oriented features; and adding positional encoding to the network-oriented features.

[0024] The shallow feature extraction module can employ two-dimensional convolution, three-dimensional convolution, layer-wise shared convolution, or residual blocks composed of multiple convolutional layers connected together. Samples are mapped to a unified feature dimension through convolutional projection or linear projection.

[0025] The position encoding is used to preserve the relative positional relationship of each focal plane. The position encoding can be a learnable position encoding, a sinusoidal position encoding, or a layer number encoding generated from the focal plane number.

[0026] The preliminary feature extraction formula is expressed as follows:

[0027]

[0028] in, Indicates multiple layers of original observation input. This indicates a preprocessing normalization operation. This represents the shallow feature extraction module and the projection layer. Indicates position code, This represents the initial structure lighting characteristics obtained.

[0029] The final obtained initial structured illumination features At the same time, the original signal characteristics of each layer and the relative position information between each layer are preserved.

[0030] S2. Extract spatial and spectral domain features from the initial structured illumination features, and construct a spatial-frequency hybrid representation through complementary fusion; specifically:

[0031] The spatial domain features are obtained by applying local window attention, convolutional coding, or residual coding to X, as shown in the formula:

[0032]

[0033] Where Bs(·) represents the spatial domain coding module, This represents spatial domain features, which are used to preserve the spatial differences between filament edges, particle morphology, pore ring structure, and background and structure, expressing the local morphology and spatial continuity of organelle structures.

[0034] The spectral domain features are obtained by extracting the amplitude-frequency information of X through discrete frequency domain transformation. The amplitude-frequency information includes amplitude features, phase features, high-frequency components, and low-frequency components. The discrete frequency domain transformation can employ two-dimensional Fourier transform or three-dimensional Fourier transform, as shown in the following formula:

[0035]

[0036] Where FFT(·) represents frequency domain transform, and this embodiment uses Fourier transform, and Bf(·) represents the spectrum domain coding module. It represents the spectral domain characteristics, used to express periodic stripes, high-frequency structures, aliasing components, and noise distribution in SIM data;

[0037] High-frequency components are used to describe details and edges, while low-frequency components are used to describe overall brightness and background trends. By modeling the high-frequency region, the stripe-related frequency region, and the low-frequency background region separately, the network can better distinguish between real structures and stripe artifacts or aliasing artifacts.

[0038] To align the spectral domain features with the spatial domain features, an inverse frequency domain transformation or convolution mapping is used to make the dimensions of the spectral domain features consistent with those of the spatial domain features. The formula is as follows:

[0039]

[0040] right and Complementary fusion is performed to obtain a hybrid spatial-frequency domain characterization.

[0041] Specifically, the complementary fusion first involves splicing... and The concatenated features are processed through convolutional mapping and nonlinear activation, and then used as gated fusion weights; based on these gated fusion weights... and Gated fusion is performed; complementary fusion design allows the network to adaptively select spatial domain information or spectral domain information based on the structure and noise conditions of different regions. The formula is:

[0042]

[0043]

[0044] in, Indicates to and Feature concatenation is performed, where Conv(·) represents convolution mapping, σ(·) represents the Sigmoid function, G represents the gate weights, and ⊙ represents element-wise multiplication. This represents a hybrid spatial-frequency domain characterization. By distinguishing between real structures and stripe / aliasing artifacts in the spatial and frequency domains, it enhances the suppression of SIM-specific noise and artifacts.

[0045] S3. Perform feature fusion and residual learning on the spatial-frequency domain hybrid representation at different scales to obtain the fused representation.

[0046] For features at different scales, shallow features typically retain more texture and edge information from the spatial-frequency domain hybrid representation, while deep features usually have a larger receptive field and stronger noise suppression capabilities. Through multi-scale fusion, the model can suppress noise while preserving details, taking into account the recovery of various microscopic structures such as point-like, line-like, and ring-like structures. Specifically, multi-scale feature fusion can be achieved through a combination of downsampling, upsampling, window attention, convolutional residual blocks, and cross-scale connection modules. In this embodiment, the formulas for feature fusion and residual learning are as follows:

[0047]

[0048] Among them, F k R represents the feature of the k-th processing stage. k ( ) represents the residual convolution module, T k ( ) represents the Transformer feature modeling module. Residual connections are used to preserve the original details, alleviate the degradation of deep network training, and thus avoid the problem of overfitting or loss of fidelity in structured light micrograph details.

[0049] Furthermore, information at different depths and at different focal planes has a mapping relationship. To enable mutual compensation and correction of information between different depths and at different focal planes, and to improve axial consistency and interlayer resolution, this application also sets up cross-layer flow paths between different feature extraction layers. Taking the form of adjacent layer feature correction as an example, each feature extraction stage receives both the spatial-frequency domain fusion features of the current stage and the compensation features from adjacent focal planes or the previous stage; for the features of the d-th focal plane... Supplementary information is extracted from layers d-1 and d+1, and fused using attention or convolution to obtain the cross-layer compensated features of layer d. The formula is as follows:

[0050]

[0051] in, Represents the cross-layer flow function. This represents the feature of the d-th layer after cross-layer compensation.

[0052] Cross-layer flow paths can be achieved based on adjacent focal plane feature exchange, cross-layer attention, 3D convolution, and cross-depth residual connections. In this way, the model can simultaneously utilize local texture, deep structural information, and inter-layer related information. This design can reduce the impact of single-layer noise on the reconstruction results and improve the structural continuity in the axial direction.

[0053] S4. Based on the fusion representation, a layer-by-layer reconstruction mechanism or a joint reconstruction mechanism is adopted to generate the final reconstructed image.

[0054] The layer-by-layer reconstruction mechanism refers to predicting the reconstruction results of each focal plane separately and combining them into a three-dimensional stack according to the layer sequence. Layer-by-layer reconstruction makes it easier to control the reconstruction quality of each layer.

[0055] The joint reconstruction mechanism refers to outputting multi-layer reconstruction results or multi-channel reconstruction results at the end of the model at one time. Joint reconstruction can make fuller use of inter-layer correlation.

[0056] The formula for outputting the reconstruction result is: ,in, The final fusion representation is represented by Rθ(·), which represents the reconstruction head, which can be composed of convolutional layers, upsampling layers, subpixel convolutional layers, and deconvolutional layers. Ŷ represents the reconstruction result output by the network. If the improved Transformer-based DRCT network is cropped into multiple image patches before inference, the network reconstructs each image patch separately and obtains the complete image by averaging the overlapping regions. Finally, the reconstruction result can be saved as a multi-layer TIFF file as needed.

[0057] Furthermore, the improved Transformer-based DRCT network is pre-trained before use, and the model is trained using a composite loss function; the composite loss function includes: pixel reconstruction loss, structure preservation loss, and high-frequency / texture enhancement loss. For the network output... References and annotations voxels or number of pixels :

[0058] The pixel reconstruction loss is used to measure numerical deviation and employs the L1 loss function, with the following formula:

[0059]

[0060] The structure preservation loss, used to measure structural similarity, is formulated as follows:

[0061]

[0062] In some embodiments, the structural retention loss can also be replaced by the layer-by-layer SSIM average value, in which case the structural retention loss can be expressed as:

[0063] The high-frequency enhancement loss, used to improve detail and texture recovery, is expressed by the following formula:

[0064]

[0065] in, This represents a high-frequency operator used to preserve high-frequency information; it can be a Laplacian operator, a Gaussian high-pass filter, a Sobel gradient operator, a Laplacian pyramid, or a frequency domain high-pass mask.

[0066] Furthermore, the total loss function can be written as:

[0067]

[0068] in, , and These represent the pixel reconstruction loss weights, structure preservation loss weights, and high-frequency enhancement loss weights, respectively.

[0069] For pixel reconstruction loss, it can be maintained It is a constant, and can also be appropriately reduced when structural losses and high-frequency losses increase. To avoid prematurely amplifying noise and artifacts in the early stages of training, the weights can be adjusted according to the training phase. As a preferred embodiment, a large pixel reconstruction loss weight is maintained in the early stages of training. The perceptual loss weight is gradually increased, initially focusing on pixel loss to ensure pixel-level convergence. After the network output stabilizes, the structure preservation loss weight and high-frequency enhancement loss weight are gradually increased to avoid prematurely amplifying artifacts. This is achieved based on the following linear ramp function with a saturation region, the formula of which is:

[0070]

[0071] Where t represents the current training round or iteration number, Indicates the length of the weight preheating stage. Indicates the first The initial weights of the loss terms, Indicates the first The target weight of the loss item.

[0072] Training optimization employs an adaptive first-order optimization algorithm. Preferably, the Adam or AdamW optimization algorithm is used. The learning rate can employ a cosine annealing, multi-step decay, or linear warm-up followed by decay strategy.

[0073] To improve efficiency and stability, a hybrid precision training and gradient norm control strategy can be employed during training. Gradient norm control can be written as... ,in Indicates the current gradient. This indicates a preset threshold.

[0074] Training data can consist of low signal-to-noise ratio (SNR) raw observation inputs and corresponding reference labels. Reference labels can be obtained through classical reconstruction, Wiener filtering reconstruction, multiple acquisition averaging reconstruction, or manual correction reconstruction from high SNR structured illumination microscopy data. These reference labels are used for supervised learning training and are not limited to absolute true values ​​in a real physical sense.

[0075] Furthermore, comparative tests were conducted on the methods involved in this application, and the test results are as follows: Figure 2 As shown, compared to the traditional Wiener reconstruction, the method designed in this application achieves better reconstruction results at low signal-to-noise ratios. Figure 3 In the image, the left side shows a schematic diagram of traditional layer-by-layer reconstruction, while the right side shows a schematic diagram of multi-layer joint reconstruction used in this application. It can be seen that the image reconstructed by the method designed in this application has stronger noise resistance and higher resolution. The inherent textures such as tissue structure are free from breaks, distortions, and artifacts, and the texture integrity is better.

[0076] Finally, it should be noted that the above description only depicts some embodiments of this application. For those skilled in the art, various changes, modifications, substitutions, and variations can be conceived of these embodiments without departing from the principles and spirit of this application. The scope of protection of this application is defined by the appended claims and their equivalents, and all the above-mentioned behaviors should be covered within the scope of protection of this application.

Claims

1. A super-resolution reconstruction method based on structured light illumination microscopic 3D images, characterized in that, include: An improved Transformer-based DRCT network was used to process the multi-layer observation stack of structured illumination microscopy images to obtain the corresponding reconstruction stack. The improved Transformer-based DRCT network processes data in the following ways: S1. Obtain multi-layer raw observation input of structured illumination microscopic image, preprocess the multi-layer raw observation input, and perform preliminary feature extraction to obtain initial structured illumination features; S2. Extract spatial and spectral features from the initial structured illumination features, and construct a spatial-frequency hybrid representation through complementary fusion. S3. Perform feature fusion and residual learning on the spatial-frequency domain hybrid representation at different scales to obtain the fused representation; S4. Based on the fusion representation, a layer-by-layer reconstruction mechanism or a joint reconstruction mechanism is adopted to generate the final reconstructed image.

2. The super-resolution reconstruction method based on structured light illumination microscopic three-dimensional images according to claim 1, characterized in that, The preliminary feature extraction includes: passing the preprocessed sample sequentially through a shallow feature extraction module and a projection layer to obtain network-oriented features; and adding positional encoding to the network-oriented features.

3. The super-resolution reconstruction method based on structured light microscopic three-dimensional images according to claim 2, characterized in that, The shallow feature extraction module maps samples to a unified feature dimension through convolutional projection or linear projection.

4. The super-resolution reconstruction method based on structured light microscopic three-dimensional images according to claim 2, characterized in that, The position encoding is used to preserve the relative positional relationship of each focal plane.

5. The super-resolution reconstruction method based on structured light microscopic three-dimensional images according to claim 4, characterized in that, The position coding employs learnable position coding, sinusoidal position coding, or layer number coding generated from the focal plane number.

6. The super-resolution reconstruction method based on structured light microscopic three-dimensional images according to claim 1, characterized in that, The construction of a spatial-frequency domain hybrid representation through complementary fusion includes: transforming spectral domain features through inverse frequency domain transformation or convolutional mapping. Size and spatial domain characteristics Consistency; To and Complementary fusion is performed to obtain a spatial-frequency domain hybrid characterization; the complementary fusion includes: splicing and The concatenated features, after convolutional mapping and nonlinear activation, are used as gated fusion weights. Based on these gated fusion weights, the system... and Perform gating fusion.

7. The super-resolution reconstruction method based on structured light illumination microscopic three-dimensional images according to claim 1, characterized in that, The feature fusion process involves establishing cross-layer flow paths between different feature extraction layers, for features of the d-th focal plane. Supplementary information is extracted from the (d-1)th and (d+1)th layers and fused through attention or convolution to obtain the cross-layer compensated features of the dth layer.

8. The super-resolution reconstruction method based on structured light microscopic three-dimensional images according to claim 1, characterized in that, If the input to the improved Transformer-based DRCT network is cropped into multiple image patches before inference, the network reconstructs each image patch separately and obtains the complete image by averaging the overlapping regions.

9. The super-resolution reconstruction method based on structured light microscopic three-dimensional images according to claim 1, characterized in that, The improved Transformer-based DRCT network is pre-trained before use and employs a composite loss function, as shown in the formula: ; in, Represents the total loss function. This represents the pixel reconstruction loss, used to measure numerical deviation. This represents the structure preservation loss, used to measure structural similarity. This indicates high-frequency enhancement loss, used to improve detail and texture recovery. , and These represent the pixel reconstruction loss weights, structure preservation loss weights, and high-frequency enhancement loss weights, respectively.

10. The super-resolution reconstruction method based on structured light microscopic three-dimensional images according to claim 9, characterized in that, During training, after the network output stabilizes, the weights of structure preservation loss and high-frequency reinforcement loss are gradually increased.