A method for super-resolution reconstruction of microscopic images under distorted structured light illumination

CN122780102APending Publication Date: 2026-09-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611218809.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-12
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]有鉴于此,本发明的目的在于提供一种面向畸变结构光照明显微图像的超分辨重建方法,用于解决现有SIM重建过程中原始图像存在畸变、错位或配准误差时,照明相位先验失效、参数估计不稳定以及重建结果易出现条纹伪影和细节损失的问题

Benefits of technology

本发明面向存在平移、旋转、尺度变化、剪切及透视等几何畸变的结构光照明显微原始图像进行超分辨重建,通过对多帧SIM图像进行统一预处理、三级残差单应性配准、有效区域筛选以及幅相联合注意力融合,实现帧间几何误差校正与互补结构信息的有效利用。相较于依赖精确照明相位、条纹方向、调制度等参数估计的传统SIM重建方法,本发明通过建立多帧畸变SIM图像与高质量超分辨图像之间的映射关系,降低了照明参数误差以及帧间错位对重建结果的影响,提高了复杂成像条件下的适应能力。通过三级残差单应性配准对较大范围几何畸变及剩余配准误差进行逐级校正,并利用有效区域掩膜排除空间变换后形成的无效边界,能够减少边缘拉伸、黑边及错误特征进入后续重建过程。进一步地,通过从频域幅值和频域相位两个方面提取特征并进行联合注意力加权,同时结合各帧有效区域及可靠性进行多帧特征融合,有利于充分利用不同照明方向和相位下的互补高频信息,并抑制残余错位、异常频谱及噪声造成的条纹伪影、蜂窝状伪影和结构断裂现象。配合像素一致性和结构相似性约束,使输出图像在保持灰度一致性的同时具有较好的细节恢复能力和结构连续性。由此,本发明能够减少传统SIM重建过程中复杂的人工参数调节和反复迭代处理,提高畸变SIM图像重建的自动化程度和稳定性,适用于复杂成像条件下结构光照明显微图像的快速超分辨重建。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780102A_ABST
    Figure CN122780102A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of distortion structure light illumination microscopic image super-resolution reconstruction method, belong to structure light illumination microscopic imaging and super-resolution image reconstruction technical field.The method is by obtaining and pre-processing multi-direction, multi-phase SIM original image, constructs training sample with interframe non-uniform geometric distortion;Select reference frame, carry out three-level residual homography registration to the rest frame and generate effective area mask;Then the spectral amplitude and phase feature of registration image are extracted, and the effective area and frame reliability are combined to carry out multi-frame feature fusion;Finally, based on registration loss and reconstruction loss, the model is trained and jointly fine-tuned, and the image to be reconstructed is registered, feature fusion and spatial upsampling, and output super-resolution image.The present application can improve the registration stability, detail recovery ability and structure continuity of SIM image under complex distortion condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of structured illumination micro-imaging and super-resolution image reconstruction technology, and relates to a super-resolution reconstruction method for distorted structured illumination micro-images. Background Technology

[0002] Structured illumination microscopy (SIM) typically uses multi-directional, multi-phase structured light to modulate high-frequency information of the sample into the detectable frequency band of the microscope system. Super-resolution images are then obtained through steps such as phase estimation, modulation estimation, spectral shifting, Wiener filtering, and spectral fusion. Existing traditional reconstruction algorithms such as OpenSIM have systematically streamlined this process, but they usually require spatial consistency between original images and rely on accurate illumination fringe period, direction, and phase parameters. In actual imaging, optical path aberrations, changes in sample refractive index, mechanical errors, or multi-channel acquisition errors can cause distortion, misalignment, and illumination pattern shifts in the original SIM images, rendering fixed phase priors and parameter estimations unreliable. This leads to fringe artifacts, honeycomb artifacts, and loss of detail in the reconstruction results. For unknown or distorted illumination, methods such as PE-SIMS attempt to reduce reliance on precise illumination patterns through pattern estimation and statistical priors, but these methods still suffer from problems such as complex iterative calculations, strong prior assumptions, and insufficient adaptability to complex distorted original frames.

[0003] In recent years, learning-based SIM reconstruction methods such as ML-SIM and AL-SIM have improved reconstruction speed, noise resistance, and artifact suppression capabilities. However, they may still suffer from insufficient generalization under conditions of real distortion, registration error, and changes in training distribution. Therefore, there is an urgent need for a stable super-resolution reconstruction method for micro-images with significant distortion and illumination. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a super-resolution reconstruction method for micro-images with significant structural illumination distortion. This method addresses the problems of illumination phase prior failure, unstable parameter estimation, and the tendency for fringe artifacts and detail loss in reconstruction results when the original image has distortion, misalignment, or registration errors. This method models the mapping relationship between the distorted original SIM image and the super-resolution image, achieving stable reconstruction of SIM images under complex distortion conditions and improving the structural continuity and detail recovery capability of the reconstructed image.

[0005] To achieve the above objectives, the present invention provides the following technical solution: A super-resolution reconstruction method for micro-images with significant structural distortion under illumination, comprising the following steps: S1. Acquire multiple frames of low-resolution SIM original images of the same sample under multiple illumination directions and multiple illumination phases. Sort the multiple frames of low-resolution SIM original images according to illumination direction and illumination phase, and perform background subtraction and intensity normalization respectively to form network input data. S2. Based on the ideal SIM images corresponding to the original low-resolution SIM images of multiple frames, construct distorted SIM training samples. Apply geometric transformations to each frame of ideal SIM images and superimpose intensity perturbation, background offset and simulated acquisition noise to obtain multi-frame training images with non-uniform distortion between frames. Use the corresponding distortion-free super-resolution reference image as the training label. S3. Determine the reference frame from the multi-frame training images, and take the remaining frames as the frames to be registered. Extract the image features of the reference frame and each frame to be registered respectively, and use three-level residual homography registration to estimate the geometric deviation of each frame to be registered relative to the reference frame step by step. Based on the residual homography matrices at each level, obtain the final homography matrix corresponding to each frame to be registered, and perform spatial transformation on each frame to be registered according to the final homography matrix to obtain multi-frame registered images. S4. Based on the reverse mapping sampling coordinates of each frame to be registered when performing spatial transformation, determine the effective sampling area in each registered image, and generate an effective area mask corresponding to each registered image to distinguish the effective image area and the invalid boundary area after spatial transformation. S5. Perform feature extraction on the multi-frame registered images, convert the obtained intermediate features to the frequency domain and extract the spectral amplitude features and spectral phase features respectively, generate amplitude-phase joint attention weights based on the spectral amplitude features and spectral phase features, and determine the multi-frame fusion weights by combining the effective region mask and the reliability score of each frame. Perform weighted fusion of the features of each frame based on the multi-frame fusion weights to obtain the fused features. S6. The registration network and the super-resolution reconstruction network are trained based on the distorted SIM training samples. The registration network is constrained by photometric consistency loss, normalized cross-correlation loss, gradient consistency loss, multi-frame template consistency loss and homography matrix regularization loss. The super-resolution reconstruction network is constrained by pixel absolute error loss and structural similarity loss. The registration network and the super-resolution reconstruction network are jointly fine-tuned to obtain the trained super-resolution reconstruction model. S7. Input the original multi-frame distorted SIM images to be reconstructed into the trained super-resolution reconstruction model. After background subtraction and intensity normalization, three-level residual homography registration, effective region mask generation, amplitude-phase joint attention feature recalibration, and multi-frame feature weighted fusion, the obtained deep fusion features are spatially upsampled to output the corresponding super-resolution reconstructed image.

[0006] Furthermore, in step S1, the first The lighting direction, the first The actual SIM observation image corresponding to each illumination phase is represented as follows:

[0007] In the formula, Indicates the first The lighting direction, the first Actual SIM observation images corresponding to each illumination phase; This represents an ideal SIM image formed under corresponding lighting conditions; This indicates the geometric distortion mapping applied to the image frame. The geometric distortion mapping includes one or more of translation, rotation, scaling, shearing, or perspective distortion. This represents the background component in the observed image; This represents the noise component in the observed image; All image frames are numbered using a combination of illumination direction index and illumination phase index, with each illumination direction... Phase with illumination The corresponding index pair number is Establish a multi-frame SIM raw image set The original SIM images from multiple frames are stacked sequentially along the frame channel dimension to obtain the network input tensor. ; Then, background subtraction is performed on each frame of the original image. The background subtraction result is represented as follows:

[0008] In the formula, Indicates the first Frame of original SIM image at pixel location The grayscale value at that location; This represents the grayscale value after background subtraction. Indicates the first Background estimation of the frame image; , These represent the x and y coordinates of a pixel, respectively. Then, normalization is performed based on the grayscale range of the current frame:

[0009] In the formula, Represents the normalized i-th Frame image; and They represent the first The minimum and maximum gray values ​​of the image within the effective area after background subtraction; This indicates a positive number set to prevent the denominator from being zero.

[0010] Furthermore, in step S2, a homography transformation is constructed for each frame of the ideal SIM image, then the... The homography matrix corresponding to the frame is represented as follows:

[0011] In the formula, Indicates the action on the first The homography transformation matrix of a frame SIM image; This represents the translation matrix, which consists of the horizontal translation amount. and longitudinal translation Sure; This represents a rotation matrix, which is composed of rotation angles. Sure; This represents the scaling transformation matrix, which consists of the horizontal scaling factor. and longitudinal scale factor Sure; The shear transformation matrix is ​​represented by the shear parameters. and Sure; This represents the perspective transformation matrix, used to control the degree of projection distortion; Image resampling is performed based on the homography matrix, and inverse mapping is applied to the output pixel coordinates; let the homogeneous coordinates of the output pixels be... The corresponding original image sampling position is:

[0012]

[0013]

[0014] In the formula, This represents the homogeneous coordinates obtained after inverse mapping using the homography matrix; , This represents the actual sampling coordinates corresponding to the original image; based on the actual sampling coordinates Perform bilinear interpolation to obtain the first... Frame geometric distortion image ; The training image is obtained by further superimposing intensity perturbation, background shift, and simulated acquisition noise on the geometrically distorted image:

[0015] In the formula, The first term represents the result after geometric distortion, intensity variation, and noise degradation. Training images; Indicates the light intensity perturbation coefficient; Indicates the additional background offset; This represents the simulated acquisition noise; This means that the pixel value is limited to the range of 0 to 1.

[0016] Furthermore, in step S3, one frame is selected from the processed training images as a reference frame, and the remaining images are used as frames to be registered. The reference frame and any frame to be registered are mapped to the same feature space through a feature extraction network with shared parameters, forming paired features.

[0017] In the formula, Indicates a reference frame; Indicates the first Frame of images to be registered; This represents a feature extraction network with shared parameters; Indicates splicing along the channel; Represents paired features used in homography matrix regression; A three-level residual homography registration is used to perform geometric correction on the frames to be registered, from coarse to fine. The first level is used to estimate the global homography transformation over a larger range, while the second and third levels estimate the remaining geometric errors based on the registration results of the previous level. The cumulative homography matrix is ​​updated level by level according to the following formula:

[0018] In the formula, Indicates completion of the first The cumulative homography matrix after level residual correction; Indicates the first The residual homography matrix output by the homography matrix regressor; After completing the third-level residual correction, the first The final homography matrix and the final registered image of the frame are represented as follows:

[0019] In the formula, Indicates the first The final homography matrix obtained by the frame after three-level residual correction; This represents a spatial transformation operation performed based on homography and bilinear interpolation. Indicates the first The final registered image of the frame.

[0020] Furthermore, in step S4, the binary effective region mask generated based on the reverse mapping coordinates is as follows:

[0021] In the formula, Indicates the first Frame-registered image at pixel location The valid area is marked at the location; , This represents the corresponding reverse-mapped sampling coordinates; and These represent the width and height of the original image, respectively; the effective region is set to 1, the invalid region is set to 0, and the effective region mask corresponding to the reference frame is always set to 1.

[0022] Furthermore, in step S5, the nine registered images are combined along the channel dimension in a predetermined order, and initial reconstructed features are extracted through shallow convolutional units:

[0023]

[0024] In the formula, This represents the joint input consisting of nine registered images. This represents the total number of image frames. This represents a shallow convolutional feature extraction unit; This represents the initial reconstructed features obtained; For the intermediate features in the reconstructed network, perform a two-dimensional Fourier transform on each feature channel:

[0025] In the formula, Indicates the first Spatial domain characteristics of each channel; Represents a two-dimensional Fourier transform; Indicates the first Frequency domain characteristics corresponding to each channel; Extract the spectral amplitude and spectral phase based on the frequency domain characteristics:

[0026]

[0027] In the formula, Indicates the first The spectral amplitude of each channel; Indicates the first The spectral phase of each channel; Represents the modulus of a complex number; Represents complex phase; Statistical and nonlinear mappings are performed on the spectral amplitude and spectral phase features respectively to generate amplitude attention weights and phase attention weights, and a joint attention weight is obtained:

[0028]

[0029]

[0030] In the formula, Represents the set of spectral amplitude features for multiple characteristic channels; Represents the set of spectral phase features for multiple characteristic channels; and These represent the statistical and nonlinear mappings of the magnitude branch and the phase branch, respectively. Indicates the magnitude attention weight; Indicates phase attention weights; Represents the Sigmoid function; Indicates the amplitude-phase joint attention weight; The generated effective region mask is scaled to the current feature size, and pixel-level frame fusion weights are constructed by combining the frame reliability score predicted by the network:

[0031] In the formula, Indicates the first Frame at pixel position The fusion weight at the location; This indicates scaling to the current feature size. Frame effective area mask; The network prediction of the first The reliability score of the frame at the corresponding location; Represents an exponential function; Finally, a fusion weight is used to perform weighted fusion of the features from the nine frames:

[0032] In the formula, Indicates the first Features of frames after amplitude-phase joint attention recalibration; This represents the features after weighted fusion of nine frames.

[0033] Furthermore, in step S6, the total loss function of the registration network is expressed as:

[0034] In the formula, Indicates the total registration loss; This indicates a loss of photometric uniformity; This represents the normalized cross-correlation loss; This represents the gradient consistency loss; This represents the template consistency loss across multiple frames; This represents the regularization loss of the homography matrix; , , , and These represent the weights of the corresponding loss terms; The total reconstruction loss function of the super-resolution reconstruction network is expressed as:

[0035] In the formula, Indicates super-resolution reconstruction loss; This represents the mean absolute error between the model's output image and the reference image. This represents the super-resolution image output by the model. This represents a distortion-free, high-quality super-resolution reference image. Indicators representing structural similarity; and These represent the weights of the pixel absolute error loss and the structural similarity loss, respectively.

[0036] Furthermore, the training process in step S6 includes, in sequence, training the registration network separately, freezing the registration network to train the super-resolution network, and performing end-to-end joint fine-tuning. The total loss of the joint fine-tuning is expressed as:

[0037] In the formula, This represents the total loss during the end-to-end joint fine-tuning phase; This indicates the weight of the registration loss in the joint training.

[0038] Furthermore, in step S7, for the obtained deep fusion features Channel mapping convolution is used to map the deep fusion features to the number of channels required for pixel rearrangement, and then spatial upsampling is completed through pixel rearrangement operation:

[0039] In the formula, This represents the channel-mapped convolution before upsampling; This indicates a pixel rearrangement operation with a magnification factor of 2. This represents the features after spatial upsampling; Finally, the super-resolution reconstructed image is obtained by output convolution:

[0040] In the formula, Indicates the output convolutional layer; This represents the final super-resolution reconstructed image.

[0041] The beneficial effects of this invention are as follows: This invention addresses super-resolution reconstruction of micro-images with significant structured illumination distortions, including translation, rotation, scale changes, shearing, and perspective distortions. By performing unified preprocessing on multiple SIM images, three-level residual homography registration, effective region selection, and amplitude-phase joint attention fusion, it achieves inter-frame geometric error correction and effective utilization of complementary structural information. Compared to traditional SIM reconstruction methods that rely on precise estimation of parameters such as illumination phase, fringe direction, and modulation, this invention establishes a mapping relationship between multiple distorted SIM images and high-quality super-resolution images, reducing the impact of illumination parameter errors and inter-frame misalignment on the reconstruction results and improving adaptability under complex imaging conditions. Three-level residual homography registration progressively corrects large-scale geometric distortions and residual registration errors, and the use of effective region masks to exclude invalid boundaries formed after spatial transformations reduces edge stretching, black borders, and erroneous features from entering the subsequent reconstruction process. Furthermore, by extracting features from both frequency domain amplitude and frequency domain phase and applying joint attention weighting, while simultaneously fusing multi-frame features based on the effective region and reliability of each frame, this approach effectively utilizes complementary high-frequency information under different illumination directions and phases, and suppresses residual misalignment, anomalous spectra, and noise-induced stripe artifacts, honeycomb artifacts, and structural breaks. Combined with pixel consistency and structural similarity constraints, the output image maintains grayscale consistency while exhibiting good detail recovery and structural continuity. Therefore, this invention reduces the complex manual parameter adjustments and iterative processing required in traditional SIM reconstruction, improves the automation and stability of distorted SIM image reconstruction, and is suitable for rapid super-resolution reconstruction of micro-images with significant structured illumination under complex imaging conditions.

[0042] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0043] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 This is a schematic diagram of the overall process of the super-resolution reconstruction method for micro-images with obvious illumination of distorted structures according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the three-level residual homography registration process according to an embodiment of the present invention. Detailed Implementation

[0044] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0045] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0046] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0047] Please see Figures 1-2 This is a super-resolution reconstruction method for micro-images with obvious distorted structures and illumination.

[0048] Example This embodiment details the implementation process of a super-resolution reconstruction method for micro-images with significant structural distortion and illumination. Figure 1As shown, the overall process includes SIM raw image acquisition and preprocessing, distortion training sample construction, residual homography registration, effective region selection, multi-frame feature fusion, model joint training, and super-resolution image output. Specifically, firstly, multiple low-resolution SIM raw images of the same sample under multiple illumination directions and phases are acquired and arranged according to a predetermined illumination direction and phase order. Background subtraction and intensity normalization are performed on each frame to form unified network input data. Subsequently, based on the corresponding ideal SIM image, translation, rotation, scaling, shearing, and perspective transformations are applied independently to each frame, and intensity perturbation, background shift, and noise are superimposed to construct training samples that can characterize the non-uniform distortion between frames. For the multi-frame SIM images to be processed, a reference frame is selected, and the remaining images are used as registration frames. Large-scale geometric distortion and residual registration errors are corrected step by step through three-level residual homography registration. At the same time, an effective region mask is generated according to the effective sampling range after spatial transformation to eliminate the influence of invalid boundaries on the subsequent reconstruction process. After geometric registration, features are extracted from each frame, obtaining feature representations from both frequency domain amplitude and frequency domain phase. Adaptive recalibration of different features is performed through amplitude-phase joint attention. Then, multi-frame features are weighted and fused using an effective region mask and the reliability scores of each frame to fully utilize complementary information between images with different illumination directions and phases, while suppressing residual misalignment and noise. During training, photometric consistency, normalized cross-correlation, gradient consistency, multi-frame template consistency, and homography matrix regularization are used to constrain the registration network, while pixel error and structural similarity are used to constrain the super-resolution reconstruction network. Staged training and joint fine-tuning enable collaborative optimization of the registration and super-resolution reconstruction modules. Finally, the fused deep features are spatially upsampled through channel mapping and pixel rearrangement, and then convolved with the output to obtain the final super-resolution reconstructed image, thus forming a complete reconstruction process from multi-frame distortion correction, effective region constraint, amplitude-phase feature fusion to super-resolution output.

[0049] This embodiment addresses potential translation, rotation, scaling, shearing, and perspective distortions in nine frames of original SIM images across three directions and three phases. It constructs a complete technical processing chain, sequentially performing distortion sample generation, three-level residual homography registration, effective region selection, and amplitude-phase joint attention fusion, ultimately achieving end-to-end super-resolution reconstruction. This method first explicitly corrects inter-frame geometric distortions, then directly reconstructs the super-resolution image using the complementary information from the nine frames, thereby suppressing residual misalignments, invalid boundaries, noise, and anomalous spectral features.

[0050] In this embodiment, nine frames of original SIM images are first acquired and ordered input is constructed. Nine low-resolution SIM images of the same sample are acquired under three illumination directions and three illumination phases. The lighting direction, the first The actual SIM observation image corresponding to each illumination phase is represented as follows:

[0051] In the formula, Indicates the first The lighting direction, the first Actual SIM observation images corresponding to each illumination phase; This represents an ideal SIM image formed under corresponding lighting conditions; This indicates the geometric distortion mapping applied to the image frame. The geometric distortion mapping includes one or more of translation, rotation, scaling, shearing, or perspective distortion. This represents the background component in the observed image; This represents the noise component in the observed image; in this embodiment, .

[0052] The nine original SIM images are uniformly numbered according to a fixed "illumination direction - illumination phase" sequence:

[0053] In the formula, This represents the uniform frame index, with values ​​ranging from 1 to 9. This represents a set of nine SIM images. The three illumination directions are 0°, 60°, and 120°, respectively, and the three illumination phases under each illumination direction are 0°, 120°, and 240°, respectively.

[0054] Stack the nine images along the frame channel dimension in the above order to obtain the network input tensor:

[0055] In the formula, This represents the network input tensor composed of nine SIM images; This indicates that the input channels are stacked along the channel dimension in a predetermined frame order. Through the above processing, each input channel maintains a fixed mapping relationship with its corresponding illumination direction and illumination phase.

[0056] This clarifies the correspondence between each input channel and the illumination direction and phase, rather than simply treating multiple frames of images as channel inputs in general.

[0057] Then, preprocessing and intensity normalization are performed on the original images. Background subtraction and intensity normalization are performed on each frame of the original image. The background subtraction result is expressed as:

[0058] In the formula, Indicates the first Frame of original SIM image at pixel location The grayscale value at that location; This represents the grayscale value after background subtraction. Indicates the first Background estimation of the frame image; , These represent the horizontal and vertical coordinates of the pixel, respectively.

[0059] Then, normalization is performed based on the grayscale range of the current frame:

[0060] In the formula, Represents the normalized i-th Frame image; and They represent the first The minimum and maximum gray values ​​of the image within the effective area after background subtraction; This indicates a positive number set to prevent the denominator from being zero. This step is only used to unify the background and intensity range of the nine frames and does not assume or recover the precise structured light phase during the preprocessing stage.

[0061] Next, distortion SIM training samples are generated based on homography transformation. To ensure that the training samples cover the inter-frame inconsistencies in geometric errors that may occur during actual acquisition, a homography transformation is constructed for each frame of the ideal SIM image. The homography matrix corresponding to the frame is represented as follows:

[0062] In the formula, Indicates the action on the first The homography transformation matrix of a frame SIM image; This represents the translation matrix, which consists of the horizontal translation amount. and longitudinal translation Sure; This represents a rotation matrix, which is composed of rotation angles. Sure; This represents the scaling transformation matrix, which consists of the horizontal scaling factor. and longitudinal scale factor Sure; The shear transformation matrix is ​​represented by the shear parameters. and Sure; This represents the perspective transformation matrix, used to control the degree of projection distortion.

[0063] When resampling an image based on the homography matrix, a reverse mapping is performed on the output pixel coordinates. Let the homogeneous coordinates of the output pixel be... The corresponding original image sampling position is:

[0064] In the formula, This represents the homogeneous coordinates obtained after inverse mapping using the homography matrix; , This represents the actual sampling coordinates corresponding to the original image. Based on the actual sampling coordinates... Perform bilinear interpolation to obtain the first... Frame geometric distortion image .

[0065] The training image is obtained by further superimposing intensity perturbation, background shift, and simulated acquisition noise on the geometrically distorted image:

[0066] In the formula, The first term represents the result after geometric distortion, intensity variation, and noise degradation. Training images; Indicates the light intensity perturbation coefficient; Indicates the additional background offset; This represents the simulated acquisition noise, which includes at least one of Poisson noise and Gaussian noise. This means that the pixel value is limited to the range of 0 to 1.

[0067] Nine distorted images form the training input, labeled with the corresponding distortion-free super-resolution reference image. The homography parameters of each frame are sampled independently or partially correlated to cover situations where different frames have different geometric errors in actual acquisition. Compared to applying the same global enhancement transform to all nine frames, this method can specifically simulate non-uniform distortion between frames.

[0068] like Figure 2 As shown, three-level residual homography registration and effective region mask generation are then performed. One frame is selected from the nine images as the reference frame; in this embodiment, the fifth frame from the middle of the sequence is preferred, and the remaining eight frames are used as the frames to be registered. The reference frame and any frame to be registered are mapped to the same feature space through a feature extraction network with shared parameters, forming paired features:

[0069] In the formula, Indicates a reference frame; Indicates the first Frame of images to be registered; This represents a feature extraction network with shared parameters; Indicates splicing along the channel; This represents the paired features used in homography matrix regression.

[0070] A three-level residual homography registration method is used to perform geometric correction on the frames to be registered, from coarse to fine. The first level is used to estimate the global homography transformation over a larger range, while the second and third levels estimate the remaining geometric errors based on the registration results of the previous level. The cumulative homography matrix is ​​updated level by level according to the following formula:

[0071] And order:

[0072] In the formula, Indicates completion of the first The cumulative homography matrix after level residual correction; Indicates the first The residual homography matrix output by the homography matrix regressor; This represents a 3×3 identity matrix. The output range of each residual parameter can be limited according to the allowable variation range of the corresponding geometric parameters to reduce the instability caused by direct regression of large homography parameters.

[0073] After completing the third-level residual correction, the first The final homography matrix and the final registered image of the frame are represented as follows:

[0074] In the formula, Indicates the first The final homography matrix obtained by the frame after three-level residual correction; This represents a spatial transformation operation performed based on homography and bilinear interpolation. Indicates the first The final registered image of the frame. For the reference frame, no geometric transformation is performed; simply set... ,in Indicates the reference frame index.

[0075] Since homography may cause the backsampled coordinates of some output positions to fall outside the range of the original image, a binary effective region mask is generated based on the backmapped coordinates.

[0076] In the formula, Indicates the first Frame-registered image at pixel location The valid area is marked at the location; , This represents the corresponding reverse-mapped sampling coordinates; and These represent the width and height of the original image, respectively. The effective region is set to 1, the invalid region to 0, and the effective region mask corresponding to the reference frame is always set to 1. The effective region mask is used for both registration loss calculation and subsequent feature fusion, preventing black borders, mirror filling, or boundary stretching caused by perspective transformation from entering the reconstruction result. Compared to the conventional method of estimating a single geometric transformation at once, the three-level residual estimation corrects large-scale distortions and residual errors from coarse to fine, reducing the instability of directly regressing large homography parameters.

[0077] Then, nine-frame joint super-resolution reconstruction, amplitude-phase attention fusion, and artifact suppression are performed. The registered nine-frame images are combined along the channel dimension in a predetermined order, and initial reconstruction features are extracted through shallow convolutional units.

[0078]

[0079] In the formula, This represents the joint input consisting of nine registered images; This represents a shallow convolutional feature extraction unit; This represents the initial reconstructed features obtained.

[0080] For the intermediate features in the reconstructed network, perform a two-dimensional Fourier transform on each feature channel:

[0081] In the formula, Indicates the first Spatial domain characteristics of each channel; Represents a two-dimensional Fourier transform; Indicates the first The frequency domain characteristics corresponding to each channel.

[0082] Extract the spectral amplitude and spectral phase based on the frequency domain characteristics:

[0083] In the formula, Indicates the first The spectral amplitude of each channel; Indicates the first The spectral phase of each channel; Represents the modulus of a complex number; Represents complex phase. Spectral amplitude is used to characterize frequency energy and high-frequency details in different channels, while spectral phase is used to characterize edge, contour, and spatial structure location information.

[0084] Statistical and nonlinear mappings are performed on the spectral amplitude and spectral phase features respectively to generate amplitude attention weights and phase attention weights, and a joint attention weight is obtained:

[0085]

[0086]

[0087] In the formula, Represents the set of spectral amplitude features for multiple characteristic channels; Represents the set of spectral phase features for multiple characteristic channels; and These represent the statistical and nonlinear mappings of the magnitude branch and the phase branch, respectively. Indicates the magnitude attention weight; Indicates phase attention weights; Represents the Sigmoid function; This represents the amplitude-phase joint attention weight. The amplitude-phase joint attention weight is applied to the corresponding intermediate features to adaptively recalibrate different feature channels.

[0088] Simultaneously, the generated effective region mask is scaled to the current feature size, and pixel-level frame fusion weights are constructed by combining the frame reliability score predicted by the network:

[0089] In the formula, Indicates the first Frame at pixel position The fusion weight at the location; This indicates scaling to the current feature size. Frame effective area mask; The network prediction of the first The reliability score of the frame at the corresponding location; This represents an exponential function.

[0090] The features of the nine frames are weighted and fused using fusion weights:

[0091] In the formula, Indicates the first Features of frames after amplitude-phase joint attention recalibration; This represents the features after weighted fusion of nine frames. Therefore, the amplitude branch is used to suppress abnormal high-frequency energy and noise, the phase branch is used to enhance the consistency of spatial structure positions, the effective region mask is used to prevent invalid boundaries from obtaining higher fusion weights, and the frame reliability score is used to reduce the impact of residual misalignment or strong noise frames on the fusion result.

[0092] Therefore, this step can simultaneously suppress double edges, perspective boundary artifacts, and strong noise in a single frame caused by residual misalignment. This mechanism differs from direct stitching, simple averaging, or ordinary channel attention; its weights are jointly determined by frequency domain amplitude and phase information and the effective registration area.

[0093] Next, registration loss and reconstruction loss are constructed, and joint training is performed in stages. The registration network is constrained by multi-scale photometric consistency, normalized cross-correlation, gradient consistency, multi-frame template consistency, and homography matrix regularization. The total registration loss is expressed as:

[0094] In the formula, Indicates the total registration loss; This indicates a loss of photometric uniformity; This represents the normalized cross-correlation loss; This represents the gradient consistency loss; This represents the template consistency loss across multiple frames; This represents the regularization loss of the homography matrix; , , , and These represent the weights of the corresponding loss terms. The photometric consistency loss, normalized cross-correlation loss, and gradient consistency loss are calculated at multiple image scales and aggregated across eight frames to be registered, excluding the reference frame. The photometric consistency loss constrains the consistency of pixel grayscale; the normalized cross-correlation loss reduces the interference of different illumination phases and brightness variations on geometric registration; the gradient consistency loss prioritizes aligning structural edges; the multi-frame template consistency loss ensures a stable common structure in the multi-frame registration results; and the homography matrix regularization loss limits unreasonable excessive deformation of the prediction matrix.

[0095] The super-resolution reconstruction network employs pixel absolute error loss and structural similarity loss, and its reconstruction loss is expressed as:

[0096] In the formula, Indicates super-resolution reconstruction loss; This represents the mean absolute error between the model's output image and the reference image. This represents the super-resolution image output by the model. This represents a distortion-free, high-quality super-resolution reference image. Indicators representing structural similarity; and These represent the weights of the pixel absolute error loss and the structural similarity loss, respectively.

[0097] The training process consists of training the registration network separately, freezing the registration network to train the super-resolution network, and then performing end-to-end joint fine-tuning using a relatively small registration learning rate. The total loss of the joint fine-tuning is expressed as:

[0098] In the formula, This represents the total loss during the end-to-end joint fine-tuning phase; This indicates the weight of the registration loss in the joint training.

[0099] During the joint fine-tuning phase, the registration network adopts a smaller learning rate than the super-resolution reconstruction network to avoid significant shifts in the geometric correction capabilities already acquired by the registration network. This also allows the registration and super-resolution reconstruction modules to co-optimize while maintaining their respective functionalities. This joint constraint prevents the registration network from losing correct geometric correction capabilities simply to reduce the final reconstruction error, and enables the registration and super-resolution modules to co-optimize while maintaining their respective functionalities.

[0100] Finally, after processing through multiple amplitude and phase attention residual blocks, deep fused features are obtained through global residual connections. The registered and fused features are then processed through multiple amplitude and phase joint attention residual blocks, and deep fused features are obtained through global residual connections. For 2x super-resolution reconstruction, the deep fusion features are first mapped to the number of channels required for pixel rearrangement using channel-mapping convolution, and then spatial upsampling is completed through a PixelShuffle pixel rearrangement operation with a magnification of 2.

[0101] In the formula, This represents the deep fusion feature obtained by combining amplitude and phase joint attention residual blocks and global residual connections; This represents the channel-mapped convolution before upsampling; This indicates a PixelShuffle pixel rearrangement operation with a magnification of 2. This represents the features after spatial upsampling.

[0102] Finally, the super-resolution reconstructed image is obtained by output convolution:

[0103] In the formula, Indicates the output convolutional layer; This represents the final super-resolution reconstructed image.

[0104] In the actual inference process, the following steps are executed sequentially: sorting nine SIM images, background subtraction and normalization, reference frame determination, third-level residual homography registration of the remaining eight frames, effective region mask generation, amplitude-phase joint attention feature recalibration, multi-frame reliability weighted fusion, and pixel rearrangement upsampling, directly outputting a super-resolution image. The above processing does not require estimation of Zernike coefficients, nor does it rely on active optical correction devices such as spatial light modulators or deformable mirrors, thus forming an end-to-end super-resolution reconstruction pipeline for distorted SIM raw images.

[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A super-resolution reconstruction method for micro-images with significant structural distortion and illumination, characterized in that: The method includes the following steps: S1. Acquire multiple frames of low-resolution SIM original images of the same sample under multiple illumination directions and multiple illumination phases. Sort the multiple frames of low-resolution SIM original images according to illumination direction and illumination phase, and perform background subtraction and intensity normalization respectively to form network input data. S2. Based on the ideal SIM images corresponding to the original low-resolution SIM images of multiple frames, construct distorted SIM training samples. Apply geometric transformations to each frame of ideal SIM images and superimpose intensity perturbation, background offset and simulated acquisition noise to obtain multi-frame training images with non-uniform distortion between frames. Use the corresponding distortion-free super-resolution reference image as the training label. S3. Determine the reference frame from the multi-frame training images, and take the remaining frames as the frames to be registered. Extract the image features of the reference frame and each frame to be registered respectively, and use three-level residual homography registration to estimate the geometric deviation of each frame to be registered relative to the reference frame step by step. Based on the residual homography matrices at each level, obtain the final homography matrix corresponding to each frame to be registered, and perform spatial transformation on each frame to be registered according to the final homography matrix to obtain multi-frame registered images. S4. Based on the reverse mapping sampling coordinates of each frame to be registered when performing spatial transformation, determine the effective sampling area in each registered image, and generate an effective area mask corresponding to each registered image to distinguish the effective image area and the invalid boundary area after spatial transformation. S5. Perform feature extraction on the multi-frame registered images, convert the obtained intermediate features to the frequency domain and extract the spectral amplitude features and spectral phase features respectively, generate amplitude-phase joint attention weights based on the spectral amplitude features and spectral phase features, and determine the multi-frame fusion weights by combining the effective region mask and the reliability score of each frame. Perform weighted fusion of the features of each frame based on the multi-frame fusion weights to obtain the fused features. S6. The registration network and the super-resolution reconstruction network are trained based on the distorted SIM training samples. The registration network is constrained by photometric consistency loss, normalized cross-correlation loss, gradient consistency loss, multi-frame template consistency loss and homography matrix regularization loss. The super-resolution reconstruction network is constrained by pixel absolute error loss and structural similarity loss. The registration network and the super-resolution reconstruction network are jointly fine-tuned to obtain the trained super-resolution reconstruction model. S7. Input the original multi-frame distorted SIM images to be reconstructed into the trained super-resolution reconstruction model. After background subtraction and intensity normalization, three-level residual homography registration, effective region mask generation, amplitude-phase joint attention feature recalibration, and multi-frame feature weighted fusion, the obtained deep fusion features are spatially upsampled to output the corresponding super-resolution reconstructed image.

2. The super-resolution reconstruction method for micro-images with obvious distorted structure and illumination according to claim 1, characterized in that: In step S1, the first The lighting direction, the first The actual SIM observation image corresponding to each illumination phase is represented as follows: In the formula, Indicates the first The lighting direction, the first Actual SIM observation images corresponding to each illumination phase; This represents an ideal SIM image formed under corresponding lighting conditions; This indicates the geometric distortion mapping applied to the image frame. The geometric distortion mapping includes one or more of translation, rotation, scaling, shearing, or perspective distortion. This represents the background component in the observed image; This represents the noise component in the observed image; All image frames are numbered using a combination of illumination direction index and illumination phase index, with each illumination direction... Phase with illumination The corresponding index pair number is Establish a multi-frame SIM raw image set The original SIM images from multiple frames are stacked sequentially along the frame channel dimension to obtain the network input tensor. ; Then, background subtraction is performed on each frame of the original image. The background subtraction result is represented as follows: In the formula, Indicates the first Frame of original SIM image at pixel location The grayscale value at that location; This represents the grayscale value after background subtraction. Indicates the first Background estimation of the frame image; , These represent the x and y coordinates of a pixel, respectively. Then, normalization is performed based on the grayscale range of the current frame: In the formula, Represents the normalized i-th Frame image; and They represent the first The minimum and maximum gray values ​​of the image within the effective area after background subtraction; This indicates a positive number set to prevent the denominator from being zero.

3. The super-resolution reconstruction method for micro-images with obvious distorted structure and illumination according to claim 2, characterized in that: In step S2, a homography transformation is constructed for each frame of the ideal SIM image, then the... The homography matrix corresponding to the frame is represented as follows: In the formula, Indicates the action on the first The homography transformation matrix of a frame SIM image; This represents the translation matrix, which consists of the horizontal translation amount. and longitudinal translation Sure; This represents a rotation matrix, which is composed of rotation angles. Sure; This represents the scaling transformation matrix, which consists of the horizontal scaling factor. and longitudinal scale factor Sure; The shear transformation matrix is ​​represented by the shear parameters. and Sure; This represents the perspective transformation matrix, used to control the degree of projection distortion; Image resampling is performed based on the homography matrix, and inverse mapping is applied to the output pixel coordinates; let the homogeneous coordinates of the output pixels be... The corresponding original image sampling position is: In the formula, This represents the homogeneous coordinates obtained after inverse mapping using the homography matrix; , This represents the actual sampling coordinates corresponding to the original image; based on the actual sampling coordinates Perform bilinear interpolation to obtain the first... Frame geometric distortion image ; The training image is obtained by further superimposing intensity perturbation, background shift, and simulated acquisition noise on the geometrically distorted image: In the formula, The first term represents the result after geometric distortion, intensity variation, and noise degradation. Training images; Indicates the light intensity perturbation coefficient; Indicates the additional background offset; This represents the simulated acquisition noise; This means that the pixel value is limited to the range of 0 to 1.

4. The super-resolution reconstruction method for micro-images with obvious distorted structure and illumination according to claim 3, characterized in that: In step S3, one frame is selected from the processed training images as a reference frame, and the remaining images are used as frames to be registered. The reference frame and any frame to be registered are mapped to the same feature space through a feature extraction network with shared parameters, forming paired features. In the formula, Indicates a reference frame; Indicates the first Frame of images to be registered; This represents a feature extraction network with shared parameters; Indicates splicing along the channel; Represents paired features used in homography matrix regression; A three-level residual homography registration is used to perform geometric correction on the frames to be registered, from coarse to fine. The first level is used to estimate the global homography transformation over a larger range, while the second and third levels estimate the remaining geometric errors based on the registration results of the previous level. The cumulative homography matrix is ​​updated level by level according to the following formula: In the formula, Indicates completion of the first The cumulative homography matrix after level residual correction; Indicates the first The residual homography matrix output by the homography matrix regressor; After completing the third-level residual correction, the first The final homography matrix and the final registered image of the frame are represented as follows: In the formula, Indicates the first The final homography matrix obtained by the frame after three-level residual correction; This represents a spatial transformation operation performed based on homography and bilinear interpolation. Indicates the first The final registered image of the frame.

5. The super-resolution reconstruction method for micro-images with obvious distorted structure and illumination according to claim 4, characterized in that: In step S4, the binary effective region mask generated based on the reverse mapping coordinates is as follows: In the formula, Indicates the first Frame-registered image at pixel location The valid area is marked at the location; , This represents the corresponding reverse-mapped sampling coordinates; and These represent the width and height of the original image, respectively. The effective region is set to 1, the invalid region is set to 0, and the effective region mask corresponding to the reference frame is always set to 1.

6. The super-resolution reconstruction method for micro-images with obvious distorted structure and illumination according to claim 5, characterized in that: In step S5, the nine registered images are combined along the channel dimension in a predetermined order, and the initial reconstructed features are extracted through shallow convolutional units: In the formula, This represents the joint input consisting of nine registered images. This represents the total number of image frames. This represents a shallow convolutional feature extraction unit; This represents the initial reconstructed features obtained; For the intermediate features in the reconstructed network, perform a two-dimensional Fourier transform on each feature channel: In the formula, Indicates the first Spatial domain characteristics of each channel; Represents a two-dimensional Fourier transform; Indicates the first Frequency domain characteristics corresponding to each channel; Extract the spectral amplitude and spectral phase based on the frequency domain characteristics: In the formula, Indicates the first The spectral amplitude of each channel; Indicates the first The spectral phase of each channel; Represents the modulus of a complex number; Represents complex phase; Statistical and nonlinear mappings are performed on the spectral amplitude and spectral phase features respectively to generate amplitude attention weights and phase attention weights, and a joint attention weight is obtained: In the formula, Represents the set of spectral amplitude features for multiple characteristic channels; Represents the set of spectral phase features for multiple characteristic channels; and These represent the statistical and nonlinear mappings of the magnitude branch and the phase branch, respectively. Indicates the magnitude attention weight; Indicates phase attention weights; Represents the Sigmoid function; Indicates the amplitude-phase joint attention weight; The generated effective region mask is scaled to the current feature size, and pixel-level frame fusion weights are constructed by combining the frame reliability score predicted by the network: In the formula, Indicates the first Frame at pixel position The fusion weight at the location; This indicates scaling to the current feature size. Frame effective area mask; The network prediction of the first The reliability score of the frame at the corresponding location; Represents an exponential function; Finally, a fusion weight is used to perform weighted fusion of the features from the nine frames: In the formula, Indicates the first Features of frames after amplitude-phase joint attention recalibration; This represents the features after weighted fusion of nine frames.

7. The super-resolution reconstruction method for micro-images with obvious distorted structure and illumination according to claim 6, characterized in that: In step S6, the total loss function of the registration network is expressed as: In the formula, Indicates the total registration loss; This indicates a loss of photometric uniformity; This represents the normalized cross-correlation loss; This represents the gradient consistency loss; This represents the template consistency loss across multiple frames; This represents the regularization loss of the homography matrix; , , , and These represent the weights of the corresponding loss terms; The total reconstruction loss function of the super-resolution reconstruction network is expressed as: In the formula, Indicates super-resolution reconstruction loss; This represents the mean absolute error between the model's output image and the reference image. This represents the super-resolution image output by the model. This represents a distortion-free, high-quality super-resolution reference image. Indicators representing structural similarity; and These represent the weights of the pixel absolute error loss and the structural similarity loss, respectively.

8. The super-resolution reconstruction method for micro-images with obvious distorted structure and illumination according to claim 7, characterized in that: The training process in step S6 includes training the registration network separately, freezing the registration network to train the super-resolution network, and performing end-to-end joint fine-tuning. The total loss of the joint fine-tuning is expressed as: In the formula, This represents the total loss during the end-to-end joint fine-tuning phase; This indicates the weight of the registration loss in the joint training.

9. The super-resolution reconstruction method for micro-images with obvious distorted structure and illumination according to claim 8, characterized in that: In step S7, for the obtained deep fusion features Channel mapping convolution is used to map the deep fusion features to the number of channels required for pixel rearrangement, and then spatial upsampling is completed through pixel rearrangement operation: In the formula, This represents the channel-mapped convolution before upsampling; This indicates a pixel rearrangement operation with a magnification factor of 2. This represents the features after spatial upsampling; Finally, the super-resolution reconstructed image is obtained by output convolution: In the formula, Indicates the output convolutional layer; This represents the final super-resolution reconstructed image.