Frequency-spatial domain alignment based super-resolution processing for microscopic image sequence

By using a frequency-spatial aligned neural network architecture to process low-resolution fluorescence microscopy image sequences, the resolution limitations and SISR noise amplification problems of traditional microscopy techniques are solved, enabling high-resolution and temporally continuous microscopy image reconstruction, which is suitable for live cell imaging.

WO2025232673A1PCT designated stage Publication Date: 2025-11-13INSTITUTE OF BIOPHYSICS CHINESE ACADEMY OF SCIENCES

Patent Information

Application Number
PCT/CN2025/092292
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-04
Filing Date
2025-04-30
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Traditional microscopy techniques are limited by the diffraction limit, making it difficult to achieve high-resolution imaging. Furthermore, SISR-based computational methods suffer from noise amplification and temporal continuity issues in live cell imaging, affecting image quality and data analysis.

Method used

A frequency-spatial aligned neural network architecture is adopted to process low-resolution fluorescence microscopy image sequences through feature extraction, frequency-spatial alignment and reconstruction modules, and deformable convolution is used for feature matching and optical flow correction to generate high-resolution image sequences.

Benefits of technology

It improves the spatial resolution and temporal continuity of microscopic images, reduces phototoxicity in live imaging, and enhances imaging speed and image quality, making it suitable for long-term video data processing of live cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025092292_13112025_PF_FP_ABST
    Figure CN2025092292_13112025_PF_FP_ABST
Patent Text Reader

Abstract

A frequency-spatial domain alignment based method for super-resolution processing for a microscopic image time series. The method comprises: providing a low-resolution fluorescence microscopic image sequence, which consists of P low-resolution fluorescence microscopic images, for a biological sample, wherein P is a positive integer greater than 3; and feeding the low-resolution fluorescence microscopic image sequence into a super-resolution processing system consisting of a feature extraction module (220), a feature propagation and alignment module (230) and a reconstruction module (240).
Need to check novelty before this filing date? Find Prior Art

Description

Super-resolution processing of microscopic image sequences based on frequency-spatial alignment Technical Field

[0001] This application relates to super-resolution reconstruction of microscopic images using deep learning, particularly to neural network systems for super-resolution processing of video (time series of microscopic images) based on frequency-spatial alignment techniques. Background Technology

[0002] Limited by optical diffraction, traditional fluorescence microscopy can only achieve a maximum resolution of around 200 nanometers, insufficient for resolving intricate organelle structures such as mitochondrial cristae and cellular microfilaments, thus severely hindering the observation of related life science phenomena. Super-resolution microscopy has broken through the long-standing diffraction-limited resolution, fundamentally changing the ability of microscopes to visualize fine biological structures. For example, structured light illumination super-resolution microscopy (SIM) uses multiple modulated (e.g., sinusoidal moiré fringes) excitation lights to illuminate the sample, and after reconstruction using specific algorithms, it can achieve a 2-fold increase in resolution, enabling the resolution of even finer structures. Structured light microscopy systems have wide applications in live-cell imaging due to their fast imaging speed, low phototoxicity, and wide sample applicability. However, it is worth noting that any super-resolution technique using hardware (such as the aforementioned super-resolution microscope) to improve spatial resolution requires long acquisition times and / or high illumination intensity. Although SIM is becoming increasingly popular in live-cell imaging due to its low phototoxicity and high temporal resolution, super-resolution images from structured light microscopy need to be obtained from a series of original images through reconstruction algorithms. Because the reconstruction algorithm involves multiple complex steps such as Fourier transform, inverse Fourier transform, frequency domain shift, frequency domain filtering, and frequency domain stitching, noise in the original image is amplified. When the signal-to-noise ratio (SNR) is low, severe reconstruction artifacts occur, significantly degrading the quality of the reconstructed image. Therefore, compared to traditional microscopy techniques, structured light microscopy has higher SNR requirements, which limits its application scope. In particular, in live cell imaging, due to low fluorescent labeling efficiency, cell sensitivity to phototoxicity and photobleaching, and the high speed of biological movement, the effective number of photons collected in the original image is very limited. Under these conditions, the artifact phenomenon in structured light microscopy images is very severe, making it impossible to correctly distinguish them from sample information, hindering the observation of biological phenomena and subsequent data analysis.

[0003] On the other hand, the rapid development of artificial intelligence has broken through many limitations of traditional hardware. Various deep learning networks have performed exceptionally well in single-image super-resolution (SISR) processing tasks. This task typically involves converting a single low-resolution (LR) image into a high-resolution (HR) image. Recently, neural networks popular in the field of computer vision have been used to improve the resolution of microscopic images, achieving good results. Examples include Deep Fourier Channel Attention Network (DFCAN) and Deep Fourier Generative Adversarial Network (DFGAN). After proper training, they can process a single noisy low-resolution image into a noise-free super-resolution image, thus greatly reducing the acquisition requirements of the original image for super-resolution imaging, thereby improving the imaging speed and acquisition time of live super-resolution imaging.

[0004] However, SISR-based computational methods also face several limitations. First, the inherent quantum properties of photons introduce randomness into optical measurements, meaning that photon noise is unavoidable in all microscopic images, leading to errors in the neural network output. Second, since the noise in each original wide-field image is independently sampled, even if these images are continuous imaging of the same cell structure or biological sample, differences in sampling noise will cause the neural network to give different predictions, making it impossible to guarantee the temporal continuity of the output image.

[0005] These factors reduce the fidelity and reliability of the SISR model, hindering its widespread application in super-resolution imaging, especially in video imaging data processing. Summary of the Invention

[0006] To address the aforementioned issues, this application proposes a neural network architecture for super-resolution reconstruction of microscopic video data. This neural network architecture employs an algorithm based on "frequency-spatial alignment using deformable convolution" to achieve higher spatial resolution, more accurate structure, and better temporal continuity in super-resolution reconstruction results when processing wide-field fluorescence microscopy image sequences after prior training.

[0007] According to one aspect of this application, a method for super-resolution processing of a sequence of microscopic images based on frequency-spatial alignment is provided, comprising:

[0008] Provides a sequence of P low-resolution fluorescence microscopy images (such as wide-field fluorescence microscopy images or images captured by a low numerical aperture microscopy system) for a biological sample, where P is a positive integer greater than or equal to 3; and

[0009] The low-resolution fluorescence microscopy image sequence is processed through a super-resolution processing system consisting of a feature extraction module, a feature propagation and alignment module, and a reconstruction module. In the feature extraction module, multiple sequentially connected basic neural network sub-modules are used to extract features from the low-resolution fluorescence microscopy image sequence to generate a feature map sequence.

[0010] In the feature propagation and alignment module, before feature propagation, the feature map sequence is aligned in the frequency domain to the spatial domain. Frequency domain-spatial alignment includes frequency domain alignment and spatial domain alignment. During frequency domain alignment...

[0011] Fourier transform is performed on the i-th feature map group corresponding to the i-th frame low-resolution fluorescence microscopy image in the feature map group sequence to obtain the phase spectrum of the i-th feature map group, where i is a positive integer and less than P. Then, while ensuring that ik is greater than 0 and not greater than P, Fourier transform is performed on the ik-th feature map group corresponding to the ik-th frame low-resolution fluorescence microscopy image in the feature map group sequence to obtain the phase spectrum and amplitude spectrum of the ik-th feature map group, where k = 1, 2, -1 or -2.

[0012] The phase spectrum of the i-th feature map group is superimposed with the phase spectrum of the ik-th feature map group hk through multiple sequentially connected basic neural network submodules, then added to the phase spectrum of the ik-th feature map group, and finally subjected to an inverse Fourier transform together with the amplitude spectrum of the ik-th feature map group to obtain the ik-th frequency domain enhanced feature map group.

[0013] During spatial alignment, the ik-th frequency-domain enhanced feature map set, the ith-th feature map set, and the optical flow calculated from the ith-th and ik-th low-resolution fluorescence microscopy images are used as inputs. The output of the feature propagation and alignment process is then processed by a deformable convolutional alignment neural network architecture with optical flow correction.

[0014] In the reconstruction module, the feature map sequence processed by the feature propagation and alignment module is reconstructed into a super-resolution image sequence.

[0015] Optionally, the low-resolution fluorescence microscopy image sequence is upsampled and then added to the super-resolution image sequence.

[0016] Optionally, when training the neural network of the super-resolution processing system, a sequence of low-resolution fluorescence microscopy images composed of multiple low-resolution fluorescence microscopy images obtained for biological samples is used as the training input, and a sequence of high-resolution images obtained for the same biological sample is used as the training ground value.

[0017] Optionally, when training the neural network of the super-resolution processing system, a structured light image sequence obtained by using structured light imaging technology to obtain biological samples and then super-resolution processing to obtain a structured light super-resolution image sequence is used as the training ground value, and an image sequence obtained by selectively combining the structured light image sequences is used as the training input, wherein the selective combination depends on the phase difference between the structured light stripes.

[0018] Optionally, in the feature propagation and alignment module, propagation is performed in the form of primary propagation, and / or secondary propagation, and / or higher-order propagation.

[0019] Optionally, the feature propagation is implemented using multiple sequentially connected basic neural network sub-modules; and / or, the reconstruction module includes multiple sequentially connected basic neural network sub-modules and a pixel rearrangement sub-module for super-resolution processing.

[0020] Optionally, the underlying neural network submodule for feature extraction, the underlying neural network submodule for frequency domain alignment, the underlying neural network submodule for feature propagation, and / or the underlying neural network submodule for reconstruction may be the same or different.

[0021] Optionally, the basic neural network submodule can select a residual neural network, a U-Net, a residual channel attention neural network (RCAN), a residual dense neural network (RDN), or a self-attention neural network (Transformer) of any depth.

[0022] Optionally, in each training cycle, pixel blocks at the same position in each image sequence are randomly selected, randomly rotated and flipped, and then used as the input image and target image of the neural network, respectively. The error between the network output and the target image is calculated, and its gradient is backpropagated to update the network parameters. After the network input error converges, training is stopped and the network parameters are stored.

[0023] Optionally, the low-resolution fluorescence microscopy image is a wide-field fluorescence microscopy image or an image captured by a low numerical aperture microscopy system.

[0024] Optionally, the high-resolution image sequence is a structured light ultra-high-resolution image sequence obtained by super-resolution processing of a structured light image sequence obtained using structured light imaging technology for the same biological sample.

[0025] Optionally, the high-resolution image sequence is a sequence of fluorescence microscopic images taken using a high numerical aperture microscope for the same biological sample.

[0026] Optionally, both the Fourier transform and the inverse Fourier transform are performed in the real number domain.

[0027] According to another aspect of this application, a computer program product is also provided, comprising a computer program or instructions, characterized in that the computer program or instructions, when executed by a processor, are the aforementioned method or modification.

[0028] Using the method of this application, a neural network of a super-resolution processing system can be trained by taking a series of noisy, low-resolution fluorescence microscopy image sequences (such as wide-field fluorescence microscopy image sequences or fluorescence microscopy image sequences taken by low numerical aperture microscopes) as input and the corresponding high-resolution microscopy image sequences (such as structured illumination super-resolution microscopy image sequences or fluorescence microscopy image sequences taken by high numerical aperture microscopes) as ground truth. The trained neural network can directly generate clear, high-resolution image sequences from low-resolution image sequences with extremely low signal-to-noise ratios. In particular, for long-term video data of live cells, the method of this application can perform super-resolution processing on long-term wide-field illumination imaging results based on the trained neural network, effectively improving the fidelity and temporal continuity of super-resolution reconstruction compared to single-image super-resolution algorithms, i.e., SISR (Single Image Super-resolution). Using the method of this application, extremely low excitation light power and extremely short exposure time can be used during the imaging process compared to conventional imaging. Subsequent image restoration and super-resolution processing via neural networks significantly reduce phototoxicity in live-cell super-resolution imaging and enable the capture of high-speed dynamic processes, thus improving its applicability to live-cell imaging. In summary, this invention is of great significance for improving the quality of microscopic images and expanding the applications of super-resolution microscopy. Attached Figure Description

[0029] A more comprehensive understanding of the principles and aspects of this application will be gained from the detailed description below, in conjunction with the accompanying drawings. It should be noted that the scale of the drawings may vary for clarity, but this will not affect the understanding of this application. In the drawings:

[0030] Figure 1 schematically illustrates a basic block diagram of a microscopic imaging system according to an embodiment of this application;

[0031] Figure 2 schematically illustrates a super-resolution processing system according to an embodiment of this application, wherein the system processes a wide-field fluorescence image sequence;

[0032] Figure 3 schematically illustrates an example of a basic neural network according to an embodiment of this application, wherein the basic neural network is shown as a residual neural network; and

[0033] Figure 4 schematically illustrates an example of a frequency domain-spatial domain alignment submodule according to an embodiment of this application. Detailed Implementation

[0034] In the accompanying drawings of this application, features with the same structure or similar function are indicated by the same reference numerals.

[0035] Figure 1 schematically illustrates a basic block diagram of a microscopic imaging system, which generally includes an optical imaging system 100 and a control and data processing system 200. The optical imaging system 100 can be, but is not limited to, a wide-field microscopic imaging system, a scanning confocal microscopic imaging system, a light-panel illumination microscopic imaging system, a two-photon scanning imaging system, and a structured illumination microscopic imaging system. Taking a structured illumination microscopic imaging system as an example, the optical imaging system 100 includes an excitation optical path and a detection optical path. The excitation optical path includes an excitation objective and other optical components for generating excitation light. The excitation beam can pass through the excitation objective to excite fluorescence on a biological sample. The detection optical path includes a detection objective and other optical components for imaging, used to receive and detect the excited fluorescence. Those skilled in the art will understand that, depending on the configuration of the microscopic imaging system, the excitation objective and the detection objective can be the same objective or different objectives.

[0036] According to one embodiment of this application, the same optical imaging system 100 can switch between low-resolution fluorescence microscopy imaging modes (such as wide-field fluorescence microscopy imaging mode or low numerical aperture microscopy imaging mode) and high-resolution microscopy imaging modes (such as structured illumination super-resolution microscopy imaging mode or high numerical aperture microscopy imaging mode). The following description of low-resolution fluorescence microscopy imaging mode will only use wide-field microscopy imaging mode as an example, and the following description of high-resolution microscopy imaging mode will only use structured illumination super-resolution microscopy imaging mode as an example. However, those skilled in the art should understand that low-resolution fluorescence microscopy imaging modes are not limited to wide-field microscopy imaging mode, and high-resolution microscopy imaging modes are not limited to structured illumination super-resolution microscopy imaging mode. Subject to the purpose of this application, those skilled in the art can use any suitable imaging mode for corresponding substitutions. When the optical imaging system 100 is in wide-field microscopy imaging mode, two-dimensional fluorescence microscopy imaging can be performed on biological samples, especially live biological samples, to obtain a two-dimensional wide-field fluorescence image time series (hereinafter referred to as wide-field image time series) for the biological sample. When the optical imaging system 100 is in structured illumination super-resolution microscopy mode, it can perform two-dimensional fluorescence microscopy imaging on biological samples, especially live biological samples. The images are then processed by the control and data processing system 200 to obtain a two-dimensional super-resolution fluorescence image time series for the biological sample. Within the scope of this application, the optical imaging system 100 in structured illumination super-resolution microscopy mode can take any suitable form known in the art. For example, as an example only, the optical imaging system described in the publication *Visualizing Intracellular Organelle and Cytoskeletal Interactions at Nanoscale Resolution on Millisecond Timescales* by Guo, Y. et al., Cell 175, 1430-1442e1417 (2018), is described, and its specific construction will not be elaborated upon in this application.

[0037] The control and data processing system 200 mainly includes a computer and related components (such as a data storage device), which can control the operation of the optical imaging system 100 and receive image data from the optical imaging system 100 and perform corresponding post-processing. For example, when the optical imaging system 100 is in structured light illumination microscopic imaging mode, the acquired structured light fluorescence image time series (hereinafter referred to as structured light fluorescence image sequence) is provided to the control and data processing system 200, and after a series of structured light super-resolution data processing, it is reconstructed into a high signal-to-noise ratio, super-resolution microscopic image sequence.

[0038] For this purpose, the control and data processing system 200 may include a structured light super-resolution reconstruction module 210 capable of performing structured light super-resolution reconstruction on a structured light fluorescence image sequence using a standard structured light super-resolution reconstruction algorithm. Within the scope of this application, the standard structured light super-resolution reconstruction algorithm can be considered as an algorithm already known in the field of microscopic imaging. As an example, the standard structured light super-resolution reconstruction algorithm can be found in the published paper by Gustafsson, MG et al., *Three-dimensional resolution doubling in wide-field fluorescence microscopy by structured illumination*. *Biophys J* 94, 4957-4970 (2008).

[0039] It should be noted that, within the scope of this application, the modules and / or submodules described herein can be understood to include data storage, such as computer-readable storage media, in which programs or subroutines and neural network models can be stored and executed by a computer. These programs or subroutines and neural network models, when executed by a computer, can implement the methods / steps described below. Specific programming methods for the programs and / or subroutines are not discussed in this application; those skilled in the art can implement the relevant functions using any well-known programming software and / or commercial software. Therefore, the context of this application, when describing the operation of related systems or the operation or methods of modules, should be understood to mean that they can also be programmed for computer invocation and execution.

[0040] The control and data processing system 200 may further include a feature extraction module 220, a feature propagation and alignment module 230, and a reconstruction module 240. The feature extraction module 220, the feature propagation and alignment module 230, and the reconstruction module 240 constitute a video (microscopic image time series or microscopic image sequence) super-resolution processing (or super-resolution reconstruction) system according to an embodiment of this application. Figure 2 schematically illustrates an example of this video (microscopic image time series or microscopic image sequence) super-resolution processing (or super-resolution reconstruction) system processing a low-resolution fluorescence image time series (hereinafter referred to as a low-resolution image sequence). In the context of this application, the term video can be considered as consisting of images grouped according to the order of time they were captured, and can be referred to as an image time series, or simply an image sequence.

[0041] The feature extraction module 220 includes multiple sequentially connected basic neural network sub-modules. In the following explanation of the technical solution of this application, the basic neural network is illustrated using a residual (ResNet) neural network as an example. However, those skilled in the art should understand that the basic neural network used in the embodiments of this application can be any other suitable neural network, such as, but not limited to, U-Net of arbitrary depth, Residual Channel Attention Neural Network (RCAN), Residual Dense Neural Network (RDN), Transformer self-attention neural network, etc., as long as the basic neural network used can achieve the purpose of this application.

[0042] According to one embodiment of this application, the feature extraction module 220 includes N sequentially connected residual neural network sub-modules 221, where N is an integer greater than 1, for example, five sequentially connected residual neural network sub-modules 221. Each residual neural network sub-module may include, for example, two convolutional layers and an activation function, as shown in FIG3. Furthermore, the feature extraction module 220 may also include one or more convolutional layers 222 (shown as one in FIG2) to adjust the number of channels input to the sequentially connected residual neural network sub-modules 221, for example, adjusting from a single input channel to a commonly used number of channels for feature extraction in convolutional layers (e.g., 64 channels). According to an embodiment of this application, the feature extraction module 220 is configured to perform feature extraction processing on low-resolution fluorescence microscopy images or image sequences (e.g., wide-field fluorescence microscopy images or wide-field fluorescence microscopy image sequences or low numerical aperture microscopy images or image sequences) acquired by the optical imaging system 100 to output a set of feature maps or a sequence of feature maps. It should be noted that, in the context of this application, an image, low-resolution image, super-resolution image, or graph can be understood as a two-dimensional digital matrix that can be processed by a computer; an image sequence or graph group can be understood as a group of two-dimensional digital matrices; and a graph group sequence can be understood as a group of graphs. For example, for the feature extraction module 220, when a single low-resolution fluorescence microscopy image (hereinafter referred to as a low-resolution image) is input, feature extraction processing can output a feature graph group sequence; and when a sequence of low-resolution fluorescence microscopy images (hereinafter referred to as a low-resolution image sequence) is input, feature processing can output a feature graph group sequence. When considering an image sequence, based on the temporal propagation order, the first image can be called the previous frame image, and the second image can be called the next frame image.

[0043] It is important to understand that the feature extraction module 220, feature propagation and alignment module 230, and reconstruction module 240 of this application should be based on the fundamental principles of neural networks. For example, when referring to residual neural network submodules or convolutional layers in other modules below, it may refer to the same residual neural network submodule 221 or convolutional layer 222 in feature extraction module 220, but the usage of these residual neural network submodules 221 is reconfigured simply because the functions of the mentioned modules are different. In alternative embodiments, when referring to residual neural network submodules or convolutional layers in other modules below, it may also refer to residual neural network submodules or convolutional layers that are different from those in feature extraction module 220.

[0044] The feature propagation and alignment module 230 may also include M sequentially connected basic neural network sub-modules, such as M sequentially connected residual neural network sub-modules 231, where M is an integer greater than 1, for example, five sequentially connected residual neural network sub-modules 231. Furthermore, a convolutional layer 232 is provided before the M sequentially connected residual neural network sub-modules 231 to adjust the number of channels input to the sequentially connected residual neural network sub-modules 231. According to embodiments of this application, the residual neural network sub-modules 231 and / or convolutional layers 231 may be the same as or different from the residual neural network sub-modules 221 and / or convolutional layers 221.

[0045] The feature propagation and alignment module 230 is configured to take the feature map sequence output by the feature extraction module 220 as input. After processing by the frequency-spatial alignment submodule 233, the feature map corresponding to each frame of the low-resolution image (processed by the feature extraction module 220) is further processed by the convolutional layer 232 and the aforementioned M sequentially connected residual neural network submodules 231. The output constitutes a new feature map for the next frame. This new feature map is then processed by the frequency-spatial alignment submodule 233, the convolutional layer 232, and the aforementioned M sequentially connected residual neural network submodules 231, forming a new feature map. This process is recursively propagated to the last frame, thus updating all feature map sets corresponding to all frames. Next, this propagation process is repeated in the reverse direction to update the feature map sets corresponding to all frames again. After this, the forward and backward propagation processes are repeated one or more times.

[0046] The propagation process of the aforementioned feature propagation can be found in the published literature BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment. in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition 5972-5981 (2022). In Figure 2, the solid arrows in the feature propagation and alignment module 230 represent primary propagation, and the dashed arrows represent secondary propagation. Of course, those skilled in the art should understand that feature propagation can also be implemented in a higher-order manner.

[0047] Figure 4 schematically illustrates an example of a frequency-spatial alignment submodule 233 according to an embodiment of this application. As shown, the frequency-spatial alignment submodule 233 may include a frequency domain alignment portion and a spatial domain alignment portion, such that the input feature map set is first aligned in the frequency domain and then aligned in the spatial domain. For example, the frequency domain alignment portion of the frequency-spatial alignment submodule 233 may include M' sequentially connected basic neural network submodules, such as M' sequentially connected residual neural network submodules 23311, where M' is an integer greater than 1, for example, five sequentially connected residual neural network submodules 23311. Furthermore, a convolutional layer 23312 is provided before the M' sequentially connected residual neural network submodules 23311 to adjust the number of channels input to the sequentially connected residual neural network submodules 23311.

[0048] The operation of the frequency-spatial alignment submodule 233 according to this application is described below with reference to FIG4. For a low-resolution image frame (hereinafter referred to as the current frame) in the low-resolution image sequence acquired by the optical imaging system 100, the image frame is processed by the feature extraction module 220 to obtain the feature map group h. i Here, i represents the temporal order of the selected low-resolution images in the low-resolution image sequence. For example, when a low-resolution image sequence consists of P low-resolution images (P is a positive integer greater than or equal to 3), then i can be any integer less than P. Therefore, this low-resolution image can be called the i-th frame low-resolution image. Furthermore, the preceding or two preceding frames, or the following or two following frames of this image are also processed by the feature extraction module 220 to obtain the feature map set h. i-k , where k = -1, 1, -2, or 2. Therefore, the corresponding low-resolution image can be called the low-resolution image of the ikth frame (ensuring that ik > 0 and is less than or equal to P).

[0049] Then, in the frequency domain alignment part, Fourier transform is used to align the feature map group h. i and feature map group h i-k Processing is performed to generate feature map set h i Phase spectrum and feature map h i-k The phase spectrum and amplitude spectrum are described below:

[0050] Here, FFT stands for Fast Fourier Transform, and Angle and Abs represent the phase spectrum and amplitude spectrum after the Fast Fourier Transform, respectively. Those skilled in the art should understand that the Fast Fourier Transform (or Inverse Transform) mentioned in this paper is a fast implementation algorithm for the Fourier Transform (or Inverse Transform).

[0051] As shown in the figure, feature map group h i Phase spectrum and characteristic map group h i-k The phase spectra are then channel-stacked. Here, channel stacking can be implemented using any stacking method familiar to those skilled in the art, such as stacking two phase spectra sequentially or interleaving channels. The result of the channel stacking is then processed by the convolutional layer 23312 and the residual neural network submodule 23311 to obtain the phase residual, which is then compared with the feature map group h. i-k The following phase spectrum is obtained by adding the phase spectra together. For example, this process can be represented by the following formula:

[0052] Wherein, RB represents the residual neural network submodule 23311, f represents the convolutional layer 23312, and ⊕ represents channel stacking. In the formula shown, the superscript (2) of RB can mean that the data processed by the convolutional layer is then passed through two residual neural networks in sequence, but it can also be passed through M' residual neural networks as described above as needed.

[0053] Then, the phase spectrum obtained above With feature map group h i-k The amplitude spectrum is re-inverse Fourier transformed and used as the output of the frequency domain alignment part, as follows:

[0054] Here, iFFT represents the inverse fast Fourier transform, thus obtaining the feature map group after frequency domain enhancement processing. This is the output of the frequency domain alignment part. It is worth noting that the above-mentioned Fast Fourier Transform and Inverse Transform are both performed in the real number domain. Those skilled in the art should understand that the results produced by the two-dimensional Fourier Transform have a conjugate symmetry relationship. If the Fourier Transform and Inverse Transform are used directly, non-real number transformation results will be produced, resulting in highly uncertain results. Therefore, the technical solution of this application adopts the real number domain Fourier Transform, that is, considering the conjugate symmetry of the Fourier Transform result, only half of the information is saved, and the other half is restored through the form of conjugate symmetry, thereby ensuring that the result of the Inverse Fourier Transform is still a real number.

[0055] As shown in Figure 4, the feature map group after the above frequency domain enhancement processing With feature map group h i Together with the optical flow calculated from the i-th low-resolution image and the ik-th low-resolution image, it serves as the input to the spatial alignment portion of the frequency-spatial alignment submodule 233. In embodiments of this application, the spatial alignment portion can be configured as a deformable convolutional alignment neural network architecture for optical flow correction in a suitable manner known to those skilled in the art. For example, the optical flow prediction neural network portion architecture in the deformable convolutional alignment neural network for optical flow correction can be implemented with reference to the published document A Ranjan, MJ Black: Optical Flow Estimation Using a Spatial Pyramid Network. IEEE, 2017, and the deformable convolutional alignment neural network portion architecture can be implemented with reference to the published document BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment. in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition 5972-5981 (2022), which will not be elaborated upon herein.

[0056] During the spatial alignment process of the frequency-spatial alignment submodule 233, the feature map group h is processed... i and feature map group h i-kSpatial alignment is performed. Specifically, an optical flow prediction neural network (e.g., Spynet) is used to predict the optical flow between the i-th and ik-th low-resolution images. This neural network combines the classic spatial pyramid method with deep learning to calculate optical flow, avoiding large shifts. The entire video sequence (original image sequence) can be input into the optical flow prediction neural network for prediction, yielding pixel-level shifts along the horizontal and vertical axes. Although the accuracy of this optical flow prediction is not high, it can serve as the basis for subsequent deformable convolutions. The optical flow is then used to first perform spatial optical flow correction on the already frequency-domain aligned feature maps. After optical flow correction, the optically corrected feature maps are channel-overlayed with the feature maps of the current frame, and then processed by a convolutional neural network to generate the offset required for deformable convolution. The difference between deformable convolution and regular convolution is that its sampling position can be flexibly determined according to the offset size, giving it stronger spatial alignment capabilities. Using the offset generated by the convolutional neural network, deformable convolution is performed on the optically corrected feature maps to generate a frequency-domain and spatially aligned feature map set.

[0057] The reconstruction module 240 may include M” sequentially connected basic neural network sub-modules, such as M” sequentially connected residual neural network sub-modules 241, where M” is an integer greater than 1, for example, five sequentially connected residual neural network sub-modules 241. Furthermore, a convolutional layer 242 is provided before the M” sequentially connected residual neural network sub-modules 241 to adjust the number of channels input to the sequentially connected residual neural network sub-modules 241. According to embodiments of this application, the residual neural network sub-modules 241 and / or the convolutional layer 241 may be the same as or different from the residual neural network sub-modules 221 and / or the convolutional layer 221. In addition, the reconstruction module 240 also includes a pixel shuffle sub-module 243 for super-resolution processing. In the reconstruction module 240, a convolutional layer 244 is also configured after the pixel shuffle sub-module 243 to adjust the number of channels accordingly. According to an embodiment of this application, the input to the reconstruction module 240 is a sequence of feature maps corresponding to the low-resolution image sequence input to the feature extraction module 220, which has been processed by the feature propagation and alignment module 230. After processing by the reconstruction module 240, the feature map sequence outputs a super-resolution image sequence corresponding to the low-resolution image sequence.

[0058] As shown in Figure 4, according to a preferred embodiment of this application, the super-resolution image sequence output by the reconstruction module 240 is added to the image sequence after upsampling (the sampling factor corresponds to the pixel rearrangement processing) the low-resolution image sequence input to the feature extraction module 220, and the output is used as the final super-resolution image sequence.

[0059] The advantages of the above-mentioned technical solution of this application are as follows:

[0060] (1) This application proposes a frequency domain-spatial domain alignment algorithm based on deformable convolution, which can achieve better alignment and matching of microscopic image feature groups.

[0061] (2) This application proposes a video super-resolution reconstruction system that utilizes the above-mentioned deformable convolution frequency-spatial alignment algorithm, which can realize super-resolution processing of multiple continuous time series low-resolution, noisy original microscopic images.

[0062] (3) The method proposed in this application can be applied to the original microscopic data of different modalities, and can reconstruct and restore the data of different modalities.

[0063] (4) The image reconstructed by the method proposed in this application has good fidelity and temporal continuity that surpasses other existing technologies.

[0064] When training the neural network of the video super-resolution processing system of this application, the relevant neural network model is optimized using a loss function, which includes, but is not limited to, mean squared error (MSE), mean absolute error (MAE), structural similarity (SSIM), or a weighted sum thereof.

[0065] For each group of original structured light image sequences, they are normalized and included in the training dataset. The video super-resolution neural network is trained using the training dataset generated by the above method. For example, in each training cycle, a matching low-resolution image sequence and a high-resolution image sequence (e.g., the corresponding structured light super-resolution image sequence or high numerical aperture microscope image sequence) are randomly selected from the dataset, and pixels at the same location are randomly rotated and flipped to serve as the input image and target image of the neural network, respectively. The error between the network output and the target image is calculated, and its gradient is backpropagated to update the network parameters. Training is stopped and the network parameters are stored after the network input error converges. Those skilled in the art should understand that the training methods for neural networks are not limited to those listed.

[0066] According to embodiments of this application, during training, the low-resolution image sequence as input can be obtained by the optical imaging system 100 for the biological sample to be detected in a low-resolution imaging mode, while simultaneously, a high-resolution image sequence can be obtained by the optical imaging system 100 for the biological sample to be detected in a high-resolution microscopy imaging mode as training ground values. For example, as a non-limiting example, the optical imaging system 100 obtains a structured light image sequence for the biological sample to be detected in a structured light illumination imaging mode, and performs structured light super-resolution processing to obtain a super-resolution image sequence, which is used as training ground values. In an alternative embodiment, the structured light image sequence used for super-resolution processing can be selectively combined (based on the phase difference of the structured light fringes, for example, if the structured light fringes are 2π / 3, three adjacent structured light images are combined into a wide-field image) into a wide-field image sequence as training input. In an alternative embodiment, the high-resolution image sequence can be directly obtained by the optical imaging system 100 for the biological sample to be detected in a high numerical aperture microscopy imaging mode as training ground values.

[0067] After training, the trained neural network and its parameters can be used for super-resolution processing of long-term low-resolution image time series. The neural network can take multiple consecutive images as input at one time and output corresponding image sequences that are upsampled by 2 or 3 times. It can complete the super-resolution processing of multiple images within one second, improving the quality and temporal continuity of imaging. It can be widely used for video observation and recording of living cells.

[0068] The training and inference process described above is merely a non-limiting example of the neural network system for video super-resolution processing based on frequency-spatial alignment technology in this application. Those skilled in the art should understand that the "frequency-spatial alignment algorithm based on deformable convolution" in this application can also be used in other neural network architectures for video, such as denoising, registration, and segmentation of multiple images for microscopic image tasks.

[0069] After training the neural network of the video (microscopic image time series) super-resolution processing system, super-resolution processing (or prediction) is performed on the original fluorescence image sequence Y obtained by the optical imaging system. This involves processing the noisy fluorescence image y... i(i = 1, 2, ..., N, where N is an integer) Super-resolution reconstruction is performed using a video (microscopic image time series) super-resolution processing system to obtain a super-resolution image Z, which is the final super-resolution image. Alternatively, if processing a long sequence of raw fluorescence images using this system, the sequence can be segmented, with each segment input separately into the system. Finally, a sliding window process is used to recover the long-term data; this process is repeated for each raw fluorescence image sequence.

[0070] As an example only, the reference values ​​of the main parameters of the residual neural network used as an example in the technical solutions described in this application are shown in Table 1. It should be noted that the basic neural network module that can be used in this application does not specify the use of a particular neural network structure. Typical network structures such as U-Net, Residual Channel Attention Neural Network (RCAN), Residual Dense Neural Network (RDN), and Transformer can all achieve the described functions.

[0071] Table 1 Main parameters of the neural network

[0072] In addition, the main parameters used in the neural network training process of this application need to be adjusted according to the specific situation of the dataset (signal-to-noise ratio, structural complexity, etc.). A set of parameters for reference is shown in Table 2.

[0073] Table 2 Reference parameters for network training process

[0074] This application proposes a microscopic video reconstruction system based on a frequency-spatial alignment algorithm. This system can effectively process continuous time-series microscopic data from raw images with extremely low signal-to-noise ratios (SNR) and reconstruct high-resolution, noise-free images with high fidelity. Applying this method significantly reduces the SNR requirements of the raw image during microscopic imaging, thereby reducing light damage to biological samples during imaging, increasing imaging speed, significantly improving image quality in structured light microscopy, and expanding its application scope.

[0075] Although specific embodiments of this application are described in detail herein, they are provided for illustrative purposes only and should not be construed as limiting the scope of this application. Furthermore, those skilled in the art will understand that the various embodiments described herein can be used in combination with each other. Various substitutions, modifications, and alterations can be conceived without departing from the spirit and scope of this application.

Claims

1. A method for super-resolution processing of microscopic image sequences based on frequency-spatial alignment, comprising: Provides a sequence of P low-resolution fluorescence microscopy images for a biological sample, where P is a positive integer greater than or equal to 3; and The low-resolution fluorescence microscopy image sequence is processed through a super-resolution processing system consisting of a feature extraction module, a feature propagation and alignment module, and a reconstruction module. In the feature extraction module, multiple sequentially connected basic neural network sub-modules are used to extract features from the low-resolution fluorescence microscopy image sequence to generate a feature map sequence. In the feature propagation and alignment module, before feature propagation, the feature map sequence is aligned in the frequency domain to the spatial domain. Frequency domain-spatial alignment includes frequency domain alignment and spatial domain alignment. During frequency domain alignment... Perform a Fourier transform on the i-th feature map group corresponding to the i-th frame of the low-resolution fluorescence microscopy image to obtain the phase spectrum of the i-th feature map group. Where i is a positive integer and less than P, and then, ensuring that ik is greater than 0 and not greater than P, a Fourier transform is performed on the ikth feature map group corresponding to the ikth frame low-resolution fluorescence microscopy image in the feature map group sequence to obtain the phase spectrum of the ikth feature map group. and amplitude spectrum Where k = 1, 2, -1 or -2; The phase spectrum of the i-th feature map group is compared with that of the ik-th feature map group h. k After channel superposition of the phase spectrum, it is processed by multiple sequentially connected basic neural network submodules and compared with the phase spectrum of the i-th feature map group. Add and combine with the amplitude spectrum of the i-th feature map group Perform inverse Fourier transform together to obtain the ik-th frequency domain enhanced feature map set. During spatial alignment, the ik-th frequency-domain enhanced feature map group, the ith-th feature map group, and the optical flow calculated from the ith-th and ik-th low-resolution fluorescence microscopy images are used as inputs. The output of the feature propagation and alignment process is then processed by a deformable convolutional alignment neural network architecture with optical flow correction. In the reconstruction module, the feature map sequence processed by the feature propagation and alignment module is reconstructed into a super-resolution image sequence.

2. The method according to claim 1, characterized in that, The low-resolution fluorescence microscopy image sequence was upsampled and then added to the super-resolution image sequence.

3. The method according to claim 1 or 2, characterized in that, When training the neural network of the super-resolution processing system, a sequence of low-resolution fluorescence microscopy images composed of multiple low-resolution fluorescence microscopy images obtained for biological samples is used as the training input, and a sequence of high-resolution images obtained for the same biological sample is used as the training ground value.

4. The method according to claim 1 or 2, characterized in that, When training the neural network of the super-resolution processing system, the structured light image sequence obtained by structured light imaging technology for biological samples and the structured light super-resolution image sequence obtained by super-resolution processing are used as the training ground values. The image sequence obtained by selectively combining the structured light image sequences is used as the training input, where the selective combination depends on the phase difference between the structured light stripes.

5. The method according to any one of claims 1 to 4, characterized in that, In the feature propagation and alignment module, propagation is performed in the form of first-order propagation, and / or second-order propagation, and / or higher-order propagation.

6. The method according to claim 5, characterized in that, The feature propagation is implemented using multiple sequentially connected basic neural network sub-modules; and / or, the reconstruction module includes multiple sequentially connected basic neural network sub-modules and a pixel rearrangement sub-module for super-resolution processing.

7. The method according to claim 6, characterized in that, The underlying neural network submodules for feature extraction, frequency domain alignment, feature propagation, and / or reconstruction may be the same or different.

8. The method according to claim 7, characterized in that, The basic neural network submodule can select residual neural networks, U-Net, residual channel attention neural networks (RCAN), residual dense neural networks (RDN), or self-attention neural networks (Transformer) of arbitrary depth.

9. The method according to claim 8, characterized in that, In each training cycle, pixel blocks at the same position in each image sequence are randomly selected, and after random rotation and flipping, they are used as the input image and target image of the neural network, respectively. The error between the network output and the target image is calculated, and its gradient is backpropagated to update the network parameters. After the network input error converges, training is stopped and the network parameters are stored.

10. The method according to claim 1 or 2, characterized in that, The low-resolution fluorescence microscopy image is a wide-field fluorescence microscopy image or an image captured by a low numerical aperture microscopy system.

11. The method according to claim 3, characterized in that, The high-resolution image sequence is a structured light ultra-high-resolution image sequence obtained by super-resolution processing of a structured light image sequence obtained using structured light imaging technology on the same biological sample.

12. The method according to claim 3, characterized in that, The high-resolution image sequence is a sequence of fluorescence microscopic images taken using a high numerical aperture microscope for the same biological sample.

13. The method according to claim 1 or 2, characterized in that, Both the Fourier transform and the inverse Fourier transform are performed in the real number domain.

14. A computer program product comprising a computer program or instructions, characterized in that, The computer program or instructions, when executed by a processor, implement the method of claims 1 to 13.

Citation Information

Patent Citations

  • MRI (Magnetic Resonance Imaging) accelerated reconstruction system guided by multi-scale space-frequency domain feature information

    CN116563409A

  • Face counterfeit video detection method based on Fourier domain adaptation

    CN116563957A

  • Self-supervised microscopic image super-resolution processing method and system

    CN116721017A

  • Microscopic image sequence super-resolution processing based on frequency domain-spatial domain alignment

    CN119205502A

  • Method for aligning multiple image frames, apparatus for aligning multiple image frames, and storage medium

    WO2023245383A1

Cited By

  • Mine image enhancement method and device and electronic equipment

    CN121599863A

  • Bridge crack three-dimensional reconstruction method and system based on depth feature fusion and storage medium

    CN122336154A