Super-resolution processing of microscopic image sequences based on frequency-spatial alignment

By adopting a frequency domain-space-aligned neural network architecture in microscopic image super-resolution processing, the problems of high signal-to-noise ratio requirements and serious reconstruction artifacts are solved, and microscopic image sequence processing with high resolution and time continuity is achieved, reducing phototoxicity and expanding the application range.

CN119205502BActive Publication Date: 2025-05-13INSTITUTE OF BIOPHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410544992.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-04
Publication Date
2025-05-13
Estimated Expiration
2044-05-04

AI Technical Summary

Technical Problem

The prior art has problems such as high signal-to-noise ratio requirements, serious reconstruction artifacts, and inability to ensure time continuity in microscopic image super-resolution processing, especially in vivo cell imaging.

Method used

A neural network architecture based on frequency domain-space alignment is proposed, and the microscopic image sequence is super-resolution processed through feature extraction, frequency domain-space alignment and reconstruction modules to improve spatial resolution and temporal continuity.

Benefits of technology

It realizes the generation of high-resolution, noise-free microscopic image sequences under extremely low signal-to-noise ratio conditions, reduces phototoxicity, improves imaging speed and time continuity, and expands the application range of super-resolution microscopy technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119205502B_ABST
    Figure CN119205502B_ABST
Patent Text Reader

Abstract

This application discloses a method for super-resolution processing of time series of microscopic images based on frequency-spatial alignment, comprising: providing a low-resolution fluorescence microscopic image sequence consisting of P low-resolution fluorescence microscopic images for a biological sample, wherein P is a positive integer greater than 3; and subjecting the low-resolution fluorescence microscopic image sequence to a super-resolution processing system consisting of a feature extraction module, a feature propagation and alignment module, and a reconstruction module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to super-resolution reconstruction of microscopic images using deep learning, and in particular to a neural network system for super-resolution processing of videos (microscopic image time series) based on frequency-spatial domain alignment technology. Background Art

[0002] Limited by optical diffraction, the resolution of traditional fluorescence microscopy can only reach about 200 nanometers at most, which is not enough to analyze fine organelle structures such as mitochondrial ridges and cell microfilaments, thus seriously hindering the observation of related life science phenomena. Super-resolution microscopy has broken through the long-standing limitation of diffraction-limited resolution and completely changed the visualization ability of microscopy for fine biological structures. For example, structured light illumination super-resolution microscopy (SIM) uses multiple modulated (such as sinusoidal moiré) excitation lights to illuminate the sample, and then reconstructs it through a specific algorithm to obtain a 2-fold increase in resolution, enabling the analysis of more delicate structures. Structured light microscopy systems have very wide applications in the field of living cell imaging due to their fast imaging speed, low phototoxicity, and wide sample applicability. However, it is worth noting that any super-resolution technology that improves spatial resolution through hardware (such as the above-mentioned super-resolution microscope) requires a long acquisition time and / or high illumination intensity. Although SIM is becoming more and more popular in living cell imaging with its low phototoxicity and high temporal resolution, super-resolution images of structured light microscopy need to be obtained from a series of original images through reconstruction algorithms. Since the reconstruction algorithm involves multiple complex steps such as Fourier transform, inverse Fourier transform, frequency domain shift, frequency domain filtering, and frequency domain splicing, the noise in the original image will be amplified. When the signal-to-noise ratio is low, serious reconstruction artifacts will occur, which will seriously reduce the quality of the reconstructed image. Therefore, compared with traditional microscopy technology, structured light microscopy technology has higher requirements for signal-to-noise ratio, which limits the scope of application of structured light microscopy technology. In particular, in live cell imaging, the number of effective collected photons of the original image is very limited due to reasons such as low fluorescence labeling efficiency, cell sensitivity to phototoxicity and photobleaching, and fast movement speed of organisms. Under these conditions, the artifact phenomenon of structured light microscopy images is very serious, making it impossible to correctly distinguish them from sample information, hindering the observation of biological phenomena and subsequent data analysis.

[0003] On the other hand, the rapid development of artificial intelligence has broken through the limitations of many traditional hardware. Various deep learning networks have performed well in single image super-resolution (SISR) processing tasks. This processing task usually converts a single low-resolution (LR) image into a high-resolution (HR) image. Recently, popular neural networks in the field of computer vision have been used to improve the resolution of microscopic images and have achieved good results, such as deep Fourier channel attention network (DFCAN) and deep Fourier generative adversarial network (DFGAN). After appropriate training, a single noisy low-resolution image can be processed into a noise-free super-resolution image, which can greatly reduce the acquisition requirements of the original image for super-resolution imaging, thereby improving the imaging speed and shooting time of in vivo super-resolution imaging.

[0004] However, SISR-based computational methods also face many limitations. First, the inherent quantum properties of photons introduce randomness in optical measurements, so photon noise is inevitably present in all microscopic images, which brings errors to the output results of the neural network; second, since the noise of each original wide-field image is independently sampled, even if these images are continuous imaging of the same cell structure or biological sample, the difference in sampling noise will cause the neural network to give different prediction results, and the temporal continuity of the output image cannot be guaranteed.

[0005] The above factors reduce the fidelity and credibility of the SISR model, hindering its widespread application in super-resolution imaging, especially video imaging data processing. Summary of the invention

[0006] In response to the above problems, the present application aims to propose a neural network architecture for super-resolution reconstruction of microscopic video data. The neural network architecture adopts the "frequency domain-spatial domain alignment based on deformable convolution" algorithm, so that the super-resolution reconstruction neural network architecture after preliminary training can have higher spatial resolution, more accurate structure and better temporal continuity when processing wide-field fluorescence microscopy image sequences. Super-resolution reconstruction results.

[0007] According to one aspect of the present application, a method for super-resolution processing of a microscopic image sequence based on frequency domain-spatial domain alignment is provided, comprising:

[0008] Providing a low-resolution fluorescence microscopic image sequence consisting of P low-resolution fluorescence microscopic images (such as wide-field fluorescence microscopic images or images taken by a low numerical aperture microscopic imaging system) for a biological sample, wherein P is a positive integer greater than or equal to 3; and

[0009] The low-resolution fluorescence microscopic image sequence is passed through a super-resolution processing system consisting of a feature extraction module, a feature propagation and alignment module, and a reconstruction module, wherein in the feature extraction module, a plurality of sequentially connected basic neural network submodules are used to extract features from the low-resolution fluorescence microscopic image sequence to generate a feature map group sequence,

[0010] In the feature propagation and alignment module, before feature propagation, the feature graph group sequence is aligned in frequency domain and space domain, and the frequency domain and space domain alignment includes frequency domain alignment and space domain alignment in sequence. When performing frequency domain alignment,

[0011] Performing Fourier transform on the i-th feature image group corresponding to the i-th frame of low-resolution fluorescence microscopy image in the feature image group sequence to obtain the phase spectrum of the i-th feature image group, wherein i is a positive integer and is less than P, and then, while ensuring that ik is greater than 0 and not greater than P, performing Fourier transform on the ik-th feature image group corresponding to the ik-th frame of low-resolution fluorescence microscopy image in the feature image group sequence to obtain the phase spectrum and amplitude spectrum of the ik-th feature image group, wherein k=1, 2, -1 or -2;

[0012] Combine the phase spectrum of the i-th feature map group with the ik-th feature map group h k After the phase spectrum of the ikth feature map group is channel superimposed, it is processed by multiple sequentially connected basic neural network submodules and added to the phase spectrum of the ikth feature map group and inverse Fourier transformed together with the amplitude spectrum of the ikth feature map group to obtain the ikth frequency domain enhanced feature map group.

[0013] When performing spatial domain alignment, the ikth frequency domain enhanced feature map group, the ith feature map group, and the optical flow calculated from the ith frame low-resolution fluorescence microscopy image and the ikth frame low-resolution fluorescence microscopy image are used as input, and the deformable convolutional alignment neural network architecture processed by optical flow correction is used as the output of feature propagation and alignment processing.

[0014] In the reconstruction module, the feature map group sequence processed by the feature propagation and alignment module is reconstructed into a super-resolution image sequence.

[0015] Optionally, the low-resolution fluorescence microscopy image sequence is upsampled and then added to the super-resolution image sequence.

[0016] Optionally, when training the neural network of the super-resolution processing system, a low-resolution fluorescence microscopy image sequence consisting of multiple low-resolution fluorescence microscopy images acquired for a biological sample is used as a training input, and a high-resolution image sequence acquired for the same biological sample is used as a training truth value.

[0017] Optionally, when training the neural network of the super-resolution processing system, a structured light image sequence obtained by using structured light imaging technology for biological samples and a structured light super-resolution image sequence obtained by super-resolution processing is used as a training true value, and an image sequence obtained by selectively combining the structured light image sequence is used as a training input, wherein the selective combination depends on the phase difference between the structured light fringes.

[0018] Optionally, in the feature propagation and alignment module, propagation is performed in the form of primary propagation, and / or secondary propagation, and / or higher-order propagation.

[0019] Optionally, the feature propagation is implemented using multiple basic neural network sub-modules connected in sequence; and / or, the reconstruction module includes multiple basic neural network sub-modules connected in sequence and a pixel rearrangement sub-module to perform super-resolution processing.

[0020] Optionally, the basic neural network submodule for feature extraction, the basic neural network submodule for frequency domain alignment, the basic neural network submodule for feature propagation, and / or the basic neural network submodule for reconstruction are the same or different.

[0021] Optionally, the basic neural network submodule can select a residual neural network of any depth, a U-net, a residual channel attention neural network (RCAN), a residual dense neural network (RDN) or a self-attention neural network (Transformer).

[0022] Optionally, in each training cycle, pixel blocks at the same position in each image sequence are randomly selected, randomly rotated and folded, and used as neural network input images and target images respectively. The error between the network output and the target image is calculated, and its gradient is back-propagated to update the network parameters. After the network input error converges, the training is stopped and the network parameters are stored.

[0023] Optionally, the low-resolution fluorescence microscopy image is a wide-field fluorescence microscopy image or an image captured by a low numerical aperture microscopy imaging system.

[0024] Optionally, the high-resolution image sequence is a structured light image sequence obtained by using structured light imaging technology for the same biological sample and then subjected to super-resolution processing to obtain a structured light super-high-resolution image sequence.

[0025] Optionally, the high-resolution image sequence is a sequence of fluorescence microscopic images taken with a high numerical aperture microscope for the same biological sample.

[0026] Optionally, the Fourier transform and the inverse Fourier transform are both performed in the real number domain.

[0027] According to another aspect of the present application, a computer program product is provided, including a computer program or instructions, characterized in that the computer program or instructions perform the aforementioned method or modification when executed by a processor.

[0028] Using the method of the present application, a continuous sequence of multiple low-resolution fluorescence microscopy images containing noise (such as a wide-field fluorescence microscopy image sequence or a fluorescence microscopy image sequence taken by a low numerical aperture microscope) can be used as input, and a corresponding high-resolution microscopy image sequence (such as a structured light illumination super-resolution microscopy image sequence or a fluorescence microscopy image sequence taken by a high numerical aperture microscope) can be used as the true value to train the neural network of the super-resolution processing system. The trained neural network can directly generate a clear high-resolution image sequence from a low-resolution image sequence with an extremely low signal-to-noise ratio. In particular, for long-term video data of living cells, the method of the present application can perform super-resolution processing on the long-term wide-field illumination imaging results based on the trained neural network, which effectively improves the fidelity and temporal continuity of super-resolution reconstruction compared to the single image super-resolution algorithm, namely SISR (Single Image Super-resolution). By using the method of the present application, during the imaging process, extremely low excitation light power and extremely short exposure time can be used compared to normal shooting, and image restoration and super-resolution processing can be performed through a neural network, which significantly reduces the phototoxicity of in vivo super-resolution imaging, and can capture high-speed dynamic processes, thereby improving the applicability of in vivo cell imaging. In summary, the present invention is of great significance for improving the quality of microscopic images and expanding the application of super-resolution microscopy technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The principles and various aspects of the present application can be more fully understood from the following detailed description in conjunction with the following drawings. It should be noted that the scales of the drawings may be different for the purpose of clear description, but this will not affect the understanding of the present application. In the drawings:

[0030] Figure 1 The basic block diagram of a microscopic imaging system according to an embodiment of the present application is schematically shown;

[0031] Figure 2 A super-resolution processing system according to an embodiment of the present application is schematically shown, wherein the system processes a wide-field fluorescence image sequence;

[0032] Figure 3 An example of a basic neural network according to an embodiment of the present application is schematically shown, wherein the basic neural network is shown as a residual neural network; and

[0033] Figure 4An example of a frequency domain-spatial domain alignment submodule according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0034] In the various figures of the present application, features with the same structure or similar functions are represented by the same reference numerals.

[0035] Figure 1 The basic block diagram of a microscopic imaging system is schematically shown, which generally includes an optical imaging system 100 and a control and data processing system 200. The optical imaging system 100 can be, but is not limited to, a wide-field microscopic imaging system, a scanning confocal microscopic imaging system, a light sheet illumination microscopic imaging system, a two-photon scanning imaging system, and a structured light illumination microscopic imaging system. Taking the structured light illumination microscopic imaging system as an example, the optical imaging system 100 includes an excitation light path and a detection light path, wherein the excitation light path includes an excitation objective lens and other optical components for generating excitation light, and the excitation light beam can be emitted through the excitation objective lens so as to excite fluorescence on a biological sample, and the detection light path includes a detection objective lens and other optical components for imaging, which are used to receive and detect the excited fluorescence. It should be clear to those skilled in the art that, depending on the configuration of the microscopic imaging system, the excitation objective lens and the detection objective lens can be the same objective lens or different objective lenses.

[0036] According to one embodiment of the present application, the same optical imaging system 100 can switch between a low-resolution fluorescence microscopy imaging mode (such as a wide-field fluorescence microscopy imaging mode or a low numerical aperture microscopy imaging mode) and a high-resolution microscopy imaging mode (such as a structured light illumination super-resolution microscopy imaging mode or a high numerical aperture microscopy imaging mode). The following is an explanation of the low-resolution fluorescence microscopy imaging mode using only the wide-field microscopy imaging mode as an example, and the following is an explanation of the high-resolution microscopy imaging mode using only the structured light illumination super-resolution microscopy imaging mode as an example. However, those skilled in the art should be aware that the low-resolution fluorescence microscopy imaging mode is not limited to the wide-field microscopy imaging mode and the high-resolution microscopy imaging mode is not limited to the structured light illumination super-resolution microscopy imaging mode. Under the premise of meeting the purpose of the present application, those skilled in the art can use any suitable imaging mode for corresponding replacement. When the optical imaging system 100 is in the wide-field microscopy imaging mode, two-dimensional fluorescence microscopy imaging can be performed on biological samples, especially living biological samples, to obtain a two-dimensional wide-field fluorescence image time series (hereinafter referred to as the wide-field image time series) for the biological sample. When the optical imaging system 100 is in the structured light illumination microscopy imaging mode, two-dimensional fluorescence microscopy imaging can be performed on biological samples, especially living biological samples, and then a two-dimensional super-resolution fluorescence image time series for the biological samples is obtained through processing by the control and data processing system 200. Within the scope of discussion of the present application, the optical imaging system 100 in the structured light illumination super-resolution microscopy mode can adopt any suitable form known in the art. For example, as just an example, the optical imaging system described in the public document Visualizing Intracellular Organelle and Cytoskeletal Interactions at Nanoscale Resolution on Millisecond Timescales. Cell 175, 1430-1442e1417 (2018) by Guo, Y. et al. can be referred to, so its specific structure is not repeated in this application.

[0037] The control and data processing system 200 mainly includes a computer and related components (such as a data storage device, etc.), which can control the operation of the optical imaging system 100 and can receive image data from the optical imaging system 100 and perform corresponding post-processing. For example, when the optical imaging system 100 is in the structured light illumination microscopic imaging mode, the acquired structured light fluorescence image time series (hereinafter referred to as structured light fluorescence image sequence) is provided to the control and data processing system 200, and is reconstructed into a high signal-to-noise ratio, super-resolution microscopic image sequence through a series of structured light super-resolution data processing.

[0038] For this purpose, the control and data processing system 200 may include a structured light super-resolution reconstruction module 210 capable of selecting a standard structured light super-resolution reconstruction algorithm to perform structured light super-resolution reconstruction on the structured light fluorescence image sequence. In the scope of the present application, the standard structured light super-resolution reconstruction algorithm can be considered as an algorithm already known in the field of microscopic imaging. As an example, the standard structured light super-resolution reconstruction algorithm can refer to the public document Three-dimensional resolution doubling in wide-field fluorescence microscopy by structured illumination. Biophys J 94, 4957-4970 (2008) by Gustafsson, MG et al.

[0039] It should be pointed out that within the scope of the present application, the modules and / or submodules described herein can be understood to include data storage devices, such as computer-readable storage media, in which programs or subroutines and neural network models that are called and run by computers can be stored. These programs or subroutines and neural network models can implement the methods / steps described below when called and executed by a computer. The specific programming methods for programs and / or subroutines are not discussed in this application, and those skilled in the art can implement the relevant functions with any well-known programming software and / or commercial software. Therefore, the context of this application should be understood as describing the operation of the relevant system or the operation or method of the module, as they can also be written as programs to be called and executed by a computer.

[0040] The control and data processing system 200 may further include a feature extraction module 220, a feature propagation and alignment module 230, and a reconstruction module 240. The feature extraction module 220, the feature propagation and alignment module 230, and the reconstruction module 240 constitute a video (microscopic image time series or microscopic image series) super-resolution processing (or super-resolution reconstruction) system according to an embodiment of the present application. Figure 2 The example of the video (microscopic image time series or microscopic image sequence) super-resolution processing (or super-resolution reconstruction) system processing a low-resolution fluorescence image time series (referred to as low-resolution image sequence) is schematically shown. In the context of the present application, the term related to video can be considered to be composed of images grouped according to the order of shooting in time, which can be called image time series, or image sequence for short.

[0041] The feature extraction module 220 includes a plurality of basic neural network submodules connected in sequence. In the following explanation of the technical solution of the present application, the basic neural network is described by taking the residual (ResNet) neural network as an example. However, it should be clear to those skilled in the art that the basic neural network used in the embodiments of the present application can be any other suitable neural network, for example, including but not limited to a U-net of any depth, a residual channel attention neural network (RCAN), a residual dense neural network (RDN), a self-attention neural network (Transformer), etc., as long as the basic neural network used can achieve the purpose of the present application.

[0042] According to one embodiment of the present application, the feature extraction module 220 includes N sequentially connected residual neural network submodules 221, where N is an integer greater than 1, for example, five sequentially connected residual neural network submodules 221. Each residual neural network submodule may include, for example, two convolutional layers and an activation function, such as Figure 3 In addition, the feature extraction module 220 may also include one or more convolutional layers 222 (in Figure 2 (shown as one in the figure) to adjust the number of channels of the residual neural network submodules 221 connected in sequence to the input, for example, from the single channel of the input to the commonly used number of channels for feature extraction in the convolution layer (for example, 64 channels). According to an embodiment of the present application, the feature extraction module 220 is configured to perform feature extraction processing on a low-resolution fluorescence microscopy image or image sequence (for example, a wide-field fluorescence microscopy image or a wide-field fluorescence microscopy image sequence or a low numerical aperture microscopy image or image sequence) acquired by the optical imaging system 100 to output a feature map group or a feature map group sequence. It should be pointed out that in the context of the present application, an image, a low-resolution image, a super-resolution image or a map can be understood as a two-dimensional digital matrix that can be processed by a computer; an image sequence or a map group can be understood as a group of two-dimensional digital matrices; and a map group sequence can be understood as a group of map groups. For example, for the feature extraction module 220, when a single low-resolution fluorescence microscopic image (hereinafter referred to as a low-resolution image) is input, a feature map group can be output after feature extraction processing; and when a sequence of low-resolution fluorescence microscopic images (hereinafter referred to as a low-resolution image sequence) is input, a sequence of feature map groups can be output after feature processing. When considering the image sequence, based on the time propagation order, the first image can be called the previous frame image, and the next image can be called the next frame image.

[0043] It should be clear that the basic principles of neural network should be used as the basis for understanding the feature extraction module 220, feature propagation and alignment module 230, and reconstruction module 240 of the present application. For example, when referring to residual neural network submodules or convolutional layers in other modules as follows, it may refer to the same residual neural network submodules 221 or convolutional layers 222 in the feature extraction module 220, and the usage of these residual neural network submodules 221 is reconfigured only because the functions of the modules mentioned are different. In an alternative embodiment, when referring to residual neural network submodules or convolutional layers in other modules as follows, it may also refer to residual neural network submodules or convolutional layers that are different from those in the feature extraction module 220.

[0044] The feature propagation and alignment module 230 may also include M sequentially connected basic neural network submodules, for example, M sequentially connected residual neural network submodules 231, where M is an integer greater than 1, for example, five sequentially connected residual neural network submodules 231. In addition, a convolution layer 232 is provided before the M sequentially connected residual neural network submodules 231 to adjust the number of channels input to the sequentially connected residual neural network submodules 231. According to an embodiment of the present application, the residual neural network submodule 231 and / or the convolution layer 231 may be the same or different residual neural network submodule and / or convolution layer as the residual neural network submodule 221 and / or the convolution layer 221.

[0045] The feature propagation and alignment module 230 is configured to take the feature map group sequence output by the feature extraction module 220 as input, and the feature map group corresponding to each frame of the low-resolution image (processed by the feature extraction module 220) is processed by the frequency domain-spatial domain alignment submodule 233, and then processed by the convolution layer 232 and the above-mentioned M sequentially connected residual neural network submodules 231, and its output constitutes a new feature map group for the next frame of image. The new feature map of the next frame of image continues to be processed by the frequency domain-spatial domain alignment submodule 233 and then processed by the convolution layer 232 and the above-mentioned M sequentially connected residual neural network submodules 231 to form a new feature map group, and recursively propagates to the last frame in sequence. At this point, all feature map groups corresponding to all frames have been updated. Next, repeat this propagation process in the opposite direction to update the feature map groups corresponding to all frames again. After this, repeat the above forward and reverse propagation processes once or multiple times.

[0046] The propagation process of the above feature propagation can be referred to the public document BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment. in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition 5972-5981 (2022). Figure 2 In FIG. 2 , the solid arrow in the feature propagation and alignment module 230 represents the primary propagation, and the dashed arrow represents the secondary propagation. Of course, it should be clear to those skilled in the art that feature propagation can also be implemented in a higher-order manner.

[0047] Figure 4 An example of a frequency domain-spatial domain alignment submodule 233 according to an embodiment of the present application is schematically shown. As shown in the figure, the frequency domain-spatial domain alignment submodule 233 may include a frequency domain alignment part and a spatial domain alignment part, so that the input feature map group is first aligned in the frequency domain and then aligned in the spatial domain. For example, the frequency domain alignment part of the frequency domain-spatial domain alignment submodule 233 may include M' sequentially connected basic neural network submodules, such as M' sequentially connected residual neural network submodules 23311, where M' is an integer greater than 1, such as five sequentially connected residual neural network submodules 23311. In addition, a convolutional layer 23312 is set before the M' sequentially connected residual neural network submodules 23311 to adjust the number of channels of the input sequentially connected residual neural network submodules 23311.

[0048] Refer to the following Figure 4 The working method of the frequency domain-spatial domain alignment submodule 233 according to the present application is introduced. For a frame of low-resolution image (hereinafter referred to as the current frame) in the low-resolution image sequence collected by the optical imaging system 100, the frame image is processed by the feature extraction module 220 to obtain a feature map group h i , where i represents the temporal order of the low-resolution images selected in the low-resolution image sequence. For example, when a low-resolution image sequence is composed of P low-resolution images (P is a positive integer greater than or equal to 3), i can be any integer less than P. Therefore, the low-resolution image can be called the i-th low-resolution image. In addition, the previous frame or the previous two frames of the image or the next frame or the next two frames of the image are also processed by the feature extraction module 220 to obtain the feature map group h i-k , where k = -1, 1, -2 or 2. Therefore, the corresponding low-resolution image can be called the ikth frame low-resolution image (ensuring that ik>0 and is less than or equal to P).

[0049] Then, in the frequency domain alignment part, Fourier transform is used to align the feature map groups h i And feature map group h i-k Processing is performed to generate a feature map group h i The phase spectrum and characteristic diagram group h i-k The phase spectrum and amplitude spectrum of are expressed as follows:

[0050]

[0051] Wherein, FFT represents fast Fourier transform, Angle and Abs represent phase spectrum and amplitude spectrum after fast Fourier transform, respectively. It should be clear to those skilled in the art that the fast Fourier transform (or inverse transform) mentioned herein is a fast implementation algorithm of the Fourier transform (or inverse transform).

[0052] As shown in the figure, the feature map group h i The phase spectrum and characteristic diagram group h i-k Here, channel superposition can be achieved by any superposition method familiar to those skilled in the art, such as front-to-back channel superposition or interleaved channel superposition of two phase spectra. Then, the result of channel superposition is processed by the convolution layer 23312 and the residual neural network submodule 23311 to obtain the phase residual, and is combined with the feature map h i-k The phase spectrum is added to obtain the following phase spectrum For example, the process can be represented by the following formula:

[0053]

[0054] Wherein, RB represents the residual neural network submodule 23311, f represents the convolution layer 23312, and ⊕ represents channel superposition. In the formula shown, the superscript (2) of RB may mean that the data processed by the convolution layer then passes through two residual neural networks in sequence, but may also pass through M' residual neural networks as described above as needed.

[0055] Then, the phase spectrum obtained above is With the feature map group h i-k The amplitude spectrum of the re-Fourier inverse transform is used as the output of the frequency domain alignment part, as follows

[0056]

[0057] Among them, iFFT represents the inverse fast Fourier transform, so that the feature map group after frequency domain enhancement processing is obtained As the output of the frequency domain alignment part. It is worth noting that the above-mentioned fast Fourier transform and inverse transform are both performed in the real domain. Those skilled in the art should be aware that the results produced by the two-dimensional Fourier transform have a conjugate symmetric relationship. If the Fourier transform and inverse transform are used directly, non-real transformation results will be produced, resulting in a strong uncertainty in the result. Therefore, the technical solution of the present application adopts the real domain Fourier transform, that is, considering the conjugate symmetry of the Fourier transform result, only half of the information is saved, and the other half is restored in the form of conjugate symmetry, thereby ensuring that the inverse Fourier transform result is still a real number.

[0058] like Figure 4 As shown, the feature map group after the above frequency domain enhancement processing With the feature map group h i , together with the optical flow calculated from the i-th frame low-resolution image and the ik-th frame low-resolution image, as the input of the spatial alignment part of the frequency-spatial alignment submodule 233. In an embodiment of the present application, the spatial alignment part can be set as a deformable convolutional alignment neural network framework for optical flow correction in a suitable manner known to those skilled in the art. For example, the optical flow prediction neural network part framework in the deformable convolutional alignment neural network for optical flow correction can be implemented with reference to the public document A Ranjan, MJ Black: Optical Flow Estimation Using a Spatial Pyramid Network. IEEE, 2017, and the deformable convolutional alignment neural network part framework can be implemented with reference to the public document BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment. in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition 5972-5981 (2022), which will not be described in detail in this article.

[0059] After the spatial alignment part of the frequency domain-spatial domain alignment submodule 233 is processed, the feature map group h i And feature map group h i-kSpatial alignment is performed. For example, specifically, an optical flow prediction neural network (such as Spynet) is used to predict the optical flow between the i-th frame low-resolution image and the ik-th frame low-resolution image. This neural network combines the classic spatial pyramid method with deep learning to calculate the optical flow, which can avoid large movements. The entire video sequence (original image sequence) can be input into the optical flow prediction neural network for prediction, and the pixel-level horizontal and vertical axis movements can be obtained. Although the optical flow prediction accuracy is not high, it can be used as the basis for subsequent deformable convolution, and the optical flow can be used to first perform spatial optical flow correction on the feature map group that has been aligned in the frequency domain. After the optical flow correction, the feature map group that has been corrected by the optical flow can be channel-superimposed with the feature map group of the current frame, and a convolutional neural network is used to generate the offset (Offset) required for the deformable convolution. The difference between the deformable convolution and the general convolution is that its sampling position can be flexibly determined according to the offset size, giving it a stronger spatial alignment capability. Using the offset generated by the convolutional neural network, the feature map group that has been corrected by the optical flow is deformed convolved to generate a feature map group that is dually aligned in the frequency domain and the spatial domain.

[0060] The reconstruction module 240 may include M” sequentially connected basic neural network submodules, for example, M” sequentially connected residual neural network submodules 241, where M” is an integer greater than 1, for example, five sequentially connected residual neural network submodules 241. In addition, a convolution layer 242 is set before the M” sequentially connected residual neural network submodules 241 to adjust the number of channels input to the sequentially connected residual neural network submodules 241. According to an embodiment of the present application, the residual neural network submodule 241 and / or the convolution layer 241 may be the same or different residual neural network submodule and / or convolution layer as the residual neural network submodule 221 and / or the convolution layer 221. In addition, the reconstruction module 240 also includes a pixel shuffle submodule 243 for super-resolution processing. In the reconstruction module 240, a convolution layer 244 is also configured after the pixel shuffle submodule 243 to adjust the number of channels accordingly. According to an embodiment of the present application, the input of the reconstruction module 240 is a feature map group sequence corresponding to the low-resolution image sequence input to the feature extraction module 220, which has been processed by the feature propagation and alignment module 230. After the feature map group sequence is processed by the reconstruction module 240, a super-resolution image sequence corresponding to the low-resolution image sequence is output.

[0061] like Figure 4 As shown, according to a preferred embodiment of the present application, the super-resolution image sequence output by the reconstruction module 240 is added to the image sequence after upsampling the low-resolution image sequence input to the feature extraction module 220 (the sampling multiple corresponds to the pixel rearrangement processing), and the output is used as the final super-resolution image sequence.

[0062] The advantages of the above technical solution of this application are:

[0063] (1) This application proposes a frequency domain-spatial domain alignment algorithm based on deformable convolution, which can achieve better alignment and matching of microscopic image feature map groups.

[0064] (2) The present application proposes a video super-resolution reconstruction system that utilizes the above-mentioned deformable convolution-based frequency-domain-spatial domain alignment algorithm, which can achieve super-resolution processing of multiple continuous time series of low-resolution, noisy original microscopic images.

[0065] (3) The method proposed in this application can be applied to original microscopic data of different modalities and can achieve reconstruction and restoration of data of different modalities.

[0066] (4) The image reconstructed by the method proposed in this application has good fidelity and temporal continuity that exceeds other existing technologies.

[0067] When training the neural network of the video super-resolution processing system of the present application, the relevant neural network model is optimized using a loss function, and the loss function includes but is not limited to mean square error (MSE), mean absolute error (MAE), structural similarity (SSIM) or their weighted sum, etc.

[0068] For each group of original structured light image sequences, they are normalized and assigned to the training data set. The video super-resolution neural network is trained using the training data set generated by the above method. For example, in each training cycle, a matching low-resolution image sequence and a high-resolution image sequence as a true value (such as a corresponding structured light super-resolution image sequence or a high numerical aperture microscopic image sequence) are randomly taken out from the data set, and pixel blocks at the same position are randomly extracted, and randomly rotated and folded as neural network input images and target images respectively. The error between the network output and the target image is calculated, and its gradient is back-propagated to update the network parameters. After the network input error converges, the training is stopped and the network parameters are stored. It should be clear to those skilled in the art that the training method of the neural network is not limited to that listed.

[0069] According to an embodiment of the present application, during training, a low-resolution image sequence as input can be obtained by the optical imaging system 100 in a low-resolution imaging mode for a biological sample to be detected, and at the same time, a high-resolution image sequence can be obtained by the optical imaging system 100 in a high-resolution microscopic imaging mode for a biological sample to be detected as a training truth value. For example, as a non-limiting example, the optical imaging system 100 obtains a structured light image sequence in a structured light illumination imaging mode for a biological sample to be detected, and processes the structured light super-resolution into a super-resolution image sequence as a training truth value. In an alternative embodiment, the structured light image sequence used for super-resolution processing can be selectively combined into a wide-field image sequence (based on the phase difference of the structured light fringes, for example, when the structured light fringes are 2π / 3, three adjacent structured light images are combined into a wide-field image) as a training input. In an alternative embodiment, a high-resolution image sequence can be directly obtained by the optical imaging system 100 in a high numerical aperture microscopic imaging mode for a biological sample to be detected as a training truth value.

[0070] After the neural network is trained, the trained neural network and its parameters can be used to perform super-resolution processing on long-term low-resolution image time series. The neural network can input multiple consecutive pictures at one time and output the corresponding 2x or 3x upsampled image sequence. It can complete super-resolution processing of multiple images within one second, which improves the imaging quality and time continuity. It can be widely used for video observation and recording of living cells.

[0071] The above training and reasoning process is only a non-limiting example of the neural network system for video super-resolution processing based on the frequency domain-spatial domain alignment technology of the present application. Those skilled in the art should be aware that the "frequency domain-spatial domain alignment algorithm based on deformable convolution" of the present application can also be used in other neural network architectures for videos, such as microscopic image tasks such as denoising, registration, and segmentation of multiple images.

[0072] After the neural network training of the video (microscopic image time series) super-resolution processing system is completed, the original fluorescence image sequence Y obtained by the optical imaging system is super-resolution processed (or predicted). i(i=1,2,…,N, where N is an integer) The video (microscopic image time series) super-resolution processing system is used to perform super-resolution reconstruction processing to obtain a super-resolution image Z as the final super-resolution image. Of course, if the video (microscopic image time series) super-resolution processing system is used to process a long-term original fluorescence image sequence, the long-term original fluorescence image sequence can be segmented, and each segment is input into the video (microscopic image time series) super-resolution processing system for processing, and finally the sliding window processing is performed to restore the long-term data, that is, the processing process is repeated for each original fluorescence image sequence.

[0073] Just as an example, in the technical solution described in this application, the reference values ​​of the main parameters of the residual neural network as an example are shown in Table 1. It should be noted that the basic neural network module that can be used in this application does not require the use of a specific neural network structure. Typical network structures such as U-Net, residual channel attention neural network (RCAN), residual dense neural network (RDN), self-attention neural network (Transformer) can all achieve the functions described.

[0074] name Main parameters Convolutional Layer Kernel size 3×3 Activate Layer LeakyRelu function with a slope of 0.1

[0075] Table 1 Main parameters of the neural network

[0076] In addition, the main parameters used in the neural network training process of the present application need to be adjusted according to the specific situation of the data set (signal-to-noise ratio, structural complexity, etc.). A set of parameters for reference is shown in Table 2.

[0077]

[0078] Table 2 Reference parameters of network training process

[0079] This application proposes a microscopic video reconstruction system based on a frequency domain-spatial domain alignment algorithm, which can effectively process continuous time series microscopic data on original images with extremely low signal-to-noise ratios, and reconstruct super-resolution noise-free images with high fidelity. The application of this method significantly reduces the requirements for the signal-to-noise ratio of the original image during microscopic imaging, thereby reducing the light damage to biological samples during imaging and increasing the imaging rate, significantly improving the image quality of structured light microscopy and expanding its application range.

[0080] Although the specific embodiments of the present application are described in detail herein, they are provided only for the purpose of explanation and should not be considered to limit the scope of the present application. In addition, it should be clear to those skilled in the art that the various embodiments described in this specification can be used in combination with each other. Various replacements, changes and modifications can be conceived without departing from the spirit and scope of the present application.

Claims

1. A method for super-resolution processing of a microscopic image sequence based on frequency domain-spatial domain alignment, comprising: Providing a low-resolution fluorescence microscopy image sequence consisting of P low-resolution fluorescence microscopy images for a biological sample, wherein P is a positive integer greater than or equal to 3; and The low-resolution fluorescence microscopic image sequence is passed through a super-resolution processing system consisting of a feature extraction module, a feature propagation and alignment module, and a reconstruction module, wherein in the feature extraction module, a plurality of sequentially connected basic neural network submodules are used to extract features from the low-resolution fluorescence microscopic image sequence to generate a feature map group sequence, In the feature propagation and alignment module, before feature propagation, the feature graph group sequence is aligned in frequency domain and space domain, and the frequency domain and space domain alignment includes frequency domain alignment and space domain alignment in sequence. When performing frequency domain alignment, Perform Fourier transform on the i-th feature image group corresponding to the i-th frame of low-resolution fluorescence microscopy image in the feature image group sequence to obtain the phase spectrum of the i-th feature image group Where i is a positive integer and is less than P. Then, while ensuring that ik is greater than 0 and not greater than P, Fourier transform is performed on the ikth feature image group corresponding to the ikth frame of low-resolution fluorescence microscopy image in the feature image group sequence to obtain the phase spectrum of the ikth feature image group. and Amplitude Spectrum Wherein, k = 1, 2, -1 or -2; Combine the phase spectrum of the i-th feature map group with the ik-th feature map group h k After the phase spectrum of the ikth feature map group is superimposed on the channel, it is processed by multiple sequentially connected basic neural network submodules and combined with the phase spectrum of the ikth feature map group Add and combine with the amplitude spectrum of the ikth feature map group Perform inverse Fourier transform together to obtain the ikth frequency domain enhanced feature map group When performing spatial domain alignment, the ikth frequency domain enhanced feature map group, the ith feature map group, and the optical flow calculated from the ith frame low-resolution fluorescence microscopy image and the ikth frame low-resolution fluorescence microscopy image are used as input, and the deformable convolutional alignment neural network architecture processed by optical flow correction is used as the output of feature propagation and alignment processing. In the reconstruction module, the feature map group sequence processed by the feature propagation and alignment module is reconstructed into a super-resolution image sequence.

2. The method according to claim 1, characterized in that The low-resolution fluorescence microscopy image sequence is upsampled and added to the super-resolution image sequence.

3. The method according to claim 1 or 2, characterized in that: When training the neural network of the super-resolution processing system, a low-resolution fluorescence microscopic image sequence consisting of multiple low-resolution fluorescence microscopic images obtained for a biological sample is used as a training input, and a high-resolution image sequence obtained for the same biological sample is used as a training truth value.

4. The method according to claim 1 or 2, characterized in that: When training the neural network of the super-resolution processing system, a structured light image sequence obtained by using structured light imaging technology for biological samples and a structured light super-resolution image sequence obtained by super-resolution processing is used as a training truth value, and an image sequence obtained by selectively combining the structured light image sequence is used as a training input, wherein the selective combination depends on the phase difference between the structured light fringes.

5. The method according to claim 1 or 2, characterized in that: In the feature propagation and alignment module, propagation is performed in the form of primary propagation, and / or secondary propagation, and / or higher-order propagation.

6. The method according to claim 5, characterized in that The feature propagation is implemented using multiple basic neural network sub-modules connected in sequence; and / or, the reconstruction module includes multiple basic neural network sub-modules connected in sequence and a pixel rearrangement sub-module to perform super-resolution processing.

7. The method according to claim 6, characterized in that The basic neural network submodule for feature extraction, the basic neural network submodule for frequency domain alignment, the basic neural network submodule for feature propagation, and / or the basic neural network submodule for reconstruction are the same or different.

8. The method according to claim 7, characterized in that The basic neural network submodule can select a residual neural network of any depth, a U-net, a residual channel attention neural network (RCAN), a residual dense neural network (RDN), or a self-attention neural network.

9. The method according to claim 8, characterized in that In each training cycle, pixel blocks at the same position in each image sequence are randomly selected, randomly rotated and folded, and used as the neural network input image and target image respectively. The error between the network output and the target image is calculated, and its gradient is back-propagated to update the network parameters. After the network input error converges, the training is stopped and the network parameters are stored.

10. The method according to claim 1 or 2, characterized in that: The low-resolution fluorescence microscopic image is a wide-field fluorescence microscopic image or an image taken by a low numerical aperture microscopic imaging system.

11. The method according to claim 3, characterized in that The high-resolution image sequence is a structured light image sequence obtained by using structured light imaging technology for the same biological sample and then subjected to super-resolution processing to obtain a structured light super-high-resolution image sequence.

12. The method according to claim 3, characterized in that The high-resolution image sequence is a sequence of fluorescence microscopic images taken with a high numerical aperture microscope for the same biological sample.

13. The method according to claim 1 or 2, characterized in that: The Fourier transform and the inverse Fourier transform are both performed in the real number domain.

14. The method according to claim 3, characterized in that In the feature propagation and alignment module, propagation is performed in the form of primary propagation, and / or secondary propagation, and / or higher-order propagation.

15. The method according to claim 4, characterized in that In the feature propagation and alignment module, propagation is performed in the form of primary propagation, and / or secondary propagation, and / or higher-order propagation.

16. A computer program product comprising a computer program or instructions, characterized in that The computer program or instructions implement the methods of claims 1 to 15 when executed by a processor.

Citation Information

Patent Citations

  • Rapid method and system for directly reconstructing structured light illumination super-resolution image in airspace

    CN111077121A

  • Wide-field illumination fluorescence super-resolution microscopic imaging method based on deep learning

    CN115841423A