Self-supervised structured light microscopy reconstruction method and system based on pixel rearrangement

The self-supervised multimodal structured light microscope reconstruction method addresses noise amplification issues by training a denoising neural network on rearranged pixel images, enhancing image quality and expanding live cell imaging capabilities.

JP7866807B2Active Publication Date: 2026-05-28INSTITUTE OF BIOPHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INSTITUTE OF BIOPHYSICS CHINESE ACADEMY OF SCIENCES
Filing Date
2024-03-29
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Structured light microscopy suffers from noise amplification in super-resolution reconstructed images due to low signal-to-noise ratios in original images, especially in live cell observations, leading to artifacts that hinder effective microscopy observation.

Method used

A self-supervised multimodal structured light microscope reconstruction method using pixel rearrangement and a denoising neural network trained on a single set of fluorescence images, eliminating the need for high signal-to-noise ratio images and reducing experimental costs.

Benefits of technology

Improves image quality and expands the applicability of structured light microscopy by effectively denoising super-resolution images without damaging living cells, enabling rapid dynamic process capture and reducing phototoxicity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007866807000003
    Figure 0007866807000003
  • Figure 0007866807000004
    Figure 0007866807000004
  • Figure 0007866807000005
    Figure 0007866807000005
Patent Text Reader

Abstract

This application discloses a self-supervised multimodal structured light microscope reconstruction method and system, comprising the steps of: exciting a biological sample using structured light; obtaining J original fluorescence image sequences (Y) generated by the excitation of the biological sample, where J is an integer of 1 or more, each original fluorescence image sequence (Y) contains S fluorescence images, where S is an integer of 2 or more, each fluorescence image has a pixel size of M*N, and M and N are even numbers; generating a training set for each original fluorescence image sequence (Y) of the J original fluorescence image sequences (Y); training a denoising neural network based on the training set; and performing super-resolution reconstruction on the S fluorescence images in each original fluorescence image sequence (Y) of the J original fluorescence image sequences (Y) using a standard structured light super-resolution reconstruction algorithm to form a super-resolution image, and obtaining a final super-resolution reconstructed image by using the super-resolution image as input to the denoising neural network. [Selected Figure] Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to a self-supervised structured light microscope reconstruction method and system, and more particularly to self-supervised structured light microscope reconstruction trained using a pixel rearrangement technique. [Background technology]

[0002] Structured light microscopy is a technique that involves irradiating a sample with multiple modulated excitation beams (e.g., sinusoidal moiré fringes) and then reconstructing it using a specific algorithm. Compared to conventional fluorescence microscopy, it achieves significantly higher resolution microscopic imaging results, enabling more detailed microstructural analysis. Structured light microscopy is widely applied in the field of live cell imaging due to its fast imaging speed, low phototoxicity, and broad applicability to various samples.

[0003] However, unlike conventional fluorescence microscopy techniques where the collected original images become the final visualization result, super-resolution reconstructed images using structured light microscopy techniques must be obtained from a series of original images using a reconstruction algorithm. Here, the reconstruction algorithm involves multiple complex steps such as Fourier transform, inverse Fourier transform, frequency domain translation, frequency domain filtering, and frequency domain splicing. Therefore, if noise is present in the original images, the noise will be significantly amplified in the images processed by the reconstruction algorithm. If the signal-to-noise ratio of the original images is low, it will have a significant impact on the quality of the final super-resolution reconstructed image, resulting in serious reconstruction "artifacts." The presence of "artifacts" significantly affects the quality of the final super-resolution reconstructed image, making it impossible to distinguish between actual sample information and "artifacts" generated during reconstruction, thus affecting the effectiveness of microscopy observation.

[0004] When imaging and observing living cells using structured light microscopy, the signal-to-noise ratio of the resulting original fluorescence image is typically low due to the following factors: 1. Because the types of fluorescent dyes and fluorescent proteins that can be applied to fluorescent labeling of living cells are limited, the fluorescent labeling efficiency is usually low, resulting in low emission efficiency of the fluorescent signal when the sample is excited by light irradiation. 2. To avoid damaging living cells, only low-intensity excitation light sources can usually be used, which also reduces fluorescence intensity. 3. Since living cells typically move at high speeds, capturing the migration process of living cells requires minimizing the exposure time for imaging, which in turn reduces the intensity of the acquired fluorescence signal accordingly.

[0005] For these reasons, how to remove image noise, especially in live cell observation, is a crucial breakthrough in structured light microscopy imaging technology. Prior research on image noise reduction in structured light microscopy imaging technology has mainly used noise reduction algorithms such as Wiener filterlin and BM3D to directionally suppress noise based on the statistical difference between image noise and sample information. However, when the signal-to-noise ratio is extremely low in the original fluorescence image obtained from live cell observation, this algorithm using statistical difference cannot effectively remove noise due to the deep coupling between image noise and sample information.

[0006] Furthermore, with the rapid development of deep learning technology in recent years, image noise reduction using neural network methods has been attracting increasing attention. However, when training a noise reduction neural network using original fluorescence images of living cells, it is usually necessary to use a "supervised" training mechanism, which requires the prior collection of a large number of "high signal-to-noise ratio - low signal-to-noise ratio" image pairs to construct a training set. However, collecting a large number of "high signal-to-noise ratio - low signal-to-noise ratio" image pairs requires an increase in excitation light intensity, which not only damages living cells but also significantly increases experimental costs due to the increased selection requirements for the types of fluorescent dyes and fluorescent proteins. In addition, in living cell observation, it is necessary to sample the same biological sample multiple times, which reduces temporal resolution and makes it impossible to process video data. [Overview of the project] [Problems that the invention aims to solve]

[0007] To address the above problem, this application proposes a novel self-supervised multimodal structured light microscope image super-resolution reconstruction noise reduction technique. When training the neural network used in this technique, the training set is generated using a novel "pixel rearrangement technique" on a single group of fluorescence images collected only once (eliminating the need to use high signal-to-noise ratio images as training images or to sample the same raw biological sample multiple times). This allows for "self-supervised" training of the neural network, ensuring that the trained neural network can reconstruct denoised super-resolution images based on the original fluorescence images with a low signal-to-noise ratio. This significantly improves the image quality of structured light microscopes and expands their range of applications. [Means for solving the problem]

[0008] According to one aspect of this application, a self-supervised multimodal structured light microscope reconstruction method, The step of exciting a biological sample using structured light and obtaining J original fluorescence image sequences (Y) generated by the excitation of the biological sample, wherein J is an integer greater than or equal to 1, each original fluorescence image sequence (Y) contains S fluorescence images, S is an integer greater than or equal to 2, and each fluorescence image has a pixel size of M*N, where M and N are even numbers. For each of the J original fluorescence image sequences (Y), a program is executed on a computer, 1) The i-th fluorescence image (y i By rearranging pixels in ), the first, second, third, and fourth sub-images with a pixel size of (M / 2)*(N / 2) are extracted, and the two-dimensional pixel matrix y of the first sub-image is obtained. i,(2m-1)(2n-1) , the two-dimensional pixel matrix y of the second sub-image i,(2m-1)(2n) , the two-dimensional pixel matrix y of the third sub-image i,(2m)(2n-1) , and the two-dimensional pixel matrix y of the fourth sub-image i,(2m)(2n) These are the i-th fluorescence images (y i ) 2D pixel matrix y i,MN A step in which a selection is made from, m is an integer and m=1, 2, 3, ..., M / 2, and n is an integer and n=1, 2, 3, ..., N / 2, 2) The first sub-image, the second sub-image, the third sub-image, and the fourth sub-image are subjected to a 2x pixel upsampling process to obtain a first upsampled sub-image, a second upsampled sub-image, a third upsampled sub-image, and a fourth upsampled sub-image with a pixel size of M*N; 3) Performing a pixel translation of [0.5,0.5] on the 2D pixel matrix of the first upsampled subimage to obtain the first subimage after 2D subpixel translation; performing a pixel translation of [-0.5,0.5] on the 2D pixel matrix of the second upsampled subimage to obtain the second subimage after 2D subpixel translation; performing a pixel translation of [0.5,-0.5] on the 2D pixel matrix of the third upsampled subimage to obtain the third subimage after 2D subpixel translation; and performing a pixel translation of [-0.5,-0.5] on the 2D pixel matrix of the fourth upsampled subimage to obtain the fourth subimage after 2D subpixel translation. 4) The steps of forming a first sub-image group from S 2D subpixel translations obtained from the S fluorescence images, S second sub-images obtained from the S fluorescence images, S second sub-images obtained from the S fluorescence images, S second sub-image group from the S fluorescence images, S third sub-image group from the S fluorescence images, S fourth sub-image group from the S fluorescence images, S fourth sub-image group from the S fluorescence images, 5) Using a standard structured light reconstruction algorithm, structured light reconstruction is performed on the first subimage group, the second subimage group, the third subimage group, and the fourth subimage group to obtain the first super-resolution subimage, the second super-resolution subimage, the third super-resolution subimage, and the fourth super-resolution subimage; 6) The steps of selecting two different super-resolution subimages from the first super-resolution subimage, the second super-resolution subimage, the third super-resolution subimage, and the fourth super-resolution subimage to form a training super-resolution image pair and adding it to the training set are performed, When a program is executed on a computer, 7) A step of training a denoising neural network based on the training set, 8) A self-supervised multimodal structured light microscope reconstruction method is provided, which involves performing super-resolution reconstruction on S fluorescence images in each of the J original fluorescence image sequences (Y) using a standard structured light super-resolution reconstruction algorithm to form a super-resolution image, using the super-resolution image as input to the denoising neural network to obtain a final super-resolution reconstructed image.

[0009] When selectively training a denoising neural network, one of the super-resolution subimages from the training super-resolution image pair is used as training input data, and the other width super-resolution subimage is used as training target data.

[0010] Selectively, if the maximum number of pixels in the row direction and / or column direction of the fluorescence image in the original fluorescence image sequence is odd, the maximum number of pixels in the row direction and / or column direction of the fluorescence image is ensured to be even by deleting the corresponding number of pixels in the rows or columns.

[0011] When selectively training a denoising neural network based on the training set, in each training cycle, one pair of training super-resolution images is randomly selected from the training set, and pixel blocks at the same location in the two super-resolution sub-images are randomly selected, rotated and folded, and used as the input image and target image of the denoising neural network, respectively. The error between the network output and the target image is calculated, and the gradient is backpropagated to update the network parameters.

[0012] Selectively, the denoising neural network includes, but is not limited to, a U-type neural network model, a residual neural network model, a residual channel attention convolutional neural network model, or a Fourier channel attention convolutional neural network model.

[0013] Selectively, a biological sample is excited using structured light by an optical imaging system (100), and j original fluorescence image sequences (Y) generated by the excitation of the biological sample are obtained, the optical imaging system (100) includes, but is not limited to, a two-dimensional structured light system (2D-SIM), a three-dimensional structured light system (3D-SIM), a lattice light sheet structured light system (LLS-SIM), and a grazing incidence illumination structured light system (GI-SIM).

[0014] Selectively, the 2x pixel upsampling process is implemented using nearest neighbor interpolation, bilinear interpolation, or biquadratic spline interpolation.

[0015] According to another aspect of this application, a self-supervised multimodal structured optical microscope reconstruction system, A super-resolution reconstruction module (210) is configured to perform super-resolution reconstruction on an image using a standard structured optical super-resolution reconstruction algorithm, A noise reduction module (220) is provided with a noise reduction neural network, and the following is included: The self-supervised multimodal structured optical microscope reconstruction system is The system is configured to excite a biological sample based on structured light and to obtain J original fluorescence image sequences (Y) generated by the excitation of the biological sample, where J is an integer greater than or equal to 1, each original fluorescence image sequence (Y) contains S fluorescence images, where S is an integer greater than or equal to 2, and each fluorescence image has a pixel size of M*N, where M and N are even numbers. The aforementioned self-supervised multimodal structured light microscope reconstruction system further includes: For each of the J original fluorescence image sequences (Y) in the aforementioned J original fluorescence image sequences (Y), 1) For the i-th fluorescence image (yi) among the S fluorescence images, extract the first sub-image, the second sub-image, the third sub-image, and the fourth sub-image with a pixel size of (M / 2) * (N / 2) by means of pixel rearrangement. The two-dimensional pixel matrix yi,(2m - 1)(2n - 1) of the first sub-image, the two-dimensional pixel matrix yi,(2m - 1)(2n) of the second sub-image, the two-dimensional pixel matrix yi,(2m)(2n - 1) of the third sub-image, and the two-dimensional pixel matrix yi,(2m)(2n) of the fourth sub-image are respectively selected from the two-dimensional pixel matrix y i ) of the i-th fluorescence image (y i,MN . m is an integer and m = 1, 2, 3,..., M / 2, n is an integer and n = 1, 2, 3,..., N / 2, and 2) Perform a 2-fold pixel upsampling process on the first sub-image, the second sub-image, the third sub-image, and the fourth sub-image to obtain the first upsampled sub-image, the second upsampled sub-image, the third upsampled sub-image, and the fourth upsampled sub-image with a pixel size of M * N. 3) Perform a pixel translation of [0.5, 0.5] on the two-dimensional pixel matrix of the first upsampled sub-image to obtain the first sub-image after two-dimensional sub-pixel translation. Perform a pixel translation of [-0.5, 0.5] on the two-dimensional pixel matrix of the second upsampled sub-image to obtain the second sub-image after two-dimensional sub-pixel translation. Perform a pixel translation of [0.5, -0.5] on the two-dimensional pixel matrix of the third upsampled sub-image to obtain the third sub-image after two-dimensional sub-pixel translation. Perform a pixel translation of [-0.5, -0.5] on the two-dimensional pixel matrix of the fourth upsampled sub-image to obtain the fourth sub-image after two-dimensional sub-pixel translation. 4) The S sub-images obtained from the S fluorescence images after 2D sub-pixel translation are made into a first sub-image group, the S sub-images obtained from the S fluorescence images after 2D sub-pixel translation are made into a second sub-image group, the S sub-images obtained from the S fluorescence images after 2D sub-pixel translation are made into a third sub-image group, and the S sub-images obtained from the S fluorescence images after 2D sub-pixel translation are made into a fourth sub-image group. 5) Using a standard structured light reconstruction algorithm, structured light image reconstruction is performed on the first subimage group, the second subimage group, the third subimage group, and the fourth subimage group to obtain the first super-resolution subimage, the second super-resolution subimage, the third super-resolution subimage, and the fourth super-resolution subimage. 6) The system is configured to select two different super-resolution subimages from the first super-resolution subimage, the second super-resolution subimage, the third super-resolution subimage, and the fourth super-resolution subimage to form a training super-resolution image pair and add it to the training set, The aforementioned self-supervised multimodal structured light microscope reconstruction system further includes: 7) Training a denoising neural network based on the aforementioned training set, 8) A self-supervised multimodal structured light microscope reconstruction system is provided, which is configured to perform super-resolution reconstruction on S fluorescence images in each of the J original fluorescence image sequences (Y) using a standard structured light super-resolution reconstruction algorithm to form a super-resolution image, and to use the super-resolution image as input to the denoising neural network to obtain a final super-resolution reconstructed image.

[0016] When selectively training a denoising neural network, one of the super-resolution subimages from the training super-resolution image pair is used as training input data, and the other width super-resolution subimage is used as training target data.

[0017] Selectively, if the maximum number of pixels in the row direction and / or column direction of the fluorescence image in the original fluorescence image sequence is odd, the maximum number of pixels in the row direction and / or column direction of the fluorescence image is ensured to be even by deleting the corresponding number of pixels in the rows or columns.

[0018] When selectively training a denoising neural network based on the training set, in each training cycle, one pair of training super-resolution images is randomly selected from the training set, and pixel blocks at the same location in the two super-resolution sub-images are randomly selected, rotated and folded, and used as the input image and target image of the denoising neural network, respectively. The error between the network output and the target image is calculated, and the gradient is backpropagated to update the network parameters.

[0019] The noise reduction neural network includes, but is not limited to, a U-type neural network model, a residual neural network model, a residual channel attention convolutional neural network model, or a Fourier channel attention convolutional neural network model.

[0020] An optical imaging system (100) is used to excite a biological sample using structured light, and j original fluorescence image sequences (Y) generated by the excitation of the biological sample are obtained. The optical imaging system (100) includes, but is not limited to, a two-dimensional structured light system (2D-SIM), a three-dimensional structured light system (3D-SIM), a lattice light sheet structured light system (LLS-SIM), and a grazing incidence illumination structured light system (GI-SIM).

[0021] The aforementioned 2x pixel upsampling process is achieved using nearest neighbor interpolation, bilinear interpolation, or biquadratic spline interpolation.

[0022] By employing the technical means of this application, particularly in the case of long-term video data of living cells, it is possible to train a neural network by constructing a training set from the original data itself and to denoise the data itself, without requiring any extra data or pre-trained models. This effectively reduces experimental costs and improves the versatility of the technology.

[0023] Furthermore, by employing the technical means of this application, it becomes possible to use extremely low excitation light power and extremely short exposure times in the fluorescence imaging process of cells, especially living cells, significantly reducing phototoxicity, enabling the capture of rapid dynamic processes, and improving the applicability of living cell imaging. In summary, this has great significance for improving the quality of structured light microscopy images and expanding the applications of structured light microscopy technology. [Brief explanation of the drawing]

[0024] The principles and embodiments of this application can be more comprehensively understood by referring to the following detailed description and drawings. The scale of the drawings may be changed for clarity but will not affect the understanding of this application.

[0025] [Figure 1] A schematic diagram of the basic block structure of a structured light microscopy imaging system is shown. [Figure 2A] This invention schematically illustrates the process of training the neural network of the denoising module of the self-supervised multimodal structured optical microscope reconstruction system according to one embodiment of this application. [Figure 2B] This invention schematically illustrates the process of reconstructing an original fluorescence image sequence using a neural network-trained, self-supervised, multimodal structured optical microscope reconstruction system according to one embodiment of this application. [Figure 3] A schematic diagram of a two-dimensional pixel matrix representing a fluorescence image is shown. [Figure 4A]A schematic diagram of the two-dimensional pixel matrix representing the subimage extracted from the two-dimensional pixel matrix in Figure 3 using "pixel rearrangement" is schematically shown. [Figure 4B] A schematic diagram of the two-dimensional pixel matrix representing the subimage extracted from the two-dimensional pixel matrix in Figure 3 using "pixel rearrangement" is schematically shown. [Figure 4C] A schematic diagram of the two-dimensional pixel matrix representing the subimage extracted from the two-dimensional pixel matrix in Figure 3 using "pixel rearrangement" is schematically shown. [Figure 4D] A schematic diagram of the two-dimensional pixel matrix representing the subimage extracted from the two-dimensional pixel matrix in Figure 3 using "pixel rearrangement" is schematically shown. [Figure 5] A schematic diagram illustrating the processing of the 2D pixel matrix of the sub-image in Figure 4A using "2x pixel upsampling" is shown. [Figure 6] A schematic flowchart of a self-supervised multimodal structured optical microscope reconstruction method according to one embodiment of this application is shown. [Figure 7A] Schematic diagrams illustrating exemplary microscopic image processing for the training and reconstruction processes related to this application are shown. [Figure 7B] Schematic diagrams illustrating exemplary microscopic image processing for the training and reconstruction processes related to this application are shown. [Modes for carrying out the invention]

[0026] In the drawings of this application, features having the same structure or similar function are indicated by the same reference numerals.

[0027] Figure 1 schematically shows a basic block diagram of a structured light microscope imaging system, which basically includes an optical imaging system 100 and a control and data processing system 200. The optical imaging system 100 includes an excitation optical path and a detection optical path. The excitation optical path includes an excitation objective lens and other optical assemblies for generating excitation light, and the excitation light beam is emitted through the excitation objective lens in the manner of structured light with periodic fringes to excite fluorescence on a biological sample. The detection optical path includes a detection objective lens and other optical assemblies for imaging, and is used to receive and detect the excited fluorescence. As will be seen by those skilled in the art, depending on the configuration of the structured light microscope imaging system, the excitation objective lens and the detection objective lens may be the same objective lens or may be different objective lenses. When performing three-dimensional fluorescence microscopy imaging on a biological sample, especially a live biological sample, a multilayer fluorescence image is sampled by continuously scanning along the axial direction which is the optical axis of the detection objective lens. Thus, the multilayer fluorescence images acquired after completing each scan and sampling constitute a single fluorescence image stack (also called a "sequence").

[0028] The control and data processing system 200 mainly comprises a computer and related components (e.g., data storage) and can control the operation of the optical imaging system 100 and receive image data from the optical imaging system 100 and perform corresponding post-processing. For example, an acquired fluorescence image stack is provided to the control and data processing system 200, and after a series of data processing, a three-dimensional microscopic image with a high signal-to-noise ratio is reconstructed. For this purpose, the control and data processing system 200 may include a self-supervised multimodal structured optical microscope image reconstruction module or system (which may be referred to as the "self-supervised multimodal structured optical microscope image reconstruction module or system"). The self-supervised multimodal structured optical microscope image reconstruction module or system includes a super-resolution reconstruction submodule 210 and a denoising submodule 220. Within the scope of this application, the modules and / or submodules described herein may be understood to include data storage such as computer-readable storage media capable of storing programs or subprograms and denoising neural network models that can be invoked and executed by a computer, particularly the computer of the control and data processing system 200. When these programs or subprograms and denoising neural network models are invoked and executed by a computer, the methods / steps described below, in particular the self-supervised multimodal structured optical microscope image reconstruction methods / steps, can be realized. Specific programming methods for the programs and / or subprograms are omitted from this application. Those skilled in the art can implement the relevant functions using any well-known programming software and / or commercially available software. Therefore, when this application describes the operation or methods of the relevant systems or modules below, it should be understood that they can also be programmed as programs invoked and executed by a computer.

[0029] The super-resolution reconstruction submodule 210 can perform super-resolution reconstruction on fluorescence images acquired by the optical imaging system 100 using a standard structured light super-resolution reconstruction algorithm. Within the scope of this application, the standard structured light super-resolution reconstruction algorithm can be considered an algorithm already known in the field of microscopy imaging. For example, a standard structured light super-resolution reconstruction algorithm can be found in the disclosure by Gustafsson, MG et al., Three-dimensional resolution doubling in wide-field fluorescence microscopy by structured illumination. Biophys J 94, 4957-4970 (2008).

[0030] The denoising submodule 220 can use any neural network architecture implemented in a manner known to those skilled in the art for image denoising. For example, the neural network models used in the denoising submodule 220 include, but are not limited to, U-type neural network models, residual neural network models, residual channel attention convolutional neural network models, or Fourier channel attention convolutional neural network models. When training the neural network of the training denoising submodule 220, the relevant network model is optimized using a loss function, which includes, but is not limited to, mean squared error (MSE), mean absolute error (MAE), structural similarity (SSIM), or their weights and sums.

[0031] Therefore, the super-resolution reconstruction submodule 210 (or referred to as the "super-resolution reconstruction module") and the denoising submodule 220 (or referred to as the "denoising module") constitute the self-supervised multimodal structured optical microscope reconstruction system according to this application. Figure 2A schematically shows the process of training the neural network of the denoising module 220 of the self-supervised multimodal structured optical microscope reconstruction system using the original fluorescence image sequence Y acquired by the optical imaging system 100 according to one embodiment of this application. Figure 2B schematically shows the process of reconstructing the original fluorescence image sequence Y acquired by the optical imaging system 100 using the self-supervised multimodal structured optical microscope reconstruction system with a neural network trained according to one embodiment of this application.

[0032] The self-supervised multimodal structured light microscope reconstruction system according to this application is applicable to an optical imaging system 100, which may include, but is not limited to, a two-dimensional structured light system (2D-SIM), a three-dimensional structured light system (3D-SIM), a lattice light sheet structured light system (LLS-SIM), and a grazing incidence illumination structured light system (GI-SIM). Therefore, the term "multimodal structured light microscope reconstruction" means that this "structured light microscope reconstruction" is applicable to various optical imaging systems. Furthermore, to overcome the various shortcomings mentioned in the background art section, the multimodal structured light microscope reconstruction system according to this application employs a unique self-supervised approach to train a neural network. The basic principle for training the neural network of the denoising module 220 of the self-supervised multimodal structured light microscope reconstruction system according to this application will be described below with reference to Figure 2A.

[0033] First, for example, a two-dimensional structured light system is used as an example of the optical imaging system 100 to perform a fluorescence image scan on a biological sample and obtain the original image sequence Y. For example, the original image sequence Y collected by the optical imaging system 100 is a series of noisy fluorescence images y under a series of different illumination modes. i It can be expressed as (i=1, 2, ..., S), where S represents the number of illumination modes of the optical imaging system 100, and S is an integer greater than or equal to 2.

[0034] As those skilled in the art will see, each fluorescence image can be represented as a two-dimensional pixel matrix that can be read and processed by the computer of the control and data processing system 200. For example, fluorescence image y i is, y i,MN It can be expressed as y, where M is the number of pixels in the row of the fluorescence image and N is the number of pixels in the column of the fluorescence image. In the case of the fluorescence image relating to this application, M and N may be the same or different, and both are even numbers. For example, the fluorescence image is y i,512*512 This may also be the case, indicating that the size of the i-th image y is 512*512 pixels.

[0035] Of course, in the case of the original fluorescence image where M or N is an odd number, a person skilled in the art should recognize that, in order to apply the proposed method of this application, one row or one column of pixels in the fluorescence image (e.g., at the boundary) may be appropriately removed to ensure that the maximum number of pixels in both the row and column directions of the pixel size of the original fluorescence image used in the proposed method of this application is even.

[0036] Figure 3 shows the i-th fluorescence image y of the original image sequence Y. i A 2D pixel matrix y representing i,MNThis is schematically shown. According to the embodiment of this application, a training set used to train the neural network of the denoising module 220 is obtained by a method of "rearranging pixels" on the two-dimensional pixel matrix of all fluorescence images in the original image sequence Y. Below, using Figure 3 as an example, the i-th fluorescence image y i The method of "pixel rearrangement" for this will be explained. As those skilled in the art will see, the same method can be applied to other fluorescence images of the original image sequence Y. (1) The fluorescence image y is obtained by a method that performs 2x pixel downsampling. i 2D pixel matrix y i,MN Four sub-images y i,(2m-1)(2n-1) , y i、(2m-1)(2n) , y i,(2m)(2n-1) , y i,(2m)(2n) Extract and generate such that m is an integer and m=1, 2, 3, ..., M / 2, and n is an integer and n=1, 2, 3, ..., N / 2. Therefore, the 2D pixel matrix y i,MN If the pixel size of is M*N, then the four sub-images y i,(2m-1)(2n-1) , y i,(2m-1)(2n) , y i,(2m)(2n-1) , y i,(2m)(2n) The pixel size is (M / 2)*(N / 2) in all cases, as shown in Figures 4A-4D. In the context of this application, when referring to a fluorescence image, a two-dimensional pixel matrix of a fluorescence image, a subimage, or a two-dimensional pixel matrix of a subimage, the pixels of the image are counted starting with the leftmost pixel in the row direction as 1 and the topmost pixel in the column direction as 1. (2) Four sub-images y i,(2m-1)(2n-1) , y i,(2m-1)(2n) , y i,(2m)(2n-1) , y i,(2m)(2n) The image ys1' is upsampled to four sub-images ys1' by applying a 2x pixel upsampling process to the image, resulting in a pixel size of M*N. i,MN , ys2' i,MN , ys3' i,MN , ys4' i,MN Obtain one of the subimages y'. i,(2m-1)(2n-1) Upsampled subimage ys1'i,MN Only this is schematically shown in Figure 5. As those skilled in the art will know, when performing a 2x pixel upsampling operation on an image, available methods include, but are not limited to, known computer image processing methods such as nearest neighbor interpolation, bilinear interpolation, and biquadratic spline interpolation. Therefore, a detailed description of these methods is omitted in this specification. (3) Four upsampled subimages ys1' i,MN , ys2' i,MN , ys3' i,MN , ys4' i,MN The pixel sizes are all from the original 2D pixel matrix y i,MN Since it is the same, considering the differences in the previous downsampling method, each sub-image ys1' i,MN , ys2' i,MN , ys3' i,MN , ys4' i,MN The center of the image is the original 2D pixel matrix y i,MN It has a different offset. Therefore, in order to eliminate the effect of the offset between the image center of such sub-images and the image center of the original image, each sub-image ys1' i,MN , ys2' i,MN , ys3' i,MN , ys4' i,MN Sub-pixel alignment is performed for each, i.e., sub-image ys1' i,MN For this, perform a 2D subpixel translation at [0.5,0.5] to create the subimage ys1 i,MN Obtain the subimage ys2'. i,MN For this, perform a 2D subpixel translation at [-0.5,0.5] to create the subimage ys2 i,MN Obtain subimage ys3'. i,MN For this, perform a 2D subpixel translation of [0.5,-0.5] to create the subimage ys3. i,MN Obtain subimage ys4'. i,MN For this, perform a 2D subpixel translation of [-0.5,-0.5] to create the subimage ys4. i,MN Obtain it.

[0037] Within the scope of this application, when performing image translation (subpixel alignment), translation to the right in the row direction is positive, translation to the left is negative, translation downward in the column direction is positive, and translation upward is negative. For example, [0.5,0.5] should be understood as translating the subimage to be processed by half a pixel to the right in the row direction and by half a pixel downward in the column direction. [-0.5,-0.5] should be understood as translating the subimage to be processed by half a pixel to the left in the row direction and by half a pixel upward in the column direction. The image translation processing according to the proposed technology of this application can be implemented using any known software, program, or instruction in computer image processing technology, so a detailed explanation is omitted here.

[0038] As those skilled in the art will know, in a process of performing two-dimensional subpixel translation on a subimage, invalid pixel values ​​(e.g., pixel values ​​cleared to zero by the translation) exist in the corresponding boundary rows and columns of the subimage. As shown in Figure 5, subimage ys1' i,MN As an example, after performing a 2D subpixel translation of [0.5,0.5], the translated subimage ys1 i,MN The leftmost column and top row of pixels may contain null values ​​(pixel values ​​cleared to zero), as indicated by "x" in the figure. However, since null values ​​at the boundaries do not affect the subsequent overall image processing and recognition results, the effects of such null values ​​may be ignored in the proposed art of this application.

[0039] (4) For each fluorescence image in the original image sequence Y, repeat the steps in (1) to (3) above, then create the subimage group ys1 i,MN , ys2 i,MN ys3 i,MN ys4 i,MN We can obtain the following, where i=1, 2, ..., S. Then, using a standard structured light reconstruction algorithm, we can reconstruct the subimage group ys1 i,MN , ys2 i,MN ys3i,MN ys4 i,MN Structured light reconstruction is performed on the image to obtain four super-resolution subimages Y1, Y2, Y3, and Y4. As mentioned above, a standard structured light super-resolution reconstruction algorithm can be found in the disclosure by Gustafsson, MG et al., Three-dimensional resolution doubling in wide-field fluorescence microscopy by structured illumination. Biophys J 94, 4957-4970 (2008) or other known reconstruction algorithms.

[0040] (5) If four super-resolution sub-images Y1, Y2, Y3, and Y4 are obtained from the original image sequence Y, two different super-resolution sub-images are selected to form a training super-resolution image pair. For example, for one original image sequence, there are 12 possible combinations of training super-resolution image pairs, namely (Y1,Y2), (Y1,Y3), (Y1,Y4), (Y2,Y3), (Y2,Y4), (Y3,Y4), (Y2,Y1), (Y3,Y1), (Y4,Y1), (Y3,Y2), (Y4,Y2), and (Y4,Y3). In each training super-resolution image pair consisting of two super-resolution images, the first super-resolution image is used as training input data, and the second super-resolution image is used as training target data. In other words, assuming that the order of the two super-resolution images in a training super-resolution image pair is ignored, it is also possible to use one super-resolution image from a training super-resolution image pair as training input data and the other super-resolution image as training target data.

[0041] When J original fluorescence image sequences are acquired from living cells using the optical imaging system 100 (where J is an integer greater than or equal to 1), 12 combinations of training super-resolution image pairs can be obtained for each of the J original fluorescence image sequences (for example, each original fluorescence image sequence consists of S fluorescence images) using the method described above. Ultimately, J*12 training super-resolution image pairs are obtained for all J original fluorescence image sequences. These training super-resolution image pairs constitute a training set for training the neural network of the denoising module 220. The neural network of the denoising module 220 can be trained using any one of the training super-resolution image pairs selected from the training set. For example, in each training cycle, one training super-resolution image pair is randomly selected from the training set, pixel blocks at the same location are randomly extracted and rotated, folded, and then used as the input image and target image of the neural network, respectively. The error between the network output and the target image is calculated, and the gradient is backpropagated to update the network parameters. After the network's input error converges, training is stopped and the network parameters are stored. As those skilled in the art will see, the neural network training method is not limited to those listed herein. It will be seen that the neural network training described in this application is indeed self-supervised training.

[0042] As those skilled in the art will see, when J=1, the optical imaging system 100 can be considered to be performing static observation of the biological sample. In this case, the optical imaging system 100 acquires one static fluorescence imaging sequence consisting of multiple original fluorescence images of the biological sample. When J≧2, the optical imaging system 100 can be considered to be performing dynamic observation of the biological sample (for example, continuous video recording). In this case, multiple fluorescence imaging sequences are acquired, and each fluorescence imaging sequence contains multiple original fluorescence images.

[0043] Once the neural network training of the denoising module 220 is complete, the self-supervised multimodal structured optical microscopy imaging system reconstructs (or predicts) the original fluorescence image sequence acquired by the optical imaging system 100, as shown in Figure 2B. For the noisy fluorescence image in each original fluorescence image sequence, the super-resolution reconstruction submodule 210 performs super-resolution reconstruction to acquire a super-resolution image. The acquired super-resolution image is then used as input to the neural network of the denoising module 220 to obtain the final denoised super-resolution image.

[0044] Figure 6 schematically shows one embodiment of the self-supervised multimodal structured optical microscope reconstruction method according to this application. Assume that J original fluorescence image sequences Y of a biological sample or living cell are obtained by the optical imaging system 100, where J is an integer greater than or equal to 1. The j-th original fluorescence image sequence Y j (j=1, 2, ..., J) contains the fluorescence image y i (i=1, 2, ..., S) is included, S is an integer greater than or equal to 2, and each fluorescence image y i The pixel size is M*N, where M and N are integers greater than 2 and are even.

[0045] In step S10, starting from j=1, the j-th original fluorescence image sequence Y j Fluorescence image y iFor example, starting from i = 1, using a computer, the fluorescent image y is generated by 2x pixel downsampling i from the two-dimensional pixel matrix y i,MN into four sub-images y i,(2m-1)(2n-1) y i,(2m-1)(2n) y i,(2m)(2n-1) y i,(2m)(2n) where m is an integer and m = 1, 2, 3, …, M / 2, and n is an integer and n = 1, 2, 3, …, N / 2.

[0046] In step S20, using a computer, 2x pixel upsampling processing is performed on the four sub-images y i,(2m-1)(2n-1) y i,(2m-1)(2n) y i,(2m)(2n-1) y i,(2m)(2n) to obtain four upsampled sub-images ys1’ i,MN ys2’ i,MN ys3’ i,MN ys4’ i,MN with a pixel size reaching M*N. When performing 2x pixel upsampling processing on an image, the methods that can be adopted include, but are not limited to, any known method for processing images on a computer, such as nearest neighbor interpolation, bilinear interpolation, bicubic spline interpolation, etc.

[0047] In step S30, using a computer, sub-pixel alignment is performed on each sub-image ys1’ i,MN ys2’ i,MN ys3’ i,MN ys4’ i,MN respectively, that is, for the sub-image ys1’ i,MN a two-dimensional sub-pixel translation of [0.5, 0.5] is performed to obtain the sub-image ys1 i,MN For the sub-image ys2’ i,MN a two-dimensional sub-pixel translation of [-0.5, 0.5] is performed to obtain the sub-image ys2 i,MN For the sub-image ys3’ i,MN a two-dimensional sub-pixel translation of [0.5, -0.5] is performed to obtain the sub-image ys3 i,MNObtain subimage ys4'. i,MN For this, perform a 2D subpixel translation of [-0.5,-0.5] to create the subimage ys4. i,MN Obtain it.

[0048] In step S31, the computer is used to return the original image sequence Y j For each fluorescence image in the system, it is determined whether steps S10 to S30 above have been repeated. For example, if i ≠ S, then i = i + 1, and steps S10 to S30 are repeated. If i = S, the process proceeds to step S32.

[0049] In step S32, the original image sequence Y j Steps S10 to S30 above were repeated for each fluorescence image in ys1, resulting in four sub-image groups ys1 i,MN , ys2 i,MN ys3 i,MN ys4 i,MN We can obtain the following, where i=1, 2, ..., S. Using the super-resolution reconstruction module 210, each sub-image group ys1 i,MN , ys2 i,MN ys3 i,MN ys4 i,MN Super-resolution reconstruction is performed on the image to obtain four super-resolution sub-images Y1, Y2, Y3, and Y4.

[0050] In step S33, two different super-resolution subimages are randomly selected from the four super-resolution subimages Y1, Y2, Y3, and Y4 acquired in step S32 to form a training super-resolution image pair. The first super-resolution image in the training super-resolution image pair is used as training input data, and the second super-resolution image is used as training target data. For example, since four super-resolution subimages can be acquired for one original fluorescence image sequence, there are 12 possible combinations of training super-resolution image pairs. These training super-resolution image pairs can then be used as part of the training set used to train the neural network.

[0051] In step S40, it is determined whether steps S10 to S33 above were performed for all original fluorescence image sequences. For example, it can be determined whether j has reached the maximum value J. If NO, j = j + 1, and steps S10 to S32 are repeated. If YES, proceed to step S50.

[0052] In step S50, the neural network of the denoising module 220 is trained using the acquired training set. As those skilled in the art will see, training can be performed in any suitable manner. For example, in each training cycle, any number of training super-resolution image pairs are randomly selected from the training set, and then two super-resolution images from the same position in the training super-resolution image pair are randomly extracted, rotated randomly, folded, and used as the input image and target image of the neural network, respectively. The error between the network output and the target image is calculated, and the gradient is backpropagated to update the network parameters. After the network input error converges, training is stopped and the network parameters are stored.

[0053] In step S60, starting from j=1, the j-th original fluorescence image sequence Y j Fluorescence image y j For (j=1, 2, 3, ..., S), super-resolution reconstruction is performed using the super-resolution reconstruction module 210 to obtain a super-resolution image.

[0054] In step S70, the super-resolution image acquired in step S70 is used as input to the neural network of the noise reduction module 220, and the acquired output image becomes the final noise-reduced super-resolution image.

[0055] In step S80, it is determined whether steps S60 to S70 above were performed for each of the original fluorescence sequences. For example, it can be determined whether j has reached the maximum value J. If NO, j = j + 1, and steps S60 and S70 are repeated. If YES, proceed to step S100. In step S100, the self-supervised multimodal structured photolimiting reconstruction method is terminated.

[0056] According to this application, as an example, a typical neural network that can be used as the neural network of the noise reduction module 220 is a U-net, and the reference values ​​of its main parameters are shown in Table 1. It should be noted that the neural network usable in this application does not need to be a specific neural network structure; typical network structures such as residual networks (ResNet), residual channel attention networks (RCAN), residual dense neural networks (RDN), and self-attention networks (Transformer) can all achieve the aforementioned functions.

[0057] [Table 1]

[0058] Furthermore, the main parameters used in the neural network training process of this application need to be adjusted according to the specific circumstances of the dataset (signal-to-noise ratio, structural complexity, etc.). For reference, the parameters for one group are shown in Table 2.

[0059] [Table 2]

[0060] Figures 7A and 7B schematically illustrate the process of training and reconstructing using the original fluorescence image sequence with the method described above. Figure 7A shows the process of training the neural network, and Figure 7B shows the process of reconstructing the original fluorescence image sequence using the neural network-trained reconstruction system.

[0061] This application employs a unique "pixel rearrangement" method combined with super-resolution reconstruction to obtain a training set for a denoising module of a neural network. The main advantages are as follows: 1. When training a denoising neural network using the technology of this application, a sufficient training set for training the denoising neural network can be obtained based on the premise of an original fluorescence image sequence with an extremely low signal-to-noise ratio. This eliminates the need to sample the same biological sample multiple times and the need to irradiate the same biological sample with high-intensity excitation light to obtain an original fluorescence image with a high signal-to-noise ratio, thereby reducing the possibility of damage to the biological sample. 2. This application enables the division of an original fluorescence image into four sub-images using a unique "pixel rearrangement" technique, and then obtains training super-resolution image pairs through upsampling, sub-pixel alignment, and super-resolution reconstruction, thereby realizing a complete self-supervised neural network training mechanism. 3. The proposed technology of this application can be directly applied to long-duration video data processing. In this application, a training set can be directly constructed from video data to train a neural network of a denoising module, and then denoising can be performed on super-resolution microscopic images using the trained denoising module.

[0062] This application provides a self-supervised structured light microscopy reconstruction method and system based on pixel rearrangement, which can effectively denoise structured light microscopy data from original images with extremely low signal-to-noise ratios and reconstruct noise-free super-resolution images with high fidelity. By applying this method, the signal-to-noise ratio requirements for the original image in the structured light microscopy imaging process are significantly reduced, photodamage to biological samples in the imaging process is reduced, imaging speed is improved, the image quality of the structured light microscopy is significantly improved, and its range of applications is expanded.

[0063] While specific embodiments of this application are described in detail herein, these are for interpretation purposes only and should not be considered to limit the scope of this application. Furthermore, as will be apparent to those skilled in the art, the various embodiments described herein can be used in combination with one another. Various substitutions, modifications, and improvements can be made without departing from the spirit and scope of this application.

Claims

1. A self-supervised multimodal structured light microscope reconstruction method, The steps include: exciting a biological sample using structured light and obtaining J original fluorescence image sequences Y generated by the excitation of the biological sample, wherein J is an integer of 1 or more, each original fluorescence image sequence Y contains S fluorescence images, S is an integer of 2 or more, and each fluorescence image has a pixel size of M*N, where M and N are even numbers; For each of the J original fluorescence image sequences Y, a program is executed on a computer, 1) The i-th fluorescence image y among the S fluorescence images i By rearranging the pixels, the first, second, third, and fourth sub-images with a pixel size of (M / 2) * (N / 2) are extracted, and the two-dimensional pixel matrix y of the first sub-image is obtained. i,(2m-1)(2n-1) , the two-dimensional pixel matrix y of the second sub-image i,(2m-1)(2n) , the two-dimensional pixel matrix y of the third sub-image i,(2m)(2n-1) , and the two-dimensional pixel matrix y of the fourth sub-image i,(2m)(2n) These are the i-th fluorescence image y, respectively. i 2D pixel matrix y i,MN A step in which a selection is made from, m is an integer and m = 1, 2, 3, ..., M / 2, and n is an integer and n = 1, 2, 3, ..., N / 2, 2) The first sub-image, the second sub-image, the third sub-image, and the fourth sub-image are subjected to a 2x pixel upsampling process to obtain a first upsampled sub-image, a second upsampled sub-image, a third upsampled sub-image, and a fourth upsampled sub-image with a pixel size of M*N; 3) The steps of obtaining the first sub-image after two-dimensional sub-pixel translation by performing a pixel translation of [0.5, 0.5] on the two-dimensional pixel matrix of the first upsampled sub-image, performing a pixel translation of [-0.5, 0.5] on the two-dimensional pixel matrix of the second upsampled sub-image to obtain the second sub-image after two-dimensional sub-pixel translation, performing a pixel translation of [0.5, -0.5] on the two-dimensional pixel matrix of the third upsampled sub-image to obtain the third sub-image after two-dimensional sub-pixel translation, and performing a pixel translation of [-0.5, -0.5] on the two-dimensional pixel matrix of the fourth upsampled sub-image to obtain the fourth sub-image after two-dimensional sub-pixel translation, 4) The steps of forming a first sub-image group from the S fluorescent images obtained by translating S two-dimensional subpixels, a second sub-image group from the S fluorescent images obtained by translating S two-dimensional subpixels, a third sub-image group from the S fluorescent images obtained by translating S two-dimensional subpixels, and a fourth sub-image group from the S fluorescent images obtained by translating S two-dimensional subpixels, 5) Using a standard structured light reconstruction algorithm, structured light reconstruction is performed on the first subimage group, the second subimage group, the third subimage group, and the fourth subimage group to obtain the first super-resolution subimage, the second super-resolution subimage, the third super-resolution subimage, and the fourth super-resolution subimage; 6) The steps of selecting two different super-resolution subimages from the first super-resolution subimage, the second super-resolution subimage, the third super-resolution subimage, and the fourth super-resolution subimage to form a training super-resolution image pair and adding it to the training set are performed. When a program is executed on a computer, 7) A step of training a denoising neural network based on the training set, 8) Using a standard structured optical super-resolution reconstruction algorithm, super-resolution reconstruction is performed on S fluorescence images in each of the J original fluorescence image sequences Y to form a super-resolution image, and this super-resolution image is used as input to the denoising neural network to obtain the final super-resolution reconstructed image. A self-supervised multimodal structured optical microscope reconstruction method characterized by the following:

2. When training a denoising neural network, one super-resolution subimage of the training super-resolution image pair is used as training input data, and the other width super-resolution subimage is used as training target data. The self-supervised multimodal structured optical microscope reconstruction method according to feature 1.

3. If the maximum number of pixels in the row direction and / or column direction of the fluorescence image in the original fluorescence image sequence is odd, ensure that the maximum number of pixels in the row direction and / or column direction of the fluorescence image becomes even by deleting the corresponding number of pixels in the rows or columns. The self-supervised multimodal structured optical microscope reconstruction method according to feature 2.

4. When training a denoising neural network based on the aforementioned training set, in each training cycle, one pair of training super-resolution images is randomly selected from the training set, and pixel blocks at the same position in the two super-resolution sub-images are randomly selected, rotated and folded randomly, and then used as the input image and target image of the denoising neural network, respectively. The error between the network output and the target image is calculated, and the network parameters are updated so that the error converges. A self-supervised multimodal structured optical microscope reconstruction method according to any one of claims 1 to 3.

5. The noise reduction neural network includes, but is not limited to, a U-shaped neural network model, a residual neural network model, a residual channel attention convolutional neural network model, or a Fourier channel attention convolutional neural network model. The self-supervised multimodal structured optical microscope reconstruction method according to feature 4.

6. An optical imaging system (100) excites a biological sample using structured light, and acquires j original fluorescence image sequences Y generated by the excitation of the biological sample. The optical imaging system (100) includes, but is not limited to, a two-dimensional structured light system, a three-dimensional structured light system, a lattice light sheet structured light system, and a grazing incidence illumination structured light system. The self-supervised multimodal structured optical microscope reconstruction method according to feature 5.

7. The aforementioned 2x pixel upsampling process is implemented using nearest neighbor interpolation, bilinear interpolation, or biquadratic spline interpolation. The self-supervised multimodal structured optical microscope reconstruction method according to feature 6.

8. A self-supervised multimodal structured light microscope reconstruction system, A super-resolution reconstruction module (210) is configured to perform super-resolution reconstruction on an image using a standard structured optical super-resolution reconstruction algorithm, The system includes a noise reduction module (220) equipped with a noise reduction neural network, The self-supervised multimodal structured optical microscope reconstruction system is The system is configured to excite a biological sample based on structured light and to obtain J original fluorescence image sequences Y generated by the excitation of the biological sample, where J is an integer greater than or equal to 1, each original fluorescence image sequence Y contains S fluorescence images, where S is an integer greater than or equal to 2, and each fluorescence image has a pixel size of M*N, where M and N are even numbers. The aforementioned self-supervised multimodal structured light microscope reconstruction system further includes: For each of the J original fluorescence image sequences Y in the aforementioned J original fluorescence image sequences Y, 1) For the i-th fluorescence image y among the S fluorescence images, a method of performing pixel rearrangement is used to extract a first sub-image, a second sub-image, a third sub-image, and a fourth sub-image with a pixel size of (M / 2)*(N / 2). The two-dimensional pixel matrix y of the first sub-image i , the two-dimensional pixel matrix y of the second sub-image i,(2m-1)(2n-1) , the two-dimensional pixel matrix y of the third sub-image i,(2m-1)(2n) , and the two-dimensional pixel matrix y of the fourth sub-image i,(2m)(2n-1) are each selected from the two-dimensional pixel matrix y of the i-th fluorescence image y i,(2m)(2n) . Here, m is an integer and m = 1, 2, 3,..., M / 2, and n is an integer and n = 1, 2, 3,..., N / 2, and i the two-dimensional pixel matrix y of the i-th fluorescence image y i,MN where m and n are integers such that m = 1, 2, 3,..., M / 2 and n = 1, 2, 3,..., N / 2 2) Perform a 2x pixel upsampling process on the first sub-image, the second sub-image, the third sub-image, and the fourth sub-image to obtain a first upsampled sub-image, a second upsampled sub-image, a third upsampled sub-image, and a fourth upsampled sub-image with a pixel size of M*N, 3) Perform a pixel translation of [0.5, 0.5] on the 2D pixel matrix of the first upsampled subimage to obtain the first subimage after 2D subpixel translation; perform a pixel translation of [-0.5, 0.5] on the 2D pixel matrix of the second upsampled subimage to obtain the second subimage after 2D subpixel translation; perform a pixel translation of [0.5, -0.5] on the 2D pixel matrix of the third upsampled subimage to obtain the third subimage after 2D subpixel translation; and perform a pixel translation of [-0.5, -0.5] on the 2D pixel matrix of the fourth upsampled subimage to obtain the fourth subimage after 2D subpixel translation. 4) The S first sub-images obtained from the S fluorescence images after two-dimensional sub-pixel translation are made into a first sub-image group, the S second sub-images obtained from the S fluorescence images after two-dimensional sub-pixel translation are made into a second sub-image group, the S third sub-images obtained from the S fluorescence images after two-dimensional sub-pixel translation are made into a third sub-image group, and the S fourth sub-images obtained from the S fluorescence images after two-dimensional sub-pixel translation are made into a fourth sub-image group. 5) Using a standard structured light reconstruction algorithm, structured light image reconstruction is performed on the first subimage group, the second subimage group, the third subimage group, and the fourth subimage group to obtain the first super-resolution subimage, the second super-resolution subimage, the third super-resolution subimage, and the fourth super-resolution subimage. 6) The system is configured to select two different super-resolution subimages from the first super-resolution subimage, the second super-resolution subimage, the third super-resolution subimage, and the fourth super-resolution subimage to form a training super-resolution image pair and add it to the training set, The aforementioned self-supervised multimodal structured light microscope reconstruction system further includes: 7) Training a noise reduction neural network based on the aforementioned training set, 8) Using a standard structured optical super-resolution reconstruction algorithm, super-resolution reconstruction is performed on S fluorescence images in each of the J original fluorescence image sequences Y to form a super-resolution image, and this super-resolution image is used as input to the denoising neural network to obtain the final super-resolution reconstructed image. A self-supervised multimodal structured optical microscope reconstruction system characterized by the following features.

9. When training a denoising neural network, one super-resolution subimage of the training super-resolution image pair is used as training input data, and the other width super-resolution subimage is used as training target data. The self-supervised multimodal structured optical microscope reconstruction system according to feature 8.

10. If the maximum number of pixels in the row direction and / or column direction of the fluorescence image in the original fluorescence image sequence is odd, ensure that the maximum number of pixels in the row direction and / or column direction of the fluorescence image becomes even by deleting the corresponding number of pixels in the rows or columns. The self-supervised multimodal structured optical microscope reconstruction system according to feature 9.

11. When training a denoising neural network based on the aforementioned training set, in each training cycle, one pair of training super-resolution images is randomly selected from the training set, and pixel blocks at the same position in the two super-resolution sub-images are randomly selected, rotated and folded randomly, and then used as the input image and target image of the denoising neural network, respectively. The error between the network output and the target image is calculated, and the network parameters are updated so that the error converges. A self-supervised multimodal structured optical microscope reconstruction system according to any one of claims 8 to 10.

12. The noise reduction neural network includes, but is not limited to, a U-shaped neural network model, a residual neural network model, a residual channel attention convolutional neural network model, or a Fourier channel attention convolutional neural network model. The self-supervised multimodal structured optical microscope reconstruction system according to feature 11.

13. An optical imaging system (100) excites a biological sample using structured light, and acquires j original fluorescence image sequences Y generated by the excitation of the biological sample. The optical imaging system (100) includes, but is not limited to, a two-dimensional structured light system, a three-dimensional structured light system, a lattice light sheet structured light system, and a grazing incidence illumination structured light system. The self-supervised multimodal structured optical microscope reconstruction system according to feature 12.

14. The aforementioned 2x pixel upsampling process is implemented using nearest neighbor interpolation, bilinear interpolation, or biquadratic spline interpolation. The self-supervised multimodal structured optical microscope reconstruction system according to feature 13.