Method and system for fluorescence image acquisition of live cell biological samples
By generating high-quality 2D projection images of live cell biological samples through deep learning and convolutional neural networks, the problems of long acquisition time and high phototoxicity in existing technologies are solved, and high-throughput, low-phototoxicity image acquisition and analysis are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-30
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies for acquiring three-dimensional fluorescence images of live cell biological samples suffer from problems such as long acquisition time, high risk of phototoxicity, and dependence on expert users, especially when using the Z-axis stacking method.
By employing deep learning and convolutional neural networks, one or more long-exposure Z-axis swept images are acquired, and a high-quality focused 2D projection image is generated using a trained neural network, thus avoiding the shortcomings of traditional Z-axis stacking methods.
It achieves higher acquisition throughput, reduces phototoxicity risk, generates high-quality images, enables standard image analysis, is applicable to traditional fluorescence microscopes and rotating confocal microscopes, and simplifies data processing.
Smart Images

Figure CN115428037B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. nonprovisional patent application No. 16 / 935,326, filed on July 22, 2020, which is incorporated herein by reference. Background Technology
[0003] This invention relates to the field of methods and systems for acquiring fluorescence images of live cellular biological samples, such as three-dimensional (3D) cell cultures, including organoids and tumor spheroids.
[0004] The use of live cell samples is applied across a wide range of research fields, including immuno-oncology, oncology and metabolism, neuroscience, immunology, infectious diseases, toxicology, stem cells, cardiology, and inflammation. Within these fields, studies have been conducted on cell health and proliferation, cell function, cell motility, and morphology, including investigations into complex immune-tumor cell interactions, synaptic activity, and cancer cell metabolism.
[0005] Fluorescence microscopy is a method for observing the light emission spectrum of samples, including live cell samples. Observing fluorescence from a sample requires an optical system, including an incident light source and an optical filter with an out-of-band rejection parameter. This setup allows researchers to analyze the fluorescence quality of a sample in real time. Fluorescence microscopy, with its imaging capabilities (via various designs of cameras or imagers), has become an integral part of modern live cell research laboratories.
[0006] In fluorescence microscopy, samples are typically processed, stained, or chemically combined with one or more fluorophores. Fluorescent molecules are microscopic molecules that may be proteins, small organic compounds, organic dyes, or synthetic polymers; they absorb light of a specific wavelength and emit light of a longer wavelength. Certain semiconductor metal nanoparticles can also act as fluorophores, emitting light relative to their geometric composition.
[0007] Live-cell biological samples can be microscopically imaged in a variety of ways under a fluorescence microscope, typically at magnifications of 10x or 20x, to assess the sample's growth, metabolism, morphology, or other properties at one or more time points. Microscopic imaging can include fluorescence imaging, in which fluorophores in the sample are excited by light at the fluorophore's excitation wavelength, causing them to emit fluorescence at the fluorophore's emission wavelength. In surface fluorescence imaging, the excitation light is provided through the same objective lens used to collect the emitted light.
[0008] The general trend in biological research has been toward studying the importance of 3D models (such as tumor spheres, organs, and 3D cell cultures) rather than their two-dimensional (2D) counterparts (such as a single image in a given focal plane), because it is believed that 3D models can better reproduce the physiological conditions present in real in vivo systems.
[0009] Standard methods for acquiring three-dimensional information of samples using fluorescence microscopy or confocal microscopy include a stepwise "Z-stacking" method, in which a series of sample images are captured, each located at a specific depth or Z-coordinate of the sample in a three-dimensional coordinate system. This method is as described in this invention. Figure 2C As shown, this will be described in more detail later. The generated series of images can then be manipulated using specialized software to project or combine them into a single 2D image. Timothy Jackson et al., in their U.S. patent application filed April 21, 2020, entitled "Image Processing and Segmentation of Sets of Z-Stacked Images of Three-Dimensional Biological Samples," Serial No. 16 / 854710, proposed a method for creating this Z-projection, which is assigned to the assignee of this invention, the description of which is incorporated herein by reference. An example of an open-source image analysis software package for projecting Z-axis stacked images into 2D images is ImageJ, described in NIH Image to ImageJ: 25 years of image analysis. Nature methods, 9(7), 671-675. Several Z-projection schemes exist, including maximum projection, average projection, Sobel filter-based projection, and wavelet-based projection.
[0010] However, the Z-stack method may have some inherent challenges and problems, especially for live cell samples. First, fine-tuning different acquisition parameters, including exposure time, start Z-position, end Z-position, and Z-step (ΔZ increment between image acquisition positions), requires prior knowledge of the sample, making it suitable only for expert users. Second, acquiring such images can become very slow, limiting throughput and exposing the sample to significant amounts of fluorescence from the light source, potentially leading to photobleaching or phototoxicity, which is highly undesirable in live cell research. Finally, acquiring 3D Z-axis stacks requires advanced software for data visualization and analysis. Summary of the Invention
[0011] This document describes a method and system, including image acquisition strategies and image processing procedures that use deep learning to address the above imaging challenges. In particular, described herein is a method to acquire single focal 2D projections of 3D live cell biological samples in a high-throughput manner, leveraging recent advances in deep learning and convolutional neural networks.
[0012] Rather than using known Z-stack strategies, one or more long-exposure "Z-scan" images are acquired, i.e., one or a series of sequential long-exposure sequential image acquisitions in which a camera in a fluorescence microscope is exposed to a sample while moving the Z focal plane through the sample, thereby integrating the fluorescence intensity of the sample over the Z dimension. The acquisition method is much faster than Z-stacking, thus enabling higher throughput and reducing the risk of sample exposure to excessive fluorescence, thereby avoiding phototoxicity and photobleaching issues. The long-exposure images are then input to a trained neural network that is trained to generate high-quality focal 2D projection images that represent the projection of the 3D sample. With these high-quality 2D projection images, it is possible to obtain biological relevant analysis metrics that describe the fluorescence signal using standard image analysis techniques, such as fluorescent object counts and other fluorescence intensity metrics (e.g., mean intensity, texture, etc.).
[0013] In one particular aspect, a method for generating a focal two-dimensional projection of a fluorescence image of a three-dimensional live cell sample is provided. The method includes the steps of: acquiring one or more long-exposure images of the sample with a camera, the camera thereby integrating the fluorescence intensity from the sample over the Z dimension by moving the focal plane of the camera through the sample along the Z direction, providing the one or more long-exposure images to a neural network model trained from a plurality of training images; and generating a focal 2D projection image using the neural network model.
[0014] There can be several different types of neural network model. In one embodiment, the neural network model is a convolutional neural network, in another embodiment an encoder-decoder based supervised model. The encoder-decoder based supervised model can take the form of a U-net, as described below. Alternatively, the neural network model is trained using an adversarial approach, for example a generative adversarial network (GAN). See Goodfellow et al., “Generative Adversarial Nets,” in Advances in Neural Information Processing Systems 27, Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, Eds. Curran Associates, Inc., 2014, pp. 2672-2680. This adversarial approach means that two neural networks are trained simultaneously: a generator that predicts a high-quality focus projection image, and a discriminator that attempts to distinguish between the predicted focus projection image and a real or actually obtained image. The advantage of this approach is that, since the generator attempts to fool the discriminator, adversarial training can produce more realistic outputs than ordinary supervised training. The conditional GAN is another possibility, in which the generator is conditioned on one or more long-exposure images. The neural network model can also be trained using a cycle-consistency loss approach, so that it can be trained using unpaired data, for example using a GAN or a CycleGAN (GAN with cycle-consistency loss), which is also described in the literature. Unpaired data means that there are Z-sweep images and focus projection images, but not necessarily a one-to-one pixel matching of images. The advantage of the cycle-consistency approach is that it does not require perfect registration between the Z-sweep and focus projection images. For example, there can be imperfect registration when there is a slight movement in the X-Y plane as the camera is moved along the Z axis.
[0015] Three different approaches or paradigms are considered for generating one or more long-exposure Z-sweep images. In one configuration, a single long-exposure Z-sweep image is obtained, referred to below as “paradigm 1”. Alternatively, a series of long-exposure sequential Z-sweep images are obtained, which are then summed, referred to below as “paradigm 2”. As a further alternative, a series of long-exposure sequential Z-sweep images are obtained, which are not summed, referred to below as “paradigm 3”.
[0016] In another aspect, a method of training a neural network model to generate focused 2D projection images of live cell samples is disclosed. The method includes the step of acquiring a training set in the form of a plurality of images. Such images can be paired images or unpaired images. In practice, pairing is beneficial, especially if the neural network model uses supervised training. However, unpaired images can be used, for example using training with cycle consistency loss or generative adversarial networks.
[0017] The training images can include (1) one or more long-exposure Z-scan images of a live cell sample obtained by passing the focal plane of a camera through the sample in the Z direction, the camera thereby integrating the fluorescence intensity from the sample over the Z dimension, and (2) an associated two-dimensional Z-stack projection pixel-scale ground truth image, wherein the pixel-scale ground truth image is obtained from a set of Z-stack images of the same live cell sample, each image of the Z-stack images obtained at a different Z focal plane position of the sample, and wherein the Z-stack images are combined using a Z-projection algorithm. The method further includes the step of performing a training process of the neural network from the training set, thereby generating a model that is trained to ultimately generate 2D projection images of live cell samples from input in the form of one or more long-exposure Z-scan images.
[0018] As mentioned above, the long-exposure Z-scan images in the training set can take the form of a single Z-scan image or a set of consecutive Z-scan images (optionally summed), or any combination of the three acquisition paradigms.
[0019] In another aspect, a live cell imaging system is described for use in conjunction with a sample holding device (e.g. adapted to hold a live cell sample micro-well plate). The system includes a fluorescence microscope having one or more excitation light sources, one or more objective lenses, and a camera operable to acquire fluorescence images from a live cell sample held within the sample holding device. The fluorescence microscope includes a motor system configured to move the fluorescence microscope relative to the sample holding device, including in the Z direction, such that the camera acquires one or more long-exposure Z-scan images of the sample, and one or more Z-scan images obtained by passing the focal plane of the camera through the sample in the Z direction. The system further includes a processing unit including a trained neural network model for generating a focused two-dimensional projection of fluorescence images of the live cell sample from the one or more long-exposure Z-scan images.
[0020] In another aspect, a method for generating a training set for training a neural network is provided. The method comprises the steps of: (a) using a camera, acquiring one or more long-exposure fluorescence images of a three-dimensional sample by passing the focal plane of the camera through the sample in the Z direction, the camera thereby integrating the fluorescence intensity of the sample over the Z dimension; (b) generating a pixel-scale ground truth image of the same sample from one or more different images of the sample obtained from the camera; (c) repeating steps (a) and (b) for a plurality of different samples; and (d) providing the images obtained by performing steps (a), (b), and (c) as a training set for training a neural network.
[0021] The method of the present invention has a number of advantages:
[0022] (1) The method allows one to obtain biologically relevant analytical metrics from high-quality projection images using standard image analysis techniques, such as fluorescent object counting or other fluorescence intensity metrics.
[0023] (2) The one or more long-exposure images are a true representation of the sample fluorescence integrated over the Z dimension, thereby generating true, accurate data.
[0024] (3) The method of the present invention can be implemented without changing the traditional epifluorescence microscope hardware.
[0025] (4) While the workflow of the present invention is designed to accommodate traditional wide-field fluorescence microscopes, it can also be applied to spinning-disk confocal microscopes to improve the throughput of that modality.
[0026] (5) The method provides a single, high-quality 2D projection image as output, eliminating the burden on the user to process 3D data sets using complex software and analysis tools. The 2D projection image can be input into standard 2D image visualization, segmentation, and analysis pipelines, such as those currently implemented on the most advanced fluorescence microscope platforms on the market.
[0027] In yet another aspect, a computer-readable medium is provided, storing non-transitory instructions for a live-cell imaging system comprising a camera and a processing unit implementing a neural network model, the instructions causing the system to perform the method of the present invention, e.g., capturing one or more long-exposure Z-scan images of a sample, providing the images to a trained neural network model, and generating a two-dimensional projection image using the trained network model. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 is a schematic illustration of a method and system for acquiring 3D fluorescence images of live-cell biological samples.
[0029] Figure 2A is a schematic illustration of a method and system for acquiring 3D fluorescence images of live-cell biological samples. Figure 1Method diagram for Z-sweep image acquisition for a motorized fluorescence microscope; Figure 2B Method diagram for continuous Z-sweep image acquisition; Figure 2C Method diagram for prior art Z-axis stack image acquisition.
[0030] Figure 3 Method diagram for generating a training dataset for training a neural network model. Figure 3 Also described is a method for using the trained neural network model to take new input images for inference, and a method for outputting in the form of high-quality 2D projection images.
[0031] Figure 4 Model training program flow diagram for training a neural network model for Figure 3 .
[0032] Figure 5 Embodiment diagram of a network in which a trained neural network model is implemented on a remote computer and communicates with a remote workstation that receives fluorescence images from a connected fluorescence microscope.
[0033] Figure 6 High-quality focused 2D projection image diagram obtained using a trained neural network model from Figure 2A or Figure 2B long-exposure image acquisition methods.
[0034] Figure 7 Live cell imaging system diagram for biological samples loaded into the sample wells of a microplate. The system includes a fluorescence optical module for acquiring fluorescence images of the samples in the microplate using features of the invention.
[0035] Figure 8 Schematic diagram of a fluorescence optical module located within Figure 7 a live cell imaging system.
[0036] Figure 9 A set of three examples of input images and model outputs, compared to ground truth pixel-scale projections, demonstrating the utility of the current method. In Figure 9 , the neural network model is a trained conditional generative adversarial network (GAN), and the input images are generated under the second image acquisition paradigm of Figure 3 .
[0037] Figure 10 High-level architecture of the conditional GAN used to generate the model outputs (predicted projections) shown in Figure 9 . DETAILED DESCRIPTION
[0038] Reference is now made to Figure 1A live-cell biological sample (which may be in the form of any type of live-cell sample used in the current study, such as, as described in the background, a live three-dimensional cell culture) is placed on or within a suitable holding structure 12 (e.g., a microplate or glass slide). One or more fluorescent reagents (fluorophores) may be added to the live-cell sample 10, as shown in 14; the type or nature of the fluorescent reagent depends, of course, on the details of the sample 10 and the nature of the study being conducted. The holding structure 12 with the sample 10 and reagents is then supplied to an electric fluorescence microscope 16 equipped with a camera, which may be in the form of any conventional fluorescence microscope with imaging capabilities known in the art or available from many manufacturers (including the assignee of this invention).
[0039] according to Figure 2A or Figure 2B The method involves using a fluorescence microscope 16 to generate one or more long-exposure Z-scan images (or a set of such images) 18. Long-exposure Z-scan images of the three-dimensional live-cell biological sample 10 are obtained by passing the focal plane of the camera in the microscope along the Z-direction through the sample, such as... Figure 2A 25 and Figure 6 As shown in Figure 25, the camera integrates the fluorescence intensity of the sample in the Z dimension. Long-exposure Z-sweep images generate dataset 30, as shown in Figure 25. Figure 2A and Figure 6 As shown. Then, this dataset 30 is provided as input to the computer or processor implementing the trained neural network model 22. Figure 1 As described below, the model is trained from multiple sets of training images and generates high-quality focused 2D projection images, such as... Figure 3 and Figure 6 As shown in 140. This 2D projection diagram 140 ( Figure 3 and Figure 6 The image contains 3D fluorescence information (because image acquisition is achieved by integrating fluorescence intensity over the Z dimension), but in the form of a projection onto a two-dimensional focused image 140. Then, the resulting focused 2D projected image 140 (…) Figure 3 , Figure 6 It can be used for any standard and known image analysis technique, such as fluorescent object counting or other fluorescence intensity measurement.
[0040] exist Figure 2A In the process, as the focal plane passes through the sample along the Z-direction, as shown by arrow 25, a single long-exposure Z-sweep image is obtained by continuously exposing the camera in a fluorescence microscope, thus yielding a two-dimensional image dataset 30. This is Figure 3 The “paradigm 1”. Image dataset 30 has limited use for image analysis techniques because of the lack of focus. Therefore, a trained neural network model is used to convert the out-of-focus images into more useful focused projection images, such as... Figure 3 andFigure 6 As shown in 140. Figure 9 More examples are shown below, which will be discussed later in this article.
[0041] Another long-exposure Z-sweep image acquisition method is as follows Figure 2B As shown. In this alternative, image acquisition includes acquiring a series of consecutive long-exposure Z-sweep images, in Figure 2B Numbered 1, 2, 3, and 4, the camera in the microscope is exposed, and as the focal plane moves through a Z-dimensional increment, the fluorescence intensity is integrated in the Z-dimensional dimension, and the fluorescence is captured. Figure 2B The series of four corresponding images or image datasets shown in Figures 1, 2, 3, and 4 is then summed to generate a single long-exposure Z-sweep dataset or image 30. This is Figure 3 Paradigm 2. (and) Figure 2A Compared to other methods, this image acquisition technique may have certain advantages, including avoiding image saturation and allowing the application of traditional or deep learning-based 2D fluorescence deconvolution algorithms to each smaller sweep (1, 2, 3, 4) before summation. This can eliminate out-of-focus fluorescence and ultimately achieve more accurate fluorescence intensity measurements.
[0042] As a third implementation scheme for acquiring long-exposure Z-sweep images, Figure 2B The procedure may also change, such as Figure 3 As shown in "Paradigm 3". Images 1-4 are according to Figure 2B The images obtained as described above are not summed. These four images are fed into a trained neural network to predict a corresponding focused image from all images. In this embodiment, the four images can be combined into a four-channel image. The advantages described above still apply to this embodiment. Furthermore, it has the advantage that the neural network model can more easily infer a well-focused projected image from the images obtained in this paradigm.
[0043] Figure 2C This paper describes a prior art Z-axis stacked image acquisition method in which images of sample 10 are acquired at different focal planes (e.g., positions A, B, C, and D) to generate the corresponding image datasets A, B, C, and D shown in the figure. Note that in this method, the Z-dimensional coordinates are fixed or stationary during camera exposure.
[0044] Figure 3 The training of the convolutional neural network model is shown in more detail. This model is used to generate a convolutional neural network model from input in the form of one or more long-exposure Z-sweep images. Figure 6 2D projected image 140.
[0045] The training dataset of model inputs is prepared by generating a Z-stack 104 of the same 3D live cell training sample (100) Figure 2C and a Z-scan (100) Figure 2A or Figure 2B of any of the three paradigms discussed above. In addition, a pixel-scale ground truth image 108 of the same training sample is obtained from the Z-stack 104 and the Z-projection algorithm 106. Thus, using the selected Z-projection algorithm, e.g. the ImageJ algorithm referenced above, the Z-stack 104 is used to generate high quality 2D projections 108 that will be used as the“pixel-scale ground truth” for model training. The Z-scan images 102, 102A or 105 are used as input for model training along with the associated pixel-scale ground truth 108. These pairs of data (in fact, thousands of such pairs of data) are used to train a convolutional neural network (CNN) model, as shown at 110, in order to enable the network to learn how to generate high quality 2D projection images (140, Figure 6 ) from long-exposure Z-scan images (30, Figure 6 ). The trained neural network model is shown at 22 in Figure 1 and Figure 3 As noted above, the inputs to the model training 110 can be pairs of images as just described, e.g. in a supervised model training exercise, but as noted above the model inputs need not be paired, e.g. in the case of a cycle-consistency loss or generative adversarial network model, using simple sets of images from two domains (pixel-scale ground truth Z-projections and Z-scans).
[0046] The neural network model 22 can be a supervised model, e.g. an encoder-decoder based model such as a U-net, see Ronneberger et al.,“U-net: Convolutional networks for biomedical image segmentation”, published by Springer in International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI) pp. 234-241, also published in arXiv: 1505.04597 (2015), the contents of which are incorporated herein by reference. This supervised model is trained to directly predict high quality projection images from corresponding long-exposure Z-scan images.
[0047] The neural network model 22 can also be designed and trained using an adversarial approach, e.g. GAN, see Goodfellow et al., “Generative Adversarial Nets,” in Advances in Neural Information Processing Systems 27, Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, and K. Q. Weinberger, Eds. Curran Associates, Inc., 2014, pp. 2672-2680, where one network, the generator, is trained according to the other network (the discriminator). The task of the discriminator is to distinguish between real high-quality projections and the output of the generator, and the task of the generator is to make the output indistinguishable from real projections (i.e. the pixel scale ground truth).
[0048] Another alternative model architecture is the conditional GAN, to be combined with Figure 10 Further details are described in further detail.
[0049] The trained neural network model 22 can also be trained using a cycle-consistency loss approach, e.g. CycleGAN, see Zhu, Jun Yan, et al. “Unpaired image-to-image translation using cycle-consistent adversarial networks.” Proceedings of the IEEE international conference on computer vision (2017), also published as arXiv: 1703.10593 (2017), the contents of which are hereby incorporated by reference, which means that unpaired Z-scan images are translated into high-quality projections, and then translated back again into Z-scan images, and the network is trained to minimize the error of the reverse reconstruction of the cycle. The advantage of using cycle-consistency is that this training does not require perfect registration between the Z-scan images and the high-quality projections. It also provides the possibility of training on unpaired data from two domains (in this case Z-scan images and Z-stack projection images).
[0050] Once the neural network model has been trained from a set of training samples (there can be hundreds or thousands of such 3D live cell samples), at inference time, there is no longer a need to acquire Figure 2C Z-stack images, which is a slower process and more destructive to the sample. Instead, the trained neural network model 22 is used to Figure 2AAlternatively, a 2B technique may be used, collecting only one or more high-throughput long-exposure Z-sweep images and then feeding these images into a trained model 22 to generate a high-quality 2D projected image 140. In one implementation shown in 120 (“Paradigm 1”), only a single Z-sweep is performed, and the resulting image dataset 132 is fed into the trained model 22. Figure 3 In another implementation, denoted as "Paradigm 2", position 134 indicates the execution of a series of consecutive Z-sweeps ( Figure 2B The generated images 136 are summed to generate an image dataset 32A. The sequential Z-sweep images 136 offer several advantages without sacrificing acquisition speed: (1) they avoid image saturation, and (2) they allow the application of a conventional or deep learning-based 2D fluorescence deconvolution algorithm to each smaller sweep before summing, which can eliminate defocused fluorescence and ultimately achieve more accurate fluorescence intensity measurement. The generated images 136 are selectively deconvolved with fluorescence, then summed as shown in 138, and the resulting image dataset 132A is then fed to the trained model 22.
[0051] The third alternative ( Figure 3 In Paradigm 3), a series of consecutive Z-sweeps 135 ( Figure 2B The procedure (132B) yields a set of resulting image datasets from each sweep. The image datasets 132B are then provided to the trained model 22. Alternatively, if, in this alternative, the image datasets are combined into a single multi-channel input image as described above, the neural network output can be improved by providing complementary features.
[0052] The model trained on the input image obtained from paradigm 1, 2, or 3 is used for model inference so as to match the paradigm type used to obtain the input image. In other words, for example, if the input image is obtained according to paradigm 2 at the time of inference, then the inference is performed by the model trained on the image obtained according to paradigm 2.
[0053] Due to model interference, the trained model 22 produces model output 140, which is a high-quality focused two-dimensional projection image.
[0054] Model training will be Figure 4 A more detailed description follows. The model training process 200 includes training on live cells (e.g., 3D training samples) (using...). Figure 2A Step 202: Using 2B image acquisition technology for long exposure Z-sweep Figure 2CThe same steps 204 of acquiring Z-stack images of training samples Z are taken by the technology of the present application, and the Z-stack acquired in step 206 is projected into a 2D image using Z-stack projection technology (e.g. ImageJ). This process is repeated for many different training samples (e.g. thousands or tens of thousands of live cells) as shown at 208, spanning all expected sample types used in current live cell research applications of fluorescence microscopy. Then, the multiple images (steps 202 and 206) from the paired or unpaired datasets (Z-scan images and pixel scale ground truth images) from both domains are provided as a training set 210, and used to train a neural network or generative adversarial network approach as shown at 212, according to a CycleGAN procedure, in a supervised learning model training exercise or a self-supervised model training exercise. The result of this exercise is a trained neural network model 22.
[0055] In one possible variation, there can be several discrete models trained using this general approach, one for each type of live cell sample, e.g. one for stem cells, one for oncology, another for brain cells, etc. Each model is trained using hundreds or thousands of images (paired or unpaired) obtained by the procedures specified in Figure 3 and Figure 4
[0056] Referring to Figure 1 and Figure 5 , the trained neural network model 22 can be implemented in a workstation 24 or in a processor associated with or part of the fluorescence microscope / imager 16. Alternatively, the trained model can be implemented on a remote server, as shown at 300 in Figure 5 In this configuration, the fluorescence microscope / imager 16 is connected to the workstation 24, which is connected to the server 300 over a computer network 26. The workstation transmits the long-exposure Z-scan images to the server 300. In turn, the trained neural network model in the server performs inference as shown at Figure 3 and generates a 2D projected image (140, Figure 6 ), which is transmitted back to the workstation 24 over the network 26, where it is displayed to the user or provided as input to image analysis and quantification software pathways running on the workstation 24.
[0057] Fluorescence Microscope System
[0058] One possible implementation of the features of the present application is in a live cell imaging system 400, which includes a fluorescence microscope for acquiring fluorescence images in three-dimensional live cell research applications. Figure 7 A live cell imaging system 400 is shown having a housing 410, which, during use, can be placed in a temperature and humidity controlled incubator, not shown. The live cell imaging system 400 is designed to receive a microplate 12, which includes a plurality of sample holding wells 404, each receiving a 3D live cell sample. The system includes a set of fluorescent reagents 406, one or more of which are added to each sample well 404 to enable fluorescence assays from the live cell samples. The system includes an associated workstation 24, which implements image analysis software and has display capabilities that enable the researcher to see the results of live cell experiments performed on the samples. The live cell imaging system 400 includes a tray 408 that slides out of the system and allows the microplate 12 to be inserted onto the tray 408, which is then retracted and closed to place the microplate 12 inside the live cell imaging system housing 410. The microplate 12 remains stationary inside the housing while a fluorescence optical module 402 (see Figure 8 ) moves relative to the plate 12 and acquires a series of fluorescence images during the course of an experiment. The fluorescence optical module 402 implements the long-exposure Z-sweep image acquisition techniques shown in Figure 2A and / or 2B (paradigms 1, 2, or 3 as previously described) to generate image datasets that are input to the trained neural network model as shown in the inference section of Figure 3 .
[0059] Figure 8 is a more detailed optical diagram of the fluorescence optical module 402 in Figure 7 . For more detailed information about the fluorescence optical module 422 shown in Figure 8 , see U.S. Patent Application No. 16 / 854,756, filed April 21, 2020, entitled “Optical module with three or more color fluorescent light sources and methods for use thereof,” assigned to the assignee of the present invention, the contents of which are incorporated by reference herein.
[0060] The module 402 includes LED excitation light sources 450A and 450B that emit light at different wavelengths, such as 453-486 nm and 546-568 nm, respectively. The optical module 402 can be configured with a third LED excitation light source (not shown) that emits light at a third wavelength, such as 648-674 nm, and even a fourth LED excitation light source that emits at a fourth different wavelength. The light emitted by the LEDs 450A and 450B passes through narrow band filters 452A and 452B, respectively, that pass light at specific wavelengths designed to excite fluorophores in the sample. The light that passes through the filter 452A is reflected off of dichroic mirror 454A and dichroic mirror 454B and directed toward an objective lens 460, such as a 20x magnification lens. The light emitted by the LED 450B also passes through the filter 452B and also passes through the dichroic mirror 454B and is directed toward the objective lens 460. The excitation light that passes through the objective lens 460 then impinges on the bottom of the sample plate 12 and enters the sample 10. In turn, the emission from the fluorophores in the sample passes through the objective lens 460, reflects off of the mirror 454B, passes through the dichroic mirror 454A, and passes through a narrow band emission filter 462 (which filters out non-fluorescent light) and then impinges on a digital camera 464, which can take the form of a charge-coupled device (CCD) or other type of camera known in the art and used in fluorescence microscopy. Then, when the light source 450A or 450B is in the on state, the motor system 418 operates to move the Z dimension of the entire optical module 402 to obtain a long exposure Z-scan image (or 2B). It will be appreciated that typically only one optical channel is activated at a time, e.g., the LED 450A is turned on and a long exposure Z-scan image is captured, then the LED 450A is turned off, the LED 450B is activated, and a second long exposure Z-scan image is captured. Figure 2A
[0061] It will be appreciated that the objective lens 460 can be mounted on a turret that can be rotated about a vertical axis to place a second objective lens of a different magnification into the optical path to obtain a second long exposure Z-scan image at a different magnification. In addition, the motor system 418 can be configured to move in the X and Y directions beneath the sample plate 12 so that the optical path of the fluorescence optical module 402 and the objective lens 460 are directly beneath each well 404 of the sample plate 12 and the fluorescence assay described above is performed from each well (and thus each live cell sample) that is fixed in the plate 12.
[0062] The details of the motor system 418 of the fluorescence optical module 402 can vary widely and are known to those skilled in the art.
[0063] The operation of the live cell imaging system is programmed by a conventional computer or processing unit that controls the motor system and the camera in coordination to acquire images of the samples in the wells. The processing unit can implement Figure 3 A trained neural network model. In this case, as one possible implementation, the computer includes a memory storing non-transitory instructions (code) that implement the method of the invention, including the images acquired from paradigm 1, 2 and / or 3 and the previously trained neural network model as described above in connection with Figure 3 and Figure 4 The computer or processing unit also includes a memory that includes instructions for generating all training images from a set of training samples used for model training, which can be performed in the computer of the live cell imaging system or in a remote computing system that implements the model and performs the model training.
[0064] Applications
[0065] The methods herein are useful for two-dimensional projection image generation of organoids, tumor spheroids, and three-dimensional structures found in other biological samples such as cell cultures. As previously described, the use of live cell samples encompasses a wide range of research areas, including immuno-oncology, oncology, metabolism, neuroscience, immunology, infectious disease, toxicology, stem cells, cardiology, and inflammation. In these research areas, cell health and proliferation, cell function, cell motility, and morphology are studied, including complex immune-tumor cell interactions, synapse activity, and cancer cell metabolism. The methods of the invention are relevant to all of these applications.
[0066] In particular, the methods of the invention are relevant to the above applications in that they allow for high-throughput fluorescent image capture of samples, generating high-quality fluorescent 2D projection images that can be segmented and analyzed to determine how experimental conditions (e.g., drug treatments) affect organoids, tumor spheroids, or other three-dimensional biological structures. Organoids (e.g., pancreatic cell organoids, liver cell organoids, and intestinal cell organoids) and tumor spheroids are of particular interest because of their three-dimensional structure, which more closely approximates the “natural” three-dimensional environment of cultured cells. Thus, the response of organoids, tumor spheroids, or other such three-dimensional multicellular structures to drug or other applied experimental conditions can more closely approximate the response of the corresponding sample in a simulated human or other target environment.
[0067] Example 1
[0068] The methods of the invention use a trained conditional generative adversarial network (GAN) model to test on three-dimensional live cell biological samples. The trained model generates two-dimensional focal projection images, three examples of which are shown in Figure 9 The images were obtained using samples loaded from microwell plate wells using the instruments in Figure 7 and 8 The entire fluorescent imaging system and samples are placed in an incubator to culture the samples under appropriate temperature and humidity conditions for live cell culture analysis.
[0069] Specifically, Figure 9 Three examples of input images (901, 906, 912) and model outputs (902, 908, 914) are compared with pixel-scale ground truth projections (904, 910, 916) obtained through the Z-stack projection algorithm, demonstrating the practicality of the proposed method as previously mentioned. Figure 9 In the above description, the input images (901, 906, 912) are as described above. Figure 2B and Figure 3 Generated under the second image acquisition paradigm. Pixel-scale ground truth projected images 904, 910, and 916 are generated using... Figure 2C The method and Z-axis stack projection algorithm were used to obtain the results.
[0070] Figure 9 The image depicts an example of glioblastoma cell spheroids, a typical example of single tumor spheroid analysis. This analysis was provided by the assignee. One of the analyses available for live-cell imaging systems, such as the present invention Figure 7 and Figure 8 As shown. For more information on this analysis, please visit https: / / www.essenbioscience.com / en / applications / cell-health-viability / spheroids / . Figure 9 Each image shown represents the entire microwell of a microplate; the spheres have a radius of approximately 400 μm, and there is only one sphere per well. Each sphere is a 3D structure composed of thousands of cells. Fluorescent labels are applied to the microwells to label the cell nuclei. Therefore, each bright spot visible in the model output images (902, 908, 914) and the pixel-scale ground truth projections (904, 910, 916) represents a single cell.
[0071] generate Figure 9 The training model for the output image (902908914) is a conditional generative adversarial network (GAN). For more information on such networks, please see Philip Isola et al., Image-to Image Translation with Conditional Adversarial Networks, arXiv.1611.07004[cv.cs](2016), and Xudong Mao et al., Least Squares Generative Adversarial Networks, arXiv.1611.04076[cv.cs](2016), the contents of which are incorporated herein by reference.
[0072] The high-level architecture of conditional GANs used is as follows:Figure 10 As with other GANs, this model includes a generator 1004 and a discriminator 1010. The generator 1004 receives as input 1002 one of the images acquired according to one of the three paradigms described above. A "fake" dataset 1006 is generated that is a concatenation of the input image and the predicted projection image generated by the generator 1004. A "real" dataset 1008 is generated that is a concatenation of an input image and the voxel-scale ground truth projection image. These two datasets 1006 and 1008 are then provided to the discriminator that computes the loss. At each training iteration, the generator 1004 is able to better generate "fake" images (i.e. predicted projection images) and the discriminator 1010 is able to better identify real (voxel-scale ground truth) versus fake (predicted projection) images by updating the loss functions of the generator and the discriminator. After the model is trained, this enables the generator 1004 to generate high quality projection images. In the architecture of this embodiment, the generator 1004 and the discriminator 1010 are separate neural networks that can have different underlying architectures. In this embodiment, the generator 1004 is a U-net (see Ronneberger et al., "U-net: Convolutional networks for biomedical image segmentation", published by Springer jointly with the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) pages 234-241) and the discriminator 1010 is a PatchGAN with the architecture described in the Isola et al. paper cited above. Since these networks are described in detail in the scientific literature, for the sake of brevity and to not obfuscate the inventive details of the present invention, a more detailed discussion is omitted. ResNet50 is another example of a network that can be used as a discriminator. For more details on GANs, see the Goodfellow et al. paper cited earlier. Figure 3 Figure 10 As with other GANs, this model includes a generator 1004 and a discriminator 1010. The generator 1004 receives as input 1002 one of the images acquired according to one of the three paradigms described above. A "fake" dataset 1006 is generated that is a concatenation of the input image and the predicted projection image generated by the generator 1004. A "real" dataset 1008 is generated that is a concatenation of an input image and the voxel-scale ground truth projection image. These two datasets 1006 and 1008 are then provided to the discriminator that computes the loss. At each training iteration, the generator 1004 is able to better generate "fake" images (i.e. predicted projection images) and the discriminator 1010 is able to better identify real (voxel-scale ground truth) versus fake (predicted projection) images by updating the loss functions of the generator and the discriminator. After the model is trained, this enables the generator 1004 to generate high quality projection images. In the architecture of this embodiment, the generator 1004 and the discriminator 1010 are separate neural networks that can have different underlying architectures. In this embodiment, the generator 1004 is a U-net (see Ronneberger et al., "U-net: Convolutional networks for biomedical image segmentation", published by Springer jointly with the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) pages 234-241) and the discriminator 1010 is a PatchGAN with the architecture described in the Isola et al. paper cited above. Since these networks are described in detail in the scientific literature, for the sake of brevity and to not obfuscate the inventive details of the present invention, a more detailed discussion is omitted. ResNet50 is another example of a network that can be used as a discriminator. For more details on GANs, see the Goodfellow et al. paper cited earlier.
[0073] In the conditional GAN attached, the generator is conditioned on some data, such as the long-exposure (Z-sweep) input image in this embodiment. The generator is conditioned using one or more long-exposure images, both for training and inference. At the end of training, only the generator is needed since it will be used to generate the projection images at inference. The loss function minimization and discriminator-generator iterations are only relevant to the training phase.
[0074] Model training for GANs is described as follows: As mentioned above, GANs have two models, the generator (G) and the discriminator (D). G produces output from noise; the output mimics the distribution from the training dataset. D attempts to distinguish between real data and generated data. Essentially, G attempts to fool D; at each training iteration, the loss functions for D and G are updated. Model training can be done to update G or D more or less. As model training progresses, G can better generate data that mimics real data in the training dataset, and after model training, G can make inferences on input images and directly generate predicted projections, Figure 9 Three examples are shown in FIGS. 5A-5C.
[0075] Regarding Figure 10 Additional notes on the architecture and model training are as follows. (1) Soft labels. Instead of actual labels being 1, they can be set to a value between [0.9, 1]. (2) Loss function. Binary cross-entropy loss is the original loss function for GANs. Mean squared error (MSE) loss without activation function can be used to make it more stable (making it a least squares (LS) GAN). (3) No fully connected layers. If Resnet is used as the discriminator, the last fully connected layer can be removed and a flatten layer is inputted (i.e., to provide the list of labels for its loss function). (4) PatchGAN discriminator. In this implementation of GAN, a fully convolutional discriminator can be used, with the output layer having multiple neurons, each corresponding to a certain patch of the input image.
[0076] Conclusion
[0077] The method of the present invention overcomes many of the shortcomings of traditional methods of acquiring 3D fluorescence information from live 3D cell cultures. Traditional step-wise “Z-stack” fluorescence imaging methods of 3D samples are slow, require user input, and ultimately expose the sample to excess fluorescent light, leading to phototoxicity and photobleaching of the sample, both of which are highly undesirable. Other methods require specialized hardware (such as a spinning disk) or advanced optical equipment (such as light sheet microscopy). Alternative deep learning methods use methods that can compromise the integrity of the data, including ultra-low exposure times, or generate 3D data from a single focal plane.
[0078] In contrast, the method of the present application does not require specialized hardware, only a simple fluorescence microscope with an axial motor. Notably, the above techniques can be easily applied to other acquisition systems. The speed of acquiring images from live cell samples is fast compared to traditional methods and reduces the risk of phototoxicity and photobleaching. Furthermore, the raw images acquired from the camera are true images of 3D fluorescence as they originate from the integrated fluorescence over the Z dimension. Finally, the output of a single high-quality 2D projection image eliminates the burden on the user to use complex software and analysis tools to process 3D datasets, such high-quality 2D-projections (140, Figure 3 and Figure 6 , Figure 9 , 902, 908, 914) can be inputted into standard 2D image visualization, segmentation, and analysis pipelines.
Claims
1. A method for generating a focused two-dimensional projection image of a fluorescent image of a three-dimensional live cell sample, comprising the steps of: acquiring one or more long-exposure images of the sample by moving a focal plane of a camera through the sample in the Z-direction, the camera thereby integrating the fluorescent intensity from the sample in the Z-dimension, providing the one or more long-exposure images to a neural network model trained from a plurality of training images; generating the focused two-dimensional projection image with the trained neural network model; wherein the training set is input to the neural network model in the form of a plurality of images, the images comprising: (1) one or more long-exposure images of a three-dimensional live cell sample acquired by moving a focal plane of a camera through the training sample in the Z-direction, the camera thereby integrating the fluorescent intensity from the training sample in the Z-dimension, and (2) a pixel-scale ground truth image, wherein the pixel-scale ground truth image is obtained from a set of images obtained at different Z-focal plane positions of the training sample and combined into a two-dimensional projection image using a projection algorithm.
2. The method of claim 1, wherein the neural network model comprises a convolutional neural network (CNN) model.
3. The method of claim 2, wherein the CNN model comprises an encoder-decoder based model.
4. The method of claim 1, wherein the neural network model is trained using supervised learning.
5. The method of claim 1, wherein the neural network model is trained using a generative adversarial network (GAN) method.
6. The method of claim 5, wherein the GAN method comprises a conditional GAN with a generator and a discriminator.
7. The method of claim 6, wherein the generator of the conditional GAN is conditioned on the one or more long-exposure images.
8. The method of claim 1, wherein the neural network model is trained using a cycle consistency loss method.
9. The method of claim 8, wherein the cycle consistency loss method comprises CycleGAN.
10. The method of claim 1, wherein the one or more long-exposure images comprise a set of consecutive long-exposure images.
11. The method of claim 10, wherein the method further comprises the step of performing fluorescent deconvolution on each consecutive image, and a summing operation that sums the consecutive long-exposure images after fluorescent deconvolution.
12. The method of claim 1, wherein the live cell sample is contained within a well of a microplate.
13. The method of claim 1, wherein the plurality of training images comprises one or more long-exposure images of a plurality of three-dimensional live cell training samples selected for model training, and an associated pixel-scale ground truth image for each three-dimensional live cell training sample, the long-exposure images each being obtained by moving a focal plane of a camera through the training sample in the Z-direction, each pixel-scale ground truth image comprising a two-dimensional projection of a set of Z-stack images of the training sample.
14. The method of any one of claims 1-13, wherein the three-dimensional live cell sample comprises an organoid, a tumor spheroid, or a three-dimensional cell culture.
15. The method of any one of claims 1-13, wherein the camera is incorporated into a live cell imaging system, the trained neural network model is implemented in a computing platform remote from the live cell imaging system, the computing platform being in communication with the live cell imaging system over a computer network.
16. A method for training a neural network to generate two-dimensional projection images of fluorescent images of a three-dimensional live cell sample, comprising the steps of: (a) obtaining a training set in the form of a plurality of images, wherein the images comprise: (1) one or more long-exposure images of a three-dimensional live cell sample obtained by passing the focal plane of a camera through the training sample in the Z direction, the camera thereby integrating the fluorescent intensity from the training sample in the Z dimension, and (2) pixel-scale ground truth images, wherein the pixel-scale ground truth images are obtained from a set of images obtained at different Z focal plane positions of the training sample and combined into two-dimensional projection images using a projection algorithm; (b) performing a neural network model training procedure using the training set to generate a trained neural network.
17. The method of claim 16, wherein images (1) and (2) comprise a plurality of paired images.
18. The method of claim 16, wherein images (1) and (2) comprise a plurality of unpaired images, and wherein the model training procedure comprises a cycle consistency loss or a generative adversarial network model training procedure.
19. The method of claim 18, wherein the cycle consistency loss model training procedure comprises a CycleGAN.
20. The method of claim 16, wherein the neural network comprises a convolutional neural network, CNN.
21. The method of claim 20, wherein the CNN comprises an encoder-decoder based neural network.
22. The method of any one of claims 16, 17, 20, or 21, wherein the neural network is trained using supervised learning.
23. The method of claim 16, wherein the trained neural network comprises a generative adversarial network, GAN.
24. The method of claim 23, wherein the GAN comprises a conditional GAN having a generator and a discriminator.
25. The method of claim 24, wherein the generator of the conditional GAN is conditioned on one or more long-exposure images.
26. The method of claim 16, wherein the one or more long-exposure images comprise a set of sequential images.
27. The method of claim 26, wherein the method further comprises the step of performing fluorescent deconvolution on each sequential image, and a summation operation that sums the sequential images after fluorescent deconvolution.
28. A live cell imaging system for use in conjunction with a sample holding device adapted to contain a three-dimensional sample to generate two-dimensional projection images of the sample, comprising: a fluorescence microscope having one or more excitation light sources, one or more objective lenses, and a camera that can acquire one or more fluorescent images from a three-dimensional sample within the sample holding device, wherein the fluorescence microscope comprises a motor system configured to move the fluorescence microscope in the Z direction relative to the sample holding device such that the camera obtains one or more long-exposure images of the sample, the images being obtained by moving the focal plane of the camera through the sample in the Z direction, the camera thereby integrating the fluorescence intensity of the sample over the Z dimension; and a processing unit comprising a trained neural network model for generating two-dimensional projection images of a three-dimensional sample from one or more long-exposure images, wherein the method of any one of claims 16-27 trains the neural network model.
29. The system of claim 28, wherein the one or more long-exposure images comprise a set of consecutive images.
30. The system of claim 28, wherein the sample holding device comprises a microwell plate having a plurality of microwells.
31. The system of claim 28, wherein, The neural network model is trained from a plurality of training images, the training images comprising one or more long-exposure images, the long-exposure images being obtained by moving the focal plane of the camera through the training sample in the Z direction, the camera thereby integrating the fluorescence intensity from the training sample over the Z dimension, and a ground truth image of the training sample at a relevant pixel scale, the ground truth image being generated from one or more different images of the training sample obtained by the camera.
32. The system of claim 28, wherein the three-dimensional sample comprises an organoid, a tumor spheroid, or a three-dimensional cell culture.
33. The system of any one of claims 28-32, further comprising a remote computing platform that implements the trained neural network model and is in communication with the live cell imaging system over a computer network.
34. A method for generating a training set for training a neural network, comprising the steps of: (a) using a camera, obtaining one or more long-exposure fluorescence images of a three-dimensional training sample by moving the focal plane of the camera through the training sample in the Z direction, the camera thereby integrating the fluorescence intensity from the training sample over the Z dimension; (b) generating a ground truth image of the same training sample at a pixel scale from one or more different images of the training sample obtained by the camera; (c) repeating steps (a) and (b) for a plurality of different training samples; and (d) providing the images obtained by performing steps (a), (b), and (c) as a training set for training a neural network.
35. The method of claim 34, wherein, The one or more different images of the training sample obtained by the camera in step (b) comprise a set of images obtained at different Z focal plane positions of the training sample, wherein the ground truth image is generated by projecting the set of images into a two-dimensional projection image.
36. The method of claim 34, wherein the three-dimensional training sample comprises an organoid, a tumor spheroid, or a three-dimensional cell culture.
37. The method of claim 34, further comprising the step of repeating steps (a)-(d) for different types of three-dimensional training samples, thereby generating different types of training sets.
38. The method of any one of claims 34-37, wherein the one or more long-exposure images comprise a set of consecutive images.
39. The method of claim 38, wherein the method further comprises the step of performing a fluorescent deconvolution on each successive image, and a summation operation that sums the successive images after the fluorescent deconvolution.
40. The method of any one of claims 34-37, further comprising performing step (a) according to one or more image acquisition paradigms to obtain a long-exposure image.
41. A computer readable medium for storing non-transitory instructions of a live cell imaging system comprising a video camera and a processing unit implementing a neural network model, the instructions causing the system to perform the method of any one of claims 1-12.
Citation Information
Patent Citations
Optical module with three or more color fluorescent light sources and methods for use thereof
US11320380B2
Microscopy system, microscopy method, and computer-readable storage medium
US20190025213A1
Systems and methods for two-dimensional fluorescence wave propagation onto surfaces using deep learning
WO2020139835A1