Method and apparatus for Fourier lamination imaging reconstruction processing
By using GPUs to parallel process low-resolution images in Fourier stacking microscopy, the problem of long FPM reconstruction time is solved, efficient high-resolution image generation is achieved, and its application range is expanded.
Patent Information
- Application Number
- CN202380087493.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-23
- Filing Date
- 2023-12-22
- Publication Date
- 2025-09-05
AI Technical Summary
Fourier stacking microscopy (FPM) consumes too much time in the image acquisition and reconstruction process, which limits its application in imaging moving samples.
Graphics processing units (GPUs) are used to process low-resolution images in parallel and generate high-resolution images through FPM reconstruction, thus reducing reconstruction time.
This significantly reduces FPM reconstruction time, improves imaging efficiency, and is suitable for a wider range of applications such as medical diagnosis and clinical testing.
Smart Images

Figure CN120604159A_ABST
Abstract
Description
[0001] This patent document claims the benefit of Indian Patent Application No. 202211075050 filed on December 23, 2022, which is hereby incorporated by reference in its entirety. Technical Field
[0002] The present application relates to diagnostic imaging and, more particularly, to methods and apparatus for Fourier stack imaging reconstruction processing. Background Art
[0003] Fourier ptychography (FPM) is a microscopy technique that enables high-resolution imaging over a wide field of view. FPM employs an array of light sources to illuminate the sample while capturing a set of low-resolution images. Each low-resolution image is illuminated by a different light source or group of light sources from the array. The captured low-resolution images are then stitched together in the Fourier domain to generate a high-resolution image.
[0004] FPM offers several advantages over conventional microscopy, such as a significantly higher spatial-bandwidth product, a simple and low-cost setup (with minimal mechanical actuation), and a small footprint. However, due to the number of images to be captured, FPM suffers from long image acquisition times, which limits its applicability for imaging moving samples. In some applications, reconstructing a high-resolution image from the captured low-resolution images can also be very time-consuming.
[0005] Therefore, a need exists for improved methods and apparatus for FPM. Summary of the Invention
[0006] In some embodiments, a method of Fourier stacking microscopy (FPM) is provided, the method comprising: obtaining an image of a sample using an FPM system; storing the image in a memory; uploading the image to a graphics processing unit (GPU); and performing FPM reconstruction using the GPU to generate a reconstructed image, wherein the FPM reconstruction comprises performing parts of the FPM reconstruction in parallel on the GPU to reduce the FPM reconstruction time.
[0007] In some embodiments, a Fourier stack imaging system includes: a plurality of light sources configured to emit light onto a sample location; an optical system configured to image at least a portion of a sample positioned at the sample location; an image capture device configured to capture images of the sample through the optical system under different light conditions provided by the plurality of light sources; a processor; a GPU in communication with the processor; and a memory coupled to the processor. The memory includes computer-executable instructions stored therein that, when executed by the processor, cause the processor to: (a) obtain an image of the sample positioned at the sample location; (b) store the image in the memory; (c) upload the image to the GPU; and (d) initiate Fourier stacking reconstruction using the GPU to generate a reconstructed image, wherein the FPM reconstruction includes performing portions of the FPM reconstruction in parallel on the GPU to reduce FPM reconstruction time.
[0008] A system of one or more computers may be configured to perform specific operations or actions by having software, firmware, hardware, or a combination thereof installed on the system that, when in operation, causes the system to perform the actions. One or more computer programs may be configured to perform specific operations or actions by including instructions that, when executed by a data processing device, cause the device to perform the actions.
[0009] Other features and aspects of the present invention will become more fully apparent from the following detailed description, the appended claims and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1A An example Fourier ptychography microscopy (FPM) system provided in accordance with an embodiment of the present disclosure is shown.
[0011] Figure 1B Shown is a diagram of an embodiment according to the present invention. Figure 1A An example light source array for use with an FPM system.
[0012] Figure 2 An example method of Fourier stacking microscopy according to embodiments provided herein is shown.
[0013] Figures 3A to 3B An example low-resolution image and a linear array for storing image data from the low-resolution image are shown according to embodiments provided herein.
[0014] Figure 4 Shown are example low-resolution images compared to FPM-reconstructed high-resolution images according to embodiments provided herein.
[0015] Figure 5AAn example method for using a GPU for FPM reconstruction according to embodiments provided herein is shown, in which a series of kernel calls and fast Fourier transform (FFT) calls are used.
[0016] Figure 5B An example method for using a GPU for FPM reconstruction according to embodiments provided herein is shown, in which a single kernel call is used.
[0017] Figure 6 Another example method for using a GPU for FPM reconstruction according to the embodiments provided herein is shown, in which a series of kernel calls and FFT calls are used.
[0018] Figure 7 Another example method for using a GPU for FPM reconstruction according to embodiments provided herein is shown, in which a single kernel call is used. DETAILED DESCRIPTION
[0019] Independent of grammatical usage of the term, individuals who identify as male or female are also included within the term.
[0020] As previously mentioned, while FPM offers many advantages, its use may be limited in some applications due to the considerable time required to obtain results using this technique. The primary delays associated with FPM include the time required to capture a large number of low-resolution images and the time required to reconstruct a high-resolution image from the captured low-resolution images (e.g., which may be several orders of magnitude longer than the image capture time). The embodiments provided herein can significantly reduce FPM image processing time, thereby enabling FPM to be used in a wider range of applications (e.g., any application that benefits from faster results, such as clinical testing for medical diagnosis / treatment, etc.).
[0021] According to some embodiments, one or more graphics processing units (GPUs) are utilized to perform processing of low-resolution images to reconstruct high-resolution images using FPM. Using one or more GPUs can significantly reduce the time bottleneck associated with image reconstruction during FPM without compromising the image quality, scalability, or portability of the FPM system. Such an approach is scalable and portable and can provide performance improvements of up to 500 to 10,000 times compared to serial central processing unit (CPU) FPM implementations.
[0022] In an example embodiment, an image of a sample is obtained using a Fourier stacking microscopy (FPM) system. For example, the sample can be illuminated using an array of light sources (e.g., a light emitting diode (LED) array) while being imaged using low-resolution optical elements and an image capture device such as a camera. Each captured image can be stored in a memory (e.g., a memory associated with a central processing unit (CPU) that communicates with the image capture device), for example, by storing image data for each captured image in the memory. Thereafter, the image (e.g., as image data) is uploaded to a graphics processing unit (GPU). FPM reconstruction can then be performed using the GPU to generate a reconstructed image. For example, portions of the FPM reconstruction can be performed in parallel on the GPU to reduce the FPM reconstruction time.
[0023] Refer to the following Figures 1A to 7 These and other embodiments are described herein.
[0024] Figure 1A An example Fourier ptychography microscopy (FPM) system 100 is shown, provided in accordance with an embodiment of the present disclosure. Figure 1A , the FPM system 100 includes a light source array 102 having a plurality of light sources 102 a to 102 n configured to emit light onto a sample location 104 .
[0025] The optical system 106 is configured to image at least a portion of a sample 108 positioned at the sample position 104. Figure 1A As shown, the image capture device 110 is configured to capture images (eg, low resolution images 112 a - 112 n ) of the sample 108 through the optical system 106 under different light conditions provided by the plurality of light sources 102 a - 102 n of the light source array 102 .
[0026] A computer 114 having a processor 116 can be coupled to the image capture device 110 and receive images (e.g., low-resolution images) captured by the image capture device 110 for storage in a memory. In some embodiments, the images can be stored in a memory 118 (e.g., RAM, a hard drive, and / or another memory type) associated with the processor 116. Alternatively or additionally, the image data can be stored in an external memory 120 (e.g., local external memory, a remote storage device, a cloud storage device, or a combination thereof).
[0027] One or more graphics processing units (GPUs) 122a to 122n may be coupled to the processor 116, as described further below. Any suitable number of GPUs may be used (e.g., 1, 2, 5, 10, etc.). A display 124 having a user interface 126 may be coupled to the processor 116 and / or the one or more GPUs 122a to 122n, e.g., for displaying low-resolution images, reconstructed high-resolution images, and / or the like.
[0028] The light source array 102 can include a uniform or non-uniform array of light sources 102a-102n, which can be controlled by the processor 116 or another suitable processor, microprocessor, controller, microcontroller, digital signal processor (DSP), field programmable gate array (FPGA) configured to perform as a microcontroller, etc. In some embodiments, the light sources 102a-102n of the light source array 102 can be individually controlled and operated individually or in combination with one or more light sources 102a-102n. Example light sources 102a-102n can include light emitting diodes (LEDs), monochromatic or single-bandwidth emitting light sources, multi-bandwidth light sources (e.g., RGB LEDs), superluminescent LEDs, laser diodes (particularly semiconductor laser diodes), thermal emitters, fiber-based light sources, etc. All light sources 102a to 102n may be identical, or one or more light sources 102a to 102n may differ in at least one of the following characteristics: wavelength, spectral bandwidth, spatial emission characteristics, temporal emission characteristics such as continuous operation or pulsed operation, coherence parameters such as temporal coherence degree and / or spatial coherence degree, brightness or degree and / or the like.
[0029] In some embodiments, a light source array 102 having 256 individually controllable LEDs can be employed in an xy grid (e.g., 16×16 LEDs), with each LED spaced approximately 1 mm to 10 mm apart and emitting at approximately 0.4 microns to 0.7 microns, as shown. Figure 1B In one particular embodiment, the LEDs may be spaced approximately 2.5 mm to 3.5 mm apart and utilize wavelengths of 0.45 microns, 0.51 microns, and / or 0.62 microns. Other light source array arrangements, numbers of light sources, types of light sources, and / or emission wavelengths may be employed. As mentioned, although the processor 116 may be configured to Figure 1B 1 is shown as controlling the light source array 102 , but in other embodiments, a different processor or other control mechanism may be used to control the operation of the light source array 102 .
[0030] Optical system 106( Figure 1A) may include, for example, an optical objective lens 106a and a focusing lens 106b. Other optical components may be used. As described, one of the benefits of FPM is that FPM allows the use of low-cost, low-resolution optical components. In some embodiments, the optical objective lens 106a may have a numerical aperture (NA) of approximately 0.05 to 0.9. Other NA optical objective lenses may be used. In one or more embodiments, the focusing lens 106b may be a tube lens such as an achromatic tube lens or another suitable lens. The image capture device 110 may include any suitable imaging device capable of imaging a sample through the optical system 106, such as a CMOS sensor or the like. Example pixel sizes may be in the range of about 1 micron to about 10 microns, although other pixel sizes may be used.
[0031] In some embodiments, the processor 116 may be a central processing unit (CPU). In other embodiments, the processor 116 may include and / or be implemented as one or more other computing resources, such as, but not limited to, a microprocessor, a microcontroller, an embedded microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA) configured to perform as a microcontroller, etc. The computer 114 may include any suitable computing device, such as a tablet computer, a laptop computer, a desktop computer, a server, etc.
[0032] Memory 118 and / or 120 may be any suitable type of memory, such as, but not limited to, one or more of volatile memory and / or non-volatile memory (e.g., RAM, DRAM, SRAM, cache, hard drive, combinations thereof, etc.). In other words, memory 118 and / or 120 may include more than one type of memory. Memory 118 and / or 120 may have a plurality of instructions stored therein that, when executed by processor 116, cause processor 116 to perform various actions specified by one or more of the stored instructions. Code and data may be stored in a first type of memory (e.g., a hard drive) and transferred to a second type of memory for execution (e.g., RAM). In some embodiments, memory 118 and / or 120 may include one or both.
[0033] GPUs 122a through 122n may include any suitable graphics processing unit. In some embodiments, one or more of GPUs 122a through 122n may include an RTX 30 series GPU, such as an RTX 3080 or 3090 available from NVIDIA Corporation of Santa Clara, California. Other GPUs may be used. Each GPU 122a through 122n may include memory 128a through 128n, respectively, such as SRAM, DRAM, or the like.
[0034] The display 124 may include any suitable display, such as a light emitting diode (LED) display, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, etc. The user interface 126 may include, for example, a display screen or touch panel and / or screen, an audio speaker, a microphone, or any combination thereof. In some embodiments, the user interface 126 may be controlled by the processor 116, and the functionality of the user interface 126 may be implemented at least in part by computer-executable instructions (e.g., program code or software) stored in the memory 118 and / or executed by the processor 116.
[0035] Figure 2 An example method 200 of Fourier stacking microscopy according to embodiments provided herein is shown. Figure 2 In block 202, an image of a sample is obtained using the FPM system. For example, a sample 108 can be placed at a sample position 104 and illuminated using one or more light sources from the light source array 102. The sample 108 can then be imaged using the image capture device 110 through the optical system 106. Specifically, the image capture device 110 can capture a low-resolution image 112a of the sample 108. For each subsequent image 112b through 112n, the light source array 102 can be adjusted so that the sample 108 is illuminated using a different light source 102a through 102n or a different arrangement of the light sources 102a through 102n. In one or more embodiments, the intensity of the one or more light sources can be varied.
[0036] In some embodiments, for each high-resolution image generated by the FPM system 100, approximately 40 to 400 low-resolution images may be obtained and processed. Fewer or more low-resolution images may be used. The pixel count of each low-resolution image may be within a wide range. In some embodiments, for example, the pixel count may be approximately 3000 x 4000 pixels per image (e.g., depending on the sensor size employed within the image capture device 110). Larger or smaller image pixel counts may be employed.
[0037] After obtaining an image of the sample, the image is stored in memory in block 204. In some embodiments, each low-resolution image 112a-112n is stored in memory 118 (associated with processor 116). Alternatively, the low-resolution images 112a-112n can be stored in external memory 120.
[0038] To simplify FPM reconstruction of a high-resolution image from a set of low-resolution images using the GPU of the FPM system 100, the image data of the low-resolution images 112a to 112n may be flattened and aligned while being stored in memory, as described below with reference to Figures 3A to 3B Descriptive.
[0039] Figures 3A to 3B 1 and 2. Example low-resolution images 112a to 112n and linear arrays (e.g., linear arrays 302a to 302n within memory 118) for storing image data from the low-resolution images 112a to 112n are shown according to embodiments provided herein. Figure 3A As shown, each low-resolution image 112a to 112n may be divided into a plurality of tiles or regions of interest (ROIs) ( Figure 3B ROI in an example embodiment of 11 To ROI nm ). For example, in some embodiments, each low-resolution image can be divided into approximately 200 to 400 ROIs per image. Other numbers of ROIs may be used. Similarly, in some embodiments, each region of interest within an image can include a predefined number of pixels, such as 192×192 pixels per ROI or another number of pixels. Furthermore, in some embodiments, at least some adjacent regions of interest can share one or more image pixels. As will be further described below, according to embodiments provided herein, one or more of the GPUs 122a to 122n can be used to process each ROI individually and in parallel to significantly reduce the time to reconstruct a high-resolution image from a low-resolution image. For example, each pixel of each ROI can be processed in parallel using one or more GPUs.
[0040] Return to Figure 3A , four ROIs 304a, 304b, 304c, and 304d are shown with shading in the low-resolution images 112a, 112b, and 112n. These ROIs represent two-dimensional data to be stored in a one-dimensional array 302a to 302n (e.g., within the memory 118, although other memory locations such as the external memory 120 may be used). To achieve this, each ROI is flattened when stored in the memory. For example, the image data of the ROI 304a of the image 112a is shown as flattened when stored in the linear array 302a of the memory 118 (e.g., as depicted by the ROI 304a being represented as a square in the low-resolution image 112a and the corresponding image data 306a being represented as a rectangle in the linear array 302a). Additionally, the ROI image data is aligned within the linear arrays 302a to 302n. For example, the image data for each ROI is arranged contiguously within each linear array. As Figure 3A As shown, within memory 118, image data for ROI 304a is immediately adjacent to image data for ROI 304b within linear array 302a. In other words, within each linear array, the image data is stored together, without intermediate data stored between the ROI image data. Thus, in some embodiments, storing images in memory may include defining multiple ROIs within each image and flattening and aligning the image data ROI-by-ROI in memory. Other methods may be employed for arranging the image data within memory. In some embodiments, additional processing or pre-processing of the data may be performed, for example, to reduce noise in the image data, remove artifacts from the image data, and / or remove stray light intensity from the image data.
[0041] Return to Figure 2 In the method 200, after storing the low-resolution image in the memory, in block 206, the captured image is uploaded to the GPU. This may include transferring the image data from the CPU memory to the GPU memory. For example, the linear arrays 302a to 302n ( Figure 3A ) is transferred to the GPU memory 128a ( Figure 1A ) (e.g., for storage in one or more linear arrays, not shown), or, if an additional GPU is employed, transferred to one or more other GPU memories. In some embodiments, additional pre-processing of the data can be performed within GPU 122a (e.g., and / or another GPU, if employed), for example, to reduce noise in the image data, remove artifacts from the image data, and / or remove stray light intensity from the image data.
[0042] After uploading the image to the GPU, in block 208, FPM reconstruction is performed using the GPU to generate a reconstructed image, wherein portions of the FPM reconstruction are performed in parallel on the GPU to reduce FPM reconstruction time (e.g., rather than performing the FPM reconstruction serially using one or more CPUs). For example, according to embodiments provided herein, FPM reconstruction is performed for each region of interest in the low-resolution image independently of other regions of interest within the image, and a reconstructed image is generated based on the FPM reconstruction of each region of interest within each image. For example, FPM reconstruction can be performed for each region of interest in each image in parallel using one or more GPUs. For example, FPM reconstruction can be performed for each region of interest in each image in parallel using GPU 122a (and / or one or more other GPUs 122b to 122n, if desired) by performing pixel-level parallelization (as further described below).
[0043] In many cases, FPM reconstruction performed on one or more low-resolution ROIs generates one or more FPM-reconstructed ROIs having a greater number of pixels than the low-resolution ROIs employed during reconstruction. Figure 3A An example linear array 308 (e.g., of the memory 128a of the GPU 122a) is shown containing reconstructed high-resolution ROI image data 310a, 310b, 310c, and 310d generated by FPM reconstruction of the low-resolution ROIs of the low-resolution images 112a to 112n (e.g., including the low-resolution ROI image data 306a, 306b, 306c, and 306d of the linear array 302a, respectively). As can be seen in the linear array 308, the high-resolution ROI image data 310a, 310b, 310c, and 310b contain more data (e.g., more pixels) than the low-resolution ROI image data 306a, 306b, 306c, and 306d. This is also shown in FIG. Figure 4 As shown in Figure 4 An example low-resolution image 112n is shown compared to an FPM-reconstructed high-resolution image 412, in which individual ROIs can be larger than corresponding ROIs in the low-resolution image 112n. Likewise, the overall length and width of the high-resolution image 412 can be increased compared to the length and width of the low-resolution image used to generate the high-resolution image 412 via FPM reconstruction.
[0044] Refer to the following Figures 5A to 7 An example implementation for performing FPM reconstruction is described. For example, Figure 5A and Figure 6 5 and 600, respectively, illustrate example methods 500a and 600 for FPM reconstruction using a GPU, wherein a series of kernel calls and fast Fourier transform (FFT) calls (e.g., forward FFT calls and / or inverse FFT calls) are used, while Figure 5B and Figure 7 Example methods 500b and 700 are shown, respectively, for FPM reconstruction using a GPU, where a single kernel call is employed. As used herein, in some embodiments, a series of FFT calls may include both forward FFT calls and inverse FFT calls.
[0045] Many algorithms have been proposed to implement Fourier ptychographic microscopy, such as the alternating projection method described in R. W. Gerchberg and W. O. Shaxton, “A practical algorithm for the determination of phase from image and diffraction plane pictures”, Optik, Bd. 35, pp. 227-246, (1972) and X. Ou, G. Zheng and C. Yang, “Embedded pupil function recovery for Fourier ptychographic microscopy”, Optics Express, Bd. 22, Nr. 5, pp. 4960-4972, 3 (2014) (hereinafter “Ou et al.”), and L. Bian, J. Suo, G. Zheng, K. Guo, F. Chen and Q. Dai, “Fourier ptychographic reconstruction using Wirtinger flow optimization”, Optics Express, Bd. 22, Nr. 5, pp. 4960-4972, 3 (2014) (hereinafter “Ou et al.”). Express, Bd.23, Nr.4, pp.4856-4866, 2 (2015) and L. Bian, J. Suo, J. Chung, X. Ou, C. Yang, F. Chen and Q. Dai, "Fourier ptychographic reconstruction using Poissonmaximum likelihood and truncated Wirtinger gradient", Scientific Reports, Bd.6, Nr.1, p.27384, 7 (2016). Other example FPM algorithms include the maximum likelihood estimation formula in T. Nguyen, Y. Xue, Y. Li, L. Tian and G. Nehmetallah, "Deep learning approach for Fourier ptychographymicroscopy", Optics Express, Bd.26, Nr.20, pp.26470-26484, 10 (2018) and Y. Rivenson, Y. Zhang, H. D. Teng and A. Ozcan, "Phase recovery and holographic image reconstruction using deep learning in neural networks", Light: Science and Applications, Bd. 7, Nr. 2, p. 17141, 2 (2018) . Although some embodiments described herein relate to pupil function recovery algorithms, such as those described by Ou et al., it should be understood that other algorithms may be employed.
[0046] Reference Figure 5A In block 502, a low-resolution, "raw" image of the sample is obtained. For example, the sample 108 can be placed at the sample position 104 and illuminated using one or more light sources from the light source array 102. The sample 108 can then be imaged using the image capture device 110 through the optical system 106. Specifically, the image capture device 110 can capture a low-resolution image 112a of the sample 108. For each subsequent image, the light source array 102 can be adjusted so that the sample 108 is illuminated using a different light source 102a to 102n or a different arrangement of light sources 102a to 102n. In some embodiments, for each high-resolution image generated by the FPM system 100, approximately 40 to 400 low-resolution images can be obtained and processed. Fewer or more low-resolution images can be used.
[0047] Once the low resolution image is obtained, in block 504, the image is flattened and aligned in memory (e.g., image data from the image is flattened and aligned in memory). Figures 3A to 3B As described, in some embodiments, each low-resolution image (e.g., low-resolution images 112a to 112n) can be divided into multiple regions of interest (ROIs). For example, in some embodiments, each low-resolution image can be divided into approximately 200 to 400 ROIs per image. Other numbers of ROIs can be used. Generally, the size of the ROIs used can be limited by physical constraints associated with the spatiotemporal coherence of the illumination and the applicability of spatially calibrated or geometrically derived parameters. Additionally, spatially correlated aberrations associated with the optical device (e.g., optical system 106) may also impose constraints on the size of the region of interest that can be used to obtain a high-quality reconstruction.
[0048] According to the embodiments provided herein, Fourier stack imaging reconstruction can be performed independently for each ROI. This feature makes the FPM reconstruction method highly suitable for parallelization. However, the choice of parallelization strategy and memory management is important for optimizing resource usage and performance.
[0049] The ROI represents two-dimensional data to be stored in a one-dimensional array within the memory 118 (or other memory location such as the external memory 120). To achieve this, each ROI is flattened when stored in the memory. For example, Figure 3A As shown, the image data of the ROI 304a of the low-resolution image 112a is shown as flattened (e.g., reduced from 2D data to 1D data) when stored in the linear array 302a of the memory 118 (e.g., as depicted by the ROI 304a being represented as a square in the low-resolution image 112a and the corresponding image data 306a being represented as a rectangle in the linear array 302a). Additionally, the ROI image data is aligned when stored in the memory (e.g., within the linear arrays 302a through 302n). For example, the image data for each ROI is arranged consecutively within each linear array. Figure 3A As shown, within memory 118, image data for ROI 304a is immediately adjacent to image data for ROI 304b within linear array 302a, with no intermediate data stored between the ROI image data. Other methods may be employed for arranging the ROI image data within memory. As previously described, in one or more embodiments, processor 116 may be used to perform additional pre-processing of the data, for example, to reduce noise in the image data, remove artifacts from the image data, and / or remove stray light intensity from the image data.
[0050] In some embodiments, after the images are flattened and aligned, the flattened and / or aligned low-resolution images are uploaded to the GPU (e.g., for use during FPM reconstruction of a high-resolution image based on the captured low-resolution image) in block 506. For example, the image data can be transferred from a memory of the CPU (e.g., the memory 118 of the processor 116) to a memory of the GPU (e.g., the memory 128a of the GPU 122a). As a specific example, the image data stored in the linear arrays 302a to 302n of the memory 118 of the processor 116 can be transferred to the GPU memory 128a of the GPU 122a (e.g., for storage in one or more linear arrays, not shown), or to one or more other GPUs if additional GPUs are employed.
[0051] In block 508, in some embodiments, a pupil function representing an optical system (e.g., optical system 106) can be uploaded to a GPU (or multiple GPUs) employed for FPM reconstruction. Alternatively, the pupil function can be reconstructed from the low-resolution images 112a to 112n. For example, a pupil function recovery algorithm such as that described by Ou et al. can be employed to recover the pupil function. As a specific example, the embedded pupil function recovery (EPRY) algorithm of Ou et al. can be employed during FPM reconstruction to recover the optical system (e.g., Figure 1A The pupil function of the optical system 106 for imaging the sample 108 is shown in FIG.
[0052] Refer again Figure 5A , in block 510, memory within the GPU (or GPUs) is allocated for FPM reconstruction output and intermediate variables (e.g., variables that may be used during FPM reconstruction but are not output with the reconstructed image). As an example, one or more linear arrays within GPU 122a (or another GPU) may be allocated for FPM reconstruction and / or intermediate variables.
[0053] In block 512, the image data for the low-resolution image may be pre-processed within the GPU (or GPUs). For example, pre-processing of the data may be performed within GPU 122a (and / or another GPU, if employed) to reduce noise in the image data, remove artifacts from the image data, and / or remove stray light intensity from the image data (e.g., assuming that processor 116 has not already performed pre-processing of the image data).
[0054] In block 514a, FPM reconstruction is performed by the GPU (or multiple GPUs) using a series of kernel calls, batch fast Fourier transform (FFT) calls (e.g., forward FFT calls and / or inverse FFT calls), and / or memory copy calls. In some embodiments, the memory 118 ( Figure 1A ) may include computer-executable instructions that, when executed by processor 116, cause processor 116 to perform a series of kernel calls and FFT calls on GPU 122a (and / or any other GPU employed), and in some embodiments, to perform one or more memory copy calls on GPU 122a (and / or any other GPU employed). Figure 6An example sequence of kernel calls, batch FFT calls, and memory copy calls is described. The series of kernel calls, batch FFT calls, and / or memory copy calls causes GPU 122a to perform FPM reconstruction on the ROI within each low-resolution image 112a to 112n. For example, FPM reconstruction of each ROI can be performed independently of other ROIs within the image and in parallel using at least one GPU (e.g., one or more of GPUs 122a to 122n). A high-resolution reconstructed image can then be generated based on the FPM reconstruction of each ROI within each image. This parallel processing of ROIs within the GPU significantly reduces FPM reconstruction time when compared to serial processing using one or more CPUs and enables pixel-level parallelization during ROI reconstruction. An example FPM reconstruction algorithm and process are described further below.
[0055] The reconstructed ROI for the reconstructed high-resolution image may be stored in GPU memory (e.g., memory 128a of GPU 122a) before output, for example, by Figure 3A The linear array 308 is shown.
[0056] In block 516, after FPM reconstruction, the FPM reconstructed image is output. For example, the image data generated by GPU 122a (and / or any other GPU employed) during FPM reconstruction for each ROI may be combined into a high-resolution reconstructed image (e.g., such as Figure 4 and via the display 124 and / or the user interface 126 ( Figure 1A In some embodiments, GPU 122 a may output the high-resolution image directly to display 124 , while in other embodiments, GPU 122 a may transfer the high-resolution image data to memory 118 of processor 116 and processor 116 may output the high-resolution image using display 124 and / or user interface 126 .
[0057] Reference Figure 5B , blocks 502 to 512 and block 516 of the example method 500b (e.g., for employing a GPU during FPM reconstruction) can be similar or identical to blocks 502 to 512 and 516 of method 500a. However, block 514a of method 500a is replaced by block 514b in method 500b. Specifically, in block 514b of method 500b, FPM reconstruction is performed using a single kernel call. In other words, FPM reconstruction can be performed on each ROI of the low-resolution images 112a to 112n in parallel (e.g., using one or more GPUs 122a to 122n) using a single kernel call. In some embodiments, the memory 118 ( Figure 1A) may include computer executable instructions that, when executed by the processor 116, cause the processor 116 to execute a single kernel call to perform FPM reconstruction of a high resolution image from the low resolution images 112a to 112n, as described below with reference to Figure 7 Following the single kernel call of blocks 502 through 512 and block 514b, in block 516, the FPM reconstructed high-resolution image may be output.
[0058] Reference Figure 6 112n). With method 600, in block 602, FPM reconstruction can begin by making an initial guess for a sample spectrum and pupil function. The initial guess for the sample spectrum is a frequency-domain estimate of a high-resolution image that will be produced by FPM reconstruction using low-resolution images (e.g., low-resolution images 112a through 112n). Any suitable initial guess can be employed, and such initial guess can depend on the particular FPM algorithm employed during FPM reconstruction. In some embodiments, the initial guess for the sample spectrum can be based on one or more of the low-resolution images (e.g., one or a combination of the low-resolution images, an initial guess heuristically defined based on one or more low-resolution images, etc.). In other embodiments, the initial guess for the sample spectrum can be based on randomized data.
[0059] In embodiments in which a pupil function of an optical system (e.g., optical system 106) is measured, estimated, or otherwise determined prior to FPM reconstruction, an initial guess of the pupil function is not required, and the pupil function can be uploaded for use during FPM reconstruction (e.g., to GPU 122a and / or any other GPU employed). There may be instances in which different pupil functions may be employed for different low-resolution images and / or different ROIs within a low-resolution image. In such instances, FPM reconstruction may include uploading one or more pupil functions to the GPU (e.g., and employing the one or more pupil functions during FPM reconstruction).
[0060] If the pupil function is to be recovered (e.g., determined) during the FPM reconstruction process, the initial guess for the pupil function may depend on the particular pupil function recovery algorithm employed, available pupil function pre-characterization information, known aberrations of the employed optical system, etc. For example, for the Embedded Pupil Function Recovery (EPRY)-FPM algorithm of Ou et al., the initial pupil function may be a low-pass filter with an applicable shape (e.g., a circle). In general, any suitable binary mask or calculated initial pupil function may be employed.
[0061] Return to Figure 6After an initial guess of the sample spectrum and / or pupil function is made, in block 604, a memory copy is performed to retrieve pre-processed low-resolution image data for transfer to one or more GPUs and use during FPM reconstruction. For example, processor 116 may initiate a memory copy operation to transfer low-resolution image data from memory 118 of processor 116 to memory 128a of GPU 122a. The amount of pre-processed image data retrieved may depend, for example, on the size of memory available within the GPU employed (e.g., the size of memory 128a within GPU 122a). In some embodiments, the memory copy operation may retrieve all ROI data associated with a particular low-resolution image (e.g., low-resolution image 112a). In other embodiments, GPU hardware with sufficient computational power may implement optimized paging functionality to retrieve data as needed. As described, the pre-processed image data may include low-resolution image data that has been flattened, aligned, and / or processed to reduce noise, remove artifacts, or remove stray light intensity from the image data.
[0062] After the memory copy operation, in block 606, a first kernel (e.g., kernel 1) is executed in which the current estimate of the high-resolution image (e.g., a portion of the sample spectrum) and the pupil function are combined (e.g., multiplied) in the Fourier domain (for convenience, the combined high-resolution image and pupil function are referred to as the estimated "exit wave," as described in Ou et al.). In block 608, the Fourier domain estimated exit wave (e.g., the combined estimated high-resolution image and pupil function) is converted to the real domain by performing a batch inverse FFT to generate an estimated exit wave at the imaging device (e.g., image capture device 110). For example, the processor 116 may initiate a first kernel call (in block 606) to generate the estimated exit wave (e.g., referred to as a forward simulation in the Fourier domain) and then initiate a batch inverse FFT (block 608) to convert the estimated exit wave to the temporal (real) domain. During batch inverse FFT, GPU 122a (and / or any other GPUs employed) may perform an inverse FFT on each ROI of the estimated exit wave in parallel (eg, in a batch manner).
[0063] In block 610, a second kernel (e.g., kernel 2) is executed in which an intensity correction is applied to the estimated exit wave. For example, the estimated exit wave can be updated using the low-resolution image data retrieved in block 604 (e.g., the estimated exit wave is intensity corrected based on the current low-resolution image data provided to the GPU in block 604). In some embodiments, the operations in block 604 can be performed asynchronously (e.g., as long as the required data for intensity correction in block 610 exists). The correction of the exit wave depends on the FPM reconstruction algorithm used. For example, in Ou et al., the intensity correction includes replacing the modulus of the estimated (e.g., simulated) exit wave with the square root of the low-resolution image intensity. In some embodiments, the intensity correction can include identifying the estimated exit wave for intensity correction and each pixel within each ROI of the current low-resolution image, and performing the intensity correction on the estimated exit wave pixel by pixel in parallel on one or more GPUs.
[0064] After intensity correction, in block 612, the intensity-corrected estimated exit wave is converted to the Fourier domain (e.g., via a batch forward FFT initiated by processor 116 and performed on a pixel-by-pixel basis for each ROI of the intensity-corrected estimated exit wave in parallel using GPU 122a and / or another GPU), and in block 614, the sample spectrum and pupil function are updated based on the intensity-corrected estimated exit wave (e.g., via a third kernel call (kernel 3) initiated by processor 116 and executed by GPU 122a). For example, as described in Ou et al., the sample spectrum and pupil function can be updated based on the difference between the exit wave before intensity correction and the exit wave after intensity correction in block 610. As mentioned, in embodiments in which the pupil function is predetermined, updating the pupil function is optional.
[0065] Blocks 604 to 614 are repeated for each low-resolution image (e.g., each low-resolution image 112a to 112n). In other words, for each low-resolution image or a portion of each low-resolution image (e.g., a predetermined number of ROIs based on the GPU architecture and GPU memory availability), the low-resolution image data may be uploaded to the GPU (e.g., in block 604), an estimated exit wave may be created for the current estimated high-resolution image (e.g., a portion of the sample spectrum) and the pupil function (e.g., in block 606), the estimated exit wave may be transformed into the real domain (e.g., in block 608), the estimated exit wave may be intensity corrected using the low-resolution image data uploaded to the GPU (e.g., in block 610), the intensity-corrected estimated exit wave may be transformed into the Fourier domain (e.g., in block 612), and the estimated high-resolution image and pupil function may be updated based on the intensity-corrected estimated exit wave (e.g., in block 614).
[0066] In block 616 , it is determined whether all low-resolution images have been used to update the estimated high-resolution image (and pupil function). If not, blocks 604 to 614 are repeated for the remaining low-resolution images; otherwise, the method 600 proceeds to decision block 618 .
[0067] In block 618, a determination is made as to whether the estimated high-resolution image has converged to an acceptable level or whether a maximum number of iterations has been performed. For example, a simulated low-resolution image based on the estimated high-resolution image can be compared to one or more of the low-resolution images to confirm that the estimated high-resolution image accurately depicts the details of the low-resolution image. In some embodiments, this can include simulating the low-resolution image based on the estimated high-resolution image (e.g., intentionally reducing the details within the high-resolution image to approximate the level of detail within the low-resolution image). Assuming the estimated high-resolution image has not yet converged relative to the low-resolution image, the updating of the estimated high-resolution image and pupil function (e.g., in blocks 604 to 614) can be repeated using each low-resolution image. Repeated updating of the estimated high-resolution image and pupil function can continue until the estimated high-resolution image converges or until a maximum number of iterations has been reached (e.g., the maximum number of times the estimated high-resolution image can be updated with the same set of low-resolution images). In some embodiments, the maximum number of iterations can be in the range of 2 to 10 iterations, although other numbers of iterations can be used. In one or more embodiments, for example, the convergence check can be omitted, and a predetermined number of iterations can be performed based on a priori information.
[0068] Assuming that the estimated high-resolution image has reached sufficient convergence relative to the low-resolution image or the maximum number of iterations has been reached (e.g., as determined by the processor 116 in block 618), the high-resolution reconstructed image is output in block 620. For example, each of the ROIs of the estimated high-resolution image (e.g., such as stored in a linear array in the memory 128a of the GPU 122a) can be combined and output as a reconstructed high-resolution image. As described, in some embodiments, the GPU 122a or processor 116 can cause or facilitate the high-resolution image data to be stored (e.g., in the memory 118 and / or 120) and / or displayed via the user interface 126 (e.g., on the display 124) Figure 1A ) Output high-resolution images. In some embodiments, the reconstructed image data can be transmitted to an external storage device (eg, a local external storage device, a cloud storage device, etc.).
[0069] Figure 7Another example method 700 for FPM reconstruction using a GPU according to embodiments provided herein is shown, in which a single kernel call is used. Figure 7 In block 701, FPM reconstruction may begin, where a memory copy is performed to retrieve pre-processed low-resolution image data for all low-resolution images to be transferred to one or more GPUs and used during FPM reconstruction (e.g., all image data for low-resolution images 112a through 112n). For example, processor 116 may initiate a memory copy operation to transfer low-resolution image data from memory 118 of processor 116 to memory 128a of GPU 122a. The low-resolution image data may include data for a portion of the total number of ROIs that fits within the GPU memory availability. If the GPU does not have sufficient memory to store all the low-resolution image data, subsequent memory copy operations may be performed. In other embodiments, GPU hardware with sufficient computational power may implement optimized paging functionality to retrieve data as needed. As described, the pre-processed image data may include low-resolution image data that has been flattened, aligned, and / or processed to reduce noise, remove artifacts, or remove stray light intensity from the image data.
[0070] In block 702, an initial guess is made for the sample spectrum and pupil function. As described, the initial guess for the sample spectrum is a frequency domain estimate of the high-resolution image produced by FPM reconstruction using the low-resolution image. Any suitable initial guess for the sample spectrum (estimated high-resolution image) and / or pupil function may be used, and such initial guess may depend on the particular FPM algorithm and / or the particular pupil function recovery algorithm employed during the FPM reconstruction. Example initial guesses for the sample spectrum and pupil function are described above with reference to block 602 of method 600.
[0071] In embodiments in which the pupil function of an optical system (e.g., optical system 106) is measured, estimated, or otherwise determined prior to FPM reconstruction, an initial guess of the pupil function is not required, and the pupil function can be uploaded for use during FPM reconstruction (e.g., to GPU 122a and / or any other GPU employed).
[0072] After an initial guess of the sample spectrum and / or pupil function, a single kernel 703 (e.g., a giant kernel) is executed to perform FPM reconstruction based on the initial sample spectrum and pupil function. As described below, the single kernel 703 call may include blocks 704, 706, 710, and 712. In some embodiments, the processor 116 ( Figure 1A) can initiate (e.g., via a kernel call) execution of a single kernel 703 by GPU 122a and / or any other GPU employed. As will be described further below, the use of a single kernel call that processes all image data for all low-resolution images significantly reduces the time required for high-resolution image reconstruction.
[0073] Referring to the single kernel 703, in block 704, the current estimate of the high-resolution image (e.g., a portion of the sample spectrum) and the pupil function are combined (e.g., multiplied) in the Fourier domain to form an estimated exit wave as described above with reference to block 606 of method 600, and the Fourier domain estimated exit wave (e.g., the combined estimated high-resolution image and pupil function) is converted to the real domain by performing a batch inverse FFT (e.g., a custom FFT implementation on a GPU that can be embedded in a single kernel that will achieve near pixel-level parallelism) as described above with reference to block 608 of method 600 to generate the estimated exit wave at the imaging device (e.g., image capture device 110). Thus, block 704 performs similar operations to both blocks 606 and 608 of method 600.
[0074] In block 706, an intensity correction is applied to the estimated exit wave, as previously described with reference to block 610 of method 600. For example, the estimated exit wave may be updated (e.g., intensity corrected based on the low-resolution image data) using low-resolution image data from a first low-resolution image (e.g., low-resolution image 112a) of the low-resolution images. Thus, block 706 performs similar operations as block 610 of method 600.
[0075] After intensity correction, in block 708, the intensity-corrected estimated exit wave is converted to the Fourier domain (e.g., via a batch forward FFT performed pixel-by-pixel on each ROI of the intensity-corrected estimated exit wave in parallel using GPU 122a), and the sample spectrum and pupil function are updated based on the intensity-corrected estimated exit wave. For example, as described in Ou et al., the sample spectrum and pupil function can be updated based on the difference between the exit wave before intensity correction in block 706 and the exit wave after intensity correction in block 706. As mentioned, in embodiments in which the pupil function is predetermined, pupil function updating is optional. Thus, block 708 performs operations similar to both blocks 612 and 614 of method 600.
[0076] For each low-resolution image (e.g., each low-resolution image 112a to 112n), blocks 704 to 708 are repeated. In other words, for each low-resolution image or a portion of each low-resolution image (e.g., a predetermined number of ROIs based on the GPU architecture and GPU memory availability), an estimated exit wave may be created for the current estimated high-resolution image and pupil function and the estimated exit wave may be transformed into the real domain (in block 704), the estimated exit wave may be intensity corrected using the low-resolution image data uploaded to the GPU (in block 706), the intensity-corrected estimated exit wave may be transformed into the Fourier domain, and the estimated high-resolution image and pupil function may be updated based on the intensity-corrected estimated exit wave (e.g., in block 708).
[0077] In block 710, a determination is made as to whether all low-resolution images have been used to update the estimated high-resolution image (and pupil function). If not, blocks 704 through 708 are repeated; otherwise, in block 712, a determination is made as to whether the estimated high-resolution image has converged to an acceptable level or whether a maximum number of iterations has been performed (e.g., as previously described with reference to block 618 of method 600). Assuming that the simulated low-resolution image based on the estimated high-resolution image has not converged relative to the low-resolution image, the updating of the estimated high-resolution image and pupil function (e.g., in blocks 704 through 708) may be repeated using each low-resolution image. The repeated updating of the estimated high-resolution image and pupil function may continue until the estimated high-resolution image converges or until a maximum number of iterations has been reached (e.g., the maximum number of times the estimated high-resolution image may be updated with the same set of low-resolution images). As noted, in some embodiments, the maximum number of iterations may be in the range of 2 to 10 iterations, although other numbers of iterations may be used. In one or more embodiments, for example, the convergence check may be omitted, and a predetermined number of iterations may be performed based on a priori information.
[0078] Assuming the estimated high-resolution image has achieved sufficient convergence relative to the low-resolution image or the maximum number of iterations has been reached, the high-resolution reconstructed image is output in block 714. For example, each of the ROIs of the estimated high-resolution image (e.g., stored in a linear array in the memory 128a of the GPU 122a) may be combined and output as a reconstructed high-resolution image.
[0079] Implementations described herein may employ one or more GPUs to implement pixel-level parallelism during FPM reconstruction. The use of pixel-level parallelism may achieve significant performance improvements. GPU computing may use a single instruction multiple thread (SIMT) execution model, where kernels are designed to execute synchronously on each recruited thread. To further improve performance, implementations described herein may be designed to operate on all ROIs of an image (e.g., using pixel-level parallelism) in parallel within a single kernel call. Figure 6 The invention provides a large kernel (described in the method 600) to minimize memory copy operations and kernel calls. This improvement can achieve a greater than 500 times improvement in FPM rebuild time compared to a serial CPU implementation.
[0080] As described above, in some embodiments, a full physics model (e.g., via pupil function recovery) can be employed to improve high-resolution image quality. The systems and methods provided herein are portable to different GPU architectures and hardware configurations and can achieve high-throughput FPM imaging at a reduced implementation cost.
[0081] To further improve performance, the embodiments described herein can reduce memory copy operations and kernel calls by embedding the FFT calculation within another kernel and by using optimized thread synchronization, wherein the FFT calculation is performed with a single kernel call (e.g., Figure 7 The entire FPM reconstruction of all ROIs or portions of ROIs in all low-resolution images is performed using a CPU (described in method 700). Doing so also enables efficient use of shared memory, texture memory, local registers, and cache, minimizing DRAM accesses, which can have relatively high latency. This can achieve a greater than 10,000-fold improvement in FPM reconstruction time compared to a serial CPU implementation.
[0082] As described above, in some embodiments, a Fourier stack imaging system may include: a plurality of light sources configured to emit light onto a sample position; an optical system configured to image at least a portion of a sample positioned at the sample position; an image capture device configured to capture an image of the sample through the optical system under different light conditions provided by the plurality of light sources; a processor; a GPU in communication with the processor; and a memory coupled to the processor.
[0083] In one or more embodiments, the memory may include computer-executable instructions stored therein that, when executed by the processor, cause the processor to: obtain an image of a sample positioned at a sample location; store the image in the memory; upload the image to a GPU; and initiate FPM reconstruction using the GPU to generate a reconstructed image. The FPM reconstruction may include performing portions of the FPM reconstruction in parallel on the GPU to reduce FPM reconstruction time.
[0084] In some embodiments, the processor may be configured to control the operation of at least one of the plurality of light sources and the image capture device. The processor may also be configured to initiate FPM reconstruction by calling a kernel of the GPU. For example, FPM reconstruction may be initiated in response to a single kernel call or a series of kernel calls (e.g., multiple kernel calls) and FFT calls.
[0085] In some embodiments, a GPU is configured to generate a reconstructed image from a plurality of images by: for each image: defining a plurality of regions of interest within the image; performing FPM reconstruction on each region of interest independently of other regions of interest within the image; and generating a reconstructed image based on the FPM reconstruction of each region of interest within each image. Each region of interest within each image may include a predetermined number of pixels within the image. In some embodiments, at least some adjacent regions of interest may share one or more image pixels. Any suitable region of interest size may be employed. Factors that may affect the size of the regions of interest may include the optical system used, the computer architecture used, the number of available GPU cores, and the like. As described, in at least some embodiments, all regions of interest of an image and all pixels of each region of interest, and therefore all pixels of the image, may be processed synchronously during FPM reconstruction using the GPU.
[0086] The GPU can be configured to perform FPM reconstruction on each region of interest in each image in parallel (e.g., by performing pixel-level parallelization). The GPU can also be configured to perform FPM reconstruction on one or more regions of interest by selecting an FPM algorithm and applying the selected FPM algorithm to the one or more regions of interest. In some embodiments, the GPU can be configured to perform FPM reconstruction on one or more low-resolution regions of interest and generate one or more FPM-reconstructed regions of interest having a greater number of pixels (e.g., 2 to 4 times the number of pixels) than the one or more low-resolution regions of interest. Some reconstructed ROIs can have the same number of pixels as the low-resolution ROI used to generate the reconstructed ROI.
[0087] In some embodiments, a different pupil function may be used for each ROI of an image, while in other embodiments, one or more ROIs of an image may share a pupil function. In other embodiments, a common pupil function may be used for all ROIs of an image.
[0088] In one or more embodiments, a pupil function may be generated for a first reconstruction iteration and then used in subsequent reconstruction iterations, while in other embodiments, a pupil function may be generated during each reconstruction iteration (e.g., for each ROI or one or more ROIs).
[0089] The foregoing description discloses only example embodiments of the present invention; modifications to the above-disclosed apparatus and methods that fall within the scope of the present invention will be readily apparent to those skilled in the art. Therefore, while the present invention has been disclosed in conjunction with its example embodiments, it should be understood that other embodiments may fall within the spirit and scope of the present invention as defined by the following claims.
[0090] The elements and features recited in the appended claims may be combined in various ways to create new claims that also fall within the scope of the invention. Thus, although the dependent claims appended hereto depend solely on a single independent or dependent claim, it should be understood that these dependent claims may alternatively depend on alternatives to any preceding or following claim, whether independent or dependent. Such new combinations should be understood to form part of this specification.
[0091] Although the present invention has been described above with reference to various embodiments, it will be appreciated that many variations and modifications may be made to the described embodiments. Therefore, the foregoing description is intended to be illustrative rather than restrictive, and it will be appreciated that all equivalents and / or combinations of the embodiments are intended to be included in this specification.
Claims
1. A method of Fourier stacking microscopy (FPM), comprising: Obtain images of the sample using the FPM system; storing the image in a memory; uploading the image to a graphics processing unit (GPU); as well as FPM reconstruction is performed using the GPU such that a reconstructed image is generated, wherein the FPM reconstruction includes performing portions of the FPM reconstruction in parallel on the GPU such that FPM reconstruction time is reduced.
2. The method according to claim 1, wherein The storing of the image in the memory includes storing image data from the image in a memory of a central processing unit (CPU).
3. The method according to claim 2, wherein: The uploading of the image to the GPU includes transferring image data from a memory of the CPU to a memory of the GPU.
4. The method according to claim 1, wherein The execution of the FPM reconstruction includes: Execute a single kernel call; or Executes a series of kernel calls and fast Fourier transform (FFT) calls.
5. The method according to claim 1, wherein The obtaining of the image of the sample comprises: For each of the images, a light source array is employed such that the sample is illuminated using light generated by one or more light sources in the light source array.
6. The method according to claim 5, wherein: For each of the images, a different light source or a different arrangement of light sources is used to illuminate the sample.
7. The method according to claim 1, wherein The storing of the image in the memory includes flattening and aligning image data in the memory.
8. The method according to claim 7, wherein: The storing of the image in the memory comprises: defining a plurality of regions of interest within each of the images; and Image data is flattened and aligned region-of-interest in the memory.
9. The method according to claim 1, wherein The uploading of the image to the GPU includes transmitting image data to the GPU and pre-processing the image data within the GPU.
10. The method according to claim 9, wherein: Preprocessing the image data within the GPU includes reducing noise in the image data, removing artifacts from the image data, removing stray light intensity from the image data, or any combination thereof.
11. The method according to claim 1 , further comprising: The image data stored in the memory is preprocessed, the preprocessing comprising: reducing noise in the image data, removing artifacts from the image data, removing stray light intensity from the image data, or any combination thereof.
12. The method according to claim 1, wherein The performing of the FPM reconstruction comprises employing at least one pupil function representative of an optical system of the FPM system during the FPM reconstruction of the reconstructed image.
13. The method according to claim 12, wherein: The adopting of the at least one pupil function comprises: uploading the at least one pupil function to the GPU; and The at least one pupil function is employed during the FPM reconstruction of the reconstructed image.
14. The method according to claim 12, wherein: The adopting of the at least one pupil function comprises: reconstructing the at least one pupil function; and The at least one pupil function is employed during the FPM reconstruction of the reconstructed image.
15. The method according to claim 12, wherein: The employing of the at least one pupil function comprises: employing a first pupil function during the FPM reconstruction of a first region of interest of image data; and A second pupil function is employed during said FPM reconstruction of a second region of interest of the image data.
16. The method according to claim 1, wherein The execution of the FPM reconstruction includes: For each of the images: defining a plurality of regions of interest within the respective images; and performing FPM reconstruction on each region of interest in the plurality of regions of interest independently of other regions of interest in the plurality of regions of interest within the corresponding image; and The reconstructed image is generated based on the FPM reconstruction of each region of interest within each of the images.
17. The method according to claim 16, wherein Each region of interest of the plurality of regions of interest within each of the images comprises a predetermined number of pixels within the respective image.
18. The method according to claim 16, wherein At least some adjacent regions of interest among the plurality of regions of interest share one or more image pixels.
19. The method according to claim 16, wherein The performing of the FPM reconstruction of each of the plurality of regions of interest in each of the images includes performing FPM reconstruction on each of the plurality of regions of interest in parallel using the GPU.
20. The method according to claim 19, wherein The performing of the FPM reconstruction on each of the plurality of regions of interest in parallel using the GPU includes performing pixel-level parallelization.
21. The method according to claim 19, wherein The performing of the FPM reconstruction on each of the plurality of regions of interest in parallel using the GPU includes executing a single kernel call.
22. The method according to claim 16, wherein The performing of the FPM reconstruction of one or more regions of interest of the plurality of regions of interest comprises: Select the FPM algorithm; and The selected FPM algorithm is applied to the one or more regions of interest.
23. The method according to claim 22, wherein The FPM algorithm includes a pupil function recovery algorithm.
24. The method according to claim 16, wherein The FPM reconstruction performed on one or more low-resolution regions of interest among the plurality of regions of interest generates one or more FPM-reconstructed regions of interest having a greater number of pixels than the one or more low-resolution regions of interest.
25. The method according to claim 1, wherein The performing of the FPM reconstruction using the GPU to generate the reconstructed image includes employing a plurality of GPUs to generate the reconstructed image.
26. A Fourier stack imaging system comprising: a plurality of light sources configured to emit light onto the sample location; an optical system configured to image at least a portion of a sample positioned at the sample location; an image capture device configured to capture images of the sample through the optical system under different light conditions provided by the plurality of light sources; processor; a graphics processing unit (GPU) in communication with the processor; as well as a memory coupled to the processor, the memory including computer-executable instructions stored therein, the computer-executable instructions causing the processor to perform the following operations when executed by the processor: obtaining the image of the sample positioned at the sample location; storing the image in the memory; Uploading the image to the GPU; and initiating Fourier ptychography microscopy (FPM) reconstruction using the GPU such that a reconstructed image is generated, The FPM reconstruction includes executing various parts of the FPM reconstruction in parallel on the GPU, so as to reduce the FPM reconstruction time.
27. The Fourier stack imaging system according to claim 26, wherein: The processor is configured to control operation of: the plurality of light sources, the image capture device, or the plurality of light sources and the image capture device.
28. The Fourier stack imaging system according to claim 26, wherein: The processor is configured to initiate at least a portion of an FPM reconstruction by invoking a kernel of the GPU.
29. The Fourier stack imaging system of claim 26, further comprising: A display is configured to output the reconstructed image.
30. The Fourier stack imaging system according to claim 26, wherein: The memory includes computer-executable instructions that, when executed by the processor, further cause the processor to: Execute a single kernel call; or Executes a series of kernel calls and Fast Fourier Transform (FFT) calls.
31. The Fourier stack imaging system of claim 26, wherein: The GPU is configured to employ at least one pupil function representing the optical system of the Fourier ptychography system during FPM reconstruction of the reconstructed image.
32. The Fourier stack imaging system of claim 26, wherein: The GPU is configured to generate the reconstructed image, and the GPU is configured to generate the reconstructed image including the GPU being configured to perform the following operations: For each of the images: defining a plurality of regions of interest within the respective images; and performing FPM reconstruction on each region of interest in the plurality of regions of interest independently of other regions of interest in the plurality of regions of interest within the corresponding image; as well as The reconstructed image is generated based on the FPM reconstruction of each of the plurality of regions of interest within each of the images.
33. The Fourier stack imaging system according to claim 32, wherein: Each region of interest of the plurality of regions of interest within each of the images comprises a predetermined number of pixels within the respective image.
34. The Fourier stack imaging system of claim 32, wherein: At least some adjacent regions of interest among the plurality of regions of interest share one or more image pixels.
35. The Fourier stack imaging system of claim 32, wherein: The GPU is configured to perform FPM reconstruction on each of the plurality of regions of interest in parallel.
36. The Fourier stack imaging system of claim 35, wherein: The GPU is configured to perform pixel-level parallelism.
37. The Fourier stack imaging system of claim 32, wherein: The GPU is configured to perform FPM reconstruction of one or more regions of interest among the plurality of regions of interest, wherein the GPU is configured to: Select the FPM algorithm; and The selected FPM algorithm is applied to the one or more regions of interest.
38. The Fourier stack imaging system according to claim 37, wherein: The FPM algorithm includes a pupil function recovery algorithm.
39. The Fourier stack imaging system of claim 32, wherein: The GPU is configured to perform the following operations: performing FPM reconstruction on one or more low-resolution regions of interest among the plurality of regions of interest; as well as One or more FPM-reconstructed regions of interest are generated, the one or more FPM-reconstructed regions of interest having a greater number of pixels than the one or more low-resolution regions of interest.
40. The Fourier stack imaging system of claim 26, further comprising: A plurality of GPUs including the GPU.