System and method for autofocus and automated cell counting using artificial intelligence

An AI-driven imaging system uses machine learning and neural networks to rapidly and accurately count cells, addressing the limitations of conventional systems by improving autofocus and cell viability counting precision.

JP2026050361APending Publication Date: 2026-03-19LIFE TECHNOLOGIES CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Conventional cell viability counting systems are computationally intensive, require specialized equipment, and struggle to accurately distinguish between cellular and non-cellular particles, leading to inconsistent and inaccurate results.

Method used

An AI-powered imaging system that uses machine learning and convolutional neural networks to rapidly autofocus and count cells, employing methods like thresholding, morphological operations, and elliptical fitting to identify cellular components, and generate pseudo-probability maps for precise cell counting.

Benefits of technology

Enables rapid and accurate cell viability counting without additional computing resources, distinguishing between live and dead cells while ignoring artifacts, providing consistent and precise measurements within seconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026050361000001_ABST
    Figure 2026050361000001_ABST
Patent Text Reader

Abstract

We provide a system and method for autofocus and automated cell counting. [Solution] A system and method for autofocus using artificial intelligence includes (i) capturing multiple monochrome images over a nominal focal range, (ii) identifying one or more connected components within each monochrome image, (iii) sorting the identified connected components based on the number of pixels associated with each connected component, (iv) generating focus quality estimates for at least some of the sorted connected components using a machine learning module, and (iv) calculating a target focal position based on the focus quality estimates of the evaluated connected components. The target focal position can be used to perform cell counting using artificial intelligence, for example, by (i) generating seed likelihood images and whole cell likelihood images based on the output of a convolutional neural network, and (ii) generating a mask indicating the quantity and / or pixel location of objects based on the seed likelihood image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to microscopic image analysis. More specifically, the present disclosure relates to image analysis solutions for autofocus and cell count driven by artificial intelligence (AI).

Background Art

[0002] Cytometry is the counting of cells and / or the measurement of cell characteristics. Various devices and methods are used in the field of cytometry to measure characteristics such as cell number, cell size, cell morphology, and cell life cycle phase. Cytometry can also involve the measurement of various cell components, such as nucleic acids, the presence of specific proteins, cell typing and / or differentiation (e.g., viability counting), and various medical diagnostic applications.

[0003] There is a continuing need for improvement in the field of cytometry and related fields of image analysis.

Summary of the Invention

[0004] Embodiments of this disclosure provide systems and methods for autofocusing and automated cell counting that offer one or more advantages over conventional methods. For example, one or more embodiments may include a method for autofocusing an imaging system, which includes the act of capturing multiple images (e.g., monochrome images) over a nominal focal range; the act of identifying one or more binding components within each image; the act of sorting the identified binding components based on the number of pixels associated with each binding component; the act of evaluating focal quality estimates for at least some of the sorted binding components using a machine learning module; and the act of calculating a target focal position based on the focal quality estimates of the evaluated binding components. The term “monochrome image” is used in most of the examples presented herein, but it will be understood that the same principles and features may be applicable to applications involving other image types (e.g., color images). Therefore, the examples described herein are not limited to applications involving only monochrome images.

[0005] In one embodiment, sorting the identified bonded components can, additionally or alternatively, be based on the roundness and / or luminance associated with each bonded component.

[0006] In one embodiment, capturing multiple monochrome images involves capturing a monochrome 8-bit image (e.g., a 1280 × 960 monochrome 8-bit image) at each z-axis step over a nominal focal range. In one embodiment, each z-axis step is a predetermined size that is automatically determined based on the corresponding sample holder detected in the imaging system.

[0007] In one embodiment, the act of identifying one or more combined components within each monochrome image includes thresholding each monochrome image to obtain a binary image resulting from each monochrome image, applying one or more morphological operators to the resulting binary image, and defining one or more combined components within the resulting binary image. In one embodiment, thresholding is based on the difference between the minimum and maximum images of the nominal focal range. In an additional or alternative embodiment, the act of applying one or more morphological operators includes applying one or more of morphological closing, morphological opening, and / or foreground hole fill. In one embodiment, morphological closing includes 2x2 morphological closing, and / or morphological opening includes 2x2 morphological opening. As an additional or alternative, foreground hole fill includes 8-combined foreground hole fill.

[0008] In one embodiment, the act of identifying one or more combined components in each monochrome image includes measuring the first and second binary moments of each combined component, fitting corresponding ellipses having equivalent moments to each combined component, and creating a second binary image containing the corresponding ellipses. In one embodiment, the method further includes measuring each combined component in the second binary image and removing any combined components having an ellipse minor axis of less than 15 μm, or less than 10 μm, preferably less than 7 μm, more preferably less than 5 μm.

[0009] In one embodiment, the act of sorting identified combined components includes counting the number of pixels in each combined component, calculating the median pixel count across one or more combined components, and sorting the combined components based on the corresponding absolute difference of the number of pixels from the median pixel count. In one embodiment, the combined components are sorted in ascending (or descending) order by the corresponding absolute difference of the number of pixels from the median pixel count.

[0010] In one embodiment, the act of sorting one or more identified binding components includes ordering the calculated joint probabilities such that the binding component (i) represents a cell and optionally (ii) the cell is alive. The binding components with the highest joint probabilities can be sorted for further evaluation of the focus quality estimate using a machine learning module.

[0011] In one embodiment, the act of sorting one or more identified combined components includes determining one or more of the roundness or luminance of each combined component, and sorting the combined components based on the comparative roundness and / or luminance determined for each of the combined components.

[0012] In one embodiment, a method for autofocusing an imaging system may include forming a z-stack of a specific pixel size (e.g., 32×32 or other appropriate size) for each connected component. In one embodiment, a machine learning module includes an artificial neural network that receives a pixel z-stack of each connected component as input. In one embodiment, the artificial neural network includes a plurality of feature recognition layers having a design pattern of convolutional layers, linear layers, and max pooling layers. In one embodiment, the convolutional layers include a 3×3 convolutional layer, the linear layers include a ReLU nonlinear function, and / or the max pooling layer includes a 2×2 max pooling layer or an average pooling layer.

[0013] In one embodiment, the artificial neural network includes a long-short-term memory (LSTM) processing layer following multiple feature discrimination layers, the LSTM layer processing the z-stack of each connected component in a bidirectional manner. In another embodiment, the artificial neural network includes a final linear layer which is combined with the output of the LSTM layer to define a focus quality estimate.

[0014] Embodiments of the present disclosure further include a computer system configured to autofocus an imaging system. In one embodiment, the computer system configured to autofocus an imaging system includes one or more processors and one or more hardware storage devices storing computer executable instructions that, when executed by the one or more processors, constitute the computer system to perform any one or more of the methods disclosed herein.

[0015] Embodiments of the present disclosure may further include a method for performing automated cell counting, comprising the act of acquiring an image, the act of defining one or more sets of tiles based on the image, the act of processing one or more tiles using a convolutional neural network, and the act of constructing a plurality of pseudoprobability maps based on the output of the convolutional neural network. The plurality of pseudoprobability maps may include at least one or more seed likelihood images. The act may further include generating one or more masks based on one or more seed likelihood images. One or more masks may define the pixel locations of one or more objects represented in one or more seed likelihood images. One or more masks may indicate / provide cell counts.

[0016] In one embodiment, a method for performing automated cell counting includes performing joined component analysis on one or more seed likelihood images to determine the number of cells.

[0017] In one embodiment, the multiple pseudo-probability maps further include one or more whole-cell likelihood images, and the method for performing automated cell counting includes generating one or more segmented images based on at least one or more whole-cell likelihood images and one or more masks, the one or more segmented images indicating / providing the number of cells.

[0018] In one embodiment, the image is captured at the target focal position.

[0019] In one embodiment, the convolutional neural network is trained with a set of training data that includes images from multiple focal planes for focal positions, preferably in the range of -2 to +2z, in order to enhance the robustness of the method and reduce the method's sensitivity to autofocus output accuracy. The set of training data may include a ground truth output that includes one or more whole-cell binary masks and / or one or more seed masks. The images in the set of training data may include pre-processed images.

[0020] In one embodiment, the image includes a 2592 × 1944 monochrome 8-bit image with a pixel size of 0.871 μm. Other types of images (e.g., those with different dimensions, bit values, and / or pixel sizes) may also be used.

[0021] In one embodiment, a method for performing automated cell counting further includes performing one or more preprocessing operations. One or more preprocessing operations may be performed on an image to generate a preprocessed image. One or more sets of tiles may be decomposed from the preprocessed image, or one or more preprocessing operations may be performed on one or more tiles.

[0022] One or more preprocessing operations may include downsampling. Downsampling operations may utilize averaging filters. Downsampling operations may reduce the image size by at least half.

[0023] One or more preprocessing operations may include background removal operations. In some cases, background removal operations include estimating the background by calculating a local mean for each pixel within the radius of each pixel and subtracting the background from the image. In some cases, background removal operations include calculating the respective minimum values ​​in each image region of the image or downsampled image and fitting the surface through the respective minimum values.

[0024] One or more preprocessing operations may include pixel or voxel intensity normalization operations. Pixel or voxel intensity normalization operations may include global normalization operations. Pixel or voxel intensity normalization operations may include kernel-based normalization operations.

[0025] In one aspect, at least some of the one or more preprocessing operations are performed as a batch process on one or more tiles provided as input to the convolutional neural network.

[0026] In one aspect, defining a set of one or more tiles includes forming a reflected image, decomposing the reflected image into a set of one or more tiles having pixel overlap, and storing the set of one or more tiles in a tensor array. Reflecting the downsampled image can expand its edges by some pixels (e.g., 4 to 12 pixels, or 6 to 10 pixels, or 8 pixels), and the reflected image can be decomposed into a set of tiles having pixel overlap. In a non-limiting example, the reflected image is decomposed into 130 tiles of 128×128 pixels with 8-pixel overlap. The tensor array may include a 130×1×128×128 tensor array or other suitable tensor array based on the number of tiles, image size, and other image characteristics.

[0027] In one aspect, the convolutional neural network comprises a U-net convolutional neural network. The output of the convolutional neural network can be stored in a suitable tensor array (e.g., a 130×4×128×128 tensor array).

[0028] In one aspect, constructing a plurality of pseudo-probability maps includes converting the output of a convolutional neural network into a tensor (e.g., 8-bit format) and stitching the tensors into an image. Converting the output into a tensor may include multiplying a tensor array (e.g., a 130×4×128×128 tensor array) by a multiplier (e.g., 255). Stitching the tensors into an image may include constructing several (e.g., 4) full-size pseudo-probability maps from the tiles. The plurality of pseudo-probability maps may each include (i) one or more live cell seed likelihood images preferably indicating one or more central positions of live cells, (ii) one or more live cell overall likelihood images, (iii) one or more dead cell seed likelihood images preferably indicating one or more central positions of dead cells, and (iv) one or more dead cell overall likelihood images.

[0029] In one aspect, generating one or more masks includes applying a threshold greater than a 75% likelihood. The threshold may correspond to a threshold of 192 (0.75×255) in pixel intensity in the plurality of pseudo-probability maps. Other thresholds may be used, such as, by way of non-limiting example, a threshold within the range of about 50% to about 90% likelihood. Generating one or more masks may further include applying connected component labeling to detect connected regions, and the one or more masks may be generated using a deep learning algorithm.

[0030] In some cases, generating one or more segmented images includes applying a watershed transform to a distance map calculated from the pixel positions of one or more objects represented in one or more seed likelihood images and delimited by one or more whole cell likelihood images.

[0031] In one aspect, a method of performing automatic cell counting further includes determining the cell count based on the output of a convolutional neural network. A method of performing automatic cell counting may further include displaying the cell count based on the output of a convolutional neural network.

[0032] In one embodiment, a method for performing automated cell counting further includes using one or more segmented images to perform one or more feature calculation operations. One or more feature calculation operations may include elliptic fitting cells in one or more segmented images. Elliptic fitting may include measuring object size (e.g., in micrometers), constructing a histogram of object size and pixel intensity, and / or calculating object roundness.

[0033] In one embodiment, the image includes a monochrome image.

[0034] Embodiments of the present disclosure further include a computer system configured to perform automated cell counting. For example, a computer system for performing automated cell counting may include one or more processors and one or more hardware storage devices storing computer executable instructions that, when executed by the one or more processors, constitute the computer system to perform any one or more methods of performing automated cell counting as disclosed herein.

[0035] Accordingly, systems and methods for autofocusing an imaging system and / or performing automated cell counting are disclosed. In some embodiments, the systems and methods disclosed herein enable automated cell viability counting.

[0036] This summary is provided to introduce the selection of concepts in a simplified form, which will be further explained in the detailed description below. This summary is not intended to identify the main or essential features of the claimed subject matter, nor is it intended to be used to indicate the scope of the claimed subject matter.

[0037] Additional purposes and benefits of this disclosure are partially described below, some of which will become apparent from that description or can be acquired through the practice of this disclosure. The features and benefits of this disclosure can be realized and acquired through the equipment and combinations disclosed herein. These and other features of this disclosure will become more fully apparent from the following description and the appended claims or can be acquired through the practice of the disclosure as described below. [Brief explanation of the drawing]

[0038] To illustrate how the above and other advantages and features of this disclosure can be obtained, a more specific description of the disclosure, as briefly described above, will be made by reference to the specific embodiments shown in the accompanying drawings. It will be understood that these drawings only show typical embodiments of the disclosure and should therefore not be considered limiting its scope. The disclosure will be described and explained with additional specificities and details using the following accompanying drawings. [Figure 1] The image shows a perspective view of an imaging system configured to perform one or more of the methods disclosed herein, including an artificial intelligence (AI)-assisted autofocus and an automated cell viability counting method, according to one or more embodiments of the present disclosure. [Figure 2] This is an illustrative flowchart illustrating various actions that can be performed by the imaging system of Figure 1 to facilitate AI-assisted autofocus and automated cell viability counting, according to one or more embodiments of the present disclosure. [Figure 3] Figure 1 shows schematic diagrams of various exemplary components within the imaging system according to one or more embodiments of the present disclosure. [Figure 4] This is an exemplary flowchart illustrating various actions for determining a target focal position according to one or more embodiments of the present disclosure. [Figure 5A] A simplified diagram of a canonical neural network, as known in this field, is shown. [Figure 5B] Figure 5A shows the separated parts of the neural network. [Figure 6]This is a block diagram representing an exemplary design of an artificial neural network for facilitating AI-assisted autofocus, according to one or more embodiments of the present disclosure. [Figure 7] This is an exemplary flowchart illustrating various actions for performing automated AI-assisted cell viability counting using target focus locations, according to one or more embodiments of the present disclosure. [Figure 8] This provides a conceptual diagram for background removal. [Figure 9] The following is a block diagram illustrating an exemplary design of a U-net convolutional neural network for facilitating AI-assisted determination of cell viability, according to one or more embodiments of the present disclosure. [Figure 10] This provides additional conceptual diagrams of inputs and outputs associated with a U-net convolutional neural network. [Figure 11] This example demonstrates how to generate segmented images based on seed likelihood images and whole-cell likelihood images using watershed transformation. [Figure 12] This example demonstrates the use of one-object operations to generate segmented images based on seed likelihood images and whole-cell likelihood images. [Figure 13] This example demonstrates the use of a deep learning module to generate segmented images based on seed likelihood images and whole-cell likelihood images. [Figure 14] The following are illustrative images that appear after implementing a method for AI-assisted autofocus and automated cell viability counting according to one or more embodiments of the present disclosure. [Figure 15A] The same base image is shown, evaluated and annotated with the number and location of living and dead cells in each image, as determined by a biologist (Figure 15A). [Figure 15B] The same base image is shown, evaluated and annotated with the number and location of live and dead cells in each image, as determined by a conventional automated cell identification and viability method (Figure 15B). [Figure 15C]The same base image is shown, evaluated and annotated with the number and location of living and dead cells in each image, as determined by the AI-assisted autofocus and automated cell counting method disclosed herein (Figure 15C). [Modes for carrying out the invention]

[0039] Before describing embodiments of this disclosure in detail, it should be understood that this disclosure is not limited to the parameters of the systems, apparatus, methods, and / or processes, which are naturally subject to change and are particularly exemplified. Therefore, while specific embodiments of this disclosure are described in detail with reference to specific configurations, parameters, components, elements, etc., the descriptions are illustrative and should not be construed as limiting the scope of this disclosure. Furthermore, the terms used herein are intended to describe embodiments and are not necessarily intended to limit the scope of this disclosure.

[0040] It will be understood that a system, apparatus, method, and / or process according to a particular embodiment of this disclosure may include, incorporate, or otherwise include characteristics or features (e.g., components, members, elements, parts, and / or parts) described in other embodiments disclosed and / or described herein. Accordingly, various features of a particular embodiment may be compatible with, combined with, included in, and / or incorporated in other embodiments of this disclosure. Therefore, the disclosure of a particular feature relating to a particular embodiment of this disclosure should not be construed as limiting the application or inclusion of such feature to a particular embodiment. Rather, it will be understood that other embodiments may include such features, members, elements, parts, and / or parts without necessarily departing from the scope of this disclosure.

[0041] Furthermore, unless otherwise implicitly or explicitly understood or stated, it is understood that, with respect to any given component or embodiment described herein, any of the possible candidates or substitutes listed for that component may generally be used individually or in combination with each other. In addition, unless otherwise implicitly or explicitly understood or stated, it will be understood that any such list of candidates or substitutes is merely illustrative and not limiting.

[0042] In addition, unless otherwise indicated, any quantities, components, distances, or other measured numbers used in the specification and claims should be understood to be modified by the term “about.” As used herein, “about,” “approximately,” “substantially,” or their equivalents, represent quantities or conditions that are close to a specific described quantity or condition that still perform the desired function or achieve the desired result. For example, the terms “approximately,” “about,” and “substantially” may refer to quantities or conditions that deviate by less than 10%, or less than 5%, or less than 1%, or less than 0.1%, or less than 0.01%, from the specifically described quantity or condition.

[0043] Accordingly, unless otherwise indicated, the numerical parameters described in the specification and the appended claims are approximations that may vary depending on the desired characteristics to be obtained by the subject matter presented herein. At a minimum, and without intending to limit the application of the principle of equivalents to the claims, each numerical parameter should be interpreted at least in light of the number of significant figures reported and by applying ordinary rounding techniques. Although the numerical ranges and parameters describing the broad range of subject matter presented herein are approximations, the figures described in specific examples are reported as accurately as possible. However, any figure inherently contains a certain degree of error that inevitably arises from the standard deviation observed in each experimental measurement.

[0044] The headings and subheadings used herein are for structural purposes only and are not intended to be used to limit the scope of the description or claims. Overview of a System and Method for AI-Assisted Autofocus and Cell Viability Counting

[0045] Cell viability counting has traditionally been a manual, time-intensive process. Advances in the field of imaging have enabled imaging systems to automate some of these manual tasks, or, alternatively, reduce the amount of time and manual effort associated with determining cell concentration in a sample, particularly in relation to confirming the proportional number of live / dead cells in a given sample (or viability counting). While these innovations are promising, they ultimately fall short of the accuracy and precision of manual cell viability counting.

[0046] Conventional systems and methods for analyzing cell viability in samples suffer from numerous drawbacks. For example, many cell viability systems require the use of dyes, labels, or other compounds to determine the viability of cells in a sample. The use of many of these compounds often necessitates specialized, expensive, and / or bulky equipment to retrieve the results. Consequently, the equipment is unlikely to be readily available and / or to be placed in a convenient location within the laboratory space.

[0047] Some conventional systems rely on computationally intensive algorithms to identify and distinguish between living and dead cells in a sample. Unfortunately, these systems generally require access to dedicated, robust computing resources such as graphics processing units (GPUs), or distributed computing resources capable of more rapidly computing and processing the large amounts of data associated with known image recognition and processing algorithms. Unfortunately, access to distributed computing resources and network or cloud computing environments that could enable faster analysis of image-acquired cell viability data is not always available or feasible. In fact, the confidentiality of some laboratory samples and / or ongoing experiments leads to increased security measures preventing access to networked computing resources such as the cloud, including situations where such resources and analytical capabilities are provided by third parties and privacy concerns exist with respect to laboratory samples.

[0048] There is a lack of compact, relatively inexpensive systems that can be conveniently deployed and accessed from a common workspace without occupying a large footprint on a benchtop or requiring an isolated, dedicated environment. Adding a GPU, or otherwise increasing the computing power of an existing system, is impractical for some devices. Additional computing power requires additional space within the system, generates additional heat, and increases the overall cost of the system, both upfront and operationally (although the principles disclosed herein may be implemented on devices utilizing one or more GPUs). However, without additional computing resources to process image data, conventional systems cannot provide the desired results in a timely manner, if any. Known cell viability counting systems cannot focus on living cells while ignoring dust, debris, manufacturing defects at the bottom of the counting chamber, and other non-cellular particles on the microscopic image. These artifacts are confused in the results, preventing current systems from accurately and precisely distinguishing cellular particles from non-cellular particles. As a result, current cell viability counting systems are not effective in producing consistent and accurate measurements of cell viability in a sample.

[0049] Therefore, there is a need for a cost-effective, compact, and safe benchtop imaging system that can perform accurate and rapid autofocus and cell viability counting of relevant samples.

[0050] As suggested above, conventional systems and methods for autofocusing and analyzing cell viability in samples have many drawbacks. In particular, there are the high computational costs associated with image processing, which have prevented non-specialized laboratory equipment from rapidly and accurately focusing on and identifying live / dead cells in samples. The problem is that conventional known methods rely on evaluating higher-resolution images in attempts to use the additional data they provide to resolve the difference between live and dead cells. This implementation inevitably requires significant investment in additional computing resources, such as dedicated GPU batteries or access to large-scale distributed computing environments.

[0051] There is a lack of compact, relatively inexpensive systems that can be conveniently deployed and accessed from a common workspace without occupying a large footprint on a benchtop or requiring an isolated, dedicated environment. Adding a GPU, or otherwise increasing the computing power of an existing system, is impractical for some devices. Additional computing power requires additional space within the system, generates additional heat, and increases the overall cost of the system, both upfront and operationally. Embodiments described herein may be implemented on devices utilizing one or more GPUs. However, for certain applications, conventional systems, if any, cannot provide the desired results in a timely manner without additional computing resources for processing image data. Known cell viability counting systems cannot focus on living cells while ignoring dust, debris, manufacturing defects at the bottom of the counting chamber, and other non-cellular particles on the microscopic image. These artifacts are confused in the results, preventing current systems from accurately and precisely distinguishing cellular particles from non-cellular particles. As a result, current cell viability counting systems are not effective in producing consistent and accurate measurements of cell viability in a sample.

[0052] The systems and methods disclosed herein solve one or more of the problems pointed out in the art and advantageously enable rapid autofocus (for example, on mixtures of living and dead cells) using minimal processing cycles and without requiring additional processing hardware. The improved autofocus methods disclosed herein are powered by artificial intelligence and enable rapid and reliable identification of target focal locations in any given sample. This rapid and low-cost method for identifying target focal locations in a sample is incorporated into many of the disclosed cell viability counting methods as a first step enabling rapid (e.g., less than 20 seconds, preferably less than 10 seconds) identification and / or display of a cell viability count representation.

[0053] Figure 1 shows a perspective view of an imaging system 100 configured to perform one or more of the methods disclosed herein. For example, the imaging system 100 in Figure 1 is operable to facilitate a method related to the AI-assisted autofocus and / or automated cell viability counting flowchart 200 disclosed by the exemplary flowchart in Figure 2. As shown, the imaging system 100 includes a housing 102 that encloses and protects a microscope and computing system used to perform autofocus and cell viability counting. The housing 102 includes a slide port / stage assembly 106 operable to receive a cell counting slide into the imaging system (action 202). Once received therein, the imaging system 100 then determines a target focal position for imaging cells on the cell counting slide (action 204) and uses the target focal position to perform automated cell viability counting (action 206). The cell viability count is displayed in the imaging system 100, for example, using a display 104 (action 208). This representation and / or other data associated with the automated cell viability count may, in some embodiments, be removed from the imaging system 100 through user interaction with a communication module 108 which may include a USB port or other data exchange port as known in the art, and / or stored on a separate device.

[0054] In view of this disclosure, it will be understood that the principles described herein can be implemented using any suitable imaging system and / or any suitable imaging modality. Specific examples of imaging systems and imaging modalities discussed herein are provided as examples and as means of illustrating the features of the disclosed embodiments. Thus, the embodiments disclosed herein are not limited to any particular microscope system or microscopy application and may be implemented in a variety of contexts, such as bright-field imaging, fluorescence microscopy, flow cytometry, and confocal imaging (e.g., 3D confocal imaging, or any type of 3D imaging). For example, the principles discussed herein can be implemented using a flow cytometry system to provide or improve cell counting capabilities. As another example, cell count and / or viability data obtained according to the techniques of this disclosure may be used to complement fluorescence data to improve accuracy in distinguishing between different cells.

[0055] Furthermore, in consideration of this disclosure, it will be understood that any number of principles described herein can be implemented in various fields. For example, a system may implement the cell counting techniques discussed herein without necessarily implementing the autofocus, cell viability, and / or feature detection techniques described herein.

[0056] Figure 3 shows schematic diagrams of various exemplary components within the imaging system 100 of Figure 1 according to one or more embodiments of the present disclosure. For example, Figure 3 shows that the imaging system 100 may include a computer system 110 and a microscope system 120 contained therein. Figure 3 conceptually represents the computer system 110 and the microscope system 120 arranged within the housing 102 of the imaging system 100. However, in consideration of the present disclosure, it should be understood that any part of the computer system 110 or the microscope system 120 may be located at least partially outside the housing 102 within the scope of the disclosed embodiments.

[0057] Figure 3 shows that the computer system 110 of the imaging system 100 may include various components such as a processor 112, a hardware storage device 114, a controller 116, a communication module 108, and / or a machine learning module 118.

[0058] The processor 112 may comprise one or more sets of electronic circuits, including any number of logic units, registers, and / or control units, to facilitate the execution of computer-readable instructions (e.g., instructions that form a computer program). Such computer-readable instructions may be stored in a hardware storage device 114, which may include physical system memory and may be volatile, non-volatile, or some combination thereof. Further details relating to the processor (e.g., processor 112) and the computer storage medium (e.g., hardware storage device 114) are provided below.

[0059] The controller 116 may include any suitable software components (e.g., a set of computer executable instructions) and / or hardware components (e.g., application-specific integrated circuits or other dedicated hardware components) that are capable of operating to control one or more physical devices of the imaging system 100, such as a part of the microscope system 120 (e.g., a positioning mechanism 128).

[0060] The communication module 108 may include any combination of software or hardware components that can operate to facilitate communication between on-system components / devices and / or with off-system components / devices. For example, the communication module 108 may include ports, buses, or other physical connection devices for communicating with other devices (e.g., USB ports, SD card readers, and / or other devices). In addition, or alternatively, the communication module 108 may include, as a non-limiting example, a system that can operate to wirelessly communicate with external systems and / or devices through any suitable communication channel such as Bluetooth, ultra-wideband, WLAN, or infrared communication.

[0061] The machine learning module 118 may also comprise any combination of software or hardware components capable of operating to facilitate processing using machine learning models or other artificial intelligence-based structures / architectures. For example, the machine learning module 118 may include, as a non-limiting example, hardware components or computer executable instructions capable of operating to execute functional blocks and / or processing layers, such as single-layer neural networks, feedforward neural networks, radial basis function networks, deep feedforward networks, recurrent neural networks, long-term short-term memory (LSTM) networks, gated recurrent units, autoencoder neural networks, variational autoencoders, denoising autoencoders, sparse autoencoders, Markov chains, Hopfield neural networks, Boltzmann machine networks, restricted Boltzmann machine networks, deep belief networks, deep convolutional networks (or convolutional neural networks), deconvolutional neural networks, deep convolutional inverse graphics networks, generative adversarial networks, liquid state machines, extreme learning machines, echo state networks, deep residual networks, Kohonen networks, support vector machines, and neural Turing machines.

[0062] As shown in Figure 3, the imaging system 100 includes a microscope system 120 having an image sensor 122, an illumination source 124, an optical train 126, a slide port / stage assembly 106 for receiving a sample slide, and a positioning mechanism 128.

[0063] The image sensor 122 is positioned within the optical path of the microscope system and configured to capture an image of a sample, which will be used in the disclosed method to identify a target focal position and subsequently perform automated cell viability counting. As used herein, the terms “image sensor” or “camera” refer to any applicable image sensor that fits the apparatus, systems, and methods described herein, including but not limited to the aforementioned combinations such as charge-coupled devices, complementary metal-oxide-semiconductor devices, N-type metal-oxide-semiconductor devices, Quanta image sensors, and scientific complementary metal-oxide-semiconductor devices.

[0064] The optical train 126 may include one or more optical elements configured to facilitate the visibility of the cell counting slide by directing light from the illumination source 124 to the received cell counting slide. The optical train 126 may also be configured to direct light scattered, reflected, and / or emitted by the specimen in the cell counting slide towards the image sensor 122. The illumination source 124 may be configured to emit various types of light, such as white light or light in one or more specific wavelength bands. For example, the illumination source 124 may include an optical cube (e.g., Thermo Fisher EVOS® optical cube), which can be installed and / or swapped within a housing for any desired set of illumination wavelengths.

[0065] The positioning mechanism 128 may include an x-axis motor, a y-axis motor, and a z-axis motor, which are operable to adjust the components of the optical train 126 and / or image sensor 122 as appropriate.

[0066] Figure 3 further shows that in some cases the imaging system 100 includes a display 104. Figure 3 shows that the display 104 can communicate directly or indirectly with various other components of the imaging system 100, such as the computer system 110 or its microscope system 120 (for example, as indicated by the triple-headed arrow in Figure 3). For example, the imaging system 100 may use components of the microscope system 120 to capture an image, the captured image may be processed and / or stored using components of the computer system 110 (e.g., a processor 112, a hardware storage device 114, a machine learning module 118, etc.), and the processed and / or stored image may be displayed on the display 104 for observation by one or more users.

[0067] As described herein, the components of the imaging system 100 can facilitate AI-assisted autofocus on samples contained within cell counting slides imaged by the imaging system 100, as well as AI-assisted cell viability counting on samples contained within cell counting slides. In some cases, a representation of the results of AI-assisted cell viability counting (which may be performed according to the target focal position determined via AI-assisted autofocus) may be displayed on the display 104 of the imaging system 100 within a short period (e.g., within a period of about 20 seconds or less, or within a period of about 10 seconds or less) after the autofocus and cell viability counting process for the cell counting slide inserted into the imaging system 100 has been initiated.

[0068] In light of this disclosure, it should be understood that the imaging system may include additional or alternative components beyond those illustrated and described with reference to Figure 3, and such components may be organized and / or distributed in various ways. System and Method for Facilitating AI-Assisted Autofocus

[0069] As described above, facilitating AI-powered autofocus and cell viability counting includes determining a target focal position for imaging cells on a cell counting slide (act 204, as described above with reference to Figure 2). Determining a target focal position for imaging cells on a cell counting slide can be associated with various acts. Figure 4 shows an exemplary flowchart illustrating various acts related to act 204 from flowchart 200 for determining a target focal position for imaging cells on a cell counting slide. The acts shown in flowchart 200 may be illustrated and / or described in a particular order, but no particular ordering is required unless specifically stated or required, as some acts depend on other acts being completed before they are performed. Furthermore, it should be noted that not all acts represented in the flowchart are essential to facilitating the disclosed methods, including the AI-assisted autofocus and automated cell viability counting methods disclosed herein.

[0070] Act 204a related to act 204 includes capturing multiple monochrome images across a nominal focal range. In some implementations, act 204 is performed by the imaging system 100 using the processor 112, hardware storage device 114, and / or controller 116 of the computer system 110, as well as the image sensor 122, illumination source 124, optical train 126, slide port / stage assembly 106, and / or positioning mechanism of the microscope system 120.

[0071] For example, the imaging system 100 may employ a processor 112 in conjunction with one or more sensors to identify a sample holder (e.g., a cell counting slide) located within the slide port / stage assembly 106. In some cases, the imaging system 100 automatically identifies the type of sample holder (e.g., a disposable cell counting slide or a reusable cell counting slide) and automatically determines the image acquisition settings based on the detected sample holder type. For example, the imaging system 100 may identify the image z-axis step height / size and / or initial nominal focal range based on whether the cell counting slide is disposable or reusable, and / or other attributes of the cell slide (e.g., determined sample holder, coverslip, and / or other substrate thickness).

[0072] Furthermore, the imaging system 100 can use the processor 112 and / or controller 116 to cause the positioning mechanism 128 to position the optical train 126 relative to the slide port / stage assembly 106, thereby facilitating the acquisition of images of the cell counting slide (e.g., a sample within the cell counting slide). The imaging system 100 can use the image sensor 122 and illumination source 124 in combination with either of the above to acquire images of the cell counting slide. In addition, the imaging system 100 may acquire additional images of the cell counting slide under different relative positionings of the slide port / stage assembly 106 and the optical train 126 to acquire multiple monochrome images over the nominal focal range. In some implementations, the monochrome images are 8-bit images with a resolution of 1280 × 960, and each monochrome image is acquired at each z-axis step over the nominal focal range.

[0073] Any instructions for performing the actions described herein, and / or data used or generated / stored in connection with the performance of the actions described herein (e.g., cell counting slide type, nominal focal range, z-axis step, monochrome image, etc.) may be stored in the hardware storage device 114 in a volatile or non-volatile manner.

[0074] Act 204b, related to act 204, includes identifying one or more combined components within each monochrome image. In some implementations, the imaging system 100 uses one or more of the processor 112, hardware storage device 114, and / or machine learning module 118 to identify the combined components within each monochrome image.

[0075] Where used herein, “connectivity” in “connecting component” refers to which pixels are considered neighbors of the target pixel. After a suitable digitized image becomes available (for example, from multiple monochrome images acquired according to act 204a), all connecting components in the image are first identified. A connecting component is a set of pixels with a single value, e.g., black, and a path can be formed from any pixel in the set to any other pixel in the set without leaving the set, e.g., by traversing only black pixels. Generally speaking, a connecting component can be either “4-connected” or “8-connected.” In the case of 4-connected, there are four possible directions, as paths can only move horizontally or vertically. Thus, two diagonally adjacent black pixels are not 4-connected unless another horizontally or vertically adjacent black pixel acts as a bridge between the two. In the case of 8-connected, paths between pixels can also proceed diagonally. In one embodiment, an 8-connected component is used, but a 4-connected component can also be identified and used.

[0076] For example, in some implementations, identifying the combined components of each monochrome image is performed by thresholding the monochrome images to obtain the resulting binary image for each monochrome image. Thresholding may be performed based on the difference between the minimum and maximum images of the nominal focal range described above with reference to act 204a. Various methods for thresholding, such as the “upper triangle” thresholding method, are within the scope of this disclosure.

[0077] In addition to thresholding, identifying the combined components of each monochrome image may include applying one or more morphological operators to each of the resulting binary images. Morphological operations / operators include, in non-limiting examples, morphological closing and morphological opening. This may include opening, and / or foreground hole filling. For example, cells may be represented in a binary image as containing fjords (C-shaped artifacts or irregularities) due to inadequate or suboptimal illumination, focusing, image sensing, and / or post-processing (e.g., thresholding). In some cases, based on the assumption that cells should have a round shape, the morphological closing operation may fill any fiddles present in the monochrome image to approximate the cell shape in the binary image. In one embodiment, the morphological closing operation utilizes a 2x2 mask size.

[0078] Furthermore, in some cases, cells may be represented in binary images as including tendrils (e.g., side branches) extending beyond the cell wall, due to inadequate or suboptimal illumination, focusing, image sensing, and / or post-processing (e.g., thresholding). Therefore, in some implementations, the system may perform a morphological opening operation to trim or remove tendrils from the binary image in order to improve the approximation of cell shape in the binary image. In one embodiment, the morphological opening operation is performed using a 2x2 mask size.

[0079] In addition, foreground holes may appear within a set of binary images based on the appearance of cells from different focal positions (e.g., different z-height focal positions from which multiple monochrome images were captured according to act 204a). Thus, in some cases, a foreground hole fill operation may be performed to expand connected pixels in a manner that fills such foreground holes. In one embodiment, the foreground hole fill operation is an 8-coupled foreground hole fill operation, while in some implementations, the foreground hole fill operation is a 4-coupled foreground hole fill operation.

[0080] In some cases, after performing a desired morphological operation (e.g., morphological opening, morphological closing, foreground hole fill) on each binary image, the imaging system 100 may define the combined components within each binary image in preparation for further processing. However, in some cases, additional operations may be performed to identify the combined components in preparation for further processing.

[0081] For example, in some implementations, the imaging system 100 measures first and second binary moments for each binding component defined within each of the binary images described above. In some cases, the first and second binary moments of a particular binding component may correlate with the axes of an ellipse, which can serve as an approximation of the cell shape. Thus, for a particular binding component, the imaging system 100 can fit an ellipse to that specific binding component, and the ellipse has moments based on the first and second binary moments measured for that particular binding component.

[0082] In this way, the imaging system 100 can fit ellipses to each joined component defined in each of the aforementioned binary images. In some implementations, the imaging system 100 generates a second binary image from each of the aforementioned binary images (for example, each binary image generated by thresholding each monochrome image). The second binary image may include joined components generated / defined based on elliptic fitting from the first and second binary moments measured for each joined component from the aforementioned binary images. In this regard, in some cases, the ellipse-based joined components of the second binary image can help the imaging system 100 approximate the cell shape.

[0083] In some implementations, according to this disclosure, the second binary image is used for further processing to determine the target focal position. However, in some implementations, the imaging system 100 generates a third set of binary images based on the first and second binary images, and uses the third set of binary images for further processing to determine the target focal position. The third set of binary images may be generated by merging each binary image (or an initial binary image generated by thresholding each corresponding monochrome image) with its corresponding second binary image, or by taking their union.

[0084] Regardless of whether the imaging system 100 utilizes a set of first, second, or third binary images for further processing to determine the target focal position, in some implementations, the imaging system 100 may measure each binding component and remove any binding components that do not meet predetermined size criteria. For example, the imaging system 100 may remove binding components from any of the first, second, or third binary images that have a minor axis length or diameter (e.g., elliptical minor axis length) of less than 40 μm. In some implementations, the predetermined size criteria are selected based on the type of cell being counted. For example, for T cells, B cells, NK cells, and / or monocytes, the imaging system 100 may remove binding components with a minor axis diameter outside the range of approximately 2 to 40 μm, while for other cells, the imaging system 100 may remove binding components with a minor axis diameter outside the range of approximately 2 to 8 μm. Such short-axis diameters may be set, for example, to a range having endpoints defined by 2 μm, 4 μm, 6 μm, 8 μm, 10 μm, 15 μm, 20 μm, 25 μm, 30 μm, 35 μm, 40 μm, 45 μm, or 50 μm, or any two of the aforementioned values, depending on the requirements of a particular application.

[0085] Action 204c, associated with action 204, includes sorting the identified combined components based on the number of pixels associated with each combined component. As described above, the identified combined components sorted according to action 204c may be from the third set of binary images, the second binary image, or the first binary image, as described above with reference to action 204b. In some implementations, the imaging system 100 utilizes the processor 112 to sort the identified combined components based on the number of pixels associated with each combined component.

[0086] Sorting the identified combined components according to action 204c may involve various actions / steps. For example, in some cases sorting the identified combined components may involve counting the number of pixels in each combined component, calculating the median pixel count of the combined components, and sorting the combined components based on the median pixel count (for example, based on the absolute difference between the number of pixels in each combined component and the median pixel count).

[0087] In some cases, the imaging system 100 utilizes one or more other inputs in addition to, or alternatively to, the number of pixels to sort the identified combined components. For example, the imaging system 100 may utilize a measure of the pixel brightness and / or the relative roundness of the captured component. Based on the brightness and / or roundness of the captured component, the imaging system 100 can determine the probability that the captured component is a cell, and another probability that the cell is alive (e.g., a living or dead cell). These two probabilities may be combined (e.g., by multiplying the two probabilities together) to form a joint probability that the captured component is a living cell (e.g., the probability that the component based on brightness alone is a cell, and the probability that the component based on roundness alone is a cell).

[0088] The binding components can be sorted in ascending or descending order. In some cases, sorting the binding components in ascending order has the effect of moving binding components that approximate the normal size of a cell to the top of the list and binding components that are more likely to approximate other objects (e.g., aggregates and debris particles) to the bottom of the list.

[0089] Furthermore, in some cases, the imaging system 100 modifies the list by removing list elements that are unlikely to approximate the normal or expected cell size. For example, the imaging system 100 may use a Gaussian distribution or other distribution with respect to the list median to remove specific list elements, such as those outside a predetermined distance or difference from the list median. In other cases, the imaging system 100 uses the maximum list value as a starting point for determining which list elements should be removed.

[0090] Action 204d, related to action 204, includes using a machine learning module to evaluate the focal quality estimates of at least some of the sorted combined components. In some implementations, the imaging system 100 utilizes the processor 112 and / or the machine learning module 118 to evaluate the focal quality estimates of at least some of the sorted combined components.

[0091] As background, artificial intelligence, at its core, attempts to model human thought or intelligence in order to solve complex or difficult problems. Machine learning is one form of artificial intelligence that "learns" from data without a complex set of defined rules, using computer models and some form of feedback. Most machine learning algorithms can be classified based on the type of feedback used in the learning process. For example, in unsupervised learning models, unlabeled data is input into the machine learning algorithm, from which a general structure is extracted. Unsupervised learning algorithms can be powerful tools for clustering datasets. Supervised learning models, on the other hand, use labeled input data to train the machine learning algorithm, "learning" a model to reliably predict desired outcomes. Therefore, supervised machine learning models can be powerful tools for classifying datasets or performing regression analysis.

[0092] Neural networks encompass a broad category of machine learning algorithms that attempt to simulate complex thinking and decision-making processes with the help of computers by selecting suitable network topologies and processing capabilities that mimic the biological functions of the brain. Neural networks are dynamic enough to be used in a wide range of supervised and unsupervised learning models. Today, hundreds of different neural network models, as well as numerous connectivity characteristics and capabilities, exist. Nevertheless, the fundamental way in which neural networks work remains the same.

[0093] For example, Figure 5A shows a simplified schematic diagram of a canonical neural network. Figure 5B further provides a separate portion of the neural network shown in Figure 5A. As illustrated, so-called input neurons are located on the input side of the neural network and are connected to hidden neurons. Each neuron has one or more weighted inputs, either in the form of external signals or as outputs from other neurons. Positive and negative weighting is possible. The sum of the weighted inputs is transmitted via a transfer function to one or more output values, which either control other neurons or act as output values ​​themselves. Hidden neurons, shown in Figure 5A (and as single nodes in Figure 5B), are connected to output neurons (or output nodes in the case of Figure 5B). Of course, the region of hidden neurons can also have a fairly complex structure and can consist of several interconnected levels.

[0094] The entirety of all neurons performing a particular function is called a layer (e.g., an input layer). Some of the more influential parameters of a neural network, apart from its topology, include neuronal base potentials and the strength of connections between neurons. To set the parameters, a representative training set is iteratively evaluated by the network. After each evaluation cycle, the weights and base potentials are changed and set anew. This iteration continues until the mean failure rate falls below a predetermined minimum or until a predefined problem-related termination criterion is reached.

[0095] As described above, the imaging system 100 can use machine learning to evaluate focal quality estimates for at least some of the sorted (e.g., sorted according to action 204D in Figure 4) connected components. The imaging system 100 can use the identified connected components described above to generate input for a machine learning model. For example, in some implementations, the system localizes a pixel window over at least some of the connected components of a binary image (e.g., a third set of binary images). The pixel window can take various sizes, such as 32 × 32 pixels or another size. Those skilled in the art will understand, in consideration of this disclosure, that the size of the pixel window may depend on the imaging application in which the imaging system 100 is employed (e.g., the type of cells being imaged).

[0096] The imaging system 100 may also determine the z-stack for each combined component based on the pixel window of each respective combined component. For example, for a pixel window for a particular combined component identified in a particular binary image associated with a particular z-height (or focal position), the imaging system 100 can identify the corresponding pixel windows (e.g., pixel windows with the same pixel coordinates) in other binary images associated with z-heights offset from a particular z-height in the particular binary image for the particular combined component (e.g., offset by +1 z-step, -1 z-step, +2 z-step, -2 z-step, etc.), and this set of pixel windows can form the z-stack for the particular combined component.

[0097] In light of this disclosure, it will be understood that the size of a particular z-stack may vary in different implementations. For example, a z-stack may include several pixel windows (from binary images associated with adjacent z-heights) in the range of approximately 3 to 11 or more, and the number of pixel windows (similar to the size of the pixel windows) for the z-stack may also depend on the imaging application and / or the type of cells being imaged.

[0098] The imaging system 100 can form a z-stack for any number of connected components. For example, the imaging system 100 may form a z-stack for a predetermined number (e.g., 32 or another number) of connected components included in a sorted list of connected components (sorted as described above, for example, with reference to action 204d in Figure 4). The imaging system 100 can provide a z-stack for various identified connected components as input to a machine learning module, which can evaluate the quality of focus estimate based on the z-stack.

[0099] Those skilled in the art will understand, in consideration of this disclosure, that various machine learning models may be employed to evaluate focus quality estimates based on one or more sorted connected components (e.g., based on a z-stack as described above). Figure 6 shows one exemplary neural network that the imaging system 100 may use to facilitate focus quality evaluation according to this disclosure. In particular, Figure 6 shows an exemplary block diagram of an artificial neural network 600 for facilitating AI-assisted autofocus. The artificial neural network 600 may be trained in a supervised or partially supervised manner using training data that includes a z-stack for connected components as input training data and clearly identified target focal positions as ground truth outputs. In some cases, the training data may include a z-stack with smaller z-steps between pixels or images in the z-stack to improve the robustness of the artificial neural network 600 for evaluating focus quality.

[0100] Figure 6 shows input data 602 that can contain any number of z stacks (e.g., 1, 2, ... 32, or more), as described above. The artificial neural network 600 can receive the input data 602 and process it using one or more feature recognition layers 604. As shown in Figure 6, the feature recognition layer 604 may comprise various components such as a convolutional layer 606, a linear layer 608, and a max pooling layer 610 (or, in some cases, an average pooling layer). In some non-restrictive implementations, the convolutional layer 606 is a 3x3 convolutional layer, the linear layer 608 is a ReLU nonlinear function, and the max pooling layer is a 2x2 max pooling layer.

[0101] Figure 6 shows an implementation in which the artificial neural network 600 includes three substantially identical feature recognition layers, in particular feature recognition layer 604, feature recognition layer 612, and feature recognition layer 614. However, in other implementations, the artificial neural network 600 may have any number of feature recognition layers having the same or at least partially different components.

[0102] Figure 6 also shows an implementation of the artificial neural network 600 that includes a long short-term memory (LSTM) processing layer 616 following a feature recognition layer. The LSTM can enable processing of the z-stack provided as input in a bidirectional manner. Implementing the LSTM processing layer 616 in the artificial neural network 600 can improve the accuracy of the artificial neural network 600 for evaluating focus quality in some cases, but implementing the LSTM processing layer 616 can be process-intensive and / or time-consuming. Therefore, in some implementations, the artificial neural network 600 omits the LSTM processing layer 616 to save computation time and / or resources (e.g., to enable a total computation time of less than 20 seconds or about 10 seconds or less).

[0103] Figure 6 shows that the artificial neural network 600 includes a final linear layer 618 following the LSTM process layer 616 (or, in implementations where the LSTM process layer 616 is omitted, following the feature recognition layer 614). The final linear layer 618 may be configured to provide an output 620, which may comprise a focus quality estimate for each z height represented in a particular z stack provided as input to the artificial neural network 600. For each particular z stack provided as input to the artificial neural network 600, the imaging system 100 can identify a particular z height (or focus position) from the output 620, which represents the target focus position for that particular z stack.

[0104] By providing multiple z-stacks to the artificial neural network 600, the imaging system 100 can obtain the corresponding quality of focus estimates as output 620 for each z-stack, as well as the respective target focal positions for each z-stack. Action 204e, related to action 204, includes calculating the target focal positions based on the evaluated quality of focus estimates of the connected components. As described above, the quality of focus estimates (from output 620 from the artificial neural network 600) for multiple z-stacks can provide the respective target focal positions for each z-stack. The target focal positions can be generated or defined in various ways based on the quality of focus estimates (or the respective target focal positions for each z-stack). For example, in some cases, the imaging system 100 defines the median, mode, or mean of various respective target focal positions for each z-stack to select the overall target focal position for the cell counting slides imaged by the imaging system 100 according to action 204a described above. In this way, the imaging system 100 can utilize artificial intelligence to facilitate autofocus in an improved, computationally inexpensive, and / or rapid manner. System and Method for Facilitating AI-Assisted Cell Survival Counting

[0105] As described above, facilitating artificial intelligence-driven cell viability counting may include performing automated cell viability counting using target focus locations (action 206, as described above with reference to Figure 2). Performing automated cell viability counting using target focus locations can be associated with various actions. Figure 7 shows an illustrative flowchart depicting various actions related to action 206 from flowchart 200 for performing automated cell viability counting using target focus locations.

[0106] While this disclosure focuses in at least some respects on performing cell viability counting using a target focal position, it should be understood that, in consideration of this disclosure, the principles discussed herein can be implemented independently of each other. For example, the principles discussed herein in relation to cell counting may be implemented without necessarily implementing the autofocus techniques and / or cell viability analysis techniques discussed herein. Similarly, the autofocus techniques discussed herein may be implemented without necessarily implementing the cell counting and / or viability analysis techniques discussed herein. For example, a flow cytometer may, in some cases, implement at least some of the cell counting techniques discussed herein without first performing the autofocus techniques discussed herein. Embodiments of this disclosure may also utilize the autofocus techniques discussed herein periodically (e.g., once a day or once a week in a laboratory) while utilizing the cell counting method for each sample being processed, and such embodiments have the advantage of saving time and / or are applicable to applications where it may not be possible to autofocus each sample.

[0107] Act 206a related to act 206 includes acquiring an image. In some implementations, act 206 is performed by the imaging system 100 using the processor 112, hardware storage device 114, machine learning module 118, and / or controller 116 of the computer system 110, as well as the image sensor 122, illumination source 124, optical train 126, slide port / stage assembly 106, and / or positioning mechanism of the microscope system 120.

[0108] In some implementations, the image associated with action 206a includes a monochrome image. For example, the imaging system 100 can obtain the target focal position according to action 204 (and related actions 204a-204e) as described above with reference to Figure 2 (and Figure 4), using various components of the computer system 110 and the microscope system 120. The imaging system 100 can use the processor 112 and / or controller 116 to cause the positioning mechanism 128 to position the optical train 126 relative to the slide port / stage assembly 106 according to the target focal position. The imaging system 100 can use the image sensor 122 and illumination source 124 in combination with either of the above to capture an image of the cell counting slide using the identified target focal position. In some implementations, the captured image of the cell counting slide at the target focal position is a 2592 × 1944 monochrome 8-bit image with a pixel size of 0.871 μm. Apart from the examples described herein, images used for cell counting as described herein may have any preferred image size or aspect ratio and may be acquired using any preferred imaging device having any preferred pixel size and / or other image sensor characteristics. In some implementations, images have an image size of 96 × 96, 250 × 250, or another image size.

[0109] Any instructions for performing the actions described herein, and / or data used or generated / stored in connection with the performance of the actions described herein (e.g., monochrome images, cell viability count data), may be stored in the hardware storage device 114 in a volatile or non-volatile manner.

[0110] As described above, this example focuses on a monochrome image associated with a target focal position, but in other examples, the monochrome image may be associated with any focal position according to act 206a.

[0111] Act 206b, related to Act 206, includes generating a preprocessed image by performing one or more preprocessing operations. In some cases, the preprocessed image may include a preprocessed monochrome image. Various preprocessing operations are within the scope of this disclosure. For example, one or more preprocessing operations in Act 206b may include one or more of the following operations: downsampling (or down-averaging), background removal, and / or intensity normalization. In some cases, the user may be able to choose which preprocessing functions, if any, are performed in accordance with Act 206b.

[0112] In some implementations, the imaging system 100 utilizes the processor 112 to perform one or more downsampling operations in accordance with act 206b. For example, in some cases, the imaging system 100 performs downsampling via pixel decimation or by applying an averaging filter (e.g., generating each output pixel based on the average intensity of each adjacent pixel). The imaging system 100 may downsample a captured monochrome image by a predetermined downsampling coefficient in each image dimension, such as a coefficient of about 1.5 to about 4, or about 2 (e.g., reducing a 2592 × 1944 captured image to a downsampled image of 1296 × 972). While any downsampling coefficient is within the scope of this disclosure, in some cases, a high downsampling coefficient (e.g., a coefficient of 4 or higher) may result in the removal and / or distortion of certain features from the downsampled preprocessed image.

[0113] As described above, one or more preprocessing operations in action 206b may include background removal. In some cases, the acquired image may have a high or non-uniform background. Removing the background may facilitate easier identification of objects of interest in the acquired image. In some cases, background removal is suitable for low-fluorescence or non-fluorescence samples. Therefore, as described above, the user can choose whether or not to perform background removal as part of action 206b (for example, by checking or unchecking the box associated with performing background removal).

[0114] Figure 8 is a conceptual diagram of background removal, where the raw image 802 includes both the background 804 and the object 806. Background removal can be conceptualized as generating an estimated background 808 and subtracting the estimated background 808 from the raw image 802 to generate a background-removed image 810.

[0115] Background removal can be easily performed in various ways. For example, background removal can be performed according to a low-pass filtering method, such as by estimating the background by calculating a local mean within a radius r at each pixel and subtracting the estimated background from the image. The radius r can be defined as the input maximum object size (e.g., constrained between 1 and 255). The subtraction can result in a fixed mean value. Low-pass filtering background removal can be performed according to a selected object detection mode, such as light / dark, dark / light, and / or others. The fixed mean value may be 0 for the light / dark mode, the maximum value for the dark / light mode, and / or an intermediate value for other modes.

[0116] In some cases, low-pass filtering methods are adapted for use in images that have channels that play a role in identifying objects. Low-pass filtering can be performed efficiently in images with high-contrast edges (e.g., sharply stained nuclei). Low-pass filtering can also be useful for removing background caused by "dirty" samples (e.g., including out-of-plane fluorescence such as "floating cells"). In some cases, the parameter value for low-pass filtering background removal corresponds to the area sampled to determine the amount of background to remove. Lower values ​​may be considered more aggressive and may be constrained to a range of the maximum object diameter. Higher values ​​may be considered more conservative and may be within a range of multiples of the maximum object diameter.

[0117] In some cases, background removal may be performed using a surface fitting method, which may include dividing the input image into a grid, calculating the minimum value in each image region, and fitting the surface through these minimum values. In some cases, using the minimum value (e.g., rather than the mean) preserves the true intensity value (but may take longer to calculate).

[0118] In some cases, surface fitting methods are adapted to images requiring intensity-dependent measurements, such as fluorophoretic antibody labeling. Surface fitting methods may also be adapted, either additionally or alternatively, to preserve the faint edges of whole-cell staining. Similar to low-pass filtering methods, the parameter values ​​for surface fitting background removal may correspond to the sampled area to determine the amount of background to remove, with lower values ​​considered aggressive and higher values ​​considered conservative.

[0119] In some cases, a user interface is provided that allows the user to choose whether to perform background removal in accordance with act 206b, and / or what type of background removal to perform.

[0120] As described above, one or more preprocessing operations in act 206b may include a normalization operation. A normalization operation may be performed to normalize the intensity present within an image pixel (or voxel in the case of a 3D image) to reduce the impact of per-object intensity variations on subsequent processes. The normalization operation may include global normalization (e.g., implementing the same image statistics across the entire image) or kernel-based normalization.

[0121] Any combination of the aforementioned preprocessing operations (e.g., downsampling / down-averaging, background removal, normalization) may be performed according to act 206b. If multiple preprocessing operations are performed, they may be performed in any preferred order. As described above, the output of the preprocessing operations is one or more preprocessed (monochrome) images (e.g., downsampling / down-averaging image, background removal image, normalized image, downsampling / down-averaging and background removal image, downsampling / down-averaging and normalized image, background removal and normalized image, downsampling background removal and normalized image, etc.). One or more preprocessed images may be used for subsequent operations (e.g., acts 206c, act 206d, etc.).

[0122] The preprocessing operations may be performed on the monochrome images described above with reference to action 206a. In some cases, the preprocessing operations are performed on one or more tiles (e.g., via batch processing) as described below with reference to action 206c. In light of this disclosure, it will be understood that at least some of the preprocessing steps may be implemented in the functionality of a convolutional neural network (CNN) as described below with reference to action 206d, or may utilize a functional architecture that is at least partially independent of a CNN.

[0123] In some cases, no preprocessing is performed, and the raw data (e.g., the monochrome image from action 206a) is used directly in subsequent actions (e.g., actions 206c, 206d, etc.).

[0124] Action 206c, associated with action 206, includes defining one or more sets of tiles based on a preprocessed image (as mentioned above, the preprocessed image may include a preprocessed monochrome image). In some cases, action 206c is performed by the imaging system 100 utilizing the processor 112. Defining one or more sets of tiles can include various actions, such as extending the edges of a preprocessed monochrome image by several pixels (or voxels) to reflect it. Performing a reflect operation on a downsampled image may produce a reflected image with edges extended by 6, 8, 16, or another pixel value. The number of pixels / voxels to extend the downsampled image to obtain a reflected image may be selected based on various factors such as desired processing time and image size.

[0125] The imaging system 100 may then decompose the reflected image into one or more sets of tiles with predetermined pixel overlaps (e.g., 4, 6, 8, 10, 12, 14, or 16 pixel overlaps, or other pixel value overlaps as needed). In one example, the imaging system 100 decomposes the reflected image into 130 tiles (or another number of tiles of the same or different pixel sizes) having a pixel size of 128 × 128.

[0126] In some cases, the imaging system 100 may be configured to store one or more sets of tiles in a tensor array in order to generate input for a machine learning module (e.g., a CNN in act 206d) to facilitate automated cell counting. For example, the one or more sets of tiles described above with reference to act 206c may be stored in a 130 × 1 × 128 × 128 tensor array (or a tensor array with other dimensions depending on the image tile size and number). As described below, the tensor array may act as input for a machine learning module to perform automated cell counting and / or viability determination.

[0127] Action 206d, related to action 206, involves processing one or more tiles using a convolutional neural network (CNN). If one or more tiles are stored in a tensor array, action 206d may involve processing the tensor array using a CNN. In some implementations, the imaging system 100 utilizes the processor 112 and / or the machine learning module 118 to process the tensor array as described above with reference to action 206d. A brief context for artificial intelligence and machine learning is provided above with reference to action 204d in Figure 4 and Figures 5A and 5B.

[0128] Those skilled in the art will understand, in consideration of this disclosure, that various machine learning models may be employed to process the tensor array in accordance with act 206d. Figure 9 shows one exemplary neural network that can be used to facilitate the processing of one or more tiles as part of the performance of automated cell counting and / or viability determination by the imaging system 100. In particular, Figure 9 shows a U-net convolutional neural network 900. Figure 9 shows the U-net convolutional neural network 900 receiving one or more tiles or tiles 906 of a tensor array as input (e.g., from act 206d), and shows that the U-net convolutional neural network 900 in Figure 9 shows neural network processing configured to be performed for each tile of one or more tiles or a tensor array. The U-net convolutional neural network 900 may be configured to receive image inputs such as rectangular images of various sizes (e.g., 16-bit images). For example, the image size may be in the range of about 150 × 150 to about 250 × 250 in one or both dimensions. It should be understood that the range may vary for different imaging modalities. For example, in flow cytometry applications, the input image may be in the range of approximately 96 x 96 to approximately 248 x 248 pixels.

[0129] Furthermore, the U-NET convolutional neural network 900 may be configured to receive various numbers and / or types of image inputs. For example, the U-NET convolutional neural network 900 may be configured to receive batches of images (e.g., a z-stack or subset of a z-stack of images) and / or batches of multiple images from different imaging modalities (e.g., one or more conventional images / volumes and one or more phase-contrast images / volumes associated with the same z position) as simultaneous inputs.

[0130] The U-net convolutional neural network 900 in Figure 9 includes a down layer 902 and an up layer 904. Figure 9 also shows that the down layer 902 and the up layer 904 can include various components. As a non-limiting example, Figure 9 shows an implementation in which the down layer 902 includes a feature recognition layer (e.g., two feature recognition layers) that includes a convolutional layer (e.g., a 2D convolutional layer), a batch normalization layer (e.g., a 2D batch normalization layer), and a ReLU nonlinear function layer. Figure 9 also shows the down layer 902 including a max pooling layer (e.g., a 2D max pooling layer). The down layer 902 can facilitate stepwise downsampling of the input image (e.g., tile 906) (e.g., reducing the image size to one-quarter at each step). In some cases, as shown in Figure 9, the imaging system 100 can apply 2D batch normalization before applying the components of the down layer 902.

[0131] The implementation in Figure 9 also shows an up-layer 904, which includes two sets of upsampling layers: an upsampling layer (e.g., a 2D upsampling layer), a convolutional layer (e.g., a 2D convolutional layer), a batch normalization layer (e.g., a 2D batch normalization layer), and a ReLU nonlinear function layer. The up-layer 904 can facilitate stepwise upsampling of downsampled tiles (e.g., tiles downsampled according to the down-layer 902) (e.g., increasing the image size fourfold at each step), and the imaging system 100 can perform upsampling using the combined layer 908 (e.g., based on the corresponding image used or generated during the down-layer 902 processing). In some cases, as shown in Figure 9, the imaging system 100 may apply a sigmoid function after applying the components of the up-layer 904.

[0132] In some cases, the U-net convolutional neural network 900 is trained with a set of training data containing images from multiple focal planes for an identified target focal position (or other focal position). For example, in some implementations, the convolutional neural network 900 is trained for z-heights in the range of -2 to +2 z-steps for an identified target focal position to increase the robustness of the convolutional neural network 900 and reduce its sensitivity to autofocus output accuracy (for example, according to action 204 described above). The U-net convolutional neural network 900 is trained with training data containing monochrome images (and / or image tiles) as training input, and / or ground truth output containing manually localized and / or identified cell count, viability determination, cell segmentation, cell area (e.g., for individual cells and / or cell combinations, in μm). 2The model can be trained supervised or partially supervised using the following information: (units), cell pseudo-diameter (e.g., cell diameter in μm assuming the singlet was perfectly circular), whole-cell binary mask, seed mask, and / or other information (e.g., tags indicating whether the cell image was processable and / or whether the cell touched the region of interest). Ground truth may be obtained via human annotation / labeling / tagging / segmentation and / or at least partially generated using images captured using an appropriate imaging modality (e.g., bright-field, fluorescence). Training data may include images capturing various types of cells such as IMMUNO-TROL, macrophages, Jurkat, CAR-T, PBMC, and / or others. Images may be associated with one or more different Z-positions. In some cases, images in the training data set are preprocessed according to preprocessing operations that may be performed during final use (e.g., downsampling / down-averaging, background removal, normalization, etc.). The training input may, in addition or alternatively, include images with debris, enabling the model to distinguish between living / dead cells (including clusters of cells) and debris. In some cases, the training data may include control objects such as glass beads (e.g., 1 μm, 2.5 μm, 3.6 μm, 5.5 μm, 9.9 μm, 14.6 μm, 30.03 μm) that are similar in size to living or dead cells, enabling the U-NET Convolutional Neural Network 900 to robustly distinguish between cells and objects that are not cells but share similar physical properties (e.g., size and shape).

[0133] As mentioned above, the U-NET convolutional neural network 900 can be configured to receive various numbers and / or types of image inputs. Accordingly, the U-NET convolutional neural network 900 may be trained using batches of images (e.g., a z-stack or subset of a z-stack of images) and / or batches of multiple images from different imaging modalities (e.g., one or more conventional images / volumes and one or more phase-contrast images / volumes associated with the same z position) as concurrent inputs.

[0134] In light of this disclosure, it will be understood that the U-NET convolutional neural network 900 (or any other artificial intelligence module described herein) may be further trained and / or improved after the initial training on the set of training data described above. For example, the U-NET convolutional neural network 900 may be further trained on training data acquired for a specific field of cell analysis.

[0135] As shown in Figure 9, the output of the U-net convolutional neural network 900 may include or be used to generate various pseudo-probability maps 910. For example, Figure 9 shows the output pseudo-probability maps 910 of the U-net convolutional neural network 900 as including (i) a living cell location probability map showing the central location of living cells (e.g., a living cell seed likelihood image), (ii) a living cell mask probability map (a living cell whole likelihood image), (iii) a dead cell location probability map showing the central location of dead cells (a dead cell seed likelihood image), and (iv) a dead cell mask probability map (a dead cell whole likelihood image). For example, the pseudo-probability maps 910 in Figure 9 represent the predicted / possible locations and shapes / sizes of living and dead cells represented on tile 906 of the tensor array acquired according to act 206d.

[0136] As described above, the U-net convolutional neural network 900 can operate on each tile of the tensor array (e.g., tile 906) acquired according to action 206d. Thus, in the example of a 130×1×128×128 tensor array, the U-net convolutional neural network 900 can generate four pseudo-probability maps for each single tile represented in the tensor array, and these pseudo-probability maps (output from the U-net convolutional neural network 900) can be stored in the 130×4×128×128 tensor array.

[0137] The example presented in Figure 9 and described with reference to act 206d involves outputting (or generating based on the output of) a specific set of pseudo-probability maps via a U-NET convolutional neural network 900, but other ranges / types of U-NET CNN outputs are within the scope of this disclosure. For example, in some cases, the output of a U-NET CNN generates a seed likelihood image indicating cell location and / or a whole cell likelihood image indicating cell shape / size, which does not distinguish between living and dead cells and includes information on both living and dead cells (if both exist). For example, in some implementations, the cell counting function is performed independently of cell shape / size / feature detection so that the U-NET CNN outputs at least a seed likelihood image, and based on the seed likelihood image, coupled component analysis may be performed to determine the number of cells (e.g., without determining cell viability and / or cell shape / size / feature). Such functionality may be desirable in flow cytometry systems, as an example of a non-limiting feature.

[0138] Further conceptual representations of the inputs and outputs related to U-net CNNs are provided in Figure 10. Figure 10 shows an exemplary input image 1002, which may include a preprocessed monochrome image or preprocessed monochrome image tiles. The input image 1002 is used as input to the U-net CNN 1004, which in principle corresponds to the U-net convolutional neural network 900 described above. Figure 10 also shows an exemplary output 1006, which includes a seed likelihood image 1008 and a whole-cell likelihood image 1010 superimposed on each other.

[0139] While this example focuses on using a U-net CNN to obtain a pseudo-probability map in at least some respects, other modules such as machine learning-driven texture detection and / or kernel detection methods may be used in some embodiments.

[0140] Act 206e, related to act 206, involves constructing a plurality of pseudo-probability maps based on the output of a convolutional neural network, the plurality of pseudo-probability maps including at least a seed likelihood image and / or a whole-cell likelihood image. As described above, in some cases, the output from the U-net convolutional neural network 900 may consist of a 130 × 4 × 128 × 128 tensor array (or a tensor array of other size appropriate for the particular input used). In some implementations, the imaging system 100 may construct the plurality of pseudo-probability maps of act 206e by converting the output of the convolutional neural network (e.g., a 130 × 4 × 128 × 128 tensor array) into an 8-bit format tensor and image stitching the 8-bit format tensor. For example, converting a 130×4×128×128 tensor array to an 8-bit format may involve multiplying the tensor array by a suitable multiplier such as 255, and image stitching in an 8-bit tensor format may involve constructing four full-size pseudo-probability maps from the tiles in the 8-bit tensor format. In some implementations, the four resulting pseudo-probability maps may have representations similar to the pseudo-probability map 910 in Figure 9 (e.g., live cell locations, live cell mask, dead cell locations, and dead cell mask), or the output in Figure 10 (e.g., seed likelihood image and whole cell likelihood image).

[0141] The action 206f associated with action 206 includes generating one or more masks based on one or more seed likelihood images. For example, in some cases, the imaging system 100 may generate a binary location map by thresholding four full-size pseudo-probability maps (e.g., corresponding to the pseudo-probability map 910 in Figure 9). One or more masks can define or indicate the pixel locations of one or more objects (e.g., cells) represented in one or more seed likelihood images. Thus, in some cases, one or more masks may indicate the number of cells.

[0142] Various techniques for generating one or more masks, such as applying a likelihood threshold of 75% or higher, are within the scope of this disclosure. In some cases, using a threshold of 192 or higher in pixel intensity represents a pseudo-probability of 0.75 × 255 or higher. As a non-limiting example, various thresholds, such as those in the range of approximately 50% to approximately 90%, are within the scope of this disclosure. To detect joined regions in the output image, joined component labeling may be further applied. One or more masks may be generated based on the joined component labels. In addition, or alternatively, one or more masks may be generated using deep learning algorithms.

[0143] In some implementations, the imaging system 100 performs additional image processing operations, such as augmentation processing, on the binary location map in an attempt to capture latent cells that were ignored during other processing steps.

[0144] Act 206g, related to Act 206, includes generating one or more segmented images based on at least one whole-cell likelihood image and one or more masks. The one or more segmented images may indicate / provide cell counts and / or cell viability. In some implementations, the imaging system 100 generates one or more segmented images via a watershed transformation applied to distance maps calculated from the pixel / voxel locations of objects (e.g., cells) represented in one or more seed likelihood images. The one or more segmented images may be demarcated by one or more whole-cell likelihood images (e.g., by the respective masks of the objects) such that each pixel is assigned to an object or background. Binary or intensity-based watershed methods are within the scope of this disclosure. Figure 11 shows an exemplary representation utilizing the watershed transformation 1106 to generate a segmented image 1108 based on a seed likelihood image 1102 and a whole-cell likelihood image 1104. Different objects represented by the seed likelihood image 1102 are numerically labeled (e.g., "1", "2", "3"). The segmented image 1108 shows separate objects (labeled "1", "2", and "3" correspondingly in the segmented image 1108) that are extensions of different objects (objects 1, 2, and 3) from the seed likelihood image 1102.

[0145] In addition to, or as an alternative to, the watershed method, other techniques can be used to facilitate the generation of one or more segmented images. For example, in some embodiments, one or more segmented images may be generated that treat cell groupings as a single object, such as by assigning all pixels in the seed likelihood image with a likelihood above 75% to a mask structure (other thresholds may be used as described above). Such an approach may be considered a "one-object" approach, and downstream computations may be performed on a single object (rather than on individual objects). Such a technique may be beneficial when computational resources are limited. Cell morphological measurements of the unit object may be used to verify the number of seeds identified by the seed likelihood image described above. Figure 12 shows an exemplary representation that utilizes the one-object operation 1206 to generate a segmented image 1208 based on the seed likelihood image 1202 and the whole-cell likelihood image 1204. The different objects represented by the seed likelihood image 1202 are numerically labeled (e.g., "1", "2", "3"). Segmented image 1208 shows a unit object (labeled "1" in segmented image 1108) composed of different objects (objects 1, 2, and 3) from seed likelihood image 1202.

[0146] As another example, one or more segmented images can be generated using a deep learning module. The deep learning module can be trained using training data that includes seed / whole-cell likelihood image inputs and segmented image ground truth outputs. In some cases, using a deep learning module can enable accurate segmentation of overlapping cells. Figure 13 shows an exemplary representation of using a deep learning module 1306 to generate a segmented image 1308 based on seed likelihood image 1302 and whole-cell likelihood image 1304. The different objects represented by seed likelihood image 1302 are numerically labeled (e.g., "1", "2", "3"). Segmented image 1308 shows separate objects (correspondingly labeled "1", "2", and "3" in segmented image 1308) that are extended from the different objects (objects 1, 2, and 3) in seed likelihood image 1302.

[0147] One or more segmented images may, in some cases, indicate or provide a basis for cell counting and / or cell viability, and may be displayed on a user interface to inform one or more users of the cell counting and / or cell viability represented on the cell counting slide. In some cases, in addition to cell counting, one or more feature calculation operations may be performed using one or more segmented images to determine one or more features of the detected cells. Features can be extracted on a cell-by-cell basis. For example, the imaging system 100 may perform elliptic fitting on cells represented in one or more segmented images. Elliptic fitting on cells in one or more segmented images can utilize the techniques described above for fitting ellipses to joined components (see, for example, action 204b in Figure 4). Ellipses fitted to cells in one or more segmented images may facilitate the acquisition of advantageous data related to cells imaged on the cell counting slide, such as providing a measurement of object size (e.g., in μm), providing a histogram of object size, providing pixel intensity, providing a basis for calculating object roundness, and / or other. Further examples of features that can be obtained for detected cells include: number of objects, object centers (e.g., object centers for each dimension such as x, y, and z centers), object width (e.g., width of the minimum bounding box containing the object), object height (e.g., height of the minimum bounding box containing the object), pixel size (e.g., pixel size for each dimension such as x, y, and z pixel sizes), area, perimeter, perimeter-to-area ratio, fiber length (e.g., length of the object measured along its spine), fiber width (e.g., width of the object estimated from area and length), centroid (e.g., x-centroid, y-centroid), orientation (e.g., orientation of the bounding box aligned with the object), coherence (e.g., scale of the arrangement of substructures within the object), large radius, small radius, rotational radius (e.g., along the z axis), box length (e.g., length of the bounding box in which the object is placed), box width (e.g., width of the bounding box in which the object is placed), length-to-width ratio, box occlusion ratio,The number of pixels that make up each object, object intensity (e.g., maximum intensity, minimum intensity, total intensity, average intensity, standard deviation of intensity, intensity skewness, intensity kurtosis, intensity entropy, etc.), radial intensity moment (e.g., average radial intensity, standard deviation of radial intensity, radial intensity skewness, radial intensity kurtosis, radial distance, etc.), object co-occurrence (e.g., maximum probability of the intensity distribution of all pixels in the mask, contrast, entropy, angular second moment of two-dimensional co-occurrence, etc.), object size (e.g., diameter of a circle with an area equal to the area of ​​the object), equivalent spherical metric (e.g., diameter, surface area, or volume of an equivalent circle or sphere), equivalent elliptical moment This includes, for example, the intensity (e.g., the ratio of length to width of an equivalent circle or sphere, the volume of an ellipse produced by rotating an area equivalent ellipse around its major or minor axis), object distance (e.g., the distance from an object to its nearest neighboring object, the average distance from an object to all other objects, the standard deviation of the distances from an object to all other objects), object gradient ratio (e.g., the intensity gradient of an inner or outer region within an object mask, the ratio of the intensity gradients between the inner and outer regions within an object mask), surface area density (e.g., the total difference in intensity of all pixels within an object normalized by its area), and / or others. Any of the aforementioned features may be weighted appropriately (e.g., according to pixel intensity). Such data may also be obtained for living and / or dead cells represented in a cell counting slide, and in some cases, it may be stored in a consumer report and prepared / provided for export (e.g., using communication module 108).

[0148] Furthermore, as described above with reference to Figure 2, ellipses that fit one or more segmented images may allow the system to display a representation of cell viability counts (for example, according to action 208 of flowchart 200 in Figure 2).

[0149] For example, referring here to Figure 14, an image representing an exemplary result after processing by the disclosed AI-assisted autofocus and automated cell viability counting system is shown. The locations and masks of live / dead cells, obtained as output from processing individual tiles through an artificial neural network (e.g., the U-net convolutional neural network shown in Figure 9), are used to generate a pseudo-probability map, and after a series of image processing and elliptic fitting steps discussed above, the resulting elliptic map includes solid elliptic markers for live cells 1402 and dashed elliptic markers for dead cells 1404, and is overlaid on the corresponding bright-field image 1400 to clearly distinguish live and dead cells. In some embodiments, live and dead cells are visually distinguished in one or more images that can be viewed on a display associated with the imaging system (e.g., the display 104 of the imaging system 100 in Figure 1). For example, living cells may be outlined, shaded, and / or otherwise highlighted in a specific color (e.g., green), while dead cells may be outlined, shaded, and / or otherwise highlighted in a different color (e.g., red).

[0150] Since the disclosed system and method can utilize information within a z-stack that is not readily apparent from a single z-stack slice (e.g., polarity reversal phenomena), the disclosed system can advantageously autofocus on a mixture of living and dead cells using a minimal number of processing cycles and without requiring additional processing hardware. This improved method of autofocusing enables rapid and reliable identification of the target focal position in any given sample and is a useful first step to enable the disclosed method of automated cell viability counting. Similar to the autofocusing method described above, the disclosed method of automated cell counting and / or cell viability counting disclosed herein represents a significant improvement over competing prior art systems and methods.

[0151] For example, as shown in Figures 15A to 15C, the same base image was evaluated and annotated with the number and location of living and dead cells in each image. The image shown in Figure 15A was evaluated by a biologist with experience in annotating the number and location of living and dead cells. The image in Figure 15A was determined to be accurately and precisely annotated and served as a positive control for comparing the accuracy and precision between conventional methods of automated cell identification and viability and exemplary AI-assisted autofocus and automated cell viability counting methods such as those disclosed herein.

[0152] As shown in Figure 15B, Figure 15B shows an image annotated with the number and location of live and dead cells, evaluated by a conventional automated cell identification and viability method. The conventional method cannot accurately segment and identify the number of cells within a cell cluster, cannot consistently identify monodisperse cells as single cells, cannot accurately distinguish between live and dead cells, and cannot ignore debris. Instead, it identifies portions of debris within the observation area as clustered and monodisperse collections of live / dead cells.

[0153] In contrast to conventional methods (e.g., as illustrated in Figure 15B), the systems and methods disclosed herein for AI-assisted autofocus and automated cell viability counting appropriately segment and count the number of live / dead cells (compared to biologist-annotated controls) and, in addition, avoid debris as shown in Figure 15C. As illustrated in Figures 15A–15C (and repeatedly shown across 12 different cell types—data not shown), the disclosed solutions have shown significant improvements over conventional image processing methods in the accuracy of cell segmentation and identification of debris, as well as in viability counting. Such improvements have been repeatedly shown to occur across monodisperse and aggregated cell specimens.

[0154] Details of the exemplary data structure The following description provides exemplary data structures that may be associated with the various components / elements disclosed herein. An image may be defined by an image class which may include an array of pixel values ​​and associated metadata. Pixel values ​​can take various forms, such as Byte, SByte, UInt16, Int16, Int32, UInt32, Int64, UInt64, Single, Double, etc. Metadata may include, in non-limiting examples, the pixel data type, bits per pixel, bytes per pixel, pixel offset, stride (e.g., allowing cropping of an image region without copying pixels), XYZ position in the well (e.g., in micrometers), XYZ pixel size, image dimensions (width, height), acquisition time, intensity control settings (e.g., exposure time, gain, binning), and / or others. An image class may support grayscale and / or color (RGB). If color is used, each pixel may have a corresponding color value (e.g., an RGB color value, or a value following another color system). Image classes may be serializable and / or convertible to different types (e.g., between color and RGB). Image readers can be used to read / write image classes from standard file formats via any stream (file, memory, pipe).

[0155] A mask can define a list of pixel indices and / or a bounding box to which pixels belong. A mask class can be used as both input (e.g., for defining truth data) and output (e.g., for defining the position of an object). The bounding box may be the smallest bounding box containing the object, thereby allowing the mask definition to be maintained independently of the field from which it is taken. Therefore, in some implementations, a mask can be extracted from one image set (e.g., one time point, one pass) and applied to a second image set (e.g., another time point, another pass), even if the second image set is not associated with the exact same positions as the first image set (e.g., the second image set may be associated with different magnifications and / or XYZ positions). A mask class can provide methods for converting pixel lists between bounding box definitions. In some cases, images may be adjusted alternately to match the mask (e.g., instead of adjusting the mask to match the image). For example, an image class can provide a way to "crop" the bounding box to match the mask (e.g., cropping can change only the pixel offset and stride without actually copying the image pixel values). Using extension methods, mask classes can be converted to shapes (e.g., ellipses, rectangles, polygons) and / or vice versa. In some cases, converting a mask to a shape may be irreversible (e.g., "best fit"), but it can still provide a convenient way to display the mask.

[0156] The output data can include a variety of data types, such as doubles, integers, booleans, datetimes, double enumerations, integer enumerations, string enumerations, strings, binaries, and any of the above 1D, 2D, or ND arrays.

[0157] Computer System of this Disclosure It will be understood that computer systems take on an increasingly diverse range of forms. In this description and claims, the term “computer system” or “computing system” is broadly defined to include any device or system or combination thereof having at least one physical and tangible processor and physical and tangible memory capable of having computer-executable instructions that can be executed by the processor. By example, rather than by limitation, the term “computer system” or “computing system” as used herein is intended to include a part of the disclosed imaging system that communicates with associated optical and mechanical components and is operable to perform the various autofocus and / or cell viability counting methods disclosed herein. Thus, the term “computer system” or “computing system” as used herein can execute operational commands related to stage / sample movement, in addition to controlling automatic image acquisition and processing (e.g., determining target focal positions and / or performing automatic cell viability counting). Unless otherwise specified, the computing systems disclosed herein should be understood to be components of the disclosed imaging systems.

[0158] The memory components of the disclosed computer system may take any form and may depend on the nature and form of the computing system. Memory may be physical system memory including volatile memory, non-volatile memory, or a combination of the two. The term "memory" may be used herein to refer to non-volatile mass storage devices such as physical storage media.

[0159] The computing systems disclosed herein are understood to store on them a number of structures, often referred to as “executable components.” For example, the memory of a computing system may contain executable components. The term “executable component” is the name of a structure, as is well understood by those skilled in the art of computing, which may be software, hardware, or a combination thereof.

[0160] For example, when implemented in software, a person skilled in the art will understand that the structure of an executable component may include software objects, routines, methods, etc., that can be executed by one or more processors on a computing system, and whether such an executable component resides on the heap of the computing system or on a computer-readable storage medium. The structure of the executable component resides on a computer-readable medium in such a manner that, when executed by one or more processors of the computing system, it can cause the computing system to perform one or more functions, such as the functions and methods described herein. Such a structure may be directly computer-readable by the processor, as in the case of a binary. Alternatively, the structure may be structured to be sequentially interpretable and / or compiled (in a single or multiple stages) so that the processor can directly generate a sequentially interpretable binary.

[0161] The term “executable component” will also be well understood by those skilled in the art to include structures that are implemented exclusively or nearly exclusively within hardware logic components, such as field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), program-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), or other specialized circuits. Therefore, the term “executable component” is a term well understood by those skilled in the art to describe structures, whether implemented in software, hardware, or a combination thereof.

[0162] The terms “component,” “service,” “engine,” “module,” “control,” and “generator” may also be used in this description. As used in this description, these terms are intended to be synonymous with the term “executable component,” with or without modifying clauses, and therefore have a structure that is well understood by those skilled in computing.

[0163] Not all computing systems require a user interface, but in some embodiments, a computing system includes a user interface for use in communicating information with a user. A user interface may include not only input mechanisms but also output mechanisms. The principles described herein are not limited to specific output or input mechanisms and depend on the nature of the device. However, output mechanisms may include, for example, speakers, displays, haptic outputs, etc. Examples of input mechanisms include, for example, microphones, touchscreens, cameras, keyboards, styluses, mice, or other pointer inputs, and any type of sensor.

[0164] Accordingly, embodiments described herein may include or utilize special-purpose or general-purpose computing systems. Embodiments also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media may be any available media accessible by a general-purpose or special-purpose computer system. A computer-readable medium that stores computer-executable instructions is a physical storage medium. A computer-readable medium that carries computer-executable instructions is a transmission medium. Accordingly, embodiments disclosed or assumed herein may include, by example (but not limited to), at least two distinctly different types of computer-readable media, namely storage media and transmission media.

[0165] Computer-readable storage media include RAM, ROM, EEPROM, solid-state drives ("SSD"), flash memory, and phase-change memory. This includes memory ("PCM"), CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or other physical and tangible storage media that can be used to store desired program code in the form of computer-executable instructions or data structures and that can be accessed and executed by a general-purpose or special-purpose computing system to implement the functions disclosed in the present invention. For example, computer-executable instructions may be embodied on one or more computer-readable storage media to form a computer program product.

[0166] The transmission medium may carry the desired program code in the form of computer-executable instructions or data structures and may include networks and / or data links that can be accessed and executed by general-purpose or special-purpose computers. The aforementioned combinations also fall within the scope of computer-readable media.

[0167] Furthermore, upon reaching various computer system components, program code in the form of computer executable instructions or data structures can be automatically transferred from the transmission medium to the computer storage medium (or vice versa). For example, computer executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., a "NIC") and ultimately transferred to the computer system RAM and / or the computer system's less volatile storage medium. Thus, it should be understood that the storage medium can also be included in the computing system components that the transmission medium (or rather, primarily it) utilizes.

[0168] Those skilled in the art will further understand that a computing system may also include communication channels that enable it to communicate with other computing systems, for example, over a network. However, as provided above, the computing systems of this disclosure are preferably components of the imaging system disclosed. Thus, while a computing system may include communication channels that enable network communication (for example, for file and / or data transfer, configuration or updating of firmware and / or software associated with the computing system), it should be understood that the computing systems disclosed herein are intended to implement the methods disclosed locally, rather than in a distributed system environment linked over a network (by hardwired data links, wireless data links, or a combination of hardwired and wireless data links) where processing and / or memory can be distributed among various networked computing systems.

[0169] While the subject matter described herein is provided in a language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. Rather, the described features and actions are disclosed as exemplary forms for implementing the claims. Additional terms and definitions

[0170] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to whom this disclosure relates. The terms and expressions used herein are for illustrative purposes only and are not limiting, and in using such terms and expressions, there is no intention to exclude any features or equivalents thereof shown or described, but it is recognized that various modifications are possible within the scope of the claimed disclosure. Accordingly, while the invention has been partially and concretely disclosed by preferred embodiments, exemplary embodiments, and optional features, it should be understood that modifications and variations of the concepts disclosed herein are possible to those skilled in the art, and such modifications and variations are considered to be within the scope of the disclosure. The specific embodiments provided herein are examples of useful embodiments of the invention, as well as various modifications and / or changes to the features of the invention shown herein, and additional uses of the principles shown herein that may arise for those skilled in the art and who own this disclosure are possible with respect to the illustrated embodiments and should be considered within the scope of the disclosure.

[0171] Furthermore, unless a feature is described as requiring another feature in combination with it, any feature herein can be combined with any other feature of the same or different embodiments disclosed herein. Moreover, various well-known embodiments of exemplary systems, methods, apparatus, etc., are not described in particular detail herein to avoid obscuring the embodiments of the exemplary models. However, such embodiments are contemplated herein as well.

[0172] Where used herein, unless implicitly or explicitly understood or stated otherwise, a singular word encompasses its plural equivalents, and a plural word encompasses its singular equivalents. Therefore, note that the singular forms “a,” “an,” and “the” as used herein and in the appended claims include multiple subjects unless explicitly indicated otherwise. For example, a reference to a single referent (e.g., “widgets”) includes one, two, or more referents unless implicitly or explicitly understood or stated otherwise. Similarly, references to multiple referents should be interpreted as including one and / or multiple referents unless explicitly indicated otherwise by the content and / or context. For example, a reference to a plural referent (e.g., “widgets”) does not necessarily require multiple such referents. Instead, it will be understood that, regardless of the presumed number of referents, one or more referents are contemplated herein unless specifically stated.

[0173] All references cited herein are incorporated herein by reference in their entirety to the extent that they do not conflict with the disclosures herein. It will be apparent to those skilled in the art that methods, devices, device elements, materials, procedures, and techniques other than those specifically described herein can be applied to the implementation of the invention as broadly disclosed herein without relying on excessive experimentation. All known functional equivalents in the art of the methods, devices, device elements, materials, procedures, and techniques specifically described herein are intended to be incorporated within this disclosure.

[0174] Where a group of components or other elements is disclosed herein, it is understood that all individual members of the disclosed group and all of its subgroups are disclosed separately. Where a Markush group or any other group is used herein, all individual members of the group, and all possible combinations and subcombinations of the group, are intended to be included in this disclosure individually. All formulations or combinations of components described or illustrated herein may be used to carry out preferred and / or alternative embodiments of this disclosure unless otherwise specified. Wherever a range is indicated in the specification, all intermediate and subranges, and all individual values ​​within the given range, are intended to be included in this disclosure.

[0175] All modifications that fall within the equivalent meaning and scope of the claims should be included within those scopes.

Claims

[Claim 1] A method for performing automated cell viability counting, The capture involves capturing an image at a target focal position, where the target focal position is obtained based on the output of a machine learning module. The captured image is spatially downsampled to form a downsampled image, Decomposing the downsampled image into multiple tiles, The aforementioned multiple tiles are stored in a tensor array, Processing the tensor array using a convolutional neural network, Based on the output of the aforementioned convolutional neural network, multiple pseudo-probability maps are constructed, The process involves thresholding the aforementioned multiple pseudo-probability maps to generate a binary-formatted location map and a binary-formatted mask map. To generate a segmented downsampled image by segmenting the downsampled image using the binary location map for seeding and the binary mask map for demarcating cell regions, wherein the segmented downsampled image indicates / provides the number of viable cells. A method comprising displaying the aforementioned representation of the number of viable cells.