Method for particle classification using images based on filtering layers, and machine learning model and system therefor
By applying image processing and machine learning models to adjust and classify cytometry image data, the method enhances the accuracy and automation of particle classification, addressing inefficiencies in current imaging flow cytometer techniques.
Patent Information
- Application Number
- JP2025522265
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-17
- Filing Date
- 2023-07-26
- Publication Date
- 2025-11-12
AI Technical Summary
Current imaging flow cytometer techniques do not fully utilize imaging characteristics across multiple channels for particle classification, leading to inefficiencies in identifying and classifying particle populations.
Implementing image processing techniques and machine learning models to iteratively adjust and classify cytometry image data, enhancing the accuracy and automation of particle classification by leveraging morphological and spatial fluorescence intensity information.
Improves the accuracy and efficiency of particle classification by effectively utilizing multiple imaging channels, automating the process, and identifying populations missed by traditional methods.
Smart Images

Figure 2025536933000001_ABST
Abstract
Description
[Background technology]
[0001] Flow-based particle detection, analysis, and imaging systems are useful for detecting, analyzing, and imaging particles in fluid samples. Analysis of cytometry data can play an important role in understanding populations by helping to classify particles into one or more types of particles, e.g., classifying cells into cell types.
[0002] The development of imaging-enabled flow cytometry has added new capabilities for identifying and analyzing single particles, e.g., single cells, as demonstrated by techniques for sorting cells based on spatial image parameters calculated in real time. See, e.g., Schraivogel et al., Science 375, 315-320 (2022), the entire contents of which are incorporated herein. Using existing tools, flow cytometer researchers can utilize scatter, fluorescence, and image-derived parameters to perform gating, clustering, and / or statistical analysis to identify populations of particles, e.g., cells. Summary of the Invention [Problem to be solved by the invention]
[0003] Furthermore, imaging flow cytometer techniques introduce additional data sets for analysis, particularly single-cell images, from which image-derived parameters may be calculated. However, not all aspects of such data, such as imaging characteristics across multiple imaging channels, are fully utilized when analyzing cytometry image data according to current techniques in the art.
[0004] Thus, in connection with particle, e.g., cell, classification, the inventors have recognized that there remains a need for continued improvement in particle, e.g., cell, classification techniques, including analytical techniques that exploit image features (e.g., by utilizing additional information about features detected across multiple imaging channels). Embodiments of the present invention fulfill this need. [Means for solving the problem]
[0005] Embodiments of the present invention improve the utility of flow-based particle detection and analysis systems and the data obtained from such systems, e.g., cytometry image data, by introducing a new technique for more effectively classifying cytometry images in that particles accurately and consistently identify whether they belong to a particle category, type, or classification. Furthermore, embodiments of the present invention advantageously automate the classification of particles into various types, categories, or classifications (i.e., particle subtypes, particle subcategories, or particle subclassifications), thereby identifying populations that may be missed by traditional gating and speeding up the identification of particle, e.g., cell, populations. Existing techniques in the art involve manual gating of subpopulations, e.g., using geometric gates or invoking clustering algorithms based on predetermined parameters. The present invention utilizes image data, i.e., cytometry image data, to flexibly classify based on morphology or spatial fluorescence intensity present in the image data. Image processing techniques, i.e., preprocessing images using filters or other image adjustment techniques, such as image blurring or smoothing, can highlight important information within the image data and between different images. The inventors have discovered that repeated, i.e., iterative, application of image adjustment techniques to cytometry image data in conjunction with the application of a model, such as a machine learning model, e.g., a neural network, improves the accuracy and effectiveness of cytometry image analysis and classification techniques. Similarly, the inventors have discovered that applying image adjustment techniques to cytometry image data in conjunction with the training of a model, such as a machine learning model, e.g., a neural network, improves the accuracy and effectiveness of applying the model to classify cytometry image data.
[0006] Aspects of the present disclosure include methods for classifying cytometry image data using images of single particles, e.g., images of single cells. The method for classifying cytometry image data according to certain embodiments includes receiving unclassified cytometry image data including a plurality of images corresponding to image channels, adjusting aspects of at least one of the plurality of images of the cytometry image data, and applying a model to the adjusted cytometry image data to classify the cytometry image data in an iterative manner, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data. Aspects of the present disclosure also include methods for training a model to classify cytometry image data, the method including receiving flow cytometry data including unclassified cytometry image data including a plurality of images, each instance of the cytometry image data corresponding to an image channel, classifying each instance of the cytometry image data of the flow cytometry data to establish ground truth data, and adjusting aspects of at least one of the plurality of images of each instance of the cytometry image data of the ground truth data to train the model to classify the cytometry image data as data including particles belonging to the first category of particles.
[0007] Aspects of the present disclosure further include a system for classifying cytometry image data. According to an embodiment, the system for classifying cytometry image data comprises a processor operatively coupled to a memory, the memory storing instructions that, when executed by the processor, cause the processor to iteratively receive unclassified cytometry image data including a plurality of images corresponding to an image channel, adjust aspects of at least one of the plurality of images of the cytometry image data, and apply a model to the adjusted cytometry image data to classify the cytometry image data, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data.
[0008] Aspects of the present disclosure further include a non-transitory computer-readable storage medium having stored thereon instructions for classifying cytometry image data. According to an embodiment, the non-transitory computer-readable storage medium includes instructions including an algorithm for receiving unclassified cytometry image data, the unclassified cytometry image data including a plurality of images corresponding to an image channel, and an algorithm for iteratively adjusting aspects of at least one image of the plurality of images of the cytometry image data and applying a model to the adjusted cytometry image data to classify the cytometry image data, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data. [Brief explanation of the drawings]
[0009] The invention can be best understood from the following detailed description when read in conjunction with the accompanying drawings, in which:
[0010] [Figure 1A] FIG. 1 illustrates exemplary cytometry image data according to an embodiment of the present invention. [Figure 1B] FIG. 1 illustrates exemplary cytometry image data according to an embodiment of the present invention. [Figure 2A] 1 is a flowchart illustrating a method for classifying cytometry image data using images of single particles according to an embodiment of the present invention. [Figure 2B] 1 is a flowchart illustrating a method for classifying cytometry image data using images of single particles according to an embodiment of the present invention. [Figure 3A] FIG. 10 illustrates exemplary image-adjusted cytometry image data according to an embodiment of the present invention. [Figure 3B] FIG. 10 illustrates exemplary image-adjusted cytometry image data according to an embodiment of the present invention. [Figure 4] FIG. 1 is a block diagram illustrating a computing system according to an embodiment. [Figure 5A] 3 shows an excerpt of a user interface of a computer-implemented embodiment of the method according to the invention; FIG. [Figure 5B] 3 shows an excerpt of a user interface of a computer-implemented embodiment of the method according to the invention; FIG. [Figure 5C] 3 shows an excerpt of a user interface of a computer-implemented embodiment of the method according to the invention; FIG. [Figure 5D] 3 shows an excerpt of a user interface of a computer-implemented embodiment of the method according to the invention; FIG. [Figure 5E] 3 shows an excerpt of a user interface of a computer-implemented embodiment of the method according to the invention; FIG. [Figure 5F] 3 shows an excerpt of a user interface of a computer-implemented embodiment of the method according to the invention; FIG. [Figure 5G] 3 shows an excerpt of a user interface of a computer-implemented embodiment of the method according to the invention; FIG. [Figure 5H] 3 shows an excerpt of a user interface of a computer-implemented embodiment of the method according to the invention; FIG. DETAILED DESCRIPTION OF THE INVENTION
[0011] Aspects of the present disclosure include methods, systems, and non-transitory computer-readable storage media for classifying cytometry image data using images of single particles, e.g., single cells. A method for classifying cytometry image data according to certain embodiments includes iteratively receiving unclassified cytometry image data including a plurality of images corresponding to an image channel, adjusting aspects of at least one of the plurality of images of the cytometry image data, and applying a model to the adjusted cytometry image data to classify the cytometry image data. The model is trained to estimate the presence of particles belonging to a first category of particles in the cytometry image data. Aspects of the present disclosure further include a method for training a model to classify cytometry image data, the method including receiving flow cytometry data including unclassified cytometry image data including a plurality of images, each instance of the cytometry image data corresponding to an image channel, classifying each instance of the cytometry image data of the flow cytometry data to establish ground truth data, and adjusting aspects of at least one of the plurality of images of each instance of the cytometry image data of the ground truth data to train the model to classify the cytometry image data as data including particles belonging to the first category of particles. Systems for implementing the subject methods are further provided.A non-transitory computer-readable storage medium is further described.
[0012] Before the present invention is described in more detail, it is to be understood that the invention is not limited to the particular embodiments described, as such may, of course, vary. The scope of the present invention will be limited only by the appended claims, and it is to be further understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.
[0013] When a range of values is given, it is understood that each intervening value between the upper and lower limits of that range, to one-tenth of the unit of the lower limit unless the context clearly indicates otherwise, and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also encompassed within the invention.
[0014] In this specification, a range is presented with the term "about" before the numerical values. The term "about" is used herein to literally support the exact number that it precedes, as well as a number that is close to or approximately the number that it precedes. When determining whether a number is close to or approximately a specifically stated number, the unstated number that is close or approximately the number may be a number that, in the context in which the specifically stated number is presented, provides a substantial equivalent to the specifically stated number.
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are described.
[0016] All publications and patents cited herein are incorporated by reference to the same extent as if each individual publication or patent was specifically and individually indicated to be incorporated by reference, and are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.
[0017] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a predicate for use of exclusive terminology such as "solely," "only," and the like, or for use of a "negative" limitation in connection with the recitation of claim elements.
[0018] As will be apparent to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein comprises separate components and features which may be readily separated from or combined with any of the features of the other multiple embodiments without departing from the scope or spirit of the invention. Any recited method may be carried out in the order of events recited or in any other order which is logically possible.
[0019] Although the apparatus and methods have been or will be described for grammatical fluidity with functional descriptions, it should be clearly understood that the claims, unless expressly recited under 35 U.S.C. 112, should not be construed as necessarily limited in any way by limitations of "means" or "step" construction, but should be accorded the full scope of the meaning and equivalents of the definition given by the claims under the judicial theory of equivalents, and that if a claim is expressly recited under 35 U.S.C. 112, it should be accorded the full statutory equivalents under 35 U.S.C. 112.
[0020] The present disclosure provides a method for classifying cytometry image data using images of a single particle, e.g., a single cell. In further describing embodiments of the present disclosure, we first describe in more detail a method that receives unclassified cytometry image data including a plurality of images corresponding to image channels, adjusts aspects of at least one of the plurality of images in the cytometry image data, and iteratively classifies the cytometry image data by applying a model to the adjusted cytometry image data, where the model is trained to estimate the presence of particles belonging to a first category of particles in the cytometry image data. We then describe a system for implementing the subject method. We also describe a non-transitory computer-readable storage medium.
[0021] method As summarized above, a method for classifying cytometry image data using images of a single particle, e.g., a single cell, is provided. According to one embodiment, the method for classifying cytometry image data includes receiving unclassified cytometry image data including a plurality of images corresponding to an image channel, adjusting an aspect of at least one of the plurality of images of the cytometry image data, and applying a model to the adjusted cytometry image data to classify the cytometry image data in an iterative manner, where the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data. A method for training a model for classifying cytometry image data is further provided, where the method includes receiving flow cytometry data including unclassified cytometry image data including a plurality of images, each instance of the cytometry image data corresponding to an image channel, classifying each instance of the cytometry image data of the flow cytometry data to establish ground truth data, and adjusting an aspect of at least one of the plurality of images of each instance of the cytometry image data of the ground truth data, thereby training the model to classify the cytometry image data as data including particles belonging to a first category of particles.
[0022] Image data Image data is data obtained from particles, e.g., cells, using any convenient imaging technique. The term "image" is used in its conventional sense to refer to a representation of an object, e.g., a particle, such as a cell, produced by, for example, irradiation with light or irradiation with electromagnetic radiation or electromagnetic stimulation of molecules associated with the particle to emit light. In embodiments, image data refers to one or more images corresponding to the same field of view, and thus the same particle, e.g., cell, to the extent present in the field of view of the imaging technique, e.g., the field of view of an imaging flow cytometer.
[0023] Each of the one or more images of the image data may correspond to a different image channel, which refers to a range of frequencies detectable by an imaging technique, such as imaging flow cytometry, as described in more detail below with respect to the photodetectors of exemplary flow cytometry techniques.
[0024] Image data is data that collectively constitutes a representation of a subject and may be data obtained using any convenient protocol. In some embodiments, the image data obtained in the methods of the present invention is cytometry image data obtained by an imaging flow cytometry technique. An imaging flow cytometry technique may capture one or more images of particles, e.g., cells, of a sample by flowing the sample through a flow channel and imaging the particles as they pass through the channel. Further details regarding imaging techniques, including imaging flow cytometers, are provided below.
[0025] Exemplary cytometry image data according to embodiments of the present invention are shown schematically in FIG. 1A. Cytometry image data 100 is a collection of individual images 110a, 110b, 110c, and 110d. Although four individual images 110a, 110b, 110c, and 110d are shown in association with cytometry image data 100, in embodiments, the cytometry image data may include any convenient number of images. In some cases, different images in the image data correspond to different channels available in an imaging technology, such as an imaging flow cytometer. For example, the cytometry image data may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 16, 32, 64, 128, or more images. In such cases, each image may correspond to one, two, three, four, five, six, seven, eight, 9, 10, 16, 32, 64, or 128, or more channels available in an imaging technology, such as an imaging flow cytometer. The number of channels, and therefore the number of available images, may correspond to the underlying technology used to generate the image data, e.g., the capabilities or configuration of the imaging flow cytometer. In some cases, the imaging flow cytometer may be configured to acquire image data including one or more bright-field images (corresponding substantially to forward scatter (FSC)), one or more dark-field images (corresponding substantially to side scatter (SSC)), and one or more channels corresponding to, for example, fluorescence generated by one or more fluorescent labels stably bound to particles of the sample. In some cases, the imaging flow cytometer may be configurable such that the range of optical frequencies associated with each channel is adjustable. In embodiments, the optical frequencies detected in each channel, and therefore the optical frequencies associated with the image data images, may include any convenient range of optical frequencies. Embodiments of the present invention may receive images of any number of image channels having any range or bandwidth of optical frequencies detected in each channel, respectively.
[0026] FIG. 1A also shows schematic representations 120a, 120b, 120c, and 120d of cells. That is, images 110a, 110b, 110c, and 110d each include different representations 120a, 120b, 120c, and 120d of the same cell, with each representation varying based on the detected light frequency associated with each channel. While cells are shown as representations 120a, 120b, 120c, and 120d in image data 100, generally, any type of particle of interest may be imaged and subsequently classified in accordance with embodiments of the present invention, and such particles may vary. While each representation 120a, 120b, 120c, and 120d appears substantially similar in schematic form in FIG. 1A, they are not necessarily similar. In practice, the representation of a particle such as a cell across various channels may vary (e.g., in size, shape, texture, color, bright spots, dark spots, or the like) depending on the underlying morphology of the particle such as a cell and the frequency of light emitted across the particle within the field of view of the image data (e.g., one region of an imaged cell may emit, reflect, or scatter light of certain wavelengths, but not emit, reflect, or scatter light of other wavelengths).
[0027] 1A further illustrates how each image 110a, 110b, 110c, 110d has the same field of view. That is, each representation 120a, 120b, 120c, 120d of a cell is in substantially the same position within each image 110a, 110b, 110c, 110d. In other words, differences between images 110a, 110b, 110c, 110d relate to the light frequencies detected in each image (e.g., different channels of an imaging device used to generate the different images) and generally do not reflect, for example, different particles or different fields of view per particle. In embodiments of the present invention, the image size and / or orientation of the image data may be adjusted, e.g., cropped, resized, centered, or rotated, so that particles represented by the image data are substantially centered or at a substantially uniform position and orientation throughout the image data. That is, in some cases, the image data may be adjusted so that any particle represented in the image data appears in substantially the same position and orientation across multiple image data associated with multiple particles imaged in the sample. Such adjustments to the image data may, in some cases, facilitate more accurate classification of particles, e.g., cells, represented in the image data by more accurate or more effective training of models, e.g., machine learning models, to classify particles based on characteristics of the image data (as opposed to image artifacts introduced by inconsistent orientations of the imaged particles).
[0028] FIG. 1B shows nine rows of exemplary image data (i.e., one cell per row) of nine different cells for classification using embodiments of the present invention. Each image data corresponds to an event number 175 (i.e., event numbers 2, 3, 5, 7, 8, 13, 16, 20, and 25), which provides a reference number for the image data corresponding to an imaged event or particle, e.g., a cell, in the sample. In FIG. 1B, nine examples of image data 130, 135, 140, 145, 150, 155, 160, 165, and 170 are stacked vertically in rows. The exemplary image data has an event number 175 in the column and five different individual images displayed in horizontal columns 180, 185, 190, 193, and 196, respectively, as previously described. A FACSVulcan™ Fluorescence Imaging Activated Cell Sorter (Becton, Dickinson and Company) was used to capture the images in Figure 1B using FACSChorus 1.3.82.1 software.
[0029] FIG. 1B shows how each instance of exemplary image data 130, 135, 140, 145, 150, 155, 160, 165, 170 includes a collection of five images (e.g., image data 130 includes images 130a, 130b, 130c, 130d, 130e), each showing the same cell across a different image channel, i.e., a different range of wavelengths of light emitted, scattered, or reflected from the same cell.
[0030] Acquisition of image data and flow cytometry data As described above, embodiments of the present invention may be applied to image data collected using any convenient imaging technique, for example, an imaging flow cytometer. In some embodiments, image data of particles, e.g., cells, may be collected in conjunction with collecting light scattering data about such particles by analyzing such particles with flow cytometry (i.e., by conventional flow cytometry). Such imaging and light scattering-based techniques are described further below.
[0031] Flow cytometry A flow cytometer typically includes a sample reservoir for receiving a fluid sample, e.g., a sample containing particles, e.g., cells, for sorting or analysis, and a sheath reservoir containing a sheath fluid. The flow cytometer transports particles in the fluid sample (e.g., including cells from the sample) as a cell stream to a flow cell, while also directing the sheath fluid through the flow cell. To characterize the components of the flow stream, the flow stream is illuminated with light. Changes in the material in the flow stream, such as morphology or the presence of fluorescent labels, can alter the observed light, allowing for characterization and, in some cases, separation. Particles, e.g., molecules in fluid suspension, analyte-bound beads, or individual cells, pass through a detection region where the particles are exposed to excitation light, typically from one or more lasers, and the particles' light scattering and fluorescence properties are measured. Particles or their components are typically labeled with fluorescent dyes for ease of detection. Labeling multiple different particles or components with spectrally distinct fluorescent dyes may allow for simultaneous detection of different particles or components. In some embodiments, the analyzer is provided with multiple detectors, i.e., one detector for each scattering parameter to be measured and one or more detectors for each different dye to be detected. For example, some embodiments include a spectral configuration in which two or more sensors or detectors are used per dye. The resulting data includes the signal measured by each light scattering detector and the fluorescence emission. In certain embodiments, the flow cytometry assay may detect a signal indicative of the presence of a labeled secondary antibody in the sample.
[0032] light source As summarized above, a sample (e.g., in a flow stream of a flow cytometer) may be illuminated with light from a light source. In some embodiments, the light source is a broadband light source, emitting light having a wide range of wavelengths, e.g., 50 nm or greater, e.g., 100 nm or greater, e.g., 150 nm or greater, e.g., 200 nm or greater, e.g., 250 nm or greater, e.g., 300 nm or greater, e.g., 350 nm or greater, e.g., 400 nm or greater, e.g., 500 nm or greater. For example, one suitable broadband light source emits light having a wavelength between 200 nm and 1500 nm. Another example of a suitable broadband light source includes a light source that emits light having a wavelength between 400 nm and 1000 nm. When the method involves illumination with a broadband light source, broadband light source protocols of interest may include, but are not limited to, halogen lamps, deuterium arc lamps, xenon arc lamps, stabilized fiber-coupled broadband light sources, broadband LEDs with continuous spectra, superluminescent light emitting diodes, semiconductor light emitting diodes, broad spectrum LED white light sources, multi-LED integrated white light sources, or any combination thereof, among other broadband light sources.
[0033] In other embodiments, the methods involve irradiating using a narrowband light source that emits at a specific wavelength or narrow range of wavelengths, e.g., irradiating using a light source that emits light over a narrow range of wavelengths, such as 50 nm or less, for example 40 nm or less, for example 30 nm or less, for example 25 nm or less, for example 20 nm or less, for example 15 nm or less, for example 10 nm or less, for example 5 nm or less, for example 2 nm or less, e.g., irradiating using a light source that emits light of a specific wavelength (i.e., monochromatic light). When the methods involve irradiating using a narrowband light source, narrowband light source protocols of interest include, but are not limited to, narrow wavelength LEDs, laser diodes, or broadband light sources coupled to one or more optical bandpass filters, diffraction gratings, monochromators, or any combination thereof.
[0034] In some embodiments, the method uses one or more lasers to irradiate the sample. As discussed above, the type and number of lasers will vary depending on the sample and the desired light to be collected, and may be a gas laser, such as a helium-neon laser, an argon laser, a krypton laser, a xenon laser, a nitrogen laser, a CO laser, a CO laser, an argon-fluorine (ArF) excimer laser, a krypton-fluorine (KrF) excimer laser, a xenon-chlorine (XeCl) excimer laser, or a xenon-fluorine (XeF) excimer laser, or a combination thereof. In other cases, the method uses a dye laser, such as a stilbene laser, a coumarin laser, or a rhodamine laser, to irradiate the flowstream. In still other cases, the method irradiates the flowstream using a metal vapor laser, such as a helium-cadmium (HeCd) laser, a helium-mercury (HeHg) laser, a helium-selenium (HeSe) laser, a helium-silver (HeAg) laser, a strontium laser, a neon-copper (NeCu) laser, a copper laser, or a gold laser, and combinations thereof. In still other cases, the method irradiates the flowstream using a solid-state laser, such as a ruby laser, a Nd:YAG laser, a NdCrYAG laser, an Er:YAG laser, a Nd:YLF laser, a Nd:YVO4 laser, a Nd:YCa4O(BO3)3 laser, a Nd:YCOB laser, a titanium sapphire laser, a thulium YAG laser, a ytterbium YAG laser, a Yb2O3 laser, or a cerium-doped laser, and combinations thereof.
[0035] The sample may be illuminated using one or more of the light sources described above, e.g., two or more light sources, e.g., three or more light sources, e.g., four or more light sources, e.g., five or more light sources, e.g., ten or more light sources. The light sources may include any combination of light source types. For example, in some embodiments, the method uses an array of lasers to illuminate the sample in the flowstream, e.g., an array having one or more gas lasers, one or more dye lasers, and one or more solid state lasers.
[0036] The sample may be irradiated using a wavelength in the range of 200 nm to 1500 nm, e.g., 250 nm to 1250 nm, e.g., 300 nm to 1000 nm, e.g., 350 nm to 900 nm, e.g., 400 nm to 800 nm. For example, if the light source is a broadband light source, the sample may be irradiated using a wavelength of 200 nm to 900 nm. In other cases, if the light source includes multiple narrowband light sources, the sample may be irradiated using a specific wavelength in the range of 200 nm to 900 nm. For example, the light source may be multiple narrowband LEDs (1 nm to 25 nm), each independently emitting light having a wavelength in the range of 200 nm to 900 nm. In other embodiments, the narrowband light source includes one or more lasers (e.g., a laser array), and the sample is irradiated with a specific wavelength in the range of 200 nm to 700 nm using a laser array including, for example, gas lasers, excimer lasers, dye lasers, metal vapor lasers, and solid-state lasers as described above.
[0037] When two or more light sources are used, the light sources may be used to illuminate the sample simultaneously, sequentially, or a combination thereof. For example, each of the light sources may be used to illuminate the sample simultaneously. In other embodiments, each of the light sources is used to illuminate the flow stream sequentially. When two or more light sources are used to sequentially illuminate the sample, the time for which each light source illuminates the sample may independently be 0.001 microseconds or more, such as 0.01 microseconds or more, such as 0.1 microseconds or more, such as 1 microsecond or more, such as 5 microseconds or more, such as 10 microseconds or more, such as 30 microseconds or more, such as 60 microseconds or more. For example, the method may involve irradiating the sample using a light source (e.g., a laser) for a duration in the range of 0.001 microseconds to 100 microseconds, such as 0.01 microseconds to 75 microseconds, such as 0.1 microseconds to 50 microseconds, such as 1 microsecond to 25 microseconds, such as 5 microseconds to 10 microseconds. In embodiments in which two or more light sources are used to sequentially illuminate the sample, the duration for which the sample is illuminated by each light source may be the same or different.
[0038] The time between illumination by each light source may further vary as desired, and may be independently separated by a delay of 0.001 microseconds or more, e.g., 0.01 microseconds or more, e.g., 0.1 microseconds or more, e.g., 1 microsecond or more, e.g., 5 microseconds or more, e.g., 10 microseconds or more, e.g., 15 microseconds or more, e.g., 30 microseconds or more, e.g., 60 microseconds or more. For example, the time between illumination by each light source may be in the range of 0.001 microseconds to 60 microseconds, e.g., 0.01 microseconds to 50 microseconds, e.g., 0.1 microseconds to 35 microseconds, e.g., 1 microsecond to 25 microseconds, e.g., 5 microseconds to 10 microseconds. In some embodiments, the time between illumination by each light source is 10 microseconds. In embodiments in which the sample is illuminated sequentially by more than two (i.e., three or more) light sources, the delay between illumination by each light source may be the same or different.
[0039] The sample may be illuminated continuously or at discrete intervals. In some cases, the method involves illuminating the sample continuously with the light source. In other cases, the method involves illuminating the sample at discrete intervals with the light source, such as every 0.001 milliseconds, 0.01 milliseconds, 0.1 milliseconds, 1 millisecond, 10 milliseconds, 100 milliseconds, 1000 milliseconds, or other intervals.
[0040] Depending on the light source, the sample may be illuminated from a variety of distances, such as 0.01 mm or more, for example 0.05 mm or more, for example 0.1 mm or more, such as 0.5 mm or more, for example 1 mm or more, such as 2.5 mm or more, for example 5 mm or more, such as 10 mm or more, for example 15 mm or more, such as 25 mm or more, for example 50 mm or more. Furthermore, the angle or illumination may vary within the range of 10° to 90°, such as 15° to 85°, for example 20° to 80°, such as 25° to 75°, for example 30° to 60°, for example an angle of 90°.
[0041] In certain embodiments, the method irradiates the sample with two or more beams of frequency-shifted light. A light beam generator having a laser and an acousto-optical device for frequency-shifting the laser light may be used. In these embodiments, the method irradiates the acousto-optical device using a laser. Depending on the desired wavelength of light generated in the output laser beam (e.g., for use in irradiating a sample in a flow stream), the laser may have a specific wavelength within a range of 200 nm to 1500 nm, e.g., 250 nm to 1250 nm, e.g., 300 nm to 1000 nm, e.g., 350 nm to 900 nm, e.g., 400 nm to 800 nm. The acousto-optical device may be irradiated with one or more lasers, e.g., two or more lasers, e.g., three or more lasers, e.g., four or more lasers, e.g., five or more lasers, e.g., ten or more lasers. The lasers may include any combination of laser types. For example, in some embodiments, the method irradiates the acousto-optical device using an array of lasers, e.g., an array having one or more gas lasers, one or more dye lasers, and one or more solid-state lasers.
[0042] When two or more lasers are used, the lasers may be used to illuminate the acousto-optic device simultaneously, sequentially, or a combination thereof. For example, each of the lasers may be used to illuminate the acousto-optic device simultaneously. In other embodiments, each of the lasers may be used to illuminate the acousto-optic device sequentially. When two or more lasers are used to illuminate the acousto-optic device sequentially, the time for which each laser illuminates the acousto-optic device may independently be 0.001 microseconds or more, e.g., 0.01 microseconds or more, e.g., 0.1 microseconds or more, e.g., 1 microsecond or more, e.g., 5 microseconds or more, e.g., 10 microseconds or more, e.g., 30 microseconds or more, e.g., 60 microseconds or more. For example, the method may illuminate the acousto-optic device with a laser for a duration in the range of 0.001 microseconds to 100 microseconds, e.g., 0.01 microseconds to 75 microseconds, e.g., 0.1 microseconds to 50 microseconds, e.g., 1 microsecond to 25 microseconds, e.g., 5 microseconds to 10 microseconds. In embodiments in which two or more lasers are used to sequentially illuminate the acousto-optic device, the duration of illumination of the acousto-optic device by each laser may be the same or different.
[0043] The time between illumination by each laser may further vary as desired, and may be independently separated by a delay of 0.001 microseconds or more, e.g., 0.01 microseconds or more, e.g., 0.1 microseconds or more, e.g., 1 microsecond or more, e.g., 5 microseconds or more, e.g., 10 microseconds or more, e.g., 15 microseconds or more, e.g., 30 microseconds or more, e.g., 60 microseconds or more. For example, the time between illumination by each light source may be in the range of 0.001 microseconds to 60 microseconds, e.g., 0.01 microseconds to 50 microseconds, e.g., 0.1 microseconds to 35 microseconds, e.g., 1 microsecond to 25 microseconds, e.g., 5 microseconds to 10 microseconds. In some embodiments, the time between illumination by each laser is 10 microseconds. In embodiments in which the acousto-optic device is illuminated sequentially by more than two (i.e., three or more) lasers, the delay between illumination by each laser may be the same or different.
[0044] The acousto-optic device may be illuminated continuously or at discrete intervals. In some cases, the method uses a laser to illuminate the acousto-optic device continuously. In other cases, the laser illuminates the acousto-optic device at discrete intervals, such as every 0.001 milliseconds, 0.01 milliseconds, 0.1 milliseconds, 1 millisecond, 10 milliseconds, 100 milliseconds, e.g., 1000 milliseconds, or other intervals.
[0045] Depending on the laser, the acousto-optic device may be illuminated from a variety of distances, such as 0.01 mm or more, for example 0.05 mm or more, for example 0.1 mm or more, for example 0.5 mm or more, for example 1 mm or more, for example 2.5 mm or more, for example 5 mm or more, for example 10 mm or more, for example 15 mm or more, for example 25 mm or more, for example 50 mm or more, etc. Furthermore, the angle or illumination may vary within a range of 10° to 90°, such as 15° to 85°, for example 20° to 80°, for example 25° to 75°, for example 30° to 60°, for example an angle of 90°.
[0046] In an embodiment, a method applies high frequency drive signals to an acousto-optic device to generate angularly deflected laser beams. Two or more high frequency drive signals, such as three or more high frequency drive signals, for example four or more high frequency drive signals, for example five or more high frequency drive signals, for example six or more high frequency drive signals, for example seven or more high frequency drive signals, for example eight or more high frequency drive signals, for example nine or more high frequency drive signals, for example ten or more high frequency drive signals, for example fifteen or more high frequency drive signals, for example twenty-five or more high frequency drive signals, for example fifty or more high frequency drive signals, for example one hundred or more high frequency drive signals may be applied to the acousto-optic device to generate an output laser beam comprising a desired number of angularly deflected laser beams.
[0047] The angularly deflected laser beams generated by the high frequency drive signals each have an intensity based on the amplitude of the applied high frequency drive signal. In some embodiments, the methods apply high frequency drive signals having amplitudes sufficient to generate angularly deflected laser beams of a desired intensity. In some cases, the applied high frequency drive signals independently each have an amplitude within a range of about 0.001 V to about 500 V, e.g., about 0.005 V to about 400 V, e.g., about 0.01 V to about 300 V, e.g., about 0.05 V to about 200 V, e.g., about 0.1 V to about 100 V, e.g., about 0.5 V to about 75 V, e.g., about 1 V to about 50 V, e.g., about 2 V to about 40 V, e.g., about 3 V to about 30 V, or e.g., about 5 V to about 25 V. The applied high frequency drive signal in some embodiments has a frequency within the range of about 0.001 MHz to about 500 MHz, for example, about 0.005 MHz to about 400 MHz, for example, about 0.01 MHz to about 300 MHz, for example, about 0.05 MHz to about 200 MHz, for example, about 0.1 MHz to about 100 MHz, for example, about 0.5 MHz to about 90 MHz, for example, about 1 MHz to about 75 MHz, for example, about 2 MHz to about 70 MHz, for example, about 3 MHz to about 65 MHz, for example, about 4 MHz to about 60 MHz, for example, about 5 MHz to about 50 MHz.
[0048] In these embodiments, the angularly deflected laser beams of the output laser beam are spatially separated. Depending on the applied high frequency drive signal and the desired illumination profile of the output laser beam, the angularly deflected laser beams may be separated by 0.001 μm or more, such as 0.005 μm or more, such as 0.01 μm or more, such as 0.05 μm or more, such as 0.1 μm or more, such as 0.5 μm or more, such as 1 μm or more, such as 5 μm or more, such as 10 μm or more, such as 100 μm or more, such as 500 μm or more, such as 1000 μm or more, such as 5000 μm or more. In some embodiments, the angularly deflected laser beams overlap, for example, with adjacent angularly deflected laser beams along the horizontal axis of the output laser beam. The overlap of adjacent angularly deflected laser beams (e.g., overlap of beam spots) may be 0.001 μm or more, for example, 0.005 μm or more, for example, 0.01 μm or more, for example, 0.05 μm or more, for example, 0.1 μm or more, for example, 0.5 μm or more, for example, 1 μm or more, for example, 5 μm or more, for example, 10 μm or more, for example, 100 μm or more.
[0049] Photodetector In aspects of the method, scattered light or fluorescence is collected using a photodetector, such as a fluorescence photodetector. The fluorescence photodetector may optionally be configured to detect fluorescent emission from fluorescent molecules associated with particles in the flow cell, e.g., labeled specific binding members (e.g., labeled antibodies that specifically bind to a marker of interest). In certain embodiments, the method detects fluorescence from the sample using one or more fluorescence photodetectors, e.g., two or more, e.g., three or more, e.g., four or more, e.g., five or more, e.g., six or more, e.g., seven or more, e.g., eight or more, e.g., nine or more, e.g., ten or more, e.g., fifteen or more, e.g., twenty-five or more fluorescence photodetectors. In embodiments, each of the fluorescence photodetectors is configured to generate a fluorescence data signal. Fluorescence from the sample may be independently detected by each fluorescence photodetector over one or more wavelength ranges from 200 nm to 1200 nm. In some cases, the methods detect fluorescence from the sample over a wavelength range, e.g., 200 nm to 1200 nm, e.g., 300 nm to 1100 nm, e.g., 400 nm to 1000 nm, e.g., 500 nm to 900 nm, e.g., 600 nm to 800 nm. In other cases, the methods detect fluorescence at one or more specific wavelengths using each fluorescence detector. For example, depending on the number of different fluorescence photodetectors in the subject light detection system, fluorescence may be detected at one or more of 450 nm, 518 nm, 519 nm, 561 nm, 578 nm, 605 nm, 607 nm, 625 nm, 650 nm, 660 nm, 667 nm, 670 nm, 668 nm, 695 nm, 710 nm, 723 nm, 780 nm, 785 nm, 647 nm, 617 nm, or any combination thereof. In certain embodiments, the methods detect light at wavelengths corresponding to the fluorescence peak wavelengths of certain fluorophores present in the sample. In an embodiment, fluorescent flow cytometry data is received from one or more fluorescent light detectors (e.g., one or more detection channels), for example, two or more, for example, three or more, for example, four or more, for example, five or more, for example, six or more, for example, eight or more fluorescent light detectors (e.g., eight or more detection channels).
[0050] Light from the sample may be measured at one or more wavelengths, for example 5 or more different wavelengths, such as 10 or more different wavelengths, for example 25 or more different wavelengths, such as 50 or more different wavelengths, for example 100 or more different wavelengths, such as 200 or more different wavelengths, for example 300 or more different wavelengths, for example light collected at 400 or more different wavelengths may be measured.
[0051] The collected light may be measured continuously or at discrete intervals. In some cases, the method measures the light continuously. In other cases, the method measures the light at discrete intervals, such as every 0.001 milliseconds, 0.01 milliseconds, 0.1 milliseconds, 1 millisecond, 10 milliseconds, 100 milliseconds, 1000 milliseconds, or other intervals.
[0052] Measurements of collected light may be made one or more times during the subject methods, such as two or more times, such as three or more times, such as five or more times, such as ten or more times, In some embodiments, light propagation is measured two or more times, and the data is optionally averaged.
[0053] In certain embodiments, a method spectrally resolves light from each fluorophore of a fluorophore-biomolecule pair in a sample. In some embodiments, the overlap between each different fluorophore is determined and the contribution of each fluorophore to the overlapping fluorescence is calculated. In some embodiments, when spectrally resolving the light from each fluorophore, a spectral unmixing matrix of the fluorescence spectrum is calculated for each of multiple fluorophores whose fluorescence in the sample detected by the light detection system overlaps. In some cases, the process of spectrally resolving the light from each fluorophore and calculating a spectral unmixing matrix for each fluorophore may be used to estimate the abundance of each fluorophore, for example, to solve for the abundance of target cells in a sample.
[0054] In some embodiments, a method spectrally decomposes the light detected by the plurality of photodetectors, e.g., as described in U.S. Pat. No. 11,009,400, U.S. Patent Application Publication No. 2021 / 0247293, and U.S. Patent Application Publication No. 2021 / 0325292, the entire disclosures of which are incorporated herein by reference. For example, when spectrally decomposing the light detected by the plurality of photodetectors of the second set of photodetectors, the spectral unmixing matrix may be solved using one or more of: 1) a weighted least squares algorithm; 2) a Sherman-Morrison iterative inverse updater; 3) an LU matrix decomposition, e.g., decomposing a matrix into a product of a lower triangular (L) matrix and an upper triangular (U) matrix; 4) a modified Cholesky decomposition; 5) a QR decomposition; and 6) a weighted least squares algorithm calculation via singular value decomposition. In some embodiments, the method further characterizes the extent of light spillover detected by the plurality of photodetectors, for example, as described in U.S. Patent Application Publication No. 2021 / 0349004, the disclosure of which is incorporated herein by reference.
[0055] In some cases, the abundance of fluorophores associated with a targeted particle (e.g., chemically (i.e., covalently, ionically), or physically) is calculated from the spectrally resolved light from each fluorophore associated with the particle. For example, in one example, the relative abundance of each fluorophore associated with a targeted particle is calculated from the spectrally resolved light from each fluorophore. In another example, the absolute abundance of each fluorophore associated with a targeted particle is calculated from the spectrally resolved light from each fluorophore. In certain embodiments, particles may be identified or classified based on the relative abundance of each fluorophore determined to be associated with the particle. In these embodiments, particles may be identified or classified by any convenient protocol, such as by comparing the relative or absolute abundance of each fluorophore associated with the particle to a control sample having known particles, or by performing spectroscopic or other assay analysis of a population of particles (e.g., cells) having the calculated relative or absolute abundance of the associated fluorophore.
[0056] In certain embodiments, the method may select one or more particles (e.g., cells) of a sample identified based on the estimated abundance of a fluorophore associated with the particle. The term "selecting" is used herein in its conventional sense to refer to separating components of a sample (e.g., droplets containing cells, droplets containing non-cellular particles such as biopolymers) and, in some cases, directing the separated components to one or more sample collection vessels. For example, the method may select two or more components of a sample, e.g., three or more components, e.g., four or more components, e.g., five or more components, e.g., ten or more components, e.g., fifteen or more components, e.g., selecting twenty-five or more components of a sample. When selecting particles identified based on the abundance of a fluorophore associated with the particle, the method may acquire, analyze, and record data, e.g., using a computer, where multiple data channels record data from each detector used to obtain overlapping spectra of multiple fluorophore-biomolecule reagent pairs associated with the particle. In these embodiments, when analyzed, light from multiple fluorophores of a spectrally overlapping fluorophore-biomolecule reagent pair bound to a particle is spectrally resolved (e.g., by calculating a spectral unmixing matrix) and the particle is identified based on the estimated abundance of each fluorophore bound to the particle. The results of this analysis may be communicated to a sorting system configured to generate a set of digitized parameters based on the particle classification. In some embodiments, the methods for sorting components of a sample involve sorting particles (e.g., cells in a biological sample) as described in U.S. Pat. Nos. 3,960,449, 4,347,935, 4,667,830, 5,245,318, 5,464,581, 5,483,469, 5,602,039, 5,643,796, 5,700,692, 6,372,506, and 6,809,804, the disclosures of which are incorporated herein by reference.In some embodiments, the method sorts components of the sample using a particle sorting module such as those described in U.S. Patent No. 9,551,643, U.S. Patent No. 10,324,019, U.S. Patent Application Publication No. 2017 / 0299493, and WO 2017 / 040151, the disclosures of which are incorporated herein by reference. In certain embodiments, cells of the sample are sorted using a sorting determination module having a plurality of sorting determination units such as those described in U.S. Patent No. 11,085,868, the disclosure of which is incorporated herein by reference.
[0057] Flow cytometry assays are well known in the art. See, for example, Ormerod (ed.), Flow Cytometry: A Practical Approach, Oxford Univ. Press (1997); Jaroszeski et al. (eds.), Flow Cytometry Protocols, Methods in Molecular Biology No. 91, Humana Press (1997); Practical Flow Cytometry, 3rd ed., Wiley-Liss (1995); Virgo, et al. (2012) Ann Clin Biochem. Jan;49(pt 1):17-28; Linden, et. al., Semin Thromb Hemost. 2004 Oct;30(5):502-11; Alison, et al. J Pathol, 2010 Dec;222(4):335-344; and Herbig, et al. (2007) Crit Rev Ther Drug Carrier Syst., the disclosures of which are incorporated herein by reference. 24(3):203-255. In certain embodiments, compositions are assayed by flow cytometry using a flow cytometer capable of simultaneously exciting and detecting multiple fluorophores, such as a BD Biosciences FACSCanto™ flow cytometer used substantially according to the manufacturer's instructions. In the methods of the present disclosure, image cytometry may be performed as described in Holden et al. (2005) Nature Methods 2:773 and Valet, et al. 2004 Cytometry 59:167-171, the disclosures of which are incorporated herein by reference.
[0058] Suitable flow cytometry systems include those described in Ormerod (ed.), Flow Cytometry: A Practical Approach, Oxford Univ. Press (1997); Jaroszeski et al. (eds.), Flow Cytometry Protocols, Methods in Molecular Biology No. 91, Humana Press (1997); Practical Flow Cytometry, 3rd ed., Wiley-Liss (1995); Virgo, et al. (2012) Ann Clin Biochem. Jan;49(pt 1):17-28; Linden, et. al., Semin Thromb Hemost. 2004 Oct;30(5):502-11; Alison, et al. J Pathol, 2010 Dec;222(4):335-344; and Herbig, et al. (2007) Crit Rev Ther Drug Carrier Syst., the disclosures of which are incorporated herein by reference. 24(3):203-255.In some cases, flow cytometry systems of interest include the BD Biosciences FACSCanto™ flow cytometer, BD Biosciences FACSCanto™ II flow cytometer, BD Accuri™ flow cytometer, BD Accuri™ C6 Plus flow cytometer, BD Biosciences FACSCelesta™ flow cytometer, BD Biosciences FACSLyric™ flow cytometer, BD Biosciences FACSVerse™ flow cytometer, BD Biosciences FACSymphony™ flow cytometer, BD Biosciences LSRFortessa™ flow cytometer, BD Biosciences LSRFortessa™ X-20 flow cytometer, BD Biosciences FACSPresto™ flow cytometer, BD Biosciences FACSVia™ flow cytometer, and BD Biosciences FACSCalibur™ cell sorter, BD Biosciences FACSCount™ cell sorter, BD Biosciences FACSLyric™ cell sorter ...Verse™ flow cytometer, BD Biosciences FACSSymphony™ flow cytometer, BD Biosciences LSRFortessa™ flow cytometer, BD Biosciences LSRFor These include the BD Biosciences Via™ cell sorter, BD Biosciences Influx™ cell sorter, BD Biosciences Jazz™ cell sorter, BD Biosciences Aria™ cell sorter, BD Biosciences FACSAria™ II cell sorter, BD Biosciences FACSAria™ III cell sorter, BD Biosciences FACSAria™ Fusion cell sorter, BD Biosciences FACSMelody™ cell sorter, and BD Biosciences FACSymphony™ S6 cell sorter.
[0059] In some embodiments, the subject methods may be performed using methods such as those disclosed in U.S. Pat. Nos. 10,663,476, 10,620,111, 10,613,017, 10,605,713, 10,585,031, 10,578,542, 10,578,469, 10,481,074, 10,302,545, 10,145,793, 10,113,967, 10,006,852, 9,952,076, 9,933,341, 9,726,527, 9,453,789, 9,200,334, 9,097,640, 9,095, 94, U.S. Pat. No. 9,092,034, U.S. Pat. No. 8,975,595, U.S. Pat. No. 8,753,573, U.S. Pat. No. 8,233,146, U.S. Pat. No. 8,140,300, U.S. Pat. No. 7,544,326, U.S. Pat. No. 7,201,875, U.S. Pat. No. 7,129,505, U.S. Pat. No. 6,821,740, U.S. Pat. No. 6,813,017, U.S. Pat. No. 6,809,804, U.S. Pat. No. 6,372,506, U.S. Pat. No. 5,700,692, U.S. Pat. No. 5,643,796, U.S. Pat. No. 5,627,040, U.S. Pat. No. 5,620,842, U.S. Pat. No. 5,602,039, U.S. Pat. No. 4,987,086, U.S. Pat. No. 4,498,766 are used.
[0060] In some cases, the flow cytometry system of the present invention may be implemented using the techniques described in Diebold, et al. Nature Photonics Vol. 7(10); 806-810 (2013), as well as U.S. Pat. Nos. 9,423,353, 9,784,661, 9,983,132, 10,006,852, 10,078,045, 10,036,699, 10,222,316, 10,288,546, 10,324,019, 10,408,758, 10,451,538, 10,620,111, U.S. Patent Application Publication No. 200900222, and the like. The image data may be configured to image particles in a flow stream by fluorescence imaging using radio frequency tag emission (FIRE), as described in U.S. Patent Application Publication Nos. 2017 / 0133857, 2017 / 0328826, 2017 / 0350803, 2018 / 0275042, 2019 / 0376895, and 2019 / 0376894 (the disclosures of which are incorporated herein by reference). In accordance with embodiments of the present invention, the image data may include flow cytometry data of labeled particles, e.g., cells, obtained by a FIRE protocol using a FACSDiscover flow cytometer, e.g., as described in Schraivogel et al., Science Vol. 375(6578); 315-320 (2022). Image data may be obtained in part from fluorophores that have little effect on other detectors, such as conjugated polymer dyes BB515, BB550, BB790 (BD Biosciences).
[0061] Imaging Flow Cytometer In connection with obtaining cytometry image data for analysis, the flow stream may optionally be illuminated with multiple beams of frequency-shifted light, as described above, in Diebold, et al. Nature Photonics Vol. 7(10);806-810 (2013), as well as U.S. Pat. Nos. 9,423,353, 9,784,661, 9,983,132, 10,006,852, 10,078,045, 10,036,699, 10,222,316, 10,288,546, 10,324,019, 10,408,758, 10,451,538, 10,620,111, and U.S. Patent Application Publication No. 2009 / 0020909. Particles, e.g., cells, in a flow stream can be imaged by fluorescence imaging using radio frequency tag emission (FIRE) to generate a frequency-encoded image, as described in U.S. Patent Application Publication Nos. 17 / 0133857, 2017 / 0328826, 2017 / 0350803, 2018 / 0275042, 2019 / 0376895, and 2019 / 0376894 (the disclosures of which are incorporated herein by reference). In such cases, the flow cytometry data can include image data of particles, e.g., cells, present in the sample. See, for example, Schraivogel et al., Science Vol. 375(6578); 315-320 (2022), the disclosure of which is incorporated herein in its entirety, and U.S. Provisional Patent Application No. 63 / 256974, the disclosure of which is incorporated herein in its entirety.
[0062] Particle Classification FIG. 2A shows a flowchart 200 of a method for classifying cytometry image data using images of single cells in accordance with an embodiment of the present invention. Flowchart 200 is an exemplary embodiment of the present invention provided for illustrative purposes. Flowchart 200 is described in terms of classifying particles that are cells into cell types. However, embodiments of the present invention are not limited to classifying particles that are cells. Any convenient particle that can be imaged (e.g., by an imaging flow cytometer using the techniques described above), including but not limited to cells present in a sample, may be classified by applying embodiments of the present invention, and such particles may vary. Further, as described, flowchart 200 relates to classifying cells into cell types. However, embodiments of the present invention are not limited to classifying cells into cell types. Any convenient classification of particle image data for which a model (e.g., a convolutional neural network), such as the model described in the present invention, can be configured or trained to perform classification, including but not limited to cell types, may be applied to embodiments of the present invention, and such classifications may vary. For example, a cell image may be classified based on whether it represents an image of a singlet (single cell) or an image of a doublet (two cells). In other cases, a cell image may be classified based on whether it represents a cell that exhibits certain intracellular activity. In still other cases, a cell image may be classified based on whether it represents a cell that exhibits certain intercellular activity.
[0063] Flowchart 200 begins at step 205. Processing proceeds from start step 205 to next step 210.
[0064] At step 210, cytometry data is received. As described in detail above, the cytometry data of interest includes cytometry image data, such as cytometry image data including image data similar to the exemplary image data 100 shown in FIG. 1A or image data 130, 135, 140, 145, 150, 155, 160, 165, and 170 shown in FIG. 1B. The cytometry image data may be generated, i.e., detected, by an imaging flow cytometer, such as the exemplary flow cytometers capable of imaging particles described herein. The cytometry image data received at step 210 includes cytometry image data corresponding to multiple cells present in the sample. In other words, the cytometry image data includes multiple instances of the cytometry image data, each instance corresponding to a similarly imaged cell, and each instance of the cytometry image data includes multiple images corresponding to, for example, different channels of the imaging flow cytometer.
[0065] The cytometry data received in step 210 further includes light scatter and / or fluorescence characteristics of detected cells in the sample, e.g., data generated by a conventional flow cytometer or a non-imaging flow cytometer as described herein. The light scatter and / or fluorescence characteristics of detected cells received in step 210 correspond to each imaged cell; e.g., for each detected cell, the cytometry data available for such cell includes both light scatter and / or fluorescence characteristics and cytometry image data. Exemplary light scatter and / or fluorescence characteristics of detected cells in the sample include forward scatter (FSC) data, e.g., pulse height (FSC-H), pulse width (FSC-W), or pulse area (FSC-A), and side scatter (SSC) data, e.g., pulse height (SSC-H), pulse width (SSC-W), or pulse area (SSC-A). Exemplary light scattering and / or fluorescence properties of detected cells in a sample may further include data corresponding to fluorescence detected from cells in the sample (i.e., light emitted by one or more molecules, such as fluorescent dyes, stably bound to the detected cells).
[0066] The cytometry data received in step 210 includes unsorted cytometry data, meaning that the cytometry data corresponds to particles that have not yet been sorted as desired (e.g., including image data of cells that have not yet been sorted into cell types).
[0067] Once the flow cytometry data has been received in step 210, the flowchart 200 then proceeds to step 215.
[0068] At step 215, ground truth data is established. In embodiments, at step 215, such ground truth data is used to train a model for classifying particles in the cytometry image data. Ground truth data refers to using one or more of light scattering properties, fluorescence properties, or cytometry image data of the detected cells to establish characteristics of the detected cells. The ground truth data may be established, for example, using a subset of the cytometry data received at step 210, or may be established using a separate training data set including flow cytometry data that shares characteristics of the cytometry data received at step 210, as described above. For example, one or more of the light scattering properties, fluorescence properties, or imaging properties may be used to perform one or more of gating, clustering, or statistical analysis on the cytometry data to identify cell populations. The cytometry image data may be used to establish the ground truth data, for example, by using spatial image parameters of the cytometry image data, such as radial moment or eccentricity.
[0069] In some cases, light scatter characteristics may be used to identify ground truth data relevant to identifying images representing singlets versus doublets. For example, establishing ground truth data relevant to identifying singlets versus doublets may involve (i) plotting forward scatter pulse area (FSC-A) versus forward scatter height (FSC-H), and then (ii) applying a geometric gate to the data so plotted, where the gate separates singlets from non-singlets.
[0070] In some cases, establishing ground truth data for single cells versus doublets may involve (independently or in combination with light scattering data as described above) (i) plotting measured radial moments of images of cells observed in cytometry image data against measured eccentricity of images of cells observed in cytometry image data, and then (ii) applying gating or clustering methods to such plots to separate singlets from doublets. Radial moment may be calculated as the mean squared distance of the signal from the center of mass, and eccentricity may be calculated as the ratio of the magnitude of spread along the two principal components of the image.
[0071] The results of gating on the light scatter data (e.g., FSC-A vs. FSC-H as described above) and the results of gating or clustering on the cytometry image data (e.g., radial moment vs. eccentricity as described above), as well as the results of gating or clustering on the fluorescence data or applying statistical analysis to the fluorescence data (if any), may be combined to establish results regarding the cells detected in the sample from which the cytometry data is derived, i.e., ground truth data.
[0072] In embodiments, ground truth data, e.g., ground truth particle populations, e.g., ground truth cell populations (i.e., ground truth data for various cell populations in a sample), can be determined manually, for example, by gating or directed clustering. In other embodiments, ground truth data, e.g., ground truth cell populations, can be further determined without a priori knowledge of which particles, e.g., which cell subsets, are present. For example, ground truth data can be determined by invoking a dimensionality reduction algorithm and then invoking a clustering algorithm to identify distinct subsets. Such an approach identifies different subsets of particles, e.g., cell populations, as distinct from one another, but allows for the identification of different subsets of particles, e.g., cell populations, without actually knowing what such particles are (i.e., without knowing what the cell populations are). In still other embodiments, for machine learning applications such as neural network computations, ground truth data is established when a sufficient number of data points (images, i.e., subsets of images of the cytometry data received in step 210) are acquired per population such that subsets can be distinguished (by machine learning methods such as neural networks) in the presence of inherent variability in particles (e.g., cells). In such embodiments, at least 10, or at least 100, or at least 500, or at least 1,000, or at least 5,000 or more images per population are preferred, as the case may be.
[0073] In embodiments of the present invention, conventional flow cytometry data may be present and may be used in establishing ground truth data. However, the use of such conventional flow cytometry data is optional, as image-derived data (e.g., eccentricity or radial moment derived from image data) is available that may optionally be used in connection with establishing ground truth data. In some embodiments, data sets may be combined simply by classifying data sets with the same label. For example, singlets versus doublets may be determined by multiple user analyses using multiple gating methods, and then combined to provide input training data to a machine learning method, such as a neural network or other deep learning method.
[0074] In some cases, when establishing ground truth data, an operator or analyst may view a subset of the cytometry data and the associated ground truth for each detected particle, e.g., cell, in the cytometry data to visually confirm that the gating, clustering, or statistical analysis applied to establish such ground truth resulted in the expected image (i.e., the expected classification or segmentation of the image). For example, an operator or analyst may view a subset of the cytometry image data and the associated results of establishing the ground truth data to confirm that cells were correctly identified and labeled. That is, in the above example of identifying a single cell (singlet) versus an image of a doublet, an operator or analyst may review the subset of the cytometry image data to confirm that the single cell (singlet) was distinguished (appropriately labeled) from the image of the doublet.
[0075] In embodiments, any convenient technique may be used to train the model using ground truth data, and such training techniques are described in more detail herein.
[0076] Once the ground truth data has been established in step 215 and, in some cases, the model has been trained, the flowchart 200 then proceeds to step 220 .
[0077] At least one image of the image data is selected in step 220. The image of the image data may be, for example, any one of image 110a, image 110b, image 110c, or image 110d of the exemplary cytometry image data 100 shown in FIG. 1A. Selecting an image of the cytometry image data identifies the image for subsequent processing, as described in more detail below.
[0078] The images of the cytometry image data may be selected using any convenient technique, and such images may be varied. In some cases, the images of the image data may be selected randomly. In some cases, the images of the image data may be selected based on image characteristics, such as an estimate of the effect of image adjustments applied to the images of the image data. That is, the images may be selected based on an estimate of how the images may be affected by certain image adjustment techniques. For example, the images may be selected based on an estimate that the images may be substantially affected, i.e., altered, by the application of a "sharpening" image filter or a "blurring" image filter. In other cases, the images of the image data may be selected based on image characteristics that estimate how different they are from other images of the image data. That is, the images may be selected because they are different from the other images of the image data (i.e., the most unique of the images of the image data).
[0079] As described in more detail below, in some cases, flowchart 200 may loop (i.e., repeat) after step 235 back to step 220, where the same or a different image of the image data may be selected in step 220. In some cases, the image selected may be based in part on whether an image was previously selected in conjunction with the results of applying the model to the image data (i.e., the results in step 230), as described in more detail below.
[0080] At least one image of the cytometry image data selected in step 220 may be identified as the selected image using any convenient technique, for example, using a label based on one or more channels associated with the selected image or images.
[0081] Upon completion of the selection of at least one image of the cytometry image data in step 220 , flowchart 200 then proceeds to step 225 .
[0082] In step 225, at least one aspect of the at least one image selected in step 220 is adjusted. Adjusting an aspect of an image means modifying the image in any convenient manner, and such methods may vary. In general, in embodiments, desirable changes to the at least one image are changes that result in improving the effectiveness of an image classification model, such as those described below. Adjusting at least one image of the cytometry image data may improve the effectiveness of a classification model, e.g., by highlighting, enhancing, or clarifying features of the image data that represent important distinguishing characteristics of the imaged particles, e.g., cells. For example, image adjustment techniques may be applied to the image to focus on the shape or other characteristics of the imaged cells, where such shape or other features are indicative of cell type. Concurrently or selectively, adjusting at least one image of the cytometry image data may improve the effectiveness of a classification model, e.g., by reducing distinctive features of the at least one image of the cytometry image data that correspond only to, e.g., artifacts of, the image or image acquisition technique, as opposed to features that truly represent distinctive characteristics of the particles, e.g., cells. For example, if aspects of the shape or other characteristics of an imaged cell do not accurately reflect the imaged cell but are indicative of a different cell type, applying an image adjustment technique to the image may reduce or eliminate such aspects of the shape or other characteristics of the imaged cell. In other words, an image adjustment may have the effect of amplifying true salient characteristics of the imaged particle, e.g., cell, or reducing (i.e., minimizing) false salient characteristics of the imaged particle, e.g., cell. Furthermore, adjusting at least one image of the cytometry image data may improve the effectiveness of a classification model, e.g., by highlighting features or characteristics present in two or more images of the cytometry image data. That is, the adjustment technique may take advantage of image features by highlighting additional information about features detected across multiple imaging channels.
[0083] Image adjustments that have the effect of improving the effectiveness of a model may vary widely, may depend on the characteristics of the underlying model and the cytometry image data, and are unlikely to be fully known a priori. Thus, as described in detail below, adjusting one or more aspects of at least one image of cytometry image data comprises an iterative process, where repeated iterations enable an empirical approach to discovering effective adjustments to the image (e.g., by searching or sampling a space of possible image adjustments). In embodiments, image adjustment techniques may achieve image enhancement, i.e., improving the detectability of image features, such as salient characteristics of cell images that correspond to classification criteria for classifying cells by cell type. In other embodiments, image adjustment techniques may achieve image restoration, i.e., restoring a degraded image by reversing the degradation process. In still other embodiments, adjusting an image may improve the effectiveness of a model by removing noise, sharpening boundaries, or highlighting distinct regions of image pixels.
[0084] In some cases, adjusting an aspect of the at least one image involves applying a digital signal processing technique to the at least one image. Digital signal processing refers to performing one or more signal processing operations using digital processing (by a computer or other digital system, e.g., a dedicated digital signal processor, e.g., a graphics processor). In other cases, adjusting an aspect of the at least one image involves applying a noise reduction technique to the at least one image. In still other cases, adjusting an aspect of the at least one image involves applying one or more of the following techniques to the at least one image: adding color, removing color, blurring the image, sharpening the image, or defining edges. In still other cases, adjusting an aspect of the at least one image involves applying an image filter to the at least one image. Image filters of interest include, for example, one or more of a linear filter, a nonlinear filter, a low-pass filter, a high-pass filter, a spatial filter, or a frequency filter. Image filters of interest may also include, for example, one or more of a smoothing filter, a box filter, a Gaussian filter, or a Laplacian filter. Further details regarding image adjustment techniques can be found in Shapiro, LG and Stockman, GC (2001). Computer Vision. New Jersey: Prentice Hall, and Bishop, CM (2006). Pattern Recognition and Machine Learning. New York: Springer, the disclosures of each of which are incorporated herein in their entirety. In embodiments, the image filter may further perform thresholding (e.g., ignoring pixels above or below any desired threshold, e.g., intensity or color or grayscale value threshold) or inversion (e.g., turning black pixels into white pixels or vice versa, light gray pixels into dark gray pixels or vice versa).
[0085] A particular image adjustment technique may include a type of image adjustment (e.g., a technique that has a blurring or sharpening effect on an image) and one or more corresponding parameters, such as the degree, magnitude, or extent to which the image adjustment technique is applied. For example, a noise reduction technique may be applied along with parameters that specify the number of neighbors of each pixel to observe and consider when performing the noise reduction technique. In embodiments, the image adjustment technique, and any associated parameters of the image adjustment technique, may be selected using any convenient technique, and such techniques and parameters may vary. Selecting an image adjustment technique means selecting the adjustment technique itself and any associated parameters necessary to configure the technique.
[0086] In some cases, an image adjustment technique may be selected randomly. In other cases, an image adjustment technique may be selected based on characteristics of the image to which the image adjustment technique is applied, such as an estimate of the effect of the image adjustment applied to the image. That is, an image adjustment technique may be selected based on an estimate of how the image may be affected by a certain image adjustment technique. For example, an image adjustment technique may be selected based on an estimate that the image may be substantially affected, i.e., altered, by the application of a "sharpening" image filter or a "blurring" image filter. In other cases, an image adjustment technique may be selected based on image characteristics present in the image that are not present in other images in the image data. That is, an image adjustment technique may be selected to further accentuate or amplify, or otherwise minimize, differences (i.e., potential salient features) of the image compared to other images in the image data.
[0087] As described in more detail below, in some cases, flowchart 200 may loop (i.e., repeat) after step 235 back to step 225, where the same or a different image adjustment technique may be selected in step 225. In some cases, the image adjustment technique selected may be based in part on whether an image adjustment technique was previously selected in conjunction with the results of applying the model to the image data (i.e., the results in step 230), as described in more detail below.
[0088] In some cases, one image adjustment technique is applied to at least one image of the image data selected in step 220. In other cases, more than one image adjustment technique is applied to at least one image of the image data selected in step 220. That is, in some cases, more than one image adjustment technique may be applied to one of the images of the image data selected in step 220. The application of the image adjustments in step 225 results in at least one image being adjusted by at least one image adjustment technique. Such one or more images, together with other images of the cytometry image data, may be referred to as adjusted image data or adjusted cytometry image data.
[0089] Once adjustment of at least one aspect of the selected image in step 225 is complete, flowchart 200 then proceeds to step 230 .
[0090] In step 230, a model is applied to the adjusted cytometry image data (i.e., the cytometry image data includes at least one image adjusted in step 225). Applying a model refers to applying a model configured to estimate or predict, for example, cell classification, based on the adjusted image data generated in step 225. That is, when the model is applied to the cytometry image data, it may estimate or predict whether a particle, e.g., a cell, belongs to a first category of particles, e.g., a first cell type. In some cases, when the model is applied to the cytometry image data, it may estimate or predict whether a particle, e.g., a cell, belongs to one of multiple categories of particles, e.g., multiple cell types. In embodiments, applying the model involves providing input to the model including the adjusted image data generated in step 225 and receiving as output an estimate as to whether the adjusted image data represents particles, e.g., cells, belonging to one or more categories of particles, e.g., cell types. In some cases, the estimate provided by applying the model includes an indication as to whether a particle represented in the adjusted image data belongs to a first category of particles, or an identification of the category to which the particle represented in the adjusted image data belongs.
[0091] In such embodiments, the estimate provided by applying the model may further include an indication of the accuracy, confidence, or likelihood that the estimated category applies to the particle represented in the adjusted image data. That is, the result of applying the model to the adjusted image data may include a category-related result (e.g., that a cell belongs to a first category of cell type) and a score representing the degree of confidence associated with such category-related result. The latter result may be referred to, for example, as an accuracy score or confidence score. In some cases, the result of the model's classification for any one event is returned as a probability that the event belongs to class 1, class 2, etc., and such probability may be a confidence score for such application of the model. In such cases, such confidence scores may be expressed as additional parameters for each class and can thus be used further in the analysis, i.e., inspected, plotted, exported, etc.
[0092] Exemplary results of applying the model to the adjusted cytometry image data may include, for example, (x) a result that the adjusted cytometry image data represents particles, e.g., cells, belonging to a first category, e.g., a first cell type, with a confidence score of 60%, or (y) a result that the adjusted cytometry image data represents particles, e.g., cells, belonging to a first category, e.g., a first cell type, with a confidence score of 90%. In the case of example (x) and example (y), example (y) includes a higher confidence score and therefore represents a more reliable estimate for the adjusted image data insofar as the model exhibits a higher confidence or precision relative to the category-based estimates.
[0093] Exemplary results of applying the model to the adjusted cytometry image data may further include, for example, (m) a result in which the adjusted cytometry image data does not represent (i.e., does not include images of) particles, e.g., cells, belonging to a first category, e.g., a first cell type, with a confidence score of 70%, or (n) a result in which the adjusted cytometry image data does not represent (i.e., does not include images of) particles, e.g., cells, belonging to a first category, e.g., a first cell type, with a confidence score of 80%. In examples (m) and (n), example (n) includes a higher confidence score and therefore represents a more reliable estimate for the adjusted image data insofar as the model exhibits a higher confidence or precision associated with the category-based estimates.
[0094] Exemplary results of applying the model to the adjusted cytometry image data may further include, for example, (j) the adjusted cytometry image data represents particles, e.g., cells, belonging to a first category, e.g., a first cell type, with a 90% confidence score, or (k) the adjusted cytometry image data does not represent particles, e.g., cells, belonging to a first category, e.g., a first cell type, with a 80% confidence score (i.e., does not include images of such particles, e.g., cells). For examples (j) and (k), even though in example (j) the model estimates particles belonging to the first category and in example (k) the model estimates particles not belonging to the first category, example (j) includes a higher confidence score and therefore represents a more reliable estimate for the adjusted image data insofar as the model exhibits a higher confidence or precision associated with the category-based estimates.
[0095] Exemplary results of applying the model to the adjusted cytometry image data may further include, for example, (e) a result in which the adjusted cytometry image data represents particles, e.g., cells, belonging to a first category, e.g., a first cell type, with a 60% confidence score, or (f) a result in which the adjusted cytometry image data does not represent particles, e.g., cells, belonging to a first category, e.g., a first cell type, with a 60% confidence score (i.e., does not include images of such particles, e.g., cells). In examples (e) and (f), even though in example (e) the model estimates particles belonging to the first category and in example (f) the model estimates particles not belonging to the first category, examples (e) and (f) include equal confidence scores and are therefore equally reliable estimates insofar as the model exhibits equal confidence or precision associated with the category-based estimates.
[0096] In embodiments, it is possible to measure the accuracy of a model with respect to training data, such as the ground truth data described herein (e.g., using loss of other accuracy statistics). In other embodiments, it is possible to measure how well a model performs with respect to new data (i.e., data other than the training data or data other than the ground truth data described herein). Measuring how well a model performs with respect to training data is well established because the correct, expected, or desired results are known in advance, and accurate statistics can be calculated based on how well the model performs with respect to the correct results. In embodiments, measuring how well a model performs with respect to new, unseen data may, in some cases, be somewhat subjective or impossible. In some cases, it may be possible to manually validate a subset of the results of applying the model to the new, unseen data.
[0097] Any convenient model capable of classifying the cytometry image data, i.e., capable of estimating the presence of particles belonging to the first category of particles in the cytometry image data, may be applied, and such models may vary. Models of interest include models capable of estimating the presence of particles belonging to the first category of particles and providing a confidence score associated with the estimate, where a higher (lower) confidence score corresponds to a higher (lower) likelihood that the estimate of the presence of particles belonging to the first category of particles in the cytometry image data is accurate. In embodiments, the model includes one or more of a statistical model, a linear model, a computational model, or a machine learning model. Machine learning models of interest include one or more of a tree-based model, or an artificial neural network (i.e., neural network), such as a convolutional neural network, or a deep learning model.
[0098] An artificial neural network of interest (i.e., a computing system inspired by biological neural networks in animal brains) may include a collection of interconnected units, or nodes, called artificial neurons. The connections (or edges) between nodes (such connections are analogous to synapses in biological brains) can send signals to other nodes. Each node may receive and process signals, and then send signals to other nodes connected to it. The "signals" at the connections may be real numbers, and each node's output may be calculated by a nonlinear function of the sum of its inputs. Nodes and connections typically have weights that are adjusted in conjunction with training the neural network model, which involves a learning process. These weights increase or decrease the strength of the signals at the connections. Nodes may have thresholds such that a signal is sent to the connected node only if the summed input signals exceed such thresholds. Nodes may be organized into layers. Different layers may perform different transformations on the input. Signals proceed from the first layer (input layer) to the last layer (output layer) after possibly passing through multiple hidden layers, possibly multiple times.
[0099] Other models of interest include support vector machines, random forest algorithms, or other decision tree algorithms, or other models that have been trained using, for example, supervised learning techniques.
[0100] Further details regarding models of interest, including artificial neural networks (and training models of interest), can be found in Mitchell, T.M. (1997). Machine Learning. San Francisco, California: McGraw-Hill, and Goodfellow, I., Bengio, Y. and Courville, A. (2016). Deep Learning. Cambridge, Massachusetts: The MIT Press, the disclosures of each of which are incorporated herein in their entirety.
[0101] In an embodiment of the present invention, a model is trained to estimate the presence of particles belonging to a first category of particles in cytometry image data. By training, we mean configuring the model, if necessary, based on an underlying model to generate such estimates based on cytometry image data. For example, as described above, when training a model that is an artificial neural network, weights corresponding to connections between nodes may be identified, and such weights influence the signals transmitted between the nodes.
[0102] In embodiments, the model may be trained to predict the presence of particles belonging to a first category of particles in the cytometry image data using one or more of unsupervised learning techniques, semi-supervised learning techniques, supervised learning techniques, or round robin training techniques.
[0103] Unsupervised learning is a model training technique in which a model is trained to identify or recognize patterns (e.g., whether cytometry image data represents particles belonging to a first category of particles). In unsupervised learning, a model is trained without pre-assigned labels for the data used to train the model. That is, in embodiments of the present invention, labels indicating whether particles belonging to a first category of particles are present in the cytometry image data are not provided in connection with training the model. As a result, when unsupervised learning is applied to train a model, the model itself discovers patterns in the training data.
[0104] In another embodiment, a supervised learning technique is applied to the model when training the model. Supervised learning is a model training technique in which a model is trained to map inputs to outputs based on examples of input-output pairs (i.e., labeled training data). As a result of applying supervised learning techniques to a model, the model infers a function from labeled training data that includes a set of training examples. Such examples include a pair of input objects (e.g., cytometry image data) and a desired output value (e.g., whether a particle represented in the cytometry image data belongs to a first category of particles). Such a desired output value may be referred to as a monitoring signal.
[0105] In other embodiments, the model is trained using semi-supervised learning techniques. Semi-supervised learning is a machine learning technique that trains a model to, for example, identify or recognize patterns. In semi-supervised learning, the model is trained using both labeled and unlabeled training data.
[0106] In some cases, when training a model, a round-robin training technique is applied to the model. Round-robin training technique means that the input data to the model (i.e., the data used to train the model) is divided into multiple partitions, at least a first partition is used to train the model, and one or more remaining partitions are used to generate predictions using the trained model. This process may be repeated, and partitions of data already used to train the model are then used to generate predictions using the trained model. The round-robin training approach may have advantages such as identifying which datasets used for training lead to more accurate predictions.
[0107] In embodiments of the present invention, a model may be trained to estimate the presence of particles belonging to a first category of particles in the cytometry image data using all or a subset of the ground truth data generated in step 215. Such data may be used in conjunction with the application of any convenient training technique, including supervised, semi-supervised, or unsupervised learning, as described above. In embodiments, a model may be trained using the ground truth data established in step 215 using any convenient training technique, as described herein. That is, in some embodiments, step 215 may further train a model using the ground truth data established in step 215.
[0108] In some cases, the model may be updated (i.e., further trained or refined) based on the results of applying the model to the adjusted cytometry image data in step 225. In other cases, the model may be updated (i.e., further trained or refined) based on the results of repeatedly applying step 225 (i.e., returning to step 225 one or more times via step 235, as described below). In still other cases, the model may not be updated when applied to non-training data (e.g., data other than data used in connection with establishing the ground truth data in step 215). That is, the model may be trained in step 215 as described above, and then the model may remain unchanged when applied to data other than the ground truth data described herein.
[0109] The application of the model to the adjusted cytometry data in step 230 results in an estimate of whether a particle, e.g., a cell, belongs to a first category, and optionally an associated confidence score corresponding to such category estimate. Such results may be associated with the particular adjusted cytometry data and applied to the model. That is, the output from applying the model may be associated with information about the adjusted cytometry image data, including images containing the cytometry image data, information identifying which of the at least one image was selected for image adjustment in step 220, and information identifying which of at least one aspect of the selected image was adjusted (i.e., information identifying details about the image adjustment technique, such as the type and configuration of the image adjustment technique). Such information, in connection with subsequently identifying at least one image for further image adjustment in step 220 and adjusting at least one aspect of the image so selected in step 225, may in either case be recorded for reference when further iterating through such steps, i.e., returning to such steps via step 235, as described below.
[0110] Once the model has been applied to the adjusted image data in step 230 , the flowchart 200 then proceeds to step 235 .
[0111] In step 235, the results of applying the model in step 220 are evaluated to determine whether such results are sufficient, or alternatively, to determine whether further iterations are necessary by selecting at least one image of the cytometry image data in step 220, adjusting aspects of at least one image in step 225, and applying the model to the further adjusted image data in step 230. By sufficient, we mean any convenient technique for evaluating the results of applying the model in step 230. In some cases, the results of applying the model in step 230 are evaluated to determine whether a confidence score associated with the results of applying the model in step 230 meets a particular threshold. In some cases, the threshold may reflect a predetermined confidence score (i.e., a reference confidence score) associated with the results of applying the model. In some cases, the predetermined confidence score corresponds to the confidence score associated with applying the model to the cytometry image data without the image adjustments performed in steps 220 and 225; i.e., qualitatively, the reference score may relate to whether the image adjustments performed in steps 220 and 225 improved the results of applying the model in step 230. In other cases, the threshold may reflect an adaptive confidence score associated with the results of applying the model (e.g., a moving average or the maximum of previously obtained confidence scores). In some cases, such a threshold may be a confidence score obtained by applying the model in a previous iteration, e.g., the immediately preceding iteration, via steps 220, 225, 230, and 235. In other cases, the threshold may be related to the number of times the process has iterated through steps 220, 225, 230, and 235 (e.g., whether a minimum number of iterations has been performed). In still other cases, the threshold may be related to the difference between the confidence score and one or more previously obtained confidence scores based on the results of applying the model in step 230 during a previous iteration.In some cases, the determination in step 235 of whether the results of applying the model in step 230 are sufficient may relate to whether the results appear to have plateaued (i.e., an indication that iterations through steps 220, 225, 230, and 235 no longer provide an advantage in terms of improved or more accurate results from applying the model in step 230). In some cases, the determination in step 235 of whether the results of applying the model in step 230 are sufficient may be determined randomly, for example, by flipping a coin. As described above, any convenient technique for determining whether to apply the results of applying the model to the adjusted image data in step 230, and various techniques for making such a determination, may be applied in various iterations through steps 220, 225, 230, and 235. For example, the technique for determining whether the results are sufficient may vary based on, for example, the number of times the process has iterated through steps 220, 225, 230, and 235, or the maximum accuracy of the results obtained during prior iterations. In embodiments, the technique used to make such a determination may be selected to refine the estimate of the presence of particles belonging to the first category of particles in the cytometry image data. That is, the technique used to make such a determination may be selected to facilitate exploring the maximum effectiveness of applying the model across a search space that includes all possible image adjustments of all combinations of images in the cytometry image data. Because such a search space may be unwieldy, heuristic-based techniques may be applied to determine whether the results are sufficient. For example, in determining whether the results of applying the model in step 230 are sufficient in this manner, simulated annealing or other probabilistic techniques may be applied to approximate a global optimum of the results of applying the model in step 230 for all available image selections in step 220 and image adjustments in step 225.
[0112] In some embodiments, it may be desirable for flowchart 200 to repeat steps 220, 225, 230, and 235 a certain number of times (each iteration referred to as an epoch) so that, when evaluating the results of applying the model in step 235, it can be determined whether flowchart 200 has repeated steps 220, 225, 230, and 235 a desired, i.e., predetermined, number of times (which may be any convenient number of iterations, e.g., one or more iterations).
[0113] In other embodiments, when evaluating the results of applying the model in step 235, it may be determined that a certain confidence score, e.g., a threshold confidence score, which may be any convenient confidence score, has been achieved, or that a certain accuracy of the model, e.g., a threshold accuracy, which may be any convenient accuracy, has been achieved. When such a threshold confidence score or threshold accuracy has been achieved, the results of evaluating the results of applying the model in step 235 include comparing the results of applying the model to the predetermined threshold confidence score or accuracy result of applying the model.
[0114] In yet other embodiments, when evaluating the results of applying the model in step 235, it may be determined that the confidence score or accuracy of the results of applying the model has decreased. That is, when evaluating the results of applying the model, it may be determined that the most recent iteration of applying the model achieved a lower confidence score or other determination of accuracy of applying the model than previous applications of the model, or that the results of applying the model may be trending downward. In such cases, the results of applying the model in step 235 may indicate that further application of the model would not be beneficial for classifying the cytometry data, and therefore further iterations of applying the model may be terminated.
[0115] In yet another embodiment, when evaluating the results of applying the model in step 235, it may be determined whether each iteration of applying the model appears to have resulted in overfitting the data. Such a scenario may occur when the model "memorizes" the training data, i.e., when the model performs very well on the training data but worsens on the validation data. In some cases, the image data selected in step 220 may include the training data, and the model may be updated based on the results of applying the model to such training data, either at each interval that the model is applied to the training data, or at another convenient interval. In such a case, when the model appears to have "memorized" the training data, the results of applying the model in step 230 may indicate that further application of the model would not be useful for classifying the cytometry data, and further iterations of applying the model may be terminated.
[0116] If the results of applying the model in step 230 are determined to be satisfactory, flowchart 200 then proceeds to step 240, where the process ends. The final result of applying the model to the adjusted image data includes an estimate of whether the adjusted image data contains particles, e.g., cells, belonging to a first category of particles, e.g., a cell type. That is, the final result of applying flowchart 200 includes a final estimate of whether the cytometry image data represents particles belonging to the first category of particles. In embodiments, such final result is the result of applying the model in step 230, and such result is deemed to be the most likely to be accurate among the results obtained in each iteration of step 230. That is, the final result may be the result of applying the model in step 230 along with the corresponding maximum confidence score for each iteration of applying the model in step 230.
[0117] If step 230 determines that the results of applying the model are not satisfactory, flowchart 200 proceeds to step 220 and performs subsequent iterations through steps 220, 225, 230, and 235.
[0118] As described above, each iteration through step 220 may select the same or different image(s) for adjustment. Any convenient technique for determining whether the same or different image(s) are selected for a subsequent iteration may be applied, and such techniques may vary. In some cases, the same (different) image(s) may be selected based on whether the results of applying the model over previous iterations improved (decreased) the accuracy of the results of applying the model in step 230. For example, if the same image(s) were selected for adjustment in each of multiple previous iterations of applying the model in step 230, and the accuracy of applying the model to the adjusted image data in step 230 improved, the same image(s) may again be selected in step 220 on the basis that continuing to adjust the selected image(s) is a path to a local maximum (e.g., a local maximum confidence score) in the results of applying the model in step 230. In some cases, the same (different) image(s) may be selected based on the number of times such images were already selected during previous iterations of step 220, i.e., each image of the cytometry image data is selected at least a minimum number of times and / or no more than a maximum number of times.
[0119] As described above, each iteration through step 225 may adjust the same or different aspects of the selected image or images. Any convenient technique for determining whether the same or different image adjustment technique is selected for a subsequent iteration may be applied, and such techniques may vary. In some cases, the same (different) image adjustment technique is selected based on whether the results of applying the model over previous iterations improved (decreased) the accuracy of the results of applying the model in step 230. For example, if the same or similar image adjustment technique was applied in each of multiple previous iterations of applying the model in step 230, improving the accuracy of applying the model to the adjusted image data in step 230, the same or similar image adjustment technique may again be selected in step 220 on the theory that continuing to adjust the selected images in a similar manner is a path to a local maximum (e.g., a local maximum confidence score) in the results of applying the model in step 230. In some cases, when applying the same or similar image adjustment technique, the same type of technique is applied, but such technique is configured in a different manner. For example, the image blurring technique may be reapplied in subsequent iterations of step 225, blurring the selected image to a greater extent with each iteration of step 225. Alternatively, the image blurring technique may be reapplied to the same extent in subsequent iterations of step 225 to another image of the cytometry image data selected in step 220. In some cases, the same (or different) image adjustment technique is selected based on the number of times such image adjustment technique has already been selected during prior iterations of step 225, i.e., each available image adjustment technique is selected at least a minimum number of times and / or no more than a maximum number of times.
[0120] As described above, after subsequent iterations of steps 220, 225, 230, and 235, if step 235 determines that the results of applying the model in step 230 are satisfactory, flowchart 200 proceeds from step 235 to step 240, where flowchart 200 ends. Upon completion, the flowchart has completed an estimation of whether a particle, e.g., a cell, represented in the cytometry image data belongs to a first category of particle, e.g., a cell type. Such a final result may be an estimation obtained from any of the iterations through steps 220, 225, 230, and 235. That is, the final result obtained upon completion of flowchart 200 does not necessarily have to be the result corresponding to, e.g., the last iteration in time through steps 220, 225, 230, and 235, but instead may be the result corresponding to the iteration that yielded the most likely accurate result of each of the iterations (e.g., the result with the highest confidence score). In other words, the results obtained upon completion of flowchart 200 correspond to a search for maximum accuracy of results over a search space that includes all available image adjustments of all available images of the cytometry image data.
[0121] FIG. 2B shows a flowchart 250 of a method for training a model to classify cytometry image data using an image of a single cell according to another embodiment of the present invention. Flowchart 250 is an exemplary embodiment of the present invention provided for illustrative purposes. Flowchart 250 relates to classifying an image of a cell into a classification of singlets or doublets. However, embodiments of the present invention are not limited to cell classification. Any convenient particle that can be imaged (e.g., by an imaging flow cytometer), including but not limited to cells present in a sample, may be classified by applying embodiments of the present invention, including the embodiment shown in FIG. 2B, and such particles may vary. Furthermore, as described, flowchart 250 relates to classifying an image of a cell into a single cell versus a doublet (singlet versus doublet). However, embodiments of the present invention, including the embodiment shown in FIG. 2B, are not limited to such classification of cells. Any convenient classification of particle image data that a model, such as a model described herein (e.g., a convolutional neural network), can be configured or trained to classify (e.g., into cell types) may be applied to embodiments of the present invention, and such classifications may vary.
[0122] Flowchart 250 provides additional details regarding training a model using single-cell images to classify cytometry image data, and in particular training a model using conditioned image data to classify cytometry image data, in connection with embodiments of the present invention. Certain image processing techniques described below in connection with training a model (i.e., image processing as applied to ground truth data, as described herein) are also applicable to embodiments of the present invention in connection with applying a model using single-cell images to classify cytometry image data.
[0123] Flowchart 250 begins at step 252. In step 252, flow cytometry data, including cytometry image data, is received in the form of a computer file "stain.fcs." That is, the results of previously collected cytometry data are recorded in the file "stain.fcs," which is accessed in step 252. While any convenient format may be used to store the flow cytometry data, the illustrated embodiment uses the Flow Cytometry Standard (FCS) file format. The experimental results recorded in the file "stain.fcs" may be collected by one or more convenient flow cytometry systems capable of generating conventional flow cytometry data (e.g., scatter data) and cytometry image data, such as the flow cytometry systems and imaging flow cytometry systems described herein.
[0124] Once flow cytometry data acquisition is complete in step 252, flowchart 250 then proceeds to step 254.
[0125] In step 254, one or more gating methods are applied to each of the flow cytometry data, including the light scatter data (i.e., conventional flow cytometry data) and the cytometry image data (i.e., data collected by the imaging flow cytometer). Any convenient gating pattern or gating method may be applied. In the case of flowchart 250 for identifying singlets versus doublets, the gating methods considered in step 254 include methods intended to separate or split singlets from doublets. Different gating specifications may be applied to the conventional flow cytometry data and the imaging flow cytometry data received in step 252.
[0126] Applicable gating methods may be determined based on any convenient data analysis technique. By data analysis technique, we mean the study, understanding, and / or characterization of cytometry data. For example, when analyzing cytometry data, at least a first data analysis algorithm may be used to analyze such data. In other cases, when analyzing cytometry data, a first data analysis algorithm and one or more additional data analysis algorithms may be used to analyze such data. The data analysis algorithm may be any convenient or useful algorithm used to draw inferences from cytometry data. For example, in some embodiments, the data analysis algorithm may involve a data clustering algorithm, such as those described below. In other embodiments, the data analysis algorithm may include a dimensionality reduction algorithm, a feature extraction algorithm, a pattern recognition algorithm, etc., for example, to facilitate visualization of multidimensional data.
[0127] In some embodiments, the gating method may be further specified as a data analysis technique that clusters the cytometry data by applying a clustering algorithm to the cytometry data, where clustering refers to any algorithm, technique, or method used to identify subpopulations of data within the cytometry data, where each element of the subpopulation shares certain characteristics with each other element of the subpopulation.
[0128] Any convenient clustering algorithm may be applied to identify clusters within the cytometry data. In some cases, population clusters (and gates defining the limits of the populations) can be automatically identified and determined. Examples of automated gating methods are described, for example, in U.S. Pat. Nos. 4,845,653, 5,627,040, 5,739,000, 5,795,727, 5,962,238, 6,014,904, and 6,944,338, and U.S. Patent Application Publication No. 2012 / 0245889, each of which is incorporated herein by reference.
[0129] In some cases, a technique called k-means clustering is applied to assign particles of cytometry data to clusters. "k-means clustering" refers to a partitioning technique that divides the data points per event or per cell of a test sample into k clusters, with each data point belonging to the cluster with the closest mean. The technique of k-means clustering, including various general embodiments utilizing k-means clustering, is further described in LM Weber and MD Robinson, "Comparison of Clustering Methods for High-Dimensional Single-Cell Flow and Mass Cytometry Data," Cytometry, Part A, Journal of Quantitative Cell Science, Vol. 89, Issue 12, pp. 1084-96, the disclosure of which is incorporated herein by reference.
[0130] In other cases, self-organizing maps are applied to assign particles of cytometry data to clusters. "Self-organizing maps" refers to the application of a type of artificial neural network algorithm that generates a map as a result of a neural network training step, where the map includes a collection of clusters that define the data points or cells of the sample. FlowSOM, a technique for applying self-organizing maps, including a general embodiment of a self-organizing map, is further described in LM Weber and MD Robinson, "Comparison of Clustering Methods for High-Dimensional Single-Cell Flow and Mass Cytometry Data," Cytometry, Part A, Journal of Quantitative Cell Science, Vol. 89, Issue 12, pp. 1084-96, the disclosure of which is incorporated herein by reference. Other known or yet-to-be-discovered clustering techniques or algorithms may also be applied as desired.
[0131] In still other cases, techniques known in the art, such as the X-Shift ensemble search algorithm, are applied to assign cells from cytometry data to clusters. The X-Shift algorithm and its practical applications are further described in N. Samusik, Z. Good, MH Spitzer, KL Davis & GP Nolan (2016), "Automated mapping of phenotype space with single-cell data." Nature Methods, Vol. 13, Issue 6, p. 493, the disclosure of which is incorporated herein by reference. Other known or yet-to-be-discovered clustering techniques or algorithms may also be applied as desired.
[0132] Once the relevant gating method has been identified in step 254, flowchart 250 then proceeds to step 256.
[0133] Step 256 illustrates the results of applying the gating method identified in step 254 by applying geometric gating to each of the scatter and image data generated from the stain.fcs file accessed in step 252. For example, plot 256a illustrates the gating method identified in step 254 for conventional flow cytometry data (scatter data). In plot 256a, various aspects of forward scatter are plotted and gates are applied. Specifically, FSC-A (area of the forward scatter pulse) is plotted against FSC-H (height of the forward scatter pulse). Gate 256b is applied to plot 256a to identify and separate singlets from doublets on plot 256a.
[0134] Similarly, the gating method identified in step 254 for the cytometry image data is shown in plot 256m. In plot 256m, various aspects of the shape of the imaged cells are plotted and one or more gates are applied. In particular, a measure of radial moment (defined as the mean squared distance of the signal from the center of mass) is plotted against a measure of eccentricity (defined as the ratio of the magnitude of spread along the two principal components of the image). Gates 256n and 256o are applied to plot 256m to identify and separate singlets from doublets.
[0135] Once the gating method has been applied to the plot in step 256, the flowchart 250 then proceeds to step 258.
[0136] In step 258, the results of applying the gates in step 256 are processed to identify which flow cytometry data correspond to singlets and which flow cytometry data correspond to doublets. That is, in step 256, ground truth data, as described above, is established for flow cytometry data corresponding to singlets versus doublets based on the results of applying the gating method to plots 256a and 256m. Flow cytometry data corresponding to singlets is identified in data collection portion 258a, and flow cytometry data corresponding to doublets is identified in data collection portion 258b. Data collection portion 258a and data collection portion 258b are configured in the form of traditional flow cytometry data (i.e., scatter data obtained from plot 256a) and imaging flow cytometry data (i.e., image data obtained from plot 256m), respectively. In step 258, each data element (i.e., each instance of the cytometry image data) is labeled according to whether the data element was gated as a singlet or a doublet by the gating methods applied in steps 254 and 256. A separate label may be associated with each data element (i.e., each instance of the cytometry image data). The resulting dataset of labeled cytometry image data may be referred to as ground truth data as described herein.
[0137] Once ground truth data has been established for flow cytometry data corresponding to singlets versus doublets, flowchart 250 then proceeds to step 260 .
[0138] In connection with performing step 260, image files 262 of cytometry image data are accessed and retrieved. Image files 262 include a plurality of files, e.g., approximately 100,000 files or approximately 8 GB worth of image data. Image files 262 are in the form of .tiff image files, and the images may correspond to the cytometry image data for each flow cytometry event. That is, image files 262 may correspond to the cytometry image data for each singlet or doublet that may have been identified in connection with establishing ground truth data in step 258 (each instance of the cytometry image data includes multiple images, each corresponding to a different channel).
[0139] Continuing with step 260, optionally, an operator or data analyst may inspect, e.g., view, image data 262 to confirm the accuracy of the gating methods identified and applied in steps 254, 256, and 258 and establish ground truth data. An example of image data that may be inspected by an operator or analyst in step 260 is shown in image 260a. Image 260a is made up of multiple images, with each row corresponding to a different flow cytometry event (i.e., a detected cell, which may be either a singlet or a doublet) and each column corresponding to a different channel of the cytometry image data.
[0140] Image processing in step 260 then further involves applying one or more different filters (i.e., image adjustment techniques) to each channel. Step 260 processes eight different image channels corresponding to channels 1, 2...8 of cytometry image data image 262. Each channel, 1, 2...8, may correspond to a different collection of wavelengths of light or bright-field or dark-field images detected by a basic flow cytometer, for example.
[0141] Any convenient filter 260z may be applied to each channel at step 260. Exemplary filters shown associated with the various channels include "color" (i.e., adding, subtracting, or altering color in the image), "blur" (i.e., blurring aspects of the image), "sharpening" (i.e., sharpening aspects of the image), or "edge" (i.e., enhancing or de-emphasizing edges in the image). Various filters 260b, 260c, and 260d are applied to the various channels. Other related image adjustment techniques, such as those described above, may also be applied as part of applying different filters 260z to each channel. Applying the image adjustment techniques results in various adjusted cytometry image data 260e, 260f, and 260g. Images 260e, 260f, and 260g show the results of applying various filters to channels 260b, 260c, and 260d. Each exemplary image 260e, 260f, and 260g shows nine different cells. Image 260e shows channel 1 for nine cells and is a grayscale smoothed image. Image 260f shows channel 2 for the same nine cells in green. Image 260g shows channel 3 for these nine cells in blue.
[0142] With respect to the ground truth data described above in connection with steps 254, 256, and 258, all such image data associated with the ground truth data, e.g., images 260e, 260f, and 260g, are equally input to image processing step 260. In other words, the data labeling that occurs in connection with establishing the ground truth data is unrelated to the image processing applied in step 260. In some embodiments, it is important to apply the same image filter 260z to all input ground truth data in exactly the same way for consistent model training and model evaluation.
[0143] Once image processing is complete in step 260 , flowchart 250 then proceeds to step 264 .
[0144] In step 264, a model is applied to the adjusted image data to train the model. Any convenient model, as described above, may be applied. In step 264, for example, a neural network model is applied to adjusted cytometry image data 260e, 260f, and 260g generated by image processing in step 260. That is, as seen in steps 260 and 264, a model may be trained using multiple adjusted image data, such as adjusted cytometry image data 260e, 260f, and 260g. Inputs to the neural network model applied in step 264 may include such adjusted cytometry image data. Results of applying the neural network model in step 264 may include an estimate of whether such adjusted cytometry image data represents singlet cells or doublet cells. The neural network model applied in step 264 is trained to estimate the presence of singlets versus doublets using any convenient technique, as described above. That is, in step 264, the neural network model is trained to ultimately generate a fully trained model that can be used, for example, to predict unknown data, in step 266. The input to the neural network model in step 264 includes the ground truth data established in connection with steps 254, 256, and 260 (i.e., a set of images labeled as singlets and a set of images labeled as doublets), with such images being conditioned in step 260. The quality of the model is assessed by accuracy statistics determined during training, e.g., statistics describing whether the application of the model corresponds to the classifications established in connection with the generation of the ground truth data, such as via gating in steps 254, 256, and 258.
[0145] The output of the neural network model applied in step 264 further includes a confidence score indicating the confidence in the estimated output (i.e., the confidence in whether the singlet or doublet was correctly identified by the model). After estimating the presence or absence of singlet or doublet cells in the adjusted image data, the neural network applied in step 264 may be updated (i.e., further trained based on such results), e.g., by adjusting the weights of node connections in the model depending, in part, on the associated confidence score of the model's estimate of the presence of singlets versus doublets and / or depending on whether the model correctly predicted whether an instance in the cytometry image data corresponds to a gating result in the ground truth data. That is, the neural network training step 264 involves iterating over an input training set (e.g., the ground truth data established in connection with steps 254, 256, and 258) and adjusting the model weights as the model improves its ability to classify previously labeled data. More specifically, in embodiments, the model takes a subset of input training data (e.g., ground truth data established in connection with steps 254, 256, and 258), iterates over that subset, adjusts the model weights, and compares them to known classifications (i.e., classifications determined by the gating method described above). The model then repeats this training step using different subsets of training data (e.g., ground truth data established in connection with steps 254, 256, and 258), ultimately terminating with a final model. In embodiments, the iteration occurs within neural network training step 264 in connection with training the model. In the embodiment shown in FIG. 2B, the workflow of determining ground truth input data and applying image processing filters is substantially linear (i.e., adjusted image data is generated in advance).
[0146] Once the neural network model has been applied to train the model in step 264, the flowchart 250 then proceeds to step 266.
[0147] Flowchart 250 ends at step 266. Application of flowchart 250, upon completion at step 266, results in a model, i.e., neural network, being trained to distinguish between singlets and doublets using ground truth data as training data, and the ground truth data images being adjusted. In some cases, as a result of training the model with adjusted image data, such a neural network may be trained to identify additional cell populations not originally anticipated, for example, in connection with the identification of the gating method at step 254. In some embodiments, the model obtained at step 266 is considered complete, i.e., fully trained. The model generated at step 266 may be used for future classification (i.e., prediction) of new, as-yet-unseen data. In some cases, the model generated at step 266 may be used with new data sets, so long as the new data sets can be image-processed using the same filter 260z as used during generation of the model obtained at step 266. In other cases, the model generated in step 266 can be used on new data sets, and the cytometry image data present in such new data sets is iteratively adjusted in connection with iteratively applying the model to the cytometry image data.
[0148] 3A is a schematic illustration of an exemplary intermediate result of applying a method according to the present invention to cytometry image data. FIG. 3A shows exemplary cytometry image data generated using an embodiment of the present invention and is provided for illustrative purposes. FIG. 3A relates to cytometry image data of cells. However, embodiments of the present invention are not limited to cell classification. Any convenient particle present in a sample that can be imaged (e.g., by an imaging flow cytometer), including but not limited to cells, may be classified by applying embodiments of the present invention.
[0149] The exemplary cytometry image data 300 includes multiple images 310a, 310b, 310c, and 310d, each representing an image from a substantially identical field of view of cells 320a, 320b, 320c, and 320d, and corresponding to a different image channel of a basic flow cytometer. The cytometry image data 300 corresponds to a single flow cytometry event, such as the detection of a single cell (e.g., singlet vs. doublet) for classification into a first class of cells.
[0150] The unadjusted cytometry image data 300 may then be subjected to embodiment 350 of the present invention. Specifically, when applying embodiment 350 of the present invention, aspects of at least one of the images in the cytometry image data 300 are adjusted and a model is applied to the adjusted cytometry image data to classify the cytometry image data in an iterative manner. As described above, such a model is trained to estimate the presence of cells belonging to a first category of cells in the cytometry image data. The result of applying such an embodiment (i.e., the iterative adjustment of aspects of the cytometry image data 300 and application of the model, collectively 350) ultimately produces adjusted image data 360. The adjusted image data 360 corresponds to image data to which adjustments have been applied that result in a more accurate estimation by the model of whether the cytometry image data represents an image of a cell belonging to the first category of cells.
[0151] The adjusted image data 360 includes multiple images 370a, 370b, 370c, and 370d, each representing an image from substantially the same field of view of cells 380a, 380b, 380c, and 380d, each corresponding to a different image channel of the underlying flow cytometer, possibly adjusted differently or to a greater or lesser extent. As can be seen in the schematic representation of cells 380a, 380b, 380c, and 380d, the boundaries of cells 380a, 380b, 380c, and 380d are different from the boundaries shown for cells 320a, 320b, 320c, and 320d in unadjusted images 310a, 310b, 310c, and 310d. These differences in cell boundaries are intended to highlight the results of image adjustments applied by aspects of the model 350, which highlight or emphasize or clarify aspects of the cell image that are used to improve the model's estimation of whether the imaged cell belongs to a first category of cell. In some embodiments, obtaining final adjusted image data 360 for use in connection with a publication, or obtaining images for use to illustrate cellular aspects that serve as a basis for classification, is a useful aspect of the present invention.
[0152] 3B illustrates cytometry image data 390 showing one channel of image data from nine events, the images adjusted in accordance with an embodiment of the present invention. Cytometry image data 390 includes multiple image data corresponding to nine different events. For image data 390, a smoothing filter was applied to each individual image. In embodiments of the present invention, each image of cytometry image data, such as cytometry image data 390, may be individually and independently adjusted using any convenient image adjustment technique, and such techniques and iterative applications of such techniques are described above.
[0153] biological samples As described above, particles, e.g., cells, to be imaged and / or analyzed by a flow cytometer, e.g., an imaging flow cytometer, may be present in a sample. In some examples, the sample is a biological sample. The term "biological sample" is used in its conventional sense to refer to a whole organism, a plant, a fungus, or a subset of animal tissues, cells, or components, as may be found in blood, mucus, lymph, synovial fluid, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, amniotic fluid, amniotic cord blood, urine, vaginal fluid, or semen, as the case may be. Thus, a "biological sample" refers to both a naturally occurring organism or a subset of its tissues, and homogenates, lysates, or extracts prepared from an organism or a subset of its tissues, including, but not limited to, plasma, serum, spinal fluid, lymph, skin sections, respiratory tract, gastrointestinal tract, cardiovascular, genitourinary tract, tears, saliva, milk, blood cells, tumors, and organs. The biological sample may be tissue from any type of organism, including both healthy and diseased tissue (e.g., cancerous, malignant, necrotic, etc.). In certain embodiments, the biological sample is a liquid sample, such as blood or a derivative thereof, e.g., plasma, tears, urine, semen, etc.; in some cases, the sample is a blood sample, including whole blood, such as blood obtained from venipuncture or finger stick (which may or may not be combined with any reagents, such as preservatives, anticoagulants, etc., prior to assay).
[0154] In some embodiments, the sample source is a "mammal" or "mammalian," which terms are used broadly to describe organisms belonging to the class Mammalia, including the orders Carnivora (e.g., dogs and cats), Rodentia (e.g., mice, guinea pigs, and rats), and Primates (e.g., humans, chimpanzees, and monkeys). In some cases, the subject is a human. The methods may be applied to cytometry data for samples obtained from human subjects of both genders and at any developmental stage (i.e., newborn, infant, juvenile, adolescent, adult), and in certain embodiments, the human subject is a juvenile, adolescent, or adult. While the present invention may be applied to cytometry data for samples from human subjects, it should be understood that the methods may also be implemented with respect to cytometry data for samples from other animal subjects (i.e., "non-human subjects"), including, but not limited to, birds, mice, rats, dogs, cats, livestock, and horses.
[0155] Computer-Implemented Embodiments The steps of the various methods and algorithms described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, various exemplary steps have been described above generally in terms of functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system applying the methods of the present disclosure. The described functionality may be implemented in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0156] The various exemplary steps, elements, and computing systems (e.g., devices, databases, interfaces, and engines) described in connection with the embodiments disclosed herein may be implemented or performed by machines such as general-purpose processors, graphic processor units, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but may alternatively be a controller, microcontroller, state machine, or combination thereof, or the like. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. While described herein primarily with reference to digital technology, a processor may also include primarily analog components. The computing environment may include any type of computer system, including, but not limited to, microprocessor-based computer systems, graphic processor units, mainframe computers, digital signal processors, portable computing devices, personal organizers, device controllers, and computational engines within appliances, to name a few.
[0157] The steps of a method, process, or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or a combination of the two. The software modules, engines, and associated databases may reside in memory resources such as RAM memory, FRAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of non-transitory computer-readable storage medium, medium, or physical computer storage device known in the art. An external storage medium may be connected to the processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integrated with the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. Alternatively, the processor and the storage medium may reside as discrete components in a user terminal.
[0158] system As summarized above, aspects of the present disclosure include systems for implementing the subject methods. In one embodiment, the system includes a processor operatively coupled to a memory, the memory storing instructions that, when executed by the processor, cause the processor to iteratively receive unclassified cytometry image data including a plurality of images corresponding to an image channel, adjust aspects of at least one of the plurality of images of the cytometry image data, and apply a model to the adjusted cytometry image data to classify the cytometry image data, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data.
[0159] Systems according to some embodiments may include a display and an operator input device. The operator input device may be, for example, a keyboard, a mouse, etc. The processing module includes at least one general-purpose processor and multiple parallel processing units, all of which have access to memory in which instructions for performing the steps of the subject methods are stored. The processing module may include an operating system, a graphical user interface (GUI) controller, system memory, memory storage devices, input / output controllers, cache memory, a data backup unit, and many other devices. Each of the general-purpose processor and parallel processing units may be a commercially available processor or one of other processors that are or become available. The processor executes an operating system, which interfaces with firmware and hardware in a well-known manner to facilitate the processor's coordination and execution of functions of various computer programs, which may be written in a variety of programming languages, such as Java, Perl, Python, R, Go, JavaScript, .NET, CUDA, Verilog, C++, other high-level or low-level languages, and combinations thereof, as is known in the art. The operating system typically cooperates with the processor to coordinate and execute functions of the other elements of the computer. The operating system further provides scheduling, input / output control, file and data management, memory management, and communication control and related services, all in accordance with known techniques. The processor may be any suitable analog or digital system. In some embodiments, the one or more general-purpose processors and parallel processing units include analog electronics that provide feedback control, such as negative feedback control.
[0160] The system memory may be any of a variety of known or future memory storage devices. Examples include any commonly available random access memory (RAM), magnetic media such as a resident hard disk or tape, optical media such as a read-write compact disk, flash memory devices, or other memory storage devices. The memory storage device may be any of a variety of known or future devices, including a compact disk drive, tape drive, removable hard disk drive, or diskette drive. These types of memory storage devices typically read from and / or write to a program storage medium (not shown), such as a compact disk, magnetic tape, removable hard disk, or magnetic disk. Any of these program storage media, or other program storage media now in use or that may be developed in the future, may be considered a computer program product. As will be appreciated, these program storage media typically store computer software programs and / or data. Computer software programs, also referred to as computer control logic, are typically stored in the system memory and / or on program storage devices used in conjunction with the memory storage devices.
[0161] In some embodiments, a computer program product is described that includes a computer usable medium having stored thereon control logic (a computer software program including program code). The control logic, when executed by a processor of a computer, causes the processor to perform the functions described herein. In other embodiments, some functions are implemented primarily in hardware, for example, using hardware state machines. Implementation of hardware state machines to perform the functions described herein will be apparent to one skilled in the relevant art.
[0162] The memory may be any suitable device, such as a magnetic, optical, or solid-state storage device (including a magnetic or optical disk, or tape, or RAM, or any other suitable device, fixed or portable), from which one or more general-purpose processors and multiple parallel processing units, e.g., graphics processors, can store and retrieve data. The general-purpose processor may include a general-purpose digital microprocessor suitably programmed from a computer-readable medium storing the necessary program code. The parallel processing units may include one or more graphics processors suitably programmed from a computer-readable medium storing the necessary program code. The program may be provided to the processor remotely via one or more communication channels, or may be pre-recorded on a computer program product, such as a memory, or on another portable or fixed computer-readable storage medium using one of these devices connected to the memory. For example, a magnetic or optical disk may store the program and be read by a disk writer / reader. The system of the present invention further comprises a program, e.g., in the form of a computer program product, an algorithm for use in implementing the method as described above. The program of the present invention may be recorded on a computer-readable medium, e.g., any medium that can be directly read and accessed by a computer. Such media include, but are not limited to, magnetic storage media, such as magnetic disks, hard disk storage media, and magnetic tape; optical storage media, such as CD-ROMs; electrical storage media, such as RAM and ROM; portable flash drives; and hybrids of these categories, such as magnetic / optical storage media.
[0163] The one or more general-purpose processors may also access a communication channel to communicate with a user at a remote location, meaning that the user does not have direct contact with the system but instead relays input information to the input manager from an external device, such as a computer connected to a wide area network ("WAN"), telephone network, satellite network, or any other suitable communication channel, including a mobile phone (i.e., a smartphone).
[0164] In some embodiments, a system according to the present disclosure may be configured to include a communications interface. In some embodiments, the communications interface includes a receiver and / or a transmitter for communicating with a network and / or another device. The communications interface may be configured for wired or wireless communications, including, but not limited to, radio frequency (RF) communications such as radio frequency identification (RFID), ZigBee communications protocol, WiFi, infrared, wireless universal serial bus (USB), ultra-wideband (UWB), Bluetooth® communications protocol, and cellular communications such as code division multiple access (CDMA) or global system for mobile communications (GSM).
[0165] In one embodiment, the communication interface is configured to include one or more communication ports, e.g., physical ports or interfaces, such as a USB port, an RS-232 port, or any other suitable electrical connection port, to enable data communication between the subject system and other external devices, such as computer terminals (e.g., in a clinic or hospital environment) configured for similar complementary data communication.
[0166] In one embodiment, the communication interface is configured for infrared communication, Bluetooth® communication, or any other suitable wireless communication protocol, allowing the subject system to communicate with other devices, such as computer terminals and / or networks, communication-enabled mobile phones, personal digital assistants, or any other communication devices that a user may use in conjunction with the device.
[0167] In one embodiment, the communication interface is configured to provide a connection for data transfer utilizing Internet Protocol (IP) over a cellular network, short message service (SMS), a wireless connection to a personal computer (PC) on a local area network (LAN) connected to the Internet, or a WiFi connection to the Internet at a WiFi hotspot.
[0168] In one embodiment, the subject system is configured to communicate wirelessly with a server device via a communications interface using common standards such as 802.11 or Bluetooth® RF protocols or the IrDA infrared protocol. The server device may be another portable device, such as a smartphone, personal digital assistant (PDA), or notebook computer; or a larger device, such as a desktop computer, appliance, etc. In some embodiments, the server device has a display, such as a liquid crystal display (LCD), and input devices, such as buttons, a keyboard, a mouse, or a touchscreen.
[0169] In some embodiments, the communications interface is configured to automatically or semi-automatically communicate data stored in the subject system, e.g., any data storage unit, with a network or server device using one or more of the communications protocols and / or mechanisms described above.
[0170] The output controller may include a controller for any of a variety of known display devices for presenting information to a user, whether human or machine, local or remote. When one of the display devices provides visual information, this information may typically be logically and / or physically organized as an array of pixels. A graphical user interface (GUI) controller may include any of a variety of known or future software programs for providing a graphical input / output interface between the system and a user and for processing user input. The functional elements of the computer may communicate with each other via a system bus. Some of these communications may be achieved using a network or other type of remote communication in alternative embodiments. The output manager may also provide information generated by the processing module to a user at a remote location, for example, via the Internet, telephone, or satellite network, according to known techniques. Presentation of data by the output manager may be performed according to various known techniques. In some examples, the data may include SQL, HTML, or XML documents, emails or other files, or other forms of data. The data may include Internet URL addresses so that the user can obtain additional SQL, HTML, XML, or other documents or data from remote sources. The platform or platforms present in the subject system are typically a class of computers commonly referred to as servers, but may be any type of computer platform now known or later developed. However, the platforms may also be mainframe computers, workstations, or other computer types. The platforms may be networked or not, and may be connected via any type of cabling, now known or later, or other communication systems, including wireless systems. The platforms may be co-located or physically separated.Various operating systems may be used on any of the computer platforms, depending in some cases on the type and / or configuration of the computer platform selected. Suitable operating systems include Windows 10, Windows NT, Windows XP, Windows 7, Windows 8, iOS, Oracle Solaris, Linux, OS / 400, Compaq Tru64 Unix, SGI IRIX, Siemens Reliant Unix, Ubuntu, Zorin OS, and the like.
[0171] FIG. 4 illustrates a general configuration of an exemplary computing device 400 according to one embodiment. The general configuration of the computing device 400 illustrated in FIG. 4 includes the arrangement of computer hardware and software components. The computing device 400 may include more (or fewer) components than those illustrated in FIG. 4 . However, not all of these typical conventional components need be shown to provide a useful disclosure. As illustrated, the computing device 400 includes a processing unit 410, a network interface 420, a computer-readable medium drive 430, an input / output device interface 440, a display 450, and input devices 460, all of which may communicate with each other via a communications bus. The network interface 420 may provide connectivity to one or more networks or computing systems. Thus, the processing unit 410 may receive information and instructions from other computing systems or services via a network. The processing unit 410 may further communicate with a memory 470 and may further provide output information for an optional display 450 via the input / output device interface 440. The input / output device interface 440 may further accept input from any input device 460, such as a keyboard, a mouse, a digital pen, a microphone, a touch screen, a gesture recognition system, a voice recognition system, a gamepad, an accelerometer, a gyroscope, or other input device.
[0172] Memory 470 may include computer program instructions (grouped in some embodiments as modules or components) that processing unit 410 executes to implement one or more embodiments. Memory 470 typically includes RAM, ROM, and / or other persistent, secondary, or non-transitory computer-readable media. Memory 470 may store an operating system 472 that provides computer program instructions for use by processing unit 410 in the general management and operation of computing device 400. Memory 470 may further include computer program instructions and other information for implementing aspects of the present disclosure.
[0173] For example, in one embodiment, memory 470 includes an image processing module 474 for adjusting one or more aspects of the cytometry image data, and a model processing module 476 for applying a model to estimate the presence of particles belonging to a first category of particles in the cytometry image data.
[0174] computer-readable storage medium Aspects of the present disclosure further include non-transitory computer-readable storage media containing instructions for implementing the subject methods. The computer-readable storage medium may be used by one or more computers to fully or partially automate a system for implementing the methods described herein. In certain embodiments, instructions for the methods described herein may be encoded on a computer-readable medium in the form of a "program," and the term "computer-readable medium" as used herein refers to any non-transitory storage medium that participates in providing instructions and data to a computer for execution and processing. Examples of suitable non-transitory storage media include magnetic disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, DVD-ROMs, Blu-ray® disks, solid-state disks, and network-attached storage devices (NAS), regardless of whether such devices are internal or external to the computer. A file containing information may be "stored" on a computer-readable medium, where "storing" means recording information so that the information can be subsequently accessed and retrieved by a computer. The computer-implemented methods described herein may be performed using programs that may be written in one or more of any number of computer programming languages, including, for example, Java (Sun Microsystems, Inc., Santa Clara, CA), Visual Basic (Microsoft Corp., Redmond, WA), C++ (AT&T Corp., Bedminster, NJ), Python, and many others.
[0175] In some embodiments, a computer-readable storage medium of interest includes a computer program stored on the computer-readable storage medium, the computer program including instructions having an algorithm, when loaded into a computer, for receiving unclassified cytometry image data including a plurality of images corresponding to an image channel, and an algorithm for iteratively adjusting aspects of at least one of the plurality of images of the cytometry image data and applying a model to the adjusted cytometry image data to classify the cytometry image data, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data.
[0176] usefulness The subject systems, methods, and non-transitory computer-readable storage media find use in a variety of applications in which it is desirable to identify, analyze, or classify particles, such as cells, in a sample, such as a biological sample. In embodiments, the systems and methods described herein are used for imaging flow cytometry characterization of biological samples, including biological samples labeled with fluorescent tags. In addition, the subject systems and methods are used in analyzing or classifying particles of a sample, for example, by identifying categories or classifications of particles not previously known to be present in the sample. Furthermore, the subject systems and methods are used in analyzing or classifying particles of a sample based on additional information, such as cytometry image data corresponding to image features or characteristics across various image channels. As a result, in some cases, the subject systems and methods may be used in distinguishing between different particle types in a sample, for example, different cell types in a biological sample.
[0177] The following are given by way of example and not limitation.
[0178] experiment 5A-5H show excerpts from a computer-implemented embodiment of a method according to the present invention.
[0179] Figure 5A shows an excerpt of a user interface for such an embodiment, along with images of particle subpopulations. Figure 5B shows an excerpt of a user interface for creating image filter stacks for each channel and inspecting their aspects in a graph window using the mouse. Figure 5C shows an excerpt of a user interface for invoking a neural network on a training population.
[0180] FIG. 5D shows an excerpt from an alternative user interface for such an embodiment, along with an image of a particle subpopulation, illustrating how associations between events and image files can be specified. FIG. 5E shows an excerpt from such an alternative user interface for creating multiple image filters per channel and examining their aspects in a graph window using a mouse. FIG. 5F shows an excerpt from such an alternative user interface for selecting a training population, a neural network, and invoking a neural network on a training population by setting hyperparameters. FIG. 5G shows an excerpt from such an alternative user interface for selecting a population to test the accuracy of an existing model and classifying example images. FIG. 5H shows an excerpt from such an alternative user interface for viewing the results of classifying a population using a neural network model.
[0181] Regardless of the scope of the appended claims, the present disclosure is further defined by the following notes.
[0182] Appendix 1. A method for classifying cytometry image data, comprising: receiving unsorted cytometry image data including a plurality of images corresponding to image channels; iteratively adjusting an aspect of at least one image of the plurality of images of the cytometry image data and applying the model to the adjusted cytometry image data to classify the cytometry image data; A method in which the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data.
[0183] Appendix 2. The method of Appendix 1, wherein, when applying the model to the cytometry image data to classify the cytometry image data, an estimate of the presence of particles belonging to a first category of particles in the cytometry image data is obtained.
[0184] Appendix 3. The method of Appendix 2, wherein, upon applying the model to the cytometry image data, a confidence score is obtained associated with an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0185] Appendix 4. The method of Appendix 3, wherein adjusting aspects of at least one image among a plurality of images of unclassified cytometry image data and applying a model to the adjusted cytometry image data is performed iteratively, obtaining a plurality of estimates and confidence scores, each corresponding to an iteration of applying the model.
[0186] Appendix 5. The method of Appendix 3 or 4, wherein a higher confidence score corresponds to a higher likelihood that the estimate of the presence of particles belonging to the first category of particles in the cytometry image data is accurate.
[0187] Appendix 6. The method of any one of Appendixes 3-5, wherein a lower confidence score corresponds to a lower likelihood that the estimate of the presence of particles belonging to the first category of particles in the cytometry image data is accurate.
[0188] Appendix 7. Iteratively adjusting at least one aspect of the image and applying a model to the adjusted cytometry image data, applying an image adjustment process to at least one image of the plurality of images of the cytometry image data; applying the model to the unclassified cytometry image data including at least one adjusted image; As a result of applying the model, we obtain a confidence score, Compare the confidence score to a baseline confidence score The method according to any one of appendices 3 to 6, wherein the above steps are repeatedly performed.
[0189] Clause 8. The method of Clause 7, wherein the baseline confidence score comprises a confidence score when the model is applied to unclassified cytometry image data without any image adjustments.
[0190] Appendix 9. The method of appendix 7 or 8, wherein subsequent image adjustment processing is determined based on a comparison of the confidence score with a reference confidence score.
[0191] Appendix 10. The method of any one of Appendixes 3-9, wherein adjusting at least one image aspect and applying a model to the adjusted cytometry image data is iteratively repeated to refine an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0192] Appendix 11. The method of Appendix 10, wherein adjusting at least one image aspect and applying a model to the adjusted cytometry image data is iteratively repeated to increase a confidence score associated with each estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0193] Appendix 12. The method of Appendix 11, wherein, during the iterative process of adjusting at least one image aspect and applying the model to the adjusted cytometry image data, the method adjusts at least one image aspect and optimizes a confidence score associated with each estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0194] Appendix 13. The method of Appendix 12, wherein when optimizing the confidence score associated with each estimate of the presence of particles belonging to a first category of particles in the cytometry image data, a local maximum of the confidence score is found.
[0195] Appendix 14. The method of any one of Appendixes 1-13, wherein the method iteratively adjusts multiple aspects of a first image of the multiple images of cytometry image data and applies a model to the adjusted cytometry image data.
[0196] Appendix 15. The method of any one of Appendixes 1-14, wherein the method iteratively comprises adjusting a first aspect of two or more images of the plurality of images of cytometry image data and applying a model to the adjusted cytometry image data.
[0197] Appendix 16. The method of any one of Appendixes 1-15, wherein the iterative adjusting of at least one image aspect and applying the model to the adjusted cytometry image data involves applying a simulated annealing technique.
[0198] Appendix 17. The method of any one of Appendixes 1-16, wherein the iterative adjustment of at least one image aspect and applying the model to the adjusted cytometry image data comprises applying a predetermined set of image filters to the image data.
[0199] Appendix 18. The method of any one of Appendixes 1-17, wherein the iterative adjustment of at least one image aspect and application of a model to the adjusted cytometry image data first applies a non-configurable image filter to the image data.
[0200] Appendix 19. The method of Appendix 18, wherein, when iteratively adjusting at least one image aspect and applying a model to the adjusted cytometry image data, a configurable image filter is then applied to the image data.
[0201] Addendum 20. The method of Addendum 19, wherein input parameters are provided to the image filter when the image filter is configured.
[0202] Addendum 21. The method of Addendum 20, wherein when applying the configurable image filter, input parameters are provided to the configurable image filter in regular increments.
[0203] Addendum 22. The method of any one of Addendums 1-21, identifying adjusted cytometry image data corresponding to an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0204] Appendix 23. The method of any one of Appendixes 1 to 22, wherein adjusting an aspect of the at least one image comprises applying digital signal processing techniques to the at least one image.
[0205] Addendum 24. The method of any one of Addendums 1 to 23, wherein adjusting an aspect of the at least one image comprises applying a noise reduction technique to the at least one image.
[0206] Addendum 25. The method of any one of Addendums 1 to 24, wherein adjusting at least one aspect of the image comprises applying one or more of the following techniques to the at least one image: adding color, removing color, blurring the image, sharpening the image, or defining edges.
[0207] Addendum 26. The method of any one of Addendums 1 to 25, wherein adjusting an aspect of the at least one image comprises applying an image filter to the at least one image.
[0208] Clause 27. The method of clause 26, wherein the image filter comprises one or more of a linear filter, a nonlinear filter, a low-pass filter, a high-pass filter, a spatial filter, a frequency filter, a threshold filter, an inverse filter, or an intensity filter.
[0209] Addendum 28. The method of Addendum 26, wherein the image filter comprises one or more of a smoothing filter, a box filter, a Gaussian filter, or a Laplacian filter.
[0210] Addendum 29. The method of any one of Addendums 1 to 28, wherein the model is updated based on the results of iteratively applying the model to the adjusted cytometry image data.
[0211] Appendix 30. The method of Appendix 29, wherein when updating the model, the model is further trained to classify cytometry image data.
[0212] Appendix 31. The method of any one of Appendixes 1 to 30, wherein the model comprises a statistical model.
[0213] Appendix 32. The method of any one of Appendixes 1 to 31, wherein the model comprises a linear model.
[0214] Addendum 33. The method of any one of Addendums 1 to 32, wherein the model comprises a computational model.
[0215] Addendum 34. The method of any one of Addendums 1 to 33, wherein the model comprises a machine learning model.
[0216] Addendum 35. The method of Addendum 34, wherein the machine learning model comprises a tree-based model.
[0217] Appendix 36. The method of Appendix 34, wherein the model comprises an artificial neural network.
[0218] Clause 37. The method of clause 36, wherein the artificial neural network comprises a convolutional neural network.
[0219] Addendum 38. The method of Addendum 36, wherein the artificial neural network comprises a deep learning model.
[0220] Addendum 39. The method of any one of Addendums 1-38, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data using one or more of unsupervised learning techniques, semi-supervised learning techniques, supervised learning techniques, or round robin training techniques.
[0221] Appendix 40. The method of any one of appendices 1-39, wherein the cytometry image data comprises data collected by applying an imaging flow cytometer to a sample containing the particles.
[0222] Appendix 41. The method of any one of Appendixes 1-40, wherein an image channel of the cytometry image data includes multiple images, each corresponding to a particular wavelength range of light.
[0223] Appendix 42. The method of any one of Appendixes 1 to 41, wherein an image channel of the cytometry image data includes multiple images each corresponding to the same field of view.
[0224] Appendix 43. The method of any one of appendices 1 to 42, wherein aspects of the cytometry image data are normalized.
[0225] Addendum 44. The method of Addendum 43, wherein normalizing the aspect of the cytometry image data includes one or more of centering the cytometry image data, adjusting the aspect ratio of the cytometry image data, or rotating the cytometry image data.
[0226] Addendum 45. The method of any one of Addendums 1 to 44, wherein an image channel of the cytometry image data includes multiple images, each corresponding to a particular property of the particle.
[0227] Addendum 46. The method of Addendum 45, wherein an image channel of the cytometry image data includes multiple images, each corresponding to a particular range of fluorescence.
[0228] Appendix 47. The method of any one of appendices 1 to 46, wherein the particles are cells.
[0229] Item 48. The method of item 47, wherein the first category of particles is a cell type.
[0230] Appendix 49. The method of any one of appendices 1 to 46, wherein the particle is a singlet or a doublet.
[0231] Item 50. The method of item 49, wherein the first category of particles is singlets.
[0232] Appendix 51. The method of Appendix 49, wherein the first category of particles is doublets.
[0233] Appendix 52. The method of any one of appendices 1 to 46, which is a method for classifying intercellular activity.
[0234] Item 53. The method of item 52, wherein the first category of particles comprises cells that exhibit intercellular activity.
[0235] Addendum 54. The method of any one of Addendums 1 to 46, wherein the first category of particles comprises cells exhibiting intracellular activity.
[0236] Item 55. The method of item 54, wherein the first category of particles comprises cells exhibiting intracellular activity.
[0237] Clause 56. A method of training a model to classify cytometry image data, comprising: receiving flow cytometry data including unsorted cytometry image data including a plurality of images, each instance of the cytometry image data corresponding to an image channel; classifying each instance of the cytometry image data of the flow cytometry data to establish ground truth data; A method of training a model to classify cytometry image data as data including particles belonging to a first category of particles by adjusting aspects of at least one image among a plurality of images of each instance of the cytometry image data for ground truth data.
[0238] Addendum 57. The method of Addendum 56, wherein a gating algorithm, or a direct clustering algorithm, or a dimensionality reduction algorithm followed by a clustering algorithm is applied when classifying each instance of cytometry image data of flow cytometry data to establish ground truth data.
[0239] Addendum 58. The method of Addendum 56 or 57, wherein supervised learning based on ground truth data is applied when training a model to classify cytometry image data.
[0240] Addendum 59. The method of any one of Addendums 56 to 58, wherein the ground truth data includes, for each instance of cytometry image data in the ground truth data, a label indicating whether the cytometry image data includes particles belonging to a first category of particles.
[0241] Appendix 60. A system for classifying cytometry image data, comprising: a processor operatively coupled to a memory; The memory, when executed by the processor, causes the processor to receiving unsorted cytometry image data including a plurality of images corresponding to image channels; Iteratively adjusting an aspect of at least one of the plurality of images of the cytometry image data and applying the model to the adjusted cytometry image data to classify the cytometry image data. It remembers the commands, The system, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data.
[0242] Addendum 61. The system of Addendum 60, wherein, when applying the model to the cytometry image data to classify the cytometry image data, an estimate of the presence of particles belonging to a first category of particles in the cytometry image data is obtained.
[0243] Addendum 62. The system of Addendum 61, wherein, upon applying the model to the cytometry image data, a confidence score associated with an estimate of the presence of particles belonging to the first category of particles in the cytometry image data is obtained.
[0244] Addendum 63. The system of Addendum 62, wherein adjusting aspects of at least one image among a plurality of images of unclassified cytometry image data and applying a model to the adjusted cytometry image data in an iterative manner obtains a plurality of estimates and confidence scores, each corresponding to an iteration of applying the model.
[0245] Addendum 64. The system of Addendum 62 or 63, wherein a higher confidence score corresponds to a higher likelihood that the estimate of the presence of particles belonging to the first category of particles in the cytometry image data is accurate.
[0246] Addendum 65. The system of any one of Addendums 62-64, wherein a lower confidence score corresponds to a lower likelihood that the estimate of the presence of particles belonging to the first category of particles in the cytometry image data is accurate.
[0247] Addendum 66. Iteratively adjusting at least one image aspect and applying a model to the adjusted cytometry image data, applying an image adjustment process to at least one image of the plurality of images of the cytometry image data; applying the model to the unclassified cytometry image data including at least one adjusted image; As a result of applying the model, we obtain a confidence score, Compare the confidence score to a baseline confidence score The system of any one of appendices 62 to 65, wherein the above steps are repeated.
[0248] Addendum 67. The system of Addendum 66, wherein the baseline confidence score includes a confidence score when the model is applied to unclassified cytometry image data without any image adjustments.
[0249] Addendum 68. The system of Addendum 66 or 67, wherein subsequent image adjustment processing is determined based on the result of comparing the confidence score to a reference confidence score.
[0250] Addendum 69. The system of any one of Addendums 62-68, wherein adjusting at least one image aspect and applying a model to the adjusted cytometry image data is iteratively performed to refine an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0251] Addendum 70. The system of Addendum 69, wherein, when iteratively adjusting at least one image aspect and applying a model to the adjusted cytometry image data, iteratively iterates to increase a confidence score associated with each estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0252] Addendum 71. The system of Addendum 70, wherein, when iteratively adjusting at least one image aspect and applying a model to the adjusted cytometry image data, the system adjusts at least one image aspect and optimizes a confidence score associated with each estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0253] Addendum 72. The system of Addendum 71, wherein when optimizing the confidence score associated with each estimate of the presence of particles belonging to a first category of particles in the cytometry image data, a local maximum of the confidence score is found.
[0254] Addendum 73. The system of any one of Addendums 60-72, wherein iteratively adjusting aspects of at least one image includes iteratively adjusting aspects of a first image of the plurality of images of cytometry image data and applying a model to the adjusted cytometry image data.
[0255] Addendum 74. The system of any one of Addendums 60-73, wherein iteratively adjusting an aspect of at least one image includes iteratively adjusting a first aspect of two or more images of the plurality of images of the cytometry image data and applying a model to the adjusted cytometry image data.
[0256] Addendum 75. The system of any one of Addendums 60-74, wherein the iterative adjustment of at least one image aspect and application of the model to the adjusted cytometry image data applies a simulated annealing technique.
[0257] Addendum 76. The system of any one of Addendums 60-75, wherein the system applies a predetermined set of image filters to the image data when iteratively adjusting at least one image aspect and applying a model to the adjusted cytometry image data.
[0258] Addendum 77. A system described in any one of Addendums 60 to 76, wherein when iteratively adjusting at least one image aspect and applying a model to the adjusted cytometry image data, a non-configurable image filter is first applied to the image data.
[0259] Addendum 78. The system of Addendum 77, wherein, when iteratively adjusting at least one aspect of the image and applying a model to the adjusted cytometry image data, a configurable image filter is then applied to the image data.
[0260] Addendum 79. The system of Addendum 78, wherein when configuring the image filter, input parameters are provided to the image filter.
[0261] Addendum 80. The system of Addendum 79, wherein when applying the configurable image filter, the input parameters are provided to the configurable image filter in fixed increments.
[0262] Addendum 81. The system of any one of Addendums 60-80, which identifies adjusted cytometry image data corresponding to an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0263] Addendum 82. The system of any one of Addendums 60-81, wherein adjusting an aspect of the at least one image involves applying digital signal processing techniques to the at least one image.
[0264] Addendum 83. The system of any one of Addendums 60-82, wherein adjusting an aspect of the at least one image involves applying a noise reduction technique to the at least one image.
[0265] Addendum 84. The system of any one of Addendums 60-83, wherein adjusting at least one aspect of the image applies one or more of the following techniques to the at least one image: adding color, removing color, blurring the image, sharpening the image, or defining edges.
[0266] Addendum 85. The system of any one of Addendums 60-84, wherein adjusting an aspect of the at least one image comprises applying an image filter to the at least one image.
[0267] Clause 86. The system of clause 85, wherein the image filter comprises one or more of a linear filter, a nonlinear filter, a low-pass filter, a high-pass filter, a spatial filter, a frequency filter, a threshold filter, an inverse filter, or an intensity filter.
[0268] Clause 87. The system of Clause 85, wherein the image filter includes one or more of a smoothing filter, a box filter, a Gaussian filter, or a Laplacian filter.
[0269] Addendum 88. A system described in any one of Addendums 60 to 87, wherein the model is updated based on the results of iteratively applying the model to the adjusted cytometry image data.
[0270] Addendum 89. The system of Addendum 88, wherein when updating the model, the model is further trained to classify the cytometry image data.
[0271] Addendum 90. The system of any one of Addendums 60 to 89, wherein the model includes a statistical model.
[0272] Addendum 91. The system of any one of Addendums 60 to 90, wherein the model includes a linear model.
[0273] Addendum 92. The system of any one of Addendums 60-91, wherein the model includes a computational model.
[0274] Addendum 93. The system of any one of Addendums 60-92, wherein the model includes a machine learning model.
[0275] Addendum 94. The system of Addendum 93, wherein the machine learning model includes a tree-based model.
[0276] Addendum 95. The system of Addendum 93, wherein the model includes an artificial neural network.
[0277] Clause 96. The system of Clause 95, wherein the artificial neural network comprises a convolutional neural network.
[0278] Addendum 97. The system of Addendum 95, wherein the artificial neural network includes a deep learning model.
[0279] Addendum 98. The system of any one of Addendums 60-97, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data using one or more of unsupervised learning techniques, semi-supervised learning techniques, supervised learning techniques, or round robin training techniques.
[0280] Addendum 99. The system of any one of Addendums 60-98, wherein the cytometry image data comprises data collected by applying an imaging flow cytometer to a sample containing particles.
[0281] Addendum 100. The system of any one of Addendums 60-99, wherein an image channel of the cytometry image data includes multiple images, each corresponding to a particular wavelength range of light.
[0282] Addendum 101. The system of any one of Addendums 60-100, wherein an image channel of the cytometry image data includes multiple images each corresponding to the same field of view.
[0283] Addendum 102. A system according to any one of Addendums 60-101, which normalizes aspects of cytometry image data.
[0284] Addendum 103. The system of Addendum 102, wherein when normalizing the aspects of the cytometry image data, one or more of centering the cytometry image data, adjusting the aspect ratio of the cytometry image data, or rotating the cytometry image data are performed.
[0285] Addendum 104. A system according to any one of Addendums 60-103, wherein an image channel of the cytometry image data includes multiple images, each corresponding to a particular characteristic of the particle.
[0286] Addendum 105. The system of Addendum 104, wherein an image channel of the cytometry image data includes multiple images, each corresponding to a particular range of fluorescence.
[0287] Addendum 106. A system described in any one of Addendums 60 to 105, wherein the particles are cells.
[0288] Addendum 107. The system of Addendum 106, wherein the first category of particles is a cell type.
[0289] Addendum 108. A system described in any one of Addendums 60 to 105, wherein the particle is a singlet or a doublet.
[0290] Addendum 109. The system of Addendum 108, wherein the first category of particles is a singlet.
[0291] Addendum 110. The system of Addendum 108, wherein the first category of particles is a doublet.
[0292] Appendix 111. The system of any one of appendices 60-105, which is a system for classifying intercellular activity.
[0293] Addendum 112. The system of Addendum 111, wherein the first category of particles includes cells that exhibit intercellular activity.
[0294] Addendum 113. The system of any one of Addendums 60-105, wherein the first category of particles comprises cells exhibiting intracellular activity.
[0295] Addendum 114. The system of Addendum 113, wherein the first category of particles includes cells exhibiting intracellular activity.
[0296] Clause 115. A system for training a model to classify cytometry image data, comprising: a processor operatively coupled to a memory; The memory, when executed by the processor, causes the processor to receiving flow cytometry data including unsorted cytometry image data including a plurality of images, each instance of the cytometry image data corresponding to an image channel; classifying each instance of the cytometry image data of the flow cytometry data to establish ground truth data; and training the model to classify the cytometry image data as data including particles belonging to a first category of particles by adjusting an aspect of at least one image of the plurality of images of each instance of the cytometry image data for the ground truth data. A system that remembers instructions.
[0297] Addendum 116. The system of Addendum 115, wherein a gating algorithm, or a direct clustering algorithm, or a dimensionality reduction algorithm followed by a clustering algorithm is applied when classifying each instance of cytometry image data of flow cytometry data to establish ground truth data.
[0298] Addendum 117. The system of Addendum 115 or 116, wherein supervised learning based on ground truth data is applied when training a model to classify cytometry image data.
[0299] Addendum 118. The system of any one of Addendums 115-117, wherein the ground truth data includes, for each instance of cytometry image data in the ground truth data, a label indicating whether the cytometry image data includes particles belonging to a first category of particles.
[0300] Clause 119. A non-transitory computer-readable storage medium storing instructions for classifying cytometry image data, the instructions comprising: an algorithm for receiving unsorted cytometry image data including a plurality of images corresponding to image channels; and An algorithm for iteratively adjusting aspects of at least one image of a plurality of images of cytometry image data and applying a model to the adjusted cytometry image data to classify the cytometry image data. Including, A non-transitory computer-readable storage medium in which the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data.
[0301] Appendix 120. The non-transitory computer-readable storage medium of Appendix 119, wherein, when applying the model to the cytometry image data to classify the cytometry image data, an estimate of the presence of particles belonging to a first category of particles in the cytometry image data is obtained.
[0302] Appendix 121. The non-transitory computer-readable storage medium of Appendix 120, wherein, upon applying the model to the cytometry image data, a confidence score associated with an estimate of the presence of particles belonging to a first category of particles in the cytometry image data is obtained.
[0303] Addendum 122. The non-transitory computer-readable storage medium of Addendum 121, wherein adjusting aspects of at least one image among a plurality of images of unclassified cytometry image data and applying a model to the adjusted cytometry image data in an iterative manner results in obtaining a plurality of estimates and confidence scores, each of which corresponds to an iteration of applying the model.
[0304] Addendum 123. The non-transitory computer-readable storage medium of Addendum 121 or 122, wherein a higher confidence score corresponds to a higher likelihood that the estimate of the presence of particles belonging to the first category of particles in the cytometry image data is accurate.
[0305] Addendum 124. The non-transitory computer-readable storage medium of any one of Addendums 121-123, wherein a lower confidence score corresponds to a lower likelihood that the estimate of the presence of particles belonging to the first category of particles in the cytometry image data is accurate.
[0306] Addendum 125. Iteratively adjusting at least one image aspect and applying a model to the adjusted cytometry image data, applying an image adjustment process to at least one image of the plurality of images of the cytometry image data; applying the model to the unclassified cytometry image data including at least one adjustment image; As a result of applying the model, we obtain a confidence score, Compare the confidence score to a baseline confidence score A non-transitory computer-readable storage medium according to any one of appendices 121 to 124, which repeatedly performs the above.
[0307] Appendix 126. The non-transitory computer-readable storage medium of Appendix 125, wherein the baseline confidence score comprises a confidence score when the model is applied to unclassified cytometry image data without any image adjustments.
[0308] Appendix 127. The non-transitory computer-readable storage medium of Appendix 125 or 126, wherein subsequent image adjustment processing is determined based on a result of comparing the confidence score to a reference confidence score.
[0309] Addendum 128. A non-transitory computer-readable storage medium according to any one of Addendums 121 to 127, wherein adjusting at least one image aspect and applying a model to the adjusted cytometry image data is iteratively repeated to refine an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0310] Appendix 129. The non-transitory computer-readable storage medium of Appendix 128, wherein adjusting at least one image aspect and applying a model to the adjusted cytometry image data is iteratively repeated to increase a confidence score associated with each estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0311] Appendix 130. The non-transitory computer-readable storage medium of Appendix 129, wherein, during iterative adjusting of at least one image aspect and applying a model to the adjusted cytometry image data, the at least one image aspect is adjusted and a confidence score associated with each estimate of the presence of particles belonging to a first category of particles in the cytometry image data is optimized.
[0312] Appendix 131. The non-transitory computer-readable storage medium of Appendix 130, wherein optimizing the confidence score associated with each estimate of the presence of particles belonging to a first category of particles in the cytometry image data involves finding a local maximum in the confidence score.
[0313] Addendum 132. A non-transitory computer-readable storage medium described in any one of Addendums 119 to 131, wherein the algorithm for iteratively adjusting aspects of at least one image includes iteratively adjusting multiple aspects of a first image of multiple images of cytometry image data and applying a model to the adjusted cytometry image data.
[0314] Addendum 133. A non-transitory computer-readable storage medium according to any one of Addendums 119-132, wherein the algorithm for iteratively adjusting an aspect of at least one image includes iteratively adjusting a first aspect of two or more images of a plurality of images of cytometry image data and applying a model to the adjusted cytometry image data.
[0315] Addendum 134. A non-transitory computer-readable storage medium according to any one of Addendums 119 to 133, wherein the iterative adjustment of at least one image aspect and application of the model to the adjusted cytometry image data applies a simulated annealing technique.
[0316] Addendum 135. A non-transitory computer-readable storage medium according to any one of Addendums 119 to 134, wherein adjusting at least one image aspect and applying a model to the adjusted cytometry image data in an iterative manner includes applying a predetermined set of image filters to the image data.
[0317] Addendum 136. A non-transitory computer-readable storage medium described in any one of Addendums 119 to 135, wherein the iterative adjustment of at least one image aspect and application of a model to the adjusted cytometry image data first applies a non-configurable image filter to the image data.
[0318] Addendum 137. The non-transitory computer-readable storage medium of Addendum 136, wherein adjusting at least one image aspect and applying a model to the adjusted cytometry image data in an iterative manner includes subsequently applying a configurable image filter to the image data.
[0319] Clause 138. The non-transitory computer-readable storage medium of Clause 137, for providing input parameters to an image filter when configuring the image filter.
[0320] Clause 139. The non-transitory computer-readable storage medium of Clause 138, wherein when applying the configurable image filter, the input parameters are provided to the configurable image filter in regular increments.
[0321] Addendum 140. A non-transitory computer-readable storage medium according to any one of Addendums 119 to 139, which identifies adjusted cytometry image data corresponding to an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
[0322] Addendum 141. The non-transitory computer-readable storage medium of any one of Addendums 119-140, wherein adjusting at least one aspect of the image involves applying digital signal processing techniques to at least one image.
[0323] Addendum 142. The non-transitory computer-readable storage medium of any one of Addendums 119-141, wherein adjusting an aspect of the at least one image comprises applying a noise reduction technique to the at least one image.
[0324] Addendum 143. The non-transitory computer-readable storage medium of any one of Addendums 119-142, wherein adjusting at least one aspect of the image involves applying one or more of the following techniques to the at least one image: adding color, removing color, blurring the image, sharpening the image, or defining edges.
[0325] Addendum 144. The non-transitory computer-readable storage medium of any one of Addendums 119 to 143, wherein adjusting at least one aspect of the image comprises applying an image filter to at least one image.
[0326] Addendum 145. The non-transitory computer-readable storage medium of Addendum 144, wherein the image filter comprises one or more of a linear filter, a non-linear filter, a low-pass filter, a high-pass filter, a spatial filter, a frequency filter, a threshold filter, an inverse filter, or an intensity filter.
[0327] Clause 146. The non-transitory computer-readable storage medium of Clause 144, wherein the image filter includes one or more of a smoothing filter, a box filter, a Gaussian filter, or a Laplacian filter.
[0328] Addendum 147. A non-transitory computer-readable storage medium according to any one of Addendums 119 to 146, wherein the model is updated based on the results of iteratively applying the model to the adjusted cytometry image data.
[0329] Clause 148. The non-transitory computer-readable storage medium of Clause 147, wherein updating the model further trains the model to classify cytometry image data.
[0330] Addendum 149. The non-transitory computer-readable storage medium of any one of Addendums 119 to 148, wherein the model comprises a statistical model.
[0331] Addendum 150. The non-transitory computer-readable storage medium of any one of Addendums 119 to 149, wherein the model includes a linear model.
[0332] Addendum 151. The non-transitory computer-readable storage medium of any one of Addendums 119-150, wherein the model includes a computational model.
[0333] Addendum 152. The non-transitory computer-readable storage medium of any one of Addendums 119 to 151, wherein the model comprises a machine learning model.
[0334] Clause 153. The non-transitory computer-readable storage medium of Clause 152, wherein the machine learning model comprises a tree-based model.
[0335] Clause 154. The non-transitory computer-readable storage medium of Clause 152, wherein the model comprises an artificial neural network.
[0336] Clause 155. The non-transitory computer-readable storage medium of Clause 154, wherein the artificial neural network comprises a convolutional neural network.
[0337] Clause 156. The non-transitory computer-readable storage medium of Clause 154, wherein the artificial neural network includes a deep learning model.
[0338] Addendum 157. A non-transitory computer-readable storage medium according to any one of Addendums 119 to 156, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data using one or more of unsupervised learning techniques, semi-supervised learning techniques, supervised learning techniques, or round-robin training techniques.
[0339] Addendum 158. The non-transitory computer-readable storage medium of any one of Addendums 119 to 157, wherein the cytometry image data comprises data collected by applying an imaging flow cytometer to a sample containing particles.
[0340] Addendum 159. A non-transitory computer-readable storage medium according to any one of Addendums 119 to 158, wherein an image channel of the cytometry image data includes multiple images, each corresponding to a particular wavelength range of light.
[0341] Addendum 160. A non-transitory computer-readable storage medium according to any one of Addendums 119 to 159, wherein an image channel of the cytometry image data includes multiple images each corresponding to the same field of view.
[0342] Addendum 161. A non-transitory computer-readable storage medium according to any one of Addendums 119 to 160, for normalizing aspects of cytometry image data.
[0343] Addendum 162. The non-transitory computer-readable storage medium of Addendum 161, wherein normalizing aspects of the cytometry image data includes one or more of centering the cytometry image data, adjusting the aspect ratio of the cytometry image data, or rotating the cytometry image data.
[0344] Addendum 163. A non-transitory computer-readable storage medium according to any one of Addendums 119 to 162, wherein an image channel of the cytometry image data includes multiple images, each corresponding to a particular characteristic of a particle.
[0345] Addendum 164. The non-transitory computer-readable storage medium of Addendum 163, wherein an image channel of the cytometry image data includes multiple images, each corresponding to a particular range of fluorescence.
[0346] Addendum 165. The non-transitory computer-readable storage medium of any one of Addendums 119 to 164, wherein the particle is a cell.
[0347] Clause 166. The non-transitory computer-readable storage medium of Clause 165, wherein the first category of particles is a cell type.
[0348] Addendum 167. The non-transitory computer-readable storage medium of any one of Addendums 119 to 164, wherein the particle is a singlet or a doublet.
[0349] Clause 168. The non-transitory computer-readable storage medium of Clause 167, wherein the first category of particles is a singlet.
[0350] Clause 169. The non-transitory computer-readable storage medium of Clause 167, wherein the first category of particles is a doublet.
[0351] Addendum 170. The non-transitory computer-readable storage medium of any one of Addendums 119-164, wherein the instructions stored on the non-transitory computer-readable storage medium are instructions for classifying intercellular activity.
[0352] Clause 171. The non-transitory computer-readable storage medium of Clause 170, wherein the first category of particles includes cells exhibiting intercellular activity.
[0353] Addendum 172. The non-transitory computer-readable storage medium of any one of Addendums 119-164, wherein the first category of particles includes cells exhibiting intracellular activity.
[0354] Clause 173. The non-transitory computer-readable storage medium of Clause 172, wherein the first category of particles includes cells exhibiting intracellular activity.
[0355] Clause 174. A non-transitory computer-readable storage medium storing instructions for training a model to classify cytometry image data, the instructions comprising: an algorithm for receiving flow cytometry data comprising unsorted cytometry image data comprising a plurality of images, each instance of the cytometry image data corresponding to an image channel; an algorithm for classifying each instance of the cytometry image data of the flow cytometry data to establish ground truth data; and and an algorithm for training a model to classify cytometry image data as data including particles belonging to a first category of particles by adjusting an aspect of at least one image among a plurality of images of each instance of the cytometry image data for ground truth data. 1. A non-transitory computer-readable storage medium comprising:
[0356] Clause 175. The non-transitory computer-readable storage medium of Clause 174, wherein a gating algorithm, or a direct clustering algorithm, or a dimensionality reduction algorithm followed by a clustering algorithm is applied when classifying each instance of cytometry image data of flow cytometry data to establish ground truth data.
[0357] Addendum 176. The non-transitory computer-readable storage medium of Addendum 174 or 175, wherein supervised learning based on ground truth data is applied when training a model to classify cytometry image data.
[0358] Addendum 177. The non-transitory computer-readable storage medium of any one of Addendums 174 to 176, wherein the ground truth data includes, for each instance of cytometry image data in the ground truth data, a label indicating whether the cytometry image data includes particles belonging to a first category of particles.
[0359] Although the foregoing invention has been described in some detail by way of illustration and example, for purposes of clarity of understanding, it will be readily apparent to those skilled in the art that, in light of the teachings of the invention, certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.
[0360] Accordingly, the foregoing merely illustrates the essence of the present invention. It is clear that those skilled in the art will be able to devise various configurations that embody the essence of the present invention and are within the spirit and scope of the present invention, although not explicitly described or shown herein. Furthermore, all examples and conditional language set forth herein are intended essentially to aid the reader in understanding the essence of the present invention and the concepts provided by the inventors to advance the art, and should not be construed as limiting the scope of the present invention to the specifically set forth examples and conditions. Furthermore, all statements herein that describe the essence, aspects, and embodiments of the present invention, as well as specific examples of the present invention, are intended to encompass both structural and functional equivalents of the present invention. Additionally, such equivalents are intended to include both currently known equivalents and future-developed equivalents, i.e., all elements developed that perform the same function, regardless of structure. Furthermore, the descriptions disclosed herein are not intended to be publicly disclosed, regardless of whether such disclosure is explicitly recited in the claims.
[0361] Accordingly, it is not intended that the scope of the present invention be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the present invention are embodied by the appended claims. With respect to claims, 35 U.S.C. 112(f) or 35 U.S.C. 112(6) are expressly provided to be invoked with respect to a limitation in a claim only when the precise phrase "means for" or "step for" appears at the beginning of such limitation in the claim; if such precise phrase is not used in a claim limitation, 35 U.S.C. 112(f) or 35 U.S.C. 112(6) is not invoked.
[0362] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the filing date of U.S. Provisional Patent Application No. 63 / 416,674, filed October 17, 2022, the disclosure of which is incorporated herein by reference.
Claims
1. 1. A method for classifying cytometry image data, comprising: receiving unsorted cytometry image data including a plurality of images corresponding to image channels; iteratively adjusting an aspect of at least one of the plurality of images of the cytometry image data and applying a model to the adjusted cytometry image data to classify the cytometry image data; The method, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data.
2. 2. The method of claim 1, wherein applying the model to the cytometry image data to classify the cytometry image data results in an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
3. 3. The method of claim 2, wherein applying the model to the cytometry image data results in a confidence score associated with an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
4. adjusting aspects of at least one image of a plurality of images of unclassified cytometry image data and applying the model to the adjusted cytometry image data in an iterative manner, obtaining a plurality of estimates and confidence scores; The method of claim 3 , wherein the estimate and the confidence score each correspond to an iteration of applying the model.
5. 5. The method of claim 1, wherein adjusting a plurality of aspects of a first image of a plurality of images of the cytometry image data and applying the model to the adjusted cytometry image data is performed iteratively.
6. 6. The method of claim 1, further comprising iteratively adjusting a first aspect of two or more images of the plurality of images of the cytometry image data and applying the model to the adjusted cytometry image data.
7. 7. The method of claim 1, wherein the iterative adjustment of at least one image aspect and applying the model to the adjusted cytometry image data involves applying a simulated annealing technique.
8. 8. The method of claim 1, further comprising applying a predetermined set of image filters to the cytometry image data during the iterative process of adjusting at least one image aspect and applying the model to the adjusted cytometry image data.
9. 9. The method of claim 1, wherein the iterative adjustment of at least one image aspect and application of the model to the adjusted cytometry image data first includes applying a non-configurable image filter to the cytometry image data.
10. 10. The method of claim 1, further comprising identifying adjusted cytometry image data corresponding to an estimate of the presence of particles belonging to a first category of particles in the cytometry image data.
11. The method of any one of claims 1 to 10, wherein adjusting aspects of at least one image comprises applying digital signal processing techniques to said at least one image.
12. The method of any one of claims 1 to 11, wherein the model is updated based on the results of iteratively applying the model to adjusted cytometry image data.
13. The method of any one of claims 1 to 12, wherein the model comprises a machine learning model.
14. 1. A method of training a model to classify cytometry image data, comprising: receiving flow cytometry data including unsorted cytometry image data including a plurality of images, each instance of the cytometry image data corresponding to an image channel; classifying each instance of cytometry image data of said flow cytometry data to establish ground truth data; and training a model to classify the cytometry image data as data including particles belonging to a first category of particles by adjusting aspects of at least one image among a plurality of images of each instance of the cytometry image data of the ground truth data.
15. 1. A system for classifying cytometry image data, comprising: a processor operatively coupled to a memory; The memory, when executed by the processor, causes the processor to: receiving unsorted cytometry image data including a plurality of images corresponding to image channels; iteratively adjusting an aspect of at least one of the plurality of images of the cytometry image data and applying a model to the adjusted cytometry image data to classify the cytometry image data. It remembers the commands, The system, wherein the model is trained to predict the presence of particles belonging to a first category of particles in the cytometry image data.