Method and system for singlet identification in flow cytometry data and system therefor

Multiparameter analysis with distance-based classification models and DBSCAN clustering improves singlet identification in flow cytometry data, enhancing accuracy and reproducibility by 10-95% over manual methods.

JP2026004224APending Publication Date: 2026-01-14BECTON DICKINSON & CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025091302
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-31
Filing Date
2025-05-30
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Manual gating techniques for distinguishing singlets from multiplets in flow cytometry data are subjective, leading to inconsistencies, errors, and reduced reproducibility, and increase workflow complexity.

Method used

Implement multiparameter analysis using distance-based classification models and density-based clustering algorithms, such as DBSCAN, to automatically distinguish singlets from multiplets and debris, enhancing reproducibility and accuracy.

Benefits of technology

The method significantly increases the accuracy and reproducibility of singlet identification by 10-95% compared to manual gating, reducing errors and simplifying data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026004224000010
    Figure 2026004224000010
  • Figure 2026004224000011
    Figure 2026004224000011
  • Figure 2026004224000012
    Figure 2026004224000012
Patent Text Reader

Abstract

To provide an objective and repeatable process based on established parameters with little or no user input (or bias). It can be automated (e.g., using machine learning) and provides consistent, objective, and rapid singlet identification for complex samples.SOLUTION: Aspects of the present disclosure include methods for classifying analyte data. A method according to certain embodiments includes applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space, applying a density-based clustering algorithm to separate analyte data into a high density cluster and a low density cluster based on the density threshold, and classifying the analyte data based on the high density cluster and the low density cluster based on the size-based analyte feature space. Systems and non-transitory computer-readable storage media configured to perform the subject methods are also provided.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Characterization of analytes in biological fluids has become an important part of biological research, medical diagnosis, and the assessment of a patient's overall health and wellness. Detecting analytes in biological fluids, such as human blood or blood-derived products, can provide results that can play a role in determining treatment protocols for patients with various disease states.

[0002] Flow cytometry is a technique used to characterize and frequently sort biological materials, such as cells in a blood sample or particles of interest in another type of biological or chemical sample. Flow cytometers typically include a sample reservoir for receiving a fluid sample, such as a blood sample, and a sheath reservoir containing a sheath fluid. The flow cytometer directs the sheath fluid toward the flow cell while transporting particles (including cells) in the fluid sample as a stream of cells toward the flow cell. To characterize components of the flow stream, light is illuminated onto the flow stream. Variations in materials within the flow stream, such as the form or presence of fluorescent labels, can cause variations in the observed light, which enable characterization and separation. To characterize components of the flow stream, light must impinge on the flow stream and be collected. The light source in a flow cytometer can vary and include one or more broad-spectrum lamps, light-emitting diodes, and single-wavelength lasers. The light source is aligned with the flow stream, and the optical response from the illuminated particles is collected and quantified.

[0003] Isolation of biological particles has been achieved by adding sorting or collection capabilities to flow cytometers. Particles in the separated stream are detected as possessing one or more desired properties and are individually isolated from the sample stream by mechanical or electrical removal. A common flow sorting technique utilizes droplet sorting, in which a fluid stream containing linearly separated particles is split into droplets. Droplets containing the particles of interest are charged and deflected into a collection tube by passing through an electric field. Typically, linearly separated particles in the stream are characterized as they pass an observation point positioned directly below the nozzle tip. Once a particle is identified as meeting one or more desired criteria, the time at which it will reach its breakoff point and break away from the stream can be predicted. Ideally, a charge is applied for a short period of time just before the droplet containing the selected particle breaks off from the fluid stream, and then grounded immediately after the droplet breaks off. The droplets to be sorted retain their charge as they break off from the fluid stream, while all other droplets remain uncharged.

[0004] Parameters measured using a flow cytometer typically include light at the excitation wavelength scattered by particles primarily at narrow angles along the forward direction, called forward scatter (FSC), excitation light scattered by particles in a direction orthogonal to the excitation laser, called side scatter (SSC), and light emitted by fluorescent molecules at one or more detectors that measure signals across a range of spectral wavelengths, or light emitted by fluorescent dyes that are primarily detected at a specific detector or detector array. Different cell types can be distinguished by their light scattering properties and fluorescence emissions, which arise from labeling various cellular proteins or other components with fluorochrome-conjugated antibodies or other fluorescent probes.

[0005] The flow cytometer may further comprise a means for recording the measured data and analyzing the data. For example, data storage and analysis may be performed using a computer connected to the detection electronics. For example, the data may be stored in a table format, with each row corresponding to the data of one particle and each column corresponding to each measured feature. The use of a standard file format, such as the "FCS" file format, for storing data from a particle analyzer facilitates analysis of the data using a separate program and / or machine. Using current analysis methods, data is typically displayed as a one-dimensional histogram or two-dimensional (2D) plot for ease of visualization, although other methods may be used to visualize multidimensional data.

[0006] While flow cytometer data typically contain numerous data points (i.e., events), only a subset of the data is often of interest to users. For example, it may be desirable to identify the best parameters for distinguishing debris / small particles from single cells and multiplets. Debris is essentially cellular debris broken up during processing. Multiplets are two or more cells joined together. Cell debris can be considered "junk," i.e., data that the user does not wish to collect or process for further analysis. Multiplets are also events whose fluorescent signal is twice that observed from single cells, making them aberrant events and therefore desirable to remove from analysis. Removal of such debris and multiplets is often the first step in analyzing flow data. In other words, doublet and clump exclusion is an important preliminary step in establishing purity sorts. Doublets can disrupt purity and potentially lead to undesired results or unnecessary costs when performing downstream functional and / or genomic analyses. Summary of the Invention

[0007] The inventors have recognized that distinguishing singlets from multiplets in analyte data by manual gating techniques on two-dimensional dot plots based on scattering parameters is a subjective process, often based on user experience. Singlet identification by this process exhibits inconsistencies and errors due to inter-user subjectivity, limiting reproducibility and consistency in data analysis. Furthermore, manual gating of data on dot plots limits analysis to only two dimensions. Manual singlet identification also increases workflow complexity and introduces additional sources of error into data analysis. Embodiments of the present disclosure address, among other issues, the above-mentioned problems. The inventors have confirmed that singlet identification can be facilitated using multiparameter analysis, such as that described herein, of, for example, 50 or more different parameters, providing significantly greater sensitivity beyond what can be extracted using just two dimensions. In some instances, the multiparameter algorithms for distinguishing, and in some instances, sorting, singlets from multiplets and debris in samples described in the present disclosure provide an objective and repeatable process based on established parameters with little or no user input (or bias). Furthermore, the described embodiments can be automated (such as using machine learning) to provide consistent, objective, and rapid singlet discrimination for complex samples.

[0008] The present disclosure provides improvements to the process of classifying analyte data (e.g., flow cytometer data), for example, in the process of removing data associated with undesired analytes (e.g., debris, multiplets, etc.). In certain embodiments, the algorithms described herein can automatically distinguish single cells (singlets) from debris or other undesired events, such as multiplets. The methods described herein provide a higher accuracy rate for identifying singlets in flow cytometry, enhance data preprocessing / cleaning, and ensure that multiplets and debris are not included in downstream data analysis. The subject methods can also reduce errors associated with distinguishing singlets from multiplets by manual gating (e.g., by a user) and increase the reproducibility of the generated dataset, e.g., by increasing reproducibility by 10% or more, such as 20% or more, for example 30% or more, for example 40% or more, for example 50% or more, for example 60% or more, such as 70% or more, for example 80% or more, for example 90% or more, for example 95% or more (including 99% or more) compared to distinguishing singlets from multiplets by manual gating.

[0009] Aspects of the present disclosure include methods for classifying analyte data. The method, according to certain embodiments, includes applying a distance-based classification model to determine a density distinction threshold in a size-based analyte feature space, applying a density-based clustering algorithm to separate the analyte data into high-density clusters and low-density clusters based on the density threshold, and classifying the analyte data based on the high-density clusters and low-density clusters based on the size-based analyte feature space. Systems and non-transitory computer-readable storage media configured to perform the subject methods are also provided.

[0010] In some embodiments, the distance-based classification model is a nearest neighbor algorithm. In some examples, the density-based clustering algorithm is a density-based spatial clustering of applications with noise (DBSCAN) algorithm. In some examples, low-density data clusters are discarded. In some examples, the low-density data clusters include one or more multiplets. In some examples, the multiplets are doublets. In some examples, the multiplets are triplets. In some examples, the applied density-based clustering algorithm further distinguishes high-density clusters between debris clusters and singlet clusters. In certain examples, the debris clusters are discarded. In some examples, the applied density-based clustering algorithm further distinguishes high-density clusters by ordering the singlet clusters against a size-based analyte feature space. In some examples, the size-based analyte feature space includes one or more of a light-loss analyte feature, a major axis moment analyte feature, and a radial moment analyte feature. In certain examples, the size-based analyte feature space includes an imaging analyte feature. In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Major Axis Moment (SSC(Imaging)), Radial Moment (FSC), Radial Moment (Light Loss(Imaging)), and Radial Moment (SSC(Imaging)). In some examples, the size-based analyte feature space includes Light Loss(Violet)-A and Light Loss(Violet)-H. In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Light Loss(Violet)-W, Size(FSC), FSC-A, Radial Moment (Light Loss(Imaging)), Radial Moment (SSC(Imaging)). In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Minor Axis Moment (Light Loss(Imaging)), SSC(Imaging)-A, Light Loss(Imaging)-A, SSC(Violet)-A.In some examples, the size-based analyte feature space includes FSC-A and FSC-H. In some examples, the size-based analyte feature space includes 2-10 analyte features, e.g., 3-8 analyte features, e.g., 4-6 analyte features. In some examples, the classification model is an unsupervised algorithm. In some embodiments, classifying the analyte data is based solely on feature density, and not on identified features of the cells themselves. In some examples, features of the cells are not used for classification.

[0011] In some embodiments, the method includes training a model (e.g., a machine learning algorithm) to classify the analyte data. In some examples, training the model includes determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data and predicting a classification of the analyte dataset based on the ground truth analyte data. In some examples, the supervised learning algorithm is a random forest classifier. In some examples, training the model includes discarding predicted classifications below a confidence level and iteratively predicting a classification of the analyte dataset. In some examples, the confidence level is in the range of 60% to 100%, e.g., 70% to 95%, including 80% to 90%.

[0012] In some embodiments, the method further includes characterizing the classification of the population clusters. In some examples, a precision statistic is calculated for the classification of the population clusters. In some examples, a sensitivity statistic is calculated for the classification of the population clusters.

[0013] Aspects of the present disclosure include a system for classifying analyte data (e.g., flow cytometry data). The system according to certain embodiments includes a memory operably coupled to a processor, the memory having stored therein instructions that, when executed by the processor, cause the processor to: apply a distance-based classification model to determine a density distinction threshold in a size-based analyte feature space; apply a density-based clustering algorithm to separate the analyte data into high-density clusters and low-density clusters based on the density threshold; and classify the analyte data based on the high-density clusters and low-density clusters based on the size-based analyte feature space. In some examples, the system is or includes (e.g., is operably connected to) a flow cytometer. In some examples, the flow cytometer is an imaging-capable flow cytometer. Methods of the present disclosure may, in some cases, include providing analyte data to a system described herein and receiving classified analyte data from the system.

[0014] In some embodiments, the distance-based classification model is a nearest neighbor algorithm. In some embodiments, the distance-based classification model is a nearest neighbor algorithm. In some examples, the density-based clustering algorithm is a density-based spatial clustering of applications with noise (DBSCAN) algorithm. In some examples, the memory includes instructions for discarding low-density data clusters. In some examples, the low-density data clusters include one or more multiplets. In some examples, the multiplets are doublets. In some examples, the multiplets are triplets. In some examples, the memory includes instructions for distinguishing high-density clusters by distinguishing between debris clusters and singlet clusters. In some examples, the memory includes instructions for discarding debris clusters. In some examples, the memory includes instructions for distinguishing high-density clusters by ordering multiple singlet clusters against a size-based analyte feature space. In some examples, the size-based analyte feature space includes one or more of a light-loss analyte feature, a major axis moment analyte feature, and a radial moment analyte feature. In certain examples, the size-based analyte feature space includes an imaging analyte feature. In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Major Axis Moment (SSC(Imaging)), Radial Moment (FSC), Radial Moment (Light Loss(Imaging)), and Radial Moment (SSC(Imaging)). In some examples, the size-based analyte feature space includes Light Loss(Violet)-A and Light Loss(Violet)-H. In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Light Loss(Violet)-W, Size(FSC), FSC-A, Radial Moment (Light Loss(Imaging)), Radial Moment (SSC(Imaging)). In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Minor Axis Moment (Light Loss(Imaging)), SSC(Imaging)-A, Light Loss(Imaging)-A, SSC(Violet)-A.In some examples, the size-based analyte feature space includes FSC-A, FSC-H, hi some examples, the size-based analyte feature space includes 2 to 10 analyte features, e.g., 3 to 8 analyte features, e.g., 4 to 6 analyte features.

[0015] In some embodiments, the memory includes instructions for training a model to classify analyte data. In some examples, the memory includes instructions for training a model, such as having instructions for determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data, and instructions for predicting a classification of the analyte dataset based on the ground truth analyte data. In some examples, the supervised learning algorithm is a random forest classifier. In some examples, the memory includes instructions for discarding predicted classifications below a confidence level and for iterating the prediction of a classification of the analyte dataset. In some examples, the confidence level is in the range of 60% to 100%, e.g., 70% to 95%, including 80% to 90%. In some embodiments, the memory further includes instructions for characterizing the classification of the population clusters. In some examples, the memory includes instructions for calculating a precision statistic of the classification of the population clusters. In some examples, the memory includes instructions for calculating a sensitivity statistic of the classification of the population clusters.

[0016] In certain embodiments, the system includes a display for displaying a graphical user interface. In some examples, the display is configured to output the classified analyte data.

[0017] A non-transitory computer-readable storage medium having instructions with algorithms for classifying analyte data is also described. The non-transitory computer-readable storage medium according to certain embodiments includes an algorithm for applying a distance-based classification model to determine a density distinction threshold in a size-based analyte feature space, an algorithm for applying a density-based clustering algorithm to separate the analyte data into high-density clusters and low-density clusters based on the density threshold, and an algorithm for classifying the analyte data based on the high-density and low-density clusters based on the size-based analyte feature space.

[0018] In some embodiments, the distance-based classification model is a nearest neighbor algorithm. In some embodiments, the distance-based classification model is a nearest neighbor algorithm. In some examples, the density-based clustering algorithm is a density-based spatial clustering of applications with noise (DBSCAN) algorithm. In some examples, the non-transitory computer-readable storage medium includes an algorithm for discarding low-density data clusters. In some examples, the low-density data clusters include one or more multiplets. In some examples, the multiplets are doublets. In some examples, the multiplets are triplets. In some examples, the non-transitory computer-readable storage medium includes an algorithm for distinguishing high-density clusters by distinguishing between debris clusters and singlet clusters. In some examples, the non-transitory computer-readable storage medium includes an algorithm for discarding debris clusters. In some examples, the non-transitory computer-readable storage medium includes an algorithm for distinguishing high-density clusters by ordering multiple singlet clusters against a size-based analyte feature space. In some examples, the size-based analyte feature space includes one or more of a light-loss analyte feature, a major axis moment analyte feature, and a radial moment analyte feature. In certain examples, the size-based analyte feature space includes imaging analyte features. In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Longitudinal Moment (SSC(Imaging)), Radial Moment (FSC), Radial Moment (Light Loss(Imaging)), and Radial Moment (SSC(Imaging)). In some examples, the size-based analyte feature space includes Light Loss(Violet)-A and Light Loss(Violet)-H. In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Light Loss(Violet)-W, Size(FSC), FSC-A, Radial Moment (Light Loss(Imaging)), Radial Moment (SSC(Imaging)).In some examples, the size-based analyte feature space includes Light Loss (Violet)-A, Minor Axis Moment (Light Loss (Imaging)), SSC(Imaging)-A, Light Loss (Imaging)-A, SSC(Violet)-A. In some examples, the size-based analyte feature space includes FSC-A, FSC-H. In some examples, the size-based analyte feature space includes 2 to 10 analyte features, e.g., 3 to 8 analyte features, e.g., 4 to 6 analyte features.

[0019] In some embodiments, the non-transitory computer-readable storage medium includes an algorithm for training a model to classify analyte data. In some examples, the non-transitory computer-readable storage medium includes an algorithm for training a model, such as having an algorithm for determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data, and an algorithm for predicting a classification of the analyte dataset based on the ground truth analyte data. In some examples, the supervised learning algorithm is a random forest classifier. In some examples, the non-transitory computer-readable storage medium includes an algorithm for discarding predicted classifications below a confidence level and an algorithm for iteratively predicting a classification of the analyte dataset. In some examples, the confidence level is in the range of 60% to 100%, e.g., 70% to 95%, including 80% to 90%. In some embodiments, the non-transitory computer-readable storage medium includes an algorithm for characterizing the classification of the population clusters. In some examples, the non-transitory computer-readable storage medium includes an algorithm for calculating a precision statistic for the classification of the population clusters. In some examples, the non-transitory computer-readable storage medium includes an algorithm for calculating a sensitivity statistic for classification of population clusters.

[0020] The present disclosure can be best understood from the following detailed description when read in conjunction with the accompanying drawings, which include the following figures: [Brief explanation of the drawings]

[0021] [Figure 1] 1 shows a flowchart for implementing a method for classifying analyte data in accordance with certain embodiments. [Figure 2] 1 illustrates a flow cytometry system according to certain embodiments. [Figure 3-1] 1 illustrates an image-enabled particle sorter in accordance with certain embodiments. [Figure 3-2] 1 illustrates an image-enabled particle sorter in accordance with certain embodiments. [Figure 4] FIG. 1 illustrates a functional block diagram of a particle analysis system in accordance with certain embodiments. [Figure 5] FIG. 1 illustrates a functional block diagram of an example control system in accordance with certain embodiments. [Figure 6A] 1 illustrates a schematic diagram of a particle sorter system in accordance with certain embodiments. [Figure 6B] 1 illustrates a schematic diagram of a particle sorter system in accordance with certain embodiments. [Figure 7] 1 illustrates aspects of a computer control system according to certain embodiments. [Figure 8] 1 illustrates particle label classification for flow cytometry data according to certain embodiments. [Figure 9] Figures 1A-C show separation of singlets and multiplets by a density-based algorithm, according to certain embodiments. A shows images of singlets and multiplets generated based on light loss parameters. B shows a density plot of the data using a light loss feature set. C shows separation of the data into density-based clusters using the DBSCAN algorithm. [Figure 10] 1A and 1B illustrate determining a density threshold based on a k-nearest neighbor method according to certain embodiments. A shows a histogram of k-distances for event data in a dataset. B shows a Gaussian-smoothed histogram of k-distances. [Figure 11]Figure 1 shows clustering of data using a light loss feature set using the density-based DBSCAN algorithm. A shows the original data clustered based on the analyte features (light loss (violet)-A, light loss (violet)-H). B shows the dense data clustered based on the analyte features (light loss (violet)-A, light loss (violet)-H). [Figure 12] Cluster distances for separating singlets from debris are shown. A. Histogram of distances from the origin. B. Gaussian-smoothed histogram of cluster distances. [Figure 13] Figure 1 shows clustering of dense data using the optical loss feature set. A shows two separate clusters of data corresponding to debris and singlets. B shows the classification of the two clusters into a debris cluster and a singlet cluster. [Figure 14] 1A shows calculated precision and sensitivity statistics for different feature sets according to certain embodiments. A shows calculated precision and sensitivity statistics for singlet classification. B shows precision and sensitivity statistics calculated for multiplet classification using different feature sets. C shows precision and sensitivity statistics calculated for debris classification using different feature sets. [Figure 15] 1 shows a comparison of precision and sensitivity of different feature sets for training and test data according to certain embodiments. (A) shows precision and sensitivity statistics for training data for five different feature sets. (B) shows precision and sensitivity statistics for test data for five different feature sets. [Figure 16] 10 shows a summary of singlet classification accuracy and singlet classification sensitivity for different feature sets compared to manual gating, according to certain embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0022] Aspects of the present disclosure include methods for classifying analyte data. The method, according to certain embodiments, includes applying a distance-based classification model to determine a density distinction threshold in a size-based analyte feature space, applying a density-based clustering algorithm to separate the analyte data into high-density clusters and low-density clusters based on the density threshold, and classifying the analyte data based on the high-density clusters and low-density clusters based on the size-based analyte feature space. Systems and non-transitory computer-readable storage media configured to perform the subject methods are also provided.

[0023] Before describing the present disclosure in more detail, it is to be understood that this disclosure is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure will be limited only by the appended claims.

[0024] When a range of values ​​is presented, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limits of that range, and any other stated or intervening value in that stated range, is encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limits in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.

[0025] Certain ranges are described herein by numerical values ​​preceded by the term "about." The term "about" is used herein to literally support the exact number it precedes, as well as a number that is near or approximately the number preceded by the term. In determining whether a number is near or approximately a specifically stated number, a number not stated to be near or approximately may be a number that, in the context in which it is presented, represents a substantial equivalent to the specifically stated number.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of this disclosure, representative exemplary methods and materials are now described.

[0027] All publications and patents cited herein are incorporated by reference to disclose and describe the methods and / or materials for which the publications are cited, as if each individual publication or patent was specifically and individually indicated to be incorporated by reference. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.

[0028] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a predicate for use of exclusive terminology, such as "solely," "only," and the like, in connection with the recitation of claim elements or the use of a "negative" limitation.

[0029] As will be apparent to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features that may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the disclosure. Any recited method can be carried out in the order of events recited or in any other order that is logically possible.

[0030] Although the systems and methods have been or will be described for grammatical fluidity with functional descriptions, it is to be clearly understood that the claims should not be construed as necessarily limited by "means" or "step" limitation constructions unless expressly formulated under 35 U.S.C. § 112, but rather should be given the full scope of meaning and equivalents of the definitions provided by the claims under the doctrine of equivalents, and that if a claim is expressly formulated under 35 U.S.C. § 112, the full legal equivalents under 35 U.S.C. § 112 should be given.

[0031] How to classify analyte data Aspects of the present disclosure include methods for classifying analyte data, e.g., into data clusters. "Analyte data" refers to data obtained by evaluating a particular analyte for a particular characteristic. "Classifying" analyte data means designating analyte data (e.g., a group of analyte data) as belonging to a particular type among one or more possible different types to which the data may belong. The methods of the present disclosure may, in some cases, be sufficient to improve the classification of analyte data compared to traditional classification methods, such as when analyte data are manually classified by a user (e.g., by creating gates on flow cytometer data). For example, the subject methods may, in certain embodiments, increase the accuracy rate of classification. The accuracy rate may be determined by evaluating whether each analyte or related data point / event actually belongs to the particular type being classified. In certain cases, the methods of the present disclosure may increase the accuracy rate of classification by 1% or more, e.g., 5% or more, e.g., 10% or more, e.g., 15% or more (including 20% ​​or more), compared to traditional methods (e.g., creating manual gates). In embodiments, performing the subject methods is sufficient to increase the speed and / or efficiency with which analyte data is classified by 1% or more, e.g., 5% or more, e.g., 10% or more, e.g., 15% or more (including 20% ​​or more) compared to conventional methods (e.g., manual gate creation). In some embodiments, the methods of the present disclosure are computer-implemented methods. In other words, the method steps described herein may be performed via a processor associated with a computer system. Any suitable processor, such as those described below with respect to the systems of the present disclosure, may be used in the subject methods.

[0032] In some cases, the analyte data is flow cytometer data. "Flow cytometer data" refers to information about the characteristics of sample particles collected by any number of detectors in a particle analyzer. As discussed herein, a "particle analyzer" is an analytical tool (e.g., a flow cytometer) that enables characterization of particles based on specific (e.g., optical) parameters. "Particle" refers to a discrete component of a biological sample, such as a molecule, an analyte-bound bead, or an individual cell. While the present disclosure is primarily described with respect to flow cytometer data, the applicability of the present disclosure is not limited to flow cytometer data. In certain cases, the present disclosure may be applicable to other types of data, such as nucleic acid data.

[0033] The flow cytometer data may be received from any suitable source. In some embodiments, the flow cytometer data is received from a memory of a storage device. In such embodiments, the flow cytometer data may be pre-generated and stored in the memory of a storage device for subsequent retrieval and analysis. In other embodiments, the flow cytometer data is received in real time. Stated differently, the flow cytometer data generated during the operation of the flow cytometer may then (e.g., immediately) be loaded into a data space (e.g., a two-dimensional plot). In embodiments, the flow cytometer data is received from a forward scatter detector. The forward scatter detector may, in some instances, provide information regarding the overall size of the particle. In embodiments, the flow cytometer data is received from a side scatter detector. The side scatter detector may, in some instances, be configured to detect refracted and reflected light from the surface and internal structure of the particle, which tends to increase as the particle becomes more complex in structure.

[0034] In certain embodiments, particles are detected and uniquely identified by exposing them to excitation light and, as desired, measuring the fluorescence of each particle in one or more detection channels. Fluorescence emitted in the detection channels used to identify particles and their associated binding complexes can be measured after excitation by a single light source or separately after excitation by separate light sources. When separate excitation light sources are used to excite particle labels, the labels can be selected so that all labels are excitable by each of the excitation light sources used. In embodiments, flow cytometer data is received from a fluorescence detector. The fluorescence detector can, in some instances, be configured to detect fluorescent emissions from fluorescent molecules, such as labeled specific binding members (e.g., labeled antibodies that specifically bind to markers of interest) associated with particles in the flow cell. In certain embodiments, the method includes detecting fluorescence from the sample using one or more fluorescence detectors, e.g., two or more, e.g., three or more, e.g., four or more, e.g., five or more, e.g., six or more, e.g., seven or more, e.g., eight or more, e.g., nine or more, e.g., ten or more, e.g., fifteen or more (including 25 or more) fluorescence detectors. In embodiments, each of the fluorescence detectors is configured to generate a fluorescence data signal. Fluorescence from the sample can be independently detected by each fluorescence detector over one or more wavelength ranges of 200 nm to 1200 nm. In some examples, the method includes detecting fluorescence from the sample over wavelength ranges of, for example, 200 nm to 1200 nm, for example, 300 nm to 1100 nm, for example, 400 nm to 1000 nm, or for example, 500 nm to 900 nm (including 600 nm to 800 nm). In other examples, the method includes detecting fluorescence using each fluorescence detector at one or more specific wavelengths. For example, fluorescence may be detected at one or more of 450 nm, 518 nm, 519 nm, 561 nm, 578 nm, 605 nm, 607 nm, 625 nm, 650 nm, 660 nm, 667 nm, 670 nm, 668 nm, 695 nm, 710 nm, 723 nm, 780 nm, 785 nm, 647 nm, 617 nm, and any combination thereof, depending on the number of different fluorescence detectors in the subject optical detection system.In certain embodiments, the method includes detecting wavelengths of light that correspond to the fluorescence peak wavelengths of particular fluorophores present in the sample. In embodiments, the flow cytometer data is received from one or more photodetectors (e.g., one or more detection channels), such as two or more, such as three or more, such as four or more, such as five or more, such as six or more (including eight or more) photodetectors (e.g., eight or more detection channels).

[0035] In some cases, the methods may include classifying multiple analyte data sets, e.g., two or more sets, three or more sets, four or more sets, five or more sets, seven or more sets, eight or more sets, nine or more sets (including ten or more sets). In such cases, the sets may be from the same source or from different sources. The number of data points (e.g., events, observations) classified by the subject methods may also vary. In some cases, the number of data points ranges from 1k to 100k, e.g., 10k to 80k, e.g., 20k to 60k, including 25k to 50k. In some embodiments the number of data points is 1k or more, such as 5k or more, for example 10k or more, such as 15k or more, for example 20k or more, such as 25k or more, for example 30k or more, such as 35k or more, for example 40k or more, such as 45k or more, for example 50k or more, such as 55k or more, for example 60k or more, such as 65k or more, for example 70k or more, such as 75k or more, for example 80k or more, such as 85k or more, for example 90k or more, such as 95k or more (including 100k or more).

[0036] In some cases, before clustering the analyte data, the method includes preprocessing the data, e.g., so that it is in a form more suitable for operation by a different model. Any suitable preprocessing protocol may be used. In some embodiments, the method includes standardizing the analyte features, e.g., so that the analyte features are centered around the mean and scaled to unit variance.

[0037] A method according to certain embodiments includes applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space, applying a density-based clustering algorithm to separate the analyte data into high-density clusters and low-density clusters based on the density threshold, and classifying the analyte data based on the high-density clusters and low-density clusters in the size-based analyte feature space. "Analyte feature" refers to one or more characteristics (e.g., optical, impedance, and / or temporal characteristics) associated with each individual analyte (e.g., particle) such that each analyte is present in the analyte data as a set of digitized feature values. Depending on the requirements of a given experiment, the number of analyte features present in the data may vary, for example, including 10 or more features, for example, 20 or more features, for example, 30 or more features, for example, 40 or more features, or for example, 50 or more features (including 60 or more features). In certain examples, the analyte feature is selected from a size feature, an imaging feature, and a scattering feature. In some such examples, the analyte feature is a scattering feature selected from a side scatter (SSC) feature and a forward scatter (FSC) feature. When the analyte data is flow cytometer data, the analyte features may also be related to and / or derived from fluorescence, axial light loss (ALL), and the like. Exemplary features include, but are not limited to, size, center of mass, minor axis moment, diffusivity, major axis moment, radial moment, maximum intensity, and eccentricity. In some examples, the size-based analyte feature space includes light loss (violet)-A, major axis moment (SSC(imaging)), radial moment (FSC), radial moment (light loss (imaging)), and radial moment (SSC(imaging)). In some examples, the size-based analyte feature space includes light loss (violet)-A and light loss (violet)-H. In some examples, the size-based analyte feature space includes light loss (violet)-A, light loss (violet)-W, size (FSC), FSC-A, radial moment (light loss (imaging)), radial moment (SSC(imaging)).In some examples, the size-based analyte feature space includes Light Loss (Violet)-A, Minor Axis Moment (Light Loss (Imaging)), SSC(Imaging)-A, Light Loss (Imaging)-A, and SSC(Violet)-A. In some examples, the size-based analyte feature space includes FSC-A and FSC-H. In some examples, the size-based analyte feature space includes 2-10 analyte features, for example, 3-8 analyte features, for example, 4-6 analyte features. In some examples, the classification model is an unsupervised algorithm. In some embodiments, classifying the analyte data is based solely on feature density and not on identified features of the cells themselves. In some examples, features of the cells are not used in the classification. In some embodiments, the analyte features include one or more image parameters. In some examples, one or more image parameters are calculated from the generated particle images. In some examples, a centroid image parameter is calculated from the generated images. In some examples, a delta centroid image parameter is calculated from the generated images. In some examples, a diffuse image parameter is calculated from the generated images. In some examples, decentration image parameters are calculated from the generated images. In some examples, major axis moment image parameters are calculated from the generated images. In some examples, maximum intensity image parameters are calculated from the generated images. In some examples, radial moment image parameters are calculated from the generated images. In some examples, minor axis moment image parameters are calculated from the generated images. In some examples, particle size image parameters are calculated from the generated images. In some examples, total intensity image parameters are calculated from the generated images. In some examples, particle light loss image parameters are calculated from the generated images. In some examples, forward scatter image parameters are calculated from the generated images. In some examples, side scatter image parameters are calculated from the generated images. In some examples, image moments are calculated from the generated images. The term "image moment" is used herein in its conventional sense to refer to a weighted average of pixel intensities in an image. In some examples, a center of mass may be calculated from the image moments of an image.In other examples, the orientation of the cell can be calculated from the image moments of the image. In yet other examples, the eccentricity of the cell can be calculated from the image moments of the image.

[0038] In certain embodiments, imaging parameters (described below) calculated from the images generated for use in generating the gating strategy are summarized in Table 1.

[0039] [Table 1-1]

[0040] [Table 1-2]

[0041] The "cluster criteria" discussed herein can be any suitable standard by which analyte data can be evaluated and classified into clusters. For example, the cluster criteria can be the association of analyte data with a particular parameter of interest. Analyte data can be considered "associated" with a parameter of interest if the analyte corresponding to the data can be said to correspond to the parameter of interest, i.e., is positive for the parameter. Such analyte data can be classified into a particular cluster, and analyte data that do not correspond to the cluster criteria can be classified into one or more other clusters (e.g., clusters that are negative for the parameter of interest). In some embodiments, analyte data that do not correspond to the cluster criteria can be classified into a single cluster (i.e., so that there are two total clusters). Alternatively, analyte data that do not correspond to the cluster criteria can be classified into multiple clusters according to some other criteria, e.g., two or more clusters, three or more clusters, four or more clusters, five or more clusters, six or more clusters, seven or more clusters, eight or more clusters, nine or more clusters (including ten or more clusters). The cluster criteria can vary according to the nature of the analytes being observed and / or the nature of the experiment being performed. In some cases, the cluster criteria is the association of analyte data with singlets. In such cases, flow cytometer data associated with singlets may be classified into a singlet cluster, while flow cytometer data not associated with singlets may be classified into one or more non-singlet clusters. Non-singlets that may be classified into non-singlet clusters may include, but are not limited to, multiplets / aggregates (e.g., doublets, triplets, quadruplets, quintuplets, etc.), debris (e.g., components of lysed cells), etc. In accordance with the above, non-singlets may be classified into a single non-singlet cluster, or may be further segmented into multiple non-singlet clusters (e.g., one cluster corresponding to doublets and one corresponding to debris, etc., as desired). In some cases, non-singlet clusters include doublets or aggregates.Other cluster criteria may include association with specific fluorescent markers that may themselves be associated with particular phenotypes of the analytes, depending on the nature of the experiment being performed. In some embodiments, the cluster criteria is the size or shape of the analytes.

[0042] In embodiments, the method includes applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space. In some examples, the method includes identifying high-density and low-density regions within the data based on a provided set of features. In some examples, multiplets, which can vary in size (e.g., from 2 to 20 or more cells), are more spread out in feature space targeting the size of the event. In some examples, singlets exhibit less variance and fall within the high-density regions of the data. In some embodiments, the distance-based classification model is a nearest neighbor algorithm.

[0043] In some examples, low-density data clusters are discarded. In some examples, low-density data clusters include one or more multiplets. In some examples, the multiplets are doublets. In some examples, the multiplets are triplets. In certain examples, the multiplets include aggregates of four or more particles, e.g., five or more particles, e.g., ten or more particles (including aggregates of twenty or more particles). In some examples, the applied density-based clustering algorithm further distinguishes high-density clusters between debris clusters and singlet clusters. In certain examples, the debris clusters are discarded. In some examples, the applied density-based clustering algorithm further distinguishes high-density clusters by ordering multiple singlet clusters with respect to a size-based analyte feature space.

[0044] In some examples, the density-based clustering algorithm is a density-based spatial clustering for applications involving noise (DBSCAN) algorithm. DBSCAN groups data points by density, with points with higher densities forming clusters. Details regarding DBSCAN can be found, for example, in Ester et al. (1996) Proceedings of the Second International Conference on Knowledge Discovery and Data. 96(34):226-231, which is incorporated herein by reference. In some cases, the density-based clustering algorithm is a K-means clustering algorithm. K-means clustering involves dividing observations into clusters, with each observation belonging to the cluster (e.g., centroid) with the closest mean. Details regarding K-means clustering can be found, for example, in Lloyd, Stuart P. (1967) IEEE Transactions on Information Theory. 28(2):129-137, which is incorporated herein by reference. In other cases, the density-based clustering algorithm is a balanced iterative reducing and clustering using hierarchies (BIRCH) algorithm. BIRCH is an unsupervised data mining algorithm used for hierarchical clustering. Details regarding BIRCH can be found, for example, in Zhang et al. (1996) ACM sigmod record. 25(2):103-114. In certain implementations, BIRCH can be used to complement one or more of the other algorithms discussed herein. For example, in some embodiments, BIRCH is used to accelerate K-means clustering. In further embodiments, BIRCH is used to accelerate Gaussian mixture modeling as described herein. In certain versions, the density-based clustering algorithm is a spectral clustering algorithm.Spectral clustering involves the use of eigenvalues ​​of a similarity matrix. More information on spectral clustering can be found, for example, in Von Luxburg, U. (2007) Statistics and computing 17:395-416, which is incorporated herein by reference.

[0045] In some embodiments, the method includes training a model (e.g., a machine learning algorithm) to classify analyte data. The training data may be received from any suitable source. In some embodiments, the training data is received from a memory of a storage device. In such embodiments, the training data may be pre-generated and stored in the memory of a storage device for subsequent retrieval and analysis. In embodiments, the analyte data in the training dataset are of known classifications. For example, in some cases where the training dataset includes flow cytometer data, each individual analyte may be confirmed to correspond to one class or another by some other means. In certain examples, an expert user manually provides classifications for the training dataset. These may include, for example, manual gating on a two-dimensional plot of the flow cytometer data. These classifications, in addition to analyte features from the training dataset, may be provided for training purposes. In some embodiments, the method includes using multiple training datasets, e.g., two or more training datasets, e.g., three or more training datasets, e.g., four or more training datasets, e.g., five or more training datasets, e.g., ten or more training datasets, e.g., twenty-five or more training datasets (including fifty or more training datasets). In some examples, training the model includes determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data and predicting a classification of the analyte dataset based on the ground truth analyte data. In some examples, the supervised learning algorithm is a random forest classifier. In some examples, training the model includes discarding predicted classifications below a confidence level and iteratively predicting classifications of the analyte dataset. In some examples, the confidence level is in the range of 60% to 100%, e.g., 65% to 95%, e.g., 70% to 90%, e.g., 75% to 85%, inclusive. In other words, the training data may be considered "ground truth" data.In some versions, such ground truth data is obtained by using a specific dye or pigment in a flow cytometer experiment known to correspond to the analyte characteristic of interest. The dye or pigment may be selected depending on the nature of the characteristic. For example, if it is desired to cluster singlets and non-singlets, ground truth data can be obtained using a DNA-intercalating dye. Cells (at least eukaryotic types) generally have a single nucleus containing DNA. Thus, by staining the DNA, a user can reliably determine whether a given event / observation involves one cell (i.e., a singlet) or multiple cells (i.e., non-singlets). DNA-intercalating dyes that can be used include, but are not limited to, ethidium bromide, SYBR Green, propidium iodide, acridine orange, DAPI, and DRAQ5.

[0046] The disclosed methods also include applying a classification model to classify the analyte data into clusters. Following classification of the analyte data, some versions of the method may include evaluating the classification model, for example, by comparing the classification to ground truth data. In some cases, classifying the analyte data using methods described herein includes including 90% or more, e.g., 90% or more (including 97% or more), of the analyte data associated with the cluster criteria within a cluster associated with the cluster criteria. Furthermore, classifying the analyte data using methods described herein may include excluding 85% or more, e.g., 90% or more (including 92% or more) of the analyte data not associated with the cluster criteria from a cluster associated with the cluster criteria. In some embodiments, the method includes generating one or more population clusters based on analyte characteristics in the sample. As used herein, a "population" or "subpopulation" of analytes, such as cells, nucleic acids, or other particles, generally refers to a group of analytes having properties (e.g., optical properties, impedance properties, or temporal properties) related to one or more measured parameters such that the measured parameter data form a cluster in data space. In embodiments, the data are comprised of signals from any given number of different parameters, such as two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more (including twenty or more). Thus, populations are recognized as clusters in the data. Conversely, each data cluster is generally interpreted as representing a population of a particular type of cell or analyte, although clusters representing noise or background are also typically observed. Clusters can be defined in terms of a subset of dimensions, for example, with respect to subsets of measured parameters that represent populations that differ only in a subset of the measured parameters or features extracted from measurements of cells, particles, or nucleic acids.

[0047] In some embodiments, the method includes receiving data, calculating parameters for each analyte, and clustering the analytes based on the calculated parameters. For example, if the data is flow cytometer data, the experiment may include particles labeled with several fluorophores or fluorescently labeled antibodies, and groups of particles may be defined by populations corresponding to one or more fluorescence measurements. In this example, a first group may be defined by a certain range of light scattered by the first fluorophore, and a second group may be defined by a certain range of light scattered by the second fluorophore. If the first fluorophore and the second fluorophore are represented on the x-axis and y-axis, respectively, two differently colored populations may appear to define each group of particles when the information is displayed graphically. Any number of analytes may be assigned to a cluster containing five or more analytes, for example, ten or more analytes, for example, fifty or more analytes, for example, one hundred or more analytes, or even 500 analytes (including 1000 analytes). In certain embodiments, the methods group rare events (e.g., rare cells in a sample, such as cancer cells) detected in a sample into clusters. In these embodiments, the generated clusters of analytes may include 10 or fewer assigned analytes, such as 9 or fewer assigned analytes (including 5 or fewer assigned analytes).

[0048] In some embodiments, the method further includes characterizing the classification of the population clusters. In some examples, a precision statistic is calculated for the classification of the population clusters. In some examples, the precision statistic represents the correctness of the predicted label for the target class.

[0049] TIFF2026004224000003.tif12170

[0050] In some examples, a sensitivity statistic is calculated for the classification of the population clusters. In some examples, the sensitivity statistic represents the proportion of the target class that is captured by the prediction.

[0051] TIFF2026004224000004.tif12170

[0052] 1 shows a flowchart for implementing a method according to certain embodiments. As shown in FIG. 1, step 101 involves applying a distance-based classification model to determine a density distinction threshold in a size-based analyte feature space. In step 102, a density-based clustering algorithm is applied to separate the analyte data into high-density clusters and low-density clusters based on the density threshold. In step 103, the analyte data is classified into high-density clusters and low-density clusters based on the size-based analyte feature space.

[0053] In certain cases, the methods of the present disclosure can be performed in conjunction with the methods described in U.S. Provisional Patent Application No. 63 / 569,559, filed March 25, 2024, the disclosure of which is incorporated herein by reference in its entirety. In such cases, the computer-implemented method can include categorizing analyte data based on analyte features associated therewith by using an ensemble of decision trees to generate predicted classes for the analyte data, and refining the categorized analyte data based on the analyte features and predicted classes using a distance-based classification model to classify the analyte data.

[0054] In some examples, the sample analyzed in the present method is a biological sample. The term "biological sample" is used in its conventional sense, and in particular refers to a whole organism, whole plant, whole fungus, or a subset of animal tissues, cells, or component parts that can be found in blood, mucus, lymph, synovial fluid, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, amniotic fluid, amniotic fluid, umbilical cord blood, urine, vaginal fluid, and semen. Thus, a "biological sample" refers to both an intact organism or a subset of its tissues, as well as homogenates, lysates, or extracts prepared from an organism or a subset of its tissues, including, but not limited to, plasma, serum, cerebrospinal fluid, lymph, skin, respiratory, gastrointestinal, cardiovascular, and genitourinary tract sections, tears, saliva, milk, blood cells, tumors, and organs. A biological sample may be any type of biological tissue, including both healthy and diseased tissues (e.g., cancerous, malignant, necrotic, etc.). In certain embodiments, the biological sample is a liquid sample such as blood or a derivative thereof, e.g., plasma, tears, urine, semen, etc., and in some instances, the sample is a blood sample, including whole blood, such as blood obtained from a venipuncture or finger stick (which may or may not be combined with any reagents, such as preservatives, anticoagulants, etc., prior to assay).

[0055] In certain embodiments, the source of the sample is a "mammal" or "mammalian," a term used broadly to describe organisms belonging to the class Mammalia, including the orders Carnivora (e.g., dogs and cats), Rodentia (e.g., mice, guinea pigs, and rats), and Primates (e.g., humans, chimpanzees, and monkeys). In some examples, the subject is a human. The present methods may be applied to samples obtained from human subjects of both genders and at any developmental stage (i.e., newborn, infant, juvenile, adolescent, adult), and in certain embodiments, the human subject is a juvenile, adolescent, or adult. It should be understood that while the present disclosure may be applied to samples from human subjects, the methods may also be performed on samples from other animal subjects (i.e., "non-human subjects"), such as, but not limited to, birds, mice, rats, dogs, cats, livestock, and horses.

[0056] Cells of interest may be targeted for characterization according to various parameters, such as phenotypic characteristics identified through the attachment of specific fluorescent labels to the cells of interest. In some embodiments, the system is configured to deflect analyzed droplets determined to contain target cells. A variety of cells may be characterized using the subject methods. Target cells of interest include, but are not limited to, stem cells, T cells, dendritic cells, B cells, granulocytes, leukemia cells, lymphoma cells, viral cells (e.g., HIV cells), NK cells, macrophages, monocytes, fibroblasts, epithelial cells, endothelial cells, and erythroid cells. Target cells of interest include cells bearing favorable cell surface markers or antigens that can be captured or labeled by favorable affinity factors or conjugates thereof. For example, target cells may comprise cell surface antigens such as CD11b, CD123, CD14, CD15, CD16, CD19, CD193, CD2, CD25, CD27, CD3, CD335, CD36, CD4, CD43, CD45RO, CD56, CD61, CD7, CD8, CD34, CD1c, CD23, CD304, CD235a, T cell receptor alpha / beta, T cell receptor gamma / delta, CD253, CD95, CD20, CD105, CD117, CD120b, Notch4, Lgr5 (N-terminus), SSEA-3, TRA-1-60 antigen, disialoganglioside GD2, and CD71. In some embodiments, the target cells are selected from HIV-containing cells, Treg cells, antigen-specific T cell populations, tumor cells or hematopoietic progenitor cells (CD34+) obtained from whole blood, bone marrow, or umbilical cord blood.

[0057] In practicing the subject methods, a volume of an initial fluid sample is injected into a flow cytometer. The volume of sample injected into the flow sorting module can be in the range of, for example, 0.001 mL to 1000 mL, e.g., 0.005 mL to 900 mL, e.g., 0.01 mL to 800 mL, e.g., 0.05 mL to 700 mL, e.g., 0.1 mL to 600 mL, e.g., 0.5 mL to 500 mL, e.g., 1 mL to 400 mL, e.g., 2 mL to 300 mL (including samples of 5 mL to 100 mL).

[0058] Methods according to embodiments of the present disclosure include counting and, optionally, sorting labeled particles (e.g., target cells) in a sample. In performing the subject methods, a fluid sample containing particles is first introduced into a flow nozzle of the system. Upon exiting the flow nozzle, the particles pass substantially one at a time through a sample interrogation region, where each of the particles is illuminated by a light source, and measurements of light scattering parameters, and in some instances, desired fluorescence emissions (e.g., measurements of two or more light scattering parameters and one or more fluorescence emissions), are recorded separately for each particle. Depending on the characteristics of the flow stream being investigated, the light may be illuminated into 0.001 mm or more of the flow stream, e.g., 0.005 mm or more, 0.01 mm or more, 0.05 mm or more, 0.1 mm or more, 0.5 mm or more (including illuminating 1 mm or more of the flow stream). In certain embodiments, the method includes illuminating a planar cross-section of the flow stream within the sample interrogation region, such as with a laser (as described above). In other embodiments, the method includes illuminating a predetermined length of the flow stream within the sample interrogation region, such as corresponding to the illumination profile of a diffuse laser beam or lamp.

[0059] In certain embodiments, the method comprises irradiating the flow stream at or near the flow cell nozzle orifice. For example, the method can comprise irradiating the flow stream at a location about 0.001 mm or more, e.g., 0.005 mm or more, e.g., 0.01 mm or more, e.g., 0.05 mm or more, e.g., 0.1 mm or more, e.g., 0.5 mm or more from the nozzle orifice (including 1 mm or more from the nozzle orifice). In certain embodiments, the method comprises irradiating the flow stream directly adjacent to the flow cell nozzle orifice.

[0060] In embodiments of the present method, detectors such as photomultiplier tubes (PMTs) are used to record the light passing through each particle (called forward light scatter in certain cases), the light reflected perpendicular to the direction of the particle's stream passing through the detection region (called orthogonal or side light scatter in some cases), and, if the particle is labeled with a fluorescent marker, the fluorescence emitted by the particle as it passes through the detection region and is illuminated by an energy source. Forward light scatter (FSC), side scatter (SSC), and fluorescence emission each comprise a separate parameter for each particle (or each particle "event"). Thus, for example, two, three, or four parameters can be collected (and recorded) from particles labeled with two different fluorescent markers. The data recorded for each particle can be analyzed in real time or, if desired, stored in a data storage and analysis means, such as a computer.

[0061] In certain embodiments, particles are detected and uniquely identified by exposing them to excitation light and measuring the fluorescence of each particle in one or more detection channels as needed.The fluorescence emitted in the detection channels used to identify particles and their associated binding complexes can be measured after excitation by a single light source, or can be measured separately after excitation by separate light sources.When separate excitation light sources are used to excite particle labels, the labels can be selected so that all labels can be excited by each of the excitation light sources used.

[0062] In certain embodiments, the method also includes data acquisition, analysis, and recording using a computer or the like, where multiple data channels record data from each detector for emitted light scattering and fluorescence as each particle passes through the sample interrogation region of the particle sorting module. In these embodiments, analysis includes classifying and counting particles so that each particle is represented as a set of digitized parameter values. The system of interest can be configured to trigger on selected parameters to distinguish particles of interest from background and noise. "Trigger" refers to a preset threshold for the detection of a parameter and may be used as a means for detecting the passage of a particle through a light source. Detection of an event exceeding the threshold for the selected parameter triggers the acquisition of light scattering and fluorescence data for the particle. Data is not acquired for particles or other components in the medium being assayed that cause a response below the threshold. The trigger parameter can be the detection of forward scattered light caused by a particle passing through the light beam. The flow cytometer then detects and collects the light scattering and fluorescence data of the particle.

[0063] Specific subpopulations of interest are then further analyzed by "gating" based on the data collected for the entire population. To select the appropriate gate, the data is plotted to obtain the best possible subpopulation separation. This procedure may be performed by plotting forward light scatter (FSC) versus side (i.e., orthogonal) light scatter (SSC) on a two-dimensional dot plot. A subpopulation of particles is then selected (i.e., those cells within the gate), and particles not within the gate are excluded. Optionally, a gate may be selected by drawing a line around the desired subpopulation using a cursor on the computer screen. Only those particles within the gate are then further analyzed by plotting other parameters of these particles, such as fluorescence. If desired, the above analysis can be configured to result in counting the particles of interest in the sample.

[0064] The subject methods may further include using the particles in research, laboratory testing, or therapy. In some embodiments, the subject methods include obtaining individual cells prepared from a biological sample of a target fluid or tissue. For example, the subject methods include obtaining cells from a fluid or tissue sample used as a research or diagnostic specimen for a disease such as cancer. Similarly, the subject methods include obtaining cells from a fluid or tissue sample used for therapy. Cell therapy protocols are protocols in which viable cellular material, including, for example, cells and tissue, can be prepared and introduced into a subject as a therapeutic treatment. Conditions that can be treated by administering flow cytometry-sorted samples include, but are not limited to, blood disorders, immune system disorders, organ damage, and the like.

[0065] A typical cell therapy protocol may include the following steps: sample collection, cell isolation, genetic modification, culture and in vitro expansion, cell harvesting, sample volume reduction and washing, biopreservation, storage, and introduction of cells into a subject. The protocol may begin with collecting viable cells and tissues from a subject's source tissue to produce a cell and / or tissue sample. The sample may be collected by any suitable procedure, including, for example, administering a cell mobilizing agent to the subject, withdrawing blood from the subject, removing bone marrow from the subject, etc. After collecting the sample, cell enrichment may be performed by several methods, including, for example, centrifugation-based methods, filter-based methods, elution, magnetic separation, fluorescence-activated cell sorting (FACS), etc. In some cases, the enriched cells may be genetically modified by any convenient method, such as nuclease-mediated gene editing. The genetically modified cells may be cultured, activated, and expanded in vitro. In some cases, the cells may be preserved, e.g., cryopreserved, and stored for future use, whereupon the cells may be thawed and then administered to a patient, e.g., the cells may be infused into a patient.

[0066] system Aspects of the present disclosure also include systems (e.g., computer-implemented methods) for performing the above-described subject methods. The system according to certain embodiments includes a memory operatively coupled to a processor, the memory having stored therein instructions that, when executed by the processor, cause the processor to: apply a distance-based classification model to determine a density distinction threshold in a size-based analyte feature space; apply a density-based clustering algorithm to separate the analyte data into high-density clusters and low-density clusters based on the density threshold; and classify the analyte data based on the high-density clusters and low-density clusters based on the size-based analyte feature space.

[0067] The system may include a display and an operator input device. The operator input device may be, for example, a keyboard, a mouse, etc. The processing module includes a processor that accesses a memory in which instructions for performing the steps of the subject method are stored. The processing module may include an operating system, a graphical user interface (GUI) controller, a system memory, a memory storage device, an input / output controller, a cache memory, a data backup unit, and many other devices. The processor may be a commercially available processor or one of other processors that are or become available. The processor executes an operating system that interfaces with firmware and hardware in well-known ways and facilitates the processor's coordination and execution of the functions of various computer programs, which may be written in various programming languages, such as Java, Perl, C++, Python, other high-level or low-level languages, and combinations thereof, as known in the art. The operating system typically cooperates with the processor to coordinate and execute the functions of the other components of the computer. The operating system also provides scheduling, input / output control, file and data management, memory management, and communication control and related services, all in accordance with known techniques. In some embodiments, the processor includes analog electronics that provide feedback control, such as negative feedback control.

[0068] The system memory may be any of a variety of known or future memory storage devices. Examples include any commonly available random access memory (RAM), magnetic media such as a resident hard disk or tape, optical media such as a read-and-write compact disc, a flash memory device, or other memory storage device. The memory storage device may be any of a variety of known or future devices, including a compact disc drive, tape drive, or diskette drive. Such types of memory storage devices typically read from and / or write to a program storage medium (not shown), such as a compact disc. Any of these program storage media, or others now in use or that may later be developed, may be considered a computer program product. As will be appreciated, these program storage media typically store computer software programs and / or data. Computer software programs, also referred to as computer control logic, are typically stored in the system memory and / or in program storage devices used in conjunction with the memory storage devices.

[0069] In some embodiments, a computer program product is described that includes a computer usable medium having stored thereon control logic (a computer software program including program code). The control logic, when executed by a processor of a computer, causes the processor to perform the functions described herein. In other embodiments, some functions are implemented primarily in hardware, for example, using hardware state machines. Implementing a hardware state machine to perform the functions described herein will be apparent to one skilled in the art.

[0070] The subject programmable logic may be implemented in any of a variety of devices, such as a specifically programmed event processing computer, a wireless communication device, an integrated circuit device, etc. In some embodiments, the programmable logic may be executed by a specially programmed processor, which may include one or more processors, such as one or more digital signal processors (DSPs), configurable microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Combinations of computing devices, such as a DSP with a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration in at least partial data connection, may implement one or more of the features described herein.

[0071] The memory may be any suitable device from which a processor can store and retrieve data, such as a magnetic, optical, or solid-state storage device (including a magnetic or optical disk, or tape, or RAM, or any other suitable device, fixed or portable). The processor may include a general-purpose digital microprocessor that is appropriately programmed from a computer-readable medium carrying the necessary program code. The programming may be provided to the processor remotely through a communications channel or may be pre-stored in a computer program product, such as memory or some other portable or fixed computer-readable storage medium using any of these devices in conjunction with the memory. For example, a magnetic or optical disk may carry the program and be read by a disk writer / reader. The system of the present disclosure also includes programming in the form of a computer program product, e.g., algorithms for use in implementing the above-described methods. Programming according to the present disclosure may be recorded on a computer-readable medium, e.g., any medium that can be directly read and accessed by a computer. Such media include, but are not limited to, magnetic storage media such as floppy disks, hard disk storage media, and magnetic tape; optical storage media such as CD-ROM; electrical storage media such as RAM, ROM, portable flash drives, and hybrids of these categories such as magnetic / optical storage media.

[0072] The processor may also have access to a communication channel for communicating with a remote user, where remote means that the user is not in direct contact with the system but relays input information to the input manager from an external device, such as a computer connected to a wide area network ("WAN"), a telephone network, a satellite network, or any other suitable communication channel, including a mobile phone (i.e., a smartphone).

[0073] In some embodiments, a system according to the present disclosure may be configured to include a communications interface. In some embodiments, the communications interface includes a receiver and / or a transmitter for communicating with a network and / or another device. The communications interface may be configured for wired or wireless communications, including, but not limited to, radio frequency (RF) communications (e.g., radio frequency identification (RFID), Zigbee communications protocol, Wi-Fi, infrared, wireless universal serial bus (USB), ultra-wideband (UWB), Bluetooth® communications protocol, and cellular communications such as code division multiple access (CDMA) or global system for mobile communications (GSM).

[0074] In one embodiment, the communication interface is configured to include one or more communication ports, e.g., a physical port or interface such as a USB port, a USB-C port, an RS-232 port, or any other suitable electrical connection port that enables data communication between the system of interest and other external devices, such as a computer terminal (e.g., in a doctor's office or hospital environment) configured for similar complementary data communication.

[0075] In one embodiment, the communication interface is configured for infrared communication, Bluetooth® communication, or any other suitable wireless communication protocol to enable the target system to communicate with computer terminals and / or other devices such as networks, communication-enabled mobile phones, personal digital assistants, or any other communication device that a user can integrate and use.

[0076] In one embodiment, the communication interface is configured to provide connectivity for data transfer using Internet Protocol (IP) over a cellular network, Short Message Service (SMS), a wireless connection to a personal computer (PC) in a local area network (LAN) connected to the Internet, or a Wi-Fi connection to the Internet at a Wi-Fi hotspot.

[0077] In one embodiment, the target system is configured to wirelessly communicate with a server device via a communications interface using common standards such as, for example, 802.11 or Bluetooth® RF protocols, or the IrDA infrared protocol. The server device may be another portable device, such as a smartphone, personal digital assistant (PDA), or notebook computer, or a larger device, such as a desktop computer, appliance, etc. In some embodiments, the server device has a display, such as a liquid crystal display (LCD), and input devices, such as buttons, a keyboard, a mouse, or a touchscreen.

[0078] In some embodiments, the communication interface is configured to automatically or semi-automatically communicate data stored in the target system, e.g., the optional data storage unit, with a network or server device using one or more of the communication protocols and / or mechanisms described above.

[0079] The output controller may include a controller for any of a variety of known display devices for presenting information to a user, whether human or machine, local or remote. When one of the display devices provides visual information, this information may typically be logically and / or physically organized as an array of pixels. A graphical user interface (GUI) controller provides a graphical input / output interface between the system and the user and may include any of a variety of known or future software programs for processing user input. The functional elements of the computer may communicate with each other via a system bus. Some of these communications may be achieved in alternative embodiments using a network or other type of remote communication. The output manager may also provide information generated by the processing modules to a user at a remote location, for example, via the Internet, telephone, or satellite network, in accordance with known techniques. Presentation of data by the output manager may be implemented in accordance with various known techniques. As some examples, the data may include SQL, HTML, or XML documents, email or other files, or data in other formats. The data may also include Internet URL addresses so that the user can retrieve additional SQL, HTML, XML, or other documents or data from remote sources. The one or more platforms present in the subject system can be any type of known or future-developed computer platform, but they are typically a class of computers commonly referred to as servers. However, they may also be mainframe computers, workstations, or other computer types. They may be connected via any known or future type of cabling or other communication system, including wireless systems, and may or may not be networked. They may be co-located or physically separated.In some cases, various operating systems may be used for any computer platform, depending on the type and / or manufacturer of the computer platform selected. Suitable operating systems include Windows NT, Windows XP, Windows 7, Windows 8, Windows 10, iOS, macOS, Linux, Ubuntu, Fedora, OS / 400, i5 / OS, IBM i, Android™, SGI IRIX, Oracle Solaris, etc.

[0080] In embodiments, the system further includes a flow cytometer operably connected to the processor. Flow cytometers of interest generally include a flow cell. A flow cell of interest includes a cuvette configured to transport particles in a flow stream. As discussed herein, "flow cell" is described in its conventional sense, referring to a component that includes a flow channel for a liquid flow stream for transporting particles of sheath fluid. A cuvette of interest has a passageway (i.e., a flow channel) therethrough. The flow stream of which the flow channel is configured may include a liquid sample injected from a sample tube. In certain examples, the flow cell includes an optically accessible flow channel. The cuvette may be constructed of, for example, quartz, glass, clear plastic, or the like. In some embodiments, the cuvette is formed from silica, such as fused silica. In some cases, the flow cell is configured to be illuminated with light from a light source at one or more interrogation points. As discussed herein, "interrogation point" refers to an area within the flow cell where particles are illuminated by light from a light source, for example, for analysis. The size of the interrogation point may vary as needed. For example, if 0 μm represents the optical axis of the light emitted by the light source, the interrogation point may be in the range of -50 μm to 50 μm, e.g., -25 μm to 40 μm (including -15 μm to 30 μm). Depending on specific considerations (e.g., number and placement of lasers), multiple illumination points may be present within the flow cell.

[0081] In some embodiments, the flow cell includes or is configured for use with a sample injection port configured to provide a sample to the flow cell, hi embodiments, the sample injection system is configured to provide a suitable flow of sample into the flow cell internal chamber (e.g., flow channel). Depending on the desired characteristics of the flow stream, the rate at which the sample is delivered to the flow cell chamber by the sample injection port may be 1 μL / min or more, for example 2 μL / min or more, for example 3 μL / min or more, for example 5 μL / min or more, for example 10 μL / min or more, for example 15 μL / min or more, for example 25 μL / min or more, for example 50 μL / min or more (including 100 μL / min or more), and in some examples the rate at which the sample is delivered to the flow cell chamber by the sample injection port is 1 μL / sec or more, for example 2 μL / sec or more, for example 3 μL / sec or more, for example 5 μL / sec or more, for example 10 μL / sec or more, for example 15 μL / sec or more, for example 25 μL / sec or more, for example 50 μL / sec or more (including 100 μL / sec or more).

[0082] The sample injection port can be an orifice located in the wall of the internal chamber, or a conduit located at the proximal end of the internal chamber. When the sample injection port is an orifice located in the wall of the internal chamber, the sample injection port orifice can have any suitable cross-sectional shape, including, but not limited to, rectilinear cross-sectional shapes such as square, rectangular, trapezoidal, triangular, and hexagonal, curvilinear cross-sectional shapes such as circular and elliptical, as well as irregular shapes such as a parabolic bottom joined to a flat top. In certain embodiments, the sample injection port has a circular orifice. The size of the sample injection port orifice can vary depending on the shape, with specific examples ranging from 0.1 mm to 5.0 mm, e.g., 0.2 mm to 3.0 mm, e.g., 0.5 mm to 2.5 mm, e.g., 0.75 mm to 2.25 mm, e.g., 1 mm to 2 mm (including 1.25 mm to 1.75 mm), e.g., having an opening of 1.5 mm.

[0083] In certain examples, the sample injection port is a conduit located at the proximal end of the flow cell internal chamber. For example, the sample injection port may be a conduit positioned so that the orifice of the sample injection port is aligned with the flow cell orifice. When the sample injection port is a conduit aligned with the flow cell orifice, the cross-sectional shape of the sample injection tube may be any suitable shape, including, but not limited to, linear cross-sectional shapes such as square, rectangular, trapezoidal, triangular, and hexagonal, curved cross-sectional shapes such as circular and elliptical, as well as irregular shapes such as a parabolic bottom joined to a flat top. In certain examples, the orifice of the conduit may have an opening ranging from 0.1 mm to 5.0 mm, e.g., 0.2 mm to 3.0 mm, e.g., 0.5 mm to 2.5 mm, e.g., 0.75 mm to 2.25 mm, e.g., 1 mm to 2 mm (including 1.25 mm to 1.75 mm), e.g., 1.5 mm, and may vary depending on the shape. The shape of the tip of the sample injection port may be the same as or different from the cross-sectional shape of the sample injection tube. For example, the orifice of the sample injection port may include a beveled tip having an inclination angle in the range of 1° to 10°, e.g., 2° to 9°, e.g., 3° ​​to 8°, e.g., 4° to 7° (including an inclination angle of 5°).

[0084] In some embodiments, the flow cell also includes a sheath fluid injection port configured to supply sheath fluid to the flow cell. In embodiments, the sheath fluid injection system is configured to supply a flow of sheath fluid to the flow cell interior chamber, e.g., in conjunction with the sample, to produce a stacked flow stream of sheath fluid surrounding the sample flow stream. Depending on the desired characteristics of the flow stream, the velocity of sheath fluid delivered to the flow cell by the sheath fluid injection port can be 25 μL / sec or more, e.g., 50 μL / sec or more, e.g., 75 μL / sec or more, e.g., 100 μL / sec or more, e.g., 250 μL / sec or more, e.g., 500 μL / sec or more, e.g., 750 μL / sec or more, e.g., 1000 μL / sec or more (including 2500 μL / sec or more).

[0085] In some embodiments, the sheath fluid injection port is an orifice disposed in the wall of the internal chamber. The sheath fluid injection port orifice may have any suitable cross-sectional shape, including, but not limited to, rectilinear cross-sectional shapes such as square, rectangular, trapezoidal, triangular, and hexagonal, curvilinear cross-sectional shapes such as circular and elliptical, as well as irregular shapes such as a parabolic bottom joined to a flat top. The size of the sheath fluid injection port orifice may vary depending on the shape, with specific examples ranging from 0.1 mm to 5.0 mm, e.g., 0.2 mm to 3.0 mm, e.g., 0.5 mm to 2.5 mm, e.g., 0.75 mm to 2.25 mm, e.g., 1 mm to 2 mm (including 1.25 mm to 1.75 mm), e.g., having an opening of 1.5 mm.

[0086] The disclosed flow cytometer includes a light source configured to illuminate particles in the flow stream at an interrogation point within the flow cell. The number of light sources within a flow cytometer can vary. In some embodiments, the flow cytometer includes a single light source. Alternatively, the flow cytometer may include multiple light sources in some cases. In some such examples, the number of light sources ranges from 2 to 10, e.g., 2 to 5, including 2 to 4. Any convenient light source may be used as the light source described herein. In some embodiments, the light source is a laser. In embodiments, the laser may be any convenient laser, such as a continuous wave laser. For example, the laser may be a diode laser, such as an ultraviolet diode laser, a visible diode laser, or a near-infrared diode laser. In other embodiments, the laser may be a helium-neon (HeNe) laser. In some examples, the laser is a gas laser such as a helium-neon laser, an argon laser, a krypton laser, a xenon laser, a nitrogen laser, a CO laser, a CO laser, an argon-fluorine (ArF) excimer laser, a krypton-fluorine (KrF) excimer laser, a xenon-chlorine (XeCl) excimer laser, or a xenon-fluorine (XeF) excimer laser, or a combination thereof. In other examples, the flow cytometer of interest includes a dye laser such as a stilbene, coumarin, or rhodamine laser. In still other examples, the laser of interest includes a metal vapor laser such as a helium-cadmium (HeCd) laser, a helium-mercury (HeHg) laser, a helium-selenium (HeSe) laser, a helium-silver (HeAg) laser, a strontium laser, a neon-copper (NeCu) laser, a copper laser, or a gold laser, and combinations thereof. In yet another example, a subject flow cytometer includes a solid-state laser, such as a ruby ​​laser, a Nd:YAG laser, a NdCrYAG laser, an Er:YAG laser, a Nd:YLF laser, a Nd:YVO4 laser, a Nd:YCa4O(BO3)3 laser, a Nd:YCOB laser, a titanium sapphire laser, a thulium YAG laser, a ytterbium YAG laser, a ytterbium2O3 laser, or a cerium-doped laser, and combinations thereof.

[0087] The laser light source according to certain embodiments may also include one or more optical adjustment components. In certain embodiments, the optical adjustment component may include any device located between the light source and the flow cell that can change the spatial width of the illumination or some other characteristic of the illumination from the light source, such as the illumination direction, wavelength, beam width, beam intensity, and focus. The optical adjustment protocol may include any convenient device that adjusts one or more characteristics of the light source, including, but not limited to, lenses, mirrors, filters, optical fibers, wavelength separators, pinholes, slits, collimation protocols, and combinations thereof. In certain embodiments, the target flow cytometer includes one or more focusing lenses. In one example, the focusing lens may be a reduction lens. In yet other embodiments, the target flow cytometer includes optical fibers.

[0088] The light source may be positioned at any suitable distance from the flow cell, for example, the light source and the flow cell are separated by 0.005 mm or more, for example, 0.01 mm or more, for example, 0.05 mm or more, for example, 0.1 mm or more, for example, 0.5 mm or more, for example, 1 mm or more, for example, 5 mm or more, for example, 10 mm or more, for example, 25 mm or more (including a distance of 100 mm or more). Furthermore, the light source may be positioned at any suitable angle relative to the flow cell, for example, an angle in the range of 10 to 90 degrees, for example, 15 to 85 degrees, for example, 20 to 80 degrees, for example, 25 to 75 degrees (including 30 to 60 degrees), for example, an angle of 90 degrees.

[0089] In some embodiments, the intended light source includes multiple lasers, e.g., two or more lasers, e.g., three or more lasers, e.g., four or more lasers, e.g., five or more lasers, e.g., ten or more lasers (including fifteen or more lasers configured to provide laser light for discrete illumination of the flowstream), configured to provide laser light for discrete illumination of the flowstream. Depending on the desired wavelength of light for illuminating the flowstream, each laser may have a specific wavelength that varies from 200 nm to 1500 nm, e.g., 250 nm to 1250 nm, e.g., 300 nm to 1000 nm, e.g., 350 nm to 900 nm (including 400 nm to 800 nm). In certain embodiments, the intended lasers may include one or more of a 405 nm laser, a 488 nm laser, a 561 nm laser, and a 635 nm laser.

[0090] In certain embodiments, the light source is an optical beam generator configured to generate two or more frequency-shifted optical beams. In some examples, the optical beam generator includes a laser and a radio-frequency generator configured to apply a radio-frequency drive signal to an acousto-optic device to generate two or more angularly deflected laser beams. In these embodiments, the laser may be a pulsed laser or a continuous-wave laser. For example, the laser in the optical beam generator of interest may be a gas laser, such as a helium-neon laser, an argon laser, a krypton laser, a xenon laser, a nitrogen laser, a CO laser, a CO laser, an argon-fluorine (ArF) excimer laser, a krypton-fluorine (KrF) excimer laser, a xenon-chlorine (XeCl) excimer laser, a xenon-fluorine (XeF) excimer laser, or a combination thereof; a dye laser, such as a stilbene, coumarin, or rhodamine laser; a helium-cadmium (HeCd) laser, a helium-mercury (HeHg) laser, a helium-selenium (HeSe) laser, or a combination thereof. ) laser, helium-silver (HeAg) laser, strontium laser, neon-copper (NeCu) laser, copper laser or gold laser, and combinations thereof; solid-state lasers such as ruby ​​laser, Nd:YAG laser, NdCrYAG laser, Er:YAG laser, Nd:YLF laser, Nd:YVO4 laser, Nd:YCa4O(BO3)3 laser, Nd:YCOB laser, titanium sapphire laser, thulium YAG laser, ytterbium YAG laser, ytterbium2O3 laser, cerium-doped laser, and combinations thereof.

[0091] The acousto-optical device may be any convenient acousto-optical protocol configured to frequency-shift laser light using applied acoustic waves. In a specific embodiment, the acousto-optical device is an acousto-optical deflector. The acousto-optical device in the target system is configured to generate an angularly deflected laser beam from light from a laser and an applied high-frequency drive signal. The high-frequency drive signal may be applied to the acousto-optical device using any suitable high-frequency drive signal source, such as a direct digital synthesizer (DDS), an arbitrary waveform generator (AWG), or an electrical pulse generator.

[0092] In an embodiment, the controller is configured to apply high frequency drive signals to the acousto-optic device to produce a desired number of angularly deflected laser beams in the output laser beam, for example configured to apply three or more high frequency drive signals, for example four or more high frequency drive signals, for example five or more high frequency drive signals, for example six or more high frequency drive signals, for example seven or more high frequency drive signals, for example eight or more high frequency drive signals, for example nine or more high frequency drive signals, for example ten or more high frequency drive signals, for example fifteen or more high frequency drive signals, for example twenty-five or more high frequency drive signals, for example fifty or more high frequency drive signals (including being configured to apply 100 or more high frequency drive signals).

[0093] In some examples, to produce an intensity profile of the angularly deflected laser beam within the output laser beam, the controller is configured to apply a high frequency drive signal having an amplitude that varies, for example, from about 0.001 V to about 500 V, for example, from about 0.005 V to about 400 V, for example, from about 0.01 V to about 300 V, for example, from about 0.05 V to about 200 V, for example, from about 0.1 V to about 100 V, for example, from about 0.5 V to about 75 V, for example, from about 1 V to about 50 V, for example, from about 2 V to about 40 V, for example, from 3 V to about 30 V (including from about 5 V to about 25 V). In some embodiments, each of the applied high frequency drive signals has a frequency of about 0.001 MHz to about 500 MHz, for example, about 0.005 MHz to about 400 MHz, for example, about 0.01 MHz to about 300 MHz, for example, about 0.05 MHz to about 200 MHz, for example, about 0.1 MHz to about 100 MHz, for example, about 0.5 MHz to about 90 MHz, for example, about 1 MHz to about 75 MHz, for example, about 2 MHz to about 70 MHz, for example, about 3 MHz to about 65 MHz, for example, about 4 MHz to about 60 MHz (including about 5 MHz to about 50 MHz).

[0094] In certain embodiments, the controller includes a processor having a memory operatively coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the processor to generate an output laser beam having an angularly deflected laser beam with a desired intensity profile. For example, the memory may include instructions for generating two or more, e.g., three or more, e.g., four or more, e.g., five or more, e.g., ten or more, e.g., twenty-five or more, e.g., fifty or more, angularly deflected laser beams having the same intensity (including where the memory may include instructions for generating 100 or more angularly deflected laser beams having the same intensity). In other embodiments, the memory may include instructions for generating two or more, e.g., three or more, e.g., four or more, e.g., five or more, e.g., ten or more, e.g., twenty-five or more, e.g., fifty or more, angularly deflected laser beams having different intensities (including where the memory may include instructions for generating 100 or more angularly deflected laser beams having different intensities).

[0095] In certain embodiments, the controller includes a processor having a memory operatively coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the processor to produce an output laser beam having an intensity that decreases from the edge of the output laser beam to the center along a horizontal axis. In these examples, the intensity of the angularly deflected laser beam at the center of the output beam may be in a range of 0.1% to about 99%, e.g., 0.5% to about 95%, e.g., 1% to about 90%, e.g., about 2% to about 85%, e.g., about 3% to about 80%, e.g., about 4% to about 75%, e.g., about 5% to about 70%, e.g., about 6% to about 65%, e.g., about 7% to about 60%, e.g., about 8% to about 55% (including about 10% to about 50% of the intensity of the angularly deflected laser beam at the edge of the output laser beam along the horizontal axis). In another embodiment, the controller includes a processor having a memory operatively coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the processor to produce an output laser beam that increases in intensity from the edge to the center of the output laser beam along a horizontal axis. In these examples, the intensity of the angularly deflected laser beam at the edge of the output beam may be in the range of 0.1% to about 99%, e.g., 0.5% to about 95%, e.g., 1% to about 90%, e.g., about 2% to about 85%, e.g., about 3% to about 80%, e.g., about 4% to about 75%, e.g., about 5% to about 70%, e.g., about 6% to about 65%, e.g., about 7% to about 60%, e.g., about 8% to about 55% (including about 10% to about 50% of the intensity of the angularly deflected laser beam at the center of the output laser beam along the horizontal axis). In yet another embodiment, the controller comprises a processor having a memory operatively coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the processor to produce an output laser beam having an intensity profile with a Gaussian distribution along a horizontal axis.In yet another embodiment, the controller comprises a processor having a memory operatively coupled to the processor, the memory storing instructions that, when executed by the processor, cause the processor to produce an output laser beam having a top-hat intensity profile along a horizontal axis.

[0096] In embodiments, the objective optical beam generator may be configured to produce angularly polarized laser beams within the spatially separated output laser beam. Depending on the applied high frequency drive signal and the desired irradiance profile of the output laser beam, the angularly polarized laser beams may be separated by 0.001 μm or more, e.g., 0.005 μm or more, e.g., 0.01 μm or more, e.g., 0.05 μm or more, e.g., 0.1 μm or more, e.g., 0.5 μm or more, e.g., 1 μm or more, e.g., 5 μm or more, e.g., 10 μm or more, e.g., 100 μm or more, e.g., 500 μm or more, e.g., 1000 μm or more (including 5000 μm or more). In some embodiments, the system is configured to produce angularly polarized laser beams within the output laser beam that overlap with adjacent angularly polarized laser beams along the horizontal axis of the output laser beam, such as 5000 μm or more. The overlap between adjacent angularly deflected laser beams (e.g., beam spot overlap) may be an overlap of 0.001 μm or more, such as an overlap of 0.005 μm or more, for example an overlap of 0.01 μm or more, for example an overlap of 0.05 μm or more, for example an overlap of 0.1 μm or more, for example an overlap of 0.5 μm or more, for example an overlap of 1 μm or more, for example an overlap of 5 μm or more, for example an overlap of 10 μm or more (including an overlap of 100 μm or more).

[0097] In certain examples, the optical beam generator configured to generate two or more frequency-shifted optical beams includes a laser excitation module as described in U.S. Patent Nos. 9,423,353, 9,784,661, and 10,006,852, and U.S. Patent Application Publication Nos. 2017 / 0133857 and 2017 / 0350803, the disclosures of which are incorporated herein by reference.

[0098] The flow cytometer further includes a detector configured to collect light emitted by the illuminated particles. The photodetector is configured to detect the particle-modulated light carried by the fiber optic light collection element and generate a signal based on a characteristic (e.g., intensity) of the light. For example, the one or more particle-modulated light detectors may include one or more side-scattered light detectors for detecting side-scattered wavelengths of light (i.e., light refracted and reflected from the surface and internal structures of the particle). In some embodiments, the flow cytometer includes a single side-scattered light detector. In other embodiments, the flow cytometer includes multiple, e.g., two or more, e.g., three or more, e.g., four or more (including five or more), side-scattered light detectors.

[0099] Any convenient detector for detecting collected light may be used in the side-scattered light detector described herein. Detectors of interest may include, but are not limited to, optical sensors or detectors such as active pixel sensors (APS), avalanche photodiodes, image sensors, charge-coupled devices (CCDs), intensified charge-coupled devices (ICCDs), light-emitting diodes, photon counters, bolometers, pyroelectric detectors, photoresistors, photocells, photodiodes, photomultiplier tubes (PMTs), phototransistors, quantum dot photoconductors, or photodiodes, and combinations thereof, among other detectors. In certain embodiments, the collected light is measured with a charge-coupled device (CCD), a semiconductor charge-coupled device (CCD), an active pixel sensor (APS), a complementary metal-oxide semiconductor (CMOS) image sensor, or an N-type metal-oxide semiconductor (NMOS) image sensor. In certain embodiments, the detector has a resolution of 0.01 cm. 2 ~10cm 2 , e.g. 0.05cm 2 ~9cm 2 , e.g. 0.1cm 2 ~8cm 2 , e.g. 0.5cm 2 ~7cm 2 Range (1cm 2 ~5cm 2 and a photomultiplier tube such as a photomultiplier tube having an active detection surface area for each region including

[0100] In embodiments, a subject flow cytometer also includes a fluorescence detector configured to detect one or more fluorescent wavelengths of light, hi other embodiments, the flow cytometer includes a plurality of fluorescence detectors, e.g., two or more, e.g., three or more, e.g., four or more, five or more, ten or more, fifteen or more (including twenty or more).

[0101] Any convenient detector for detecting collected light may be used in the fluorescence detectors described herein. Detectors of interest may include, but are not limited to, optical sensors or detectors such as active pixel sensors (APS), avalanche photodiodes, image sensors, charge-coupled devices (CCDs), intensified charge-coupled devices (ICCDs), light-emitting diodes, photon counters, bolometers, pyroelectric detectors, photoresistors, photocells, photodiodes, photomultiplier tubes (PMTs), phototransistors, quantum dot photoconductors, or photodiodes, and combinations thereof, among other detectors. In certain embodiments, the collected light is measured with a charge-coupled device (CCD), a semiconductor charge-coupled device (CCD), an active pixel sensor (APS), a complementary metal-oxide semiconductor (CMOS) image sensor, or an N-type metal-oxide semiconductor (NMOS) image sensor. In certain embodiments, the detector has a resolution of 0.01 cm. 2 ~10cm 2 , e.g. 0.05cm 2 ~9cm 2 , e.g. 0.1cm 2 ~8cm 2 , e.g. 0.5cm 2 ~7cm 2 Range (1cm 2 ~5cm 2 and a photomultiplier tube such as a photomultiplier tube having an active detection surface area for each region including

[0102] When a subject flow cytometer includes multiple fluorescence detectors, each fluorescence detector may be the same, or the collection of fluorescence detectors may be a combination of different types of detectors. For example, when a subject flow cytometer includes two fluorescence detectors, in some embodiments, the first fluorescence detector is a CCD-type device and the second fluorescence detector (or imaging sensor) is a CMOS-type device. In other embodiments, both the first fluorescence detector and the second fluorescence detector are CCD-type devices. In still other embodiments, both the first fluorescence detector and the second fluorescence detector are CMOS-type devices. In still other embodiments, the first fluorescence detector is a CCD-type device and the second fluorescence detector is a photomultiplier tube (PMT). In still other embodiments, the first fluorescence detector is a CMOS-type device and the second fluorescence detector is a photomultiplier tube. In still other embodiments, both the first fluorescence detector and the second fluorescence detector are photomultiplier tubes.

[0103] In embodiments of the present disclosure, the subject fluorescence detectors are configured to measure collected light at one or more wavelengths, e.g., two or more wavelengths, e.g., five or more different wavelengths, e.g., ten or more different wavelengths, e.g., twenty-five or more different wavelengths, e.g., fifty or more different wavelengths, e.g., one hundred or more different wavelengths, e.g., two or more different wavelengths, e.g., three hundred or more different wavelengths (including measuring light emitted by a sample in a flow stream at four hundred or more different wavelengths). In some embodiments, two or more detectors of a module described herein are configured to measure the same or overlapping wavelengths of collected light.

[0104] In some embodiments, the intended fluorescence detector is configured to measure light collected over any range of wavelengths (e.g., 200 nm to 1000 nm). In certain embodiments, the intended detector is configured to collect a spectrum of light over any range of wavelengths. For example, a flow cytometer may include one or more detectors configured to collect a spectrum of light over one or more wavelength ranges from 200 nm to 1000 nm. In still other embodiments, the intended detector is configured to measure light emitted by a sample in the flow stream at one or more specific wavelengths. For example, a module may include one or more detectors configured to measure light at one or more of the following wavelengths: 450 nm, 518 nm, 519 nm, 561 nm, 578 nm, 605 nm, 607 nm, 625 nm, 650 nm, 660 nm, 667 nm, 670 nm, 668 nm, 695 nm, 710 nm, 723 nm, 780 nm, 785 nm, 647 nm, 617 nm, and any combination thereof. In certain embodiments, one or more detectors may be configured to be paired with a particular fluorophore, such as one used with a sample in a fluorescence assay.

[0105] The flow cytometer may include any suitable mechanism for supplying sheath fluid and sample fluid to the sample fluid input coupler and sheath fluid input coupler. For example, the sample fluid input coupler may be fluidly connected to a sample fluid line (e.g., tubing) that is fluidly connected to a sample fluid reservoir. Similarly, the sheath fluid input coupler may be fluidly connected to a sheath fluid line that is fluidly connected to a sheath fluid reservoir. Similarly, the flow cytometer may include any suitable mechanism for managing waste from the flow stream. The fluid output coupler may be fluidly connected to a waste line that is fluidly connected to a waste reservoir. A fluid management system that may be adapted for use with the subject flow cytometers is provided in U.S. Patent Application Publication No. 2022 / 0341838, the disclosure of which is incorporated herein by reference in its entirety.

[0106] Suitable flow cytometry systems include those described in Ormerod (ed.), Flow Cytometry: A Practical Approach, Oxford University Press (1997); Jaroszeski et al. (eds.), Flow Cytometry Protocols, Methods in Molecular Biology No. 91, Humana Press (1997); Practical Flow Cytometry, 3rd ed., Wiley-Liss (1995); Virgo, et al. (2012) Ann Clin Biochem. Jan; 49(pt1):17-28; Linden, et al., Semin Thromb Hemost. 2004 Oct; 30(5):502-11; Alison, et al. J Pathol. 2010 Dec; 222(4):335-344; and Herbig, et al. (2007) Crit Rev Ther Drug Carrier Syst. 24(3):203-255, the disclosures of which are incorporated herein by reference.In certain instances, flow cytometry systems of interest include a BD Biosciences FACSCanto™ flow cytometer, a BD Biosciences FACSCanto™ II flow cytometer, a BD Accuri™ flow cytometer, a BD Accuri™ C6 Plus flow cytometer, a BD Biosciences FACSCelesta™ flow cytometer, a BD Biosciences FACSLyric™ flow cytometer, a BD Biosciences FACSVerse™ flow cytometer, a BD Biosciences FACSymphony™ flow cytometer, a BD Biosciences LSRFortessa™ flow cytometer, a BD Biosciences LSRFortessa™ X-20 flow cytometer, a BD Biosciences FACSPresto™ flow cytometer, a BD Biosciences FACSVia™ flow cytometer, as well as a BD Biosciences FACSCalibur™ cell sorter, a BD Biosciences FACSCount™ cell sorter, a BD Biosciences These include the FACSLyric™ cell sorter, BD Biosciences Via™ cell sorter, BD Biosciences Influx™ cell sorter, BD Biosciences Jazz™ cell sorter, BD Biosciences Aria™ cell sorter, BD Biosciences FACSAria™ II cell sorter, BD Biosciences FACSAria™ III cell sorter, BD Biosciences FACSAria™ Fusion cell sorter, and BD Biosciences FACSMelody™ cell sorter, BD Biosciences FACSymphony™ S6 cell sorter, BD Biosciences FACSDiscover™ cell sorter, and others.

[0107] In some embodiments, the subject systems may be implemented using the same or similar technologies as those described in U.S. Patent Nos. 10,663,476, 10,620,111, 10,613,017, 10,605,713, 10,585,031, 10,578,542, and 10,578,469, the entire disclosures of which are incorporated herein by reference. No. 10,481,074, No. 10,302,545, No. 10,145,793, No. 10,113,967, No. 10,006,852, No. No. 9,952,076, No. 9,933,341, No. 9,726,527, No. 9,453,789, No. 9,200,334, No. 9,097,6 No. 40, No. 9,095,494, No. 9,092,034, No. 8,975,595, No. 8,753,573, No. 8,233,146, No. 8 ,140,300, No. 7,544,326, No. 7,201,875, No. 7,129,505, No. 6,821,740, No. 6,813,017 Nos. 6,809,804, 6,372,506, 5,700,692, 5,643,796, 5,627,040, 5,620,842, 5,602,039, 4,987,086, and 4,498,766.

[0108] In some embodiments, flow cytometer is configured as imaging flow cytometer.For example, in certain cases, the target system is a flow cytometry system configured to image particles in flow stream by fluorescence imaging using radiofrequency tagged emission (FIRE), as described in, for example, Diebold, et al.Nature Photonics Vol.7(10),806-810(2013) and U.S. Patent Nos. 9,423,353, 9,784,661 and 10,006,852, and U.S. Patent Application Publication Nos. 2017 / 0133857 and 2017 / 0350803, the disclosure of which is incorporated herein by reference.In some embodiments where flow cytometer is particle sorter, particle sorter is an image-enabled particle sorter. Image-enabled particle sorters are described in US Provisional Patent Applications Nos. 63 / 431,803 and 63 / 465,057, the disclosures of which are incorporated herein by reference in their entireties.

[0109] FIG. 2 illustrates a system 200 for flow cytometry according to an exemplary embodiment of the present disclosure. The system 200 includes a laser 201 configured to illuminate particles 211 in a flow stream 214 at an interrogation point 215 within a flow cell 210. While the example of FIG. 2 shows a single laser, it is understood that multiple lasers can also be used. The laser beam from the laser 201 is directed toward a focusing lens 202, which focuses the beam onto a portion of the fluid stream where the sample particles 211 are located within the flow cell 210. The flow cell 210 is part of a fluidics system that directs particles, typically one at a time, within the stream toward the focused laser beam for interrogation. Alternatively, if the flow cytometer is a stream-in air cytometer, a nozzle top can be used.

[0110] As shown in FIG. 2 , the flow cell 210 is fluidly connected to a sheath fluid reservoir 203 containing sheath fluid and a sample fluid reservoir 204 containing sample fluid. The sheath fluid from the sheath fluid reservoir 203 is supplied to at least one sheath fluid injection port 208 via a conduit (i.e., sheath fluid line) 207. Additionally, the sample fluid containing particles 211 from the sample fluid reservoir 204 is supplied to a sample injection port 206 via a conduit (i.e., sample fluid line) 205. The sample injection port 206 is fluidly connected to a sample injector 213 (e.g., a sample injection needle) configured to introduce the particles 211 into the interior of the flow cell 210. The particles 211 are hydrodynamically focused via the sheath fluid entering the sheath fluid injection port 208 such that a flow stream 214 is formed downstream of the tapered portion 212 of the flow cell 210. Particles emitting at the distal end of the flow cell 210 can be disposed of and / or collected via any suitable protocol. For example, depending on the type of flow cytometry being performed, particles may be collected at the distal end of flow cell 210, for example, via a waste line. Alternatively, particles may be sorted.

[0111] Light from the laser beam interacts with sample particles 211 through diffraction, refraction, reflection, scattering, and absorption by re-emission at a variety of different wavelengths, depending on particle characteristics such as particle size, internal structure, and the presence of one or more fluorescent molecules attached to or naturally present on or within the particle. The fluorescent light and diffracted, refracted, reflected, and scattered light may be sent to one or more detectors. In particular, forward-scattered light (FSC) is sent to a forward-scattered light detector 223. The forward-scattered light detector 223 is positioned slightly off-axis from the direct beam passing through the flow cell 210 and is configured to detect diffracted light and excitation light traveling primarily in the forward direction through or around the particle. The intensity of the light detected by the forward-scattered light detector 223 depends on the overall size of the particle. The forward-scattered light detector may include, for example, a photodiode. An optical filter 221a and a scattering bar 222 are positioned between the forward-scattered light detectors 223. The optical filter 221a may be configured to remove at least one wavelength of non-FSC light, while the scattering bar 222 may be configured to prevent the incident beam from the laser 201 (i.e., non-scattered light) from being detected by the forward scattered light detector 223.

[0112] Additionally, side-scattered light (SSC) is detected by a side-scattered light detector 224. In other words, the side-scattered light detector 224 is configured to detect refracted and reflected light from the surface and internal structure of the particle 211, which tends to increase as the particle's structure becomes more complex. In the example of FIG. 2, the flow cytometer 200 includes a dichroic mirror 220a configured to reflect SSC light to the side-scattered light detector 224 while passing non-SSC (e.g., fluorescent) light. An optical filter 221b is configured to prevent at least one wavelength of non-SSC light from being detected by the side-scattered light detector 224. Also shown are fluorescence detectors 225a-225c, each configured to detect fluorescence of a different wavelength. For example, the dichroic mirror 220b may be configured to reflect fluorescence (FL) corresponding to a first wavelength (or wavelength range) to the fluorescence detector 225a while passing light of other wavelengths. Optical filter 221c may be configured to prevent at least one wavelength of light that does not correspond to the first wavelength (or wavelength range) from being detected by fluorescence detector 225a. Similarly, dichroic mirror 220c is configured to reflect FL light corresponding to the second wavelength (or wavelength range) to fluorescence detector 225b, while passing light of a third wavelength (or wavelength range) for detection by fluorescence detector 225c. Optical filter 221d is configured to prevent at least one wavelength of light that does not correspond to the second wavelength (or wavelength range) from being detected by fluorescence detector 225b. Furthermore, optical filter 221e is configured to prevent at least one wavelength of light that does not correspond to the third wavelength (or wavelength range) from being detected by fluorescence detector 225c.

[0113] Those skilled in the art will recognize that flow cytometers according to embodiments of the present disclosure are not limited to the flow cytometer shown in FIG. 2 , but may include any flow cytometer known in the art. For example, a flow cytometer may have any number of lasers, beam splitters, filters, and detectors of various wavelengths and in a variety of different configurations. For example, while the embodiment of FIG. 2 shows three fluorescence detectors for illustrative purposes, it will be understood that any suitable number of fluorescence detectors may be used.

[0114] During operation, the operation of the cytometer is controlled by controller / processor 290, and measurement data from the detectors may be stored in memory 295 and processed by controller / processor 290. While not explicitly shown, controller / processor 290 is coupled to the detectors to receive output signals from the detectors, and may also be coupled to electrical and electromechanical components of the flow cytometer to control laser 201, fluid flow parameters, etc. Input / output (I / O) functionality 297 may also be provided within the system. Memory 295, controller / processor 290, and I / O 297 may be provided entirely as an integral part of the flow cytometer. In such embodiments, a display may form part of I / O functionality 297 for presenting experimental data to a user of cytometer 200. Alternatively, memory 295 and controller / processor 290 and some or all of the I / O functionality may be part of one or more external devices, such as a general-purpose computer. In some embodiments, some or all of memory 295 and controller / processor 290 may be in wireless or wired communication with cytometer 200. In conjunction with memory 295 and I / O 297, controller / processor 290 can be configured to perform a variety of functions related to the preparation and analysis of flow cytometer experiments.

[0115] Different fluorescent molecules in a panel of fluorescent dyes used in a flow cytometer experiment emit light in their own characteristic wavelength bands. The particular fluorescent labels used in the experiment and their associated fluorescence emission bands can be selected to generally match the filter window of the detector. I / O 297 can be configured to receive data regarding a flow cytometer experiment having a panel of fluorescent labels and multiple cell populations having multiple markers, with each cell population having a subset of the multiple markers. I / O 297 can also be configured to receive biological data assigning one or more markers to one or more cell populations, marker concentration data, emission spectrum data, data assigning labels to one or more markers, and cytometer configuration data. Flow cytometer experiment data, such as label spectral characteristics and flow cytometer configuration data, can also be stored in memory 295. Controller / processor 290 can be configured to evaluate one or more assignments of labels to markers.

[0116] In some embodiments, the subject system is a particle sorting system configured to sort particles using an enclosed particle sorting module, such as that described in U.S. Patent Application Publication No. 2017 / 0299493, filed March 28, 2017, the disclosure of which is incorporated herein by reference. In certain embodiments, particles (e.g., cells) of a sample are sorted using a sorting determination module having multiple sorting determination units, such as that described in U.S. Patent Application Publication No. 2020 / 0256781, filed December 23, 2019, the disclosure of which is incorporated herein by reference. In some embodiments, a system for sorting components of a sample includes a particle sorting module with deflection plates, such as that described in U.S. Patent Application Publication No. 2017 / 0299493, filed March 28, 2017, the disclosure of which is incorporated herein by reference.

[0117] In certain embodiments, the system is a fluorescence imaging using a radio frequency tagged luminescence imaging enabled particle sorter as shown in FIG. 3. Particle sorter 300 includes a light illumination component 300a including a light source 301 (e.g., a 488 nm laser) that generates an output beam of light 301a that is split into beams 302a and 302b by a beam splitter 302. Light beam 302a propagates through an acousto-optic device (e.g., an acousto-optic deflector, AOD) 303 to generate an output beam 303a having one or more angularly deflected light beams. In some examples, output beam 303a generated from acousto-optic device 303 includes a local oscillator beam and multiple radio frequency comb beams. Light beam 302b propagates through an acousto-optic device (e.g., an acousto-optic deflector, AOD) 304 to generate an output beam 304a having one or more angularly deflected light beams. In some examples, output beam 304a generated from acousto-optic device 304 includes a local oscillator beam and multiple radio frequency comb beams. Output beams 303a and 304a generated from acousto-optical devices 303 and 304, respectively, are combined with beam splitter 305 to generate output beam 305a, which is conveyed through optical component 306 (e.g., an objective lens) to illuminate particles in flow cell 307. In certain embodiments, acousto-optical device 303 (AOD) splits a single laser beam into an array of beamlets, each having a different optical frequency and angle. A second AOD 304 adjusts the optical frequency of a reference beam, which is then overlapped with the array of beamlets at beam combiner 305. In certain embodiments, the light illumination system having a light source and acousto-optical device may also include those described in Schraivogel et al. ("High-speed fluorescence image-enabled cell sorting," Science (2022), 375(6578), 315-320), and U.S. Patent Application Publication No. 2021 / 0404943, the disclosure of which is incorporated herein by reference.

[0118] Output beam 305a illuminates sample particles 308 propagating through flow cell 307 (e.g., with sheath fluid 309) in illumination region 310. As shown in illumination region 310, multiple beams (e.g., angularly deflected, high-frequency shifted optical beams shown as dots across illumination region 310) overlap with a reference local oscillator beam (shown as hatched across illumination region 310). Due to their different optical frequencies, the overlapping beams exhibit beating behavior, whereby each beamlet emits a distinct frequency f 1~n carries a sinusoidal modulation.

[0119] Light from the illuminated sample is conveyed to a light detection system 300b, which includes multiple photodetectors. The light detection system 300b includes a forward-scattered light photodetector 311 for generating a forward-scattered image 311a and a side-scattered light photodetector 312 for generating a side-scattered image 312a. The light detection system 300b also includes a bright-field photodetector 313 for generating a light loss image 313a. In some embodiments, the forward-scattered light detector 311 and the side-scattered light detector 312 are photodiodes (e.g., avalanche photodiodes, APDs). In some examples, the bright-field photodetector 313 is a photomultiplier tube (PMT). Fluorescence from the illuminated sample is also detected by fluorescence photodetectors 314-317. In some examples, the photodetectors 314-317 are photomultiplier tubes. The light from the illuminated sample is directed through a beam splitter 320 to the side-scattered light detection channel 312 and the fluorescence detection channels 314-317. Light detection system 300b includes bandpass optical components 321, 322, 323, and 324 (e.g., dichroic mirrors) for transmitting light of predetermined wavelengths to light detectors 314-317. In some examples, optical component 321 is a 534 nm / 40 nm bandpass. In some examples, optical component 322 is a 586 nm / 42 nm bandpass. In some examples, optical component 323 is a 700 nm / 54 nm bandpass. In some examples, optical component 324 is a 783 nm / 56 nm bandpass. The first number represents the center of the spectral band. The second number indicates the range of the spectral band. Thus, a 510 / 20 filter extends 10 nm on either side of the center of the spectral band, i.e., from 500 nm to 520 nm.

[0120] Data signals generated in response to light detected in scattered light detection channels 311 and 312, bright-field light detection channel 313, and fluorescence detection channels 314-317 are processed by real-time digital processing by processors 350 and 351. Images 311a-317a can be generated in each light detection channel based on the data signals generated by processors 350 and 351. Image-enabled sorting is performed in response to a sorting signal generated by sorting trigger 352. Sorting component 300c includes deflection plates 331 for deflecting particles into a sample container 332 or to a waste stream 333. In some examples, sorting component 300c is configured to sort particles using an enclosed particle sorting module, such as that described in U.S. Patent Application Publication No. 2017 / 0299493, filed March 28, 2017, the disclosure of which is incorporated herein by reference. In certain embodiments, the sorting component 300c includes a sorting determination module having multiple sorting determination units, such as those described in U.S. Patent Application Publication No. 2020 / 0256781, the disclosure of which is incorporated herein by reference.

[0121] In some embodiments, the system is a particle analyzer, and particle analysis system 401 (FIG. 4) can be used to analyze and characterize particles, with or without physically sorting the particles into a collection vessel. FIG. 4 shows a functional block diagram of a particle analysis system for computation-based sample analysis and particle characterization. In some embodiments, particle analysis system 401 is a flow system. Particle analysis system 401 includes a fluidics system 402. Fluidics system 402 can include or be coupled to a sample tube 405 and a moving fluid column within the sample tube through which particles 403 (e.g., cells) of the sample move along a common sample path 409.

[0122] The particle analysis system 401 includes a detection system 404 configured to collect a signal from each particle as it passes through one or more detection stations along a common sample path. The detection stations 408 generally refer to monitoring areas 407 of the common sample path. Detection, in some implementations, may include detecting light or one or more other characteristics of the particles 403 as they pass through the monitoring area 407. FIG. 4 shows one detection station 408 with one monitoring area 407. Some implementations of the particle analysis system 401 may include multiple detection stations. Additionally, some detection stations may monitor more than one area.

[0123] Each signal is assigned a signal value to form a data point for each particle. This data may be referred to as event data, as described above. The data points may be multidimensional data points that include values ​​for each property measured for the particle. The detection system 404 is configured to collect such data points continuously over a first time interval.

[0124] The particle analysis system 401 may also include a control system 406. The control system 406 may include one or more processors, amplitude control circuitry, and / or frequency control circuitry. The illustrated control system may be operatively associated with the fluidics system 402. The control system may be configured to generate a calculated signal frequency for at least a portion of the first time interval based on the Poisson distribution and the number of data points collected by the detection system 404 during the first time interval. The control system 406 may further be configured to generate an experimental signal frequency based on the number of data points in the portion of the first time interval. The control system 406 may further compare the experimental signal frequency to the calculated signal frequency or a predetermined signal frequency.

[0125] 5 shows a functional block diagram of an example particle analyzer control system for analyzing and displaying biological events, such as an analysis controller (e.g., processor) 500. Analysis controller 500 can be configured to implement various processes for controlling the graphical display of biological events.

[0126] The particle analyzer or sorting system 502 can be configured to acquire biological event data. For example, a flow cytometer can generate flow cytometry event data. The particle analyzer 502 can be configured to provide the biological event data to the analysis controller 500. A data communication channel can be included between the particle analyzer or sorting system 502 and the analysis controller 500. The biological event data can be provided to the analysis controller 500 via the data communication channel. The analysis controller 500 can be a processor configured to perform the methods of the present invention, for example, by applying a distance-based classification model to determine a density distinction threshold in a size-based analyte feature space, applying a density-based clustering algorithm to separate the analyte data into high-density clusters and low-density clusters based on the density threshold, and classifying the analyte data based on the high-density clusters and low-density clusters based on the size-based analyte feature space.

[0127] The analysis controller 500 can be configured to receive biological event data from a particle analyzer or sorting system 502. The biological event data received from the particle analyzer or sorting system 502 can include flow cytometry event data. The analysis controller 500 can be configured to provide a graphical display including a first plot of the biological event data on a display device 506. The analysis controller 500 can be further configured to render a region of interest, for example, as a gate around a population of the biological event data shown by the display device 506, overlaid on the first plot. In some embodiments, the gate can be a logical combination of one or more graphical regions of interest depicted on a histogram or bivariate plot of a single parameter. In some embodiments, the display can be used to display particle parameters or saturation detector data.

[0128] Analysis controller 500 can be further configured to display biological event data within the gate on display device 506 differently from other events in the biological event data outside the gate. For example, analysis controller 500 can be configured to render the color of the biological event data contained within the gate differently from the color of the biological event data outside the gate. Display device 506 can be implemented as a monitor, tablet computer, smartphone, or other electronic device configured to present a graphical interface.

[0129] The analysis controller 500 can be configured to receive a gate selection signal identifying a gate from a first input device. For example, the first input device can be implemented as a mouse 510. The mouse 510 can initiate a gate selection signal to the analysis controller 500 that identifies a gate to be displayed on or operated via the display device 506 (e.g., by clicking the desired gate when a cursor is positioned there). In some implementations, the first device can be implemented as a keyboard 508 or other means for providing input signals to the analysis controller 500, such as a touchscreen, a stylus, a photodetector, or a voice recognition system. Some input devices may include multiple input functions. In such implementations, each input function can be considered an input device. For example, as shown in FIG. 5, the mouse 510 can include a right mouse button and a left mouse button, each capable of generating a trigger event.

[0130] The trigger event can cause the analysis controller 500 to change how the data is displayed, what portions of the data are actually displayed on the display device 506, and / or provide input for further processing, such as selecting a population for particle sorting purposes.

[0131] In some embodiments, the analysis controller 500 can be configured to detect when a gate selection is initiated by the mouse 510. The analysis controller 500 can be further configured to automatically modify the visualization of the plot to facilitate the gating process. The modification can be based on a particular distribution of the biological event data received by the analysis controller 500.

[0132] The analysis controller 500 can be connected to a storage device 504. The storage device 504 can be configured to receive and store biological event data from the analysis controller 500. The storage device 504 can also be configured to receive and store flow cytometry event data from the analysis controller 500. The storage device 504 can be further configured to enable retrieval of biological event data, such as flow cytometry event data, by the analysis controller 500.

[0133] The display device 506 can be configured to receive display data from the analysis controller 500. The display data can include a plot of the biological event data and a gate that delineates a section of the plot. The display device 506 can be further configured to modify the information presented according to input received from the analysis controller 500, along with input from the particle analyzer 502, the storage device 504, the keyboard 508, and / or the mouse 510.

[0134] In some implementations, the analysis controller 500 can generate a user interface for receiving exemplary events for sorting. For example, the user interface can include controls for receiving exemplary events or exemplary images. The exemplary events or images or exemplary gates can be provided prior to collection of event data for the sample or based on an initial set of events for a subset of the sample.

[0135] FIG. 6A is a schematic diagram of a particle sorter system 600 (e.g., particle analyzer or sorting system 502) according to one embodiment presented herein. In some embodiments, the particle sorter system 600 is a cell sorter system. As shown in FIG. 6A, a droplet-forming transducer 602 (e.g., a piezoelectric oscillator) is coupled to a fluid conduit 601, which can be coupled to, include, or be a nozzle 603. Within the fluid conduit 601, a sheath fluid 604 hydrodynamically focuses a sample fluid 606 containing particles 609 into a moving fluid column 608 (e.g., a stream). Within the moving fluid column 608, the particles 609 (e.g., cells) are aligned in single file across a monitoring area 611 (e.g., where laser streams intersect) illuminated by an illumination source 612 (e.g., a laser). Vibration of droplet-forming transducer 602 causes moving fluid column 608 to break up into multiple droplets 610 , some of which contain particles 609 .

[0136] During operation, the detection station 614 (e.g., an event detector) identifies when a particle (or cell) of interest crosses the monitoring area 611. The detection station 614 is fed to a timing circuit 628, which in turn feeds a flash charge circuit 630. At a drop breakoff point, signaled by a timed drop delay (Δt), a flash charge can be applied to the moving fluid column 608 so that the droplets of interest carry a charge. The droplets of interest may contain one or more particles or cells to be sorted. The charged droplets can then be sorted by activating a deflection plate (not shown) to deflect the droplets into a collection tube or a container, such as a multi-well or microwell sample plate, and a well or microwell can be associated with the particular droplet of interest. As shown in FIG. 6A, the droplets can be collected in a waste receptacle 638.

[0137] Detection system 616 (e.g., a droplet boundary detector) helps automatically determine the phase of the droplet drive signal when a particle of interest passes through monitoring area 611. An exemplary droplet boundary detector is described in U.S. Patent No. 7,679,039, which is incorporated herein by reference in its entirety. Detection system 616 allows the instrument to accurately calculate the location of each detected particle in the droplet. Detection system 616 can provide amplitude signal 620 and / or phase 618 signals, which then (via amplifier 622) provide to amplitude control circuit 626 and / or frequency control circuit 624. Amplitude control circuit 626 and / or frequency control circuit 624 then control droplet forming transducer 602. Amplitude control circuit 626 and / or frequency control circuit 624 can be included in a control system.

[0138] In some implementations, the sorting electronics (e.g., detection system 616, detection station 614, and processor 640) can be coupled to a memory configured to store detected events and sorting decisions based thereon. The sorting decisions can be included in the particle's event data. In some implementations, the detection system 616 and detection station 614 can be implemented as a single detection unit or can be communicatively coupled such that event measurements can be collected by either the detection system 616 or the detection station 614 and provided to a non-collecting element.

[0139] FIG. 6B is a schematic diagram of a particle sorter system according to one embodiment presented herein. The particle sorter system 600 shown in FIG. 6B includes deflection plates 652 and 654. An electric charge can be applied via stream charging wires within the barbs. This creates a stream of droplets 610 containing particles 609 for analysis. The particles can be illuminated with one or more light sources (e.g., lasers) to generate light scattering and fluorescence information. The information about the particles is analyzed, such as by sorting electronics or another detection system (not shown in FIG. 6B). Deflection plates 652 and 654 can be independently controlled to attract or repel charged droplets, directing them toward a destination collection receptacle (e.g., any of 672, 674, 676, or 678). As shown in FIG. 6B, deflection plates 652 and 654 can be controlled to direct particles along a first path 662 toward receptacle 674 or along a second path 668 toward receptacle 678. If the particle is not of interest (e.g., does not exhibit scattering or illumination information within the specified sorting range), the deflector may allow the particle to continue along flow path 664. Such uncharged droplets may enter a waste receptacle, such as via aspirator 670.

[0140] Sorting electronics may be included to initiate the collection of measurements, receive fluorescent signals about the particles, and determine how to adjust the deflection plates to cause particle sorting. An exemplary implementation of the embodiment shown in Figure 6B includes the BD FACSAria™ line of flow cytometers commercially offered by Becton, Dickinson and Company (Franklin Lakes, NJ).

[0141] FIG. 7 illustrates the general architecture of an exemplary computing device 700 according to certain embodiments. The general architecture of computing device 700 illustrated in FIG. 7 includes an arrangement of computer hardware and software components. However, not all of these generally conventional elements need be shown to constitute an enabling disclosure. As illustrated, computing device 700 includes a processing unit 710, a network interface 720, a computer-readable medium drive 730, an input / output device interface 740, a display 750, and input devices 760, all of which can communicate with each other via a communication bus. Network interface 720 can provide connectivity to one or more networks or computing systems. Thus, processing unit 710 can receive information and instructions from other computing systems or services via a network. Processing unit 710 also communicates with memory 770 and can further communicate output information to optional display 750 via input / output device interface 740. For example, analysis software (e.g., data analysis software or a program such as FlowJo®) stored as executable instructions in the analysis system's non-transitory memory can display flow cytometry event data to a user. The input / output device interface 740 can also accept input from optional input devices 760 such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0142] Memory 770 may include computer program instructions (grouped in some embodiments as modules or components) that processing unit 710 executes to implement one or more embodiments. Memory 770 generally includes RAM, ROM, and / or other persistent, secondary, or non-transitory computer-readable media. Memory 770 may store an operating system 772 that communicates computer program instructions used by processing unit 710 in the general management and operation of computing device 700. Data may be stored in data storage device 790. Memory 770 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0143] Non-transitory computer-readable storage medium Aspects of the present disclosure further include non-transitory computer-readable storage media having instructions for implementing the subject methods. The computer-readable storage medium may be used by one or more computers to fully or partially automate a system for implementing the methods described herein. In certain embodiments, instructions according to the methods described herein may be encoded on a computer-readable medium in the form of "programming," and the term "computer-readable medium" as used herein refers to any non-transitory storage medium involved in providing instructions and data to a computer for execution and processing. Examples of suitable non-transitory storage media include floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, DVD-ROMs, Blu-ray disks, solid-state disks, and network-attached storage (NAS), whether such devices are internal or external to the computer. Files containing information may be "stored" on a computer-readable medium, where "storing" refers to recording information so that it can be accessed and retrieved at a later date by a computer. The computer-implemented methods described herein may be implemented using programming that may be written in one or more of any number of computer programming languages. Such languages ​​include, for example, Python, Java, JavaScript, C, C#, C++, Go, R, Swift, PHP, and many more.

[0144] A non-transitory computer-readable storage medium having instructions with algorithms for classifying analyte data is also described. The non-transitory computer-readable storage medium according to certain embodiments includes an algorithm for applying a distance-based classification model to determine a density distinction threshold in a size-based analyte feature space, an algorithm for applying a density-based clustering algorithm to separate the analyte data into high-density clusters and low-density clusters based on the density threshold, and an algorithm for classifying the analyte data based on the high-density and low-density clusters based on the size-based analyte feature space.

[0145] In some embodiments, the distance-based classification model is a nearest neighbor algorithm. In some embodiments, the distance-based classification model is a nearest neighbor algorithm. In some examples, the density-based clustering algorithm is a density-based spatial clustering of applications with noise (DBSCAN) algorithm. In some examples, the non-transitory computer-readable storage medium includes an algorithm for discarding low-density data clusters. In some examples, the low-density data clusters include one or more multiplets. In some examples, the multiplets are doublets. In some examples, the multiplets are triplets. In some examples, the non-transitory computer-readable storage medium includes an algorithm for distinguishing high-density clusters by distinguishing between debris clusters and singlet clusters. In some examples, the non-transitory computer-readable storage medium includes an algorithm for discarding debris clusters. In some examples, the non-transitory computer-readable storage medium includes an algorithm for distinguishing high-density clusters by ordering multiple singlet clusters against a size-based analyte feature space. In some examples, the size-based analyte feature space includes one or more of a light-loss analyte feature, a major axis moment analyte feature, and a radial moment analyte feature. In certain examples, the size-based analyte feature space includes imaging analyte features. In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Longitudinal Moment (SSC(Imaging)), Radial Moment (FSC), Radial Moment (Light Loss(Imaging)), and Radial Moment (SSC(Imaging)). In some examples, the size-based analyte feature space includes Light Loss(Violet)-A and Light Loss(Violet)-H. In some examples, the size-based analyte feature space includes Light Loss(Violet)-A, Light Loss(Violet)-W, Size(FSC), FSC-A, Radial Moment (Light Loss(Imaging)), Radial Moment (SSC(Imaging)).In some examples, the size-based analyte feature space includes Light Loss (Violet)-A, Minor Axis Moment (Light Loss (Imaging)), SSC(Imaging)-A, Light Loss (Imaging)-A, SSC(Violet)-A. In some examples, the size-based analyte feature space includes FSC-A, FSC-H. In some examples, the size-based analyte feature space includes 2 to 10 analyte features, e.g., 3 to 8 analyte features, e.g., 4 to 6 analyte features. In some examples, the classification model is an unsupervised algorithm. In some embodiments, a non-transitory computer-readable storage medium includes an algorithm for classifying analyte data based solely on feature density and not on identified features of the cells themselves. In some examples, features of the cells are not used for classification.

[0146] In some embodiments, the non-transitory computer-readable storage medium includes an algorithm for training a model to classify analyte data. In some examples, the non-transitory computer-readable storage medium includes an algorithm for training a model, such as having an algorithm for determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data, and an algorithm for predicting a classification of the analyte dataset based on the ground truth analyte data. In some examples, the supervised learning algorithm is a random forest classifier. In some examples, the non-transitory computer-readable storage medium includes an algorithm for discarding predicted classifications below a confidence level and an algorithm for iteratively predicting a classification of the analyte dataset. In some examples, the confidence level is in the range of 60% to 100%, e.g., 70% to 95%, including 80% to 90%. In some embodiments, the non-transitory computer-readable storage medium includes an algorithm for characterizing the classification of the population clusters. In some examples, the non-transitory computer-readable storage medium includes an algorithm for calculating a precision statistic for the classification of the population clusters. In some examples, the non-transitory computer-readable storage medium includes an algorithm for calculating a sensitivity statistic for classification of population clusters.

[0147] The non-transitory computer-readable storage medium may be used in one or more computer systems having a display and an operator input device. The operator input device may be, for example, a keyboard, a mouse, etc. The processing module includes a processor that accesses a memory in which instructions for performing the steps of the subject method are stored. The processing module may include an operating system, a graphical user interface (GUI) controller, a system memory, a memory storage device, and an input / output controller, a cache memory, a data backup unit, and many other devices. The processor may be a commercially available processor or one of other processors that are or become available. The processor executes an operating system, which interfaces with firmware and hardware in well-known manners and facilitates the processor's coordination and execution of the functions of various computer programs, which may be written in various programming languages, such as those mentioned above, other high-level or low-level languages, and combinations thereof, as known in the art. The operating system typically cooperates with the processor to coordinate and execute the functions of the other components of the computer. The operating system also provides scheduling, input / output control, file and data management, memory management, and communication control and related services, all in accordance with known techniques.

[0148] kit Aspects of the present disclosure further include kits, which contain storage media such as magneto-optical disks, CD-ROMs, CD-Rs, magnetic tape, non-volatile memory cards, ROMs, DVD-ROMs, Blu-ray discs, solid-state disks, and network-attached storage (NAS). Some of these program storage media, or others now in use or that may later be developed, may be included in the subject kits. In embodiments, the program storage media contain instructions for classifying flow cytometer data. In embodiments, the instructions contained in the computer-readable media provided in the subject kits, or portions thereof, may be implemented as software components of software for analyzing data. In these embodiments, a computer-controlled system according to the present disclosure may function as a software "plug-in" for an existing software package (e.g., FlowJo®).

[0149] In addition to the above components, the subject kits may further include instructions (in some embodiments). These instructions may be present in a variety of forms, one or more of which may be present in the subject kits. One form in which these instructions may be present is information printed on a suitable medium or substrate, such as one or more pieces of paper on which the information is printed, kit packaging, a package insert, etc. Another form in which these instructions may be present is a computer-readable medium on which the information is recorded, such as a diskette, a compact disc (CD), a portable flash drive, etc. Another form in which these instructions may be present is a website address that can be used via the Internet to access the information at the removed site.

[0150] Utilities The subject particle analyzers, methods, and computer systems find use in a variety of applications where it is desirable to analyze, and optionally separate, particle components in a sample in a fluid medium, such as a biological sample, and then store the separated product for later use, such as in therapeutic applications. The present disclosure finds particular use where it is desirable to classify flow cytometer data. For example, the subject particle analyzers, methods, and computer systems can be used to facilitate the determination of suitable gates for particular populations or subpopulations of flow cytometer data, particularly in data sets where such suitable gates are not readily identifiable. Embodiments of the present disclosure also find use where it is desirable to provide a flow cytometer with improved cell sorting accuracy, enhanced particle collection, particle charging efficiency, accurate particle charging, and enhanced particle deflection during cell sorting.

[0151] Embodiments of the present disclosure find use in applications in which cells prepared from biological samples may be desirable for research, laboratory testing, or therapeutic use. In some embodiments, the subject methods and devices may facilitate obtaining and / or analyzing individual cells prepared from target fluid or tissue biological samples. For example, the subject methods and systems may facilitate obtaining cells from fluid or tissue samples used as research or diagnostic samples for diseases such as cancer. Similarly, the subject methods and systems may facilitate obtaining cells from fluid or tissue samples used in therapeutic use.

[0152] The following are offered by way of example and not by way of limitation.

[0153] experiment Experiment 1 The flow cytometry data were classified into debris, singlet, and multiplet clusters according to the method of interest. A subset of the flow cytometry dataset was manually labeled. In this dataset of 10,000 events, 100 events were labeled based on particle images. A supervised learning algorithm (random forest classifier) ​​was trained on 100 labeled samples, and the trained model was used to predict the labels of the next 100 samples. If the prediction had a confidence level of 90% or less, the event was discarded from the dataset. The supervised learning algorithm was then retrained on the original subset of labeled events and an additional subset. For example, if 7 of the 100 predictions had a confidence level of less than 80%, the additional training set would include the 100 original events and 93 newly labeled high-confidence events, generating a new training set of 193 events. The model trained on 193 events was then used to predict the next 100 event labels, with low-confidence predictions discarded and high-confidence predictions added to the training set. Training the learning algorithm on a larger training set increases the accuracy rate of the trained model with each iteration. This was continued until the amount of low-confidence predictions no longer decreased with each iteration to generate a final trained model. The final trained model was used to predict the classification of the remainder of the dataset. Low-confidence predictions (e.g., less than 90% confidence) were discarded. Figure 8 shows particle label classification for flow cytometry data according to certain embodiments. Clusters of particles were classified based on the imaging parameters Light Loss (Violet)-H and Light Loss (Violet)-A.

[0154] Experiment 2 - SingletSeeker Algorithm for Automated Singlet Identification in Cytometry Manual singlet identification is a subjective and time-consuming process. As discussed above, machine learning can improve singlet identification by utilizing higher dimensional data and automating the process, which, as demonstrated by the present disclosure, is more accurate, objective, and consistent singlet identification.

[0155] method Twenty-two imaging datasets (15 training; 7 testing) of various cell types were used for algorithm development and testing. Images were used to establish ground truth. Five feature sets and the DBSCAN algorithm were tested for precision and sensitivity. Data were collected on a BD FACSDiscover S8 imaging cytometer.

[0156] The programming environment used to develop the singlet seeker algorithm included Python and associated machine learning libraries. Applications included Scikit-Learn, Scipy, Pandas, and Numpy. Density-Based Spatial Clustering of Applications with Noise (DBSCAN) was used as the clustering algorithm.

[0157] Using a semi-supervised iterative approach, ground truth was established based on images associated with each event.

[0158] The feature selection is summarized in Table 2, which included five feature sets selected for testing to best discriminate singlets.

[0159] [Table 2]

[0160] Manual gating comparison - In addition to comparing results to ground truth, an expert cytometrist performed manual gating.

[0161] Clusters were classified as multiplets, singlets, and debris, and clustering was evaluated using precision and sensitivity statistics. Precision and sensitivity statistics were calculated for each feature set. Precision statistics represent the accuracy of the predicted label for the target class.

[0162] TIFF2026004224000006.tif12170

[0163] The sensitivity statistic represents the proportion of the target class that is captured by the prediction.

[0164] TIFF2026004224000007.tif12170

[0165] Algorithm Development Initial development of the algorithm involved determining the separation of singlets and multiplets by cluster density. Figures 9A-9C illustrate the separation of singlets and multiplets using a density-based algorithm, according to certain embodiments. Figure 9A shows images of singlets and multiplets generated based on light loss parameters. As shown in the images, singlets are resolved as individual cells passing through the detection region of the flow stream. Multiplets contain cellular components, such as non-resolvable cells and cellular debris. Figure 9B shows a density plot of the data using the light loss feature set. The light loss feature set is based on the parameters Light Loss (Violet)-A and Light Loss (Violet)-H, and the density plot shows high-variance, low-density clusters and low-variance, high-density clusters. The density-based clustering algorithm, DBSCAN, is used to separate the data into high-density and low-density clusters based on a density threshold. Figure 9C illustrates the separation of data into density-based clusters using the DBSCAN algorithm. The DBSCAN algorithm separates the data based on a density threshold based on a predetermined minimum sample number parameter.

[0166] We developed a two-step process for identifying and classifying singlets. As a first step, we determined a density threshold based on k-nearest neighbor (kNN) to separate high-density and low-density data (singlets and multiplets). Figures 10A-B illustrate determining a density threshold based on k-nearest neighbor (kNN) in accordance with certain embodiments. Figure 10A shows a histogram of k-distances for event data in a dataset. The data in Figure 10A indicates that there are a large number of data points with k-distances in the range of 0.05 to 1.5, and a subset of data points with k-distances in the range of 0.25 to 0.35. Figure 10B shows a Gaussian-smoothed histogram of k-distances. As can be seen from the histogram plot, there is a first peak of high-density data with k-distances in the range of 0.05 to 1.5 (peaking at approximately 0.1), and a second peak set at approximately 0.3. The k-distance histogram exhibits a bilinear shape with a minimum between the peaks having an ε value of approximately 0.2. Figures 11A and 11B show clustering of data using a light loss feature set using a density-based DBSCAN algorithm. Figure 11A shows the original data clustered based on the analyte features (Light Loss (Violet)-A, Light Loss (Violet)-H). Figure 11B shows the high-density data clustered based on the analyte features (Light Loss (Violet)-A, Light Loss (Violet)-H). The DBSCAN algorithm removes low-density multiplet data using a selected density threshold.

[0167] The second step in identifying and classifying singlets was to determine an appropriate size threshold by finding the minimum value of the forward scatter / light loss parameter. Singlets and debris were classified based on their location relative to the minimum. Figure 12A and 12B show the cluster distance to the high-density data to separate singlets from debris. Figure 12A shows a histogram of the distance from the origin. The histogram shows two peaks in the high-density data: the first peak corresponds to cellular debris in the sample, and the second, larger peak corresponds to singlets in the sample. Figure 12B shows a Gaussian-smoothed histogram of the cluster distance. As shown in the histogram, the first, smaller peak corresponds to debris, and the second, larger peak corresponds to singlets. The minimum between singlets and debris is indicated at a distance of approximately 1.5. Figure 13A and 13B show the clustering of the high-density data using the light loss feature set. Figure 13A shows two separate data clusters corresponding to debris and singlets. Figure 13B shows the classification of the two clusters into debris clusters and singlet clusters. As shown in Figure 13, the developed algorithm successfully classifies singlets and removes debris and multiplets based on a density threshold.

[0168] Figures 14A-C show calculated precision and sensitivity statistics for different feature sets according to certain embodiments. Figure 14A shows calculated singlet classification precision and sensitivity statistics. Singlets were classified using five different feature sets: light loss, recursive feature elimination (RFE), random forest classifier (RFC), sequential feature selection (SFS), and forward scatter (FSC), and precision and sensitivity statistics were calculated. Figure 14B shows precision and sensitivity statistics calculated for multiplet classification using different feature sets. Figure 14C shows precision and sensitivity statistics calculated for debris classification using different feature sets. Precision was approximately comparable for singlet classification using different feature sets, but sensitivity for singlet classification was greatest for the RFC feature set. For multiplet classification, the RFC feature set showed the greatest precision statistic, and the light loss feature set showed the greatest sensitivity statistic. Similar to singlet classification, debris classification showed similar precision statistics with light loss having the largest precision statistic. Regarding the sensitivity statistic for debris classification, the SFS feature set showed the highest sensitivity.

[0169] result The developed feature sets provide high precision and sensitivity across the feature sets. Figure 15A and B show a comparison of precision and sensitivity of different feature sets for training and test data. Figure 15A shows precision and sensitivity statistics for the training data for five different feature sets. For the training data, light loss provided the best precision (approximately 99%), and RFC provided the best recall (approximately 91%). A summary of the precision and sensitivity of the training data for each feature set is shown in Table 3.

[0170] [Table 3]

[0171] Figure 15B shows the precision and sensitivity statistics for the test data for the five different feature sets. For the training data, light loss provided the best precision (approximately 98%), and SFS provided the best recall (approximately 87%). A summary of the precision and sensitivity for the test data by feature set is shown in Table 4.

[0172] [Table 4]

[0173] The developed feature sets provide high precision and sensitivity across the feature sets. Figure 15A and B show a comparison of precision and sensitivity of different feature sets for training and test data. Figure 15A shows precision and sensitivity statistics for the training data for five different feature sets. For the training data, light loss provided the best precision (approximately 99%), and RFC provided the best recall (approximately 91%). A summary of the precision and sensitivity of the training data for each feature set is shown in Table 3.

[0174] Figure 16 shows a summary of singlet classification precision and sensitivity for different feature sets compared to manual gating according to certain embodiments. Manual gating and automatic gating yield similar precision and sensitivity across all feature sets. Precision statistics were similar for each of the different feature sets, while sensitivity was greatest for the optical loss and RFE feature sets.

[0175] Consideration Singlets and multiplets can be successfully distinguished using an unsupervised, density-based algorithm. While all feature sets yield high precision with slightly lower sensitivity, using only the light loss parameter yields the best precision overall, while RFE appears to offer a good balance between precision and sensitivity. The density factor is a useful user input that allows for tuning of the balance between precision and sensitivity. The algorithm yields results comparable to manual gating with better reproducibility and less run-to-run variability. As an unsupervised algorithm, it should work for any cell type as is, based on detecting differences in feature density rather than identifying features of the cells themselves. The algorithm can facilitate more objective, consistent, and user-independent singlet gating.

[0176] Notwithstanding the scope of the appended claims, the present disclosure is also defined by the following clauses. 1. A computer-implemented method for classifying analyte data, comprising: applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space; applying a density-based clustering algorithm to separate the analyte data into high-density and low-density clusters based on a density threshold; Classifying the analyte data based on high-density clusters and low-density clusters based on a size-based analyte feature space; 11. A computer-implemented method comprising: 2. The computer-implemented method of clause 1, wherein the distance-based classification model is a nearest neighbor algorithm. 3. The computer-implemented method of clause 1 or 2, wherein the density-based clustering algorithm is a density-based spatial clustering of applications with noise (DBSCAN) algorithm. 4. The computer-implemented method of any one of clauses 1-3, further comprising discarding low-density data clusters. 5. The computer-implemented method of any one of clauses 1-4, wherein the analyte data is flow cytometer data. 6. The computer-implemented method of clause 5, wherein the sparse clusters are composed of multiplets. 7. The computer-implemented method of clause 6, wherein the multiplet is composed of doublets. 8. The computer-implemented method of clause 6 or 7, wherein the multiplet is composed of triplets. 9. The computer-implemented method of any one of clauses 5 to 8, wherein the applied density-based clustering algorithm further distinguishes high-density clusters between debris clusters and singlet clusters. 10. The computer-implemented method of clause 9, further comprising discarding the debris clusters. 11. The computer-implemented method of clause 5, wherein the applied density-based clustering algorithm further distinguishes high-density clusters by ordering the multiple singlet clusters with respect to a size-based analyte feature space. 12. The computer-implemented method of any one of clauses 5-11, wherein the size-based analyte feature space includes one or more of a light loss analyte feature, a major axis moment analyte feature, and a radial moment analyte feature. 13. The computer-implemented method of clause 12, wherein the size-based analyte feature space includes imaging analyte features. 14. The computer-implemented method of clause 13, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Major Axis Moment (SSC(Imaging)), Radial Moment (FSC), Radial Moment (Light Loss (Imaging)), and Radial Moment (SSC(Imaging)). 15. The computer-implemented method of clause 13, wherein the size-based analyte feature space comprises Light Loss (Violet)-A and Light Loss (Violet)-H. 16. The computer-implemented method of clause 13, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Light Loss (Violet)-W, Size (FSC), FSC-A, Radial Moment (Light Loss (Imaging)), Radial Moment (SSC (Imaging)). 17. The computer-implemented method of clause 13, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Minor Axis Moment (Light Loss (Imaging)), SSC(Imaging)-A, Light Loss (Imaging)-A, SSC(Violet)-A. 18. The computer-implemented method of clause 13, wherein the size-based analyte feature space comprises FSC-A, FSC-H. 19. The computer-implemented method of any one of clauses 5 to 18, wherein the applied density-based clustering algorithm further distinguishes high-density clusters for forward scatter (FSC) analyte features within the size-based analyte feature space. 20. The computer-implemented method of any one of clauses 1-19, wherein the size-based analyte feature space consists of 2 to 10 analyte features. 21. The computer-implemented method of clause 20, wherein the size-based analyte feature space consists of 3 to 8 analyte features. 22. The computer-implemented method of clause 21, wherein the size-based analyte feature space consists of 4 to 6 analyte features. 23. The computer-implemented method of any one of clauses 1-22, further comprising training a model to classify the analyte data. 24. Training a model is determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data; Predicting a classification of an analyte dataset based on ground truth analyte data; 24. The computer-implemented method of claim 23, comprising: 25. The computer-implemented method of clause 24, wherein the supervised learning algorithm is a random forest classifier. 26. Discarding predicted classifications below a confidence level; and Repeated prediction of the classification of the analyte dataset; 26. The computer-implemented method of clause 24 or 25, further comprising: 27. The computer-implemented method of clause 26, wherein the confidence level is in the range of 60% to 100%. 28. The computer-implemented method of clause 27, wherein the confidence level is in the range of 70% to 95%. 29. The computer-implemented method of clause 28, wherein the confidence level is in the range of 80% to 90%. 30. The computer-implemented method of any one of clauses 1-29, further comprising calculating a precision statistic for the classification of the analyte clusters. 31. The computer-implemented method of any one of clauses 1-30, further comprising calculating a sensitivity statistic for the classification of analyte clusters. 32. A system comprising a memory operatively coupled to a processor, the memory storing instructions that, when executed by the processor, cause the processor to: applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space; applying a density-based clustering algorithm to separate the analyte data into high-density and low-density clusters based on a density threshold; Classifying the analyte data based on high-density clusters and low-density clusters based on a size-based analyte feature space; A system that allows the following to be performed. 33. The system of clause 32, wherein the distance-based classification model is a nearest neighbor algorithm. 34. The system of clause 32 or 33, wherein the density-based clustering algorithm is a density-based spatial clustering for applications with noise (DBSCAN) algorithm. 35. The system of any one of clauses 32-34, wherein the memory includes instructions for discarding low-density data clusters. 36. The system of any one of clauses 32 to 35, wherein the analyte data is flow cytometer data. 37. The system of clause 36, wherein the low-density clusters are composed of multiplets. 38. The system of clause 37, wherein the multiplet is composed of doublets. 39. A system according to clause 37 or 38, wherein the multiplet is composed of triplets. 40. The system of any one of clauses 36-39, wherein the memory includes instructions for distinguishing high density clusters by distinguishing between debris clusters and singlet clusters. 41. The system of clause 40, wherein the memory includes instructions for discarding debris clusters. 42. The system of clause 36, wherein the memory includes instructions for distinguishing high-density clusters by ordering a plurality of singlet clusters with respect to a size-based analyte feature space. 43. The system of any one of clauses 36-42, wherein the size-based analyte feature space includes one or more of a light loss analyte feature, a major axis moment analyte feature, and a radial moment analyte feature. 44. The system of clause 43, wherein the size-based analyte feature space includes imaging analyte features. 45. The system of clause 44, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Major Axis Moment (SSC(Imaging)), Radial Moment (FSC), Radial Moment (Light Loss (Imaging)), and Radial Moment (SSC(Imaging)). 46. ​​The system of clause 44, wherein the size-based analyte feature space comprises Light Loss (Violet)-A and Light Loss (Violet)-H. 47. The system of clause 44, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Light Loss (Violet)-W, Size (FSC), FSC-A, Radial Moment (Light Loss (Imaging)), Radial Moment (SSC (Imaging)). 48. The system of clause 44, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Minor Axis Moment (Light Loss (Imaging)), SSC(Imaging)-A, Light Loss (Imaging)-A, SSC(Violet)-A. 49. The system of clause 44, wherein the size-based analyte feature space includes FSC-A, FSC-H. 50. The system of any one of clauses 36-45, wherein the memory includes instructions for distinguishing high density clusters for forward scatter (FSC) analyte features within a size-based analyte feature space. 51. The system of any one of clauses 32-50, wherein the size-based analyte feature space consists of 2-10 analyte features. 52. The system of clause 51, wherein the size-based analyte feature space consists of 3 to 8 analyte features. 53. The system of clause 52, wherein the size-based analyte feature space consists of 4 to 6 analyte features. 54. The system of any one of clauses 32-53, wherein the memory includes instructions for training a model to classify analyte data. 55. Training a model is determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data; Predicting a classification of an analyte dataset based on ground truth analyte data; 54. The system of claim 54, comprising: 56. The system of clause 55, wherein the supervised learning algorithm is a random forest classifier. 57. The system of clause 55 or 56, wherein the memory includes instructions for discarding predicted classifications that fall below a confidence level and repeating the prediction of classifications for the analyte dataset. 58. The system of clause 57, wherein the confidence level is in the range of 60% to 100%. 59. The system of clause 58, wherein the confidence level is in the range of 70% to 95%. 60. The system of clause 59, wherein the confidence level is in the range of 80% to 90%. 61. The system of any one of clauses 32-60, wherein the memory further comprises instructions for calculating precision statistics for the classification of analyte clusters. 62. The system of any one of clauses 32-61, wherein the memory further comprises instructions for calculating sensitivity statistics for the classification of analyte clusters. 63. The system of any one of clauses 32-62, further comprising a display configured to output the classified analyte data. 64. A system according to any one of clauses 32 to 63, wherein the processor is operably coupled to a flow cytometer. 65. A non-transitory computer-readable storage medium having stored thereon instructions for classifying analyte data, comprising: an algorithm for applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space; applying a density-based clustering algorithm to separate the analyte data into high-density and low-density clusters based on a density threshold; and an algorithm for classifying analyte data based on high-density and low-density clusters based on a size-based analyte feature space. 1. A non-transitory computer-readable storage medium comprising: 66. The non-transitory computer-readable storage medium of clause 65, wherein the distance-based classification model is a nearest neighbor algorithm. 67. The non-transitory computer-readable storage medium of clause 65 or 66, wherein the density-based clustering algorithm is a density-based spatial clustering for applications with noise (DBSCAN) algorithm. 68. A non-transitory computer-readable storage medium according to any one of clauses 65 to 67, comprising an algorithm for discarding low-density data clusters. 69. The non-transitory computer-readable storage medium of any one of clauses 65 to 68, wherein the analyte data is flow cytometer data. 70. The non-transitory computer-readable storage medium of clause 69, wherein the low-density clusters are composed of multiplets. 71. The non-transitory computer-readable storage medium of clause 70, wherein the multiplet is comprised of doublets. 72. The non-transitory computer-readable storage medium of clause 70 or 71, wherein the multiplet is composed of triplets. 73. The non-transitory computer-readable storage medium of any one of clauses 69-72, further comprising an algorithm for distinguishing high-density clusters by distinguishing between debris clusters and singlet clusters. 74. The non-transitory computer-readable storage medium of clause 73, further comprising an algorithm for discarding debris clusters. 75. The non-transitory computer-readable storage medium of claim 69, comprising an algorithm for distinguishing high-density clusters by ordering multiple singlet clusters with respect to a size-based analyte feature space. 76. The non-transitory computer-readable storage medium of any one of clauses 69-75, wherein the size-based analyte feature space includes one or more of a light loss analyte feature, a major axis moment analyte feature, and a radial moment analyte feature. 77. The non-transitory computer-readable storage medium of clause 76, wherein the size-based analyte feature space includes imaging analyte features. 78. The non-transitory computer-readable storage medium of clause 77, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Major Axis Moment (SSC(Imaging)), Radial Moment (FSC), Radial Moment (Light Loss (Imaging)), and Radial Moment (SSC(Imaging)). 79. The non-transitory computer-readable storage medium of clause 77, wherein the size-based analyte feature space comprises Light Loss (Violet)-A and Light Loss (Violet)-H. 80. The non-transitory computer-readable storage medium of clause 77, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Light Loss (Violet)-W, Size (FSC), FSC-A, Radial Moment (Light Loss (Imaging)), Radial Moment (SSC (Imaging)). 81. The non-transitory computer-readable storage medium of clause 77, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Minor Axis Moment (Light Loss (Imaging)), SSC(Imaging)-A, Light Loss (Imaging)-A, SSC(Violet)-A. 82. The non-transitory computer-readable storage medium of clause 77, wherein the size-based analyte feature space includes FSC-A, FSC-H. 83. The non-transitory computer-readable storage medium of any one of clauses 69-82, further comprising an algorithm for distinguishing high-density clusters for forward scatter (FSC) analyte features within a size-based analyte feature space. 84. The non-transitory computer-readable storage medium of any one of clauses 65 to 83, wherein the size-based analyte feature space is composed of 2 to 10 analyte features. 85. The non-transitory computer-readable storage medium of clause 84, wherein the size-based analyte feature space is comprised of between three and eight analyte features. 86. The non-transitory computer-readable storage medium of clause 85, wherein the size-based analyte feature space is comprised of four to six analyte features. 87. The non-transitory computer-readable storage medium of any one of clauses 65 to 86, further comprising an algorithm for training a model to classify analyte data. 88. Determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data; and Predicting classification of analyte dataset based on ground truth analyte data 88. A non-transitory computer-readable storage medium as described in clause 87, comprising an algorithm for training a model by 89. The non-transitory computer-readable storage medium of clause 88, wherein the supervised learning algorithm is a random forest classifier. 90. The non-transitory computer-readable storage medium of clause 88 or 89, further comprising an algorithm for discarding predicted classifications below a confidence level and repeating the prediction of classifications for the analyte dataset. 91. The non-transitory computer-readable storage medium of clause 90, wherein the confidence level is in the range of 60% to 100%. 92. The non-transitory computer-readable storage medium of clause 91, wherein the confidence level is in the range of 70% to 95%. 93. The non-transitory computer-readable storage medium of clause 92, wherein the confidence level is in the range of 80% to 90%. 94. The non-transitory computer-readable storage medium of any one of clauses 65-93, further comprising an algorithm for calculating precision statistics for classification of analyte clusters. 95. The non-transitory computer-readable storage medium of any one of clauses 65 to 93, further comprising an algorithm for calculating sensitivity statistics for the classification of analyte clusters. 96. A computer-implemented method for classifying analyte data, comprising: (a) providing analyte data to a system having instructions stored thereon, the instructions, when executed by a processor, causing the processor to: applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space; applying a density-based clustering algorithm to separate the analyte data into high-density and low-density clusters based on a density threshold; and Classifying analyte data based on high-density clusters and low-density clusters based on a size-based analyte feature space and (b) receiving the classified analyte data from the processor; 11. A computer-implemented method comprising: 97. The computer-implemented method of clause 96, wherein the distance-based classification model is a nearest neighbor algorithm. 98. The computer-implemented method of clause 96 or 97, wherein the density-based clustering algorithm is a density-based spatial clustering for applications with noise (DBSCAN) algorithm. 99. The computer-implemented method of any one of clauses 96-98, wherein the memory includes instructions for discarding sparse data clusters. 100. The computer-implemented method of any one of clauses 96-99, wherein the analyte data is flow cytometer data. 101. The computer-implemented method of clause 100, wherein the sparse clusters are composed of multiplets. 102. The computer-implemented method of clause 101, wherein the multiplet is composed of doublets. 103. The computer-implemented method of clause 101 or 102, wherein the multiplet is composed of triplets. 104. The computer-implemented method of any one of clauses 100-103, wherein the memory includes instructions for distinguishing high-density clusters by distinguishing between debris clusters and singlet clusters. 105. The computer-implemented method of clause 104, wherein the memory includes instructions for discarding debris clusters. 106. The computer-implemented method of any one of clauses 100-103, wherein the memory includes instructions for distinguishing high-density clusters by ordering a plurality of singlet clusters with respect to a size-based analyte feature space. 107. The computer-implemented method of any one of clauses 100-106, wherein the size-based analyte feature space includes one or more of a light loss analyte feature, a major axis moment analyte feature, and a radial moment analyte feature. 108. The computer-implemented method of clause 107, wherein the size-based analyte feature space includes imaging analyte features. 109. The computer-implemented method of clause 108, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Major Axis Moment (SSC(Imaging)), Radial Moment (FSC), Radial Moment (Light Loss (Imaging)), and Radial Moment (SSC(Imaging)). 110. The computer-implemented method of clause 108, wherein the size-based analyte feature space comprises Light Loss (Violet)-A and Light Loss (Violet)-H. 111. The computer-implemented method of clause 108, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Light Loss (Violet)-W, Size (FSC), FSC-A, Radial Moment (Light Loss (Imaging)), Radial Moment (SSC (Imaging)). 112. The computer-implemented method of clause 108, wherein the size-based analyte feature space comprises Light Loss (Violet)-A, Minor Axis Moment (Light Loss (Imaging)), SSC(Imaging)-A, Light Loss (Imaging)-A, SSC(Violet)-A. 113. The computer-implemented method of clause 108, wherein the size-based analyte feature space comprises FSC-A, FSC-H. 114. The computer-implemented method of any one of clauses 100-113, wherein the memory includes instructions for distinguishing high-density clusters for forward scatter (FSC) analyte features within a size-based analyte feature space. 115. The computer-implemented method of any one of clauses 96 to 114, wherein the size-based analyte feature space consists of 2 to 10 analyte features. 116. The computer-implemented method of clause 115, wherein the size-based analyte feature space consists of 3 to 8 analyte features. 117. The computer-implemented method of clause 116, wherein the size-based analyte feature space consists of four to six analyte features. 118. The computer-implemented method of any one of clauses 96-117, further comprising training a model to classify the analyte data. 119. The computer-implemented method of clause 118, wherein training the model includes determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data, and predicting a classification of the analyte dataset based on the ground truth analyte data. 120. The computer-implemented method of clause 119, wherein the supervised learning algorithm is a random forest classifier. 121. The computer-implemented method of clause 119 or 120, wherein the processor is configured to discard predicted classifications that fall below a confidence level and repeat predicting the classification of the analyte dataset. 122. The computer-implemented method of clause 121, wherein the confidence level is in the range of 60% to 100%. 123. The computer-implemented method of clause 122, wherein the confidence level is in the range of 70% to 95%. 124. The computer-implemented method of clause 123, wherein the confidence level is in the range of 80% to 90%. 125. The computer-implemented method of any one of clauses 96-124, further comprising calculating a precision statistic for the classification of the analyte clusters. 126. The computer-implemented method of any one of clauses 96-125, further comprising calculating a sensitivity statistic for the classification of analyte clusters. 127. The computer-implemented method of any one of clauses 96-102, comprising receiving the classified analyte data on a display. 128. The computer-implemented method of any one of clauses 96-127, wherein the processor is operably coupled to a flow cytometer.

[0177] Although the foregoing disclosure has been described in some detail by way of illustration and example for clarity of understanding, it will be readily apparent to those skilled in the art that certain changes and modifications can be made in light of the teachings of the present invention without departing from the spirit or scope of the appended claims.

[0178] Thus, the foregoing merely illustrates the principles of the present disclosure. It will be appreciated that those skilled in the art will be able to devise various configurations, not explicitly described or shown herein, that embody the principles of the present disclosure and are within its spirit and scope. Furthermore, all examples and conditional language recited herein are intended primarily to aid the reader in understanding the principles of the present disclosure and the concepts that the present disclosure has contributed to advancing the art, and should not be construed as being limited to such specifically recited examples and conditions. Furthermore, all statements herein reciting principles, aspects, and embodiments of the present disclosure, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Furthermore, such equivalents are intended to include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. Furthermore, nothing disclosed herein is intended as a dedication to the public, regardless of whether such disclosure is expressly recited in the claims.

[0179] Accordingly, the scope of the present disclosure is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the present disclosure are embodied by the appended claims. In the claims, 35 U.S.C. §112(f) or 35 U.S.C. §112(6) are expressly defined as being invoked for a limitation in a claim only if the exact phrase "means for" or the exact phrase "step" appears at the beginning of such limitation in the claim. If such exact phrases are not used in a claim limitation, 35 U.S.C. §112(f) or 35 U.S.C. §112(6) is not invoked.

[0180] CROSS-REFERENCE TO RELATED APPLICATIONS Pursuant to 35 U.S.C. § 119(e), this application claims priority to the filing date of U.S. Provisional Patent Application No. 63 / 654,726, filed May 31, 2024, the disclosure of which is incorporated herein by reference in its entirety.

Claims

1. 1. A computer-implemented method for classifying analyte data, comprising, via a processor: applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space; applying a density-based clustering algorithm to separate the analyte data into high-density and low-density clusters based on the density threshold; classifying the analyte data based on the high-density clusters and the low-density clusters based on the size-based analyte feature space; 11. A computer-implemented method comprising:

2. The computer-implemented method of claim 1 , wherein the density-based clustering algorithm is a Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm.

3. The computer-implemented method of claim 1 or 2, further comprising discarding sparse data clusters that contain multiplets.

4. The computer-implemented method of claim 1 , wherein the applied density-based clustering algorithm further distinguishes the high-density clusters between debris clusters and singlet clusters.

5. The computer-implemented method of claim 4 , wherein the applied density-based clustering algorithm further distinguishes high-density clusters by ordering a plurality of singlet clusters with respect to the size-based analyte feature space.

6. 6. The computer-implemented method of claim 1, wherein the size-based analyte feature space comprises between 2 and 10 analyte features, and the size-based analyte feature space comprises one or more of a light loss analyte feature, a major axis moment analyte feature, and a radial moment analyte feature.

7. the size-based analyte feature space comprising: a) Light Loss (Violet) - A, Longitudinal Moment (SSC(Imaging)), Radial Moment (FSC), Radial Moment (Light Loss(Imaging)), and Radial Moment (SSC(Imaging)); b) Light Loss (Violet)-A and Light Loss (Violet)-H; c) Light Loss (Violet)-A, Light Loss (Violet)-W, Size (FSC), FSC-A, Radial Moment (Light Loss (Imaging)), Radial Moment (SSC (Imaging)), d) Light Loss (Violet)-A, Minor Axis Moment (Light Loss (Imaging)), SSC (Imaging)-A, Light Loss (Imaging)-A, SSC (Violet)-A, and e) FSC-A, FSC-H The computer-implemented method of claim 6 , comprising one or more of:

8. further comprising training a model to classify the analyte data; training the model, determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data; predicting a classification of an analyte dataset based on the ground truth analyte data; The computer-implemented method of claim 1 , comprising:

9. Discarding predicted classifications that fall below a confidence level; and repeating the prediction of the classification of the analyte dataset; The computer-implemented method of claim 8 further comprising:

10. 1. A system comprising a memory operatively coupled to a processor, the memory having instructions stored therein that, when executed by the processor, cause the processor to: applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space; applying a density-based clustering algorithm to separate the analyte data into high-density and low-density clusters based on the density threshold; classifying the analyte data based on the high-density clusters and the low-density clusters based on the size-based analyte feature space; A system that allows the following to be performed.

11. The system of claim 10 , wherein the density-based clustering algorithm is a density-based spatial clustering of applications with noise (DBSCAN) algorithm.

12. the memory includes instructions for training a model to classify the analyte data; training the model, determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data; predicting a classification of an analyte dataset based on the ground truth analyte data; 12. The system of claim 10 or 11, comprising:

13. A non-transitory computer-readable storage medium having stored thereon instructions for classifying analyte data, the non-transitory computer-readable storage medium comprising: an algorithm for applying a distance-based classification model to determine a density discrimination threshold in a size-based analyte feature space; applying a density-based clustering algorithm to separate the analyte data into high-density and low-density clusters based on the density threshold; an algorithm for classifying the analyte data based on the high-density clusters and the low-density clusters based on the size-based analyte feature space; 1. A non-transitory computer-readable storage medium comprising:

14. 14. The non-transitory computer-readable storage medium of claim 13, wherein the density-based clustering algorithm is a Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm.

15. further comprising an algorithm for training a model to classify the analyte data; determining ground truth analyte data by training a supervised learning algorithm with manually labeled analyte data; and predicting a classification of an analyte dataset based on the ground truth analyte data; 15. The non-transitory computer-readable storage medium of claim 13 or 14, comprising an algorithm for training the model by: