Methods and systems for classifying fluorescent flow cytometry data

By using supervised algorithms and adjusting the spillover diffusion matrix, the problem of spillover diffusion noise in flow cytometry data analysis was solved, resulting in more accurate data classification and phenotypic identification.

CN115335680BActive Publication Date: 2026-05-01BECTON DICKINSON & CO
View PDF 44 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BECTON DICKINSON & CO
Filing Date
2021-01-20
Publication Date
2026-05-01

Smart Images

  • Figure CN115335680B_ABST
    Figure CN115335680B_ABST
Patent Text Reader

Abstract

The invention provides methods for classifying fluorescent flow cytometry data. In some cases, the methods include processing the flow cytometry data with a supervised algorithm configured to cluster the fluorescent flow cytometry data into different populations based on the relationship of data points to associated thresholds. In embodiments, the methods include calculating and combining spill-over diffusion coefficients in a spill-over diffusion matrix to determine the extent of spill-over diffusion. In some embodiments, the fluorescent flow cytometry data populations are adjusted to reduce the extent of spill-over diffusion. In embodiments, the spill-over diffusion adjusted populations are partitioned after evaluating potential partitions relative to the thresholds. In embodiments, the partitioned fluorescent flow cytometry data populations are classified (i.e., phenotyped) according to a hierarchy. The invention also provides systems and computer readable media for classifying fluorescent flow cytometry data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] Pursuant to Section 119(e) of Chapter 35 of the United States Code, this patent application claims priority over the filing dates of U.S. Provisional Patent Application No. 62 / 968,516, filed January 31, 2019, and U.S. Provisional Patent Application No. 63 / 053,108, filed July 17, 2019, the contents of which are incorporated herein by reference. Background Technology

[0003] Flow cytometry is a technique used to characterize and typically sort biological materials, such as cells in a blood sample or target particles in another biological or chemical sample. A flow cytometer typically includes a sample container for receiving a fluid sample (e.g., blood) and a sheath fluid container. The flow cytometer transports particles (including cells) in the fluid sample to a flow cell as a flow stream, while simultaneously guiding the sheath fluid into the flow cell. To characterize the components of the fluid medium, the fluid medium is irradiated with light. Changes in the material in the fluid medium (e.g., morphology or the presence of fluorescent markers) may cause changes in the observed light, which can be used for characterization and separation. For example, particles (e.g., molecules, microbeads bound to analytes, or individual cells) in a fluid suspension pass through a detection zone where the particles are exposed to excitation light, typically from one or more lasers, and the light scattering and fluorescence properties of the particles are measured. Particles or their components are typically labeled with fluorescent dyes for detection. Various different particles or components can be detected simultaneously by labeling them with fluorescent dyes with different spectral characteristics. In some embodiments, the analyzer includes multiple detectors, one for each scattering parameter to be measured, and one or more for each different dye to be detected. For example, some embodiments include spectral structures in which more than one sensor or detector is used for each dye. The obtained data includes the signal measured for each of the light scattering detector and the fluorescence emission.

[0004] Flow cytometers may further include means for recording and analyzing the measured data. For example, a computer connected to the detection electronics can be used for data storage and analysis. For instance, the data may be stored in list form, where each row corresponds to the data for one particle, and the columns correspond to each measured feature. Storing data from the particle analyzer using a standard file format (e.g., the “FCS” file format) facilitates data analysis using separate procedures and / or machines. When using current analytical methods, the data is typically presented as a one-dimensional histogram or a two-dimensional (2D) graph for visualization purposes, but other methods can also be used to visualize multidimensional data.

[0005] For example, parameters measured using flow cytometry typically include light scattered primarily forward at a narrow angle by the particle at the excitation wavelength (referred to as forward scattering (FSC)); the excitation light scattered by the particle in a direction orthogonal to the excitation laser (referred to as side scattering (SSC)); and light emitted by fluorescent molecules in one or more detectors (for measuring signals within a certain spectral wavelength range), or by fluorescent dyes primarily detected in said specific detectors or detector arrays. Different cell types can be identified by the light scattering characteristics and fluorescence emission obtained / produced by labeling various cellular proteins or other components with antibodies or other fluorescent probes labeled with fluorescent dyes.

[0006] Flow cytometers and scanning cytometers are available from sources such as BD Biosciences (San Jose, California). Descriptions of flow cytometry can be found in the following publications, for example: Landy et al. (eds.), Clinical Flow Cytometry, New York Academy of Sciences Annals, Vol. 677 (1993); Bauer et al. (eds.), Clinical Flow Cytometry: Principles and Applications, Williams & Wilkins (1993); Ormerod (ed.), Flow Cytometry: Practical Methods, Oxford University Press (1994); Jaroszeski et al. (eds.), Flow Cytometry Protocols, Methods in Molecular Biology, Vol. 91, Humana Press (1997); and Shapiro, Practical Flow Cytometry, 4th Edition, Wiley-Liss (2003); all content in these publications is incorporated herein by reference. For a description of fluorescence imaging microscopy, please see the following publications, for example, Pawley (ed.), Handbook of Biological Confocal Microscopy, 2nd Edition, Plenum Press (1989), which is incorporated herein by reference.

[0007] After receiving flow cytometry data from one or more detectors, a data analysis process is typically performed for user understanding. In some cases, flow cytometry data analysis involves determining the phenotype associated with the analyte (e.g., cells, particles) that was irradiated in the flow cytometer. For most samples in all cell counting methods, it is necessary to identify a subset of cell types (i.e., phenotypes) before proceeding with any further analysis. In a typical workflow, a pre-planned gating hierarchy (which summarizes generally accepted cell type definitions) must be manually adjusted to accommodate the data for each sample (to account for biological effects, batch effects, and any other characteristics). This adjustment can be done using either an "unsupervised" or "supervised" approach. When using an unsupervised approach, the flow cytometry data are first clustered into populations, and then, using some prior knowledge about the association between specific analytes and certain parameters (e.g., fluorescence), an attempt is made to assign cell type labels (i.e., phenotypes) to each population cluster. On the other hand, when using supervised methods, a gating hierarchy is typically established first, i.e., a set of rules governing how individual cells are classified (i.e., phenotypic analysis) based on their state relative to one or more parameters. Flow cytometry data is then “fitted” to the gating hierarchy such that each data point falls somewhere within the hierarchy. For example, the DAFi algorithm can fit event clusters to an existing gating tree based on event centroids to classify cell types along the natural boundaries of the data rather than fixed gate boundaries. A description of the DAFi algorithm can be found, for example, Lee et al. (2018). DAFI: A Direct Recursive Data Filtering and Clustering Approach for Improving and Interpreting Data Clustering Identification of Cell Populations from Multicolor Flow Cytometry Data. Cell Counting, Part A, 93(6), 597-610; the contents of which are incorporated herein by reference. On the other hand, FlowDensity can find density troughs in the data, which can be used as gate boundaries for a predefined hierarchy. For a description of the FlowDensity algorithm, please refer to the following publications, such as Malek et al. (2015), flowDensity: Manual gating for reproducing flow cytometry data by density-based automated cell population identification, Bioinformatics, 31(4), 606-607; the contents of which are incorporated herein by reference. However, the DAFi algorithm can only consider small inter-sample variations, while the FlowDensity algorithm requires manual expert tuning of automated phenotypic analysis based on the specifics of each assay kit.

[0008] The Ek'Balam algorithm is a flow cytometry data tuning method that improves upon the DAFi and FlowDensity algorithms because it is easier to tune and can handle larger inter-sample variations. However, flow cytometry data analysis protocols involving the measurement of two or more parameters (e.g., fluorescence signals), such as the Ek'Balam algorithm, become complicated by spillover, a phenomenon where light modulated by particles indicating a particular fluorophore is received by one or more detectors not configured to measure that parameter. Thus, light may "spill over" and be detected by non-target detectors. Spillover can be corrected by unmixing, where a new full-fluorescent pigment intensity value is calculated by solving a system of equations that correlate the fluorophore intensity value with the resulting measured detector value based on the observed spillover level. Unmixing is often referred to as "compensation" when the number of detectors equals the number of fluorophores unmixed. While unmixing corrects the intensity contribution of each fluorophore to the others, it cannot correct for noise contributions—the errors caused by spillover in the flow cytometry data. This noise is called "spillover diffusion." In some cases, spillover noise is constructive, making the signal strength higher than observed in other cases, while in others, the noise is destructive, making the strength lower. For example, Figure 1 This demonstrates how population 101 diffuses due to noise from fluorescent dye A on the detector (which measures fluorescent dye B). Because of population diffusion, classification (i.e., phenotypic analysis) of fluorescence flow cytometry data according to data analysis protocols can become inaccurate. For example, Figure 2A and Figure 2B This demonstrates the impact of spillover diffusion on the Ek'Balam algorithm's ability to correctly distinguish clusters of different flow cytometry data populations. For example... Figure 2B As shown, population 101 was incorrectly partitioned, and a portion of the population was therefore assigned an incorrect phenotype (i.e., A). + B + Instead of A + B - Therefore, a solution is needed to address spillover and diffusion in flow cytometry data analysis. Summary of the Invention

[0009] Aspects of the present invention include a method for classifying fluorescence flow cytometry data. In some embodiments, the method includes processing the flow cytometry data with a supervised algorithm configured to cluster the fluorescence flow cytometry data into populations. In embodiments, the fluorescence flow cytometry data is clustered based on the state of each data point relative to a hierarchical structure. In such embodiments, the fluorescence flow cytometry data is clustered into populations based on whether the fluorescence flow cytometry data is positive or negative relative to a specific fluorophore. In some embodiments, the fluorescence flow cytometry data is determined to be positive or negative for a specific fluorophore based on its relationship to a threshold. After clustering the data, embodiments of the method include determining the degree of spillover diffusion. In some embodiments, determining spillover diffusion includes calculating a spillover diffusion coefficient, i.e., quantifying the degree to which the intensity of light collected by a first detector from a first fluorophore is affected by light simultaneously collected by the same detector from a second fluorophore. In embodiments, the spillover diffusion coefficient is calculated for each possible combination of fluorescence detector and fluorophore composition to determine the amount of light emitted by a specific fluorophore collected at a given detector. In some embodiments, overflow diffusion coefficients are combined in the overflow diffusion matrix. In some embodiments, the flow cytometry data population is adjusted to account for the overflow diffusion magnitude determined in the overflow diffusion matrix. The method of implementing the invention may further include partitioning the different overflow diffusion-adjusted flow cytometry data populations. In embodiments, partitioning the populations involves calculating Matthews correlation coefficients and establishing partitions between different populations (i.e., populations exhibiting different fluorescence parameters) in a manner that optimizes the Matthews correlation coefficients for each correlation threshold. In embodiments, the partitioned flow cytometry data populations are then classified such that the different populations are associated with corresponding subtypes (i.e., phenotypes).

[0010] The invention also includes a system for classifying fluorescence flow cytometry data. In some embodiments, the system includes a particle analyzer configured to generate fluorescence flow cytometry data. The system may further include a processor having a memory operatively coupled thereto, wherein the memory includes instructions stored thereon that, when executed by the processor, cause the processor to process the flow cytometry data using a supervised algorithm configured to cluster the fluorescence flow cytometry data into populations. In embodiments, the fluorescence flow cytometry data is clustered based on the state of each data point relative to a hierarchical structure. In such embodiments, the fluorescence flow cytometry data is clustered into populations based on whether the fluorescence flow cytometry data is positive or negative relative to a specific fluorescent dye. In some embodiments, the fluorescence flow cytometry data is determined to be positive or negative for a specific fluorescent dye based on the relationship between the fluorescence flow cytometry data and a threshold. After clustering the data, the processor is configured to determine the degree of spillover diffusion. In some embodiments, determining spillover diffusion includes calculating spillover diffusion coefficients, i.e., quantifying the degree to which the intensity of light collected by a first detector from a first fluorophore is affected by light simultaneously collected by the same detector from a second fluorophore. In embodiments, spillover diffusion coefficients are calculated for each possible combination of the fluorescent detector and the fluorophore to determine the amount of light emitted by a particular fluorophore collected at a given detector. In some embodiments, the processor is configured to combine spillover diffusion coefficients in a spillover diffusion matrix. In some embodiments, the flow cytometry data population is adjusted to take into account the spillover diffusion magnitude determined in the spillover diffusion matrix. The processor may be further configured to partition the different spillover diffusion-adjusted flow cytometry data populations. In embodiments, partitioning the populations involves calculating Matthews correlation coefficients and establishing partitions between different populations (i.e., populations exhibiting different fluorescence parameters) in a manner that optimizes the Matthews correlation coefficients for each correlation threshold. In embodiments, the partitioned flow cytometry data populations are then classified such that the different populations are associated with corresponding subtypes (i.e., phenotypes).

[0011] The invention further includes a computer control system, wherein the system further includes one or more computers for achieving full or partial automation. In some embodiments, the system includes a computer having a computer-readable storage medium storing a computer program, wherein the computer program includes instructions, when loaded onto the computer, to perform the following operations: clustering fluorescence flow cytometry data into populations based on one or more different parameters (i.e., fluorophores); determining spillover diffusion between detector-parameter pairs (i.e., determined by calculating spillover diffusion coefficients); creating a spillover diffusion matrix showing how the detection of a particular fluorophore by a corresponding detector is affected by spillover of other fluorophores; adjusting the fluorescence flow cytometry data to compensate for spillover diffusion by reducing the spillover diffusion magnitude determined by the spillover diffusion matrix; evaluating the quality of separating different partitions of different fluorescence flow cytometry data populations by calculating a Matthews correlation coefficient relative to a threshold (used to distinguish between populations positive and negative for a given parameter); and classifying the adjusted fluorescence flow cytometry data populations (i.e., phenotypic analysis). Attached Figure Description

[0012] A full understanding of the invention can be obtained by reading the following detailed description in conjunction with the accompanying drawings. The drawings include the following illustrations:

[0013] Figure 1 The spillover diffusion of fluorescence flow cytometry data was described.

[0014] Figure 2A It describes how spillover diffusion suppresses the algorithm's ability to correctly distinguish clusters of different fluorescence flow cytometry data populations.

[0015] Figure 2B The paper describes how spillover diffusion suppresses the algorithm's ability to classify (i.e., phenotype) populations of fluorescence flow cytometry data.

[0016] Figure 3 The hierarchical structure for phenotypic analysis of T cells is depicted, where phenotypic analysis is accomplished by determining whether cells are positive or negative relative to the presence of CD4 or CD8.

[0017] Figure 4 The case of clustering fluorescence flow cytometry data relative to a threshold is described.

[0018] Figure 5 The spillover diffusion matrix is ​​depicted.

[0019] Figure 6 The process of adjusting flow cytometry data populations based on spillover diffusion is described.

[0020] Figure 7AThe flow cytometry data have been adapted based on spillover diffusion.

[0021] Figure 7B The paper describes the classification (i.e., phenotypic analysis) of flow cytometry data that has been adjusted based on spillover diffusion.

[0022] Figure 8 A flow cytometer according to certain embodiments is described.

[0023] Figure 9 A functional block diagram of an example processor according to certain embodiments is depicted.

[0024] Figure 10 A block diagram of a computing system according to certain embodiments is depicted. Detailed Implementation

[0025] This invention provides a method for classifying fluorescence flow cytometry data. In some cases, the method includes processing the flow cytometry data with a supervised algorithm configured to cluster the data into different populations based on the relationship between data points and relevant thresholds. In embodiments, the method includes calculating spillover diffusion coefficients and combining them in a spillover diffusion matrix to determine the degree of spillover diffusion. In some embodiments, the flow cytometry data populations are adjusted to reduce the spillover diffusion magnitude determined by the spillover diffusion matrix. In embodiments, the spillover diffusion-adjusted populations are partitioned after evaluating potential partitions relative to the thresholds. In embodiments, the partitioned flow cytometry data populations are classified according to a hierarchical structure (i.e., phenotypic analysis). This invention also provides a system and computer-readable medium for classifying fluorescence flow cytometry data.

[0026] Before describing the invention in more detail, it should be understood that the invention is not limited to the specific embodiments described, as differences will certainly exist in actual implementation. It should also be understood that the terminology used herein is for describing specific embodiments only and is not intended to limit the inventive concept; the scope of the invention will be defined only by the appended claims.

[0027] When a numerical range is provided, it should be understood that every intermediate value between the upper and lower limits of the range, as well as any other specified value or intermediate value within that range, is included within the scope of this invention. Unless the context explicitly specifies otherwise, each intermediate value should be as low as one-tenth of the lower limit unit. The upper and lower limits of these smaller ranges may be independently included within the smaller range and also within this invention, subject to the requirements of any specifically excluded limits within the range. Where the range includes one or two limits, the range excluding any one or both of the included limits is also included within this invention.

[0028] Certain ranges presented in this paper are preceded by the term "approximately". The purpose of using the term "approximately" in this paper is to provide textual support for the precise figures that follow and for figures that are close to or approximate to the figures following the term. In determining whether a figure is close to or approximate to a specifically listed figure, an unlisted figure that is close to or approximates can be substantially equivalent to the specifically listed figure in its context.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Although similar or equivalent methods and materials may be used in the implementation or testing of this invention, representative exemplary methods and materials are described below.

[0030] All publications and patents referenced in this specification are incorporated herein by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference, and the purpose of their inclusion is to disclose and describe methods and / or materials relating to the cited publications. References to any publication refer to its content disclosed prior to the filing date and should not be construed as an admission that the present invention is not entitled to precede such publication by prior invention. Furthermore, the publication dates provided may differ from the actual publication dates and may require separate verification.

[0031] It should be noted that, as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly indicates otherwise. It should also be noted that claims may be drafted to exclude any optional elements. Therefore, this statement is intended as a precondition for the use of specialized terms related to the stated elements of a claim, such as “alone,” “only,” or the use of the limiting word “negative.”

[0032] Upon reading this invention, the following will be apparent to those skilled in the art: each individual embodiment described and listed herein has hierarchical components and features that can be quickly decomposed or combined with features of any of the other embodiments without departing from the scope and spirit of the invention. Any of the stated methods may be implemented in the order of the stated events or in any other logically possible order.

[0033] Although the system and method have been or will be described and their functions explained for grammatical fluency, it should be clearly understood that, unless expressly provided for in Chapter 35 of the United States Code, a claim shall not in any case be construed as necessarily being limited to “method” or “step”, but shall conform to the judicial principles of equivalence and the meaning and full scope of the equivalent as defined in the claim. When a claim is explicitly drafted in accordance with Section 112 of Chapter 35 of the United States Code, the claim shall be in full conformity with the legal equivalents in Section 112 of Chapter 35 of the United States Code.

[0034] Methods for classifying fluorescence flow cytometry data

[0035] As described above, aspects of the present invention include a method for classifying fluorescence flow cytometry data. "Fluorescence flow cytometry data" refers to information about parameters of a sample (e.g., cells, particles) in a flow cell, collected by any number of fluorescence detectors in a particle analyzer. In embodiments, fluorescence flow cytometry data includes signals from a variety of different fluorescent dyes, for example, 2 to 20 different fluorescent dyes, and includes 3 to 5 different fluorescent dyes. In some embodiments, the variety of different fluorescent dyes includes 2 or more different fluorescent dyes, including 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, and 20 or more different fluorescent dyes. Fluorescence flow cytometry data can be obtained by any suitable method, including those described below.

[0036] In some embodiments, the method includes generating one or more population clusters based on determined parameters (e.g., fluorescence) of analytes (e.g., cells, particles) in the sample. As used herein, an analyte “population” or “subpopulation” (e.g., cells or other particles) generally refers to a group of analytes that possess characteristics (e.g., optical, impedance, or temporal characteristics) relative to one or more measured fluorescence parameters, causing the measured parameter data to form clusters in the data space. Thus, a population is identified as a cluster in the data. Conversely, while clusters corresponding to noise or background are also commonly observed, each data cluster is typically interpreted as corresponding to a specific type of cell or analyte population. Clusters can be defined in subsets of the dimension (e.g., subsets of measured fluorescence parameters (i.e., fluorophores)) that correspond to populations that differ only in subsets of the measured parameters or features (extracted from the sample measurements).

[0037] An aspect of this invention includes supervised clustering of fluorescence flow cytometry data. This step can employ any suitable supervised phenotypic analysis algorithm. In some embodiments, the supervised phenotypic analysis algorithm is the Ek'Balam algorithm. A description of the Ek'Balam algorithm can be found in the following publication: Amir et al. (2019), Developing a Comprehensive Antibody Staining Database Using a Standardized Analysis Pipeline, Frontiers in Immunology, 10, 1315; the contents of which are incorporated herein by reference. Therefore, in some embodiments, fluorescence flow cytometry data are clustered based on the state of each data point relative to a hierarchical structure. The hierarchical structure, as described herein, defines the criteria upon which fluorescence flow cytometry data are grouped into specific populations. In some embodiments, the hierarchical structure establishes a way in which data points that are positive or negative for the same parameter are grouped together. For example, Figure 3 The diagram illustrates a hierarchical structure for clustering T cells, where clustering is accomplished by determining whether cells are positive or negative in relation to the presence of CD4 or CD8. Cells that are CD4 positive but CD8 negative are classified as "CD4 T cells," while cells that are positive for both markers are classified as "double-positive T cells," and so on.

[0038] In some embodiments, clustering fluorescence flow cytometry data includes comparing the data to a threshold. A “threshold” refers to a value that distinguishes positive / negative fluorescence flow cytometry data (relative to a specific fluorescent dye). In other words, fluorescence flow cytometry data above the threshold can be described as positive for a specific fluorescent dye, while fluorescence flow cytometry data below the threshold can be described as negative for the same fluorescent dye. For example, Figure 4This paper describes how flow cytometry data can be clustered based on thresholds that distinguish between positive and negative data. Threshold 401 distinguishes between CD8-positive and CD8-negative flow cytometry data. Similarly, threshold 402 distinguishes between CD4-positive and CD4-negative flow cytometry data. Clustering can be performed using any suitable method. In some embodiments, clustering is performed using the FlowSOM algorithm. A description of FlowSOM can be found in Van Gassen et al. (2015), FlowSOM: Visualizing and Interpreting Flow Cytometry Data Using Self-Organizing Maps, Cell Counting, Part A, 87(7), 636-645; the contents of which are incorporated herein by reference. In other embodiments, clustering is performed using algorithms selected from the following: SOM – see Kohonen, “Self-Organizing Maps,” IEEE Transactions (1990), 78: 1464-1480; K-means – see Hartigan and Wong, “Algorithm AS136: K-Means Clustering Algorithm,” Proceedings of the Royal Statistical Society, C Series (Applied Statistics) (1979), 28: 100-108; Gaussian Mixture Model – see Sanderson and Curtin, “Open Source C++ Implementation of Multithreaded Gaussian Mixture Models, k-means and Expectation Maximization,” 11th Annual Conference on Gaussian Mixture Models, 2017. International Conference on Signal Processing and Communication Systems (ICSPCS) (https: / / doi.org / 10.1109 / ICSPCS.2017.8270510); FlowGrid - see the following publication for a related description: Ye and Ho, “Ultra-fast clustering of single-cell flow cytometry data using FlowGrid”, BMC Systems Biology, 13, 35 (2019), https: / / doi.org / 10.1186 / s12918-019-0690-2; or X-shift - see the following publication for a related description: Samusik et al., “Automatic mapping of phenotypic space to single-cell data”, Nature: Methods (2016), 13: 493-496.

[0039] After clustering the flow cytometry data, embodiments of the present invention include determining the extent of spillover diffusion of the fluorescence flow cytometry data population. As described in the introduction, and as... Figure 1 and Figure 2A-2BAs shown, fluorescence flow cytometry data at collection points (i.e., points received by one or more fluorescence detectors) are affected by spillover diffusion. Spillover is a phenomenon where light modulated by particles indicating a specific fluorescent dye is received by one or more detectors not configured to measure the parameter. Therefore, light may "spill over" and be detected by non-target detectors. Spillover diffusion is thus noise present in fluorescence flow cytometry data caused by spillover. Therefore, in some embodiments, unadjusted flow cytometry data contains errors because one or more detectors inadvertently detect light of certain wavelengths. In some embodiments, determining the degree of spillover diffusion includes quantifying the extent to which fluorescence flow cytometry data collected by a first detector from a first fluorescent dye is affected by light simultaneously collected by the same detector from a second fluorescent dye. In some cases, fluorescence flow cytometry data affected by spillover diffusion is affected by signal strength that is higher than signal strength observed in other cases (i.e., spillover diffusion noise is constructive). In other cases, fluorescence flow cytometry data affected by spillover diffusion is affected by signal strength that is lower than signal strength observed in other cases (i.e., spillover diffusion noise is destructive). In some embodiments, determining the extent of spillover diffusion includes calculating a spillover diffusion coefficient. A detailed description of the spillover diffusion coefficient can be found in the following publication: Nguyen et al. (2013), Quantifying spillover diffusion to compare instrument performance and aiding in the design of multicolor assay kits, Cell Counting, Part A, 83(3): 306-315, the contents of which are incorporated herein by reference. In some embodiments, the spillover diffusion coefficient is calculated according to Formula 1:

[0040]

[0041] As shown in Formula 1, SS is the spillover diffusion coefficient, Δσ f Δd is the incremental standard deviation, representing the emission diffusion between the positive and negative fluorescence flow cytometry data collected from the fluorophore; Δd ​​is the difference in fluorescence intensity between the positive and negative fluorescence flow cytometry data received by the fluorescence detector. Therefore, the spillover diffusion coefficient measures the degree to which the fluorescence flow cytometry data acquired by a given fluorescence detector is affected by the presence of light associated with a specific fluorophore. In other words, the spillover diffusion coefficient estimates the error (i.e., noise) in the fluorescence flow cytometry data based on the light emitted by the relevant fluorophore collected by a given detector. In embodiments, for a given fluorophore-detector pair, a higher spillover diffusion coefficient indicates greater spillover diffusion.

[0042] In some embodiments, determining the extent of spillover diffusion further includes calculating spillover diffusion coefficients for each possible fluorescence detector-fluorochrome pair to determine how the fluorescence flow cytometry data collected at each detector is affected by the presence of light associated with each fluorochrome. In embodiments, the calculated spillover diffusion coefficients of each fluorescence detector-fluorochrome pair are combined in a spillover diffusion matrix. In some embodiments, the spillover diffusion matrix shows how the detection of a particular fluorochrome by its respective detector is affected by spillover from other fluorochromes. For example, Figure 5 An embodiment of an overflow diffusion matrix providing overflow diffusion coefficients for 23 different fluorescent dyes is illustrated. Each column of the matrix corresponds to a detector configured to detect one of the 23 different fluorescent dyes, and each row of the matrix corresponds to parameters of the detected flow cytometry data. Cells where columns and rows intersect are filled with the calculated overflow diffusion coefficients of the fluorescent detector-fluorescent dye pairs, indicating the extent to which the investigated fluorescent dye contributes to detection errors in the associated detector. The total influence of the fluorescent dye on overflow diffusion can be estimated by summing all values ​​in its rows, while the total influence of overflow diffusion on the detector can be calculated by summing all values ​​in its columns. In some embodiments, the overflow diffusion coefficients are summed to calculate the total diffusion effect (i.e., the cumulative effect of overflow diffusion on a specific subset of flow cytometry data).

[0043] In some embodiments, an overflow diffusion matrix is ​​calculated using the AutoSpread algorithm. The AutoSpread algorithm, created by Becton Dickinson and described in U.S. Provisional Patent Application No. 63 / 020,758, is incorporated herein by reference. The AutoSpread algorithm is configured to create an overflow diffusion matrix (e.g., as described above) without needing to distinguish between positive and negative flow cytometry data populations relative to a given fluorophore. AutoSpread characterizes the diffusion of the detected signal from the first fluorophore by incorporating a second fluorophore into the same flow cytometry assay kit. AutoSpread generates a coefficient for each interaction between the fluorescence detector and the fluorophore and arranges these coefficients in a matrix similar to the overflow diffusion matrix described above. In embodiments, calculating the overflow diffusion coefficients involves assuming that the fluorescence intensity acquired by the fluorescence detector for the negative flow cytometry data population is zero and the corresponding standard deviation is unknown. In some embodiments, the overflow diffusion coefficients are calculated according to Equation 2:

[0044]

[0045] As shown in Equation 2, SS is the spillover diffusion coefficient, σ 2It is the standard deviation of the positive fluorescence flow cytometry data population. is an estimate of the standard deviation of the negative fluorescence flow cytometry data population, where d is the intensity of light collected by the fluorescence detector. In some embodiments, to obtain an estimate of the standard deviation of the negative fluorescence flow cytometry data population, assuming that the fluorescence intensity collected by the fluorescence detector for the negative fluorescence flow cytometry data population is zero. The spillover diffusion coefficient was calculated using a series of linear regression methods. First, the fluorescence flow cytometry data were categorized by quantiles based on the intensity values ​​detected by the fluorescence detector. The default number of quantiles was 256, but this was reduced to 8 to ensure each quantile had a sufficient number of data points for reliable estimation of the standard deviation. Next, regression analysis was performed on the corresponding robust standard deviation of the light emitted by the fluorophore relative to the square root of the median intensity of light detected for each quantile. Assuming the intensity of light detected for the negative population was zero, the y-intercept of the ordinary least squares fit was considered an estimate of the standard deviation of the negative flow cytometry data population. A new zero-adjusted standard deviation was obtained using the estimate of the corresponding standard deviation of the light emitted by the fluorophore. Regression analysis was then performed on the zero-adjusted standard deviation of the fluorophore relative to the square root of the median fluorescence intensity detected by the fluorescence detector for each quantile. The slope of the ordinary least squares fit (calculated using Equation 2) was considered the spillover diffusion coefficient.

[0046] The invention further includes adjusting fluorescence flow cytometry data to account for spillover diffusion. "Adjustment" means changing the data to more accurately quantify the fluorescent dyes present in the irradiated sample (e.g., cells, particles) in the flow cell. In some embodiments, the fluorescence flow cytometry data is adjusted to substantially eliminate all constructive errors caused by spillover diffusion. In embodiments, adjusting the fluorescence flow cytometry data includes generating different spillover-diffusion-adjusted populations. In some embodiments, generating different spillover-diffusion-adjusted populations includes reducing the spillover diffusion magnitude of the flow cytometry data-related populations, i.e., counteracting the constructive effect of the signal affected by spillover diffusion. In some embodiments, the spillover diffusion magnitude is determined by the spillover diffusion matrix. In some embodiments, adjusting the flow cytometry data includes reducing the total diffusion effect of the relevant portions of the flow cytometry data. For example, Figure 6 The diagram illustrates adjustments to flow cytometry data to account for spillover propagation. Arrow 601 indicates how the median of the population cluster is shifted downwards. In other words, reducing the spillover propagation of data points within the population cluster alters the cluster's position relative to other population clusters. Figure 7AThe flow cytometry data, adjusted through the process described above, are depicted. In some embodiments, the fluorescence flow cytometry data are adjusted by calculating the probability that the corresponding value of each cell is higher than a threshold based on all known noise sources, including spillover diffusion.

[0047] After adjusting the flow cytometry data population for spillover diffusion (e.g., as described above), aspects of the invention further include partitioning the spillover diffusion-adjusted flow cytometry data population. "Partitioning" means dividing flow cytometry data populations that are determined to be different based on their positivity or negativity relative to a variety of different fluorescent dyes. In other words, partitions are created to formally distinguish flow cytometry data populations that are to be classified in different ways (e.g., representing different phenotypes).

[0048] In embodiments, partitioning different spillover-diffusion-adjusted flow cytometry data cohorts involves calculating the Matthews correlation coefficient (MCC). This coefficient is a measure of binary classification quality and quantifies the accuracy with which positive and negative flow cytometry data cohorts can be distinguished using a given threshold. A relevant description of the Matthews correlation coefficient can be found in the following publication: Matthews, BW (1975), Comparison of prediction and observational secondary structures of T4 phage lysozyme, Acta Biochimica et Biophysica Sinica (BBA) - Protein Structure, 405(2), 442-451; the contents of which are incorporated herein by reference. In embodiments, potential partitions between different spillover-diffusion-adjusted fluorescence flow cytometry data cohorts are evaluated relative to various thresholds. In such embodiments, the Matthews correlation coefficient is calculated for each threshold to assess the level of consistency between the threshold and the partition. Therefore, embodiments of the present invention involve partitioning cohorts to optimize the Matthews correlation coefficient relative to said correlation threshold, which distinguishes flow cytometry data cohorts that are positive and negative relative to a particular fluorescent dye. In other words, partitioning fluorescence flow cytometry data involves maximizing the ability to distinguish different populations (i.e., populations exhibiting different combinations of fluorescence parameters), depending on their relationship to a relevant threshold (e.g., quantified by the Matthews correlation coefficient). In this embodiment, the process is iterative to determine the optimal partitioning for distinguishing between populations that are positive and negative for the relevant parameters of each flow cytometry data. In this embodiment, the Matthews correlation coefficient is calculated according to Equation 3:

[0049]

[0050] As shown in Formula 3, MCC is the Matthews correlation coefficient, TP represents a true positive event, TN represents a true negative event, FP represents a false positive event, and FN represents a false negative event. According to the present invention, a true positive event refers to flow cytometry data that is assessed as positive for a specific fluorescent dye based on a threshold and partition; a true negative event refers to flow cytometry data that is assessed as negative for a specific fluorescent dye based on a threshold and partition; a false positive event refers to flow cytometry data that is assessed as positive for a specific fluorescent dye based on partition but negative based on the threshold; a false negative event refers to flow cytometry data that is assessed as negative for a specific fluorescent dye based on partition but positive based on the threshold.

[0051] In some embodiments of the invention, the fluorescence flow cytometry data do not contain a signal that is positive for a specific fluorescent dye. In other words, no fluorescence emitted from said fluorescent dye is detected. In such embodiments, partitioning the different overflow-diffusion-adjusted fluorescence flow cytometry data may include calculating the balance accuracy of each partition as an alternative metric for determining the optimal partition. The balance accuracy is the average of the accuracy in determining positive events and the accuracy in determining negative events, calculated according to Equation 4:

[0052]

[0053] As shown in Formula 4, BA is the balanced accuracy, TP represents true positive events, TN represents true negative events, FP represents false positive events, and FN represents false negative events.

[0054] Embodiments of the invention further include classifying the partitioned flow cytometry data populations, i.e., determining the subtypes of cells or particles specified by each distinct spillover-diffusion-adjusted flow cytometry data population. In embodiments, classification is determined based on the hierarchical structure (as described above). Therefore, classifications (i.e., phenotypic analysis) are assigned to the partitioned flow cytometry data populations (e.g., as described above) based on the combinations of fluorescence parameters they exhibit. For example, Figure 7B Depicting Figure 7A The diagram illustrates how spillover diffusion-adjusted fluorescence flow cytometry data are partitioned to correctly assign their respective phenotypes to different population clusters. For example... Figure 2B As shown, and as described in the introduction, a portion of population cluster 101 was assigned the wrong phenotype without any data adjustment. However, regarding spillover effects (e.g., as...), Figure 6 (As shown) After adjusting the population (arrow 601), the correct phenotype (i.e., A) is assigned to the entire population 701. + B - Instead of assigning the correct phenotype to only a subset of the 701 population.

[0055] As described above, any suitable method can be used to obtain the fluorescence data employed in the method of the present invention. In some embodiments, a sample containing particles is irradiated with a light source, and light from the sample is detected to generate a relevant particle population based at least in part on the measurements of the detected light. In some cases, the sample is a biological sample. According to its conventional meaning, the term "biological sample" refers to a subset, cell, or component of a whole organism, plant, fungus, or animal tissue, and in some cases may be found in blood, mucus, lymph, synovial fluid, cerebrospinal fluid, saliva, bronchoalveolar lavage fluid, amniotic fluid, amniotic cord blood, urine, vaginal fluid, and semen. Therefore, "biological sample" refers both to a natural organism or a subset of its tissues and to homogenates, lysates, or extracts prepared from said organism or a subset of its tissues, including but not limited to: plasma; serum; cerebrospinal fluid; lymph; slices of skin, respiratory tract, gastrointestinal tract, cardiovascular, and genitourinary tract; tears; saliva; breast milk; blood cells; tumors; organs, etc. Biological samples can be any type of biological tissue, including healthy tissue and pathological tissue (e.g., cancerous tissue, malignant tissue, necrotic tissue, etc.). In some embodiments, the biological sample is a liquid sample, such as blood or its derivatives (e.g., plasma, tears, urine, semen, etc.), wherein, in some cases, the sample is a blood sample, including whole blood samples, such as blood obtained by venipuncture or finger puncture (wherein the blood may or may not be mixed with any reagents (e.g., preservatives, anticoagulants, etc.) before analysis).

[0056] In some embodiments, the sample source is "mammal," a term widely used to describe organisms belonging to the mammal class, including carnivores (e.g., dogs and cats), rodents (e.g., mice, guinea pigs, and rats), and primates (e.g., humans, chimpanzees, and monkeys). In some cases, the subject is human. The method is applicable to samples obtained from male and female subjects at any developmental stage (i.e., newborns, infants, adolescents, teenagers, and adults), in which, in some embodiments, the human subject is an adolescent, teenager, or adult. Although the invention is applicable to samples taken from human subjects, it should be understood that the method can also be performed on blood taken from other animal subjects (i.e., in "non-human subjects"), including but not limited to birds, mice, rats, dogs, cats, livestock, and horses.

[0057] In implementing the target method, a sample containing particles is irradiated with light from a light source (e.g., in the fluid medium of a flow cytometer). In some embodiments, the light source is a broadband light source that emits light with a wide range of wavelengths, such as 50 nm or more, 100 nm or more, 150 nm or more, 200 nm or more, 250 nm or more, 300 nm or more, 350 nm or more, 400 nm or more, including a span of 500 nm or more. For example, a suitable broadband light source emits light with a wavelength range of 200 nm to 1500 nm. Another example of a suitable broadband light source includes a light source that emits light with a wavelength range of 400 nm to 1000 nm. When the method includes irradiation with a broadband light source, the desired broadband light source scheme may include, but is not limited to: halogen lamps, deuterium arc lamps, xenon arc lamps, stable fiber-coupled broadband light sources, broadband LEDs with a continuous spectrum, superluminescent diodes, semiconductor light-emitting diodes, broadband LED white light sources, multi-LED integrated light sources, and other broadband light sources or any combination thereof.

[0058] In other embodiments, the method includes irradiation with a narrowband light source that emits light of a specific wavelength or a narrow range of wavelengths. For example, irradiation with a light source emitting light of a narrow range of wavelengths, such as 50 nm or less, 40 nm or less, 30 nm or less, 25 nm or less, 20 nm or less, 15 nm or less, 10 nm or less, 5 nm or less, 2 nm or less, including light sources emitting light of a specific wavelength (i.e., monochromatic light). When the method includes irradiation with a narrowband light source, the desired narrowband light source may include, but is not limited to: narrow-wavelength LEDs, laser diodes, or broadband light sources coupled to one or more optical bandpass filters, diffraction gratings, monochromators, or any combination thereof.

[0059] Aspects of the invention include collecting fluorescence using a fluorescence detector. In some cases, the fluorescence detector may be configured to detect fluorescence emission from fluorescent molecules, such as a label-specific binding member (e.g., a labeled antibody that specifically binds to a target marker) associated with the particles in the flow cell. In some embodiments, the method includes detecting fluorescence from the sample using one or more fluorescence detectors, such as two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, and including 25 or more fluorescence detectors. In embodiments, each of the fluorescence detectors is configured to generate a fluorescence data signal. Each fluorescence detector may independently detect fluorescence from the sample at one or more wavelengths in the wavelength range of 200 nm–1200 nm. In some cases, the method includes detecting fluorescence from the sample within a wavelength range, such as 200 nm to 1200 nm, 300 nm to 1100 nm, 400 nm to 1000 nm, 500 nm to 900 nm, including 600 nm to 800 nm. In other cases, the method includes detecting fluorescence at one or more specific wavelengths using various fluorescence detectors. For example, the fluorescence can be detected at one or more wavelengths selected from 450 nm, 518 nm, 519 nm, 561 nm, 578 nm, 605 nm, 607 nm, 625 nm, 650 nm, 660 nm, 667 nm, 670 nm, 668 nm, 695 nm, 710 nm, 723 nm, 780 nm, 785 nm, 647 nm, 617 nm, and any combination thereof, depending on the number of different fluorescence detectors in the target light detection system. In some embodiments, the method includes detecting the wavelength of light corresponding to the fluorescence peak wavelength of certain fluorescent dyes present in the sample. In embodiments, fluorescence flow cytometry data are received through one or more fluorescence detectors (e.g., one or more detection channels), such as two or more, three or more, four or more, five or more, six or more, and including eight or more fluorescence detectors (e.g., eight or more detection channels).

[0060] A system for classifying fluorescence flow cytometry data.

[0061] The invention also includes a system for classifying fluorescence flow cytometry data. In embodiments, the fluorescence flow cytometry data is clustered, adjusted for spillover diffusion, and partitioned to classify different populations in different ways. In some embodiments, the system includes a particle analyzer configured to generate fluorescence flow cytometry data and a processor configured to analyze the fluorescence flow cytometry data.

[0062] In some embodiments, the target particle analyzer includes a flow cell and a laser configured to irradiate particles in the flow cell. In embodiments, the laser can be any suitable laser, such as a continuous-wave laser. For example, the laser can be a diode laser, such as an ultraviolet diode laser, a visible diode laser, and a near-infrared diode laser. In other embodiments, the laser can be a helium-neon (HeNe) laser. In some cases, the laser is a gas laser, such as a helium-neon laser, an argon laser, a krypton laser, a xenon laser, a nitrogen laser, a CO2 laser, a CO laser, an argon fluoride (ArF) excimer laser, a krypton fluoride (KrF) excimer laser, a xenon chloride (XeCl) excimer laser, or a xenon fluoride (XeF) excimer laser, or a combination thereof. In other cases, the target flow cytometer includes a dye laser, such as a stilbene, coumarin, or rhodamine laser. In other cases, the target laser includes metal vapor lasers, such as helium-cadmium (HeCd) lasers, helium-mercury (HeHg) lasers, helium-selenium (HeSe) lasers, helium-silver (HeAg) lasers, strontium lasers, neon-copper (NeCu) lasers, copper lasers, or gold lasers and combinations thereof. In other cases, the target flowing laser includes solid-state lasers, such as ruby ​​lasers, Nd:YAG lasers, NdCrYAG lasers, Er:YAG lasers, Nd:YLF lasers, Nd:YVO4 lasers, Nd:YCa4O(BO3)3 lasers, Nd:YCOB lasers, Ti:sapphire lasers, thulim YAG lasers, YAG ytterbium lasers, ytterbium trioxide lasers, or cerium-doped lasers and combinations thereof.

[0063] The invention also includes a forward-scattering detector configured to detect forward-scattered light. The number of forward-scattering detectors in the target flow cytometer can vary as needed. For example, the target particle analyzer may include one or more forward-scattering detectors, such as two or more, three or more, four or more, and including five or more. In some embodiments, the flow cytometer includes one forward-scattering detector. In other embodiments, the flow cytometer includes two forward-scattering detectors.

[0064] Any suitable detector used to detect the collected light can be used in the forward scattering detector described herein. The target detector may include, but is not limited to, optical sensors or detectors, such as active pixel sensors (APS), avalanche photodiodes, image sensors, charge-coupled devices (CCDs), enhancement-mode charge-coupled devices (ICCDs), light-emitting diodes, photon counters, radiative thermal analyzers, thermoelectric detectors, photoresistors, photovoltaic cells, photodiodes, photomultiplier tubes (PMTs), phototransistors, quantum dot photoconductors, or combinations thereof, and other detectors. In some embodiments, the collected light is measured using a charge-coupled device (CCD), semiconductor charge-coupled device (CCD), active pixel sensor (APS), complementary metal-oxide-semiconductor (CMOS), image sensor, or N-type metal-oxide-semiconductor (NMOS) image sensor. In some embodiments, the detector is a photomultiplier tube, for example, with an effective detection surface area range of 0.01 cm² for each region. 2 Up to 10cm 2 (e.g., 0.05cm) 2 Up to 9cm 2 0.1cm 2 Up to 8cm 2 0.5cm 2 Up to 7cm 2, And including 1cm 2 up to 5cm 2 ) photomultiplier tube.

[0065] In the case where the target particle analyzer includes multiple forward scattering detectors, the detectors can be identical, or the detector set can be a combination of detectors of different types. For example, in the case where the target particle analyzer includes two forward scattering detectors, in some embodiments, the first forward scattering detector is a CCD-type device, while the second forward scattering detector (or imaging sensor) is a CMOS-type device. In other embodiments, both the first and second forward scattering detectors are CCD-type devices. In other embodiments, both the first and second forward scattering detectors are CMOS-type devices. In other embodiments, the first forward scattering detector is a CCD-type device, while the second forward scattering detector is a photomultiplier tube (PMT). In other embodiments, the first forward scattering detector is a CMOS-type device, while the second forward scattering detector is a photomultiplier tube. In other embodiments, both the first and second forward scattering detectors are photomultiplier tubes.

[0066] In an embodiment, the forward scattering detector is configured to measure light continuously or at discrete time intervals. In some cases, the target detector is configured to measure the collected light continuously. In other cases, the target detector is configured to perform measurements at discrete time intervals, for example, the time intervals for measuring light are 0.001 milliseconds, 0.01 milliseconds, 0.1 milliseconds, 1 millisecond, 10 milliseconds, 100 milliseconds, including 1000 milliseconds or some other time interval.

[0067] Embodiments of the present invention also include an optical dispersion / splitter module located between the flow cell and the forward scattering detector. The intended optical dispersion device includes, but is not limited to, colored glass, bandpass filters, interference filters, dichroic mirrors, diffraction gratings, monochromators and combinations thereof, and other wavelength separation devices. In some embodiments, the bandpass filter is located between the flow cell and the forward scattering detector. In other embodiments, more than one bandpass filter is located between the flow cell and the forward scattering detector, for example, two or more, three or more, four or more, and including five or more. In embodiments, the minimum bandwidth range of the bandpass filter is 2 nm to 100 nm, for example, 3 nm to 95 nm, 5 nm to 95 nm, 10 nm to 90 nm, 12 nm to 85 nm, 15 nm to 80 nm, including bandpass filters with a minimum bandwidth range of 20 nm to 50 nm. The device reflects light of other wavelengths to the forward scattering detector.

[0068] Some embodiments of the invention include a side-scattering detector configured to detect the side-scattering wavelength of light (e.g., light refracted and reflected from the surface and internal structure of the particle). In other embodiments, the flow cytometer includes a plurality of side-scattering detectors, such as two or more, three or more, four or more, and including five or more.

[0069] Any suitable detector used to detect the collected light can be used in the side-scattering detector described herein. The target detector may include, but is not limited to, optical sensors or detectors, such as active pixel sensors (APS), avalanche photodiodes, image sensors, charge-coupled devices (CCDs), enhancement-mode charge-coupled devices (ICCDs), light-emitting diodes, photon counters, radiative thermal analyzers, thermoelectric detectors, photoresistors, photovoltaic cells, photodiodes, photomultiplier tubes (PMTs), phototransistors, quantum dot photoconductors, or combinations thereof, and other detectors. In some embodiments, the collected light is measured using a charge-coupled device (CCD), semiconductor charge-coupled device (CCD), active pixel sensor (APS), complementary metal-oxide-semiconductor (CMOS), image sensor, or N-type metal-oxide-semiconductor (NMOS) image sensor. In some embodiments, the detector is a photomultiplier tube, for example, with an effective detection surface area range of 0.01 cm² for each region. 2 Up to 10cm 2 (e.g., 0.05cm) 2 Up to 9cm 2 0.1cm 2 Up to 8cm 2 0.5cm 2 Up to 7cm 2 And including 1cm 2 up to 5cm 2 ) photomultiplier tube.

[0070] In the case where the target particle analyzer includes multiple side-scattering detectors, each side-scattering detector can be identical, or the set of side-scattering detectors can be a combination of detectors of different types. For example, in the case where the target particle analyzer includes two side-scattering detectors, in some embodiments, the first side-scattering detector is a CCD-type device, while the second side-scattering detector (or imaging sensor) is a CMOS-type device. In other embodiments, both the first and second side-scattering detectors are CCD-type devices. In other embodiments, both the first and second side-scattering detectors are CMOS-type devices. In other embodiments, the first side-scattering detector is a CCD-type device, while the second side-scattering detector is a photomultiplier tube (PMT). In other embodiments, the first side-scattering detector is a CMOS-type device, while the second side-scattering detector is a photomultiplier tube. In other embodiments, both the first and second side-scattering detectors are photomultiplier tubes.

[0071] Embodiments of the present invention also include an optical dispersion / splitter module located between the flow cell and the side-scattering detector. The intended optical dispersion device includes, but is not limited to, colored glass, bandpass filters, interference filters, dichroic mirrors, diffraction gratings, monochromators and combinations thereof, and other wavelength separation devices.

[0072] In one embodiment, the target particle analyzer further includes a fluorescence detector configured to detect light of one or more fluorescence wavelengths. In other embodiments, the particle analyzer includes multiple fluorescence detectors, such as two or more, three or more, four or more, five or more, ten or more, fifteen or more, and including twenty or more.

[0073] Any suitable detector used to detect the collected light can be used in the fluorescence detector described herein. The target detector may include, but is not limited to, optical sensors or detectors, such as active pixel sensors (APS), avalanche photodiodes, image sensors, charge-coupled devices (CCDs), enhancement-mode charge-coupled devices (ICCDs), light-emitting diodes, photon counters, radiative thermal analyzers, thermoelectric detectors, photoresistors, photovoltaic cells, photodiodes, photomultiplier tubes (PMTs), phototransistors, quantum dot photoconductors, or combinations thereof, and other detectors. In some embodiments, the collected light is measured using a charge-coupled device (CCD), semiconductor charge-coupled device (CCD), active pixel sensor (APS), complementary metal-oxide-semiconductor (CMOS), image sensor, or N-type metal-oxide-semiconductor (NMOS) image sensor. In some embodiments, the detector is a photomultiplier tube, for example, with an effective detection surface area range of 0.01 cm² for each region. 2 Up to 10cm 2 (e.g., 0.05cm) 2 Up to 9cm 2 0.1cm 2 Up to 8cm 2 0.5cm 2 Up to 7cm 2 And including 1cm 2 up to 5cm 2 ) photomultiplier tube.

[0074] In the case where the target particle analyzer includes multiple fluorescence detectors, the fluorescence detectors may be identical, or the set of fluorescence detectors may be a combination of detectors of different types. For example, in the case where the target particle analyzer includes two fluorescence detectors, in some embodiments, the first fluorescence detector is a CCD-type device, and the second fluorescence detector (or imaging sensor) is a CMOS-type device. In other embodiments, both the first and second fluorescence detectors are CCD-type devices. In other embodiments, both the first and second fluorescence detectors are CMOS-type devices. In other embodiments, the first fluorescence detector is a CCD-type device, and the second fluorescence detector is a photomultiplier tube (PMT). In other embodiments, the first fluorescence detector is a CMOS-type device, and the second fluorescence detector is a photomultiplier tube. In other embodiments, both the first and second fluorescence detectors are photomultiplier tubes.

[0075] Embodiments of the present invention also include an optical dispersion / splitter module located between the flow cell and the fluorescence detector. The optical dispersion device includes, but is not limited to, colored glass, bandpass filters, interference filters, dichroic mirrors, diffraction gratings, monochromators and combinations thereof, and other wavelength separation devices.

[0076] In embodiments of the invention, the target fluorescence detector is configured to measure light collected at one or more wavelengths, such as two or more wavelengths, five or more different wavelengths, ten or more different wavelengths, 25 or more different wavelengths, 50 or more different wavelengths, 100 or more different wavelengths, 200 or more different wavelengths, 300 or more different wavelengths, including measuring light emitted by a sample in a fluid medium at 400 or more different wavelengths. In some embodiments, two or more detectors in a flow cytometer as described herein are configured to measure collected light having the same or overlapping wavelengths.

[0077] In some embodiments, the target fluorescence detector is configured to measure light collected within a wavelength range (e.g., 200 nm–1000 nm). In some embodiments, the target detector is configured to collect the spectrum of light within a wavelength range. For example, a particle analyzer may include one or more detectors configured to collect the spectrum of light at one or more wavelengths within the wavelength range of 200 nm–1000 nm. In other embodiments, the target detector is configured to measure light emitted by a sample in a fluid medium at one or more specific wavelengths. For example, a particle analyzer may include one or more detectors configured to measure light at one or more wavelengths of 450 nm, 518 nm, 519 nm, 561 nm, 578 nm, 605 nm, 607 nm, 625 nm, 650 nm, 660 nm, 667 nm, 670 nm, 668 nm, 695 nm, 710 nm, 723 nm, 780 nm, 785 nm, 647 nm, 617 nm, and any combination thereof. In some embodiments, one or more detectors may be configured to pair with a specific fluorophore, such as a fluorophore used in conjunction with a sample in fluorescence analysis.

[0078] Suitable flow cytometry systems may include, but are not limited to, those described in the following publications: Ormerod (ed.), Flow Cytometry: A Practical Approach, Oxford University Press (1997); Jaroszeski et al. (eds.), Flow Cytometry Protocols, Methods of Molecular Biology, Vol. 91, Humana Press (1997); Practical Flow Cytometry, 3rd Edition, Wiley-Liss (1995); Virgo et al. (2012), Annals of Clinical Biochemistry, Jan. 49 (pt 1): 17-28; Linden et al., Symposium on Thrombosis and Hemostasis, Oct. 2004; 30 (5): 502-11; Alison et al., Journal of Pathology, Dec. 2010; 222 (4): 335-344; and Herbig et al. (2007), Reviews of Therapeutic Drug Carrier Systems, 24 (3): 203-255; the contents of these publications are incorporated herein by reference. In some cases, the target flow cytometry system includes BDBiosciences FACSCanto TM Flow cytometer, BD Biosciences FACSCanto TM II flow cytometer, BDAccuri TM Flow cytometer, BD Accuri TM C6 Plus flow cytometer, BD Biosciences FACSCelesta TMFlow cytometer, BD Biosciences FACSLyric TM Flow cytometer, BD Biosciences FACSVerse TM Flow cytometer, BD Biosciences FACSymphony TM Flow cytometer, BD Biosciences LSRFortessa TM Flow cytometer, BD Biosciences LSRFortessa TM X-20 flow cytometer, BD Biosciences FACSPresto TM Flow cytometer, BD Biosciences FACSVia TM Flow cytometer and BD Biosciences FACSCalibur TM Cell sorter, BD Biosciences FACSCount TM Cell sorter, BD Biosciences FACSLyric TM Cell sorter, BD Biosciences Via TM Cell sorter, BD Biosciences Influx TM Cell sorting instrument, BDBiosciences Jazz TM Cell sorter, BD Biosciences Aria TM Cell sorting instrument, BD Biosciences FACSAria TM II Cell Sorter, BD Biosciences FACSAria TM III Cell Sorter, BD Biosciences FACSAria TM Fusion Cell Sorter and BD Biosciences FACSMelody TM Cell sorting instrument, BDBiosciences FACSymphony TM S6 cell sorter, etc.

[0079] In some embodiments, the target system is a flow cytometry system, such as the system described in the following U.S. patents: 10,663,476; 10,620,111; 10,613,017; 10,605,713; 10,585,031; 10,578,542; 10,578,469; 10,481,074; 10,302,545; 10,145,793; 10,113,967; 10,006,852; 9,952,076; 9,933,341; 9,726,527; 9,453,789; 9,200,334; 9,09 The contents of these documents are incorporated herein by reference in their entirety: 7,640; 9,095,494; 9,092,034; 8,975,595; 8,753,573; 8,233,146; 8,140,300; 7,544,326; 7,201,875; 7,129,505; 6,821,740; 6,813,017; 6,809,804; 6,372,506; 5,700,692; 5,643,796; 5,627,040; 5,620,842; 5,602,039; 4,987,086 and 4,498,766.

[0080] In some cases, the flow cytometry system of the present invention is configured to image particles in a fluid medium by fluorescence imaging using radio frequency labeled emission (FIRE), for example, the system described in the following publication: Diebold et al., Nature Photonics, Vol. 7(10); 806-810 (2013); 9,423,353, 9,784,661, 9,983,132, 10,006,852, 10,078,045, 10,036,699, 10,2 U.S. Patents Nos. 22,316, 10,288,546, 10,324,019, 10,408,758, 10,451,538, and 10,620,111; and U.S. Patent Publications Nos. 2017 / 0133857, 2017 / 0328826, 2017 / 0350803, 2018 / 0275042, 2019 / 0376895, and 2019 / 0376894, the contents of which are incorporated herein by reference.

[0081] In some embodiments, the target system further includes a processor having a memory operatively coupled thereto, wherein the memory includes instructions stored thereon that, when executed by the processor, cause the processor to perform the following operations: clustering fluorescence flow cytometry data to determine the degree of spillover diffusion; adjusting the fluorescence flow cytometry data for spillover diffusion; and partitioning the spillover-adjusted fluorescence flow cytometry data.

[0082] In some embodiments, the processor is configured to generate one or more population clusters based on determined parameters of the analyte (e.g., cells, particles) in the sample. In embodiments, the fluorescence flow cytometry data includes signals from a variety of different fluorescent dyes, for example, 2 to 20 different fluorescent dyes, and includes 3 to 5 different fluorescent dyes. In some embodiments, the variety of different fluorescent dyes includes 2 or more different fluorescent dyes, including 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, and includes 20 or more different fluorescent dyes. Therefore, populations are identified as clusters in the data. Conversely, while clusters corresponding to noise or background are also typically observed, each data cluster is generally interpreted as corresponding to a specific type of cell or analyte population. Clusters can be defined in subsets of the dimension (e.g., subsets of the fluorescence parameters being measured (i.e., fluorescent dyes)) that correspond to groups that differ only in subsets of the measured parameters or features (extracted from the sample measurements).

[0083] In some embodiments, the processor includes instructions for supervised clustering of fluorescence flow cytometry data. Therefore, in some embodiments, the fluorescence flow cytometry data is clustered based on the state of each data point relative to a hierarchical structure. In some embodiments, the hierarchical structure establishes a way in which data points that are positive or negative for the same fluorescent dye are grouped together. In some embodiments, clustering the fluorescence flow cytometry data includes comparing the data to a threshold. In other words, fluorescence flow cytometry data above a threshold can be described as positive for a specific fluorescent dye, while fluorescence flow cytometry data below a threshold can be described as negative for the same fluorescent dye.

[0084] After clustering the flow cytometry data, the processor is configured to determine the extent of spillover diffusion in the fluorescence flow cytometry data population. In some embodiments, unadjusted flow cytometry data contains noise (i.e., spillover diffusion) because one or more detectors inadvertently detect light of certain wavelengths. In some embodiments, determining the extent of spillover diffusion includes quantifying the degree to which the intensity of light collected by a first detector from a first fluorophore is affected by light simultaneously collected by the same detector from a second fluorophore. In some embodiments, determining the extent of spillover diffusion includes calculating a spillover diffusion coefficient. In some embodiments, the spillover diffusion coefficient is calculated according to Formula 1:

[0085]

[0086] As shown in Formula 1, SS is the spillover diffusion coefficient, Δσ f Δd is the incremental standard deviation, representing the emission diffusion between the positive and negative fluorescence flow cytometry data collected from the fluorophore; Δd ​​is the difference in fluorescence intensity between the positive and negative fluorescence flow cytometry data received by the fluorescence detector. Therefore, the spillover diffusion coefficient measures the degree to which the fluorescence flow cytometry data acquired by a given fluorescence detector is affected by the presence of light associated with a specific fluorophore. In other words, the spillover diffusion coefficient estimates the error (i.e., noise) in the fluorescence flow cytometry data based on the light emitted by the relevant fluorophore collected by a given detector. In embodiments, for a given fluorophore-detector pair, a higher spillover diffusion coefficient indicates greater spillover diffusion.

[0087] In some embodiments, determining the extent of spillover diffusion further includes calculating spillover diffusion coefficients for each possible fluorescence detector-fluorochrome combination to determine how the fluorescence flow cytometry data collected at each detector is affected by the presence of light associated with each fluorochrome. In embodiments, the calculated spillover diffusion coefficients of each fluorescence detector-fluorochrome pair are combined in a spillover diffusion matrix. In some embodiments, the spillover diffusion matrix shows how the detection of a particular fluorochrome by its corresponding detector is affected by spillover from other fluorochromes. Cells where columns and rows intersect are filled with the calculated spillover diffusion coefficients of the fluorescence detector-fluorochrome pair, indicating the extent to which the investigated fluorochrome contributes to detection errors at the associated detector. The total influence of the fluorochrome on spillover diffusion can be estimated by summing all values ​​in its rows, while the total influence of spillover diffusion on the detector can be calculated by summing all values ​​in its columns. In some embodiments, the spillover diffusion coefficients are summed to calculate the total diffusion effect (i.e., the cumulative effect of spillover diffusion on a particular subset of fluorescence flow cytometry data).

[0088] In some embodiments, an overflow diffusion matrix is ​​calculated using the AutoSpread algorithm. The AutoSpread algorithm is configured to create an overflow diffusion matrix without needing to distinguish between positive and negative flow cytometry data populations relative to a given fluorophore (e.g., as described above). AutoSpread characterizes the diffusion of the detected signal from the first fluorophore by incorporating a second fluorophore into the same flow cytometry assay kit. AutoSpread generates a coefficient for each interaction between the fluorescence detector and the fluorophore and arranges these coefficients in a matrix similar to the overflow diffusion matrix described above. In embodiments, calculating the overflow diffusion coefficients involves assuming that the fluorescence intensity acquired by the fluorescence detector for the negative flow cytometry data population is zero and the corresponding standard deviation is unknown. In some embodiments, the overflow diffusion coefficients are calculated according to Equation 2:

[0089]

[0090] As shown in Equation 2, SS is the spillover diffusion coefficient, σ 2 It is the standard deviation of the positive fluorescence flow cytometry data population. is an estimate of the standard deviation of the negative fluorescence flow cytometry data population, where d is the intensity of light collected by the fluorescence detector. In some embodiments, to obtain an estimate of the standard deviation of the negative fluorescence flow cytometry data population, assuming that the fluorescence intensity collected by the fluorescence detector for the negative fluorescence flow cytometry data population is zero. The spillover diffusion coefficient was calculated using a series of linear regression methods. First, the fluorescence flow cytometry data were categorized by quantiles based on the intensity values ​​detected by the fluorescence detector. The default number of quantiles was 256, but this was reduced to 8 to ensure each quantile had a sufficient number of data points for reliable estimation of the standard deviation. Next, regression analysis was performed on the corresponding robust standard deviation of the light emitted by the fluorophore relative to the square root of the median intensity of light detected for each quantile. Assuming the intensity of light detected for the negative population was zero, the y-intercept of the ordinary least squares fit was considered an estimate of the standard deviation of the negative flow cytometry data population. A new zero-adjusted standard deviation was obtained using the estimate of the corresponding standard deviation of the light emitted by the fluorophore. Regression analysis was then performed on the zero-adjusted standard deviation of the fluorophore relative to the square root of the median fluorescence intensity detected by the fluorescence detector for each quantile. The slope of the ordinary least squares fit (calculated using Equation 2) was considered the spillover diffusion coefficient.

[0091] In some embodiments, the processor is configured to adjust fluorescence flow cytometry data to account for spillover diffusion. In some embodiments, the flow cytometry data is adjusted to eliminate all constructive errors caused by spillover diffusion. In embodiments, adjusting the fluorescence flow cytometry data includes generating different spillover-diffusion-adjusted populations. In some embodiments, generating different spillover-diffusion-adjusted populations includes reducing the spillover diffusion magnitude of the relevant populations of the flow cytometry data, i.e., offsetting the effect of increased signal due to constructive spillover diffusion errors. In some embodiments, the spillover diffusion magnitude is determined by the spillover diffusion matrix. In some embodiments, adjusting the flow cytometry data includes reducing the total diffusion effect of relevant portions of the flow cytometry data.

[0092] After adjusting a flow cytometry data population for spillover diffusion (e.g., as described above), the processor can be configured to partition the spillover diffusion-adjusted flow cytometry data population. Partitions are created to formally distinguish flow cytometry data populations to be classified differently (e.g., representing different phenotypes). In embodiments, partitioning different spillover diffusion-adjusted flow cytometry data populations includes calculating Matthews correlation coefficients. In embodiments, potential partitions between different spillover diffusion-adjusted flow cytometry data populations are evaluated relative to various thresholds. In such embodiments, Matthews correlation coefficients are calculated for each threshold to assess the level of consistency between the threshold and the partition. Therefore, embodiments of the invention involve partitioning populations to optimize the Matthews correlation coefficients relative to said correlation thresholds, which are capable of distinguishing flow cytometry data populations that are positive and negative relative to a particular fluorescent dye. In other words, partitioning flow cytometry data involves maximizing the ability to distinguish different populations (i.e., populations exhibiting different combinations of fluorescence parameters) depending on their relationship to a correlation threshold (e.g., quantified by the Matthews correlation coefficient). In one embodiment, the processor iterates this process to determine the optimal partitioning for distinguishing between populations that are positive and negative for relevant parameters of each flow cytometry data. In another embodiment, the Matthews correlation coefficient is calculated according to Formula 3:

[0093]

[0094] As shown in Formula 3, MCC is the Matthews correlation coefficient, TP represents a true positive event, TN represents a true negative event, FP represents a false positive event, and FN represents a false negative event. According to the present invention, a true positive event refers to flow cytometry data that is assessed as positive for a specific fluorescent dye based on a threshold and partition; a true negative event refers to flow cytometry data that is assessed as negative for a specific fluorescent dye based on a threshold and partition; a false positive event refers to flow cytometry data that is assessed as positive for a specific fluorescent dye based on partition but negative based on the threshold; a false negative event refers to flow cytometry data that is assessed as negative for a specific fluorescent dye based on partition but positive based on the threshold.

[0095] In some embodiments of the invention, the fluorescence flow cytometry data does not contain a signal that is positive for a specific fluorescent dye. In other words, no fluorescence emitted from said fluorescent dye is detected. In such embodiments, partitioning the different overflow-diffusion-adjusted populations of fluorescence flow cytometry data may include calculating the balance accuracy of each partition as a proxy for determining the optimal partition. The balance accuracy is the average of the accuracy in determining positive events and the accuracy in determining negative events, calculated according to Equation 4:

[0096]

[0097] As shown in Formula 4, BA is the balanced accuracy, TP represents true positive events, TN represents true negative events, FP represents false positive events, and FN represents false negative events.

[0098] Embodiments of the invention further include classifying the partitioned flow cytometry data populations, i.e., determining the subtypes of cells or particles specified by each distinct spillover-diffusion-adjusted flow cytometry data population. In embodiments, classification is determined based on the hierarchical structure (as described above). Therefore, classifications (i.e., phenotypic analysis) are assigned to the partitioned flow cytometry data populations (e.g., as described above) based on the combinations of fluorescence parameters they exhibit.

[0099] Figure 8A system 800 for flow cytometry according to an illustrative embodiment of the present invention is shown. System 800 includes a flow cytometer 810, a controller / processor 890, and a memory 895. The flow cytometer 810 includes one or more excitation lasers 815a-815c, a focusing lens 820, a flow chamber 825, a forward scatter detector 830, a side scatter detector 835, a fluorescence collecting lens 840, one or more beam splitters 845a-845g, one or more bandpass filters 850a-850e, one or more long-pass (“LP”) filters 855a-855b, and one or more fluorescence detectors 860a-860f.

[0100] The 815a-c laser is excited to emit light in the form of a laser beam. Figure 8 In the example system shown, the laser beams emitted from excitation lasers 815a-815c have wavelengths of 488 nm, 633 nm, and 325 nm, respectively. The laser beams are first guided through one or more beam splitters 845a and 845b. Beam splitter 845a transmits light with a wavelength of 488 nm and reflects light with a wavelength of 633 nm. Beam splitter 845b transmits UV light (light with a wavelength range of 10 to 400 nm) and reflects light with wavelengths of 488 nm and 633 nm.

[0101] The laser beam is then directed to a focusing lens 820, which focuses the beam onto a portion of the fluid medium containing the sample particles within a flow chamber 825. The flow chamber is part of a fluid system that guides particles (typically one at a time in the flow) to the focused laser beam for interrogation. The flow chamber may comprise a flow cell in a benchtop cytometer or a nozzle head in an air-flow cytometer.

[0102] Light from the laser beam interacts with the particles in the sample through diffraction, refraction, reflection, scattering, and absorption, and is re-emitted at various wavelengths depending on the characteristics of the particles (e.g., their size, internal structure, and the presence of one or more fluorescent molecules attached to or naturally present on or within the particles). The fluorescence emission, as well as the diffracted, refracted, reflected, and scattered light, can be directed to one or more of the following: a beam splitter 845a-845g, a bandpass filter 850a-850e, a long-pass filter 855a-855b, and a fluorescence collecting lens 840. The light is then directed to one or more of the following: a forward scattering detector 830, a side scattering detector 835, and one or more fluorescence detectors 860a-860f.

[0103] A fluorescence collecting lens 840 collects light emitted through particle-laser beam interaction and directs the light to one or more beam splitters and filters. Bandpass filters (e.g., bandpass filters 850a-850e) allow a narrow range of wavelengths to pass through. For example, bandpass filter 850a is a 510 / 20 filter. The first number represents the center of the spectral band. The second number provides the range of the spectral band. Thus, on each side of the center of the spectral band, the 510 / 20 filter increases the wavelength by 10 nm, or increases the wavelength from 500 nm to 520 nm. Short-pass filters transmit light with wavelengths equal to or less than a specified wavelength. Long-pass filters (e.g., long-pass filters 855a-855b) transmit light with wavelengths equal to or longer than a specified wavelength. For example, long-pass filter 855a (a 670 nm long-pass filter) transmits light with wavelengths equal to or longer than 670 nm. Filters are typically selected to optimize the detector's specificity for a particular fluorescent dye. The filter can be configured such that the spectral band of the light transmitted to the detector is close to the emission peak of the fluorescent dye.

[0104] Beam splitters direct light of different wavelengths in different directions. Beam splitters can be characterized by filter properties, such as short-pass and long-pass. For example, beam splitter 805g is a 620SP beam splitter, meaning that beam splitter 845g transmits light with wavelengths of 620 nm or shorter and reflects light with wavelengths longer than 620 nm in different directions. In one embodiment, beam splitters 845a-845g may include optical mirrors, such as dichroic mirrors.

[0105] A forward scattering detector 830 is positioned away from the direct beam of light passing through the flow cell and is configured to detect diffracted light, primarily forward-directed excitation light traveling through or around the particle. The intensity of the light detected by the forward scattering detector depends on the overall size of the particle. The forward scattering detector may include a photodiode. A side scattering detector 835 is configured to detect refracted and reflected light from the particle surface and internal structure, and its number tends to increase with increasing particle structural complexity. One or more fluorescence detectors 860a-860f can detect fluorescence emission from fluorescent molecules associated with the particle. The side scattering detector 835 and the fluorescence detector may include photomultiplier tubes. The signals detected at the forward scattering detector 830, the side scattering detector 835, and the fluorescence detector can be converted into electronic signals (voltages) by the detectors. This data can provide information about the sample.

[0106] During operation, the cytometer is controlled by a controller / processor 890, and measurement data from the detector can be stored in memory 895 and processed by the controller / processor 890. Although not explicitly shown, the controller / processor 890 is coupled to the detector to receive the output signal therefrom, and can also be coupled to the electrical and electromechanical components of the flow cytometer 800 to control the laser, fluid flow parameters, etc. Input / output (I / O) functionality 897 can also be provided in the system. Memory 895, controller / processor 890, and I / O 897 can be provided as integral components of the flow cytometer 810. In such embodiments, a display can also be part of the I / O functionality 897 for presenting experimental data to the user of the cytometer 800. Alternatively, some or all of memory 895 and controller / processor 890 and I / O functionality can be part of one or more external devices (e.g., a general-purpose computer). In some embodiments, some or all of memory 895 and controller / processor 890 can communicate wirelessly or wiredly with the cytometer 810. The controller / processor 890, together with the memory 895 and I / O 897, can be configured to perform various functions related to the preparation and analysis of flow cytometry experiments.

[0107] Figure 8The system shown includes six different detectors capable of detecting fluorescence in six different wavelength bands (which may be referred to herein as “filter windows” for a given detector), as defined by the configuration of filters and / or separators in the beam path between flow cell 825 and each detector. Different fluorescent molecules used in flow cytometry experiments emit light within their own characteristic wavelength bands. The specific fluorescent label used for the experiment and its associated fluorescence emission band can be selected to generally correspond to the filter windows of the detectors. However, with more detectors and more labels, a perfect correspondence between filter windows and fluorescence emission spectra cannot be achieved. In fact, while the peak of the emission spectrum of a particular fluorescent molecule may lie within the filter window of a particular detector, some emission spectra of the label may also overlap with the filter windows of one or more other detectors. This can be referred to as spillover. I / O 897 can be configured to receive data for flow cytometry experiments having a set of fluorescent labels and multiple cell populations (with multiple markers), each cell population having a subset of multiple markers. I / O 897 can also be configured to receive biological data (involving the assignment of one or more markers to one or more cell populations), marker density data, emission spectral data, data relating to the assignment of markers to one or more markers, and cytometer configuration data. Flow cytometry experimental data (e.g., marker spectral characteristics and flow cytometry configuration data) can also be stored in memory 895. Controller / processor 890 can be configured to evaluate the assignment of one or more markers to the markers.

[0108] Those skilled in the art will recognize that the flow cytometer according to embodiments of the present invention is not limited to Figure 8 The flow cytometer shown can include any flow cytometer known in the art. For example, a flow cytometer can have any number of lasers, beam splitters, filters, and detectors, which have various wavelengths and various different configurations.

[0109] Figure 9A functional block diagram illustrating an example of a processor 900 for analyzing and displaying data is shown. The processor 900 may be configured to implement various processes for controlling the graphical display of biological events. A flow cytometer 902 may be configured to acquire fluorescence flow cytometry data by analyzing biological samples (e.g., as described above). The device may be configured to provide biological event data to the processor 900. A data communication channel may be incorporated between the flow cytometer 902 and the processor 900. The data may be provided to the processor 900 via the data communication channel. The processor 900 may be configured to provide a graphical display including graphs (e.g., as described above) to a display 906. For example, the processor 900 may be further configured to gate a population of fluorescence flow cytometry data shown on the display device 906, overlaying the graph. In some embodiments, the gate may be a logical combination of one or more target graphical regions plotted on a single-parameter histogram or a bivariate graph. In some embodiments, the display may be used to display analyte parameters or saturation detector data.

[0110] The processor 900 can be further configured to display the fluorescence flow cytometry data on the display device 906 inside the gate in a manner different from other events in the fluorescence flow cytometry data outside the gate. For example, the processor 900 can be configured to make the color of the fluorescence flow cytometry data contained inside the gate different from the color of the fluorescence flow cytometry data outside the gate. In this way, the processor 900 can be configured to display different colors to represent each unique data group. The display device 906 can be implemented in the form of a monitor, tablet computer, smartphone, or other electronic device configured to display a graphical interface.

[0111] Processor 900 may be configured to receive a door selection signal from a first input device to identify the door. For example, the first input device may be implemented in the form of a mouse 910. Mouse 910 may issue a door selection signal to processor 900 to determine a group to be displayed on or manipulated via display device 906 (e.g., clicking on or inside the desired door when the cursor is present). In some embodiments, the first device may be implemented in the form of a keyboard 908 or other means for providing input signals to processor 900 (e.g., a touchscreen, stylus, optical detector, or voice recognition system). Some input devices may include multiple input functions. In such embodiments, each of the input functions may be considered as an input device. For example, such as... Figure 9 As shown, mouse 910 may include a right mouse button and a left mouse button, both of which can generate trigger events.

[0112] The triggering event may cause the processor 900 to change the way it displays fluorescence flow cytometry data (actually displaying a portion of the data on the display device 906), and / or provide input for further processing, such as selecting a target population for analysis.

[0113] In some embodiments, processor 900 may be configured to detect the timing of mouse 910 initiating a gate selection. Processor 900 may be further configured to automatically modify the graph visualization to facilitate the gate selection process. The modification may be based on a specific distribution of data received by processor 900.

[0114] Processor 900 may be connected to storage device 904. Storage device 904 may be configured to receive and store data from processor 900. Storage device 904 may also be configured to allow processor 900 to retrieve data, such as fluorescence flow cytometry data.

[0115] Display device 906 may be configured to receive display data from processor 900. The display data may include a fluorescence flow cytometry data graph and gates summarizing portions of the graph. Display device 906 may be further configured to change the displayed information based on input received from processor 900 and input received from device 902, storage device 904, keyboard 908, and / or mouse 910.

[0116] In some implementations, the processor 900 may generate a user interface to receive example events for sorting. For example, the user interface may include controls for receiving example events or example images. Example events, images, or example gates may be provided before acquiring event data for the sample or based on an initial set of events for a portion of the sample.

[0117] Computer control system

[0118] The invention further includes a computer control system, wherein the system further includes one or more computers for achieving full or partial automation. In some embodiments, the system includes a computer having a computer-readable storage medium storing a computer program, wherein the computer program includes instructions, when loaded onto the computer, to perform the following operations: clustering fluorescence flow cytometry data into populations based on one or more different parameters (i.e., fluorescent dyes); determining spillover diffusion between detector-parameter pairs (i.e., determined by calculating spillover diffusion coefficients); creating a spillover diffusion matrix showing how the detection of a particular parameter by a corresponding detector is affected by spillover of other parameters; modifying the fluorescence flow cytometry data to compensate for spillover diffusion by reducing the spillover diffusion magnitude determined by the spillover diffusion matrix; evaluating the quality of separating different partitions of different fluorescence flow cytometry data populations by calculating a Matthews correlation coefficient relative to a threshold (used to distinguish between populations positive and negative for a given parameter); and classifying the adjusted fluorescence flow cytometry data populations (i.e., phenotypic analysis).

[0119] In an embodiment, the system is configured to analyze data within analysis software or tools used to analyze flow cytometry data or nucleic acid sequence data, for example. (Ashland, Oregon). FlowJo is a software package developed by FlowJo LLC (a subsidiary of Becton Dickinson) for analyzing flow cytometry data. The software is configured to manage flow cytometry data and generate graphical reports on it (https: / / www.flowjo.com / learn / flowjo-university / flowjo). It can be integrated into data analysis software or tools (e.g., ...) in an appropriate manner. The initial data is analyzed within the system, using methods such as manual gating, cluster analysis, or other computational techniques. The system of the present invention, or parts thereof, can be implemented as software components for data analysis, for example... In these embodiments, the computer control system according to the invention can be used as a software package suitable for existing software packages (e.g. ) software "plugins".

[0120] In an embodiment, the system includes an input module, a processing module, and an output module. The target system may include hardware and software components, wherein the hardware components may be one or more platforms, such as servers, that enable the functional elements of the system (i.e., elements in the system that perform specific tasks (e.g., managing the input and output of information, processing information, etc.)) to function by executing software applications on one or more computer platforms on which the system is equipped.

[0121] The system may include a display and operator input devices. Operator input devices may be a keyboard, mouse, etc. The processing module includes a processor that can access memory having instructions stored thereon for performing the target method steps. The processing module may include an operating system, a graphical user interface (GUI) controller, system memory, memory storage devices, input / output controllers, cache memory, data backup units, and many other devices. The processor may be a commercially available processor or one of other processors that are already available or will be available in the future. As is known in the art, the processor executes an operating system, which is connected to firmware and hardware in a well-known manner and helps the processor coordinate and execute the functions of various computer programs written in various programming languages, such as Java, Perl, C++, other high-level or low-level languages, and combinations thereof. The operating system typically works with the processor to coordinate and execute the functions of other computer components. The operating system also provides scheduling, input / output control, file and data management, memory management, communication control, and related services according to known techniques. The processor may be any suitable analog or digital system. In some embodiments, the processor includes analog electronics that allow a user to manually align the light source with the fluid medium based on the first and second optical signals. In some embodiments, the processor includes analog electronics that provide feedback control (e.g., negative feedback control).

[0122] The system memory can be any of a variety of known or future memory storage devices. Examples include any generally available random access memory (RAM), magnetic media (e.g., resident hard disks or magnetic tapes), optical media (e.g., optical discs), flash memory devices, or other memory storage devices. The memory storage device can be any of a variety of known or future devices, including optical disc drives, magnetic tape drives, removable hard disk drives, or floppy disk drives. Memory storage devices of this type typically read content from and / or write content to program storage media (not shown), including optical discs, magnetic tapes, removable hard disks, or floppy disks. Any of these program storage media, or other media currently in use or that may be developed in the future, can be considered a computer program product. It is understood that these program storage media typically store computer software programs and / or data. Computer software programs, also known as computer control logic, are typically stored in system memory and / or program storage devices used in conjunction with memory storage devices.

[0123] In some embodiments, a computer program product is described that includes a computer-usable medium storing control logic (computer software program, including program code). When executed by a computer processor, the control logic enables the processor to perform the functions described herein. In other embodiments, some functions are implemented primarily in hardware using a hardware state machine. Enabling a hardware state machine to perform the functions described herein will be apparent to those skilled in the art.

[0124] The memory can be any suitable device in which the processor can store and retrieve data, such as magnetic, optical, or solid-state storage devices (including disks, optical discs, magnetic tapes, or RAM, or any other suitable fixed or portable device). The processor may include a general-purpose digital microprocessor that has been appropriately programmed based on a computer-readable medium carrying the necessary program code. The programmed information can be provided to the processor remotely via a communication channel or pre-stored in a computer program product using any device connected to the memory, such as memory or certain other portable or fixed computer-readable storage media. For example, a disk or optical disc can carry the programmed information and can be read using a disk writer / reader. The system of the present invention also includes a programmed information for implementing the methods described above, such as a computer program product or algorithm. The programmed information according to the present invention can be recorded in a computer-readable medium, such as any medium that can be directly read and accessed by a computer. Such media include, but are not limited to: magnetic storage media, such as floppy disks, hard disk storage media, and magnetic tapes; optical storage media, such as CD-ROMs; electrical storage media, such as RAM and ROMs; portable flash drives; and hybrids of these categories, such as magnetic / optical storage media.

[0125] The processor can also access communication channels to communicate with users located in remote locations. A remote location refers to a location where the user has no direct contact with the system but instead forwards input information from external devices (e.g., connected to a wide area network (“WAN”), telephone network, satellite network, or any other suitable communication channel, including mobile phones (i.e., smartphones)) to the input manager.

[0126] In some embodiments, the system according to the invention can be configured to include a communication interface. In some embodiments, the communication interface includes a receiver and / or transmitter for communicating with a network and / or another device. The communication interface can be configured for wired or wireless communication, including but not limited to: radio frequency (RF) communication (e.g., radio frequency identification (RFID), Zigbee communication protocol, WiFi, infrared communication, wireless universal serial bus (USB), ultra-wideband (UWB)). Communication protocols and cellular communications, such as Code Division Multiple Access (CDMA) or Global System for Mobile Communications (GSM).

[0127] In one embodiment, the communication interface is configured to include one or more communication ports, such as physical ports or interfaces (e.g., USB ports, RS-232 ports) or any other suitable electrical connection ports, to enable data communication between the target system and any external device (e.g., a computer terminal configured to enable similar complementary data communication (e.g., in a physician's office or hospital environment)).

[0128] In one embodiment, the communication interface is configured for infrared communication. Communication or any other suitable wireless communication protocol that enables the target system to communicate with other devices, such as computer terminals and / or networks, communication-enabled mobile phones, personal digital assistants, or any other communication devices that the user can use in conjunction with them.

[0129] In one embodiment, the communication interface is configured to provide data transmission connectivity via mobile phone networks or SMS service using Internet Protocol (IP); to provide wireless connectivity to personal computers (PCs) within a local area network (LAN) connected to the Internet; or to provide WiFi connectivity to a WiFi hotspot for connecting to the Internet.

[0130] In one embodiment, the target system is configured to wirelessly communicate with a server device via a communication interface, for example, using a common standard such as 802.11 or... The server device may be an RF protocol or an IrDA infrared protocol. It can also be another portable device, such as a smartphone, personal digital assistant (PDA), or laptop; or a larger device, such as a desktop computer, instrument, etc. In some embodiments, the server device includes a display (e.g., a liquid crystal display (LCD)) and input devices (e.g., buttons, keyboard, mouse, or touchscreen).

[0131] In some embodiments, the communication interface is configured to automatically or semi-automatically communicate data stored in the target system (e.g., stored in an optional data storage unit) with a network or server device using one or more of the communication protocols and / or mechanisms described above.

[0132] The output controller may include a controller for any of a variety of known display devices to provide information to local or remote users (whether human or machine). If one of the display devices provides visual information, the information may typically be logically and / or physically organized into an array of image elements. The graphical user interface (GUI) controller may include any of a variety of known or future-developed software programs to provide a graphical input and output interface between the system and the user, and to process user input. Functional elements of the computer may communicate with each other via a system bus. Some of this communication may be implemented using a network or other types of remote communication in alternative embodiments. The output manager may also provide information generated by the processing module to a user at a remote location, for example, via the Internet, telephone, or satellite networks, according to known technologies. The output manager may implement data display according to a variety of known technologies. In some examples, the data may include SQL, HTML, or XML documents, emails, or other files or other forms of data. The data may include Internet URLs so that the user can retrieve other SQL, HTML, XML, or other documents or data from a remote source. One or more platforms in the target system may be any type of known or future-developed computer platform, although they typically belong to a certain class of computers (often referred to as servers). However, the platform can also be a mainframe computer, workstation, or other computer type. They can be connected via any known or future-to-be-developed cable or other communication system (including wireless systems connected by networking or other means). They can be located in the same location or physically separated. Various operating systems can be used on any computer platform, depending on the type and / or brand of the chosen platform. Suitable operating systems include Windows NT, Windows XP, Windows 7, Windows 8, iOS, Sun Solaris, Linux, OS / 400, Compaq Tru64 Unix, SGI IRIX, Siemens Reliant Unix, etc.

[0133] Figure 10 The overall architecture of an example computing device 1000 according to certain embodiments is described. Figure 10The overall architecture of the computing device 1000 shown includes the arrangement of computer hardware and software components. However, it is not necessary to show all such conventional components when providing implementation disclosure. As shown, the computing device 1000 includes a processing unit 1010, a network interface 1020, a computer-readable media drive 1030, an input / output device interface 1040, a display 1050, and an input device 1060, all of which can communicate with another device via a communication bus. The network interface 1020 can be connected to one or more networks or computing systems. The processing unit 1010 can therefore receive information and instructions from other computing systems or services via the network. The processing unit 1010 can also communicate with a memory 1070 and can further provide output information to an optional display 1050 via the input / output device interface 1040. For example, analysis software (e.g., data analysis software or programs, such as...) stored as executable instructions in the non-transitory memory of an analysis system. It can display flow cytometry event data to the user. The input / output device interface 1040 can also accept input from optional input devices 1060, such as keyboards, mice, digital pens, microphones, touch screens, gesture recognition systems, voice recognition systems, game controllers, accelerometers, gyroscopes, or other input devices.

[0134] Memory 1070 may contain computer program instructions (grouped into modules or components in some embodiments) executed by processing unit 1010 to implement one or more embodiments. Memory 1070 typically includes RAM, ROM, and / or other persistent, auxiliary, or non-transitory computer-readable media. Memory 1070 may store operating system 1072, which provides computer program instructions for use by processing unit 1010 in the routine management and operation of computing device 1000. Data may be stored in data storage device 1090. Memory 1070 may further include computer program instructions and other information for implementing aspects of the invention.

[0135] Computer-readable storage media

[0136] The invention further includes a non-transitory computer-readable storage medium having instructions for implementing the target method. The computer-readable storage medium can be employed on one or more computers, enabling a system for implementing the methods described herein to be fully or partially automated. In some embodiments, instructions according to the methods described herein can be encoded in a “programmed” form into a computer-readable medium, wherein the term “computer-readable medium” as used herein refers to any non-transitory storage medium that participates in providing instructions and data to a computer for execution and processing. Examples of suitable non-transitory storage media include floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, DVD-ROMs, Blu-ray discs, solid-state drives, and network-attached storage devices (NAS), whether such devices are internal or external to a computer. In some cases, instructions can be provided on an integrated circuit device. In some cases, the target integrated circuit device may include a reconfigurable field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or a complex programmable logic device (CPLD). Files containing information can be “stored” on a computer-readable medium, where “stored” means recording information so that a computer can access and retrieve it later. The computer implementations described herein can be implemented using programming languages ​​that can be written in one or more of any number of computer programming languages. Such languages ​​include, for example, Java (Sun Microsystems, Inc., Santa Clara, California), Visual Basic (Microsoft Corp., Redmond, Washington), and C++ (AT&T Corp., Bedminster, New Jersey), among many others.

[0137] In some embodiments, the intended computer-readable storage medium includes a computer program stored thereon, wherein the computer program includes instructions, when loaded onto the computer, to perform the following operations: clustering fluorescence flow cytometry data into populations based on one or more different parameters; determining spillover diffusion between detector-parameter pairs (i.e., determined by calculating spillover diffusion coefficients); creating a spillover diffusion matrix showing how the detection of a particular parameter by a corresponding detector is affected by spillover from other parameters; modifying the fluorescence flow cytometry data to compensate for spillover diffusion by reducing the spillover diffusion magnitude determined by the spillover diffusion matrix; evaluating the quality of separating different partitions of different fluorescence flow cytometry data populations by calculating a Matthews correlation coefficient relative to a threshold (used to distinguish between populations positive and negative for a given parameter); and classifying the adjusted fluorescence flow cytometry data populations (i.e., phenotypic analysis).

[0138] In an embodiment, the system is configured to analyze data within analysis software or tools used to analyze flow cytometry data or nucleic acid sequence data, for example. Appropriate methods can be used in data analysis software or tools (e.g., The initial data is analyzed within the system, using methods such as manual gating, cluster analysis, or other computational techniques. The system of the present invention, or parts thereof, can be implemented as software components for data analysis, for example... In these embodiments, the computer control system according to the invention can be used as a software package suitable for existing software packages (e.g. ) software "plugins".

[0139] The computer-readable storage medium can be used in more or more computer systems having displays and operator input devices. Operator input devices may be keyboards, mice, etc. The processing module includes a processor that can access memory having instructions stored thereon for performing the target method steps. The processing module may include an operating system, a graphical user interface (GUI) controller, system memory, memory storage devices, input / output controllers, cache memory, data backup units, and many other devices. The processor may be a commercially available processor or one of other processors that are already available or will be available in the future. As is known in the art, the processor executes an operating system, which is connected to firmware and hardware in a well-known manner and helps the processor coordinate and execute the functions of various computer programs written in various programming languages, such as Java, Perl, Python, C++, other high-level or low-level languages, and combinations thereof. The operating system also provides scheduling, input / output control, file and data management, memory management, communication control, and related services according to known techniques.

[0140] utility

[0141] The target apparatus, method, and computer system can be used in a variety of applications requiring improved resolution and accuracy when determining parameters of analytes (e.g., cells, particles) in biological samples. For example, the invention can be used to analyze data affected by spillover diffusion. Since flow cytometry typically involves collecting multiple fluorescence parameters through multiple detectors, the detected fluorescence intensity can be erroneously increased due to multiple detectors detecting the same light. Therefore, the invention can be used during the analysis of flow cytometry data containing signals from multiple fluorescent dyes. The target apparatus, method, and computer system can also be used to classify (i.e., phenotypic analyze) flow cytometry data populations that are typically mischaracterized due to spillover diffusion. In some embodiments, the target method and system provide a fully automated scheme, requiring minimal data adjustment and, if any, manual input.

[0142] This invention can be used to characterize many types of analytes, especially those related to medical diagnostics or patient care protocols, including but not limited to: proteins (including free proteins and surface-bound proteins, such as those found in cells), nucleic acids, viral particles, etc. Furthermore, samples can be from in vitro or in vivo sources, and samples can be diagnostic samples.

[0143] kit

[0144] The invention further includes a kit, wherein the kit includes storage media such as floppy disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, DVD-ROMs, Blu-ray discs, solid-state drives, and network-attached storage devices (NAS). Any of these program storage media, or other media currently in use or that may be developed later, may be included in the target kit. In an embodiment, the program storage media includes instructions for performing the following: clustering fluorescence flow cytometry data into populations; determining spillover diffusion of the populations; adjusting the flow cytometry data based on spillover diffusion; and determining partitions between the adjusted flow cytometry data (e.g., as described above). In an embodiment, the instructions contained in a computer-readable medium provided with the target kit or a portion thereof may be implemented in the form of a software component for analyzing data, such as… In these embodiments, the computer control system according to the invention can be used as a software package suitable for existing software packages (e.g. ) software "plugins".

[0145] In addition to the components described above, the target kit may further include (in some embodiments) features such as those for installing plugins into existing software packages (e.g., The instructions may be in various forms within the target kit, one or more of which may be present in the kit. One form of these instructions may be printed information on suitable media or substrates (e.g., a single sheet or several sheets of paper with information printed on them), kit packaging, packaging instructions, etc. Another form of these instructions may be computer-readable media on which information is already recorded, such as floppy disks, optical discs (CDs), portable flash drives, etc. Yet another form of these instructions may be a URL, allowing access to information on a remote website via the Internet.

[0146] Although the claims are attached herein, the scope of the invention is also limited by the following terms:

[0147] 1. A method for classifying fluorescence flow cytometry data, the method comprising:

[0148] The flow cytometry data are processed using a supervised algorithm, which is configured to:

[0149] The fluorescence flow cytometry data were clustered into a population;

[0150] Determine the extent of spillover diffusion of the fluorescence flow cytometry data population;

[0151] Based on the determined spillover diffusion, the fluorescence flow cytometry data population is adjusted to generate different spillover diffusion-adjusted populations; and

[0152] Partitions are established between the different overflow-diffusion-adjusted fluorescence flow cytometry data populations to classify the different overflow-diffusion-adjusted fluorescence flow cytometry data populations.

[0153] 2. The method according to Clause 1, wherein the fluorescence flow cytometry data are clustered into groups based on whether the fluorescence flow cytometry data are positive or negative relative to a specific fluorescent dye.

[0154] 3. The method according to Clause 1 or 2, wherein the fluorescence flow cytometry data are determined to be positive or negative for a specific fluorescent dye based on the relationship between the fluorescence flow cytometry data and a threshold.

[0155] 4. The method according to any one of clauses 1 to 3, wherein determining spillover diffusion comprises quantifying the extent to which fluorescence flow cytometry data acquired by a fluorescence detector increases due to the collection of light emitted by a specific fluorescent dye.

[0156] 5. The method according to any one of the preceding clauses, wherein determining the spillover diffusion comprises calculating the spillover diffusion coefficient of the fluorescence detector-fluorescent dye pair.

[0157] 6. The method according to Clause 5, wherein the spillover diffusion coefficient is calculated according to Formula 1:

[0158]

[0159] in:

[0160] SS is the spillover diffusion coefficient;

[0161] Δσ f It is the incremental standard deviation, representing the emission diffusion between the positive and negative fluorescence flow cytometry data collected from the fluorescent dye; and

[0162] Δd is the difference in fluorescence intensity between the positive and negative fluorescence flow cytometry data received by the fluorescence detector.

[0163] 7. The method according to Clause 5, wherein calculating the spillover diffusion coefficient includes assuming that the fluorescence intensity acquired by the fluorescence detector for the negative flow cytometry data population is zero.

[0164] 8. The method according to Clause 7, wherein the spillover diffusion coefficient is calculated according to Formula 2:

[0165]

[0166] in:

[0167] SS is the spillover diffusion coefficient;

[0168] σ 2 It is the standard deviation of the positive fluorescence flow cytometry data population;

[0169] It is an estimate of the standard deviation of the negative fluorescence flow cytometry data population; and

[0170] d is the intensity of the light collected by the fluorescence detector.

[0171] 9. The method according to Clause 8, wherein the calculation is performed by linear regression.

[0172] 10. The method according to any one of the preceding clauses, wherein the fluorescence flow cytometry data are collected from light emitted by a variety of different fluorescent dyes.

[0173] 11. The method according to Clause 10, wherein the number of the plurality of different fluorescent dyes ranges from 2 to 20 different fluorescent dyes.

[0174] 12. The method according to clause 10 or 11, wherein the number of the plurality of different fluorescent dyes ranges from 3 to 5 different fluorescent dyes.

[0175] 13. The method according to any one of clauses 10 to 12, wherein the overflow diffusion coefficient of each fluorescence detector-fluorescent dye pair is calculated.

[0176] 14. The method according to any one of clauses 10 to 13, wherein the overflow diffusion coefficients of each fluorescent detector-fluorescent dye pair calculated in the overflow diffusion matrix are combined.

[0177] 15. The method according to Clause 14, wherein the method further comprises calculating the spillover diffusion magnitude of each of the plurality of different fluorescent dyes based on the spillover diffusion matrix.

[0178] 16. The method according to Clause 15, wherein adjusting the fluorescence flow cytometry data comprises reducing the spillover diffusion magnitude of each of the plurality of different fluorescent dyes in the fluorescence flow cytometry data population corresponding to the fluorescent dye.

[0179] 17. The method according to any one of the preceding clauses, wherein establishing partitions among different spillover-diffusion-adjusted fluorescence flow cytometry data populations comprises evaluating the differences between the different spillover-diffusion-adjusted populations relative to a threshold.

[0180] 18. The method according to Clause 17, wherein evaluating the differences between different spillover-adjusted populations relative to a threshold includes calculating the Matthews correlation coefficient.

[0181] 19. The method according to Clause 18, wherein the fluorescence flow cytometry data does not contain data showing positivity for at least one fluorescent dye.

[0182] 20. The method according to Clause 19, wherein evaluating the differences between different spillover-adjusted populations relative to a threshold includes calculating equilibrium accuracy.

[0183] 21. The method according to any one of the preceding clauses, wherein the different overflow-diffusion-adjusted fluorescence flow cytometry data populations are partitioned according to a hierarchical structure.

[0184] 22. The method according to Clause 21, wherein the hierarchy details the association between spillover diffusion-adjusted populations of fluorescence flow cytometry data that are positive or negative for a particular fluorescent dye and the corresponding phenotype.

[0185] 23. A system comprising:

[0186] A particle analyzer component configured to acquire fluorescence flow cytometry data; and

[0187] A processor includes memory operatively coupled thereto, wherein the memory contains instructions stored thereon, which, when executed by the processor, cause the processor to perform the following operations:

[0188] The fluorescence flow cytometry data were clustered into a population;

[0189] Determine the extent of spillover diffusion of the fluorescence flow cytometry data population;

[0190] Based on the determined spillover diffusion, the fluorescence flow cytometry data population is adjusted to generate different spillover diffusion-adjusted populations; and

[0191] Partitions are established between the different overflow-diffusion-adjusted fluorescence flow cytometry data populations to classify the different overflow-diffusion-adjusted fluorescence flow cytometry data populations.

[0192] 24. The system according to Clause 23, wherein the fluorescence flow cytometry data are clustered into groups based on whether the fluorescence flow cytometry data are positive or negative relative to a specific fluorescent dye.

[0193] 25. The system according to Clause 23 or 24, wherein the fluorescence flow cytometry data is determined to be positive or negative for a specific fluorescent dye based on the relationship between the fluorescence flow cytometry data and a threshold.

[0194] 26. The system according to any one of clauses 23 to 25, wherein determining spillover diffusion comprises quantifying the extent to which fluorescence flow cytometry data acquired by a fluorescence detector increases due to the collection of light emitted by a specific fluorescent dye.

[0195] 27. The system according to any one of clauses 23 to 26, wherein determining spillover diffusion comprises calculating the spillover diffusion coefficient of the fluorescence detector-fluorescent dye pair.

[0196] 28. The system according to Clause 27, wherein the spillover diffusion coefficient is calculated according to Formula 1:

[0197]

[0198] in:

[0199] SS is the spillover diffusion coefficient;

[0200] Δσ f It is the incremental standard deviation, representing the emission diffusion between the positive and negative fluorescence flow cytometry data collected from the fluorescent dye; and

[0201] Δd is the difference in fluorescence intensity between the positive and negative fluorescence flow cytometry data received by the fluorescence detector.

[0202] 29. The system according to Clause 27, wherein calculating the spillover diffusion coefficient includes assuming that the fluorescence intensity acquired by the fluorescence detector for the negative flow cytometry data population is zero.

[0203] 30. The system according to Clause 29, wherein the spillover diffusion coefficient is calculated according to Formula 2:

[0204]

[0205] in:

[0206] SS is the spillover diffusion coefficient;

[0207] σ 2 It is the standard deviation of the positive fluorescence flow cytometry data population;

[0208] It is an estimate of the standard deviation of the negative fluorescence flow cytometry data population; and

[0209] d is the intensity of the light collected by the fluorescence detector.

[0210] 31. The system according to Clause 30, wherein the calculation is performed by linear regression.

[0211] 32. The system according to any one of clauses 23 to 31, wherein the fluorescence flow cytometry data are collected from light emitted by a variety of different fluorescent dyes.

[0212] 33. The system according to Clause 32, wherein the number of the plurality of different fluorescent dyes ranges from 2 to 20 different fluorescent dyes.

[0213] 34. The system according to clause 32 or 33, wherein the number of the various fluorescent dyes ranges from 3 to 5 different fluorescent dyes.

[0214] 35. The system according to any one of clauses 32 to 34, wherein the overflow diffusion coefficient of each fluorescence detector-fluorescent dye pair is calculated.

[0215] 36. The system according to any one of clauses 32 to 35, wherein the overflow diffusion coefficients of each fluorescence detector-fluorescent dye pair calculated in the overflow diffusion matrix are combined.

[0216] 37. The system according to Clause 36 further comprises calculating the spillover diffusion magnitude of each of the plurality of different fluorescent dyes based on the spillover diffusion matrix.

[0217] 38. The system according to Clause 37, wherein adjusting the fluorescence flow cytometry data comprises reducing the spillover diffusion magnitude of each of the various fluorescent dyes in the fluorescence flow cytometry data population corresponding to the fluorescent dye.

[0218] 39. The system according to any one of clauses 23 to 38, wherein establishing partitions between different spillover-diffusion-adjusted fluorescence flow cytometry data populations comprises evaluating the differences between the different spillover-diffusion-adjusted populations relative to a threshold.

[0219] 40. The system according to Clause 39, wherein evaluating the differences between different spillover-adjusted populations relative to a threshold includes calculating the Matthews correlation coefficient.

[0220] 41. The system according to Clause 40, wherein the fluorescence flow cytometry data does not contain data showing positivity for at least one fluorescent dye.

[0221] 42. The system according to Clause 41, wherein evaluating the differences between different spillover-adjusted populations relative to a threshold includes calculating balance accuracy.

[0222] 43. The system according to any one of clauses 23 to 42, wherein the different overflow-diffusion-adjusted fluorescence flow cytometry data populations are partitioned according to a hierarchical structure.

[0223] 44. The system according to Clause 43, wherein the hierarchy details the association between spillover diffusion-adjusted populations of fluorescence flow cytometry data that are positive or negative for a particular fluorescent dye and the corresponding phenotype.

[0224] 45. A non-transitory computer-readable storage medium comprising instructions stored thereon for classifying flow cytometry data by a method comprising:

[0225] The flow cytometry data are processed using a supervised algorithm, which is configured to:

[0226] The fluorescence flow cytometry data were clustered into a population;

[0227] Determine the extent of spillover diffusion of the fluorescence flow cytometry data population;

[0228] Based on the determined spillover diffusion, the fluorescence flow cytometry data population is adjusted to generate different spillover diffusion-adjusted populations; and

[0229] Partitions are established between the different overflow-diffusion-adjusted fluorescence flow cytometry data populations to classify the different overflow-diffusion-adjusted fluorescence flow cytometry data populations.

[0230] 46. ​​The non-transitory computer-readable storage medium as described in Clause 45, wherein the fluorescence flow cytometry data are clustered into groups based on whether the fluorescence flow cytometry data are positive or negative relative to a specific fluorescent dye.

[0231] 47. A non-transitory computer-readable storage medium as described in Clause 45 or 46, wherein the fluorescence flow cytometry data is determined to be positive or negative for a specific fluorescent dye based on the relationship between the fluorescence flow cytometry data and a threshold.

[0232] 48. The non-transitory computer-readable storage medium according to any one of clauses 45 to 47, wherein determining spillover diffusion includes quantifying the extent to which fluorescence flow cytometry data acquired by a fluorescence detector increases due to the collection of light emitted by a specific fluorescent dye.

[0233] 49. The non-transitory computer-readable storage medium according to any one of clauses 45 to 48, wherein determining spill diffusion comprises calculating the spill diffusion coefficient of the fluorescence detector-fluorescent dye pair.

[0234] 50. The non-transitory computer-readable storage medium as described in Clause 49, wherein the spillover diffusion coefficient is calculated according to Formula 1:

[0235]

[0236] in:

[0237] SS is the spillover diffusion coefficient;

[0238] Δσ f It is the incremental standard deviation, representing the number of positive and negative fluorescent flow cytometers collected from the fluorescent dye.

[0239] According to the launch diffusion between; and

[0240] Δd is the difference in fluorescence intensity between the positive and negative fluorescence flow cytometry data received by the fluorescence detector.

[0241] 51. The non-transitory computer-readable storage medium as described in Clause 49, wherein calculating the spillover diffusion coefficient includes assuming that the fluorescence intensity acquired by the fluorescence detector for the negative flow cytometry data population is zero.

[0242] 52. The non-transitory computer-readable storage medium as described in Clause 51, wherein the spillover diffusion coefficient is calculated according to Formula 2:

[0243]

[0244] in:

[0245] SS is the spillover diffusion coefficient;

[0246] σ 2 It is the standard deviation of the positive fluorescence flow cytometry data population;

[0247] It is an estimate of the standard deviation of the negative fluorescence flow cytometry data population; and

[0248] d is the intensity of the light collected by the fluorescence detector.

[0249] 53. The non-transitory computer-readable storage medium as described in Clause 52, wherein the linear regression method is used to calculate...

[0250] 54. A non-transitory computer-readable storage medium according to any one of clauses 45-53, wherein said fluorescence flow cytometry data are collected from light emitted from a variety of different fluorescent dyes.

[0251] 55. The non-transitory computer-readable storage medium according to Clause 54, wherein the number of the various fluorescent dyes ranges from 2 to 20 different fluorescent dyes.

[0252] 56. The non-transitory computer-readable storage medium according to clause 54 or 55, wherein the number of the various fluorescent dyes ranges from 3 to 5 different fluorescent dyes.

[0253] 57. A non-transitory computer-readable storage medium according to any one of clauses 54 to 56, wherein the spillover diffusion coefficient of each fluorescence detector-fluorescent dye pair is calculated.

[0254] 58. The non-transitory computer-readable storage medium according to any one of clauses 54 to 57, wherein the overflow diffusion coefficients of each fluorescence detector-fluorescent dye pair calculated in combination in the overflow diffusion matrix.

[0255] 59. The non-transitory computer-readable storage medium as described in Clause 58, further comprising calculating the spillover diffusion magnitude of each of the plurality of different fluorescent dyes based on the spillover diffusion matrix.

[0256] 60. The non-transitory computer-readable storage medium according to Clause 59, wherein adjusting fluorescence flow cytometry data comprises reducing the spillover diffusion magnitude of each of the various fluorescent dyes in the population of fluorescence flow cytometry data corresponding to the fluorescent dye.

[0257] 61. A non-transitory computer-readable storage medium according to any one of clauses 45 to 60, wherein establishing partitions among different spillover-diffusion-adjusted fluorescence flow cytometry data populations comprises evaluating the differences of the different spillover-diffusion-adjusted populations relative to a threshold.

[0258] 62. The non-transitory computer-readable storage medium as described in Clause 61, wherein evaluating the differences between different spillover-adjusted populations relative to a threshold includes calculating the Matthews correlation coefficient.

[0259] 63. The non-transitory computer-readable storage medium as described in Clause 62, wherein the fluorescence flow cytometry data does not contain data showing positivity for at least one fluorescent dye.

[0260] 64. The non-transitory computer-readable storage medium as described in Clause 63, wherein evaluating the differences between different spillover-adjusted populations relative to a threshold includes calculating balance accuracy.

[0261] 65. The non-transitory computer-readable storage medium according to any one of clauses 45 to 64, wherein the different overflow-diffusion-adjusted populations of fluorescence flow cytometry data are partitioned according to a hierarchical structure.

[0262] 66. The non-transitory computer-readable storage medium as described in Clause 65, wherein the hierarchy details the association between spillover-diffusion-adjusted populations of fluorescence flow cytometry data that are positive or negative for a particular fluorescent dye and the corresponding phenotype.

[0263] Although the invention has been described in detail with illustrations and examples for the purpose of clarity, it will be apparent to those skilled in the art, given the edifying significance of the invention, that certain changes and modifications may be made without departing from the spirit or scope of the appended claims.

[0264] Therefore, the foregoing only illustrates the principles of the invention. It should be understood that those skilled in the art can design various structures, although not explicitly stated or shown herein, but these designs reflect the principles of the invention and do not depart from its spirit and scope. Furthermore, all examples and conditional language listed herein are primarily intended to help the reader understand the principles of the invention and the inventors' ideas for further expanding the field, and should be interpreted as not being limited by these specifically listed examples and conditions. Moreover, all statements herein referencing the principles, aspects, and embodiments of the invention and their specific examples are intended to cover their structural and functional equivalents. Furthermore, the equivalents are intended to include currently known equivalents and those to be developed in the future, i.e., any functionally identical elements developed regardless of their structure. Moreover, no part of the invention will be disclosed to the public, whether or not it is explicitly stated in the claims.

[0265] Therefore, the scope of the invention is not limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the invention are embodied in the appended claims. In the claims, 35 U.S.C. 112(f) or 35 U.S.C. 112(6) is explicitly invoked only when the limiting phrase “means for…” or “steps for…” is explicitly used at the beginning of the limiting phrase of the claims; if such phrase is not used in the limiting phrase of the claims, then 35 U.S.C. 112(f) or 35 U.S.C. 112(6) is not invoked.

Claims

1. A method for classifying fluorescence flow cytometry data, the method comprising: The flow cytometry data are processed using a supervised algorithm, which is configured to: The fluorescence flow cytometry data were clustered into a population; Determine the extent of spillover diffusion of the population from the fluorescence flow cytometry data; The populations of the fluorescence flow cytometry data are adjusted based on the determined spillover diffusion level to generate different populations of fluorescence flow cytometry data adjusted for spillover diffusion level; and Partitions are established between the different populations of fluorescence flow cytometry data adjusted for spillover diffusion to classify the different populations of fluorescence flow cytometry data adjusted for spillover diffusion.

2. The method of claim 1, wherein the fluorescence flow cytometry data are clustered into groups based on whether the fluorescence flow cytometry data are positive or negative relative to a specific fluorescent dye.

3. The method according to claim 1 or 2, wherein the fluorescence flow cytometry data is determined to be positive or negative for a specific fluorescent dye based on the relationship between the fluorescence flow cytometry data and a threshold, i.e., positive fluorescence flow cytometry data or negative fluorescence flow cytometry data.

4. The method of claim 1, wherein determining the extent of spillover diffusion comprises quantifying the degree to which fluorescence flow cytometry data acquired by a fluorescence detector increases due to the collection of light emitted by a specific fluorescent dye.

5. The method according to claim 3, wherein determining the degree of spillover diffusion includes calculating the spillover diffusion coefficient of the fluorescence detector-fluorescent dye pair.

6. The method of claim 5, wherein the spillover diffusion coefficient is calculated according to formula 1: ; in: SS is the spillover diffusion coefficient; ∆σ f It is the incremental standard deviation, representing the emission diffusion between the positive fluorescence flow cytometry data and the negative fluorescence flow cytometry data collected from the fluorescent dye; as well as ∆d is the difference in fluorescence intensity between the positive fluorescence flow cytometry data and the negative fluorescence flow cytometry data received by the fluorescence detector.

7. The method of claim 5, wherein calculating the spillover diffusion coefficient includes assuming that the fluorescence intensity acquired by the fluorescence detector for the population of negative fluorescence flow cytometry data is zero.

8. The method of claim 7, wherein the spillover diffusion coefficient is calculated according to formula 2: ; in: SS is the spillover diffusion coefficient; It is the standard deviation of the population of the positive fluorescence flow cytometry data; It is an estimate of the population standard deviation of the negative fluorescence flow cytometry data; as well as It is the intensity of the light collected by the fluorescence detector.

9. The method of claim 8, wherein the calculation is performed using linear regression. .

10. The method according to claim 5, wherein the fluorescence flow cytometry data are collected from light emitted by a variety of different fluorescent dyes.

11. The method of claim 10, wherein the overflow diffusion coefficients of each fluorescent detector-fluorescent dye pair calculated in the overflow diffusion matrix are combined.

12. The method of claim 11, wherein the method further comprises calculating the overflow diffusion magnitude of each of the plurality of different fluorescent dyes based on the overflow diffusion matrix.

13. The method of claim 12, wherein adjusting the fluorescence flow cytometry data comprises subtracting the spillover diffusion magnitude of each of the plurality of different fluorescent dyes in a population of fluorescence flow cytometry data corresponding to the fluorescent dye.

14. A system for classifying fluorescence flow cytometry data, comprising: A particle analyzer component configured to acquire fluorescence flow cytometry data; and A processor includes memory operatively coupled thereto, wherein the memory contains instructions stored thereon, which, when executed by the processor, cause the processor to perform the following operations: The fluorescence flow cytometry data were clustered into a population; Determine the extent of spillover diffusion of the population from the fluorescence flow cytometry data; The populations of the fluorescence flow cytometry data are adjusted based on the determined spillover diffusion level to generate different populations of fluorescence flow cytometry data adjusted for spillover diffusion level; and Partitions are established between the different populations of fluorescence flow cytometry data adjusted for spillover diffusion to classify the different populations of fluorescence flow cytometry data adjusted for spillover diffusion.

15. A non-transitory computer-readable storage medium comprising instructions stored thereon for classifying fluorescence flow cytometry data by a method comprising: The fluorescence flow cytometry data are processed using a supervised algorithm, which is configured to: The fluorescence flow cytometry data were clustered into a population; Determine the extent of spillover diffusion of the population from the fluorescence flow cytometry data; The populations of the fluorescence flow cytometry data are adjusted based on the determined spillover diffusion level to generate different populations of fluorescence flow cytometry data adjusted for spillover diffusion level; and Partitions are established between the different populations of fluorescence flow cytometry data adjusted for spillover diffusion to classify the different populations of fluorescence flow cytometry data adjusted for spillover diffusion.

Citation Information

Patent Citations

  • Flow cytometer with optical equalization

    US10006852B2

  • Parallel flow cytometer using radiofrequency multiplexing

    US10036699B2

  • Multi-modal fluorescence imaging flow cytometry system

    US10078045B2

  • Parallel flow cytometer using radiofrequency multiplexing

    US10222316B2

  • Multi-modal fluorescence imaging flow cytometry system

    US10288546B2