Systems, methods, and articles of manufacture for detecting abnormal cells using multidimensional analysis
Patent Information
- Application Number
- CN202310383534.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-09-19
- Filing Date
- 2017-09-19
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2037-09-19
Smart Images

Figure CN116359503B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention application filed on September 19, 2017, with application number 201780071450.1 (international application number PCT / US2017 / 052311) and entitled "System, method and article for detecting abnormal cells using multidimensional analysis".
[0002] Cross-reference to related applications
[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 396,621, filed September 19, 2016, pursuant to Section 119(e) of U.S. Patent Application No. 35U.SC, which is incorporated herein by reference in its entirety. [Technical Field]
[0004] This disclosure relates to multidimensional analysis of measured cellular characteristics, and more particularly to systems, methods, and articles for detecting abnormal cells in a cell test set using multidimensional analysis of cellular characteristics measured by flow cytometry. [Background Technology]
[0005] Related technical descriptions
[0006] One method for characterizing heterogeneous cell populations is flow cytometry, originally developed by Herzenberg and colleagues (Science. 1969 166(906):747-9; J Histochem Cytochem. 1976 24(1):284-91; Clin Chem. 1973 19(8):813-6; Ann. NY Acad. of Sci. 1975 254:163-171). Using this technique, cells are labeled with antibodies conjugated to a dye. Flow cytometry can routinely detect three, four, or more immunofluorescent markers simultaneously in a quantitative manner. By combining multiple immunofluorescent markers with the light scattering properties of cells, it is possible not only to distinguish cell lineages but also to differentiate cells at different stages of maturity within these lineages. This is determined based on the expression patterns of unique cell surface antigens (see, for example, Loken MR et al., in Flow Cytometry in Hematology. Laerum OD, edited by Bjerksnes R., Academic Press, New York, pp. 31–42, 1992; Civin CI et al., in "Concise Reviews in Clinical and Experimental Hematology," edited by Martin J. Murphy, AlphaMed Press, Dayton OH, 1992, pp. 149–159). The population identified by flow cytometry can then be separated using on-instrument cell sorting electronics.
[0007] Multiparameter flow cytometry is currently used to detect various types of leukemia. However, current technology requires time-consuming data analysis by professionals (i.e., those proficient in both flow cytometry and hematology, such as physicians). Educating professionals to differentiate between normal and abnormal cell populations requires a lengthy learning process. Furthermore, when using flow cytometry to monitor a patient's response to treatment, conventional techniques require the use of patient-specific groups to detect residual disease.
[0008] Therefore, there is still a need in the field for techniques to improve detection accuracy and simplify data analysis. This disclosure addresses this need, as well as other needs. [Summary of the Invention]
[0009] In one implementation, flow cytometry is used to characterize the normal cell group. The centroid and radius of the cluster group are defined in n-dimensional space, corresponding to the normal maturation of cell lineages within the normal cell group. The cell test group is characterized using flow cytometry, and this characterization is compared to the cluster group. This method facilitates the detection of low-level tumor cells based on phenotypic differences between the test group and its normal counterparts (e.g., assessed through analysis of complex data from normal and abnormal cell populations).
[0010] In one aspect, one embodiment includes a method for diagnosing cancer in a biological cell test group in n-dimensional space, the method comprising: exposing each cell in a normal biological cell group to multiple of four or more reagents using a first protocol; measuring corresponding multiple fluorescence intensities of each cell in the normal biological cell group using a second protocol; mapping each cell in the normal biological cell group to a corresponding point in n-dimensional space based at least in part on the multiple fluorescence intensities of the cells measured in the normal biological cell group, wherein the corresponding points form a normal point group; defining a normal cluster group in n-dimensional space by defining a centroid line and radius based on the mapping of the normal point group in n-dimensional space, wherein each cluster in the normal cluster group corresponds to a maturation level within a cell lineage; exposing each cell in the biological cell test group to multiple reagents using the first protocol; measuring corresponding multiple fluorescence intensities of each cell in the biological cell test group using the second protocol; mapping each cell in the test cells of the biological cell group to a corresponding point in n-dimensional space based at least in part on the multiple fluorescence intensities of the cells measured in the biological cell test group, wherein the corresponding points form a test point group; and comparing the test point group with a normal cluster group.
[0011] In one aspect, one method involves exposing cells to a variety of reagents in any number. Some instruments are capable of producing nine or more colors. The use of additional reagents and colors aids in the characterization of cells.
[0012] In another embodiment, one implementation includes a method for characterizing a biological cell test group in n-dimensional space, the method comprising: mapping each cell in a normal biological cell group to a corresponding point in n-dimensional space using a first scheme, wherein the corresponding points form a normal point group; defining a centroid and radius for a normal cluster group in n-dimensional space based on the mapping of the normal point group in n-dimensional space, wherein the cluster corresponds to a maturity level within a cell lineage; mapping each cell in the biological cell test group to a corresponding point in n-dimensional space using the first scheme, wherein the corresponding points form a test point group; and comparing the test point group with the normal cluster group.
[0013] In another embodiment, one implementation includes a method for diagnosing a biological cell test set, the method comprising: mapping each cell in the biological cell test set to a corresponding point in an n-dimensional space using a defined scheme, the corresponding points forming a test point set; and comparing the test point set with a defined normal cluster set in the n-dimensional space, wherein the clusters in the defined normal cluster set correspond to the maturity level within the cell lineage, and the clusters are defined by a centroid and a radius.
[0014] In another embodiment, one implementation includes a method for characterizing a biological cell test set, the method comprising: mapping each cell in the biological cell test set to a corresponding point in an n-dimensional space using a defined scheme, the corresponding points forming a test point set; representing the test point set in a Cartesian coordinate display including a first axis corresponding to cell maturation within a cell lineage and a second axis corresponding to the frequency of occurrence; and representing a normal cluster set in an n-dimensional space in the Cartesian coordinate display, wherein the cluster is defined by a centroid and a radius and corresponds to the cell maturation level within a cell lineage.
[0015] In another embodiment, one implementation includes a method for characterizing a normal cell lineage in n-dimensional space, the method comprising: exposing each cell in a normal biological cell group to multiple reagents using a first protocol; measuring corresponding multiple features of each cell in the normal biological cell group using a second protocol; mapping each cell in the normal biological cell group to a corresponding point in n-dimensional space based at least in part on the multiple features of the cells measured in the normal biological cell group, wherein the corresponding points form a normal point group; and defining the centroid and radius of a cluster group based on the mapping of the normal point group in n-dimensional space, wherein each cluster corresponds to a maturity level within the normal cell lineage.
[0016] In another embodiment, one embodiment includes a computer-readable medium storing instructions that enable a diagnostic system to facilitate the detection of cancer cells in a biological cell test group by: retrieving a first data set including indications of a plurality of three or more fluorescence intensities for each cell in a normal biological cell group measured using a defined protocol; mapping each cell in the normal biological cell group to a corresponding point in an n-dimensional space based at least in part on the first data set, wherein the corresponding points form a normal point group; defining a centroid line and radius for a normal cluster group in an n-dimensional space based on the mapping of the normal point group in the n-dimensional space, wherein the cluster corresponds to a maturation level within a cell lineage; retrieving a second data set including indications of a plurality of corresponding fluorescence intensities for each cell in the biological cell test group measured using a defined protocol; mapping each cell in the test cells of the biological cell group to a corresponding point in an n-dimensional space based at least in part on the second data set, wherein the corresponding points form a test point group; and comparing the test point group with the normal cluster group.
[0017] In another embodiment, one implementation includes a computer-readable medium storing instructions that enable a diagnostic system to facilitate the detection of cancer cells in a biological cell group by: retrieving a first data set; defining the centroid and radius of a normal cluster group in n-dimensional space based on the first data set, wherein clusters in the normal cluster group correspond to normal maturation levels within a cell lineage; retrieving a second data set; and comparing the second data set with the normal cluster group.
[0018] In another embodiment, one embodiment includes a computer-readable medium storing instructions that cause the control system to facilitate the detection of cells in a biological cell test group by: receiving a first set of data corresponding to a plurality of fluorescence intensities of a normal biological cell group measured using a defined protocol; defining a normal cluster group in a multidimensional space based on the first set of data; wherein the cluster is defined by a centroid line and a radius and corresponds to the cell maturation level within a cell lineage; receiving a second set of data corresponding to an indication of a plurality of fluorescence intensities for each cell in the biological cell test group measured using the defined protocol; and comparing the second set of data with the defined normal cluster group.
[0019] In another embodiment, one embodiment includes a computer-readable medium containing a data structure for characterizing a test group of biological cells, the data structure including: a header portion; a text portion; and a data portion, wherein the text portion contains information about the data portion, and the data portion contains information for defining the centroid and radius of a normal cluster group, and wherein clusters in the normal cluster group correspond to normal maturation levels within a cell lineage.
[0020] On the other hand, an implementation of the diagnostic system includes: a controller; a memory; a data interface; a control interface; and a graphics engine, wherein the diagnostic system is configured to compare a set of test data with a set of normal clusters in an n-dimensional space defined by a centroid and a radius, and wherein the clusters in the set of normal clusters correspond to normal maturation levels within a cell lineage.
[0021] On the other hand, an implementation of a system for diagnosing cell test groups includes: means for defining normal cluster groups corresponding to normal cell lineages; and means for comparing cell test groups with normal cluster groups. [Attached Image Description]
[0022] This patent or application document contains at least one color drawing. Upon request and payment of the necessary fees, this Office will provide a published copy of this patent or application with the color drawing.
[0023] Figure 1 This is a functional block diagram of a system implementation scheme for a method to realize a diagnostic cell test set.
[0024] Figure 2 This is a schematic diagram of a data structure suitable for storing data related to biological cell groups.
[0025] Figure 3 It is suitable for storage and processing Figure 2 The diagram shows a data structure containing information related to the data.
[0026] Figures 4A to 9A And 4B to 9B are projected onto the system (such as Figure 1 The system shown in the figure generates a pseudo-two-dimensional display of multidimensional data.
[0027] Figure 10A and 10B It is projected onto the system (such as) Figure 1 The system shown in the figure is a diagram of multidimensional data in a pseudo-3D display generated by the system.
[0028] Figure 11A and 11B This shows the system (such as) Figure 1 The system shown in the figure generates a menu for a graphical user interface.
[0029] Figures 12A to 17A And 12B to 17B are projected onto the system (such as Figure 1 The system shown in the figure generates a pseudo-two-dimensional display of multidimensional data.
[0030] Figures 18A to 18C This is a flowchart illustrating the system operations used to define the normal centroid and radius of a normal cluster group corresponding to a normal cell lineage.
[0031] Figure 19 This is a flowchart illustrating the system operation for defining the centroid line of a normal cluster group corresponding to a normal cell lineage.
[0032] Figure 20A and 20B It is projected onto the system (such as) Figure 1 The system shown in the figure is a diagram of multidimensional data in a pseudo-3D display generated by the system.
[0033] Figure 21 This is a flowchart illustrating the system operations used to define the normal centroid line and radius of a normal cluster group corresponding to a normal cell lineage.
[0034] Figure 22 This is a flowchart illustrating the system operations used to determine whether points in a test data set are included within a normal cluster group in n-dimensional space.
[0035] Figure 23A and 23B It is projected onto the system (such as) Figure 1The system shown in the figure is a diagram of multidimensional data in a pseudo-3D display generated by the system.
[0036] Figure 24 This is a flowchart illustrating the system operations used to compare the defined centroids of a test data set with those of a normal cluster set.
[0037] Figure 25 It is a schematic diagram of a data structure suitable for storing information, which is used to define the centroid and radius of a normal cluster group in n-dimensional space.
[0038] Figure 26 It is projected onto the system (such as) Figure 1 The system shown in the figure generates a pseudo-two-dimensional display of multidimensional data.
[0039] Figure 27 and 28 These are flowcharts showing the training and implementation phases of the subroutine used to define the reference group in n-dimensional space.
[0040] Figure 29 This is a flowchart illustrating the system operations used to determine the standard reference mean and normalized test data sets for patient data.
[0041] Figures 30-33 It is projected onto the system (such as) Figure 1 The system shown in the figure is a diagram of multidimensional data in a pseudo-3D display generated by the system.
[0042] Figure 34 The representation of the standard reference average vector in n-dimensional space is shown.
[0043] Figure 35 and 36 It is projected onto the system (such as) Figure 1 The system shown in the figure is a diagram of multidimensional data in a pseudo-3D display generated by the system.
[0044] Figure 37 The representation of the average intensity vector of the test patient data set in n-dimensional space is shown.
[0045] Figure 38 The normalized vector representation of the test patient data set in n-dimensional space is shown.
[0046] Figure 39 The application of normalized vectors to normalize the strength of test patient data sets is shown.
[0047] Figure 39A This is a diagram illustrating an example application of normalizing vectors to cells in a data set.
[0048] Figure 40 This shows the methods used to characterize systems (such as...) Figure 1The flowchart shown is a system operation flowchart for setting the radius of the normal patient data group in the system shown.
[0049] Figure 40A This is a flowchart illustrating the implementation scheme of the cell cluster subroutine.
[0050] Figure 41 This is a diagram illustrating the tangential intersection between the cells (data points) of a data group and the centroid line defining a normal cluster group.
[0051] Figure 41A This is a diagram illustrating an example calculation of the tangential intersection point.
[0052] Figure 42 It is a graph of the parametric distance values between the cells (data points) of a data group and the corresponding tangential intersection points between the centroid lines defining the normal cluster group, and it is in the form of a floating-point array.
[0053] Figure 43 This is a diagram showing the standard deviation of parameters in a cluster of combined normal patient data sets, in the form of a floating-point array.
[0054] Figure 43A It is a diagram showing the average distance between the tangential intersections of cells and the centroid line, in the form of a floating-point array.
[0055] Figure 44 This is a diagram of a standardized normal patient data set, presented as a floating-point array.
[0056] Figure 45 This is a graph of the scaled standard deviation in a cluster of combined normal patient data sets, presented as a floating-point array.
[0057] Figure 45A It is a diagram of the standardized average distance of clusters combining normal patient data sets, in the form of a floating-point array.
[0058] Figure 46 This shows the normal distance between cells (data points) used to characterize normal cell groups in a cluster and the centroid line, as well as the distance between the data points and the centroid line determined by the system (e.g., ...). Figure 1 The flowchart shows the system operation of normal cell (data point) frequency decomposition.
[0059] Figure 47 It is a graph of the parametric distance values between the cells (data points) of a data group and the corresponding tangential intersection points between the centroid lines that define a normal cluster group. It is in the form of a floating-point array, including Euclidean distance and the number of clusters.
[0060] Figure 48 This is a diagram of the average distance matrix of cells (data points) in a cluster of normal patient data, which is in the form of a floating-point array.
[0061] Figure 49 This is a diagram of the frequency matrix of cells (data points) in a cluster of normal patient data, which is in the form of a floating-point array.
[0062] Figure 50 This is a diagram of the normal position matrix of cell (data point) combinations from a normal patient, which is in the form of a floating-point array.
[0063] Figure 51 This is a diagram of the normal percentage matrix of cells (data points) in a cluster of cells (data points) from normal patients, presented in the form of a floating-point array.
[0064] Figure 52 This shows the method used to pass through the system (such as...) Figure 1 The system shown is a flowchart of the system operation that compares a group of patient cells (data points) with a defined normal cluster group.
[0065] Figure 53 It is a graph of the parametric distance values between the corresponding tangential intersection points of the cells (data points) in the test patient data group and the centroid line defining the normal cluster group. It is in the form of a floating-point array, including Euclidean distance and cluster number.
[0066] Figure 54 This is a diagram of the average distance matrix of cells (data points) in a cluster of test patient data sets, which is in the form of a floating-point array.
[0067] Figure 55 This is a diagram of the frequency matrix of cells (data points) in a cluster of test patient data sets, which is in the form of a floating-point array.
[0068] Figure 56 and 57 It is projected onto the system (such as) Figure 1 The illustration shows the multidimensional data in the pseudo-two-dimensional image generated by the system shown.
[0069] Figure 58 This shows the method used to pass through the system (such as...) Figure 1 The system shown is a flowchart of the system operation that compares a group of patient cells (data points) with a defined normal cluster group.
[0070] Figure 59 It is a graph of the parametric distance values between the corresponding tangential intersection points of the cells (data points) in the test patient data group and the centroid line defining the normal cluster group. It is in the form of a floating-point array, including Euclidean distance and cluster number.
[0071] Figure 60 and 60AThis is a diagram of the normal radius table of the combined normal cluster group, which is in the form of a floating-point array.
[0072] Figure 61 , 61A 62 and 62A are illustrations used to filter cells (data points) to represent information in the images of the test patient data set, and are in the form of floating-point arrays.
[0073] Figure 63 and 63A It is a diagram that includes (or excludes) images representing a group of test patient data, which is in the form of a floating-point array.
[0074] Figure 64 , 64A Figures 64B, 65, 65A, and 65B show example user interfaces for controlling the generation and display of images representing cells (data points) of test patient data groups.
[0075] Figure 66 A reference cell population is shown.
[0076] Figure 67 and 68 It represents the results of the test patient data set.
[0077] Figure 69 It is a representation of the predicted population of premyelocytes.
[0078] Figure 70 and 71 This is a graph of the data set results.
[0079] Figure 72 It is a representation of the predicted population of monocytes.
[0080] Figure 73 and 74 This is a graph of the data set results.
[0081] Figure 75 It represents a predicted population of undifferentiated progenitor cells.
[0082] Figure 76 and 77 This is a graph of the data set results.
[0083] Figure 78 This is a representation of the reference population relative to CD45 and SSC.
[0084] Figure 79 It represents the movement of a lymphocyte population to a fixed point.
[0085] Figure 80 It represents the CD34 intensity of CD34++.
[0086] Figure 81 It represents the CD14 intensity of monocytes.
[0087] Figure 82 It represents the CD33 intensity of CD14++ monocytes.
[0088] Figure 83 It is a representation of a comparison of cell populations.
Detailed Implementation Methods
[0089] In the following description, certain details are set forth to provide a comprehensive understanding of various embodiments of the apparatus, system, method, and article of manufacture. However, those skilled in the art will understand that other embodiments may be practiced without these details. In other instances, well-known structures and methods associated with, for example, flow cytometers, controllers, etc., such as power supplies, transistors, memories, logic gates, buses, etc., are not shown or described in detail in some of the accompanying drawings to avoid unnecessarily obscuring the description of the embodiments.
[0090] Unless the context otherwise requires, throughout the specification and the following claims, the word “comprise” and its variations, such as “comprises” and “comprising”, shall be understood in an open, inclusive sense, that is, as “including but not limited to”.
[0091] Throughout this specification, the reference to "one embodiment" or "an embodiment" means that a specific feature, structure, or characteristic described with respect to that embodiment is included in at least one embodiment. Therefore, the phrase "in one embodiment" or "in an embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment or all of the embodiments. Furthermore, specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments to obtain further embodiments.
[0092] The headings are provided for convenience only and do not explain the scope or meaning of this disclosure.
[0093] The size and relative positions of the elements in the accompanying drawings are not necessarily drawn to scale. For example, the shapes and angles of various elements are not drawn to scale, and some of these elements have been enlarged and positioned to improve readability. Furthermore, the specific shapes of the elements drawn are not necessarily intended to convey any information about the actual shape of the specific element, and are chosen solely for ease of identification in the accompanying drawings.
[0094] Gene products can be identified on the cell surface or in the cytoplasm of cells using specific monoclonal antibodies. Flow cytometry can be used to simultaneously detect multiple immunofluorescent markers in a quantitative manner. Immunofluorescence staining techniques are well known and can be performed according to any of a variety of protocols, such as those described in Current Protocols in Cytometry (John Wiley & Sons, NY, NY, edited by J. Paul Robinson et al.). Typically, biological samples (such as peripheral blood, bone marrow, lymph node tissue, umbilical cord blood, thymus tissue, tissue from sites of infection, spleen tissue, tumor tissue, etc.) are collected from the subject using techniques known in the art, and cells are isolated from them. In one embodiment, blood is collected from the subject, and any mature red blood cells are lysed using a buffer solution (such as buffered NH4Cl). The remaining leukocytes were washed and then incubated with an antibody (e.g., a monoclonal antibody) conjugated to any of the various dyes (fluorophores) known in the art (see, for example, http: / / colondoubleslashwww.glenspectra.co.uk / glen / filters / fffluorpn.htm or http: / / colondoubleslashcellscience.bio-rad.com / fluorescence / fluorophoradata.htm). Representative dyes mentioned herein include, but are not limited to, FITC (fluorescein isothiocyanate), R-phycoerythrin (PE), and allophycocyanin (APC). And Texas Red.
[0095] Various antibodies known in the art and specific antibodies generated using techniques well known in the art can be used in the context of embodiments disclosed herein. Generally, antibodies used in the methods described herein are specific to cell markers of interest (such as any of the CD cell surface markers) (see, for example, the CD index at http: / / colon.double.slash.www.ncbi.nlm.nih.gov / PROW / guide / 45277084.html; or Current Protocols in Immunology, John Wiley & Sons, NY, NY), cytokines, adhesion proteins, developmental cell surface markers, tumor antigens, or other proteins expressed by the cell population of interest. In fact, antibodies specific to any protein expressed by cells can be used in the context of this disclosure. Exemplary antibodies include, but are not limited to, antibodies that specifically bind to CD3, CD33, CD34, CD8, CD4, CD56, CD19, CD14, CD15, CD16, CD13, CD38, CD71, CD11b, HLA-DR, blood group glycoprotein, CD45, CD20, CD5, CD7, CD2, CD10, and TdT.
[0096] After incubation with dye-conjugated antibodies for a period of time, usually about 20 minutes in the dark (incubation time and conditions may vary depending on the specific protocol), leukocytes are washed with buffered saline and resuspended in protein-containing buffered saline for introduction into flow cytometers.
[0097] Flow cytometry analyzes heterogeneous cell populations one cell at a time and can classify cells based on the binding of immunofluorescent monoclonal antibodies and the light scattering characteristics of each cell (see, for example, Immunol Today. 2000 21(8):383-90). Fluorescence detection is performed using photomultiplier tubes; the number of detectors (channels) determines the number of optical parameters the instrument can examine simultaneously, while bandpass filters ensure that only the intended wavelengths are collected. Therefore, flow cytometry can routinely detect a variety of immunofluorescent markers quantitatively and can measure other parameters such as forward light scattering (an indicator of cell size) and right-angle light scattering (an indicator of cell granularity). Thus, immunofluorescence and flow cytometry can be used to differentiate and sort various cell populations.
[0098] For example, by combining four colors of immunofluorescence with physical parameters of forward light scattering (for cell size measurement) and right-angle light scattering (for cell particle size measurement), a six-dimensional data space can be generated, in which the detection of a specific cell population in normal blood or bone marrow is limited to a small fraction of this data space. As those skilled in the art will recognize upon reading this specification, more or fewer immunofluorescent markers of different colors may also be used. Excitation of fluorophores is not limited to light in the visible spectrum; several dyes, such as the Indo series (for measuring intracellular calcium) and the Hoechst series (for cell cycle analysis), are excitable in the ultraviolet range. Therefore, some instruments currently available in the art are equipped with ultraviolet emission sources, such as four-laser, 10-color Becton Dickinson LSR II. Furthermore, commercially available fluorescence-activated cell sorters, such as FACSVANTAGE, can be used. TM (Becton Dickinson, San Jose, CA), ALTRA TM (Beckman Coulter, Fullerton, CAA) or The sorting instrument (DakoCytomation, Carpinteria, CA) can also sort cell populations into purified fractions.
[0099] Gene expression observed during the development of blood cells, from hematopoietic stem cells to mature cells found in the blood, is a highly regulated process. See Civin CI, Loken MR: Cell Surface Antigens on Human Marrow Cells:Dissection of Hematopoietic Development Using Monoclonal Antibodies and Multiparameter Flow Cytometry Int'l J. Cell Cloning 5:1-16 (1987), which is incorporated herein by reference in its entirety. Therefore, the specific, tightly controlled expression of genes occurs not only within different lineages of blood cells, but also within different stages of maturation within these lineages. See Loken, MR, Terstappen LWMM, Civin Cl, Fackler, MJ: Flow Cytometry Characterization of Erythroid,Lymphoid and Monomyeloid Lineages in Normal Human Bone MarrowFlow Cytometry in Hematology, Laerum OD, Bjerksnes R. (eds.), Academic Press, New York, pp. 31-42 (1992), which is incorporated herein by reference in its entirety. These gene products not only appear and / or disappear at precise stages of maturation, but the amount of these glycoproteins is also regulated within very tight limits in normal cells. These antigenic relationships have been shown to be established early in fetal development and to remain constant throughout adulthood in blood cells undergoing constant renewal and replenishment. See also LeBein TW, Wormann B, Villablanca JG, Law CL, Shah VO, Loken MR. Multiparameter Flow Cytometric Analysis of Human Fetal Bone Marrow B Cells Leukemia 4:354-358 (1990), which is incorporated herein by reference in its entirety. These patterns and gene expression relationships during the maintenance of normal cell maturation after chemotherapy or even bone marrow transplantation. See Wells DA, Sale GE, Shulman HE, Myerson D, Bryant E, Gooley T, Loken MR: Multidimensional Flow Cytometry of Marrow Can Differentiate Leukemic Lymphoblasts From Normal Lymphoblasts and Myeloblasts Following Chemotherapy and / or Bone Marrow Transplant Am. J. Clin. Path. 110:84-94 (1998), which is incorporated herein by reference in its entirety. Therefore, there is a very close-coordinated multigene regulation during normal blood cell development in terms of the timing of expression and the amount of gene products expressed on the cell surface.
[0100] A comparison of normal antigen expression with that of the tumor process reveals that the regulation of gene expression is disrupted in tumor cells. This disruption results in antigenic relationships that differ from those observed during normal cell maturation. See Hurwitz, CA; Locken, MR; Graham, ML; Karp, JE; Borowitz, MJ; Pullen, DJ; Civin, CI. Asynchronous Antigen Expression in B Lineage Acute Lymphoblastic Leukemia Blood, 72:299-307 (1998). These are not new antigens, but rather gene products that are normally expressed but have lost the coordinated regulation found in normal cells. Both acute lymphoblastic leukemia (“ALL”) and acute myeloid leukemia (“AML”) abnormally express antigens. See Terstappen LWMM, Loken MR: Myeloid Cell Differentiation in Normal Bone Marrow and Acute Myeloid Leukemia Assessed by Multi-Dimensional Flow CytometryAnal. Cell Path. 2:229-240 (1990), which is incorporated herein by reference in its entirety. Types of exceptions include:
[0101] (1) Lineage distortion, defined as the expression of non-lineage antigens;
[0102] (2) Antigens are expressed asynchronously, for example, antigens that are usually found on immature cells are expressed on mature cells;
[0103] (3) Antigen absence; and
[0104] (4) Quantitative abnormalities.
[0105] See Terstappen LWMM, Konemann S, Safford M, Loken MR, Zurlutter K, BuchnerTh, Hiddemann W, Wormann B: Flow Cytometric Characterization of Acute Myeloid Leukemia, Part II. Phenotypic Heterogeneity at Diagnosis ,Leukemia 6:70-80 (1991), which is incorporated herein by reference in its entirety.
[0106] Not only do leukemia cells differ from normal phenotypes, but the relationships between antigens also vary from case to case, suggesting that each leukemic transformation results in a loss of coordinating gene regulation, leading to a unique phenotypic pattern for each type of leukemia. In 120 pediatric ALL cases and 86 adult AML cases, each detailed phenotype differed from normal and from one another. See ibid.; Hurwitz, ibid. Thus, tumor transformation affects the regulation of primary DNA sequences (genotype) and normal genes, causing them to be expressed inappropriately at the wrong time during development, in the wrong amount, and / or in the background of other genes not observed in normal cells (phenotype). The loss of coordinating gene regulation appears to be a hallmark of tumor transformation leading to aberrant phenotypes, where each leukemia clone differs from normal and from other leukemias of the same type.
[0107] It should be noted that the implementation plan is not limited to the analysis of leukemia cells (e.g., acute and chronic lymphocytic leukemia (ALL, CLL) and acute and chronic myeloid leukemia (AML, CML)) and other hematopoietic and lymphoma cells. The implementation plan can be applied to the analysis of any of a variety of malignancies, such as lymphoma, myeloma, or premalignant tumors (e.g., myelodysplastic syndromes), and other diseases, including any of a variety of hematologic disorders.
[0108] Flow cytometry can be used to utilize this phenotypic difference from normal to aid in the diagnosis of leukemia and to monitor response to treatment. Flow cytometry has been used in hematology to characterize tumors, for example, to differentiate AML from ALL. However, conventional methods require that the cells of interest form the major portion of the total cells examined, and that the expected disease process is known before analysis, as when morphological examination identifies a population of leukemia cells of indeterminate subtype. Focus on tumor cells can be extended to the detection of residual disease. However, conventional residual disease detection techniques using flow cytometry require patient-specific reagent sets to identify the specific phenotype observed at the time of diagnosis. See Reading CI, Estey EH, Huh YO, Claxton DF, Sanchez G, Terstappen LW, O'Brien MC, Baron S, Deisseroth AB, Expression of Unusual Immunophenotype Combinations in Acute Myelogenous Leukemia Blood 81:3083-3090 (1993), which is incorporated herein by reference in its entirety. This patient-specific group has been used to detect residual ALL and AML at levels reduced to 0.03%–0.05%. See Coustan-Smith E, Sancho J, Hancock ML, Boyett JM, Behm FG, Raimondi SC, Sandlund JT, Rivera GK, Rubnitz JE, Ribeiro RC, PuiCH, Campana D. Clinical Importance of Minimal Residual Disease in Childhood Acute Lymphoplastic Leukemia ,Blood 96:2691-2696(2001);San Miguel JF,VidrialesMB,Lopez-Berges C,Diaz-Mediavilla J,Gutierrez N,Canizo C,Ramos F,CalmunitiaMJ,Perez J,Gonzalez M,Orfao A, Early Immunophenotypical Evaluation of Minimal Residual Disease in Acute Myeloid Leukemia Identifies Different Patient Risk Groups and may Contribute to Postinduction Treatment Stratification ,Blood 98:1746-1751 (2002), which is incorporated herein by reference in its entirety.
[0109] However, routine testing for residual disease using patient-specific reagent sets has the following limitations:
[0110] 1. Diagnostic samples with anomalous phenotypes are required for group construction. In 25% of cases, anomalous phenotypes may not be identified. See Vidriales, ibid.
[0111] 2. The processing time is long because technicians must review previous analyses for a particular patient to determine the reagent combination to use in each case.
[0112] 3. A phenotype that differs from the initially diagnosed leukemia cell population may not be detectable. For example, the phenotype may change from diagnosis to relapse due to clonal evolution or the growth of small chemoresistant subclones. See San Miguel, ibid.
[0113] 4. Unexpected or unforeseen abnormalities, such as secondary spinal dysplasia or other spectrum abnormalities, may be overlooked.
[0114] Using patient-specific groups to assess residual disease works well in controlled settings, such as studies where all consecutive samples are available and adherence is high, with specimens obtained at specific times of treatment. However, in clinical practice, residual disease analysis may be required by flow cytometry labs when a preliminary diagnosis has not been made. Detailed immunophenotypes are often unavailable or incomplete.
[0115] Residual disease detection can also be performed using standardized groups and differences from normal cells as tumor-specific markers. Coordination of gene expression is so precise that a difference of half a decade's divergence in antigen expression is sufficient to distinguish normal from abnormal tumor cells. In this approach, specific reagent sets are used for each suspected lineage, such as B-lineage ALL; T-lineage ALL; AML; B-lineage non-Hodgkin lymphoma (“B-NHL”) and T-lineage NHL (“T-NHL”), as well as MDS and myeloma. Tumor populations can be identified by first identifying the expected pattern of normal cells and then focusing on cells that do not match the expected pattern of normal cells. This method for detecting residual disease has been used for several years at the Fred Hutchinson Cancer Research Center and has successfully predicted outcomes for hematopoietic tumors. For example:
[0116] 1. In hematopoietic stem cell transplantation for ALL, flow cytometry has been shown to be more sensitive and specific than morphology, cytogenetics, or a combination of both in predicting relapse in 120 patients. See Wells, DA, ibid.
[0117] 2. In pediatric AML, flow cytometry analysis of residual disease was the best predictor of outcomes in a study of 252 patients. (Sievers, EL, Lange, BJ, Alonzo, TA, Gerbing, RB, Bernstein, ID, Smith, FO, Arceci, RJ, Woods, WG, Locken, MR) Immunophenotypic evidence of leukemia after induction therapy predicts relapse: results from a prospective Children’s Cancer Group study of 252 patients with acute myeloid leukemia Blood 101:3398-3406 (2003). Patients with detectable tumors at any time during treatment are 4 times more likely to relapse and 3 times more likely to die than those with undetectable tumors.
[0118] 3. In hematopoietic stem cell transplantation, flow cytometry can differentiate between normal regenerative blasts and recurrent tumors based on abnormal antigen expression. See Shulman H, Wells D, Gooley T, Myerson D, Bryant E, Locken M. The biologic significance of rare peripheral blasts after hematopoietic cell transplant is predicted by multidimensional flow cytometry, Am J Clin Path 112:513-523 (1999). In the absence of detectable tumor cells, patients may exhibit 20% normal blast cells in the blood or up to 50% regenerative blast cells in the bone marrow.
[0119] Detecting abnormal phenotypes of small cell populations in the blood or bone marrow expands the applicability of flow cytometry beyond simple phenotypes of leukemia. Flow cytometry has been used to confirm that a significant proportion (10%) of patients diagnosed with myelodystrophy are misdiagnosed and have lymphoid abnormalities rather than bone marrow abnormalities. See Wells DA, Hall MC, Shulman HE, Loken MR. Occult B cell malignancies can be detected by three-color flow cytometry in patients with cytopenias, Leukemia 12:2015-2023 (1998). Flow cytometry also allows for the development of scoring systems to stratify patients with spinal dysplasia based on the degree of abnormalities detected in mature bone marrow cells. See Wells, D., Benesch, M., Locken, M., Vallejo, C., Myerson, D., Leisenring, W., Deeg, H. Myeloid and monocytic dyspoiesis as determined by flow cytometric scoring in myelodysplastic syndrome correlates with the IPSS and with outcome after hematopoietic stem cell transplantationBlood 102:394-403 (2003). Patients with more abnormal bone marrow cells, demonstrated by abnormal immunophenotypes and exhibiting greater gene expression abnormalities, had higher relapse rates and post-stem cell transplantation mortality compared to patients with fewer detectable abnormalities. This also showed a high correlation with the International Prognostic Score System (IPSS). Furthermore, high flow cytometry scores classified patients in the intermediate group I of the IPSS system into the statistically significant group based on post-stem cell transplantation relapse.
[0120] Based on the differences from normal, tumor detection has several advantages.
[0121] 1. This technique does not require diagnostic samples for creating specific groups.
[0122] 2. This method allows for rapid sample processing in high-capacity laboratories, where the same group is used for different patients.
[0123] 3. The results are not affected by changes in phenotype after treatment.
[0124] 4. Appropriate selection of standardized groups allows for the detection of unexpected or unforeseen findings caused by hematological abnormalities.
[0125] The routine distinction between normal and abnormal cell populations does indeed present significant limitations. Typically, data analysis must be performed by professionals (MDs or PhDs proficient in both flow cytometry and hematology) rather than technicians, as various clinical situations may indicate whether an observed abnormality is normal or abnormal. Educating professionals to differentiate between normal and abnormal cell populations requires a lengthy learning process. A trained hematologist may require six months to a year to learn these techniques. Currently, professionals' assessment of normal relative to abnormal is based on experience with all the inherent difficulties of subjective analysis, similar to training in diagnostic microscopy. It is difficult to extend the analysis to other sites while maintaining the same sensitivity and specificity. In challenging situations, two or more professionals must reach a consensus on the final diagnosis.
[0126] For example, Weir et al. described a normal “template” generated by four-color flow cytometry analysis of normal B-cell precursors, to which tumor samples could be compared. See Weir, EG et al., Leukemia (1999) 13:558-567. However, unlike this disclosure, the template is a specific, fixed set of geometric regions drawn around the displayed dot plot events, which is then used as the boundary of normal. As Weir et al. noted, the presence of unpredictable isolated events in normal samples falling outside the normal boundaries defined by the template presents a serious problem that has not yet been addressed by their method, particularly in the setting of minimal residual disease detection. Furthermore, as with other prior methods, the analysis requires a trained individual to compare patient samples with the template.
[0127] Furthermore, populations identified by multiple monoclonal antibodies in normal bone marrow do not appear as distinct spherical clouds in multidimensional space. Instead, the data can be described as a series of tubes or snakes whose size and location vary from head to tail in the multidimensional data space as cell lineages progress from immature to mature forms. Therefore, cluster analysis procedures that process data as spherical clouds produce results with the aforementioned limitations.
[0128] In contrast, the implementation described further herein provides a method for determining the centroid and radius, etc., of one or more event clusters corresponding to a normal cell maturation lineage. In this way, statistical analysis can be used to determine whether the event represents an anomalous event (i.e., cancer).
[0129] Normal bone marrow consists of multiple lineages, each undergoing continuous homeostatic maturation. By first assessing normal cells, statistical measures of normal and abnormal composition can be defined. This definition then becomes the standard for analysis. Automating the identification of which cells are in their expected, defined normal locations will help new professionals and technicians teach what phenotypic abnormalities are. It also allows for analysis standardization at multiple sites, thus providing consistency between analyses in identifying abnormal populations.
[0130] Automating the identification of aberrant cells also allows for increased sensitivity. Current manual assessments use three antibodies combined with forward and right-angle light scattering, collecting 10,000 events per tube. A group consists of 7 to 14 different tubes, each with a different antibody combination. Using this current system, tumors can be detected with near 100% specificity. See Am.J.Clin.Path.110:84-94, ibid.; Blood 98:1746-1751, ibid.; Blood 101:3398-3406, ibid. A single professional can analyze and report 20-30 such cases per day. Increased sensitivity is a limiting factor for conventional methods because professionals must spend more time analyzing each case. Automating the identification of aberrant cells will allow for larger datasets (counting more cells) and the application of more antibodies without increasing the time analysts must spend on each sample.
[0131] Statistical analysis can be used to identify more subtle changes in hematopoietic abnormalities. This is particularly important for the analysis of myelodysplastic syndromes (“MDS”), where abnormalities are observed in more mature cells rather than just immature blast cells. Statistical analysis will identify bulges or shifts in centroid lines within tubes that can represent abnormal cellular regulation. It can also define regulatory points and progression rates through developmental processes, leading to a better understanding of the loss of coordinated gene regulation observed during tumor transformation.
[0132] Figure 1 This is a functional block diagram of system 100, which is an implementation scheme for detecting abnormal cells using multidimensional analysis. System 100 includes a measurement system 102 and a diagnostic system 104.
[0133] Measurement system 102 measures the characteristics of cells in a cell sample and includes, as shown, a flow cytometer 106 and a data formatter 108. More than one flow cytometer 106 may be used, although typically one instrument is used for measurements of a specific sample. For example, as discussed in more detail below, measurements from a normal cell group may be obtained using one flow cytometer, while measurements from a test cell group may be obtained using another flow cytometer. Other measuring devices may be used in measurement system 102, such as microscopes (e.g., high-throughput microscopes).
[0134] Measurement system 102 may include a separate data formatter 108 to format the data collected by measurement system 102. Alternatively, data formatter 108 may be part of another component of system 100, such as flow cytometer 106 or diagnostic system 104. For example, data formatter 108 may format data collected by flow cytometer 106 into the flow cytometry standard FCS 2.0 format or another data file format. Measurement system 102 may include additional components such as controllers, memory, discrete circuitry and hardware, and various combinations thereof.
[0135] The diagnostic system 104 analyzes the data received from the measurement system 102, as discussed in more detail below. Figure 1 In the illustrated embodiment, the diagnostic system 104 includes a controller 110, a memory 112, a parser 114, a control input / output interface 116, a data input / output interface 118, a graphics engine 120, a statistics engine 122, a display 124, a printer 126, and a diagnostic system bus 130. In addition to the data bus, the diagnostic system bus 130 may also include a power bus, a control bus, and a status signal bus. However, for clarity, the various diagnostic system buses are... Figure 1 The diagram shows the diagnostic system bus 130.
[0136] Diagnostic system 104 may be physically located away from measurement system 102. Measurement system 102 may be coupled to diagnostic system 104 via one or more communication links (such as the Internet, extranet, and / or intranet, or other local or wide area networks). Similarly, components of diagnostic system 104 may be physically located away from each other and may be coupled together via communication links (such as the Internet, extranet, and / or intranet, or other local or wide area networks). One or more diagnostic systems may exist, each of which may be coupled to one or more measurement systems. Communication links may be wired, wireless, or various combinations thereof.
[0137] The diagnostic system 104 can be implemented in various ways, including as a separate subsystem. The diagnostic system 104 can be implemented as a digital signal processor (DSP), an application-specific integrated circuit (ASIC), etc., or as a series of instructions stored in memory (such as memory 112) and executed by a controller (such as controller 110). Therefore, software modifications to existing hardware can allow the implementation of the diagnostic system 104. Various subsystems (such as the parser 114 and the control input / output interface 116) in... Figure 1The functional block diagrams identify them as separate blocks because they perform specific functions, which will be described in more detail below. These subsystems may not be discrete units, but may be functions of software routines, which may, but are not necessarily, individually callable and therefore identifiable elements. The diagnostic system 104 can be implemented using any suitable software or combination of software, including, for example, WinList implemented with a Java runtime environment or a 3-D Java runtime environment and / or Java.
[0138] While the illustrated embodiment represents a single controller 110, other embodiments may include multiple controllers. Memory 112 may include, for example, registers, read-only memory (“ROM”), random access memory (“RAM”), flash memory, and / or electrically erasable readable programmable read-only memory (“EEPROM”), and may provide instructions and data used by the diagnostic system 104.
[0139] The diagnostic system 104 may include additional components such as controllers, memory, discrete circuits and hardware, and various combinations thereof.
[0140] This article describes an implementation plan for conducting research on B lymphocyte lineages. Where appropriate, [the plan will be implemented / implemented]. Figure 1 The citations are incorporated into the description of this study. The implementation methods described herein can be used to study, characterize, and diagnose other normal and diseased lineages, such as erythrocytes, T lymphocytes, and others, including those with multiple lineages (such as bone marrow lineages) (see Shulman H, 1999, ibid.; Wells DA, 1998, ibid.; and Loken MR and Wells DA, Normal). Antigen Expression in Hematopoiesis:Basis for Interpreting Leukemia Phenotypes , in Immunophenotyping, Eds Carleton Stewart and Janel KANicholson, 2000, Wiley-Liss Corporation).
[0141] The B lymphocyte lineage is a single lineage and is well-defined as four developmental stages within the bone marrow, with multiple antigenic differences between the well-characterized stages. Identifying the entire B lineage by expressing the single antigen CD19 allows for the detection of all four stages of B lineage cells. The earliest B lineage cells (stage I) are identified by CD34 expression, high levels of CD10, and low levels of CD45. During stage II, CD34 is lost, CD10 intensity decreases by 2-fold, CD45 intensity increases, and CD20 expression begins. Once CD20 reaches its maximum, CD45 further increases, while the loss of CD10 indicates stage III. The final stage of B lymphocyte development (IV) is characterized by the lack of CD10 and CD22 expression and high levels of CD45.
[0142] As those skilled in the art will understand, other cell lineages that can be characterized using the methods described herein may comprise multiple lineages or branched lineages, and a lineage may be defined as a different number of developmental stages. For example, bone marrow lineages include erythrocyte and granulocyte-monocyte lineages, etc. Granulocyte-monocyte lineages branch into monocyte and neutrophil lineages.
[0143] Neutrophils can be divided into five identifiable stages. Stage I myeloblasts, identified by CD34 expression, also exhibit high levels of HLA-DR, CD13, and CD33, but not CD11b, CD15, and CD16. These myeloblasts are of medium size but exhibit low side scattering (SSC) in forward light scattering (FSC). Progression to stage II is characterized by loss of CD34 and HLA-DR, high-level acquisition of CD15, and a significant increase in SSC expression without CD11b expression (see Loken MR and Wells DA, 2000, ibid.). Stage II is accompanied by a slight decrease in CD33. Stage III of neutrophil development is characterized by intermediate-level acquisition of CD11b, loss of CD13, and a decrease in SSC associated with secondary granule appearance. Stage IV is indicated by a related increase in CD13 and CD16 and a further slight decrease in CD33 expression. Stage V corresponds to mature neutrophils found in peripheral blood. This cell has the highest levels of CD16, CD13, and CD45, with increased density.
[0144] Based on the expression of cell surface antigens, the monocyte lineage has three detectable stages. Monocyte development has two maturation stages following the myelogenic stage (indistinguishable from stage I of neutrophil development). These cells retain HLA-DR throughout their development, unlike neutrophils which rapidly lose this antigen during the promyelocyte stage. Monocyte maturation (stage II) is initially identified by the rapid appearance of CD11b, while maintaining intermediate levels of CD45. Stage II of monocyte development is accompanied by increased expression of CD13 and CD33 and low expression of CD15. Stage III of development is defined by the coordinated growth of both CD45 and CD14 (see Loken MR and Wells DA, 2000, ibid.).
[0145] Erythrocytes undergo only two stages (see Locken M, 1992, ibid.). Stage I is characterized by the loss of CD45 and the increase of CD71. Stage II is marked by the expression of blood group glycoproteins and the presence of hemoglobin. The final steps of erythrocyte maturation are observed by the loss of the nucleus in reticulocytes, the reduction of CD71, and subsequently the loss of RNA (see Locken, MR, Shah VO, Dattilio KL, Civin CI (1987) Flow cytometric analysis of human bone marrow. I. Normal erythroid development. Blood 69:255-263).
[0146] As described above in Loken MR and Wells DA, 2000, T lymphocytes can be classified into four stages of thymus development based on their reactivity patterns to 10 antigens (CD1a, CD2, CD3, CD4, CD5, CD7, CD8, CD10, CD34, and CD45). Three stages are clearly defined by differences in multiple antigens, while the fourth stage is distinguished by size.
[0147] Therefore, as those skilled in the art will understand upon reading this specification, the methods described herein using B lymphocyte lineages as examples can be used to characterize other cell lineages in n-dimensional space, as described herein and those known in the art.
[0148] In the implementation method described herein regarding B lymphocyte lineages, all four stages of B cell development are identified using two reagent tubes with four colors:
[0149] Tube 1: CD20 FITC, CD10 PE, CD45 PerCP and CD19 APC.
[0150] Tube 2: CD22 FITC, CD34 PE, CD45 PerCP and CD19 APC.
[0151] Redundancy of markers (CD19 and CD45) in the two tubes allowed for comparison of data between different tubes. In this study, datasets of 200,000 events were collected on a FACS Calibur flow cytometer (Becton Dickinson, San Jose, CA). Sample preparation procedures were standard and followed a fixed protocol. See Am.J.Clin.Path.110:84-94, ibid. List pattern data from two phenotypically normal patients were collected in FCS format for analysis. Clusters identified by individuals proficient in both flow cytometry and hemopathology (e.g., physicians) were compared with those identified by a diagnostic system 104 using a clustering algorithm. Visual centers of clusters identified by professionals were compared with those generated by the diagnostic system 104. The process was iterative, as the user modified the identified clusters based on the results from the clustering algorithm and ran additional clustering algorithms using the modified cluster definitions.
[0152] A cohort of normal bone marrow B lymphocytes in tube 1 underwent a four-color analysis. Samples were collected to obtain 200,000 events for analysis. Cells were placed in tubes and stained with reagents CD20-fluorescein (FITC), CD10 phycoerythrin (PE), CD45 polydinophytin-chlorophyll protein (PerCP), and CD19 allophycocyanin (APC). Characteristics of exposed cells were measured using flow cytometry (see [link to relevant documentation]). Figure 1 Flow cytometer 106). System (e.g.) Figure 1 The system 100 shown uses a combination of data received from measurement results and input from users (such as professionals or technicians) to measure and analyze samples, as discussed in more detail below.
[0153] The publicly available flow cytometry standard FCS 2.0 specification can be used to store the measured characteristics of cells in a sample. Other data formats and structures can also be used, such as FCS 1.0 or FCS 3.0 formats. Example data structure 200 for storing data sets is provided in... Figure 2 As shown in the image. (Reference) Figure 1 and 2Parser 114 parses the header section 202, text section 204, data section 206, and analysis section 208, along with collected information including parameter names, the total number of data points, and data type details. Header section 202 describes the location of the other sections in data structure 200. Header section 202 contains offset information for the start and end points of text section 204, data section 206, and analysis section 208. Text section 204 contains a series of ASCII-encoded key-value pairs describing various aspects of data structure 200. For example, $TOT / 5000 / is a key-value pair indicating that the total number of events in the file is 5000, while $PAR gives the total parameter number. Data section 206 contains raw data. This data is typically in one of three modes (list, related, or unrelated) described in text section 204 by, for example, the $MODE key value. For example, data can be written to data section 206 in one of four formats (binary, floating-point, double-precision floating-point, or ASCII) as described by the $DATATYPE key value. A common data storage format is a list-mode storage in binary integer form ($DATATYPE / I / $MODE / L / ). The $PnB keyword group specifies the storage bit width for each parameter. The PnR keyword group specifies the channel number range for each parameter. For example, $PnB / 16 / $PnR / 1024 / , where n is an integer, specifies a 16-bit field for parameter n and a value for parameter n ranging from 0 to 1023, corresponding to 10 bits. Analysis section 208 is an optional segment that, when present, may contain the results of data processing. Analysis can also be performed offline after data is collected and stored in a data structure (such as data structure 200). In the test study, analysis section 208 was not used. However, analysis section 208 can be used to store information defining the centroid and radius of the data set.
[0154] The data offset in FCS 2.0 format is given in the property file. Example property file 300 is shown in... Figure 3 As shown in the diagram, the attribute file contains a header section 302 containing information on how to read the attribute file 300; a format section 304 containing information on the format of the data structure 200; and a filter section 306 containing information that the parser 114 can use to filter the data stored in the data structure 200. The parser 114 uses the information extracted from the attribute file 300 to parse the loaded data structure 200. The attribute file 300 can be easily modified to allow the use of various data file formats, such as various flow cytometry standard formats.
[0155] System 100 can use the fluorescence intensity corresponding to CD19 as the initial gate. Therefore, it is not necessary to evaluate all 200,000 cells in the 200,000-cell event list; only CD19-positive cells (which include all B lineage cells) can be evaluated. This enhances statistics by increasing the number of B lineage cells to be analyzed without increasing the computational time required to distinguish B lymphocytes from most other cells in the bone marrow. If no such gate is present on the cells of interest, it may take 6–8 hours of computation to identify clusters in the 200,000-cell event list. The proportion of immature B lymphocytes (stages I–III) is, on average, less than 2% of all nucleated cells in normal bone marrow. See Loken, MR, Shah, VO, Dattilo, KL, Civin, CL. Flow Cytometry Analysis of Human Bone Marrow:II.Normal B Lymphoid Development Blood 70:1316 (1987). Therefore, by increasing the total count to 200,000 and gating relatively infrequent CD19-positive cells, the cells of interest can be analyzed while maintaining the entire dataset and avoiding artifacts introduced for CD19 by electronic gating during data collection. However, in an alternative implementation, electronic gating for CD19 can be used during data collection.
[0156] Interested groups from the example normal data set collected as described above regarding tube 1 Figures 4A to 9A The image shows a series of four-color analyses generated using WinList. Interested groups can also view the data in other ways, such as the corresponding four-color analysis display. Figures 4B to 9B As shown in the image. Figures 4A to 9A Figures 4B to 9B are collectively referred to as Figures 4 to 9 in this paper.
[0157] Event clusters were initially identified across multiple 2×2 display projections of 6-dimensional data (4 colors and 2 light scattering parameters). The display can be, for example, a representation of the data in a Cartesian coordinate system. The display projections can be... Figure 1 The graphics engine 120 shown generates the image. Users (such as those proficient in both flow cytometry and hematology) identify ML regions in a 2×2 display projection in a coordinate system (where the horizontal axis corresponds to forward light scattering and the vertical axis corresponds to side light scattering), as shown in Figure 4. ML regions correspond to nucleated cells. Users identify lymphocytes, monocytes, bone marrow cells, and blastocytes in a 2×2 display projection in a coordinate system (such as a Cartesian coordinate system, where the horizontal axis corresponds to side light scattering and the vertical axis corresponds to the fluorescence intensity level of CD45), as shown in Figure 5. Users identify B lymphocytes in a 2×2 display projection in a coordinate system (where the horizontal axis corresponds to side light scattering and the vertical axis corresponds to the fluorescence intensity level of CD19), as shown in Figure 6.
[0158] The user identified stage I, stage II, and stage III / IV clusters in a 2×2 display projection in a coordinate system (the horizontal axis corresponds to the fluorescence intensity level of CD19 and the vertical axis corresponds to the fluorescence intensity level of CD45), as shown in Figure 7. These stages correspond to the maturation levels of B lymphocytes. The user also identified stage I, stage II, stage III, and stage IV clusters in a 2×2 display projection in a coordinate system (the horizontal axis corresponds to the fluorescence intensity level of CD10 and the vertical axis corresponds to the fluorescence intensity level of CD45), as shown in Figure 8. Finally, the user identified stage I, stage II, stage III, and stage IV clusters in a 2×2 display projection in a coordinate system (the horizontal axis corresponds to the fluorescence intensity level of CD20 and the vertical axis corresponds to the fluorescence intensity level of CD10), as shown in Figure 9.
[0159] Based on the user assessment in Figure 4-9, accessed cells were assigned to initial clusters. This resulted in a seven-dimensional normal data set, with dimensions corresponding to: forward light scattering; side light scattering; CD19 fluorescence intensity level; CD45 fluorescence intensity level; CD20 fluorescence intensity level; CD10 fluorescence intensity level; and clusters, corresponding to the maturation stage within the B cell population. A color was assigned to each cluster for identification, and the data was mapped to a six-dimensional space. The data was processed by a diagnostic system (such as...). Figure 1 The diagnostic system 104 shown is displayed in a rotatable pseudo-3D graphic display (which has color coding based on cluster identification).
[0160] The diagnostic system 104 maps normal data sets to a three-axis coordinate system (such as a Cartesian coordinate system) and displays the data for user viewing. Each axis corresponds to a dimension of the data set, while color indicates the cluster to which a particular cell is assigned. The data sets can also be displayed tabularly or in a combined display. Figure 10A and 10B (Collectively referred to as Figure 10) shows an example display 400 that combines pseudo-3D graphic representation 402 with tabular representation 404. Figure 10A It is a color display and Figure 10B This corresponds to the shadow display.
[0161] The graphical representation 402 includes an x-axis 406 corresponding to the fluorescence intensity of CD20, a y-axis 408 corresponding to the fluorescence intensity of CD10, and a z-axis 410 corresponding to the fluorescence intensity of CD45. Data in the first cluster 412 are assigned red and correspond to the stage I maturity level. Data in the second cluster 414 are assigned green and correspond to the stage II maturity level. Data in the third cluster 416 are assigned blue and correspond to the stage III maturity level. Data in the fourth cluster 418 are assigned yellow and correspond to the stage IV maturity level.
[0162] The table representation 404 includes a first column 420 indicating the cluster number, a second column 422 indicating the number of points in the cluster, a third column 424 indicating the color or shading assigned to the cluster, a fourth column 426 indicating the cluster radius, a sixth column 428 indicating the percentage of anomalous events or points in the total group of events or points, and a seventh column 430 indicating whether the logarithmic distance between the cluster's centroid and the cluster's statistical centroid is greater than a threshold. Display 400, as shown, can be an interactive computer display. Users can update the information used to generate display 400 using data input fields 432 and 434. As shown, the threshold is set to 2.5 in field 434.
[0163] The diagnostic system 104 allows users to select the three axes for which data to be mapped using a menu in the graphical user interface (GUI). Figure 11A and 11B This demonstrates what can be achieved by diagnostic systems (such as...) Figure 1 The diagnostic system 104 shown employs an example menu 436. The diagnostic system 104 also allows the user to select other settings via the menu. For example, the menu may include options for: selecting between different stored filter parameters, editing stored filter parameters, and specifying new filter parameters. For example, high-resolution data may be filtered to exclude data with values corresponding to more than 10... 2 The side scattering parameters and data corresponding to CD19 parameters less than 10 to 1.6989701 are displayed. Menu selection also allows selection of planes in the coordinate system to be filtered. Multiple filtering criteria can be used, and the filtering criteria can be greater than or less than a specified threshold. The menu system also allows selection of specific clusters to apply various filtering criteria. This allows the user to view various pseudo-3D displays of the normal data set to help the user select initial data for the diagnostic system 104 to use when defining the centroid line and radius of the normal data set. The diagnostic system 104 also allows menu selection of a standard deviation method or a fixed value and rotation of the displayed image. The diagnostic system 104 can also display the cluster boundaries of the data set based on the selected centroid and radius.
[0164] The normal data set may also include separate data files corresponding to individual samples. For example, a user can examine and manipulate a data set that includes cells drawn from a single individual and a single tube, or a user can combine samples drawn from multiple individuals and / or tubes into a single normal data set. If a sample drawn from an individual is deemed abnormal, that sample can be excluded from the normal data set.
[0165] Referring to this study, in the example B lymphocyte dataset from tube 1, the value n equals 6. Each n-dimensional point is mapped to an n-dimensional space, which can be represented in a floating-point array using n+1 floating-point parameters. Table 1 shows the floating-point array of the example six-dimensional B lymphocyte dataset, where P1PR1 is the value of the first parameter of the first point, P2PR1 is the value of the first parameter of the second point, ..., P... n PR1 is the value of the first parameter of the nth point, etc., where a seventh parameter P is added for the cluster with assigned points. n C#. Floating-point arrays can be generalized to any number of dimensions. Diagnostic system 104 improves point-to-cluster allocation by applying one or more clustering algorithms to normal data sets in n-dimensional space.
[0166] <![CDATA[P2PR1]]> <![CDATA[P2PR2]]> <![CDATA[P2PR3]]> <![CDATA[P2PR4]]> <![CDATA[P2PR5]]> <![CDATA[P2PR6]]> <![CDATA[P2C#]]> <![CDATA[P3PR1]]> <![CDATA[P3PR2]]> <![CDATA[P3PR3]]> <![CDATA[P3PR4]]> <![CDATA[P3PR5]]> <![CDATA[P3PR6]]> <![CDATA[P3C# <!-- 16 -->]]> ... ... ... ... ... ... ... <![CDATA[P n PR1]]> <![CDATA[P n PR2]]> <![CDATA[P n PR3]]> <![CDATA[P n PR4]]> <![CDATA[P n PR5]]> <![CDATA[P n PR6]]> <![CDATA[P n C#]]>
[0167] Table 1: Floating-point arrays of six-dimensional data sets
[0168] The diagnostic system 104 allows the user to cluster data using a selected clustering algorithm. For example, the user can specify multiple clusters, k, and use the K-means algorithm to cluster the data. For instance, the diagnostic system 104 can divide the data into k clusters and assign a center to each cluster. The center can be randomly assigned to one of the points, or based on the user's observation input. The distance between two points in n-dimensional space can be defined as follows:
[0169] D(P1,P2)=SQRT[(K1(P1PR1-P2PR1)) 2 +(K2(P1PR2-P2PR2)) 2 +
[0170] +(K3(P1PR3-P2PR3)) 2 +...+(K n (P1PR n -P2PR n )) 2 Equation 1
[0171] Where D(P1,P2) is the distance between two points in n-dimensional space, and P1PR1 is the value of the first parameter of the first point, P2PR1 is the value of the first parameter of the second point, ..., P n PR1 is the value of the first parameter of the nth point, and K1, K2, K3...K nThis is a weighting constant. In this study, the weighting constant is set to 1. In other words, no weighting is used in this study. These centers are updated iteratively until a convergence criterion is met. In each iteration, each data point is assigned to its nearest center, and the center is recalculated using the average parameter value of all points belonging to the cluster. A typical convergence criterion used in this study is that points are not (or are rarely) reassigned to new cluster centers. See Forgy, E. Cluster Analysis of Multivariate Data:Efficiency vs.Interpretability of Classifications ,Biometrics,21:768(1965), used to discuss k-means clusters.
[0172] Another example clustering algorithm is the DBSCAN clustering algorithm. Define the neighborhood radius E. ps The threshold number of points in the neighborhood, minPts, is used, and the diagnostic system 104 employs the DBSCAN clustering algorithm. The neighborhood radius and threshold number of points are user-defined. Density-based clustering is based on the fact that the cluster density is higher than its surrounding environment. DBSCAN automatically discovers dense clusters at a given density threshold. See Ester, M., Kriegel, H., Sander, J., Xu, X., A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise In the Proceedings of the 2d International Conference on KDD (1996), DBSCAN clusters were discussed. By definition, the density threshold is specified by two parameters: the neighborhood radius (E... ps The threshold number of points (minPts) in the ∈-neighborhood. If 'p' is in the ∈-neighborhood of 'q', then point 'p' is the density of points directly reachable from point 'q'. If there exists point 'p'... i 'chain, where i = 1…n and 'p i+1 'From 'p i 'Directly accessible density,' 'q' is 'p1' and 'p' is 'p' i+1 If point 'p' is a density reachable from point 'q', then point 'p' is density-connected to point 'q'. If there exists a point 'o' such that both 'p' and 'q' are density-connected from 'o', then point 'p' is density-connected to another point 'q'. In this study, the diagnostic system 104 begins by introducing points into a temporary storage (tempStore, e.g., a list) and finding their ∈-neighborhoods. If a data point's ∈-neighborhood contains points smaller than 'minPts', it is marked as noise and another point is introduced into the tempStore. Otherwise, all ∈-neighborhood points are introduced into the tempStore. This process is repeated until all points have been considered. In short, the DBSCAN cluster groups density-connected points together as dense clusters and removes undeniably density-connected points as noise.
[0173] The diagnostic system 104 can also cluster data by using, for example, bridge clustering. Bridge clustering combines K-means clustering with DBSCAN clustering. See Dash, M., Liu, H., Xu, X., '1+1>2':Merging Distance and Density Based Clustering , Proceedings of the IEEE 7th International Conference on Database Systems for Advanced Applications (DASFAA’01), April 18-21, 2001, Hong Kong, China, for discussion of bridge clustering. K-means is performed first, followed by density-based clustering on each K-means cluster, and finally, the K-means clusters are refined by removing noise found in the density-based clusters. For effective merging, each data point has the following three columns to store clustering results: <K-means_ID>, <DBSCAN_ID> and <core / ∈-core / non-core>, wherein:
[0174] K-means_ID is the cluster assigned to each point when K-means is run on data points;
[0175] DBSCAN_ID is the cluster assigned to each point when DBSCAN is run on each K-means cluster; and
[0176] The values of core / ∈-core / non-core are assigned based on the following definitions:
[0177] Definition 1 (Core distance): For each cluster, the core distance is half the distance between the center of the cluster and its nearest cluster center.
[0178] Definition 2 (Core point): A core point is not far from its cluster center within "core distance - ∈". The core region of a cluster is a region where every data point is a core point.
[0179] Definition 3 (+∈ core point): The distance between the point and its cluster center is between "core distance" and "core distance + ∈".
[0180] Definition 4 (-∈ core point): The distance between the point and its cluster center is between "core distance" and "core distance - ∈". For convenience, when +∈ and -∈ core points are considered, they are collectively referred to as ∈-core points. An ∈-core region is a region where every point is an ∈-core point.
[0181] Definition 5 (Non-core point): A point that is neither a core point nor an ∈-core point. A non-core region is a region where every point is a non-core point.
[0182] The diagnostic system 104 can also employ wavelet clustering. Wavelet transform is a special form of Fourier transform. See Press, WH, Flannery, BP, Teukiosky, SA. Numerical Recipes ln C:The Art of Scientific Computing Ch. 13.10, Cambridge University Press (1992). This technique has been well established in the fields of image processing and data mining for pattern and edge recognition. See Sheikholeslami, G., Chatterjee, S., Zhang, A. WaveCluster:A Multi-Resolution Clustering Approach for Very Large Spatial Databases, Proceedings of the 24th VLDB Conference, New York, USA, 1998. For example, standard Daubechies wavelet filtering and N-dimensional discrete wavelet transform (NDDFT) can be used.
[0183] Similarly, the same populations (phases) in the second normal dataset were identified in the second tubes (CD22, CD34, CD45, CD19), as... Figures 12A to 17A The colors in Figures 12B to 17B The shaded areas are shown in Figures 12 through 17 (collectively referred to herein as Figures 12 to 17). The user identifies the ML region in a 2×2 display projection in a coordinate system (where the horizontal axis corresponds to forward light scattering and the vertical axis corresponds to side light scattering), as shown in Figure 12. The ML region corresponds to nucleated cells. The user identifies B lymphocytes in a 2×2 display projection in a coordinate system (where the horizontal axis corresponds to side light scattering and the vertical axis corresponds to the fluorescence intensity level of CD19), as shown in Figure 13.
[0184] The user identified stage I, stage II, and stage III / IV clusters in a 2×2 display projection in a coordinate system (the horizontal axis corresponds to the fluorescence intensity level of CD19 and the vertical axis corresponds to the fluorescence intensity level of CD45), as shown in Figure 14. These stages correspond to the maturation levels of B lymphocytes. The user also identified stage I, stage II / III, and stage IV clusters in a 2×2 display projection in a coordinate system (the horizontal axis corresponds to the fluorescence intensity level of CD22 and the vertical axis corresponds to the fluorescence intensity level of CD34), as shown in Figure 15. The user further identified stage I, stage II / III, and stage IV clusters in a 2×2 display projection in a coordinate system (the horizontal axis corresponds to the fluorescence intensity level of CD34 and the vertical axis corresponds to the fluorescence intensity level of CD45), as shown in Figure 16. Finally, the user identified stage I, stage II / III, and stage IV clusters in a 2×2 display projection in a coordinate system (the horizontal axis corresponds to the fluorescence intensity of CD22 and the vertical axis corresponds to the fluorescence intensity of CD45), as shown in Figure 17. The results from tube 1 and tube 2 are combined to produce a single normal data set, as described in more detail below.
[0185] Once the user selectively uses the clustering software to identify and refine the clusters, a centroid line and radius are defined for the normal clusters, where each cluster corresponds to the cell maturity level within the cell lineage. Figures 18A to 18C An embodiment of subroutine 500 is shown, which can be used to define a normal cell population, as described above, relative to... Figure 1 The system 100 shown and the B lymphocytes collected in tubes 1 and 2 are discussed. The entire process of defining the normal cell population should be considered as an iterative process. Other cell lineages (such as bone marrow lineages) may contain multiple lineages or branched lineages. In such cases, multiple centroids may be defined, or the defined centroids may have branches.
[0186] Subroutine 500 begins at 502 and proceeds to 504. At 504, system 100 generates a first normal data set by gating CD19-positive cells, and proceeds to 506 to filter the data set collected based on cell characteristics in measurement tube 1. At 506, system 100 generates a second normal data set by gating CD19-positive cells, and proceeds to 508 to filter the data set collected based on cell characteristics in measurement tube 2.
[0187] At 508, system 100 distinguishes mature and immature cells in the first data set. This can be done, for example, by plotting the fluorescence intensity of CD45 against the fluorescence intensity of CD19 and clustering the first data set based on user input and automated clustering techniques. The system proceeds to 510, where it determines whether to modify the distinction between mature and immature cells in the first data set. This decision can be based on the results of automated clustering techniques, statistical analysis of the data, and / or the display of data sets generated based on this distinction, and can be automatic and / or based on user input. If system 100 determines that the distinction should be modified, system 100 returns to 508. If system 100 determines that the distinction should not be modified, system 100 proceeds to 512.
[0188] At 512, system 100 identifies clusters representing stages I, II, III, and IV in the first data set. This can be done, for example, by plotting the fluorescence intensity of CD45 against the fluorescence intensity of CD10 and CD20 and clustering the data based on user input and automated clustering techniques. System 100 proceeds to 514, where it determines whether to modify the identification of the clusters in the first data set. This decision may be based on the results of automated clustering techniques, statistical analysis of the data, and / or on the display of the data set generated by the identification, and may be automatic and / or based on user input. If system 100 determines that the identification should be modified, system 100 returns to 512. If system 100 determines that the identification should be accepted, the system proceeds to 516.
[0189] At 516, system 100 identifies clusters representing stage I in the second dataset. This can be done, for example, by plotting the fluorescence intensity of CD34 against the fluorescence intensity of CD45 and clustering the data based on user input and automated clustering techniques. The system proceeds to 518, where it determines whether to modify the identification of the stage I clusters in the second dataset. This decision can be based on the results of automated clustering techniques, statistical analysis of the data, and / or on the display of the datasets generated by the identification, and can be automatic and / or based on user input. If system 100 determines that the identification should be modified, system 100 returns to 516. If system 100 determines that the identification should be accepted, the system proceeds to 520.
[0190] At 520, system 100 identifies clusters representing stage IV in the second dataset. This can be done, for example, by plotting the fluorescence intensity of CD22 against the fluorescence intensity of CD34 and clustering the data based on user input and automated clustering techniques. The system proceeds to 522, where it determines whether to modify the identification of stage IV clusters in the second dataset. This decision can be based on the results of automated clustering techniques, statistical analysis of the data, and / or on the display of the dataset generated by the identification, and can be automatic and / or based on user input. If system 100 determines that the identification should be modified, system 100 returns to 520. If system 100 determines that the identification should be accepted, the system proceeds to 524.
[0191] At 524, system 100 identifies clusters representing stages II and III in the second dataset. This can be accomplished, for example, by plotting the fluorescence intensity of CD34 against the fluorescence intensity of CD45 based on user input and automated clustering techniques. The system proceeds to 526, where it determines whether to modify the identification of the stage II / III clusters in the second dataset. This decision can be based on the results of automated clustering techniques, statistical analysis of the data, and / or on the display of the dataset generated by the identification, and can be automatic and / or based on user input. If system 100 determines that the identification should be modified, system 100 returns to 524. If system 100 determines that the identification should be accepted, the system proceeds to 528.
[0192] At 528, system 100 defines a centroid line for each cluster identified at actions 512, 516, 520, and 524. The centroid line used for the cluster can be fractal and can be determined based on user input and automated clustering techniques. The centroid line used for the cluster can be defined, for example, by combining the geometric mean in n-dimensional space with the centroid point determined by the clustering algorithm. The system proceeds to 530, where it determines whether to modify the defined centroid line of the identified cluster. This decision can be based on the results of automated clustering techniques, statistical analysis of the data, and / or on the display of the data set generated by the identification, and can be automatic and / or based on user input. If system 100 determines that the identification should be modified, system 100 returns to 528. If system 100 determines that the identification should be accepted, the system proceeds to 532.
[0193] At position 532, System 100 defines a normal centroid line corresponding to a normal maturation lineage based on the combined dataset. This can be accomplished, for example, by using geometrical bending to connect the defined centroid line of the identified clusters. System 100 can also combine user input with automated clustering techniques to define the normal centroid line. The distance along this centroid line, compared to the start and end, is a measure of the maturation of those cells in a given lineage, as assessed by a specific combination of monoclonal reagents. It should be noted that different antibody combinations can be used to amplify certain parts of the maturation process, while other combinations focus on other maturation stages or other lineages.
[0194] The system proceeds to step 534, where it determines whether to modify the definition of the normal centroid line. This decision may be based on the results of automated clustering techniques, statistical analysis of data, and / or on the display of the data set generated by this determination, and may be automatic and / or based on user input. If system 100 determines that the definition should be modified, system 100 returns to step 532. If system 100 determines that the definition should be accepted, the system proceeds to step 536.
[0195] At position 536, system 100 defines a boundary or normal radius around a defined normal centroid line. The normal radius or boundary can be a fixed radius or it can be variable. For example, it can be a fixed distance, such as 10, or it can be a function of the position on the defined normal centroid line or in n-dimensional space. One definition can be used for the first portion of the defined normal centroid line, and a second definition can be used for the remaining portions. Statistical algorithms (such as wavelet clustering techniques and / or K-means edge envelope techniques (using cluster density)) and / or based on input from the user can be used to determine the normal radius. A smoothing algorithm for defining a specific 3D pattern can also be employed and compared with observations from a statistically determined number of documents.
[0196] System 100 proceeds to 538, where it determines whether to modify the defined normal radius. This decision may be based on the results of automated clustering technology, statistical analysis of data, and / or on the display of the data set generated by this determination, and may be automatic and / or based on input from the user. If system 100 determines that the definition should be modified, system 100 returns to 536. If system 100 determines that the definition should be accepted, the system proceeds to 540, where subroutine 500 stops.
[0197] In some implementations, system 100 can perform Figures 18A to 18C Other actions not shown may be omitted. Figures 18A to 18C All the actions shown in the diagram may be performed in different orders. Figures 18A to 18CThe actions can be iterated. For example, subroutines can be made more iterative. For instance, subroutine 500 can be modified so that system 100 determines whether to modify the defined normal centroid line after action 538, and if so, returns to 532. Subroutine 500 can also call other subroutines to perform various functions, as shown in the following reference. Figure 19 The description is for subroutine 600. Subroutine 500 can also return the value of any desired variable, such as data entered by the user.
[0198] Figure 19 This is a flowchart of example subroutine 600, which can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104 shown is used to define the normal centroid line of the cluster group. Figure 20A and 20B (Collectively referred to as Figure 20) shows a graphical representation of the data from this study, the initial reference centroid 702, and the calculated normal centroid 704.
[0199] Subroutine 600 begins at 602 and proceeds to 604. At 604, diagnostic system 104 identifies a set of reference points. For example, diagnostic system 104 may identify ten reference points selected by the user after reviewing various representations of the data set. Alternatively, diagnostic system 104 may identify multiple statistically selected reference points, or may combine user input with statistical analysis. In this study, the user selected ten reference points after reviewing various representations of the data.
[0200] The diagnostic system 104 proceeds to 606, where it defines a reference centroid line based on the identified set of reference points. Figure 20 shows an example initial reference centroid line 702 defined based on ten reference points identified by the user in the study.
[0201] The diagnostic system 104 proceeds to 608, where it determines the number of clusters into which the data is grouped. For example, in this study, the diagnostic system 104 groups the data into four clusters based on input from the user. Alternatively, the number of clusters can be determined statistically (by using, for example, dbscan clusters) or by using input from the user combined with statistical analysis.
[0202] Diagnostic system 104 proceeds to step 610, where it identifies the centroids of the corresponding number of clusters. This can be done by assigning each point to a cluster based on user input, a statistical algorithm, or a combination thereof. See the discussion of clustering algorithms above. The parameter values of all points assigned to a cluster are summed together, and the result is then divided by the number of points in the cluster to obtain the parameter values of the centroids. For example, if diagnostic system 104 determines at step 608 that the data is grouped into four clusters, then diagnostic system 104 will identify four centroids, each corresponding to one cluster. Table 2, generated below, shows an example calculation of the centroids of a cluster containing 5 data points in 3D space.
[0203]
[0204]
[0205] Table 2: Example calculations for the centroid point
[0206] For ease of explanation, the number of points, number of dimensions, and parameter values in Table 2 are selected.
[0207] The diagnostic system 104 proceeds to 612, where it determines the nearest point on the reference centroid line for each identified centroid point.
[0208] Diagnostic system 104 proceeds to 614, where it calculates the difference between each centroid and the nearest point on the reference centroid line. In this study, this is done using the squared distance formula discussed above, without weighting. See Equation 1.
[0209] The diagnostic system proceeds to point 616, where it adjusts the reference point based on the centroid and the nearest reference point. In this study, this is done by adding the difference between the centroid and the nearest point in the cluster to the reference point within that cluster.
[0210] The diagnostic system proceeds to step 618, where it redefines the reference centroid line using adjusted reference points and the centroid point of each cluster. In this study, this is accomplished by connecting the centroid lines of each cluster using geometrical bending. An example of the redefined reference centroid line is shown as line 704 in Figure 20. The reference centroid line can be further refined using statistical analysis. For example, statistically insignificant points or points outside the defined radius can be removed from the data set. Calculations performed by the diagnostic system 104 when employing subroutine 600 can be stored for later use. For example, when clustering the data during this study, the diagnostic system 104 determines the squared distance between the centroid point and the reference point. This data is stored for calculating the standard deviation value.
[0211] The diagnostic system proceeds to step 620, where it returns the redefined centroid line and the values of any desired variables, as input by the user. The diagnostic system proceeds to step 622, where it stops.
[0212] In some implementations, system 100 can perform Figure 19 Other actions not shown may be omitted. Figure 19 All the actions shown in the diagram may be performed in different orders. Figure 19 Actions can be made more iterative. For example, subroutine 600 can be modified so that system 100 determines whether the number of clusters should be modified after action 616, and if so, returns to action 608. Subroutine 600 can also call other subroutines to perform various functions.
[0213] Figure 21 This is a flowchart illustrating an example subroutine 800, which can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104 shown is used to define the normal centroid line and normal radius of the cluster group.
[0214] Subroutine 800 begins at 802 and proceeds to 804. At 804, diagnostic system 104 identifies a set of reference points. For example, diagnostic system 104 may identify 10 reference points selected by the user after viewing various representations of the data set. Alternatively, diagnostic system 104 may identify multiple statistically selected reference points, or may identify reference points based on statistical analysis combined with user input. In this study, the user selected reference points after viewing various display representations of the data.
[0215] The diagnostic system 104 proceeds to 806, where it defines a reference centroid line based on the identified set of reference points. Figure 20 shows an example initial reference centroid line 702 defined based on ten points identified by the user in the study.
[0216] The diagnostic system 104 proceeds to step 808, where it determines the number of clusters into which the data is grouped. For example, in this study, the diagnostic system 104 grouped the data into four clusters based on input from the user.
[0217] The diagnostic system 104 proceeds to step 810, where it identifies the centroids of the corresponding number of clusters. This can be done by assigning each point to a cluster based on user input or a statistical algorithm, or, as in this study, based on a combination thereof. See the discussion of clustering algorithms above. The parameter values of all points assigned to a cluster are summed together, and the result is then divided by the number of points in the cluster to obtain the parameter values of the centroids. For example, if the diagnostic system 104 determines at step 808 that the data is grouped into four clusters, then the diagnostic system 104 will identify four centroids, each corresponding to one cluster.
[0218] The diagnostic system 104 proceeds to 812, where it determines the corresponding nearest point on the reference centroid line for each identified centroid point.
[0219] Diagnostic system 104 proceeds to 814, where it calculates the difference between each centroid and the nearest point on the reference centroid line. In this study, this is done using the squared distance formula discussed above, without weighting. See Equation 1.
[0220] The diagnostic system 104 proceeds to 816, where it adjusts the reference point based on the centroid and the nearest reference point using input from the user, statistical analysis, or a combination thereof. In this study, the difference between the centroid and the nearest point in the cluster is added to the reference point in the cluster.
[0221] Diagnostic system 104 proceeds to 818, where it redefines the reference centroid line using adjusted reference points and the centroid point of each cluster. In this study, this is accomplished by connecting the centroid lines of each cluster using geometric bending, interpolation, etc. For example, two centroid points can be enclosed in a bend, and additional secondary points can be added in the gap based on the analysis of the normal patient data group, such as based on the mean of normal patients. An example of the redefined reference centroid line is shown as line 704 in Figure 20.
[0222] The diagnostic system 104 proceeds to 820, where it defines the radius of the cluster group. As mentioned above, the radius can be a function of the reference centroid or the position in n-dimensional space. The reference centroid and the radius can form various cluster shapes. For example, spherical clusters, hyperspheres, or hyperellipsoids can be defined by the reference centroid and the radius. The shape of the cluster can resemble a sausage, a barbell, or various other shapes. In this study, the user inputs the radius of each cluster in the normal cluster group, which is the distance from the nearest point on the reference centroid.
[0223] Diagnostic system 104 proceeds to 822, where it determines whether an error criterion is met. For example, diagnostic system 104 may determine whether a statistically insignificant number of points are outside the cluster defined by the reference centroid and radius. If the error criterion is met, diagnostic system 104 proceeds to 824, where the subroutine returns the defined centroid and radius of the data set, along with any other desired variables. If the error criterion is not met, diagnostic system 104 proceeds to 826, where it adjusts the data set. For example, diagnostic system 104 may determine that statistically insignificant points in the data set should be ignored. Diagnostic system 104 returns to 810 for further processing of the adjusted data set.
[0224] Some implementation schemes for System 100 are available. Figure 21 Other actions not shown may be omitted. Figure 21 All the actions shown in the diagram may be performed in different orders. Figure 21 The subroutine can perform actions such as iterating. For example, it can make the subroutine more iterative. Subroutine 800 can also call other subroutines to perform various functions. For example, subroutine 800 can call a subroutine to determine whether the identified cluster should be re-clustered, such as... Figure 22 Subroutine 900 is shown in the diagram.
[0225] Figure 22 This is a flowchart illustrating an example subroutine 900, which can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104 shown uses this information to determine whether points in a data set are contained within a cluster group defined by the centroid line and radius. This information can be used by the diagnostic system 104 to, for example, determine whether the defined normal cluster group should be redefined because too many cells are classified as abnormal, or to detect abnormal cells in a test cell group.
[0226] The subroutine begins at 902 and proceeds to 904. At 904, the diagnostic system 104 retrieves the data set and proceeds to 906. At 906, the diagnostic system 104 sets the data fields associated with each point in the data set to indicate that the subroutine has not yet classified that point and proceeds to 908.
[0227] At 908, the diagnostic system 104 retrieves points related to the selected cluster from the data set and proceeds to 910. At 910, the diagnostic system 104 determines whether the unclassified points related to the selected cluster are within the centroid line and radius of the selected cluster. This can be done, for example, by calculating the distance between the unclassified point and the nearest point on the cluster's centroid line; if the distance is less than the cluster radius at the nearest point on the centroid line, the point is classified as normal; and if the distance is not less than the cluster radius at the nearest point on the centroid line, the point is classified as abnormal.
[0228] If diagnostic system 104 determines at 910 that the point is within the selected cluster, then diagnostic system 104 proceeds to 912, where it classifies the cell as normal and indicates that the cell has been classified. If diagnostic system 104 determines at 910 that the point is not within the selected cluster, then diagnostic system 104 proceeds to 914, where it classifies the cell as abnormal and indicates that the cell has been classified. The same data field can be used to indicate whether a cell is unclassified, classified as normal, or classified as abnormal. Alternatively, two or more data fields can be used to separately indicate whether a cell has been classified, and if so, whether the cell is normal or abnormal.
[0229] Diagnostic system 104 proceeds from 912 or 914 to 916, where it determines whether all cells associated with the selected cluster have been classified. If the answer at 916 is NO, diagnostic system 104 returns to 910. If the answer at 916 is YES, diagnostic system 104 proceeds to 918. At 918, diagnostic system 104 determines whether all clusters in the cluster group have been processed. If the answer at 918 is NO, diagnostic system 104 returns to 908. If the answer at 918 is YES, diagnostic system 104 proceeds to 920, where subroutine 900 stops.
[0230] Some implementation schemes for System 100 are available. Figure 22 Other actions not shown may be omitted. Figure 22 All the actions shown in the diagram may be performed in different orders. Figure 22 The subroutine 900 can be modified to process data groups sequentially instead of processing data clusters at once, and it can also omit an indication of whether data points have been classified. Subroutine 900 can also call other subroutines; for example, it can call a subroutine to calculate the distance between a point and the nearest point on the centroid line.
[0231] The data generated by System 100 (including data generated for defining normal cell lineages and data from test cell groups) can be represented in various formats and used for various purposes. For example, as mentioned above, the data can be displayed as multiple 2×2 projections of multidimensional data in a Cartesian coordinate system, or as pseudo-three-dimensional projections of multidimensional data in a Cartesian coordinate system. See Figures 4-10 and 12-17 and 20 discussed above. Color or shading can be used to indicate additional dimensions. These methods of displaying data are particularly useful for user-defined and redefinition of the normal centroid and radius of a given mature lineage.
[0232] The data can also be displayed as a two-dimensional plot of continuous cell frequencies along a defined centroid line. Positions along the centroid line correspond to time measurements during maturation. Therefore, a histogram can be generated showing the group distribution of cells throughout maturation. Figure 23A and 23B (Collectively referred to as Figure 23) shows a plot of the sampling frequencies of continuous cells for B lymphocyte lineages along a defined normal centroid line. Horizontal axis 1 corresponds to the location along the defined centroid line. Four clusters 2, 3, 4, and 5 corresponding to the maturation stage are identified along horizontal axis 1. Vertical axis 6 corresponds to the number of points in the data set at each sampling point along the centroid line. In Figure 23, 108 sampling points are selected for the centroid line as follows. Ten reference points are identified along the centroid line. The midpoints of the ten reference points along the centroid line are calculated, resulting in 19 points. Then, the six midpoints of these 19 points are calculated, resulting in 108 points. The percentage of total data points sampled for each cluster is also shown.
[0233] Additional samples can be used to define the normal centroid and radius. For example, the dual-tube 4-color plate process described above can be used to stain large quantities of bone marrow samples exhibiting normal antigen expression. These samples can be selected from standard workflows and may include samples from bone marrow donors, patients without hematologic malignancies, and post-transplant patients with 100% donor chimerism who have received transplants for non-ALL disease. Samples may include both pediatric and adult samples. Additional samples can be randomized or selected according to desired criteria (such as sex, age, or minority group). Selection by sex, age, or minority group is not expected to result in significant differences in the normal centroid and radius for the definition of B lymphocyte maturation lineage.
[0234] The expanded dataset can be used to assess variability in cluster location among individuals from whom samples were collected, as well as compositional differences expected in routine analysis of the samples. This dataset may also include and / or be compared with data from patients with aberrant bone marrow samples that are not the result of clonal or tumorigenesis processes, such as samples from early post-stem cell transplant patients containing only the most immature cells, or from patients treated with Rituxan (anti-CD20). In these patients, B lymphocyte development in the bone marrow is truncated at the onset of stage II, and any cells expressing CD20 are eliminated by the drug. The dataset can also be compared with peripheral blood samples containing only stage IV cells.
[0235] Figure 24 This is a flowchart of example subroutine 1000, which can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104 shown employs this method to compare the test data set with a normal cluster set defined by the centroid. This information can be used by the diagnostic system to, for example, determine whether the defined normal cluster set should be redefined.
[0236] The subroutine begins at 1002 and proceeds to 1004. At 1004, diagnostic system 104 retrieves the test data set and proceeds to 1006. At 1006, diagnostic system 104 assigns the points in the test data set to clusters, as discussed elsewhere in this document (e.g., using gating, using clustering algorithms, using support vector machines, etc., and various combinations thereof), and proceeds to 1008. At 1008, diagnostic system 104 determines the centroid of each cluster in the test data set, as described above. For example, the diagnostic system can determine the parameter value of the centroid of the cluster by adding a corresponding parameter value to each point in the cluster and dividing the result by the number of points in the cluster. Alternatively, the diagnostic system can use statistically adjusted centroids for the test data set. Diagnostic system 104 proceeds from 1008 to 1010.
[0237] At point 1010, diagnostic system 104 determines the corresponding statistical centroid for each cluster based on the previously analyzed data sets. For example, the parameter values for the statistical centroids can be determined by adding the corresponding parameter values for the centroids used to define a set of previously analyzed data sets and dividing the result by the number of data sets. Diagnostic system 104 proceeds from 1010 to 1012.
[0238] At point 1012, the diagnostic system 104 determines whether the cluster's error criteria in the test data set are met. For example, the diagnostic system 104 can compare the logarithm of the distance between the cluster's centroid and its corresponding statistical centroid with a threshold (e.g., 2.5). If the logarithm of the distance is greater than the threshold, the diagnostic system 104 can determine that the error criteria are not met. Other error criteria may be used.
[0239] If the diagnostic system 104 determines at 1012 that the error criteria for the cluster in the test data set are not met, then the diagnostic system 104 proceeds to 1014, where an error indication is set for the cluster in the test data set. If the diagnostic system 104 determines at 1012 that the error criteria for the cluster in the test data set are met, then the diagnostic system 104 proceeds to 1014, where an error-free indication is set for the cluster in the test data set.
[0240] Diagnostic system 104 proceeds from 1014 or 1016 to 1018, where it determines whether all clusters in the test data group have been evaluated. If diagnostic system 104 determines at 1018 that not all clusters have been processed, it returns to 1012. If diagnostic system 104 determines at 1018 that all clusters in the test group have been evaluated, it proceeds to 1020, where subroutine 1000 stops.
[0241] Some implementation schemes for System 100 are available. Figure 24 Other actions not shown may be omitted. Figure 24 All the actions shown in the diagram may be performed in different orders. Figure 24 The action. For example, subroutine 1000 can be modified to sequentially compare all data groups in the normal data group to determine which data groups should be removed from the normal data group.
[0242] Once the cluster boundaries (normal centroid and radius) of the normal maturation lineage are defined, the test samples can be analyzed by subjecting them to the same reagent exposures and measurement protocols used to define the normal maturation lineage. The results from the test data samples can then be compared to the defined normal maturation lineage, allowing for the characterization and diagnosis of the test samples. Systems (such as...) Figure 1The system 100 shown only needs to provide a definition of normal cluster boundaries to diagnose test samples. Alternatively, the system 100 can provide a defined normal data set and defined centroid lines and radii, or the system 100 can provide a defined normal data set and define the normal cluster boundaries.
[0243] Figure 25 A data structure 1100 suitable for defining normal boundaries for cell lineages is shown. The data structure 1100 and corresponding instructions can be stored in a computer-readable medium, such as memory, which may include… Figure 1 The memory 112 shown, or portable memory such as CD ROM, floppy disk and / or flash memory, and / or signal transmission in a signal transmission medium (such as a wired or wireless medium). Data structure 1100 has a header portion 1102 describing the location of other parts of data structure 1100. Text portion 1104 contains information describing various aspects of data structure 1100, such as the number of clusters and how centroids and radii are defined. For example, a centroid can be defined by providing parameters for interpolation equations or by providing reference points to be connected together or a combination thereof. Similarly, a radius can be defined by providing parameters for interpolation equations or fixed radius values for clusters or a combination thereof. For example, a radius can have a fixed value within a cluster and can be a function of a position within a second cluster. Centroid data portion 1106 of data structure 1100 contains information defining centroids, and radius data portion 1108 contains information defining radii. If desired, normal data sets for defining normal centroids and radii can be provided as data structure 1100 or separate data structures (such as...). Figure 2 The data field in the data structure 200 shown.
[0244] Individual clusters can be further decomposed into sub-clusters, which can be defined and analyzed using processes similar to those discussed above. For example, they can be modified. Figure 21 The subroutine 800 shown defines the centroid line or point and radius of the subcluster, and can be modified. Figure 22 The subroutine 900 shown determines whether a cell test set contains a subcluster corresponding to a defined subcluster. It is anticipated that dbscan clusters will be particularly useful for identifying subclusters corresponding to sub-maturity levels within a cluster, which in turn correspond to maturity levels within a cell lineage.
[0245] System 100 can be used to diagnose a test dataset by comparing it to the normal centroid line and radius defined by the cell lineage. The entire test dataset can be compared to the defined normal and diagnosed by a diagnostic system (such as...). Figure 1The diagnostic system 100 displays the data on a suitable display device or medium (such as raster scanning, active or passive matrix display) or on a passive medium (such as paper or kraft paper). Alternatively, data events in the test data set located in the "normal" position, particularly B-lineage lymphoblasts, can be subtracted from the test data set, leaving an "abnormal" data set corresponding to the residual population of potentially "abnormal" cells (leukemia lymphoblasts). This data can then be processed by the diagnostic system (such as...) Figure 1 The diagnostic system 104 shown herein analyzes and displays the remaining anomalies to the user. The remaining anomalies can define anomaly subgroups within the test data set. Clustering techniques (such as those discussed above) can be used to identify clusters using anomaly subgroups of the test data set, and statistical analysis can be employed to determine whether any identified clusters within anomaly subgroups are significant.
[0246] System 100 can be tested before being used to diagnose cancer. For example, numerous samples from patients with obvious ALL can be stained and data collected for comparison with normal samples. These samples are expected to have identifiable normal cells that System 100 will recognize, as well as CD19-positive leukemia cells that will not fall within the boundaries defined by the normal centroid line and radius. It should be noted that B-lineage ALL leukemia cells all express CD19 and, therefore, will be included in the original gating strategy.
[0247] The testing of System 100 may include mixing data from ALL patients with different proportions of normal samples to simulate residual disease detection. For example, System 100 may process 25 normal samples and generate a defined centroid line and radius for a normal mature lineage, which System 100 may store as a digital object in memory 112. This information can be looped back using a statistical algorithm on a data file containing aberrant cell clusters. Cell events confined to regions of normal clusters can be removed, while the remaining events represent “abnormal” clusters. The number and location of expected tumor cells in the mixture can be compared with those identified. This can be done before and after subtracting “normal” cells from the test data set.
[0248] Smoothing algorithms (including averaging and filtering algorithms) can be used to smooth the representation of data. For example, a portion of a cluster can be averaged. For instance, it may be known that the average maturity level of a portion of a specific cluster is an important indicator of whether a test sample is normal, but the individual differences within that portion of the cluster are not significant.
[0249] Data from two datasets can be displayed simultaneously in this way. For example, data from the test sample can be overlaid on the data used to define the normal centroid. A first color or other indicator can be used to represent the normal distribution, and a second color or other indicator can be used to represent the distribution of the test sample.
[0250] Simplified data display can be employed to compare the visual impact and ease of interpreting normal and / or abnormal development. For example, the proportion of cells in each of the four B lymphocyte lineage stages can be plotted to represent identifiable clusters in the data space. Total events in each of the four clusters can be displayed to represent the maturation of cells within normal bone marrow and / or test samples against a normal representation. Parameters for depicting abnormal cells include: the number of abnormal events, distance from normal, dispersion within the abnormal population, and cellular markers that distinguish abnormal cells from normal cells.
[0251] Figure 26 A simplified representation of data collected from test samples is shown, superimposed on a representation of a defined normal data set. The horizontal axis 1 corresponds to an indication of the maturity level of the cell lineage, and this indication corresponds to four stages 2, 3, 4, and 5 within the maturity level clusters within the cell lineage. The vertical axis 6 corresponds to an indication of the number of cells at various maturity levels. This indication can be, for example, a percentage or logarithmic indication of the total number of cells within a stage. The bars 7 show the defined normal range for the sample. Bar 7 can correspond to, for example, the standard deviation of the normal cell group, or it can correspond to the centroid and radius of the defined normal cell set. The dashed line 8 shows the results for the test sample.
[0252] Quality control processes can be employed. For example, bead formulations can be used to evaluate instrument performance, such as rainbow beads (RCP and RFP, Spherotech, Libertyville, IL), which are plastic microspheres in which dye is embedded inside the particles to ensure fluorescence stability. RFP beads have only a single peak in each of the four fluorescence channels and are used as the primary standard. RCP beads (a mixture of six intensity beads observed in all channels) are used as a secondary standard and provide linear data about each fluorescence detector. Fluorescence emission spectral compensation is established and monitored by staining normal blood with anti-CD4 antibodies (FITC, PE, PerCP, and APC) conjugated to each chromophore used. Cells stained with these antibodies are analyzed separately to ensure that fluorescence from the expected chromophore is detected only in the appropriate fluorescence channel (24). Each batch of reagents used in cell evaluation is titrated and then placed in stock. The antibody titer that produces the maximum fluorescence intensity is selected, and the reagent specificity of each new batch of antibodies is checked.
[0253] Using these quality control procedures, two flow cytometers experimentally produced identical results for the same sample. In studies using normal adult blood with these procedures, the intensity of CD4 on lymphocytes was found to be virtually unchanged across 21 individuals measured on both instruments over an 8-month period. The mean fluorescence intensity of CD4 in these 21 individuals was 1596 + / - 116 standard deviation fluorescence units, resulting in a CV of 7%. These results suggest that the individual-to-individual biological variability of this antigen is essentially zero within a data space with a dynamic range of forty years. The amount of CD4 expressed on lymphocytes is itself a biological standard. The quantification of centroid location (measured on immature bone marrow cells) can be compared with the variability of antigen expression on normal mature blood cells, which will provide a basis for understanding the biological variability between individuals regarding the intensity of antigen expression during the maturation of blood cells (not just mature cells).
[0254] The system (e.g.,) can be determined by changing known quantities (factors of 2 and 4) the target value of the primary standard fluorescence quality control beads. Figure 1 The tolerance of the system (100) shown is such that the system can be demodulated using known quantities. After appropriate compensation is established, each channel can be tested individually or together. For example, bone marrow cells stained with four color combinations can be collected in each setting, and the data can be analyzed using the system under test. This will assess how far the system is from the optimal standard setting for operation while still allowing the system to correctly identify cells at developmental stages. This performance is then used to define the tolerance required for quality control procedures based on the system's ability to identify appropriate cell populations.
[0255] As mentioned above, multidimensional analysis can be used to detect abnormal cells. Typically, the centroid line is defined to simulate cell lineage maturation based on a normal patient dataset (which usually includes data from multiple normal patients). The normal radius of variation around this line is defined based on the normal patient dataset.
[0256] Subsequently, the cell test group (e.g., cells from a patient) can be characterized based on the identification of cells outside that radius. Such cells can be identified, for example, in terms of percentage, location, etc., and the cell test group can be classified as normal or abnormal based on the identification of cells outside that radius. In the following description, references to cells in a cell group may refer to cells in the cell group to be exposed to the defined protocol, or may refer to the corresponding data point in a set of data points generated based on flow cytometry, as indicated in the context in which the reference is used.
[0257] Individual variation in flow cytometry data defined based on an internal reference population.
[0258] The use of flow cytometry to detect hematologic malignancies is based on identifying cells that do not exhibit the antigens or physical properties expected of normal hematopoietic cells. This method depends on how accurately normal blood and bone marrow cells can be identified using quantitative antibody binding and physical characteristics such as light scattering. A powerful combination of features has been used to classify cells into different lineages, known as CD45 gating, which displays data from a sample combining CD45 intensity and light scattering (side scattering, right-angle light scattering, or orthogonal light scattering). See, for example, Stezler GT, Shuls KE, Loken MR. “CD45 gating for routine flowcytometric analysis of human bone marrow specimens.” Ann NY Acad Sci 1993;677:265-80.
[0259] Determining the composition of bone marrow using this technology is daunting because the tissue comprises at least 11 different cell types and a full range of immature cells, from hematopoietic stem cells to mature cells in the blood. The applicant has recognized that understanding the individual-to-individual variability for all cellular characteristics is helpful in identifying, quantifying, characterizing the immunophenotype and physical features of abnormal cells, and then classifying abnormal cells based on the most recent normal cell composition. The limitation of this method depends on the variability observed between individuals for these assay parameters over a period of time. Based on this understanding, the applicant has developed methods to reduce the variability observed in normal cells for each lineage between individuals and to reduce the analytical variance used to detect those characteristics (due to differences arising from sample processing, multiple reagent batches, different analytical instruments, different analysts, etc.).
[0260] The first step in analyzing the composition of a bone marrow sample is to identify key reference cell populations within the sample. Selected reference populations form clusters in the multidimensional data space and can be definitively identified using appropriate reagent combinations. Subjective bias in identifying these specific cell populations can be reduced by using automated analytical procedures. These reference populations can then be used to improve the identification and analysis of other cell populations in the multidimensional data space. The applicant has recognized that the relationships of cellular characteristics among these reference populations are surprisingly constant and that variability can be reduced by normalizing the data sets relative to individual cell populations. This facilitates the standardization of bone marrow analysis and allows for centroidal analyses comparing data from different patients for each mature cell lineage.
[0261] Certain reference populations can be reproducibly identified in bone marrow samples from normal individuals. These populations include: mature lymphocytes, amorphous progenitor cells, promyelocytes, mature monocytes, and mature neutrophils. Each of these populations can be specifically identified by a combination of antibody and light scattering characteristics, as listed in Table 3 below.
[0262] Mature lymphocytes High CD45, Low FSC, Low SSC Undifferentiated progenitor cells Bright CD34, CD33 positive, low SSC intermediate CD45 Promyelocytes HLA-DR- / CD11b-, High SSC, Middle CD45 Mature monocytes CD14+, High CD33+, High CD45, SSC in between Mature neutrophils High CD13, Medium CD33, High CD45, High SSC
[0263] Table 3: Phenotypic and light scattering characteristics of the reference population
[0264] Reference populations can be automatically identified. This analysis can be used to eliminate or reduce subjective bias from technicians or others analyzing the data. Furthermore, automated analysis simplifies the process. One implementation uses a machine learning analysis called Support Vector Machine (SVM), as discussed in more detail below. The SVM can be taught by providing a series of examples where “experts” manually identify each reference population of interest. The SVM identifies mathematical characteristics associated with these expert identifications of cell populations and uses these mathematical characteristics to find such populations in subsequent data from different patients. Thus, the SVM provides a reproducible method for mathematically identifying reference populations of interest, which reduces subjective bias in the analysis and facilitates the automation of the process. The SVM method was tested on pediatric patients (n=50) who had recently undergone treatment for acute myeloid leukemia (AML). These “stressed” bone marrow samples, recovered after chemotherapy, were randomly selected from those patients who did not show residual AML, with additional criteria being that the samples were of high quality in terms of sufficient cell count, lack of blood dilution, and minimal dead cells.
[0265] Once reference populations are identified using SVM, the mean and standard deviation of antigen intensity for cells (data points) within these reference populations can be calculated. These mean antigen intensities exhibit inherent (but small) variability among patients. This positional variability within each of these reference populations can be further reduced by normalizing the data relative to the position of individual cell populations. For example, the CD45 and SSC positions of progenitor cells can be normalized for each patient to the CD45 and SSC positions of individual mature lymphocytes. In another instance, the CD33 and CD45 intensities of monocytes can be normalized to the CD33-to-CD45 intensities of each patient's respective progenitor cell population, as referenced below. Figure 29 This will be discussed in more detail. While the absolute amount of CD33 antigen was not highly modulated, this normalization showed that differences in CD33 intensity were consistent across cell populations within an individual patient. Overall, this data normalization approach significantly reduced positional variability for each reference population, thus allowing for even more precise and standardized assessments of the disease.
[0266] Use support vector machines to select the reference group of interest.
[0267] As discussed elsewhere in this document, users can filter certain flow cytometry parameters (e.g., CD19, SSC, etc.) to select on which mature cell populations (e.g., reference cells) are obtained.
[0268] Traditionally, algorithms used to analyze flow cytometry data are unsupervised because they search for patterns in the data. For example, many automated flow cytometry analysis algorithms search for clusters of cells with similar fluorescence characteristics. It's important to note that for these algorithms, the quantified fluorescence characteristics are not crucial; all that matters is the existence of groups of “similar” cells that can be clustered together. Current automated flow cytometry analysis procedures use this type of unsupervised approach because fluorescence intensity data is inconsistent—and therefore, finding similar cell groups allows for the discovery of homologous cell populations.
[0269] In one implementation, a Support Vector Machine (“SVM”) can be used to select cell populations (data) to achieve maturity within a cell lineage. SVM can be used to define a multidimensional boundary to select either a cell population of interest or a reference population. This boundary is independent of the frequency of the reference population. If the reference population exists at a low frequency, clustering algorithms may struggle to identify it. SVM locates positions in space to detect the reference population rather than its frequency.
[0270] Support Vector Machines (SVMs) can be considered supervised machine learning algorithms, meaning that SVMs are fed data to "learn" how to identify cell populations of interest. In one study, SVM classification of cell types was fed to an SVM, and the SVM identified mathematical features in the classification based on quantitative fluorescence characteristics (rather than on, for example, statistical clustering techniques to identify cell clusters or cell groups). In this study, a meticulously constructed and quality-controlled flow cytometry assay produced incredibly stable fluorescence intensity measurements, which helped predict whether an individual cell belonged to a population based on its fluorescence intensity characteristics. This facilitated the identification of entire cell lineages (at all stages of maturity) without any user input (gating) in the absence of a single antibody defining the group of interest.
[0271] Previously, multiple lineages had to be defined manually by analysts. Most flow cytometry experts struggled to separate cells from different lineages, particularly monocytes and neutrophils. In one implementation, using SVM helps automate this process because SVM automates cell classification into different lineages, even without specific lineage markers. This facilitates assessing the maturity of each lineage separately and classifying each lineage as normal or abnormal.
[0272] Figure 27 and 28This is a flowchart of an example SVM subroutine 2000, which can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104) employs a method to identify cell populations of interest in a dataset using multidimensional boundary definitions generated using SVM. For convenience, reference will be made to... Figure 1 The diagnostic system 104 discusses subroutines 2000. These definitions can be used by the diagnostic system to, for example, identify or refine groups of cell populations of interest to achieve maturation within the centroid (e.g., assigning cells to clusters in a normal cluster group), identify reference populations for defining or refining centroid lines or the radius of centroid lines or segments thereof, perform vector normalization, classify cell test groups, etc.
[0273] As shown in the figure, the SVM subroutine 2000 includes a training phase 2700 and an implementation phase 2730. In the training phase 2700, the SVM subroutine 2000 is taught to automatically identify cell populations. In the implementation phase 2730, the SVM subroutine 2000 identifies cell populations of interest (such as a reference cell population, a cell population from a new test patient, etc.), which can be used for, for example, to assign cells to clusters of normal cell groups, perform vector normalization, obtain population maturity via centroid lines, refine the radius, characterize cell test groups, etc.
[0274] The subroutine begins at 2702 and proceeds to 2704. At 2704, the diagnostic system 104 selects the test cell population of interest from the data set. This can be accomplished, for example, using an existing software platform (e.g., Winlist, Java implemented with a Java runtime environment, 3D Java runtime environment, etc.) to manually set a series of one or more gates to select the cell population of interest from the data set. The diagnostic system proceeds from 2704 to 2706.
[0275] At 2706, the diagnostic system 104 generates a dataset for identifying the selected cells. For example, the dataset for identifying the selected cells can be exported in text file format to the floating-point array 2708, which has one column for each measurement parameter and an additional classification column, as well as one row for each cell in the dataset. As shown, the classification column contains a binary evaluation for each cell in the test dataset, -1 if the cell is not included in the defined population, and +1 if the cell is included in the defined population. Other data formats can be used, such as comma-separated value files, etc.
[0276] Subroutine 2000 proceeds to 2710, where the diagnostic system 104 determines whether there is another set of data to be processed. For example, multiple sets of data corresponding to cells from multiple normal patients may be used to train the SVM. When it is determined at 2710 that there is another set of normal patient data to be processed, subroutine 2000 returns to 2704 to process the next set of normal patient data. Otherwise, subroutine 2000 proceeds to 2712.
[0277] At 2712, the diagnostic system 104 combines data sets indicating the cell population of interest. In one study, this was done by reading floating-point arrays (e.g., floating-point array 2708) from each normal data set and merging the floating-point arrays into a combined data set. Subroutine 2000 proceeds to 2714.
[0278] At 2714, diagnostic system 104 generates an SVM that identifies a multidimensional decision boundary for the cells of interest in the combined dataset (e.g., an SVM that separates cells evaluated as +1 from those evaluated as -1). The decision boundary can typically be a complex multidimensional shape. Cells on one side of the boundary are classified as belonging to the group of interest (e.g., classified as +1), while cells on the other side are classified as not belonging to the group of interest (e.g., classified as -1). How to generate SVMs as prediction algorithms is known, and these known techniques can be applied to combined normal datasets to generate SVMs that identify multidimensional boundaries. See, for example, Chang, Chih-Chung, and Chih-Jen Lin. "LIBSVM: a library for support vector machines." ACM Transactions on Intelligent Systems and Technology (TIST) 2.3 (2011):27. Subroutine 2000 proceeds to 2716.
[0279] At 2716, the diagnostic system 104 optionally evaluates the predictive performance of the SVM, and may adjust, for example, the cost and gamma factor, the number of normal patient data sets used to generate the combination for training the SVM, etc., based on this evaluation. For example, leave-one-out cross-validation may be used. See, for example, Golub G, Heath M, Wahba G. Generalized cross-validation as a method for choosing a good ridge parameter. Technometrics 1979; 21(2):215-23. The process is typically a two-step process. The result is optimized and then evaluated. Cross-validation is a method used to optimize the input parameters in an algorithm. Assume a training data set of 25 patients. For a fixed combination of input parameters (such as the cost of the SVM, the gamma input parameter), the algorithm is cross-validated by training on a subgroup of the training data (e.g., 24 patients instead of 25 patients) and then testing on the remaining patients (e.g., 1 patient). The errors generated by the SVM are summed, and the process is repeated such that each patient is a test patient exactly once. For a specific combination of input variables, the total error (from 25 repetitions with the test patient) is calculated and stored. Then, the input variables (cost, gamma) are adjusted, and the training (n=24) and testing (n=1) processes are repeated in the same manner for new combinations of input variables. During the evaluation phase, the total error from each combination of input variables is compared, and the combination of input variables with the lowest total error is used to train the resulting SVM. Subroutine 2000 proceeds to 2718, where the training phase ends.
[0280] The implementation phase 2730 of the subroutine begins at 2732. The subroutine proceeds from 2732 to 2734. At 2734, the diagnostic system 104 classifies each cell in the test patient cell group (e.g., cells of test patients who may or may not be normal patients) using the multidimensional decision boundary defined in the training phase. In other words, each cell in the test cell group is classified as +1 or -1 based on the side of the decision boundary in which the cell resides. Subroutine 2000 proceeds from 2734 to 2736. At 2736, the diagnostic system 104 optionally applies additional filtering criteria (e.g., filtering based on certain flow cytometry parameters (such as CD19, SSC, etc.), which may be done based on default settings, in response to user input, etc.). Subroutine 2000 proceeds to 2738, where it stops.
[0281] Some implementation schemes for System 100 are available. Figure 27 and 28 Other actions not shown may be omitted. Figure 27 and 28All the actions shown in the diagram may be performed in different orders. Figure 27 and 28 The subroutine 2000 can be modified in some embodiments to perform only the training phase 2700 and in other embodiments to perform only the implementation phase 2730 (e.g., the first diagnostic system 100 can be used to train the SVM, and the second diagnostic system can be used to apply the defined multidimensional boundaries to the test dataset; the SVM can be stored for reuse instead of being generated, etc.). In another instance, the subroutine can determine after 2716 that an additional normal dataset should be used, for example, in response to an indication that the generated boundaries are not sufficiently reliable as predictors of the normal cell population, and thus proceed from 2716 to 2704 to add an additional normal patient training dataset. In another instance, a temporary dataset from the training phase, such as a floating-point array generated at 2706, can be stored for validation at 2716. In yet another instance, the subroutine can be modified to identify multiple populations of interest within a dataset or a combined dataset, or to identify a subpopulation of interest within an identified reference population. For example, a first reference population or subpopulation can be identified, corresponding to the first maturation stage within a cell lineage; a second reference population or subpopulation can be identified, corresponding to the second maturation stage within a cell lineage; and so on. In another example, a first reference population can be identified, corresponding to all maturation stages within a first cell lineage (e.g., B lymphocyte lineage cells), and a second reference population can be identified, corresponding to all maturation stages within a second cell lineage (e.g., monocyte lineage cells); and so on.
[0282] SVMs are trained in a two-stage setup. Multiple SVMs can be used to identify multiple reference populations. For example, to identify B lymphocytes and undifferentiated progenitor cells, two separate SVMs can be used and then merged. In another example, a primary SVM can be trained to identify all B lymphocytes (CD19+) cells, and then several secondary SVMs can be trained to identify the stages (stages 1-4) of these B lymphocytes within the B lymphocyte (CD19+) population.
[0283] In this study, multidimensional boundary definitions for SVM generation have been created for B lymphocyte lineage cells and other cell lineages (e.g., monocytes, lymphocytes, erythrocytes, neutrophils, dendritic cells, eosinophils, basophils, NK cell lineages, plasma cells, and mast cell lineages). The plan is to investigate T cell subgroups. For example, the identified reference cell populations and subpopulations can be used to identify clusters of normal cells (data points), which are then used to define the centroid of the normal cluster group in vector normalization (as described below). Figures 18A-18C Subroutine 500 can be modified from... Figure 18BThe cluster identification begins with subroutine 2000 at position 528, where subroutine 500 continues to identify the centroid line of each cluster, and then, based on the reference population identified using subroutine 2000, defines the normal centroid line and radius of the cluster group; this can be modified. Figure 19 Subroutine 600 identifies a reference point at 604 based on the reference population or subpopulation identified by subroutine 2000; it can be modified. Figure 21 Subroutine 800 is used to identify reference points based on a reference population or subpopulation identified by subroutine 2000; etc. For example, a defined multidimensional decision boundary can be used to classify cells in a test cell group (e.g., it can be modified). Figure 22 Subroutine 900 applies the multidimensional decision boundary defined by subroutine 2000 at step 910 to determine whether points in the retrieved data set (e.g., the test patient data set) are in a cluster; etc. SVM can also be used for quality control of measurement systems. For example, mature lymphocytes are stable and generally unaffected by chemotherapy in patients with acute myeloid leukemia. Identification of lymphocytes and calculation of lymphocyte characteristics can provide an indication of whether the instrument is set up correctly. For each patient, a reference population can be used for quality control. If a normal reference population requires too much normalization (e.g., if the normalization vector, as described below, is too large), the data set may be flagged, which could indicate a problem with the entire data set. If multiple separate data sets are flagged, it may indicate a problem with the instrument.
[0284] Vector normalization of antigen intensity
[0285] The above-described implementation scheme captures the distance of normal cells from the centroid. The radius captures biologically variable components, which may be due to the variable expression of surface antigens and thus variations in distance from normal cells to the centroid. The radius also captures technical variations in the analytical components, which may be due to, for example, different instrument settings and tolerances, different antibody transport, etc. In one implementation scheme, vector normalization is employed to account for and reduce technical and fluid variations. This helps to define a tighter and more specific radius around the centroid, which more accurately represents the biologically variable components, and facilitates easier and more accurate identification of aberrant cells, as well as promoting focus on specific groups of aberrant cells, etc.
[0286] In flow cytometry, the intensity of a marker depends on the cell's position as it passes through the laser beam. If the cell trajectory deviates slightly from the center, the cell reading will be lower than if the cell were centered within the laser beam. Due to the cell's position relative to the laser flow, the cluster may widen as the flow rate increases.
[0287] In one implementation of the vector normalization process, a standard reference mean is determined for the normal patient group. The standard reference mean is the average multidimensional intensity of this normal patient group. The cell intensities of new patients are then normalized to the standard reference mean. SVM can be used to identify cell populations of interest in determining the standard reference mean and the normalization of cell intensities in new patients.
[0288] Figure 29 This is a flowchart of example vector normalization subroutine 3000. Subroutine 3000 can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104) was used to normalize the cell population of the test cell group. For convenience, the reference cell population will be used. Figures 29-39 and Figure 1 The diagnostic system 104 discusses subroutine 3000. As shown in the figure, subroutine 3000 performs calculations in logarithmic space rather than linear space.
[0289] Subroutine 3000 includes a first phase 3100 for determining a standard reference mean and a second phase 3200 for normalizing the cells of the new patient to the determined standard reference mean.
[0290] The first phase 3100 of subroutine 3000 begins at 3102 and proceeds to 3104, where the SVM is trained to identify a reference population. For example, training phase 2700 of subroutine 2000 could be used. The user selects a cell population to use as a reference for subsequent data normalization and defines a multidimensional boundary for identifying the population of interest (see, for example,...). Figure 27 2704 to 2718). Selection process (e.g., Figure 27 (2704) can be based on biological knowledge, such as the knowledge that certain cell populations may exhibit different advantages in intensity normalization. For example, lymphocytes are easily identified in a sample and can be used to accurately normalize CD45 intensity; monocytes are slightly more challenging to identify but can be used to more accurately normalize CD33 intensity, which varies more between individuals, etc. The variability in CD33 intensity is a result of specific single nucleotide polymorphisms (differences) in DNA, called SNPs. Normalization using a reference population identified using SVM can reduce the general variability in CD33 intensity. Depending on the cell type the user wants to assess (e.g., for leukemia), the user can choose different reference populations to normalize the data. In one study, six dimensions were used in the selection process.
[0291] Subroutine 3000 proceeds from 3104 to 3106. At 3106, a multidimensional boundary is applied to identify a reference population of interest among known normal patients. Diagnostic system 104 can be used, for example, for known normal patients. Figure 28 The implementation plan of subroutine 2000 in the implementation stage 2730. Figure 30 Example lymphocyte populations of interest, indicated by purple, are shown for the SSC and CD45 parameter intensities. Reference populations of interest may contain hundreds of thousands of cells or a relatively small number of cells (e.g., undifferentiated progenitor cells). For ease of illustration, illustrations of populations of interest with other intensities (e.g., other pairs of the six parameters employed) are omitted. Subroutines 3000 proceed from 3106 to 3108.
[0292] At 3108, the average intensity of each parameter in the reference population of known normal patients is calculated by summing the cell intensities of the populations of interest for a given parameter and dividing by the total number of cells identified in the populations of interest. Figure 31 An example plot shows the mean intensity of the SSC and CD45 parameters (represented by purple dots) for the lymphocyte population of interest in a single normal patient in the study. For clarity, plots showing the mean intensity of other intensity pairs or populations of interest in multiple dimensions are omitted (e.g., six parameters were used in the study as described). The mean intensity of each of the six parameters is calculated. The mean reference intensity for each of the six parameters in a normal patient is stored.
[0293] At point 3110, the diagnostic system determines whether there is another set of normal patient data to be processed. If an additional set of normal patient data to be processed is determined at point 3110, the subroutine returns to point 3106 to process the normal patient data set. If no additional set of normal patient data to be processed is determined at point 3110, the subroutine proceeds to point 3112. Figure 32 The example mean intensity of the 27 normal patient data groups in the study relative to the SSC and CD45 parameters is shown, with each intensity indicated by a purple dot. For clarity, plots of the mean intensity of the 27 normal patient data groups in other intensity pairs or other dimensions (e.g., using 6 parameters as described) are omitted.
[0294] At position 3112, the standard reference mean is calculated for the normal reference group. The standard reference mean is the average of all mean reference intensities for each parameter in the normal patient data group. Figure 33 The standard reference average for the SSC and CD45 parameter intensities is shown as purple dots in the study. For clarity, illustrations of standard reference averages for other intensity pairs or other dimensions are omitted. In this study, the standard reference average is a vector with six parameters, such as... Figure 34 As shown in the figure, these results are rounded.
[0295] The first-stage procedure can be repeated to calculate the standard reference mean intensity vector for each desired reference population (e.g., monocytes, erythrocytes, lymphocytes, amorphous progenitor cells, neutrophils, promyelocytes, etc.), ending at 3114. In this study, the first-stage procedure was performed on lymphocyte lineages and other cell lineages (including monocytes, amorphous progenitor cells, neutrophils, and promyelocyte lineages).
[0296] The second phase 3200 of subroutine 3000 begins at 3202 and proceeds to 3204. At 3204, a reference population is selected based on biological knowledge (e.g., one of the reference populations whose standard reference mean intensity was determined in the first phase), and a corresponding multidimensional boundary (e.g., the boundary determined at 3104) is applied to identify the reference population of interest for the test patient. The number of cells in the reference population of interest can be hundreds of thousands or more, or a relatively small number of cells (e.g., undifferentiated progenitor cells, mast cells, plasma dendritic cells). The diagnostic system 104 may employ, for example... Figure 28 The implementation plan of subroutine 2000 in the implementation stage 2730. Figure 35 The example lymphocyte populations of interest in this study are shown in purple for the SSC and CD45 parameter intensities. For clarity, illustrations of populations of interest with other intensities (e.g., other pairs of the six parameters employed) are omitted. Subroutine 3000 proceeds from 3204 to 3206.
[0297] At 3206, the average intensity of each parameter in the test patient reference population is calculated by summing the cell intensities of the populations of interest for a given parameter and dividing by the total number of cells identified in the population of interest of the test patient. Figure 36 An example plot shows the average intensity of the SSC and CD45 parameters (represented by purple dots) in a study testing the lymphocyte population of interest from patients. For clarity, plots showing the average intensity of other intensity pairs or multiple dimensions of the population of interest are omitted (e.g., using 6 parameters as shown). The average intensity of each of the 6 parameters is calculated. In this study, the result is a vector with 6 parameters, each corresponding to the average of the corresponding parameter of the population of interest, such as... Figure 37 As shown. The displayed result is rounded.
[0298] Subroutine 3000 proceeds from 3206 to 3208. At 3208, the patient's normalized vector is calculated by determining the difference between the patient's mean parameter intensity and the standard reference mean determined in the first phase 3100. The resulting normalized vectors from the patients in this study are shown... Figure 38 In the last column, as shown in the figure, the patient's mean parameter intensity is subtracted from the standard reference mean.
[0299] Subroutine 3000 proceeds from 3208 to 3210. At 3210, the normalized vector of the patient calculated at 3208 is used to normalize the intensity of each cell in the test patient. In one implementation, the normalized vector can be used to normalize selected cells of the test patient, such as cells from an identified reference population. Figure 39 The diagram illustrates the application of a normalized vector to the first cell of a test patient. As shown, the normalized vector is added to the cell's unnormalized strength. Examples of applying normalized vectors to individual cells in the data set are illustrated in [the diagram / illustration]. Figure 39A In Chinese, it can plot, visualize, and evaluate the normalization strength of test patients for disease applications. For example, it can compare the normalized cells of a test patient with the centroid line and radius of a normal cell population that defines a diagnosis of cancer (such as residual cancer). For example, it can modify... Figure 22 Subroutine 900 is used to classify the cells (data points) of the retrieved data set before... Figure 29 The subroutine 3000 normalizes the reference population of the retrieved data set. When the test patient data set of patients with residual disease is normalized, both normal cells and tumor cells are typically normalized.
[0300] Some implementation schemes for System 100 are available. Figure 29 Other actions not shown may be omitted. Figure 29 All the actions shown in the diagram may be performed in different orders. Figure 29 The actions involved. For example, in some implementations, subroutine 3000 can be modified to determine the difference between the patient's mean parameter intensity and the standard reference mean by subtracting the standard reference mean from the patient's mean parameter intensity, and to determine the normalized intensity of the cell by subtracting the normalized vector from the unnormalized intensity of the cell. In another instance, other or additional normal cluster configurations can be defined and modeled for determining the standard reference mean and mean parameter intensity. For example, in one study, a first tube group of normal patient cells is subjected to a first protocol that produces a first normal patient data set (which may correspond to, for example, ten clusters) with parameters FSC, SSC, CD20 (FITC), CD10 (PE), CD45, and CD19, and a second tube group of normal patient cells is subjected to a second protocol that produces a second normal data set (which may correspond to, for example, four clusters) with parameters FSC, SSC, CD22 (FITC), CD34 (PE), CD45, and CD19. Other protocols and cluster configurations can be employed. The selected protocol can then be applied to the cell group of the test patient to generate the test patient data set.
[0301] Radius definition
[0302] As mentioned above, statistical algorithms can be used to define the normal radius around the centroid line. See, for example, Figures 18A-18C And its discussion. For each parameter measured by flow cytometry, the dimensions of the multidimensional radius can be statistically characterized. These statistical characteristics can be used for data visualization in images (e.g., linear, normalized subtraction, etc.). In a series of studies, cells from 25 or more groups of normal patients were used to mathematically characterize how far normal B cells are from the defined centroid line in each dimension of the radius of each cluster. The larger the number of cell groups from normal patients, the greater the statistical confidence in the definition of the normal data group. This variation is then normalized in the form of a z-score, which can also be called a chi-square analysis.
[0303] After defining normal centroid lines for a series of normal patients (see, for example, Figures 18A-18B (and its description), which can statistically characterize the radius of a normal data set.
[0304] Figure 40 This is a flowchart of the example radius characterization subroutine 4000, which can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104) uses a radius to characterize the normal clustering group in n-dimensional space, which, together with the normal centroid line, defines the normal clustering group. In one study, six parameters were used to define the normal clustering group of the B-cell lineage. For convenience, reference will be made to... Figures 40-45 and Figure 1 The diagnostic system 104 discusses subroutine 4000. Radius characterization of the data set is predicted in two steps. First, cells belonging to the lineage whose radius will be characterized are selected from the data set. Experts can manually identify cells belonging to the lineage. Alternatively, SVM can be used to identify cells belonging to the lineage, as described above in subroutine 2732 (…). Figure 28 As shown in the diagram. Secondly, each cell in the lineage clusters to the nearest corresponding reference point. This can be achieved, for example, using subroutine 1000 (…). Figure 24 The process is completed using action 1006 or a similar process. In other instances, the flowchart of the implementation scheme for the clustering process is shown in... Figures 18A to 18C and Figure 40A As shown in the image.
[0305] After identifying cells belonging to the lineage and clustering them to the reference point, subroutine 4000 begins at 4002 and proceeds to 4004. At 4004, the tangential intersections between the cells of the normal patient data group and the centroid line are identified. The identification of tangential intersections on the centroid line can be accomplished using dot product. An example illustration of tangential intersections is shown in... Figure 41 The example calculation of the tangential intersection point is shown in the diagram. Figure 41A As shown in the diagram, the centroid line can be segmented, and candidate segments can be used to identify tangential intersections of cells, which reduces the number of computations required.
[0306] In one study, 10 mature B lymphocyte clusters were identified. The centroid line was divided into nine segments: the first segment extended from the center of cluster 1 to the center of cluster 2; the second segment from the center of cluster 2 to the center of cluster 3; the third segment from the center of cluster 3 to the center of cluster 4; the fourth segment from the center of cluster 4 to the center of cluster 5; the fifth segment from the center of cluster 5 to the center of cluster 6; the sixth segment from the center of cluster 6 to the center of cluster 7; the seventh segment from the center of cluster 7 to the center of cluster 8; the eighth segment from the center of cluster 8 to the center of cluster 9; and the ninth segment from the center of cluster 9 to the center of cluster 10. Cells in the first cluster would have one candidate segment, i.e., the first segment; cells in the second cluster would have two candidate segments, i.e., the first and second segments; cells in the third cluster would have two candidate segments, i.e., the second and third segments; and so on, with cells in the tenth cluster having one candidate segment, i.e., the ninth segment. In this study, when there was no true tangential intersection between a cell and a candidate segment, the center of the cluster to which the cell belonged was considered the tangential intersection point for that cell.
[0307] Other segmentation schemes can be adopted. For example, the centroid line of 10 clusters of mature B lymphocytes in the study can be divided into 10 segments that roughly correspond to the clusters, and the candidate segments of a cell can be defined as the segments of the cluster to which the cell belongs and the segments of the adjacent clusters of the cluster to which the cell belongs.
[0308] Subroutine 4000 proceeds from 4004 to 4006. At 4006, the diagnostic system 104 calculates the one-dimensional distance between the identified tangential intersection points of the cell and the cell on the centroid line for each parameter. This can be accomplished, for example, using Equation 2 as listed below:
[0309] Parameter distance = parameter 细胞 -parameter TCP [Equation 2]
[0310] Where parameters 细胞 These are the values of the corresponding parameters of the cell, and the parameters TCP These are the parameter values of the tangential intersection points of cells along the centroid line. The parameter distance values of cells can be stored in a floating-point array (see [link to documentation]). Figure 42 Subroutine 4000 proceeds from 4006 to 4008. At 4008, the diagnostic system determines whether there are more cells to be processed for a normal patient. If more cells are determined to be processed for a normal patient at 4008, the subroutine returns to 4004 to process the next cell. If no more cells are determined to be processed for a normal patient at 4008, the subroutine proceeds to 4010. Figure 42 An example is shown of a floating-point array storing the parameter distances of six parameters for a normal patient data set, where each row corresponds to a patient's cell and each column corresponds to a parameter.
[0311] At 4010, the diagnostic system 104 determines whether there is another set of normal patient data to be processed. When it is determined at 4010 that there is another set of normal patient data to be processed, subroutine 4000 returns to 4004 to process the cell for the next set of normal patient data. When it is not determined at 4010 that there is another set of normal patient data to be processed, subroutine 4000 proceeds to 4012.
[0312] At 4012, for the combined normal patient data set, the mean and standard deviation of the distances between cell-to-cell intersections are determined for each parameter within each cluster. For example, this can be achieved using... Figure 24 Action 1006 of subroutine 1000 Figures 18A-18C The clustering process, Figure 40A Clustering processes (described below) are used to identify clusters. A combined set of normal patient data can be, for example, a combined floating-point array of normal patient data generated by merging individual data sets. The combined floating-point array can be used to determine and store the mean and standard deviation of the distances between cell-to-cell intersections for each parameter in each cluster within the combined floating-point array. Figure 43 An example floating-point array is shown that stores the determined standard deviation. Figure 43A An example floating-point array is shown, storing the average distance determined between the tangential intersections of cells.
[0313] Subroutine 4000 proceeds from 4012 to 4014, where the diagnostic system 104 performs a z-score transformation on each normal patient data set based on the standard deviation. This can be accomplished using a floating-point array used to generate the normal patient data set at 4006 based on the standard deviation determined at 4012. The standard deviation is used to normalize the distances in the test patient floating-point array using the z-score transformation. Different parameters measured by flow cytometry have different normal variability. For example, CD45 protein expression is incredibly consistent (with low variability) on B lymphocytes, while FSC (size) has high variability. Normalizing each column of the normal data set with the z-score transformation helps to statistically compare the variability of different parameters, resulting in a standard deviation in each individual column that is by definition equal to 1 (see [link to relevant documentation]). Figure 45 (Discussed below). (Rounding errors may cause slight bias). The distance from each cell to the tangential intersection point is standardized by using a z-score transformation to standardize each column. The mean distance for each cluster is zero and the standard deviation is 1, so a cell located one standardized unit away from the centroid in the CD45 dimension is biologically important for a cell located one standardized unit away from the centroid in the FSC dimension.
[0314] Subroutine 4000 proceeds from 4014 to 4016. At 4016, diagnostic system 104 determines the multidimensional Euclidean distance between the scaled position of each cell and the centroid intersection point. This can be accomplished, for example, using Equation 3 listed below:
[0315] Euclidean distance = SQRT(a 2 +b 2 +c 2 +d 2 +e 2 +f 2 Equation 3
[0316] Where af is the normalized distance from the cell to the centroid line, determined at 5208. A normalized floating-point array can be generated for the normal patient dataset, indicating the distribution of the z-transform, with columns added for the Euclidean distance. Figure 44 An example normalized floating-point array for a normal patient data set is shown. The Euclidean distance calculation can be modified to weight the distances between cell-to-cell intersections differently. Subroutines run from 4016 to 4018.
[0317] At 4018, the diagnostic system 104 determines whether there is another set of normal patient data to be processed. When an additional set of normal patient data to be processed is determined at 4018, subroutine 4000 returns to 4014 to process the cell for the next set of normal patient data. If no additional set of normal patient data to be processed is determined at 4018, subroutine 4000 proceeds to 4020.
[0318] At 4020, the normalized floating-point arrays are merged, and the mean and standard deviation of the normalized Euclidean distances are calculated for each cluster. An example (standard deviation) of the merged and normalized floating-point arrays is provided for the normal data set. Figure 45 As shown in the figure. Since the distance of a single parameter has been standardized, the standard deviation is equal to 1. Figure 45 The values in the table represent the single-component dimension of the radius of the normal data set, as well as the multidimensional Euclidean distance radius. Note that the values in the last column may not be equal to the square root of the sum of the squares of the individual parameters and can vary within each cluster. An example (average) normalized floating-point array for the normal data set is shown in [the table / instructions]. Figure 45A As shown in the diagram. The distances for individual parameters have been standardized, and the mean of each parameter is equal to 0. Note that the values in the last column may not be equal to the square root of the sum of the individual parameters. Subroutine 4000 proceeds from 4020 to 4022, where it ends.
[0319] Some implementation schemes for System 100 are available. Figure 40 Other actions not shown may be omitted. Figure 40All the actions shown in the diagram may be performed in different orders. Figure 40 The action. For example, subroutine 4000 may be modified in some implementations to generate a floating-point array at 4020, omitting a single component dimension of the radius of the normal data set. In another instance, other or additional normal cluster configurations may be defined and modeled. For example, a first tube set of normal patient cells may be subjected to a first scheme that produces a first normal patient data set (which may correspond to, for example, ten clusters) with parameters FSC, SSC, CD20 (FITC), CD10 (PE), CD45, and CD19, and a second tube set of normal patient cells may be subjected to a second scheme that produces a second normal data set (which may correspond to, for example, four clusters) with parameters FSC, SSC, CD22 (FITC), CD34 (PE), CD45, and CD19, and so on. The selected scheme may be applied to the cell set of the test patient to generate the test patient data set.
[0320] Notice, Figure 18A The implementation of subroutine 500 of -C can be modified to... Figure 18C 536 actions were adopted Figure 40 The subroutine 4000 is used to define the radius of the defined normal cluster group.
[0321] Figure 40A This is a flowchart of example cluster subroutine 4000a. Subroutine 4000a can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104) employs a cluster test to examine cells or data points in a cell group. The subroutine begins at 4002a and proceeds to 4004a. At 4004a, reference points for a reference population are identified, such as clusters in a normal cell group. Example reference points are shown in Table 4 below. Reference points can be, for example, cluster centroids or points selected after reviewing the representation of the data. The identification of reference points is discussed in more detail elsewhere in this document.
[0322] Reference point 1 1.2 1.1 0.5 3.0 1.8 2.0 Reference point 2 1.0 1.1 0.4 2.6 2.2 2.4 … … … … … … … Reference point n … … … … … …
[0323] Table 4
[0324] Subroutine 4000a proceeds from action 4004a to action 4006a. At 4006a, the distance from the cell to each reference point is determined and stored. These distances can be determined using the following equation 4:
[0325] Equation 4:
[0326] Where par is the first parameter (e.g., CD45), par.2 is the second parameter, and so on.
[0327] Table 5 shows example distances of cells stored in a floating-point array.
[0328] Cell 1 0.7 0.9 2.0 1
[0329] Table 5
[0330] Subroutine 4000a proceeds from 4006a to 4008a. At 4008a, subroutine 4000a identifies the minimum distance between a reference point and a cell associated with a reference population (e.g., a cluster). Subroutine 4000a proceeds from 4008a to 4010a. At 4010a, the index corresponding to the reference population associated with the reference point is appended to a floating-point array, which is the minimum distance from the cell (see, for example, ...). Figure 42 Subroutine 4000a proceeds from 4010a to 4012a, where it determines whether there are more cells to be processed. If more cells are determined to be processed at action 4012a, subroutine 4000a proceeds from 4012a to 4006a to process the next cell. If no more cells are determined to be processed at action 4012a, subroutine 4000a proceeds from 4012a to 4014a, where it terminates.
[0331] Some implementation schemes for System 100 are available. Figure 40A Other actions not shown may be omitted. Figure 40A All the actions shown in the diagram may be performed in different orders. Figure 40A The actions of the subroutine. For example, subroutine 4000a may be modified in some implementations to use data storage formats other than floating-point arrays.
[0332] Identify and test patients for cells outside the normalized radius (percentage, etc.).
[0333] In one implementation, an image comparing the test cell group to a defined normal cluster group is generated, which helps to quickly and visually communicate the presence of abnormal cell populations in the test cell group using the diagnostic methods disclosed herein. The generated image is referred to herein as a summary plot and may include image pixels. The summary plot graphically summarizes the cell frequency and location information of cells in each cluster / maturation stage. Currently, visual assessment of potential abnormalities is limited to the analysis of a series of dot plots, which can only be understood by flow cytometry experts. The implementation of the generated summary plot image facilitates interpretation by non-flow cytometry experts and quickly communicates that the cell maturation stage is abnormal, such as, for example, regarding the following... Figures 46-57 The above is discussed. In one implementation, potential abnormalities can be further investigated by subtracting cells within the normal radius from the centroid line, such as regarding the following. Figures 58-65 This can facilitate the rapid and specific identification of cells with a significantly different centroid from the normal centroid.
[0334] Figure 46 This is a flowchart of example subroutine 5000, which can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104) employs a normal range defining how far each cluster / maturation stage of normal cells is from the centroid line, and a normal cell frequency decomposition defining the proportion of cells in each cluster / maturation stage. These definitions can then be used as a reference, where the cell set of the test patient is compared to the definitions of the normal cluster / maturation stage group. For convenience, the reference will be... Figures 46-51 and Figure 1 The diagnostic system 104 discusses subroutine 5000.
[0335] Subroutine 5000 begins at 4600 and proceeds to 4602. At 4602, the normal radius is obtained or determined by the diagnostic system 104. For example, a stored normal radius can be retrieved, or the normal radius can be determined, for example, as described above with reference to subroutine 4000 and... Figures 40-45 The normal radius of the cluster of normal cell lineages is determined, and this information can be stored for future use. Additionally, cells in the normal data group are clustered to a reference point (see, for example, action 1006 of subroutine 1000). Figure 24 A, Figure 40A Subroutine 5000 proceeds from 4602 to 4604.
[0336] At position 4604, the tangential intersection point between the cells of the normal patient data group and the centroid line was identified on the centroid line. The identification of the tangential intersection point on the centroid line can be accomplished using dot product. See above for more information. Figure 40 and 41 The discussion focuses on determining the tangential intersection points. As mentioned above, the centroid line can be segmented, and candidate segments can be used to identify the tangential intersection points of cells, which reduces the number of computations required.
[0337] Subroutine 5000 proceeds from 4604 to 4606. At 4606, diagnostic system 104 calculates the one-dimensional distance between the identified tangential intersection points of the cell and the cell on the centroid line. This can be done, for example, using Equation 2, which is listed above and repeated below for convenience:
[0338] Parameter distance = parameter 细胞 -parameter TCP [Equation 2]
[0339] Where parameters 细胞 These are the values of the corresponding parameters of the cell, and the parameters TCP These are the parameter values of the tangential intersection points of cells along the centroid line. The parameter distance values of cells can be stored in a floating-point array (see, for example,...). Figure 47 (Single-dimensional distance in the context). Subroutine 5000 proceeds from 4606 to 4608.
[0340] At 4608, the diagnostic system 104 performs a z-score transformation on a floating-point array on the normal patient data set (e.g., the data set generated at 4006) based on the radius definition retrieved or determined at 4602.
[0341] Subroutine 5000 proceeds from 4608 to 4610. At 4610, the diagnostic system 104 determines the multidimensional Euclidean distance between the scaled position (4608) of each cell and the centroid intersection point. This can be accomplished, for example, using Equation 5 listed below:
[0342] Euclidean distance = SQRT(a 2 +b 2 +c 2 +d 2 +e 2 +f 2 Equation 5
[0343] Where af is the normalized distance from the cell to the centroid line determined at 4608. A normalized floating-point array can be generated, storing Euclidean distances for normal patient data sets to indicate the z-transform distribution. The subroutine proceeds from 4610 to 4612.
[0344] At 4612, the diagnostic system 104 determines whether there are additional cells to be processed in the normal data set. When it is determined at 4612 that there are additional cells to be processed in the normal data set, subroutine 5000 returns to 4604 to process the next cell. Figure 47 An example floating-point array of normal patients is shown, which includes columns for Euclidean distance and cluster number. The subroutine proceeds to 4614 when no additional cells to be processed are identified at 4612.
[0345] At 4614, the diagnostic system 104 determines whether another set of normal patient data to be processed exists. When it is determined at 4614 that another set of normal patient data to be processed exists, subroutine 5000 returns to 4604 to process the cell for the next set of normal patient data. A single floating-point array can store cells for a single set of normal patient data (see [link]). Figure 47 When no other group of normal patient data to be processed is identified at 4614, subroutine 5000 proceeds to 4616.
[0346] At position 4616, the average Euclidean distance from the centroid line is calculated for each cluster / maturity stage of each normal patient. For example, refer to... Figure 47 , about Figure 47 The average Euclidean distance column determines the average distance for each cluster of patients. The calculated average can be stored in the distance matrix. Figure 48An example distance matrix is shown, where each row corresponds to a patient in the normal patient group of the study, each column corresponds to a cluster, and each value corresponds to the mean Euclidean distance of cells in that cluster from the centroid line. Other mean distances may be determined, either in place of or in addition to the mean Euclidean distance from the centroid line. For example, the mean distance from the centroid line to each cluster of each patient may be determined relative to any combination of other individual parameters (such as CD10(PE)) or parameters of the floating-point array (e.g., FSC, SSC, CD20(FITC), CD45, CD34, etc.).
[0347] Subroutine 5000 proceeds from 4616 to 4618. At 4618, the percentage of cells belonging to each cluster for each normal patient is calculated. For example, the percentage of patient cells in a cluster of normal patients can be determined by dividing the number of patient cells in the patient data set that are in the cluster by the total number of patient cells in the patient data set, and then multiplying the result by 100. Some implementations may determine the ratio of cells in the patient data set to the total number of cells in the patient data set in other ways, such as determining a ratio rather than a percentage. These percentages can be stored in a frequency matrix, such as using data from... Figure 49 The data from the study shows that, theoretically, Figure 49 Each row should add up to 100%. However, in some implementations, rounding values can be used, which may introduce small rounding errors. Subroutine 5000 proceeds from 4618 to 4620.
[0348] At 4620, the diagnostic system 104 determines whether there is another group of normal patient data to be processed. If an additional group of normal patient data to be processed is determined at 4620, subroutine 5000 returns to 4616 to process the next patient data group in the normal patient data group. If no additional group of normal patient data to be processed is determined at 4620, subroutine 5000 proceeds to 4622.
[0349] At 4622, the diagnostic system 104 determines the mean and standard deviation of the mean Euclidean distance determined at 4616 for each cluster. In other words, the mean Euclidean distance for each cluster starts from the centroid line, and the variation of this mean Euclidean distance is determined for the normal patient group. This can be achieved, for example, by referring to... Figure 48 This is accomplished by determining the mean and standard deviation of each column of the stored distance matrix. These results can be stored for later use. Figure 50 An example normal position matrix showing the determined mean and standard deviation for each cluster in a storage study is shown. Some implementations may determine and store additional or different mean and standard deviation information. For example, if the average Euclidean distance relative to another parameter (column) is determined at 4616, the mean and standard deviation relative to that parameter can be determined and stored.
[0350] Subroutine 5000 proceeds from 4622 to 4624. At 4624, diagnostic system 104 determines the average cell frequency for each cluster. This can be done by averaging the columns, for example, by referring to... Figure 49 The frequency matrix. These results can be stored for later use, for example, to compare the cells of a test patient with a defined normal patient dataset. Figure 51 An example normal percentage matrix used in the study is shown, which stores a defined average or mean percentage of cells in each cluster. Subroutine 5000 proceeds from 4624 to 4626, where subroutine 5000 ends.
[0351] Some implementation schemes for System 100 are available. Figure 46 Other actions not shown may be omitted. Figure 46 All the actions shown in the diagram may be performed in different orders. Figure 46 The actions can be described as follows. For example, subroutine 5000 may be modified in some implementations to determine additional and / or different averages at 4616, and the data may be stored in data structures other than floating-point arrays and matrices. In another instance, other or additional normal cluster configurations can be defined and modeled. For example, a first tube group of normal patient cells may be subjected to a first scheme that produces a first normal patient data set corresponding to ten clusters with parameters FSC, SSC, CD20 (FITC), CD10 (PE), CD45, and CD19; a second tube group of normal patient cells may be subjected to a second scheme that produces a second normal data set corresponding to four clusters with parameters FSC, SSC, CD22 (FITC), CD34 (PE), CD45, and CD19, and so on. The selected scheme can be applied to the cell groups of test patients to generate test patient data sets to be compared with the defined cluster groups.
[0352] Figure 52 This is a flowchart of example subroutine 6000, which can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104) employs a method to compare a set of test patient data with a defined normal cluster group, which is generated using any of the methods described above or various combinations of the disclosed methods to facilitate rapid and efficient communication of potential maturity stage abnormalities that may exist in the test patients. For convenience, reference will be made to... Figures 52-57 and Figure 1 The diagnostic system 104 discusses subroutine 6000.
[0353] Subroutine 6000 begins at 5200 and proceeds to 5202. At 5202, the diagnostic system 104 retrieves or determines the normal radius for the normal data group. For example, the normal radius of the cell group within the normal data group can be determined, for example, as described above with reference to subroutine 4000 and... Figures 40-45 The procedure does not require "repetition" for additional patient data sets (e.g., at actions 4010 and 4018, it will be uncertain whether additional patient data sets need to be processed), nor does it require merging normalized floating-point arrays, etc. Subroutine 6000 proceeds from 5202 to 5204.
[0354] At position 5204, the tangential intersection points between the cells of the test patient data group and the centroid line defined for the normal patient data group are identified. The identification of tangential intersection points on the centroid line can be accomplished using dot product. See above for more information. Figure 40 and 41 The discussion focuses on determining the tangential intersection points. As mentioned above, the centroid line can be segmented, and candidate segments can be used to identify the tangential intersection points of cells, which reduces the number of computations required.
[0355] Subroutine 6000 proceeds from 5204 to 5206. At 5206, diagnostic system 104 calculates the one-dimensional distance between the identified tangential intersection points of the cell and the cell on the centroid line. This can be done, for example, using Equation 2 listed above and repeated below for convenience:
[0356] Parameter distance = parameter 细胞 -parameter TCP [Equation 2]
[0357] Where parameters 细胞 These are the values of the corresponding parameters of the cell, and the parameters TCP These are the parameter values of the tangential intersection points of cells along the centroid line. The parameter distance values of cells can be stored in a floating-point array (see [link to documentation]). Figure 53 Subroutine 6000 proceeds from 5206 to 5208.
[0358] At 5208, the diagnostic system 104 performs a z-score transformation on the distance data of the test patient dataset (e.g., a floating-point array of test patient datasets generated at 5206) based on the radius definition retrieved or determined at 5202.
[0359] Subroutine 6000 proceeds from 5208 to 5210. At 5210, the diagnostic system 104 determines the multidimensional Euclidean distance between the scaled position (5208) of each cell and the centroid intersection point. This can be accomplished, for example, using Equation 6 listed below:
[0360] Euclidean distance = SQRT(a 2 +b2 +c 2 +d 2 +e 2 +f 2 Equation 6
[0361] Where af is the normalized distance from the cell to the centroid line, determined at 5208. A normalized floating-point array can be generated for the test patient dataset indicating the z-transform distribution. Subroutine 6000 proceeds from 5210 to 5212.
[0362] At 5212, the diagnostic system 104 determines whether there are additional cells to be processed in the test patient data set. When it is determined at 5212 that there are additional cells of interest to be processed in the test patient data set, subroutine 6000 returns to 5204 to process the next cell. If it is not determined at 5212 that there are additional cells of interest to be processed in the test patient data set, the subroutine proceeds to 5216.
[0363] At position 5216, the average Euclidean distance from the centroid line is calculated for each cluster / maturity stage of the test patients. For example, refer to... Figure 53 , about Figure 53 The average Euclidean distance column determines the average distance for each cluster of patients. The calculated average can be stored in the test patient distance matrix. Figure 54 An example test patient distance matrix from the study is shown, where each column corresponds to a cluster. Other average distances can be determined, either in place of or in addition to the average Euclidean distance derived from the centroid line. For example, the average distance for each cluster of the test patient can be determined relative to any of other parameters (such as CD10(PE)) or other parameters of the floating-point array (e.g., FSC, SSC, CD20(FITC), CD45, CD34, etc.).
[0364] Subroutine 6000 proceeds from 5216 to 5218. At 5218, the percentage of cells belonging to each cluster of test patients is calculated. For example, the percentage of patient cells in a cluster of test patients can be determined by dividing the number of patient cells in the patient data set that are in the cluster by the total number of patient cells in the patient data set, and then multiplying the result by 100. Some implementations may determine the ratio of cells in the patient data set to the total number of cells in the patient data set in other ways, such as by determining a ratio. The percentage can be stored in a test patient frequency matrix, as in this study. Figure 55 As shown. Theoretically, Figure 55 The percentages should add up to 100%. However, in some implementations, rounding values can be used, which may introduce small rounding errors. Subroutine 6000 proceeds from 5218 to 5220.
[0365] At 5220, the diagnostic system 104 generates an image representing the differences between the test patient data group and a defined normal cluster group. For example, the diagnostic system 104 may generate the pixels of the image. This image may display, for example, a comparison of cellular characteristics in the test patient group and the normal data group, such as the frequency of cells at each maturation stage and the average distance of cells from the centroid line.
[0366] Figure 56 Example image 5600 of a first tube set of normal patient cells undergoing a first protocol in a study that generated a first normal patient data set corresponding to ten clusters with parameters FSC, SSC, CD20 (FITC), CD10 (PE), CD45 (PerCP), and CD19 (APC). Figure 56 In the diagram, immature clusters are on the left and mature clusters are on the right.
[0367] like Figure 56 As shown, image 5600 includes a corresponding indication of the range 5602 of the normal distance from each defined normal cluster to the centroid line. For clarity, reference numeral 5602 is used only for identification. Figure 56 One of the range indicators in the range indicator (the range indicator for the third cluster from the left). For example, the average location of each cluster (e.g., Figure 50 The average position row of the normal position matrix is used to determine the corresponding range for each cluster. As shown, the lower limit of each range is 0, and the upper limit of each range is the average position of the cluster plus the standard deviation of the average position of the cluster (e.g., ...). Figure 50 The standard deviation of the normal position matrix (position row) is twice that of the normal position matrix. For example, referring to the third cluster from the left, the mean position of the clusters in the normal position matrix is 3.04, and the standard deviation of the position is 0.75. Therefore, for example, for the third cluster, the indication of range 5602 is from 0 to 4.54. For example, the indication of range 5602 represents a statistical range where 97.5% of the normal centers of each defined normal cluster lie on the centroid line. Indicators other than lines may be used.
[0368] Image 5600 includes individual indicators 5604 of the expected frequencies of multiple cells in each defined normal cluster. This is in Figure 56 The middle is based on the average frequency of cells in each defined normal cluster (e.g., from Figure 51 The normal percentage matrix is represented by row 1) with the black outline of the scaled circle. For clarity, reference numeral 5604 is used only for identification. Figure 56One of the expected frequency indicators (the expected frequency indicator for the seventh cluster from the left). Clusters with higher average frequencies are proportionally larger than clusters with lower average frequencies. For example, the circle indicating the expected frequency corresponding to cluster 7 is smaller than the circle indicating the average frequency of cluster 3, indicating that cluster 7 has a lower expected cell frequency than cluster 3. Shapes other than circles may be used.
[0369] Image 5600 includes various indicators 5606 of the distance of cells in the test patient dataset from the centroid line and the frequency of cells in clusters within the test patient dataset. As shown, colored circles summarize the cells in each cluster of the test patients. The size of the circle corresponds to the test frequency (from the test frequency matrix) of that particular cluster. For clarity, reference numeral 5606 is used only for identification. Figure 56 The distance and frequency indicators for cells in the test patient data group are shown in the diagram (the distance and frequency indicators for cells in the second cluster from the left in the test patient data group). If the colored circle indicating 5606 is larger than the black outline of the corresponding defined normal cluster indicating 5604, this indicates that the number of cells in the test patient cluster exceeds the average frequency of that cluster in the normal data group. If the colored circle indicating 5606 is smaller than the black outline of the corresponding defined normal cluster indicating 5604, this indicates that the number of cells in the test patient cluster is less than the average frequency of that cluster in the normal data group.
[0370] The size of the black circle 5604 can vary depending on certain patient characteristics, such as age. For example, pediatric patients have more immature cells than older patients. Therefore, the size of the black circle may differ for different patient groups. The location of the circle will generally be the same for different groups.
[0371] The position of circle 5606 corresponds to the average distance of cells in the cluster of the tested patient from the centroid line. If the circle falls outside the indicated range 5602, it indicates that the cells in the cluster of the tested patient are located further away from the average centroid line, which may indicate an underlying abnormality (cancer). Note that in Figure 56 In this context, the position of indicator 5604, representing the expected frequency of multiple cells within each defined normal cluster, coincides with the position of indicator 5606, representing the distance from the centroid line in the test patient data set. In other words, the position of indicator 5604 does not indicate the distance or range of distance between the defined normal cluster and the centroid line. Consistent positioning of indicator 5604 with indicator 5606 facilitates comparison of the expected frequencies of multiple cells within the defined normal cluster with the cell frequencies of clusters in the test patient data set.
[0372] Figure 57Another example image 5700 of a second tube set of normal patient cells undergoing a second protocol is shown, which generates a second normal data set corresponding to four clusters with parameters FSC, SSC, CD22 (FITC), CD34 (PE), CD45 (PerCP), and CD19 (APC). Figure 57 In the diagram, immature clusters are on the left and mature clusters are on the right.
[0373] like Figure 57 As shown, image 5700 includes an indication 5702 of the range 5702 of the normal distance 5702 from the centroid line to each defined normal cluster. For example, this range can be determined using the mean position of each cluster plus or minus twice the standard deviation of the cluster mean position. Image 5700 includes an indication 5704 of the expected frequency of multiple cells within each defined normal cluster. This is in Figure 57 The black outline of a circle is scaled according to the average frequency of cells in each defined normal cluster. Clusters with higher average frequencies are proportionally larger than clusters with lower average frequencies. Shapes other than circles can be used.
[0374] Image 5700 includes indicators 5706 showing the distance of cells in the test patient data set from the centroid line and the frequency of cells in clusters within the test patient data set. As shown, colored circles summarize the cells in each cluster of the test patients. The size of the circle corresponds to the test frequency (from the test frequency matrix) of that particular cluster. If the colored circle of indicator 5607 is larger than the black outline of indicator 5704 corresponding to the defined normal cluster, this indicates that the number of cells in the test patient cluster exceeds the average frequency of that cluster in the normal data set.
[0375] The size of the black circle 5704 can vary depending on certain patient characteristics, such as age. For example, pediatric patients have more immature cells than older patients. Therefore, the size of the black circle may differ for different patient groups. The location of the circle will generally be the same for different groups.
[0376] The position of circle 5706 corresponds to the average distance of cells in the cluster of cells tested from the centroid line. If the circle falls outside the indicated range 5702, it indicates that the cells in the cluster of cells tested are located further away from the average centroid line, which may indicate a potential abnormality (cancer). Interpretation should typically be made by a medical expert (e.g., a physician), which may include consideration of other information about the patient. Note that in Figure 57In this context, the position of indicator 5704, representing the expected frequency of multiple cells in each defined normal cluster, coincides with the position of indicator 5706, representing the distance from the centroid line in the test patient data set. In other words, the position of indicator 5704 does not indicate the distance or range of distance between the defined normal cluster and the centroid line. Consistent positioning of indicator 5704 and indicator 5706 facilitates the comparison of the expected frequencies of multiple cells in the defined normal cluster with the cell frequencies of the clusters in the test patient data set. Subroutine 6000 proceeds from 5220 to 5222, where it terminates.
[0377] Some implementation schemes for System 100 are available. Figure 52 Other actions not shown may be omitted. Figure 52 All the actions shown in the diagram may be performed in different orders. Figure 52 The subroutine 6000 can be modified in some implementations to determine additional and / or different average values at 5216, and the data can be stored in data structures other than floating-point arrays and matrices. In another instance, other or additional normal cluster configurations can be defined and modeled. In yet another instance, the subroutine 6000 can be modified to store, display, or print the generated images.
[0378] Figure 58 This is a flowchart of example subroutine 7000, which can be generated by a diagnostic system (such as...). Figure 1 The diagnostic system 104) employs a method to compare a set of test patient data with a defined normal cluster group, which is generated using any of the methods described above or various combinations of the disclosed methods to facilitate rapid and efficient communication of potential maturity stage abnormalities that may exist in the test patients. For convenience, reference will be made to... Figures 58-65 and Figure 1 The diagnostic system 104 discusses subroutine 7000.
[0379] Subroutine 7000 begins at 5800 and proceeds to 5802. At 5802, the diagnostic system 104 retrieves or determines the normal radius for the cell group of the normal patient group. Subroutine 7000 proceeds from 5802 to 5804. At 5804, the cell clusters of the test patient data group are aligned to the corresponding reference point, and the tangential intersections between the cells of the test patient data group and the centroid line defined for the normal patient data group are identified. The test patient data group can be, for example, a lineage or reference population data group identified using SVM. The identification of the tangential intersections on the centroid line can be accomplished using dot product. See above regarding... Figure 40 , 41The discussion of determining tangential intersections is similar to that in section 41A. As mentioned above, the centroid line can be segmented, and candidate segments can be used to identify tangential intersections of cells, which reduces the number of computations required.
[0380] Subroutine 7000 proceeds from 5804 to 5806. At 5806, diagnostic system 104 calculates the one-dimensional distance between the identified tangential intersection points of the cell and the cell on the centroid line. This can be done, for example, using Equation 2, which is listed above and repeated below for convenience:
[0381] Parameter distance = parameter 细胞 -parameter TCP [Equation 2]
[0382] Where parameters 细胞 These are the values of the corresponding parameters of the cell, and the parameters TCP These are the parameter values of the tangential intersection points of cells along the centroid line. The parameter distance values of cells can be stored in a floating-point array (see [link to documentation]). Figure 59 Subroutine 7000 proceeds from 5806 to 5808.
[0383] At 5808, the diagnostic system 104 performs a z-score transformation on the distance data of the test patient data set (e.g., a floating-point array of test patient data sets generated at 5806) based on the radius definition retrieved or determined at 5802.
[0384] Subroutine 7000 proceeds from 5808 to 5810. At 5810, the diagnostic system 104 determines the multidimensional Euclidean distance between the scaled position (5808) of each cell and the centroid intersection point. This can be accomplished, for example, using Equation 7 listed below:
[0385] Euclidean distance = SQRT(a 2 +b 2 +c 2 +d 2 +e 2 +f 2 Equation 7
[0386] Where af is the normalized distance from the cell to the centroid line, determined at 5808. A normalized floating-point array can be generated for a normal patient dataset indicating a z-transform distribution. Subroutine 7000 proceeds from 5810 to 5812.
[0387] At point 5812, the diagnostic system 104 determines whether there are additional cells to be processed in the test patient data set. If an additional cell of interest to be processed is determined to exist in the test patient data set at point 5812, subroutine 7000 returns to point 5804 to process the next cell. If no additional cell of interest to be processed is determined to exist in the test patient data set at point 5812, the subroutine proceeds to point 5814.
[0388] At 5814, select the radius parameter / component to use as the filtering criterion for subtraction. This can be done based on default parameters, user selection, etc. For example, if the user wants to identify cells that differ from normal cells while considering all parameter combinations, they can choose Euclidean distance. In another instance, if the user wants to identify only those cells with different CD10 expression than normal cells, they can choose CD10. Figure 45 A table of normal radii is shown, which can be used, for example, with the reference above. Figures 40-45 The discussion uses subroutine 4000 to generate it. For example... Figures 60-63 As shown, in one implementation of this study, Euclidean distance is chosen as the basis for filtering. Figures 60A-63A As shown, in one implementation of this study, CD10 is chosen as the basis for filtering. When the selected filtering implementation is a single parameter, the absolute value of the test parameter can be compared with the subtraction vector to determine whether the cell should be subtracted (e.g., Figure 63A Cell 3). Subroutine 7000 proceeds from 5814 to 5816.
[0389] At point 5816, a multiplication factor is selected. This can be done, for example, based on a default multiplication factor (which can vary based on the selected filter parameters, clusters, etc.), or based on user selection. A lookup table can be used. Subroutine 7000 proceeds from 5816 to 5818. At 5818, a subtraction vector is generated, for example, by multiplying the standard deviation of the selected radius of each cluster by the selected cluster multiplication factor and adding it to the average of the selected radii of the cluster. Figure 61 and 61A Example radius tables are shown, storing the selected radius (from 5814), the selected multiplication factor (from 5816), and the subtraction vector (from 5818). Subroutine 7000 proceeds from 5818 to 5820.
[0390] At 5820, such as Figure 62 and 62A As shown, the subtraction vector for each cluster is appended to... Figure 59The floating-point array. Subroutine 7000 proceeds from 5820 to 5822. At 5822, subroutine 7000 determines the cells to be included in the representation (e.g., images) of the test patient data set. This can be done, for example, by comparing the variable of interest selected as a filtering criterion (see 5814) for each cell with the subtraction vector. Cells less than or equal to the subtraction vector can be marked as not to be included in the representation of the test patient data set, to be subtracted, not to be displayed, semi-transparent, etc. This data can be stored in a floating-point array, such as... Figure 63 As shown, it demonstrates that the Euclidean distance is used as a filtering criterion, which is compared with the subtraction vector to determine whether cells should be excluded from the representation of the test patient data group.
[0391] Subroutine 7000 proceeds from 5822 to 5824. At 5824, the subroutine generates one or more representations (images) of the test patient data set. This can be accomplished, for example, using one or more floating-point arrays generated at 5822 to generate a pixel display representing the test patient data set.
[0392] Figure 64 and 65 An example of a user interface for controlling the generation and display of representations of test patient data sets is shown. As illustrated, Figure 64 The user interface includes user-selectable controls and / or data input fields 6402, 6404, and 6406, as well as representations of test patient data sets 6410 and 6412. The first user-selectable control 6402 allows the user to select a radius parameter, on which the test patient data set is filtered (see [reference]). Figure 58 (5814), and as shown in the figure is the drop-down menu selector 6402. As shown in the figure, the drop-down menu selector displays "Euclidean distance" as a selection, which corresponds to selecting Euclidean distance, i.e., the six-dimensional radius in this study. Other filter options can be selected to help represent cells beyond a single parameter (e.g., CD10, SSC, etc.). The second user-selectable control 6404 allows the user to set the subtraction vector (see 5814). Figure 58 (5818), and a slider is shown in the figure. Some implementations allow the user to select the multiplication factor to be used to define the subtraction vector. Data input field 6406 allows the user to input multiple events to be displayed. The number of events to be displayed can be used to indicate the number of cells in the test patient data set to be processed when generating the display (e.g., the first 5000 cells, the first 5000 cells before subtraction, the 5000 cells in the cluster (mature stage), etc.). More cells may need to be visualized to detect low levels of leukemia. In one implementation, a default number of cells can be selected, a maximum number of cells can be set, etc. Figure 64 The diagram shows a display of the test patient data set when the subtraction vector magnitude was set to zero. Figure 65 The display shows the test patient data set when the subtraction vector magnitude is set to 2. Figures 64A to 65B The diagram illustrates an example user interface and display representation when the interface includes controls for selecting plotting groups. A plotting group might simply be a cell lineage to be visualized. For example, a user could select to visualize neutrophils, monocytes, lymphocytes, etc., or any combination thereof, to facilitate, for example, visualization of any cell that differs from the normal maturation pattern.
[0393] In one implementation, the pixels representing subtracted cells may be semi-transparent, and the pixels representing non-subtracted cells may be colored based on the cluster to which the cell is assigned. In some implementations, cells of other lineages or other reference populations may also be displayed (e.g., plasma cells may be displayed alongside B lymphocytes), and controls to facilitate the selection of such lineages may be provided. As shown in the figure, Figure 64 and 65 The display includes pixel display 6410 representing a test patient data group, with CD10 on the vertical axis and CD20 on the horizontal axis; and pixel display 6412 representing a test patient data group, with CD45 on the vertical axis and SSC on the horizontal axis. The centroid line 6414 defining the normal cluster group is represented in a display with black pixels. In one embodiment, subtracted cells can be represented using transparent pixels, and non-subtracted cells can be represented using pixel color based on the cluster to which the cell is assigned. In one embodiment, subtracted cells may not be included in the display. Figure 64 In this context, the subtraction vector is set to zero. Therefore, the subtraction vector is zero, and pixel displays 6410 and 6412 include colored pixels representing all cells in the test patient data set (up to any limitations set via user-selectable control 6406). Figure 65 In this representation, the subtraction vector is set to 2. Therefore, the subtraction vector is not zero, and pixels 6410 and 6412 include colored pixels representing cells from the test patient data group that are more than 2 standard deviations from the centroid line 6414 in the normalized six-dimensional Euclidean space, while pixels representing cells from the test patient are translucent and are no more than 2 standard deviations from the centroid line 6414 (or are not included in the representation). Figure 64A -B represents the pixel display depicting cells located outside the radius of the neutrophil centroid line. Figure 65A -B is the pixel display depicting cells outside the centroid radius of neutrophils, erythrocytes, monocytes, and dendritic cells.
[0394] Subroutine 7000 proceeds from 5824 to 5826, where it terminates. Some implementation schemes of System 100 are available. Figure 58 Other actions not shown may be omitted. Figure 58All the actions shown in the diagram may be performed in different orders. Figure 58 The subroutine 7000 can be modified in some implementations to combine actions 5822 and 5824, instead of separately determining the cells to be included in the representation of the test patient data set and generating the representation of the test patient data set. In another instance, the subroutine may include a loop to facilitate the dynamic generation of the representation, for example, by adjusting the multiplication factor, filtering parameters, or the number of axis parameters dynamically represented. In another instance, the implementation may be modified to generate a floating-point array excluding subtracted cells, which can facilitate remote display of the representation by reducing the amount of data used to generate the display. In yet another instance, subroutine 7000 may be modified to store or print the generated image.
[0395] Implementations using this method simplify the complexity of analyzing, for example, six-dimensional or more dimensional data and relate it to the statistical analysis of normal cells. Furthermore, the data can be divided into different lineages, further distinguishing between maturation stages and lineages. This helps physicians familiar with the concept of cell maturation from progenitor cells to provide a more intuitive explanation of how blood cells mature through various developmental stages. Importantly, it is possible to distinguish regenerating bone marrow from an increase in the number of immature normal cells (referred to as leftward shift), which differs from the presence of abnormal cells. This demonstration combines the concept of a single lineage, the maturity of the lineage, and the frequency of test cells referencing the expected frequency of each developmental stage, as well as whether these cell populations fall within the expected statistical position range in N-dimensional space. This information can be combined with patient knowledge, the patient's clinical data (such as complete blood counts, cytogenetics, and treatment history). Therefore, the data can be interpreted in medical practice. This illustration simplifies the analysis for physicians, rather than requiring them to think in six-dimensional space. This can be facilitated by subtracting events close to the normal centroid. The human eye is very good at identifying clusters of events and distinguishing them from random events that slightly exceed expected boundaries.
[0396] The results of statistical characterization studies on the bright reference populations of lymphocytes, promyelocytes, monocytes, and CD34 confirmed that the location of the reference populations could be identified, remaining constant in six-dimensional space even after chemotherapy, and thus could be used to define the variability of normal cells in stressed bone marrow samples. Therefore, using support vector machines to identify the reference populations, as discussed in this paper, can provide a basis for determining differences from normal cells.
[0397] In one study, the internal reference population considered was based on a discrete cell cluster definition, including mature lymphocytes, amorphous progenitor cells, promyelocytes, mature neutrophils, and mature monocytes. This study is discussed in more detail below. The representation of the reference cell population is as follows: Figure 66 It appeared in the middle.
[0398] The data set consisted of 77 randomly selected, phenotypically normal pediatric AML patients who completed induction (day 28) and participated in AAML1031. Three years of patient data were collected using three flow cytometers. Multiple batches of reagents for each antigen were investigated. Changes in surface gene expression levels during HSC maturation to mature hematopoietic cells in each lineage were investigated to characterize and compare the data with a reference population in a multidimensional space.
[0399] The first reference population in the study was lymphocytes. Lymphocytes were present in every sample and used as a reference for CD45, SSC, and FSC. Lymphocytes were negative for CD34. CD45 intensity was used to demonstrate an adequate antibody-to-cell ratio. Using the above references... Figure 27 The implementation scheme of the discussed method involves training a support vector machine on a manually selected lymphocyte population of 27 normal patients (with 8 tubes = 216 normal data sets). Using the above reference... Figure 28 The implementation of the discussed method applies defined multidimensional boundaries to normal patients and 50 test patients (8 tubes = 400 test patient data set) to predict lymphocyte populations. Manual gating references were not applied. Figure 30 The lymphocyte population of a normal patient was shown, and Figure 35 The lymphocyte populations of the test patients are shown. The above reference values were used for each of the 27 healthy patients and for each of the 50 test patients. Figure 29 The implementation scheme of the discussed method determines the mean and standard deviation of antigen intensity for each marker in each tube. For normal patient data, the mean of the mean is determined. The concordance of antigen intensity among tested patients is compared with the determined mean of the normal mean. The results of 50 test patient data groups are represented in... Figure 67 and 68 It reappears in the middle.
[0400] The second population studied was promyelocytes, which are the most mature bone marrow cells to be observed in AML. This population was identified as HLA-DR negative, CD11b negative, and high SSC. The predicted population of premyelocytes from patients was tested in... Figure 69 The comparison results are indicated in blue. Figure 70 and 71 The results are shown in the diagram. The results are stable for the instrument. Statistical analysis of the SSC can be used to determine whether promyelocytes exhibit granular changes. For example, granular changes in promyelocytes, as measured by the SSC, have been observed in stressed bone marrow and in patients with myelodysplastic syndromes (MDS). Hypothalamic promyelocytes can be identified in MDS.
[0401] The third population studied was monocytes. This population was identified as CD14-positive, CD33-positive, with high levels of CD45 and intermediate SSCs. The predicted population of monocytes in... Figure 72 The middle is represented by green, followed by Figure 73 and 74 The comparison results are shown. Black dots (overlapping with green dots) identified as non-monocytes in the high CD14 region are non-living cells / bimodal.
[0402] The fourth population in the study was amorphous progenitor cells, identified as bright CD34. A combination of CD33 and CD34 was selected to identify this cell population. The predicted population of amorphous progenitor cells was... Figure 75 The diagram shows: red indicates consistency between SVM and expert assessment of the predicted amorphous progenitor cell population; purple indicates cells predicted by SVM to be amorphous progenitor cells, which experts did not predict; and blue indicates cells predicted by experts to be amorphous progenitor cells, but which SVM did not predict. The results of using SVM prediction (red and purple in the diagram above) to determine the mean and standard deviation of the test patient data groups are presented in... Figure 76 and 77 It is generated in the middle.
[0403] exist Figure 78 The results were compared between the reference population and CD45 and SSC.
[0404] exist Figure 79 The representation of lymphocyte populations moving to a fixed point is reproduced (e.g., using the above reference). Figure 29 Subroutine 3200 discusses the vector normalization process. The amount of movement for the CD34 bright population, monocytes, and promyelocytes is the same as the amount used to move the lymphocyte population. When the position of the lymphocyte population moves to a fixed point, the positions of the data points for the other reference populations tighten, and the CD34 bright population, promyelocyte population, and lymphocyte population appear to move up and down together. This indicates that variability is reflected within individuals, rather than between reference populations as identified by the parameters.
[0405] Further studies of other cell surface antigens on these reference populations showed that the intensity of surface gene products (CD) on undifferentiated progenitor cells and mature monocytes was also largely constant between individuals. Figure 80 It represents the CD34 intensity of CD34++. Figure 81 It represents the CD14 intensity of monocytes.
[0406] Interestingly, the CD33 intensity on mature monocytes was not constant between individuals. However, the ratio of CD33 intensity between monocytes and amorphous progenitor cells was essentially constant. Figure 82 It represents the CD33 intensity of CD14++ monocytes.
[0407] On the left Figure 83 In the representation, monocytes (green) and amorphous progenitor cells (red) exhibit uneven amounts of CD33 (and less variability in CD45). This variability was reduced by normalizing the data in a manner similar to that used in previous CD45 / SSC studies using mature lymphocytes as a reference population. Figure 83 In the right-hand portion, the position of amorphous progenitor cells shifts to a single location, accompanied by a shift in the position of monocytes. This normalization results in a more compact distribution of monocyte populations, indicating that the ratio of CD33 levels among these cell populations is preserved even if the absolute amounts vary from person to person. In this way, individual variability is reduced.
[0408] Some embodiments may take the form of a computer program product or include a computer program product. For example, according to one embodiment, a computer-readable medium is provided that includes a computer program adapted to perform one or more of the methods or functions described above. The medium may be a physical storage medium, such as a read-only memory (ROM) chip, or a disk (such as a DVD-ROM, CD-ROM, or hard disk), a memory, network, or portable medium that is read by a suitable drive or via a suitable connection (including being encoded with one or more barcodes or other relevant codes stored on one or more such computer-readable media and readable by a suitable reader device).
[0409] Furthermore, in some implementations, some or all of these methods and / or functions may be implemented or provided in other ways (e.g., at least in part in firmware and / or hardware), including but not limited to one or more application-specific integrated circuits (ASICs), digital signal processors, discrete circuits, logic gates, standard integrated circuits, controllers (e.g., by executing appropriate instructions and including microcontrollers and / or embedded controllers), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), state machines, and devices employing RFID technology, and various combinations thereof.
[0410] As those skilled in the art will recognize, the above methods can be used in a variety of settings, including but not limited to diagnostic and disease and treatment monitoring.
[0411] All of the above U.S. patents, U.S. patent applications, U.S. patent applications, foreign patents, foreign patent applications and non-patent publications mentioned in this specification and / or listed in the application data sheet are incorporated herein by reference in their entirety.
[0412] Based on the foregoing, it should be understood that while specific embodiments have been described herein for illustrative purposes, various modifications may be made without departing from the spirit and scope of this disclosure. The above embodiments are provided for illustrative purposes only and are not intended to be limiting.
Claims
1. A method for characterizing a biological cell test set in n-dimensional space, comprising: The first protocol involves exposing each cell in a normal biological cell group to multiple of four or more reagents; The second protocol was used to measure the corresponding multiple fluorescence intensities for each cell in the normal biological cell group; Based at least in part on multiple fluorescence intensities of cells measured in the normal biological cell group, each cell in the normal biological cell group is mapped to a corresponding point in n-dimensional space, where these corresponding points form a normal point group; Multiple reference populations are defined in the normal point group using a support vector machine, including a lymphocyte reference population and a monocyte reference population. The reference clusters in the n-dimensional space are defined by defining the centroid lines and radii based at least in part on a plurality of defined reference populations, wherein each cluster in the reference clusters corresponds to a level of maturity within a cell lineage; Using this first protocol, each cell in the biological cell test group is exposed to multiple reagents; Using this second protocol, multiple fluorescence intensities were measured for each cell in the biological cell test group. Based at least in part on multiple fluorescence intensities of cells measured in this biological cell test set, each cell in the biological cell test set is mapped to a corresponding point in n-dimensional space, where these corresponding points form a test point set; and The test point set is compared with a reference cluster set, wherein defining the reference cluster set and comparing the test point set with the defined reference cluster set includes defining one or more multidimensional boundaries in the n-dimensional space using a support vector machine, wherein the method includes: Determining the standard reference mean for each reference group, wherein determining the standard reference mean for the reference group includes: The support vector machine is trained to generate multidimensional boundaries in the n-dimensional space, thereby identifying the reference group of interest. For each normal patient in a group of normal patients, the generated multidimensional boundary is applied to identify a reference population of interest for the normal patients, and the average intensity of one or more parameters of the reference population is determined by summing the intensities of the populations of interest for a given parameter and dividing by the total number of cells in the populations of interest; and The standard reference mean of the reference population is calculated by determining the average of all mean reference strengths for each parameter of the reference population, wherein the standard reference mean is a vector; and The test point set is normalized, and the normalized test point set includes: The lymphocyte reference population of the test point group was identified by using the generated multidimensional boundary of the lymphocyte reference population of interest used to determine the standard reference mean. The average intensity of one or more parameters of the lymphocyte reference population for the test point group is determined by summing the intensities of points with a given parameter in the lymphocyte reference population and dividing by the total number of cells in the lymphocyte reference population for the test point group. A normalized vector for the lymphocyte reference population used in the test point group is generated by determining the difference between the mean parameter intensity of the lymphocyte reference population used in the test point group and the standard reference mean vector. The intensity of each point in the reference population of the test point group is normalized based on the normalized vector generated for the lymphocyte reference population of the test point group. Determine whether the normalized vector of the lymphocyte reference population used for the test point group is greater than a predetermined value; In response to determining that the normalized vector of the lymphocyte reference population used for the test point group is greater than the predetermined value, the test point group is marked as having a problem with the quality control of the measurement.
2. A method for characterizing a biological cell test set in n-dimensional space, comprising: The first scheme is used to map each cell in the normal biological cell group to a corresponding point in n-dimensional space, where these corresponding points form a normal point group; Based on the mapping of the normal point group in the n-dimensional space, the centroid line and radius are defined for the reference cluster group in the n-dimensional space, where the cluster corresponds to the maturity level within the cell lineage; Using this first scheme, each cell in the biological cell test group is mapped to a corresponding point in the n-dimensional space, and these corresponding points form a test point group; The test point set is compared with a reference cluster set, wherein defining the centroid and radius of the reference cluster set and comparing the test point set with the reference cluster set includes defining one or more multidimensional boundaries in the n-dimensional space using a support vector machine, wherein the method includes: Determine the standard reference mean for each of a plurality of reference populations, including a lymphocyte reference population and a monocyte reference population. Determining the standard reference mean for each reference population includes: The support vector machine is trained to generate multidimensional boundaries in the n-dimensional space, thereby identifying the reference group of interest. For each normal patient in a group of normal patients, the generated multidimensional boundary is applied to identify a reference population of interest for the normal patients, and the average intensity of one or more parameters of the reference population is determined by summing the intensities of the populations of interest for a given parameter and dividing by the total number of cells in the populations of interest; and The standard reference mean of the reference population is calculated by determining the average of all mean reference strengths for each parameter of the reference population, wherein the standard reference mean is a vector; and The test point set is normalized, and the normalized test point set includes: The lymphocyte reference population of the test point group was identified by using the generated multidimensional boundary of the lymphocyte reference population of interest used to determine the standard reference mean. The average intensity of one or more parameters of the lymphocyte reference population for the test point group is determined by summing the intensities of points with a given parameter in the lymphocyte reference population and dividing by the total number of cells in the lymphocyte reference population for the test point group. A normalized vector for the lymphocyte reference population used in the test point group is generated by determining the difference between the mean parameter intensity of the lymphocyte reference population used in the test point group and the standard reference mean vector. The intensity of each point in the reference population of the test point group is normalized based on the normalized vector generated for the lymphocyte reference population of the test point group. Determine whether the normalized vector of the lymphocyte reference population used for the test point group is greater than a predetermined value; In response to determining that the normalized vector of the lymphocyte reference population used for the test point group is greater than the predetermined value, the test point group is marked as having a problem with the quality control of the measurement.
3. A method for characterizing a biological cell assay set, comprising: Using a defined scheme, each cell in the biological cell test group is mapped to a corresponding point in n-dimensional space, and these corresponding points form a test point group; and The test point set is compared with a reference cluster set defined in the n-dimensional space, wherein the clusters in the defined reference cluster set correspond to maturity levels within a cell lineage, and the clusters are defined by centroid lines and radii. The comparison includes adjusting and classifying points in the test point set based on one or more multidimensional boundaries defined in the n-dimensional space using a support vector machine. The method includes: Determine the standard reference mean for each of a plurality of reference populations, including a lymphocyte reference population and a monocyte reference population. Determining the standard reference mean for each reference population includes: The support vector machine is trained to generate multidimensional boundaries in the n-dimensional space, thereby identifying the reference group of interest. For each normal patient in a group of normal patients, the generated multidimensional boundary is applied to identify a reference population of interest for the normal patients, and the average intensity of one or more parameters of the reference population is determined by summing the intensities of the populations of interest for a given parameter and dividing by the total number of cells in the populations of interest; and The standard reference mean of the reference population is calculated by determining the average of all mean reference strengths for each parameter of the reference population, wherein the standard reference mean is a vector; and The test point set is normalized, and the normalized test point set includes: The lymphocyte reference population of the test point group was identified by using the generated multidimensional boundary of the lymphocyte reference population of interest used to determine the standard reference mean. The average intensity of one or more parameters of the lymphocyte reference population for the test point group is determined by summing the intensities of points with a given parameter in the lymphocyte reference population and dividing by the total number of cells in the lymphocyte reference population for the test point group. A normalized vector for the lymphocyte reference population used in the test point group is generated by determining the difference between the mean parameter intensity of the lymphocyte reference population used in the test point group and the standard reference mean vector. The intensity of each point in the reference population of the test point group is normalized based on the normalized vector generated for the lymphocyte reference population of the test point group. Determine whether the normalized vector of the lymphocyte reference population used for the test point group is greater than a predetermined value; In response to determining that the normalized vector of the lymphocyte reference population used for the test point group is greater than the predetermined value, the test point group is marked as having a problem with the quality control of the measurement.
4. A method for characterizing a biological cell assay set, comprising: The test point set is compared with a normal cluster set in an n-dimensional space defined by a centroid line and a radius, the test point set representing each cell in the biological cell test set mapped to a corresponding point in the n-dimensional space using a defined scheme, wherein a support vector machine is used to determine at least one of the centroid line and radius based on one or more n-dimensional boundaries defined in the n-dimensional space; and A digital image is generated based on the comparison, wherein the method includes: Determine the standard reference mean for each of a plurality of reference populations, including a lymphocyte reference population and a monocyte reference population. Determining the standard reference mean for each reference population includes: The support vector machine is trained to generate multidimensional boundaries in the n-dimensional space, thereby identifying a reference population of cells of interest. For each normal patient in a group of normal patients, the generated multidimensional boundary is applied to identify a reference population of interest for the normal patients, and the average intensity of one or more parameters of the reference population is determined by summing the intensities of the populations of interest for a given parameter and dividing by the total number of cells in the populations of interest; and The standard reference mean of the reference population is calculated by determining the average of all mean reference strengths for each parameter of the reference population, wherein the standard reference mean is a vector; and The test point set is normalized, and the normalized test point set includes: The lymphocyte reference population of the test point group was identified by using the generated multidimensional boundary of the lymphocyte reference population of interest used to determine the standard reference mean. The average intensity of one or more parameters of the lymphocyte reference population for the test point group is determined by summing the intensities of points with a given parameter in the lymphocyte reference population and dividing by the total number of cells in the lymphocyte reference population for the test point group. A normalized vector for the lymphocyte reference population used in the test point group is generated by determining the difference between the mean parameter intensity of the lymphocyte reference population used in the test point group and the standard reference mean vector. The intensity of each point in the reference population of the test point group is normalized based on the normalized vector generated for the lymphocyte reference population of the test point group. Determine whether the normalized vector of the lymphocyte reference population used for the test point group is greater than a predetermined value; In response to determining that the normalized vector of the lymphocyte reference population used for the test point group is greater than the predetermined value, the test point group is marked as having a problem with the quality control of the measurement.
5. The method as described in any one of claims 1-4, wherein at least one aspect comprises: The cells were exposed to four reagents; and The fluorescence intensity and light scattering of the cells were measured at four levels using flow cytometry.
6. The method as described in any one of claims 1-4, wherein the scheme includes staining cells with a marker of CD10, a marker of CD19, a marker of CD20, and a marker of CD45.
7. The method as described in any one of claims 1-4, wherein the scheme comprises staining cells with a marker of FSC, a marker of SSC, a marker of CD20 FITC, a marker of CD10 PE, a marker of CD45, and a marker of CD19.
8. The method as described in any one of claims 1-4, wherein the scheme comprises staining cells with a marker of FSC, a marker of SSC, a marker of CD22 FITC, a marker of CD34 PE, a marker of CD45, and a marker of CD19.
9. The method as described in any one of claims 1-4, wherein the normal biological cell group is a subgroup of the normal biological cell sample.
10. The method of any one of claims 1-4, wherein the normal biological cell group comprises a plurality of subgroups, and each subgroup comprises a cell group selected from samples taken from an individual.
11. The method as described in any one of claims 1-4, comprising representing the defined reference cluster group in the n-dimensional space in a Cartesian coordinate display.
12. The method of claim 11, wherein the color is used to represent an additional dimension.
13. The method of claim 11, wherein the defined reference cluster group in the n-dimensional space corresponds to different maturation stages within the cell lineage.
14. The method as described in any one of claims 1-4, wherein the centroid line comprises a plurality of branches.
15. The method of claim 11, wherein the defined reference cluster group comprises a group of hyperellipsoids in the n-dimensional space defined by the centroid line and the radius.
16. The method as described in any one of claims 1-4, comprising: Cells in the test cell group are classified based on the multidimensional boundary in the n-dimensional space generated using a support vector machine.
17. The method of claim 16, wherein the multidimensional boundary in the n-dimensional space used to classify the cells in the test cell group is used to define the centroid line.
18. The method of claim 17, wherein the multidimensional boundary in the n-dimensional space is used to define the radius.
19. The method as described in any one of claims 1-4, comprising: If multiple test point groups are flagged as having problems with measurement quality control, it indicates that the instrument is not set up correctly.
20. A computer-readable medium storing instructions for enabling a system to facilitate the characterization of biological cell test groups by performing the method as described in any one of claims 1 to 19.
21. A system for characterizing biological cells as a test cell group, comprising: One or more memory units; as well as A digital signal processing circuit coupled to the one or more memories, wherein the digital signal processing circuit implements the method as described in any one of claims 1 to 19 in operation.
Citation Information
Patent Citations
Method and device for classifying, displaying, and exploring biological data
CN102144153A
Method and system for characterizing cell populations
US20150087240A1