Detection of early stage lung cancer in sputum using automated flow cytometry and machine learning
The combination of flow cytometry and machine learning with TCPP labeling in sputum analysis addresses the limitations of existing lung cancer screening by providing a highly accurate and non-invasive method for early detection.
Patent Information
- Application Number
- JP2024576395
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-20
- Filing Date
- 2023-06-22
- Publication Date
- 2025-08-20
AI Technical Summary
Current lung cancer screening methods, such as LDCT, have high false positive rates and involve invasive follow-up procedures, while sputum cytology is time-consuming and prone to observer bias, necessitating a more accurate and non-invasive test for early lung cancer detection.
A flow cytometry method using tetra(4-carboxyphenyl)porphyrin (TCPP) labeling and automated gating to analyze sputum samples, combined with machine learning, to classify cells and identify lung cancer with high sensitivity and specificity.
The method achieves 82% sensitivity and 88% specificity in distinguishing cancerous from non-cancerous sputum samples, reducing unnecessary procedures and improving early lung cancer detection.
Smart Images

Figure 2025527109000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and benefit of U.S. Provisional Patent Application No. 63 / 357,994, entitled "Sputum Analysis by Flow Cytometry; an Effective Platform to Analyze the Lung Environment," filed July 1, 2022, and U.S. Provisional Patent Application No. 63 / 390,826, entitled "Detection of Early-Stage Lung Cancer in Sputum Using Automated Flow Cytometry and Machine Learning," filed July 20, 2022. The specification and claims therein are incorporated herein by reference. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] Not applicable. INCORPORATION-BY-REFERENCE OF MATERIAL SUBMITTED ON A COMPACT DISC
[0003] Not applicable. Statement regarding prior disclosure by the inventor or co-inventors
[0004] Not applicable. copyrighted work
[0005] Not applicable. [Background technology]
[0006] It should be noted that the following description refers to a number of publications by author(s) and publication years, and that a particular publication should not be considered prior art to the present invention simply because of its most recent publication date. The mention of such publications herein is intended to provide more background and should not be construed as an admission that such publications are prior art for purposes of determining patentability.
[0007] An estimated 1.8 million people worldwide died of lung cancer in 2020. In 2022, an estimated 130,180 people will die of lung cancer in the United States alone. Overall, the 5-year lung cancer survival rate remains low at 22.9%, as most patients have advanced disease. The American National Lung Screening Trial (NLST) demonstrated that LDCT screening detects 93.8% of lung cancers in high-risk individuals (i.e., individuals aged 55–74 years with a smoking history of more than 30 pack years who currently smoke or have quit smoking within the past 15 years). Low-dose computed tomography (LDCT) is the standard of care for lung cancer screening in the United States (US). While LDCT has a sensitivity of 93.8%, its specificity of 73.4% leads to potentially harmful follow-up procedures for patients without lung cancer. Therefore, there is a need for additional, more accurate tests that can be used as an adjunct to LDCT for lung cancer diagnosis. Low-dose spiral computed tomography may not provide a clear treatment strategy when the identified nodules are small.
[0008] The NLST demonstrated that LDCT screening reduced lung cancer-specific mortality by 20% overall compared with screening with chest radiography. However, in this study, 96.4% of positive LDCT scans were false positives, and approximately 90% of LDCT-positive patients underwent additional procedures to determine whether the nodules seen on the LDCT scan were cancerous. These procedures, including imaging, biopsy, and surgical resection, can cause serious adverse effects, including death. New guidelines for interpreting LDCT scans and models for estimating the probability that a nodule is cancerous have improved the false positive rate (FPR). Nevertheless, only a portion of eligible patients undergo LDCT screening. Lack of information about the benefits and potential harms of screening (whether due to lack of knowledge or time), the costs associated with LDCT, difficulty accessing LDCT, and repeated radiation exposure from serial LDCT scans all likely contribute to the low adoption rate.
[0009] A simple, noninvasive, radiation-free, cost-effective test that could help physicians more reliably confirm or rule out a lung cancer diagnosis would likely reduce unnecessary follow-up procedures and increase lung cancer screening. Sputum has long been a part of lung cancer diagnosis and is a readily available bodily fluid. The PAP sputum cytology test, developed by Papanicolaou and optimized by Saccomanno, was the first lung cancer diagnostic method, beginning in the 1960s. In this test, two sputum smear slides are labeled with PAP stain and read by a pathologist specializing in lung cytology. While the sensitivity of sputum cytology varies widely, its specificity is extremely high. A review of 16 published studies of sputum cytology involving over 28,000 patients reported sensitivity ranging from 42% to 97%, with an average sensitivity of 66% and an average specificity of 99%.
[0010] The low sensitivity of sputum cytology is due in part to inadequate specimens and limited analysis of some samples. Incompatibility can occur because the sample obtained is saliva, or because mucus, debris, or red blood cells in the smear obscure the cellular components necessary for accurate analysis. Over the years, modifications to the original sputum cytology test have improved its sensitivity. Nebulizers and auxiliary devices, such as the Acapella and Langflute, and patient adherence to proper instructions on how to obtain lung sputum samples have been shown to improve patients' sputum production. Liquid cytology tests and automated slide preparation devices can reduce background contamination in sputum smears, thereby improving slide quality. Increasing the number of sample reads has also been shown to increase the likelihood of detecting abnormal cells indicative of lung cancer.
[0011] Porphyrins, such as TCPP, are currently used as diagnostic reagents for bladder cancer and for identifying the margins of cancerous tissue during surgery. Using microscopy, we demonstrated that labeling sputum cells with the fluorescent porphyrin TCPP allows for highly accurate differentiation between study participants with lung cancer and those without the disease using slide examination (a cytology-based method) and human reviewers. Cytology-based methods are of limited utility due to the time-consuming nature of slide reading and the need for highly specialized personnel. Furthermore, the presence of large amounts of debris and excessive numbers of squamous epithelial cells (SECs) or buccal cells often renders samples unsuitable for diagnosis. Because slide examination is time-consuming, which often interferes with the analysis of the entire sample, potentially resulting in missed important events, and because human reviewers introduce subjective bias (also known as operator bias) that results in inconsistency between different human reviewers, alternative methods for analyzing sputum to determine the likelihood of cancer would be useful.
[0012] It is shown in accordance with one embodiment of the present invention that it is feasible to use a flow cytometry platform to analyze sputum samples to identify significant differences between samples obtained from individuals diagnosed with lung cancer and those without the disease, without interfering with the operation of the instrument.
[0013] Early detection of lung cancer through screening can increase survival and reduce morbidity. Certain regions of the United States and the United Kingdom currently advocate annual low-dose computed tomography (LDCT) screening for high-risk individuals. Therefore, a positive LDCT result requires follow-up testing to determine whether the nodule is benign or malignant. These medical procedures carry inherent morbidity and mortality risks. 6 This can place a significant burden on screening participants and their families, and the associated costs represent a significant economic burden on patients and society.
[0014] Therefore, efforts have been directed toward developing noninvasive tests that can be used in conjunction with LDCT or as stand-alone tests to identify those at high risk of having lung cancer and who should undergo LDCT. In either case, the goal of these tests is to identify lung cancer patients early while simultaneously eliminating unnecessary medical procedures in low-risk patients. One readily accessible material from the lung is sputum, which contains a variety of blood cells and exfoliated bronchial epithelial cells, including precancerous and malignant cells in lung cancer patients. 7We previously reported a slide-based test capable of classifying cancer and non-cancer patients from sputum stained with tetra(4-carboxyphenyl)porphyrin (TCPP). Although the accuracy was 81%, the reading of labeled slides was time-consuming, subject to observer bias, and under-sampling could result in missing important low-frequency events. A high-throughput approach using automated flow cytometry (FCM) for sputum sample analysis may ameliorate the shortcomings of slide-based analysis.
[0015] The disclosed embodiments combine flow cytometry and machine learning to develop a sputum-based test that can assist physician decision-making in such cases. Summary of the Invention [Means for solving the problem]
[0016] One embodiment of the present invention provides a flow cytometry method for analyzing a sputum sample from a subject suspected of having lung cancer. A plurality of cells, e.g., a suspension of single cells from the sputum sample, is obtained from the subject suspected of having lung cancer. The plurality of cells is marked with i) a plurality of cell lineage-specific marker compositions, ii) a cell viability composition, and iii) a tetra(4-carboxyphenyl)porphyrin (TCPP) composition. For example, i) may include at least three, at least four, at least five, or at least six of CD206, CD3, CD19, CD66b, CD45, EpCAM, PanCK, and any combination thereof, where any combination may explicitly exclude any of CD206, CD3, CD19, CD66b, CD45, EpCAM, and PanCK. In a further example, ii) the cell viability composition preferentially labels dead cells over live cells and may include FVS510. The plurality of cells marked by i to iii are analyzed using a flow cytometer to obtain a cell size-selected subpopulation from the plurality of cells based on an automatically selected bead size exclusion gate. For example, the bead size exclusion gate is set to 5 μm to approximately 30 μm, and events smaller than approximately 5 μm and larger than approximately 30 μm are not further analyzed. For example, the analyzing step includes obtaining flow cytometry values for side scatter, forward scatter, TCPP-derived fluorescence, fluorescence from a cell viability composition, and fluorescence from multiple lineage-specific marker compositions from the plurality of cells. A viable singlet cell population is selected from the cell size-selected subpopulation using an automated non-debris gate (e.g., the automated non-debris gate excludes a majority of dead cells from the non-debris population) and an automated singlet gate (e.g., the automated singlet gate is applied to the cell population selected by the automated non-debris gate). Flow cytometry values are obtained from the viable singlet cell population based on the multiple lineage-specific marker compositions, the viability marker, and the TCPP marker. A trained classifier is applied to the metadata obtained from the subject (e.g., age) and the obtained flow cytometry values.A classification of the sputum sample is generated based on application of the trained classifier, where the classification is selected from a plurality of classification options including cancer and non-cancer.
number
[0017] In one embodiment, the lineage-specific marker composition comprises fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19, and fluorescent anti-CD66b. In another embodiment, the classifier is a linear equation with coefficients b0 through b5 determined by fitting the classifier model to the particular set of samples used to construct the classifier.
[0018] Another embodiment provides a system for automated analysis of flow cytometry data, the system comprising a computer processor in communication with a memory storing flow cytometry data from a plurality of markers in a plurality of cells from a sputum sample of a subject, the plurality of markers including: i) a plurality of cell lineage-specific marker compositions; ii) a cell viability composition; and iii) a tetra(4-carboxyphenyl)porphyrin (TCPP) composition. For example, i) may include at least three, at least four, at least five, or at least six of CD206, CD3, CD19, CD66b, CD45, EpCAM, PanCK, and any combination thereof, where any combination may explicitly exclude any of CD206, CD3, CD19, CD66b, CD45, EpCAM, and PanCK. In a further example, ii) the cell viability composition preferentially labels dead cells over live cells and may include FVS510. The system further provides a computer program product embodied in a non-transitory computer-readable medium, the computer program product including instructions for causing a computer processor to: receive flow cytometry data acquired from a plurality of cells from a sputum sample; select from the plurality of cells in the sputum sample an automatically selected subpopulation of cells based on application of an automated gate selected from a bead size exclusion gate, a viability gate, and a singlet gate; for example, set the bead size exclusion gate between 5 μm and about 30 μm, and events less than about 5 μm and greater than about 30 μm are not further analyzed; for example, an automated no-debris gate excludes a majority of dead cells from the no-debris population; for example, an automated singlet gate is applied to the cell population selected by the automated no-debris gate; determine flow cytometry values from this subpopulation for a plurality of lineage-specific marker compositions, viability markers, and TCPP markers of interest; apply a classifier to the flow cytometry values of interest and the subject metadata, for example, the trained classifier:
number
[0019] Another embodiment of the present invention provides a non-transitory computer-readable medium including program code that, when executed, causes a processing circuit to: obtain flow cytometry values (e.g., side scatter, forward scatter, TCPP-derived fluorescence, cell viability composition-derived fluorescence, and cell lineage-specific marker composition-derived fluorescence) for a viable singlet population of sputum cells from the subject based on a plurality of cell lineage-specific marker compositions (e.g., the plurality of cell lineage-specific marker compositions selected from fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19, and fluorescent anti-CD66b), a viability marker (e.g., but not limited to, FVS510), and a TCPP marker; apply a trained classifier to the metadata obtained from the subject and the obtained flow cytometry values; generate a classification of the sputum sample based on application of the trained classifier, the classification being selected from a plurality of classification options including cancer and non-cancer; and the method of claim 1, wherein the cell viability composition preferentially labels dead cells over live cells. In one embodiment, a bead size exclusion gate is set to exclude events less than about 5 μm and greater than about 30 μm to select a subpopulation of cells from the sputum sample having a cell size that is not excluded from further analysis by the bead size exclusion gate, and from this population, a viable singlet population is selected. In another embodiment, an automated non-debris gate that excludes the majority of dead cells from the non-debris population is used to select the viable singlet population for further analysis. In one embodiment, the trained classifier is
number
[0020] One aspect of an embodiment of the present invention provides a flow cytometry method, e.g., automated FCM, for analyzing sputum from a subject suspected of having lung cancer, the method comprising: 1) isolating debris and squamous epithelial cells (SECs) from the analyzed sputum sample using a gating strategy defined, e.g., by bead standards and viability dyes; The method includes one or more of the following: 1) removing both pulmonary and non-pulmonary contaminants (e.g., pulmonary leukocytes and common oral contaminants) to generate a singlet cell population; 2) including quality control parameters to detect alveolar macrophages in the sample, thereby verifying that each sputum sample is of pulmonary origin; 3) establishing a numerical cutoff for the cell population of interest for sample suitability to provide a reliable analysis; 4) obtaining optical properties from sputum-derived cells labeled with one or more of a fluorescent-specific antibody or fragment thereof and a leukocyte-specific marker and / or an epithelial cell lineage-specific marker, such as TCPP, to identify significant differences between samples obtained from individuals diagnosed with lung cancer and those without lung cancer; 5) obtaining subject metadata, such as information about age and / or years of smoking; 6) applying a classifier based on the properties selected from the output of items 1-5; and 7) determining whether the analyzed sputum sample is above or below a numerical value representing the likelihood of cancer. If cancer or the likelihood of cancer is confirmed, the subject is treated for further testing.
[0021] One embodiment of the present invention provides a computer-implemented method for classifying a lung sputum sample from a subject at risk for lung cancer, the method comprising receiving data from the subject on at least one processor. The at least one processor is used to evaluate the data using a classifier, which is an electronic representation of a classification system, the classifier being trained using a plurality of electronically stored training datasets, each of the plurality of training datasets representing a separate training dataset, each separate training dataset representing an individual subject and data for each subject, each training dataset further comprising a determination of a lung cancer characterization, if present for each subject, wherein the classification system comprises a cancer or non-cancerous identification for the lung sputum sample. The at least one processor is used to evaluate a classification of the test sputum sample from the subject based on the evaluating step. In one embodiment, the data comprises flow cytometry data, subject metadata, or a combination thereof. For example, the flow cytometry data is obtained from a sputum sample labeled with a plurality of markers in a plurality of cells from the sputum sample of the subject, the plurality of markers including: i) a plurality of cell lineage-specific marker compositions; ii) a cell viability composition; and iii) a tetra(4-carboxyphenyl)porphyrin (TCPP) composition. The subject metadata includes one or more of: gender, age, genetic information, biomarker data, smoking status, medical history, or a combination thereof. In a further embodiment, a non-transitory computer-readable medium storing an executable program includes instructions for performing a computer-implemented classification method. [Brief explanation of the drawings]
[0022] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate one or more embodiments of the invention and, together with the description, serve to explain the principles of the invention. The drawings are only for purposes of illustrating one or more embodiments of the invention and are not to be construed as limiting the invention.
[0023] [Figure 1] 1 shows a flow chart of an identified sample according to one embodiment of the present invention.
[0024] [Figures 2A-2C] Figures 2A and 2B show automated gating of FCM data according to one embodiment of the present invention. Figure 2A shows that the bead size exclusion (BSE) gate parameters are set for the entire sputum sample. Figure 2B shows that events exceeding the BSE gate threshold, set at 2.5 x 10, based on observation of both forward scatter (FSC) area and side scatter (SSC) area, are dead cells. Figure 2C shows a dot plot of NIST beads obtained from the automated flowClust analysis, illustrating how it is used to set the lower left threshold of the BSE gate.
[0025] [Figure 3A] Figure 3A shows automated gating of FCM data according to one embodiment of the present invention. Figure 3A shows that for some samples, events appear in the lower right corner ("debris") of a forward scatter-height (FSC-H) versus side scatter-height (SSC-H) plot. [Figure 3B] Figure 3B shows automated gating of FCM data according to one embodiment of the present invention. The events identified as debris in Figure 3A have extremely low optical SSC-A signatures that are extremely rare in human cell populations. [Figure 3C-3I]Figures 3C-3I show automated gating of FCM data according to one embodiment of the present invention. Figures 3C-3I are histograms with the determined density on the Y-axis and the determined fluorescence intensity on the X-axis. Figure 3C shows that the events identified as debris in Figure 3A are identified as live because they are not stained with a viability dye (FVS510), a dye that preferentially stains dead cells. Figure 3D shows that the debris identified in Figure 3A shows low levels of TCPP when labeled with TCPP. Figure 3E shows that the debris events identified in Figure 3A are negative for CD45. Figure 3F shows that the debris events identified in Figure 3A are negative for CD66b / CD3 / C19 when incubated with a compound that specifically recognizes CD66b, e.g., an antibody, a compound that specifically recognizes CD3, e.g., an antibody, and a compound that specifically recognizes CD19, e.g., an antibody. Figure 3H shows that the debris events identified in Figure 3A are negative for CD206 when incubated with a compound that recognizes CD206, e.g., an antibody. Figure 3G shows that the debris events identified in Figure 3A are negative for Pan-CK when incubated with a compound that specifically recognizes Pan-CK. Figure 31 shows that the debris events identified in Figure 3A are negative for EpCAM when incubated with a compound that specifically recognizes EpCAM. In accordance with one embodiment of the present invention, the debris events identified by the profile in Figure 3A are filtered out to avoid false positive results.
[0026] [Figure 4A] 1 shows a dot plot analysis according to one embodiment of the present invention. [Figure 4B] 1 shows a dot plot analysis according to one embodiment of the present invention. [Figures 4C-4F] Figure 4F shows a dot plot analysis according to one embodiment of the present invention.Figure 4F shows a histogram analysis of the non-debris events identified as shown in Figure 3A according to one embodiment of the present invention.
[0027] [Figure 5A]2A and 3B show dot plots of events identified when debris is excluded as in FIGS. 2A and 3A. [Figure 5B] 2A and 3B show dot plots of events identified when debris is excluded as in FIGS. 2A and 3A. [Figure 5C] 2A and 3B show dot plots of events identified when debris is excluded as in FIGS. 2A and 3A. [Figure 5D] 2A and 3B show dot plots of events identified when debris is excluded as in FIGS. 2A and 3A. [Figure 5E] 2A and 3B show dot plots of events identified when debris is excluded as in FIGS. 2A and 3A. [Figure 5F] Figure 5F shows a dot plot of events identified when debris is excluded as in Figures 2A and 3A. Figure 5F shows a histogram of events identified when debris is excluded as in Figures 2A and 3A. [Figure 5G] 2A and 3B show dot plots of events identified when debris is excluded as in FIGS. 2A and 3A. [Figure 5H] 2A and 3B show dot plots of events identified when debris is excluded as in FIGS. 2A and 3A.
[0028] [Figure 6A] A histogram showing cells stained with TCPP is shown, with TCPP / log10SSC on the x-axis and density on the y-axis. [Figure 6B] A histogram showing cells stained with TCPP is shown, with TCPP / log10SSC on the x-axis and density on the y-axis.
[0029] [Figure 6C] A histogram showing cells stained with FVS510 (a viability dye) is shown, with FVS510 / log10FSC on the x-axis and density on the y-axis. [Figure 6D]A histogram showing cells stained with FVS510 (a viability dye) is shown, with FVS510 / log10FSC on the x-axis and density on the y-axis.
[0030] [Figure 6E] Dot plots are shown with CD206 on the x-axis and CD66b / CD3 / CD19 on the y-axis. [Figure 6F] Dot plots are shown with CD206 on the x-axis and CD66b / CD3 / CD19 on the y-axis.
[0031] [Figure 7] 1 shows a flowchart of data analysis for sputum-based pulmonary testing according to one embodiment of the present invention.
[0032] [Figure 8A] Receiver operator (ROC) graphs of false positive rate versus true positive rate calculated by varying the model response threshold are shown.
[0033] [Figure 8B] 8B is a graph of the model value (probability of cancer) based on the ROC curve of FIG. 8A.
[0034] [Figure 9] A correlation graph of age on the x-axis versus model value on the y-axis is shown when samples were analyzed using an LSRII flow cytometer or a NaviosEX cytometer, respectively.
[0035] [Figure 10] 1 is an exemplary system according to one embodiment of the systems disclosed herein.
[0036] [Figure 11] 1 illustrates an embodiment of classifier development, according to an embodiment of the present invention.
[0037] [Figures 12A-12D]Figure 12 shows cells selected by a size exclusion gate, a live cell gate, and a doublet discrimination gate using flow cytometry analysis according to one embodiment of the present invention. Cells in a sputum sample were either unstained with PE (Figure 12A) or stained with an anti-CD45 antibody fluorescently labeled with PE ("CD45-PE") (Figure 12B). The upper panel indicates the cell population that is CD45+ staining cells with a "+" and the lower panel indicates a "-" where CD45- cells are located. Forward scatter (FSC) is on the x-axis. The absence (-) and presence (+) of CD45 staining were determined in samples exposed to the anti-CD45 antibody fluorescently labeled with PE, as indicated by a "+" square (cells staining positive for CD45; "CD45+ cells") and a "-" square (cells not staining with CD45; "CD45- cells"). The CD45 positivity cutoff is based on the unstained sample. Figure 12C shows a representative profile of CD45+ cells from a sputum sample in a "blood tube" stained with antibodies that distinguish between granulocytes and lymphocytes (CD66b, which binds granulocytes; CD3, which binds B cells; and CD19, which binds T cells; see Gate 1) and alveolar macrophages (CD206; see Gate 2) and interstitial macrophages (CD206; see Gate 3), according to one embodiment of the present invention. As used herein, the cells in Figure 12C may be referred to as "blood cells," and the cells in Figure 12D may be referred to as "non-blood cells." Figure 12D shows a representative profile of CD45- cells stained with antibodies against epithelial markers, pancytokeratin (panCK) ("Alexa488-labeled panCK antibody" on the "y" axis) and EpCAM ("EpCA-PE-CF-594-labeled EpCAM antibody" on the "x" axis), according to one embodiment of the present invention. Gate 4 represents a subpopulation of cells within the CD45 − cell population that stained positive for both epithelial markers.
[0038] [Figures 13A-13C]Figures 13A and 13B show dot plots showing TCPP versus CD66b / CD3 / CD19-FITC / Alexa488 (using a "blood tube") (Figure 13A) and TCPP versus panCK-Alexa488 (using an "epithelial tube") (Figure 13B), according to one embodiment of the present invention. The upper box labeled "H" indicates the TCPP-high gate used to define the TCPP-high cutoff. Figure 13C shows a histogram of TCPP fluorescence intensity versus relative cell number on the "y" axis, according to one embodiment of the present invention, where the TCPP-high cutoff is derived from the gate shown in Figure 13A or Figure 13B. The TCPP-low population is defined as the intersection of the histogram of the unstained sputum sample and the TCPP-stained sample. The intermediate TCPP staining population, TCPP intermediate, is defined as the population between the TCPP-high and TCPP-low populations.
[0039] [Figures 14A-14J] Figures 14A-14C show sputum cell populations selected from viable singlet sputum cells subdivided by the level of TCPP staining (shown in Figures 13A-13C) and further analyzed by the "blood cell" and "epithelial cell" markers shown in Figure 12, according to one embodiment of the present invention. Figures 14A-14D show the analysis of TCPP-high cells. Figures 14E-14H show the analysis of TCPP-intermediate cells. Figures 14I-14L show the analysis of TCPP-low cells. The first profile in each row, Figures 14A, 14E, and 14I, shows the light scatter profile of the respective TCPP subpopulation. The second profile in each row, Figures 14B, 14F, and 14J, shows the distribution of CD45 staining, or lack thereof, of cells in the respective TCPP subpopulation. The CD45+ fraction ("+") of each TCPP subpopulation was further analyzed and shown in the respective profiles in column 3 (Figures 14C, 14G, and 14K), which show the distribution of CD66b / CD3 / CD19 staining versus CD206 staining. The fourth column (Figures 14D, 14H, and 14L) shows the respective CD45- ("-") fraction of each TCPP subpopulation, which shows pan-cytokeratin staining versus EpCAM staining.
[0040] [Figures 15A-15C] Figures 15A, 15B, and 15D show the differences between cancer (CA) and non-cancer (Non-CA) samples obtained from lineage marker analysis and size from the TCPP-high cell populations described in Figures 14A, 14B, and 14D, according to one embodiment of the present invention. Each dot (CA) and square (Non-CA) represents one sample. Figure 15A shows that the TCPP-high population in cancer samples has a smaller SSC than the TCPP-high population in non-cancer samples (**p<0.01). Figure 15B shows that in cancer samples, the proportion of EpCAM+panCK+ cells in the CD45- fraction of the TCPP-high subpopulation is greater than the corresponding CD45- fraction in the non-cancer samples (**p<0.01). Figure 15C shows that the EpCAM mean fluorescence intensity (MFI) of TCPP-high CD45-EpCAM+panCK+ cells is higher in cancer samples than in the corresponding cell population in the non-cancer samples (*p<0.05). The thick black horizontal bars indicate the median values for each sample group.
[0041] [Figures 16A-16C] Figures 16A-16C show the differences between cancer (CA) and non-cancer (Non-CA) samples obtained from the blood cell populations described in Figures 12A-12C, according to one embodiment of the present invention. Each dot (CA) and square (non-CA) represents one sample. Figure 16A shows that the proportion of CD45+ cells in sputum samples from cancer (CA) samples is significantly higher than that in non-cancer (Non-CA) samples (**p=0.0099). Figure 16B shows that among CD45+ cells, the granulocyte / lymphocyte cell subpopulation (Gate 1 in Figure 12C) is significantly higher in sputum samples from cancer patients compared to non-cancer patients (*p=0.0378). Figure 16C shows that the CD45+ subpopulation of interstitial macrophages (Gate 3 in Figure 12C) was also significantly greater in sputum samples from cancer patients compared to non-cancer patients (**p=0.0031). The thick black horizontal bars indicate the median values for each sample group. DETAILED DESCRIPTION OF THE INVENTION
[0042] Before describing the present invention in more detail, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0043] Detailed Description of the Invention Sputum is a bodily fluid that can be obtained noninvasively and separated to release its cellular contents, providing a snapshot of the lung environment. Sputum can be made into a single-cell suspension and stained with both TCPP and fluorochrome-conjugated antibodies for manual and automated flow cytometry (FCM) analysis. Automated FCM uses TCPP and a trained classifier to interrogate the sputum sample with a series of compounds specific for cell lineage markers to obtain information about the lung environment from the subject who provided the sputum sample and capture predictive features of the sputum sample, allowing the sputum sample to be analyzed for cancer or cancer-associated cells. In one embodiment, the sputum sample analyzed by flow cytometry contains approximately 10,000 cells in subpopulations selected by automated gating and analyzed and classified in real time by the systems and methods disclosed herein. In other embodiments, the total number of cells in the subpopulation can be about 1,000 to 5,000, about 5,000 to 10,000, or about 10,000 to 100,000. Data acquired by automated FCM can be stored and analyzed at a later time, instead of being analyzed in real time as the data is acquired.
[0044] Over the past two decades, the field of automated flow cytometry analysis has produced powerful software tools for identifying cell populations that correlate with clinical outcomes and managing increasingly complex FCM datasets. For example, in human immune profiling, significant efforts have been devoted to replicating expert analysis of FCM data and automating cell population identification. Such data-driven algorithms can now match or exceed human expertise and fully automate the analysis of acquired FCM data, thereby eliminating potential operator bias. However, applying automated flow cytometry analysis of sputum samples to determine lung health has proven challenging due to the complexity of the lung environment, the interplay between inflammatory markers and disease, and the frequent absence of cells of interest in sputum samples. Complicating factors in analyzing sputum samples for lung health include the variability in sputum sample size: too few samples contain too few cells to analyze, while too many samples can dilute events of interest when they are rare. Cytological testing of sputum samples on slides is limited to a portion of the sample, which reduces the sensitivity of slide testing, and sensitivity is dependent on the observer's ability to see meaningful events on the slide. When looking for rare events in a sample, smearing on the slide and overlapping cells on the slide reduce the usefulness of cytology to provide meaningful analysis. Further cytology does not provide a complete picture of non-malignant cells that may be present in the sample and their incidence in the sample, and what this information means for lung health in relation to lung cancer or other lung diseases.
[0045] Another aspect of an embodiment of the present invention provides a supervised learning approach to develop a test that combines automated FCM data acquisition from induced sputum to isolate viable single-cell events with machine learning techniques to classify patient samples as cancerous or non-cancer. In another embodiment, sputum is not induced with saline and / or is not collected by lavage. In another embodiment, sputum is induced by sound waves or vibrations within the lungs or by use of a flute medical device. In another embodiment, sputum is collected by spontaneous expectoration and / or with the assistance of a positive expiratory pressure therapy device.
[0046] In one embodiment of the present invention, the developed lung cancer / non-cancer classifier performs well, with a sensitivity of 82% and a specificity of 88%. Furthermore, the classifier exhibits comparable sensitivity and specificity when applied to an independent set of collected samples using a different flow cytometer platform (Navios EX) than that used for test development (LSRII). One aspect of the tan automated FCM lung test system and method according to one embodiment of the present invention is that the test is accurate even for early stages (I and II) and small lung nodules (<20 mm in diameter). The system and method according to one embodiment of the present invention is robust to differences in sample handling and processing and captures important predictors of early lung carcinogenesis.
[0047] In one embodiment of the present invention, sputum was obtained from current or former smokers, e.g., with a 20-pack-year smoking history or greater, who were identified as having lung cancer or at high risk for developing the disease. Separated sputum cells were counted, determined to be viable or dead, and labeled with a series of markers to determine cell type, e.g., anti-CD45 to separate leukocytes from non-leukocytes; other markers are possible, as discussed herein. After excluding debris and dead cells, including squamous epithelial cells, we identified reproducible population signatures, confirming the pulmonary origin of the samples. For example, in addition to labeling sputum samples with leukocyte-specific and epithelial-specific fluorescent antibodies, we also labeled sputum samples with fluorescent meso-tetra(4-carboxyphenyl)porphyrin (TCPP), which is known to preferentially stain cancer-associated cells. Differences in cellular characteristics, population size and fluorescence intensity, useful for distinguishing cancer samples from high-risk samples were identified.
[0048] In one embodiment of the present invention, we developed an analytical pipeline that combines automated flow cytometry processing with machine learning to distinguish cancerous from non-cancerous cells in sputum samples. Flow data and patient characteristics were evaluated to identify predictors of lung cancer. The training set was used to fit the model, and the remaining samples were used for independent validation (test set). We further validated the approach on a second set of samples processed on a different flow cytometry platform.
[0049] Referring now to Figure 1, a flowchart illustrates the utilization of sputum samples. Of the 171 samples (136 non-cancer; 31 cancer; 4 undetermined health status) initially considered and run on the LSRII flow cytometer, 168 were used for model building and analysis pipeline development. Because the addition of unlabeled samples is useful for model building, we included four samples with undetermined disease status. In addition, 14 samples flagged as ineligible based on cell count (see below) were also used in model development to better capture the underlying data distribution and help make model generalization more robust to sample noise. Three samples could not be used at all due to technical issues during acquisition.
[0050] A final set of 150 samples was used for the model validation phase (122 non-cancer; 28 cancer). Of the 168 samples, 18 were excluded: 13 had too few cells for accurate analysis, 1 had too few alveolar macrophages to be confirmed as a lung sample, and 4 samples were excluded because their cohort status could not be confirmed. Independent validation of the automated analysis was performed using 32 new samples. Participants met the same enrollment criteria, and samples were processed using the same protocol as the previous set of samples. A different flow cytometer (Navios EX) was used to run the second set of samples, but the same model and coefficients were used to analyze data from both instruments. The initial set of 171 samples run on the LSRII was considered (136 high-risk; 31 confirmed cancer; 4 inconclusive), and 150 samples were ultimately used to develop the automated analysis pipeline (122 high-risk; 28 confirmed cancer). Of the 21 excluded samples, 13 had too few cells for accurate analysis, 1 had too few alveolar macrophages to be confirmed as a lung sample, and 3 had technical issues during acquisition. An additional 4 samples were excluded because their cohort status could not be confirmed. Of the 45 samples processed on the Navios EX, 7 were excluded because they had too few cells, 1 was excluded because they had too few macrophages, and 5 were excluded because of technical issues (1 during processing and 4 with the flow cytometer). The remaining samples consisted of 26 high-risk samples and 6 cancer samples.
[0051] Referring now to Figure 2A, the bead size exclusion (BSE) gate parameters were set for the entire sputum sample. Figure 2B shows a 2.5 x 10 sputum size exclusion (SSC) gating parameter based on observation of both forward scatter (FSC) and side scatter (SSC) areas. 5Events above the BSE gate threshold set in ( ) indicate dead cells. Figure 2C shows a dot plot of NIST beads obtained from automated flowClust analysis and illustrates how it is used to set the lower left threshold of the BSE gate. The lower left threshold of the BSE gate is derived from automated NIST bead flowClust analysis. Single-cell suspensions from 3-day sputum samples were labeled with viability dyes to exclude dead cells, antibodies to distinguish cell types, and porphyrins to label cancer-associated cells and run on a flow cytometer.
[0052] A step toward the development of a computer-automated lung cancer test is the identification of viable single cells by an automated flow cytometer (FCM), which involves the sequence of events depicted by FIGS. 2A, 3A, 4A, and 4B according to one embodiment of the present invention. In one embodiment of the test, the sample preparation component of the method consists of multiple test tubes, e.g., one test tube may be labeled with a blood cell marker (the "blood tube") and another with an epithelial cell marker (the "epithelial tube"). For example, both tubes may contain a fluorescent anti-CD45 antibody that selectively binds to blood cells, and a viability dye to facilitate viability gating (see viability gate in FIG. 4A) to further analyze and identify cancer or cancer-associated cells and to remove dead cell populations, including squamous epithelial cells (SECs), from the TCPP. Cells in the blood tube may be stained with anti-CD206 (a marker for lung macrophages) and / or anti-CD3 (a marker for T cells), and / or anti-CD19 (a marker for B cells), and / or anti-CD66b (a marker for granulocytes) in addition to or instead of CD45. Cells in the epithelial tube can be stained with pan-cytokeratin (panCK) antibodies and / or antibodies against epithelial cell adhesion molecule (EpCAM). Viability dyes, TCPP cell markers, and blood cell marker(s) and epithelial cell marker(s) can be distinguished from one another based on their optical properties, e.g., fluorescence. In one embodiment of the present invention, the fluorescence intensity, forward scatter, and side scatter of cell subpopulations selected from the more comprehensive cell population based on selection criteria including one or more of bead size exclusion gating, singlet gating, viability dye gating, lineage marker gating, and TCPP gating of flow cytometry sputum data were used for downstream numerical analysis.It should be noted that multiple lineage-specific markers may be selected to distinguish different cell subpopulations in a sputum sample, such as macrophages (e.g., by using anti-CD206), T cells (e.g., by using anti-CD3), B cells (e.g., by using anti-CD19), granulocytes (e.g., by using anti-CD66b), leukocytes (e.g., by using anti-CD45), and epithelial cells (e.g., by using anti-EpCAM and / or anti-PanCK). In one embodiment, the specific lineage marker is not limited to the specific example provided, as are other cluster of differentiation (CD) markers, such as CD64 to identify interstitial macrophages. 高 / CD11b 高 / MHCII / CD11c 高 / siglec F 陰性 and CD64 on alveolar macrophages 高 / CD11c 高 / F4 / 80 陽性 / MerTK 陽性 / siglec F 低 and combinations thereof may be used, and CD4 and CD8 may be used to recognize T cells.
[0053] Additionally, in one embodiment, the control samples may include one or more of the following: polystyrene beads of known diameter (approximately 5-30 μm NIST beads), calibration samples for each fluorescent dye channel used, an unstained sputum sample, and a cell lineage / cell type marker sample (e.g., antibody isotype sputum control). Each sample tube corresponds to a single Flow Cytometry Standard (fcs) file, which contains sample metadata and per-event values for each acquired optical and fluorescent channel, as well as time parameters recorded as the contents of the sample tube are interrogated and / or acquired by the flow cytometer.
[0054] SECs are highly autofluorescent and can result in false-positive events when sputum samples are examined and / or analyzed by flow cytometry. Therefore, removing SECs from the cell population analyzed during sputum sample analysis is a step in one embodiment of the lung test. The inventors have confirmed that neither physical removal of SECs by filtration prior to analysis nor negative size selection during analysis results in the exclusion of SEC cells. A live / dead cell discriminator (FVS510) for removing SECs from the cell population derived from the sputum sample being analyzed was one solution for excluding SECs from the population of events acquired by the flow cytometer and / or analyzed in downstream analyses.
[0055] When the sputum cells of interest and SECs fell within this gated region, we analyzed the viability of isolated sputum cells within the size parameters of 5–30 μm. The FVS510 positivity cutoff was based on unstained controls. Backgating dead cells on the sputum light scatter profile revealed that these cells generally had a high SSC, as would be expected for SECs. To confirm the viability of SECs, sputum cells were sorted into dead and live cell populations. Aliquots of the pre-sorted samples and the sorted populations were transferred to cytospins and stained with Wright-Giemsa. These slides revealed that SECs were primarily present among dead cells, whereas live cells sorted from the same samples contained hematopoietic and nonhematopoietic cells, as well as a small number of contaminating SECs. Therefore, we concluded that sputum samples can be analyzed by flow cytometry while excluding contaminating SECs using the viability gate.
[0056] Figure 2A shows the process of restricting events to be analyzed downstream based on forward scatter area (FSC-A) and side scatter area (SSC-A), both of which are valid surrogates for cell size. A two-dimensional cluster gate was used to find the major peak of 5 μm NIST beads in FSC-A vs. SSC-A (Figure 2C). The lower FSC-A limit of the bead clusters was set as the minimum sample FSC-A to exclude small particles and debris. Events above this threshold were classified as dead (FVS510). + ) cells (Figures 2A and 2B), and therefore, the upper limit of 2 × 10 for both FSC-A and SSC-A was set. 5 was set.
[0057] Figures 3A-3B illustrate dot plot data setup according to one embodiment of the present invention. Referring now to Figure 3A, some samples display events in the lower right corner of the FSC-H vs. SSC-H plot. As shown in Figures 3C-3I, these events are excluded from further analysis to avoid including live, marker-negative, or low-TCPP cell events / cells. Figure 3A shows that a temporary flowClust gate (non-debris gate) is set on FSC-H vs. SSC-H non-debris events such that mostly live cells are retained for the final FVS510 tail gating ("core viability gate"). Figure 3B shows dot plots of low SSC-A and low FSC-A debris, illustrating the identification of additional debris (Figure 3A), which can be seen to exhibit a light scattering profile that is not typical of the cell population. Further analysis (Figures 3A-I) shows that this debris can be mistaken for live cells (Figure 3C) that are low or negative for other markers (Figures 3D-I), and therefore must be excluded.
[0058] Events within a bead size exclusion (BSE) gate were then restricted to exclude populations with abnormal FSC-height and SSC-height profiles (Figure 3A) and staining profiles that may result in inappropriate inclusion in downstream analyses. Exclusion of potentially abnormal populations is generally justified.
[0059] Referring now to Figure 4A, a viability gate (solid rectangle) was set for the remaining events based on FVS510-A fluorescence. Cells within the viability gate (set at the FVS510 positivity threshold) were retained (Figure 4A). For some samples, setting the viability gate was difficult due to variability in sputum cell composition between patients. In these cases, a heuristic-based viability gate setting was used (Figure 4C-F).
[0060] Referring now to Figure 4B, a singlet gate (solid polygon) was set for all surviving events (events within the viability gate in Figure 4A). The identified (surviving) singlets were used for downstream numerical analysis. For some singlet gates, setting the viability gate was difficult due to the variability in sputum cellular composition between patients. In these cases, a heuristic-based singlet gate setting was used (Figure 5A-H).
[0061] Figures 4C-4E show dot plots of heuristic-based viability gate settings. The subpopulation most likely to contain viable singlet cells (i.e., cells with relatively small area and height in the light scatter channel) is used to guide the viability gate positioning (Figures 4C and 4D). Figure 4C shows a sample with a low percentage of high SSC / FSC cells. The temporary flowClust gate was set on non-debris events in FSC-H vs. SSC-H to ensure that most viable cells were retained for the final FVS510 tail gating (the "core viability gate"). Figure 4D shows that for samples with less than 10% of events within the core viability gate, flowClust was re-run more comprehensively by increasing the "quantile" parameter to 0.99. Figure 4E shows the temporary singlet gate set on core viability events, which is calculated by placing the top right point of both axes at 2.5x10. 5 The upper diagonal line is complemented by setting the viability cutoff to 0. Figure 4F shows a histogram of FVS510 staining (viability dye) for events identified by the temporal singlet gate in Figure 4E. The viability cutoff is automatically set for the core surviving singlets (black histogram). The line identified as the "blue curve" is the entire non-debris FVS510 profile for comparison. The vertical line identified as the "red bar" indicates the viability gate cutoff. Surviving events are to the left of the viability gate cutoff. Once the viability cutoff is determined, all temporal gates are removed. Using the viability gate cutoff thus determined, a viability gate is set for the entire sample (shown in Figures 4C and 4D).
[0062] Referring now to Figures 5A-5D, heuristic-based singlet gating is shown. In some cases, high SSC-A cells (see, e.g., Figure 3A) miss the singlet gate for the entire viable cell population. Figure 5A shows a dot plot with FSC-A on the x-axis and FSC-W on the y-axis in which a singlet gate (solid polygon) was automatically assigned to a sample with too many high SSC-A cells. A broad singlet gate is an inaccurate gate. Such a gate can be corrected by gating a transient subpopulation of viable cells to exclude high SSC-A cells (Figure 5B). Figure 5B shows a singlet gating pattern with FSC-A on the x-axis and SSC-A on the y-axis, approximately 5x10 4 Figure 5C shows a dot plot of SSC-A temporally gated to exclude events exceeding 2.5 × 10. The gate was automatically applied to the restricted population (solid polygon) and the upper right corner of both axes was set to 2.5 × 10. 5 Figure 5D shows the refined singlet gate (solid polygon) placed on the surviving population when the temporary gate is removed (compare the refined polygon in Figure 5D with the singlet gate (solid polygon) in Figure 5A).
[0063] Referring now to Figures 5E-5H, heuristic-based singlet gating is shown for different types of cases that deviate from the singlet gate setting. In these cases, a population accounting for more than 10% of the singlets falls between approximately 2.5 (log-linear scale) and the viability threshold that throws the viability and singlet gate settings. This can be corrected by re-setting the viability gate and applying the singlet gate to a more restricted surviving population. Figure 5E shows a dot plot of such a difficult case, with FVS510-A on the x-axis and FSC-A on the y-axis. The solid oval represents the singlets greater than 10%, and the dashed line represents the viability threshold. Figure 5F shows the population mixture analysis shown as a density histogram of FVS510 staining. This analysis highlights the difference in signal distribution between the majority of events located to the left of the ellipse in Figure 5E (identified as the black curve) and the rightmost population identified by the solid ellipse in Figure 5E (identified as the blue curve). A natural cutoff (dashed line) is indicated at 2.5 for these aberrant cases. Figure 5G shows a dot plot with an adjusted viability gate (solid rectangle) replacing the viability gate found by automated tail gating (dashed line). Figure 5H shows a dot plot with FSC-A on the x-axis and FSC-W on the y-axis, where a new singlet gate (solid polygon) was calculated for the refined live cell population identified by the adjusted viability gate shown in Figure 5G.
[0064] A "singlet" gate (FSC area vs. FSC width) was used to exclude cell doublets or small aggregates (Figure 5A). Due to sample variability between patients, it was difficult to automatically set the singlet gate. In some samples, high SSC-A cells were included in the live cell population, skewing the results of the singlet gating algorithm. This could be corrected by gating a transient subpopulation to exclude most of these high SSC-A events (Figure 5B). In other samples, two populations were observed within the live cell gate: one with low FVS510 staining and the other slightly below the viability threshold and with a high side scatter profile. Correction in this case involved redefining the viability gate and applying the singlet gate to the more restricted viable population. Light scatter and fluorescence signal values for each single event were recorded and used for downstream model development and validation.
[0065] Based on the results of our previous slide examination, we predicted that smoking history (or correlates such as age) and TCPP signal density (rather than fluorescence intensity itself) would be important predictors. Therefore, we calculated the fluorescence signal of all channels as log 10 FSC-A or log 10 The resulting density distribution was divided by SSC-A and divided into three regions (R1 = less than approximately 0.25, R2 = approximately 0.25 to 0.6, and R3 = more than approximately 0.6, Figures 6A to 6D). Two such density signals, i.e., TCPP / log 10 SSC-A (Figure 6A, region 3 [R3]) and FVS510-A / log 10 FSC-A (Figure 6C, region 2 [R2]) was found to be informative for the classifier: TCPP / log 10 The predictive value of SSC-A signal density emerged spontaneously, rather than being arbitrarily set by stepwise regression. The fact that FVS510-A signal density was also found to be informative is intriguing and may be related to the fact that apoptotic cells can take up this dye at intermediate levels.
[0066] Figure 6A shows a histogram of live single cells stained with TCPP, with the R1, R2, and R3 regions shown, with the R3 region identified by a solid rectangle. Region R3 (shaded) represents live singlets with high TCPP signal relative to side scatter (log10 transformed). Figure 6B plots the x-axis as TCPP / log 10 Figure 6C shows a histogram showing cells without TCPP (unstained control) with SSC as the y-axis and density as the x-axis, with no events observed in region R3. 10 Histograms of viable single cells are shown with FSC as the y-axis and density as the y-axis, showing cells negative for FVS510 (R1) or stained with low levels of FVS510 (R2; solid shaded rectangle). While both R1 and R2 cells are below the FVS510 positivity cutoff, R2 cells nevertheless have a relatively high FVS510 signal relative to forward scatter (log10 transformed) compared to the unstained control sample in Figure 6D. Figure 6D shows the x-axis as FVS510 / log 10 Histograms are shown with FSC as the y-axis and density as the y-axis, showing cells that were not stained with FVS510 (unstained control).
[0067] Combinations of lineage markers can identify subpopulations that cannot be captured by a single lineage marker alone. Careful examination of patient sputum samples by FCM revealed complex patterns of lineage marker expression in blood and epithelial tubes, but it was unclear what information this revealed about lung health.
[0068] Figure 6E shows a dot plot with CD206 on the x-axis and CD66b / CD3 / CD19 on the y-axis. The solid shaded box in the lower quadrant of CD206 identifies non-macrophages, while the solid box in the middle to upper quadrants of CD206 identifies lung macrophages. Anti-CD206 is a macrophage cell marker specific for a macrophage population present in lung tissue but not in the blood circulation. Anti-CD66b is a granulocyte cell marker, anti-CD3 is a T cell marker, and anti-CD19 is a B cell marker. Figure 6F shows a dot plot with CD206 on the x-axis and CD66b / CD3 / CD19 on the y-axis. The solid shaded box in the lower quadrant of CD206 identifies the absence of non-macrophages, and the solid shaded box in the middle to upper quadrants of CD206 identifies the absence of lung macrophages from the control sample. Figure 6E shows the absence of non-macrophages (CD206 低 ) White blood cells (CD66b / CD3 / CD19 中間 The shaded area (CD206) is shown. 中間 / 高 CD66b / CD3 / CD19 低~中間 ) contains lung macrophages, which are used as a marker for samples deemed to adequately capture the lung environment. Compare with the unstained control sample in Figure 6F, where quadrant labeling also applies to Figure 6E. Both panels are from the same illustrative sample.
[0069] Pairwise analysis of cell markers was performed by partitioning the fluorescence based on signal distribution in the blood tube (Figures 6E-6F) and epithelial tube (data not shown). Signal intensity on a log-linear scale was quantified into windows of low (<1.5), low-to-intermediate (1.5-2.5), intermediate (2.5-3), and high (>3). The resulting number of events per 10,000 was tallied for each region of a 4 x 4 grid of CD206 vs. CD3 / CD19 / CD66b (blood tube) or EpCAM vs. panCK (epithelial tube, data not shown). Regions of the blood signal intensity grid were found to be informative for the predictive classifier lung testing model (tan-shaded rectangle; low levels for CD206 and intermediate levels for CD3 / CD19 / CD66b, Figure 6E). This population may indicate the presence of an immune or inflammatory process in the lung.
[0070] The development of automated processing of flow cytometer data features of sputum samples acquired on an LSRII flow cytometer according to one embodiment of the present invention provides one or more setup and / or quality control steps prior to automated analysis of patient tubes / samples as follows: For each sample tube, remove outlier events using, for example, the time vs. fluorescence channel, e.g., [flowCut] Using compensation tubes to automatically derive leakage matrices, e.g., using [flowStats]; It should be noted that on the Navios flow cytometer, automated processing begins with an operator spill matrix of the unstained sample, and the alignment of the medians in the "off" channel (i.e., the non-FITC channel if FITC is being compensated) is checked and fine-tuned as necessary based on the expectation that in the off channel (e.g., the non-FITC channel) the means of the positive and negative populations (as defined in the channel being compensated, e.g., FITC) should be aligned. Correct the fluorescence signal and convert it to a log-linear scale, for example using [flowWorkspace]; It should be noted that the automatic processing in Navious converts 20-bit signals (0-1,048,575) to 6 decades to roughly match the LSRII range. Automatically find the smallest NIST bead peak and set a lower FSC-A threshold (low) using, for example, [openCyto] to exclude most of the debris. For example, when using an LSRII flow cytometer, set a threshold of 10 4 (low)~5×10 4 (High) window is the region of interest; For Navios, 10 4 (low)~10 5 It should be noted that you should look at the (High) window. Further automated analysis of patient tubes / samples according to one embodiment of the present invention provides: Patient Tube: Events within the rectangular size gate are retained for further analysis, followed by "BSE" (FSC-A: low ~2.5 x 10 5 , SSC-A: 0~2.5x10 5 ). [flowWorkspace] Navios: "BSE" (FSC-A: low to 10 6 SSC-A: 5 x 10 3 ~10 6 ) A secondary "NON-DEBRIS" gate was applied to the retained BSE population to identify occasional high FSC-H (2x10 5 Ultra) SSC-H is medium / high (1x10 5 Exclude hyper-abnormal cell populations. [flowWorkspace]; In Navios: The equivalent of the NON-DEBRIS gate in LSRII is not the exclusion gate for FSC-H vs. SSC-H, but the selection polygon "cleanSSC" defined for SSC-W vs. SSC-A (lower left = 50, 50; upper left = 50, 10 6 ;Bottom right=200, 50;Top right=10 3 , 10 6) The choice of channel depends on how well debris can be visualized and excluded. This gate is applied at the end of Navios processing for the retained viable singlets. Use a subset of cells (a preliminary gating set) that are most likely to contain the viable singlet population of interest to set the coordinates of the viability gate, which helps reduce artifacts and distortions caused by contaminating squamous epithelial cells (SECs) and other dead cells. For LSRII: Use FSC-H vs. SSC-H "hxh" to remove most of the contaminating SEC, and measure both channels at 1 x 10 5 Set an automatic flowClust gate for NON-DEBRIS events, restricting them to < 0.99 quantile cutoff. [openCyto, flowClust] For some samples, if less than 10% of events are retained in hxh (low.viable=true), the channel limit should be set to 1.5×10 5 and the quantile to 0.9. · Set auto singletGate "singlet" in hxh using FSC-A vs. FSC-W channels and "wider_gate=TRUE" setting. [openCyto, flowStats]; Force the top right FSC-A coordinate to be the same as the bottom right one to avoid downward distortion of the gate due to lingering SECs. [flowWorkspace]; Using the FVS510-A channel, limit min=2, max=3 (min and max on a log-linear scale), tolerance=0.1, and set the auto tailgate "VIABLE" to singlet. [openCyto] If low.viable=true, first identify singlet events with SSC-A<5×10 4 and set the adjustment parameter adjust=1;[flowCore] Alternatively, singlet events were initially screened for FSC-A <5 x 104 and set the adjustment parameter adjust=1.2; [flowCore] It should be noted that the processing steps are somewhat different for Navios, as Navios resolves the BV510-A signal in more detail than LSRII. BV510-A [0,3.5] vs. SSC-W [50,200]. Set a rectangular gate on NON-DEBRIS events using "quickLive" to exclude most of the contaminating SECs. ·Set up automatic singletGate "SINGLETS" in quickLive using FSC-A vs. FSC-W channels and "wider_gate=FALSE" setting. Force the rightmost FSC-A coordinate to the BSE limit (10 6 ) and the maximum FSC-W limit is 2 × 10 5 to counter the case where lingering SECs cause the gate to warp downwards. ·Using the FVS510-A channel, set the automatic tailGate "VIABLE" for BSE with the arguments (side="left", max=3.5, tolerance=0.1). It should be noted that the BV510-A cutoff found in most samples is reasonable, around 4, but that heuristic "fine-tuning" may be required to correct for cases where the signal does not fall into two well-resolved populations. Fine-tuning is performed on the VIABLE tailGate parameter. It should be noted that the isotype sample does not contain TCPP, resulting in a higher BV510-A signal. Fine-tuning threshold: reason.cut=4, high.cut=5 (patient samples excluding isotype) reason.cut=4.5, high.cut=5.5 (isotype) If (number of events where BV510-A>VIABLE cutoff) is more than 1.5 times (number of events where BV510-A>reason.cut): The VIABLE cutoff is too low and needs to be moved to the right. Set the tailGate "min" argument to the larger of the reason.cut and the BV510-A mean expression value, and increase by 0.5 if more than 25% of the events above the reason.cut also exceed the high.cut. If Ff (the number of events where BV510-A>VIABLE cutoff) is more than twice the number of events where BV510-A>reason.cut): Set tailGate to tolerance=0.2 and adjust=0.5. If (number of events where BV510-A>VIABLE cutoff) is less than 1.5 times (number of events where BV510-A>reason.cut): The VIABLE cutoff is too high and needs to be moved to the left. Set the tailGate "max" argument to the smaller of the high.cut and the BV510-A mean expression value, and increase it by an additional 0.5 if >25% of events are BV510-A>. · Set the tailGate "min" argument to 3.5 and increase it by another 0.5 if more than 25% of events have BV510-A>high.cut. If (the number of events where BV510-A>VIABLE cutoff) is less than 2.25 times (the number of events where BV510-A>reason.cut): There is an overshift to the left, so increase the tailGate "min" by another 0.5 and set adjust=2. ·Create a temporary gating set using events from both the BSE gate and the SINGLETS gate, and recalculate the VIABLE of the automatic tailGate using the fine-tuned arguments. Apply the VIABLE gates determined from the provisional gating set to the NON-DEBRIS events in the complete gating set. [flowWorkspace] · The automated singletGate "SINGLETS" set to VIABLE events gives good results in most cases, but some samples still contain confounding events at this point. SEC contamination may map close to larger viable cells, resulting in a bottom-left coordinate less than 0 or a bottom-right coordinate less than the top-left coordinate. In either case, set VIABLE to SSC-A < 5 x 10 before setting up an automatic singletGate with "wider_gate=FALSE". 4 [openCyto, flowStats] In both cases, force the top right FSC-A coordinate to be the same as the bottom right one to avoid downward distortion of the gate due to lingering SECs. [flowWorkspace] - Apply a SINGLETS gate to a VIABLE event. [flowWorkspace] When viable cells are relatively scarce, the proportion of events retained in the SINGLETS gate in the gap between FVS510-A>2.5 (log-linear scale) and the viability cutoff can be substantial. If the "gap" population is greater than 10% of SINGLETS, a forced survival cutoff of FVS501-A<2.5 will be used. -Remove "VIABLE" from the complete gating set. [flowWorkspace] Set "VIABLE" rectangleGate gate for NON-DEBRIS of FVS510-A<2.5. Recalculate the automated singletGate "SINGLETS" as described above. [openCyto, flowStats, flowWorkspace] - Complete gating set. Apply SINGLETS gate to VIABLE events. [flowWorkspace] For Navios, remove the intermediate gates from the patient samples and keep only the BSE. -Set a VIABLE gate on BSE. ·Apply a SINGLETS gate to VIABLE. · Apply cleanSSC gate to SINGLETS. For each patient tube gate, a complete matrix of events retained by SINGLETS (LSRII) or cleanSSC (Navios) is written out. These values, along with patient metadata, serve as input to the cancer / non-cancer classifier. It should be noted that the retained events are quantized to generate model variable values. Due to differences in the output values of the LSRII and Navios detectors, the ranges will not be identical. See 5.5 Grid Settings below.
[0071] According to one embodiment of the present invention, the analysis steps of the LSRII flow cytometer [R Package(s)] with details of the embedded Navios are revealed: Note 1: The process for the Navios EX flow cytometer is essentially the same, but adjusted for the different file format (i.e., LMD files for Navios EX instead of FCS files for LSRII) and detector sensitivity and dynamic range (LSRII = 18 bit, Navios = 20 bit). 1.1: The LSRII utilizes two configuration files: one to match file names with compensation fluorescent controls (matchfile.csv), and one to identify the patient tube(s) and flow cytometer channel names (samplematch.csv). For specific examples, see 5. Additional Materials "LSRII Configuration Files" (5.1 and 5.2) below. 1.2: Rather than using a configuration file, Navios sample collection now uses a controlled tube naming protocol. See 5.3 Controlled Tube Naming Protocol for Navios below. 1.3 The names and order of fluorescence channels in flow cytometer files are different between LSRII and Navios, but can be mapped. See 5.4 Channel Equivalence and Detector Maximums below. Note 2: Gate names are shown in bold capital letters. Note 3: The names of temporary gates are shown in bold lower case. Note 4: Calculated thresholds are in bold italics. Note 5: The main package(s) used in this process are indicated in [square brackets]. Note 6: Some heuristic adjustments were necessary to handle a wide range of sample compositions and viability.
[0072] An embodiment of a sample processing pipeline according to one embodiment of the present invention is provided below: Analysis Tube: NIST beads Compensation tubes (one per fluorescence channel) Patient single cell suspension (4 tubes): Unstained as a control for PE-anti-CD45 and TCPP Isotype controls stained with FVS510 (viability detection), PE-anti-CD45, PE-CF594 isotype, A488 isotype, and FITC isotype Blood stained with TCPP, FVS510, PE-anti-CD45, PE-CF594-anti-CD206, A488-anti-CD3, FITC-anti-CD66b, and A488-anti-CD19 (Note: A488 and FITC are read in the same channel, resulting in a combined CD66b / CD19 / CD3 signal) "Epithelium" stained with TCPP, FVS510, PE-anti-CD45, PE-CF594-anti-EpCAM (epithelial cell adhesion molecule), and A488-anti-PanCK (pan-cytokeratin). 5.5 Grid Settings JPEG2025527109000005.jpg111151
[0073] 5. Additional Materials and Methods for Entering Data into the System 5.1 Sample matchfile.csv configuration file for LSRII Specimen_001_BL-18-148 Unstained_011.fcs,unstained Specimen_001_A549 FVS510_007.fcs,BV510-A Specimen_001_PE comp beads_003.fcs,PE-A Specimen_001_PE-TxRed comp beads_005.fcs,PE-Texas Red-A Specimen_001_FITC comp beads_004.fcs,FITC-A Specimen_001_A549 TCPP_008.fcs,APC-A 5.2 Sample samplematch.csv configuration file for LSRII # multiple collection tubes separated by / / variable,detail blood,s_BL-19-151_Blood_016.fcs / / s_BL-19-151_Blood-tube2_017.fcs epithelial,s_BL-19-151_Epithelial_018.fcs isotype,s_BL-19-151_Isotype_015.fcs stub,BL-19-151 fitc_name,FITC-A nist,s_NIST_001.fcs 5.3 Controlled Tube Naming Protocol for Navios 01_NIST 03_FITC 04_PE 05_PECF594 (Texas Red) 06_APCpos (must be combined with tube 13 pre-correction) 07_FVS510 08_Accession#_Unstained (patient sample) 09_Accession#_Isotype (patient sample) 10_Accession#_Blood (patient sample) 11_Accession#_Epithelial (patient sample) 13_APCneg (must be combined with tube 6 pre-correction)
[0074] 5.4 An example of channel equivalence and detector maxima for different instruments used in data conversion, shown in the grid below. JPEG2025527109000006.jpg97150
[0075] We identified a narrowed list of lineage markers that showed the most promise as potential predictors, and investigated the pairwise interactions between them for classifiers that yielded high sensitivity and high selectivity in tests to identify likely and unlikely cancerous cells. Surprisingly, the correlation between age and FVS510-A / log 10 They found that adding a negative value proportional to the "number of FSC-A R2 events" improved the performance of the classifier. One interpretation of this interaction term is that it helps to mitigate the possible age-related accumulation of stressed cells as a result of smoking or health history in high-risk patient groups.
[0076] After developing the two stages of the lung test, a complete pipeline was assembled, including quality control steps, determination of predictor variable values, and sample classification (Figure 7).
[0077] Referring now to FIG. 7, a lung sputum data processing pipeline is shown in accordance with one embodiment of the present invention. This schematic illustrates the following key elements: Quality Control (QC) measurements 701 and 703 are performed on the data and instrumentation; data acquisition files are reviewed for suitability; sputum samples containing approximately 100 million viable singlet events and approximately 10 or more lung macrophages are determined to be suitable; subject-specific data, such as the information in Table 1, are obtained and combined with sputum sample characteristics obtained from the FCM-acquired population 705 from FCM; and classifier 706 is a linear equation with coefficients b0 through b5 determined by fitting the classifier model 706 to the particular set of samples used to build the classifier. The classifier 706 is applied to new samples to determine a cancer likelihood value based on the cutoff value 707. The classifier model is applied to the FCM-acquired population obtained from the sputum sample and the subject-specific data input into the classifier model to determine whether the sample is likely 708 or not 709 to be cancerous (bottom diamond). The coefficients b0-b5 derived from fitting the classifier 706 to the particular set of samples used to build the classifier are as follows: b0:53.515414 b10.701153 b2:0.001545 b3:58.071495 b4:4.258709 b5:0.784982 Those skilled in the art will appreciate that fitting the classifier 706 to a different set of sample values will change the values of the coefficients, but will not change the model itself.
[0078] In one embodiment of the present invention, data acquisition suitability 701 and sample acceptability assessment 703 begin by verifying that each collection tube's data file is readable and its encoded data matrix is complete. The time signature is used to inspect each tube's fluorescence channel to remove flow anomalies resulting from air bubbles or clogs during sample acquisition. A fluorescence compensation tube is then used to derive a compensation matrix de novo (rather than using the compensation matrix encoded in the sample file metadata). The fluorescence signal is corrected and converted to a log-linear scale to generate a sample data matrix used by automated FCM gating to isolate viable singlet events. To ensure reliability for downstream numerical analysis, samples containing a threshold number of singlets, e.g., at least 10,000 viable singlets, were analyzed, taking into account the possibility of low event counts within some analysis windows. To confirm that the sputum samples were derived from the lung, the shaded region in Figure 6E (CD206) where lung macrophages are found was analyzed. 中間及び高 CD3 / CD19 / CD66b 低~中間 The threshold was set so that there were at least 10 cells in each region.
[0079] A further step in the testing pipeline of Figure 7 is to provide a classifier model with values based on age and flow from surviving singlets. In one embodiment of the present invention, the classifier for determining the likelihood of having cancer uses four variables: i) age; ii) TCPP / log(s) per 10,000 surviving singlets (per 10K); and iii) TCPP / log(s) per 10,000 surviving singlets (per 10K). 10 number of events with SSC-A, iii) FVS510-A / log per 10K in region 2 ( Figure 6B , R2); 10 Number of events with FSC-A, and iv) CD206 per 10K 低 CD3 / CD19CD66b 中間The correlation coefficients are determined using the number of events in the region (Figure 6C, shaded rectangle). Additionally, the classifier includes a negative interaction term between age and the FVS510 density variable, as well as an "intercept" term (b0) that helps avoid forcing our multifactorial model to zero when all variables are set to zero but cannot be directly interpreted as biologically meaningful components of the classifier. The values of the coefficients (b1, b2, b3, b4, b5) depend on the training set used for model fitting and provide weights to the variables. We did not normalize the data provided to the model to facilitate model interpretation. For example, the model equation shows that increasing the number of events with high TCPP density increases the likelihood of cancer, consistent with our previous results.
[0080] The decision step in the pipeline in Figure 7 is to assign cancer / non-cancer. The model returns a value in the range [0, 1]. Whether a given sample is classified as cancer or not depends on whether the model returns a value greater than a predetermined cutoff. If the value is less than or equal to the cutoff, it is classified as non-cancer. A reasonable cutoff can be selected by setting a stepwise cutoff value between 0 and 1 and measuring true and false positive determinations at each step compared to known group categories.
[0081] The development of a GLM classifier according to one embodiment of the present invention is further iterated below in Figure 11. The input to this process is shown by the output of the FCM step (see, e.g., 1101-1102 in Figure 11). Evaluate combinations of potential predictors (1103-1105 in Figure 11). 1) Combinations of potential predictors were evaluated (see, e.g., 1103 in FIG. 11 ) to develop a classifier using a generalized linear model (GLM) for semi-supervised machine learning as follows: 1.1 Clinical parameters available for all samples 1.2. Measurements obtained from patient blood and epithelial samples (light and fluorescence) and quantized using equally spaced or heuristically placed split points (3x3 and 4x4 grids per channel tested) individually and in pairwise combinations. 1.3. Quantified readings as in #2 above minus background counts 1.4. Quantized Fluorescence / log10(Light Scattering) 1.5. Frequencies of specific subpopulations identified by prior manual analysis 2) Using separate training and test groups (approximately 2 / 3 training and 1 / 3 test, randomly selected and non-overlapping (see, e.g., 1104-1105 in Figure 11 ), including sequential step-forward and step-backward parameters in the GLM; 3) Evaluate the GLM based on the Akaike information criterion (AIC), which estimates the prediction error and provides a comparative measure of model quality (see, e.g., 1105 in Figure 11). 4) retaining parameters that are repeatedly retained in different iterations of model testing with various combinations of predictor variables tested individually or in combination to detect potential interactions (see, e.g., 1104-1105 in Figure 11 ); 5) Controlling overfitting of the final model by repeated random sampling (n=10) of training / test sets to verify predictive robustness independent of the specific samples used to train the model; (e.g., see 1106 in Figure 11); and 6) Validate the processing pipeline and classifier predictions on samples not used in model building and testing (see, e.g., 1107-1110 in Figure 11). After the classifier is developed, it is applied to new samples (see 1101 and 1111 in Figure 11) that have not been previously analyzed and whose cancer / non-cancer classification is unknown.
[0082] Referring now to Figure 8A, the results of this process are shown as a receiver operating characteristic (ROC) curve with an AUC of 0.89. In testing, a threshold of 0.28 performed best in distinguishing between cancer and non-cancer (Figure 8B, solid vertical line).
[0083] Referring now to Figure 8A, receptor operating characteristic (ROC) curves are shown, with lines indicated by asterisks, showing the false positive rate versus true positive rate calculated by varying the model response threshold. For comparison, the ROC curve for a previous version of the slide-based lung test is shown as an inverted triangle. Dashed lines indicate 80% and 90% sensitivity and specificity (1 - false positive rate). Figure 8B shows a model based on the ROC curve of Figure 8A. The threshold for the model value (probability of cancer) was set at 0.28 (solid vertical line), corresponding to a sensitivity of 82.1% and 87.7%. Dashed lines indicate the level at which cancer prediction falls short of either a very low probability (leftmost dashed line) or a very high probability (rightmost dashed line). Filled circles represent cancer, and open circles represent high risk. CyPath Lung performance for LSRII samples is shown.
[0084] In one embodiment of the present invention, automated analysis of samples by flow cytometry combined with machine learning yielded a predictive model that was highly sensitive (82%) and specific (88%) and robust to differences in sample handling and disease stage. Importantly, the test was 92% sensitive and 87% specific in difficult-to-manage cases with no nodules or nodules 20 mm or smaller in diameter.
[0085] Referring now to Figures 9A-9B, correlation graphs show age on the x-axis versus modeled values on the y-axis when samples are analyzed on an LSRII flow cytometer or a NaviosEX cytometer, respectively.
[0086] Referring now to FIG. 10 , an exemplary system and method 1000 for receiving input from a flow cytometer 1002 and subject metadata information 1020 based on a sample 1001 analyzed by the flow cytometer is shown. A computer-implemented device 1003 includes a processor 1010 (e.g., processing circuitry) capable of executing program instructions or software, causing the computer to perform various methods or tasks, such as executing techniques for generating and / or using analytical models of classifier models described herein. The processor 1010 is coupled to memory 1030 via a bus (not shown), which is used to store information such as program instructions and / or other data during computer operation. A storage device 1040, such as a hard disk drive, nonvolatile memory, or other non-transitory storage device, stores information such as program instructions, data files of multidimensional data and reduced datasets, and other information. The computer also includes various input / output elements 1050, including parallel or serial ports, USB, Firewire or IEEE 1394, Ethernet, and other such ports for connecting the computer to external devices such as printers, video cameras, displays, medical imaging equipment, monitoring equipment, etc. Other input / output elements include wireless communication interfaces such as Bluetooth, Wi-Fi, and cellular data networks. The computer itself may be a traditional personal computer, a rack-mounted or business computer or server, or any other type of computerized system. In a further example, the computer may include fewer than all of the elements listed above, such as a thin client or mobile device having only some of the above elements. In another example, the computer is distributed among multiple computer systems, such as a distributed server having many computers working together to provide various functions.
[0087] Referring now to FIG. 11 , an exemplary process flowchart 1100 is shown for use in semi-supervised machine learning of a GLM classifier for cancer / non-cancerous status and subsequent classification of biological samples, such as sputum, in accordance with one embodiment of the present invention. A processed sputum sample is provided to the flow cytometer for analysis as outlined herein 1100. Flow cytometer data 1101 from multiple sputum samples is processed by automated FCM to extract a data matrix of viable singlet populations from each sample 1102. For all samples, values of potential predictor variables associated with the collected flow cytometry data matrix and associated clinical data are calculated 1103. The multiple samples are randomly divided into a training set and a test set 1104. The potential predictor values from the training set, along with the known cancer / non-cancerous status of the training set, are used to construct a GLM classifier. The resulting classifier is evaluated on the test set for its ability to correctly classify the test set into cancer / non-cancerous 1105. Step 1105 is repeated with various combinations of potential predictor variables, and only the truly predictive variables obtained from each iteration are retained. A GLM classifier is fitted to only the retained predictor variables to obtain the final model 1106. A single sputum sample of unknown status to be classified is processed through automated FCM 1101 (single) to obtain the final predictor variable values from the sputum and subject metadata (e.g., age), which are input into the final model 1111 to generate a classification value 1107. Based on the classification value cutoff determined by the ROC analysis, an output classification decision of cancer 1109 or non-cancer 1110 is generated for the sputum sample 1108.
[0088] According to one aspect of the present invention, the system and method accurately classify study participants as cancer or non-cancer with high accuracy, including participants at various stages of disease and participants with nodules less than 20 mm. Thus, this test has the potential to improve the process of early lung cancer diagnosis.
[0089] Materials and Methods Clinical Trials and Collection Facilities The minimal-risk study was registered with ClinicalTrials.gov, reviewed and approved by the Sterling Institutional Review Board (Atlanta, Georgia), and conducted in accordance with the ethical principles of the Declaration of Helsinki (v1996) and Good Clinical Practice guidelines. Sample collection took place at five study sites: Atlantic Health System (New Jersey); Mt. Sinai Hospital (New York); Radiology Associates of Albuquerque (New Mexico); South Texas Veterans Healthcare System; and Waterbury Pulmonary Associates (Connecticut). Each site received institutional approval to participate in the study. Each potential participant was presented with an informed consent form, and only those who signed it were enrolled.
[0090] Participant Information: Participants (male and female) were eligible for enrollment in one of two groups. The non-cancer group included participants (50-80 years old) who were either current smokers with a smoking history of at least 20 pack-years or current non-smokers with a smoking history of at least 30 pack-years who had quit smoking within the past 15 years. There were two exceptions: one patient who quit smoking 25 years ago and the other who quit smoking 26 years ago. Most participants in the non-cancer group had received LDCT results or other forms of imaging diagnostics that were not suspicious for cancer and were recommended to undergo LDCT screening again in 12 months. In a small number of cases, participants initially in the non-cancer group underwent follow-up LDCT, PET / CT, or biopsy. These participants were followed until their health status was confirmed. If a lung cancer diagnosis was made, the participant was transferred to the cancer group.
[0091] Each participant in the cancer group was assessed by a physician as likely to have lung cancer based on their medical history and the results of LDCT or other imaging studies. After providing a sputum sample, the diagnosis was confirmed by biopsy. The exception was a patient who developed a new 24-mm nodule and was unable to undergo biopsy due to excessive fragility. If the biopsy did not reveal cancer, the participant was transferred to the non-cancer group. There were no age or smoking history restrictions for enrollment in the cancer group.
[0092] For each participant, the following demographic data were collected: sex (male or female); age (years); ethnicity (Hispanic / non-Hispanic Latino / Latino); and race (American Indian / Alaska Native; Asian; Black / African American; Native Hawaiian / Other Pacific Islander; White; Other). Data on smoking history, as well as comorbidities (asthma, COPD, emphysema, chronic bronchitis), and previous cancer history were collected. All participants were required to provide contact information for their primary care physician and to consent to the release of their medical information upon request. Exclusion criteria included severe obstructive pulmonary disease, an inability to cough forcefully enough to obtain a sputum sample, angina with minimal exertion, and pregnancy.
[0093] Exclusion criteria included severe obstructive pulmonary disease, inability to cough forcefully enough to obtain a sputum sample, angina with minimal exertion, and pregnancy.
[0094] Sputum samples: Donors were trained to use the Acapella assistive device (Smiths Medical, St. Paul, MN) and repeated this procedure at home for three consecutive days, storing their specimen cups in a cool, dark place or in the refrigerator. Within one day of completing collection, samples were shipped overnight to the bioAffinity laboratory, where further processing and FCM analysis were performed by individuals blinded to the sample's origin.
[0095] Sputum processing: Sputum was isolated and labeled. For example, sputum samples were incubated with a mixture of 0.1% dithiothreitol and 0.5% N-acetyl-L-cysteine at room temperature for 15 minutes and neutralized with Hank's balanced salt solution. Cells were then filtered through a 100-micron nylon strainer, washed, and resuspended in HBSS. Total cell yield was determined using the trypan blue exclusion method. Sputum was liquefied using pre-warmed 0.1% dithiothreitol (DTT) at a 1:4 ratio (w / v) and pre-warmed 0.5% N-acetyl-L-cysteine (NAC) at a 1:1 ratio (w / v) based on the sputum weight. The resulting cell suspension was filtered through a 100 μm nylon cell strainer (Falcon, Corning) to remove larger debris while minimizing cell loss. Cells were collected in a 50 mL conical tube, washed, and centrifuged at 800 × g for 10 minutes. The separated sputum pellets were combined into one 15 mL conical tube per sputum sample. Total cell yield and viability were determined with a Neubauer hemocytometer using the trypan blue exclusion method.
[0096] In one embodiment, a small portion of the cells was reserved for control use, while the majority was split into two tubes for main analysis. For example, both tubes were labeled with Fixable Viability Stain 510 (FVS510) and CD45-PE. One tube, the so-called "blood tube," contained CD66b-FITC, CD3-Alexa-Fluor-488, CD19-Alexa-Fluor-488, and CD206-PE-CF594. In the other tube, the "epithelial tube," cells were labeled with pan-cytokeratin-Alexa-Fluor-488 and EpCAM-PE-CF594. The cells were incubated on ice for 35 minutes. After washing with HBSS, the cells were fixed and stored on ice until the next day, and TCPP solution (20 μg / mL) was added (3.3 × 10 6 After incubation on ice for 1 h, cells were washed twice with cold HBSS and kept on ice until analysis.
[0097] In another embodiment, the sample is divided into at least two tubes: white blood cells (CD45 + ) one tube containing a marker for investigating the cellular compartment and one tube containing the epithelial (CD45 - Cell labeling was performed by separating the sputum into separate tubes for cell compartments. Each tube contained anti-CD45 antibody, FVS510 (to exclude dead cells, including SECs), and porphyrin TCPP (to identify cancer-associated cells). To identify leukocyte populations, anti-CD206 antibody, which labels macrophages, and a cocktail of antibodies labeling granulocytes (anti-CD66b) and lymphocytes (anti-CD3 and anti-CD19) were added. Anti-cytokeratin (panCK) and anti-EpCAM were used to identify epithelial cells. Because the initial DTT and NAC treatments for sputum processing were sufficient for intracellular cytokeratin staining, a permeabilization step for cytokeratin labeling was not performed.
[0098] Isolated sputum cells were incubated with antibodies and FVS510 for 35 minutes. After washing once with cold HBSS, cells were fixed with paraformaldehyde on ice for 1 hour, after which the cells were washed again and stored on ice until labeling with TCPP the next day. TCPP was added to the cells for 1 hour. After incubation, cells were washed twice with cold HBSS and then stored on ice until flow cytometry analysis. Cells were kept on ice and protected from light throughout the labeling procedure until analysis. For further reagent details, see Table 7.
[0099] Flow cytometry: Sputum samples were acquired on a BD LSR II flow cytometer (BD Biosciences) equipped with four lasers (404 nm, 488 nm, 561 nm, and 633 nm) or a Navios EX (Beckman Coulter Life Sciences) equipped with three lasers (405 nm, 488 nm, and 638 nm). Post-acquisition data analysis was performed using FlowJo software (Tree Star, Ashland, OR).
[0100] Sample Characteristics: Of the 171 LSRII patient samples collected, 150 were suitable for analysis by the full testing pipeline, consisting of 122 from high-risk patients without cancer and 28 from lung cancer patients (Table 1). Because the addition of unlabeled samples proved useful for model building, four additional samples with undetermined disease status were included in the pipeline development phase. Additionally, 14 samples flagged as ineligible based on counts (see below for implementation of the lung testing pipeline) were also used in the model fitting phase to better capture the underlying data distribution and help make model generalization more robust to sample noise. Only three samples could not be used at all due to technical issues during acquisition.
[0101] Table 1. Patient characteristics of LSRII lung test validation samples [Table 1]
[0102] Traditionally, the presence of "large numbers" of macrophages in a sputum smear is considered indicative of a sample derived from the lung. In one embodiment of the present invention, the FCM lung test utilized a quality control measure using CD206, a cell surface antigen specific to a macrophage population present in lung tissue but not in the blood circulation. In one embodiment, sputum cells were stained with i) a CD45-specific cell marker, e.g., an antibody against CD45 (to identify white blood cells), ii) a CD206-specific cell marker, iii) a CD66b-specific cell marker, iv) a CD3-specific cell marker, and v) a CD19-specific cell marker. In one embodiment, any combination of i) through v) can be added to the sputum sample to further separate macrophages from other hematopoietic cells, e.g., an antibody cocktail composed of anti-CD66b, anti-CD3, and anti-CD19 compounds, e.g., where the compounds are antibodies or fragments thereof. In one embodiment, FVS510 was used as a viability dye to exclude dead cells. A portion of viable sputum cells specifically expresses CD45, as revealed by the anti-CD45PE signal. Cytospins of sorted CD45+ sputum cells confirmed their hematopoietic origin.
[0103] Further analysis of the sputum samples analyzed by FCM revealed that sputum-derived leukocytes (CD45 + It becomes clear that the sputum cells (cells) contain distinct subpopulations of macrophages. Cells selected by the size exclusion gate, live cell gate, and doublet discrimination gate show a representative light scatter profile of unstained single viable sputum cells, defining both the CD45+ and CD45- gates for sorting and further analysis. Cells that fell within the gate of single viable sputum cells were stained with a blood panel of antibodies for further analysis.
[0104] Sputum-derived leukocyte profiles of FVS510-CD45+ cells derived from different samples and stained with a blood panel of antibodies were further analyzed. Gates based on lineage-specific markers (fluorescence) and / or optical properties of specific cell types (FFS / SSC) were used to identify lymphocytes / granulocytes (Gate 1), as well as alveolar (Gate 2) and interstitial macrophages (Gate 3).
[0105] In this embodiment, fluorescence minus one (FMO) controls were analyzed using the same gates for leukocyte subpopulations defined by the blood panel of antibodies. Both FMO controls included the viability dyes CD45 and TCPP.
[0106] From leukocyte antibody panels Sputum cells stained with the CD66b, CD3, and CD19 antibodies were analyzed. Sputum cells stained with a leukocyte antibody panel excluding the CD206 antibody were also analyzed. Wright-Giemsa-stained cytospins from the sorted CD45+ Gate 2 and Gate 3 populations were examined under a microscope for visual identification and measurement. Cells in the Gate 2 cell population were approximately 20 μm in size, and cells in the Gate 3 population were approximately 10 μm in size.
[0107] Cell type was confirmed by a pathologist, as was cell size measurement of the macrophage populations sorted in Gate 2 and Gate 3. At least 100 cells were measured for each population. The mean cell size for Gate 2 was 16 μm + / - standard deviation approximately 3-5 μm (****p<0.0001).
[0108] We captured FCM profiles of CD45+ sputum cells labeled with anti-CD206 antibody and a cocktail of anti-CD66b, anti-CD3, and anti-CD19 antibodies (see Figure 12C). The isotype control exhibits higher background staining than the unstained or fluorescence minus-one (FMO) controls. Because the use of isotype control antibodies comes with its own set of problems, we used FMO controls to identify major sputum subpopulations. By comparing the CD66b / CD3 / CD19 cocktail FMO control with the stained sample containing all antibodies, we were able to set gate 1 to identify lymphocytes and granulocytes together. Similarly, by comparing the FMO control for the CD206 antibody, we were able to identify two populations of CD206-positive cells.
[0109] After sorting cells from gates 2 and 3, cytological analysis revealed cell populations with morphology consistent with that of macrophages. However, cells sorted from gate 3 were significantly smaller in size compared to cells sorted from gate 2. Our calculated sizes for the alveolar macrophage (gate 2) and interstitial macrophage (gate 3) populations are consistent with previously reported size ranges.
[0110] Alveolar macrophages are identified as being strongly positive for CD206 and autofluorescence in FITC channel gate 2, while interstitial lung macrophages are smaller in size and have lower CD206 expression in gate 3.
[0111] The mean background staining for the CD206 FMO control was 0.0023% (+ / - SD 0.0021%) for both gates combined. A positivity threshold based on two standard deviations (SD) above the mean background staining resulted in a threshold of 0.0065%, or approximately 6 macrophages per 100,000 cells, for both gates combined. Due to concerns that a lower threshold would not fall within the linear detection range of the PE-CF594 fluorochrome, an arbitrary threshold of 0.05% was instead chosen, which included alveolar and interstitial macrophages. This threshold could not be based solely on interstitial macrophages. The 0.05% threshold was well within the linear range of detection of the flow cytometer, and this threshold meets the Papanicolaou Institute criteria for a "high number of macrophages" for an adequate sample.
[0112] 179 samples were analyzed for macrophage content. Based on the criteria outlined above, 15 samples were found to have inadequate macrophage counts. However, 6 of these samples (3.4%) had fewer than 1000 CD45+ events for analysis, too few cells for adequate analysis based on the threshold established in accordance with one embodiment of the present invention. Five of these 6 samples had a total sputum cell count of 1.5 x 10 before antibody staining. 6 The remaining 9 samples (5.0%) had more than 10,000 CD45+ cells (range 11648 to 463382), all of which were 1.7 × 10 6 Furthermore, of the 164 eligible samples, more than 1.5 × 10 cells were present at the start of the antibody staining process. 6 Only four had fewer than 10,000 CD45+ cells. All of these samples contained sufficient numbers of macrophages, but three of the four had fewer than 10,000 CD45+ cells (with a range of 1327 to 2908). This data is consistent with a mean cell count of 1.5 × 10 6 This suggests that sputum samples containing less than 1 x 10 sputum are too few for reliable diagnostic flow cytometry analysis. In one embodiment of the present invention, the sputum sample used in one method contains less than 1 x 10 sputum. 5More than 1 x 10 6 More than 10 or 2 x 10 6 More than 10 or 10 x 10 6 Cytology slides typically contain more than 3 × 10 cells. 5 The cells remain on the slide and therefore cannot accommodate the number of cells from a sputum sample required to characterize the cell types and characteristics of the cells as detailed in the methods of the present invention.
[0113] The total number of sputum cells (excluding SECs) was calculated for each individual sample before antibody labeling. Overall, all eligible samples (n=164) were found to contain >0.05% macrophages (alveolar and interstitial combined). Of the eligible samples, 18 contained >50 million cells. The median number of cells in eligible samples was 14.6 x 10 6 In unsuitable samples (n=15), either no alveolar macrophages were observed or the combined alveolar and interstitial macrophage gate events were less than 0.05%. The median number of cells in unsuitable samples was 6.9 × 10 6 Some of the unsuitable samples were "too sparse" (<1000 CD45+ events) for a reliable profile, whereas the remainder contained sufficient cells but did not meet the QC macrophage criteria to be considered suitable. The median cell count for the sparse samples was 1.1 × 10 6 It is a cell.
[0114] Further analysis of 164 eligible sputum samples revealed differences between cancer ("CA") and non-cancer ("Non-CA") (also referred to as high-risk) individuals. This set included 32 samples from individuals diagnosed with lung cancer and 132 samples from high-risk individuals without cancer. The cancer group included 40.6% current smokers, and the high-risk group included 44.7% current smokers. There was no significant difference in pack-year smoking history between groups. There was also no significant difference in the average number of years since quitting smoking among former smokers. There were fewer women in the cancer group than in the high-risk group (21.9% vs. 54.5%, respectively). The mean age of participants in the cancer group was 69.8 years, compared with 64.8 years in the high-risk group (p<0.0002).
[0115] Referring now to FIG. 12B, in the first stage of the analysis, CD45 expression in each compartment was measured without the TCPP marker. + (Upper rectangle "+") vs. CD45 - The percentages of CD45 cells (bottom box "-") and various subpopulations were examined. Sputum samples from cancer patients were significantly higher in CD45 cells than in sputum from non-cancer ("non-CA") / high-risk patients without the disease. + A significantly higher proportion of CD45 cells was found in both samples (49.64% vs. 38.95%; p = 0.0099). + Although distinct subpopulations of the compartment were recognizable, the relative contribution of each population varied between samples and groups. + Comparing the relative sizes of cell subpopulations between cancer and high-risk samples revealed that cancer samples contained significantly more granulocytes / lymphocytes (see Figure 12C, gate 1 p=0.0378) and interstitial macrophages (see Figure 12C, gate 3 p=0.0031), as well as CD45 in gate 2 of Figure 12C. + A subpopulation of cells was found to be alveolar macrophages and CD206 positive.
[0116] CD45 is identified in the box indicated by "-" in Figure 12B. - The cell population identified in the compartment contains cells of epithelial origin, which is indicated by CD45 -The cells were sorted and their morphology was confirmed by the presence of goblet cells and ciliated epithelial cells when visualized on cytospins. Using antibodies against EpCAM and cytokeratin, we were able to detect the CD45 expression of cells by flow cytometry. - The populations could be further delineated. FMO controls show low background for each antibody used. - The relative contribution of cell subpopulations varied from sample to sample, and no significant differences were observed between the cancer and high-risk groups. + EpCAM + CD45 of cells that were positive for - Subpopulations were identified.
[0117] Viable single CD45 from different samples - Sputum cells were stained with epithelial cell markers, e.g., PanCK and EpCAM (epithelial profile) / epithelial antibody panel. Fluorescence minus one (FMO) controls of the profile were obtained by FCM. FMO controls included viability dyes, CD45, and TCPP. FVS510-CD45 stained with isotype controls of the antibodies used. - Sputum-derived epithelial profiles of cells were obtained (unstained sputum cells). FMO control FVS510-CD45 stained with EpCAM antibody but without panCK antibody profile. - Cells were obtained from FMO control FVS510-CD45 stained with panCK but without EpCAM antibody. - Cells were obtained.
[0118] In the second stage of FCM analysis, TCPP fluorescence was examined. Based on the TCPP staining intensity, single live cells were divided into three cell populations: TCPP-high cells, TCPP-intermediate (IM) cells, and TCPP-low cells (see Figures 13A and 13B). No difference in the relative proportions of these cell populations was observed between the high-risk and cancer groups. Next, CD45 + Leukocyte populations and CD45- The content of the epithelial cell population was further investigated (see "+" and "-" compartments in Figures 14B, 14F and 14J).
[0119] The CD45+ compartment of TCPP-high cells (see the "+" cell population in Figure 14B, further analyzed in Figure 14C) contains alveolar macrophages (CD45 + ;CD206 ++ cells) were abundant (see Figure 14C), and CD45 - The compartment (see the "-" cell population in Figure 14B, which is further analyzed in Figure 14D) contained EpCAM + ;panCK + Double-positive cells are abundant (upper right panel). TCPP-intermediate cells account for the majority of sputum cells, and therefore the profile of this subpopulation resembles that of the whole sample (see Figures 14E-14H). TCPP-low cells exhibit relatively low light scattering characteristics compared to TCPP-high cells (see Figure 14I), and are predominantly CD45 - (See Figure 14J) and do not express the epithelial markers EpCAM or panCK (See Figure 14L). 中間 TCPP cells make up the majority of sputum cells, and therefore the profile of this subpopulation is similar to that of the whole sample (Figures 14E-H). 高 Cells and TCPP 低 The cells were collected from the whole sample (or TCPP 中間 TCPP shows a different profile when compared to other soluble forms of soluble phospholipids (e.g., soluble phospholipids). 高 The cells exhibited a wide range of light scattering profiles (see Figure 14A), and TCPP 高 Cellular CD45 + compartment (see Figure 14C) contains CD45 + ;CD206 ++ Cells (i.e., alveolar macrophages) were abundant, and TCPP 高 Cellular CD45 - The compartment contains EpCAM + ;panCK +Abundant double-positive cells are observed (see Figure 14D). 低 The cells were treated with TCPP 高 Cells and TCPP 中間 They exhibit relatively low light scattering properties compared to cells (see Figure 14I), and are predominantly CD45 - (See FIG. 14J) and do not express the epithelial markers EpCAM or panCK (See FIG. 14L).
[0120] Based on gating for different cell lineage markers by FCM, we identified sputum cell populations with different TCPP fluorescence intensities. Using dot plot analysis showing TCPP fluorescence versus FITC / Alexa488 fluorescence (i.e., CD66b / CD3 / CD19 in the "blood tube" (see Figure 13A) and pan-cytokeratin (CK) in the "epithelial tube"), we defined a TCPP-high cutoff (identified by the upper bold rectangle / zone marked "H"). While a dot plot of TCPP fluorescence versus PE-CF594 fluorescence could also be used for this purpose, the former method more easily identifies cells with the highest TCPP FI.
[0121] Referring now to Figure 13C, the TCPP-high ("H") cutoff is obtained from a gate placed on a dot plot of sputum cells with TCPP on the y-axis and CD66b / CD3 / CD19 on the x-axis. The TCPP-low ("L") population is defined as the intersection point of the unstained sputum and TCPP-stained samples superimposed. The TCPP-intermediate ("IM") population is defined as the population between the TCPP-high and TCPP-low populations.
[0122] Several significant differences were observed in the characteristics specific to the TCPP-high population between the high-risk and cancer groups. First, TCPP-high cells from cancer samples had lower side scatter values than those from the high-risk group (see Figure 15A). Second, the CD45 expression levels of the TCPP-high population were significantly higher than those of the cancer group. - The compartment contains EpCAM + panCK +In addition, this double-positive population from cancer samples contained a higher proportion of cells in the same quadrant from high-risk samples (see Figure 15B). Furthermore, this double-positive population from cancer samples expressed higher levels of EpCAM, but not panCK, compared to cells in the same quadrant from high-risk samples (see Figure 15C).
[0123] The TCPP-high population in cancer samples had a smaller SSC than the TCPP-high population in high-risk samples (**p<0.01), confirming the difference in sputum cell characteristics between cancer and high-risk sputum samples. - Fractionation of EpCAM + panCK + The percentage of cells in high-risk samples was compared with the corresponding CD45 - The TCPP-high CD45-EpCAM fraction was significantly higher than that of the EpCAM-high CD45- + panCK + The mean fluorescence intensity (MFI) of EpCAM cells was higher in cancer samples than in the corresponding cell populations of high-risk samples (*p<0.05).
[0124] Further analysis of the cancer groups revealed that the mean EpCAM fluorescence intensity was significantly higher in early-stage cancer samples (stage I / II) compared with late-stage cancer samples (stage III / IV) (p=0.047). No significant differences were identified based on cancer type (squamous cell carcinoma vs. adenocarcinoma) or smoking history (current vs. former smokers). Interestingly, when high-risk smokers were separated based on smoking history, the profile of high-risk current smokers was higher compared with former smokers. + ;panCK + There were significantly more abundant cells (p=0.0008) and macrophage populations, including both alveolar (p=<0.0001) and interstitial macrophages (p=0.0141).
[0125] Referring now to Figure 16, significant differences between cancer (CA) and non-cancer (non-CA) samples from the blood cell populations described in Figures 12A-12C are shown. Each dot (CA) and square (non-CA) represents one sample. Figure 16A shows the CD45 expression level in sputum samples from cancer (CA) samples. + Figure 16B shows that the proportion of CD45 cells was significantly higher than that in non-cancer samples (**p=0.0099). + Among the cells, the granulocyte / lymphocyte subpopulation (Gate 1 in Figure 12C) was significantly greater in sputum samples from cancer patients compared to non-cancer patients (*p=0.0378). Figure 16C shows that CD45 expression in interstitial macrophages was significantly greater in sputum samples from cancer patients compared to non-cancer patients. + The subpopulation (Gate 3 in Figure 12C) also shows that sputum samples from cancer patients are significantly larger than those from non-cancer patients (**p=0.0031). The thick black horizontal bars indicate the median values for each sample group.
[0126] Current methodologies used for sputum analysis have challenges that limit their clinical use. Sputum cytology suffers from low sensitivity due to the sophisticated techniques required to identify subtle nuclear changes. The need to screen a large number of slides is time-consuming, hindering its clinical use. While imaging and molecular techniques can assess genetic alterations in sputum-derived cells, screening methods that detect genetic abnormalities based on nuclear ploidy or in situ hybridization only use a few hundred cells per sputum sample, and microchip analysis of enriched epithelial cells obtained from sputum analysis for genetic abnormalities only screens 2,000 cells per slide. Excluding a large proportion of sputum cells from analysis can mask important disease parameters and reduce sensitivity to the point of clinical usefulness. The limitations of these various techniques should not be confused with the extremely useful properties of sputum as a biological fluid, which can provide an important cellular snapshot of the lung environment.
[0127] Flow cytometry platforms are well suited to analyzing exfoliated cells isolated from sputum to identify tumor-associated alterations in leukocyte and nonleukocyte populations that would otherwise go undetected by traditional cytological methods. FCM's ability to detect and analyze cells based on physical properties (i.e., size and granularity) and cell surface molecules is significant. Unlike microscopy or cytology, flow cytometry can analyze large numbers of cells in a short time. Variability in autofluorescence and nonspecific binding properties of cell populations within and between sputum samples precludes the use of commercially available biological controls commonly used for immunophenotyping of well-characterized hematopoietic populations. For this reason, a positive threshold for the macrophage gate has been established using an internal FMO control. The ability to identify alveolar macrophages as a distinct leukocyte subpopulation allowed us to include built-in flow cytometry quality control parameters to determine the pulmonary origin of each sputum sample. To ensure quality control according to one embodiment of the present invention, sample quality needed to be confirmed based on cytology.
[0128] The lungs are constantly exposed to pathogens and harmful particulates. Alveolar macrophages are the primary innate defense for maintaining a healthy lung environment. Alveolar macrophages express CD206 (CD206 ++ ) and their autofluorescence gives clear CD45 signals on the granulocyte / lymphocyte axis, ranging from intermediate to high. + Furthermore, the results show that The light scattering profile of alveolar macrophages This confirmed previous observations of overlap with contaminating SECs, highlighting the need to isolate SECs from further analysis.
[0129] CD - 206 intermediate to positive cells (CD206 + ) are also macrophages, but they express alveolar CD206 ++They are smaller than macrophages and have minimal FITC autofluorescence, indicating that this macrophage population may represent interstitial macrophages. Although interstitial macrophages (in contrast to alveolar macrophages) are not normally in contact with the airway lumen, the proinflammatory environment caused by chronic smoking is ideal for interstitial macrophages to penetrate into the airways. Therefore, the presence of interstitial macrophages in sputum obtained from heavy smokers is not surprising. This is further supported by our finding that high-risk current smokers have significantly more macrophages in their sputum than high-risk former smokers.
[0130] In one embodiment of the present invention, the minimum number of sputum-derived cells in a sputum sample to obtain an appropriate profile in automated FCM to determine the presence of macrophages was approximately 1.5 million. The cutoff for determining sample suitability, 5 macrophages per 10,000 cells (0.05%), was well within the detection range of the flow cytometer. The 0.05% macrophage cutoff for sample suitability included interstitial macrophages, because their combined presence is a cell population specific to lung tissue. The rare presence of interstitial macrophages without the absence of alveolar macrophages makes biological interpretation difficult, so samples without any alveolar macrophages were considered unsuitable.
[0131] A comparative multiparameter analysis of sputum samples from individuals with confirmed lung cancer versus those at high risk of developing the disease revealed significant differences between the two groups: cancer samples contained significantly more CD45+ cells, specifically granulocytes / lymphocytes and interstitial macrophages, than did high-risk samples.
[0132] By adding the porphyrin TCPP to the staining protocol, we were able to identify several significant differences in the most brightly stained subpopulation (TCPP-high) between the cancer and high-risk groups. TCPP-high cells from the cancer group, regardless of their CD45 lineage, had lower side scatter characteristics than TCPP-high cells from the high-risk group, suggesting the reduced cytoplasmic content, organelle degranulation, and vacuolization reported in malignant tumors.
[0133] TCPP-rich non-leukocyte (CD45 - Analysis of the subpopulations revealed that the cancer group contained a greater proportion of cells staining for the epithelial markers panCK and EpCAM. This difference from the high-risk group was primarily due to significantly fewer of these cells in the sputum of high-risk former smokers compared with current smokers. Epithelial cell subpopulations from the cancer group also expressed higher levels of EpCAM, but similar levels of panCK. This was most pronounced in the stage I / II subgroup.
[0134] Detection of epithelial-derived cancers and circulating tumor cells has traditionally relied on detecting both EpCAM and cytokeratin expression. Our flow cytometry-based analysis, which confirmed increased expression of EpCAM in samples from confirmed stage I / II cancers and from high-risk participants who continued to smoke, suggests that EpCAM expression may be particularly important in early lung cancer detection.
[0135] According to one embodiment of the lung test, analysis of light scattering and fluorescence signals from viable single cells identified by automated FCM is determined. Logistic regression models the relationship between predictor variables and a categorical (in our case, binary cancer / non-cancer) response variable. Stepwise regression is a supervised machine learning process that adds and removes potential predictor variables and examines the fit of the resulting model. Clinical factors for which complete data were available (Table 1) were included as potential predictors. During the forward and backward stepwise regression process, age was the clinical parameter that repeatedly scored as significant.
[0136] The lung test's performance was evaluated on 122 high-risk samples (also referred to herein as non-cancer NC) and 28 cancer samples listed in Table 1, as well as 32 additional samples processed on a different FCM instrument (Navios EX) (Table 2). These 32 samples represent a different set of patients than those used for test development using the LSRII cytometer. The same model with the same coefficients was used for both instruments, but the cutoff for the Navios samples was 0.5 instead of 0.28. The results, shown in Table 3, indicate that the lung test performed extremely well, with sensitivity, specificity, and accuracy all exceeding 80% for the LSRII samples and very similar figures for the smaller set of Navios EX samples. Both platforms achieved highly reliable negative predictive values (NPVs) of over 95%. Table 2. Patient characteristics of Navios EX validation samples [Table 2]
[0137] Table 3. Lung test performance [Table 3]
[0138] The lung test also performed remarkably well when nodules ≥20 mm in diameter were not detected by LDCT, with a sensitivity of 92%, a specificity of 87%, and an area under the receiver operating characteristic curve of 94% (Table 3, "All nodules <20 mm"). Furthermore, the lung test performed well for all tumor types shown and for all disease stages, including I and II (Tables 4, 5).
[0139] Table 4. Performance of lung tests by tumor type and stage (LSRII) [Table 4]
[0140] Table 5. Performance of lung tests by tumor type and stage (Navios EX) [Table 5]
[0141] Each retained predictor contributed significantly to the model (Wald test p-value < 0.05), and their individual removal negatively affected the ability to accurately classify cancer and high-risk samples (Table 6). Age is a well-established clinical correlate of lung cancer. 31 This is also the case with our model, but nevertheless the correlation between age and model value was not robust in either the LSRII or Navios EX samples (Figure 8), with some younger patients being classified as "cancer" and many older patients being classified as "non-cancer." Indeed, CD206 低 CD3 / CD19CD66b 中間 The number of misclassified samples due to signal exclusion was significantly reduced by age and FVS510-A / log 10 The results were comparable to those of the exclusion of the interaction with FSC-A R2 (Table 6).
[0142] Table 6. Impact of model predictors on classification [Table 6]
[0143] Table 7. Reagents used for sputum staining and flow cytometry analysis according to one embodiment of the present invention. [Table 7] Both reagents were titrated using sputum from individuals at high risk of developing lung cancer.
[0144] One aspect of the present invention provides an automated flow cytometry system and method for machine learning-based analysis to predict the presence of lung cancer from sputum samples. One non-limiting hypothesis is that sputum as diagnostic material provides a snapshot of the tumor itself, its microenvironment (ME), and its field of cancerization (FoC). While expert cytological analysis of sputum can detect cancerous and precancerous cells, this analysis method is labor-intensive and not well suited for large-scale, non-automated screening due to the limited number of cells analyzed due to the slide size of the cytology sample. It is prone to observer bias and cannot examine a large number of cells from a sample in a few seconds. Automated image processing has been used to capture malignancy-associated changes in cells and has met with some success, but is still hampered by its complex technology and the limited number of cells analyzed.
[0145] Another aspect of the present invention provides a system and method for analyzing biological samples, such as sputum, for the presence of cancer cells using a high-throughput, automated flow cytometry-based approach combined with machine learning, offering the following advantages: a) the test can be used in routine laboratories without requiring expert evaluation of the sample or being subject to operator bias; b) the entire sputum sample can be rapidly analyzed; and c) numerical analysis can capture the complex interactions between lung cancer, ME, and FoC cells, which may be difficult to reliably detect in individuals. During the development of the lung test, for example, it was unexpectedly discovered that the predictive value of viability staining density suggests some relationship to apoptosis. Furthermore, it was unexpectedly observed that specific markers of immune function would be beneficial.
[0146] One aspect of the present invention provides an automated flow cytometry-based test that investigates three aspects of tumorigenesis: TCPP staining, programmed cell death, and immune response. Other researchers have shown that combining different types of measurements, such as cytology with genetic mutations or microRNA and methylation biomarkers, significantly improves the performance of sputum-based tests for early lung cancer detection. While we used the same technology platform to measure different cancer-related processes, these additional parameters may contribute to the performance improvement from slide-based tests to flow cytometry-based tests (Figure 7). Furthermore, flow cytometry tests read the entire sample, which was also predicted to improve test performance.
[0147] All but one study participant met the criteria for lung cancer screening most recently published by the US Preventive Services Task Force. Our study group can be considered a sample of individuals eligible for lung cancer screening (one of the lung's target populations), but sampling was small and underrepresented, including minorities, including women in the cancer group. Furthermore, the prevalence of cancer in our study was slightly below 19% in both datasets, a figure significantly higher than that of lung cancer screening populations or patients with lung nodules between 7 and 19 mm (the other target populations for lung testing).
[0148] In its 2017 official policy statement, the American Thoracic Society stated that for a molecular biomarker to be considered clinically useful, it should influence clinical management decisions to improve clinical outcomes. The authors discuss the use case of expanding screening to include participants currently ineligible for LDCT screening, reducing the cancer prevalence from the NLST level of 1 / 120 to a hypothetical 1 / 500. Based on the NLST data, they assume a plausible harm threshold of 0.83%, resulting in a minimum positive diagnostic likelihood ratio (PDLR) of 4.18, a level met by the larger LSRII group (Table 3). Using a hypothetical prevalence of 1 / 400 instead of 1 / 500 at the same harm threshold would result in a PDLR of 3.35, met by both the LSRII and Navios groups. It is believed that the lung testing systems and methods disclosed herein may help expand early lung cancer screening to relatively underserved populations, such as young female and male African American smokers.
[0149] Lung testing could also be potentially combined with a risk calculator (e.g., Brock University's Lung Cancer Risk Calculator) to aid clinical decision-making in LDCT-positive patients with intermediate-sized nodules. Although a pan-Canadian study found that the largest nodules were nonmalignant in 20% of participants, only 2% of NLST patients underwent invasive follow-up at sizes smaller than 7 mm, and those larger than 20 mm may erroneously prompt immediate follow-up. However, intermediate-sized nodules are notoriously troublesome. If we estimate the risk threshold (R), i.e., the threshold above which invasive follow-up is considered worthwhile, to be the frequency of cancer in the NLST population with nodules 7–19 mm in diameter (4.8%) and assume a cancer prevalence of 3.8% in the LDCT-positive population, then the sensitivity / (1–specificity) must be equal to or greater than [(1–prevalence) / prevalence] × R / (1–R) = [(1–0.038) / 0.038] × 0.048 / (1–0.48) = 1.28, a threshold our test comfortably meets. 44 (Table 3, PDLR).
[0150] One aspect of the lung test according to one embodiment of the present invention is a non-invasive sputum-based test for the detection of early-stage lung cancer. This test uses a flow cytometry platform to analyze the cellular content of sputum, and the analysis is fully automated and therefore unbiased. Test performance in cases with small nodules (<20 mm) was 92% sensitive and 87% specific.
[0151] When a range of values is described, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range, and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, unless specifically excluded in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0152] Certain ranges are described herein by the use of the term "about" before the numerical values. The term "about" is used herein to literally support the exact number described thereafter, as well as a number that is close to or approximately the number described thereafter. When determining whether a number is close to or approximately a specifically recited number, the unrecited close or approximate number may be a number that, in the context in which it is presented, provides a substantial equivalent to the specifically recited number.
[0153] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative exemplary methods and materials are described herein.
[0154] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a foundation prior to the use of exclusive terminology such as "solely," "only," or the use of a "negative" limitation in connection with the recitation of claim elements.
[0155] As will be apparent to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has distinct components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible.
[0156] Although the systems and methods have been or will be described with functional descriptions, taking into account grammatical flow, it is expressly understood that the claims should not be construed as necessarily limited by "means" or "step" limitation constructions unless expressly formulated under 35 U.S.C. § 112; rather, the meaning of the definitions provided by the claims and their equivalents should be given their full scope under the doctrine of judicial equivalents, and that if the claims are expressly formulated under 35 U.S.C. § 112, they should be given their full statutory equivalents under 35 U.S.C. § 112. How to classify flow cytometer data
[0157] As described above, aspects of the present invention include methods for classifying flow cytometer data. "Flow cytometer data" refers to information about the characteristics of sample particles (e.g., beads, cells, or debris) collected by any number of detectors in a particle analyzer. As used herein, a "particle analyzer" refers to an analytical tool (e.g., a flow cytometer) that enables characterization of particles based on specific parameters (e.g., optical parameters). "Particle" refers to a discrete component of a biological sample, such as a molecule, an analyte-bound bead, or an individual cell.
[0158] Methods of interest include classifying one or more population clusters based on determined parameters (e.g., fluorescence) of events (e.g., particles) in a sample. As used herein, a "population" or "subpopulation" of events, such as cells or other particles, generally refers to a group of events that have characteristics (e.g., optical, impedance, or temporal characteristics) related to one or more measured parameters such that the measured parameter data form a cluster in data space. Data obtained from analyzing cells (or other particles) by flow cytometry is often multidimensional, with each cell corresponding to a point in the multidimensional space defined by the measured parameters. In various embodiments, the data is composed of signals from multiple different parameters, e.g., two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, etc. Accordingly, populations are recognized as clusters in the data. Conversely, each data cluster is generally interpreted as corresponding to a population of a particular type of cell or particle, although clusters corresponding to noise or background are also typically observed. Clusters may be defined within a subset of dimensions, for example, with respect to a subset of measurement parameters (eg, fluorescent dyes) corresponding to populations that differ only in a subset of measurement parameters or features extracted from the sample measurements.
[0159] Aspects of the methods of the present invention include receiving a first gate having a defined boundary. As described herein, a "gate" generally refers to a boundary of a classifier that identifies a subset of data of interest (the data representing features or characteristics of particles / cells in a sample). In cytometry, a gate can be the boundary of a group (i.e., a population) of events of particular interest. In other words, a gate defines a boundary for classifying a population of flow cytometry data. In various embodiments, a gate identifies flow cytometry events that exhibit the same or similar set of parameters. An event is a cell or particle detected by a sensor as it passes between the sensor and the interrogation light source of the flow cytometer. The optical features or characteristics of the event detected by the detector / sensor can be analyzed for each event. Flow cytometry data analysis is built on the principle of gating. Gates and regions are placed around populations of events that exhibit common characteristics, typically forward scatter (FSC), side scatter (SSC), and / or expression of cell surface or intercellular markers, to further examine and quantify these populations. Gating refers to the selection of a sequential subpopulation of cells for analysis by flow cytometry and is the process of isolating a specific population of interest within a heterogeneous sample. This allows the light scattering (FSC and SSC) and fluorescence characteristics of the population of interest to be highlighted across the available dot plots, increasing the specificity of the analysis.
[0160] In some embodiments, the first gate is a gate drawn by a trained algorithm. In such embodiments, the trained algorithm can demarcate a region (e.g., in two-dimensional space) within which a particular classification can be assigned to the flow cytometer data. For example, drawing the first gate can include overlaying a polygon on a two-dimensional plot representing the flow cytometer data. For example, the first gate can be received from a database of gates used in previous attempts to classify flow cytometer data.
[0161] In various embodiments, the method includes receiving flow cytometer data, calculating parameters for each population, and gating the populations for further analysis based on the target population of interest. For example, an experiment may include particles / cells labeled with several fluorophores or fluorescently labeled antibodies, and groups of particles may be defined by populations corresponding to one or more fluorescence measurements. In this example, a first group may be defined by a specific range of light scatter for the first fluorophore, and a second group may be defined by a specific range of light scatter for a selected population from the first group. A third group may be defined by a third fluorophore based on one or more selected populations from the first group, the second group, or a combination thereof.
[0162] The flow cytometer data may be received from any suitable source. In some embodiments, the flow cytometer data is received from a storage device's memory. In such embodiments, the flow cytometer data may have been previously generated and stored in the storage device's memory for later redetection and analysis. In other embodiments, the flow cytometer data is received in real time. In other words, flow cytometer data generated during the operation of the flow cytometer may subsequently (e.g., immediately) be entered into a first gated data space (e.g., a two-dimensional plot). In some cases, the flow cytometer may be operated to generate data until a recording criterion is met. As used herein, a "recording criterion" refers to a condition that, when met, terminates the operation of the flow cytometer and data collection. Any suitable recording criterion may be used. In certain cases, the recording criterion is a time limit. When the recording criterion is a time limit, flow cytometer data collection stops after a specified amount of time (e.g., ranging from a few seconds to three hours). In other cases, the recording criterion is a total number of events. In such cases, flow cytometer data collection stops after a certain number of particles (e.g., user-defined) have been analyzed. In yet other cases, the recording criterion is the number of events in a population. In such cases, flow cytometer data collection may stop after a certain number of particles (e.g., user-defined) in a particular population (e.g., a population exhibiting a particular phenotype) have been analyzed.
[0163] In certain embodiments, particles are detected and uniquely identified by exposing them to excitation light and measuring the fluorescence of each particle in one or more detection channels as desired. The fluorescence emitted in the detection channels used to identify the particle and its associated binding complex may be measured after excitation by a single light source, or may be measured separately after excitation by separate light sources. When separate excitation light sources are used to excite particle labels, the labels may be selected so that all labels are excitable by each of the excitation light sources used.
[0164] In various embodiments, the flow cytometer data is received from a forward scatter detector. The forward scatter detector of interest provides information regarding the overall size of the particle. In various embodiments, the flow cytometer data is received from a side scatter detector. The side scatter detector of interest detects refracted and reflected light from the particle's surface and internal structure, which tends to increase as the complexity of the particle structure (e.g., particle size) increases. In various embodiments, the flow cytometer data is received from a fluorescence detector. The fluorescence detector of interest is configured to detect fluorescent emissions from fluorescent molecules associated with the particles in the flow cell, such as labeled specific binding members (e.g., labeled antibodies that specifically bind to a marker of interest). In certain embodiments, the method includes detecting fluorescence from the sample with one or more fluorescence detectors, including two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, etc., and twenty-five or more fluorescence detectors.
[0165] In certain embodiments, the method also includes data acquisition, analysis, and recording, such as by a computer, where multiple data channels record data from each detector for light scattering and fluorescence emitted by each particle as it passes through the sample interrogation region of the flow cytometer. In these embodiments, analysis includes classifying and counting particles so that each particle is represented by a set of digitized parameter values. The system can be configured to trigger on selected parameters to distinguish particles of interest from background and noise or cell populations that are not of interest. A "trigger" refers to a preset threshold for detecting a parameter and can be used as a means for detecting the passage of a particle through a light source. Detection of an event exceeding the selected parameter threshold activates the acquisition of light scattering and fluorescence data for the particle. Data is not acquired for particles or other components in the medium under test that cause a response below the threshold. The trigger parameter can be the detection of forward scattered light caused by a particle passing through the light beam. The flow cytometer then detects and collects the light scattering and fluorescence data for the particle. The recorded data for each particle is analyzed in real time or, if desired, stored in a data storage and analysis means, such as a computer.
[0166] In at least one embodiment, as will be readily appreciated by those skilled in the art, an apparatus according to the present invention includes a general-purpose or special-purpose computer or distributed system programmed with computer software to perform the above-described steps, which may be in any suitable computer language, including R, Python, C++, C#, Perl, Java, PHP, HTML, MySQL®, distributed programming languages, etc. The apparatus may also include a plurality of such computers / distributed systems (e.g., connected via the Internet and / or one or more intranets) in various hardware implementations. For example, data processing may be performed by a suitably programmed microprocessor, computing cloud, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), etc., along with appropriate memory, network, and bus elements. Alternatively, containers may be used.
[0167] Embodiments of the present invention provide a technology-based solution that overcomes existing problems for those who may have early-stage cancer, healthcare providers, insurance companies, and diagnostic laboratories in a professional manner using current state-of-the-art technology. One embodiment of the present invention is necessarily based on computer technologies, such as computer learning. Embodiments of the present invention offer significant advantages over current state-of-the-art technology, such as increased flexibility, faster results, non-invasive procedures, and automated screening of samples. For example, using flow cytometry and automated analysis of data, thousands of cells from a biological specimen can be analyzed and characterized with high sensitivity and specificity within a time frame of minutes to hours, a speed and accuracy that is not possible with analysis by a human observer to obtain the same data in the same amount of time.
[0168] The foregoing examples can be re-performed with equal success by substituting the reactants and / or operating conditions used in the foregoing examples with the reactants and / or operating conditions of the present invention as generally or specifically described.
[0169] In this specification and claims, "about" or "approximately" means within twenty percent (20%) of the cited numerical value. Any computer software disclosed herein may be embodied on any computer-readable medium (including combinations of media), including but not limited to CD-ROM, DVD-ROM, hard drive (local or network storage), USB key, other removable drive, ROM, virtual machine, software container (e.g., Docker), and firmware.
[0170] Although the invention has been described in detail with particular reference to these embodiments, other embodiments can be used with the same results. Variations and modifications of the present invention will be obvious to those skilled in the art, and it is intended to cover in the appended claims all such modifications and equivalents.
[0171] All publications and patents cited herein are incorporated by reference to the same extent as if each individual publication or patent was specifically and individually indicated to be incorporated by reference, and are incorporated by reference to disclose and describe the methods and / or materials for which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the publication dates provided may be different from the actual publication dates, which may require independent confirmation.
Claims
1. 1. A flow cytometric method for automated computer-assisted analysis of sputum samples from subjects suspected of having lung cancer, comprising: obtaining a plurality of cells from a sputum sample of a subject suspected of having lung cancer; marking the plurality of cells with i) a plurality of cell lineage-specific marker compositions, ii) a cell viability composition, and iii) a tetra(4-carboxyphenyl)porphyrin (TCPP) composition; analyzing the plurality of cells marked in i-iii with the flow cytometer to obtain a cell size-selected subpopulation from the plurality of cells based on an automatically selected bead size exclusion gate; automatically selecting, by the computer, a viable singlet cell population from the cell size-selected subpopulation using an automated non-debris gate and an automated singlet gate; automatically obtaining flow cytometer values from the viable singlet cell population based on the plurality of cell lineage-specific marker compositions, the viability marker, and the TCPP marker by the computer; automatically applying, by the computer, a trained classifier to the metadata obtained from the subject and the obtained flow cytometry values; automatically generating, by the computer, a classification of the sputum sample based on application of the trained classifier, wherein the classification is selected from a plurality of classification options including cancer and non-cancer; A method comprising:
2. 10. The method of claim 1, wherein the sputum sample is a single cell suspension.
3. 2. The method of claim 1, wherein the lineage-specific marker is selected from CD206, CD3, CD19, CD66b, CD45, EpCAM, PanCK, and any combination thereof.
4. The method of claim 1 , wherein the cell viability determining composition preferentially labels dead cells over live cells.
5. The method of claim 1, wherein the cell viability determining composition is FVS510.
6. The method of claim 1, wherein the analyzing step comprises obtaining flow cytometry values for side scatter, forward scatter, fluorescence from TCPP, fluorescence from the cell viability determination composition, and fluorescence from the plurality of cell lineage-specific marker compositions from the plurality of cells.
7. 2. The method of claim 1, wherein the plurality of cell lineage-specific marker compositions are selected from fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19, fluorescent anti-CD66b, fluorescent anti-CD45, fluorescent anti-pan-cytokeratin, and fluorescent anti-EpCAM, and any combination thereof.
8. 10. The method of claim 1, wherein the bead size exclusion gate is set between 5 μm and about 30 μm, and events less than about 5 μm and greater than about 30 μm are not further analyzed.
9. 10. The method of claim 1, wherein the automated non-debris gate excludes a majority of dead cells from the non-debris population.
10. 2. The method of claim 1, wherein the automated singlet gate is applied to the cell population selected with the automated non-debris gate.
11. The method of claim 7 , wherein the subject metadata includes age.
12. 2. The method of claim 1, wherein the sputum sample contains a minimum number of CD206-expressing cells that makes the sputum sample acceptable for assessing lung health.
13. the trained classifier: [Equation 1] where the coefficient b 0 ~b 5 is determined by fitting the trained classifier to a plurality of sputum samples used to build the classifier. The method of claim 11.
14. 1. A system for automated analysis of flow cytometry data, comprising: a computer processor in communication with a memory storing flow cytometry data from a plurality of markers in a plurality of cells from a sputum sample of a subject, the plurality of markers comprising: i) a plurality of cell lineage-specific marker compositions; ii) a cell viability composition; and iii) a tetra(4-carboxyphenyl)porphyrin (TCPP) composition; A computer program product embodied in a non-transitory computer-readable medium, the computer program product comprising: receiving flow cytometry data obtained from the plurality of cells from a sputum sample; selecting, from the plurality of cells in the sputum sample, an automatically selected subpopulation of cells based on application of an automated gate selected from a bead size exclusion gate, a viability gate, and a singlet gate; determining flow cytometry values for the plurality of lineage-specific marker compositions of interest, the viability markers, and the TCPP marker from the subpopulation; applying a classifier to the flow cytometry values of a subject of interest and the subject metadata; outputting, to a display device, an identification of one or more classifications of the sputum sample, including cancer or non-cancerous; a computer program product including instructions for automatically performing the steps of: Including, the system.
15. The system of claim 14 , wherein the cell viability determining composition preferentially labels dead cells over live cells.
16. 15. The system of claim 14, wherein the cell viability determination composition is FVS510 and the metadata of the subject is age.
17. The system of claim 14, wherein the flow cytometry values are obtained from the plurality of cells for side scatter, forward scatter, fluorescence from TCPP, fluorescence from the cell viability determination composition, and fluorescence from the plurality of cell lineage-specific marker compositions.
18. 17. The system of claim 16, wherein the plurality of lineage-specific marker compositions are selected from fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19, fluorescent anti-CD66b, fluorescent anti-CD45, fluorescent anti-EpCAM, and fluorescent anti-pan-cytokeratin.
19. 15. The system of claim 14, wherein the bead size exclusion gate is set to exclude events having a size less than about 5 μm and greater than about 30 μm.
20. 15. The system of claim 14, wherein the automated non-debris gate excludes a majority of dead cells from the non-debris population.
21. the trained classifier: [Equation 2] where the coefficient b 0 ~b 5 is determined by fitting the trained classifier to a plurality of sputum samples used to build the classifier.
20. The system of claim 18.
22. A non-transitory computer-readable medium, comprising: At runtime, the processing circuit automatically obtaining flow cytometer values for a viable singlet population of the subject's sputum sample based on side scatter, forward scatter, fluorescence from TCPP, fluorescence from the cell viability composition, and fluorescence from the plurality of cell lineage-specific marker compositions; automatically applying a trained classifier to the metadata obtained from the subject and the obtained flow cytometry values; automatically generating a classification of the sputum sample based on application of the trained classifier, the classification being selected from a plurality of classification options including cancer and non-cancer. A non-transitory computer-readable medium containing program code for causing a
23. 23. The non-transitory computer-readable medium of claim 22, wherein the cell viability determining composition preferentially labels dead cells over live cells.
24. 23. The non-transitory computer-readable medium of claim 22, wherein the cell viability determining composition is FVS510.
25. 25. The non-transitory computer-readable medium of claim 24, wherein the plurality of lineage-specific marker compositions are selected from fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19, fluorescent anti-CD66b, fluorescent anti-CD45, fluorescent anti-EpCAM, and fluorescent anti-pan-cytokeratin, or any combination thereof.
26. 23. The non-transitory computer readable medium of claim 22, wherein a minimum number of CD206 positive cells is present in the sputum sample analyzed.
27. 23. The non-transitory computer readable medium of claim 22, wherein the bead size exclusion gate is set between 5 μm and 30 μm.
28. 23. The non-transitory computer readable medium of claim 22, wherein the automated non-debris gate excludes a majority of dead cells from the non-debris population.
29. 26. The non-transitory computer-readable medium of claim 25, wherein the metadata of the subject is age.
30. the trained classifier: [Equation 3] where the coefficient b 0 ~b 5 is determined by fitting the trained classifier to a plurality of sputum samples used to build the classifier.
30. The non-transitory computer-readable medium of claim 29.