Detection of early lung cancer in sputum using automated flow cytometry and machine learning

The analysis of sputum samples through flow cytometry and machine learning technology solves the problems of low sensitivity and high false positive rates of lung cancer screening in the prior art, and achieves high-precision, non-invasive lung cancer screening, reducing unnecessary medical procedures and radiation exposure.

CN119948340APending Publication Date: 2025-05-06BIOAFFINITY TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380062360.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-07-20
Filing Date
2023-06-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology has low sensitivity and high false positive rates in lung cancer screening, resulting in unnecessary follow-up procedures for many cancer-free patients and high cost and radiation exposure for low-dose computed tomography (LDCT) screening.

Method used

The sputum samples are automatically analyzed by flow cytometry combined with machine learning. By labeling multiple cell lineage-specific marker compositions, cell vitality compositions and tetracarboxyphenyl)porphyrin (TCPP) compositions, they identify and classify cell subpopulations in sputum samples, thereby assisting doctors in making diagnostic decisions.

Benefits of technology

It improves the sensitivity and specificity of lung cancer screening, reduces unnecessary follow-up procedures, reduces the economic and radiation exposure burden of patients, and provides a high-precision, non-invasive, radiation-free testing method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948340A_ABST
    Figure CN119948340A_ABST
Patent Text Reader

Abstract

A system and method for analyzing a sputum sample from a subject suspected of having lung cancer comprising obtaining a plurality of cells from the sputum sample from the subject, labeling the plurality of cells with i) a plurality of cell lineage specific marker compositions, ii) a cell viability composition, and iii) a tetra (4-carboxyphenyl) porphyrin (TCPP) composition; analyzing the plurality of cells labeled with i-iii with a flow cytometer to obtain a subpopulation selected by cell size from the plurality of cells based on an automatically selected bead size exclusion gate; selecting a viable singlet cell population from the selected cell size subset using an automatic fragment-free gate and an automatic singlet gate; obtaining flow cytometry values from the viable singlet cell population based on the plurality of cell lineage specific marker compositions, viability markers, and TCPP markers; applying the trained classifier to the metadata from the subject and the obtained flow cytometry values; and generating a classification of the sputum sample based on the application of the trained classifier, where the classification is selected from a plurality of classification options including cancer and non-cancer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 357,994, filed on July 1, 2022, entitled “Sputum Analysis by Flow Cytometry; An Efficient Platform for Analyzing the Lung Environment” and U.S. Provisional Application No. 63 / 390,826, filed on July 20, 2022, entitled “Detection of Early Lung Cancer in Sputum Using Automated Flow Cytometry and Machine Learning”, the specifications and claims of which are incorporated herein by reference.

[0003] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0004] not applicable.

[0005] Incorporation by reference of information submitted on CD-ROM

[0006] not applicable.

[0007] Statement regarding prior publications by the inventor or co-inventors

[0008] not applicable.

[0009] Copyrighted Material

[0010] not applicable. Background Art

[0011] Please note that the following discussion cites a large number of publications by author and year of publication, and that some publications should not be considered prior art to the present invention due to their recent publication dates. The discussion of such publications herein is intended to provide a more complete background and should not be construed as an admission that such publications are prior art for purposes of determining patentability.

[0012] In 2020, lung cancer caused approximately 1.8 million deaths worldwide. It is estimated that in 2022, 130,180 people will die from lung cancer in the United States alone. Overall, the five-year survival rate for lung cancer remains low at 22.9% because most patients present with advanced disease. The National Lung Screening Trial (NLST) showed that LDCT screening detected 93.8% of lung cancers in a high-risk population (i.e., those aged 55-74 years, who smoked >30 pack-years, and who currently smoked or had quit within the past 15 years). Low-dose computed tomography (LDCT) is the standard of care for lung cancer screening in the United States (US). LDCT has a sensitivity of 93.8%, but a specificity of 73.4%, which results in potentially harmful follow-up procedures for patients without lung cancer. Therefore, there is a need for additional high-precision assays that can be used as an adjunct to LDCT for the diagnosis of lung cancer. When the nodules found are small, low-dose spiral computed tomography may not lead to a clear treatment path.

[0013] The NLST showed that LDCT screening resulted in a 20% overall reduction in lung cancer-specific mortality compared with screening with chest radiography. Unfortunately, in this trial, 96.4% of positive LDCT scans were false positives, resulting in approximately 90% of patients with positive LDCT undergoing additional procedures to determine whether the nodules seen on their LDCT scans were cancerous. These procedures, which include imaging, biopsy, and surgical resection, can result in serious adverse effects, including death. New guidelines for interpreting LDCT scans and models to estimate the probability that a nodule is cancerous have improved the false-positive rate (FPR). Despite this, only a small proportion of eligible patients undergo LDCT screening. Failure to communicate the benefits and potential harms of screening (whether due to lack of knowledge or time), costs associated with LDCT, lack of access to LDCT, and repeated radiation exposure from serial LDCT scans may contribute to low screening adoption.

[0014] A simple, non-invasive, radiation-free and cost-effective test that can help doctors make or rule out a lung cancer diagnosis with greater certainty, thereby reducing unnecessary follow-up procedures and increasing lung cancer screening. Sputum is an easily accessible body fluid that has long been part of lung cancer diagnosis. The PAP sputum cytology test, developed by Papanicolaou and optimized by Saccomanno, was the first lung cancer diagnostic method dating back to the 1960s. For this test, two sputum night smears are marked with PAP stain and then read by a pathologist specializing in lung cytology. The sensitivity of sputum cytology varies widely, but the specificity is high. A review of 16 published studies on sputum cytology, including more than 28,000 patients, showed that its sensitivity ranged from 42% to 97%, with an average sensitivity of 66%, while the specificity showed an average of 99%.

[0015] Sputum cytology has poor sensitivity, in part due to insufficient samples and analyzing only a small portion of the sample. Insufficient samples may be produced because the sample is saliva or because mucus / debris / red blood cells within the smear obscure the cellular components needed for accurate analysis. Over time, modifications to the original sputum cytology tests have improved their sensitivity. Nebulizers and auxiliary devices (such as throat clearing machines (acapella) and lung flutes) and patient compliance with correct instructions on how to produce lung sputum samples have been shown to improve patients' ability to produce sputum. Liquid cytology tests and automated slide preparation devices can reduce background contaminants in sputum smears and thereby improve the quality of the slides. Increasing the number of sample reads has been shown to increase the likelihood of finding abnormal cells indicative of lung cancer.

[0016] Porphyrins, such as TCPP, are currently used as diagnostic agents for bladder cancer and surgery to identify the margins of cancerous tissue. Using microscopy, we demonstrated that by labeling sputum cells with the fluorescent porphyrin TCPP, we could distinguish study participants with lung cancer from those without lung cancer with high accuracy using a slide-based assay (cytology-based approach) and a human grader. Cytology-based approaches are of limited utility because reading slides is time-consuming and requires highly specialized personnel. In addition, large amounts of debris and the presence of excessive squamous epithelial cells (SEC) or cheek cells often result in samples that are insufficient for diagnosis. Because slide-based assays are time-consuming, often precluding analysis of the entire sample and thus potentially missing important events, and because human graders introduce subjective bias (also known as operator bias) that is inconsistent between different human graders, alternative methods of analyzing sputum to determine the likelihood of cancer would be useful.

[0017] According to one embodiment of the present invention, the feasibility of analyzing whole sputum samples using a flow cytometry platform without clogging the instrument is demonstrated to identify significant differences between samples obtained from humans diagnosed with lung cancer and those never diagnosed with the disease.

[0018] Early detection of lung cancer through screening can improve survival and reduce morbidity. The United States and some parts of the United Kingdom now advocate annual low-dose computed tomography (LDCT) screening for high-risk individuals. Therefore, positive LDCT results require follow-up testing to determine whether the nodule is benign or malignant. These medical procedures have inherent risks of morbidity and mortality. 6 , and may impose a heavy burden on screening participants and their families, while the associated costs impose a heavy economic burden on patients and society.

[0019] Therefore, efforts have begun to develop non-invasive tests that can be used in conjunction with LDCT or as a stand-alone test to identify people who are at high risk for lung cancer and should undergo LDCT. In both cases, the goal of these tests is to eliminate unnecessary medical procedures in low-risk patients while identifying patients with lung cancer at an early stage. One material that is easily obtained from the lungs is sputum, which contains a variety of blood cells and exfoliated bronchial epithelial cells. 7 , including precancerous and malignant cells in lung cancer patients. We previously reported a slide-based assay capable of classifying cancer and non-cancer patients based on sputum stained with tetrakis(4-carboxyphenyl)porphyrin (TCPP). Despite an accuracy rate of 81%, reading the labeled slides is time-consuming, subject to observer bias, and may miss critical low-frequency events due to insufficient sampling. Sample analysis of sputum using a high-throughput method using automated flow cytometry (FCM) can compensate for the shortcomings of slide-based analysis.

[0020] The disclosed embodiments combine flow cytometry and machine learning to develop a sputum-based test that can assist physicians in making decisions in this setting. Summary of the invention

[0021] One embodiment of the present invention provides a flow cytometer method for analyzing sputum samples from subjects suspected of having lung cancer. Multiple cells, such as single cell suspensions of sputum samples, are obtained from sputum samples of subjects suspected of having lung cancer. The multiple cells are labeled with i) multiple cell lineage-specific marker compositions, ii) cell viability compositions, and iii) tetra(4-carboxyphenyl)porphyrin (TCPP) compositions. For example, i) includes at least 3, or at least 4, or at least 5, or at least 6 of CD206, CD3, CD19, CD66b, CD45, EpCAM, PanCK, and any combination thereof, and any combination can clearly exclude any one of CD206, CD3, CD19, CD66b, CD45, EpCAM, and PanCK. In a further example, ii) the cell viability composition preferentially marks dead cells rather than live cells, and may include FVS510. Multiple cells labeled with i-iii are analyzed using a flow cytometer to obtain a subpopulation selected by cell size from the multiple cells based on an automatically selected bead size exclusion gate. For example, the bead size exclusion gate is set between 5 μm and about 30 μm, wherein no further analysis is performed on events less than about 5 μm and greater than about 30 μm. For example, the analysis step includes obtaining side scatter, forward scatter, fluorescence from TCPP, fluorescence from cell viability compositions, and flow cytometry values ​​of fluorescence from the multiple cell lineage-specific marker compositions from the multiple cells. From the subpopulation selected by cell size, an automatic non-debris gate (e.g., an automatic non-debris gate excludes most dead cells from the non-debris population) and an automatic singlet gate (e.g., an automatic singlet gate is applied to the cell population selected in the automatic non-debris gate) are used to select a live singlet cell population. From the live singlet cell population, flow cytometer values ​​are obtained based on the multiple cell lineage-specific marker compositions, viability markers, and TCPP markers. A trained classifier is applied to metadata (e.g., age) from a subject and flow cytometry values ​​are obtained. Based on the application of a trained classifier, a classification of a sputum sample is generated, wherein the classification is selected from a plurality of classification options including cancer and non-cancer. For example, the trained classifier is

[0022]

[0023] Wherein the b0-b5 coefficients are determined by fitting the trained classifier to multiple sputum samples used to construct the classifier. It should be noted that CD66b / CD3 / CD19 are immune cell markers, where anti-CD66b binds to granulocytes, anti-CD3 binds to T cells, anti-CD19 binds to B cells, and anti-CD45 binds to blood cells.

[0024] In one embodiment, the plurality of cell lineage specific marker compositions include fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19, and fluorescent anti-CD66b. In another embodiment, the classifier is a straight line equation comprising coefficients b0-b5 determined by fitting the classifier model to a specific set of samples used to construct the classifier.

[0025] Another embodiment provides a system for automatically analyzing flow cytometry data, the system comprising a computer processor in communication with a memory, the memory storing flow cytometry data of multiple markers in multiple cells of a sputum sample from a subject, wherein the multiple markers include i) multiple cell lineage-specific marker compositions, ii) cell viability compositions, and iii) tetra(4-carboxyphenyl)porphyrin (TCPP) compositions. For example, i) includes at least 3, or at least 4, or at least 5, or at least 6 of CD206, CD3, CD19, CD66b, CD45, EpCAM, PanCK, and any combination thereof, and any combination can explicitly exclude any one of CD206, CD3, CD19, CD66b, CD45, EpCAM, and PanCK. In a further example, ii) the cell viability composition preferentially marks dead cells rather than live cells, and may include FVS510. The system further provides a computer program product embodied in a non-transitory computer-readable medium, the computer program product including instructions for causing a computer processor to perform the following operations. Flow cytometry data obtained from a plurality of cells from a sputum sample is received. A subpopulation of cells automatically selected based on application of an automatic gate is selected from the plurality of cells in the sputum sample, the automatic gate being selected from a bead size exclusion gate, a viability gate, and a singlet gate. For example, a bead size exclusion gate is set between 5 μm and about 30 μm, wherein events smaller than about 5 μm and larger than about 30 μm are no longer further analyzed, such as an automatic debris-free gate excluding most dead cells from a debris-free population, such as an automatic singlet gate being applied to a population of cells selected in an automatic debris-free gate. Target flow cytometry values ​​for the plurality of cell lineage-specific marker compositions, viability markers, and TCPP markers are determined from the subpopulation. A classifier is applied to the target flow cytometry values ​​and metadata of the subject, such as a trained classifier is

[0026]

[0027] where the b0-b5 coefficients are determined by fitting the trained classifier to the multiple sputum samples used to construct the classifier.

[0028] An output is generated on a display device wherein one or more classifications of the sputum sample including cancer or non-cancer are identified.

[0029] Another embodiment of the present invention provides a non-transitory computer-readable medium, which includes program code, which, when executed, causes the processing circuit to perform the following operations. Based on multiple cell lineage-specific marker compositions (e.g., the multiple cell lineage-specific marker compositions are selected from fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19 and fluorescent anti-CD66b), viability markers (e.g., FVS510, but not limited to this) and TCPP markers, flow cytometer values ​​(e.g., side scatter, forward scatter, fluorescence from TCPP, fluorescence from cell viability compositions, and fluorescence from the multiple cell lineage-specific marker compositions) of the viability singlet population of the subject's sputum cells are obtained. The trained classifier is applied to the metadata from the subject and the obtained flow cytometry values. Based on the application of the trained classifier, a classification of the sputum sample is generated, wherein the classification is selected from a plurality of classification options including cancer and non-cancer. The method of claim 1, wherein the cell viability composition preferentially marks dead cells rather than live cells. In one embodiment, a bead size exclusion gate is set to exclude events smaller than about 5 μm and larger than about 30 μm, and a subpopulation of cells having cell sizes not excluded from further analysis by the bead size exclusion gate is selected from the sputum sample, and a viable singlet population is selected from this population. In another embodiment, the viable singlet population is selected using an automatic non-debris gate that excludes most dead cells from the non-debris population for further analysis. In one embodiment, the trained classifier is

[0030]

[0031] where the b0-b5 coefficients are determined by fitting the trained classifier to the multiple sputum samples used to construct the classifier.

[0032] An aspect of one embodiment of the present invention provides a flow cytometry method, such as automated FCM, for analyzing sputum from a subject suspected of having lung cancer, wherein the method includes one or more of the following: 1) eliminating contaminants, including debris and squamous epithelial cells (SEC, a common contaminant in the oral cavity), from the analyzed sputum sample using a gating strategy, such as defined by a bead standard and a viability dye, to produce a singlet cell population; 2) including quality control parameters to detect alveolar macrophages in the sample to verify the lung origin of each sputum sample, 3) defining cells of interest for sample adequacy 1) determining a cutoff value for a population to provide a reliable analysis, 4) obtaining optical properties from sputum-derived cells labeled with one or more of leukocyte and / or epithelial cell lineage-specific markers (e.g., fluorescent-specific antibodies or fragments thereof) and TCPP to identify significant differences between samples obtained from persons diagnosed with lung cancer and samples obtained from persons never diagnosed with the disease, 5) obtaining metadata for the subject, such as age and / or years of smoking information, 6) applying a classifier based on properties selected from the outputs of items 1-5, and 7) determining whether the analyzed sputum sample is above or below a cancer likelihood value. If cancer or cancer likelihood is identified, the subject is subjected to further testing.

[0033] One embodiment of the present invention provides a computer-implemented method for classifying lung sputum samples of test subjects with lung cancer risk, including receiving data from test subjects on at least one processor. The at least one processor is used to evaluate the data using a classifier, the classifier is an electronic representation of a classification system, the classifier is trained using multiple electronically stored training data sets, each of the multiple training data sets represents a separate training data set, wherein each separate training data set represents data of an individual subject and a corresponding subject, and each training data set further comprises a determination of the characterization of lung cancer (if present) in the corresponding subject, wherein the classification system includes the identification of cancer or non-cancer of lung sputum samples. The at least one processor is used to evaluate the classification of test sputum samples from test subjects based on the evaluation step. In one embodiment, the data includes flow cytometry data, subject metadata, or a combination thereof. For example, flow cytometry data is obtained from a sputum sample labeled with multiple markers in multiple cells from a subject's sputum sample, wherein the multiple markers include i) multiple cell lineage-specific marker compositions, ii) cell viability compositions, and iii) tetra(4-carboxyphenyl)porphyrin (TCPP) compositions. The subject metadata includes one or more of gender, age, genetic information, biomarker data, smoking status, medical history, or a combination thereof. In a further embodiment, a non-transitory computer-readable medium storing an executable program includes instructions for executing a computer-implemented classification method. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate one or more embodiments of the present invention, and together with the description, are used to explain the principles of the present invention. The accompanying drawings are only used to illustrate one or more embodiments of the present invention and should not be interpreted as limiting the present invention. In the drawings:

[0035] Figure 1 A flow chart for identifying samples according to one embodiment of the present invention is shown.

[0036] Figure 2A and 2B Automatic gating of FCM data according to one embodiment of the present invention is shown. Figure 2A Shown are the beads size exclusion (BSE) gate parameters set for the entire sputum sample. Figure 2B The results show that the forward (FSC) and side (SSC) scatter areas were observed to be above the BSE-gate threshold setting of 2.5x10 5 The event is a dead cell. Figure 2C A dot plot of NIST beads from the automated flowClust analysis is shown to illustrate how it can be used to set the lower left threshold of the BSE-gate.

[0037] Figures 3A-3I Automatic gating of FCM data according to one embodiment of the present invention is shown. Figure 3A It is shown that in some samples, events appear in the lower right corner of the forward scatter-height (FSC-H) versus side scatter-height (SSC-H) plot ("debris"). Figure 3B Show Figure 3A The fragment recognition events in the WT have very low light SSC-A characteristics, which is very unusual for human cell populations. Figure 3C -I is a histogram with the Y-axis representing density and the X-axis representing fluorescence intensity. Figure 3C Shows Figure 3A The fragmentation events shown in were identified as live due to their lack of staining with a viability dye (FVS510), a dye that preferentially stains dead cells. Figure 3D Show Figure 3A The fragments identified in expressed low levels of TCPP when labeled with TCPP. Figure 3E Show Figure 3A Fragmentation events identified in were negative for CD45. Figure 3F Show Figure 3A The fragmentation events identified in the assay were negative for CD66b / CD3 / C19 when incubated with a compound (e.g., an antibody) that specifically recognizes CD66b, a compound (e.g., an antibody) that specifically recognizes CD3, and a compound (e.g., an antibody) that specifically recognizes CD19. Figure 3H Show Figure 3A The fragmentation events identified in the assay were negative for CD206 when incubated with a compound (eg, an antibody) that recognizes CD206. Figure 3G Show Figure 3A Fragmentation events identified in , when incubated with a compound that specifically recognizes Pan-CK, are negative for Pan-CK. Fig. 3I Show Figure 3A When incubated with a compound that specifically recognizes EpCAM, the fragmentation event identified in the EpCAM is negative for EpCAM. According to one embodiment of the present invention, Figure 3A Fragmentation events identified in the spectra were excluded to avoid false positive results.

[0038] Figures 4A-4E According to one embodiment of the present invention, Figure 3A Dot plot analysis of the fragment-free events identified in , and Figure 4F A histogram analysis of the debris-free events is shown.

[0039] Figure 5A -E and Figure 5G -H shows the Figure 2A and Figure 3A A dot plot of events identified in , where debris is excluded, and Fig. 5F It shows that Figure 2A and Figure 3A Histogram of identified events in , where debris is excluded.

[0040] Fig. 6A and 6B TCPP / log 10 Histogram with SSC on the x-axis and density on the y-axis showing cells stained with TCPP. Figure 6C and 6D Shown is the FVS510 / log 10 Histogram with FSC on the x-axis and density on the y-axis showing cells stained with FVS510 (viability dye). Fig. 6E -F shows a dot plot with CD206 on the x-axis and CD66b / CD3 / CD19 on the y-axis.

[0041] Figure 7 A flow chart is shown for data analysis of sputum-based pulmonary assays according to one embodiment of the present invention.

[0042] Fig. 8A Shown is a receiver operating characteristic (ROC) plot of the false positive rate versus the true positive rate calculated as the model response threshold is varied. Figure 8B is based on Fig. 8APlot of the model value (cancer likelihood) for the ROC curve in .

[0043] Figures 9A-9B Shown are correlation plots with age on the x-axis and model values ​​on the y-axis when samples were analyzed on an LSRII flow cytometer or a NaviosEX cytometer, respectively.

[0044] Fig.10 An exemplary system according to one embodiment of the system disclosed herein is shown.

[0045] Fig.11 An implementation of classifier development according to one embodiment of the present invention is shown.

[0046] Fig. 12A -B shows the analysis of cells selected by size exclusion gate, live cell gate and dual state discrimination gate using flow cytometry according to one embodiment of the present invention, wherein the cells in the sputum sample are PE unstained ( Fig. 12A ), or stained with PE fluorescently labeled anti-CD45 antibody ("CD45-PE") ( Fig. 12B ), where the above image would be CD45 + The cell population of stained cells is indicated by "+", and the CD45 - The position of cells is indicated as "-", where forward scatter "FSC" is the x-axis. The absence "-" and presence "+" of CD45 staining were determined on samples exposed to anti-CD45 antibody fluorescently labeled with PE, as shown by the "+" box (cells positive for CD45 staining; "CD45 + cells") and “-” box (cells not stained with CD45; “CD45 - The cut-off value for CD45 positivity was based on unstained samples.

[0047] Fig. 12C CD45 in sputum samples from a "blood tube" according to one embodiment of the present invention is shown. + Representative profile of cells, this sputum sample was stained with antibodies for recognizing granulocytes and lymphocytes (CD66b bound to granulocytes, CD3 bound to B cells, and CD19 bound to T cells; see gate 1) and antibodies for recognizing alveolar macrophages (CD206; see gate 2) and interstitial macrophages (CD206; see gate 3). As used herein, Fig. 12C The cells in the Fig.12D The cells in the blood are sometimes referred to herein as "non-blood cells."

[0048] Fig.12DThe present invention shows the use of antibodies against epithelial markers pan-cytokeratin (panCK) (“panCK antibody labeled with Alexa 488” on the “y” axis) and EpCAM (“EpCAM antibody labeled with Alexa 488” on the “x” axis) according to one embodiment of the present invention. - PE-CF-594 labeled EpCAM antibody ("EpCAM antibody") was used to stain CD45 - Representative profiles of cells. Gate 4 represents CD45 - A subpopulation of cells within a cell population that stains positively for both epithelial markers.

[0049] Fig.13A -B shows a dot plot showing TCPP versus CD66b / CD3 / CD19-FITC / Alexa488 (using "blood tube") according to one embodiment of the present invention. Fig.13A ) and dot plots showing TCPP versus panCK-Alexa488 (using “epithelial tubes”) ( Fig. 13B ). The upper box marked with an "H" indicates TCPP 高 Gate, used to define TCPP 高 Cutoff value.

[0050] Fig. 13C A histogram of TCPP fluorescence intensity versus relative cell # on the "y" axis is shown, according to one embodiment of the present invention, where TCPP 高 The cut-off value is taken from Fig.13A or Fig. 13B The TCPP gate is defined at the intersection when the unstained sputum histogram is overlaid with the TCPP-stained sample histogram. 低 Group. Group TCPP with moderate TCPP staining IM Defined as in TCPP 高 Swarm and TCPP 低 Groups among groups.

[0051] Fig.14A -J shows a subdivision of the TCPP staining level according to one embodiment of the present invention (eg Fig.13A -C) Sputum cell populations selected from live, singlet, sputum cells, which were further analyzed by "blood cell" markers and "epithelial cell" markers, as shown in Figure 12. Fig.14A -D shows TCPP 高 Analysis of cells. Fig.14E -H shows TCPP IM Analysis of cells. Fig.14I -L shows TCPP 低 Analysis of cells. Each row Fig.14A , Fig.14E , Fig.14IThe first profile shows the light scattering profiles of the individual TCPP subpopulations. Fig. 14B , Fig.14F , Fig.14J The second profile shows the distribution of CD45 staining or lack of CD45 staining of cells in each TCPP subpopulation. + part ("+") and in the third column Fig. 14C , Figure 14G , Figure 14K The corresponding profile in , which shows the distribution of CD66b / CD3 / CD19 staining versus CD206 staining. Fig.14D , Fig.14H , Figure 14L The corresponding CD45 - ("-") parts, pan-cytokeratin staining versus EpCAM staining is shown.

[0052] Fig.15A -C shows the differences between cancer (CA) and non-cancer (non-CA) samples according to one embodiment of the present invention, as determined by cell lineage marker analysis and from Fig.14A , Fig. 14B and Fig.14D TCPP 高 The size of the cell population was derived. Each dot (CA) and square (non-CA) represents one sample. Fig.15A TCPP in cancer samples is shown 高 The group showed a higher level of TCPP than non-cancer samples. 高 The SSC group was smaller (**p<0.01). Fig. 15B TCPP 高 CD45 - EpCAM + panCK + The proportion of cells was greater than that of the corresponding CD45 - The proportion in the part (**p<0.01). Fig. 15C TCPP in cancer samples 高 CD45 - EpCAM + panCK + The mean fluorescence intensity (MFI) of EpCAM in cells was higher than that of the corresponding cell subpopulations in non-cancer samples (*p<0.05). The thick black horizontal bar indicates the median value of each sample group.

[0053] Fig.16A -C shows an embodiment according to the present invention, Fig. 12ADifferences between cancer (CA) and non-cancer (non-CA) samples derived from the blood cell populations described in C. Each dot (CA) and square (non-CA) represents one sample. Fig.16A CD45 in sputum samples from cancer samples (CA) is shown. + The proportion of cells was significantly higher than that in non-cancerous (non-CA) samples (**p=0.0099). Fig. 16B Shown in CD45 + In the sputum samples obtained from cancer patients, granulocyte / lymphocyte subsets ( Fig. 12C Gate 1) in was significantly larger (*p=0.0378). Fig. 16C Figure 3 shows the CD45 expression of interstitial macrophages in sputum samples obtained from cancer patients compared to sputum samples from non-cancer patients. + Subgroups ( Fig. 12C Gate 3 in was also significantly greater (**p=0.0031). The thick black horizontal bar indicates the median value for each sample group.

[0054] Before describing the present invention in more detail, it should be understood that the present invention is not limited to the specific embodiments described, as such embodiments may of course vary. It should also be understood that the terminology used herein is only used to describe specific embodiments and is not intended to be limiting, as the scope of the present invention will only be limited by the appended claims. DETAILED DESCRIPTION

[0055] Sputum is a biological fluid that can be obtained in a non-invasive manner and can be separated to release its cellular contents, thereby providing a snapshot of the lung environment. Sputum can be made into a single cell suspension and stained with TCPP and fluorescent dye-coupled antibodies for manual and automatic flow cytometry (FCM) analysis. Automatic FCM allows the use of TCPP to analyze cancer or cancer-related cells in sputum samples, while using a set of compounds specific to cell lineage markers to query sputum samples to provide information about the lung environment of the subject providing the sputum sample, thereby capturing the predictive features of sputum samples using trained classifiers. In one embodiment, a sputum sample analyzed by flow cytometry contains about 10,000 cells in a subpopulation, which is selected by automatic gating and analyzed and classified in real time by the systems and methods disclosed herein. In other embodiments, the total number of cells in the subpopulation can be about 1,000-5,000, about 5,000-10,000, or about 10,000-100,000. Data acquired through automatic FCM can be stored and analyzed later, rather than in real time as the data is acquired.

[0056] Over the past 20 years, the field of automated FCM analysis has produced powerful software tools for identifying cell populations associated with clinical outcomes and managing increasingly complex FCM datasets. Much work has focused on reproducing expert analysis of FCM data to automatically identify cell populations, such as in human immune profiling. Such data-driven algorithms can now match or exceed human expertise, and analysis of acquired FCM data can be fully automated, eliminating potential operator bias. However, the application of automated flow cytometry to analyze sputum samples to determine lung health has proven difficult due to the complexity of the lung environment, the interplay between inflammatory markers and disease, and the often rare presence of cells of interest in sputum samples. Complicating factors when analyzing sputum samples to understand lung health include: sputum samples vary in size, as samples that are too small will contain too few cells to analyze, while samples that are too large may dilute the events of interest when the events are rare. Cytological examination of sputum samples on slides is limited to a portion of the sample, reducing the sensitivity of the slide assay, and is dependent on the skill of the observer to see meaningful events on the slide. The overlap of smears on the slide and cells on the slide reduces the usefulness of cytology to provide meaningful analysis when looking for rare events in a sample. Further cytology examination does not provide a holistic picture of what non-malignant cells may be present in the sample, their frequency in the sample, and what this information means about lung health in relation to lung cancer or other diseases of the lung.

[0057] Another aspect of one embodiment of the invention provides a supervised learning method for developing an assay that combines automated FCM data collection from induced sputum to isolate viable single cell events with machine learning techniques to classify patient samples as cancer or non-cancer. In another embodiment, sputum is not induced using saline and / or sputum is not collected using lavage. In another embodiment, sputum is induced using sound waves or vibrations in the lungs or by using a flute medical device. In another embodiment, sputum is collected using spontaneous expectoration and / or with the assistance of a positive expiratory pressure therapy device.

[0058] In one embodiment of the invention, the developed lung cancer / non-cancer classifier performed well with a sensitivity of 82% and a specificity of 88%. Further, the classifier achieved comparable sensitivity and specificity when applied to a set of independent samples collected using a flow cytometer platform (Navios EX) different from the flow cytometer platform used for assay development (LSRII). One aspect of the automated FCM lung assay system and method according to one embodiment of the invention is that the assay is accurate also in early stages (I and II) and when the lung nodules are small (<20 mm in diameter). The system and method of one embodiment of the invention is robust to differences in sample handling and processing and captures important predictive factors for the development of early lung cancer.

[0059] In one embodiment of the present invention, sputum is obtained from current and previous smokers, for example, with a smoking history of 20+ pack years, and the smoker is diagnosed with lung cancer or has a high risk of suffering from the disease. The sputum cells separated are counted, vitality is determined, and a group of markers (such as anti-CD45) are marked to determine the cell type, to separate leukocytes from non-leukocytes, but other markers are also possible, which will be discussed in this article. After excluding debris and dead cells (including squamous epithelial cells), we determined repeatable group characteristics and confirmed the lung source of the sample. For example, in addition to labeling sputum samples with leukocyte and epithelial specific fluorescent antibodies, fluorescent meso-tetra (4-carboxyphenyl) porphyrin (TCPP) known to preferentially stain cancer (related) cells is used to label sputum samples. Identification can be used to distinguish the differences in cell characteristics, group size and fluorescence intensity of cancer samples and high-risk samples.

[0060] In one embodiment of the invention, an analysis process combining automated flow cytometry processing and machine learning was developed to distinguish cancer cells from non-cancer cells in sputum samples. Flow data and patient characteristics were evaluated to determine predictors of lung cancer. The training set was used to fit the model, while the remaining samples were used for independent validation (test set). The method was further validated on a second set of samples processed on different flow cytometry platforms.

[0061] Reference now Figure 1, the utilization of sputum samples is shown in the flow diagram. Of the 171 samples initially considered for run on the LSRII flow cytometer (136 non-cancer; 31 cancer; 4 health status unconfirmed), 168 samples were used for model building and analytical pipeline development. This included 4 samples for which we could not determine the disease status, as the addition of unlabeled samples aided model building. In addition, 14 samples that were flagged as failing based on cell counts (see below) were also used for model development to better capture the distribution of the underlying data and help make the generalization of the model more robust to sample noise. Three samples could not be used at all due to technical issues during acquisition.

[0062] A total of 150 samples were used in the model validation phase (122 noncancer; 28 cancer). Of the 168 samples, 18 were omitted: 13 samples included too few cells for accurate analysis, 1 sample included too few alveolar macrophages to confirm it as a lung sample, and 4 samples were excluded because their cohort status could not be confirmed. The automated analysis was independently validated using 32 new samples. Participants adhered to the same inclusion criteria, and samples were processed using the same protocol as the previous sample set. Although a different flow cytometer (Navios EX) was used to run the second set of samples, the same model and coefficients were used to analyze the data from both instruments. 171 samples were initially considered for run on the LSRII (136 high risk; 31 confirmed cancer; 4 unconfirmed); 150 samples were ultimately used for automated analysis pipeline development (122 high risk; 28 confirmed cancer). Of the 21 samples omitted, 13 samples included too few cells for accurate analysis, 1 sample included too few alveolar macrophages to confirm it was a lung sample, and 3 samples had technical issues during collection. An additional four samples were excluded because their cohort status could not be confirmed. Of the 45 samples processed for Navios EX, 7 samples were excluded due to too few cells, 1 sample was excluded due to too few macrophages, and 5 samples were excluded due to technical issues (1 sample during processing and 4 samples during flow cytometry). The remaining samples consisted of 26 high-risk and 6 cancer samples.

[0063] Reference now Figure 2A , set the beads size exclusion (BSE) gate parameters on the entire sputum sample. Figure 2B Based on the observation of forward (FSC) and side (SSC) scatter areas, the BSE gate threshold setting of 2.5x10 5 The event is a dead cell. Figure 2CA dot plot of NIST beads from automated flowClust analysis is shown to show how it can be used to set the lower left threshold of the BSE gate. The lower left threshold of the BSE gate is derived from the automated NIST beads flowClust analysis. Single cell suspensions from three-day sputum samples were labeled with viability dyes to exclude dead cells, antibodies to distinguish cell types, porphyrins to label cancer-associated cells, and run on a flow cytometer.

[0064] One step in the development of a computer-automated lung cancer assay is the automated flow cytometry (FCM) identification of viable single cells (according to one embodiment of the invention, comprising Figure 2A , 3A , 4A, 4B). In one embodiment of the assay, the sample preparation assembly of the method consists of a plurality of assay tubes, for example, one assay tube may be labeled with a blood cell marker ("blood tube") and another assay tube may be labeled with an epithelial cell marker ("epithelial tube"). For example, both tubes may contain a fluorescent anti-CD45 antibody that selectively binds to blood cells and a viability dye to facilitate viability gating (see Figure 4AThe cells in the blood tubes can be stained with anti-CD206 (a marker for lung macrophages) and / or anti-CD3 (a marker for T cells) and / or anti-CD19 (a marker for B cells) and / or anti-CD66b (a marker for granulocytes) in addition to or as an alternative to CD45. The cells in the epithelial tubes can be stained with antibodies to pan-cytokeratin (panCK) and / or epithelial cell adhesion molecule (EpCAM). Vitality dyes, TCPP cell markers, blood cell markers, and epithelial cell markers can be distinguished from each other based on optical properties (e.g., fluorescence). In one embodiment of the invention, the fluorescence intensity, forward scatter, and side scatter of a cell subpopulation selected from a more inclusive cell population is used for downstream numerical analysis based on selection criteria including one or more of bead size exclusion gating, singlet gating, viability dye gating, cell lineage marker gating, and TCPP gating of flow cytometry sputum data. It should be noted that multiple cell lineage-specific markers are selected to identify different cell subpopulations in sputum samples, such as macrophages (e.g., using anti-CD206), T cells (e.g., using anti-CD3), B cells (e.g., using anti-CD19), granulocytes (e.g., using anti-CD66b), leukocytes (e.g., using anti-CD45), and epithelial cells (e.g., using anti-EpCAM and / or anti-PanCK). In one embodiment, specific cell lineage markers are not limited to the specific examples provided, as other clusters of differentiation (CD) markers may be used, such as CD64 for identifying interstitial macrophages. 高 / CD11b 高 / MHCII 高 / CD11c 高 / siglec F 负 and CD64 for identification of alveolar macrophages 高 / CD11c 高 / F4 / 80 正 / MerTK 正 / siglec F 低 A combination of CD4 and CD8 can be used to identify T cells.

[0065] In addition, in one embodiment, control samples may include one or more of the following: polystyrene beads of known diameter (approximately 5-30 μm NIST beads), compensation samples for each fluorescent dye channel used, unstained sputum samples, and cell lineage / type marker samples (e.g., antibody isotype sputum controls). Each sample tube corresponds to a single flow cytometry standard (fcs) file, which contains sample metadata and each event value acquired for each optical and fluorescent channel, as well as time parameters recorded when the flow cytometer queries and / or acquires the contents of the sample tube.

[0066] SEC is highly autofluorescent and can potentially lead to false positive events when sputum samples are interrogated and / or analyzed by flow cytometry. Therefore, during the analysis of sputum samples, eliminating SEC from the analyzed cell population is a step in one embodiment of the pulmonary assay. The inventors have found that whether physically eliminating SEC by filtering before analysis or performing negative size selection during analysis will not result in the exclusion of SEC cells. Eliminating SEC from the cell population of the analyzed sputum sample using a live / dead cell discriminator (FVS510) is a solution to exclude SEC from the event population acquired by the flow cytometer and / or analyzed in downstream analysis.

[0067] Since both the sputum cells of interest and SEC fell into this gating region, viability analysis was performed on isolated sputum cells within the 5 to 30 μm size parameter. The cutoff for FVS510 positivity was based on an unstained control. Back gating of dead cells onto the sputum light scatter profile showed that these cells had a generally high SSC, which is expected for SEC. To confirm that SEC was dead, sputum cells were separated into dead and live populations. Aliquots of the pre-sorted sample and the post-sorted population were transferred to a cytospin and stained with Wright-Giemsa. These slides showed that SEC was primarily present in dead cells, while live cells sorted from the same sample included both hematopoietic and non-hematopoietic cells, as well as a small amount of contaminating SEC. Therefore, it was determined that sputum samples could be analyzed by flow cytometry while excluding contaminating SEC via viability gating.

[0068] Figure 2A The steps to limit the events to be analyzed downstream are shown based on forward scatter area (FSC-A) and side scatter area (SSC-A), both of which are reasonable surrogates for cell size. A 2D cluster gate is used to find the main peak of 5 μm NIST beads in FSC-A vs. SSC-A ( Figure 2C ). The lower FSC-A limit of the bead cluster was set to the minimum sample FSC-A to exclude small particles and debris. Both FSC-A and SSC-A were set to 2x10 5The upper limit of the threshold is 50%, because events above these thresholds were found to be deaths (FVS510 + )cell( Figure 2A and Figure 2B ).

[0069] Figure 3A-3B FIG. 1 shows a point diagram data setting according to an embodiment of the present invention. Figure 3A , in some samples, events appeared in the lower right corner of the FSC-H vs. SSC-H plot. These events were excluded from further analysis to avoid including events / cells as live, marker-negative, TCPP-low cells, e.g. Figure 3C-3I shown. Figure 3A Shown is a temporary flowClust gate set on debris-free events in FSC-H vs. SSC-H (Debris-Free Gate) to retain the majority of live cells for the final FVS510 tail gating ("Core Viability Gate"). Figure 3B A dot plot of fragments with low SSC-A and low FSC-A is shown, and identification of additional fragments is shown ( Figure 3A ), and this fragment showed a light scatter profile that was not typical of a cell population. This needed to be ruled out because additional analysis ( Figure 3A -I) indicates that this fragment could be mistaken for a living cell ( Figure 3C ), the viable cells were low to negative for other markers ( Figure 3D-3I ).

[0070] Events within the bead size exclusion (BSE) gate were then restricted to exclude proteins with unusual FSC and SSC height profiles that may be unsuitable for inclusion in downstream analysis ( Figure 3A ) and the groups with different staining profiles. In general, it is necessary to exclude groups with abnormal appearance.

[0071] Reference now Figure 4A , a vitality gate (solid rectangle) was set for the remaining events based on FVS510-A fluorescence. Cells within the vitality gate (set as the threshold for FVS510 positivity) were retained ( Figure 4A ). Setting a vitality gate is challenging for some samples due to differences in sputum cell composition between patients. In these cases, heuristic-guided vitality gate setting ( Figure 4C -F).

[0072] Reference now Figure 4B , for all vitality events ( Figure 4AThe singlet gate (solid polygon) is set based on the events within the liveness gate. The identified (live) singlets are used for downstream numerical analysis. Setting the singlet gate is challenging for some samples due to differences in sputum cell composition between different patients. In these cases, heuristic-guided singlet gate setting ( Figure 5A -H).

[0073] Figure 4C-4E A dot plot showing the heuristic-guided viability gate settings. The subpopulation most likely to contain viable singlet cells (i.e., cells with relatively small area and height in the light scatter channel) is used to guide the positioning of the viability gate ( Figure 4C and 4D ). Figure 4C Samples with a lower percentage of high SSC / FSC cells are shown.A temporary flowClust gate was set on all debris-free events in FSC-H vs. SSC-H to retain the majority of live cells for the final FVS510 tail gating ("core viability gate"). Figure 4D It is shown that for samples with <10% of events in the core vitality gate, flowClust can be re-run more inclusively by increasing the "quantile" parameter to 0.99. Figure 4E Shows a temporary singlet gate set on a core vitality event, by setting the upper right point to 2.5x10 on both axes 5 to force capture of the upper diagonal. Figure 4F Shows Figure 4E FVS510 staining (vitality dye) histogram of events identified in the temporary singlet gate in . The viability cutoff is automatically set at the core viability singlet state (black histogram). The line marked "blue curve" is the intact, fragment-free FVS510 profile used for comparison. The vertical line marked "red bar" indicates the viability gate cutoff. Viability events are located to the left of the viability gate cutoff. Once the viability cutoff is determined, all temporary gates are removed. The viability gate cutoff thus determined is now used to set the viability gate for the entire sample (e.g. Figure 4C and 4D shown).

[0074] Reference now Figures 5A-5D , showing heuristically guided singlet gating. In some cases, cells with high SSC-A (e.g., see Figure 3A ) to escape singlet gating on a fully viable cell population. Figure 5A A dot plot is shown with FSC-A on the x-axis and FSC-W on the y-axis, and with an automatically assigned singlet gate (solid polygon) on a sample with too many high SSC-A cells. The wide singlet gate is an erroneous gate. This gate can be corrected by fitting a gate to a temporal subpopulation of live cells that does not include high SSC-A cells ( Figure 5B ). Figure 5B A dot plot is shown with FVS510-A on the x-axis and SSC-A on the y-axis, with the exclusion of samples above approximately 5x10 4 A temporary gate was set up on SSC-A for the event. Figure 5C shows a singlet gate (using Figure 5B The singlet gate was automatically fitted to the restricted population (solid polygon) and the upper right corner was set to 2.5x10 on both axes. 5 (dashed polygon) to include the upper diagonal. Figure 5D The refined singlet gate (solid polygon) applied to the fully energetic population is shown when the temporary gate is removed ( Figure 5D The refined polygons and Figure 5A The singlet gate (solid polygon) in Figure 5 is used for comparison.

[0075] Reference now Figure 5E-5H , heuristically guided singlet gating is shown for different types of situations where the singlet gate is set. In these cases, the population representing about>10% singlets is located between about 2.5 (logiclescale) and the vitality threshold that escapes setting the vitality gate and the singlet gate. This can be corrected by resetting the vitality gate and fitting the singlet gate to a more restricted vitality population. Figure 5E A dot plot of this difficult case is shown with FVS510-A on the x-axis and FSC-A on the y-axis. The solid ellipse represents >10% singlet states and the dashed line represents the viability threshold. Fig. 5F Population mixture analysis shown in density histograms of FVS510 staining is shown. The analysis highlights Figure 5E The rightmost group (indicated by the blue curve) marked by the solid ellipse in Figure 5E The signal distribution differences for the majority of events (indicated by the black curve) are located to the left of the middle ellipse and indicate that the natural cutoff value for these unusual cases is 2.5 (dashed line). Figure 5G A dot plot is shown where the vitality gate found by automatic tail gating (dashed line) is replaced by an adjusted vitality gate (solid line box). Figure 5H A dot plot with FSC-A as the x-axis and FSC-W as the y-axis is shown for Figure 5G The refined viable cell population identified in the “Adjusted Vitality Gate” shown in , a new singlet gate (solid polygon) was calculated.

[0076] Use a "singlet" gate (FSC area vs. FSC width) to exclude cell doublets or small aggregates ( Figure 5A). Automatically setting the singlet gate is challenging due to the variability of samples between patients. In some samples, high SSC-A cells are included in the viable cell population and distort the results of the singlet gating algorithm. This can be corrected by fitting a gate to a temporary subpopulation that excludes most of these high SSC-A events ( Figure 5B ). In other samples, two populations can be seen within the live cell gate, one with low FVS510 staining and the other just below the viability threshold with a high side scatter profile. In this case, the correction involves resetting the viability gate and fitting the singlet gate to the more restricted viability population. Light scattering and fluorescence signal values ​​are recorded for each single event and used for downstream model development and validation.

[0077] Based on our earlier slide-based assay results, we expected that smoking history (or related factors such as age) and TCPP signal density (rather than fluorescence intensity per se) would be important predictors. Therefore, we divided the fluorescence signal of all channels by log 10 FSC-A or log 10 SSC-A, and the obtained density distribution was divided into three regions (R1 = about <0.25, R2 = about 0.25-0.6, R3 = about >0.6, Figures 6A-6D ). It turns out that two such density signals provide useful information to the classifier: TCPP / log 10 SSC-A( Fig. 6A , Region 3 [R3]) and FVS510-A / log 10 FSC-A( Figure 6C , Region 2 [R2]). TCPP / log 10 The predictive value of the SSC-A signal density was not imposed by the stepwise regression but emerged spontaneously. The fact that the FVS510-A signal density was also found to provide useful information is interesting and may be related to the fact that apoptotic cells can take up this dye at moderate levels.

[0078] Fig. 6A A histogram of live single cells stained with TCPP is shown, showing the R1, R2 and R3 fractions, where the R3 fraction is identified by a solid rectangular box. Region R3 (shaded) represents live singlets with high TCPP signal relative to side scatter (log10 transformed). Figure 6B TCPP / log 10 Histogram with SSC on the x-axis and density on the y-axis shows cells without TCPP (unstained control) and no events in region R3. Figure 6C A histogram of live single cells is shown, with the x-axis being FVS510 / log 10FSC and density on the y-axis show cells that are negative for FVS510 (R1) or cells that stain with low levels of FVS510 (R2; solid rectangular shaded box). Cells in R1 and R2 are below the FVS510 positive cutoff, but nonetheless have similar Fig.6D Cells in R2 showed relatively higher FVS510 signal relative to forward scatter (log10 transformed) compared to the unstained control sample in Figure 2. Fig.6D Shown is the FVS510 / log 10 Histogram with FSC on the x-axis and density on the y-axis shows cells not stained with FVS510 (unstained control).

[0079] Combinations of cell lineage markers can identify subpopulations that a single cell lineage marker alone may not capture. Careful examination of patient sputum samples by FCM revealed complex patterns of cell lineage marker expression in blood and epithelial ducts, but this information was not readily apparent for lung health.

[0080] Fig. 6E A dot plot is shown with CD206 as the x-axis and CD66b / CD3 / CD19 as the y-axis. The solid shaded box on the lower quadrant of CD206 identifies non-macrophages, while the solid box on the middle to upper quadrant of CD206 identifies lung macrophages. Anti-CD206 is a macrophage marker that is specific for a population of macrophages that is located in lung tissue and is not present in the blood circulation. Anti-CD66b is a granulocyte marker, anti-CD3 is a T cell marker, and anti-CD19 is a B cell marker. Fig. 6F A dot plot with CD206 on the x-axis and CD66b / CD3 / CD19 on the y-axis is shown, and the solid shaded box on the lower quadrant of CD206 identifies the absence of non-macrophages, while the solid boxes on the middle to upper quadrants of CD206 identify the absence of lung macrophages in the control samples. Fig. 6E The non-macrophages (CD206 低 ) Leukocytes (CD66b / CD3 / CD19 中 ) is the tan shaded area. The shaded area (CD206 中 / 高 CD66b / CD3 / CD19 低-中 ) contained lung macrophages, which were used as a marker for samples considered to adequately sample the lung environment. Fig. 6F The quadrant labels are also applicable to the unstained control samples. Fig. 6E All figures are from the same illustrative sample.

[0081] Based on blood Figure 6E-6F) and epithelial tubes (data not shown), paired cell markers were analyzed by binning fluorescence. Signal intensity on a logistic scale was quantified into low (<1.5), low-medium (1.5-2.5), medium (2.5-3), and high (>3) windows. Every 10,000 events were tabulated for each region in the resulting 4x4 grid of CD206 versus CD3 / CD19 / CD66b (blood tubes) or EpCAM versus panCK (epithelial tubes, data not shown). One region in the blood signal intensity grid was found to provide useful information for the predictive classifier lung assay model (tan shaded rectangle; low for CD206 and medium for CD3 / CD19 / CD66b, Fig. 6E ). This group may indicate the presence of immune or inflammatory processes in the lungs.

[0082] According to one embodiment of the present invention, the development of automated processing of flow cytometric data features of sputum samples collected using an LSRII flow cytometer provides one or more setup and / or quality control steps prior to automated analysis of patient tubes / samples, as follows:

[0083] For each sample tube, use e.g. time-contrast fluorescence channels to remove outlier events, e.g. using [flowCut]

[0084] The compensation tubes are used to automatically derive a spillover matrix, for example using [flowStats]; it should be noted that for Navios flow cytometers, the automated process begins with the operator spillover matrix from an unstained sample and checks the alignment of the medians in the "off" channel (i.e., the non-FITC channel if FITC is being compensated), making adjustments as needed based on the expectation that the means of the positive and negative populations (as defined in the channel being compensated, e.g., FITC) should align in the off channel (e.g., the non-FITC channel).

[0085] · Compensate the fluorescence signal and convert it to a logical scale, for example using [flowWorkspace]; · Note that for automatic processing on Navious, the 20-bit signal (0 to 1,048,575) is converted to 6 decimals to roughly match the LSRII range.

[0086] Automatically find the approximately smallest NIST Everest peak to set a lower FSC-A threshold (low) to exclude large amounts of debris, such as using [openCyto]. For example, when using an LSRII flow cytometer, start with 10 4 (Low) to 5x10 4 The (high) window is the destination area;

[0087] It should be noted that for Navios:4 (Low) to 10 5 (High) window According to one embodiment of the present invention, further automated analysis of the patient tube / sample provides:

[0088] Patient tube:

[0089] Events within the rectangular size gate were retained for further analysis, “BSE” (FSC-A: low - 2.5x10 5 , SSC-A: 0-2.5x10 5 ). [flowWorkspace]

[0090] For Navios: "BSE" (FSC-A: Low -10 6 , SSC-A: 5x10 3 -10 6 )

[0091] Apply a secondary “no debris” gate to the preserved BSE population to exclude occasional high FSC-H content (>2x10 5 ) and abnormal cell populations with medium / high SSC-H content (>1x10 5 ).[flowWorkspace];

[0092] For Navios: The equivalent of the LSRII debris-free gate is the selection polygon "cleanSSC" defined on SSC-W vs. SSC-A (lower left = 50,50; upper left = 50,10 6 ; Lower right = 200,50; Upper right = 10 3 ,10 6 ) rather than an exclusion gate on FSC-H vs. SSC-H. The choice of channel depends on how well the fragments are visualized and excluded. This gate is applied at the end of the Navios treatment for retained viable singlets.

[0093] Use the subset of cells most likely to contain the live singlet population of interest (virtual gate set) to set the viability gate coordinates. This helps reduce artifacts and distortions caused by contaminating squamous epithelial cells (SEC) and other dead cells.

[0094] For LSRII:

[0095] Using FSC-H vs SSC-H, "hxh" sets an automatic flowClust gate on debris-free events to exclude most contaminating SECs, limiting both channels to <1x10 5 And a 0.99 quantile cutoff was used. [openCyto, flowClust]

[0096] In some samples, if <10% of events were retained in hxh (low.live=true), the channel limit was relaxed to 1.5x10 5 And the quantile is relaxed to 0.9.

[0097] Set automatic singlet gate "Singlet" on hxh using FSC-A vs FSC-W channels and "wider_gate = true" setting. [openCyto,flowStats];

[0098] The upper right FSC-A coordinate is forced to the same value as the lower right FSC-A coordinate to avoid persistent SEC that would tilt the gate downward. [flowWorkspace];

[0099] Set up automatic tailgate "live" on singlets using FVS510-A channel, limits to min=2, max=3 (min and max on logic scale), tolerance=0.1. [openCyto] If low.live=True, first limit singlet events to SSC-A<5x10 4 And set smoothing parameter adjustment = 1; [flowCore]

[0100] Otherwise, first restrict singlet events to those with FSC-A < 5x10 4 And set the smoothing parameter adjustment = 1.2.

[0101] [flowCore]

[0102] For Navios, it should be noted that the BV510-A signal is better resolved on the Navios than on the LSRII, resulting in slightly different processing steps.

[0103] BV510-A [0,3.5] versus SSC-W [50,200], "quickLive", was used with a rectangular gate set on fragment-free events to exclude most contaminating SEC.

[0104] Set the automatic singlet gate "Singlet" on the QuickLife using the FSC-A vs FSC-W channels and the "wider_gate = false" setting.

[0105] Force the rightmost FSC-A coordinate to the BSE limit (10 6 ), and increase the maximum FSC-W limit to 2x10 5 , to counteract the situation where the gate is tilted downward by the continued presence of the SEC.

[0106] Set the automatic tailgate "live" on the BSE using the FVS510-A channel with the parameters (side = "left", maximum = 3.5, tolerance = 0.1).

[0107] It is worth noting that the BV510-A cutoff for most samples was reasonable and close to 4, but occasionally a heuristic "tweak" was needed to correct for cases where the signal did not fall into two well-resolved groups. The live tailgate parameters were adjusted.

[0108] Of note, the isotype samples had no TCPP, which resulted in a higher BV510-A signal

[0109] Adjust the threshold:

[0110] Reason.Cutoff = 4, High.Cutoff = 5 (excludes patient samples of isotype)

[0111] Cause.Cutoff = 4.5, High.Cutoff = 5.5 (same type)

[0112] If (BV510-A > #Events of Live Cutoff) > 1.5 times (BV510-A > #Events of Reason.Cutoff): Live Cutoff is too low, so needs to be moved to the right.

[0113] Set the Tailgate "Minimum" parameter to the greater of Cause.Cutoff and Mean BV510-A Expression, and increase by an additional 0.5 if >25% of events > Cause.Cutoff are also > High.Cutoff.

[0114] If (BV510-A > #Events of Live Cutoff) > 2 times (BV510-A > #Events of Reason. Cutoff): Set Tailgate Tolerance = 0.2 and Adjustment = 0.5.

[0115] If (BV510-A > #Events of Live Cutoff) times 1.5 < (BV510-A > #Events of Reason.Cutoff): Live Cutoff is too high, so it needs to be moved left.

[0116] Set the Tailgate "Max" parameter to the lesser of High.Cutoff and Mean BV510-A expression, then increase by 0.5 if >25% of events have BV510-A > High.Cutoff.

[0117] Set the Tailgate "Min" parameter to 3.5, and increase by another 0.5 if >25% of events have BV510-A > High.Cutoff.

[0118] If (BV510-A > # events of live cutoff) times 2.25 < (BV510-A > # events of reason.cutoff): Overshift left to increase tailgate "min" by another 0.5 and set adjust = 2. Create a temporary gating set with events in both BSE and singlet gates and recalculate automatic tailgate live using adjusted parameters.

[0119] Apply live gates determined from the virtual gate set to the full gate set without fragmentation events. [flowWorkspace]

[0120] • Automatic singlet gate "Singlet" set on live events gives good results in most cases, but some samples still contain confounding events.

[0121] SEC contamination can map near larger live cells, which can result in lower left coordinate < 0 or lower right coordinate < upper left coordinate. In both cases, limit live cells to SSC-A < 5x10 before setting the automatic singlet gate using "wider_gate = false". 4 [openCyto, flowStats]

[0122] In all cases, force the upper right FSC-A coordinate to have the same value as the lower right FSC-A coordinate to avoid situations where persistent SEC would tip the gate downward. [flowWorkspace]

[0123] Apply singlet gates to live events. [flowWorkspace]

[0124] • When viable cells are relatively rare, the proportion of events retained in the singlet gate in the gap between FVS510-A>2.5 (logistic scale) and the viability cutoff can be substantial.

[0125] If the "gap" population > 10% of singlet states, force FVS501-A < 2.5 as viability cutoff. - Remove "alive" from the full gating set. [flowWorkspace]

[0126] Set up a "live" rectangular gate for FVS510-A < 2.5 without debris. [flowWorkspace]

[0127] Recalculate the automatic singlet gate 'Singlet' as described above. [openCyto, flowStats, flowWorkspace]

[0128] -Apply the singlet gate to events that are active in the complete gating set. [flowWorkspace]

[0129] It is important to note that for Navios, the intermediate gate has been removed from the patient samples, leaving only the BSE. Apply live gate to BSE.

[0130] Apply the singlet gate to the live.

[0131] Apply the cleanSSC gate to singlet states.

[0132] A complete matrix of events preserved by singlet states (LSRII) or cleanSSC (Navios) is written out for gating of each patient tube. These values ​​along with patient metadata are the input to the cancer / non-cancer classifier.

[0133] Note that the retained events are quantized to generate the model variable values. The ranges are not the same due to the differences in the output values ​​of the LSRII and Navios detectors. See 5.5 Grid Settings below.

[0134] According to one embodiment of the present invention, the analysis steps of LSRII flow cytometer [R package] with embedded Navios details are demonstrated:

[0135] Note 1: The steps for the Navios EX flow cytometer are essentially the same but are adjusted for the different file format (i.e., LMD files for NaviosEX instead of FCS files for LSRII) and detector sensitivity and dynamic range (LSRII = 18 bits, Navios = 20 bits).

[0136] 1.1: LSRII uses 2 configuration files. One is to match the file name to the fluorescence control for compensation (matchfile.csv). The other is to identify the patient tube (possibly multiple tubes) and the flow cytometer channel name (samplematch.csv). For examples, see 5. Supplementary material "Configuration files for LSRII" (5.1 and 5.2) below.

[0137] 1.2: Instead of using configuration files, Navios Sample Collection introduces a controlled tube-naming convention. See 5.3 Navios Controlled Tube Naming Protocol below.

[0138] 1.3: The names and order of the fluorescence channels in the flow cytometer files are different, but can be found in LSRII

[0139] and Navios. See below 5.4 Channel Equivalence and Detector Maximum. NOTE 2: Gate names are in bold capital letters.

[0140] Note 3: Temporary gate names are indicated in bold lower case letters.

[0141] NOTE 4: Calculated thresholds are in bold italics.

[0142] Note 5: The main packages used in the current step are indicated by [square brackets].

[0143] Note 6: Some heuristic tuning is required to handle a wide range of sample compositions and viabilities.

[0144] The following provides an implementation of a sample processing flow according to one embodiment of the present invention:

[0145] Analysis tube:

[0146] NIST beads

[0147] Compensation tube (one for each fluorescence channel)

[0148] Patient single cell suspension (4 tubes):

[0149] Unstained as a control for PE-anti-CD45 and TCPP

[0150] Isotype controls stained with FVS510 (viability), PE-anti-CD45, PE-CF594 isotype, A488 isotype, FITC isotype

[0151] "Blood" stained with TCPP, FVS510, PE-anti-CD45, PE-CF594-anti-CD206, A488-anti-CD3, FITC-anti-CD66b, A488-anti-CD19 (Note: A488 and FITC are read on the same channel, generating a combined CD66b / CD19 / CD3 signal)

[0152] “Epithelium” stained with TCPP, FVS510, PE-anti-CD45, PE-CF594-anti-EpCAM (epithelial cell adhesion molecule), A488-anti-PanCK (pan-cytokeratin)

[0153] 5.5 Grid Settings

[0154] LSRII Navios Fluorescence CD206 CD66b / CD3 / CD19 CD206 CD66b / CD3 / CD19 Low <1.5 <1.5 <2.5 <2.5 Low Medium >=1.5,<2.5 >=1.5,<2.5 >=2.5,<3.5 >=2.5,<3.5 middle >=2.5,<3 >=2.5,<3 >=3.5,<4.5 >=3.5,<4 high >=3 >=3 >=4.5 >=4 Fluorescence / Size log10(FSC) log10(SSC) log10(FSC) log10(SSC) Low <0.3 <0.3 <0.6 <0.9 middle >=0.3,<0.6 >=0.3,<0.6 >=0.6,<0.8 >=0.9,<1.2 high >=0.6 >=0.6 >=0.8 >=1.2

[0155] 5. Additional Materials as Examples of Data Inputs to Systems and Methods

[0156] 5.1 Sample matchfile.csv configuration file for LSRII

[0157] File name, channel

[0158] Sample_001_BL-18-148Unstained_011.fcs, unstained

[0159] Sample_001_A549 FVS510_007.fcs, BV510-A

[0160] Sample_001_PE comp beads_003.fcs, PE-A

[0161] Sample_001_PE-TxRed comp beads_005.fcs, PE-Texas Red-A

[0162] Sample_001_FITC comp beads_004.fcs, FITC-A

[0163] Sample_001_A549 TCPP_008.fcs, APC-A

[0164] 5.2 Sample samplematch.csv configuration file for LSRII

[0165] #Multiple collection tubes are separated by / /

[0166] Variables, details

[0167] Blood, s_BL-19-151_blood_016.fcs / / s_BL-19-151_blood-tube_017.fcs

[0168] Epithelial, s_BL-19-151_Epithelial_018.fcs

[0169] Isotype, s_BL-19-151_isotype_015.fcs

[0170] Stub, BL-19-151

[0171] fitc_name, FITC-A

[0172] nist,s_NIST_001.fcs

[0173] 5.3 Navios Controlled Naming Agreement

[0174] 01_NIST

[0175] 03_FITC

[0176] 04_PE

[0177] 05_PECF594 (Texas Red)

[0178] 06_APCpos (requires pre-compensation of combination tube 13)

[0179] 07_FVS510

[0180] 08_Accession#_Unstained (patient sample)

[0181] 09_Accession#_Isotype (patient sample)

[0182] 10_Registration #_Blood (patient sample)

[0183] 11_Registration#_Epithelial (patient sample)

[0184] 13_APCneg (requires combination tube 6 pre-compensation)

[0185] 5.4 Examples of channel equivalence and detector maximums for different instruments used to convert data are shown in the grid below.

[0186]

[0187]

[0188] *As found in FCS (LSRII) or LMD (Navios) metadata.

[0189] We identified a reduced list of cell lineage markers as the most promising potential predictors and interrogated their pairwise interactions to derive classifiers that provide high sensitivity and selectivity in the assay for identifying cells that are likely to be cancerous and those that are unlikely to be cancerous. 10 Negative values ​​proportional to the number of events in FSC-A R2 improve the performance of the classifier. One interpretation of this interaction term is that it serves to moderate the accumulation of stress cells in the high-risk patient group that may be age-related due to smoking or health history.

[0190] After developing the two phases of the lung assay, a complete pipeline was assembled, including quality control steps, determination of predictive variable values, and sample classification ( Figure 7 ).

[0191] Reference now Figure 7, according to one embodiment of the present invention, a lung sputum data processing flow is shown. The schematic diagram represents the following major elements: Quality control (QC) measures 701 and 703 are performed on the data and instruments. The adequacy of the data acquisition file is reviewed. Sputum samples with approximately 10,000 live singlet events and approximately 10 or more lung macrophages in the sample are determined to be sufficient. A retrieval of subject-specific data (such as the information in Table 1) is obtained and combined with sputum sample features obtained from the FCM population 705 obtained from the FCM. The classifier 706 is a straight line equation including coefficients b0-b5, which are determined by fitting the 706 classifier model to the specific sample set used to build the classifier. The classifier 706 is applied to a new sample to determine the value of the possibility of cancer based on the cutoff value 707. The classifier model is applied to the FCM population obtained from the sputum sample, and the subject-specific data is input into the classification model to determine whether the sample may be cancer 708 or not 709 (bottom diamond). The coefficients b0-b5 derived from fitting the classifier 706 to the particular set of samples used to build the classifier are as follows:

[0192] b0:53.515414

[0193] b1: 0.701153

[0194] b2: 0.001545

[0195] b3:58.071495

[0196] b4:4.258709

[0197] b5: 0.784982

[0198] It will be appreciated by one of ordinary skill in the art that fitting the classifier 706 to a different set of sample values ​​will change the values ​​of the coefficients but will not change the model itself.

[0199] In one embodiment of the invention, data collection adequacy 701 and sample acceptability assessment 703 first ensure that the data file for each collection tube is readable and that its encoded data matrix is ​​complete. Time stamps are used to check the fluorescence channel in each tube and eliminate flow rate anomalies caused by bubbles or blockages during sample collection. The compensation matrix is ​​then derived from scratch using a fluorescence compensation tube (rather than using the compensation matrix encoded in the sample file metadata). The fluorescence signal is compensated and converted to a logical scale to produce a sample data matrix used by automatic FCM gating to isolate live singlet events. In order to have confidence in downstream numerical analysis, samples containing a threshold number of singlets are analyzed, such as at least 10,000 live singlets, given that the number of events in some analysis windows may be small. The threshold is set to Fig. 6EThere are approximately at least 10 cells in the shaded area, among which we found lung macrophages (CD206 中和高 CD3 / CD19 / CD66b 低-中 ) and confirmed that the sputum sample originated from the lungs.

[0200] Figure 7 A further step in the assay flow is to provide the age (in years) and the flow-based value from the live singlet to the classifier model. In one embodiment of the invention, four variables are used to determine the classifier for determining the likelihood of having cancer: i) age, ii) the number of cells with TCPP / log in region 3, and iii) the number of cells with TCPP / log in region 4. 10 Number of events per 10,000 live singlet states (per 10K) of SSC-A ( Fig. 6A , R3), iii) FVS510-A / log in zone 2 10 Number of events per 10K of FSC-A ( Figure 6B , R2), and iv) CD206 低 CD3 / CD19CD66b 中 Number of events per 10K in the region ( Figure 6C , shaded box). In addition, the classifier contains a negative term representing the interaction between age and the FVS510 density variable and an “intercept” term (b0) that is used to avoid forcing our multifactorial model to zero if all variables are set to zero, but cannot be directly interpreted as a biologically meaningful component of the classifier. The values ​​of the coefficients (b1, b2, b3, b4, and b5) depend on the training set used to fit the model and provide weights for the variables. To make the interpretation of the model easier, we did not normalize the data provided to the model. For example, the model formula tells us that increasing the number of events with high TCPP increases the likelihood of cancer, which is consistent with our previous results.

[0201] Figure 7 The determining step in the pipeline is making the cancer / non-cancer assignment. The model returns a value in the range [0, 1]. Whether a given sample is classified as cancer depends on whether the value returned by the model is greater than a predetermined cutoff. If the value is less than or equal to the cutoff, the sample is classified as non-cancer. A reasonable cutoff value can be chosen by stepping through the cutoff values ​​between 0 and 1 and measuring the true positive and false positive calls compared to the known group class at each step.

[0202] The development of a GLM classifier according to one embodiment of the invention is described below and Fig.11 The input to the process is described by the output of the FCM step (see e.g. Fig.11 1101-1102 in the text).

[0203] Evaluating combinations of potential predictive factors ( Fig.11 1103-1105)

[0204] 1) A combination of potential predictive factors was evaluated (see e.g. Fig.11 1103) to develop a classifier using a generalized linear model (GLM) for semi-supervised machine learning as follows:

[0205] 1.1. Clinical parameters available for all samples

[0206] 1.2. Individual and pairwise combined quantitative measurements from patient blood and epithelial samples (light and fluorescence) using equally spaced or heuristically positioned breakpoints (3x3 and 4x4 grids per test channel)

[0207] 1.3. Quantitative measurement, as in #2 above minus background counts

[0208] 1.4. Quantification of fluorescence / log10 (light scattering)

[0209] 1.5. Frequencies of specific subpopulations determined by prior manual analysis

[0210] 2) Use different training and testing groups (roughly 2 / 3 training, 1 / 3 testing, randomly chosen and non-overlapping), iteratively stepping forward and stepping back parameter inclusion in the GLM (see e.g. Fig.11 1104-1105);

[0211] 3) Evaluate GLMs based on the Akaike Information Criterion (AIC), which estimates prediction error and provides a comparator of model quality; (see e.g. Fig.11 1105);

[0212] 4) Repeat the retention of parameters in different iterations of model testing and test them individually and in combination with different combinations of potential predictive variables to detect potential interactions (see e.g. Fig.11 1104-1105);

[0213] 5) Control overfitting of the final model by repeated random sampling of the training / test sets (n=10) to verify the predictive robustness independent of the specific samples used to train the model; (see e.g. Fig.11 1106); and

[0214] 6) Validate the processing pipeline and classifier predictions on samples not used in model construction and testing (see e.g. Fig.11 1107-1110 in the text).

[0215] Once a classifier is developed, it is applied to new samples that have not been previously analyzed and for which the cancer / non-cancer classification is unknown (see Fig.111101 and 1111 in it).

[0216] Reference now Fig. 8A , shows the results of this process as a receiver operating characteristic (ROC) curve with an AUC of 0.89. The assay achieved the best performance in distinguishing cancer from non-cancer at a threshold of 0.28 ( Figure 8B , solid vertical line).

[0217] Reference now Fig. 8A , shows a receiver operating characteristic (ROC) curve showing the false positive rate versus the true positive rate calculated as the model response threshold changes, the line indicated by the asterisk in the figure. For comparison, the ROC curve of the previous version of the slide-based lung assay is shown as an inverted triangle curve. The dotted lines indicate 80% and 90% sensitivity and specificity (1-false positive rate). Figure 8B Shown based on Fig. 8A Model of the ROC curve in . The model value (cancer likelihood) threshold was set to 0.28 (solid vertical line), corresponding to 82.1% sensitivity and 87.7% sensitivity. The dashed lines indicate the levels below which cancer is predicted to be very unlikely (leftmost dashed line) or very likely (rightmost dashed line). Dark circles represent cancer and light circles represent high risk. CyPath Lung performance on LSRII samples is shown.

[0218] In one embodiment of the invention, automated analysis of samples using flow cytometry combined with machine learning yielded a predictive model that was sensitive (82%), specific (88%), and robust to differences in sample processing and disease stage. Importantly, for difficult-to-treat cases with no nodules or nodules ≤ 20 mm in diameter, the test had a sensitivity of 92% and a specificity of 87%.

[0219] Reference now Fig. 9A -B, Correlation plot is a graphic representation of age on the x-axis versus model value on the y-axis when samples were analyzed on an LSRII flow cytometer or a NaviosEX cytometer, respectively.

[0220] Reference now Fig.10, an exemplary system and method 1000 is shown for receiving input from a flow cytometer 1002 based on a sample 1001 analyzed by the flow cytometer and subject metadata information 1020. A computer-implemented device 1003 includes a processor 1010 (e.g., a processing circuit) that is operable to execute program instructions or software to cause the computer to perform various methods or tasks, such as performing techniques for generating and / or using analysis models of classifier models as described herein. The processor 1010 is coupled to a memory 1030 via a bus (not shown), which is used to store information such as program instructions and / or other data while the computer is running. A storage device 1040 (e.g., a hard drive non-volatile memory or other non-transient storage device) stores data files such as program instructions, multidimensional data and reduced data sets, and other information. The computer also includes various input-output elements 1050, including parallel or serial ports, USB, Firewire or IEEE 1394, Ethernet, and other such ports to connect the computer to external devices, such as printers, video cameras, display devices, medical imaging devices, monitoring devices, etc. Other input-output elements include wireless communication interfaces, such as Bluetooth, Wi-Fi, and cellular data networks. The computer itself can be a traditional personal computer, a rack-mounted or commercial computer or server, or any other type of computerized system. In a further example, the computer can include less than all of the elements listed above, such as a thin client or mobile device having only some of the elements shown. In another example, the computer is distributed in multiple computer systems, such as a distributed server where many computers work together to provide various functions.

[0221] Reference now Fig.11, according to one embodiment of the present invention, an exemplary processing flow chart 1100 is shown for semi-supervised machine learning of a GLM classifier for cancer / non-cancer status and subsequent classification of biological samples such as sputum. The sputum sample processed for processing in a flow cytometer is provided to the flow cytometer for analysis as described herein 1100. Flow cytometric data from multiple sputum samples is processed using automated FCM 1101 to extract a data matrix of live singlet populations from each sample 1102. For all samples 1103, the values ​​of potential predictor variables associated with the collected flow cytometry data matrix and related clinical data are calculated. Multiple samples are randomly divided into training sets and test sets 1104. A GLM classifier is constructed using the values ​​of the potential predictors from the training set and the known cancer / non-cancer status of the training set. The ability of the resulting classifier to correctly classify the test set as cancer / non-cancer 1105 is evaluated on the test set. Step 1105 is repeated using different combinations of potential predictive variables, retaining only the true predictive variables in each iteration. A GLM classifier is fitted with only the retained predictive variables to produce a final model 1106. A single sputum sample of unknown status requiring classification is processed using automated FCM 1101 (single) to obtain the values ​​of the final predictive variables from the sputum and subject metadata (e.g., age) 1111 and input into the final model to generate classification values ​​1107. Based on the cutoff values ​​of the classification values ​​determined by ROC analysis, an output classification call is generated 1108 for the sputum sample of cancer 1109 or non-cancer 1110.

[0222] According to one aspect of the invention, the system and method correctly classified study participants as cancer or non-cancer with high accuracy, including participants at different stages of the disease and with nodules smaller than 20 mm. Therefore, the test has the potential to improve the process of early lung cancer diagnosis.

[0223] Materials and methods

[0224] Clinical trials and collection sites

[0225] The minimal risk study was registered with ClinicalTrials.gov, reviewed and approved by the Sterling Institutional Review Board (Atlanta, GA), and conducted in accordance with the ethical principles of the Declaration of Helsinki (v 1996) and good clinical practice guidelines. Sample collection was performed at five study centers: Atlantic Health System, New Jersey; Mount Sinai Hospital, New York; Radiology Associates, Albuquerque, New Mexico; South Texas Veterans Health Care System; and Pulmonary Associates, Waterbury, Connecticut. Each site received institutional approval to participate in the study. Each potential participant received an informed consent form, and only those who signed the consent form were included.

[0226] Participant Information: Participants (males and females) were eligible for inclusion in one of two groups. The non-cancer group included participants (aged 50-80 years) who were either current smokers with a smoking history of at least 20 pack-years or current nonsmokers with a smoking history of at least 30 pack-years but who had quit within the past 15 years. The exceptions were two patients who had quit smoking 25 and 26 years ago, respectively. Most participants in the non-cancer group received LDCT results or other forms of imaging indicating that cancer was not suspected, and they were advised to return for LDCT screening in 12 months. In a few cases, participants who were initially assigned to the non-cancer group underwent follow-up LDCT, PET / CT, or biopsy. These participants will be followed until their health status is confirmed. If they are diagnosed with lung cancer, they will be switched to the cancer group.

[0227] Each participant in the cancer group had been evaluated by a physician as having a high suspicion of lung cancer based on medical history and LDCT or other imaging results. The diagnosis was confirmed by biopsy after providing a sputum sample. The exception was one patient who developed a new nodule of 24 mm that was too fragile for biopsy. If the biopsy did not show cancer, the participant was transferred to the non-cancer group. There were no restrictions on age or smoking history for inclusion in the cancer group.

[0228] For each participant, we collected the following demographic data: sex (male or female); age (years); race (Hispanic / non-Hispanic Latino / Latino); and ethnicity (American Indian / Alaska Native; Asian; Black / African American; Native Hawaiian / Other Pacific Islander; White; Other). Data on smoking history and on comorbidities (asthma, COPD, emphysema, chronic bronchitis) and previous cancer history were collected. All participants were required to be willing to provide contact information for their primary care physician and to consent to the disclosure of medical information upon request. Exclusion criteria included the presence of severe obstructive pulmonary disease, inability to cough with sufficient force to produce a sputum sample, angina on slight exertion, and pregnancy.

[0229] Exclusion criteria included the presence of severe obstructive pulmonary disease, inability to cough with sufficient force to produce a sputum sample, angina with minimal effort, and pregnancy.

[0230] Sputum samples: Sample donors were trained on how to use a throat clearing machine assistive device (Smiths Medical, St. Paul, MN), repeating this process at home for three consecutive days and storing the sample cup in a cool, dark place or refrigerator. Within one day of collection, samples were shipped overnight to bioAffinity laboratories for further processing and FCM analysis by personnel who were blinded to the source of the samples.

[0231] Sputum processing: Sputum is separated and labeled. For example, sputum samples are incubated with a mixture of 0.1% dithiothreitol and 0.5% N-acetyl-L-cysteine ​​for 15 minutes at room temperature and then neutralized with Hank's balanced salt solution. The cells are then filtered out through a 100-micron nylon filter, washed and resuspended in HBSS. The total cell yield is determined using the trypan blue exclusion method. Sputum is liquefied using preheated 0.1% dithiothreitol (DTT) (ratio of 1:4 (w / v) to sputum weight) and preheated 0.5% N-acetyl-L-cysteine ​​(NAC) (ratio of 1:1 (w / v)). The resulting cell suspension is filtered through a 100 μm nylon cell strainer (Falcon, Corning Inc.) to eliminate larger debris while minimizing cell loss. The cells are collected in a 50 mL conical tube, washed and centrifuged at 800 x g for 10 minutes. For each sputum sample, separate sputum pellets were combined into a 15 mL conical tube. Total cell yield and viability were determined using a Neubauer hemacytometer and trypan blue exclusion.

[0232] In one embodiment, a small aliquot of cells is kept as a control, and most of the cells are divided into two tubes for primary analysis. For example, both tubes are labeled with fixable viability stain 510 (FVS510) and CD45-PE. One tube (the so-called "blood tube") received CD66b-FITC, CD3-Alexa-Fluor-488, CD19-Alexa-Fluor-488, and CD206-PE-CF594. In another tube (the "epithelial tube"), cells were labeled with pan-cytokeratin-Alexa-Fluor-488 and EpCAM-PE-CF594. The cells were incubated on ice for 35 minutes. After washing with HBSS, the cells were fixed and stored on ice until the next day, and then TCPP solution (20 μg / mL) (3.3x10 6 After incubation on ice for 1 h, cells were washed twice with cold HBSS and kept on ice until analysis.

[0233] In another embodiment, cell labeling is performed by dividing the sample into at least two tubes: one tube includes a sample for examining white blood cells (CD45 + ) cell compartment markers, one tube for epithelial (CD45 -) cell compartment. Each tube contains anti-CD45 antibody, FVS510 (for exclusion of dead cells, including SEC), and porphyrin TCPP (for identification of cancer (related) cells). To identify leukocyte populations, anti-CD206 antibody was added to mark macrophages, and a cocktail of antibodies was added to mark granulocytes (anti-CD66b) and lymphocytes (anti-CD3 and anti-CD19). For epithelial cell identification, we used anti-cytokeratin (panCK) and anti-EpCAM. No permeabilization step was performed for cytokeratin labeling because the initial DTT and NAC treatment used for sputum processing was sufficient for intracellular cytokeratin staining.

[0234] Isolated sputum cells were incubated with antibodies and FVS510 for 35 minutes. After washing once with cold HBSS, cells were fixed with paraformaldehyde on ice for one hour, after which cells were washed again and stored on ice until TCPP labeling the next day. TCPP was added to cells for one hour. After incubation, cells were washed twice with cold HBSS and then stored on ice until flow cytometric analysis. Cells were kept on ice and protected from light throughout the labeling process until analysis. See Table 7 for more detailed information on reagents.

[0235] Flow cytometry: Sputum samples were collected on a BD LSR II flow cytometer (BD Biosciences) equipped with 4 lasers (404 nm, 488 nm, 561 nm, and 633 nm) or a Navios EX (Beckman Coulter Life Sciences) equipped with 3 lasers (405 nm, 488 nm, and 638 nm). Post-collection data analysis was performed using FlowJo software (Tree Star, Inc., Ashland, OR).

[0236] Sample characteristics: Of the 171 patient LSRII samples collected, 150 were sufficient for analysis by the full assay pipeline, 122 were from high-risk patients without cancer, and 28 were from lung cancer patients (Table 1). An additional 4 samples for which we did not have clear disease status were included in the pipeline development phase, as the addition of unlabeled samples has been shown to aid in model building. In addition, 14 samples that were flagged as failing based on counts (see Lung Assay Pipeline below) were used in the model fitting phase to better capture the distribution of the underlying data and help the generalization of the model be more robust to sample noise. Only 3 samples could not be used at all due to issues with the collection process.

[0237] Table 1. Patient characteristics of the LSRII lung assay validation sample

[0238]

[0239]

[0240] n = number of samples

[0241] SD = Standard Deviation

[0242] Traditionally, the presence of "a large number" of macrophages in a sputum smear indicates that the sample is from the lungs. According to one embodiment of the present invention, a quality control measure using the cell surface antigen CD206 is utilized in the FCM lung assay, which is specific to a population of macrophages present in lung tissue but not found in the blood circulation. In one embodiment, sputum cells are stained with i) a cell marker specific for CD45, such as an antibody to CD45 (to identify white blood cells), ii) a cell marker specific for CD206, iii) a cell marker specific for CD66b, iv) a cell marker specific for CD3, and v) a cell marker specific for CD19. In one embodiment, any combination of i)-v) can be combined and added to the sputum sample, such as a cocktail of antibodies consisting of anti-CD66b, anti-CD3, and anti-CD19 compounds, wherein, for example, the compound is an antibody or fragment thereof, to further separate macrophages from other hematopoietic cells. In one embodiment, FVS510 was used as a viability dye to exclude dead cells. Anti-CD45PE signal demonstrated that a certain proportion of live sputum cells specifically expressed CD45. Cytospin of sorted CD45+ sputum cells confirmed their hematopoietic origin.

[0243] Further analysis of sputum samples by FCM showed that sputum-derived leukocytes (CD45 + Cells) include different macrophage subsets. Representative light scatter plots of unstained single live sputum cells are shown for cells selected by size exclusion gates and live cell gates and dual state discrimination gates, and CD45+ and CD45- gates are defined for sorting and further analysis. Cells that fall into the live, single sputum cell gate and are stained with a blood antibody panel are further analyzed.

[0244] The sputum-derived leukocyte profiles of FVS510-CD45+ cells from different samples stained with a blood antibody panel were further analyzed. Gates were used to identify lymphocytes / granulocytes (Gate 1), as well as alveolar macrophages (Gate 2) and interstitial macrophages (Gate 3) based on cell lineage-specific markers (fluorescence) and / or optical properties of specific cell types (FFS / SSC).

[0245] In this embodiment, fluorescence minus one (FMO) controls were analyzed using the same gates as those used for leukocyte subsets defined by the blood antibody panel. All FMO controls included viability dye, CD45, and TCPP.

[0246] Sputum cells stained with the leukocyte antibody panel minus CD66b, CD3, and CD19 antibodies were analyzed. Sputum cells stained with the leukocyte antibody panel minus CD206 antibodies were also analyzed. Wright-Giemsa-stained cytospins from the sorted CD45+ gate 2 and gate 3 populations were examined under a microscope for visual identification and measurement. Cells in the gate 2 population measured approximately 20 μm and cells in the gate 3 population measured approximately 10 μm.

[0247] The pathologist confirmed the cell type and measured the cell size of the sorted macrophage populations in gates 2 and 3. For each population, at least 100 cells were measured. The mean cell size in gate 2 was 16 μm + / - a standard deviation of approximately 3-5 μm (****p<0.0001).

[0248] FCM profiles of CD45+ sputum cells captured with anti-CD206 antibody and a mixture of anti-CD66b, anti-CD3, and anti-CD19 antibodies (see Fig. 12C ). The background staining of the isotype control was higher than that of the unstained or fluorescence minus one (FMO) control. Since the use of isotype control antibodies presents its own set of problems, we used FMO controls to identify the major subpopulations in sputum. By comparing the FMO control of the CD66b / CD3 / CD19 mixture to the stained sample including all antibodies, Gate 1 can be set to identify combined lymphocytes and granulocytes. Similarly, by comparing the FMO control of the CD206 antibody, two populations of CD206-positive cells can be identified.

[0249] After sorting cells from gates 2 and 3, cytological analysis revealed populations with morphology consistent with macrophages. However, the size of cells sorted from gate 3 was significantly smaller than that of cells sorted from gate 2. The sizes we calculated for the alveolar macrophage population (gate 2) and interstitial macrophage population (gate 3) were consistent with previously reported size ranges.

[0250] Alveolar macrophages were identified in gate 2 of the FITC channel as strongly positive for CD206 and autofluorescent, whereas interstitial lung macrophages were smaller in size and lower in gate 3 for CD206 expression.

[0251] In the combined two gates, the average background staining in the CD206FMO control was 0.0023% (+ / - SD 0.0021%). The positive threshold for the combined two gates was set to 0.0065%, or approximately 6 macrophages per 100,000 cells, based on a positive threshold of 2 standard deviations (SD) above the average background staining. Due to concerns that the low threshold would not fall within the linear detection range of the PE-CF594 fluorescent dye, we chose an arbitrary threshold of 0.05% that included both alveolar macrophages and interstitial macrophages. This threshold cannot be based solely on interstitial macrophages. The 0.05% threshold is well within the detection linear range of the flow cytometer and meets the criteria for "abundant macrophages" for an adequate sample set by the Papanicolaou Society.

[0252] 179 samples were analyzed for macrophage content. Based on the above criteria, 15 samples were found to have insufficient macrophage counts. However, six of these samples (3.4%) had fewer than 1000 CD45+ events for analysis, which was too few cells to adequately analyze based on the limits set according to one embodiment of the present invention. Five of these six samples had less than 1.5 x 10 total sputum cells before antibody staining. 6 The remaining 9 samples (5.0%) contained more than 10,000 CD45+ cells (range 11648-463382), and all samples showed more than 1.7 x 10 6 Furthermore, only 4 of the 164 adequate samples showed fewer than 1.5 x 10 6 Although these samples all included robust macrophage counts, three of the four showed fewer than 10,000 D45+ cells (range 1327-2908). This data suggests that the samples contained fewer than 1.5 x 10 6 Sputum samples containing more than 10 cells are too small to be analyzed by reliable flow cytometry. In one embodiment of the present invention, a sputum sample used in a method comprises more than 1x10 5 cells, or 1x 10 6 cells or greater than 2 x 10 6 cells or greater than 10 x 10 6 Slides used for cytological assays typically accommodate only 3 x 10 5 cells / slide and therefore cannot accommodate the number of cells in a sputum sample required to characterize the cell types and features as detailed in the methods of the present invention.

[0253] Before antibody labeling, the total number of sputum cells (excluding SEC) was counted for individual samples. All adequate samples (n=164) showed >0.05% macrophages (combined alveolar and interstitial). Among the adequate samples, 18 samples had cell counts exceeding 50 million cells. The median cell count in the adequate samples was 14.6 x 10 6 Insufficient samples (n=15) either showed no alveolar macrophages or had <0.05% combined events in the alveolar and interstitial macrophage gates. The median number of cells in the insufficient samples was 6.9 x 10 6 A subset of insufficient samples contained “too few cells” to obtain a reliable profile (<1000 CD45+ events), while the remaining samples contained sufficient cells but did not meet the QC macrophage criteria, so they could not be considered sufficient samples. The median cell count for the insufficient samples was 1.1 x 10 6 cells.

[0254] 164 adequate sputum samples were further analyzed to determine the difference between cancer "CA" and non-cancer "non-CA" (also known as high risk). This group included 32 samples obtained from individuals diagnosed with lung cancer and 132 samples from high-risk individuals without cancer. The cancer group included 40.6% current smokers, and the high-risk group included 44.7%. There were no significant differences in smoking pack years between the groups. There was also no significant difference in the average number of years that former smokers had quit smoking. The proportion of women in the cancer group was less than that in the high-risk group (21.9% and 54.5%, respectively). The average age of participants in the cancer group was 69.8 years old, while in the high-risk group it was 64.8 years old (p<0.0002).

[0255] Reference now Fig. 12B In the first phase of the analysis, CD45 + (Top box "+") vs. CD45 - The proportion of cells (bottom box “-”) and the various subpopulations without TCPP markers in each compartment. We found that CD45 + The proportion of cells was significantly higher in sputum from non-cancer (“non-CA”) / disease-free high-risk patients (49.64% vs. 38.95%; p=0.0099). CD45 was identifiable in all samples. + Different subpopulations of the compartment, however, the relative contribution of each population varied between samples and between groups. + The relative sizes of the cell subsets were analyzed and we found that cancer samples contained significantly more granulocytes / lymphocytes (see Fig. 12C , gate 1 p = 0.0378) and interstitial macrophages (see Fig. 12C, gate 3p = 0.0031), and Fig. 12C CD45 in Gate 2 + A subset of cells were alveolar macrophages and were positive for CD206.

[0256] Fig. 12B The CD45 - The cell populations identified in the compartment included cells of epithelial origin, and when CD45 - This was confirmed by the presence of goblet cells and ciliated epithelial cells when the cells were sorted and the morphology was observed on a cytospin. The use of antibodies against EpCAM and cytokeratin allowed us to further characterize CD45 by flow cytometry. - Cell populations. FMO controls show low background for the corresponding antibodies used. - The relative contributions of cell subsets varied between samples, and no significant differences were observed between the cancer and high-risk groups. Fig.12D Compartment 4 identified CD45 cells that were positive for EpCAM and panCK - Cell subsets (“panCK + EpCAM + ”).

[0257] Live single CD45 from different samples - Sputum cells were stained with epithelial cell markers such as PanCK and EpCAM (epithelial profile) / epithelial antibody panel. Fluorescence minus one (FMO) controls for this profile were obtained by FCM. FMO controls included viability dyes, CD45, and TCPP. FVS510-CD45 staining with isotype controls for the antibodies used was obtained. - Sputum derived epithelial profile of cells (unstained sputum cells). FMO control FVS510-CD45 stained with EpCAM but not with panCK antibody profile was obtained. - FMO control FVS510-CD45 cells stained with panCK but not EpCAM antibody were obtained - cell.

[0258] TCPP fluorescence was observed in the second stage of FCM analysis. Single live cells were divided into three cell subsets based on TCPP staining intensity: TCPP-high cells, TCPP-intermediate (IM) cells, and TCPP-low cells (see Fig.13A and Fig. 13B The relative ratios of these cell subsets did not differ between the high-risk group and the cancer group. We then further examined the CD45 + Leukocyte populations and CD45 -Epithelial cell population content (see Fig. 14B , Fig.14F and Fig.14J "+" and "-" compartments).

[0259] CD45+ compartment of TCPP-high cells (see Fig. 14B "+" cell population, in Fig. 14C further analyzed) are enriched in alveolar macrophages (CD45 + ;CD206 ++ Cells) (see Fig. 14C ), while CD45 - Compartment (see Fig. 14B “-” cell population, in Fig.14D EpCAM was enriched in + ;panCK + Double positive cells (upper right compartment). TCPP-IM cells represent the majority of sputum cells, so the profile of this subpopulation is similar to that of the entire sample (see Fig.14E -H). TCPP-low cells display relatively lower light scattering properties compared to TCPP-high cells (see Fig.14I ), and most of them are CD45 - (See Fig.14J ), do not express the epithelial markers EpCAM or panCK (see Figure 14L ). TCPP IM cells represent the majority of sputum cells, so the profile of this subpopulation is similar to that of the entire sample ( Fig.14E -H). With the whole sample (or TCPP IM cells), TCPP 高 and TCPP 低 Cells showed different profiles. TCPP 高 Cells display a broad light scatter profile (see Fig.14A ), and TCPP 高 CD45 + Compartment (see Fig. 14C ) is rich in CD45 + ;CD206 ++ cells (i.e. alveolar macrophages), while TCPP 高 CD45 - EpCAM-enriched compartment + ;panCK + Double positive cells (see Fig.14D ). and TCPP 高 and TCPP IM Compared with cells, TCPP 低 Cells exhibit relatively low light scattering properties (see Fig.14I), and most of them are CD45 - (See Fig.14J ), do not express the epithelial markers EpCAM or panCK (see Figure 14L ).

[0260] By FCM, sputum cell populations with different TCPP fluorescence intensities were identified based on gating of different cell lineage markers. Fig.13A ) and pan-cytokeratin (panCK) in "epithelial tubes" to define a TCPP-high cutoff (identified by the upper bold box / compartment labeled "H"). Dot plots of TCPP versus PE-CF594 fluorescence can also be used for this purpose, but in the former, cells with the highest FI for TCPP are more easily identified.

[0261] Reference now Fig. 13C , the TCPP-high ("H") cutoff value was taken from the gate on the sputum cell dot plot, where the y-axis is TCPP and the x-axis is CD66b / CD3 / CD19. When unstained sputum was overlaid with TCPP-stained samples, the TCPP-low ("L") population was defined at the intersection. The TCPP-intermediate ("IM") population was defined as a population between the TCPP-high and TCPP-low populations.

[0262] The unique properties of the TCPP-high population revealed some significant differences between the high-risk and cancer groups. First, TCPP-high cells from the cancer group samples showed lower side scatter values ​​than those of the high-risk group (see Fig.15A ). Secondly, CD45 - Compartments contain a higher percentage of EpCAM + panCK + Cells (see Fig. 15B ). In addition, the double positive population from the cancer group samples expressed higher levels of EpCAM, but not panCK, compared to cells belonging to the same quadrant of the high-risk group samples (see Fig. 15C ).

[0263] Differences in sputum cell characteristics between cancer and high-risk sputum samples were identified, with the TCPP-high population in cancer samples showing a smaller SSC than the TCPP-high population in high-risk samples (**p<0.01). In cancer samples, the CD45 - EpCAM + panCK + The proportion of cells was greater than the corresponding CD45 -Partial (**p<0.01). TCPP-high CD45-EpCAM in cancer samples + panCK + The mean fluorescence intensity (MFI) of EpCAM in cells was higher than that in the corresponding cell subsets of high-risk samples (*p<0.05).

[0264] Upon further analysis of the cancer groups, significantly higher mean fluorescence intensity of EpCAM was observed in early stage cancer samples (stage I / II) compared to late stage cancer samples (stage III / IV) (p=0.047). No significant differences based on cancer type (squamous cell carcinoma vs. adenocarcinoma) were identified, nor were any differences based on smoking history (current smokers vs. former smokers). Interestingly, when we separated high-risk smokers based on smoking history, data from current high-risk smokers showed the presence of significantly more TCPP-high; EpCAM+; panCK+ cells (p=0.0008) as well as macrophage populations (including alveolar (p=<0.0001) and interstitial (p=0.0141)) compared to former smokers.

[0265] Referring now to FIG. 16, a diagram depicting Fig. 12A - Significant differences between cancer (CA) and non-cancer samples (non-CA) derived from blood cell populations described in C. Each dot (CA) and square (non-CA) represents one sample. Fig.16A CD45 in sputum samples from cancer samples (CA) is shown. + The proportion of cells was significantly higher than that in non-cancerous samples (**p=0.0099). Fig. 16B Shown in CD45 + In the sputum samples obtained from cancer patients, granulocyte / lymphocyte subsets ( Fig. 12C Gate 1) in was significantly larger (*p=0.0378). Fig. 16C Figure 3 shows the expression of CD45 in interstitial macrophages in sputum samples from cancer patients compared with sputum samples from non-cancer patients. + Subgroups ( Fig. 12C The median value for each sample group is indicated by the thick black horizontal bar.

[0266] Current methods for sputum analysis present challenges that limit its clinical application. Sputum cytology has low sensitivity due to the high skill required to identify subtle nuclear changes. The need to screen a large number of slides makes it very time-consuming, which also hinders its clinical application. Imaging and molecular techniques can assess genetic changes in sputum-derived cells, but screening methods based on nuclear ploidy or in situ hybridization to detect genetic abnormalities use only a few hundred cells per sputum sample, while microarray analysis of enriched epithelial cells derived from genetic aberrations analyzed in sputum screens only 2000 cells per slide. Excluding most sputum cells from analysis may hide important disease parameters, resulting in sensitivity that is lower than clinically useful. The limitations of these different techniques should not be confused with the highly useful nature of sputum as a biological fluid that can provide an important cellular snapshot of the lung environment.

[0267] Flow cytometry platforms are well suited for analyzing exfoliated cells isolated from sputum to identify tumor-related changes in leukocyte and non-leukocyte populations that cannot be detected by traditional cytological methods. FCM is very powerful in its ability to detect and analyze cells based on the physical properties of cells (i.e., size and granularity) and cell surface molecules. Unlike microscopy or cytology, flow cytometry can analyze a large number of cells in a short period of time. The variability of autofluorescence and nonspecific binding properties of cell populations within and between sputum samples prohibits the use of commercially available biological controls, which are typically used for the immunophenotype of highly characterized hematopoietic populations. Therefore, internal FMO controls have been used to establish positive thresholds for macrophage gates. The ability to identify alveolar macrophages as a unique leukocyte subset allows us to incorporate built-in flow cytometry quality control parameters to determine the lung origin of each sputum sample. According to one embodiment of the present invention, sample quality confirmation based on cytology is required to ensure quality control.

[0268] The lungs are constantly exposed to pathogens and harmful particles. Alveolar macrophages are the main primary innate defense that maintains a healthy lung environment. Alveolar macrophages are characterized by a unique CD45 + The group with high CD206 expression (CD206 ++ ), and have a moderate to high signal on the granulocyte / lymphocyte axis due to their autofluorescence. Furthermore, the results confirm previous observations where the light scatter profile of alveolar macrophages overlapped with that of contaminating SEC, highlighting the necessity to isolate SEC from further analysis.

[0269] CD206 - Intermediate positive cells (CD206 + ) are also macrophages, although they are smaller than alveolar CD206 ++The macrophages are small and show minimal FITC autofluorescence, suggesting that this population of macrophages likely represents interstitial macrophages. Although interstitial macrophages (in contrast to alveolar macrophages) do not normally contact the airway lumen, the pro-inflammatory environment induced by chronic smoking is ideal for interstitial macrophages to infiltrate the airways. Therefore, their presence in sputum obtained from heavy smokers is not unexpected. This is further supported by our finding that current high-risk smokers had significantly more macrophages in their sputum than former high-risk smokers.

[0270] In one embodiment of the invention, the minimum number of sputum-derived cells in a sputum sample for automated FCM is about 1.5 million cells to provide an adequate overview so that the presence of macrophages can be determined. A cutoff of 5 macrophages per 10,000 cells (0.05%) for determining sample adequacy is well within the detection range of flow cytometry. Interstitial macrophages are included in the 0.05% macrophage cutoff for sample adequacy because both alveolar and interstitial macrophages are lung tissue-specific cell populations. The presence of interstitial macrophages without the presence of alveolar macrophages (rare) is difficult to explain biologically, so samples that do not contain any alveolar macrophages are considered inadequate.

[0271] Comparison of sputum samples from people diagnosed with lung cancer with those obtained from people at high risk for the disease, multi-parameter analysis revealed significant differences between the two groups. The cancer samples contained significantly more CD45+ cells than the high-risk samples, specifically more granulocytes / lymphocytes and interstitial macrophages.

[0272] Addition of the porphyrin TCPP to the staining protocol allowed identification of multiple significant differences in the brightest staining subset (TCPP-high) between the cancer and high-risk groups. TCPP-high cells in the cancer group, regardless of their CD45 lineage, exhibited lower side scatter characteristics than TCPP-high cells from the high-risk group, suggesting reduced cytoplasmic content, organelle degranulation, and vacuolization, which provide evidence of malignancy.

[0273] TCPP-high cells of non-leukocytes (CD45 - Analysis of the 2019-nCoV subpopulations showed that the cancer group contained a greater percentage of cells stained with the epithelial markers panCK and EpCAM. This difference from the high-risk group was mainly due to significantly fewer of these cells in the sputum of former high-risk smokers compared to current smokers. The epithelial cell subpopulation from the cancer group also expressed higher levels of EpCAM despite identical panCK levels. This was most pronounced in the stage I / II subgroup.

[0274] Historically, the detection of cancers of epithelial origin and circulating tumor cells has relied on the detection of EpCAM and cytokeratin expression. Our flow cytometry-based analysis found increased EpCAM expression in samples diagnosed with stage I / II cancer and in samples from high-risk participants who continued to smoke, suggesting that EpCAM expression may be of particular importance in the detection of early lung cancer.

[0275] According to one embodiment of the lung assay, analysis of light scatter and fluorescence signals from live single cells identified by automated FCM was determined. A logistic regression model established the relationship between the predictor variables and the categorical (in our case binary cancer / non-cancer) response variable. Stepwise regression is a supervised machine learning process by which potential predictive variables are added and removed, and the resulting model is examined for goodness of fit. Clinical factors for which complete data were available (Table 1) were included as potential predictors. Age was a clinical parameter that was repeatedly assessed as significant in the forward and backward stepwise regression processes.

[0276] The performance of the lung assay was evaluated on the 122 high-risk samples (also referred to herein as non-cancer NCs) and 28 cancer samples described in Table 1, as well as an additional 32 samples processed on a different FCM instrument (Navios EX) (Table 2). These 32 samples comprised a different set of patients than those used for assay development using the LSRII flow cytometer. The same model with the same coefficients was used for both instruments, but the cutoff for the Navios samples was 0.5 instead of 0.28. The results shown in Table 3 show that the lung assay performed very well, with sensitivity, specificity, and accuracy all >80% for the LSRII samples and very similar numbers for the smaller Navios EX sample set. Very robust negative predictive values ​​(NPVs) were obtained for both platforms. > 95%.

[0277] Table 2. Patient characteristics of the Navios EX validation samples

[0278]

[0279]

[0280] n = number of samples

[0281] a Show individual values ​​instead of mean (SD)

[0282] Table 3. Lung Assay Performance

[0283]

[0284] a0.83% reported in NLST 2013. 1

[0285] b If the assay was used only for NLST 2013 LDCT-positive cases, 2.9%.

[0286] c Sensitivity / (1-specificity), see Pepe et al. 2

[0287] Using R package bdpv 3 According to Mercaldo et al. 4 Calculate confidence intervals for prevalence.

[0288] For cases in which LDCT did not detect nodules ≥20 mm in diameter (Table 3, “all nodules <20 mm”), the lung assay also performed very well, with a sensitivity of 92%, a specificity of 87%, and an area under the ROC curve of 94%. In addition, the lung cancer assay performed well for all tumor types and in all disease stages, including I and II (Tables 4, 5).

[0289] Table 4. Performance of lung assay by tumor type and stage (LSRII)

[0290]

[0291] n = number of samples

[0292] N / A = Information not available

[0293] Table 5. Performance of lung assay by tumor type and stage (Navios EX)

[0294]

[0295] NA = Information not available

[0296] a A biopsy was not performed due to comorbidities. However, the patient was considered to have lung cancer.

[0297] Each of the retained predictors contributed significantly to the model (Wald test p-value < 0.05), and removing them individually had a negative impact on the ability to correctly classify cancer and high-risk samples (Table 6). Age is a well-established clinically relevant factor for lung cancer. 31 As in our model; nevertheless, the correlation between age and model values ​​was not overwhelming in either the LSRII or Navios EX samples (Figure 8), with some younger patients being called “cancer” and many older patients being called “non-cancer.”低 CD3 / CD19CD66b 中 Number of samples misclassified by signal vs. exclusion age and its relationship with FVS510-A / log 10 The FSC-A R2 interaction resulted in the same number of misclassified samples (Table 6).

[0298] Table 6. Effect of model predictors on classification

[0299]

[0300] a The complete model is shown in Figure 6

[0301] b 150LSRII samples from Table 1

[0302] c Including the interaction term age: FVS510-A / log10FSC-A R2

[0303] Table 7. Reagents for sputum staining and flow cytometry analysis according to one embodiment of the present invention

[0304]

[0305]

[0306] h = human; panCK = pan-cytokeratin; m = mouse; all antibodies are monoclonal *For research purposes, 10 μm, 40 μm, and 50 μm large bead NIST particle size standards were also used (S2 Fig).

[0307] **Concentration not determined; used according to manufacturer's protocol. All reagents were titrated using sputum from persons at high risk for lung cancer.

[0308] One aspect of the present invention provides an automated flow cytometry system and method for analyzing with machine learning to predict the presence of lung cancer from a sputum sample. One hypothesis is (but not limited to) that sputum as diagnostic material provides a snapshot of the tumor itself, its microenvironment (ME), and its field of canceration (FoC). Expert cytological analysis of sputum can detect cancer cells and precancerous cells, but this is an extremely laborious method that is not suitable for large-scale screening without automation, is prone to observer bias, and cannot examine a large number of cells in a sample in a few seconds because cytological samples are limited to the size of the slide, thereby limiting the number of cells to be analyzed. Automated image processing has been used to capture changes in cells associated with malignancies with some success, but is still hampered by technical complexity and the small number of cells analyzed.

[0309] Another aspect of the invention provides systems and methods for analyzing the presence of cancer cells in biological samples such as sputum by high-throughput, automated flow cytometry-based methods combined with machine learning to provide the following benefits: a) the assay can be put into routine laboratory use without the need for specialized sample evaluation or subject to operator bias; b) the entire sputum sample can be analyzed rapidly; and c) the numerical analysis can capture the complex interactions between lung cancer, ME, and FoC cells that are difficult to reliably detect individually. For example, it was unexpectedly discovered during the development of the lung assay that the predictive value of the density of viability staining indicated an association with apoptosis. In addition, it was unexpectedly observed that specific markers of immune function provided useful information.

[0310] One aspect of one embodiment of the invention provides an automated, flow cytometry-based test that examines three aspects of tumorigenesis: TCPP staining, programmed cell death, and immune response. Other studies have shown that the performance of sputum-based tests for early lung cancer detection can be significantly improved when different types of measurements are combined, such as combining cytology with genetic mutations or microRNA and methylation biomarkers. Although we used the same technology platform to measure different cancer-related processes, these additional parameters may contribute to the performance improvement from slide-based assays to flow cytometry-based assays ( Figure 7 ). In addition, flow cytometry-based assays can read the entire sample, which can also improve test performance.

[0311] All study participants except one met the most recent lung cancer screening criteria published by the U.S. Preventive Services Task Force. Although our study group can be considered a sample of a population eligible for lung cancer screening (one of the target populations for lung cancer), the sample size was small and minorities were underrepresented, as were women in the cancer group. In addition, the cancer prevalence in both datasets in our study was slightly lower than 19%, which is much higher than that in the lung cancer screening population or the group of patients with lung nodules between 7 and 19 mm (another target population for lung determination).

[0312] In an official 2017 policy statement, the American Thoracic Society stated that molecular biomarkers should influence clinical management decisions in a manner that improves clinical outcomes to be considered clinically useful. The authors discuss a use case in which screening was expanded to include participants who are currently ineligible for LDCT screening, thereby reducing the prevalence of cancer from the NLST level of 1 / 120 to a hypothetical 1 / 500. They assumed a reasonable hazard threshold of 0.83% based on the NLST data, resulting in a minimum positive diagnostic likelihood ratio (PDLR) of 4.18, which was achieved by the larger LSRII group (Table 3). Using an assumed prevalence of 1 / 400 instead of 1 / 500 and using the same hazard threshold would yield a PDLR of 3.35, a criterion met by both the LSRII and Navios groups. The pulmonary assay systems and methods disclosed herein can help expand early lung cancer screening to relatively underserved populations, such as young women and male African American smokers.

[0313] Lung determination can also support clinical decision making in patients with a positive LDCT and an intermediate-sized nodule, potentially in conjunction with a risk calculator such as the one from Brock University. Below 7 mm, only 2% of NLST patients underwent invasive follow-up, and above 20 mm, immediate follow-up may be warranted out of an abundance of caution, although a pan-Canadian study found that the largest nodules were not malignant in 20% of participants. Intermediate-sized nodules, however, are very challenging to follow. If we estimate the risk threshold (R) above which invasive follow-up is worthwhile to be the frequency of cancer in the NLST population with nodules 7–19 mm in diameter (4.8%), and assume a 3.8% prevalence of cancer in the LDCT-positive population, then sensitivity / (1-specificity) needs to be ≥[(1-prevalence) / prevalence] x R / (1-R) ​​= [(1-0.038) / 0.038] x 0.048 / (1-0.48) = 1.28, 44 Our assay can easily meet this threshold (Table 3, PDLR).

[0314] One aspect of the pulmonary assay according to one embodiment of the present invention is a non-invasive, sputum-based test for detecting early-stage lung cancer. It uses a flow cytometer platform to analyze the cellular content of sputum, and the analysis process is fully automated and therefore unbiased. Test performance in cases of small nodules (<20 mm) showed 92% sensitivity and 87% specificity.

[0315] When providing a numerical range, it should be understood that unless the context clearly dictates otherwise, each intermediate value between the upper and lower limits of the range (to one tenth of the lower limit unit) and any other specified value or intermediate value within the specified range are included in the present invention. The upper and lower limits of these smaller ranges may be independently included in the smaller range and are also included within the scope of the present invention, but are subject to any explicitly excluded limitations within the specified range. When a specified range includes one or two limits, the range excluding one or both of the limits is also included in the present invention.

[0316] Certain ranges are indicated herein by numerical values ​​preceded by the term "about." The term "about" is used herein to provide literal support for an exact number followed by it and for a number that is close to or approximately the number followed by the term. In determining whether a number is close to or approximately a specifically recited number, the close or approximately unrecited number may be a number that provides a substantial equivalent to the specifically recited number in the context in which it is presented.

[0317] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are now described.

[0318] It is noted that, as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should also be noted that the claims can be drafted to exclude any optional elements. As such, this statement is intended to serve as antecedent basis for the use of exclusive terminology such as "only," "only" and the like in connection with the recitation of claim elements or the use of a "negative" limitation.

[0319] It will be clear to those skilled in the art after reading this disclosure that each individual embodiment described and illustrated herein has discrete components and features that can be easily separated or combined with the features of any of the other several embodiments without departing from the scope or spirit of the invention. Any method described can be performed in the order of events described or in any other order that is logically possible.

[0320] Although the systems and methods have been or will be described for purposes of grammatical fluency and functional explanation, it should be expressly understood that unless expressly expressed pursuant to 35 U.S.C. §112, the claims should not in any way be interpreted as necessarily subject to any limitations of the “method” or “step” limitation construction, but should be given the full scope of meaning and equivalents provided by the claims in accordance with the doctrine of judicial equivalence, and in the event that a claim is expressly expressed pursuant to 35 U.S.C. §112, it should be given full legal equivalents pursuant to 35 U.S.C. §112.

[0321] Methods for classifying flow cytometry data

[0322] As described above, aspects of the invention include methods for classifying flow cytometric data. "Flow cytometric data" refers to information about the properties of sample particles (e.g., beads or cells or debris) collected by any number of detectors in a particle analyzer. As described herein, a "particle analyzer" is an analytical tool (e.g., a flow cytometer) that enables the characterization of particles based on certain (e.g., optical) parameters. "Particle" refers to discrete components of a biological sample, such as molecules, analyte-bound beads, single cells, etc.

[0323] The purpose method includes determining parameters (such as fluorescence) based on events (such as particles) in the sample, and one or more colony clusters are classified. As used herein, "group" or "subgroup" of events (such as cells or other particles) generally refers to an event group with properties (such as optics, impedance or time properties) about one or more measurement parameters, so that the parameter data measured form a cluster in the data space. The data obtained by analyzing the flow cytometry from cells (or other particles) are generally multidimensional, wherein each cell corresponds to a point in the multidimensional space defined by the measured parameters. In a plurality of embodiments, data are composed of signals from a plurality of different parameters, such as 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, and include or more. Therefore, the group is regarded as a cluster in the data. On the contrary, each data cluster is usually interpreted as a group corresponding to a cell or particle of a particular type, although clusters corresponding to noise or background are usually also observed. A cluster may be defined in a subset of dimensions, such as a subset with respect to measured parameters (eg fluorescent dyes), corresponding to groups that differ only in a subset of measured parameters or features extracted from the sample measurements.

[0324] Aspects of the present method include receiving a first gate with clear boundaries. As described herein, "gate" generally refers to the classifier boundary of identifying a subset of target data (data represents the characteristics or properties of particles / cells in a sample). In flow cytometry, a gate can limit the group (i.e., group) of specific target events. In other words, a gate defines the boundary for classifying a flow cytometry data group. In multiple embodiments, gate identification exhibits flow cytometry events of the same or similar parameter set. An event is a cell or particle detected by a sensor when a cell or particle passes between a sensor and an interrogation light source of a flow cytometer. The optical characteristics or properties of an event detected by a detector / sensor can be analyzed for each event. Flow cytometry data analysis is based on the gating principle. Gates and regions are placed around event groups with common characteristics, typically forward scatter (FSC), side scatter (SSC) and / or cell surface or intercellular marker expression, to further investigate and quantify these groups. Gating refers to selecting a continuous cell subpopulation to be analyzed in flow cytometry, and is the process of separating a specific target group in a heterogeneous sample. This allows the light scatter (FSC and SSC) and fluorescence properties of the population of interest to be highlighted in all available dot plots, thereby increasing the specificity of the analysis.

[0325] In some embodiments, the first gate is a gate drawn by a trained algorithm. In such embodiments, the trained algorithm can define the boundaries of an area (e.g., in two-dimensional space) within which flow cytometry data can be assigned a particular classification. For example, drawing the first gate can include superimposing a polygon onto a two-dimensional graph representing the flow cytometry data. For example, the first gate can be received from a database of gates that have been used in previous attempts to classify flow cytometry data.

[0326] In various embodiments, methods include receiving flow cytometer data, calculating parameters for each population, and gating the populations for further analysis based on a population of interest. For example, an experiment may include particles / cells labeled with multiple fluorophores or fluorescently labeled antibodies, and groups of particles may be defined by populations corresponding to one or more fluorescence measurements. In this example, a first group may be defined by a range of light scatter for a first fluorophore, and a second group may be defined by a range of light scatter for a selected population from the first group; and a third group may be defined by a third fluorophore based on a selected population of one or more of the first group, the second group, or a combination thereof.

[0327] Flow cytometer data can be received from any suitable source. In some embodiments, flow cytometer data is received from the memory of a storage device. In such embodiments, flow cytometer data may have been previously generated and stored in the memory of the storage device for subsequent call and analysis. In other embodiments, flow cytometer data is received in real time. In other words, the flow cytometer data generated during the operation of the flow cytometer may then (e.g., immediately) fill the data space (e.g., two-dimensional graph) with the first gate. In some cases, the flow cytometer can be operated to generate data until the recording criteria are met. The "recording criteria" discussed herein is a condition that, when met, causes the flow cytometer operation and data collection to terminate. Any suitable recording criteria can be used. In some cases, the recording criteria is a time limit. In the case where the recording criteria is a time limit, after a specified time (e.g., ranging from a few seconds to 3 hours), the flow cytometer data collection stops. In other cases, the recording criteria is the total number of events. In this case, after a certain number of particles (e.g., specified by the user) are analyzed, the flow cytometer data collection stops. In yet other cases, the recording criteria is the number of events within the group. In this case, flow cytometric data collection may stop after a certain number of particles (eg, specified by the user) in a particular population (eg, exhibiting a certain phenotype) have been analyzed.

[0328] In certain embodiments, particles are detected and uniquely identified by exposing the particles to excitation light and measuring the fluorescence of each particle in one or more detection channels, as desired. Fluorescence emitted in a detection channel used to identify a particle and its associated binding complex can be measured after excitation with a single light source, or can be measured separately after excitation with different light sources. If separate excitation light sources are used to excite particle tags, the tags can be selected so that all tags can be excited by each excitation light source used.

[0329] In a number of embodiments, flow cytometer data is received from a forward scattered light detector. The purpose forward scattered light detector can provide information about the overall size of the particles. In a number of embodiments, flow cytometer data is received from a side scattered light detector. The purpose side scattered light detector detects refracted and reflected light from the surface and internal structure of the particles, which tend to increase with the increase of particle structure complexity (e.g., particle size). In a number of embodiments, flow cytometer data is received from a fluorescence detector. The purpose fluorescence detector is configured to detect fluorescence emission from fluorescent molecules, such as a labeled specific binding member associated with the particles in the flow cell (e.g., a labeled antibody that specifically binds to a marker of interest). In certain embodiments, the method includes using one or more fluorescence detectors to detect fluorescence from a sample, such as 2 or more, such as 3 or more, such as 4 or more, such as 5 or more, such as 6 or more, such as 7 or more, such as 8 or more, such as 9 or more, such as 10 or more, such as 15 or more and including 25 or more fluorescence detectors.

[0330] The method in certain embodiments also includes data acquisition, analysis and recording, such as using a computer, wherein multiple data channels record data from each detector about the light scattering and fluorescence emitted by each particle when passing through the sample interrogation area of ​​the flow cytometer. In these embodiments, the analysis includes classifying and counting the particles so that each particle exists as a set of digital parameter values. The subject system can be set to trigger according to the selected parameters so that the target particles are distinguished from the background and noise or non-target cell populations. "Trigger" refers to a preset threshold for detecting a parameter and can be used as a means of detecting particles passing through a light source. Detection of an event exceeding the selected parameter threshold triggers the collection of light scattering and fluorescence data of the particles. Data of particles or other components in the detected medium that cause the response to be below the threshold will not be collected. The trigger parameter can be the detection of forward scattered light caused by the particle passing through the light beam. Then, the flow cytometer detects and collects the light scattering and fluorescence data of the particles. As needed, the data recorded for each particle is analyzed in real time or stored in a data storage device and an analysis device (such as a computer).

[0331] In at least one embodiment, and as will be readily appreciated by those of ordinary skill in the art, the apparatus according to the present invention will include a general or special purpose computer or distributed system programmed with computer software that implements the above steps, which may be in any suitable computer language, including R, Python, C++, C#, Perl, Java, PHP, HTML, MySQL, distributed programming languages, etc. The apparatus may also include a plurality of such computers / distributed systems in various hardware implementations (e.g., connected via the Internet and / or one or more intranets). For example, data processing may be performed by appropriately programmed microprocessors, computing clouds, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc., in combination with appropriate memory, network, and bus elements. Alternatively, containers may also be used.

[0332] Embodiments of the present invention provide a technology-based solution that overcomes the problems existing in the prior art in a technological manner, thereby solving existing problems for patients who may have early-stage cancer, healthcare providers, insurance companies, and diagnostic laboratories. Embodiments of the present invention are necessarily rooted in computer technology, such as computer learning. Embodiments of the present invention achieve important benefits over the prior art, such as greater flexibility, faster results, non-invasive procedures, automated screening of samples, etc. For example, thousands of cells from a biological sample can be analyzed and characterized using flow cytometry, and the data can be automatically analyzed in a matter of minutes to hours with high sensitivity and specificity, while it would be impossible for a human observer to obtain the same data in the same amount of time to achieve such speed and accuracy.

[0333] The previous examples can be repeated with similar success by substituting the generically or specifically described reactants and / or operating conditions of this invention for those used in the previous examples.

[0334] Please note that in the specification and claims, "about" or "approximately" means within twenty percent (20%) of the cited numerical value. All computer software disclosed herein may be embodied on any computer-readable medium (including combinations of media), including but not limited to CD-ROM, DVD-ROM, hard disk (local or network storage device), USB key, other removable drive, ROM, virtual machine, software container (such as Docker) and firmware.

[0335] Although the invention has been described in detail with particular reference to these embodiments, other embodiments can achieve the same results. Variations and modifications of the invention will be apparent to those skilled in the art, and the invention is intended to cover all such modifications and equivalents in the appended claims.

[0336] All publications and patents cited in this specification are incorporated herein by reference as if each individual publication or patent was specifically and individually indicated to be incorporated herein by reference, and are incorporated herein by reference to disclose and describe the methods and / or materials related to the publication cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. In addition, the publication dates provided may be different from the actual publication dates, which may need to be independently confirmed.

Claims

1. A flow cytometric method for automatically analyzing a sputum sample using a computer, the sputum sample being from a subject suspected of having lung cancer, the method comprising: obtaining a plurality of cells from the sputum sample from the subject suspected of having lung cancer; labeling the plurality of cells with i) a plurality of cell lineage-specific marker compositions, ii) a cell viability composition, and iii) a tetrakis(4-carboxyphenyl)porphyrin (TCPP) composition; analyzing the plurality of cells labeled with i-iii using the flow cytometer to obtain a subpopulation selected by cell size from the plurality of cells based on an automatically selected bead size exclusion gate; automatically selecting a live singlet cell population from the cell size-selected subpopulation by the computer using an automatic debris-free gate and an automatic singlet gate; automatically obtaining flow cytometry values ​​from the live singlet cell population by the computer based on the plurality of cell lineage-specific marker compositions, viability markers, and TCPP markers; automatically applying, by the computer, the trained classifier to the metadata from the subject and the obtained flow cytometry values; and Based on the application of the trained classifier, a classification of the sputum sample is automatically generated by the computer, wherein the classification is selected from a plurality of classification options including cancer and non-cancer.

2. The method according to claim 1, wherein: The sputum samples were single cell suspensions.

3. The method according to claim 1, wherein: The multiple cell lineage-specific markers are selected from CD206, CD3, CD19, CD66b, CD45, EpCAM, PanCK and any combination thereof.

4. The method according to claim 1, wherein: The cell viability composition preferentially labels dead cells over live cells.

5. The method according to claim 1, wherein: The cell viability composition is FVS510.

6. The method according to claim 1, wherein: The analyzing step includes obtaining the following flow cytometry values ​​from the plurality of cells: side scatter, forward scatter, fluorescence from TCPP, fluorescence from the cell viability composition, and fluorescence from the plurality of cell lineage-specific marker compositions.

7. The method according to claim 1, wherein: The plurality of cell lineage specific marker compositions are selected from fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19, fluorescent anti-CD66b, fluorescent anti-CD45, fluorescent anti-pan-cytokeratin, fluorescent anti-EpCAM and any combination thereof.

8. The method according to claim 1, wherein: The bead size exclusion gate was set between 5 μm and about 30 μm, wherein events smaller than about 5 μm and larger than about 30 μm were not further analyzed.

9. The method according to claim 1, wherein: The automated debris-free gate excludes most dead cells from the debris-free population.

10. The method according to claim 1, wherein: The automatic singlet gate is applied to the cell population selected in the automatic no-debris gate.

11. The method according to claim 7, wherein: The subject's metadata includes age.

12. The method according to claim 1, wherein: The sputum sample includes a minimum number of CD206 expressing cells such that the sputum sample is acceptable for determining lung health.

13. The method according to claim 11, wherein: The trained classifier is Wherein, the b0-b5 coefficients are determined by fitting the trained classifier to a plurality of sputum samples used to construct the classifier.

14. A system for automatically analyzing flow cytometry data, the system comprising: a computer processor in communication with a memory having stored therein flow cytometry data from a plurality of markers in a plurality of cells in a sputum sample of a subject, wherein the plurality of markers comprises i) a plurality of cell lineage-specific marker compositions, ii) a cell viability composition, and iii) a tetrakis(4-carboxyphenyl)porphyrin (TCPP) composition; A computer program product embodied in a non-transitory computer readable medium, the computer program product comprising instructions for causing the computer processor to automatically perform: receiving the flow cytometry data collected from the plurality of cells from a sputum sample; selecting a subpopulation of cells from the plurality of cells in the sputum sample, the subpopulation of cells being automatically selected based on applying an automated gate selected from a bead size exclusion gate, a viability gate, and a singlet gate; determining target flow cytometry values ​​of the plurality of cell lineage-specific marker compositions, the viability marker, and the TCPP marker from the subpopulation; applying a classifier to the flow cytometry value of interest and the subject's metadata; An output is generated on a display device wherein one or more classifications of the sputum sample are identified, including cancer or non-cancer.

15. The system of claim 14, wherein: The cell viability composition preferentially labels dead cells over live cells.

16. The system of claim 14, wherein: The cell viability composition is FVS510, and the metadata of the subject is age.

17. The system of claim 14, wherein: The flow cytometry values ​​are the flow cytometry values ​​obtained for side scatter, forward scatter, fluorescence from TCPP, fluorescence from the cell viability composition, and fluorescence from the plurality of cell lineage-specific marker compositions.

18. The system of claim 16, wherein: The plurality of cell lineage specific marker compositions are selected from fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19, fluorescent anti-CD66b, fluorescent anti-CD45, fluorescent anti-EpCAM and fluorescent anti-pan-cytokeratin.

19. The system of claim 14, wherein: The bead size exclusion gate was set to exclude events with a size less than about 5 μm and greater than about 30 μm.

20. The system of claim 14, wherein: The automated debris-free gate excludes most dead cells from the debris-free population.

21. The system of claim 18, wherein: The trained classifier is Wherein, the b0-b5 coefficients are determined by fitting the trained classifier to a plurality of sputum samples used to construct the classifier.

22. A non-transitory computer readable medium comprising program code which, when executed, causes a processing circuit to: automatically obtaining flow cytometry values ​​from a viable singlet population of a sputum sample from a subject based on side scatter, forward scatter, fluorescence from TCPP, fluorescence from a cell viability composition, and fluorescence from a plurality of cell lineage-specific marker compositions; automatically applying the trained classifier to the metadata from the subject and the obtained flow cytometry values; and Based on the application of the trained classifier, a classification of the sputum sample is automatically generated, wherein: The classification is selected from a plurality of classification options including cancer and non-cancer.

23. The non-transitory computer readable medium of claim 22, wherein: The cell viability composition preferentially labels dead cells over live cells.

24. The non-transitory computer readable medium of claim 22, wherein: The cell viability composition is FVS510.

25. The non-transitory computer readable medium of claim 24, wherein: The plurality of cell lineage specific marker compositions are selected from fluorescent anti-CD206, fluorescent anti-CD3, fluorescent anti-CD19, fluorescent anti-CD66b, fluorescent anti-CD45, fluorescent anti-EpCAM and fluorescent anti-pan-cytokeratin or any combination thereof.

26. The non-transitory computer readable medium of claim 22, wherein: A minimum number of CD206 positive cells were present in the sputum samples to be analyzed.

27. The non-transitory computer readable medium of claim 22, wherein: The bead size exclusion gate was set between 5 μm and 30 μm.

28. The non-transitory computer readable medium of claim 22, wherein: The automated debris-free gate excludes most dead cells from the debris-free population.

29. The non-transitory computer readable medium of claim 25, wherein: The metadata of the subject is age.

30. The non-transitory computer readable medium of claim 29, wherein: The trained classifier is Wherein, the b0-b5 coefficients are determined by fitting the trained classifier to a plurality of sputum samples used to construct the classifier.