Method for immunoprofiling and predicting a disease by using a staining kit and identifying immune cell subsets characterized for the disease

The staining kit and machine learning-based method effectively address the inefficiencies in NPC diagnosis by characterizing immune cell subsets and predicting NPC likelihood, achieving high diagnostic accuracy.

JP7696971B2Active Publication Date: 2025-06-23FULLHOPE BIOMEDICAL
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023172839
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-05
Filing Date
2023-10-04
Publication Date
2025-06-23
Estimated Expiration
2043-10-04

AI Technical Summary

Technical Problem

Current diagnostic methods for nasopharyngeal carcinoma (NPC) are inefficient, often leading to late-stage diagnoses due to the diverse symptoms and limitations of nasopharyngeal imaging methods.

Method used

A staining kit and a method that comprehensively compare immune cell subsets between NPC patients and healthy controls using flow cytometry and machine learning algorithms to identify disease-characterized immune cell subsets and predict NPC likelihood.

Benefits of technology

The method enables rapid characterization of disease-characterized immune cell subsets and efficient prediction of NPC likelihood, achieving high sensitivity, specificity, and AUC values in clinical classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696971000016
    Figure 0007696971000016
  • Figure 0007696971000017
    Figure 0007696971000017
  • Figure 0007696971000018
    Figure 0007696971000018
Patent Text Reader

Abstract

To comprehensively compare immune cell subsets between patients having a disease, such as nasopharyngeal carcinoma (NPC), and healthy controls in order to identify characterized immune cell subsets of the disease.SOLUTION: A staining kit is provided, including: a first pattern including antibodies against T cells, B cells, NK cells, monocytes, regulatory cells, CD8, CD45, and CTLA4; a second pattern including antibodies against T cells, B cells, NK cells, monocytes, regulatory cells, dendritic cells, and CD45; a third pattern including antibodies against T cells, B cells, NK cells, monocytes, CD8, CD45, CD45RA, CD62L, CD197, CX3CR1 and TCRαβ; and a fourth pattern including antibodies against B cells, CD23, CD38, CD40, CD45 and IgM, where the antibodies of each pattern are labeled with fluorescent dyes. Also provided are a method of identifying characterized immune cell subsets of a disease, and a method of predicting the likelihood of NPC in a subject in the need thereof using the staining kit.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a staining kit and its use, and in particular, to a staining kit, as well as a method for identifying immunocyte subsets with determined disease characteristics using the same and a method for predicting the likelihood of nasopharyngeal carcinoma (NPC) in a subject in need thereof.

Background Art

[0002] According to overall statistics on cancer incidence and mortality, nasopharyngeal carcinoma (NPC) is the 24th most common and lethal cancer in 2020, including 133,000 newly diagnosed subjects (0.7% of newly diagnosed subjects) and 80,000 new cancer deaths (0.8% of new cancer deaths). Although NPC is not a very lethal cancer, early diagnosis of NPC is still difficult. The diagnostic algorithm for NPC starts from symptoms in the neck, nose, or pharynx and is confirmed using nasopharyngeal imaging methods (such as magnetic resonance imaging, positron emission tomography, or computed tomography) or nasopharyngeal biopsy. However, due to the very diverse symptoms and the blind spots of nasopharyngeal imaging methods in the pharyngeal recess, patients with NPC are often diagnosed at an advanced stage (stage III or IV). Furthermore, advanced NPC has distant metastases to the brainstem, bone marrow, liver, and lungs, and although NPC cells are very sensitive to radiotherapy, there are concerns about the toxicity of chemotherapy and radiotherapy to these sites. Therefore, the prognosis of advanced NPC is much worse than that of early NPC.

[0003] Based on the practice guidelines from the European Society for Medical Oncology, the recommended diagnostic tool for NPC is nasopharyngeal imaging combined with biopsy after onset. [2] However, due to the high cost of fully complying with the research projects in the guidelines, more than two-thirds of recurrent NPC subjects are not diagnosed clinically. Therefore, oncologists are trying to use deep learning, convolutional neural networks, and narrow-band light observation to improve the sensitivity of nasopharyngeal imaging. [1]In addition to improvements in imaging techniques, the presence of anti-Epstein-Barr virus (EBV) antibodies or EBV cDNA is closely related to the risk of NPC, but not related to human papillomavirus (HPV). However, although about 95% of the human population is infected with EBV, less than 0.01% of the infected humans develop NPC. In addition, EBV infection is also associated with the incidence of Burkitt lymphoma, Hodgkin lymphoma, and gastric cancer. Therefore, neither anti-EBV antibodies nor EBV cDNA has the faithful accuracy to assist in NPC diagnosis or be an important indicator for NPC. That is, the development of a supported tool for NPC diagnosis is important.

[0004] The crosstalk between tumors and immune cells is a topic of interest that explains how tumor cells escape immune surveillance. Therefore, the comparison of immune profiles between cancer patients and healthy controls (HC) in terms of composition, metabolism, and transcriptome has been widely studied. For patients with breast cancer, colorectal cancer, and advanced hepatocellular carcinoma, their immune profiles are different from those of HC and are promising as leading indicators in cancer diagnosis. Regarding NPC, compared with HC, the high proportion of EBV antigen-specific CD8 + regulatory T cells and the low proportion of CD4 + T cells, IL-17-producing CD8 + T cells, and naive B lymphocytes. Except for this, the characterization of the immune profiles between NPC and HC is still unclear. In addition, recent studies on the changes in immunoprofiling in NPC have focused on transcriptome changes using single-cell sequencing or comparative immunoprofiling between peripheral blood mononuclear cells (PBMC) and tumor-infiltrating lymphocytes. [2] 。

Summary of the Invention

Problems to be Solved by the Invention

[0005] In view of the above description, the present invention aims to comprehensively compare immune cell subsets between patients with diseases such as nasopharyngeal carcinoma (NPC) and healthy controls in order to identify immune cell subsets with disease-defined characteristics. Furthermore, the present invention uses the characterized immune cell subsets to predict the likelihood of diseases such as NPC in a subject in need thereof.

Means for Solving the Problems

[0006] In one aspect, the present disclosure provides a staining kit. The staining kit includes a first pattern including antibodies against T cells, B cells, NK cells, monocytes, regulatory cells, CD8, CD45, and CTLA4; a second pattern including antibodies against T cells, B cells, NK cells, monocytes, regulatory cells, dendritic cells, and CD45; a third pattern including antibodies against T cells, B cells, NK cells, monocytes, CD8, CD45, CD45RA, CD62L, CD197, CX3CR1, and TCR αβ ; and a fourth pattern including antibodies against B cells, CD23, CD38, CD40, CD45, and IgM, wherein the antibodies in each pattern are labeled with a fluorescent dye.

[0007] Preferably, the T cells include CD3, CD4, CD25, CD45RO, CCR7, or any combination thereof; the B cells include CD10, CD19, CD21, CD127, IgG, or any combination thereof; the NK cells include CD56; the monocytes include CD14; the regulatory cells include PD-1, PD-L1, FoxP3, or any combination thereof; and the dendritic cells include CD11c, HLA-DR, or any combination thereof.

[0008] In another aspect, the present disclosure includes the following steps: (a) obtaining peripheral blood mononuclear cells (PBMC) and / or white blood cells (WBC) from a plurality of healthy controls and a plurality of patients with a disease, respectively; (b) staining the PBMC and / or WBC of the healthy controls and the patients, respectively, by using the staining kit, The staining kit includes a first pattern containing antibodies against T cells, B cells, NK cells, monocytes, regulatory cells, CD8, CD45, and CTLA4; a second pattern containing antibodies against T cells, B cells, NK cells, monocytes, regulatory cells, dendritic cells, and CD45; a third pattern containing antibodies against T cells, B cells, NK cells, monocytes, CD8, CD45, CD45RA, CD62L, CD197, CX3CR1, and TCR αβ ; and a fourth pattern containing antibodies against B cells, CD23, CD38, CD40, CD45, and IgM, wherein the antibodies in each pattern are labeled with a fluorescent dye, step (c) obtaining data on the fluorescence intensity of each antibody bound to PBMCs and / or WBCs of healthy controls and patients, respectively, by using flow cytometry, (d) identifying immune cell subsets in PBMCs and / or WBCs of healthy controls and patients, respectively, by using a phylogenetic method, and obtaining a data set containing data on the type and proportion of the immune cell subsets, and (e) evaluating the data set by using machine learning software and obtaining the type of immune cell subsets of patients that can be distinguished from those of healthy controls as disease-characterized immune cell subsets A method for identifying disease-characterized immune cell subsets is provided, which includes the above steps.

[0009] In yet another aspect, step (e) includes, by machine learning software, the following steps: (i) performing data preprocessing on the data set, (ii) performing feature selection from the preprocessed data by using the Boruta algorithm and obtaining pre-determined immune cell subsets of the disease as the selected features, and (iii) applying the data of the selected features and training a machine learning model by using at least one of a random forest (RF) algorithm, a logistic regression (LR) algorithm, and a support vector machine (SVM) algorithm and further includes performing the above steps.

[0010] The following steps: (a) Staining peripheral blood mononuclear cells (PBMCs) from a subject by using the above staining kit; (b) Obtaining data on the fluorescence intensity of each antibody bound to the characterized immune cell subsets of NPC by using flow cytometry, and obtaining a data set including data on the type and proportion of the characterized immune cell subsets; and (c) Evaluating the data set by using machine learning software to predict whether the subject has nasopharyngeal carcinoma A method for predicting the likelihood of nasopharyngeal carcinoma (NPC) in a subject in need thereof, comprising: The characterized immune cell subsets of NPC are selected from the group consisting of memory B cells, monocytes, T cells, naive CD4αβ T cells, PD-1 + CD4 T cells, PD-L1 + CD4 T cells, PD-1 + PD-L1 + Monocytes, CD4 NKTreg cells, MHC II + CD4 T cells and MHC II + CD4 NKT cells.

[0011] Preferably, step (c) further comprises the following steps performed by machine learning software: (i) Applying a holdout set having the characterized immune cell subsets to test the trained machine learning model; (ii) Predicting that the subject has nasopharyngeal carcinoma if the value of the predicted probability obtained from the RF algorithm or the LR algorithm is greater than a first threshold, or if the value of the decision function obtained from the SVM algorithm is greater than a second threshold.

[0012] Therefore, the present disclosure provides at least the following advantages: 1. The staining kit of the claimed invention can be applied to readily and rapidly characterize a disease-characterized subset of immune cells. 2. The staining kit of the claimed invention can be applied to efficiently predict the likelihood of a disease, such as NPC, in a subject in need thereof. 3. The present invention provides a novel disease diagnosis platform by constructing a powerful tool for supporting disease diagnosis by combining a staining kit with machine learning software such as flow cytometry and Python programs using RF, LR, and / or SVM algorithm analysis. 4. The performance of the method for predicting / diagnosing the disease of the claimed invention can reach high AUC (e.g., up to 0.98), sensitivity (e.g., up to 100%), and specificity (e.g., up to 90%) using RF, LR, and / or SVM algorithm analysis.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2a

Figure 2b1

Figure 2b2

Figure 2c1

Figure 2c2

Figure 2d1

Figure 2d2

Figure 2e1

Figure 2e2

Figure 2e3

Figure 2f

Figure 2g

Figure 2h1

Figure 2h2

Figure 2h3

Figure 2i

Figure 2j

Figure 2k1

Figure 2k2

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 4C

Figure 5A

Figure 5B

Figure 5C

Figure 6A

Figure 6B

Figure 6C

DETAILED DESCRIPTION OF THE INVENTION

[0014] Definition As used herein, the term "healthy control" means a subject who does not have the disease being studied but may have other conditions that indirectly affect the outcome.

[0015] As used herein, the term "characteristically determined immune cell subset" means a group of immune cell subsets having comparative immune profiling between a patient and a healthy control (HC), such as the amount of an immune cell subset in a patient that is significantly higher or lower than that in an HC.

[0016] As used herein, the term "predicted probability" means the probability of a subject with respect to each group in a trained machine learning model.

[0017] As used herein, the term "decision function" means a function that calculates the distance of a subject to the separating hyperplane of an SVM classifier.

[0018] As used herein, the term "holdout set" means a dataset that is not used in a machine learning model training process.

[0019] Embodiments Embodiment 1: A staining kit comprising a first pattern including antibodies against T cells, B cells, NK cells, monocytes, regulatory cells, CD8, CD45, and CTLA4; a second pattern including antibodies against T cells, B cells, NK cells, monocytes, regulatory cells, dendritic cells, and CD45; a third pattern including antibodies against T cells, B cells, NK cells, monocytes, CD8, CD45, CD45RA, CD62L, CD197, CX3CR1, and TCR αβ ; and a fourth pattern including antibodies against B cells, CD23, CD38, CD40, CD45, and IgM, wherein the antibodies in each pattern are labeled with a fluorescent dye.

[0020] Embodiment 2: The staining kit of Embodiment 1, wherein T cells contain CD3, CD4, CD25, CD45RO, CCR7 or any combination thereof; B cells contain CD10, CD19, CD21, CD127, IgG or any combination thereof; NK cells contain CD56; monocytes contain CD14; regulatory cells contain PD-1, PD-L1, FoxP3 or any combination thereof; and dendritic cells contain CD11c, HLA-DR or any combination thereof.

[0021] Embodiment 3: A first pattern containing antibodies against CD3, CD4, CD8, CD14, CD19, CD25, CD45, CD56, CTLA4, Foxp3 and PD-1; a second pattern containing antibodies against CD3, CD4, CD11c, CD14, CD19, CD45, CD56, HLA-DR, PD-1 and PD-L1; a third pattern containing antibodies against CD4, CD8, CD14, CD19, CD45, CD45RA, CD45RO, CD56, CD62L, CD197, CX3CR1 and TCR αβ and a fourth pattern containing antibodies against CD10, CD19, CD21, CD23, CD38, CD40, CD45, CD127, IgG and IgM, wherein the antibodies of each pattern are labeled with a fluorescent dye, the staining kit of Embodiment 1 or 2.

[0022] Embodiment 4: In the first pattern, the antibodies against CD3, CD4, CD8, CD14, CD25, CD45, CD56, CTLA4, Foxp3 and PD-1 are labeled with different fluorescent colors, and the antibodies against CD14 and CD19 are labeled with the same fluorescent color, the staining kit of Embodiment 3.

[0023] Embodiment 5: In the second pattern, the antibodies against CD3, CD4, CD11c, CD14, CD19, CD45, CD56, HLA-DR, PD-1 and PD-L1 are labeled with different fluorescent colors, the staining kit of Embodiment 3 or 4.

[0024] Embodiment 6: In the third pattern, antibodies against CD4, CD8, CD14, CD45, CD45RA, CD45RO, CD62L, CD197, CX3CR1, and TCR αβ are labeled with different fluorescent colors, and antibodies against CD14, CD19, and CD56 are labeled with the same fluorescent color, a staining kit according to any one of Embodiments 3 to 5.

[0025] Embodiment 7: In the fourth pattern, antibodies against CD10, CD19, CD21, CD23, CD38, CD40, CD45, CD127, IgG, and IgM are labeled with different fluorescent colors, a staining kit according to any one of Embodiments 3 to 6.

[0026] Embodiment 8: The following patterns: a fifth pattern including antibodies against CD3, CD4, CD8, CD14, CD19, CD25, CD45, CD56, CD69, PD-1, TCR αβ and TCR γδ ; a sixth pattern including antibodies against CD3, CD11b, CD11c, CD13, CD14, CD19, CD33, CD39, CD45, CD56, and HLA-DR; a seventh pattern including antibodies against CD3, CD4, CD8, CD14, CD19, CD27, CD28, CD45, CD56, PD-1, TCR αβ and TCR γδ ; an eighth pattern including antibodies against CD3, CD11b, CD14, CD16, CD19, CD45, CD56, CD64, CD66b, CD123, CD193, CD203c, and Siglec-8; and a ninth pattern including antibodies against CD4, CD8, CD14, CD19, CD45, CD45RO, CD56, CX3CR1, CD197, LAG-3, TCR αβ and TIM-3, a staining kit according to any one of Embodiments 3 to 7, wherein the antibodies in each pattern are labeled with fluorescent colors.

[0027] Embodiment 9: In the fifth pattern, CD3, CD4, CD8, CD25, CD45, CD56, CD69, PD-1, TCRαβ and TCR γδ An antibody against γδ is labeled with a different fluorescent color, and antibodies against CD3, CD14, and CD19 are labeled with the same fluorescent color. A staining kit according to any one of Embodiments 3 to 8.

[0028] Embodiment 10: In the sixth pattern, antibodies against CD3, CD11b, CD11c, CD13, CD14, CD33, CD39, CD45, and HLA-DR are labeled with different fluorescent colors, and antibodies against CD3, CD19, and CD56 are labeled with the same fluorescent color. A staining kit according to any one of Embodiments 3 to 9.

[0029] Embodiment 11: In the seventh pattern, CD3, CD4, CD8, CD27, CD28, CD45, CD56, PD-1, TCR αβ and TCR γδ An antibody against γδ is labeled with a different fluorescent color, and antibodies against CD3, CD14, and CD19 are labeled with the same fluorescent color. A staining kit according to any one of Embodiments 3 to 10.

[0030] Embodiment 12: In the eighth pattern, antibodies against CD3, CD11b, CD16, CD45, CD64, CD66b, CD123, CD193, CD203c, and Siglec-8 are labeled with different fluorescent colors, and antibodies against CD3, CD14, CD19, and CD56 are labeled with the same fluorescent color. A staining kit according to any one of Embodiments 3 to 11.

[0031] Embodiment 13: In the ninth pattern, antibodies against CD4, CD8, CD14, CD45, CD45RO, CX3CR1, CD197, LAG-3, TCR αβ and TIM-3 are labeled with different fluorescent colors, and antibodies against CD14, CD19, and CD56 are labeled with the same fluorescent color. A staining kit according to any one of Embodiments 3 to 12.

[0032] Embodiment 14: The following steps: (a) Obtaining peripheral blood mononuclear cells (PBMC) and / or white blood cells (WBC) from a plurality of healthy controls and a plurality of patients having a disease, respectively. (b) Staining the PBMC and / or WBC of the healthy controls and the patients, respectively, by using any one of the staining kits of Embodiments 1 to 13. (c) Obtaining data on the fluorescence intensity of each antibody bound to the PBMC and / or WBC of the healthy controls and the patients, respectively, by using flow cytometry. (d) Identifying immune cell subsets in the PBMC and / or WBC of the healthy controls and the patients, respectively, by using a systematic method, and obtaining a data set including data on the type and proportion of the immune cell subsets. (e) Evaluating the data set by using machine learning software, and obtaining immune cell subsets of the patients that are distinguishable from those of the healthy controls as immune cell subsets with determined disease characteristics. A method for identifying immune cell subsets with determined disease characteristics, including the above steps.

[0033] Embodiment 15: Step (e) is the following steps performed by machine learning software: (i) Performing data preprocessing on the data set. (ii) Performing feature selection from the preprocessed data by using the Boruta algorithm, and obtaining predetermined immune cell subsets of the disease as the selected features. (iii) Applying the data of the selected features and training a machine learning model by using at least one of a random forest (RF) algorithm, a logistic regression (LR) algorithm, and a support vector machine (SVM) algorithm. The method of Embodiment 14 further including the above steps.

[0034] Embodiment 16: The method of Embodiment 14 or 15, wherein the disease includes cancer, immune diseases, and infectious diseases.

[0035] Embodiment 17: A method according to any one of Embodiments 14 to 16, wherein the immunological disease includes, but is not limited to, idiopathic thrombocytopenic purpura, Guillain-Barré syndrome, myasthenia gravis, multiple sclerosis, optic neuritis, Kawasaki disease, rheumatoid arthritis, systemic lupus erythematosus, atopic dermatitis, atherosclerosis, coronary artery disease, cardiomyopathy, reactive arthritis, Crohn's disease, ulcerative colitis, graft-versus-host disease, and type 1 diabetes mellitus.

[0036] Embodiment 18: A method according to any one of Embodiments 14 to 16, wherein the infectious disease includes, but is not limited to, candidiasis, candidal sepsis, aspergillosis, streptococcal pneumonia, streptococcal skin and oropharyngeal conditions, gram-negative bacterial sepsis, tuberculosis, mononucleosis, influenza, respiratory diseases caused by respiratory syncytial virus, human immunodeficiency virus, hepatitis B, hepatitis C, malaria, schistosomiasis, methicillin-resistant Staphylococcus aureus, vancomycin-resistant Enterococcus, carbapenem resistance and carbapenemase-producing Enterobacteriaceae, mycobacteriosis, and trypanosomiasis.

[0037] Embodiment 19: The method according to any one of Embodiments 14 to 16, wherein the cancer includes, but is not limited to, brain cancer, bone cancer, skin cancer, esophageal cancer, gastric cancer, bile duct cancer, colorectal cancer, head and neck cancer, kidney cancer, fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colon cancer, pancreatic cancer, breast cancer, ovarian cancer, prostate cancer, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland cancer, sebaceous gland cancer, papillary carcinoma, papillary adenocarcinoma, cystadenocarcinoma, medullary carcinoma, bronchial carcinoma, renal cell carcinoma, hepatocellular carcinoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms tumor, cervical cancer, testicular tumor, lung cancer, small cell lung cancer, bladder cancer, epithelial cancer, glioma, astrocytoma, medulloblastoma, craniopharyngioma, epithelioma, pinealoma, hemangioblastoma, acoustic neuroma, anaplastic glioma, meningioma, melanoma, neuroblastoma, retinoblastoma, leukemia, lymphoma, multiple myeloma, Waldenström macroglobulinemia, myelodysplastic disease, heavy chain disease, neuroendocrine tumor, and schwannoma (Schwanoma).

[0038] Embodiment 20: The method according to Embodiment 19, wherein the head and neck cancer includes nasopharyngeal carcinoma (NPC).

[0039] Embodiment 21: The method according to any one of Embodiments 14 to 20, wherein the immune cell subset with determined NPC characteristics is selected from the group consisting of memory B cells, monocytes, T cells, naive CD4αβ T cells, PD-1 + CD4 T cells, PD-L1 + CD4 T cells, PD-1 + PD-L1 + monocytes, CD4 NKTreg cells, MHC II + CD4 T cells and MHC II + CD4 NKT cells.

[0040] Embodiment 22: The following steps: (a) A step of staining peripheral blood mononuclear cells (PBMC) from a subject by using any one of the staining kits according to Embodiments 1 to 13. (b) By using flow cytometry, data acquisition of the fluorescence intensity of each antibody bound to the characterized immune cell subsets of NPC is performed to obtain a data set including data on the type and its proportion of the characterized immune cell subsets, and (c) By using machine learning software, evaluating the data set and predicting whether the subject has nasopharyngeal carcinoma A method for predicting the likelihood of nasopharyngeal carcinoma (NPC) in a subject in need thereof, comprising: The characterized immune cell subsets of NPC are memory B cells, monocytes, T cells, naive CD4αβ T cells, PD-1 + CD4 T cells, PD-L1 + CD4 T cells, PD-1 + PD-L1 + Monocytes, CD4 NKTreg cells, MHC II + CD4 T cells and MHC II + Selected from the group consisting of CD4 NKT cells, the method.

[0041] Embodiment 23: The step (c) is the following step performed by machine learning software: (i) Applying a holdout set having the characterized immune cell subsets and testing the trained machine learning model (ii) If the value of the predicted probability obtained from the RF algorithm or the LR algorithm is greater than the first threshold, or the value of the decision function obtained from the SVM algorithm is greater than the second threshold, predicting that the subject has nasopharyngeal carcinoma The method according to Embodiment 22, further comprising:

[0042] Embodiment 24: The method according to Embodiment 22 or 23, wherein the first threshold of the predicted probability is equal to 0.5 and the second threshold of the decision function is equal to 0.

Example

[0043] Identification of the characterized immune cell subsets of NPC Regarding the flowchart for identifying immune cell subsets determined by NPC characteristics, please refer to Figure 1. The related procedures are described in detail below.

[0044] NPC and HC Regarding NPC, the eligibility criteria included: over 20 years old, over 50 kilograms, newly diagnosed according to practice guidelines, no history of severe infectious diseases such as human immunodeficiency virus and syphilis, not having received radiotherapy, chemotherapy, or autoimmune treatment within one month. Patients with central nervous system metastasis, pulmonary fibrosis or fibrotic pneumonia, and pulmonary edema or ascites conforming to the Common Terminology Criteria for Adverse Events (CTCAE) level 2 or above were ineligible. Subjects enrolled in HC had the same eligibility criteria as NPC, except for the diagnosed NPC. The outcome of this trial depended on the comparison of the classification and proportion of immune cell subsets between NPC and HC. Therefore, during the patient visit, the clinician collected 20 mL of anticoagulated peripheral blood from the subject for PBMC and granulocyte isolation and identification. The study design of this clinical trial followed the Declaration of Helsinki and the protocol of this clinical trial approved by the Institutional Review Board of Far Eastern Memorial Hospital (approval code 108170-E).

[0045] Reagents and antibodies All reagents and antibodies used in this study were listed in Table 1 and Table 2. The reagents were obtained from Cytiva (Marlborough, MA, USA), Lonza (Basel, Switzerland), and Sigma-Aldrich (Merck KGaA, Darmstadt, Germany). All antibodies were obtained from Beckman-Coulter (Brea, CA, USA), Biolegend (San Diego, CA, USA), and Thermo-Fisher (Waltham, MA, USA).

[0046] [Table 1]

[0047]

Table 2A

Table 2B

Table 2C

[0048] PBMC and granulocyte isolation and immunostaining EDTA-anticoagulated peripheral blood (hereinafter abbreviated as leukocytes) was divided into two aliquots, one for PBMC isolation and the other for granulocyte isolation. For PBMC isolation, leukocytes mixed with an aliquot of PBS were loaded into a centrifuge tube pre-loaded with Ficoll and centrifuged at 930×g for 30 minutes at room temperature (X-15R, Beckman-Coulter). PBMCs were recovered from the buffy coat, washed with PBS, subsequently centrifuged at 750×g for 7 minutes, and then resuspended in staining buffer (0.5% bovine serum albumin / PBS containing 0.02% (w / v) sodium azide) for immunostaining.

[0049] The ammonium chloride-potassium chloride (ACK) lysis method was applied to the lysis of red blood cells (RBC) in granulocyte isolation based on the protocol in the manual of the lysis buffer. Briefly, a portion of whole blood was mixed with 20 parts of ACK lysis buffer and subsequently gently shaken at room temperature for 10 minutes. The mixture was then centrifuged at 400×g for 5 minutes to remove the lysed red blood cells. The remaining white blood cells (WBC) were washed with PBS and subsequently immunostained.

[0050] All PBMCs were stained using cell surface markers, and some of them were further stained using intracellular cell markers. Each pattern refers to the cell staining pattern page. All staining procedures were maintained in the dark. All staining procedures were maintained in the dark. For surface marker staining, PBMCs were directly incubated with the desired antibodies labeled with fluorescent dyes in the staining kit (see Tables 2 and 3) at 4°C for 10 minutes. For intracellular marker staining, a portion of the surface marker-labeled PBMCs was fixed and permeabilized using the Foxp3 / Transcription Factor Staining Buffer Set (eBioscience™) according to the recommended protocol from the manual. Subsequently, the permeabilized PBMCs were stained with CTLA4 or Foxp3 at room temperature for 30 minutes. Thereafter, the stained PBMCs were washed using the staining buffer and then analyzed by flow cytometry.

[0051]

Table 3A

Table 3B

[0052] Data acquisition and data adjustment Regarding data acquisition, the fluorescence intensity of PBMCs was measured by a flow cytometer (Navios, Beckman Coulter), and the raw data was collected by Kaluza analysis software V1.3 (Beckman Coulter). Immunocyte subsets were defined by a systematic method using filtration with a 2-marker set (parameters for the X-axis and Y-axis). The definition of the cell subsets is listed in Table 4. As shown in Figures 2a - 2k, the detailed filtration process was performed based on Table 4 to obtain a raw dataset related to at least 82 types of immunocyte subsets and their ratios.

[0053]

Table 4A

Table 4B

Table 4C

Table 4D

Table 4E

Table 4F

[0054] Machine Learning Using Python Programs The above raw dataset consisted of HC (n = 34) and NPC (n = 15) subjects. Data preprocessing was performed through the following steps. First, the target label was converted to numerical value 0 for HC and numerical value 1 for NPC. Since one of the NPC subjects had 7 missing values of pattern PT - 87, missing value imputation was applied based on the mean of the other features of NPC (PT - 87). 82 types of immune cell subsets were saved for each subject after data preprocessing. The hold - out set consisted of some HC subjects (n = 20) separated from 34 HC subjects, and NPC subjects (n = 10) based on the percentiles and mean of the original 15 NPC subjects. The remaining HC subjects (n = 14) and the original NPC subjects (n = 15) were used as the training set. Feature selection in the training set was performed using the Boruta algorithm, a wrapper - based technique based on the random forest classification algorithm. Boruta compared the Z - scores of the shuffled shadow features and the original features to determine the feature importance in all iterations. After a predefined number of iterations, the predefined immune cell subsets of the disease as the selected features were obtained, which were significantly more relevant to classification than the randomly rearranged features. The selected features of this example are shown in Table 5. The training set of the selected features was applied to train models using three different types of machine learning (ML) algorithms. Random Forest (RF), Logistic Regression (LR), and Support Vector Machine (SVM) were used to classify the flow cytometry data from two classes of HC and NPC subjects. To minimize the influence of features with a relatively high magnitude for distance calculation, Min - Max scaling was applied to SVM to ensure that all features had a similar effect when the classifier constructed the hyperplane. Min - Max scaling is a normalization technique that converts the minimum feature value to 0 and the maximum feature value to 1. Subsequently, the discrimination ability of the model was evaluated by the area under the curve (AUC) of the Receiver Operating Characteristic (ROC) curve. The ROC curve is often used to compare the model performance in clinical classification problems.This shows the relationship between the true positive rate (sensitivity) and the false positive rate (1 - specificity) for each possibility threshold. If the discrimination ability of the model is good, the ROC curve approaches the upper left corner of the plot. Finally, the discrimination ability of the model was quantified by calculating the area under the ROC curve using the trapezoidal formula to obtain the AUC result. To compare the model performance, the ROC curves and AUC results of the training set were visualized. The plots are shown in FIGS. 3A - 3C. To explain the model by calculating the contribution of each selected feature to the prediction, Shapley Additive exPlanations (SHAP) was applied. To visualize the ranking of feature importance and the values of features per subject with SHAP values, SHAP summary plots were drawn. The color of the data points for each feature represents high or low feature values, and each data point represents one subject. Red indicates high feature values and blue indicates low feature values. The y - axis of the plot is the ranking of feature importance of the selected feature, and the x - axis is the range of SHAP values. According to the trained model, the higher the SHAP value, the higher the risk of NPC. The summary plots of the training set are shown in FIGS. 4A - 4C.

[0055]

Table 5

Example

[0056] Prediction of the likelihood of NPC in a subject by machine learning using a Python program The hold - out set consisted of HC (n = 20) and pseudo - NPC (n = 10) based on the percentiles and mean of the original 15 NPC subjects. 82 types of immune cell subsets were preserved for each subject. For each subject, 10 selected features shown in Table 5 were used.

[0057] To test the trained models of three different ML algorithms (RF, LR, and SVM), a holdout set containing the selected features was applied. The trained models were used to classify flow cytometry data from two classes of HC and NPC subjects.

[0058] For SVM, to ensure that all features have similar effects when there is no data leakage and the classifier constructs a hyperplane, the Min-Max scaling range obtained from the training process was applied to the holdout set.

[0059] Subsequently, the sensitivity and specificity of the trained models were tested using the holdout set. When the predicted probabilities of the trained RF and LR models were greater than 0.5, the models predicted the subjects as NPC. When the decision function of the trained SVM model was greater than 0, the models predicted the subjects as NPC. The prediction results of the holdout set are shown in Table 6.

[0060]

Table 6

[0061] Finally, to evaluate the discrimination ability of the models, the AUC of the ROC curve was used. The ROC curve is often used to compare the performance of models in clinical classification problems. This shows the relationship between the true positive rate (sensitivity) and the false positive rate (1 - specificity) for each possible threshold. If the discrimination ability of the model is good, the ROC curve approaches the upper left corner of the plot.

[0062] After forming the ROC curve, the discrimination ability of the model was quantified by calculating the area under the ROC curve using the trapezoidal formula to obtain the AUC result. To compare the model performance, the ROC curve and AUC results of the holdout set were visualized. The plots are shown in Figures 5A - 5C.

[0063] To explain the model by calculating the contribution of each selected feature to the prediction, the SHAP method was applied. To visualize the ranking of feature importance and the values of features per subject with SHAP values, SHAP summary plots were depicted. The color of the data points (subjects) for each feature represents high or low feature values, and each data point represents one subject. Red indicates high feature values and blue indicates low feature values. The y-axis of the plot is the ranking of feature importance of the selected features, and the x-axis is the SHAP value range. According to the trained model, the higher the SHAP value, the higher the risk of NPC. The summary plots of the holdout set are shown in FIGS. 6A-6C.

[0064] In summary, the present invention can be used not only for the diagnosis of diseases such as NPC, but also based on the states of PD-L1, PD-1, and T cells in the 10 types of immune cell subsets selected by the machine learning, to provide predictions regarding immunotherapy. For example, for a subject, during the period when the amount of PD-L1 expression of immune cells increases, can atezolizumab be administered for treatment; during the period when the amount of PD-1 expression of immune cells increases, can nivolumab or pembrolizumab be administered for treatment; or during the period when T cells decrease, can immune cells such as NK, DC, cytokine-induced killer (CIK), and T cells be replenished for treatment.

[0065] Unless otherwise defined, all technical and scientific terms and any acronyms used herein have the same meaning as commonly understood by one of ordinary skill in the art of the present invention. Any compositions, methods, kits, and means for conveying information similar or equivalent to those described herein can be used to practice the present invention, but the preferred compositions, methods, kits, and means for conveying information are those described herein.

[0066] All references cited in this specification are hereby incorporated by reference to the fullest extent permitted by law. The discussion of those references is intended solely to summarize the assertions made by their authors. No admission is made that any reference (or any portion of any reference) is prior art of significance. Applicants reserve the right to challenge the accuracy and appropriateness of any cited reference. (Reference) TIFF0007696971000015.tif60168

Claims

1. Pattern 1 comprising antibodies against CD3, CD4, CD8, CD14, CD19, CD25, CD45, CD56, CTLA4, FoxP3 and PD-1; Pattern 2 comprising antibodies against CD3, CD4, CD11c, CD14, CD19, CD45, CD56, HLA-DR, PD-1 and PD-L1; Pattern 3 comprising antibodies against CD4, CD8, CD14, CD19, CD45, CD45RA, CD45RO, CD56, CD62L, CD197, CX3CR1 and TCR αβ Pattern 4 comprising antibodies against CD10, CD19, CD21, CD23, CD38, CD40, CD45, CD127, IgG and IgM, wherein the antibodies of each pattern are labeled with a fluorescent dye, a staining kit for predicting the likelihood of nasopharyngeal carcinoma (NPC).

2. (a) A step of staining peripheral blood mononuclear cells (PBMC) and / or white blood cells (WBC) of a plurality of healthy controls and a plurality of patients with a disease, respectively, by using the staining kit according to Claim 1, (b) A step of obtaining data on the fluorescence intensity of each antibody bound to PBMC and / or WBC of healthy controls and patients, respectively, by using flow cytometry, (c) A step of identifying immune cell subsets in PBMC and / or WBC of healthy controls and patients, respectively, by using a systematic method, and obtaining a data set including data on the type and proportion of the immune cell subsets, and (d) A step of evaluating the data set by using machine learning software and obtaining an immune cell subset of a patient that can be distinguished from that of a healthy control as an immune cell subset with determined disease characteristics A method for identifying an immune cell subset with determined disease characteristics, comprising wherein the disease is nasopharyngeal carcinoma (NPC).

3. The step (d) is the following steps performed by machine learning software: (i) A step of performing data preprocessing on the data set of the data set, (ii) Performing feature selection from the preprocessed data by using the Boruta algorithm, and obtaining a predetermined subset of immune cells of the disease as the selected features, and (iii) Applying the data of the selected features and training a machine learning model by using at least one of a random forest (RF) algorithm, a logistic regression (LR) algorithm, and a support vector machine (SVM) algorithm The method according to claim 2, further comprising:

4. The predetermined subset of immune cells of NPC is a memory B cell, a monocyte, a T cell, a naive CD4αβ T cell, PD-1 + CD4 T cell, PD-L1 + CD4 T cell, PD-1 + PD-L1 + Monocyte, CD4 NKTreg cell, MHC II + CD4 T cell and MHC II + The method according to claim 3, comprising a memory B cell, a monocyte, a T cell, a naive CD4αβ T cell, PD-1

5. (a) Staining peripheral blood mononuclear cells (PBMCs) from a subject by using the staining kit according to claim 1; (b) Obtaining data on the fluorescence intensity of each antibody bound to the predetermined subset of immune cells of NPC by using flow cytometry, and obtaining a data set including data on the type and proportion of the predetermined subset of immune cells; and (c) Evaluating the data set by using machine learning software and predicting whether the subject has nasopharyngeal carcinoma (NPC) A method for predicting the likelihood of nasopharyngeal carcinoma (NPC) in a subject in need thereof, comprising: The predetermined subset of immune cells of NPC is a memory B cell, a monocyte, a T cell, a naive CD4αβ T cell, PD-1 + CD4 T cell, PD-L1 + CD4 T cell, PD-1 + PD-L1 +Single cells, CD4 NKTreg cells, MHC II + CD4 T cells and MHC II + A method comprising CD4 NKT cells. Claim 6 Step (c) is the following step performed by machine learning software: (i) Applying a holdout set having a characterized subset of immune cells and testing the trained machine learning model; (ii) Predicting that the subject has nasopharyngeal cancer if the value of the predicted probability obtained from the RF algorithm or the LR algorithm is greater than a first threshold, or if the value of the decision function obtained from the SVM algorithm is greater than a second threshold; The method according to claim 5, further comprising:

Citation Information

Patent Citations

  • System and method for determining lung health

    CN112424341A

  • Methods, reagents and kits for flow cytometric immunophenotyping test

    JP2017207519A

  • Automated flow cytometry analysis method and system

    JP2018505392A

  • Systems and methods for determining lung health

    JP2021521466A

  • Methods and compositions for determining the composition of the tumor microenvironment

    JP2022512973A