Microbiome-assisted diagnostic markers for nasopharyngeal carcinoma
By detecting the abundance of specific bacterial markers in the nasopharyngeal region and combining logistic regression models, kits and electronic devices for diagnosis and prognosis judgment of nasopharyngeal carcinoma are developed, the problem of early diagnosis of nasopharyngeal carcinoma is solved, the diagnostic accuracy and prognosis judgment ability are improved, and the late stage of tumor detection is reduced.
Patent Information
- Application Number
- CN202311341498.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-10-17
AI Technical Summary
Early diagnosis of nasopharyngeal carcinoma is difficult to achieve. The existing technology lacks effective biomarkers for early detection and distinction between nasopharyngeal carcinoma patients and healthy people, resulting in the tumor staging later when the patient is discovered, which increases the health and economic burden.
Sperm markers such as Cutibacterium acnes, Corynebacterium accolens, Cutibacterium granulosum, Limnobacter thiooxidans, Peptostreptococcus stomatis, Streptococcus mitis, Gemella haemolysans, Gemella morbillorum, Klebsiella pneumoniae, Neisseria sicca, Ralstonia insidiosa and Granulicateella adiacins were used to quantitatively detect the abundance of these strains, combined with logistic regression models, kits and electronic devices for diagnosis, monitoring or prognosis judgment were developed, and nasopharyngeal microbial sampling was used for diagnosis and prognosis judgment.
It improves the accuracy of early diagnosis and prognosis judgment ability of nasopharyngeal carcinoma, can effectively distinguish between patients with nasopharyngeal carcinoma from healthy people, reduces the late stage of tumor discovery, and reduces the health and economic burden.
Smart Images

Figure CN117535407B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedicine and relates to a nasopharyngeal carcinoma microbiota-assisted diagnostic marker. Background Art
[0002] Nasopharyngeal carcinoma is a malignant tumor originating from the nasopharyngeal epithelium, which is prevalent in southern my country and Southeast Asian countries. Although the mortality rate of nasopharyngeal carcinoma has declined in recent years with the development of diagnosis and treatment technology, the incidence rate remains high. Nasopharyngeal carcinoma is more common in people aged 40-55 years and is more common in men. This tumor is more common in young and middle-aged people and has a regional high incidence. In addition, due to the hidden nature of nasopharyngeal carcinoma, the tumor stage is usually late when the patient is discovered, which brings a large health and economic burden to high-incidence areas such as South China. Early detection and early treatment will hopefully improve the patient's long-term prognosis and enhance the quality of life after treatment.
[0003] The nasopharynx is connected to the external environment through the mouth and nose, and is colonized with a rich variety of bacteria, which are key members of the local microenvironment. In healthy individuals, the nasopharyngeal flora is mainly composed of aerobic bacteria from the genera Corynebacterium, Streptococcus, Moraxella, and Crafty Coccus. These resident flora colonized in the nasopharyngeal mucosa play an important role in maintaining the stability of the local microecology, resisting the colonization of opportunistic pathogens, and regulating the balance of mucosal immunity, and are the cornerstone of maintaining nasopharyngeal health. The nasopharynx is often the first stop for respiratory pathogens to enter the body. At present, our understanding of the nasopharyngeal flora is mostly focused on its association with viral or bacterial infections, asthma and other respiratory inflammatory diseases. A large number of studies have confirmed that dysbiosis is a key cause of tumor development, and some flora characteristics have shown their clear marker value in tumors such as colorectal cancer. Therefore, finding flora characteristics that can distinguish nasopharyngeal carcinoma patients from controls will hopefully develop new approaches for early diagnosis, treatment detection, and intervention of nasopharyngeal carcinoma. Summary of the invention
[0004] On the one hand, the present invention provides the use of a reagent for detecting bacterial species markers in preparing a product for diagnosing, monitoring or prognosticating nasopharyngeal carcinoma; the bacterial species markers include: one or more of Cutibacterium acnes, Corynebacterium accolens, Cutibacterium granulosum, Limnobacter thiooxidans, Peptostreptococcusstomatis, Streptococcus mitis, Gemella haemolysans, Gemella morbillorum, Klebsiella pneumoniae, Neisseria sicca, Ralstonia insidiosa and Granulicatellaadiacens.
[0005] In some embodiments, the bacterial species marker includes: Cutibacterium acnes; Corynebacterium accolens; or Cutibacterium granulosum.
[0006] In some embodiments, the bacterial species markers include: Cutibacterium acnes and Streptococcus mitis.
[0007] In some embodiments, the bacterial species marker includes a first bacterial species marker and a second bacterial species marker; the first bacterial species marker includes: Cutibacterium acnes and Streptococcus mitis; the second bacterial species marker includes: Corynebacterium accolens, Cutibacterium granulosum, Limnobacterthiooxidans, Peptostreptococcus stomatis, Gemella haemolysans, Gemellamorbillorum, Klebsiellapneumoniae, Neisseria sicca, Ralstonia insidiosa and Granulicatella adiacens. One or more.
[0008] In some embodiments, the second bacterial species marker includes: Limnobacter thiooxidans; in some embodiments, the second bacterial species marker includes: Limnobacter thiooxidans and Peptostreptococcus stomatis; in some embodiments, the second bacterial species marker includes: Limnobacter thiooxidans, Peptostreptococcus stomatis and Corynebacteriumaccolens; in some embodiments, the second bacterial species marker includes: Limnobacter thiooxidans, Peptostreptococcus stomatis, Corynebacterium accolens and Klebsiella pneumoniae; in some embodiments, the second bacterial species marker includes: Limnobacter thiooxidans, Peptostreptococcus stomatis, Corynebacterium accolens, Klebsiella pneumoniae and Gemella haemolysans; or Limnobacter thiooxidans, Peptostreptococcus stomatis, Corynebacterium accolens, Klebsiellapneumoniae and Gemella morbillorum; in some embodiments, the second bacterial species marker includes: Limnobacter thiooxidans, Peptostreptococcusstomatis, Corynebacterium accolens, Klebsiellapneumoniae, Gemella haemolysans and Gemella morbillorum; in some embodiments, the second bacterial species marker includes: Corynebacteriumaccolens, Limnobacter thiooxidans, Peptostreptococcus stomatis, Gemellahaemolysans, Gemella morbillorum, Klebsiella pneumoniae, Neisseria sicca, Ralstonia insidiosa and Granulicatella adiacens.
[0009] In some embodiments, the reagent for detecting bacterial species markers is a reagent for quantitatively detecting the abundance of bacterial species markers; in some embodiments, the reagent for detecting bacterial species markers is a primer, a probe, an antisense oligonucleotide, an aptamer, and / or an antibody; in some embodiments, the reagent for detecting bacterial species markers is a 16S rRNA detection reagent, a PCR detection reagent, an RNAseq detection reagent, a high-throughput sequencing reagent, and / or a metagenomic sequencing reagent; in some embodiments, the sampling site of the bacterial species includes the nasopharynx or the nasal cavity; in some embodiments, the product is a reagent, a kit, a test paper, a chip and / or an electronic device.
[0010] On the one hand, the present invention provides a kit for diagnosing, monitoring or prognosticating nasopharyngeal carcinoma, the kit comprising a detection reagent for detecting a bacterial species marker, wherein the bacterial species marker comprises one or more of Cutibacterium acnes, Corynebacterium accolens, Cutibacterium granulosum, Limnobacter thiooxidans, Peptostreptococcus stomatis, Streptococcus mitis, Gemella haemolysans, Gemellamorbillorum, Klebsiellapneumoniae, Neisseria sicca, Ralstonia insidiosa and Granulicatella adiacens.
[0011] In some embodiments, the detection reagent includes a detection reagent for detecting the content of the bacterial species marker; preferably, the reagent for detecting the bacterial species marker is a primer, a probe, an antisense oligonucleotide, an aptamer, and / or an antibody; preferably, the reagent for detecting the bacterial species marker is a 16S rRNA detection reagent, a PCR detection reagent, an RNAseq detection reagent, a high-throughput sequencing reagent, and / or a metagenomic sequencing reagent; preferably, the kit also includes a sampling reagent and / or a device, and the sampling reagent and / or the device include: an oral microbial sampling reagent and / or
[0012] Or device, nasal microorganism sampling reagent and / or device, nasopharyngeal microorganism sampling reagent and / or device, or intestinal microorganism sampling reagent and / or device; preferably, the nasopharyngeal microorganism sampling reagent and / or device includes: a nasopharyngeal brush.
[0013] In one aspect, the present invention provides an electronic device for diagnosing, monitoring or prognosing nasopharyngeal carcinoma, the electronic device comprising:
[0014] An acquisition and detection module, wherein the acquisition and detection module is configured to acquire a biological sample, detect the biological sample, and obtain the content of a bacterial species marker in the biological sample;
[0015] A diagnosis, monitoring or prognosis judgment module, wherein the diagnosis, monitoring or prognosis judgment module uses a logistic regression model and the content of the bacterial species marker to establish a regression equation, construct a model, and then outputs a diagnosis, monitoring or prognosis judgment result according to the model;
[0016] The bacterial species markers include:
[0017] Cutibacterium acnes, Corynebacterium accolens, Cutibacteriumgranulosum, Limnobacter thiooxidans, Peptostreptococcus stomatis, Streptococcusmitis, Gemella haemolysans, Gemella morbillorum, Klebsiella pneumoniae, Neisseriasicca, Ralstonia insidiosa and Granulicatella one or more of adiacens.
[0018] In some embodiments, the sampling sites of the bacterial species include the nasopharynx, oral cavity, intestinal nasal cavity and / or nasal passage; preferably, the model outputs the diagnosis, monitoring or prognosis result as follows: When the patient is diagnosed with nasopharyngeal cancer or has a high risk; When the patient is diagnosed with nasopharyngeal cancer, the patient is judged to have no nasopharyngeal cancer or have
[0019] In one aspect, the present invention provides a method for screening bacterial species markers for diagnosing, monitoring or prognosing nasopharyngeal carcinoma, the screening method comprising:
[0020] Testing the biological sample taken from the subject to obtain a bacterial flora dataset;
[0021] Obtaining the subject's condition of suffering from nasopharyngeal carcinoma, and constructing a first data set using the microbiome data set and the condition of suffering from nasopharyngeal carcinoma;
[0022] Performing multivariate statistics and data mining on the first data set, performing species difference analysis on the bacterial community data set, and obtaining the bacterial species markers associated with the nasopharyngeal carcinoma;
[0023] The biological sample includes nasal mucus, nasal washing fluid, nasal swab, nasal exfoliated cells, nasopharyngeal mucus, nasopharyngeal washing fluid, nasopharyngeal swab and / or nasopharyngeal exfoliated cells;
[0024] The subjects include healthy people and nasopharyngeal cancer patients;
[0025] The bacterial species markers include:
[0026] Cutibacterium acnes, Corynebacterium accolens, Cutibacteriumgranulosum, Limnobacter thiooxidans, Peptostreptococcus stomatis, Streptococcusmitis, Gemella haemolysans, Gemella morbillorum, Klebsiella pneumoniae, Neisseriasicca, Ralstonia insidiosa and Granulicatella one or more of adiacens.
[0027] In some embodiments, the detection includes the use of a detection reagent; preferably, the detection reagent includes a detection reagent for detecting the content of the bacterial species marker; preferably, the reagent for detecting the bacterial species marker is a primer, a probe, an antisense oligonucleotide, an aptamer, and / or an antibody; preferably, the reagent for detecting the bacterial species marker is a 16S rRNA detection reagent, a PCR detection reagent, an RNAseq detection reagent, a high-throughput sequencing reagent, and / or a metagenomic sequencing reagent; preferably, the bacterial species is detected from nasal mucus, nasal wash fluid, nasal swab, nasal exfoliated cells, nasopharyngeal mucus, nasopharyngeal wash fluid, nasopharyngeal swab and / or nasopharyngeal exfoliated cells to obtain a bacterial flora data set.
[0028] On the one hand, the present invention provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the model in the electronic device, or the screening method for bacterial markers for diagnosing, monitoring or prognosing nasopharyngeal carcinoma.
[0029] On the one hand, the present invention provides a processor, which is used to run a program, wherein the model or the screening method for bacterial markers for diagnosing, monitoring or prognosing nasopharyngeal carcinoma is executed when the program is running.
[0030] Those skilled in the art can clearly understand that the present invention can be implemented by means of software plus hardware devices such as detection devices. Based on this understanding, the data processing part of the technical solution of the present application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various embodiments of the present application, or certain parts of the embodiments.
[0031] The present invention can be used in many general or special computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.
[0032] As known to those skilled in the art, some modules or steps of the present invention can be implemented in a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented with program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0033] Based on the microbial flora of individuals associated with nasopharyngeal carcinoma presented in this application, clinicians can also provide clinical guidance based on personal microbial flora. For example, when people's microbial flora abundance approaches or exceeds a risk range, drug intervention, control or treatment of nasopharyngeal carcinoma can be recommended. In addition, for some patients with high risk or confirmed nasopharyngeal carcinoma, necessary nasopharyngeal or nasal microenvironment treatment can be given according to the flora situation. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is the quality control process of the nasopharyngeal brush flora sequencing data in Example 1.
[0035] Figure 2 This is a principal coordinate diagram of the differences in the diversity of nasopharyngeal flora structure between nasopharyngeal carcinoma patients and the control group in Example 1.
[0036] Figure 3The figure is the receiver operating characteristic curve of the model constructed by the single flora feature in Example 2 for distinguishing nasopharyngeal carcinoma patients and controls; the AUC marked in the figure is the area under the curve of the receiver operating characteristic curve; the point marked data are the optimal cutoff value and its corresponding specificity and sensitivity; the marking order is: cutoff value (specificity, sensitivity).
[0037] Figure 4 These are the 11 bacterial flora characteristics with the best screening modeling effect in Example 2.
[0038] Figure 5 The receiver operating characteristic curves and predicted values of the 11 flora characteristic models in Example 2; the AUC indicated in the figure is the area under the curve of the receiver operating characteristic curve; the point marked data are the optimal cutoff value and its corresponding specificity and sensitivity; the marking order is: cutoff value (specificity, sensitivity).
[0039] Figure 6 The receiver operating characteristic curves and predicted values of the two flora characteristic models in Example 2; the AUC indicated in the figure is the area under the curve of the receiver operating characteristic curve; the point marked data are the optimal cutoff value and its corresponding specificity and sensitivity; the marking order is: cutoff value (specificity, sensitivity).
[0040] Figure 7 It is the ROC of the Cutibacterium acnes and other bacteria combined model and the Cutibacterium acnes single bacteria model in Example 2; wherein: ROC1--Cutibacterium acnes single bacteria model; ROC2--Cutibacterium acnes and Streptococcus mitis two bacteria combined model; ROC3--Cutibacterium acnes and Gemella haemolysans two bacteria combined model; ROC4--Cutibacterium acnes and Limnobacter thiooxidans two bacteria combined model.
[0041] Figure 8 It is the number of times the 11 flora characteristics in Example 2 are included in the top ten, top twenty and top thirty models of the two-site model effects.
[0042] Fig. 9 The receiver operating characteristic curves and predicted values of the three flora characteristic models in Example 2; the AUC indicated in the figure is the area under the curve of the receiver operating characteristic curve; the point marked data are the optimal cutoff value and its corresponding specificity and sensitivity; the marking order is: cutoff value (specificity, sensitivity).
[0043] Fig.10It is the ROC of the three characteristic bacterial community joint models of Cutibacterium acnes and Streptococcus mitis and other bacteria in Example 2; wherein: ROC1--Cutibacterium acnes and Streptococcus mitis two-bacteria joint model; ROC2--Cutibacterium acnes, Streptococcus mitis and Limnobacterthiooxidans three-bacteria joint model; ROC3--Cutibacterium acnes, Streptococcus mitis and Peptostreptococcus stomatis three-bacteria joint model; ROC4--Cutibacterium acnes, Streptococcus mitis and Gemella morbillorum three-bacteria joint model.
[0044] Fig.11 It is the number of times the 11 flora characteristics in Example 2 are included in the top ten, top twenty and top thirty models of the three-point model effects.
[0045] Fig.12 The receiver operating characteristic curves and predicted values of the four flora characteristic models in Example 2; the AUC indicated in the figure is the area under the curve of the receiver operating characteristic curve; the point marked data are the optimal cutoff value and its corresponding specificity and sensitivity; the marking order is: cutoff value (specificity, sensitivity).
[0046] Fig.13It is the ROC of the four characteristic bacterial community joint models of Cutibacterium acnes, Streptococcus mitis and Limnobacter thiooxidans and other bacteria in Example 2; wherein: ROC1--Cutibacterium acnes, Streptococcus mitis and Limnobacter thiooxidans three-bacteria joint model; ROC2--Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans and Peptostreptococcusstomatis four-bacteria joint model; ROC3--Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans and Gemella morbillorum four-bacteria joint model; ROC4--Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans and Klebsiellapneumoniae four-bacteria joint model.
[0047] Fig.14 It is the number of times the 11 flora characteristics in Example 2 are included in the top ten, top twenty and top thirty models of the four-site model effects.
[0048] Fig.15 The receiver operating characteristic curves and predicted values of the five flora characteristic models in Example 2; the AUC indicated in the figure is the area under the curve of the receiver operating characteristic curve; the point-marked data are the optimal cutoff value and its corresponding specificity and sensitivity; the marking order is: cutoff value (specificity, sensitivity).
[0049] Fig.16It is the ROC of the joint model of 5 characteristic flora of Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans and Corynebacterium accolens and other bacteria in Example 2; Among them: ROC1--Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans and Peptostreptococcus stomatis four-bacteria joint model; ROC2--Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans and Peptostreptococcus stomatis and Corynebacterium accolens five-bacteria joint model; ROC3--Cutibacterium acnes, Streptococcusmitis, Limnobacter thiooxidans, Gemella morbillorum and Corynebacterium accolens five-bacteria joint model; ROC4--Cutibacterium acnes, Streptococcus mitis, Limnobacterthiooxidans, Klebsiellapneumoniae and Corynebacterium accolens five-bacteria joint model.
[0050] Fig.17 It is the number of times the 11 flora characteristics in Example 2 are included in the top ten, top twenty and top thirty models of the five-site model effects.
[0051] Fig.18 The receiver operating characteristic curves and predicted values of the six flora characteristic models in Example 2; the AUC indicated in the figure is the area under the curve of the receiver operating characteristic curve; the point-marked data are the optimal cutoff value and its corresponding specificity and sensitivity; the marking order is: cutoff value (specificity, sensitivity).
[0052] Fig.19 It is the number of times the 11 flora characteristics in Example 2 are included in the top ten, top twenty and top thirty models of the six-site model effects.
[0053] Fig. 20 The figure shows the changing trend of the area under the curve (AUC) of the optimal model incorporating different numbers of bacterial flora characteristics in Example 2.
[0054] Fig.21 The receiver operating characteristic curves and predicted values of the 8 flora characteristic models in Example 2; the AUC indicated in the figure is the area under the curve of the receiver operating characteristic curve; the point marked data are the optimal cutoff value and its corresponding specificity and sensitivity; the marking order is: cutoff value (specificity, sensitivity). DETAILED DESCRIPTION
[0055] The technical solution of the present invention is further described below by specific embodiments, which do not limit the protection scope of the present invention. Some non-essential modifications and adjustments made by others based on the concept of the present invention still fall within the protection scope of the present invention.
[0056] the term
[0057] The term "biomarker", also known as "marker" or "diagnostic marker", refers to a measurable indicator of an individual's biological state. Such biomarkers can be any substance in an individual as long as they are related to a specific biological state (e.g., disease state) of the individual being tested, for example, nucleic acid markers (also known as gene markers, such as DNA), protein markers, cytokine markers, chemokine markers, carbohydrate markers, antigen markers, antibody markers, species markers (species / genus markers) and functional markers (KO / OG markers), etc. Among them, the meaning of nucleic acid markers is not limited to existing genes that can be expressed as biologically active proteins, but also includes any nucleic acid fragments, which can be DNA or RNA, which can be modified DNA or RNA, or unmodified DNA or RNA, and a collection composed of them. In this article, nucleic acid markers are sometimes also referred to as characteristic fragments. In the present invention, biomarkers can also be represented by "microbial markers" or "bacterial markers". Biomarkers are measured and evaluated, and are often used to examine normal biological processes, pathogenic processes, or pharmacological responses to therapeutic interventions. In a specific embodiment of the present invention, the present invention uses high-throughput sequencing to analyze biological samples of NPC patients and healthy people. Based on the data of high-throughput sequencing, NPC patients and healthy people are compared to determine biomarkers related to the diagnosis and differentiation of NPC patients and healthy people.
[0058] The term "abundance" refers to the measure of the number of target microorganisms in a biological sample. "Abundance" is also referred to as "load". Generally, by molecular methods, the 16S rRNA gene copy number of the target microorganism is typically measured by, for example, fluorescence in situ hybridization (FISH), quantitative polymerase chain reaction (qPCR) or PCR / pyrophosphate sequencing, and bacteria are quantified. The quantification of target nucleic acid sequence abundance in a biological sample may be absolute or relative. "Relative quantification" is usually based on one or more internal reference genes, i.e., the 16S rRNA gene from a reference strain, such as using universal primers and expressing the abundance of the target nucleic acid sequence as the percentage of total bacterial 16S rRNA gene copies or by normalization of Escherichia coli 16S rRNA gene copies and measuring bacteria. "Absolute quantification" gives the exact number of target molecules by comparing with a DNA standard or by normalization of DNA concentration. "Quantitative level" can be concentration (DNA amount per unit volume), DNA amount or gene copy number per cell number, cycle threshold value (Ct value) or any mathematical transformation thereof, such as the Log10 of gene copy number.
[0059] In the present invention, any method well known to those skilled in the art can be used to detect microbial markers or determine the level of microbial markers. These methods include, but are not limited to, methods of sequence amplification using primers, and immunological methods using antigen-antibody reactions. Among them, the method of sequence amplification using primers can be, for example, polymerase chain reaction (PCR), reverse transcription-polymerase chain reaction (RT-PCR), multiple PCR, touchdown PCR, hot start PCR, nested PCR, boosted PCR, real-time PCR, differential PCR, cDNA end rapid amplification, reverse polymerase chain reaction, vector-mediated PCR, thermal asymmetric staggered PCR, ligase chain reaction, repair chain reaction, transcription-mediated amplification, autonomous sequence replication, selective amplification reaction of target base sequences. The immunological method using antigen-antibody reactions can be, for example, protein blotting, enzyme-linked immunosorbent assay, radioimmunoassay, radioimmunodiffusion, Euclid's immunodiffusion, rocket immunoelectrophoresis, tissue immunostaining, immunoprecipitation analysis, complement fixation analysis, fluorescence activated cell sorter, protein chip, etc., but the scope of the present invention is not limited thereto.
[0060] The term "sample" refers to a sample containing at least one substance from a subject. Examples of sample items may include cells, tissues, body fluids, biopsy specimens, blood, urine, feces, saliva, sputum, plasma, serum, cell culture supernatant, swab samples, etc. In a specific embodiment of the present invention, the sample is a nasopharyngeal swab.
[0061] The term "test subject", also referred to as "research subject" or "subject", refers to any animal subject, including but not limited to humans, laboratory animals, livestock and domestic pets. A subject may be inhabited by a variety of microorganisms. A subject may have different microbial communities in various habitats on and in its body. The subject may be diagnosed with a disease or suspected of being at risk for a disease. A subject may have a microbiome state (dysbiosis) that leads to disease. Preferably, the test subject, research subject or subject is a human.
[0062] The term "detection" is the same as "diagnosis". In addition to the early diagnosis of NPC, it also includes the diagnosis of intermediate and late stage NPC, and also includes NPC screening, risk assessment, prognosis, disease identification, diagnosis and monitoring of disease stages, and selection of therapeutic targets.
[0063] The term "diagnosis", also known as "early diagnosis", refers to confirming the presence or characteristics of a pathological condition. In a specific embodiment of the present invention, the diagnosis refers to distinguishing the presence or absence of nasopharyngeal carcinoma.
[0064] The term "16S rRNA" refers to rRNA that contains a conserved region common to all species and a hypervariable region that can classify a specific species, which constitutes the 30S subunit of the prokaryotic ribosome. Therefore, the presence of microorganisms can be identified by base sequence analysis. In particular, since there is almost no diversity between homologous species and there is diversity between different species, prokaryotes can be effectively identified by comparing the sequences of 16S rRNA. In addition, since 16S rDNA is a gene that encodes 16S rRNA, 16S rDNA can also be used to identify microorganisms.
[0065] The term "AUC (Area under the ROC curve)" refers to the area under the receiver operating characteristic curve (ROC curve), which is a tool for measuring imbalance in classification. ROC curves and AUC are often used to evaluate the quality of a binary classifier. The ROC curve is a probability curve that plots the relationship between TPR (true positive rate) and FPR (false positive rate) at different thresholds, where the horizontal axis is FPR (False positive rate) and the vertical axis is TPR (True positive rate). FPR refers to how many of all negative examples are predicted as positive examples; TPR refers to how many true positive examples are predicted; the ROC curve can essentially separate "signal" from "noise"; AUC is a probability value. When a positive sample and a negative sample are randomly selected, the probability that the current classification algorithm ranks this positive sample ahead of the negative sample based on the calculated score is the AUC value, which is used as a summary of the ROC curve. The higher the AUC, the better the model's performance in distinguishing positive and negative classes. When AUC = 1, the classifier is able to correctly distinguish all positive and negative points; however, if AUC is 0, the classifier will predict all negatives as positives and all positives as negatives. Therefore, the closer the AUC is to 1, the more sensitive and specific the test is, while an AUC close to 0.5 indicates that the test is neither sensitive nor specific.
[0066] Example 1: Based on the third-generation 16S rRNA amplicon sequencing, it was found that there were significant differences in the nasopharyngeal flora between nasopharyngeal carcinoma patients and healthy controls
[0067] In this example, a total of 84 NPC patients and 94 control subjects were recruited, nasopharyngeal brush samples were collected, and clinical and basic information of the participants was collected. Nasopharyngeal brush DNA was extracted and full-length amplicon sequencing of the 16S rRNA gene V1-V9 region was performed (PacBio sequencing technology). A total of 156 bacterial flora data that met quality control were successfully detected (70 NPC patients and 86 healthy controls), and differential bacterial flora characteristics between cases and controls were identified. The specific implementation plan is as follows:
[0068] 1.1 Research subjects inclusion and information collection
[0069] This cohort is a case-control study design, which is a group of 70 NPC patients and 86 controls recruited by the applicant's team from June to November 2020 at the Red Cross Hospital of Wuzhou City, Guangxi Zhuang Autonomous Region, China. Pacbio16S full-length sequencing technology was used to explore the characteristics of NPC-related nasopharyngeal microbiota at the species level. The inclusion criteria are as follows:
[0070] *The inclusion criteria for the case group are:
[0071] ① Patients diagnosed with nasopharyngeal carcinoma by pathology and have not received any nasopharyngeal carcinoma-related treatment;
[0072] ②Age 18-75 years old;
[0073] ③ Voluntarily participate, sign the informed consent and accept sample collection.
[0074] *The inclusion criteria for the control group are:
[0075] ①No abnormal symptoms of the head and neck were found;
[0076] ② No serious diseases in the past (such as heart, brain, liver, kidney, mental illness and malignant tumor);
[0077] ③Age 18-75 years old;
[0078] ④ Voluntarily participate, sign the informed consent form and be able to accept sample collection.
[0079] The personal information of the participants was collected by professionally trained staff, including basic personal information such as gender, age, contact information, medical history, smoking and drinking history, family history, and previous EB virus testing. The basic information characteristics of the included population are shown in Table 1.
[0080] Table 1: Basic information characteristics of the included cohort
[0081]
[0082]
[0083] 1.2 Nasopharyngeal brush collection and DNA extraction
[0084] The collection process of nasopharyngeal brush samples is as follows: insert the nasopharyngeal brush into the nasopharynx through the middle nasal meatus, determine the position of the brush with a nasopharyngeal endoscope, evenly wipe the walls and recesses of the nasopharynx for 2-3 circles, take out the brush after it has fully absorbed the wiped tissue, quickly place it in a collection tube containing preservation solution, and freeze it at -80°C.
[0085] The DNA extraction process of nasopharyngeal brushes is as follows: thaw the nasopharyngeal brushes on ice, vortex the thawed samples, fully release the exfoliated cells and other substances on the brushes into the preservation solution to make a sample suspension, and then take 200 μL of the sample suspension for DNA extraction. Use sterile glass beads to grind at the maximum speed in a vortex instrument for 10 minutes to fully lyse the bacterial cell wall. The pretreated samples were extracted with the QIAGEN tissue and blood DNA extraction kit according to the kit instructions, and the extracted DNA was frozen in a -20°C refrigerator. Blank control samples were set up during sample processing and DNA extraction.
[0086] 1.3 Sample flora sequencing experimental process
[0087] Nasopharyngeal brush DNA was quantified using a Qubit fluorometer (with a Qubit dsDNABR kit) and quality controlled using gel electrophoresis. The nasopharyngeal brush was diluted to 25 ng / μL and dispensed into a 96-well plate. A microbiome detection system based on 27F / 1492R primers to amplify the full-length region of the 16SrRNA gene and combined with PacBio third-generation sequencing. The specific process is: the bacterial universal primers 27F (5'-AGRGTTYGATYMTGGCTCAG-3', Forward primer) and 1492R (5'-RGY TACCTTGTTACGACTT-3', Reverse primer) containing a 12bp barcode sequence were used to amplify the full-length 16SrRNA gene. KAPAHiFi HotStart DNA polymerase was used to amplify the nasopharyngeal brush DNA for 32-33 cycles.
[0088] *PCR amplification system:
[0089] KAPA HiFi HotStart DNA Polymerase-------12.5μL
[0090] Nuclease-free water(invitrogen)-------4.5μL
[0091] Forwardprimer(2.5μM)---------------3μL
[0092] Reverse primer(2.5μM)---------------3μL
[0093] Nasopharyngeal brush DNA (25ng / μL)---------------2μL
[0094] *PCR reaction conditions are:
[0095]
[0096] Agarose gel electrophoresis was used to confirm the accuracy of the length of the amplified PCR product (~1500bp); the amplified product was purified using Agencourt AMPure XP (0.6×) magnetic beads, and the purified products of each sample were mixed at equimolar concentrations. The purified products were used to construct the SMRTbell DNA library according to the manufacturer's instructions (PacBio platform) and sequenced using the PacBio Sequel platform (Pacific Biosciences). The sequencing depth of each sample was 5,000 to 20,000 sequences. Negative quality control samples were set up during the PCR and library construction process, and were included in the sequencing together with the blank control samples during the sample processing process.
[0097] 1.4 Analysis process of 16S rRNA full-length amplification sequencing data
[0098] High-quality circular consensus sequences (CCS) were obtained from raw PacBio sequencing data using SMRT Link software (v9.0.0, Pacific Biosciences). Sequences were split into corresponding samples based on the 12 bp barcode sequence on the amplification primers during library construction using Lima (v2.0.0). The DADA2 (v1.22.0) workflow customized for PacBio full-length 16S rRNA gene sequencing data was used for quality control, denoising, and identification of amplicon sequence variants (ASVs) of CCS sequences. Based on the above process, the sequence quantity data for each sample in each ASV can be obtained.
[0099] The silva_nr99_v138_train_set database and the silva_species_assignment_v138 sequence database were used to annotate the bacterial species classification of the ASVs obtained in the above process. According to the "RIDE Checklist", a series of schemes were applied for sequence quality control and decontamination, specifically: 1) Removal of ASVs with annotation information as mitochondria or chloroplasts; 2) Removal of ASVs that were not annotated at the bacterial phylum level. In addition, combined with the sequence information obtained from the negative control samples designed during sample collection and processing and DNA library construction and sequencing, the R package Decontam (v1.10.0) was used to identify and filter potential environmental pollution. The sequencing depth of 2000 was used as the cutoff value, and samples with a sequencing depth lower than this cutoff value were deleted, and finally the ASV abundance table after quality control was obtained. See the sample quality control process for details. Figure 1 .
[0100] 1.5 Results: There are significant differences in the nasopharyngeal flora between NPC patients and healthy controls
[0101] First, we tested whether the nasopharyngeal flora structure of NPC patients was significantly different from that of the control group. The analysis showed that the microbial flora structure of NPC patients was significantly different from that of the control group (Bray-Curtis distance matrix, PERMANOVA analysis, P<0.001, Figure 2 ), indicating that there are significant differences in the overall composition of the nasopharyngeal flora between NPC patients and the control group.
[0102] 1.6 Results: Differential microbiome characteristics between NPC patients and healthy controls
[0103] In order to further explore the characteristics of NPC-related flora, ANCOM-BC analysis was performed to determine the differentially abundant bacterial species between NPC patients and the control group. The analysis corrected for potential confounding factors such as gender, age, smoking, drinking, and oral nasopharyngeal diseases. A total of 533 species were detected in the nasopharynx, and 97 species with a detection rate higher than 5% were tested for differences. It was found that the relative abundance of 25 species was significantly different between NPC patients and the control group (p<0.05). The results of the species-level difference test are shown in Table 1. Among them, 18 species of Corynebacterium were detected, and only Corynebacterium accolens was statistically significant; 25 species of Staphylococcus were detected, and only Staphylococcus epidermidis and Staphylococcus hominis were statistically significant; 17 species of Streptococcus were detected, and only Streptococcus mitis was statistically significant; 25 species of Prevotella were detected, and only Prevotella nanceiensis was statistically significant.
[0104] Table 1: Differential bacterial flora characteristics in the nasopharyngeal microbiota between NPC cases and controls
[0105]
[0106]
[0107] Example 2: Evaluation of the effect of significantly different nasopharyngeal flora on distinguishing nasopharyngeal carcinoma patients from healthy controls
[0108] Based on the samples and data involved in Example 1, this example evaluates the predictive effect of the model constructed by each flora feature alone and in different combinations on nasopharyngeal carcinoma patients and control populations. We randomly divided the sample into a training set and a test set in a ratio of 6:4, and independently verified the model constructed in the training set in the test set. The test set includes 39 nasopharyngeal carcinoma patients and 50 controls, and the validation set includes 31 nasopharyngeal carcinoma patients and 36 healthy controls. The following is the specific implementation content:
[0109] 2.1: The discriminative effect of a single differential bacterial flora feature on NPC and control
[0110] First, the 25 nasopharyngeal carcinoma differential microbial characteristics described in Example 1 were independently modeled using a Logistic regression model. The independent variable of the model is the relative abundance information of each microbial characteristic, and the dependent variable is the disease group (whether the patient is nasopharyngeal carcinoma). First, the model is constructed in the training set population, and the model prediction effect of the training set population is calculated. The constructed model is then verified in the test set, and its prediction effect in the independent population is calculated. A total of 25 prediction models (Model 1 to Model 25) were established based on the characteristics of each microbial community. The enrichment group, detection rate and model prediction information of each microbial community indicator are shown in Table 3.
[0111] Table 3: Prediction model and area under the curve of nasopharyngeal flora species with significant differences between NPC cases and controls
[0112]
[0113]
[0114] Table Notes: *(The x variable in the Logistic regression model equation is the relative abundance of the bacterial community characteristics represented by the row. is the model predicted probability value calculated by the regression model).
[0115] The receiver operating characteristic curves of each model, the optimal cutoff value and its corresponding sensitivity and specificity information are shown in Figure 3. There are 7 species with AUC>0.6 for individual bacterial community features in the training set, including: 1) Cutibacterium acnes, Corynebacterium accolens, Cutibacterium granulosum, Staphylococcus epidermidis, Streptococcus mitis, Gemella haemolysans and Peptostreptococcus stomatis, which are enriched in the control group. Among them, there are 3 bacterial community features with AUC>0.6 for both the training set and the test set, including: 1) Cutibacterium acnes: the AUC of the training set is 0.714, and the AUC of the test set is 0.634; when the optimal cutoff value of the training set is 0.522, the sensitivity of the training set is 76.9%, and the specificity is 66.0%. At this time, the sensitivity of the test set is 45.2%, and the specificity of the test set is 72.2%; 2) Corynebacterium accolens: The AUC of the training set is 0.702, and the AUC of the test set is 0.656; when the optimal cutoff value of the training set is 0.515, the sensitivity of the training set is 69.2%, and the specificity is 68.0%. At this time, the sensitivity of the test set is 48.4%, and the specificity of the test set is 72.2%; 3) Cutibacterium granulosum: The AUC of the training set is 0.657, and the AUC of the test set is 0.656; when the optimal cutoff value of the training set is 0.511, the sensitivity of the training set is 61.5%, and the specificity is 68.0%. At this time, the sensitivity of the test set is 45.2%, and the specificity of the test set is 80.6%.
[0116] From the AUC of a single bacterial community feature, we can see that different bacterial community features have different patterns in identifying NPC patients. For example, the AUC of bacterial community features such as Peptostreptococcus stomatis and Gemella haemolysans is significantly left-biased, showing excellent specificity (98% to 100%); the AUC of Staphylococcus hominis is significantly right-biased, with 87.4% and 96.8% sensitivity in the training set and test set. Cutibacterium acnes, Corynebacteriumaccolens and Cutibacterium granulosum have advantages in comprehensive discrimination ability. The above results suggest that the combination of multi-microbial community markers will hopefully achieve complementary advantages of markers and improve the accuracy of the bacterial community model in discriminating NPC.
[0117] 2.2: Construction of a diagnostic model for nasopharyngeal carcinoma based on 11 nasopharyngeal flora characteristics
[0118] Furthermore, we used the LASSO model for further feature selection, specifically: using the glmnet function in R language, setting the parameter family = "binomial"; at the same time, using the cv.glmnet function to perform a 5-fold cross-validation to optimize the model lamda value (lamda value is 1se), and further select 11 flora characteristics for the next stage of prediction model. The number of times the selected flora characteristics were included in the model in 100 LASSO model iterations is shown in Figure 4 The relative abundance information of the above bacterial flora characteristics was used to construct a logistic regression model to judge NPC patients and control populations.
[0119] The formula of the model constructed by 11 bacterial communities (Model 26) is as follows:
[0120] Model 26:
[0121] The constructed model Model26 showed good predictive effect in distinguishing NPC patients from control populations, with the area under the curve (AUC) of the training set = 0.922 (95% CI: 0.870-0.974), and the AUC of the test set was 0.852 (95% CI: 0.757-0.947); when the optimal cutoff value of the training set was 0.432, the sensitivity was 87.2% and the specificity was 82.0%, while the sensitivity of the test set was 71.0% and the specificity of the test set was 80.6%. The receiver operating characteristic curves of the training set and the test set are shown in the figure. Figure 5 .
[0122] 2.3: Construction of a diagnostic model for nasopharyngeal carcinoma based on two combinations of nasopharyngeal flora characteristics
[0123] In order to test the predictive effect of models with different numbers and different flora characteristics, the performance of modeling with two flora characteristics and their combination for nasopharyngeal carcinoma diagnosis was first tested, and the pairwise combination of the 11 flora characteristics screened above was modeled using a Logistic regression model. Among the 55 combinations tested, the model with the combination of Cutibacterium acnes and Streptococcus mitis flora characteristics performed best, with an AUC of 0.808 in the training set and an AUC of 0.648 in the test set; when the optimal cutoff value of the model was 0.494, the sensitivity of the training set was 77% and the specificity was 76%; the sensitivity of the test set was 48% and the specificity was 72%. The model formula is:
[0124] Model 27:
[0125] The receiver operating characteristic curves of Model27 in the training set and test set are shown in Figure 6 .
[0126] It is worth noting that among the models incorporating the two loci, the top three models in the training set AUC effect all included Cutibacterium acnes, and the effect of this bacterium in combined modeling with Streptococcus mitis and Gemella haemolysans was significantly higher than that of independent modeling ( Figure 7 ). The model constructed by Cutibacterium acnes and Gemella haemolysans had an AUC of 0.778 in the training set and an AUC of 0.683 in the test set. When the optimal cutoff value of the model was 0.466, the sensitivity of the training set was 82% and the specificity was 66%. The sensitivity of the test set was 52% and the specificity was 72%. The model constructed by Cutibacterium acnes and Limnobacter thiooxidans had an AUC of 0.765 in the training set and an AUC of 0.689 in the test set. When the optimal cutoff value of the model was 0.490, the sensitivity of the training set was 82% and the specificity was 66%. The sensitivity of the test set was 55% and the specificity was 69%.
[0127] In addition, by counting the number of times the microbial characteristic variables were included in the top ten, top twenty, and top thirty models, it was found that when two-point modeling was performed, the microbial characteristics that were most often included in the model included Cutibacteriumacnes, Corynebacterium accolens, Peptostreptococcus stomatis, and Streptococcusmitis ( Figure 8 ), indicating that the diagnostic efficacy is better when the above flora characteristics are jointly modeled.
[0128] 2.4: Construction of a diagnostic model for nasopharyngeal carcinoma based on three nasopharyngeal flora characteristics
[0129] We tested the performance of modeling the diagnosis of NPC by incorporating three bacterial community characteristics and their combination. Three of the 11 bacterial community characteristics selected above were modeled using the Logistic regression model, and a total of 165 combinations were tested. Among them, the combination modeling of the three bacterial community characteristics of Cutibacterium acnes, Streptococcus mitis and Limnobacter thiooxidans performed best among the 165 combinations, with an AUC of 0.849 in the training set and an AUC of 0.701 in the test set; when the optimal cutoff value of the model was 0.458, the sensitivity of the training set was 85% and the specificity was 76%; the sensitivity of the test set was 58% and the specificity was 69%. The model formula is:
[0130] Model 28:
[0131] The receiver operating characteristic curves of Model28 in the training set and test set are shown in Fig. 9 .
[0132] Among the models incorporating the three bacterial community characteristics, the top three models in terms of AUC of the training set all included Cutibacteriumacnes and Streptococcus mitis. On this basis, Limnobacter thiooxidans, Peptostreptococcus stomatis and Gemella morbillorum ( Fig.10 ). The model constructed by Cutibacteriumacnes, Streptococcus mitis and Peptostreptococcus stomatis had an AUC of 0.835 in the training set and an AUC of 0.670 in the test set. When the optimal cutoff value of the model was 0.446, the sensitivity of the training set was 80% and the specificity was 76%. The sensitivity of the test set was 52% and the specificity was 72%. The model constructed by Cutibacterium acnes, Streptococcus mitis and Gemella morbillorum had an AUC of 0.835 in the training set and an AUC of 0.649 in the test set. When the optimal cutoff value of the model was 0.464, the sensitivity of the training set was 79% and the specificity was 76%. The sensitivity of the test set was 48% and the specificity was 72%.
[0133] By counting the number of times the microbial characteristic variables were included in the top ten, top twenty, and top thirty models, it was found that when three microbial characteristic variables were modeled, the microbial characteristics that were included most often included Cutibacteriumacnes, Streptococcus mitis, Limnobacter thiooxidans, Corynebacterium accolens, and Gemella haemolysans ( Fig.11 ), indicating that the combination of the above microbial characteristics is expected to improve the accuracy of the model.
[0134] 2.5: Construction of a diagnostic model for nasopharyngeal carcinoma based on four nasopharyngeal flora characteristics
[0135] We tested the performance of modeling nasopharyngeal carcinoma diagnosis by incorporating four bacterial community characteristics and their combination. Four of the 11 bacterial community characteristics screened above were modeled using the Logistic regression model, and a total of 330 combinations were tested. Among them, the combined modeling of four bacterial community characteristics, Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans, and Peptostreptococcus stomatis, performed best among the 330 combinations, with an AUC of 0.876 in the training set and an AUC of 0.723 in the test set; when the optimal cutoff value of the model was 0.403, the sensitivity of the training set was 85% and the specificity was 76%; the sensitivity of the test set was 61% and the specificity was 69%. The model formula is:
[0136] Model 29:
[0137] The receiver operating characteristic curves of Model29 in the training set and test set are shown in Fig.12 .
[0138] The top three models with the best AUC results in the training set of the models incorporating the four bacterial community characteristics all included Cutibacteriumacnes, Streptococcus mitis and Limnobacter thiooxidans, and on this basis, Peptostreptococcus stomatis, Gemella morbillorum and Klebsiellapneumoniae ( Fig.13). The model constructed by Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans and Gemellamorbillorum had an AUC of 0.874 in the training set and an AUC of 0.704 in the test set. When the optimal cutoff value of the model was 0.423, the sensitivity of the training set was 85%, and the specificity was 76%. The sensitivity of the test set was 58%, and the specificity was 69%. The model constructed by Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans and Klebsiellapneumoniae had an AUC of 0.870 in the training set and an AUC of 0.738 in the test set. When the optimal cutoff value of the model was 0.430, the sensitivity of the training set was 87%, and the specificity was 76%. The sensitivity of the test set was 65%, and the specificity was 69%.
[0139] By counting the number of times the microbial characteristic variables were included in the top ten, top twenty, and top thirty models, it was found that when modeling the characteristics of the four bacterial species, the most frequently included microbial characteristics included Cutibacteriumacnes, Limnobacter thiooxidans, Streptococcus mitis, Corynebacterium accolens, Peptostreptococcus stomatis, and Gemella haemolysans ( Fig.14 ), indicating that the above bacterial community characteristics are more important in the model and the multi-point joint effect is stronger.
[0140] 2.6: Construction of a diagnostic model for nasopharyngeal carcinoma based on a combination of five nasopharyngeal flora characteristics
[0141] The performance of modeling the diagnosis of nasopharyngeal carcinoma by incorporating five bacterial community characteristics and their combination was tested. Five of the 11 bacterial community characteristics screened above were modeled using the Logistic regression model, and a total of 462 combinations were tested. Among them, the combined modeling of five bacterial community characteristics, Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans, Peptostreptococcus stomatis, and Corynebacterium accolens, performed best among the 462 combinations, with an AUC of 0.901 in the training set and an AUC of 0.768 in the test set; when the optimal cutoff value of the model was 0.447, the sensitivity of the training set was 87% and the specificity was 86%; the sensitivity of the test set was 55% and the specificity was 83%. The model formula is: Model 30:
[0142] The receiver operating characteristic curves of the training set and the test set are shown in Fig.15 .
[0143] The top three models with the best AUC results in the training set of the models incorporating the five bacterial community characteristics all included Cutibacteriumacnes, Streptococcus mitis, Limnobacter thiooxidans and Corynebacterium accolens, and on this basis, they were combined with Peptostreptococcus stomatis, Gemella morbillorum and Klebsiellapneumoniae ( Fig.16). The model constructed by combining Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans, Corynebacterium accolens and Gemella morbillorum had an AUC of 0.900 in the training set and an AUC of 0.753 in the test set. When the optimal cutoff value of the model was 0.469, the sensitivity of the training set was 87% and the specificity was 86%; the sensitivity of the test set was 48% and the specificity was 83%. The AUC of the model constructed for Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans, Corynebacterium accolens and Klebsiella pneumoniae was 0.896 in the training set, and the AUC of the test set was 0.756; when the optimal cutoff value of the model was 0.366, the sensitivity of the training set was 90% and the specificity was 76%; the sensitivity of the test set was 68% and the specificity was 69%.
[0144] By counting the number of times the microbial characteristic variables were included in the top ten, top twenty, and top thirty models, it was found that when the five microbial characteristics were modeled, the microbial characteristics that were included most often included Cutibacteriumacnes, Limnobacter thiooxidans, Streptococcus mitis, Gemella morbillorum, Corynebacterium accolens, Peptostreptococcus stomatis, and Gemella haemolysans ( Fig.17 ).
[0145] 2.7: Construction of a diagnostic model for nasopharyngeal carcinoma based on a combination of six nasopharyngeal flora characteristics
[0146] We tested the performance of modeling the diagnosis of NPC by incorporating six bacterial community characteristics and their combination. Six of the 11 bacterial community characteristics selected above were modeled using the Logistic regression model, and a total of 462 combinations were tested. Among them, the combination modeling of six bacterial community characteristics, Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans, Peptostreptococcus stomatis, Corynebacterium accolens, and Klebsiellapneumoniae, performed best among the 462 combinations, with an AUC of 0.909 in the training set and 0.783 in the test set; when the optimal cutoff value of the model was 0.430, the sensitivity of the training set was 85% and the specificity was 90%; the sensitivity of the test set was 61% and the specificity was 86%. The model formula is:
[0147] Model 31:
[0148] The receiver operating characteristic curves of the training set and the test set are shown in Fig.18 .
[0149] Among the models incorporating six bacterial community characteristics, the model with the second-highest AUC effect in the training set was the model that included Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans, Peptostreptococcus stomatis, Corynebacterium accolens, and Gemella morbillorum, with an AUC of 0.907 in the training set and an AUC of 0.759 in the test set. When the optimal cutoff value of the model was 0.415, the sensitivity of the training set was 87% and the specificity was 86%; the sensitivity of the test set was 55% and the specificity was 83%.
[0150] By counting the number of times the microbial characteristic variables were included in the top ten, top twenty, and top thirty models, it was found that when the six microbial characteristics were modeled, the microbial characteristics that were included most often included Cutibacteriumacnes, Limnobacter thiooxidans, Streptococcus mitis, Gemella morbillorum, Corynebacterium accolens, Peptostreptococcus stomatis, and Gemella morbillorum ( Fig.19 ), indicating that the diagnostic efficacy is better when the above flora characteristics are jointly modeled.
[0151] 2.8 Effect of incorporating different numbers of bacterial community features on model prediction
[0152] In order to see the changes in the prediction effect of the model incorporating different numbers of microbial features, the combination of all models incorporating 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 and 11 microbial features was calculated, and the model with the highest AUC in the test set with different numbers of features was selected for display. It can be seen that compared with incorporating only one feature, the combination of multiple features can significantly improve the accuracy of the prediction of the test set and the validation set ( Fig. 20 ).
[0153] When the number of features included is increased to 8 bacterial community features, the model performance in the training set and the test set reaches a plateau ( Fig. 20 ). The model (Model32) included Corynebacterium accolens, Cutibacterium acnes, Limnobacter thiooxidans, Peptostreptococcus stomatis, Gemella haemolysans, Klebsiella pneumoniae, Ralstonia_insidiosa and Granulicatella_adiacens, with an AUC of 0.922 in the training set and 0.816 in the test set; when the optimal cutoff value of the model was 0.420, the sensitivity of the training set was 87% and the specificity was 86%; the sensitivity of the test set was 68% and the specificity was 89%. The model formula is:
[0154] Model 32:
[0155] The receiver operating characteristic curves of the training set and the test set are shown in Fig.21 。
Claims
1. Application of reagents for detecting bacterial markers in the preparation of diagnostic products for nasopharyngeal carcinoma, It is characterized in that The bacterial marker is Granulicatella adiacens , Cutibacteriumacnes , Gemellahaemolysans , Corynebacteriumaccolens , Ralstoniainsidiosa , Neisseriasicca , Klebsiella pneumoniae 、Streptococcus mitis、 Limnobacterthiooxidans, Peptostreptococcusstomatis and Gemella morbillorum The reagent for detecting bacterial species markers is a reagent for quantitatively detecting the abundance of bacterial species markers.
2. Application of reagents for detecting bacterial markers in the preparation of diagnostic products for nasopharyngeal carcinoma, It is characterized in that The bacterial species marker is Cutibacterium acnes and Streptococcus mitis The reagent for detecting bacterial species markers is a reagent for quantitatively detecting the abundance of bacterial species markers.
3. Application of reagents for detecting bacterial markers in the preparation of diagnostic products for nasopharyngeal carcinoma, It is characterized in that The bacterial species marker is Cutibacterium acnes, Streptococcus mitis and Limnobacter thiooxidans The reagent for detecting bacterial species markers is a reagent for quantitatively detecting the abundance of bacterial species markers.
4. Application of reagents for detecting bacterial markers in the preparation of diagnostic products for nasopharyngeal carcinoma, It is characterized in that The bacterial species marker is Cutibacterium acnes, Streptococcus mitis, Limnobacter thiooxidans and Peptostreptococcus stomatis The reagent for detecting bacterial species markers is a reagent for quantitatively detecting the abundance of bacterial species markers.
5. Application of reagents for detecting bacterial markers in the preparation of diagnostic products for nasopharyngeal carcinoma. It is characterized in that The bacterial marker is Cutibacterium acnes , Streptococcus mitis , Limnobacter thiooxidans , Peptostreptococcus stomatis and Corynebacterium accolens The reagent for detecting bacterial species markers is a reagent for quantitatively detecting the abundance of bacterial species markers.
6. Application of reagents for detecting bacterial markers in the preparation of diagnostic products for nasopharyngeal carcinoma. It is characterized in that The bacterial species marker is Cutibacterium acnes , Streptococcus mitis , Limnobacter thiooxidans , Peptostreptococcus stomatis , Corynebacterium accolens and Klebsiella pneumoniae The reagent for detecting bacterial species markers is a reagent for quantitatively detecting the abundance of bacterial species markers.
7. Application of reagents for detecting bacterial markers in the preparation of diagnostic products for nasopharyngeal carcinoma. It is characterized in that The bacterial species marker is Corynebacterium accolens , Cutibacterium acnes , Limnobacter thiooxidans , Peptostreptococcus stomatis , Gemella haemolysans , Klebsiella pneumoniae , Ralstonia_insidiosa and Granulicatella_adiacens The reagent for detecting bacterial species markers is a reagent for quantitatively detecting the abundance of bacterial species markers.
8. The use according to any one of claims 1 to 7, It is characterized in that The reagents for detecting bacterial species markers are primers, probes, aptamers and / or antibodies.
9. The use according to any one of claims 1 to 7, It is characterized in that The reagent for detecting bacterial species markers is a PCR detection reagent.
10. The use according to any one of claims 1 to 7, It is characterized in that The reagent for detecting bacterial species markers is a reagent for high-throughput sequencing.
11. The use according to any one of claims 1 to 7, It is characterized in that The reagent for detecting bacterial species markers is a reagent for 16SrRNA detection.
12. The use according to any one of claims 1 to 7, It is characterized in that The reagent for detecting bacterial species markers is a reagent for RNAseq detection.
13. The use according to any one of claims 1 to 7, It is characterized in that The reagent for detecting bacterial species markers is a reagent for metagenomic sequencing.
14. The use according to any one of claims 1 to 7, It is characterized in that The sampling site of the bacteria is the nasopharynx.
15. The use according to any one of claims 1 to 7, It is characterized in that The product is a reagent.
16. The use according to any one of claims 1 to 7, It is characterized in that The product is a test kit.
17. The use according to any one of claims 1 to 7, It is characterized in that The product is a test strip, a chip and / or an electronic device.