Plasma circulating microbial markers for early screening of lung cancer and application thereof

By combining nine plasma circulating microbial biomarkers with low-depth whole-genome sequencing technology and machine learning algorithms, a high-precision lung cancer diagnostic model was constructed, overcoming the limitations of existing technologies in early lung cancer diagnosis and achieving efficient and low-cost early lung cancer screening.

CN120843702BActive Publication Date: 2026-04-21BEIJING XUTENG GENE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING XUTENG GENE TECHNOLOGY CO LTD
Filing Date
2025-07-15
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies for the early diagnosis of lung cancer suffer from problems such as low sensitivity of imaging examinations, high invasiveness of endoscopy, insufficient sensitivity of serum tumor markers, and unsuitability of tissue biopsy for large-scale screening. Liquid biopsy technologies, such as ctDNA testing, are costly, require large amounts of deep sequencing, and have high false positive and false negative rates, making it difficult to achieve large-scale early lung cancer screening.

Method used

Nine circulating microbial biomarkers (Staphylococcus aureus, Fusobacterium nucleatum, etc.) were combined with low-pass whole-genome sequencing (Low Pass WGS) technology. A high-precision lung cancer diagnostic model was constructed using machine learning algorithms to screen out specific microbial biomarkers and reduce detection costs.

Benefits of technology

It has achieved efficient identification of early-stage lung cancer, significantly reduced detection costs, and established a stable and clinically applicable diagnostic model suitable for large-scale population screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120843702B_ABST
    Figure CN120843702B_ABST
Patent Text Reader

Abstract

This invention relates to the field of molecular diagnostics and early tumor screening technology, specifically disclosing plasma circulating microbial biomarkers for early lung cancer screening and their applications. The biomarkers include: Staphylococcus aureus, Fusobacterium nucleatum, Klebsiella pneumoniae, Alternaria incomplexa, Legionella pneumophila, Neisseria mucosa, Ralstonia pickettii, Pasteurella multocida, and Haemophilus parainfluenzae. This invention screens and combines nine specific microbial biomarkers with low-pass whole-genome sequencing technology. The WGS (World Gastroenterology) assay detects nine characteristic microbial biomarkers and uses machine learning algorithms to build a high-precision diagnostic model. The established diagnostic model has excellent stability and clinical applicability, which can not only achieve efficient identification of early lung cancer, but also significantly reduce detection costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular diagnostics and early tumor screening technology, and in particular to plasma circulating microbial markers for early lung cancer screening and their applications. Background Technology

[0002] Early diagnosis of lung cancer faces significant challenges. Currently, commonly used clinical screening methods mainly include imaging examinations, endoscopy, tumor marker testing, and tissue biopsy. However, these methods all have obvious limitations: imaging examinations (such as CT scans) have a low detection rate for early small lesions (<1cm); endoscopy is invasive and expensive; serum tumor markers (such as CEA, CYFRA21-1, etc.) have insufficient sensitivity and specificity; and while tissue biopsy is accurate, it is highly invasive and unsuitable for large-scale screening.

[0003] In recent years, liquid biopsy technology has become a research hotspot for early cancer screening due to its advantages such as being non-invasive and repeatable. This technology mainly achieves early cancer diagnosis by detecting biomarkers such as circulating tumor DNA (ctDNA), circulating tumor cells (CTC), and exosomes in the blood. Among these, ctDNA detection technology is the most mature. Although PCR-based ctDNA detection methods have the advantages of low cost and short detection time, their throughput is low, and they only measure the methylation level of known genes at some target sites. They cannot detect multiple mutations, unknown mutations, or new mutations, making them unsuitable for large-scale cancer screening. The clinical application of NGS methods for ctDNA detection also faces many challenges: ctDNA content in blood is extremely low (only accounting for 0.01%-1% of cfDNA), and NGS panel detection and conventional WGBS detection require ultra-high depth sequencing (usually >10000X), resulting in high costs; the ctDNA signal released by early tumors is weak, easily leading to false negatives; and factors such as clonal hematopoiesis may cause false positive results. Finding new biomarkers for lung cancer, especially those for early warning, monitoring, and diagnosis, is of great significance for improving the early diagnosis rate of lung cancer, enabling early intervention and treatment, and reducing lung cancer mortality. Summary of the Invention

[0004] To address the problems mentioned in the background art, the present invention aims to provide plasma circulating microbial markers for lung cancer diagnosis and their applications.

[0005] This invention is achieved through the following scheme:

[0006] The plasma circulating microbial markers used for early lung cancer screening include the following nine microorganisms: Staphylococcus aureus, Fusobacterium nucleatum, Klebsiella pneumoniae, Alternaria incomplexa, Legionella pneumophila, Neisseria mucosa, Ralstonia pickettii, Pasteurella multocida, and Haemophilus parainfluenzae.

[0007] Circulating microbial DNA (cmDNA) offers unique advantages as a novel biomarker. Its concentration in blood is relatively stable, unaffected by tumor stage, and its microbial community characteristics exhibit high tumor specificity. The inventors discovered that when the biomarker contains the aforementioned nine microorganisms, its sensitivity is high. Combined with Low Pass WGS technology, detection can be achieved at low depths (typically <5X), significantly reducing costs.

[0008] Reagents for detecting the above-mentioned circulating microbial markers in plasma.

[0009] The uses of the above-mentioned circulating microbial markers in plasma or the above-mentioned reagents for detecting circulating microbial markers in plasma include the use of constructing diagnostic models for predicting lung cancer and preparing diagnostic kits for lung cancer.

[0010] A diagnostic kit for lung cancer, comprising the above-mentioned reagents for detecting circulating microbial markers in plasma.

[0011] A method for constructing a diagnostic model to predict lung cancer includes the following steps:

[0012] (1): Obtain plasma samples from several lung cancer patients and healthy individuals. All plasma samples were subjected to cfDNA extraction and Low Pass WGS sequencing. The relative abundance of the plasma circulating microbial markers described in claim 1 was detected by bioinformatics analysis.

[0013] (2): Divide the data obtained in step (1) into a training set and a validation set, input them into the machine learning model, optimize the parameters, train with the training set, validate with the validation set, and store the model.

[0014] Furthermore, the machine learning model described in step (2) is any one of Logistic Regression, Support Vector Machine, Random Forest, or Xbgoost.

[0015] Furthermore, the optimization parameters in step (2) are optimized based on cross-validation.

[0016] A diagnostic model for predicting lung cancer was constructed using the method described above.

[0017] The beneficial effects of this invention are as follows: It uses circulating plasma microorganisms as molecular markers for early lung cancer screening, screening out and combining nine specific microbial markers. These nine characteristic microbial markers are then detected using low-pass whole-genome sequencing (Low Pass WGS) technology. A high-precision diagnostic model is constructed using machine learning algorithms. The established diagnostic model exhibits excellent stability and clinical applicability, enabling efficient identification of early-stage lung cancer while significantly reducing detection costs. This method overcomes the sensitivity limitations of existing liquid biopsy techniques in early lung cancer detection, significantly reducing sequencing costs while ensuring detection accuracy, and providing a novel solution for large-scale lung cancer screening. Attached Figure Description

[0018] Figure 1 Flowchart of circulating microbial nucleic acids in plasma for lung cancer screening and diagnosis;

[0019] Figure 2 Basic Bioanalytical Procedures for Gastric Cancer Plasma Microbial Screening

[0020] Figure 3 For lung cancer screening indicators and microbial weights;

[0021] Figure 4 ROC curves for the training and validation sets. Detailed Implementation

[0022] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention are within the protection scope of the present invention.

[0023] Example

[0024] Establish and validate a screening and diagnostic model for lung cancer based on circulating microorganisms in plasma, and design the process as follows: Figure 1 As shown, the specific steps are as follows:

[0025] ① Sample selection

[0026] Plasma samples were collected from 230 pathologically confirmed lung cancer patients (who had not received preoperative treatment) and 90 healthy controls (collected using cfDNA anticoagulant tubes). All lung cancer cases were confirmed by imaging, laboratory testing, and pathology, while healthy controls were from routine physical examinations. All samples underwent cfDNA extraction and Low Pass WGS sequencing. The samples were divided into a training set (180 cancer patients / 60 healthy individuals) and a validation set (50 cancer patients / 30 healthy individuals), with no significant differences in age or sex between the two groups.

[0027] The implementation case aimed to detect early-stage cancer. Therefore, the pathologist confirmed the cancer type and stage by examining the corresponding tissue samples from the blood samples. The implementation case included a selection of early-stage cancer patients (Stage I / II). The cancer stage information for the training and validation sets is as follows:

[0028] Training set: 52 cases in stage I (28.89%), 72 cases in stage II (40.00%), and 56 cases in stage III and above (31.11%).

[0029] Validation set: 12 cases (24.00%) in stage I, 20 cases (40.00%) in stage II, and 18 cases (36.00%) in stage III and above.

[0030] ②Low-Pass WGS Experiment and Analysis Workflow

[0031] The wet experimental procedure for Low Pass WGS experiments mainly includes: plasma separation, nucleic acid extraction, library construction, and sequencing.

[0032] Plasma Separation: Peripheral blood samples should be collected using dedicated cell-free DNA collection tubes (e.g., Streck tubes), with a collection volume of 10 mL. Peripheral blood should be transported to the laboratory at a temperature controlled between 6℃ and 26℃ and processed for separation within 72 hours. Severely hemolyzed blood cannot be used for experiments. A two-step low-temperature separation method is used for plasma separation. The first step is low-speed centrifugation (4℃, 1600g, 15 min) to remove whole blood cells and collect the supernatant plasma. The second step is high-speed centrifugation (4℃, 16000g, 10-20 min) to further remove platelets, cell debris, apoptotic bodies, etc., retaining cell-free nucleic acids in the supernatant. 10 mL of peripheral blood typically yields 3 mL-5 mL of plasma.

[0033] Cell-free nucleic acid extraction: cfDNA was extracted using the QIAamp Circulating Nucleic Acid Kit (Kaijie, 55114). Vector RNA was added during extraction to improve the recovery rate of small nucleic acid fragments and increase the enrichment of microbial nucleic acids. The obtained cell-free nucleic acid was quantified using Qubit. When the total extraction volume was too high (>300 ng), fragment quality control was performed using an Agilent 2100 instrument. If only genomic DNA was extracted and no small fragments of cell-free cfDNA were found, it was not recommended to continue the experiment.

[0034] Library construction: The xGen Prism DNALibrary Prep Kit, KAPA HyperPure magnetic beads, HiFi Hotstart ReadyMix, and PCR Index Primer were used for library construction. The amount of free nucleic acid used was generally 50-200 ng. The experimental procedure included end repair, magnetic bead purification, ligation reaction 1 (20℃ for 15 min, 65℃ for 15 min, 4℃ hold), ligation reaction 2 (65℃ for 30 min, 4℃ hold), second magnetic bead purification, library amplification using KAPA HiFi Hotstart ReadyMix, and final magnetic bead purification to obtain the library sample. The number of library amplification cycles was 10-12, and the library construction included 8 bp paired-end adapters (umi).

[0035] Sequencing: The sequencing platform used was the MGISEQ-2000RS platform from BGI Genomics, and the sequencing strategy was PE150 (paired-end 150bp) sequencing. Sterile water was added as a control during each sequencing run to monitor reagent / environmental microbial contamination.

[0036] ③ Bioinformatics analysis workflow

[0037] MGISEQ-2000 sequencing data were processed using a self-developed bioinformatics analysis workflow: First, fastp was used for quality control, filtering low-quality (Q<15), short sequences (<50bp), and adapter contamination; then, Bowtie2 alignment with the GRCh38 reference genome was used to remove human sequences. To enrich microbial signals, BBmap was used to filter repetitive sequences and mitochondrial contamination, and based on a local database containing over 30,000 microorganisms, Kraken2 was used for k-mer classification and Bracken abundance correction. Key decontamination procedures included filtering low-abundance (<0.01% or reads<10) and low-frequency microorganisms (detection rate <20%), and using Decontam software in conjunction with a blank control (P<0.5) to remove exogenous contamination. This workflow effectively improved the specificity of microbial detection while ensuring data quality. The bioinformatics analysis workflow for lung cancer plasma microorganisms is described below. Figure 2 .

[0038] ④ Screening of lung cancer indicator microbial biomarker combinations

[0039] Based on a self-developed bioinformatics analysis workflow, microbiome sequencing analysis was performed on plasma samples from lung cancer patients and healthy controls. Forty differentially expressed microorganisms were identified through MaAsLin differential analysis (corrected for sex and age). Further analysis using a random forest algorithm combined with 10-fold cross-validation identified nine of the most discriminative microbial biomarkers, along with their corresponding weights, as shown in Table 1. Figure 3As shown, the weights for Staphylococcus aureus are 0.43, Fusobacterium nucleatum 0.3, Klebsiella pneumoniae 0.24, Alternaria incomplexa 0.19, Legionella pneumophila 0.12, Neisseria mucosa -0.1, Ralstonia pickettii -0.12, Pasteurella multocida -0.16, and Haemophilus parainfluenzae -0.4. Among them, Staphylococcus aureus, Fusobacterium nucleatum, Klebsiella pneumoniae, Alternaria incomplexa, and Legionella pneumophila were significantly enriched in lung cancer patients (microbial weights were positive), while Neisseria mucosa, Ralstonia pickettii, Pasteurella multocida, and Haemophilus parainfluenzae were more enriched in healthy controls (microbial weights were negative).

[0040] Microbial weights were calculated using a random forest algorithm: microbial abundance was used as the input feature, with directional labels assigning positive values ​​to microorganisms enriched in lung cancer and negative values ​​to those enriched in the healthy group. The Gini impurity method was used to calculate the weights of each microorganism. The accuracy of the analysis was then determined by randomly shuffling the microbial values; a greater decrease in accuracy indicated a higher microbial value and thus a higher level of importance. The weighted fusion of Gini importance and permutation importance was multiplied by the directional label (±1) to obtain the final signed microbial weights. The microbial weights in lung cancer plasma are shown in Table 1.

[0041] Table 1. Numerical table of plasma microbial weights for lung cancer patients.

[0042] Microbial name Microbial weight Staphylococcusaureus 0.43 Fusobacterium nucleatum 0.3 Klebsiellapneumoniae 0.24 Alternariaincomplexa 0.19 Legionella pneumophila 0.12 Neisseriamucosa -0.1 Ralstoniapickettii -0.12 Pasteurellamultocida -0.16

[0043] ⑤ Construct a diagnostic model based on lung cancer plasma microorganisms and validate its performance.

[0044] A lung cancer diagnostic model was constructed based on a combination of nine selected circulating microbial biomarkers in plasma, and systematically validated on the training and validation sets using a random forest algorithm. Figure 4 In model construction, the detection status of each microorganism (detected = 1, not detected = 0) is multiplied by its corresponding microbial weight value to calculate the total score. If the total score exceeds a preset threshold of 0.42, it is considered a positive result for lung cancer. Results show that the model has excellent diagnostic performance: in the training set, the AUC value reached 0.9987 (sensitivity 96.11%, specificity 96.67%), with positive predictive value of 98.86% and negative predictive value of 89.23%, demonstrating its accurate discrimination ability; the validation set results remained highly consistent, with an AUC value of 0.9945 (sensitivity 92.00%, specificity 93.33%), positive predictive value of 95.83%, and negative predictive value of 87.50%, further confirming the model's robustness. Statistical analysis showed that this biomarker combination was highly significant in distinguishing lung cancer patients from healthy individuals (P < 0.001), highlighting its important application value in clinical non-invasive screening.

[0045] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A plasma circulating microbial marker for early lung cancer screening, characterized in that, The following nine microorganisms are included: Staphylococcus aureus, Fusobacterium nucleatum, Klebsiella pneumoniae, Alternaria incomplexa, Legionella pneumophila, Neisseria mucosa, Ralstonia pickettii, Pasteurella multocida, and Haemophilus parainfluenzae.

2. A reagent for detecting the circulating microbial markers in plasma as described in claim 1.

3. The use of the plasma circulating microbial marker as described in claim 1, characterized in that, The plasma circulating microbial biomarkers were used to construct a diagnostic model for predicting lung cancer.

4. A diagnostic kit for lung cancer, characterized in that, Includes the reagent as described in claim 2.

5. The method for constructing a diagnostic model for predicting lung cancer as described in claim 3, characterized in that, Includes the following steps: (1): Obtain plasma samples from several lung cancer patients and healthy individuals. All plasma samples were subjected to cfDNA extraction and LowPass WGS sequencing. The relative abundance of the plasma circulating microbial markers described in claim 1 was detected by bioinformatics analysis. (2): Divide the data obtained in step (1) into a training set and a validation set, input them into the machine learning model, optimize the parameters, train with the training set, validate with the validation set, and store the model.

6. The construction method according to claim 5, characterized in that, The machine learning model mentioned in step (2) is any one of Logistic Regression, Support Vector Machine, Random Forest or Xbgoost.

7. The construction method according to claim 5, characterized in that, The optimization parameters mentioned in step (2) are optimized based on cross-validation.

Citation Information

Patent Citations

  • Lung cancer diagnosis marker and application thereof

    CN113913333A

  • Kit for simultaneously detecting lung cancer and pulmonary infection

    CN114525341A