Plasma circulating microbial markers for early screening of gastric cancer and application thereof
By combining plasma circulating microbial biomarkers and Low Pass WGS technology, a machine learning model was constructed, which solved the problems of invasiveness and low sensitivity of existing gastric cancer screening methods, achieving efficient and low-cost early gastric cancer screening and improving diagnostic accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING XUTENG GENE TECHNOLOGY CO LTD
- Filing Date
- 2025-07-11
- Publication Date
- 2026-05-26
AI Technical Summary
Existing gastric cancer screening methods, such as endoscopy, imaging examinations, and serum biomarker testing, are invasive, have low sensitivity or low specificity, and are difficult to meet the needs of large-scale early screening. Furthermore, existing ctDNA testing is costly and susceptible to wild-type DNA interference, and traditional microbial biomarkers, such as Helicobacter pylori testing, lack specificity.
A combination of circulating plasma microbial biomarkers, including 10 microorganisms such as Raulella solanaceae, Staphylococcus aureus, Alternaria alternata, and Porphyromonas gingivalis, was used in conjunction with Low Pass WGS technology to construct a machine learning model for gastric cancer diagnosis. Early screening was performed by detecting the relative abundance of these microorganisms.
It enables high-sensitivity and high-specificity early screening of gastric cancer under low-depth sequencing conditions, reduces detection costs, effectively distinguishes gastric cancer patients from healthy individuals, and improves the early diagnosis rate.
Smart Images

Figure CN120866523B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical detection technology, and in particular to plasma circulating microbial markers for early screening of gastric cancer and their applications. Background Technology
[0002] The prognosis of gastric cancer is closely related to the timing of diagnosis. The 5-year survival rate for patients with early-stage gastric cancer can reach over 95%, while that for patients with late-stage cancer drops sharply to less than 30%, highlighting the importance of early screening.
[0003] Currently, commonly used clinical methods for gastric cancer screening mainly include endoscopy, imaging examinations, and serum biomarker testing. Endoscopy, as the gold standard for diagnosis, has high sensitivity (85%-90%), but due to its invasive nature, patient acceptance is low, with a compliance rate of less than 40%. Imaging examinations such as CT and MRI have a detection rate of only 30-50% for early-stage gastric cancer, and the interpretation of results is highly subjective. While serum biomarker testing is simple to perform, traditional biomarkers such as CEA and CA72-4 have poor specificity (<65%), and the positive detection rate for advanced gastric cancer is only 40-60%, which is insufficient to meet clinical needs. These clinical testing methods are primarily used for high-risk groups or patients with existing symptoms, generally discovering the cancer at an intermediate or advanced stage, and are not suitable for large-scale early screening.
[0004] In recent years, molecular detection technologies have shown promising prospects in the field of early cancer screening. One commonly used method is liquid biopsy of ctDNA. ctDNA refers to cell-free DNA fragments derived from the tumor genome in human peripheral blood. This non-invasive sampling method involves collecting blood from the subject for testing, and can rely on next-generation sequencing (NGS) or quantitative real-time PCR (qPCR) technology. ctDNA detection based on qPCR technology has high sensitivity and specificity, primarily detecting specific gene loci. Currently, commonly used detection loci include SLC6A3, Septin9, and RNF180, but their gene and detection site coverage is relatively low, making them unsuitable for large-scale gastric cancer screening. In recent years, liquid biopsy of ctDNA based on high-throughput NGS sequencing has become increasingly widely used. Its advantages include the ability to detect numerous genes and loci, high coverage, high throughput, and the ability to process large numbers of samples at once, making it ideal for large-scale screening while also maintaining high sensitivity and specificity. The main sequencing methods for liquid biopsy of ctDNA using NGS next-generation sequencing include targeted sequencing, whole-genome sequencing, and whole-exome sequencing. However, ctDNA usually accounts for only 0.01% to 1.00% of cell-free DNA (cfDNA). High-depth sequencing data is usually required to meet the detection needs during the sequencing process, which is costly and susceptible to interference from wild-type DNA due to issues such as clonal hematopoiesis.
[0005] While Helicobacter pylori (Hp) is closely associated with the development of gastric cancer, its high carrier rate in healthy individuals (>50% in Asia) makes standalone testing lack specificity. Other gastric juice microbial tests require invasive sampling, limiting their clinical application. Therefore, identifying new gastric cancer screening biomarkers and technologies, especially those for early warning monitoring and diagnosis, is crucial for improving the early diagnosis rate of gastric cancer, enabling early intervention and treatment, and reducing gastric cancer mortality. Summary of the Invention
[0006] In view of the shortcomings of the prior art, the purpose of this invention is to provide plasma circulating microbial markers for early screening of gastric cancer and their applications.
[0007] In the first aspect, the circulating microbial markers in the plasma include: Ralstonia solanacearum, Staphylococcus aureus, Alternaria metachromatica, Porphyromonas gingivalis, Lactobacillus reuteri, Rhodococcus erythropolis, Akkermansia muciniphila, Lactobacillus gasseri, Bifidobacterium bifidum, and Neisseria gonorrhoeae.
[0008] The inventors discovered that the above 10 microorganisms have obvious tumor specificity. When the above 10 microorganisms are included, the markers have high sensitivity. Combined with Low Pass WGS technology, detection can be achieved at low depths.
[0009] Secondly, the present invention provides the use of the plasma circulating microbial markers, the uses of which include constructing diagnostic models for predicting gastric cancer and preparing gastric cancer diagnostic kits.
[0010] Thirdly, the present invention also provides the application of reagents for detecting microbial abundance in the preparation of gastric cancer diagnostic kits, wherein the microorganisms are the following 10 microorganisms: Ralstonia solanacearum, Staphylococcus aureus, Alternaria metachromatica, Porphyromonas gingivalis, Lactobacillus reuteri, Rhodococcus erythropolis, Akkermansia muciniphila, Lactobacillus gasseri, Bifidobacterium bifidum, and Neisseria gonorrhoeae.
[0011] Fourthly, the present invention provides a method for constructing the diagnostic model for predicting gastric cancer, the method comprising the following steps:
[0012] (1): Obtain plasma samples from several gastric cancer patients and healthy individuals. All plasma samples were subjected to cfDNA extraction and Low Pass WGS sequencing. The relative abundance of the plasma circulating microbial markers described in claim 1 was detected by bioinformatics analysis.
[0013] (2): Divide the data obtained in step (1) into a training set and a validation set, input them into the machine learning model, optimize the parameters, train with the training set, validate with the validation set, and store the model.
[0014] Furthermore, the machine learning model described in step (2) is any one of Logistic Regression, Support Vector Machine, Random Forest, or Xbgoost.
[0015] Furthermore, the optimization parameters in step (2) are optimized based on cross-validation.
[0016] Fifthly, the present invention provides a diagnostic model for predicting gastric cancer, wherein the input variables of the model are the relative abundance of the following 10 circulating microbial markers in plasma: Ralstonia solanacearum, Staphylococcus aureus, Alternaria metachromatica, Porphyromonas gingivalis, Lactobacillus reuteri, Rhodococcus erythropolis, Akkermansia muciniphila, Lactobacillus gasseri, Bifidobacterium bifidum, and Neisseria gonorrhoeae. The following algorithm was used: the total abundance of the 10 microorganisms was standardized to 100%, and the relative abundance of each species was calculated (0% for undetected species); then, the relative abundance value (without a percentage sign) of each microorganism was multiplied by its corresponding microbial weight (Feature Importance); finally, all products were summed to obtain a comprehensive score. ROC curve analysis was used to select the optimal cutoff value for the comprehensive score; a score ≥16 was considered positive for gastric cancer.
[0017] The beneficial effects of this invention are as follows: This invention is the first to use circulating plasma microorganisms as molecular markers for early gastric cancer screening. It screens out and combines 10 specific microbial markers for the first time. Circulating plasma microorganisms mainly originate from microorganisms within tumor tissue and have the following unique advantages: they account for a relatively high proportion (1%-3.5%) in sequencing data, their signals are stable and unaffected by tumor stage; they are completely independent of the human genome, avoiding interference from clonal hematopoiesis; they have significant tumor specificity, enabling the identification of tumor origin; combined with Low Pass WGS technology, detection can be achieved at low depths. Therefore, this solution significantly reduces screening costs while ensuring detection sensitivity, providing a completely new solution for early gastric cancer screening. Attached Figure Description
[0018] Figure 1 Flowchart of blood plasma circulating microbial nucleic acid in gastric cancer screening and diagnosis;
[0019] Figure 2 For bioinformatics analysis workflow;
[0020] Figure 3 For screening indicators and microbial weights in plasma for gastric cancer;
[0021] Figure 4 To verify the ROC curve of the dataset. Detailed Implementation
[0022] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art are within the protection scope of the present invention.
[0023] Example
[0024] Establish and validate a screening and diagnostic model for gastric cancer based on circulating microorganisms in plasma, and design the process as follows: Figure 1 As shown, the specific steps are as follows:
[0025] ① Sample selection
[0026] Plasma samples were collected from 170 pathologically confirmed gastric cancer patients (who had not received preoperative treatment) and 90 healthy controls. All cancer cases were confirmed by imaging, laboratory testing, and pathology. Healthy controls were obtained from routine physical examinations. All samples underwent cfDNA extraction and Low Pass WGS sequencing. Samples were randomly assigned to a training set (120 cancer patients / 60 healthy individuals) and a validation set (50 cancer patients / 30 healthy individuals). There were no significant differences in age or sex between the two groups. Cancer patients were staged as follows:
[0027] Training set: 28 cases in stage I (23.33%), 52 cases in stage II (43.33%), and 40 cases in stage III and above (33.33%).
[0028] Validation set: 11 cases in stage I (22.00%), 23 cases in stage II (46.00%), and 16 cases in stage III and above (32.00%).
[0029] ② WGS sequencing of circulating microbial nucleic acids in plasma
[0030] The process includes key experimental procedures such as plasma separation, nucleic acid extraction, library construction, and sequencing.
[0031] Plasma separation: Balance the blood collection tubes and centrifuge at low speed (4℃, 1600g, 15min) to remove whole blood cells, collecting the supernatant plasma. Then, centrifuge at high speed (4℃, 16000g, 10min) to further remove platelets, cell debris, apoptotic bodies, etc., retaining free microbial nucleic acids in the supernatant. Approximately 3mL-5mL of plasma is obtained from each 10mL tube of peripheral blood. Highly hemolyzed samples should not be used.
[0032] Nucleic acid extraction: Plasma was treated with 1U DNase I (RNase-free) for 10 min (4℃) to degrade free host DNA while retaining microbial nucleic acids. Extraction was then performed after the treatment was terminated. cfDNA extraction was performed using the QIAamp Circulating Nucleic Acid Kit (Kaijie, 55114). Vector RNA was added during extraction to improve the recovery rate of small nucleic acid fragments and increase the enrichment of microbial nucleic acids.
[0033] Library construction: 50-200 ng of cfDNA was used for library construction. The process included end repair, magnetic bead purification, two-step ligation, and library amplification to obtain a WGS library. KAPA HiFi was used to reduce amplification bias during library amplification, with a cycle number of 10-12. Adapters with 8 bp paired-end units were added during library construction.
[0034] Sequencing: PE150 sequencing was performed on the MGISEQ-2000 platform (manufactured by BGI Genomics) following BGI's sequencing experimental protocol. Sterile water was added to each batch as a control to monitor reagent / environmental microbial contamination.
[0035] ③ Bioinformatics analysis
[0036] The workflow includes: first, using FASTP for data quality control (filtering sequences <50bp in length, low-quality bases, and adapter sequences) to obtain clean data; second, using Bowtie2 to align with the GRCh38 / hg38 reference genome to remove human sequences; and third, using BBmap to filter low-complexity sequences to obtain high-complexity non-human sequences. Microbial identification was performed by aligning with a self-built database (containing approximately 30,000 microbial reference sequences) using Kraken2, and relative abundance of microorganisms was calculated using Bracken. The decontamination process includes: 1) endogenous filtering (removing microorganisms with relative abundance <0.01% or absolute reads <10, as well as species with a detection rate <20% in similar samples); 2) exogenous filtering (based on blank and reagent controls, using Decontam software with a P < 0.5 threshold to remove contaminants). The analysis workflow is described below. Figure 2 .
[0037] ④ Screening for indicative microbial biomarkers for gastric cancer
[0038] The plasma microbial composition of gastric cancer patients and healthy controls was systematically compared using MaAsLin differential analysis after adjusting for sex and age. The analysis revealed significant differences in 38 microbial species between the two groups, with 26 species enriched in the gastric cancer group and 12 species enriched in the healthy control group.
[0039] Feature selection based on the random forest algorithm showed that the model performed optimally when using 10 microbial biomarkers (confirmed by five repetitions of 10-fold cross-validation), as shown in Table 1 and... Figure 3 As shown, the final determined biomarker combinations and microbial weights include: Ralstonia solanacearum (weight 0.4), Staphylococcus aureus (weight 0.36), Alternaria metachromatica (weight 0.31), Porphyromonas gingivalis (weight 0.28), Lactobacillus reuteri (weight 0.26), Rhodococcus erythropolis (weight 0.2), Akkermansia muciniphila (weight -0.28), Lactobacillus gasseri (weight -0.29), Bifidobacterium bifidum (weight -0.3), and Neisseria gonorrhoeae (weight -0.36). Among them, the following bacteria were enriched in the gastric cancer group: Ralstonia solanacearum, Staphylococcus aureus, Alternaria metachromatica, Porphyromonas gingivalis, Lactobacillus reuteri, and Rhodococcus erythropolis (with positive microbial weight values); while the following bacteria were enriched in the healthy group: Akkermansia muciniphila, Lactobacillus gasseri, Bifidobacterium bifidum, and Neisseria gonorrhoeae (with negative microbial weight values).
[0040] The calculation of microbial weights is based on the random forest algorithm, and the main processes are as follows: Using standardized relative abundance of microorganisms as input features and the gastric cancer / health label of the sample as the target variable, the weight of each microorganism is calculated using the Gini impurity method, which calculates the total reduction in impurity for that microorganism when splitting across all decision tree nodes. Simultaneously, a permutation importance verification method is used, observing the decrease in accuracy by randomly shuffling microbial values (a greater decrease indicates higher importance of the microorganism, i.e., a larger value). Microbial enrichment is assigned positive values to the gastric cancer group and negative values to the healthy group, thus establishing a biologically meaningful microbial weight system. The microbial weight values in gastric cancer plasma are shown in Table 1.
[0041] Table 1. Numerical Table of Plasma Microbial Weights in Gastric Cancer
[0042] Microbial marker names Microbial weight (Feature Importance) Ralstoniasolanacearum 0.4 Staphylococcusaureus 0.36 Alternariametachromatica 0.31 Porphyromonasgingivalis 0.28 Lactobacillusreuteri 0.26 Rhodococcus erythropolis 0.2 Akkermansiamuciniphila -0.28 Lactobacillusgasseri -0.29 Bifidobacterium bifidum -0.3 Neisseriagonorrhoeae -0.36
[0043] ⑤ Build the model and verify its performance
[0044] Based on 10 selected microbial biomarkers, a gastric cancer diagnostic model was constructed on both the training and validation sets using a random forest algorithm. The following algorithm was employed: the total abundance of the 10 microorganisms was standardized to 100%, and the relative abundance of each species was calculated (0% for undetected species); then, the relative abundance value (without a percentage sign) of each microorganism was multiplied by its corresponding feature importance; finally, all products were summed to obtain a comprehensive score. ROC curve analysis was used to select the optimal cutoff value for the comprehensive score; a score ≥16 was considered a positive result for gastric cancer. Figure 4 As shown, the model exhibits excellent discriminative performance: in the training set, the area under the ROC curve (AUC) reaches 0.993, while maintaining a sensitivity of 89.17% and a specificity of 93.33%, with good positive predictive values (96.40%) and negative predictive values (81.16%). In the independent validation set, the model also maintains high discriminative power, with an AUC of 0.986, a sensitivity of 86.00%, and a specificity of 90.00%, and the positive predictive values (87.50%) and negative predictive values (79.40%) remain stable. These results indicate that the diagnostic model constructed from this combination of microbial biomarkers can effectively distinguish between gastric cancer patients and healthy controls.
[0045] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. Circulating microbial nucleic acids from plasma for early screening of gastric cancer, characterized in that, The microorganisms consist of the following 10 species: Ralstonia solanacearum, Staphylococcus aureus, Alternaria metachromatica, Porphyromonas gingivalis, Lactobacillus reuteri, Rhodococcus erythropolis, Akkermansia muciniphila, Lactobacillus gasseri, Bifidobacterium bifidum, and Neisseria gonorrhoeae.
2. The use of circulating microbial nucleic acid in plasma as described in claim 1, characterized in that, The applications include building diagnostic models to predict gastric cancer.
3. The application of a reagent for detecting the abundance of microorganisms in circulating microbial nucleic acids in plasma in the preparation of a gastric cancer diagnostic kit, characterized in that, The microorganisms mentioned are the following 10 species: Ralstonia solanacearum, Staphylococcus aureus, Alternaria metachromatica, Porphyromonas gingivalis, Lactobacillus reuteri, Rhodococcus erythropolis, Akkermansia muciniphila, Lactobacillus gasseri, Bifidobacterium bifidum, and Neisseria gonorrhoeae.
4. The method for constructing the diagnostic model for predicting gastric cancer as described in claim 2, characterized in that, Includes the following steps: (1): Obtain plasma samples from several gastric cancer patients and healthy individuals. All plasma samples were subjected to cfDNA extraction and LowPass WGS sequencing. The relative abundance of microorganisms in the circulating microbial nucleic acid of the plasma as described in claim 1 was detected by bioinformatics analysis. (2): Divide the data obtained in step (1) into a training set and a validation set, input them into the machine learning model, optimize the parameters, train with the training set, validate with the validation set, and store the model.
5. The construction method according to claim 4, characterized in that, The machine learning model mentioned in step (2) is any one of Logistic Regression, Support Vector Machine, Random Forest or Xbgoost.
6. The construction method according to claim 4, characterized in that, The optimization parameters mentioned in step (2) are optimized based on cross-validation.
Citation Information
Patent Citations
CN118726620A
WO2022104278A1