Gastric cancer diagnosis model based on circulating microorganism DNA characteristics and application thereof
By combining the DNA characteristics of circulating microorganisms with machine learning models, especially those based on the DNA characteristics of Bacillus species, the early diagnosis of gastric cancer has been solved. By testing blood or plasma samples, a highly sensitive and convenient diagnosis of gastric cancer has been achieved, solving the problem of early diagnosis in existing technologies.
Patent Information
- Application Number
- CN202510789003.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies are insufficient for the early and accurate diagnosis of gastric cancer. Traditional detection methods are highly invasive or may result in late-stage expression, and existing molecular marker detection methods are difficult to accurately identify in tumor heterogeneity, lacking convenient alternatives.
By utilizing the DNA characteristics of circulating microorganisms, especially the relative abundance of microorganisms such as Bacillus and Burkholderia, and combining them with machine learning models such as random forests and generalized linear models, a gastric cancer diagnostic model is constructed, which is then detected through blood or plasma samples.
It achieves highly sensitive diagnosis of gastric cancer, especially showing excellent diagnostic performance in patients with normal serum tumor marker levels. The test is convenient and can be repeated multiple times, simplifying the operation process.
Smart Images

Figure BDA0005447851470000081 
Figure HDA0005447851480000011 
Figure HDA0005447851480000021
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and in particular to a gastric cancer diagnostic model based on the characteristics of circulating microbial DNA and its application. Background Technology
[0002] Gastric cancer (GC) is the third leading cause of cancer-related deaths worldwide, causing more than 720,000 deaths annually. Because most patients are diagnosed at an advanced stage or with metastasis, overall survival remains low. Therefore, early diagnosis of gastric cancer plays a crucial role in improving prognosis.
[0003] Although gastroscopy has been established as the gold standard for gastric cancer screening, its invasive nature limits its widespread adoption in communities. Furthermore, currently used serum biomarkers are preferentially expressed in advanced stages, making them difficult to apply to early diagnosis. Clearly, there is an urgent clinical need for accurate, cost-effective, and easily implemented methods for early detection of gastric cancer. It is worth noting that gastric cancer, as a highly heterogeneous disease, leads to differentiated responses to chemotherapy and immunotherapy. While molecular markers for advanced gastric cancer, such as human epidermal growth factor receptor 2 (HER2) and microsatellite instability (MSI), have been successfully used in clinical management, existing detection techniques such as fluorescence in situ hybridization (FISH) and immunohistochemistry (IHC) often struggle to accurately determine HER2 and MSI status due to tumor heterogeneity. In clinical practice, these methods also have limitations in reassessing recurrent patients, thus necessitating the development of more convenient alternatives.
[0004] Recent studies have shown that various tumor tissues, including prostate cancer, breast cancer, and colorectal cancer, possess characteristic microbial communities. With advancements in sequencing technology and molecular biology, numerous studies have confirmed the carcinogenic effects and diagnostic potential of symbiotic microbiota. Some researchers have proposed that circulating microbiome DNA (cmDNA) could serve as a novel non-invasive biomarker for cancer detection. However, to date, there are no studies on the use of cmDNA for the diagnosis of gastric cancer. Summary of the Invention
[0005] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a gastric cancer diagnostic model based on the characteristics of circulating microbial DNA and its application, in order to solve the problems in the prior art.
[0006] To achieve the above and other related objectives, the present invention is obtained through the following technical solution.
[0007] In a first aspect, the invention provides the use of a substance for detecting species characteristics of circulating microbial DNA (cmDNA) in blood or plasma in the preparation of a gastric cancer diagnostic product;
[0008] The microorganisms mentioned include Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae.
[0009] In this invention, "circulating microbial DNA" or "cmDNA" has the same meaning, referring to circulating cell-free DNA from microorganisms in the host's blood or plasma.
[0010] In some embodiments of the present invention, the substances for detecting biomarkers include reagents and / or instruments for detecting the aforementioned circulating microbial DNA by sequencing.
[0011] In some embodiments of the present invention, the sequencing includes, but is not limited to, at least one of 16S sequencing, metagenomic high-throughput sequencing, whole-genome sequencing, and PCR-pyrosequencing.
[0012] In some embodiments of the present invention, the substances in the detection marker combination are selected from at least one of the following groups: probes, gene chips, and PCR primers.
[0013] In some embodiments of the present invention, the product comprises reagents, kits, test strips, chips, or systems.
[0014] In a second aspect, the present invention provides a product comprising the aforementioned substance for detecting the species characteristics of circulating microbial DNA (cmDNA) in blood or plasma; the product is used for the diagnosis of gastric cancer.
[0015] In some embodiments of the present invention, the cmDNA species characteristic is the relative abundance at the microbial family level.
[0016] A third aspect of the present invention provides a method for constructing a gastric cancer diagnostic model, the method comprising: a step of constructing a model using the species characteristics of the circulating microbial DNA, wherein the microorganisms include Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae.
[0017] In this invention, "species characteristics" specifically refers to the relative abundance of microorganisms in the host's blood (especially plasma), and in some preferred embodiments, the relative abundance refers to the relative abundance at the species level.
[0018] In some embodiments of the present invention, the model includes a machine learning model; the machine learning model includes at least one of Random Forest (RF), Generalized Linear Model (GLM), Multilayer Perceptron (MLP), and Distributed Gradient Boosting Library (XGBoost); preferably Random Forest (RF).
[0019] In some embodiments of the present invention, the method for obtaining the cmDNA species characteristics is as follows:
[0020] 1) cfDNA sequencing: Collect plasma or blood from the sample to be tested to extract cfDNA, and sequence the cfDNA to obtain the cfDNA sequence;
[0021] 2) cmDNA acquisition: Align the cfDNA sequence to the host reference genome, filter out host gene fragments to obtain the blood microbial DNA fragment cmDNA;
[0022] 3) Species feature extraction: Microbial species information is annotated by sequence alignment based on cmDNA, and the relative abundance of microorganisms is evaluated to obtain cmDNA species features.
[0023] In some embodiments, the sequencing described in step 1) of this application includes at least one of 16S sequencing, metagenomic high-throughput sequencing, whole-genome sequencing, and PCR-pyrosequencing.
[0024] In some implementations, the whole-genome sequencing may be low-depth whole-genome sequencing.
[0025] In some specific implementations, the low-depth whole-genome sequencing can be 2× to 10× whole-genome sequencing.
[0026] In some implementations, the reference genome is host hg19 (Human Genome version 19).
[0027] In a fourth aspect of the invention, a system for diagnosing gastric cancer is provided, the system comprising:
[0028] Data acquisition module: Acquires the species characteristics of circulating microbial DNA in the blood or plasma of the sample to be tested, including Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae.
[0029] The result determination module uses a diagnostic model to diagnose gastric cancer in the test sample based on the obtained characteristics of circulating microbial DNA species.
[0030] In some embodiments of the present invention, the diagnostic model is constructed using the DNA species characteristics of the circulating microorganisms as feature information.
[0031] In some embodiments of the present invention, the diagnostic model is constructed by the method described in the third aspect of the present invention.
[0032] In some embodiments of the present invention, the sample to be tested includes a human or a non-human animal; the non-human animal is not limited, and includes, but is not limited to, rats, monkeys, pigs, cattle, sheep, rabbits, etc.
[0033] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0034] Data Acquisition: Acquire the species characteristics of circulating microbial DNA in the blood or plasma of the sample to be tested. The microorganisms include Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae.
[0035] Based on the obtained characteristics of circulating microbial DNA species, a diagnostic model was used to diagnose gastric cancer in the test samples.
[0036] In some embodiments of the present invention, the diagnostic model is constructed by the method described in the third aspect of the present invention.
[0037] The present invention also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0038] Data Acquisition: Acquire the species characteristics of circulating microbial DNA in the blood or plasma of the sample to be tested. The microorganisms include Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae.
[0039] Based on the obtained characteristics of circulating microbial DNA species, a diagnostic model was used to diagnose gastric cancer in the test samples.
[0040] In some embodiments of the present invention, the diagnostic model is constructed by the method described in the third aspect of the present invention.
[0041] Beneficial effects:
[0042] This application is the first to propose the use of circulating microbial DNA (cmDNA) characteristics from blood or plasma for the diagnosis of gastric cancer. The microorganisms include Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae. This invention identifies the species characteristics of cmDNA in blood or plasma and combines this with a machine learning model for gastric cancer diagnosis. The diagnostic model of this invention exhibits high diagnostic sensitivity even in gastric cancer patients with normal serum tumor marker levels and shows stable diagnostic efficacy in clinical subgroups of gastric cancer. Furthermore, compared to traditional tissue sampling, blood cmDNA sampling is more portable and allows for repeated sampling. Compared to traditional methods, this application also has advantages such as simple operation and short detection cycle. The biomarkers and gastric cancer diagnostic model in this application have good clinical application potential and value. Attached Figure Description
[0043] Figure 1 The ROC curves of the cmDNA-MLM model are shown in the training set, internal validation queue, and external validation queue.
[0044] Figure 2 The sensitivity of the cmDNA-MLM model in gastric cancer clinical subgroups stratified by serum tumor marker levels in internal and external validation cohorts.
[0045] Figure 3To assess the sensitivity of the cmDNA-MLM model in gastric cancer subgroups stratified based on clinicopathological characteristics (including age, sex, alcohol / smoking status, differentiation degree, and molecular subtype) in both internal and external validation cohorts. Detailed Implementation
[0046] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification.
[0047] Before further describing specific embodiments of the present invention, it should be understood that the scope of protection of the present invention is not limited to the specific embodiments described below; it should also be understood that the terminology used in the embodiments of the present invention is for describing specific embodiments and not for limiting the scope of protection of the present invention. Test methods in the following embodiments that do not specify specific conditions are generally performed under conventional conditions or as recommended by the respective manufacturers.
[0048] When numerical ranges are given in the embodiments, it should be understood that, unless otherwise stated in the present invention, both endpoints of each numerical range and any value between the two endpoints may be selected. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. In addition to the specific methods, apparatus, and materials used in the embodiments, based on the knowledge of the prior art possessed by one of ordinary skill in the art and the description of this invention, any prior art methods, apparatus, and materials similar to or equivalent to those described, apparatus, and materials in the embodiments of this invention may be used to implement the present invention.
[0049] The statistical method used in this embodiment is as follows:
[0050] Wilcoxon test was used to compare relative abundance differences at different phylogenetic levels among different cohorts. A two-sided p-value < 0.05 was considered statistically significant. All statistical analyses were performed in R software (version 4.3.0). The Wilcoxon rank-sum test (Mann-Whitney U test) was used for nonparametric comparisons between groups.
[0051] Example 1
[0052] I. Sample Enrollment in this Study
[0053] This study incorporated circulating microbiome DNA (cmDNA) samples from the plasma of participants at Zhejiang Cancer Hospital to construct a comprehensive dataset, including both non-gastric cancer and gastric cancer patient groups. Subsequently, a training dataset and an internal validation cohort were formed through random allocation.
[0054] This study also enrolled relevant subjects from three medical institutions—Zhejiang Cancer Hospital, Shanghai Yangpu Hospital, and Chongqing Yubei District People's Hospital—based on the aforementioned grouping criteria, forming an external validation cohort for model performance evaluation.
[0055] All cases were confirmed by histopathology, and the specific grouping criteria are as follows:
[0056] Non-gastric cancer patient group: healthy individuals or patients with gastritis, chronic atrophic gastritis, gastric polyps, or benign gastric hyperplasia confirmed by endoscopic biopsy;
[0057] Gastric cancer patient group: Patients diagnosed with gastric adenocarcinoma (including signet ring cell carcinoma), gastric squamous cell carcinoma, and gastric neuroendocrine carcinoma.
[0058] Blood samples were collected between 2016 and 2024. Plasma samples from gastric cancer patients were collected after diagnosis and before systemic treatment (chemotherapy / surgery / radiotherapy).
[0059] Exclusion criteria included patients with other malignant tumors (including lymphoma) and samples with low microbial abundance as shown by sequencing.
[0060] All samples were collected in the morning on an empty stomach to minimize interference from dietary intake and circadian rhythms on cmDNA. Gastric cancer staging was determined based on the AJCC (8th edition) criteria, combined with imaging and pathological biopsy results.
[0061] This study passed the ethical review of three institutions: Zhejiang Cancer Hospital (IDs IRB-2024-229 and IRB-2021-68); Shanghai Yangpu Hospital (ID LL-2024-ZRKX-057); and Chongqing Yubei District People's Hospital (ID 2024B02B02).
[0062] II. Blood Collection and Plasma Processing
[0063] Peripheral blood samples were collected using cell-free DNA preservation tubes (Ardent BioMed, Guangdong, China, catalog number #BY10240301).
[0064] First, fresh blood was centrifuged at 1,600×g for 10 minutes at 4°C. The separated plasma supernatant was then transferred to a new centrifuge tube using a pipette. Second, the obtained plasma was centrifuged again at 16,000×g for 10 minutes at the same temperature to completely remove residual fragments. The final supernatant was then collected and stored in an ultra-low temperature freezer at -80°C for long-term storage.
[0065] III. Circulating Cell-Free DNA Sequencing and Data Processing
[0066] cfDNA was extracted from a median plasma volume of 1 mL using the VAMNE MagUltra circulating cell-free DNA extraction kit (Novizan Biotechnology, Nanjing, China, catalog number #N913), strictly following the instruction manual. After extraction, precise quantification was performed using a Qubit 4.0 real-time fluorescence analyzer (Thermo Fisher Scientific, Lenexa, Kansas, USA). The obtained DNA samples were temporarily stored at -80°C for further in-depth analysis.
[0067] Library construction was performed using the VAHTS Universal DNA Library Preparation Kit (Vazyme Biotech, Nanjing, China, catalog number #N610) and the VAHTS Multiplex Barcoding Kit (catalog number #N322), with starting amounts of 5-20 ng cfDNA. Library quality was rigorously controlled using an Agilent 2100 Bioanalyzer (Agilent Technologies, Santa Clara, California, USA) to ensure DNA fragment integrity and optimal size distribution (main peak concentrated in the 160-220 bp range). All libraries were sequenced at 150 bp paired ends using the Illumina NovaSeq X Plus sequencing platform (San Diego, California, USA), achieving an average coverage depth of 2×.
[0068] Raw whole-genome sequencing (WGS) data underwent quality assessment using FastQC, followed by adapter filtering and quality trimming using Cutadapt and Ktrim. High-quality sequencing reads (clean reads) were aligned to the human reference genome (hg19) using the BWA-MEM algorithm, and then PCR duplications were removed and unique alignment sequences were selected using Samtools and Picard. To effectively isolate microbial-derived data, a three-step bioinformatics analysis workflow was constructed based on the alignment results to systematically eliminate interference from human sequences.
[0069] Samtools' SAM tags "-f 12" and "-F 256" were used to remove reads from human alignments. The "-f 12" tag only retains reads whose ends are not aligned to the human genome, while the "-F 256" tag ensures that only primary alignments are included. Next, the filtered reads were aligned to the NCBI Microbial Reference Genome Database using Kraken2, discarding alignments with a k-mer match of less than 10%, and a minimum of 50 reads aligned to the reference genome was required for a valid match. Finally, the taxonomic annotations generated by Kraken2 were reprocessed using Bracken, reassigning reads along the taxonomic tree and re-estimating species abundance. QIIME2 was used to calculate the total abundance of the samples, and samples with a total abundance below 100,000 were discarded.
[0070] This study ultimately included 870 participants who constructed and validated machine learning models (MLMs), specifically including:
[0071] 1) A comprehensive dataset was constructed by including 572 plasma circulating microbiome DNA (cmDNA) samples from Zhejiang Cancer Hospital, including untreated gastric cancer patients (n=265) and non-gastric cancer patients (including healthy individuals and subjects with benign lesions) (n=307). The dataset was randomly assigned to form a training dataset and an internal validation cohort: the training set included 141 gastric cancer patients and 145 non-gastric cancer patients, whose cmDNA profiles were used to construct a machine learning model (cmDNA-MLM); the model performance was then evaluated using an internal validation cohort containing 124 gastric cancer patients and 162 non-gastric cancer patients.
[0072] 2) An external validation cohort was constructed by collecting plasma circulating microbiome DNA (cmDNA) samples from three medical institutions: Zhejiang Cancer Hospital, Shanghai Yangpu Hospital, and Chongqing Yubei District People's Hospital. The external validation cohort included 131 patients with gastric cancer and 167 patients without gastric cancer.
[0073] IV. Identification of Differential Abundance Taxonomy
[0074] The alpha diversity index (including richness and evenness) was calculated using QIIME2, and statistical comparisons between groups were performed using the Wilcoxon rank-sum test. Beta diversity analysis based on Bray-Curtis differentials was used to assess inter-group differences, and the results were visualized using non-metric multidimensional scaling (NMDS). To identify clinically significant inter-group groups, their relative abundance was assessed using the Wilcoxon rank-sum test.
[0075] In addition, the linear discriminant analysis (LDA) effect size (LEfSe) method was used to identify the significantly enriched groups in specific clinical groups. The LDA score threshold was set at 2.0 and the p-value was <0.05 as the significance criterion.
[0076] The results showed that by comparing the relative abundance of bacteria in the gastric cancer (GC) patient group and the non-GC patient group in the training cohort, the biodiversity and composition of cmDNA were significantly altered in GC patients.
[0077] Gastric cancer-related specific microbial groups were screened using the linear discriminant analysis effect size (LEfSe) method: Based on the comparison of differences at the full taxonomic level, a branching diagram was constructed to show the taxonomic differences between benign and malignant groups, and linear discriminant analysis (LDA) was used to identify significantly different groups (LDA>2, p<0.05). These significantly different groups were incorporated as significant features into the development of the cmDNA-MLM model.
[0078] V. Constructing a Predictive Model
[0079] A total of 78 microbial taxa were identified by LEfSe (LDA score > 2.0 and p < 0.05). The relative abundance of the selected microbial taxa was used as a feature to train a machine learning model, developing a cmDNA-MLM model. The caret package (version v6.0.94; https: / / cran.r-project.org / web / packages / caret / ) was used for model construction and evaluation. Random Forest (RF), Generalized Linear Model (GLM), Multilayer Perceptron (MLP), and Distributed Gradient Boosting Library (XGBoost) were used to train the model. The model was optimized for hyperparameters and evaluated using 5-fold cross-validation. To ensure model robustness, the entire process was repeated 100 times. The Random Forest model performed well on the training set (AUC: 0.805, 95% CI: 0.756–0.854). Figure 1 The performance on the internal test set was further improved (AUC: 0.864, 95% CI: 0.824–0.905); Figure 1 In external independent validation, the random forest model still exhibited high diagnostic efficacy (AUC: 0.92, 95% CI: 0.891–0.949). In the trained random forest diagnostic model, key features such as Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae were identified as biomarkers. Based on these key features, the generalized linear model (GLM), multilayer perceptron (MLP), and XGBoost model showed AUCs of 0.898, 0.92, and 0.926 respectively in the external validation set, demonstrating high sensitivity and specificity and excellent diagnostic efficacy. Figure 1 (and Table 1);
[0080] Table 1
[0081]
[0082] Example 2: The cmDNA-MLM model demonstrated high diagnostic sensitivity in gastric cancer patients with normal serum tumor marker levels.
[0083] This study further collected serum tumor marker levels in gastric cancer (GC) patients from internal and external validation cohorts and analyzed the diagnostic sensitivity of the random forest-based cmDNA-MLM model in subgroups with high and normal levels of tumor markers.
[0084] The results showed that for gastric cancer patients with elevated serum alpha-fetoprotein (AFP>8.78 ng / mL), carcinoembryonic antigen (CEA>5 ng / mL), carbohydrate antigen 125 (CA125>35 U / mL), carbohydrate antigen 199 (CA199>37 U / mL), carbohydrate antigen 242 (CA242>20 U / mL), and carbohydrate antigen 724 (CA724>6.9 U / mL), the model sensitivities were 77.8%, 70.7%, 80.0%, 80.4%, 76.0%, and 75.5%, respectively. For patients with normal levels of the above biomarkers (AFP < 8.78 ng / mL, CEA < 5 ng / mL, CA125 < 35 U / mL, CA199 < 37 U / mL, CA242 < 20 U / mL, CA724 < 6.9 U / mL), the diagnostic sensitivity remained high, at 77.3%, 78.7%, 76.4%, 76.1%, 76.6%, and 76.8%, respectively. Figure 2 ).
[0085] Further evaluation of the model's performance in patients with gastric cancer at different stages showed that, regardless of tumor marker levels, cmDNA-MLM exhibited consistently high sensitivity in both early-stage (stage I) and late-stage (stages II-IV) gastric cancer subgroups. Figure 2 It is noteworthy that for early-stage gastric cancer patients with normal tumor markers (AFP < 8.78 ng / mL, CEA < 5 ng / mL, CA125 < 35 U / mL, CA199 < 37 U / mL, CA242 < 20 U / mL, CA724 < 6.9 U / mL), the model sensitivities still reached 69%, 71.9%, 69.5%, 73.1%, 72.4%, and 71.2%, respectively. Figure 2 ).
[0086] This study further evaluated the diagnostic performance of the cmDNA-MLM model in gastric cancer subgroups stratified based on clinicopathological features (including age, sex, alcohol / smoking status, differentiation degree, and molecular subtype).
[0087] The results are as follows Figure 3 As shown, cmDNA-MLM exhibited high sensitivity across all subgroups: 78.5% in elderly patients (≥60 years), 79.4% in younger patients (<60 years), 76.7% in male patients, 81.8% in female patients, 73.6% in the alcohol-consuming group, 79.9% in the non-alcohol-consuming group, 74.4% in the smoking group, and 80.1% in the non-smoking group. Figure 3Furthermore, the model's sensitivities in the moderately differentiated group, poorly differentiated group, microsatellite highly unstable (MSI-H) group, microsatellite stable (MSS) group, HER2-positive (HER2+) group, and HER2-negative (HER2-) group were 77.8%, 82.3%, 74.4%, 80.1%, 78.4%, and 79.7%, respectively. Figure 3 ).
[0088] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. Application of substances used to detect the species characteristics of circulating microbial DNA in blood or plasma in the preparation of gastric cancer diagnostic products; The microorganisms include Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae and Stappiaceae.
2. The application according to claim 1, characterized in that, The substance includes reagents and / or instruments for detecting the species characteristics of the circulating microbial DNA through sequencing; Preferably, the sequencing includes at least one of 16S sequencing, metagenomic high-throughput sequencing, whole-genome sequencing, and PCR-pyrosequencing; Preferably, the substance is selected from at least one of the following groups: probes, gene chips, and PCR primers; Preferably, the product comprises reagents, reagent kits, test strips, chips, or systems; Preferably, the species characteristics of the circulating microbial DNA are relative abundance at the microbial family level.
3. A product for the diagnosis of gastric cancer, the product comprising the substance for detecting the species characteristics of circulating microbial DNA as described in any one of claims 1 to 2.
4. A method for constructing a gastric cancer diagnostic model, comprising the step of constructing the model using the characteristics of circulating microbial DNA species as described in any one of claims 1 to 2.
5. The construction method according to claim 4, characterized in that, The model includes a machine learning model; the machine learning model includes at least one of random forest, generalized linear model, multilayer perceptron and distributed gradient boosting library.
6. The construction method according to claim 4, characterized in that, The method for obtaining the species characteristics of the circulating microbial DNA includes: cfDNA sequencing: Collect plasma or blood from the sample to be tested to extract cfDNA, and sequence the cfDNA to obtain the cfDNA sequence; Obtaining circulating microbial DNA: The cfDNA sequence is aligned to the host reference genome, and DNA fragments derived from blood or plasma microorganisms are isolated to obtain circulating microbial DNA; Species feature extraction: Based on the circulating microbial DNA, microbial species information is annotated by sequence alignment, and the relative abundance of microorganisms is evaluated to obtain the species features of the circulating microbial DNA; Preferably, the sequencing includes at least one of 16S sequencing, metagenomic high-throughput sequencing, whole-genome sequencing, and PCR-pyrosequencing; Preferably, the whole genome sequencing is low-depth whole genome sequencing.
7. A system for diagnosing gastric cancer, the system comprising: Data acquisition module: Acquires the species characteristics of circulating microbial DNA in the blood or plasma of the sample to be tested, including Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae. The result determination module uses a diagnostic model to diagnose gastric cancer in the test sample based on the acquired circulating microbial DNA species characteristic data.
8. The system according to claim 7, characterized in that, The diagnostic model is constructed by the method described in any one of claims 4 to 6.
9. A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Data Acquisition: Acquire the species characteristics of circulating microbial DNA in the blood or plasma of the sample to be tested. The microorganisms include Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae. Based on the obtained circulating microbial DNA species characteristic data, the diagnostic model was used to diagnose gastric cancer in the test sample. Preferably, the diagnostic model is constructed by the method described in any one of claims 4 to 6.
10. A computer program product comprising a computer program that, when executed by a processor, performs the following steps: Data acquisition: acquiring DNA species characteristic data of circulating microorganisms in blood or plasma of a sample to be tested, wherein the microorganisms include Bacillaceae, Burkholderiaceae, Comamonadaceae, Hyphomicrobiaceae, Lysobacteraceae, Methylophilaceae, Nocardiaceae, Oxalobacteraceae, Rhizobiaceae, and Stappiaceae; Based on the obtained circulating microbial DNA species characteristic data, a diagnostic model is used to diagnose gastric cancer in the sample to be tested; preferably, the diagnostic model is constructed by the method described in any one of claims 4 to 6.