Methylation marker combination for detecting early gastric cancer and application thereof

By quantitatively detecting the methylation of multiple specific biomarker genes in plasma samples from patients with early gastric cancer, and combining 146 methylation sites from 11 DNA regions with a logistic regression model, the invasiveness and insufficient sensitivity of existing gastric cancer screening technologies have been addressed, achieving highly sensitive and specific non-invasive early gastric cancer detection.

CN122382202BActive Publication Date: 2026-08-25JIAXING YUNYING MEDICAL INSPECTION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610838624.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-08-25
Estimated Expiration
2046-06-11

AI Technical Summary

Technical Problem

Existing gastric cancer screening methods, such as gastroscopy, are invasive and costly. Serological tests lack sufficient sensitivity and specificity in the early detection of gastric cancer. Although next-generation sequencing (NGS) provides a panoramic view, it is limited to the qualitative detection of individual genes and lacks high sensitivity and specificity.

Method used

The NGS method was used to quantitatively detect the methylation of multiple specific biomarker genes in plasma samples from patients with early gastric cancer. By combining 146 methylation sites from 11 DNA regions with a logistic regression model, a non-invasive detection with high sensitivity and high specificity was achieved.

Benefits of technology

It achieves high sensitivity (93.26%-94.17%) and high specificity (96.15%-97.89%) in early gastric cancer detection, is suitable for large-scale screening of high-risk populations, reduces invasive examinations, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122382202B_ABST
    Figure CN122382202B_ABST
Patent Text Reader

Abstract

The application discloses a methylation marker combination for detecting early gastric cancer and application thereof. The marker combination comprises 146 methylation sites of 11 DNA regions and a transformation rate control marker region, and the methylation levels of the target regions are significantly different between early gastric cancer patients and healthy people. The kit of the application contains reagents for detecting the methylation degree of the above sites, and supports various high-throughput sequencing methods. The detection system adopts a logistic regression model to calculate a prediction value, multiplies the average methylation rate of each region by a corresponding weight coefficient, adds them together, and compares the sum with a threshold value to make a positive judgment. The application realizes high sensitivity and high specificity for non-invasive early detection, is particularly suitable for non-invasive screening of plasma free DNA samples, can effectively improve the survival rate of gastric cancer patients, can significantly reduce medical expenses, and has a wide clinical application and industrial utilization prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, and more specifically, to a combination of methylation biomarkers for detecting early gastric cancer and their applications. Background Technology

[0002] Early screening and diagnosis of gastric cancer are crucial for improving patient prognosis. Current gastric cancer screening methods mainly include gastroscopy, imaging examinations, and serological testing. Among these, gastroscopy is invasive and faces limitations in terms of compliance and implementation costs in large-scale screening scenarios. The sensitivity and specificity of non-invasive methods such as serological testing in the early detection of gastric cancer still need to be improved.

[0003] Tumor methylation often occurs in the early stages of tumor development, and cfDNA methylation variations are tissue-specific. Therefore, blood cfDNA methylation testing has become an important diagnostic tool for patients suspected of having tumors in the early stages. Currently, common first-generation methylation detection methods using Q-PCR only provide a close-up, localized view of methylation, limited to individual genes, and are only qualitatively detectable; furthermore, their sensitivity and specificity need improvement. In contrast, next-generation sequencing (NGS) provides a panoramic, high-resolution view of methylation, allowing for quantitative detection of methylation rates at specific methylation sites in specific genes. Moreover, it can simultaneously quantify multiple methylation sites across a large number of genes, providing more accurate results and becoming an important early tumor methylation detection solution, both now and in the future. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention employs NGS (Next Generation Sequencing) on ​​plasma samples from early-stage gastric cancer patients to quantitatively detect the methylation of multiple specific biomarker genes in gastric cancer, achieving highly sensitive, highly specific, rapid, and non-invasive early gastric cancer signal detection.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A methylation biomarker set, using GRch37 as a reference genome, comprising at least one of the following DNA regions:

[0007] Marker 1, chr4:42153736-42153886,

[0008] Marker 2, chr6:26240834-26240984,

[0009] Marker 3, chr14:102248080-102248230,

[0010] Marker 4, chr12:133485299-133485449,

[0011] Marker 5, chr2:73147711-73147861,

[0012] Marker 6, chr18:60985624-60985774,

[0013] Marker 8, chr20:37434666-37434816

[0014] Marker 9, chr2:182322480-182322630,

[0015] Marker 10, chr3:179169180-179169330,

[0016] Marker 12, chr16:23847489-23847639,

[0017] Marker 13, chr8:41504152-41504302,

[0018] Conversion rate control marker region chr7:5571788-5571938.

[0019] Furthermore, marker 1 includes methylation sites: chr4-42153755, chr4-42153761, chr4-42153764, chr4-42153768, chr4-42153774, chr4-42153777, chr4-42153785, chr4-42153788, chr4-42153795, chr4-4215379 7. chr4-42153801, chr4-42153804, chr4-42153810, chr4-42153813, chr4-42153842, chr4-42153846, chr4-42153855, chr4-42153858, chr4-42153861, chr4-42153864 and chr4-42153867;

[0020] Marker 2 includes methylation sites: chr6-26240855, chr6-26240860, chr6-26240863, chr6-26240874, chr6-26240881, chr6-26240888, chr6-26240920, chr6-26240922, chr6-26240930, chr6-26240939, and chr6-26240950;

[0021] Marker 3 includes methylation sites: chr14-102248201, chr14-102248195, chr14-102248189, chr14-102248169, chr14-102248164, chr14-102248159, chr14-102248148, chr14-102248140, chr14-102248138, chr14-102248133, chr14-102248131, chr14-102248127, chr14-102248106, chr14-102248104, and chr14-102248099;

[0022] Marker 4 includes methylation sites: chr12-133485324, chr12-133485336, chr12-133485339, chr12-133485341, chr12-133485349, chr12-133485354, chr12-133485382, chr12-133485385, and chr12-133485416;

[0023] Marker 5 includes methylation sites: chr2-73147841, chr2-73147838, chr2-73147830, chr2-73147828, chr2-73147826, chr2-73147823, chr2-73147817, chr2-73147815, chr2-73147812, chr2-73147804, chr2-73147799, chr2-73147790, chr2-73147772, chr2-73147768, chr2-73147756, and chr2-73147739;

[0024] Marker 6 includes methylation sites: chr18-60985741, chr18-60985733, chr18-60985720, chr18-60985713, chr18-60985706, chr18-60985702, chr18-60985691, chr18-60985688, chr18-60985676, chr18-60985666, chr18-60985663, chr18-60985660, chr18-60985657, chr18-60985655, and chr18-60985646;

[0025] Marker 8 includes methylation sites: chr20-37434793, chr20-37434790, chr20-37434766, chr20-37434749, chr20-37434738, chr20-37434731, chr20-37434728, chr20-37434724, chr20-37434702, and chr20-37434684;

[0026] Marker 9 includes methylation sites: chr2-182322501, chr2-182322503, chr2-182322530, chr2-182322537, chr2-182322545, chr2-182322549, chr2-182322564, chr2-182322569, chr2-182322574, and chr2-182322600;

[0027] Marker 10 includes methylation sites: chr3-179169204, chr3-179169206, chr3-179169212, chr3-179169215, chr3-179169224, chr3-179169237, chr3-179169249, chr3-179169252, chr3-179169268, chr3-179169290, chr3-179169292, chr3-179169295, chr3-179169297, chr3-179169300, chr3-179169306, and chr3-179169311;

[0028] Marker 12 includes methylation sites: chr16-23847507, chr16-23847513, chr16-23847519, chr16-23847522, chr16-23847525, chr16-23847529, chr16-23847535, chr16-23847547, chr16-23847551, chr16-23847556, chr16-23847560, chr16-23847568, chr16-23847575, and chr16-23847586;

[0029] Marker 13 includes methylation sites: chr8-41504171, chr8-41504182, chr8-41504197, chr8-41504217, chr8-41504252, chr8-41504254, chr8-41504258, chr8-41504263, and chr8-41504265;

[0030] Conversion rate control markers include methylation sites: chr7-5571912, chr7-5571911, chr7-5571910, chr7-5571909, chr7-5571902, chr7-5571891, chr7-5571887, chr7-5571874, chr7-5571873, chr7-5571872, chr7- 5571868, chr7-5571846, chr7-5571844, chr7-5571841, chr7-5571836, chr7-5571833, chr7-5571828, chr7-5571821, chr7-5571820, chr7-5571814, chr7-5571811 and chr7-5571810.

[0031] Furthermore, the methylation levels of DNA regions differ between early-stage gastric cancer patients and healthy individuals.

[0032] Furthermore, the biomarker combination has a corresponding primer combination, including the following primers:

[0033] The upstream primer for chr4:42153736-42153886 is GCGTTTGGGTTATATTTGA, and the downstream primer is AAACCGCCGCCGCCGCTT.

[0034] The upstream primer for chr6:26240834-26240984 is: GGTTTTTGGAGAATGTGAT, and the downstream primer is: TAAAAAACCGAAATAAAACTCAAC.

[0035] The upstream primer for chr14:102248080-102248230 is TTTGGGTGTCGTTTCGGTAT, and the downstream primer is CGCCGATCTAACCTCCTA.

[0036] The upstream primer for chr12:133485299-133485449 is GTATAGGAGGGGTGGAAT, and the downstream primer is ACCCGAAACGAAAATATAACG.

[0037] The upstream primer for chr2:73147711-73147861 is AGCGTTTCGGTTAGCGGA, and the downstream primer is GAATCTCAACRATACTCAT.

[0038] chr18:60985624-60985774 Upstream primer: CGTTTTCGTATCGGGTAT, Downstream primer: CACAAATAACACCRAACT;

[0039] The upstream primer for chr20:37434666-37434816 is TTCGGTATTGTATTTTCGGT, and the downstream primer is ACCTCRCAATTCCTT.

[0040] The upstream primer for chr2:182322480-182322630 is TTATAAYGTGGATATTGAGAG, and the downstream primer is ATTCTCACGCCGCTCTAA.

[0041] The upstream primer for chr3:179169180-179169330 is TTGGGYGTTTAGTAGTTTTTG, and the downstream primer is CGCCTCCAACCACGTCCC.

[0042] The upstream primer for chr16:23847489-23847639 is: CGCGTAAGATGGTTGATT, and the downstream primer is: ATAAACTACTTAAAAAAACGAACGA.

[0043] The upstream primer for chr8:41504152-41504302 is GAGGAYGAGGGTTTTAGG, and the downstream primer is CCATTTATTTAACTATAAAAACAAA.

[0044] The upstream primer for the conversion rate control marker chr7:5571788-5571938 is TAAGGTTAGGGATAGGATAGTTTTA, and the downstream primer is TAACCACCACCCAACACACAAT.

[0045] Furthermore, a kit for detecting early gastric cancer includes a methylation detection reagent for the degree of methylation of at least one marker region in the above-mentioned combination of markers.

[0046] Furthermore, the methylation detection reagent includes any one or more of the following methods, which include: pyrosequencing, bisulfite conversion sequencing, methylation array, qPCR, digital PCR, next-generation sequencing, third-generation sequencing, whole-genome methylation sequencing, DNA enrichment detection, simplified bisulfite sequencing, HPLC, MassArray, methylation-specific PCR, or combinations thereof.

[0047] An early gastric cancer detection system includes a calculation module that calculates a detection model prediction value based on the methylation level of the marker regions in the above-mentioned marker combination. The detection model prediction value is calculated using a logistic regression model, where the average methylation rate of each marker region is multiplied by the corresponding weight coefficient and then summed. If the result is greater than a threshold, the system is judged to be positive for gastric cancer.

[0048] By adopting the above technical solution, the beneficial effects of the present invention are as follows:

[0049] 1. High sensitivity and high specificity: Based on the combination of 146 methylation sites in 11 DNA regions and a logistic regression model, it achieves a sensitivity of 93.26%-94.17%, a specificity of 96.15%-97.89%, and an AUC value of 0.9621-0.9773 in the training and validation sets, effectively distinguishing early gastric cancer patients from healthy individuals.

[0050] 2. Non-invasive and convenient testing: NGS methylation quantitative detection is performed using plasma samples, avoiding invasive examinations. It is suitable for large-scale screening of high-risk populations, with a short testing cycle and low cost.

[0051] 3. High stability and repeatability: Through multiple rounds of biomarker screening, panel optimization, and amplification procedure adjustment, reliable test results and small batch-to-batch variation are ensured.

[0052] 4. High clinical application value: Early identification of gastric cancer risk, improvement of 5-year survival rate, reduction of unnecessary invasive examinations, and promotion of personalized medicine. Attached Figure Description

[0053] Figure 1The ROC curves are the training set results of models 1-8 of this invention.

[0054] Figure 2 This is the ROC curve of the validation set results for Model 5 of this invention. Detailed Implementation

[0055] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] Table 1. Blood Methylation Primer Table

[0057] Table 1 lists the primer information corresponding to the methylation biomarkers used in this invention for early gastric cancer detection. Specifically, it includes the amplification sequence coordinates of each biomarker in the reference genome GRCh37, the upstream and downstream primer sequences, and the corresponding methylation sites covered. All primers are designed for differentially methylated regions that have been screened and validated, and are used to specifically amplify target regions of cell-free plasma DNA after sulfite conversion. Using the primer combinations shown in Table 1, stable and simultaneous detection of multiple methylation sites can be achieved, providing a reliable experimental basis for subsequent library construction, high-throughput sequencing, and methylation rate calculation, thereby ensuring the accuracy, repeatability, and feasibility of the detection method of this invention.

[0058] Experimental steps

[0059] Blood sample collection

[0060] Whole blood collection

[0061] 1. Blood collection method: 192 positive plasma samples for gastric cancer were obtained from the plasma of patients with early gastric cancer (stage I and II) (mean age 58 years, females accounting for 46.35%) who were stored at -80 degrees Celsius by the applicant; negative plasma samples were collected and stored in the morning on an empty stomach by collecting 10 ml of blood from 173 healthy individuals (mean age 60 years, females accounting for 43.35%) using disposable vacuum blood collection tubes produced by the applicant.

[0062] 2. Preparation of plasma samples: Centrifuge the blood collection tube containing whole blood for 12 minutes at a centrifugal force of 1500 rcf. Remove the blood collection tube from the centrifuge and transfer the supernatant plasma into a labeled 2 mL centrifuge tube to obtain the sample required for the experiment.

[0063] Pretreatment of cell-free DNA in plasma

[0064] The methylation detection sample pretreatment reagent and plasma sample lysis and sulfite conversion kit produced by the applicant were used to lyse, convert and purify the samples. The operation steps are as follows.

[0065] 1. Preparation of reagents in the kit

[0066] Prepare the reagents in the above kit:

[0067] Table 2. Reagent Preparation

[0068] 2. cfDNA extraction

[0069] Take 1.8 ml of plasma, add 100 μl of 20% SDS, mix by inverting, centrifuge briefly, then add 30 μl of PK enzyme (20 mg / ml), vortex 4-5 times to mix, incubate at 60℃ for 30 min, place at room temperature for 5 min, and centrifuge briefly.

[0070] The above mixture was further extracted using an Auto-Pure-20B purifier. The program was run according to the settings. After the program was completed, the liquid was taken out from the last well and placed into a new PCR tube. The tube was labeled with the corresponding sample number, which is cfDNA. It can then be used for the next step of sulfite conversion.

[0071] (1) Sample addition: Add all the solution in the EP tube to column 1 of the purification kit (add different samples to different wells), and be sure to record the sample number;

[0072] (2) After installing the magnetic rod sleeve, place the purification kit into the fully automated nucleic acid extractor and set it up and operate it according to the prescribed procedure:

[0073] (3) After the program finishes running, the solution is transferred to a new 0.5 mL centrifuge tube. The collected solution is the cfDNA with magnetic beads and sulfite.

[0074] 3. Sulfite Conversion

[0075] 3.1 System Configuration

[0076] Prepare sulfite reaction systems in 0.5 mL PCR tubes for the lysed samples as shown in the table below.

[0077] Table 3. Sulfite Reaction System

[0078] 3.2 Sulfite Conversion

[0079] After vortexing and mixing, and briefly centrifuging, place the PCR tube in a PCR instrument and incubate at 95℃ for 5 min and 85℃ for 60 min. After incubation, cool to room temperature and briefly centrifuge to obtain the transformation product.

[0080] 3.3 Bis-DNA purification treatment

[0081] The purified product was run using the Auto-Pure-32A purifier according to the set program. After the run was completed, the liquid was removed from the well and placed into a new PCR 8-strip, labeled with the corresponding sample number. This is the purified Bis-DNA, which can be used for the next library enrichment reaction.

[0082] (1) Sample addition: Add all the solution from the PCR tube to column 1 of the purification kit (add different samples to different wells), and be sure to record the sample number;

[0083] (2) After installing the magnetic rod sleeve, place the purification kit into the fully automated nucleic acid extractor and set it up and operate it according to the prescribed procedure:

[0084] (3) After the program finishes running, aspirate the solution into a new 1.5 mL centrifuge tube. The collected solution is the DNA after sulfite conversion and should be used for library construction immediately.

[0085] Library enrichment reaction

[0086] The Bis-DNA obtained after purification in the previous step was subjected to a library enrichment reaction.

[0087] 1. Preparation of enrichment PCR reaction: Prepare the enrichment reaction solution as needed according to the formula shown in the table below. After preparation, aliquot the solution into clean 8-strip containers. The aliquoted enrichment reaction solutions are grouped according to the primer mixture (primers and physical locations of gastric cancer markers), and are respectively named MG enrichment PCR reaction strip-1 to MG enrichment PCR reaction strip-4.

[0088] Table 4. Enrichment Reaction Table

[0089] 2. Gently open the caps of the MG enrichment PCR reaction strip-1 to MG enrichment PCR reaction strip-4 prepared in the previous step, and add 1 μL of Taq enzyme to each of the MG enrichment reaction solutions-1 to MG enrichment PCR reaction solutions-4, ensuring that the pipette tip is fully inserted into the reaction solution.

[0090] 3. Add 11 μL of the Bis DNA sample after sulfite conversion to the four enrichment reaction strips in sequence along the wall of the PCR tube, and carefully cap the tubes.

[0091] 4. After vortexing the PCR reaction strip to mix well, centrifuge, taking care to avoid air bubbles;

[0092] 5. Place the above four enrichment PCR reaction strips into the PCR instrument;

[0093] 6. Open the PCR instrument settings interface, set the amplification program according to the table below, and perform PCR amplification.

[0094] Table 5. PCR reaction procedure for library enrichment

[0095] Library preparation reaction

[0096] 1. Preparation for the joint connection reaction:

[0097] Prepare the required number of adapters according to the formula shown in the table below (except for Ace Taq enzyme and the PCR reaction product enriched in the previous step). Alternatively, you can prepare them in advance and aliquot them into clean 8-strips and store them at -20°C for later use.

[0098] After removing the connector reaction solution from the refrigerator and allowing it to melt, centrifuge it briefly in a centrifuge before use.

[0099] Table 6. Reaction Components for Library Preparation

[0100] 2. Gently open the cap of the reaction strip connector and add 0.5 μL of Taq enzyme into the reaction strip connector, ensuring that the pipette tip is fully inserted into the reaction solution;

[0101] 3. Take 5 μL of the enriched PCR amplification product from the previous step and add it sequentially along the wall of the PCR tube to the adapter connection reaction strip, then carefully close the tube cap.

[0102] 4. Connect the connector to the reaction strip, shake to mix thoroughly, and then centrifuge, taking care to avoid air bubbles;

[0103] 5. Connect the above-mentioned adapter to the reaction strip and place it into the PCR instrument;

[0104] 6. Open the PCR instrument settings interface, set the amplification program according to the table below, and perform PCR amplification.

[0105] Table 7. PCR reaction procedure for library preparation

[0106] Library Product Sorting

[0107] 1. Remove the magnetic beads used for purification, vortex to mix, and incubate at room temperature for at least 30 minutes;

[0108] 2. Add 40 μL of purified water to each sample reaction tube in sequence;

[0109] 3. Take 126 μL of the incubated magnetic beads (shake well again before use) and add them sequentially to each sample reaction tube. Shake well and then incubate briefly to completely mix and resuspend the DNA and magnetic beads.

[0110] 4. Incubate at room temperature for 5 minutes;

[0111] 5. Incubate on a magnetic rack for 5 minutes until the solution is clear. Carefully aspirate and discard the supernatant, being careful not to disturb the magnetic beads.

[0112] 6. Add 200µL of freshly prepared 80% ethanol solution (the ethanol solution should just cover the magnetic bead sample), place it on a magnetic rack, and incubate it on the magnetic rack for 30 seconds until the solution is clear. Discard the supernatant.

[0113] 7. Repeat step 6 above for a second wash;

[0114] 8. Ensure that the ethanol solution in the centrifuge tube has been completely discarded, and place the eight-tube or PCR tube on a magnetic rack to air dry at room temperature for 5-10 minutes;

[0115] 9. Remove the eight-tube or PCR tube from the magnetic rack, add 100 μL of purified water to fully wet the magnetic beads, shake well to mix, and then quickly centrifuge to collect the liquid at the bottom of the tube. Let it stand at room temperature for 5 min.

[0116] 10. Place the 8-tube or PCR tube on a magnetic rack and let it stand for 5 minutes until the solution is clear. Transfer 70 μL of the supernatant to a new 8-tube or 1.5 mL centrifuge tube to obtain the sequencing library. Use the Equalbit 1×dsDNA HS Assay Kit and its matching Qubit fluorescence instrument to detect the concentration. The library concentration should be ≥1 ng / μL. Otherwise, the library preparation sample does not meet the requirements and should be reconstructed.

[0117] 11. Perform deep sequencing on an Illumina MiniSeq sequencer, with each sample being sequenced at 200M.

[0118] Methylation data processing workflow

[0119] The first step is data processing and quality control.

[0120] The methylation rate of 171 methylation sites for 13 gene markers in plasma samples from gastric cancer-positive and healthy individuals was statistically analyzed. The methylation rate of a site was calculated as: (Number of methylated reads at that site / (Number of methylated reads at that site + Number of unmethylated reads at that site)). Data from primer-located methylation sites were removed, and the methylation rate of methylation sites with excessively low sequencing depth (site sequencing depth > 10000X) was excluded. The average methylation rate of all methylation sites within each gene marker was then calculated to represent the methylation rate of that gene marker. Marker 1 contained... There are 21 methylation sites. Marker 2 contains 11 methylation sites, marker 3 contains 15 methylation sites, marker 4 contains 9 methylation sites, marker 5 contains 16 methylation sites, marker 6 contains 15 methylation sites, marker 7 contains 14 methylation sites, marker 8 contains 10 methylation sites, marker 9 contains 10 methylation sites, marker 10 contains 16 methylation sites, marker 11 contains 11 methylation sites, marker 12 contains 14 methylation sites, and marker 13 contains 9 methylation sites.

[0121] 1. The methylation conversion rate of all 22 methylation sites on the internal reference gene marker ACTB in all samples was greater than 99.0%, indicating that the methylation conversion in this experiment was successful.

[0122] 2. The criteria for removing abnormal methylation rates in normal individuals are as follows: First, check the average methylation rate. If the methylation rate is higher than 5%, then check the individual sites of the marker. If there is a site with a significantly elevated methylation rate (greater than 20%) or a jump point (10%) greater than half of all sites, remove that sample. If all are within 10%, do not remove. If there is a jump point greater than 20% but the average methylation rate is less than 5%, do not remove. The training set of this application originally included 106 normal individuals with negative samples. After removing abnormal methylation rates, blood methylation data from 95 normal individuals were ultimately retained. The methylation rates of the normal individuals in the selected validation set were all acceptable.

[0123] The second step is model training. The methylation rates of 13 gene markers selected from the training set (103 positive gastric cancer cases + 95 normal individuals) are substituted into the logistic regression model for training. The coefficients and thresholds corresponding to each gene marker are obtained. The methylation level of each gene marker is multiplied by the corresponding coefficient and then added together. If the sum is greater than the threshold, the sample is judged as positive; otherwise, it is negative.

[0124] The third step is model tuning and evaluation. The weights of the dependent variable in the model are adjusted according to the ratio of positive and negative data to balance the training set data. The regularization parameters are tuned to find the optimal parameters. The threshold is adjusted to adapt to clinical diagnosis. The sensitivity and specificity of different models are calculated. ROC curves are plotted to evaluate the excellence of different models and select the best model.

[0125] The fourth step is to use the above optimal model to perform model validation analysis on the validation set (89 cases of positive gastric cancer + 78 normal individuals), calculate its sensitivity and specificity, and plot the ROC curve.

[0126] This invention collects plasma samples from patients with early-stage gastric cancer and healthy controls, performs DNA extraction, sulfite conversion, targeted amplification, and NGS sequencing, calculates the average methylation rate of each biomarker region, and constructs a logistic regression model based on these data to detect early-stage gastric cancer. The following examples detail the biomarker screening process, data collection, model construction, reagent kit preparation, and the application of the detection system.

[0127] Example 1: Biomarker Screening and Validation

[0128] The process involved multiple rounds of screening and validation, from initial candidate biomarkers to the final 13 biomarkers. The screening strategy employed a progressive validation approach, divided into two stages: the first stage was initial biomarker screening, using leukocyte and tissue samples to identify basic biomarkers; the second stage was plasma validation, verifying the contribution rate and specificity of biomarkers in plasma samples. Each stage was based on the screening results of the previous stage, and through systematic experimental design, 13 stable and reliable biomarkers were ultimately determined.

[0129] Phase 1: Initial screening of biomarkers

[0130] Experiment 1: Baseline detection of leukocyte methylation

[0131] Objective: To establish a baseline for methylation in normal human leukocytes and to screen for candidate biomarkers with methylation rates below 10%.

[0132] Methods: Methylation levels of multiple candidate biomarkers were detected using leukocyte samples from 20 healthy individuals. Methylation rate and sequencing depth of each biomarker were analyzed by NGS sequencing, and biomarkers with an average methylation rate <10% and an average sequencing depth >100X were screened.

[0133] Results: From the 20 biomarkers, 18 biomarkers with methylation rates <10% and depths >100X were retained, while two biomarkers with methylation rates >10% in multiple samples were deleted (chr16:89008384-89008534, chr2:100720706-100720856). See the table below.

[0134] Table 8. White blood cell study data

[0135] Experiment 2: Verification of the contribution rate of gastric cancer tissue

[0136] Objective: To verify the contribution rate of biomarkers retained after leukocyte screening to the detection of gastric cancer tissue samples.

[0137] Methods: Biomarkers retained from Experiment 1 were detected using 20 gastric cancer tissue samples. The contribution rate of each biomarker in the tissue samples was analyzed by NGS sequencing. Contribution rate = number of positive samples / total number of samples. Biomarkers with an average methylation rate greater than 15% were considered positive, and biomarkers with a contribution rate <80% were excluded.

[0138] Results: Sixteen biomarkers with a contribution rate ≥ 80% were retained, and two biomarkers (chr7:69063449-69063599, chr1:171810678-171810828) were removed for further validation. The results are shown in the table below.

[0139] Table 9. Organizational Research Data

[0140] Phase Two: Plasma Verification

[0141] Experiment 3: Assessment of Plasma Contribution Rate

[0142] Objective: To evaluate the detection efficacy of biomarkers that perform well in tissue samples in plasma samples.

[0143] Methods: Plasma samples from 20 gastric cancer patients were used to detect biomarkers with a tissue contribution rate >80%. A 10% contribution rate was defined as a positive marker, and markers with a depth >100X were included in the statistical analysis. The contribution rate of biomarkers in plasma was assessed, and biomarkers with a contribution rate <50% were excluded. The contribution rate was calculated as follows: Contribution rate = Number of positive samples / Total number of samples.

[0144] Results: Fourteen biomarkers with a positive plasma contribution rate ≥50% were screened, and their effective amplification was analyzed to verify the reliability of the data. Two biomarkers with a positive plasma contribution rate <50% (chr8:72756028-72756178, chr1:107683995-107684145) were removed. The results are shown in the table below.

[0145] Table 10. Plasma Contribution Rate Data

[0146] Experiment 4: Specificity Verification of Negative Plasma

[0147] Objective: To evaluate the specificity of biomarkers in normal human plasma samples and to exclude false positives.

[0148] Methods: Using 20 negative plasma samples from healthy individuals, the specificity of each biomarker was analyzed, and biomarkers with specificity <90% were removed.

[0149] Results: One biomarker with a specificity <90% (chr5:112073377-112073527) was removed, leaving 13 biomarkers with a specificity ≥90%. The results are shown in the table below.

[0150] Table 11. Plasma-specific data

[0151] Through multiple rounds of screening, 13 biomarker combinations, totaling 171 methylation sites, were ultimately identified for subsequent model construction. The contribution rate and specificity statistics are shown in the table below.

[0152] Table 12. Statistical results of contribution rate and specificity

[0153] Example 2: Sample collection, data processing, and detection model construction

[0154] 2.1 Sample Collection

[0155] Training set: 103 positive plasma samples from early gastric cancer patients (pathological diagnoses included gastric cancer, gastric adenocarcinoma, etc.) and 95 negative plasma samples from healthy individuals were collected. Among the positive samples, gastric adenocarcinoma accounted for the highest proportion; the negative samples were from healthy individuals.

[0156] Validation set: 89 positive samples from gastric cancer and 78 negative samples from healthy individuals.

[0157] 2.2 DNA Extraction and Sulfite Conversion

[0158] cfDNA was extracted from plasma and transformed using the Zymo EZ DNA Methylation Gold Kit.

[0159] Detection site conversion rate = (1- ()×100%, requirement >99.5%.

[0160] Internal control ACTB sites: 22 sites including 5571810. In the training set, the mean non-transformation rate for positive samples was 0.005 (standard deviation 0.002), corresponding to a transformation rate of approximately 99.5%; the mean non-transformation rate for negative samples was 0.004 (standard deviation 0.002), corresponding to a transformation rate of approximately 99.6%. The transformation rate for all samples was >99.5%.

[0161] 2.3 Targeted amplification and NGS sequencing

[0162] Targeted amplification was performed using the primer combinations described above. After purification, the amplified products were deep sequenced on an Illumina MiniSeq sequencer (150 Mreads per sample). Sequencing data processing included removing primer sites and filtering sites with a sequencing depth <10000X. The methylation rate for each site was calculated as: number of methylated reads / (number of methylated reads + number of unmethylated reads).

[0163] The average methylation rate of each marker region is calculated based on its site:

[0164] Marker 1 contains 21 methylation sites, marker 2 contains 11 methylation sites, marker 3 contains 15 methylation sites, marker 4 contains 9 methylation sites, marker 5 contains 16 methylation sites, marker 6 contains 15 methylation sites, marker 7 contains 14 methylation sites, marker 8 contains 10 methylation sites, marker 9 contains 10 methylation sites, marker 10 contains 16 methylation sites, marker 11 contains 11 methylation sites, marker 12 contains 14 methylation sites, and marker 13 contains 9 methylation sites.

[0165] 2.4 Quality Control and Anomaly Removal

[0166] Negative sample removal rules: In normal samples, if the average methylation rate is higher than 5% and there are more than half of the sites with methylation rates greater than 20% or greater than 10% in that region, the sample is removed. For example, if the methylation rates of the nine methylation sites of marker 4 in a negative sample are 12.23%, 10.98%, 11.47%, 13.04%, 12.35%, 11.45%, 10.89%, 11.68%, and 12.18%, respectively, then the average rate is >10%, indicating a jump, and the sample is removed. The training set ultimately retained 95 negative samples, and the validation set retained 78 negative samples.

[0167] 2.5 Model Training and Tuning

[0168] Eight models were trained using the LogisticRegression model (solver='liblinear') from the Python-sklearn library, based on the average methylation rate of 13 markers in the training set. In the `class_weight` parameter, 0 represents negative samples and 1 represents positive samples, with the weight of positive samples being a multiple of that of negative samples.

[0169] Table 13. Model Parameter Table

[0170] Model predicted value = Σ(average methylation rate_i × weight coefficient_i). If the predicted value > threshold, it is considered positive.

[0171] Table 14. Training Set Table

[0172] like Figure 1 As shown in Table 14, statistical analysis and model establishment of 171 methylation sites of 13 gene markers from 198 cases (103 positive gastric cancer cases + 95 normal individuals) in the training set yielded 8 early gastric cancer diagnostic models. Among them, Model 5 performed best, confirming 11 markers with a threshold of 0.98, a sensitivity of 94.17%, a specificity of 97.89%, and an area under the ROC curve (AUC) of 0.9773.

[0173] Table 15. Validation Set Table

[0174] like Figure 2 As shown in Table 15, the present invention again uses the validation set (89 positive cases of gastric cancer + 78 normal individuals) to test the effectiveness of model 5. The results show that model 5 detected a total of 83 positive cases with a sensitivity of 93.26%, detected 75 negative cases with a specificity of 96.15%, and the area under the ROC curve (AUC) was 0.9621.

[0175] The training methods and processes for the above models can be found in the following literature: [1]VERHULST P F. Notification of population suitability[J]. Correspondance Mathématique et Physique,1838,10:113-120. [2]BERKSON J. Application of the Logistic Function to Bio-Assay[J].Journal of the American Statistical Association, 1944,39(227):357-365. [3]NELDER JA, WEDDERBURN RW M. Generalized Linear Models[J].Journal of the Royal Statistical Society: Series A, 1972,135(3):370-384. [4]PREGIBON D. Resistant fits for some commonly used logistic models with medical applications[J]. Biometrics, 1982,38(2):485-498. [5]GREEN P J. Iteratively reweighted least squares for maximumlikelihood estimation, and some robust and resistant alternatives[J]. Journal of the Royal Statistical Society: Series B (Methodological), 1984,46(2):149-192. The above literature is used to illustrate the theoretical sources of model training and parameter solving methods, and does not limit the specific implementation of marker screening, sample processing, sequencing procedures and model parameters in this invention.

[0176] Example 3: Preparation and Application of the Reagent Kit

[0177] Kit components: primer combination (concentration optimized, e.g., upstream primer 10-20 nM), transformation reagent, PCR reagent (Taq enzyme, dNTPs), magnetic bead purification reagent, NGS library reagent.

[0178] Application procedure: Extract DNA from 10mL of plasma, transform, amplify (1st denaturation, 60s optimization), sequence, and input the calculated rate into the model. Clinical screening: High-risk individuals; positive results are further confirmed endoscopically.

[0179] Example 4: Application of the detection system

[0180] System software: Input methylation data, calculate using Model 5 (coefficients such as marker 1:9.79, etc.). Output positive / negative results and scores. Integrated with a cloud platform for batch processing.

[0181] Applications: In screening high-risk populations, positive samples are further confirmed using endoscopy, reducing unnecessary invasive procedures. The system demonstrates high accuracy in the validation set and supports personalized risk assessment.

[0182] The above embodiments, through detailed screening, data processing, and model validation, confirm the efficacy of the present invention in non-invasive detection of early gastric cancer.

[0183] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any ordinary changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included within the protection scope of the present invention.

Claims

1. A combination of methylation markers, characterized in that, Using GRch37 as a reference genome, the methylation marker combination includes the following 11 DNA regions: Marker 1, chr4:42153736-42153886, Marker 2, chr6:26240834-26240984, Marker 3, chr14:102248080-102248230, Marker 4, chr12:133485299-133485449, Marker 5, chr2:73147711-73147861, Marker 6, chr18:60985624-60985774, Marker 8, chr20:37434666-37434816 Marker 9, chr2:182322480-182322630, Marker 10, chr3:179169180-179169330, Marker 12, chr16:23847489-23847639, Marker 13, chr8:41504152-41504302.

2. The methylation marker combination according to claim 1, characterized in that, Using GRch37 as a reference genome, the methylation biomarker combination also includes the transformation control biomarker region chr7:5571788-5571938.

3. The methylation marker combination according to claim 2, characterized in that, The marker 1 includes methylation sites: chr4-42153755, chr4-42153761, chr4-42153764, chr4-42153768, chr4-42153774, chr4-42153777, chr4-42153785. chr4-42153788, chr4-42153795, chr4-42153797, chr4-42153801, chr4-42153804, chr4-42153810, chr4-42153813, chr4-42153842, chr4-42153846, chr4-42153855, chr4-42153858, chr4-42153861, chr4-42153864 and chr4-42153867; The marker 2 includes methylation sites: chr6-26240855, chr6-26240860, chr6-26240863, chr6-26240874, chr6-26240881, chr6-26240888, chr6-26240920, chr6-26240922, chr6-26240930, chr6-26240939 and chr6-26240950; The marker 3 includes methylation sites: chr14-102248201, chr14-102248195, chr14-102248189, chr14-102248169, chr14-102248164, chr14-102248159, chr14-102248148, chr14-102248140, chr14-102248138, chr14-102248133, chr14-102248131, chr14-102248127, chr14-102248106, chr14-102248104 and chr14-102248099; The marker 4 includes methylation sites: chr12-133485324, chr12-133485336, chr12-133485339, chr12-133485341, chr12-133485349, chr12-133485354, chr12-133485382, chr12-133485385, and chr12-133485416; The marker 5 includes methylation sites: chr2-73147841, chr2-73147838, chr2-73147830, chr2-73147828, chr2-73147826, chr2-73147823, chr2-73147817, chr2-73147815, chr2-73147812, and chr2-73147804. chr2-73147799, chr2-73147790, chr2-73147772, chr2-73147768, chr2-73147756 and chr2-73147739; The marker 6 includes methylation sites: chr18-60985741, chr18-60985733, chr18-60985720, chr18-60985713, chr18-60985706, chr18-60985702, chr18-60985691, chr18-60985688, chr18-60985676, chr18-60985666, chr18-60985663, chr18-60985660, chr18-60985657, chr18-60985655, and chr18-60985646; The marker 8 includes methylation sites: chr20-37434793, chr20-37434790, chr20-37434766, chr20-37434749, chr20-37434738, chr20-37434731, chr20-37434728, chr20-37434724, chr20-37434702, and chr20-37434684; The marker 9 includes methylation sites: chr2-182322501, chr2-182322503, chr2-182322530, chr2-182322537, chr2-182322545, chr2-182322549, chr2-182322564, chr2-182322569, chr2-182322574, and chr2-182322600; The marker 10 includes methylation sites: chr3-179169204, chr3-179169206, chr3-179169212, chr3-179169215, chr3-179169224, chr3-179169237, chr3-179169249, chr3-179169252, chr3-179169268, chr3-179169290, chr3-179169292, chr3-179169295, chr3-179169297, chr3-179169300, chr3-179169306 and chr3-179169311; The marker 12 includes methylation sites: chr16-23847507, chr16-23847513, chr16-23847519, chr16-23847522, chr16-23847525, chr16-23847529, chr16-23847535, chr16-23847547, chr16-23847551, chr16-23847556, chr16-23847560, chr16-23847568, chr16-23847575, and chr16-23847586; The marker 13 includes methylation sites: chr8-41504171, chr8-41504182, chr8-41504197, chr8-41504217, chr8-41504252, chr8-41504254, chr8-41504258, chr8-41504263, and chr8-41504265.

4. The methylation marker combination according to claim 3, characterized in that, The conversion rate control markers include methylation sites: chr7-5571912, chr7-5571911, chr7-5571910, chr7-5571909, chr7-5571902, chr7-5571891, chr7-5571887, chr7-5571874, chr7-5571873, chr7-5571872, chr7 -5571868, chr7-5571846, chr7-5571844, chr7-5571841, chr7-5571836, chr7-5571833, chr7-5571828, chr7-5571821, chr7-5571820, chr7-5571814, chr7-5571811 and chr7-5571810.

5. A kit for detecting early gastric cancer, characterized in that, The invention includes a detection reagent for detecting the degree of methylation in each target region of the methylation marker combination according to any one of claims 1-4.

6. The reagent kit according to claim 5, characterized in that, The detection reagent is applicable to any one or more methods, such as bisulfite conversion sequencing, methylation chip method, or combinations thereof, to detect the degree of methylation in each target region of the methylation biomarker combination.

7. An early gastric cancer detection system, characterized in that, The detection system includes a calculation module, which calculates the methylation degree of the marker region in the marker combination according to any one of claims 1-4 as a predicted value of the detection model. The detection model predicts values ​​using a logistic regression model. The average methylation rate of each biomarker region is multiplied by its corresponding weight coefficient and then summed. If the result is greater than a threshold, it is considered a positive risk indicator for gastric cancer.

Citation Information

Patent Citations

  • Methylation level based broad-spectrum marker for detecting tumors, and applications thereof

    CN110229913A

  • Methods of amplifying DNA to maintain methylation status

    CN110741092A