Methylation marker combination for detecting early esophageal cancer and application thereof
By using NGS technology to quantitatively detect multiple methylation marker genes in plasma samples from patients with early esophageal cancer, and combining this with a logistic regression model, the problems of high equipment cost, complex operation, and low sensitivity in existing esophageal cancer detection methods are solved. This achieves a non-invasive early esophageal cancer detection method with high sensitivity and high specificity, which is suitable for large-scale screening and clinical promotion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIAXING YUNYING MEDICAL INSPECTION CO LTD
- Filing Date
- 2026-06-11
- Publication Date
- 2026-07-14
AI Technical Summary
Existing methods for esophageal cancer detection suffer from high equipment costs, demanding operational skills, high invasiveness, and low sensitivity, failing to meet the needs of early cancer screening. Furthermore, the sensitivity and specificity of existing methylation detection methods need to be improved.
NGS technology was used to quantitatively detect multiple methylation marker genes in plasma samples from patients with early esophageal cancer. Combined with a logistic regression model, specific DNA regions and primer combinations were used to detect the signals through methods such as pyrosequencing and bisulfite conversion sequencing, achieving high sensitivity and high specificity for non-invasive early esophageal cancer signal detection.
It achieves high sensitivity and high specificity in early esophageal cancer detection, avoids invasive testing, is suitable for large-scale screening, is low-cost, has standardized operation, and is suitable for clinical promotion.
Smart Images

Figure CN122382201A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and more specifically, to a combination of methylation biomarkers for detecting early esophageal cancer and their applications. Background Technology
[0002] Currently, clinical methods for esophageal cancer detection and screening include imaging endoscopy, tissue biopsy, and detection of tumor serum markers such as CEA and SCC. However, these methods are either unsuitable for early cancer screening due to high equipment costs, high operational technical requirements, high testing costs, or high invasiveness and low sensitivity, and cannot meet clinical testing needs.
[0003] DNA methylation abnormalities typically occur in the early stages of cancer and persist throughout the entire process of cancer development and progression. Once formed, the methylation state requires prolonged and continuous stimulation from the external environment to change. Therefore, DNA methylation serves as an important biomarker for early screening, auxiliary detection, and prognosis. Currently, common first-generation methylation detection methods using Q-PCR technology only provide a close-up, localized view of methylation, limited to individual genes and individual samples, offering only qualitative detection, and the sensitivity and specificity of detection need improvement. In contrast, next-generation sequencing (NGS) provides a panoramic, high-resolution view of methylation, enabling quantitative detection of methylation rates at specific methylation sites in specific genes. Moreover, it can simultaneously quantify multiple methylation sites across a large number of genes, providing more accurate results. Furthermore, NGS protocols can simultaneously perform large-scale parallel testing of hundreds or thousands of samples, significantly reducing experimental costs and testing cycles, making it an important early tumor methylation detection method now and in the future. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention employs NGS (Next Generation Sequencing) on plasma samples from patients with early-stage esophageal cancer to quantitatively detect multiple methylation marker genes specific to esophageal cancer, achieving highly sensitive, highly specific, rapid, and non-invasive early esophageal cancer signal detection.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A combination of methylation biomarkers for detecting early esophageal cancer includes at least one of the following DNA regions (human genome version number GRCh37): Biomarker 1, chr12: 133485485-133485635; Biomarker 2, chr19: 38183058-38183208; Biomarker 3, chr10: 22634386-22634536; Biomarker 4, chr1: 111217322-111217472; Biomarker 5, chr12: 6665281-6665431; Biomarker 6... chr17:75369561-75369711, Marker 7, chr17:75369627-75369777, Marker 8, chr3:157821497-157821647, Marker 9, chr7:93519341-93519491, Marker 10, chr8:57358422-57358572, Marker 11, chr2:63281133-63281283, Marker 12, chr1:44883337-44883487.
[0007] Furthermore, the biomarker set also includes the internal control biomarker region chr7:5571788-5571939, with 33 methylation sites.
[0008] Furthermore, marker 1 includes methylation sites: chr12-133485488, chr12-133485504, chr12-133485507, chr12-133485514, chr12-133485530, chr12-133485542, chr12-133485546, chr12-133485573, chr12-133485581, chr12-133485591, chr12-133485629, and chr12-133485632;
[0009] Marker 2 includes methylation sites: chr19-38183206, chr19-38183200, chr19-38183194, chr19-38183165, chr19-38183146, chr19-38183131, chr19-38183127, chr19-38183121, chr19-38183113, chr19-38183098, chr19-38183092, chr19-38183090, and chr19-38183081;
[0010] Marker 3 includes methylation sites: chr10-22634532, chr10-22634529, chr10-22634526, chr10-22634513, chr10-22634504, chr10-22634495, chr10-22634491, chr10-22634484, chr10-22634467, chr10-22634463, c hr10-22634458, chr10-22634455, chr10-22634450, chr10-22634443, chr10-22634440, chr10-22634433, chr10-22634417, chr10-22634415, chr10-22634408, chr10-22634396 and chr10-22634390;
[0011] Marker 4 includes methylation sites: chr1-111217330, chr1-111217341, chr1-111217344, chr1-111217351, chr1-111217358, chr1-111217374, chr1-111217376, chr1-111217382, chr1-111217393, chr1-1112 17396, chr1-111217399, chr1-111217402, chr1-111217406, chr1-111217421, chr1-111217425, chr1-111217433, chr1-111217436, chr1-111217447, chr1-111217453 and chr1-111217467;
[0012] Marker 5 includes methylation sites: chr12-6665425, chr12-6665400, chr12-6665388, chr12-6665371, chr12-6665336, chr12-6665331, chr12-6665312, chr12-6665306, chr12-6665301 and chr12-6665289;
[0013] Marker 6 includes methylation sites: chr17-75369563, chr17-75369585, chr17-75369592, chr17-75369594, chr17-75369601, chr17-75369603, chr17-75369624, chr17-75369631, chr17-75369658, chr17-75369662, chr17-75369664, chr17-75369668, chr17-75369687, chr17-75369700, chr17-75369702, and chr17-75369707;
[0014] Marker 7 includes methylation sites: chr17-75369775, chr17-75369767, chr17-75369749, chr17-75369736, chr17-75369728, chr17-75369715, chr17-75369707, chr17-75369702, chr17-75369700, chr17-75369687, chr17-75369668, chr17-75369664, chr17-75369662, chr17-75369658, and chr17-75369631;
[0015] Marker 8 includes methylation sites: chr3-157821604, chr3-157821597, chr3-157821595, chr3-157821586, chr3-157821555, chr3-157821545, chr3-157821537, chr3-157821529, chr3-157821519, and chr3-157821499;
[0016] Marker 9 includes methylation sites: chr7-93519474, chr7-93519471, chr7-93519436, chr7-93519434, chr7-93519419, chr7-93519417, chr7-93519408, chr7-93519405, chr7-93519402, chr7-93519386, chr7-93519373 and chr7-93519368;
[0017] Marker 10 includes methylation sites: chr8-57358564, chr8-57358557, chr8-57358543, chr8-57358536, chr8-57358534, chr8-57358520, chr8-57358508, chr8-57358506, chr8-57358471, chr8-57358465, chr8-57358454, chr8-57358450, chr8-57358446, chr8-57358441, and chr8-57358423;
[0018] The marker 11 includes methylation sites: chr2-63281139, chr2-63281156, chr2-63281158, chr2-63281161, chr2-63281167, chr2-63281169, chr2-63281177, chr2-63281188, chr2-63281190. chr2-63281199, chr2-63281202, chr2-63281208, chr2-63281221, chr2-63281223, chr2-63281237, chr2-63281243, chr2-63281248, chr2-63281255 and chr2-63281265;
[0019] Marker 12 includes methylation sites: chr1-44883346, chr1-44883354, chr1-44883362, chr1-44883364, chr1-44883368, chr1-44883372, chr1-44883383, chr1-44883386, chr1-44883392, chr1-44883434, chr1-44883436, chr1-44883444, chr1-44883446, chr1-44883459, and chr1-44883475;
[0020] Internal control markers include methylation sites: chr7-5571939, chr7-5571934, chr7-5571933, chr7-5571927, chr7-5571922, chr7-5571918, chr7-5571917, chr7-5571916, chr7-5571913, chr7-5571912, chr7-5571911, chr7-5571910, chr7-5571903, chr7-5571892, chr7-5571888, and chr7-5571875. chr7-5571874, chr7-5571873, chr7-5571869, chr7-5571847, chr7-5571845, chr7-5571842, chr7-5571837, chr7-5571834, chr7-5571829, chr7-5571822, chr7-5571821, chr7-5571815, chr7-5571812, chr7-5571811, chr7-5571809, chr7-5571802 and chr7-5571790.
[0021] Furthermore, the methylation levels of the aforementioned DNA regions differ between patients with early-stage esophageal cancer and healthy individuals.
[0022] Furthermore, the biomarker combination has a corresponding primer combination, including the following primers:
[0023] The upstream primer for chr12:133485485-133485635 is GAGCGGTATTGAAGTGAGTT, and the downstream primer is ATCGACGTTTAATTTTTAACC.
[0024] The upstream primer for chr19:38183058-38183208 is GGCGGAATCGGTTGCG, and the downstream primer is CATAAAAACATAAACTATAACCCGATAAAT.
[0025] The upstream primer for chr10:22634386-22634536 is TTTTCGTCGTCGTTTGTT, and the downstream primer is TTCCGATTCCRATAAATTACC.
[0026] The upstream primer for chr1:111217322-111217472 is TTTATTAAGYGTGTGGGTAT, and the downstream primer is AAAACGACCTAACCTCAC.
[0027] The upstream primer for chr12:6665281-6665431 is TTTAGGCGGTGGTTTT, and the downstream primer is ACAAAACCRACCAATC.
[0028] The upstream primer for chr17:75369561-75369711 is TCGTTGTTTATTAGTTATTATGTCGGATTT, and the downstream primer is CCCCCGTCCCGCGCCA.
[0029] The upstream primer for chr17:75369627-75369777 is AGCGGGAGGGCGTTTG, and the downstream primer is CTTCGAAAATAAATACTAAACTAACTACTA.
[0030] The upstream primer for chr3:157821497-157821647 is AGTTGTTGGTTGTATGAAAAG, and the downstream primer is CCGCTCTTTCTATTCTCTCTT.
[0031] The upstream primer for chr7:93519341-93519491 is TATATTTGGGAGGTTTG, and the downstream primer is CCCAAACTAAAACTTCCTAT.
[0032] The upstream primer for chr8:57358422-57358572 is AGGGATTTCGTTTTGCGT, and the downstream primer is CGCAATCCTAACTACATTC.
[0033] The upstream primer for chr2:63281133-63281283 is GGTTTAYGTGGTTATTAATG, and the downstream primer is TAAAAATATCAAAATAACGAATC.
[0034] The upstream primer for chr1:44883337-44883487 is GAAGTGATTYGGGTTGT, and the downstream primer is AAAACCCACCCCGTAAC.
[0035] The upstream primer for the internal reference region chr7:5571788-5571939 is TAAGGTTAGGGATAGGATAGTT, and the downstream primer is TAACCACCACCCAACAC.
[0036] A kit for detecting early esophageal cancer, comprising a methylation detection reagent for the degree of methylation of at least one marker region in the above-mentioned combination of markers.
[0037] Furthermore, the methylation detection reagents include reagents used in any one or more of the following methods, which include: pyrosequencing, bisulfite conversion sequencing, methylation array, qPCR, digital PCR, next-generation sequencing, third-generation sequencing, whole-genome methylation sequencing, DNA enrichment detection, simplified bisulfite sequencing, HPLC, MassArray, methylation-specific PCR, or combinations thereof.
[0038] An early esophageal cancer detection system includes a calculation module that calculates the detection model prediction value based on the methylation level of at least one marker region in the above-mentioned combination of markers.
[0039] The predicted value of the detection model is calculated using a logistic regression model. The average methylation rate of each marker region is multiplied by the corresponding weight coefficient and then summed. If the result is greater than the threshold, it is judged as esophageal cancer positive.
[0040] By adopting the above technical solution, the beneficial effects of the present invention are as follows:
[0041] High sensitivity and high specificity: NGS quantitatively detects 178 methylation sites in 12 biomarker regions. Combined with a logistic regression model, the sensitivity reaches 89.05%, specificity 96.75%, and AUC=0.982245 in the training set (Model 1); in the validation set, the sensitivity is 88.88%, specificity 96%, and AUC=0.978654. Compared to traditional imaging, endoscopic, or serum biomarkers (such as CEA and SCC), this invention is more suitable for early screening, avoiding invasiveness and low sensitivity issues.
[0042] Non-invasive rapid testing: Based on cell-free DNA in plasma, sample collection is simple, and the entire process includes lysis, sulfite conversion, library construction and sequencing, achieving non-invasive early signal detection, suitable for large-scale screening.
[0043] Accurate quantification: Calculated using average methylation rate (excluding primer sites and low-depth sites), with an internal control ACTB conversion rate >99.5% to ensure quality control, providing a panoramic high-resolution view, superior to the qualitative detection of first-generation Q-PCR.
[0044] Model optimization and applicability: The optimal model is selected through data elimination, weight adjustment and ROC evaluation. It supports multiple sample types (such as tissue and urine) and can be extended to prognostic assessment to help patients intervene in advance and improve survival rate.
[0045] Cost-effective: Using standard equipment (such as Illumina MiniSeq) and reagent kits, the cost is low, the operation is standardized, and it is suitable for clinical application. Attached Figure Description
[0046] Figure 1 This is the ROC curve of the training set results for Model 1 of this invention.
[0047] Figure 2 This is the ROC curve of the training set results for Model 2 of this invention.
[0048] Figure 3 This is the ROC curve of the training set results for Model 3 of this invention.
[0049] Figure 4 This is the ROC curve of the training set results for Model 4 of this invention.
[0050] Figure 5 This is the ROC curve of the training set results for Model 5 of this invention.
[0051] Figure 6 This is the ROC curve of the validation set results for Model 1 of this invention. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] Table 1. Blood Methylation Primers
[0054] Table 1 lists the primer information for the methylation biomarkers used in this invention for early esophageal cancer detection. Specifically, it includes the amplification sequence coordinates of each biomarker in the reference genome GRCh37, the upstream and downstream primer sequences, and the corresponding methylation sites covered. All primers are designed for differentially methylated regions that have been screened and validated, and are used to specifically amplify target regions of cell-free plasma DNA after sulfite conversion. Using the primer combinations shown in Table 1, stable and simultaneous detection of multiple methylation sites can be achieved, providing a reliable experimental basis for subsequent library construction, high-throughput sequencing, and methylation rate calculation, thereby ensuring the accuracy, repeatability, and feasibility of the detection method of this invention.
[0055] Experimental steps
[0056] Blood sample collection
[0057] 1. Whole blood collection
[0058] We used disposable vacuum blood collection tubes to collect and preserve 10 ml of blood in the morning fasting state from 381 patients with early-stage (stage I and II) esophageal cancer (mean age 56 years, females 36.75%, 201 cases were included in the training set and 180 cases in the validation set) and 273 healthy individuals (mean age 58 years, females 34.8%, 123 cases were included in the training set and 150 cases in the validation set).
[0059] 2. Preparation of plasma samples
[0060] Centrifuge the blood collection tube containing whole blood for 12 minutes at a centrifugal force of 1500 rcf. Remove the blood collection tube from the centrifuge and transfer the supernatant plasma into a labeled 2 mL centrifuge tube to obtain the sample required for the experiment.
[0061] Pretreatment of cell-free DNA in plasma
[0062] The samples were lysed, converted to sulfite, and purified using the methylation detection sample pretreatment reagent, plasma sample lysis and sulfite conversion kit produced based on the present invention. The operation steps are as follows.
[0063] 1. Preparation of reagents in the kit
[0064] Prepare the reagents in the above kit.
[0065] Table 2. Reagent Preparation
[0066] 2. cfDNA extraction: Take 1.8 ml of plasma, add 100 μl of 20% SDS, mix by inverting, centrifuge briefly, then add 30 μl of PK enzyme (20 mg / ml), vortex 4-5 times to mix; incubate at 60℃ for 30 min, place at room temperature for 5 min, and centrifuge briefly.
[0067] The above mixture was further extracted using an Auto-Pure-20B purifier. The program was run according to the settings. After the program was completed, the liquid was taken out from the last well and placed into a new PCR tube. The tube was labeled with the corresponding sample number, which is cfDNA. It can then be used for the next step of sulfite conversion.
[0068] (1) Sample addition: Add all the solution in the EP tube to column 1 of the purification kit (add different samples to different wells), and be sure to record the sample number;
[0069] (2) After installing the magnetic rod sleeve, place the purification kit into the fully automated nucleic acid extractor and set it up and operate it according to the prescribed procedure:
[0070] (3) After the program finishes running, the solution is transferred to a new 0.5 mL centrifuge tube. The collected solution is the cfDNA with magnetic beads and sulfite.
[0071] 3. Sulfite Conversion
[0072] 3.1 System preparation: Prepare sulfite reaction systems in 0.5 mL PCR tubes according to the table below for the extracted DNA samples.
[0073] Table 3. Conversion Table for Sulfite Reaction System
[0074] 3.2 Sulfite Conversion
[0075] After vortexing and mixing, and briefly centrifuging, place the PCR tube in a PCR instrument and incubate at 95℃ for 5 min and 85℃ for 60 min. After incubation, cool to room temperature and briefly centrifuge to obtain the transformation product.
[0076] 3.3 Bis-DNA purification treatment
[0077] The purified product was run using the Auto-Pure-32A purifier according to the set program. After the run was completed, the liquid was taken out from the well and placed into a new PCR 8-strip, which was then labeled with the corresponding sample number. This is the purified Bis-DNA, which can be used for the next library enrichment reaction.
[0078] (1) Sample addition: Add all the solution from the PCR tube to column 1 of the purification kit (add different samples to different wells), and be sure to record the sample number;
[0079] (2) After installing the magnetic rod sleeve, place the purification kit into the fully automated nucleic acid extractor and set it up and operate it according to the prescribed procedure:
[0080] (3) After the program finishes running, aspirate the solution into a new 1.5 mL centrifuge tube. The collected solution is the DNA after sulfite conversion and should be used for library construction immediately.
[0081] Library enrichment reaction
[0082] The Bis DNA obtained after purification in the previous step was subjected to a library enrichment reaction.
[0083] Preparation of enrichment PCR reaction
[0084] Prepare the enrichment reaction solution as needed according to the formula shown in the table below. After preparation, aliquot the solution into clean eight-slot containers. The aliquoted enrichment reaction solutions are grouped according to the mixed primers and are referred to as MEC enrichment PCR reaction strip-1 to MEC enrichment PCR reaction strip-4.
[0085] Table 4. Enrichment Reaction Table
[0086] 1. Gently peel off the caps of the MEC enrichment PCR reaction strip-1 tube to the MEC enrichment PCR reaction strip-4 tube prepared in the previous step, and add 1 μL of Taq enzyme to MEC enrichment reaction solution-1 to MEC enrichment PCR reaction solution-4 respectively, ensuring that the pipette tip is fully inserted into the reaction solution;
[0087] 2. Add 20 μL of the Bis DNA sample to be tested after sulfite conversion to the four enrichment reaction strips in sequence along the PCR tube wall, and carefully close the tube caps.
[0088] 3. After vortexing the PCR reaction strip to mix well, centrifuge, taking care to avoid air bubbles;
[0089] 4. Place the above four enrichment PCR reaction strips into the PCR instrument;
[0090] 5. Open the PCR instrument settings interface, set the amplification program according to the table below, and perform PCR amplification.
[0091] Table 5. PCR reaction procedure for library enrichment
[0092] Library preparation reaction
[0093] Preparation for the connector connection reaction:
[0094] Prepare the required number of adapters according to the formula shown in the table below (except for Ace Taq enzyme and the PCR reaction product enriched in the previous step). Alternatively, you can prepare them in advance and aliquot them into clean 8-strips and store them at -20°C for later use.
[0095] After removing the connector reaction solution from the refrigerator and allowing it to melt, centrifuge it briefly in a centrifuge before use.
[0096] Table 6. Reaction Components for Library Preparation
[0097] 1. Gently open the cap of the reaction strip connector and add 0.5 μL of Taq enzyme into the reaction strip connector, ensuring that the pipette tip is fully inserted into the reaction solution;
[0098] 2. Take 5 μL of the enriched PCR amplification product from the previous step and add it sequentially along the wall of the PCR tube to the adapter connection reaction strip, then carefully close the tube cap.
[0099] 3. Connect the connector to the reaction strip, shake to mix, and then centrifuge, taking care to avoid air bubbles;
[0100] 4. Connect the above-mentioned adapter to the reaction strip and place it into the PCR instrument;
[0101] 5. Open the PCR instrument settings interface, set the amplification program according to the table below, and perform PCR amplification.
[0102] Table 7. PCR Reaction Procedure for Library Preparation
[0103] Library Product Sorting
[0104] 1. Remove the magnetic beads used for purification, vortex to mix, and incubate at room temperature for at least 30 minutes;
[0105] 2. Add 30 μL of MEC library-1 to MEC library-2, add 30 μL of MEC library-3 to MEC library-4, and add 10 μL of purified water to each sample reaction tube in sequence;
[0106] 3. Take 126 μL of the incubated magnetic beads (shake well again before use) and add them sequentially to each sample reaction tube. Shake well and then incubate briefly to completely mix and resuspend the DNA and magnetic beads.
[0107] 4. Incubate at room temperature for 5 minutes;
[0108] 5. Incubate on a magnetic rack for 5 minutes until the solution is clear. Carefully aspirate and discard the supernatant, being careful not to disturb the magnetic beads.
[0109] 6. Add 200µL of freshly prepared 80% ethanol solution (the ethanol solution should just cover the magnetic bead sample), place it on a magnetic rack, and incubate it on the magnetic rack for 30 seconds until the solution is clear. Discard the supernatant.
[0110] 7. Repeat step 6 above for a second wash;
[0111] 8. Ensure that the ethanol solution in the centrifuge tube has been completely discarded, and place the eight-tube or PCR tube on a magnetic rack to air dry at room temperature for 3-5 minutes;
[0112] 9. Remove the eight-tube or PCR tube from the magnetic rack, add 70 μL of purified water to fully wet the magnetic beads, shake well to mix, and then quickly centrifuge to collect the liquid at the bottom of the tube. Let it stand at room temperature for 5 min.
[0113] 10. Place the 8-tube or PCR tube on a magnetic rack and let it stand for 5 minutes until the solution is clear. Transfer 60 μL of the supernatant to a new 8-tube or 1.5 mL centrifuge tube to obtain the sequencing library. Use the Equalbit 1×dsDNA HS Assay Kit and its matching Qubit fluorescence instrument to detect the concentration. The library concentration should be ≥1 ng / μL. Otherwise, the library preparation sample does not meet the requirements and should be reconstructed.
[0114] 11. Perform deep sequencing on the library using an Illumina MiniSeq sequencer, with each sample being sequenced at 150M.
[0115] Methylation data processing workflow
[0116] The first step is data processing and quality control.
[0117] 1. The methylation rate of 178 methylation sites for 12 gene markers in plasma samples from esophageal cancer-positive and normal individuals was statistically analyzed. The methylation rate was calculated as: (Number of methylated reads at that site / (Number of methylated reads at that site + Number of unmethylated reads at that site)). Data from methylation sites located on primers were removed, and the methylation rate of methylation sites with excessively low sequencing depth (site sequencing depth > 10000X) was excluded. The average methylation rate of all methylation sites within each gene marker was calculated to represent the methylation rate of that gene marker. Gene marker 1 contains 12 methylation sites, gene marker 2 contains 13 methylation sites, gene marker 3 contains 21 methylation sites, gene marker 4 contains 20 methylation sites, gene marker 5 contains 10 methylation sites, gene marker 6 contains 16 methylation sites, gene marker 7 contains 15 methylation sites, gene marker 8 contains 10 methylation sites, gene marker 9 contains 12 methylation sites, gene marker 10 contains 15 methylation sites, gene marker 11 contains 19 methylation sites, and gene marker 12 contains 15 methylation sites.
[0118] 2. The methylation conversion rate of all 33 methylation sites on the internal reference gene marker ACTB in all samples was greater than 99.0%, indicating that the methylation conversion in this experiment was successful.
[0119] The second step is model training. The methylation rates of the 12 genes selected from the training set data are substituted into the logistic regression model for training. The coefficients and thresholds corresponding to each gene marker are obtained. The methylation level of each gene marker is multiplied by the corresponding coefficient and then added together. If the sum is greater than the threshold, the sample is judged as positive; otherwise, it is negative.
[0120] The third step is model tuning and evaluation. The weights of the dependent variable in the model are adjusted according to the ratio of positive and negative data to balance the training set data. The regularization parameters are tuned to find the optimal parameters. The threshold is adjusted to adapt to clinical testing. The sensitivity and specificity of different models are calculated. ROC curves are plotted to evaluate the excellence of different models and select the best model.
[0121] The fourth step is to perform model validation analysis on the validation set using the optimal model described above, calculate its sensitivity and specificity, and plot the ROC curve.
[0122] Example 1
[0123] Biomarker Screening and Validation: This embodiment details the screening process from initial candidate biomarkers to a final selection of 12 biomarkers, validating their performance in leukocyte, tissue, and plasma samples through multiple rounds of experiments. Screening was based on metrics such as methylation rate, sequencing depth, contribution rate, and specificity.
[0124] Based on the above experimental method, the following experiment was conducted:
[0125] Background and Objectives: This series of experiments aims to establish a method for detecting methylation biomarkers in early esophageal cancer based on NGS sequencing technology through multiple rounds of screening and validation. 178 methylation sites in 12 DNA regions were screened from initial candidate biomarkers and systematically validated in leukocyte, tissue, and plasma samples. Finally, a logistic regression detection model was constructed.
[0126] Experiment 1: Preliminary Screening Experiment for White Blood Cells
[0127] Experimental objective: To conduct preliminary detection of 20 candidate biomarkers using normal human leukocyte samples, confirm the methylation rate and sequencing depth of each gene, and screen biomarkers with an average methylation rate of less than 10% for subsequent tissue sample validation.
[0128] Experimental methods:
[0129] 1. Collect white blood cell samples from 20 normal individuals;
[0130] 2. Perform PCR amplification using candidate primer combinations;
[0131] 3. Perform NGS sequencing to analyze the average methylation rate and average sequencing depth of methylation sites;
[0132] 4. Screening criteria: average methylation rate < 10%, average sequencing depth > 100×.
[0133] Experimental Results and Analysis:
[0134] 1. All biomarkers showed good sequencing depth, meeting the requirement of sequencing depth > 100×;
[0135] 2. Among the 20 biomarkers, based on the average methylation rate and average sequencing depth of each biomarker, two biomarkers (chr7:100318354-100318504, chr3:13522704-13522854) had an average methylation rate >10%, which did not meet the screening criteria.
[0136] 3. The 18 selected biomarkers entered the esophageal cancer tissue testing and verification stage.
[0137] Conclusion: Eighteen candidate biomarkers with low methylation rates and good sequencing depth in leukocytes were successfully screened, laying the foundation for subsequent tissue sample validation, as shown in the table below.
[0138] Table 8. White blood cell study data
[0139] Experiment 2: Tissue Sample Validation Experiment
[0140] Experimental objective: To evaluate the sensitivity and contribution rate of candidate biomarkers in esophageal cancer tissues through tissue sample validation, and to screen out biomarkers with high sensitivity and high contribution rate in tissues.
[0141] Experimental methods:
[0142] 1. Collected 20 esophageal cancer tissue samples;
[0143] 2. Detection was performed using primers retained after leukocyte screening;
[0144] 3. Analyze the methylation rate and contribution rate of each biomarker in tissue samples;
[0145] 4. Screening criteria: Contribution rate ≥ 80%.
[0146] Experimental Results and Analysis:
[0147] 1. A positive marker was defined as one with an average methylation rate greater than 15% for each marker, and markers with a depth > 500× were included in the statistical analysis;
[0148] 2. Among the 18 biomarkers, based on the calculation of average methylation rate and contribution rate, three biomarkers (chr17:72920346-72920496, chr14: 24616416-24616566, chr1:25255984-25256134) had a contribution rate of less than 80%, and therefore did not meet the screening criteria.
[0149] 3. The 15 biomarkers selected will enter the plasma testing and validation stage for esophageal cancer.
[0150] Conclusion: After validation with tissue samples, 15 biomarkers with high sensitivity (contribution rate ≥80%) in esophageal cancer tissues were screened, providing reliable candidate biomarkers for validation with plasma samples, as shown in the table below.
[0151] Table 9. Organizational Research Data
[0152] Table 10. Contribution rate of each biomarker in the organization
[0153] Experiment 3: Plasma Sensitivity Detection
[0154] Objective: To validate tissue-screened biomarkers in plasma samples, evaluate the contribution rate of each biomarker in plasma, and screen biomarker combinations suitable for plasma testing.
[0155] Experimental methods:
[0156] 1. Collect plasma samples from 20 patients with esophageal cancer;
[0157] 2. Detection was performed using primer combinations selected from tissue screening;
[0158] 3. Analyze the contribution rate of each biomarker in esophageal cancer plasma;
[0159] Experimental Results and Analysis:
[0160] The 15 biomarkers screened in the early stage were tested. The average methylation of each biomarker was greater than 10% as a positive marker and the contribution rate was greater than 20%. Statistical calculations were performed, and one biomarker (chr1:169396638-169396788) had a contribution rate of less than 20%, which did not meet the screening criteria.
[0161] Conclusion: Fourteen biomarkers with good contribution rates in plasma were successfully screened, as shown in the table below.
[0162] Table 11. Data on plasma contribution rate of esophageal cancer
[0163] Table 12. Contribution rate of various plasma biomarkers
[0164] Experiment 4: Plasma-specific detection
[0165] Experimental objective: To validate esophageal cancer plasma biomarkers with high contribution rates in normal human plasma samples, evaluate the specificity of each biomarker in normal human plasma, and screen out biomarkers with high specificity.
[0166] Experimental methods:
[0167] 1. Collect plasma samples from 20 healthy individuals;
[0168] 2. Detection was performed using primer combinations for esophageal cancer plasma screening;
[0169] 3. Analyze the specificity of each biomarker in normal human plasma;
[0170] Experimental Results and Analysis:
[0171] Fourteen biomarkers were screened from esophageal cancer plasma. A positive marker was defined as an average methylation level greater than 10% and a specificity ≥ 95%. Two biomarkers (chr12:124950743-124950893 and chr17:79259487-79259637) had specificities below 95% and did not meet the screening criteria. Twelve biomarkers with high specificity were selected to provide reliable data for model construction, as shown in the table below.
[0172] Table 13. Specific data from normal human plasma studies
[0173] Table 14. Specificity of plasma biomarkers
[0174] Through multiple rounds of screening, 12 biomarker combinations were finally identified, totaling 178 methylation sites, for subsequent model construction.
[0175] Example 2: Sample Collection, Data Processing, and Detection Model Construction
[0176] 2.1 Sample Collection
[0177] Training set: 201 positive plasma samples from early-stage esophageal cancer (pathological examination included esophageal cancer, esophageal squamous cell carcinoma, recurrent esophageal cancer, and small cell esophageal cancer) and 123 negative plasma samples from healthy individuals were collected. Among the positive samples, esophageal squamous cell carcinoma accounted for the highest proportion; the negative samples were from healthy individuals.
[0178] Validation set: 180 positive esophageal cancer samples and 150 negative samples from healthy individuals.
[0179] 2.2 DNA Extraction and Sulfite Conversion
[0180] Cell-free DNA was extracted from plasma samples and transformed using the Yunying sulfite conversion kit. Transformation efficiency was assessed using the conversion rate of the internal control region ACTB, calculated as follows:
[0181] Detection site conversion rate = (1- The conversion rate must be greater than 99.5% (100%), otherwise the sample will be discarded.
[0182] 2.3 Targeted amplification and NGS sequencing
[0183] Targeted amplification was performed using the primer combination described above. After purification of the amplified products, deep sequencing was performed on an Illumina MiniSeq sequencer (150M reads per sample). Sequencing data processing included primer site removal and filtering of sites with sequencing depth < 10000X. The methylation rate for each site was calculated as: number of methylated reads / (number of methylated reads + number of unmethylated reads).
[0184] The average methylation rate of each marker region is calculated based on its site:
[0185] Marker 1: 12 loci; Marker 2: 13 loci; Marker 3: 21 loci; Marker 4: 20 loci; Marker 5: 10 loci; Marker 6: 16 loci; Marker 7: 15 loci; Marker 8: 10 loci; Marker 9: 12 loci; Marker 10: 15 loci; Marker 11: 19 loci; Marker 12: 15 loci.
[0186] 2.4 Model Training and Optimization
[0187] Five models were trained using the LogisticRegression model (solver='liblinear') from the Python sklearn library, based on the average methylation rate of 12 markers in the training set. In the `class_weight` parameter, 0 represents negative samples and 1 represents positive samples, with the weight of positive samples being a multiple of that of negative samples.
[0188] Table 15. Model Parameter Table
[0189] Model predicted value = Σ (average methylation rate_i × weight coefficient_i). If the predicted value > the threshold, it is considered positive.
[0190] Table 16. Training Set Table
[0191] Combination Figure 1-5The training set shown exhibits that, among the five trained models, Model 1 identified 179 positive and 119 negative cases, with a sensitivity of 89.05% and specificity of 96.75%, a threshold of 2.97, and an area under the ROC curve (AUC) of 0.982245, demonstrating the best discriminative ability. Model 2 identified 161 positive and 118 negative cases, with a sensitivity of 80.10% and specificity of 95.93%, a threshold of 2.0, and an AUC of 0.971093. Model 3 identified 176 positive and 108 negative cases, with a sensitivity of... Model 1 achieved a sensitivity of 87.56% and a specificity of 87.8%, with a threshold of 2.05 and an area under the ROC curve (AUC) of 0.95219. Model 4 identified 170 positive and 111 negative cases, with a sensitivity of 84.58% and a specificity of 90.24%, a threshold of 2.45, and an AUC of 0.951988. Model 5 identified 166 positive and 116 negative cases, with a sensitivity of 82.59% and a specificity of 94.31%, a threshold of 2.14, and an AUC of 0.950451. Training results: Model 1 was the best, with a sensitivity of 89.05% (179 / 201), a specificity of 96.75% (119 / 123), a threshold of 2.97, and an AUC of 0.982245.
[0192] Table 17. Validation Set Table
[0193] Combination Figure 6 In the validation set (180 esophageal cancer patients + 150 normal individuals), 160 cases in Model 1 were judged as positive, with a sensitivity of 88.88%, and 144 cases were judged as negative, with a specificity of 96%. The area under the ROC curve was AUC-0.978654.
[0194] Example 3: Preparation and Application of the Reagent Kit
[0195] The kit includes: the above primer combination (after concentration optimization), sulfite conversion reagent, PCR amplification reagent (including enzyme, dNTPs, and buffer), magnetic bead purification reagent, and NGS library construction reagent.
[0196] Applications: Extracting DNA from plasma, performing transformation, amplification, and sequencing, calculating methylation rate, and inputting the results into a model for judgment. Suitable for clinical screening, testing 10mL plasma samples, with sensitivity >85% and specificity >95%.
[0197] Example 4: Application of the detection system
[0198] The detection system includes software modules that use Model 1 described above to calculate predicted values. It takes sample methylation data as input and outputs a positive / negative result and risk score. The system is integrated into a cloud platform and supports batch processing. Clinical application: In screening high-risk populations, positive samples are further confirmed endoscopically, reducing unnecessary invasive procedures.
[0199] The above embodiments verify the effectiveness of the biomarker combination, reagent kit and system of the present invention in the early detection of esophageal cancer, which can achieve non-invasive and highly accurate detection.
[0200] In summary, the model of this invention can identify the risk of early esophageal cancer with a highly sensitive, specific, rapid, and non-invasive method, so as to conduct relevant medical examinations in advance and prevent it at an early stage.
[0201] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any ordinary changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included within the protection scope of the present invention.
Claims
1. A combination of methylation markers, characterized in that, Using GRCh37 as a reference genome, the methylation marker combination includes at least one of the following DNA regions: Marker 1, chr12:133485485-133485635, Marker 2, chr19:38183058-38183208, Marker 3, chr10:22634386-22634536, Marker 4, chr1:111217322-111217472, Marker 5, chr12:6665281-6665431, Marker 6, chr17:75369561-75369711, Marker 7, chr17:75369627-75369777 Marker 8, chr3:157821497-157821647, Marker 9, chr7:93519341-93519491, Marker 10, chr8:57358422-57358572, Marker 11, chr2:63281133-63281283, Marker 12, chr1:44883337-44883487.
2. The methylation marker combination according to claim 1, characterized in that, The combination of markers also includes the internal reference marker region chr7:5571788-5571939.
3. The methylation marker combination according to claim 2, characterized in that, The marker 1 includes methylation sites: chr12-133485488, chr12-133485504, chr12-133485507, chr12-133485514, chr12-133485530, chr12-133485542, chr12-133485546, chr12-133485573, chr12-133485581, chr12-133485591, chr12-133485629, and chr12-133485632; The marker 2 includes methylation sites: chr19-38183206, chr19-38183200, chr19-38183194, chr19-38183165, chr19-38183146, chr19-38183131, chr19-38183127, chr19-38183121, chr19-38183113, chr19-38183098, chr19-38183092, chr19-38183090, and chr19-38183081; The marker 3 includes methylation sites: chr10-22634532, chr10-22634529, chr10-22634526, chr10-22634513, chr10-22634504, chr10-22634495, chr10-22634491, chr10-22634484, chr10-22634467, chr10-22634463. chr10-22634458, chr10-22634455, chr10-22634450, chr10-22634443, chr10-22634440, chr10-22634433, chr10-22634417, chr10-22634415, chr10-22634408, chr10-22634396 and chr10-22634390; The marker 4 includes methylation sites: chr1-111217330, chr1-111217341, chr1-111217344, chr1-111217351, chr1-111217358, chr1-111217374, chr1-111217376, chr1-111217382, chr1-111217393, chr1-111 217396, chr1-111217399, chr1-111217402, chr1-111217406, chr1-111217421, chr1-111217425, chr1-111217433, chr1-111217436, chr1-111217447, chr1-111217453 and chr1-111217467; The marker 5 includes methylation sites: chr12-6665425, chr12-6665400, chr12-6665388, chr12-6665371, chr12-6665336, chr12-6665331, chr12-6665312, chr12-6665306, chr12-6665301 and chr12-6665289; The marker 6 includes methylation sites: chr17-75369563, chr17-75369585, chr17-75369592, chr17-75369594, chr17-75369601, chr17-75369603, chr17-75369624, chr17-75369631, chr17-75369658, chr17-75369662, chr17-75369664, chr17-75369668, chr17-75369687, chr17-75369700, chr17-75369702, and chr17-75369707; The marker 7 includes methylation sites: chr17-75369775, chr17-75369767, chr17-75369749, chr17-75369736, chr17-75369728, chr17-75369715, chr17-75369707, chr17-75369702, chr17-75369700, chr17-75369687, chr17-75369668, chr17-75369664, chr17-75369662, chr17-75369658, and chr17-75369631; The marker 8 includes methylation sites: chr3-157821604, chr3-157821597, chr3-157821595, chr3-157821586, chr3-157821555, chr3-157821545, chr3-157821537, chr3-157821529, chr3-157821519 and chr3-157821499; The marker 9 includes methylation sites: chr7-93519474, chr7-93519471, chr7-93519436, chr7-93519434, chr7-93519419, chr7-93519417, chr7-93519408, chr7-93519405, chr7-93519402, chr7-93519386, chr7-93519373 and chr7-93519368; The marker 10 includes methylation sites: chr8-57358564, chr8-57358557, chr8-57358543, chr8-57358536, chr8-57358534, chr8-57358520, chr8-57358508, chr8-57358506, chr8-57358471, chr8-57358465, chr8-57358454, chr8-57358450, chr8-57358446, chr8-57358441, and chr8-57358423; The marker 11 includes methylation sites: chr2-63281139, chr2-63281156, chr2-63281158, chr2-63281161, chr2-63281167, chr2-63281169, chr2-63281177, chr2-63281188, chr2-63281190. chr2-63281199, chr2-63281202, chr2-63281208, chr2-63281221, chr2-63281223, chr2-63281237, chr2-63281243, chr2-63281248, chr2-63281255 and chr2-63281265; The marker 12 includes methylation sites: chr1-44883346, chr1-44883354, chr1-44883362, chr1-44883364, chr1-44883368, chr1-44883372, chr1-44883383, chr1-44883386, chr1-44883392, chr1-44883434, chr1-44883436, chr1-44883444, chr1-44883446, chr1-44883459, and chr1-44883475.
4. A combination of methylation markers according to claim 2 or 3, characterized in that, The internal reference markers include methylation sites: chr7-5571939, chr7-5571934, chr7-5571933, chr7-5571927, chr7-5571922, chr7-5571918, chr7-5571917, chr7-5571916, chr7-5571913, chr7-5571912, chr7-5571911, chr7-5571910, chr7-5571903, chr7-5571892, chr7-5571888, chr7-5571875. chr7-5571874, chr7-5571873, chr7-5571869, chr7-5571847, chr7-5571845, chr7-5571842, chr7-5571837, chr7-5571834, chr7-5571829, chr7-5571822, chr7-5571821, chr7-5571815, chr7-5571812, chr7-5571811, chr7-5571809, chr7-5571802 and chr7-5571790.
5. A methylation marker combination according to claim 4, characterized in that, The marker combination has a corresponding primer combination, including the following primers: The upstream primer for chr12:133485485-133485635 is GAGCGGTATTGAAGTGAGTT, and the downstream primer is ATCGACGTTTAATTTTTAACC. The upstream primer for chr19:38183058-38183208 is GGCGGAATCGGTTGCG, and the downstream primer is CATAAAAACATAAACTATAACCCGATAAAT. The upstream primer for chr10:22634386-22634536 is TTTTCGTCGTCGTTTGTT, and the downstream primer is TTCCGATTCCRATAAATTACC. The upstream primer for chr1:111217322-111217472 is TTTATTAAGYGTGTGGGTAT, and the downstream primer is AAAACGACCTAACCTCAC. The upstream primer for chr12:6665281-6665431 is TTTAGGCGGTGGTTTT, and the downstream primer is ACAAAACCRACCAATC. chr17:75369561-75369711 Upstream primer: TCGTTGTTTATTAGTTATTATGTCGGATTT, Downstream primer: CCCCCGTCCCGCGCCA; The upstream primer for chr17:75369627-75369777 is AGCGGGAGGGCGTTTG, and the downstream primer is CTTCGAAAATAAATACTAAACTAACTACTA. The upstream primer for chr3:157821497-157821647 is AGTTGTTGGTTGTATGAAAAG, and the downstream primer is CCGCTCTTTCTATTCTCTCTT. The upstream primer for chr7:93519341-93519491 is TATATTTGGGAGGTTTG, and the downstream primer is CCCAAACTAAAACTTCCTAT. The upstream primer for chr8:57358422-57358572 is AGGGATTTCGTTTTGCGT, and the downstream primer is CGCAATCCTAACTACATTC. The upstream primer for chr2:63281133-63281283 is GGTTTAYGTGGTTATTAATG, and the downstream primer is TAAAAATATCAAAATAACGAATC. The upstream primer for chr1:44883337-44883487 is GAAGTGATTYGGGTTGT, and the downstream primer is AAAACCCACCCCGTAAC. The upstream primer for the internal reference region chr7:5571788-5571939 is TAAGGTTAGGGATAGGATAGTT, and the downstream primer is TAACCACCACCCAACAC.
6. A kit for detecting early esophageal cancer, characterized in that, The kit includes a detection reagent for detecting the degree of methylation in each target region of the biomarker combination according to any one of claims 1-5.
7. The kit for detecting early esophageal cancer according to claim 6, characterized in that, The detection reagents include those for detecting the degree of methylation in each target region of the methylation biomarker combination using pyrosequencing, bisulfite conversion sequencing, methylation array, qPCR, digital PCR, next-generation sequencing, third-generation sequencing, whole-genome methylation sequencing, DNA enrichment detection, simplified bisulfite sequencing, HPLC, MassArray, methylation-specific PCR, or combinations thereof.
8. Use of the biomarker combination of any one of claims 1-5 in the preparation of a kit for detecting early esophageal cancer.
9. An early esophageal cancer detection system, characterized in that, The detection system includes a calculation module, which calculates the predicted value of the methylation degree of the marker region in the marker combination according to any one of claims 1-5; The detection model predicts values using a logistic regression model. The average methylation rate of each biomarker region is multiplied by its corresponding weight coefficient and then summed. If the result is greater than a threshold, it is considered a positive risk indicator for esophageal cancer.