Methylation marker for predicting nasopharynx cancer risk and screening method and application thereof
Patent Information
- Application Number
- CN202411626700.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2026-05-15
AI Technical Summary
Existing diagnostic methods for nasopharyngeal carcinoma have insufficient sensitivity, especially early screening methods such as fiberoptic nasopharyngoscopy and imaging techniques, which are difficult to detect small tumors. Furthermore, the sensitivity of EBV testing is unstable, leading to overdiagnosis and overtreatment.
By combining EB virus nucleic acid detection and methylation detection, a machine learning model was constructed by screening differentially methylated regions. Methylation biomarkers were used to predict the risk of nasopharyngeal carcinoma in blood. Methylation analysis was performed using methylation-specific PCR, MeDIP-Seq, and MBD-Seq technologies, and an early screening model was constructed by combining machine learning algorithms.
It improves the sensitivity and specificity of early screening for nasopharyngeal carcinoma, enabling earlier detection of nasopharyngeal carcinoma risk, reducing overdiagnosis, and meeting the needs of early screening.
Smart Images

Figure CN122038563A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of early cancer screening technology, specifically, it relates to a methylation biomarker for predicting the risk of nasopharyngeal carcinoma, its screening method and application. Background Technology
[0002] Nasopharyngeal carcinoma has the highest incidence rate among malignant tumors of the ear, nose, and throat, and is one of the most common malignant tumors. Nasopharyngeal carcinoma grows in the nasopharynx at the back of the nasal cavity, a relatively hidden location, and often has no obvious symptoms in its early stages, making it easily overlooked. Without early screening, most cases are discovered at an advanced stage, requiring conventional treatments such as chemotherapy, radiation therapy, or invasive procedures, with a five-year survival rate of only 10-15%. However, by implementing regular early screening, adopting cancer-suppressing measures based on early screening risk, and combining these with gene therapy, molecular biological therapy, immunotherapy, lifestyle modifications such as physical exercise and dietary adjustments, the five-year survival rate can reach 95-100%. Effective screening can prevent the occurrence of nasopharyngeal carcinoma at an early stage.
[0003] Nasopharyngeal carcinoma (NPC) is a multifactorial disease. It is widely believed that NPC results from the long-term interaction and evolution of factors such as host cell infection by Epstein-Barr virus (EBV) and the tumor microenvironment. Therefore, EBV-related biomarkers are frequently used in the clinical auxiliary diagnosis of NPC. EBV is a human gamma herpesvirus with a double-stranded linear DNA structure, approximately 172 kb in length. It is mainly transmitted through saliva and exists latently in human B lymphocytes and nasopharyngeal epithelial cells, closely associated with NPC and other diseases. EBV-DNA in plasma may be released during apoptosis of NPC cells or produced by viral replication. Therefore, EBV-DNA can be used for the early diagnosis and screening of NPC. Researchers screened asymptomatic patients with plasma EBV-DNA and retested patients who tested positive the first time four weeks later. Of the 309 subjects who remained positive, 34 were pathologically diagnosed with NPC. Currently, the sensitivity of EBV testing ranges from 53% to 94%, which may lead to overdiagnosis and overtreatment.
[0004] Current diagnostic methods still have significant limitations. For example, fiberoptic nasopharyngoscopy and pathological examination are invasive, while imaging techniques can only detect tumors larger than 0.5 cm. There is an urgent need for accessible and high-performance early screening products to advance early screening for nasopharyngeal carcinoma.
[0005] ctDNA methylation is an important epigenetic modification that can alter genetic expression without changing the gene sequence, thereby controlling gene expression. Gene methylation occurs before cancer develops and is a crucial mechanism in cancer development, acting as a "switch" that regulates gene expression. Its stability and consistency are relatively high, making it an ideal method for early cancer screening.
[0006] Currently, numerous studies have identified specific methylation genes in various cancer types, such as colorectal cancer, gastric cancer, and liver cancer, which can be detected in plasma to indicate early-stage development. However, there are very few reports on clinically applicable DNA methylation markers related to nasopharyngeal carcinoma. Summary of the Invention
[0007] To address at least one of the aforementioned technical problems, the inventors made significant efforts to establish a system for screening methylation biomarkers and constructing machine learning models to detect early nasopharyngeal carcinoma using a combined EB nucleic acid detection and methylation detection system, thereby completing this application.
[0008] The first aspect of this application provides a method for screening methylation biomarkers for predicting the risk of nasopharyngeal carcinoma, comprising the following steps: Biological samples were obtained from multiple subjects who were positive for Epstein-Barr virus infection, including nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects; Methylation sequencing was performed on biological samples from the subjects who tested positive for EB virus infection to obtain methylation sequencing data; Based on methylation sequencing data, differentially methylated regions between nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects were obtained, and the methylation expression levels of the differentially methylated regions were obtained. The biological samples were randomly divided into training and testing sets, and a methylation expression level matrix was constructed. The differentially methylated regions were ranked by importance using a feature selection tool. At least one of the top N differentially methylated regions was selected as a characteristic methylated region as a methylation region biomarker for predicting the risk of nasopharyngeal carcinoma. The gene corresponding to the methylation region biomarker is a methylation gene biomarker, where N = 1 to 100.
[0009] In this application, "predicting nasopharyngeal carcinoma risk" refers to predicting whether a subject has a risk of developing nasopharyngeal carcinoma, that is, differentiating between precancerous and asymptomatic cancers in subjects who have not yet shown abnormal symptoms, thereby completing early screening for nasopharyngeal carcinoma. Therefore, in this application, "predicting nasopharyngeal carcinoma risk," "early screening for nasopharyngeal carcinoma," and "nasopharyngeal carcinoma screening" are used interchangeably and represent the same meaning.
[0010] In this application, the non-nasopharyngeal carcinoma subject can be a normal individual or a healthy individual (i.e., an individual without any disease), or a patient with a benign disease.
[0011] In this application, the differentially methylated region (DMR) is obtained by screening nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects using the DiffBind tool (version 3.8.4). Preferably, in some specific embodiments of this application, the differentially methylated regions are obtained using both the DESeq and EdgeR algorithms, and their intersection is taken to derive the differentially methylated region.
[0012] In some specific embodiments of this application, the specific steps for screening differentially methylated regions are as follows: (1) The edgeR and DESeq2 algorithms of the DiffBind tool were used to find the differentially methylated regions shared by blood samples from nasopharyngeal carcinoma patients and blood samples from non-nasopharyngeal carcinoma subjects and / or human genome (Hg38); (2) Use the edgeR and DESeq2 algorithms of the DiffBind tool to find the differentially methylated regions shared by nasopharyngeal carcinoma tissue samples compared with blood samples and / or human genomes (Hg38) of non-nasopharyngeal carcinoma subjects; (3) Use the intersect function in the bedtools program to find the common differentially methylated regions in blood and tissue samples.
[0013] In this application, the methylated gene refers to the gene corresponding to the differentially methylated region.
[0014] In some embodiments of this application, the plurality of EB virus infection-positive subjects are detected using an EB virus nucleic acid detection method.
[0015] In some specific embodiments of this application, an EBV nucleic acid detection kit is used to detect EBV-DNA in plasma to screen for samples that are positive for EBV infection. The testing principle of the screening method is as follows: Polymerase chain reaction (PCR) combined with TaqMan fluorescent probe technology was used to quantitatively detect Epstein-Barr virus (EBV) nucleic acid DNA in whole blood samples. The target gene was a fragment of the EBV antinuclear antigen precursor protein gene (EBNA-LP), which is a highly specific conserved region. Primer-probe combinations meeting the detection requirements were screened after comparison with the NCBI database and experimental data. The probes were labeled with 5'-FAM / 3'-BHQ1. Simultaneously, primers and probes were designed based on the conserved region of the human β-actin gene, labeled with 5'-HEX / 3'-BHQ1, serving as internal standards in the kit to achieve end-to-end monitoring from sample quality and sample extraction to nucleic acid amplification.
[0016] The reaction system consists of two tubes: a PCR reaction solution (primers, probes, dNTPs, magnesium ions, etc.) and a mixed enzyme solution (DNA polymerase and UNG enzyme). Quality control samples include strong positive, weak positive, and negative controls, prepared from pseudoviruses containing the EBV gene and pseudoviruses containing the β-actin gene. Strong positive and weak positive controls should exhibit both FAM and HEX signal amplification, while negative controls should exhibit HEX signal amplification. Four known concentrations of quantitative standards are also provided to calibrate the EBV concentration in the samples. During PCR amplification, the probe binds to the target sequence, and the 5' fluorescent reporter group of the probe is cleaved by the exonuclease activity of Taq polymerase, moving away from the quencher group and generating a fluorescent signal. The quantitative real-time PCR instrument detects the fluorescence signal at the end of each annealing extension phase, automatically plots a real-time amplification curve, and calculates the sample Ct value (the number of cycles required for the fluorescence signal in each reaction tube to reach a set threshold). Four quantitative standards were set as standards and their corresponding concentrations were entered. The PCR software automatically plotted a standard curve of the amplification Ct of the quantitative standards versus the Log10 value of the concentration. Based on the amplified Ct value of the EB virus gene in the sample, the EB concentration in the sample was calculated. A uracil-N-glycosylation enzyme (UNG enzyme) / dUTP anti-contamination system was introduced into the PCR detection system to avoid false positive results. The most common and important contaminant in PCR detection is the PCR product. By replacing dTTP with dUTP / dTTP, the PCR product is a DNA strand containing dU. During the initial PCR incubation step at 42℃, the UNG enzyme degrades the uracil bases in the existing U-DNA contaminant in the reaction system. The UNG enzyme is then inactivated under the subsequent denaturation condition of 94℃ for 2 minutes, preventing further degradation of newly amplified U-DNA products and thus ensuring the specificity and accuracy of the amplification results.
[0017] DNA methylation is the addition of a methyl group at the 5-position of the cytosine ring in DNA by DNA methyltransferases, typically leading to gene silencing. Interestingly, aberrant circulating tumor DNA (ctDNA) methylation can be detected in various cancers, and these ctDNAs undergo more frequent covalent modifications, often showing numerous changes in the early stages of carcinogenesis. Therefore, these characteristics and ease of detection allow methylated ctDNA to be used as a marker for early screening. ctDNA specifically refers to DNA fragments released into the circulation system from tumor cell somatic DNA shed or after apoptosis; it is part of cell-free DNA (cfDNA) and has good tumor specificity. ctDNA detection is performed using peripheral blood samples, hence the term "liquid biopsy." The main advantages of liquid biopsy are: first, convenient sample collection, avoiding complex and invasive tumor tissue biopsies; second, peripheral blood is a strong buffering system, eliminating the tumor heterogeneity issues associated with local tissue biopsy-related tests. Recent studies have shown that it plays an important role in early tumor diagnosis, treatment monitoring, and prognosis.
[0018] In the human genome, approximately 3%–6% of cytosine (C) binds to methyl groups under the action of DNA methyltransferases, forming a covalent bond at the C' 5-carbon position of CpG disodium nucleotides. CpG disodium nucleotides in the human genome often exist in dense clusters of 300–3000 bp (CpG islands), and these CpG islands are usually located near the transcription start sites of genes, playing a role in regulating gene expression. When DNA is methylated, the binding of transcription factors to DNA is inhibited, genes are silenced, and the corresponding protein levels decrease. Cancer cells, which differ significantly from normal cells, typically have highly methylated promoter regions of tumor suppressor genes, effectively shutting them down and allowing cancer cells to grow unchecked.
[0019] Methylation enrichment is an analytical method used to study methylation modifications on DNA. DNA methylation is an important epigenetic modification involving the addition of methyl groups to the cytosine ring in the DNA molecule. This modification plays a crucial role in biological processes such as gene expression, cell differentiation, and genome stability. Therefore, understanding the state of DNA methylation is essential for comprehending biological processes and the development of diseases.
[0020] Common methylation enrichment techniques include methylation-specific PCR (MSP), methylation-sensitive restriction enzyme digestion, MeDIP-Seq (Methylated DNA Immunoprecipitation Sequencing), and MBD-Seq (Methyl-CpG Binding Domain Sequencing), among others. MSP uses methylation-specific primers to selectively amplify methylated DNA fragments via PCR. It is simple, rapid, and applicable to specific CpG sites, but it cannot provide genome-wide methylation information and is only suitable for pre-defined target regions.
[0021] Methylation-sensitive restriction enzyme cleavage utilizes the difference in sensitivity of restriction enzymes to DNA sequences to distinguish between methylated and unmethylated DNA regions. This method is based on the principle that DNA methylation affects the sensitivity of bases on the cytosine ring to restriction enzymes. This technique does not require expensive sequencing technology and can be analyzed using methods such as gel electrophoresis. However, it cannot provide high-resolution information on individual CpG sites, typically providing the methylation status of the entire region. Furthermore, it is limited by the specificity of the selected restriction enzyme, which may result in missed or overdetected methylation sites. Additionally, it cannot directly distinguish between 5-methylcytosine and other forms of DNA modification.
[0022] MeDIP-Seq uses methylated DNA antibodies to selectively enrich methylated DNA fragments, which are then analyzed using high-throughput sequencing. It can enrich entire methylated genomic regions, making it suitable for whole-genome methylation analysis. However, it cannot provide high-resolution information on individual CpG sites.
[0023] MBD-Seq uses methylated DNA-binding proteins (such as MBD2 or MBD3) to enrich methylated DNA fragments, which are then analyzed by sequencing. It offers high enrichment efficiency and is suitable for genome-wide methylation analysis. However, similar to MeDIP-Seq, it cannot provide high-resolution information on individual CpG sites.
[0024] In this application, the methylation process is also referred to as methylation transformation. Common methylation-processing sequencing technologies include bisulfite sequencing (BS-seq). BS-seq uses bisulfite to treat DNA, converting unmethylated cytosine to uracil, while methylated cytosine remains unaffected, and then analyzes the results through sequencing. It can provide high-resolution information for individual CpG sites and can perform methylation analysis of the entire genome. The bisulfite treatment process includes denaturation, deamination, and desulfonation. The DNA is first denatured into single strands, and then subjected to extreme temperatures, high salt, acidity, and alkalinity. The resulting transformed DNA is predominantly single-stranded, with a mixture of double strands, fragment nicks, gap damage, and uracil-state nucleotides. This process generally leads to the loss of 90% of the DNA template, and a large amount of methylation information cannot be detected in subsequent processes. Furthermore, during base transformation, there is a possibility of incomplete or over-conversion, resulting in human bias. Subsequent PCR amplification further amplifies this bias, causing significant data waste and inaccurate signals. To reduce the bias in WGBS, more DNA needs to be used, the number of PCR cycles reduced, and the efficiency of PCR amplification enzymes optimized. These are the main reasons why WGBS sequencing is inefficient in screening ctDNA methylation markers. Currently, methylation markers obtained based on bisulfite treatment generally suffer from low sensitivity, especially in blood samples. The already limited number of free DNA fragments becomes significantly more difficult to detect methylation levels after bisulfite treatment.
[0025] In some embodiments of this application, methylated fragments are enriched using methylated DNA immunoprecipitation (MeDIP) technology. The methylated DNA antibody is selected from one of 5-methylcytidine antibody, 5-methylcytosine (5-mC) antibody, 5-hydroxymethylcytosine (5-hmC) antibody, 5-formylcytosine (5-fC) antibody, and 5-carboxycytosine (5-caC) antibody.
[0026] In some embodiments of this application, the methylation expression level of the differentially methylated region is its relative methylation level (RML), and correspondingly, the resulting methylation expression level matrix is an RML matrix. Specifically, for a differentially methylated region, the RML is calculated as follows:
[0027] In some embodiments of this application, the inventors designed a hybrid capture screening probe panel of approximately 0.9 Mb size covering early screening of high-incidence tumors, including nasopharyngeal carcinoma, covering over 37,000 CpG sites on the human genome. It covers all possible candidate CpG sites for screening all target gene regions in nasopharyngeal carcinoma, including promoter regions, the first exon region, 1 kb upstream of the start codon, gene introns, and regions between adjacent genes.
[0028] To screen for methylation markers specific to nasopharyngeal carcinoma, a multicenter sample library for screening characteristic methylation profiles of nasopharyngeal carcinoma was constructed clinically. Blood and tissue samples were collected from nasopharyngeal carcinoma patients, and non-nasopharyngeal carcinoma subjects (including normal healthy individuals and patients with benign diseases) were used as controls.
[0029] The experimental steps for processing the collected clinical samples can be found below: (1) Extracting DNA from human plasma / tissue; (2) The DNA is repaired at the end, an "A" tail is added, and it is ligated to the adapters to obtain the ligation product; (3) After binding 5-methylcytosine antibody to magnetic beads, perform antigen-antibody reaction with the conjugation product of one or more samples; (4) The enriched methylated fragments were purified and amplified by PCR to obtain an amplified methylated DNA enriched library. (5) Hybridize and capture multiple methylated DNA enrichment libraries to obtain the target methylated DNA fragment; (6) The target fragment of the probe hybridization is purified and then amplified by PCR to obtain a hybridized amplified methylated DNA library. (7) After hybridization, the methylated DNA library was amplified, pooled, and prepared before sequencing to obtain a high-throughput sequencing library; (8) The aforementioned high-throughput sequencing library was sequenced using the Illumina next-generation sequencing platform to analyze the methylation level of genes related to the occurrence and development of nasopharyngeal carcinoma. (9) Data statistical analysis: quality control of the reads after sequencing, sequence alignment analysis with human gene sequence / target probe coverage area, peak detection of methylation enrichment region, and preliminary screening of differential DNA methylation regions in nasopharyngeal carcinoma.
[0030] (10) Use the DiffBind tool (version 3.8.4) to perform differential peak screening for nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects. Use the DESeq and EdgeR algorithms to further prioritize the screening of regions that intersect and are within the panel.
[0031] (11) Further intersect the differential DNA methylation regions of nasopharyngeal carcinoma selected from blood and tissue samples to obtain a reliable combination of nasopharyngeal carcinoma markers.
[0032] In one specific embodiment of this application, the above method is used to screen for a combination of methylation region markers for early screening of nasopharyngeal carcinoma, containing at least 56 methylation regions and 930 CpG characteristic sites.
[0033] A second aspect of this application provides the use of a reagent for detecting the methylation level of a combination of methylation region markers in the preparation of a kit for predicting the risk of nasopharyngeal carcinoma, wherein the combination of methylation region markers includes at least one of the following: chr9:87498973-87499172, chr2:181457944-181458143, chr11:43581352-43581551, chr11:443040 60-44304259, chr14:60509500-60509699, chr14:69547960-69548159, chr12:48198223-48198422, chr1:170661243-170661442, chr8:120124836-120125035 and chr6:77463420-77463619.
[0034] In some embodiments of this application, the combination of methylation region markers further includes at least one of the following: chr1:58249628-58249827, chr22:42283819-42284018, chr1:229407770-229407969, chr21:33025905-33026104, chr12:24903154-24903353, chr12:4272683-4272882, chr1:240091583-240091782, chr2:161423989-161424188, chr4:1002218-1002417, and chr20:43915894-43916093.
[0035] In some embodiments of this application, the methylation region marker combination further includes at least one of the following: chr1:153679627-153679826, chr6:10426407-10426606, chr4:121380465-121380664, chr2:72917151-72917350, chr9:137156751-137156950, chr5:50969615-50969814, chr13:107868359-107868558, chr7:93890509-938907 08. chr21:33071882-33072081, chr19:41135567-41135766, chr4:140426905-140427104, chr6:137498188-137498387, chr17:486465 27-48646726、chr4:107931901-107932100、chr1:44417990-44418189、chr1:156923961-156924160、chr12:62632009-62632208、chr10 :104640217-104640416, chr12:102958605-102958804, chr7:87628001-87628200, chr12:8018782-8018981, chr5:80960731-8096093 0. chr17:44325178-44325377, chr7:49775977-49776176, chr16:23754731-23754930, chr12:21527708-21527907, chr2:184598918-18 4599117, chr12:75207510-75207709, chr8:37965734-37965933, chr4:110622983-110623182, chr3:194487549-194487748, chr19:58440233-58440432, chr21:25640018-25640217, chr5:100903323-100903522, chr1:228458190-228458389 and chr21:45973989-45974188.
[0036] A third aspect of this application provides the use of a reagent for detecting the methylation level of a combination of methylation gene markers in the preparation of a kit for predicting the risk of nasopharyngeal carcinoma, wherein the combination of methylation gene markers includes at least one of DAPK1, ITGA4, ALX4, and MIR129-2.
[0037] In some embodiments of this application, the combination of methylation gene markers further includes at least one of the following: SIX6, CCDC177, OR10AD1, PRRX1, COL14A1, HTR1B, DAB1, TCF20, ACTA1, OLIG2, BCAT1, CCND2, FMN2, TBR1, IDUA, TOX2, NPR1, TFAP2A, QRFPR, EMX1, MIR3621, LINC02106, FAM155A, TFPI2, OLIG1, CYP2 F1, CLGN, OLIG3, MIR196A1, LOC101929595, RNF220, LRRC71, MIRLET7I, SORCS3, PAH, RUNDC3B, FOXJ2, RASGRF2-AS1, SLC 25A39, VWC2, CHP2, SPX, ZNF804A, KCNC2, ADRB3, PITX2, LINC00884, ZNF132, JAM2, ST8SIA4, HIST3H2BB and LOC101928796.
[0038] A fourth aspect of this application provides a kit for predicting the risk of nasopharyngeal carcinoma, comprising: (1) A reagent for detecting the methylation level of a combination of methylation region markers or a reagent for detecting the methylation level of a combination of methylation gene markers, wherein the combination of methylation region markers includes at least one of the following: chr9:87498973-87499172, chr2:181457944-181458143, chr11:43581352-43581551, chr11:44304060-44304259, chr14:6050950 The methylation gene marker combination includes at least one of DAPK1, ITGA4, ALX4, and MIR129-2. The methylation gene marker combinations are: 0-60509699, chr14:69547960-69548159, chr12:48198223-48198422, chr1:170661243-170661442, chr8:120124836-120125035, and chr6:77463420-77463619. (2) EB virus nucleic acid detection reagent, wherein the methylation region marker combination methylation level detection reagent or the methylation gene marker combination methylation level detection reagent is used to detect subjects who are positive for EB virus nucleic acid.
[0039] The fifth aspect of this application provides a method for constructing a model for predicting the risk of nasopharyngeal carcinoma, comprising the following steps: The methylation level data of methylation region marker combinations or methylation gene marker combinations in the population biological samples were randomly divided into two groups: one group was the training set and the other group was the test set. The population biological samples were all positive for EB virus nucleic acid and included nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects. The model for predicting nasopharyngeal carcinoma risk is constructed using the training and test sets based on machine learning methods. The methylation region marker combination includes at least one of the following: chr9:87498973-87499172, chr2:181457944-181458143, chr11:43581352-43581551, chr11:44304060-44304259, chr14:60509500-60509699, chr14:6954 The methylation gene marker combination includes at least one of DAPK1, ITGA4, ALX4, and MIR129-2. The methylation gene marker combination includes 7960-69548159, chr12:48198223-48198422, chr1:170661243-170661442, chr8:120124836-120125035, and chr6:77463420-77463619.
[0040] In some embodiments of this application, the method for constructing the model includes the following steps: The importance of the combination of methylation region markers or methylation gene markers is evaluated using the random forest model algorithm, and markers that can be used to build the model are selected. In the training set and the test set, the model is constructed using the Extra Trees Classifier, K-Nearest Neighbors Classifier, etc., to build the nasopharyngeal carcinoma early screening model algorithm. Feature selection is performed on the model, the importance scores of candidate markers are calculated, and cumulative importance curves are constructed. The candidate models were scored using the Lazy predict package, and the Random Forest model ranked first with an initial AUC value of over 0.78. Further algorithm optimization will improve the model's sensitivity and specificity.
[0041] The sixth aspect of this application provides a system for predicting the risk of nasopharyngeal carcinoma, comprising the following modules: The data input module is used to input methylation level data of methylation region marker combinations or methylation gene marker combinations obtained from the subject's biological samples. The methylation region marker combinations include at least one of the following: chr9:87498973-87499172, chr2:181457944-181458143, chr11:43581352-43581551, chr11:44304060-44304259, chr14:60 509500-60509699, chr14:69547960-69548159, chr12:48198223-48198422, chr1:170661243-170661442, chr8:120124836-120125035 and chr6:77463420-77463619; the methylation gene marker combination includes at least one of DAPK1, ITGA4, ALX4 and MIR129-2; A database storage module is used to store methylation level data of the combination of methylation region markers or the combination of methylation gene markers in the population biological samples, wherein the population biological samples are all positive for EB virus nucleic acid and include nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects; The disease prediction module is connected to the data input module and the database storage module, respectively. It is used to construct a prediction model using the methylation level data of the combination of methylation region markers or the combination of methylation gene markers of the population biological sample, and to predict whether the subject has the risk of developing nasopharyngeal carcinoma based on the methylation level data of the combination of methylation region markers or the combination of methylation gene markers of the subject obtained from the data input module.
[0042] Beneficial effects of this application Compared with the prior art, this application has the following advantages: This application combines EBV detection with DNA methylation markers, enabling more sensitive detection of nasopharyngeal carcinoma-related methylation markers in human blood samples. Furthermore, by effectively detecting changes in their methylation levels, an early screening model can be constructed, which can improve the sensitivity and specificity of early nasopharyngeal carcinoma screening and meet the urgent need for early nasopharyngeal carcinoma screening.
[0043] The methylation markers for nasopharyngeal carcinoma provided in this application enable early prediction of nasopharyngeal carcinoma. Attached Figure Description
[0044] Figure 1 An IGV visualization of the ITGA4 gene in Example 2 of this application is shown, illustrating its difference in nasopharyngeal carcinoma samples compared to non-nasopharyngeal carcinoma subject samples.
[0045] Figure 2 An IGV visualization of the MIR129-2 gene in Example 2 of this application is shown, illustrating its difference in nasopharyngeal carcinoma samples compared to non-nasopharyngeal carcinoma subject samples.
[0046] Figure 3 An IGV visualization of the DAPK1 gene in Example 2 of this application is shown, illustrating its differences in nasopharyngeal carcinoma samples compared to non-nasopharyngeal carcinoma subject samples.
[0047] Figure 4 The diagram shows a scatter plot of the differential peaks in the training / test sets for the binary classification model constructed based on 56 differentially methylated regions in Embodiment 3 of this application.
[0048] Figure 5 The ROC curve of the binary classification model constructed based on 56 differentially methylated regions in Embodiment 3 of this application is shown in the training / test set.
[0049] Figure 6 The ROC curve of the binary classification model constructed based on 56 differentially methylated regions in Example 3 of this application is shown on the validation set.
[0050] Figure 7 This is a flowchart illustrating the detection of the EB virus nucleic acid by using qPCR technology based on the EB virus nucleic acid detection method in Example 1 of this application.
[0051] Figure 8 The ROC curves of the four genes (DAPK1, ITGA4, ALX4, and MIR129-2qPCR) in Example 6 of this application are shown individually in the training / test sets.
[0052] Figure 9 The ROC curves of the gene combinations in Example 6 of this application in the training / test set are shown. Detailed Implementation
[0053] Unless otherwise stated, implied from the context, or as is customary in the art, all parts and percentages in this application are based on weight, and all testing and characterization methods used are concurrent with the filing date of this application. Where applicable, any patent, patent application, or disclosure relating to this application is incorporated herein by reference in its entirety, and its equivalent patent families are also incorporated herein by reference, in particular the definitions of relevant terms in the art disclosed in such documents. If any definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition provided in this application shall prevail.
[0054] In the specification and claims of this application, the singular forms "an," "a," and "the" include their plural forms, unless the context otherwise requires. Thus, for example, "a reagent" can be understood to include multiple reagent components.
[0055] In the description and claims of this application, the terms "individual," "subject," or "patient" are used interchangeably and refer to a vertebrate, preferably a mammal. The mammal may be a human, a non-human primate, a mouse, a rat, a dog, a cat, a horse, or a cow, but is not limited to these examples.
[0056] According to embodiments of this application, the methylation region markers provided in this application refer to chromosomal loci or regions that can be used to detect or screen whether a subject has nasopharyngeal carcinoma. These can be nucleic acid sequences of a certain length, or nucleotides at one or two specific sites. The terms "methylation marker," "methylation region marker," and "characteristic methylated region" have the same meaning, referring to a methylation level that indicates whether a subject has nasopharyngeal carcinoma. Furthermore, it should be understood that the term "marker" should include both the sense strand sequence of a gene or region and the antisense strand sequence of the marker or gene.
[0057] In some embodiments of this application, the methylated regions or genes of this application also include various variants thereof. Variants include nucleic acid sequences from the same region that have at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with it (i.e., having one or more deletions, insertions, substitutions, reverse sequences, etc.). Therefore, the content of this application should be understood to extend to such variants that achieve the same result, although in fact the actual nucleic acid sequences between individuals have minor genetic variations. The term "marker" as used is broadly interpreted to include both 1) the original marker found in a biological sample or genomic DNA (in a specific methylation state) and 2) its processed sequence (e.g., the corresponding region after immunoprecipitation treatment).
[0058] The term "AUC" is an abbreviation for "Area Under the Curve." Specifically, it refers to the area under the receiver operating characteristic (ROC) curve. An ROC curve is a plot of the true positive rate versus the false positive rate for different possible cutoff points in a diagnostic test. It shows the balance between sensitivity and specificity based on the selected cutoff point (any increase in sensitivity will be accompanied by a decrease in specificity). The area under the ROC curve (AUC) is a measure of a diagnostic test (the larger the area, the better; optimally 1); randomized trials will have an ROC curve with an area of 0.5 on the diagonal.
[0059] To make the technical problems, technical solutions and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments.
[0060] The following examples are used to illustrate preferred embodiments of this application. Those skilled in the art will understand that the techniques disclosed in the examples represent technologies discovered by the inventors that can be used to implement this application, and therefore can be considered preferred embodiments of this application. However, those skilled in the art should understand from this specification that many modifications can be made to the specific embodiments disclosed herein, still yielding the same or similar results, without departing from the spirit or scope of this application.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains, and all materials cited herein and referenced by them are incorporated herein by reference.
[0062] Those skilled in the art will recognize, or can learn through routine experimentation, many equivalents of the specific embodiments of the invention described herein. These equivalents will be included in the claims.
[0063] In the embodiments of this application, the material sources and preprocessing methods are as follows: 1. Sample Collection Blood samples were collected from subjects with nasopharyngeal carcinoma and those with normal nasopharynx, according to different regions.
[0064] 2. Sample Preprocessing When preprocessing cfDNA in blood samples and DNA in tissue samples, techniques commonly used in the field can be employed for subsequent EB virus detection and methylation detection. For example, in the following embodiment of this application: 1 mL of whole blood from a 10 mL patient blood sample is first taken and lysed with erythrocyte lysis buffer. Nucleic acid is then released under the action of surfactant and heating. Simultaneously, impurities such as metal ions in the liquid are adsorbed using Chelex-100 resin. After centrifugation, purified nucleic acid is obtained. The remaining whole blood samples are centrifuged twice: first at 4°C and 600×g for 20 min, and then at 16000×g for 10 min to obtain plasma. 1–2 mL of plasma is taken and cell-free DNA is extracted using the VAHTS Free-Circulating DNA MaxiKit (Cat. No. N903-03, Vazyme) to obtain cfDNA. The concentration of cfDNA is quantified using Qubit 4.0. The cfDNA yield for each sample should not be less than 5 ng.
[0065] Unless otherwise specified, the other experimental methods in the following examples are conventional methods. Unless otherwise specified, the instruments and equipment used in the following examples are all conventional laboratory instruments and equipment; unless otherwise specified, the experimental materials used in the following examples were all purchased from conventional biochemical reagent stores.
[0066] Example 1: Screening of EB virus-positive samples by real-time fluorescent qPCR detection The inventors used an EB virus real-time fluorescence qPCR detection method (EB virus nucleic acid detection kit (fluorescent quantitative PCR method) kit, catalog number RP140001T1050, Jiangsu Mole Biotechnology Co., Ltd.) to screen EB-positive samples from 654 human blood samples. The screening method includes the following steps: (1) Prepare a 400 μL blood sample from the person to be tested; (2) DNA was extracted using nucleic acid extraction and purification reagents from Jiangsu Mole Biotechnology Co., Ltd. (registration number: Sutai Medical Device Registration 20170120); (3) Prepare quality control products and standards: Take 100 μL of each of the strong positive quality control product (EB), weak positive quality control product (EB), negative quality control product (EB) and 4 quantitative standards (EB) in the kit and mix them with 100 μL of lysis buffer. Heat at 95°C for 10 min, centrifuge at 12000 rpm for 10 min, transfer the supernatant to a new EP tube and label it. (4) Remove the PCR reaction solution (EB) from the kit and thaw it on ice or at (2~8)℃; remove the mixed enzyme solution (EB), gently shake all components to mix, and then briefly centrifuge at low speed. Prepare the PCR reaction system for each person according to the table below:
[0067] (5) After the PCR reaction system is prepared, dispense the PCR reaction system and add 20.0 μL of PCR reaction system to each PCR reaction well; (6) Add the nucleic acid prepared in step (3) to the PCR reaction well containing the PCR reaction system. The amount of nucleic acid added is 20.0 µL / well. Add 20.0 µL of sterile RNase-free water or physiological saline to the PCR negative control well.
[0068] (7) The reaction was performed using an ABI 7500 instrument. The reaction procedure is shown in the table below:
[0069] (8) Judgment of test results The test results must meet the following 6 requirements: ①The PCR negative control FAM fluorescence Ct value should be negative, and the HEX fluorescence Ct value should be >35.00 or there should be no typical amplification curve; ② The negative control sample should have a negative FAM fluorescence Ct value and a HEX fluorescence Ct value ≤ 35.00; ③ The EB virus nucleic acid DNA quantification result of the weakly positive quality control should be approximately 2 × 10⁻⁶. 2 ~9×10 4 IU / mL, Ct value should be between 28.00 and 33.00, HEX fluorescence Ct value ≤ 35.00; ④ The EB virus nucleic acid DNA quantification result of the strong positive control should be approximately 2 × 10⁻⁶. 5 ~9×10 7 IU / mL, Ct value should be between 18.00 and 23.00, HEX fluorescence Ct value ≤ 35.00; ⑤ The correlation coefficient |r| of the standard curve should be ≥0.9800; ⑥ All sample wells should show a typical amplification curve of the internal standard (HEX fluorescence).
[0070] The following results can only be judged after the above conditions are met: ① The instrument display "Undetermined" indicates that the sample to be tested is negative.
[0071] ② For a Ct value ≤37.92 and an internal standard HEX fluorescence Ct value ≤35.00, the result is reported as EB virus DNA positive.
[0072] ③ For a Ct value > 37.92 and a HEX fluorescence Ct value ≤ 35.00, report it as EB virus DNA below the detection limit of the kit.
[0073] ④ The results showed that the concentration of EB virus nucleic acid DNA in the sample was 1×10⁻⁶. 2 IU / mL ~1.0×10 7 When the concentration is IU / mL, the quantitative result is valid and the corresponding quantitative value can be reported directly.
[0074] ⑤ The results showed that the sample concentration was greater than 1.0 × 10⁵ 7 IU / mL can be directly reported as greater than 1.0 × 10⁻⁶. 7 The concentration is IU / mL. Alternatively, the nucleic acid sample can be diluted and retested, and the result can be calibrated according to the dilution factor.
[0075] ⑥ The results show a typical amplification curve. When the concentration of EB virus nucleic acid DNA is below 100 IU / mL, the detection results are valid. The quantitative values are for reference only.
[0076] ⑦ When the concentration of EB virus DNA in the sample to be tested is less than 50 IU / mL, it cannot be detected with 100% accuracy.
[0077] (9) Analysis of clinical sample test results The results of the EB virus real-time fluorescence qPCR detection of 654 human blood samples are shown in Table 1.
[0078] Table 1. Results of real-time fluorescent qPCR detection of EB virus in 654 samples.
[0079] As shown in Table 1, using the real-time fluorescent qPCR detection method for EB virus, 428 EB-positive samples were screened from 654 human blood samples, including 230 nasopharyngeal carcinoma patients and 198 non-nasopharyngeal carcinoma subjects (including normal healthy subjects and subjects with benign diseases).
[0080] Example 2: Screening for characteristic methylation regions specific to human nasopharyngeal carcinoma 1. Design and synthesis of a large panel for methylation-targeted capture probes By surveying the PubMed database, the TCGA public database, and the publicly available Illumina methylation 450K / 850K microarray database, and tracking the latest research literature reports and bioinformatics information mining, the inventors integrated methylation regions and sites closely related to the early occurrence, development, metastasis, and tissue-specific tracing of cancer. They screened methylation indicators with high sensitivity, specificity, and tracing accuracy in relevant cancers, especially those previously reported in cfDNA sample types. Simultaneously, through whole-genome methylation sequencing (based on methylation immunoprecipitation technology) of clinical samples from different cancer types, the inventors obtained relevant differentially characterized regions, namely the 5' UTR and promoter region information from the Homer database. This information was combined with the aforementioned bioinformatics analysis and annotation results to integrate the human genome hg38 gene annotation into a local database.
[0081] The longest transcript in the local database was selected to confirm gene location. For each confirmed gene, sequence information was obtained from its promoter region, 5' UTR, first exon 1, and the 1kb upstream region of the start codon. Most cancer-related methylation information is distributed within the promoter region, 5' UTR, first exon 1, and the 1kb upstream region of the start codon. The target region covers not only methylation functional genes but also non-coding RNA regions and intergenic regions.
[0082] The personalized probe panel from Nanoda (Nanjing) Biotechnology Co., Ltd. is approximately 0.9Mb in size and covers more than 37,000 CpG sites on the human genome.
[0083] 2. Methylation capture sequencing 2.1 Sample Preparation All samples were selected from those that tested positive for EB virus nucleic acid in Example 1. These included blood samples from 230 nasopharyngeal carcinoma subjects and blood samples from 198 non-nasopharyngeal carcinoma subjects (including normal healthy individuals or those with benign diseases), totaling 428 samples.
[0084] Blood samples were collected using a 10mL cell-free DNA sample preservation tube connected to a blood collection needle. Upon arrival at the laboratory, each blood sample was centrifuged twice: first at 4℃ and 600×g for 20 minutes, and then at 16,000×g for 10 minutes to obtain plasma samples.
[0085] 2.2 Extraction of human cfDNA Cell-free DNA (cfDNA) was extracted from 1–2 mL plasma samples using the VAHTS Free-Circulating DNA Maxi Kit (Cat. No. N903-03, Vazyme). The cfDNA concentration was quantified using Qubit 4.0. The cfDNA yield for each sample was no less than 5 ng.
[0086] 2.3 Extraction of DNA from Human Tissues (1) Use a disposable blade to scrape the tissue from the slide; if it is a paraffin strip, use tweezers to pick up 2-3 strips and place them in a centrifuge tube; if it is a paraffin block, use a blade to rotate and cut the lesion tissue, remove the paraffin part, and cut it into small pieces. When there is too much paraffin tissue, the amount of FFPE slides can be reduced appropriately, or the amount of lysis buffer and other reagents can be increased proportionally; (2) Use self-prepared tissue extraction reagents or tissue genomic DNA extraction kits (DP304) from Tiangen Biotech (Beijing) Co., Ltd., etc.; (3) DNA concentration was quantified using Qubit 4.0. The DNA yield for each sample was no less than 200 ng; (4) Take 600~1000ng of DNA, add 1× TE buffer to 60μL, and use a Covaris M220 sonicator to break up the DNA. Set the parameters as follows: Average Incident Power (watt): 10; Peak Incident Power (Watt): 50; Duty Factor (percent): 20; Cycles / Burst (count): 200; Duration (seconds): 180. (5) Quality control: Take 1-2 μl of nucleic acid extraction product and perform capillary electrophoresis using the Qsep100 fully automated nucleic acid and protein analyzer (Houzhe Biotechnology) to analyze the fragment size distribution. For samples composed of more than 50% genomic DNA fragments larger than 1kb, due to poor data quality, no further analysis is performed. For samples with low genomic DNA contamination, VAHTS DNAClean Beads (Cat. No. N411-03, Vazyme) are used for fragment screening to obtain short DNA fragments of suitable size for subsequent library preparation.
[0087] 2.4 Preparation of Human cfDNA Plain Library Using the VAHTS Universal Pro DNA Library Prep Kit for illumina (Cat. No. ND608-02, Vazyme), follow the instructions in the manual to perform end repair, add an "A" tail, and ligate the cfDNA obtained in the aforementioned nucleic acid extraction steps to adapters to obtain ligation products; preferably, adapters with a unique UMI molecular identifier are used for the ligation reaction.
[0088] 2.5 MeDIP (MeMethylated DNA Immunoprecipitation) Equal amounts of EB-positive and EB-negative sample libraries were added to the same immunoprecipitation reaction, such as 12 nasopharyngeal carcinoma-positive libraries and 12 non-nasopharyngeal carcinoma-negative libraries.
[0089] When using the Methylated DNA Enrichment Kit (Cat. No. C02010021, diagenode) for methylation fragment antibody enrichment, follow the instructions for use. During implementation, some steps in the instructions were optimized and improved, significantly increasing methylation enrichment efficiency. For example, the 95°C denaturation time for the sample DNA library was reduced to 10 minutes, and a washing step was added to the cleaning of the magnetic beads used in the antibody reaction. Throughout all steps after denaturation, the sample DNA was kept at a low temperature to ensure it remained single-stranded. Repeated freeze-thaw cycles were avoided for the methylation antibodies. The immunoprecipitation sample containing the methylation antibody mixture and the cleaned magnetic beads were placed in a Ferris wheel rotator and rotated at low speed at 4°C for 4–17 hours. After the immunoprecipitation reaction, fragment purification and recovery were performed.
[0090] 2.6 Prehybridization-capture PCR amplification and purification The purified and recovered products were subjected to PCR amplification using amplification reagents from the VAHTS Universal Pro DNA Library Prep Kit for Illumina (Cat. No. ND608-02, Vazyme). After amplification, the products were purified using an equal volume of VAHTS DNA Clean Beads (Cat. No. N411-03, Vazyme) to obtain a relatively pure MeDIP library; the library concentration was quantified using Qubit4.0. 1–2 μl of the purified library product was subjected to capillary electrophoresis using a Qsep100 fully automated nucleic acid and protein analyzer (from Houze Biotechnology) to analyze fragment size distribution and perform quality control.
[0091] 2.7 Liquid-phase hybridization capture Liquid-phase hybridization capture was performed using NadPrep hybridization capture reagent (Cat. No. REF1005101, NadPrep). Hybridization capture reactions can be single-hybrid or multi-hybrid, with the total amount of MeDIP amplified library added for each reaction ranging from 300 ng to 8 μg. In Example 2, 500 ng of the purified library (if less than 500 ng, add the entire amount) was added, along with Human Cot DNA and Nad Prep® Nano Blockers. The mixture was dried in a vacuum concentrator preheated to 42°C at 1000 rpm. After drying, the prepared hybridization reaction solution (containing the probe panel) was added, followed by vortexing and brief centrifugation. Hybridization was performed at the following program: 95°C / 30 sec; 65°C / Hold (100°C hot-lid) for 4-16 hours. Then, the washed streptavidin magnetic beads were added to the hybridization system and incubated for 40 min, vortexing every 10 min to ensure complete resuspension of the magnetic beads.
[0092] After the hybridization capture reaction is complete, wash the bound magnetic beads with the four washing solutions provided in the kit, discarding any residual solution at each step; finally, add 20 μL of nuclease-free water and gently vortex to mix.
[0093] 2.8 PCR amplification and purification after hybridization capture The hybridization capture product was amplified by PCR using amplification reagents from the VAHTS Universal Pro DNA Library Prep Kit for Illumina (Cat. No. ND608-02, Vazyme). After amplification, the product was purified using an equal volume of VAHTS DNA Clean Beads (Cat. No. N411-03, Vazyme) to obtain a relatively pure hybridization capture library. Library concentration was quantified using Qubit 4.0, and fragment size was determined using a Qsep100 fully automated nucleic acid and protein analyzer.
[0094] 2.9 Pooling Dilute the library to be used to 4 nM, mix, take out 5 μL of the library, add 5 μL of 0.2 N NaOH, mix by pipetting, and denature for 5 min. Immediately after denaturation, add 990 μL of HT1 Buffer (REF: 15058251, Illumina), vortex to mix, take out 105 μL and add 1295 μL of HT1 Buffer, vortex to mix, and the resulting library is 1.5 pM.
[0095] 2.10 Sequencing The sequencer used was an Illumina NextSeq 550Dx. Reagents used included High Output Reagent Cartridge v2 (REF:15057929, Illumina) (300 cycles), High Output Flow Cell Cartridge v2.5 (REF:20022408, Illumina), and Buffer Cartridge v2 (REF:15057941, Illumina). 1300 μL of the library was added to the sample site of the High Output Reagent Cartridge v2, and each reagent was added sequentially. Sequencing could then begin. This example used paired-end sequencing, with a total sequencing time of approximately 30 hours.
[0096] 2.11 Quality Control of Sequencing Data Fastp (version 0.22.0) was used to perform quality control on the sequencing data, removing low-quality bases. The overall Q20 of the clean data was above 90%, Q30 was above 85%, and the average sequencing depth was around 300.
[0097] 3. Screening for characteristic methylation regions The methylation sequencing results of the test samples obtained in Example 1 were used to perform differential methylation region sorting and sequence alignment analysis with human gene sequences / target probe coverage areas according to the above steps. Peak detection of methylation-enriched regions was also performed to initially screen for characteristic differential methylation regions of nasopharyngeal carcinoma. The DiffBind tool (version 3.8.4) was used to screen for differential peaks between nasopharyngeal carcinoma and non-nasopharyngeal carcinoma (normal or benign) subjects. Both DESeq and EdgeR algorithms were used, and regions within the intersecting panel were further prioritized for screening. The characteristic differential methylation regions of nasopharyngeal carcinoma screened from blood and tissue samples were further intersected to obtain a combination of nasopharyngeal carcinoma-specific methylation region markers.
[0098] The specific steps are as follows: (1) Use the edgeR and DESeq2 algorithms of the DiffBind tool to find the common peak regions in nasopharyngeal carcinoma blood samples (cfDNA) compared with blood samples from non-nasopharyngeal carcinoma subjects or human genome (Hg38), namely differentially methylated regions (DMR). (2) Use the edgeR and DESeq2 algorithms of the DiffBind tool to find the common peak regions in nasopharyngeal carcinoma tissue samples compared with blood samples or human genomes (Hg38) of non-nasopharyngeal carcinoma subjects; (3) Use the intersect function in the bedtools program to find the common differentially methylated regions in blood and tissue samples; (4) Calculate a certain value for each sample The relative methylation levels (RML) are expressed by the following formula:
[0099] (5) Based on the obtained RML values, the test samples are grouped and labeled: the nasopharyngeal carcinoma group is labeled as 1, and the non-nasopharyngeal carcinoma (including normal health and benign diseases) group is labeled as 0. The training set and test set are randomly grouped in a ratio of 8:2 to construct the RML matrix.
[0100] (6) Use Feature Select to score and rank the differentially methylated regions based on their importance, construct an importance cumulative curve, and screen out several significantly different differentially methylated regions as methylation region markers. The corresponding methylated genes are methylation gene markers.
[0101] In this embodiment, immunoprecipitation technology and a tumor early screening hybrid capture panel platform were used to mine methylation region markers of nasopharyngeal carcinoma, and 56 differentially methylated regions were selected. Their coordinates in the Human Genome Database Hg38 are listed in Table 2.
[0102] Table 2 Information on 56 differentially methylated regions and gene names, CpG site information
[0103] The IGV visualization results for the ITGA4, MIR129-2, and DAPK1 genes are as follows: Figures 1-3 As shown.
[0104] Example 3: Establishing a machine learning model for early screening of nasopharyngeal carcinoma based on 56 differentially methylated regions. 1. Model Building The sci-learn random forest model algorithm was used to evaluate the importance of features, and 56 differentially methylated regions selected in Example 2 were chosen as markers for model construction. Each marker had a different feature importance index in the algorithm construction.
[0105] Another 110 plasma cfDNA samples that tested positive for EB virus were selected, including 79 cases of nasopharyngeal carcinoma and 31 cases of normal or benign lesions. The training set and test set were split (8:2) to construct a binary classification model.
[0106] The sensitivity and specificity of the binary classification model are shown in Table 3. In the training / test set, the sensitivity for nasopharyngeal carcinoma was 97.5%, and the specificity for non-nasopharyngeal carcinoma was 87.1%. Figure 4 To differentiate between the nasopharyngeal carcinoma group and the non-nasopharyngeal carcinoma group, a peak scatter plot was used. Figure 5 To differentiate between nasopharyngeal carcinoma and non-nasopharyngeal carcinoma groups, the ROC curve was plotted with AUC=0.99.
[0107] Table 3. Performance of the binary classification model based on 56 differentially methylated regions on the training / test sets.
[0108] 2. Model Validation Unlike the clinical samples used in the above model construction, the inventors selected 162 cfDNA samples that tested positive for EB virus (independent validation set), including 36 cases of nasopharyngeal carcinoma and 126 cases of non-nasopharyngeal carcinoma (including normal healthy or benign diseases) for independent validation of the binary classification model.
[0109] The results of the binary classification model constructed above are shown in Table 4 below. The sensitivity for nasopharyngeal carcinoma is 80.6%, and the specificity for normal and benign diseases is 92.1%. Figure 6 To differentiate nasopharyngeal carcinoma from normal and benign lesion groups, ROC curves were plotted with AUC=0.91.
[0110] Table 4. Performance of the binary classification model based on 56 differentially methylated regions on the independent validation set.
[0111] Example 4: Establishing a machine learning model for early screening of nasopharyngeal carcinoma based on the differentially methylated regions of TOP20 1. Model Building The model construction steps are as described in Example 3. The Hg38 coordinates of the top 20 (TOP20) characteristic differentially methylated regions selected in this example are shown in Table 5.
[0112] Table 5. Top 20 Characteristic Differential Methylation Regions
[0113] Using the samples used to build the model in Example 3, a binary classification model was constructed based on the above 20 differentially methylated regions.
[0114] The sensitivity and specificity of the binary classification model constructed based on the differentially methylated regions of TOP20 are shown in Table 6. In the training / test set, the sensitivity for nasopharyngeal carcinoma is 91.1%, and the specificity for non-nasopharyngeal carcinoma is 90.3%.
[0115] Table 6. Performance of the binary classification model based on the differentially methylated regions of TOP20 on the training and test sets.
[0116] 2. Model Validation Using the independent validation set samples from Example 3 used to validate the model, the binary classification model constructed based on the TOP20 differentially methylated regions was independently validated. The results are shown in Table 7 below: the sensitivity for nasopharyngeal carcinoma was 77.8%, and the specificity for normal and benign diseases was 93.7%.
[0117] Table 7 Performance of the binary classification model based on the differentially methylated regions of TOP20 on the independent validation set.
[0118] Example 5: Establishing a machine learning model for early screening of nasopharyngeal carcinoma based on the TOP10 differentially methylated regions. 1. Model Building The model construction steps are as described in Example 3. The Hg38 coordinates of the first 10 characteristic differentially methylated regions selected in this example are shown in Table 8.
[0119] Table 8. 10 Characteristic Differential Methylation Regions
[0120] Using the samples used to build the model in Example 3, a binary classification model was constructed based on the above 10 differentially methylated regions.
[0121] The sensitivity and specificity of the binary classification model constructed based on the differentially methylated regions of TOP10 are shown in Table 9. In the training / test set, the sensitivity for nasopharyngeal carcinoma is 89.9%, and the specificity for non-nasopharyngeal carcinoma is 96.8%.
[0122] Table 9. Performance of the binary classification model based on the TOP10 differentially methylated regions on the training / test sets.
[0123] 2. Model Validation Using the independent validation set samples from Example 3 used to validate the model, the binary classification model constructed based on the TOP10 differentially methylated regions was independently validated. The results are shown in Table 10 below: the sensitivity for nasopharyngeal carcinoma was 75%, and the specificity for normal and benign diseases was 95.2%.
[0124] Table 10 Performance of the binary classification model based on the differentially methylated regions of the TOP10 on the independent validation set.
[0125] As shown in Examples 3 to 5, selecting the Top 20 or Top 10 characteristic differential methylation regions from the 56 selected characteristic differential methylation regions to construct a machine learning model for early screening of nasopharyngeal carcinoma showed good performance on both the training / test set and the independent validation set.
[0126] In the independent validation set of 162 plasma samples, the early screening model constructed using the top 10 characteristic differentially methylated regions showed better specificity, reaching 95.2% in the independent validation set. In terms of sensitivity, the early screening model constructed using 56 characteristic differentially methylated regions maintained a high sensitivity of 80.6% in the independent validation set.
[0127] Example 6: DNA methylation detection using immunoprecipitation enrichment and qPCR method To further verify the clinical performance of differentially methylated genes corresponding to the 56 differentially methylated regions of nasopharyngeal carcinoma-related genes on nasopharyngeal carcinoma plasma samples, the inventors selected 156 additional cfDNA samples that tested positive for EBV for subsequent gene methylation immunoprecipitation enrichment tests and qPCR detection. The first four differentially methylated genes, namely DAPK1, ITGA4, MIR129-2, and ALX4, were selected for combination, and the primers for each gene are shown in Table 11. Table 11 Gene Primers and Probes
[0128] DNA methylation immunoprecipitation enrichment and qPCR detection were performed to evaluate the performance of gene combinations, as referenced. Figure 7 The specific steps are as follows: (1) DNA extraction cfDNA was extracted using a commercially available extraction kit, following the instructions in the manufacturer's manual.
[0129] (2) Treatment of methylated DNA The extracted nucleic acids were used for methylated cfDNA enrichment. Different reactions were performed. Based on the principle of 5mC antibody, a methylated DNA enrichment kit (e.g., MagMeDIP kit, Cat. No. C02010021, Diagenode) was used. After the methylation enrichment reaction, the DNA was purified according to the instructions, with an elution volume of 50 μL.
[0130] (3) qPCR detection Using the methylated DNA enriched by immunoprecipitation as a template, PCR amplification was performed, with each primer having a final concentration of 10 μM. The PCR reaction system consisted of 5 μL of enriched template DNA, 2.5 μL of premixed solution containing the primers, 17.5 μL of PCR reagent (2×Rapid Taq Master Mix), and water to a final volume of 35 μL. The PCR reaction conditions were as follows: 95℃ for 5 minutes, 95℃ for 15 seconds, 60℃ for 40 seconds, for 45 cycles.
[0131] (4) Analysis of clinical sample test results The data from the PCR test were analyzed, and the results of qPCR detection of DAPK1, ITGA4, ALX4 and MIR129-2 gene methylated DNA were summarized.
[0132] Set the Ct value of samples with a detection result of "Undetermined" to 45, and plot the ROC curves respectively, as follows: Figure 8 As shown, the areas under the ROC curves (AUCs) of the four genes screened based on methylation targets were 0.817, 0.898, 0.850, and 0.814, respectively.
[0133] Based on the ROC curves, cut-off values were set for different genes: for the DAPK1 gene, the cut-off value was set to Ct=34.26; for the ITGA4 gene, the cut-off value was set to Ct=36.39; for the ALX4 gene, the cut-off value was set to Ct=35.00; and for the MIR129-2 gene, the cut-off value was set to Ct=36.81.
[0134] If the amplification Ct value of the monitored sample is equal to or lower than the set cut-off value, the sample is judged as a positive sample; otherwise, it is judged as a negative sample. The test results of 156 samples were statistically analyzed, as shown in Tables 12-15.
[0135] Table 12 Performance of qPCR detection of DAPK1 gene methylation by immunoprecipitation enrichment method
[0136] Table 13 Performance of qPCR detection using the gene ITGA4 methylation immunoprecipitation enrichment method
[0137] Table 14 Performance of qPCR detection of ALX4 gene methylation by immunoprecipitation enrichment
[0138] Table 15 Performance of qPCR detection method for MIR129-2 methylation immunoprecipitation enrichment of gene MIR129-2 methylation
[0139] As shown in Table 4, compared with methylation immunoprecipitation enrichment qPCR, the early screening machine model for nasopharyngeal carcinoma has higher sensitivity and specificity, with a sensitivity of 80.6% and a specificity of 92.1%. Table 13 shows that the differentially methylated region of the ITGA4 gene, when validated on the qPCR detection platform, exhibits higher sensitivity (81.73%) for nasopharyngeal carcinoma and maintains high specificity (87.80%) for non-nasopharyngeal carcinoma samples, achieving an accuracy of 83.33%.
[0140] To improve the sensitivity and specificity of predictions, this application uses SPSS binary logistic regression equations to calculate the performance of the four-gene combination based on the data from the test. The detection results are applied to SPSS and calculated using the following formula: Logistic score =
[0141]
[0142] Wherein the Logistic score represents the logistic regression value. These are logistic regression systems for four genes. These are the Ct values for the four genes.
[0143] Obtain the predicted probability (P), and plot the ROC curve, as follows: Figure 9 The area under the ROC curve (AUC) of the four gene combinations screened based on methylation targets is 0.980. Based on the ROC curve, the positive prediction threshold is 0.663; a positive prediction probability greater than 0.663 is considered positive, and vice versa. The test results of 156 samples are statistically analyzed, as shown in Tables 16 and 17.
[0144] Table 16 Performance of qPCR Detection Method for Gene Combination Methylation Immunoprecipitation Enrichment
[0145] Table 17 Comparison of qPCR detection performance of single-gene and gene combination methylation immunoprecipitation enrichment methods
[0146] As shown in Table 17, the detection using gene combinations (including DAPK1, ITGA4, ALX4 and MIR129-2) has superior performance, with sensitivity, specificity and accuracy all exceeding 90%.
[0147] All references to this application are incorporated herein by reference as if each reference were individually incorporated herein by reference. Furthermore, it should be understood that after reading the foregoing teachings of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope defined by the appended claims.
Claims
1. A method for screening methylation biomarkers for predicting nasopharyngeal carcinoma risk, characterized in that, Includes the following steps: Biological samples were obtained from multiple subjects who were positive for Epstein-Barr virus infection, including nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects; Methylation sequencing was performed on biological samples from the subjects who tested positive for EB virus infection to obtain methylation sequencing data; Based on the methylation sequencing data, differentially methylated regions between nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects were obtained, and the methylation expression levels of the differentially methylated regions were obtained. The biological samples were randomly divided into training and testing sets, and a methylation expression level matrix was constructed. The differentially methylated regions were ranked by importance using a feature selection tool. At least one of the top N differentially methylated regions was selected as a characteristic methylated region as a methylation region biomarker for predicting the risk of nasopharyngeal carcinoma. The gene corresponding to the methylation region biomarker is a methylation gene biomarker, where N = 1 to 100.
2. The method according to claim 1, characterized in that, The multiple subjects who tested positive for EB virus infection were detected using EB virus nucleic acid detection methods.
3. The application of a reagent for detecting the combined methylation level of methylation region markers in the preparation of a kit for predicting the risk of nasopharyngeal carcinoma, characterized in that, The methylation region marker combination includes at least one of the following: chr9:87498973-87499172, chr2:181457944-181458143, chr11:43581352-43581551, chr11:44304060-44304259, chr14:60509500-60509699, chr14:69547960-69548159, chr12:48198223-48198422, chr1:170661243-170661442, chr8:120124836-120125035, and chr6:77463420-77463619.
4. The application according to claim 3, characterized in that, The methylation region marker combination further includes at least one of the following: chr1:58249628-58249827, chr22:42283819-42284018, chr1:229407770-229407969, chr21:33025905-33026104, chr12:24903154-24903353, chr12:4272683-4272882, chr1:240091583-240091782, chr2:161423989-161424188, chr4:1002218-1002417 and chr20:43915894-43916093.
5. The application according to claim 3 or 4, characterized in that, The methylation region marker combination further includes at least one of the following: chr1:153679627-153679826, chr6:10426407-10426606, chr4:121380465-121380664, chr2:72917151-72917350, chr9:137156751-137156950, chr5:50969615-50969814, chr13:107868359-107868558, chr7:93890509-93890708, chr21: 33071882-33072081, chr19:41135567-41135766, chr4:140426905-140427104, chr6:137498188-137498387, chr17:48646527-48646 726, chr4:107931901-107932100, chr1:44417990-44418189, chr1:156923961-156924160, chr12:62632009-62632208, chr10:10464 0217-104640416, chr12:102958605-102958804, chr7:87628001-87628200, chr12:8018782-8018981, chr5:80960731-80960930, chr 17:44325178-44325377, chr7:49775977-49776176, chr16:23754731-23754930, chr12:21527708-21527907, chr2:184598918-18459 9117, chr12:75207510-75207709, chr8:37965734-37965933, chr4:110622983-110623182, chr3:194487549-194487748, chr19:58440233-58440432, chr21:25640018-25640217, chr5:100903323-100903522, chr1:228458190-228458389 and chr21:45973989-45974188.
6. The use of a reagent for detecting methylation levels in a combination of methylation gene markers in the preparation of a kit for predicting the risk of nasopharyngeal carcinoma, wherein the combination of methylation gene markers includes at least one of DAPK1, ITGA4, ALX4, and MIR129-2.
7. The application according to claim 6, characterized in that, The methylation gene marker combination further includes at least one of the following: SIX6, CCDC177, OR10AD1, PRRX1, COL14A1, HTR1B, DAB1, TCF20, ACTA1, OLIG2, BCAT1, CCND2, FMN2, TBR1, IDUA, TOX2, NPR1, TFAP2A, QRFPR, EMX1, MIR3621, LINC02106, FAM155A, TFPI2, OLIG1, CYP2F1, CLG N, OLIG3, MIR196A1, LOC101929595, RNF220, LRRC71, MIRLET7I, SORCS3, PAH, RUNDC3B, FOXJ2, RASGRF2-AS1, SLC25A 39. VWC2, CHP2, SPX, ZNF804A, KCNC2, ADRB3, PITX2, LINC00884, ZNF132, JAM2, ST8SIA4, HIST3H2BB and LOC101928796.
8. A kit for predicting the risk of nasopharyngeal carcinoma, characterized in that, include: (1) A reagent for detecting the methylation level of a combination of methylation region markers or a reagent for detecting the methylation level of a combination of methylation gene markers, wherein the combination of methylation region markers includes at least one of the following: chr9:87498973-87499172, chr2:181457944-181458143, chr11:43581352-43581551, chr11:44304060-44304259, chr14:6050950 The methylation gene marker combination includes at least one of DAPK1, ITGA4, ALX4, and MIR129-2. The methylation gene marker combinations are: 0-60509699, chr14:69547960-69548159, chr12:48198223-48198422, chr1:170661243-170661442, chr8:120124836-120125035, and chr6:77463420-77463619. (2) EB virus nucleic acid detection reagent, wherein the methylation region marker combination methylation level detection reagent or the methylation gene marker combination methylation level detection reagent is used to detect subjects who are positive for EB virus nucleic acid.
9. A method for constructing a model for predicting the risk of nasopharyngeal carcinoma, characterized in that, Includes the following steps: The methylation level data of methylation region marker combinations or methylation gene marker combinations in the population biological samples were randomly divided into two groups: one group was the training set and the other group was the test set. The population biological samples were all positive for EB virus nucleic acid and included nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects. The model for predicting nasopharyngeal carcinoma risk is constructed using the training and test sets based on machine learning methods. The methylation region marker combination includes at least one of the following: chr9:87498973-87499172, chr2:181457944-181458143, chr11:43581352-43581551, chr11:44304060-44304259, chr14:60509500-60509699, chr14:6954 The methylation gene marker combination includes at least one of DAPK1, ITGA4, ALX4, and MIR129-2. The methylation gene marker combination includes 7960-69548159, chr12:48198223-48198422, chr1:170661243-170661442, chr8:120124836-120125035, and chr6:77463420-77463619.
10. A system for predicting the risk of nasopharyngeal carcinoma, characterized in that, Includes the following modules: The data input module is used to input methylation level data of methylation region marker combinations or methylation gene marker combinations obtained from the subject's biological samples. The methylation region marker combinations include at least one of the following: chr9:87498973-87499172, chr2:181457944-181458143, chr11:43581352-43581551, chr11:44304060-44304259, chr14:60 509500-60509699, chr14:69547960-69548159, chr12:48198223-48198422, chr1:170661243-170661442, chr8:120124836-120125035 and chr6:77463420-77463619; the methylation gene marker combination includes at least one of DAPK1, ITGA4, ALX4 and MIR129-2; A database storage module is used to store methylation level data of the combination of methylation region markers or the combination of methylation gene markers in the population biological samples, wherein the population biological samples are all positive for EB virus nucleic acid and include nasopharyngeal carcinoma patients and non-nasopharyngeal carcinoma subjects; The disease prediction module is connected to the data input module and the database storage module, respectively. It is used to construct a prediction model using the methylation level data of the methylation region combination or methylation gene combination of the population biological sample, and predict whether the subject has the risk of developing nasopharyngeal carcinoma based on the methylation level data of the methylation region marker combination or methylation gene marker combination of the subject obtained from the data input module.