Methylation marker for identifying papillary thyroid carcinoma invasion subtype, diagnosis model and application
By screening gene fragments as methylation markers and constructing diagnostic models, the problem of distinguishing invasive subtypes of papillary thyroid carcinoma in existing technologies has been solved, achieving efficient and accurate identification of invasive subtypes and improving the reliability and safety of diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-07
AI Technical Summary
Current technologies struggle to effectively differentiate between aggressive subtypes of papillary thyroid carcinoma, leading to difficulties in treatment decisions and an increase in potential complications.
Multiple gene fragments were screened as methylation markers. Combined with DNA methylation sequencing and the construction of a diagnostic model, a detection technology was developed using a logistic regression algorithm to identify invasive subtypes of papillary thyroid carcinoma.
It provides an efficient, rapid, and low-cost method that can accurately identify the invasive subtypes of papillary thyroid carcinoma, improving the reliability and safety of diagnosis and facilitating large-scale clinical application.
Smart Images

Figure CN121802045A_ABST
Abstract
Description
Technical Field
[0001] This application relates to methylation markers, diagnostic models, and applications for identifying invasive subtypes of papillary thyroid carcinoma, and belongs to the field of biomedical and molecular diagnostic technology. Background Technology
[0002] Thyroid cancer ranks ninth among malignant tumors worldwide and is the most common endocrine system cancer. Most thyroid tumors originate from follicular epithelial cells and can be classified as benign, low-risk, or malignant according to the World Health Organization (WHO) classification. Papillary thyroid carcinoma (PTC) is the most prevalent subtype, accounting for approximately 80% of all thyroid cancer cases. Its global incidence continues to rise, partly due to increased diagnostic sensitivity. Although most PTCs exhibit indolent biological behavior and a good prognosis, a significant proportion of cases still show aggressive characteristics, including extrathyroidal invasion, lymph node metastasis, and distant metastasis. These characteristics significantly increase the risk of recurrence and death. Of particular concern are variant subtypes such as tall cell type (TCV), columnar cell type, and bootie type (HV), which exhibit stronger clinicopathological aggressiveness compared to classic PTC (CPTC). Furthermore, diffuse sclerotic type (DSV) and solid type (SV) variant subtypes are also closely associated with highly aggressive behavior.
[0003] Initial treatment for PTC should be based on a comprehensive preoperative risk assessment, including physical examination, ultrasound evaluation, and cytological analysis. Treatment options range from active surveillance and minimally invasive interventions to surgery. Minimally invasive alternatives to surgery, such as active surveillance, are gaining increasing attention, particularly for low-risk thyroid cancer. Fine-needle aspiration biopsy (FNAB) remains the gold standard for preoperative diagnosis of thyroid nodules. While cytological diagnosis is crucial for identifying aggressive PTC subtypes (due to its unique prognostic and treatment guidance value), relying solely on cytology for subtype classification remains challenging. Limitations include insufficient sample size, subtle cellular morphological features, and limited experience of pathologists with rare subtypes, often leading to underreporting or misclassification. These shortcomings highlight that cytology alone is insufficient to reliably predict tumor invasiveness. Therefore, clinicians often face a decision-making dilemma when determining the extent of surgery: total thyroidectomy for low-risk cancer may cause complications such as permanent hypoparathyroidism or recurrent laryngeal nerve injury, while inadequate treatment of aggressive tumors may require supplementary surgery and increase the risk of recurrence.
[0004] DNA methylation, as a stable epigenetic modification, plays a crucial role in tumorigenesis and development. Due to its stability, cancer-specific patterns, and fundamental role in gene expression regulation, it has become a promising source of cancer biomarkers. In various cancer types, DNA methylation signatures have shown significant value in predicting tumor invasiveness and metastatic potential. In thyroid cancer research, whole-genome methylation analysis of tissue samples and peripheral blood leukocytes has revealed widespread methylation changes that can distinguish malignant from benign thyroid nodules. Multiple studies have reported differential DNA methylation profiles among different PTC subtypes and observed significant associations between DNA methylation patterns and BRAF and RAS gene mutations. However, methylation studies specifically targeting PTC subtypes (especially their invasive behavior) remain limited. To address this unmet clinical need, methylation biomarkers for invasive subtypes of papillary thyroid carcinoma have significant application value. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by screening methylation markers or targets that can distinguish between different histological subtypes of papillary thyroid carcinoma, based on the high-invasiveness and low-invasiveness of various histological subtypes. This invention aims to establish a multi-target detection technology based on DNA methylation, integrate mutation, pathological features, and imaging features, construct a detection model, and develop multi-target detection technologies based on NGS-based methylation-targeted sequencing or qPCR applicable to methylation. This will enable the identification of invasive subtypes of papillary thyroid carcinoma in a highly efficient, rapid, and low-cost manner.
[0006] To achieve the above objectives, this application adopts the following technical solution: In a first aspect, this application provides methylation markers for identifying invasive subtypes of papillary thyroid carcinoma, said methylation markers being selected from any one or more of the following gene fragments: Contains chr1:1916201:1916400 (SEQ ID NO:1) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr1:22920001:22920200 (SEQ ID NO:2) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr3:129299001:129299200 (SEQ ID NO:3) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr4:332601:332800 (SEQ ID NO:4) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr5:138730201:138730400 (SEQ ID NO:5) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr6:31148401:31148600 (SEQ ID NO:6) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr7:65959401:65959600 (SEQ ID NO:7) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr7:73638801:73639000 (SEQ ID NO:8) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr9:130587601:130587800 (SEQ ID NO:9) and any gene fragment within 5kb upstream and / or 5kb downstream of it.
[0007] Contains chr11:15963001:15963200 (SEQ ID NO:10) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr11:34176801:34177000 (SEQ ID NO:11) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr11:69407801:69408000 (SEQ ID NO:12) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr12:121437201:121437400 (SEQ ID NO:13) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr12:132857601:132857800 (SEQ ID NO:14) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr13:29060601:29060800 (SEQ ID NO:15) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr15:75470601:75470800 (SEQ ID NO:16) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr16:5198401:5198600 (SEQ ID NO:17) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr16:11017801:11018000 (SEQ ID NO:18) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr16:12354401:12354600 (SEQ ID NO:19) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr16:87712401:87712600 (SEQ ID NO:20) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr16:87904401:87904600 (SEQ ID NO:21) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr17:4560001:4560200 (SEQ ID NO:22) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr17:6797001:6797200 (SEQ ID NO:23) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr19:2254001:2254200 (SEQ ID NO:24) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr19:5146201:5146400 (SEQ ID NO:25) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr19:6223801:6224000 (SEQ ID NO:26) and any gene fragment within 5kb upstream and / or 5kb downstream; Contains chr19:12861001:12861200 (SEQ ID NO:27) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr19:44278601:44278800 (SEQ ID NO:28) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr20:2786401:2786600 (SEQ ID NO:29) and any gene fragment within 5kb upstream and / or 5kb downstream of it; Contains chr22:27895001:27895200 (SEQ ID NO:30) and any gene fragment within 5kb upstream and / or 5kb downstream of it. In some embodiments, the methylation marker is selected from any one or more of the gene fragments shown in SEQ ID NO:1-30.
[0008] In some embodiments, the methylation marker is selected from any one of the following groups: 1) Combinations of gene fragments shown in SEQ ID NO:1-30; 2) Combinations of gene fragments shown in SEQ ID NO:3, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:13, SEQ ID NO:15, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:22 and SEQ ID NO:30; 3) Combinations of gene fragments shown in SEQ ID NO:3, SEQ ID NO:10, SEQ ID NO:16, SEQ ID NO:22 and SEQ ID NO:30.
[0009] Secondly, this application provides the use of methylation markers or their detection reagents as described in the first aspect in the preparation of in vitro diagnostic products for identifying invasive subtypes of papillary thyroid carcinoma.
[0010] Thirdly, this application provides a kit for identifying invasive subtypes of papillary thyroid carcinoma, the kit comprising reagents for detecting the methylation level of methylation markers as described in the first aspect.
[0011] In some embodiments, the reagent for detecting methylation levels is selected from one or more of the following methods: bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry (e.g., mass spectrometry of flight). In some embodiments, the reagent is selected from one or more of the following: bisulfite and its derivatives, PCR buffer, polymerase, dNTP, primers, probes, methylation-sensitive or insensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporter agents, exonucleases, alkaline phosphatase, internal standards, and controls.
[0012] Fourthly, this application provides a method for constructing a diagnostic model for identifying invasive subtypes of papillary thyroid carcinoma, comprising: acquiring level data of a target biomarker in several papillary thyroid carcinoma samples as a training set, wherein the target biomarker is a methylation biomarker as described in the first aspect; A diagnostic model is constructed using a logistic regression algorithm based on the classification information of the samples and the level data of the target markers.
[0013] Fifthly, this application provides a computer-readable storage medium that stores computer instructions that, when executed by a processor, implement the method described in the fourth aspect.
[0014] In a sixth aspect, this application provides an apparatus including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the method described in the fourth aspect.
[0015] In a seventh aspect, this application provides an apparatus for identifying an invasive subtype of papillary thyroid carcinoma, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to perform the following steps: 1) obtaining the methylation level or methylation state of at least one CpG dinucleotide, a methylation marker as described in the first aspect, in a target sample, and 2) determining whether the target sample is an invasive subtype of papillary thyroid carcinoma based on the methylation level or methylation state of 1).
[0016] In some embodiments, step 2) specifically includes: inputting the methylation level or methylation state into a pre-constructed diagnostic model to obtain the probability value of the invasive subtype of papillary thyroid carcinoma, and identifying whether the target sample is an invasive subtype of papillary thyroid carcinoma based on a pre-set threshold and probability value; the diagnostic model is the diagnostic model constructed as described in the fourth aspect.
[0017] It should be understood that, within the scope of this application, the above-described technical features of this application and the technical features specifically described below (such as in the embodiments) can be combined with each other to form a preferred technical solution.
[0018] Compared with the prior art, this application has the following beneficial effects: 1) This application analyzes methylation sequencing data from tissue samples of invasive and non-invasive papillary thyroid carcinoma to identify biomarkers for distinguishing invasive subtypes of papillary thyroid carcinoma, and constructs corresponding diagnostic models based on these biomarkers. Based on the biomarkers and / or diagnostic models of this application, the presence of invasive subtypes of papillary thyroid carcinoma can be effectively identified, overcoming the shortcomings of existing identification methods that cannot distinguish invasive subtypes of papillary thyroid carcinoma, and providing a new method for identifying invasive subtypes of papillary thyroid carcinoma. 2) The detection process based on the biomarkers or diagnostic models of this application is highly safe and facilitates large-scale clinical application. Attached Figure Description
[0019] Figure 1 This is a flowchart for the screening and validation of methylation markers.
[0020] Figure 2 This represents the distribution of predicted scores in the training and test sets of the combined model built based on 30 markers in Example 3.
[0021] Figure 3 The ROC curves of the combined model built based on 30 markers in Example 3 are shown in the training and test sets.
[0022] Figure 4 This represents the distribution of predicted scores in the training and test sets for the combined model built based on 10 markers in Example 4.
[0023] Figure 5 The ROC curves of the combined model built based on 10 markers in Example 4 are shown in the training and test sets.
[0024] Figure 6 This represents the distribution of predicted scores in the training and testing sets of the combined model constructed based on five markers in Example 5.
[0025] Figure 7 The ROC curves of the combined model built based on 5 markers in Example 5 are shown in the training and test sets. Detailed Implementation
[0026] To make the technical solution of this application clearer and easier to understand, preferred embodiments are described in detail below with reference to the accompanying drawings.
[0027] It should be noted that in this application, the singular forms "an," "a," and "the" all include their plural forms unless the context otherwise requires. Therefore, for example, "an agent" can include multiple agents.
[0028] In this application, unless otherwise stated, the terms “comprising,” “including,” or “containing” mean that the listed values, steps, or ingredients are included, but do not exclude the inclusion of other values, steps, or ingredients.
[0029] I. Target landmarks and their target areas As used herein, the term "target biomarker" refers to a target nucleic acid or gene region whose methylation level indicates whether a subject has papillary thyroid carcinoma. The term "target biomarker" should be considered to include all transcriptomorphic variants of the genes described herein and all their promoters and regulatory elements. As understood by those skilled in the art, certain genes are known to exhibit allelic variations or single nucleotide polymorphisms ("SNPs") among individuals. SNPs include insertions and deletions of simple repetitive sequences of varying lengths (e.g., dinucleotide and trinucleotide repeats). Therefore, this application should be understood to extend to all forms of biomarkers / genes resulting from any other mutation, polymorphism, or allelic variation. Furthermore, it should be understood that the term "target biomarker" should include both the sense and antisense strand sequences of the biomarker or gene.
[0030] As used herein, the term "target biomarker" is broadly interpreted to include both 1) the original biomarker found in a biological sample or genomic DNA (in a specific methylation state) and 2) its processed sequence (e.g., the corresponding region after bisulfite conversion or MSRE treatment). The bisulfite-converted region differs from the target biomarker in that one or more unmethylated cytosine residues are converted to uracil, thymine, or other bases that differ from cytosine in hybridization behavior. The MSRE-treated region differs from the target biomarker in that the sequence is cleaved at one or more MSRE cleavage sites.
[0031] The target biomarkers described in this article are selected from any 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or all 30 of the following gene sequences (Hg19 coordinates): Contains chr1:1916201:1916400 (SEQ ID NO:1) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:1 is shown below: TGGGCACTCTTAGGGGTGGGGTGGCCTGAGGCCTGCACTCAGCCCTCTCTGTGCCTCCTGCCCTGCCCAGGCTGGAAAGCTTAGGTCATCCCGGGCACGTGCAGGGTCAGTGCTGTGGATGTCATGGACCTGGCTGCAGCTCCCGGGGCTGGGGGCTGTGTGGTTCGGGAGTCCTACTTGAGGTCAGCTCAGGGTCCCCT Contains chr1:22920001:22920200 (SEQ ID NO:2) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:2 is shown below: CACCGTCCCGTGGCAGGACAAGGAGATGCAGAGCTACTCCACCCTCAAGGCCGTCACCACCAGAGCCACCGTCTCCGGCCTCAAGCCGGGCACCCGCTACGTGTTCCAGGTCCGAGCCCGCACCTCAGCAGGCTGTGGCCGCTTCAGCCAGGCCATGGAGGTGGAGACCGGGAAACCCCGTGAGTGCAGGGAGGGGGCGT Contains chr3:129299001:129299200 (SEQ ID NO:3) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:3 is shown below: GTTGTCCAGACCATTTGCTGACCACGAGACCAGAGTGGTGAGTTCCCCTAGGGACAGGGTCTGGTGCAGACCTCAGGGGCTTGGCAAGAGAAGCAGCATAGCCAGGCCCGGGTGCTGCTGGCTCCCTGTGCCAGAATACACGAGCCACTGACCCTCTCCGGGCCTCCGTGTTCCTACCAAGGAGGGGCCCAGACGGAGTC Contains chr4:332601:332800 (SEQ ID NO:4) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:4 is shown below: GGGGTCGGAGCATTGACTGGTGTTGAGCAAACTCCTCACGTTTCCTGTCAGAAGAAAGACATGGCTGAGCCTGGGCTGGAGGGAGTCGCAGGTGTCCGCGGGAGACCAGCCGGGTTCTGCACATGCACTGTCCTGCTGCGCACTGTTCTGTTCCTCCAGGTCTCCTCCCGGGGAGAGAGGGCGCTGAGAACTTGGAGGAA Contains chr5:138730201:138730400 (SEQ ID NO:5) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:5 is shown below: CTCCACACACTCGAACTGCGCTGGGGCGGCAGGACTTGGCCCACGGGGCCGCAGCTCTAGGTAGGTGGCCCAGCGGGAGCCACCATCGGGGACCTGGGACTGGCGTGGGACCGCGGCGGGAGACGCTGGCCCCGGCGGCAAGGGGCTGATGAAGGCCGGCTCCGTGAACTGTTGTTGCGCCTCGCGATCGTCTGCGCCGG Contains chr6:31148401:31148600 (SEQ ID NO:6) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:6 is shown below: TTCCGGAACGAACCGTCGCCAGCAAGCACAGCAGTAGGACCAGGGGGATGCAAGAGCGGGGGCGGCCGGGGATCGTGCTTCTCGCTCAGGTCCAGATTCCCGGCAACCAGGCCGGCGGAATCACGTGCCATGCTCCAGGCCAGCGTAGTCCCGCCCATCTTCCAGCTGAGCGTACCGGGAGGCTCCCATTGGACTGGAGC Contains chr7:65959401:65959600 (SEQ ID NO:7) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:7 is shown below: GGGCACATGTAATACCAGCTATTCGGGAGGTTGAGGCAGAGGCAGACAATTGCTTGAACCCGGGAGGCAGAGGCTGCAGTTAGCCGAGATCGTGCCACTGCACTCTAGCCTGGGCGACAGACCAAGACTCCGTCTAAAAAATAAAAAACTATATATATGTGTGTGTATATATATATATATATATGCGATGGAGCCCGGGA Contains chr7:73638801:73639000 (SEQ ID NO:8) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:8 is shown below: GATCTCCGGCAAATCCCAGAGACAGACACACACCGCCAGGACCTGTCGGGCGTGGCATTTCCCGCGGGCTGTTTTGTGGAGTCTGGGGGTCTGAGTCTGGGGGAGCAGTCATGGGTGTCTCCATGTGAACCACTCGATGACCTGTCTGCCCGCTGGTCCACAGGGCAACTCCAGAGAAGCATCCCCTGGCCCG Containing chr9:130587601:130587800 (SEQ ID NO:9) and sequences within 5kb upstream and / or 5kb downstream of it. SEQ ID NO:9 is shown below: CAGCTCAGTTCCACCTTCACCGTCACCGTCCGGGGCCTGCGGGGAGACAGACGCGGATGGAACACTGAAGCGGACAGGCCAGGCGGGGAGCGAGGCCTGGCGTCGGGTGGGCGGCGGCTGCACTCTTACCTGGCCAGGTGTGGGTTTATGGGATAGGGATGGGACTGGGGCCAGGCTTGTGCCCAAAGGGAAGGGGCCCA Contains chr11:15963001:15963200 (SEQ ID NO:10) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:10 is shown below: CTTGGCAGCTGCGGCCGGGGGCGGACGGTGGGGCGGTGTGGGTTTCAGCCTCCCCGGAGGGCCCTCACGGCTGAGCAAACGTTCGGGCTGATGTCGGCAACATGCGGAATCAATTTTCGGGGAACTCAGCAGCCAAACCATCCACCTTTGGGCGGGAAGCAGGATCGCTGTAGGCCCGGGGGCTCCTTGTCTCCCGTTT Contains chr11:34176801:34177000 (SEQ ID NO:11) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:11 is shown below: GCCTTCCCCTGTTTGCACAAGGACTCCAATGTGGAGGTGAGTGCCCCATCTGGGCTCTGTCTCAGGTGTGCTTCCTCAAAAAGAAAGCGCAGAGCCCAGGCCCGGAGGCCCGGAGGCCCGGAGCTCCCCTCGACAGGCTGAGCCCTGCCTTGGGCTCCCAGATGAGGTGGTCCAGGACCGGGAACGATTTAGTTCTGCTG Contains chr11:69407801:69408000 (SEQ ID NO:12) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:12 is shown below: CTCTCTTCCTATTGAGCTGGACTGTGCTGCACGCAGAGGGCCAGAGGAGGCCACGTCTGCTCCGAGGGCCCACGGGGGCTTAGCGGGCCGCACGCGTTGCTTCCTGATCACCCGGTGCCCGCACCGGCCCCGGTATAGACAATGAGCTCTCAACACCAGGGCCTCTGCCCAGCGGCTCTGGCAGCTCCCCGGGGCAGTGA Contains chr12:121437201:121437400 (SEQ ID NO:13) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:13 is shown below: CAGGCCTGCTGGCCCTCCCTTGGCCTGTGACAGAGCCCCTCACCCCCACATCCCCCGGGCTCAGGAGGCTGCTCTGCTCCCCCAGGTCTTCACCTCAGACACTGAGGCCTCCAGTGAGTCCGGGCTTCACACGCCGGCATCTCAGGCCACCACCCTCCACGTCCCCAGCCAGGACCCTGCCAGCATCCAGCACCTGCAGC Contains chr12:132857601:132857800 (SEQ ID NO:14) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:14 is shown below: CCCCTGCAGCCAGCCCCCTGCAGCCAGCCGCCTGCAGCCAGTGGTTCTTCAGTGGCCCCAGGGAAAGGCCAGGCACTGTGGAGACGCTTCCCAGCCACACTCGCTGAGCCGCGCCGGGATAGGGTCTCTCCACCCAAGCTCCTTTTAGGAAAGTGGTTTCCTGGGAAGATGGATAAGACCCAGCGTCCACCGGGGTCTGCG Contains chr13:29060601:29060800 (SEQ ID NO:15) and its upstream and / or downstream sequences within 5 kb; SEQ ID NO:15 is shown below: ATACTGGAGGAAACCAGGGCTGGACCAAGGCACGTGGGTGCCCGGGACAGGCCAATAGTATGGTGTTTGCCAATATTTACATAGAGAAAAGACTGAGCACTCCACAGCATAAGGCAGAAGCATAAGGCAGAGAAGGAGGACCATTGTTAGTCTAGACGGGCAATTCACTTCTGCCTCTAACGTACGTACCTCACTCGGCA Contains chr15:75470601:75470800 (SEQ ID NO:16) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:16 is shown below: ATGCGGCGACTGGGGGACCGGATAGGAGGGCCCTTGTCCTGAGGCTGCGTAGGCATAGCAAGCTCCGGGATTTTACTTCGTTGGAATGTTCTTTTGAGCTACTTTGGGGGCCGTGACTGAGAGAGCTCGTGGGGGACTCAGGCCCTTTCAGAGTCACCCGTGTGGCTCGGCCACTGCGCCAGGTCTTGGAACATCTAAGG Contains chr16:5198401:5198600 (SEQ ID NO:17) and its upstream and / or downstream sequences within 5 kb; SEQ ID NO:17 is shown below: GGGCATCGGGTGAAATGGGCAGAGTGGTGCTTACCCGGGATGGCGGTGAAGTGGGACGGGGAGGTCATCGTGACAAAGGGCGGCATGAGGTACTTGGCCTTGACGCCCTCCCCAGCCAGACCATCCAGGTTGGGGGTGTCCACATCCTGATCCTAGTCTCAGTGGAAGCCCTCAAAGGAGATCAGTAGCAGCCGTGAGTG Contains chr16:11017801:11018000 (SEQ ID NO:18) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:18 is shown below: CAGGCTGTATCCCATGAGCCTCAGCATCCTGGCACCCGGCCCCTGCTGGTTCAGGGTTGGCCCCTGCCCGGCTGCGGAATGAACCACATCTTGCTCTGCTGACAGACACAGGCCCGGCTCCAGGCTCCTTTAGCGCCCAGTTGGGTGGATGCCTGGTGGCAGCTGCGGTCCACCCAGGAGCCCCGAGGCCTTCTCTGAAG Contains chr16:12354401:12354600 (SEQ ID NO:19) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:19 is shown below: GCTTCTCTTTTCTCTCATGTGGGGTGGGGGCATCCATGTCCTCCTGGTCTCTTGGAGTGCGTTCCACCTTCTCCGTCTGCTTGCTCTCCAGGGCTTTCTTAGCCAGGTCCCGGCTGGAGTGGATGTTTGCTATGAGCTTGGGTCTGTGCACGCGTCCCCGGCTGGAGTGAGTGAGTGTTGAGCTCAGGTCAGTTCACAC Contains chr16:87712401:87712600 (SEQ ID NO:20) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:20 is shown below: AAGGCCACATGAGGGCAGAGCCGAAGGCGGGAGTAAGGCCACCACGAGCCTGGGAACGCCGAGGGCGCCAGCAGCCGCCAGCAGCTGGAGAGGATGGATGGGAGGATTTTAAAGCCTTCCGATAGGGGTGTGACTTCTTGATTTTGGACTCCGGGACTGCAGGACCGTGAGAAAAGGAGTGAGTGTTAAGTCCA Contains chr16:87904401:87904600 (SEQ ID NO:21) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:21 is shown below: GGCCCCACCCCAGGGAGGGGGCGGCTGAGGAATGCCAGGGCTTGTCATTCTGGACCTTGCTCGCCCCCCGGTATGTCGGGCATTCCTCTGCGGTGATAAAACCGAGCGAAATCTCAGCAGCCCAGGCAGCCACAGGCGGGGGAGTGCGGAGGTCACACAGATACTAACTTGTAATCTGCCGTCCACATTGCTGGCTGTGG Contains chr17:4560001:4560200 (SEQ ID NO:22) and its upstream and / or downstream sequences within 5 kb; SEQ ID NO:22 is shown below: GCAGCTAAGGCCCGGCGAGAAATTGAGCACAGCAGCTGCTGGCCCAGGTACTAAGCCCCTCACTGCCCGGGCCGGGGGCGCCGGCCAGCCGCTCCGAGTGCGGGGTCCGCCGAGCCCACCGGAACTCACGCTGGCCAGCAAGCACCGCGCGCAGCCCCTGTTCCCGCTCGCGCCTCTCCCTCCACACCTCCCC Contains chr17:6797001:6797200 (SEQ ID NO:23) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:23 is shown below: GAAACCCAAATACAACAGCCACTCTGGGCTGGGCGTGGTTCCCTGGGAGCTGGGTCTCAGATCCTTTTAAAGCTGATTTCTCCCCGGATCTCCGGGCAAGATGGGCAAGTACACGGTCCGCGTAGCCACCGGGGATTTGCTCCTGGCGGGCTCTCCCAACCTGGTGCAGCTATGGCTGGTGGGCGAGCACGGGGAGGCAG Contains chr19:2254001:2254200 (SEQ ID NO:24) and its upstream and / or downstream sequences within 5 kb; SEQ ID NO:24 is shown below: GTGGCATAGCTGGTATTATTACGCATTAGTAGAGGCTCCCTGCCCGTTGACTCGCTGCCTGCCACACTTCGACGCCACTCAACCAGGACACCACCCTCCAAGATCCAAGGACCCAGCTCCGGCACCTGAGCCTCACACCAGTCCGTTCCCACCCCACACCTCACCCATGCCTCCTGCAAACCCACTCGCTCCGGCTTTCA Contains chr19:5146201:5146400 (SEQ ID NO:25) and its upstream and / or downstream sequences within 5 kb; SEQ ID NO:25 is shown below: GGCCCCGCCCGGTCACTGCTCTCGGACCTGTCACTGAGCATCCAGGACCCCCCCCGCTCCGCCGTGCAGGCCGGCCCCGCCAGTCACCACTCTCGACCTGTCACTGAGCATCCAGGACCTCCTTGCCATGCAGGCTGGCCCCACCCGGTCACCACTCTCGGACCTTGCAGGGTCTTCCCTGCTCCTTGAGAAGGGGGTG Containing chr19:6223801:6224000 (SEQ ID NO:26) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:26 is shown below: CCTGTATCACCCCACGCACCCAGGAGCACTGAGCAGGCTTCTCTGCTGTTAGCCACCAGGGGCTCCACCCTTGCTAGGAAAGGAGCAGCGGCCAGCAGGTGCACGGCGGCCCCAGAGCCTGTCCAGACAATTGTGTACCTGGCCTCCCGGCACAGAAGCCAGCAGCAGCCAGATCCCGAAGGCCCTGGGGAGGGGCTCCG Contains chr19:12861001:12861200 (SEQ ID NO:27) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:27 is shown below: GAACTGAGCTCATCAGCACCGAGCCTGCCACCGCCTTTGCCAGGTTCCGCAGCTGACGGCACCGCCCCAGCTGCCCCCTCGCCCCCATTTATGGGCTGTCCGGCAGCTGGACAGCGAAGGCATGGGTCACAGGGTGACCAGCCCAGGCGGGTTCCTAATTACTCCGGGCAGGGTGAACATCTGCCAGTCTGGGAACACGGC Contains chr19:44278601:44278800 (SEQ ID NO:28) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:28 is shown below: CGGCCAGGGCTGCGGGGAGGTCAGCGGCGCCCCTAAATCCTGCACGCACGGCGGGCCCCGCACGGGCGCCGGGTGCAGCCCACACACCACCAGCTCCAGCACGATCTGCGCCGCCTGCCGCCCGGTCAGCGCCACGCGCCAGTCCCGCAGCCCGTTGTCGGTCATGAACAGCTGCCGGTAGGGGGCCAAGAAGGGAGGGG Containing chr20:2786401:2786600 (SEQ ID NO:29) and its upstream and / or downstream sequences within 5kb; SEQ ID NO:29 is shown below: ATGGGAGGACAGAGAGACATTTCAAGGGAAGCGCTGCACAGCCGGGACTCCCTGGACGCCAGGCAGGGGCAGATGGGCAGCGCCGATGCTTCCCGAAGGTGGGAGGCCGCAGTCTGAACAGGTCCCGAGGGCTGCTAAGCGTCCCGCCAAAGCCCAATCTGGCCACCTCCAGGTCTCTGGTCCCCACCACCCATCCAGGTT Contains chr22:27895001:27895200 (SEQ ID NO:30) and its upstream and / or downstream sequences within 5 kb; SEQ ID NO:30 is shown below: CAACTGCCCTTCTGCAGTGTCCCAGGCCTAAGGTCAGAGCAGGGTACAAAGGGTTAATCCCCATGGGCAAGGTGAAGGGAGCCCTGTGTGCCGTTTTCTGACCCTGGTCTCCACAGCCAGAGCCCCGGAGGGAAGAGCCCCATGTTTGTGTTGGCACCTGGGCCAGACGGTGGGAGGAACCGGCACTCAGTGGGGCCTGG.
[0032] In some embodiments, one or more target markers described herein include: sequences containing chr3:129299001:129299200 (SEQ ID NO:3) and within 5 kb upstream and / or 5 kb downstream; sequences containing chr6:31148401:31148600 (SEQ ID NO:6) and within 5 kb upstream and / or 5 kb downstream; sequences containing chr11:15963001:15963200 (SEQ ID NO:10) and within 5 kb upstream and / or 5 kb downstream; sequences containing chr11:34176801:34177000 (SEQ ID NO:11) and within 5 kb upstream and / or 5 kb downstream; sequences containing chr12:121437201:121437400 (SEQ ID NO:3); sequences containing chr12:121437201:121437400 (SEQ ID NO:3); sequences containing chr11:129299001:1292992 ... Sequences containing chr13 (SEQ ID NO:13) and its upstream and / or downstream within 5 kb; sequences containing chr13:29060601:29060800 (SEQ ID NO:15) and its upstream and / or downstream within 5 kb; sequences containing chr16:5198401:5198600 (SEQ ID NO:17) and its upstream and / or downstream within 5 kb; sequences containing chr16:87904401:87904600 (SEQ ID NO:21) and its upstream and / or downstream within 5 kb; sequences containing chr17:4560001:4560200 (SEQ ID NO:22) and its upstream and / or downstream within 5 kb; sequences containing chr22:27895001:27895200 (SEQ ID NO:13) and its upstream and / or downstream within 5 kb; sequences containing chr22:27895001:27895200 (SEQ ID NO:15) and its upstream and / or downstream within 5 kb; sequences containing chr16:5198401:5198600 (SEQ ID NO:17 ...22:27895001:27895200 (SEQ ID NO:30) and its upstream and / or downstream sequences within 5kb.
[0033] In some embodiments, one or more target markers described herein include: sequences containing chr3:129299001:129299200 (SEQ ID NO:3) and within 5 kb upstream and / or 5 kb downstream; sequences containing chr11:15963001:15963200 (SEQ ID NO:10) and within 5 kb upstream and / or 5 kb downstream; sequences containing chr15:75470601:75470800 (SEQ ID NO:16) and within 5 kb upstream and / or 5 kb downstream; sequences containing chr17:4560001:4560200 (SEQ ID NO:22) and within 5 kb upstream and / or 5 kb downstream; sequences containing chr22:27895001:27895200 (SEQ ID NO:3) and within 5 kb upstream and / or downstream; sequences containing chr22:27895001:27895200 (SEQ ID NO:3) and within 5 kb upstream and / or downstream; sequences containing chr11:15963001:15963200 (SEQ ID NO:10) and within 5 kb upstream and / or downstream; sequences containing chr22:27895001:27895200 (SEQ ID NO:10) and within 5 kb upstream and / or downstream; sequences containing chr11:15963001:15963 ...22:27895001:27895200 (SEQ ID NO:10) and within NO:30) and its upstream and / or downstream sequences within 5kb.
[0034] The specific nucleotide sequences of the above Hg19 coordinates, as well as the upstream 5kb and downstream 5kb of each start site and each end site of each region, can be obtained from public databases such as UCSC Genome Browser, Ensemble, and NCBI website.
[0035] The target markers of the present invention also include the corresponding regions obtained after non-enzymatic conversion (such as bisulfite conversion) and the corresponding regions obtained after enzymatic conversion (such as MSRE conversion).
[0036] In some embodiments, the target markers of the present invention also include various variants of the aforementioned genes. Variants include nucleic acid sequences from the same region that have at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the genes or regions described herein (i.e., having one or more deletions, insertions, substitutions, reverse sequences, etc.). Therefore, the content of this application should be understood to extend to such variants that achieve the same results, although in reality, there is minor genetic variation in the actual nucleic acid sequences between individuals.
[0037] As used herein, the term "percentage of sequence identity (%)" refers to the percentage of amino acid (or nucleic acid) residues identical to those in a candidate sequence and a reference sequence after sequence alignment. Spacing may be introduced during alignment (if necessary) to maximize the number of identical amino acids (or nucleic acids). In other words, the percentage of sequence identity (%) of an amino acid sequence (or nucleic acid sequence) can be calculated by dividing the number of identical amino acid residues (or bases) in the reference sequence by the total number of amino acid residues (or bases) in either the candidate or reference sequence (whichever is shorter). Conserved substitutions of amino acid residues may or may not be considered identical residues. The percentage of amino acid (or nucleic acid) sequence identity can be determined using publicly available tools such as BLASTN, BLASTp (available on the website of the National Center for Biotechnology Information (NCBI), see Altschul SF et al., J. Mol.Biol., 215:403–410 (1990); Stephen F. et al., Nucleic Acids Res., 25:3389–3402 (1997)), ClustalW2 (available on the website of the European Institute of Bioinformatics), see Higgins D.G. et al., Methods in Enzymology, 266:383-402 (1996); Larkin MA et al., Bioinformatics (Oxford, England), 23(21): 2947-8 (2007)) and ALIGN or Megalign (DNASTAR) software. Those skilled in the art can use the default parameters provided by the tool, or can (e.g., by selecting a suitable algorithm) customize the parameters to suit the comparison.
[0038] The target markers of the present invention also include the corresponding regions 5kb upstream of the start site and 5kb downstream of the end site of the above-mentioned genes after non-enzymatic transformation (such as bisulfite transformation) or after enzymatic treatment (such as methylation-sensitive restriction enzyme treatment).
[0039] II. Sources and preparation of target biomarkers In this document, the target biomarker can be derived from a biological sample of any individual of interest. The term "individual" as used herein includes both human and non-human animals. Non-human animals include all vertebrates, such as mammals and non-mammals. An "individual" can also be livestock, such as cattle, pigs, sheep, poultry, and horses; or rodents, such as rats and mice; or non-human primates, such as apes, monkeys, and rhesus monkeys; or domesticated animals, such as dogs or cats. In some embodiments, the individual is a human or a non-human primate. In some embodiments, the individual is a human. In this application, "individual," "object," and "subject" are used interchangeably.
[0040] It should be understood that the sequences given in Part I above are human sequences. When non-human animal sequences are involved, the corresponding positions and sequences of the above genes in the non-human animal genomes can be easily determined using existing techniques.
[0041] As used herein, the term "biological sample" refers to a biological composition obtained or derived from an individual, comprising cells and / or other molecular entities (e.g., DNA) to be characterized or identified based on physical, biochemical, chemical, and / or physiological characteristics. Biological samples include, but are not limited to, cells, tissues, organs, and / or bodily fluids of an individual obtained by any method known to those skilled in the art. In some embodiments, the biological sample is selected from the group consisting of: histological sections, tissue biopsies, paraffin-embedded tissue, bodily fluids, surgically excised samples, isolated blood cells, cells isolated from blood, and any combination thereof. In some embodiments, the bodily fluid is selected from the group consisting of: whole blood, serum, plasma, and any combination thereof. The most suitable sample will be selected depending on the nature of the situation. In some embodiments, the biological sample is whole blood of an individual. In some embodiments, the biological sample is plasma of an individual. Various methods for preparing plasma from whole blood are known to those skilled in the art. For example, in some embodiments, plasma is obtained by centrifuging whole blood from an individual once, twice, three times, four times, five times, or more. In some embodiments, the biological sample is a thyroid cancer biopsy.
[0042] The DNA to be detected can be isolated from the biological sample. The DNA to be detected can be isolated and purified from the biological sample using various methods known in the art. Commercially available kits can be used for isolation and purification. For example, DNA can be isolated from cells and tissues by lysing the raw material under highly denaturing and reducing conditions, partially using protein-degrading enzymes, purifying nucleic acid components obtained by a phenol / chloroform extraction process, and recovering the nucleic acids from the aqueous phase by dialysis or ethanol precipitation (see, for example, Sambrook, J., Fritsch, EF in T. Maniatis, CSH, MolecularCloning, 1989). Furthermore, many reagent systems are now particularly suitable for purifying DNA fragments from agarose gels, isolating plasmid DNA from bacterial lysates, and isolating longer chains of nucleic acids (genomic DNA, total cellular RNA) from blood, tissue, or cell cultures. Many of these commercially available purification systems are based on the well-known principle of combining nucleic acids with a mineral carrier in the presence of solutions of different dissociative salts. In these systems, suspensions of finely ground glass powder, diatomaceous earth, or silica gel are used as carrier materials. Several other methods for isolating and purifying DNA from biological samples are described, for example, in US7888006B2 and EP1626085A1. The choice between methods will be influenced by several factors, including time, cost, and the amount of DNA required.
[0043] In some embodiments, the DNA contained in the biological sample includes genomic DNA. As used herein, the term “genomic DNA” refers to DNA comprising the complete genome of a cell or organism, as well as segments or portions thereof. Genomic DNA is a large segment of DNA derived from an individual (e.g., longer than approximately 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, or 300 kb) and may have natural modifications, such as DNA methylation.
[0044] In some implementations, the DNA contained in the biological sample includes cellular DNA. As used herein, the term "cellular DNA" refers to DNA present in cells, or DNA obtained from cells in vivo and isolated in vitro, or otherwise manipulated in vitro, provided that the DNA has not been removed from cells in vivo.
[0045] In some embodiments, the DNA contained in the biological sample includes extracellular cell-free DNA. As used herein, the term "extracellular cell-free DNA" refers to a DNA fragment present outside the cells of the body. This term may also be used to refer to DNA fragments obtained from extracellular sources within the body and isolated or manipulated in vitro. DNA fragments in extracellular cell-free DNA typically have a length of about 100 to 200 bp, presumably related to the length of the DNA fragment encapsulated in a nucleosome. Extracellular cell-free DNA (cfDNA) includes, for example, extracellular cell-free fetal DNA and circulating tumor DNA. Extracellular cell-free fetal DNA circulates in the body of a pregnant woman (e.g., in the blood) and represents the fetal genome, while circulating tumor DNA circulates in the body of a cancer patient (e.g., in the blood). In some embodiments, extracellular cell-free DNA may be substantially free of the individual's cellular DNA. For example, the extracellular cell-free DNA may contain less than about 1,000 ng / mL, less than about 100 ng / mL, less than about 10 ng / mL, or less than about 1 ng / mL of cellular DNA.
[0046] Extracellular cell-free DNA can be prepared using conventional techniques known in the art. For example, extracellular cell-free DNA in a blood sample can be obtained by centrifuging the blood sample at speeds of about 200-20,000 g, about 200-10,000 g, about 200-5,000 g, about 300-4000 g, etc., for about 3-30 minutes, about 3-15 minutes, about 3-10 minutes, about 3-5 minutes, etc. For example, in some embodiments, extracellular cell-free DNA in a blood sample can be obtained by centrifuging an individual's plasma or serum one, two, three, four, five, or more times. In some embodiments, in order to separate cells and their fragments from a cell-free fraction containing soluble DNA, the biological sample can be obtained by microfiltration. Generally, microfiltration can be performed using filters, such as membrane filters of 0.1 micrometers to 0.45 micrometers, such as 0.22 micrometer membrane filters.
[0047] In some implementations, commercially available DNA extraction products are used to extract extracellular cell-free DNA from whole blood, serum, or plasma for analysis. This extraction method is claimed to have high recovery rates for circulating DNA (>50%), and certain products (such as Qiagen's QIAamp Circulating Nucleic Acid Kit) are claimed to extract small DNA fragments. Typical sample volumes used are 1–5 mL of serum or plasma.
[0048] In some implementations, extracellular cell-free DNA includes circulating tumor DNA (“ctDNA”). Circulating tumor DNA (“ctDNA”) is fragmented DNA of tumor origin found in bodily fluids (e.g., blood, urine, saliva, sputum, feces, pleural fluid, cerebrospinal fluid, etc.) that are not cellularly related. Typically, ctDNA is highly fragmented, with an average length of about 150 base pairs. ctDNA typically comprises a very small fraction of extracellular cell-free DNA in bodily fluids (e.g., plasma); for example, ctDNA may constitute less than about 10% of plasma DNA. Typically, this percentage is less than about 1%, for example, less than about 0.5% or less than about 0.01%. Furthermore, the total amount of plasma DNA is typically very low, for example, about 10 ng / mL of plasma. The amount of ctDNA varies from person to person and depends on the type and location of the tumor, and for cancerous tumors, on the stage of cancer. However, ctDNA is usually very rare in bodily fluids and can only be detected using extremely sensitive and specific techniques. Detection of ctDNA can help detect and diagnose tumors, guide tumor-specific treatments, monitor treatment, and monitor cancer remission.
[0049] III. Base Transformation In this article, DNA methylation is the biological process by which a methyl group is added to a DNA molecule (e.g., to one or more cytosine bases) (e.g., through the action of DNA methyltransferases). In mammals, DNA methylation occurs at the 5' position of the cytosine-phosphate-guanine (CpG) dinucleotide (i.e., the "CpG site"), and its presence at the 5'-CpG-3' dinucleotide in the promoter or first exon of a gene leads to epigenetic inactivation. DNA methylation has been well-demonstrated to play a crucial role in regulating gene expression, tumorigenesis, and other genetic and epigenetic diseases.
[0050] As used herein, the term "methylated cytosine residue" refers to a derivative of a cytosine residue in which a methyl group is attached to a carbon atom of the cytosine ring (e.g., C5). The term "unmethylated cytosine residue" refers to an underived cytosine residue, in which, unlike "methylated cytosine residue," there is no methyl group attached to a carbon atom of the cytosine ring (e.g., C5). The CpG site where the cytosine residue is methylated is the methylated CpG site, and the CpG site where the cytosine residue is not methylated is the unmethylated CpG site.
[0051] As described herein, bases in DNA or RNA can be transformed. The terms "transformation," "cytosine transformation," or "CT transformation" refer to the process of treating DNA using non-enzymatic or enzymatic methods to convert unmodified cytosine (C) bases into bases that do not bind to guanine (G) (e.g., uracil (U)). Some reagents can distinguish between unmethylated and methylated CpG sites in DNA, thus obtaining treated DNA. These reagents can selectively act on unmethylated cytosine residues but not significantly on methylated cytosine residues. Alternatively, they can selectively act on methylated cytosine residues without significantly affecting unmethylated ones. For example, some reagents can selectively convert unmethylated cytosine residues into uracil, thymine, or another base that hybridizes differently from cytosine, while methylated cytosine residues remain unconverted; others can selectively cleave methylated residues or selectively cleave unmethylated residues. Thus, the original DNA is transformed into treated DNA in a manner that depends on whether it is methylated, thereby allowing the treated DNA to be distinguished from the original DNA through its hybridization behavior.
[0052] As used in this article, “processed DNA,” “processed sequence,” and “processed fragment” refer to DNA, nucleic acid sequences, and gene fragments that have been treated with reagents capable of distinguishing between unmethylated and methylated CpG sites in DNA, nucleic acid sequences, and gene fragments.
[0053] More specifically, cytosine conversion can be performed using either non-enzymatic or enzymatic methods. Exemplarily, non-enzymatic methods include bisulfite or bisulfite treatment. In some embodiments, the reagents used in non-enzymatic methods include bisulfite reagents. As used herein, the term "bisulfite reagent" refers to, for example, reagents disclosed herein that can be used to distinguish between methylated and unmethylated CpG dinucleotide sequences, including bisulfite, bisulfite ions, or any combination thereof. In this application, treatment of DNA with bisulfite reagents is also described as a "bisulfite reaction" or "bisulfite treatment," referring to a reaction that converts unmethylated cytosine residues, particularly in the presence of bisulfite ions, where unmethylated cytosine residues in nucleic acids are converted to uracil bases, thymine bases, or other bases that differ from cytosine in hybridization behavior, while methylated cytosine residues are not significantly converted. In other words, bisulfite treatment can be used to distinguish between methylated and unmethylated CpG dinucleotides. The bisulfite reaction for detecting methylated cytosine residues is described in detail in Frommer, M., et al., Proc Natl Acad Sci USA 89 (1992) 1827-31 and Grigg, G., Clark, S., Bioessays 16 (1994) 431-6. The bisulfite reaction includes a deamination step and a desulfonate step (see Grigg and Clark, ibid.). The statement “the methylated cytosine residues were not significantly converted” does not preclude the possibility that a very small percentage (e.g., less than 0.1%, less than 0.2%, less than 0.3%, less than 0.4%, less than 0.5%, less than 0.6%, less than 0.7%, less than 0.8%, less than 0.9%, less than 1%, less than 2%, less than 3%, less than 4%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, less than 10%, less than 11%, less than 12%, less than 13%, less than 14%, less than 15%, less than 16%, less than 17%, less than 18%, less than 19%, less than 20%) of the methylated cytosine residues were converted to uracil, thymine, or other bases that do not hybridize with cytosine, even though the intention is to convert only the unmethylated cytosine residues.
[0054] In cases such as those described in Frommer M., et al. (ibid.) or Grigg and Clark (ibid.), which disclose basic parameters for bisulfite treatment, those skilled in the art know how to perform bisulfite treatment, particularly the deamination and desulfonation steps. The effects of incubation time and temperature on deamination efficiency, as well as parameters affecting DNA degradation, have been disclosed.
[0055] In some embodiments, the bisulfite reagent is selected from the group consisting of ammonium bisulfite, sodium bisulfite, potassium bisulfite, calcium bisulfite, magnesium bisulfite, aluminum bisulfite, bisulfite ions, and any combination thereof. In some embodiments, the bisulfite reagent is sodium bisulfite. In some embodiments, the bisulfite reagent is commercially available, such as MethylCode™ Bisulfite Conversion Kit, EpiMark™ Bisulfite Conversion Kit, EpiJET™ Bisulfite Conversion Kit, EZDNAMethylation-Gold™ Kit, etc. In some embodiments, the bisulfite reaction is performed according to the kit's instructions.
[0056] Exemplary enzymatic methods include deaminase treatment, and selectively cleaving unmethylated residues without cleaving methylated residues, or selectively cleaving methylated residues without cleaving unmethylated residues, using a reagent. Preferably, the reagent is a methylation-sensitive restriction enzyme (MSRE).
[0057] The term "methylation-sensitive restriction enzyme" refers to an enzyme that selectively digests nucleic acids based on the methylation state of its recognition site. For restriction enzymes that specifically cleave when the recognition site is unmethylated or hemimethylated, cleavage does not occur or occurs with significantly reduced efficiency when the recognition site is methylated. In some embodiments, the recognition sequence of the methylation-sensitive restriction enzyme contains a CG dinucleotide (e.g., cgcg or cccggg). In some embodiments, the methylation-sensitive restriction enzyme does not cleave when the cytosine in the CG dinucleotide is methylated at the C5 carbon atom.
[0058] Exemplary MSREs are selected from the group consisting of: HpaII enzyme, SalI enzyme, SalI-HF® enzyme, ScrFI enzyme, BbeI enzyme, NotI enzyme, SmaI enzyme, XmaI enzyme, MboI enzyme, BstBI enzyme, ClaI enzyme, MluI enzyme, NaeI enzyme, NarI enzyme, PvuI enzyme, SacII enzyme, HhaI enzyme, and any combination thereof.
[0059] Using methods known in the art, methylation is determined using a methylation-sensitive restriction enzyme or a series of restriction enzyme reagents containing a methylation-sensitive restriction enzyme that can distinguish between methylated and unmethylated CpG dinucleotides in the target region, such as, but not limited to, differential methylation hybridization (“DMH”).
[0060] In some embodiments, DNA in a biological sample can be cleaved prior to treatment with a methylation-sensitive restriction enzyme. Such methods are known in the art and can include both physical and enzymatic approaches. Particularly preferred are the use of one or more restriction enzymes that are insensitive to methylation, have AT-rich recognition sites, and do not contain CG dinucleotides. Using such enzymes preserves CpG sites and CpG-rich regions in the DNA fragment. In some embodiments, such restriction enzymes are selected from MseI, BfaI, Csp6I15, Tru1I, Tru9I, MaeI, XspI, and any combination thereof.
[0061] The transformed DNA may optionally be purified. The DNA purification methods applicable to this paper are well known in the art.
[0062] IV. Quantitative Analysis This invention can detect the methylation status or level of at least one CpG dinucleotide among any 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or all 30 target biomarkers described herein, to identify whether the subject has an invasive subtype of papillary thyroid carcinoma. The detection reagents and diagnostic kits described in this invention can be used for the detection of the methylation status or methylation level.
[0063] In this document, "methylation state" refers to the presence or absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule containing methylated cytosine is considered methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered unmethylated. In some embodiments, a nucleic acid may be characterized as "unmethylated" if it is not methylated at a specific locus (e.g., a locus containing a specific single CpG dinucleotide) or a specific combination of loci, even if it is methylated at other loci of the same gene or molecule.
[0064] Therefore, methylation status describes the state of methylation of nucleic acids (e.g., genomic sequences or target biomarkers described herein). Furthermore, methylation status refers to a characteristic associated with methylation in a segment of nucleic acid at a specific genomic locus. Such characteristics include, but are not limited to, whether any cytosine (C) residues within this DNA sequence are methylated, the location of one or more methylated C residues, the frequency or percentage of methylated C throughout any particular region of the nucleic acid, and methylation allele differences due to, for example, differences in allele origins. "Methylation status" refers to the relative, absolute, or pattern of methylated or unmethylated C throughout any particular region of a nucleic acid in a biological sample. For example, if one or more cytosine (C) residues within a nucleic acid sequence are methylated, it may be termed "hypermethylated" or has "increased methylation," while if one or more cytosine (C) residues within a DNA sequence are unmethylated, it may be termed "demethylated" or has "decreased methylation." Similarly, if one or more cytosine (C) residues within a nucleic acid sequence are methylated compared to another nucleic acid sequence (e.g., from a different region or from a different individual), the sequence is considered hypermethylated or has increased methylation compared to other nucleic acid sequences. Alternatively, if one or more cytosine (C) residues within a DNA sequence are unmethylated compared to another nucleic acid sequence (e.g., from a different region or from a different individual), the sequence is considered demethylated or has decreased methylation compared to other nucleic acid sequences.
[0065] Methylation level represents the proportion (or percentage, fraction, ratio, degree) of one or more sites that are methylated. The methylation level of a region (or a group of sites) is the mean of the methylation levels of all sites in that region (or all sites in the group). Therefore, an increase or decrease in the methylation level of a region does not indicate that the methylation level of all methylated sites in that region has increased or decreased. The process of converting the results of methods for detecting DNA methylation (e.g., simplified methylation sequencing) into methylation levels is known in the art. Methylation level can be determined, for example, by quantitative analysis of the amount of intact DNA present after restriction digestion with a methylation-sensitive restriction enzyme. In this example, if a specific sequence in DNA is quantitatively analyzed using quantitative PCR, an amount of template DNA approximately equal to that in a simulated control indicates that the sequence is not highly methylated, while a significantly less template amount than in a simulated sample indicates the presence of methylated DNA in the sequence. Therefore, methylation level, as in the example above, can be used as a quantitative indicator of methylation status. This is particularly useful when it is necessary to compare the methylation level of a sequence in a sample with a threshold level.
[0066] The methylation level / state of one or more CpG dinucleotide sequences within a DNA sequence (e.g., a target marker) can be determined by various analytical methods known in the art, preferably quantitative methods. Exemplary analytical methods include: polymerase chain reaction, including real-time polymerase chain reaction, digital polymerase chain reaction, and bisulfite-conversion-based PCR (e.g., methylation-specific PCR (MSP)) and sequences within 5 kb upstream and / or downstream; nucleic acid sequencing; whole-genome methylation sequencing (RRBS) and sequences within 5 kb upstream and / or downstream; simplified methylation sequencing; mass-based separation (e.g., electrophoresis, mass spectrometry) and sequences within 5 kb upstream and / or downstream; target capture (e.g., hybridization, microarray) and sequences within 5 kb upstream and / or downstream; methylation-sensitive restriction endonuclease analysis; methylation-sensitive high-resolution melting curve analysis; chip-based methylation mapping analysis; mass spectrometry; and quantitative fluorescence methods. In this article, detection includes detecting any one strand at a gene or locus.
[0067] In some implementations, quantitative analysis is performed by real-time PCR. Non-limiting examples of real-time PCR include HeavyMethyl™ PCR described by Cottrell et al., Nucl. Acids Res. 32: e10, 2003; MethyLight™ PCR described by Eads et al., Cancer Res. 59:2302-2306, 1999; and Headloop PCR described by Rand et al., Nucl. Acids Res. 33: e 127, 2005.
[0068] As used herein, the term "HeavyMethyl™ PCR" refers to a real-time PCR technique recognized in the art in which one or more non-extendable nucleic acid (e.g., oligonucleotide) blocks bind to bisulfite-treated nucleic acids in a methylation-specific manner (i.e., the blocks bind specifically to unmutated DNA under moderate to high stringency conditions). An amplification reaction is performed using one or more primers, which may optionally be methylation-specific but with one or more blocks distributed alongside them. In the presence of unmethylated nucleic acids (i.e., mutated DNA), the blocks bind and no PCR product is produced. The methylation level of nucleic acids in the sample is determined using a TaqMan™ analytical method, substantially as described, for example, by Holland et al., Proc. Natl. Acad. Sci. USA, 88:7276-7280, 1991.
[0069] As used herein, the term "MethyLight™ PCR" refers to a fluorescence-based real-time PCR technique recognized in the art, employing a dual-labeled fluorescent oligonucleotide probe called a TaqMan™ probe, designed to hybridize with CpG-rich sequences located between forward and reverse amplification primers. The TaqMan™ probe comprises a fluorescent "reporter portion" and a "quencher portion" covalently bound to a linker portion (e.g., phosphoramide) linked to the nucleotides of the TaqMan™ oligonucleotide. During PCR amplification, the TaqMan™ probe hybridized to the CpG-rich sequence is cleaved by the 5' nuclease activity of Taq polymerase, thereby generating a signal detectable in real-time during the PCR reaction. In this method, molecular beacons can be used as detectable probes, and the system is independent of the 5'-3' exonuclease activity of the DNA polymerase used (see Mhlanga and Malmberg, Methods 25:463-471, 2001).
[0070] As used herein, the term “Headloop PCR” refers to a type of real-time PCR recognized in the art that selectively amplifies target nucleic acids but suppresses the amplification of non-target variants by extending the 3' stem loop to form a hairpin structure that cannot further provide amplification template.
[0071] In some embodiments, the real-time PCR is multiplex real-time PCR. As used herein, the term "multiplex" can refer to an analysis or other analytical method that can simultaneously determine the presence and / or amount of multiple markers (e.g., multiple nucleic acid sequences) by using more than one marker, each marker having at least one different detection feature, such as fluorescence features (e.g., excitation wavelength, emission wavelength, emission intensity, FWHM (full width at half maximum) or fluorescence lifetime) or unique nucleic acid or protein sequence features.
[0072] In some implementations, quantitative analysis is performed by nucleic acid sequencing. Exemplary methods of nucleic acid sequencing are known in the art; see, for example, Frommer et al., Proc. Natl. Acad. Sci. USA 89:1827-1831, 1992; Clark et al., Nucl. Acids Res. 22:2990-2997, 1994. For example, comparing sequences obtained from untreated samples or known nucleotide sequences of target regions with sequences obtained from bisulfite-treated samples can help identify methylated cytosine in DNA sequences. Thymine residues detected at any cytosine site in a bisulfite-treated sample, compared to an untreated sample, can be considered mutations caused by bisulfite treatment, i.e., the presence of methylated cytosine at that site.
[0073] Methods for sequencing DNA are known in the art, including, for example, dideoxy chain termination or Maxam-Gilbert methods (see Sambrook et al., Molecular Cloning, A Laboratory Manual (2nd Ed., CSHP, New York 1989)), pyrosequencing (see Uhlmann et al., Electrophoresis, 23:4072-4079, 2002), solid-phase pyrosequencing (see Landegren et al., Genome Res., 8(8):769-776, 1998), solid-phase microsequencing (see, for example, Southern et al., Genomics, 13:1008-1017, 1992), microsequencing using FRET (see, for example, Chen and Kwok, Nucleic Acids Res. 25:347-353, 1997), ligation sequencing, or ultra-deep sequencing (see Marguiles et al., Nature 437). (7057):376-80 (2005)).
[0074] In some implementations, quantitative analysis is performed by mass-based separation (e.g., electrophoresis, mass spectrometry). For example, the presence of methylated cytosine residues can be detected by combined bisulfite restriction assay (COBRA), essentially as described in Xiong and Laird, Nucl. Acids Res., 25:2532-2534, 2001. This method utilizes the difference in restriction enzyme recognition sites between methylated and unmethylated nucleic acids after treatment with a compound that can selectively mutate unmethylated cytosine residues (e.g., bisulfite). For example, the restriction endonuclease Taq1 cleaves the sequence TCGA, which, after bisulfite treatment of unmethylated nucleic acids, will be TTGA and therefore will not be cleaved. The digested and / or undigested nucleic acids are then detected using detection techniques known in the art, such as electrophoresis and / or mass spectrometry. For example, after treatment with compounds that selectively mutate unmethylated cytosine residues, different techniques are used to detect differences in nucleic acids in the amplified products based on differences in nucleotide sequence and / or secondary structure, such as methylation-specific single-strand conformation analysis (MS-SSCA) (Bianco et al., Hum. Mutat., 14:289-293, 1999), methylation-specific denaturing gradient gel electrophoresis (MS-DGGE) (Abrams and Stanton, Methods Enzymol., 212:71-74, 1992), and methylation-specific denaturing high-performance liquid chromatography (MS-DHPLC) (Deng et al., Chin. J. Cancer Res., 12:171-191, 2000).
[0075] In some embodiments, quantitative analysis is performed via target capture (e.g., hybridization, microarrays). Suitable detection methods via hybridization are known in the art, such as Southern blotting, dot blotting, slit blotting, or other nucleic acid hybridization methods (Kawai et al., Mol. Cell. Biol. 14:7421-7427, 1994; Gonzalgo et al., CancerRes. 57:594-599, 1997). In some embodiments, the probes used for hybridization analysis are detectably labeled. In some embodiments, the nucleic acid-based probes used for hybridization analysis are unlabeled. Such unlabeled probes can be immobilized on a solid support such as a microarray and can hybridize with a detectably labeled target nucleic acid molecule. An example of a microarray is a methylation-specific microarray, which can be used to distinguish sequences with transformed cytosine residues from sequences with untransformed cytosine residues (see Adorjan et al., Nucl. Acids Res., 30: e21, 2002). Hybridization-based analysis can also be used for nucleic acids treated with methylation-sensitive restriction enzymes. For example, the methylation status of CpG dinucleotide sequences within a DNA sequence can be determined using oligonucleotide probes that hybridize simultaneously with bisulfite-treated DNA using PCR amplification primers (wherein the primers can be methylation-specific primers or standard primers).
[0076] In some embodiments, the quantitative analysis is performed in the presence of a detection reagent. As used herein, the term "detection reagent" is a reagent used in a quantitative analysis step to detect the presence, absence, or amount of nucleic acid. Various detection reagents known in the art may be used in this application. In some embodiments, the detection reagent is selected from the group consisting of: fluorescent probes, intercalating dyes, chromophore-labeled probes, radioisotope-labeled probes, and biotin-labeled probes.
[0077] In some embodiments, quantitative analysis includes amplifying the treated DNA using quantitative primer pairs and DNA polymerase. As used herein, the term "quantitative primer pair" refers to one or more primer pairs used in the quantitative analysis step. Preferably, the quantitative primer pair is capable of hybridizing with at least nine consecutive nucleotides of the treated DNA under stringent, moderately stringent, or highly stringent conditions.
[0078] In some embodiments, the quantitative analysis includes determining the methylation level of one or more target markers based on the presence or level of multiple CpG dinucleotides, TpG dinucleotides, or CpA dinucleotides in the treated DNA. In some embodiments, the quantitative analysis includes determining the methylation level of cytosine residues based on the presence or level of one or more CpG dinucleotides in the treated DNA. In some embodiments, the quantitative analysis includes determining the methylation level of cytosine residues based on the presence or level of one or more TpG dinucleotides in the treated DNA. In some embodiments, the quantitative analysis includes determining the methylation level of cytosine residues based on the presence of CpA dinucleotides in the treated DNA.
[0079] In some embodiments, the quantitative analysis step is performed by fractionating the treated DNA product into multiple fractions. In some embodiments, multiple different quantitative analytical tests are performed on the multiple fractions, wherein different combinations of the treated DNA product (if present in the fraction) are quantified in one of the fractions. In some embodiments, a control marker is quantified in each fraction.
[0080] In some implementations, the methylation level of each target marker is quantitatively analyzed separately based on the pre-amplified DNA using an MSP (see Herman, ibid.). For example, by using one or more primers that specifically hybridize to the unconverted sequence under moderate and / or highly stringent conditions, amplification products are generated only if the template contains methylated cytosine at the CpG site.
[0081] In some embodiments, the quantitative primer pairs are designed to amplify at least a portion of the treated DNA product; that is, the quantitative analysis is designed as nested PCR. Nested PCR is an improvement on PCR designed to increase sensitivity and specificity. Nested PCR involves the use of two primer sets and two consecutive PCR reactions. A first round of amplification is performed to produce a first amplicon, and a second round of amplification is performed using a primer pair, where one or both primers anneal to sites within the region defined by the initial primer pair; that is, the second primer pair is considered to be "nested" within the first primer pair. In this way, background amplification products from the first PCR reaction that do not contain the correct internal sequence are no longer further amplified in the second PCR reaction.
[0082] Typically, the PCR reaction solution contains Taq DNA polymerase, PCR buffer, primers, probes, dNTPs, and Mg. 2+ Preferably, the Taq DNA polymerase is a hot-start Taq DNA polymerase. Exemplarily, Mg... 2+The final concentration is 1.0-20.0 mM; the concentration of each primer is 100-500 nM; the concentration of each probe is 100-500 nM. An exemplary PCR reaction condition is: 95℃ pre-denaturation for 5 min; 95℃ denaturation for 15 s; 60℃ annealing and extension for 60 s, for 50 cycles.
[0083] In some embodiments, the method of the present invention includes a pre-amplification step. One purpose of pre-amplifying a target biomarker is to increase the quantity of the target biomarker in the treated DNA. As used herein, the term “amplification” generally refers to any process that results in an increase in the copy number of a molecule or a group of related molecules. When “amplification” is used for polynucleotide molecules, it refers to the generation of multiple copies of a polynucleotide molecule or a portion of a polynucleotide molecule, typically starting from a small number of polynucleotides, wherein the amplified material (amplifier, PCR amplifier) is generally detectable. Polynucleotide amplification encompasses a variety of chemical and enzymatic processes. Forms of amplification include generating multiple copies of DNA from one or more copies of a template RNA or DNA molecule via polymerase chain reaction (reverse transcription PCR, PCR), strand displacement amplification (SDA) reaction, transcription-mediated amplification (TMA) reaction, nucleic acid sequence-based amplification (NASBA) reaction, or ligase chain reaction (LCR).
[0084] The target marker in the treated DNA can be pre-amplified using pre-amplification primers. As used herein, the term "primer" refers to a single-stranded oligonucleotide that can serve as the starting point for template-guided DNA synthesis under suitable conditions (e.g., buffer and temperature) in the presence of four different nucleosides and a reagent for polymerization (e.g., DNA polymerase). In any given case, the length of the primer depends on, for example, the intended use of the primer and is typically in the range of 15 to 30 nucleotides. Short primer molecules generally require lower temperatures to form a sufficiently stable hybridization complex with the template. The primer does not need to reflect the exact sequence of the template, but must be complementary enough to hybridize with it. The primer site is the region on the template that hybridizes with the primer. A primer pair is a set of primers that includes a 5' forward primer that hybridizes to the 5' end of the sequence to be amplified and a 3' reverse primer that hybridizes to the complementary strand at the 3' end of the sequence to be amplified. Those skilled in the art can design primers based on common knowledge in the art regarding the marker to be amplified (see, for example, PCR Primer: A Laboratory Manual, Cold Spring Harbor Laboratories, NY, 1995). Furthermore, several software packages for designing optimal probes and / or primers for use in a wide variety of analyses are publicly available, such as Primer 3, which is available from the Center for Genome Research, Cambridge, Mass., USA. Obviously, their potential uses should also be considered when designing probes or primers. For example, primers designed for the purposes of this invention may include at least one CpG site, or the amplification product obtained from such primers may include at least one CpG site. Tools for designing primers for detecting DNA methylation status are also known in the art, such as MethPrimer (LiLC and Dahiya R. MethPrimer: designing primers for methylation PCRs. Bioinformatics. 2002 Nov;18(11):1427-31). In this application, by using pre-amplification primers as a primer pool, any target marker in the processed DNA (at least a portion of the target marker or a subregion of the target marker) can be pre-amplified.
[0085] As used herein, the term "complementary" refers to hybridization or base pairing between nucleotides or nucleic acids, such as between the two strands of a double-stranded DNA molecule, or between a primer binding site and an oligonucleotide primer on a single-stranded nucleic acid to be sequenced or amplified. Complementary nucleotides are typically A and T (or A and U), or C and G. Two single-stranded RNA or DNA molecules are said to be complementary when the nucleotides of one strand are optimally aligned, compared, and have appropriate nucleotide insertions or deletions, and pair with at least about 80% (typically at least about 90% to 95%, more preferably about 98% to 100%) of the nucleotides of the other strand. Alternatively, complementarity exists when an RNA or DNA strand hybridizes with its complementary sequence under selective hybridization conditions. Typically, selective hybridization occurs when there is at least about 65% (preferably at least about 75%, more preferably at least about 90%) complementarity on a segment of at least 14 to 25 nucleotides. See M. Kanehisa, Nucleic Acids Res. 12:203 (1984), which is incorporated herein by reference.
[0086] In some embodiments, the pre-amplification primer pool contains at least one methylation-specific primer pair. In some embodiments, the pre-amplification primer pool contains multiple methylation-specific primer pairs. In some embodiments, the pre-amplification step is performed by methylation-specific PCR (“MSP”), which is PCR using methylation-specific primers. This technique (i.e., MSP) has been described in Herman et al., Methylation-specific PCR: a novel PCR assay for methylation status of CpGislands. Proc Natl Acad Sci USA. 1996 September 3; 93 (18): 9821-6 and United States Patent No. 6,265,171.
[0087] As used herein, the term "methylation-specific primer pair" refers to a primer pair specifically designed to recognize CpG sites to amplify a specific target marker in treated DNA by utilizing differences in methylation. Primers act only on molecules with or without a specific methylation state. For example, primers may be oligonucleotides that, under stringent, moderately stringent, or highly stringent conditions, can specifically hybridize with a specific CpG site that is methylated in a methylation-specific manner, but cannot hybridize with a specific CpG site that is not methylated. Therefore, the primers will specifically amplify the target marker that is methylated at the specific CpG site. As another example, primers may be oligonucleotides that, under stringent, moderately stringent, or highly stringent conditions, can specifically hybridize with a specific CpG site that is not methylated in a methylation-specific manner, but cannot hybridize with a specific CpG site that is methylated. Therefore, the primers will specifically amplify the target marker that is not methylated at the specific CpG site. Therefore, in this application, the use of methylation-specific primers in the pre-amplification of at least one target marker in treated DNA allows for the differentiation between methylated and unmethylated CpG sites. The methylation-specific primer pairs of this application comprise at least one primer that hybridizes to a CpG dinucleotide treated with bisulfite. Therefore, the sequence of the primers specifically targeting methylated DNA contains at least one CpG dinucleotide, and the sequence of the primers specifically targeting unmethylated DNA contains a "T" at the C position of CpG, and / or an "A" at the G position of CpG.
[0088] Methylation-specific primer pairs typically include a forward primer and a reverse primer, both of which contain an oligonucleotide sequence that hybridizes with at least nine consecutive nucleotides of one of the target markers (or a subregion of the target marker) under stringent, moderately stringent, or highly stringent conditions, wherein the at least nine consecutive nucleotides of one of the target markers (or a subregion of the target marker) contain at least one (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) CpG site.
[0089] As used herein, the term "hybridization" can refer to the process by which two single-stranded polynucleotides bind non-covalently to form a stable double-stranded polynucleotide. In one aspect, the resulting double-stranded polynucleotide can be called a "hybrid" or a "double strand." The salt concentration in "hybridization conditions" is typically less than about 1 M, often less than about 500 mM, and can be less than about 200 mM. "Hybridization buffers" include buffered salt solutions, such as 5% SSPE, or other such buffers known in the art. Hybridization temperatures can be as low as 5 °C, but are typically above 22 °C, and more typically above about 30 °C, and often above 37 °C. Hybridization is typically performed under stringent conditions, i.e., conditions under which the sequence will hybridize with its target sequence but not with other non-complementary sequences. Stringent conditions depend on the sequence and vary under different conditions. For example, longer fragments may require higher hybridization temperatures for specific hybridization than shorter fragments. Because other factors can affect the stringency of hybridization, including base composition and the length of the complementary strand, the presence of organic solvents, and the degree of base mismatch, the combination of parameters is more important than an absolute measurement of any single parameter alone. The stringent conditions are typically chosen to be approximately 5°C lower than the melting temperature (Tm) of a particular sequence at a specific ionic strength and pH. Tm can be the temperature at which half of a double-stranded nucleic acid population separates into single strands. Several equations for calculating the Tm of nucleic acids are well known in the art. As shown in standard references, a simple estimate of the Tm value can be calculated using the formula Tm = 81.5 + 0.41 (%G + C) when the nucleic acid is in a 1M NaCl aqueous solution (see, for example, Anderson and Young, Quantitative Filter Hybridization, in Nucleic Acid Hybridization (1985)). Other references (e.g., Allai and SantaLucia, Jr., Biochemistry, 36:10581-94 (1997)) include alternative calculation methods that take into account structural and environmental factors, as well as sequence characteristics, when calculating Tm.
[0090] Typically, the stability of hybrids is a function of ion concentration and temperature. Hybridization reactions are usually carried out under relatively stringent conditions, followed by washing in wash solutions with varying but higher stringency. Exemplary stringent conditions include a pH of approximately 7.0 to approximately 8.3, a temperature of at least 25°C, and a sodium ion (or other salt) concentration of at least 0.01 M to no more than 1 M. For example, 5 × SSPE (750 mM NaCl, 50 mM sodium phosphate, 5 mM EDTA, pH 7.4) and a temperature of approximately 30°C are suitable for allele-specific hybridization, although suitable temperatures depend on the length of the hybridization region and / or GC content. In one aspect, “hybridization stringency” for determining the percentage of mismatches can be defined as follows: 1) High stringency: 0.1 × SSPE, 0.1% SDS, 65°C; 2) Moderate stringency (also known as medium stringency): 0.2 × SSPE, 0.1% SDS, 50°C; 3) Low stringency: 1.0 × SSPE, 0.1% SDS, 50°C. It should be understood that the same stringency can be achieved using alternative buffers, salts, and temperatures. For example, moderately stringent hybridization can refer to conditions that allow nucleic acid molecules (e.g., probes) to bind to complementary nucleic acid molecules. The hybridized nucleic acid molecules typically have at least 60% identity, including, for example, at least 70%, 75%, 80%, 85%, 90%, or 95% identity. Moderately stringent conditions can be equivalent to hybridization at 42°C, 50% formamide, 5× Denhardt solution, 5× SSPE, and 0.2% SDS, followed by washing at 42°C, 0.2× SSPE, and 0.2% SDS. Highly stringent conditions can be provided, for example, hybridization at 42°C, 50% formamide, 5× Denhardt solution, 5× SSPE, and 0.2% SDS, followed by washing at 65°C, 0.1× SSPE, and 0.1% SDS. Low-strictness hybridization can be performed under conditions equivalent to the following: hybridization at 22°C, 10% formamide, 5 × Denhardt solution, 6 × SSPE, 0.2% SDS, followed by washing at 37°C in 1 × SSPE, 0.2% SDS. The Denhardt solution contains 1% sucrose, 1% polyvinylpyrrolidone, and 1% bovine serum albumin (BSA). 20 × SSPE (sodium chloride, sodium phosphate, EDTA) contains 3 M sodium chloride, 0.2 M sodium phosphate, and 0.025 M EDTA.Other suitable moderately and highly stringent hybridization buffers and conditions are well known to those skilled in the art and are described, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, Plainview, NY (1989) and Ausubelet et al., Short Protocols in Molecular Biology, 4th ed., John Wiley & Sons (1999).
[0091] In some implementations, the pre-amplification primer pool also includes control primer pairs for amplifying control biomarkers. Typically, control biomarkers are nucleic acids with known characteristics (e.g., known sequence, known copy number per cell) used for comparison with the experimental target (e.g., nucleic acid at unknown concentration). Controls can be endogenous, preferably invariant genes, which can be normalized to the experimental or target nucleic acids in the analysis. Such controls, normalized due to inter-sample variability, may occur for example, in sample processing, analytical efficiency, etc., and allow for precise inter-sample data comparisons, quantitative analysis of amplification efficiency and bias.
[0092] In some embodiments, the present invention employs RRBS technology to detect the methylation level of the CpG site of the target biomarker of interest, and then calculates the methylation level of the methylated region (DMR) of the biomarker, which is used as the DNA methylation level of the biomarker. The calculation of DMR can be performed as described in this application.
[0093] V. Identification of whether the subject has an invasive subtype of papillary thyroid carcinoma. This invention has discovered that the methylation levels of one or more target biomarkers described herein can be used to determine the presence of an invasive subtype of papillary thyroid carcinoma. In one or more embodiments, the methylation level of CpG sites in the target biomarkers described herein can be detected in a sample, and then the methylation level of the methylated region (DMR) of the biomarker can be calculated as the DNA methylation level of the biomarker. In this study, genomic DNA was divided into continuous 200 bp windows (step size 200 bp, meaning adjacent windows are joined end-to-end without overlap), and each window may contain several CpG sites. The methylation levels (e.g., methylation ratios) of all CpG sites within each 200 bp window were integrated to calculate the average methylation level of that window, which is the DMR level.
[0094] It should be understood that when using two or more target biomarkers, each sample can have its own DMR calculated from the methylation level of the CpG sites in each detected target biomarker. In the training set samples, the parameters of the above-mentioned prediction model formula are obtained by training with the DMRs of the two or more target biomarkers obtained from all samples. For the test sample, the prediction model score y is obtained by substituting the calculated DMR of the sample into the formula of the prediction model determined by the training set, then substituting that DMR into the above-mentioned prediction model formula obtained by training with the DMRs of the two or more target biomarkers in the training set and optionally validating with the validation set. This y is then compared with a threshold defined by the Youden index obtained from the two or more target biomarkers in the training set. If the score is higher than this threshold, it is judged to be an invasive subtype of papillary thyroid carcinoma, or a risk of having a highly invasive subtype.
[0095] VI. Compositions and Kits This invention provides a methylation detection or diagnostic kit and diagnostic reagent or composition for subtype identification, wherein the kit and composition comprise a reagent for detecting the methylation state or level of at least one CpG dinucleotide of one or more target biomarkers described herein. Depending on the target biomarker to be detected, the kit and composition may contain primers and / or probe molecules. Preferably, the primers comprise primer pairs capable of hybridizing with the target biomarker to be detected or its target region under stringent, moderately stringent, or highly stringent conditions. The primers may also include primers for detecting internal controls such as ACTB.
[0096] In some embodiments, the primers are packaged in a single container or in separate containers. In some embodiments, the kit further comprises one or more blocking oligonucleotides.
[0097] In some embodiments, the kit and composition further comprise detection reagents. In some embodiments, the detection reagents are selected from the group consisting of: fluorescent probes, probes labeled with embedded dyes or chromophores, probes labeled with radioisotopes, and biotin-labeled probes.
[0098] In some embodiments, the kit may also include DNA polymerase and / or a container suitable for storing biological samples obtained from an individual. In some embodiments, the kit further includes instructions for use and / or an interpretation of the kit's test results.
[0099] In some embodiments, the kit and composition may further include reagents for enzymatic or non-enzymatic conversion. In a preferred embodiment, the kit further includes a bisulfite reagent or a methylation-sensitive restriction enzyme (MSRE). In some embodiments, the bisulfite reagent is selected from the group consisting of ammonium bisulfite, sodium bisulfite, potassium bisulfite, calcium bisulfite, magnesium bisulfite, aluminum bisulfite, bisulfite ions, and any combination thereof. In some embodiments, the bisulfite reagent is sodium bisulfite. In some embodiments, the MSRE is selected from the group consisting of HpaII enzyme, SalI enzyme, SalI-HF® enzyme, ScrFI enzyme, BbeI enzyme, NotI enzyme, SmaI enzyme, XmaI enzyme, MboI enzyme, BstBI enzyme, ClaI enzyme, MluI enzyme, NaeI enzyme, NarI enzyme, PvuI enzyme, SacII enzyme, HhaI enzyme, and any combination thereof.
[0100] The kit and composition may also include a converted positive standard, wherein unmethylated cytosine is converted into a base that does not bind to guanine. The positive standard may be fully methylated.
[0101] The kit and composition may further include PCR reaction reagents. Preferably, the PCR reaction reagents include Taq DNA polymerase, PCR buffer, dNTPs, and Mg2+. 2+ .
[0102] In some embodiments, the kits and compositions further comprise standard reagents for performing CpG position-specific methylation assays, wherein the assays include one or more of the following techniques: MS-SNuPE, MSP, MethyLight™, HeavyMethyl™, COBRA, and nucleic acid sequencing.
[0103] In some embodiments, the kits and compositions may include additional reagents selected from the group consisting of: buffers (e.g., restriction enzymes, PCR, preservation or washing buffers), DNA recovery reagents or kits (e.g., precipitation, ultrafiltration, affinity columns), and DNA recovery components, etc.
[0104] The kit of this application may further include one or more of the following components known in the field of DNA enrichment: a protein component that selectively binds methylated DNA; a triple-stranded nucleic acid component, one or more adapters, optionally in a suitable solution; a substance or solution for ligation, such as a ligase or buffer; a substance or solution for column chromatography; a substance or solution for immunologically based enrichment (e.g., immunoprecipitation); a substance or solution for nucleic acid amplification, such as PCR; a dye or several dyes, if applicable as a coupling agent, if applicable in solution; a substance or solution for hybridization; and / or a substance or solution for a washing step.
[0105] This application also includes a medium containing the sequence of the isolated nucleic acid molecule described herein and optionally its methylation information, the medium being used to compare with gene methylation sequencing data to determine the presence, abundance, and / or methylation level of the nucleic acid molecule. Preferably, the medium is a card printed with the sequence and optionally its methylation information, such as a paper, plastic, metal, or glass card. Preferably, the medium is a computer-readable medium storing the sequence and optionally its methylation information and a computer program that, when executed by a processor, performs the following steps: comparing the methylation sequencing data of a sample with the sequence to obtain the presence, abundance, and / or methylation level of the nucleic acid molecule containing the sequence in the sample.
[0106] This application also includes an apparatus for identifying an invasive subtype of papillary thyroid carcinoma, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to perform the following steps: (1) acquiring the methylation level of one or more target markers or target regions thereof selected herein from the following, and (2) determining whether the sample is an invasive subtype of papillary thyroid carcinoma based on the methylation level of (1). Preferably, the acquisition step is performed using any of the methods described in Part IV of this application; preferably, the determination is performed using any of the methods described in Part V of this application.
[0107] VII. Uses This application also provides the application of the isolated nucleic acid molecules described in this application as detection targets in the invasive subtype of papillary thyroid carcinoma.
[0108] Compared with existing molecular diagnostic techniques for invasive subtypes of papillary thyroid carcinoma, the methylation markers and technical solutions provided by this invention effectively solve the problem of insufficient diagnostic accuracy and reliability.
[0109] The following description, in conjunction with specific embodiments, provides an explanation.
[0110] Unless otherwise specified, the experimental or testing methods described in the following examples are conventional methods; the reagents and materials described are obtained from conventional commercial sources unless otherwise specified.
[0111] Example 1: Methylation-targeted sequencing to screen methylation regions of invasive subtypes of papillary thyroid carcinoma 1. Sample Collection A total of 54 patients with the invasive subtype of papillary thyroid carcinoma and 54 patients with the non-invasive subtype of papillary thyroid carcinoma were collected. All enrolled patients signed informed consent forms. These samples were divided into training and test sets according to a certain ratio (e.g., 7:3). The training set was used to build the machine learning model described below, and the test set was used for model performance testing. Sample information is shown in Table 1. Table 1
[0112] Blood samples were collected from all 108 samples in Streck tubes. To extract plasma, the blood samples were first centrifuged at 1600g for 10 min at 4°C. A smooth braking mode was used to prevent damage to the buffy coat (erythrocyte sedimentation rate layer). The supernatant was then transferred to a new 1.5mL conical tube and centrifuged at 16000g for 10 min at 4°C. The supernatant was then transferred again to a new 1.5mL conical tube and stored at -80°C.
[0113] To extract circulating cell-free DNA (cfDNA), plasma and other samples were thawed and immediately processed using the QIAamp Circulating Nucleic Acid Extraction Kit (Qiagen 55114) according to the manufacturer's instructions. The concentration of the extracted cfDNA was quantified using qubit 3.0.
[0114] 2. Bisulfite Conversion and Library Preparation Sodium bisulfite conversion of cytosine bases was performed using a bisulfite conversion kit (ThermoFisher, MECOV50). 20 ng of genomic DNA or ctDNA was converted and purified for downstream applications according to the manufacturer's instructions.
[0115] The process includes the extraction and quality control of sample DNA, and the conversion of unmethylated cytosine on the DNA into bases that do not bind to guanine. In one or more embodiments, the conversion is performed using an enzymatic method, preferably deaminase treatment, or the conversion is performed using a non-enzymatic method, preferably with bisulfite or disulfite treatment, more preferably with calcium bisulfite, sodium bisulfite, potassium bisulfite, ammonium bisulfite, sodium disulfite, potassium disulfite, and ammonium disulfite.
[0116] Library construction was performed using the MethylTitan method, which involves dephosphorylating bisulfite-converted DNA and ligating it to a universal Illumina sequencing adapter with a molecular tag (UMI). After second-strand synthesis and purification, the transformed DNA was subjected to semi-targeted PCR to amplify the desired target region. After further purification, sample-specific barcodes and full-length Illumina sequencing adapters were added to the target DNA molecule via PCR. The resulting library was quantified using Illumina's KAPA library quantification kit (KK4844) and sequenced using an Illumina sequencer. The MethylTitan library construction method effectively enriches the desired target fragments using relatively small amounts of DNA, especially cfDNA. This method also preserves the methylation state of the original DNA well. Finally, by analyzing adjacent CpG methylated cytosine (a given target may have several to dozens of CpGs, depending on the given region), the entire methylation pattern of that specific region can be identified as a unique marker.
[0117] 3. Sequencing and Data Preprocessing (1) Paired-end sequencing was performed using an Illumina Hiseq 2500 sequencer, with a sequencing volume of 25-35M per sample. Trim_galore v 0.6.0 and cutadapt v2.1 software were used to remove adapters from the 150bp paired-end sequencing data from the Illumina Hiseq 2500 sequencer. The adapter sequence “AGATCGGAAGAGCACACGTCTGAACTCCAGTC” (SEQ ID NO: 49) was removed from the 3' end of Read 1, and the adapter sequence “AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT” (SEQ ID NO: 50) was removed from the 3' end of Read 2. Bases with sequencing quality values lower than 20 at both ends were also removed. If there was a 3bp adapter sequence at the 5' end, the entire read was removed. Reads shorter than 30 bases after adapter removal were also removed.
[0118] (2) Use Pear v0.9.6 software to merge paired-end sequences into single-end sequences. Merge reads that overlap by at least 20 bases. If the merged reads are shorter than 30 bases, discard them.
[0119] (3) Sequencing data alignment The reference genome data used came from the UCSC database (UCSC: hg19, http: / / hgdownload.soe.ucsc.edu / goldenPath / hg19 / bigZips / hg19.fa.gz).
[0120] 1) First, hg19 was transformed into cytosine to thymine (CT) and adenine to guanine (GA) using Bismark software, and the transformed genomes were indexed using Bowtie2 software.
[0121] 2) Perform CT and GA conversion on the preprocessed data as well.
[0122] 3) Use Bowtie2 software to align the transformed sequences to the transformed HG19 reference genome. The minimum seed sequence length is 20, and mismatches in the seed sequence are not allowed.
[0123] 4) Extract methylation information For each CpG site in the target region hg19, the methylation level of each site was obtained based on the alignment results described above. The nucleotide numbers of the sites involved in this application correspond to the nucleotide positions of hg19. The nucleotide numbers of the sites in this document correspond to the nucleotide positions of HG19.
[0124] (4) Methylation data matrix 1) Merge the methylation data of each sample in the training set and the test set into a data matrix, and process the missing values for each site with a depth of less than 100.
[0125] 2) Remove sites with a missing value ratio higher than 10%.
[0126] 3) For missing values in the data matrix, the KNN algorithm is used to impute the missing data.
[0127] (5) Methylation features were discovered by grouping the training set samples. 1) Randomly divide the dataset into three parts according to age matching.
[0128] 2) Set aside one portion of the dataset as test data and use the rest as training data.
[0129] 3) The training set is further divided into 3 parts for 3-fold cross-validation. The marker is selected based on the average AUC of the 3-fold cross-validation.
[0130] 4) The marker obtained in step 3 is used to train the model based on the Logistic Regression model using training data, and the model performance is verified on the test data.
[0131] 5) Gene annotation of the obtained methylation markers using Great.
[0132] The specific methylation markers for the invasive subtypes of papillary thyroid carcinoma identified through screening are as follows: SEQ ID NO:1 located within CFAP74 or upstream / downstream of this gene; SEQ ID NO:2 located within EPHA8 or upstream / downstream of this gene; SEQ ID NO:3 located within PLXND1 or upstream / downstream of this gene; SEQ ID NO:4 located within ZNF141 or upstream / downstream of this gene; SEQ ID NO:5 located within PROB1 or upstream / downstream of this gene; SEQ ID NO:6 located within PSORS1C3 or upstream / downstream of this gene; SEQ ID NO:7 located within LINC00174 or upstream / downstream of this gene; SEQ ID NO:8 located within LAT2 or upstream / downstream of this gene; SEQ ID NO:9 located within ENG or upstream / downstream of this gene; SEQ ID NO:10 located within SOX6 or upstream / downstream of this gene; SEQ ID NO:11 located within NAT10 or upstream / downstream of this gene; and SEQ ID NO:10 located within CCND1 or upstream / downstream of this gene. SEQ ID NO:12; located within C12orf43 or upstream / downstream of this gene; SEQ ID NO:13; located within GALNT9 or upstream / downstream of this gene; SEQ ID NO:14; located within FLT1 or upstream / downstream of this gene; SEQ ID NO:15; located within C15orf39 or upstream / downstream of this gene; SEQ ID NO:16; located within EEF2KMT or upstream / downstream of this gene; SEQ ID NO:17; located within DEXI or upstream / downstream of this gene; SEQ ID NO:18; located within SNX29 or upstream / downstream of this gene; SEQ ID NO:19; located within KLHDC4 or upstream / downstream of this gene; SEQ ID NO:20; located within SLC7A5 or upstream / downstream of this gene; SEQ ID NO:21; located within ALOX15 or upstream / downstream of this gene; SEQ ID NO:22; located within ALOX12P2 or upstream / downstream of this gene; SEQ ID NO:23; located within JSRP1 or upstream / downstream of this gene; SEQ ID NO:24; located within JSRP1 or upstream / downstream of this gene; SEQ ID NO:25; located within ALOX15 or upstream / downstream of this gene; SEQ ID NO:26; located within ALOX12P2 or upstream / downstream of this gene; SEQ ID NO:27; located within ALOX12P2 or upstream / downstream of this gene; SEQ ID NO:28; located within JSRP1 or upstream / downstream of this gene; SEQ ID NO:29; located within ALOX12P2 or upstream / downstream of this gene; SEQ ID NO:29; located within ALOX12P2 or upstream / downstream of this gene; SEQ ID NO:22; located within ALOX12P2 or upstream / downstream of this gene; SEQ ID NO:23; located within JSRP1 or upstream / downstream of this gene; SEQ SEQ ID NO:24; SEQ ID NO:25 located within KDM4B or upstream / downstream of the gene; SEQ ID NO:26 located within MLLT1 or upstream / downstream of the gene; SEQ ID NO:27 located within BEST2 or upstream / downstream of the gene; SEQ ID NO:28 located within KCNN4 or upstream / downstream of the gene; SEQ ID NO:29 located within CPXM1 or upstream / downstream of the gene; SEQ ID NO:30 located within MN1 or upstream / downstream of the gene.
[0133] Example 2: Discrimination performance of a single methylation marker To verify the performance of a single methylation marker in distinguishing invasive subtypes of papillary thyroid carcinoma, the model was trained on the training set data of Example 1 using methylation level data of a single marker, and the performance of the model was verified using test set samples. The specific steps are as follows ( Figure 1 ): 1. Sequence preprocessing: For each target region, calculate the methylation level (DMR) within that region.
[0134] 2. Use the logistic regression model from the sklearn (V1.0.1) package in Python (V3.9.7): model=LogisticRegression(). The formula for this model is as follows, where Tx is the methylation level data matrix after normalization of the sequencing depth of the target marker in the sample, w is the coefficient of different markers, b is the intercept value, and y is the model prediction score: y=1 / (1+e^((-w × Tx+b) ) ) 3. Use the training set samples for training: model.fit(Traindata, TrainPheno), where TrainData is the data of target methylation sites in the training set samples, TrainPheno is the phenotype of the training set samples (1 for invasive subtypes and 0 for non-invasive subtypes), and the relevant thresholds of the model are determined based on the training set samples.
[0135] 4. Use the test set samples for testing: TestPred = model.predict_proba(TestData)[:, 1], where TestData is the data of the target methylation sites in the test set samples, and TestPred is the model prediction score. Use the prediction score and the above threshold to determine whether the sample is an invasive subtype of papillary thyroid carcinoma.
[0136] 5. AUC index of statistical models.
[0137] The performance of the logistic regression model for a single target biomarker in this embodiment is shown in Table 2. As can be seen from Table 2, most target biomarkers can achieve an AUC of over 0.5 on both the test and training sets, making them good biomarkers for the invasive subtype of papillary thyroid carcinoma.
[0138] Table 2: Performance of a single marker logistic regression model
[0139] Example 3: Prediction results for all target markers This embodiment uses the methylation levels of all 30 target biomarkers to construct a logistic regression machine learning model and verifies the model's performance in accurately distinguishing whether an object is a sample of the invasive subtype of papillary thyroid carcinoma. The specific steps are basically the same as in Embodiment 2, except that data from all 30 target biomarker combinations (SEQ ID NO: 1-30) are used as input to the model. The model formula is as follows: P = 1 / (1 + e^(-(-51.410966 + 0.024486×chr16_12354401_12354600 +0.040467×chr19_12861001_12861200 + 0.053993×chr3_129299001_129299200-0.026483×chr9_130587601_130587800 + 0.018955×chr22_27895001_27895200 +0.021872×chr16_87712401_87712600-0.000448×chr7_65959401_65959600 + 0.035720×chr12_121437201_121437400 + 0.018661×chr16_11017801_11018000 + 0.030965×chr1_22920001_22920200 + 0.042765×chr20_2786401_2786600 + 0.007697×chr4_332601_332800 + 0.056953×chr7_73638801_73639000 + 0.046206×chr19_5146201_5146400 + 0.023985×chr1_1916201_1916400 + 0.025957×chr6_31148401_31148600 +0.017461×chr17_6797001_6797200-0.013130×chr19_2254001_2254200-0.018668×chr19_6223801_6224000 + 0.034953×chr11_34176801_34177000 + 0.031780×chr17_4560001_4560200 + 0.063708×chr13_29060601_29060800 + 0.038000×chr11_69407801_69408000 + 0.040449×chr11_15963001_15963200 + 0.036116×chr12_132857601_132857800 + 0.066167×chr16_5198401_5198600 + 0.064816×chr15_75470601_75470800-0.012056×chr5_138730201_138730400 + 0.002086×chr19_44278601_44278800 + 0.025609×chr16_87904401_87904600))). The distribution of model prediction scores in the training and test sets is shown below. Figure 2 The ROC curve is shown below. Figure 3 In the test set, the AUC for distinguishing between the invasive and non-invasive subtypes of papillary thyroid carcinoma reached 0.854, demonstrating good ability to differentiate between these subtypes in the sample. With a threshold set to 0.824, values above this threshold were predicted as invasive papillary thyroid carcinoma, and values below this threshold were predicted as non-invasive. The specificity in the test set was 84.6%, and the sensitivity reached 75.0%, indicating the good performance of this combined model.
[0140] Example 4: Prediction results of 10 target markers To verify the effectiveness of the combination of relevant biomarkers, this embodiment selected 10 target biomarkers (SEQ ID NO:3, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:13, SEQ ID NO:15, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:22, and SEQ ID NO:30) from all 30 methylation biomarkers to construct a logistic regression machine learning model. The model formula is as follows: P = 1 / (1 + e^(-(-33.296135 + 0.018271×chr6:31148401:31148600 +0.058311×chr11:15963001:15963200 + 0.015023×chr22:27895001:27895200 +0.050341×chr16:5198401:5198600 + 0.077265×chr12:121437201:121437400 +0.146836×chr3:129299001:129299200 + 0.046042×chr13:29060601:29060800-0.003744×chr17:4560001:4560200 + 0.051078×chr11:34176801:34177000 +0.079677×chr16:87904401:87904600))) The method for constructing the machine learning model is the same as in Example 2, but only the data of the aforementioned 10 target markers were used for the relevant samples. The model scores on the training and test sets are shown below. Figure 4 The ROC curve of this model can be found in [the image / reference]. Figure 5 It can be seen that the model shows significant differences in sample scores between the invasive and non-invasive subtypes of papillary thyroid carcinoma in the training and test sets. The model's AUC on the test set reached 0.816, indicating good performance of the combined model. When the threshold is set to 0.910, values above this value are predicted as invasive papillary thyroid carcinoma, and values below this value are predicted as non-invasive papillary thyroid carcinoma. The specificity on the test set is 76.9%, and the sensitivity on the test set reaches 70.8%, demonstrating the good performance of the combined model.
[0141] Example 5: Prediction results of 5 target markers To verify the effectiveness of the combination of relevant biomarkers, this embodiment selected five target biomarkers (SEQ ID NO:3, SEQ ID NO:10, SEQ ID NO:16, SEQ ID NO:22, and SEQ ID NO:30) from all 30 methylation biomarkers to construct a logistic regression machine learning model. The model formula is as follows: P = 1 / (1 + e^(-(-18.684188 + 0.056339×chr11:15963001:15963200 +0.041070×chr17:4560001:4560200 + 0.075599×chr15:75470601:75470800 +0.102088×chr3:129299001:129299200 + 0.021281×chr22:27895001:27895200))) The method for constructing the machine learning model is the same as in Example 2, but the relevant samples only used data from the aforementioned five target markers. The model scores on the training and test sets are shown below. Figure 6 The ROC curve of this model can be found in [the image / reference]. Figure 7 It can be seen that the model shows a significant difference in sample scores between the invasive and non-invasive subtypes of papillary thyroid carcinoma in the training and test sets. The model's AUC on the test set reached 0.731, indicating good performance of the combined model. When the threshold is set to 0.650, values above this value are predicted as invasive papillary thyroid carcinoma, and values below this value are predicted as non-invasive papillary thyroid carcinoma. The specificity on the test set is 65.4%, and the sensitivity on the test set reaches 70.8%, demonstrating the good performance of the combined model.
[0142] The above description is merely a preferred embodiment of this application and is not intended to limit this application in any form or substance. It should be noted that those skilled in the art can make several improvements and additions without departing from this application, and these improvements and additions should also be considered within the scope of protection of this application.
Claims
1. A methylation marker for identifying invasive subtypes of papillary thyroid carcinoma, characterized in that, The methylation markers are selected from any one or more of the following gene fragments: chr1:1916201:1916400 (SEQ ID NO:1) and any gene fragment within 5 kb upstream and / or 5 kb downstream of it; chr1:22920001:22920200 (SEQ ID NO:2) and any gene fragment within 5 kb upstream and / or 5 kb downstream of it; chr3:129299001:129299200 (SEQ ID NO:3) and any gene fragment within 5 kb upstream and / or 5 kb downstream of it; chr4:332601:332800 (SEQ ID NO:4) and any gene fragment within 5 kb upstream and / or 5 kb downstream of it; chr5:138730201:138730400 (SEQ ID NO:1) and any gene fragment within 5 kb upstream and / or downstream of it; chr5:138730201:138730400 (SEQ ID NO:1) and any gene fragment within 5 kb upstream and / or downstream of it; chr1:22920001:22920200 (SEQ ID NO:2) and any gene fragment within 5 kb upstream and / or 5 kb ... NO:5) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr6:31148401:31148600 (SEQ ID NO:6) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr7:65959401:65959600 (SEQ ID NO:7) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr7:73638801:73639000 (SEQ ID NO:8) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr9:130587601:130587800 (SEQ ID NO:5 ...800 (SEQ ID NO:5) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr9:130587601:130587800 (SEQ ID NO:5) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr9:130587601:130587800 NO:9) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr11:15963001:15963200 (SEQ ID NO:10) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr11:34176801:34177000 (SEQ ID NO:11) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr11:69407801:69408000 (SEQ ID NO:12) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr12:121437201:121437400 (SEQ ID NO:9) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr12:121437201:121437400 (SEQ ID NO:9) and any gene fragment within 5kb upstream and / or downstream of it; containing chr11:15963001:15963200 (SEQ ID NO:10) and any gene fragment within 5kb upstream and / or ...1:34176801:34177000 (SEQ ID NO:11) and any gene fragment within 5kb upstream and / or downstream of it; containing chr12:121437201:121437400 NO:13) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr12:132857601:132857800 (SEQ ID NO:14) and any gene fragment within 5kb upstream and / or 5kb downstream of it; containing chr13:29060601:29060800 (SEQ ID NO:15) and any gene fragment within 5kb upstream and / or 5kb downstream of it;Contains chr15:75470601:75470800 (SEQ ID NO:16) and any gene fragment within 5kb upstream and / or 5kb downstream; contains chr16:5198401:5198600 (SEQ ID NO:17) and any gene fragment within 5kb upstream and / or 5kb downstream; contains chr16:11017801:11018000 (SEQ ID NO:18) and any gene fragment within 5kb upstream and / or 5kb downstream; contains chr16:12354401:12354600 (SEQ ID NO:19) and any gene fragment within 5kb upstream and / or 5kb downstream; contains chr16:87712401:87712600 (SEQ ID NO:16) and any gene fragment within 5kb upstream and / or downstream. SEQ ID NO:20) and any gene fragment within 5kb upstream and / or 5kb downstream; containing chr16:87904401:87904600 (SEQ ID NO:21) and any gene fragment within 5kb upstream and / or 5kb downstream; containing chr17:4560001:4560200 (SEQ ID NO:22) and any gene fragment within 5kb upstream and / or 5kb downstream; containing chr17:6797001:6797200 (SEQ ID NO:23) and any gene fragment within 5kb upstream and / or 5kb downstream; containing chr19:2254001:2254200 (SEQ ID NO:24) and any gene fragment within 5kb upstream and / or 5kb downstream; containing chr19:5146201:5146400 (SEQ ID NO:20 ... SEQ ID NO:25) and any gene fragment within 5kb upstream and / or 5kb downstream; containing chr19:6223801:6224000 (SEQ ID NO:26) and any gene fragment within 5kb upstream and / or 5kb downstream; containing chr19:12861001:12861200 (SEQ ID NO:27) and any gene fragment within 5kb upstream and / or 5kb downstream; containing chr19:44278601:44278800 (SEQ ID NO:28) and any gene fragment within 5kb upstream and / or 5kb downstream; containing chr20:2786401:2786600 (SEQ ID NO:25 ... NO:29) and any gene fragment within 5kb upstream and / or 5kb downstream of it; including chr22:27895001:27895200 (SEQ ID NO:30) and any gene fragment within 5kb upstream and / or 5kb downstream of it.
2. The methylation marker according to claim 1, characterized in that, The methylation markers are selected from any one or more of the gene fragments shown in SEQ ID NO:1-30.
3. The methylation marker according to claim 1, characterized in that, The methylation markers are selected from any one of the following groups: 1) Combinations of gene fragments shown in SEQ ID NO:1-30; 2) Combinations of gene fragments shown in SEQ ID NO:3, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:13, SEQ ID NO:15, SEQ ID NO:17, SEQ ID NO:21, SEQ ID NO:22 and SEQ ID NO:30; 3) Combinations of gene fragments shown in SEQ ID NO:3, SEQ ID NO:10, SEQ ID NO:16, SEQ ID NO:22 and SEQ ID NO:
30.
4. The use of the methylation marker or its detection reagent according to any one of claims 1 to 3 in the preparation of in vitro diagnostic products for identifying invasive subtypes of papillary thyroid carcinoma.
5. A kit for identifying invasive subtypes of papillary thyroid carcinoma, characterized in that, The kit includes reagents for detecting the methylation level of the methylation markers according to any one of claims 1 to 3.
6. A method for constructing a diagnostic model for identifying invasive subtypes of papillary thyroid carcinoma, characterized in that, include: The level data of the target biomarker in several papillary thyroid carcinoma samples were obtained as a training set, wherein the target biomarker is the methylation biomarker as described in any one of claims 1 to 3; A diagnostic model is constructed using a logistic regression algorithm based on the classification information of the samples and the level data of the target markers.
7. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions, which, when executed by a processor, implement the method of claim 6.
8. An apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of claim 5.
9. A device for identifying invasive subtypes of papillary thyroid carcinoma, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: 1) obtaining the methylation level or methylation state of at least one CpG dinucleotide, a methylation marker as described in any one of claims 1 to 3, in the target sample, and 2) determining whether the target sample is an invasive subtype of papillary thyroid carcinoma based on the methylation level or methylation state of 1).
10. The apparatus according to claim 9, characterized in that, Step 2) specifically includes: inputting the methylation level or methylation state into a pre-constructed diagnostic model to obtain the probability value of the invasive subtype of papillary thyroid carcinoma, and identifying whether the target sample is an invasive subtype of papillary thyroid carcinoma based on a pre-set threshold and probability value; the diagnostic model is the diagnostic model constructed by the method described in claim 6.
Citation Information
Patent Citations
Devices and methods for isolating RNA
EP1626085A1
Method for isolating DNA from biological samples
US7888006B2