DNA methylation markers for lung cancer detection and their applications
By developing a DNA kit containing multiple target gene methylation sites and constructing a lung cancer prediction model, the challenge of non-invasive screening for early-stage lung adenocarcinoma diagnosis has been solved, achieving highly sensitive and specific lung cancer detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SINGLERA GENOMICS (SHANGHAI) LTD
- Filing Date
- 2021-06-22
- Publication Date
- 2026-05-26
AI Technical Summary
There is a lack of effective and non-invasive methods for lung cancer screening in current technologies, especially in the early diagnosis of lung adenocarcinoma. It is difficult to identify tumor-specific DNA methylation markers of ctDNA in the blood, and existing serum biomarkers have insufficient sensitivity and specificity.
Develop a DNA kit that includes the detection of methylation sites of multiple target genes, measures the methylation status or level by treating plasma samples with bisulfite, and constructs a lung cancer prediction model by combining it with a specific algorithm.
It enables non-invasive and accurate diagnosis of lung cancer, especially the early identification of lung adenocarcinoma, with high sensitivity and specificity, providing a new non-invasive screening method.
Smart Images

Figure CN122081501A_ABST
Abstract
Description
[0001] This invention application is a divisional application of Chinese patent application filed on June 22, 2021, entitled "DNA methylation markers for detecting lung cancer and their application", with application number "202110689292.4". Technical Field
[0002] This invention belongs to the field of molecular biomedical technology. Specifically, this invention relates to DNA methylation biomarkers for detecting lung cancer, a method for processing samples and a method for analyzing plasma samples obtained from subjects suspected of having lung cancer, a reagent capable of detecting the methylation level of the DNA methylation biomarker, and the application of the above-mentioned methylation biomarkers in the preparation of lung cancer diagnostic products and the construction of lung cancer prediction models. Background Technology
[0003] Lung cancer is a malignant tumor that occurs in the bronchial mucosal epithelium. It is the most common primary malignant tumor of the lung, and its high incidence and mortality rates pose a serious threat to people's health. Lung cancer is broadly divided into two categories: non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC). NSCLC mainly includes squamous cell carcinoma, adenocarcinoma, and large cell carcinoma, among which lung adenocarcinoma is the most prevalent subtype of lung cancer, and its incidence is increasing year by year. Due to the unique biological behavior of lung adenocarcinoma, most cases of lung adenocarcinoma are currently diagnosed at an advanced stage, making individualized treatment of lung adenocarcinoma a hot topic in treatment.
[0004] In the diagnosis of lung cancer, most cases are discovered at an advanced stage and are diagnosed primarily using invasive methods. Currently, there is a lack of effective and simple screening methods. Tumor markers are an important diagnostic tool that can provide effective evidence for clinical diagnosis and treatment, and reduce screening costs for patients, under simple and economical conditions. Blood is the preferred source of candidate tumor markers for lung cancer screening. Blood-based biomarkers provide a profile of the entire patient's body, including the primary tumor, metastatic disease, immune response, and the surrounding matrix. For example, autoantibodies (AAbs) generated in response to abnormal tumor antigens often appear in the preclinical stage before the onset of symptoms or imaging-based detection. Autoantibodies have been found in all histological types and stages of lung cancer. They are usually absent or found at low titers in cancer-free individuals, but are also present in patients with many other diseases. Therefore, autoantibody panel testing may be specific but not sensitive. In a clinical validation study including all lung cancer histology and stages, a combination of autoantibody tests performed well, with a specificity of 93% but a sensitivity of only 40%. Although several useful serum biomarkers have been identified in lung cancer, there is no consensus on their utility. Furthermore, despite the accumulation of substantial data on the efficacy of several serum biomarkers over the past few decades, there is a lack of guidelines and standard protocols for their use in lung cancer patients.
[0005] Circulating tumor DNA (ctDNA) molecules originate from apoptotic or necrotic tumor cells and carry tumor-specific DNA methylation markers from early-stage malignancies. In recent years, they have been investigated as promising new targets for developing non-invasive early screening tools for various cancers. However, most of these studies have not yielded effective results. Existing research indicates that the proportion of ctDNA in the plasma DNA of patients with early-stage tumors is very small, making the identification of stable and consistent lung adenocarcinoma-specific markers from plasma DNA a significant challenge. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a DNA methylation biomarker for detecting lung cancer. By performing differential methylation analysis on this biomarker, lung cancer can be effectively identified, achieving the goal of non-invasive and accurate diagnosis of lung cancer.
[0007] Specifically, according to one aspect of this application, a DNA kit for detecting lung cancer includes reagents for detecting the methylation status or level of at least one methylation site of any one or more of the following target genes: HES4, PLEKHN1, ARHGEF16, PRDM16, PTPRU, DMRTA2, ELAVL4, DMRTA2, FOXD3, LMO4, ENSG00000267561, CSF1, EPS8L3, SPAG17, TBX15, PIAS3, ITGA10, RORC, THEM5, IL6R, MUC1, BCAN, ARHGAP30, RYR2, PRKCQ, SFMBT2. , GATA3, DNAJC1, COMMD3, PTF1A, SORCS1, KIAA1598, VAX1, PDZD8, EMX2, FOXI2, MKI67, GLRX3, EBF3, PPP2R2D, ATHL1, NLRP6, DRD4, DEAF1, PAX6, ELP4, AL X4, EXT2, GLYATL2, GLYATL1, RCOR2, PKNOX2, FEZ1, IFFO1, TMTC1, IPO8, SYT10, ACVR1B, ACVRL1, DTX1, RASAL1, LHX5, SDSL, RBM19, TBX5, TBX3, RNF17, AT P12A, MAB21L1, DCLK1, LECT1, PCDH8, SOX1, SPACA7, NKX2-8, PAX9, FOXA1, MIPOL1, CLEC14A, PTGDR, OTX2, TMEM260, SIX6, SIX1, DPF3, RGS6, VRK1, GABRG 3. GABRA5, NDNL2, APBA2, ITPKA, LTK, DUOX1, SHF, PIAS1, SKOR1, LMAN1L, CSK, SOCS1, CIITA, HS3ST2, PRKCB, ZNF720, AHSP, TP53TG3B, CYLD, SALL1, NLRC 5. FAM192A, ZFHX3, CDH13, FOXF1, IRF8, PIEZO1, CTU2, LHX1, MRM1, ARHGAP23, SRCIN1, HOXB4, HOXB5, DLX4, MSI2, ENSG00000166329, MRPS23, VEZF1, TMC 6. TMC8, DCC, KLF2, AP1M1, KLF2, CILP2, PBX4, EGLN2, CYP2A6, FAM150B, TMEM18, SOX11, OSR1, POMC, DNMT3A, LBH, CD207, VAX2, C2orf40, NCK2, BCL2L11,SLC4A10, TBR1, KIAA1715, HOXD10, HOXD11, HOXD12, HOXD13, EVX2, KIAA1715, HOXD9, HOXD1, HOXD4, SDPR, TMEFF2, PTH2R, PAX3, TWIST2, CXXC11, TFAP2C, RBM38, NCAM2, PANX2, TUBGCP6, TYMP, SYCE3, TYMP, CPT1B, RAD18, SRGAP3, RARB, DLEC1, AMT, RASSF1, HEG1, SLC12A8, TR H, PLSCR1, ZIC4, VEPH1, SHOX2, SOX2, FGF12, HMX1, CPZ, SOD3, LGI2, HOPX, ARL9, SFRP2, FRG2, FRG1, IRX4, NDUFS6, C5orf38, IRX2, IRX1, MARCH11, PTGER4, PRKAA1, FOXD1, TMEM174, SRP19, APC, CDO1, SEPT8, SHROOM1, NEUROG1, PCDHGC5, DIAPH1, PHYKPL, COL23A1, GRM6, ADAMT S2, ZNF354C, TFAP2A, ID4, GLO1, DNAH8, MDFI, FOXP4, GUCA1A, TAF8, TMEM63B, VEGFA, TFAP2B, COL12A1, TBX18, PREP, PRDM1, RNF217, OLI G3, HIVEP2, GPR126, OPRM1, IPCEF1, NXPH1, HOXA9, HOXA13, EVX1, PPP1R35, TSC22D4, KCNH2, AOC1, ACTR3B, MNX1, NOM1, DNAJB6, PTPRN2, DLGAP2, TDRP, GFRA2, NKX2-6, EBF2, SOX17, BHLHE22, SLCO5A1, PRDM14, TBC1D31, ZHX1, FAM84B, OPLAH, SPATC1, DMRT2, DMRT3, NFIB, ZDHHC21, SLC24A2, CDKN2A, PAX5, MELK, DAPK1, NEK6, LHX2, NR5A1, GPR144, RAPGEF1, and MED27; at least one methylation site of the target gene is selected from methylation sites within the target gene and in the regulatory regions of the target gene.
[0008] In some embodiments, the kit includes reagents for detecting the methylation status or level of at least one methylation site of any one or more of the following target genes: PTPRU, PLEKHN1, LMO4, ENSG00000267561, GLYATL2, GLYATL1, TMEM260, DPF3, LMAN1L, CSK, NLRC5, FAM192A, ENSG00000166329, CXXC11, TYMP, SYCE3, TYMP, PRKAA1, SRP19, PHYKPL, GLO1, HIVEP2, GPR126, IPCEF1, PPP1R35, AOC1, TDRP, TBC1D31.
[0009] In some embodiments, the kit includes reagents for detecting the methylation status or level of at least one methylation site in the regulatory region of the target gene.
[0010] In some embodiments, the at least one methylation site comprises at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or even at least ten or more methylation sites.
[0011] In some implementations, the regulatory region of the target gene is any continuous region of 100bp-400bp, preferably 150bp-300bp, and more preferably 200bp-250bp.
[0012] In some embodiments, the target gene or its regulatory region is selected from any one or more sequences shown in SEQ ID No. 1-365 or their complementary sequences.
[0013] In some implementations, the number of methylation sites included in SEQ ID No. 1-365 is shown in the table below:
[0014] In some embodiments, the kit also includes bisulfite.
[0015] In some specific embodiments, the present invention also provides a DNA kit for detecting lung cancer, the kit comprising reagents for detecting the methylation status or level of at least one methylation site of any one or more of the following target genes: SORCS1, RCOR2, SYT10, SOX1, OTX2, TMEM260, CILP2, PBX4, SOX11, SLC4A10, TBR1, KIAA1715, HOXD10, HOXD11, HOXD12, HOXD13, EVX2, VEPH1, SHOX2, IRX1, GRM6, OPRM1, IPCEF1, NXPH1, DNAJB6, PTPRN2, and SLC24A2; wherein at least one methylation site of the target gene is selected from methylation sites within the target gene or in the regulatory region of the target gene.
[0016] In some preferred embodiments, the regulatory region of the target gene is selected from any or more of the sequences shown in SEQ ID NO:41, 62, 63, 71, 91, 106, 107, 108, 162, 169, 170, 182, 186, 194, 243, 244, 245, 267, 270, 271, 288, 289, 312, 313, 314, 315, 328, 329 and 354 or their complementary sequences.
[0017] In some specific embodiments, the present invention also provides that the kit includes reagents for detecting the methylation status or level of at least one methylation site of any one or more of the following target genes: SORCS1, OTX2, TMEM260, CILP2, PBX4, SOX11, IRX1, OPRM1, IPCEF1, and NXPH1; wherein the at least one methylation site of the target gene is selected from methylation sites within the target gene or in the regulatory region of the target gene.
[0018] In some preferred embodiments, the regulatory region of the target gene is selected from any one or more of the sequences shown in SEQ ID NO:41, 106, 107, 108, 162, 169, 170, 267, 313 and 315 or their complementary sequences.
[0019] In some specific embodiments, the present invention also provides that the kit includes reagents for detecting the methylation status or level of at least one methylation site of any one or more of the following target genes: DMRTA2, ELAVL4, SORCS1, GABRG3, GABRA5, TP53TG3B, SOX11, HOXD9, PLSCR1, ZIC4, VEPH1, SHOX2, HMX1, CPZ, SOD3, LGI2, IRX1, CDO1, PREP, PRDM1, OPRM1, IPCEF1, NXPH1, DNAJB6, PTPRN2, NFIB, and ZDHHC21; wherein the at least one methylation site of the target gene is selected from methylation sites within the target gene or in the regulatory region of the target gene.
[0020] In some preferred embodiments, the regulatory region of the target gene is selected from any or more of the sequences shown in SEQ ID NO:12, 41, 116, 133, 167, 193, 195, 239, 243, 251, 254, 255, 268, 281, 306, 312, 315, 328, 329 and 352 or their complementary sequences.
[0021] In some specific embodiments, the present invention also provides that the kit includes reagents for detecting the methylation status or level of at least one methylation site of any one or more of the following target genes: RORC, THEM5, RD4, SOX1, OTX2, TMEM260, GABRG3, GABRA5, DUOX1, SHF, SALL1, KLF2, AP1M1, OSR1, HOXD10, HOXD11, HOXD12, HOXD9, PAX3, TFAP2C, VEPH1, SHOX2, HMX1, CPZ, TFAP2B, OPRM1, IPCEF1, NXPH1, and SLC24A2; wherein the at least one methylation site of the target gene is selected from methylation sites within the target gene or in the regulatory region of the target gene.
[0022] In some preferred embodiments, the regulatory region of the target gene is selected from any or more of the sequences shown in SEQ ID NO:27, 55, 91, 106, 115, 120, 138, 160, 171, 189, 196, 209, 213, 244, 252, 301, 313, 314, 354 and 355 or their complementary sequences.
[0023] In some specific embodiments, the present invention also provides that the kit includes reagents for detecting the methylation status or level of at least one methylation site of any one or more of the following target genes: TBX15, FOXI2, RCOR2, OTX2, TMEM260, VRK1, SOX11, CD207, VAX2, SLC4A10, TBR1, HOXD9, TFAP2C, TYMP, SYCE3, PLSCR1, ZIC4, VEPH1, SHOX2, MARCH11, GRM6, RNF217, SOX17, DMRT2, and DMRT3; wherein the at least one methylation site of the target gene is selected from methylation sites within the target gene or in the regulatory region of the target gene.
[0024] In some preferred embodiments, the regulatory region of the target gene is selected from any or more of the sequences shown in SEQ ID NO:24, 48, 63, 107, 114, 170, 177, 181, 194, 214, 221, 238, 247, 272, 288, 289, 307, 339, 340 and 351 or their complementary sequences.
[0025] In some specific embodiments, the present invention also provides that the kit includes reagents for detecting the methylation status or level of at least one methylation site of any one or more of the following target genes: GATA3, FOXI2, GLRX3, EBF3, RCOR2, SYT10, OTX2, TMEM260, PIAS1, SKOR1, HS3ST2, CDH13, KLF2, AP1M1, SOX11, HOXD1, HOXD4, AMT, SOD3, LGI2, IRX1, TBX18, SOX17, NFIB, and ZDHHC21; wherein the at least one methylation site of the target gene is selected from methylation sites within the target gene or in the regulatory region of the target gene.
[0026] In some preferred embodiments, the regulatory region of the target gene is selected from any one or more of the sequences shown in SEQ ID NO:36, 49, 52, 62, 71, 108, 122, 127, 142, 161, 168, 169, 201, 230, 253, 270, 271, 305, 337 and 353 or their complementary sequences.
[0027] In some specific embodiments, the present invention also provides that the kit includes reagents for detecting the methylation status or level of at least one methylation site of any one or more of the following target genes: TBX15, GATA3, PAX6, ELP4, ALX4, EXT2, LHX5, RBM19, LECT1, PCDH8, DUOX1, SHF, TP53TG3B, SALL1, CILP2, PBX4, SLC4A10, TBR1, KIAA1715, HOXD10, HOXD11, HOXD12, HOXD13, EVX2, HOXD9, HOXD1, HOXD4, VEPH1, SHOX2, IRX1, and OLIG3; wherein the at least one methylation site of the target gene is selected from methylation sites within the target gene or in the regulatory region of the target gene.
[0028] In a preferred embodiment, the regulatory region of the target gene is selected from any one or more of the sequences shown in SEQ ID NO:23, 37, 57, 59, 80, 90, 121, 134, 137, 162, 182, 186, 188, 197, 200, 245, 248, 267 and 309 or their complementary sequences.
[0029] In some embodiments, the kit described in any of the above embodiments includes a reagent for detecting the methylation status or level of the regulatory region of the target gene. In a preferred embodiment, the regulatory region of the target gene is any continuous region of 100bp-400bp, preferably 150bp-300bp, more preferably 200bp-250bp. In a preferred embodiment, the regulatory region of the target gene is selected from any one or more sequences shown in SEQ ID No. 1-365 or their complementary sequences. In a preferred embodiment, the kit further includes bisulfite.
[0030] According to one aspect of this application, a method for processing a sample is provided, the method comprising: 1) Determine the amount of methylation markers in samples from subjects, wherein the methylation markers are any of the target genes or regulatory regions of the target genes mentioned above; 2) Determine the amount of reference marker in the sample; 3) Compare the amount of the methylation marker in the sample with the amount of the reference marker to determine the methylation status or level of the methylation marker.
[0031] In some embodiments, the assay includes obtaining a sample containing DNA from a subject and treating the DNA obtained from the sample with an agent that selectively modifies unmethylated cytosine residues in the obtained DNA to produce modified residues.
[0032] In some embodiments, the methylation markers of any one or more combinations include at least two of the methylation markers, preferably 2 to 100 markers, preferably 5 to 50 markers, and preferably 10 to 30 markers.
[0033] In some embodiments, the methylation markers include methylation markers selected from the group consisting of: SEQ ID No: 1-28, 30, 31, 33-87, 90-134, 137-145, 148-153, 156, 158, 159, 162-171, 173-211, 213-220, 222-234, 237-297, 299-309, 312-329, 331-357, 359-364 or their complementary sequences or variants thereof, wherein the methylation sites in the variants are not mutated.
[0034] In some embodiments, the methylation markers comprise the group consisting of: SEQ ID NO: 41, 62, 63, 71, 91, 106, 107, 108, 162, 169, 170, 182, 186, 194, 243, 244, 245, 267, 270, 271, 288, 289, 312, 313, 314, 315, 328, 329, and 354.
[0035] In some embodiments, the methylation markers include the group consisting of: SEQ ID NO:41, 106, 107, 108, 162, 169, 170, 267, 313 and 315.
[0036] In some embodiments, the methylation markers comprise the group consisting of: SEQ ID NO: 12, 41, 116, 133, 167, 193, 195, 239, 243, 251, 254, 255, 268, 281, 306, 312, 315, 328, 329 and 352.
[0037] In some embodiments, the methylation markers comprise the group consisting of: SEQ ID NO: 27, 55, 91, 106, 115, 120, 138, 160, 171, 189, 196, 209, 213, 244, 252, 301, 313, 314, 354, and 355.
[0038] In some embodiments, the methylation markers comprise the group consisting of: SEQ ID NO: 24, 48, 63, 107, 114, 170, 177, 181, 194, 214, 221, 238, 247, 272, 288, 289, 307, 339, 340, and 351.
[0039] In some embodiments, the methylation markers comprise the group consisting of: SEQ ID NO: 36, 49, 52, 62, 71, 108, 122, 127, 142, 161, 168, 169, 201, 230, 253, 270, 271, 305, 337, and 353.
[0040] In some embodiments, the methylation markers include the group consisting of: SEQ ID NO: 23, 37, 57, 59, 80, 90, 121, 134, 137, 162, 182, 186, 188, 197, 200, 245, 248, 267 and 309.
[0041] In some embodiments, the reagent used to determine the amount of methylation markers includes any one or more selected from deaminases, bisulfites, and bisulfites.
[0042] In some embodiments, the reagent used to determine the amount of the methylation marker includes any one or more selected from calcium bisulfite, sodium bisulfite, potassium bisulfite, ammonium bisulfite, sodium disulfite, potassium disulfite, and ammonium disulfite.
[0043] In some embodiments, determining the methylation state or level of a methylation marker in the sample includes determining the methylation state or level of a single base.
[0044] In some embodiments, determining the methylation state or level of methylation markers in the sample includes determining the methylation state or level of multiple bases.
[0045] In some embodiments, the methylation state or level of the methylation marker includes an increase or decrease in the methylation level of the methylation marker relative to the methylation level of the same methylation marker in one or more normal samples.
[0046] In some implementations, the reference marker is a methylation reference marker.
[0047] In some embodiments, the sample is selected from tissue samples, blood samples, serum samples, or sputum samples.
[0048] In some embodiments, the sample is a plasma sample.
[0049] In some implementations, the sample includes cfDNA isolated from plasma.
[0050] In some implementations, the subject has or is suspected of having a lung tumor.
[0051] According to another aspect of this application, a method for analyzing plasma samples obtained from a subject suspected of having a lung tumor is provided, the method comprising: (1) Extract cfDNA from the plasma sample; (2) Determine the amount of methylation markers and reference markers in cfDNA, wherein the methylation markers are any of the above target genes or the regulatory regions of target genes; (3) Compare the amount of the methylation marker with the amount of the reference marker to determine the methylation status or level of the methylation marker.
[0052] In some embodiments, the assay includes using one or more of the following methods: bisulfite-based PCR, DNA sequencing, whole-genome methylation sequencing, simplified methylation sequencing, methylation-sensitive restriction endonuclease assay, quantitative fluorescence assay, methylation-sensitive high-resolution melting curve assay, chip-based methylation mapping, and mass spectrometry.
[0053] In some embodiments, the assay includes bisulfite conversion of the methylation marker and the reference nucleic acid.
[0054] In some implementations, the determination of the methylation status or level of one or more methylation markers is achieved by using bisulfite sequencing.
[0055] According to another aspect of this application, an application of a methylation marker in the preparation of lung cancer diagnostic products is provided, wherein the methylation marker is any of the target genes or the regulatory region of the target gene.
[0056] In some preferred embodiments, the present invention provides the use of the group consisting of the following in the preparation of lung cancer diagnostic products: SEQ ID Nos: 1-28, 30, 31, 33-87, 90-134, 137-145, 148-153, 156, 158, 159, 162-171, 173-211, 213-220, 222-234, 237-297, 299-309, 312-329, 331-357 and 359-364.
[0057] In some preferred embodiments, the present invention provides the use of the group consisting of the following in the preparation of lung cancer diagnostic products: SEQ ID NO: 41, 62, 63, 71, 91, 106, 107, 108, 162, 169, 170, 182, 186, 194, 243, 244, 245, 267, 270, 271, 288, 289, 312, 313, 314, 315, 328, 329 and 354.
[0058] In some preferred embodiments, the present invention provides the use of the group consisting of the following in the preparation of lung cancer diagnostic products: SEQ ID NO: 41, 106, 107, 108, 162, 169, 170, 267, 313 and 315.
[0059] In some preferred embodiments, the present invention provides the use of the group consisting of the following in the preparation of lung cancer diagnostic products: SEQ ID NO: 12, 41, 116, 133, 167, 193, 195, 239, 243, 251, 254, 255, 268, 281, 306, 312, 315, 328, 329 and 352.
[0060] In some preferred embodiments, the present invention provides the use of the group consisting of the following in the preparation of lung cancer diagnostic products: SEQ ID NO: 27, 55, 91, 106, 115, 120, 138, 160, 171, 189, 196, 209, 213, 244, 252, 301, 313, 314, 354 and 355.
[0061] In some preferred embodiments, the present invention provides the use of the group consisting of the following in the preparation of lung cancer diagnostic products: SEQ ID NO: 24, 48, 63, 107, 114, 170, 177, 181, 194, 214, 221, 238, 247, 272, 288, 289, 307, 339, 340 and 351.
[0062] In some preferred embodiments, the present invention provides the use of the group consisting of the following in the preparation of lung cancer diagnostic products: SEQ ID NO: 36, 49, 52, 62, 71, 108, 122, 127, 142, 161, 168, 169, 201, 230, 253, 270, 271, 305, 337 and 353.
[0063] In some preferred embodiments, the present invention provides the use of the group consisting of the following in the preparation of lung cancer diagnostic products: SEQ ID NO: 23, 37, 57, 59, 80, 90, 121, 134, 137, 162, 182, 186, 188, 197, 200, 245, 248, 267 and 309.
[0064] In some preferred embodiments, the lung cancer diagnostic product is used for the early diagnosis of lung cancer.
[0065] In some preferred embodiments, the lung cancer diagnostic product is used for the diagnosis of lung adenocarcinoma.
[0066] According to another aspect of this application, an application of a methylation biomarker in constructing a lung cancer prediction model is provided, wherein the methylation biomarker is a target gene or a regulatory region of a target gene as described in any of the preceding claims.
[0067] In some preferred embodiments, the present invention provides the application of the group consisting of the following in constructing lung cancer prediction models: SEQ ID No: 1-28, 30, 31, 33-87, 90-134, 137-145, 148-153, 156, 158, 159, 162-171, 173-211, 213-220, 222-234, 237-297, 299-309, 312-329, 331-357 and 359-364.
[0068] In some preferred embodiments, the present invention provides the application of the group consisting of the following in constructing lung cancer prediction models: SEQ ID NO: 41, 62, 63, 71, 91, 106, 107, 108, 162, 169, 170, 182, 186, 194, 243, 244, 245, 267, 270, 271, 288, 289, 312, 313, 314, 315, 328, 329 and 354.
[0069] In some preferred embodiments, the present invention provides the application of the group consisting of the following in constructing lung cancer prediction models: SEQ ID NO:41, 106, 107, 108, 162, 169, 170, 267, 313 and 315.
[0070] In some preferred embodiments, the present invention provides the application of the group consisting of the following in constructing lung cancer prediction models: SEQ ID NO: 12, 41, 116, 133, 167, 193, 195, 239, 243, 251, 254, 255, 268, 281, 306, 312, 315, 328, 329 and 352.
[0071] In some preferred embodiments, the present invention provides the application of the group consisting of the following in constructing lung cancer prediction models: SEQ ID NO: 27, 55, 91, 106, 115, 120, 138, 160, 171, 189, 196, 209, 213, 244, 252, 301, 313, 314, 354 and 355.
[0072] In some preferred embodiments, the present invention provides the application of the group consisting of the following in constructing lung cancer prediction models: SEQ ID NO: 24, 48, 63, 107, 114, 170, 177, 181, 194, 214, 221, 238, 247, 272, 288, 289, 307, 339, 340 and 351.
[0073] In some preferred embodiments, the present invention provides the application of the group consisting of the following in constructing lung cancer prediction models: SEQ ID NO: 36, 49, 52, 62, 71, 108, 122, 127, 142, 161, 168, 169, 201, 230, 253, 270, 271, 305, 337 and 353.
[0074] In some preferred embodiments, the present invention provides the application of the group consisting of the following in constructing lung cancer prediction models: SEQ ID NO: 23, 37, 57, 59, 80, 90, 121, 134, 137, 162, 182, 186, 188, 197, 200, 245, 248, 267 and 309.
[0075] In some preferred embodiments, at least one of the following algorithms is used to process the methylation levels of methylation markers in multiple samples: principal component analysis, logistic regression analysis, nearest neighbor analysis, support vector machine, neural network model, and random forest.
[0076] In a preferred embodiment, the prediction model is used for the early diagnosis of lung cancer.
[0077] In a preferred embodiment, the prediction model is used for the diagnosis of lung adenocarcinoma.
[0078] According to another aspect of this application, a predictive model for lung cancer is provided, which uses at least one of the following algorithms to score methylation biomarkers: principal component analysis, logistic regression analysis, nearest neighbor analysis, support vector machine, neural network model, and random forest; wherein the methylation biomarker is any of the target genes or regulatory regions of target genes.
[0079] According to another aspect of this application, a diagnostic device for lung cancer is provided, the diagnostic device comprising: The analysis module is used to input the methylation level of the sample to be tested into the prediction model for analysis; The diagnostic module outputs the probability that the individual corresponding to the test sample has lung cancer. The prediction model uses at least one of the following algorithms to score methylation biomarkers: principal component analysis, logistic regression analysis, nearest neighbor analysis, support vector machine, neural network model, and random forest; the methylation biomarker is any of the target genes or the regulatory region of the target gene.
[0080] According to another aspect of this application, an apparatus for determining the methylation level of any of the above-mentioned target genes or regulatory regions of target genes in a biological sample is provided for use in the preparation of a kit for detecting lung cancer.
[0081] In some embodiments, the biological sample is selected from tissue samples, blood samples, serum samples, plasma samples, or sputum samples.
[0082] The beneficial effects of this invention are: The methylation biomarkers provided by this invention can effectively identify lung cancer. This invention is the first to identify biomarkers for lung adenocarcinoma based on the analysis of methylation haplotype sequencing data from plasma samples of patients with and without lung adenocarcinoma. Based on the methylation biomarker set of this invention, a predictive model for lung cancer is provided. This model can effectively identify lung adenocarcinoma with high sensitivity and specificity, providing a new method for the early identification of lung adenocarcinoma. The detection process of this model is non-invasive and highly safe. Attached Figure Description
[0083] Figure 1 This is a flowchart illustrating the technical solution of one embodiment of the present invention; Figure 2 This is the ROC curve of prediction model I of the present invention in the test set for diagnosing lung adenocarcinoma; Figure 3 This is a distribution chart of the prediction scores of the prediction model I of this invention in the test set for diagnosing lung adenocarcinoma; Figure 4 This is the ROC curve of the prediction model II of this invention in the test set for diagnosing lung adenocarcinoma; Figure 5 This is a distribution chart of the prediction scores of the prediction model II of this invention in the test set for diagnosing lung adenocarcinoma; Figure 6 This is the ROC curve of prediction model III of the present invention in the test set for diagnosing lung adenocarcinoma; Figure 7 This is a distribution chart of the prediction scores of the prediction model III of this invention in the test set for diagnosing lung adenocarcinoma; Figure 8 This is the ROC curve of the prediction model IV of this invention in the test set for diagnosing lung adenocarcinoma; Figure 9 This is a distribution chart of the prediction scores of the prediction model IV of this invention in the test set for diagnosing lung adenocarcinoma; Figure 10 The ROC curve of prediction model V of this invention in the test set for diagnosing lung adenocarcinoma; Figure 11 This is a distribution chart of the prediction scores of the prediction model V of the present invention in the test set for diagnosing lung adenocarcinoma. Figure 12 The ROC curve of the prediction model VI of this invention for diagnosing lung adenocarcinoma in the test set; Figure 13 This is a distribution chart of the prediction scores of the prediction model VI of this invention in the test set for diagnosing lung adenocarcinoma; Figure 14 The ROC curve of the prediction model VII of this invention for diagnosing lung adenocarcinoma in the test set; Figure 15 This is a distribution chart of the prediction scores of the prediction model VII of this invention in the test set for diagnosing lung adenocarcinoma; Figure 16 The ROC curve of the prediction model VIII of this invention in the test set for diagnosing lung adenocarcinoma; Figure 17 This is a distribution chart of the prediction scores of the prediction model VIII of this invention in the test set for diagnosing lung adenocarcinoma. Detailed Implementation
[0084] Unless otherwise specified, the experimental methods used in the following examples are conventional methods.
[0085] Unless otherwise specified, all materials and reagents used in the following examples are commercially available.
[0086] Example 1: Screening for methylation markers 1. Sample collection A total of 89 blood samples from patients with lung adenocarcinoma and 74 blood samples from patients without lung adenocarcinoma were collected. All enrolled patients signed informed consent forms. Sample information is shown in Table 1.
[0087] Table 1
[0088] All blood samples were collected in Streck tubes. To extract plasma, the blood samples were first centrifuged at 1600g for 10 min at 4°C. A smooth braking mode was used to prevent damage to the buffy coat layer. The supernatant was then transferred to a new 1.5ml conical tube and centrifuged at 16000g for 10 min at 4°C. The supernatant was then transferred again to a new 1.5ml conical tube and stored at -80°C.
[0089] To extract circulating cell-free DNA (cfDNA), plasma and other samples were thawed and immediately processed using the QIAamp Circulating Nucleic Acid Extraction Kit (Qiagen 55114) according to the manufacturer's instructions. The concentration of the extracted cfDNA was quantified using qubit 3.0.
[0090] 2. Bisulfite Conversion and Library Preparation Sodium bisulfite conversion of cytosine bases was performed using a bisulfite conversion kit (ThermoFisher, MECOV50). 20 ng of genomic DNA or ctDNA was converted and purified for downstream applications according to the manufacturer's instructions.
[0091] The process includes the extraction and quality control of sample DNA, and the conversion of unmethylated cytosine on the DNA into bases that do not bind to guanine. In one or more embodiments, the conversion is performed using an enzymatic method, preferably deaminase treatment, or the conversion is performed using a non-enzymatic method, preferably with bisulfite or disulfite treatment, more preferably with calcium bisulfite, sodium bisulfite, potassium bisulfite, ammonium bisulfite, sodium disulfite, potassium disulfite, and ammonium disulfite.
[0092] Library construction was performed using the MethylTitan method. Specifically, DNA converted to bisulfite was dephosphorylated and ligated to a universal Illumina sequencing adapter with a molecular tag (UMI). After second-strand synthesis and purification, the transformed DNA underwent semi-targeted PCR to amplify the desired target region. After further purification, sample-specific barcodes and full-length Illumina sequencing adapters were added to the target DNA molecule via PCR. The resulting library was then quantified using Illumina's KAPA Library Quantification Kit (KK4844) and sequenced using an Illumina sequencer. The MethylTitan library construction method effectively enriches the desired target fragment using relatively small amounts of DNA, especially cfDNA. This method also preserves the methylation state of the original DNA well. Finally, by analyzing adjacent CpG methylated cytosine (a given target may have several to dozens of CpGs, depending on the given region), the entire methylation pattern of that specific region can serve as a unique marker.
[0093] 3. Sequencing and Data Preprocessing (1) Paired-end sequencing was performed using an Illumina Hiseq 2500 sequencer, with a sequencing volume of 25-35M per sample. Trim_galore v 0.6.0 and cutadapt v2.1 software were used to remove adapters from the paired-end 150bp sequencing data from the Illumina Hiseq 2500 sequencer. The adapter sequence “AGATCGGAAGAGCACACGTCTGAACTCCAGTC” was removed from the 3' end of Read 1, and the adapter sequence “AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT” was removed from the 3' end of Read 2. Bases with sequencing quality values lower than 20 at both ends were also removed. If there was a 3bp adapter sequence at the 5' end, the entire read was removed. Reads shorter than 30 bases after adapter removal were also removed.
[0094] (2) Use Pear v0.9.6 software to merge paired-end sequences into single-end sequences. Merge reads that overlap by at least 20 bases. If the merged reads are shorter than 30 bases, discard them.
[0095] (3) Sequencing data alignment The reference genome data used in this invention comes from the UCSC database (UCSC: hg19, http: / / hgdownload.soe.ucsc.edu / goldenPath / hg19 / bigZips / hg19.fa.gz).
[0096] 1) First, hg19 was transformed into cytosine to thymine (CT) and adenine to guanine (GA) using Bismark software, and the transformed genomes were indexed using Bowtie2 software.
[0097] 2) Perform CT and GA conversion on the preprocessed data as well.
[0098] 3) Use Bowtie2 software to align the transformed sequences to the transformed HG19 reference genome. The minimum seed sequence length is 20, and mismatches in the seed sequence are not allowed.
[0099] 4) Extract methylation information For each CpG site of hg19 in the target region, the methylation level corresponding to each site is obtained based on the alignment results described above. The nucleotide numbers of the sites involved in this invention correspond to the nucleotide position numbers of hg19. MHF calculation, for each CpG site of hg19 in the target region, obtains the methylation level corresponding to each site based on the alignment results described above. The nucleotide numbers of the sites in this paper correspond to the nucleotide position numbers of HG19. A target methylation region may have multiple methylation haplotypes; this value needs to be calculated for each methylation haplotype within the target region. An example of the MHF calculation formula is as follows: MHFi,h=(Ni,h) / Ni Where i represents the target methylation region, h represents the target methylation haplotype, Ni represents the number of reads located in the target methylation region, and Ni,h represents the number of reads containing the target methylation haplotype. (4) Methylation haplotype data matrix 1) Merge the methylation haplotype data of each sample in the training set and the test set into a data matrix, and handle missing values for each site with a depth of less than 200.
[0100] 2) Remove sites with a missing value ratio higher than 10%.
[0101] 3) For missing values in the data matrix, the KNN algorithm is used to impute the missing data.
[0102] (5) Identify characteristic methylation haplotypes by grouping training set samples. 1) Perform logistic regression analysis on each methylation haplotype for the phenotype, and construct a logistic regression model. Specifically, for each target region, calculate the ratio of each MHF (Methylated Haplotype Fraction) methylation haplotype within that region. Use the statmodels package (0.12.0) of Python software (v3.6.9) to construct the logistic regression model and calculate the logistic regression coefficients. Command line: import statsmodels.api as sm logist_model = sm.Logit(Y, sm.add_constant(X)).fit pvlaue = logist_model.pvalues[1]. Where X represents the methylation haplotype value corresponding to each sample, Y represents the classification label corresponding to each sample, and pvalue represents the significance test value of logistic regression. For each amplified target region, the methylation marker corresponding to the MHF with the smallest regression coefficient is selected to form a candidate methylation haplotype.
[0103] 2) Randomly divide the training set into ten parts and perform incremental feature selection with tenfold cross-validation.
[0104] 3) Set aside one set of data from the training set as test data, and use the remaining training data as training data. Sort the candidate methylation haplotypes in each region from largest to smallest according to the significance of the regression coefficient. Add one methylation haplotype at a time. Use the 9 sets of training data to build a multinomial kernel SVM model to predict the test data.
[0105] 4) Repeat step 3 10 times to iterate through all the data. Calculate the AUC of the test data each time, and calculate the average AUC after 10 repetitions. If the AUC of the training data increases, retain the candidate methylation haplotype as a characteristic methylation marker; otherwise, discard it. Use the GREAT tool (http: / / great.stanford.edu / great / public-3.0.0 / html / index.php) to annotate the genes (as shown in Table 2).
[0106] The target genes in the methylation markers were annotated using the GREAT tool (http: / / great.stanford.edu / great / public-3.0.0 / html / 3.0.0 / html / index.php). During GREAT analysis, marker regions were associated with adjacent genes, and these adjacent genes were used to annotate the regions. The association process involved two steps: first, identifying the regulatory domain of each gene; then, associating genes covering the regulatory domain of that region with that region. For example, SORCS1 (+93) represents a marker located 93 bp downstream of the transcription start site (TSS) of the SORCS1 gene; HES4 (+1991) represents a marker located 1991 bp downstream of the transcription start site (TSS) of the HES4 gene; ARHGEF16 (-60184) represents a marker located 60184 bp upstream of the transcription start site (TSS) of the ARHGEF16 gene; and NONE indicates that the marker was not annotated.
[0107] 5) The feature combination corresponding to the median of the average AUC under different feature counts in the training set is taken as the final methylation biomarker group I.
[0108] Table 2
[0109] The methylation levels of methylation marker regions increased or decreased in the cfDNA of lung adenocarcinoma patients (as shown in Table 3). The sequences of the 365 methylation markers obtained are shown in SEQ ID NO:1-365. The methylation levels of all CpG sites for each methylation marker could be obtained by MethylTitan methylation sequencing. The mean methylation levels of all CpG sites in each region, as well as the methylation levels of individual CpG sites, can serve as biomarkers for diagnosing lung adenocarcinoma.
[0110] Table 3 shows the mean methylation levels within the methylation marker regions in the training and testing sets for both normal lung adenocarcinoma and non-lung adenocarcinoma populations. As can be seen from Table 3, the distribution of mean methylation levels within the methylation marker regions differs significantly between lung adenocarcinoma and normal individuals without lung adenocarcinoma, demonstrating good discriminatory power.
[0111] Table 3. Average methylation levels within methylation marker regions in the training and test sets.
[0112] Example 2 The method of Example 1 was used to screen characteristic methylation haplotype markers. The difference from Example 1 is as follows: 3. Sequencing and data preprocessing: (5) Identify characteristic methylation haplotypes by grouping the training set samples: 1) Perform logistic regression analysis on the phenotype for each methylation haplotype and construct a logistic regression model. Specifically, for each target region, calculate the ratio of each MHF (Methylated Haplotype Fraction) methylation haplotype in that region. Use the statmodels package (0.12.0) of the Python software (v3.6.9) to construct a logistic regression model and calculate the logistic regression coefficient. Command line: import statsmodels.api as sm logist_model = sm.Logit(Y, sm.add_constant(X)).fit pvlaue = logist_model.pvalues[1]. Where X represents the methylation haplotype value corresponding to each sample, Y represents the classification label corresponding to each sample, and pvalue represents the significance test value of logistic regression. For each amplified target region, the methylation marker with the smallest regression coefficient is selected. Then, the marker groups of all target regions are used for feature screening in the training set using 10-fold cross-validation. The number of times each marker appears in 10 model feature screenings is recorded. Methylation markers with more than 80% occurrences are formed into methylation marker group II. The information of this methylation marker group is shown in Table 4.
[0113] Example 3 The method of Example 1 was used to screen characteristic methylation haplotype markers. The difference from Example 1 is as follows: 3. Sequencing and data preprocessing: (5) Identify characteristic methylation haplotypes by grouping the training set samples: 1) Perform logistic regression analysis on each methylation haplotype according to the phenotype and construct a logistic regression model. Specifically, for each target region, calculate the ratio of each MHF (Methylated Haplotype Fraction) methylation haplotype in that region. Use the statmodels package (0.12.0) of the Python software (v3.6.9) to construct a logistic regression model and calculate the logistic regression coefficient. Command line: import statsmodels.api as sm logist_model = sm.Logit(Y, sm.add_constant(X)).fit pvlaue = logist_model.pvalues[1]. Where X represents the methylation haplotype ratio value corresponding to each sample, Y represents the classification label corresponding to each sample, and pvalue represents the significance test value of logistic regression. For each amplified target region, the marker corresponding to the MHF with the smallest regression coefficient is selected. Then, the marker groups of all target regions are used for feature screening in the training set using 10-fold cross-validation. The number of times each marker appears in 10 model feature screenings is recorded. Methylation markers with more than 90% occurrences are formed into methylation marker group III. The information of this methylation marker group is shown in Table 5.
[0114] Example 4 The method of Example 1 was used to screen characteristic methylation haplotype markers. The difference from Example 1 is as follows: 3. Sequencing and data preprocessing: (5) Identify characteristic methylation haplotypes by grouping the training set samples: 1) Perform logistic regression analysis on the phenotype for each methylation haplotype and construct a logistic regression model. Specifically, for each target region, calculate the ratio of each MHF (Methylated Haplotype Fraction) methylation haplotype in that region. Use the statmodels package (0.12.0) of the Python software (v3.6.9) to construct a logistic regression model and calculate the logistic regression coefficient. Command line: import statsmodels.api as sm logist_model = sm.Logit(Y, sm.add_constant(X)).fit pvlaue = logist_model.pvalues[1]. Where X represents the methylation haplotype value corresponding to each sample, Y represents the classification label corresponding to each sample, and pvalue represents the significance test value of logistic regression. For each amplified target region, the marker corresponding to the MHF with the smallest regression coefficient is selected. Then, the marker groups of all target regions are used for feature screening in the training set using 10-fold cross-validation. The number of times each marker appears in 10 model feature screenings is recorded. Ten methylation markers that appear more than 90% of the time and ten methylation markers that appear more than 80% but less than 90% of the time are randomly selected to form methylation marker group IV. The information of this methylation marker group is shown in Table 5.
[0115] Example 5 The method of Example 1 was used to screen characteristic methylation haplotype markers. The difference from Example 1 is as follows: 3. Sequencing and data preprocessing: (5) Identify characteristic methylation haplotypes by grouping the training set samples: 1) Perform logistic regression analysis on the phenotype for each methylation haplotype and construct a logistic regression model. Specifically, for each target region, calculate the ratio of each MHF (Methylated Haplotype Fraction) methylation haplotype in that region. Use the statmodels package (0.12.0) of the Python software (v3.6.9) to construct a logistic regression model and calculate the logistic regression coefficient. Command line: import statsmodels.api as sm logist_model = sm.Logit(Y, sm.add_constant(X)).fit pvlaue = logist_model.pvalues[1]. Where X represents the methylation haplotype value corresponding to each sample, Y represents the classification label corresponding to each sample, and pvalue represents the significance test value of logistic regression. For each amplified target region, the methylation marker with the smallest regression coefficient is selected. Then, the marker groups of all target regions are used for feature screening in the training set using 10-fold cross-validation. The number of times each marker appears in 10 model feature screenings is recorded. Ten methylation markers with more than 80% occurrences but less than 90% occurrences and ten methylation markers with more than 70% occurrences but less than 80% occurrences are randomly selected to form methylation marker group V. The information of this methylation marker group is shown in Table 5.
[0116] Example 6 The method of Example 1 was used to screen characteristic methylation haplotype markers. The difference from Example 1 is as follows: 3. Sequencing and data preprocessing: (5) Identify characteristic methylation haplotypes by grouping the training set samples: 1) Perform logistic regression analysis on the phenotype for each methylation haplotype and construct a logistic regression model. Specifically, for each target region, calculate the ratio of each MHF (Methylated Haplotype Fraction) methylation haplotype in that region. Use the statmodels package (0.12.0) of the Python software (v3.6.9) to construct a logistic regression model and calculate the logistic regression coefficient. Command line: import statsmodels.api as sm logist_model = sm.Logit(Y, sm.add_constant(X)).fit pvlaue = logist_model.pvalues[1]. Where X represents the methylation haplotype value corresponding to each sample, Y represents the classification label corresponding to each sample, and pvalue represents the significance test value of logistic regression. For each amplified target region, the methylation marker with the smallest regression coefficient is selected. Then, the marker groups of all target regions are used for feature screening in the training set using 10-fold cross-validation. The number of times each marker appears in 10 model feature screenings is recorded. Ten methylation markers with more than 80% occurrences but less than 90% occurrences and ten methylation markers with more than 60% occurrences but less than 70% occurrences are randomly selected to obtain methylation marker group VI. The information of this methylation marker group is shown in Table 5.
[0117] Example 7 The method of Example 1 was used to screen characteristic methylation haplotype markers. The difference from Example 1 is as follows: 3. Sequencing and data preprocessing: (5) Identify characteristic methylation haplotypes by grouping the training set samples: 1) Perform logistic regression analysis on the phenotype for each methylation haplotype and construct a logistic regression model. Specifically, for each target region, calculate the ratio of each MHF (Methylated Haplotype Fraction) methylation haplotype in that region. Use the statmodels package (0.12.0) of the Python software (v3.6.9) to construct a logistic regression model and calculate the logistic regression coefficient. Command line: import statsmodels.api as sm logist_model = sm.Logit(Y, sm.add_constant(X)).fit pvlaue = logist_model.pvalues[1]. Where X represents the methylation haplotype value corresponding to each sample, Y represents the classification label corresponding to each sample, and pvalue represents the significance test value of logistic regression. For each amplified target region, the methylation marker with the smallest regression coefficient is selected. Then, the marker groups of all target regions are used for feature screening in the training set using 10-fold cross-validation. The number of times each marker appears in 10 model feature screenings is recorded. Ten methylation markers with more than 90% occurrences and ten methylation markers with more than 60% but less than 70% occurrences are randomly selected to obtain methylation marker group VII. The information of this methylation marker group is shown in Table 4.
[0118] Example 8 The method of Example 1 was used to screen characteristic methylation haplotype markers. The difference from Example 1 is as follows: 3. Sequencing and data preprocessing: (5) Identify characteristic methylation haplotypes by grouping the training set samples: 1) Perform logistic regression analysis on the phenotype for each methylation haplotype and construct a logistic regression model. Specifically, for each target region, calculate the ratio of each MHF (Methylated Haplotype Fraction) methylation haplotype in that region. Use the statmodels package (0.12.0) of the Python software (v3.6.9) to construct a logistic regression model and calculate the logistic regression coefficient. Command line: import statsmodels.api as sm logist_model = sm.Logit(Y, sm.add_constant(X)).fit pvlaue = logist_model.pvalues[1]. Where X represents the methylation haplotype value corresponding to each sample, Y represents the classification label corresponding to each sample, and pvalue represents the significance test value of logistic regression. For each amplified target region, the methylation marker with the smallest regression coefficient is selected. Then, the marker groups of all target regions are used for feature screening in the training set using 10-fold cross-validation. The number of times each marker appears in 10 model feature screenings is recorded. Ten methylation markers that appear more than 90% of the time and ten methylation markers that appear more than 50% but less than 60% of the time are randomly selected to obtain methylation marker group VIII. The information of this methylation marker group is shown in Table 4.
[0119] Table 4
[0120] Example 9: Building and Testing the Predictive Model To verify the potential of using methylation haplotype markers for lung adenocarcinoma classifiers, logistic regression classification models I-VIII were constructed based on the aforementioned methylation marker groups I-VIII in the training set. The classification prediction performance of the methylation markers in these groups was then verified in the test set. The training and test sets were first divided proportionally, with 114 cases in the training set (samples 1-114) and 49 cases in the test set (samples 115-163).
[0121] Random forest classification models were constructed on the training set using the discovered methylation haplotype markers for the two groups of samples.
[0122] 1) Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0123] 2) To explore the potential of using methylation haplotype biomarkers for the identification of lung adenocarcinoma, a disease classification system was developed based on these biomarkers. A random forest model was trained using the methylation haplotype biomarker levels in the training set. The specific training process is as follows: a) Use the sklearn package (0.23.1) in Python (v3.6.9) to build the training mode for cross-validation training of the model. The command line is: model = RandomForestClassifier (max_depth=5, n_estimators=100). Where max_depth=5 means that the maximum depth of the trees used in the random forest model is 5, and n_estimators=100 means that the number of trees in the random forest model is 100.
[0124] b) Using the sklearn package (0.23.1), input the differential methylation haplotype matrix and construct a random forest model, model.fit(x_train, y_train), where x_train represents the differential methylation haplotype matrix and y_train represents the phenotypic information of the training set.
[0125] In the process of building the prediction model, cancer type was encoded as 1 and non-cancer type as 0. The threshold was set to 0.5 by default using the sklearn package (0.23.1). The constructed prediction models I-VIII ultimately used a scoring threshold of 0.5 to distinguish between benign and malignant samples. The prediction scores of prediction models I-VIII for the training set samples are shown in Table 5.
[0126] Table 5. Prediction scores of the prediction model in diagnosing lung adenocarcinoma on the training set.
[0127] Predictive model testing Based on the methylation biomarker group I-VIII of this invention, predictions were made on the test set using the prediction model I-VIII established via random forest as described above. A prediction function was used to predict the test set, outputting the predicted disease probability (the default scoring threshold is 0.5; a score greater than 0.5 indicates the subject is considered to have lung adenocarcinoma). The test set sample consisted of 49 cases (samples 115-163), and the calculation process is as follows: Command line: test_pred = model.predict(test_df) Where test_pred represents the predicted score of the test set samples obtained by the random forest prediction model constructed in the above classification prediction model, model represents the constructed random forest prediction model, and test_df represents the test set data.
[0128] Table 6 shows the prediction scores of models I-VIII for diagnosing lung adenocarcinoma in the test set, and the ROC curves are shown in Figure 6. Figure 2 , Figure 4 , Figure 6 , Figure 8 , Figure 10 , Figure 12 , Figure 14 and Figure 16 The predicted score distribution is as follows Figure 3 , Figure 5 , Figure 7 , Figure 9 , Figure 11 , Figure 13 , Figure 15 and Figure 17 The area under the population AUC was 0.955 for prediction model I; 0.947 for prediction model II; 0.939 for prediction model III; 0.934 for prediction model IV; 0.942 for prediction model V; 0.937 for prediction model VI; 0.937 for prediction model VII; and 0.930 for prediction model VIII. As shown in the figure, the prediction models constructed from the selected methylation markers all exhibited good discriminative power.
[0129] Table 6. Prediction scores of the predictive model in diagnosing lung adenocarcinoma on the test set.
[0130] This invention is the first to investigate the differences between plasma samples from individuals without lung adenocarcinoma and those with lung adenocarcinoma by analyzing the methylation levels of methylation haplotype markers in plasma cell-free DNA (cfDNA). 365 methylation markers with significant differences were identified. Based on these methylation marker groups, a lung adenocarcinoma risk prediction model was established using a random forest method. This model effectively identifies lung adenocarcinoma with high sensitivity and specificity (see Table 7), making it suitable for the screening and diagnosis of lung adenocarcinoma.
[0131] Table 7
[0132] Contents not described in detail in this specification are prior art known to those skilled in the art. The above descriptions are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A DNA methylation marker for detecting lung cancer, characterized in that, The methylation marker includes at least one of the following: A1. The target genes of the methylation markers include HES4, PLEKHN1, ARHGEF16, PRDM16, PTPRU, DMRTA2, ELAVL4, FOXD3, LMO4, CSF1, ENSG00000267561, EPS8L3, SPAG17, TBX15, PIAS3, ITGA10, RORC, THEM5, IL6R, MUC1, BCAN, ARHGAP30, RYR2, PRKCQ, SFMBT2, GATA3, DNAJC1, COMMD3, PTF1A, SORCS1, KIAA1598, VAX1, PDZD8, EMX2, FOXI2, M KI67, GLRX3, EBF3, PPP2R2D, ATHL1, NLRP6, DRD4, DEAF1, ELP4, ALX4, EXT2, GLYATL2, GLYATL1, RCOR2, PKNOX2, FEZ1, IFFO1, TMTC1, IPO8, SYT10, ACVR1B , ACVRL1, DTX1, RASAL1, LHX5, SDSL, RBM19, TBX5, TBX3, RNF17, ATP12A, MAB21L1, DCLK1, LECT1, PCDH8, SOX1, SPACA7, NKX2-8, PAX9, FOXA1, MIPOL1, CLE C14A, PTGDR, TMEM260, SIX6, SIX1, DPF3, RGS6, VRK1, GABRG3, GABRA5, NDNL2, APBA2, ITPKA, LTK, DUOX1, SHF, PIAS1, SKOR1, LMAN1L, CSK, SOCS1, CIITA, HS3ST2, PRKCB, ZNF720, AHSP, TP53TG3B, CYLD, SALL1, NLRC5, FAM192A, ZFHX3, CDH13, FOXF1, IRF8, PIEZO1, CTU2, LHX1, MRM1, ARHGAP23, SRCIN1, HOXB4 , HOXB5, DLX4, MSI2, ENSG00000166329, MRPS23, VEZF1, TMC6, TMC8, DCC, KLF2, AP1M1, CILP2, PBX4, EGLN2, CYP2A6, FAM150B, TMEM18, SOX11, POMC, DNMT 3A, LBH, CD207, VAX2, C2orf40, NCK2, BCL2L11, SLC4A10, TBR1, KIAA1715, HOXD10, HOXD11, HOXD12, HOXD13, EVX2, KIAA1715, HOXD9, HOXD1, HOXD4, SDPR,TMEFF2, PTH2R, PAX3, TWIST2, CXXC11, TFAP2C, RBM38, NCAM2, PANX2, TUBGCP6, TYMP, SYCE3, TYMP, CPT1B, RAD18, SRGAP3, RARB, DLEC1, AMT, RASSF1, HEG1, SLC12A8, TRH, PLSCR1, ZIC4, VEPH1, SHOX2, SOX2, FGF12, HMX1, CPZ, SOD3, LGI2, HOPX, ARL 9. SFRP2, FRG2, FRG1, IRX4, NDUFS6, C5orf38, IRX1, MARCH11, PTGER4, PRKAA1, FOXD1, TMEM174, SRP19, APC, CDO1, SEPT8, SHROOM1, NEUROG1, PCDHGC5, DIAPH1, PHYKPL, COL23A1, GRM6, ADAMTS2, ZNF354C, ID4, GLO1, DNAH8, MDFI, FOXP4, GUCA1A, T AF8, TMEM63B, VEGFA, TFAP2B, COL12A1, TBX18, PREP, PRDM1, RNF217, OLIG3, HIVEP2, GPR126, OPRM1, IPCEF1, NXPH1, HOXA 13. EVX1, PPP1R35, TSC22D4, KCNH2, AOC1, ACTR3B, MNX1, NOM1, DNAJB6, PTPRN2, DLGAP2, TDRP, GFRA2, NKX2-6, EBF2, SOX17 One or more combinations of the following genes are included: BHLHE22, SLCO5A1, PRDM14, TBC1D31, ZHX1, FAM84B, OPLAH, SPATC1, DMRT2, DMRT3, NFIB, ZDHHC21, SLC24A2, CDKN2A, PAX5, MELK, DAPK1, NEK6, LHX2, NR5A1, GPR144, RAPGEF1, and MED27, wherein at least one methylation site of the target gene is selected from methylation sites within the target gene or in the regulatory region of the target gene; A2. The methylation sites of the methylation markers are located at chr1:933461:933661, chr1:3310706:3310906, chr1:3310740:3310940, chr1:29586145:29586345, chr1:29586305:29586505, chr1:29586580:29586780, chr1:50881931:50882131, chr1:50882097:50882297, chr1:50884300:50884500, chr1:50884536:50884736, chr1:50 884547:50884747、chr1:50886795:50886995、chr1:50886858:50887058、 chr1:63785793:63785993、chr1:63785890:63786090、chr1:87598124:87 598324, chr1:110334699:110334899, chr1:110334797:110334997, chr1: 119526811:119527011, chr1:119527250:119527450, chr1:119532703:11 9532903、chr1:119532788:119532988、chr1:119535692:119535892、chr1 :119535830:119536030、chr1:145562746:145562946、chr1:145562922:1 45563122, chr1:151811354:151811554, chr1:154375663:154375863, chr 1:155162584:155162784、chr1:156611838:156612038、chr1:156611899: 156612099, chr1:161039528:161039728, chr1:237205919:237206119, ch r1:237206012:237206212, chr10:7449719:7449919, chr10:8097474:809 7674, chr10:8097637:8097837, chr10:22542122:22542322, chr10:23480 625:23480825、chr10:23481046:23481246、chr10:108924091:108924291、chr10:110672177:110672377、chr10:118892523:118892723、chr10:119292117:119292317、chr10:119292290:119292490、chr10:119295843:119296043、chr10:119295949:119296149、chr10:129534694:129534894、chr10:129534735:129534935、chr10:130084908:130085108、chr10:131770965:131771165、chr10:131771113:131771313、chr10:133110769:133110969、chr11:281297:281497、chr11:636862:637062、chr11:640184:640384、chr11:31837311:31837511、chr11:31837411:31837611、chr11:44326261:44326461、chr11:44326290:44326490、chr11:58672824:58673024、chr11:63687058:63687258、chr11:63687159:63687359、chr11:125036402:125036602、chr11:125036413:125036613、chr12:6665072:6665272、chr12:6665295:6665495、chr12:6665353:6665553、chr12:30354393:30354593、chr12:30354424:30354624、chr12:33592774:33592974、chr12:52311647:52311847、chr12:52311685:52311885、chr12:52311699:52311899、chr12:52311791:52311991、chr12:113515300:113515500、chr12:113515340:113515540、chr12:113901298:113901498、chr12:113917423:113917623、chr12:113917466:113917666、chr12:115111826:115112026、chr12:115112017:115112217、chr12:115112199:115112399、chr12:115112263:115112463、chr13:25320388:25320588、chr13:25320456:25320656、chr13:25320480:25320680、chr13:36703403:36703603、chr13:36703431:36703631、chr13:53421052:53421252、chr13:112717238:112717438、chr13:112758741:112758941、chr13:112758754:112758954、chr14:37116133:37116333、chr14:37116288:37116488、chr14:37126872:37127072、chr14:37126878:37127078、chr14:38061327:38061527、chr14:38061506:38061706、chr14:38724555:38724755、chr14:38724657:38724857、chr14:38724773:38724973、chr14:38724835:38725035、chr14:52735051:52735251、chr14:52735129:52735329、chr14:57265398:57265598、chr14:57275744:57275944、chr14:57275823:57276023、chr14:60976665:60976865、chr14:60976752:60976952、chr14:73180907:73181107、chr14:73180980:73181180、chr14:97499688:97499888、chr14:97499865:97500065、chr15:27113022:27113222、chr15:27113157:27113357、chr15:27113277:27113477、chr15:29395897:29396097、chr15:41795038:41795238、chr15:45427262:45427462、chr15:45427338:45427538、chr15:68114350:68114550、chr15:75081248:75081448、chr16:11326998:11327198、chr16:11327011:11327211、chr16:11327045:11327245、chr16:22825689:22825889、chr16:22825962:22826162、chr16:23847490:23847690、chr16:23847558:23847758、chr16:31580122:31580322、chr16:31580153:31580353、chr16:33964856:33965056、chr16:33964877:33965077、chr16:50875086:50875286、chr16:50875166:50875366、chr16:51189898:51190098、chr16:51190057:51190257、chr16:57025884:57026084、chr16:57025993:57026193、chr16:73098547:73098747、chr16:82660460:82660660、chr16:82660574:82660774、chr16:86321495:86321695、chr16:86321635:86321835、chr16:88812168:88812368、chr16:88812249:88812449、chr17:35165517:35165717、chr17:36666092:36666292、chr17:46666849:46667049、chr17:46666888:46667088、chr17:48042351:48042551、chr17:48042487:48042687、chr17:55520563:55520763、chr17:55520580:55520780、chr17:55952088:55952288、chr17:76126991:76127191、chr18:49867020:49867220、chr18:49867062:49867262、chr19:16394438:16394638、chr19:16394477:16394677、chr19:19650947:19651147、chr19:41317811:41318011、chr2:468096:468296、chr2:468143:468343、chr2:468144:468344、chr2:5832783:5832983、chr2:5832891:5833091、chr2:5833431:5833631、chr2:5833466:5833666、chr2:19556785:19556985、chr2:25499956:25500156、chr2:30453572:30453772、chr2:71115853:71116053、chr2:71116009:71116209、chr2:71116129:71116329、chr2:71116151:71116351、chr2:106402858:106403058、chr2:106403016:106403216、chr2:111876461:111876661、chr2:162280397:162280597、chr2:162280466:162280666、chr2:176936260:176936460、chr2:176936280:176936480、chr2:176945337:176945537、chr2:176945519:176945719、chr2:176956558:176956758、chr2:176964760:176964960、chr2:176964879:176965079、chr2:176969332:176969532、chr2:176969529:176969729、chr2:176987322:176987522、chr2:176987451:176987651、chr2:176987506:176987706、chr2:176987518:176987718、chr2:176987530:176987730、chr2:176987577:176987777、chr2:176987656:176987856、chr2:177023006:177023206、chr2:177024290:177024490、chr2:177024578:177024778、chr2:177025034:177025234、chr2:177025100:177025300、chr2:177036795:177036995、chr2:193059315:193059515、chr2:209271252:209271452、chr2:209271 370:209271570、chr2:223163219:223163419、chr2:223163395:22316359 5, chr2:239755167:239755367, chr2:239755212:239755412, chr2:242824342:242824542, chr20:55202107:55202307, chr20:55202208:55202408 、chr20:55965081:55965281、chr20:55965104:55965304、chr21:223702 51:22370451、chr21:22370432:22370632、chr22:50623352:50623552、ch r22:50987078:50987278, chr22:50987203:50987403, chr22:51016318:51016518, chr22:51016376:51016576, chr3:9178082:9178282, chr3:9178 152:9178352、chr3:25469781:25469981、chr3:25469875:25470075、chr 3:38080591:38080791、chr3:38080707:38080907、chr3:49459532:49459 732、chr3:50377975:50378175、chr3:50378126:50378326、chr3:5037824 4:50378444、chr3:50378364:50378564、chr3:124860723:124860923、chr 3:124860887:124861087、chr3:129693578:129693778、chr3:147109862 :147110062、chr3:147109970:147110170、chr3:147109980:147110180、c hr3:147110112:147110312、chr3:147110159:147110359、chr3:15781213 1:157812331、chr3:157812236:157812436、chr3:157812282:157812482、chr3:157821224:157821424、chr3:157821404:157821604、chr3:181444 089:181444289、chr3:192126117:192126317、chr3:192126124:19212632 4、chr4:8863209:8863409、chr4:8863234:8863434、chr4:24801639:248 01839、chr4:24801688:24801888、chr4:24801787:24801987、chr4:24801 872:24802072、chr4:57521292:57521492、chr4:57521380:57521580、ch r4:57521683:57521883、chr4:57521812:57522012、chr4:154709519:154 709719、chr4:154709555:154709755、chr4:154709599:154709799、chr4: 190940255:190940455、chr5:1876269:1876469、chr5:2749400:2749600、 chr5:3599720:3599920、chr5:3599734:3599934、chr5:3602162:3602362、chr5:3606438:3606638、chr5:3606470:3606670、chr5:16179948:1618 0148、chr5:40681817:40682017、chr5:40681884:40682084、chr5:72677 312:72677512、chr5:72677376:72677576、chr5:72677387:72677587、chr 5:112073279:112073479、chr5:112073423:112073623、chr5:115152406 :115152606、chr5:115152437:115152637、chr5:132161430:132161630、c hr5:134870613:134870813、chr5:134870790:134870990、chr5:14089282 4:140893024、chr5:140892833:140893033、chr5:178003891:178004091、chr5:178421448:178421648、chr5:178421497:178421697、chr5:178770 871:178771071、chr5:178770944:178771144、chr6:10417560:10417760 、chr6:19691753:19691953、chr6:19691998:19692198、chr6:38683053: 38683253、chr6:38683180:38683380、chr6:41528461:41528661、chr6:4 1528742:41528942、chr6:42072440:42072640、chr6:44002224:4400242 4、chr6:50818154:50818354、chr6:75917956:75918156、chr6:75918095 :75918295、chr6:85476974:85477174、chr6:85477096:85477296、chr6: 106429583:106429783、chr6:125283715:125283915、chr6:125283869:12 5284069、chr6:137814694:137814894、chr6:143234709:143234909、chr 6:143234843:143235043、chr6:154360602:154360802、chr6:154360617 :154360817, chr7:8482114:8482314, chr7:8482213:8482413, chr7:27204968:27205168, chr7:27204978:27205178, chr7:27206030:27206230, c hr7:27244589:27244789, chr7:27244649:27244849, chr7:100075176:100075376, chr7:100075253:100075453, chr7:150655260:150655460, ch r7:150655362:150655562, chr7:152622494:152622694, chr7:152622512:152622712, chr7:156798388:156798588, chr7:157481934:157482134chr7:157482034:157482234, chr7:158110713:158110913, chr8:686927:687127, chr8:21647523:21647723, chr8:23564023:2356 4223, chr8:23564106:23564306, chr8:25907762:25907962, chr8:25907783:25907983, chr8:55370383:55370583, chr8:55370409 :55370609、chr8:55370821:55371021、chr8:55370874:55371074、chr8:65282197:65282397、chr8:65282231:65282431、chr8:709 81887:70982087、chr8:70981921:70982121、chr8:124173191:124173391、chr8:124173217:124173417、chr8:127569252:1275694 52. chr8:145105569:145105769, chr8:145105784:145105984, chr9:1042644:1042844, chr9:1042768:1042968, chr9:14346823:1 4347023, chr9:14346937:14347137, chr9:19788555:19788755, chr9:19789033:19789233, chr9:21974601:21974801, chr9:21974 Any one or more combinations of the following ranges: 779:21974979, chr9:36986323:36986523, chr9:90112714:90112914, chr9:90112885:90113085, chr9:126778279:126778479, chr9:126778444:126778644, chr9:127257997:127258197, chr9:127258138:127258338, chr9:134609102:134609302; A3. The nucleotide sequence of the methylation marker includes at least one nucleotide sequence as shown in SEQ ID NO. 1-365 or a complementary sequence of at least one nucleotide sequence as shown in SEQ ID NO. 1-365.
2. The methylation marker according to claim 1, characterized in that, The methylation markers described in A1 also include at least one of PAX6, OTX2, OSR1, IRX2, TFAP2A, and HOXA9.
3. Primers or probes for detecting the methylation markers according to any one of claims 1 or 2, characterized in that, The primers target the nucleotide sequence containing the methylation marker for specific amplification of the target sequence; the probes specifically capture the nucleotide sequence containing the methylation marker.
4. A DNA kit for detecting lung cancer, characterized in that, The kit includes reagents for detecting the methylation status or level of at least one methylation site of the methylation marker of claim 1.
5. The reagent kit according to claim 4, characterized in that, The regulatory region of the target gene is any continuous region of 100bp-400bp. Preferably, the regulatory region of the target gene is any continuous region of 150bp-300bp. More preferably, the regulatory region of the target gene is any continuous region of 200bp-250bp.
6. A DNA kit for detecting lung cancer, characterized in that, The kit includes at least one reagent for determining the methylation status or level of at least one methylation site of any one or more of the following target genes: OTX2, SORCS1, RCOR2, SYT10, SOX1, TMEM260, CILP2, PBX4, SOX11, SLC4A10, TBR1, KIAA1715, HOXD10, HOXD11, HOXD12, HOXD13, EVX2, VEPH1, SHOX2, IRX1, GRM6, OPRM1, IPCEF1, NXPH1, DNAJB6, PTPRN2, and SLC24A2; wherein at least one methylation site of the target gene is selected from methylation sites within the target gene or in the regulatory region of the target gene.
7. The reagent kit according to claim 6, characterized in that, The regulatory region of the target gene includes at least chr14:57265398-57576023, and genes selected from chr10:108924091-108924291, chr11:63687058-63687258, chr11:63687159-63687359, chr12:33592774-33592974, and chr13:112717238-112717438. chr19:19650947-19651147, chr2:5833431-5833631, chr2:5833466-5833666, chr2:162280466-16 2280666、chr2:176945519-176945719、chr2:176987506-176987706、chr3:157812131-157812331、 chr3:157812236-157812436, chr3:157812282-157812482, chr5:3599720-3599920, chr5:3606438 -3606638、chr5:3606470-3606670、chr5:178421448-178421648、chr5:178421497-178421697、chr 6:154360602-154360802, chr6:154360617-154360817, chr7:8482114-8482314, chr7:8482213-8482413, chr7:157481934-157482134, chr7:157482034-15748223, chr9:19788555-19788755, any one or a combination thereof.
8. The reagent kit according to claim 6, characterized in that, The regulatory region of the target gene includes at least any or a combination of chr14:57265398-57265598, chr14:57275744-57275944, chr14:57275823-57276023, and chr10:108924091-108924291, chr11:63687058-63687258, chr11:63687159-63687359, and chr12:335 92774-33592974, chr13:112717238-112717438, chr19:19650947-19651147, chr2:5833431-5833631, chr2 :5833466-5833666、chr2:162280466-162280666、chr2:176945519-176945719、chr2:176987506-17698770 6. chr3:157812131-157812331, chr3:157812236-157812436, chr3:157812282-157812482, chr5:3599720- 3599920, chr5:3606438-3606638, chr5:3606470-3606670, chr5:178421448-178421648, chr5:178421497- Any or a combination of 178421697, chr6:154360602-154360802, chr6:154360617-154360817, chr7:8482114-8482314, chr7:8482213-8482413, chr7:157481934-157482134, chr7:157482034-15748223, and chr9:19788555-19788755.
9. The reagent kit according to claim 6, characterized in that, The regulatory region of the target gene includes at least one sequence or its complementary sequence from the sequences shown in SEQ ID NO: 106, 107, and 108, and a sequence or its complementary sequence selected from any or more of the sequences shown in SEQ ID NO: 41, 62, 63, 71, 91, 162, 169, 170, 182, 186, 194, 243, 244, 245, 267, 270, 271, 288, 289, 312, 313, 314, 315, 328, 329, and 354.
10. The use of the kit according to any one of claims 4-9 in constructing a lung cancer prediction system.