Methylation markers for detecting lung nodule malignancy and applications thereof
By screening 35 methylation markers associated with pulmonary nodules in plasma and constructing a machine learning model, the problem of poor diagnostic accuracy of pulmonary nodules in existing technologies was solved, and efficient and low-cost differentiation between benign and malignant pulmonary nodules was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SINGLERA GENOMICS (SHANGHAI) LTD
- Filing Date
- 2022-03-16
- Publication Date
- 2026-07-31
AI Technical Summary
Existing methods for diagnosing pulmonary nodules are insufficient to accurately distinguish between benign and malignant pulmonary nodules, resulting in high false positive rates, increased patient suffering, and waste of medical resources. Current cfDNA methylation tests also perform poorly in the diagnosis of pulmonary nodules.
The methylation levels of specific methylation markers in plasma were detected using various detection methods such as next-generation sequencing and quantitative real-time PCR. Thirty-five methylation markers associated with benign and malignant lung nodules were screened out, and a machine learning model was constructed for identification.
It achieves highly sensitive and specific non-invasive diagnosis of benign and malignant pulmonary nodules, reduces detection costs, and provides an accurate risk assessment model.
Smart Images

Figure CN116804218B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, and in particular to methylation markers for detecting benign or malignant pulmonary nodules and their applications. Background Technology
[0002] Lung cancer is a malignant tumor of the lungs, mostly originating from lung epithelial cells. It is the cancer with the highest incidence and mortality rate, with the majority (approximately 85%) being non-small cell lung cancer (NSCLC) (Sun D et al., 2020). Although current treatments such as immunotherapy and targeted therapy have significantly improved the survival rate of NSCLC patients, the prognosis remains poor. Advanced NSCLC is prone to metastasis, with a five-year survival rate of only 4.5% after metastasis, compared to 55.6% for stage I carcinoma in situ (Weichert W et al., 2014). Unfortunately, early-stage NSCLC patients often have no obvious symptoms, and more than 80% of patients are diagnosed with NSCLC at an advanced stage, accompanied by lymph node spread or distant metastasis, resulting in low survival rates. Therefore, developing effective methods for early diagnosis of NSCLC is of great significance.
[0003] Currently, the most widely used early lung cancer screening method in clinical practice is low-dose CT (LDCT). LDCT can detect early-stage NSCLC patients and reduce lung cancer mortality by 20% (National Lung Screening Trial Research T et al., 2011). LDCT has high sensitivity (93.7%), but its specificity is low. In LDCT detection, stage I lung cancer usually presents as pulmonary nodules, and its image characteristics are not significantly different from benign pulmonary nodules. It is impossible to accurately distinguish between benign and malignant pulmonary nodules from LDCT images, and the false positive rate for early diagnosis reaches 96.4%. Therefore, patients diagnosed with pulmonary nodules in the early stage by LDCT usually need long-term follow-up and multiple LDCT re-examinations, or invasive tests such as aspiration or surgical biopsy for further confirmation. These subsequent diagnosis and treatment significantly increase patient suffering and waste medical resources due to over-diagnosis.
[0004] Many auxiliary detection methods can improve the accuracy of pulmonary nodule diagnosis: the Herder model (AUC = 0.92) uses fluorodeoxyglucose positron emission tomography (FDG-PET) to assist LDCT detection (Herder GJ et al., 2005); the Mayo model (AUC = 0.83) uses image features such as nodule size, margins, morphology, and location, along with patient characteristics such as age, smoking history, and family history of genetic diseases, to provide supplementary assessment (Swensen SJ et al., 1997). In addition, many molecular biological detection methods are increasingly being applied to pulmonary nodule diagnosis, such as the detection of circulating tumor cells and cancer markers (Liu QX et al., 2021). Currently, there is still a need to develop more effective, convenient, and accurate methods to assist in the diagnosis of pulmonary nodules.
[0005] DNA methylation is closely related to the occurrence, development, diagnosis, and treatment of cancer. Benign and malignant pulmonary nodules have different DNA methylation characteristics, which can be used for the diagnosis of pulmonary nodules (Kerr KM et al., 2007). Liquid biopsy refers to a non-invasive diagnostic technique that utilizes biomarkers in body fluids. Current research has confirmed that body fluids (blood, urine, bronchoalveolar lavage fluid, etc.) contain a large amount of cell-free DNA (cfDNA), which has tissue-specific DNA methylation characteristics, enabling tissue tracing and cancer detection (Lo YMD et al., 2021). cfDNA methylation detection is now applied to the detection of various cancers, such as colorectal cancer and liver cancer, and can detect early-stage cancer with high specificity and sensitivity. Some literature has found that cfDNA methylation can be used in the early diagnosis and screening of lung cancer (AUC=0.90) (Liang N et al., 2021), but it performs poorly in the diagnosis of benign and malignant lung nodules (AUC=0.72-0.81) (Liang W et al., 2021). Therefore, there is an urgent need to find more accurate liquid biopsy methylation biomarkers for the diagnosis of benign and malignant lung nodules.
[0006] In view of this, the present invention is proposed. Summary of the Invention
[0007] The purpose of this invention is to provide a methylation biomarker for detecting benign and malignant pulmonary nodules and its application. This invention utilizes the detection of methylation levels of methylation biomarkers in plasma to differentiate between patients with benign and malignant pulmonary nodules, achieving a higher accuracy and lower cost for non-invasive and precise diagnosis of pulmonary nodules.
[0008] In order to achieve the above-mentioned objectives of the present invention, the following technical solution is adopted:
[0009] In one aspect, the present invention provides a biomarker for benign and malignant methylation of pulmonary nodules, the biomarker comprising the following biomarkers:
[0010] (a1) A fragment of at least one of the sequences shown in Seq ID No:1 to Seq ID No:35 or their complete complement sequences;
[0011] (a2) A fragment comprising the sequence described in (a1) and its upstream and downstream 1kb sequences;
[0012] DNA methylation sites in the regions of sequences shown in (a3), (a1), or (a2);
[0013] (a4), variants that share at least 90% sequence identity with (a1) or (a2) and have the same methylation sites; or
[0014] DNA methylation haplotypes covered in the regions of sequences shown in (a5), (a1), or (a2), and the abundance of said DNA methylation haplotypes.
[0015] In one embodiment, the methylation markers for benign and malignant pulmonary nodules include one or more or all of the following sequences, or one or more complementary sequences of the following sequences: Seq ID No: 1, Seq ID No: 2, Seq ID No: 3, Seq ID No: 4, Seq ID No: 5, Seq ID No: 6, Seq ID No: 7, Seq ID No: 8, Seq ID No: 9, Seq ID No: 10, Seq ID No: 11, Seq ID No: 12, Seq ID No: 13, Seq ID No: 14, Seq ID No: 15, Seq ID No: 16, Seq ID No: 17, Seq ID No: 18, Seq ID No: 19, Seq ID No: 20, Seq ID No: 21, Seq ID No: 22, Seq ID No: 23, Seq ID No: 24, Seq ID No: 25, Seq ID No: 26, Seq ID No: 27, Seq ID No: 28, Seq ID No: 2 ...0, Seq ID No: 21, Seq ID No: 2 No: 28, Seq ID No: 29, Seq ID No: 30, Seq ID No: 31, Seq ID No: 32, Seq ID No: 33, Seq ID No: 34 and Seq ID No: 35.The sequences are located in the genome at the following positions: chr1:119527250-119527450, chr2:19556785-19556985, chr2:71115853-71116351, chr2:176936260-176936480, chr2:176945337-176945719, chr2:176956558-176956758, chr2:176969332-176969729, chr2:223163219-223163595. chr3:50377975-50378564, chr3:157812131-157812482, chr5:1876269-1876469, chr5:3599720-3599934, chr5:3602162-360 2362, chr5:115152406-115152637, chr6:85476974-85477296, chr6:137814694-137814894, chr7:152622494-152622712, chr8 :23564023-23564306、chr8:65282197-65282431、chr8:124173191-124173417、chr10:7449719-7449919、chr10:8097474-809 7837、chr10:129534694-129534935、chr10:130084908-130085108、chr11:636862-637062、chr11:125036402-125036613、chr 13:53421052-53421252, chr14:37126872-37127078, chr14:38724555-38725035, chr14:52735051-52735329, chr14:5726539 8-57265598, chr16:51189898-51190257, chr16:82660460-82660774, chr17:46666849-46667088, chr17:48042351-48042687.
[0016] In one embodiment, the benign or malignant methylation markers for lung nodules include fragments comprising the sequences shown in Seq ID No:1 to Seq ID No:35 and their upstream and downstream 1kb sequences; or DNA methylation sites in the region of the shown sequences; variants having at least 90% sequence identity with the shown sequences and having the same methylation sites; or DNA methylation haplotypes covered in the region of the shown sequences and the abundance of said DNA methylation haplotypes.
[0017] In this invention, the methylation biomarker uses the methylation status / level of the CpG island-containing regions / fragments / methylation sites involved in this invention to distinguish between patients with benign and malignant pulmonary nodules. The methylation level of the methylation biomarker can be obtained by a variety of detection methods known in the art, such as, but not limited to, obtaining the methylation level through next-generation sequencing.
[0018] In this invention, the sequence number in the genome corresponds to its position on the UCSC (https: / / genome.ucsc.edu / cgi-bin / hgTracks?db=hg19) HG19 genome.
[0019] The methylation biomarkers described in this invention can be used alone or in combination as lung cancer-related methylation molecular markers for the detection or auxiliary detection or identification of benign or malignant lung nodules and / or lung cancer.
[0020] In another aspect, the present invention includes reagents for detecting the degree of methylation of the aforementioned methylation markers; for example, primers or probes for the aforementioned methylation markers, wherein the primers target the nucleotide sequence in which the methylation marker is located for specific amplification of the target sequence; and the probes specifically capture the nucleotide sequence in which the methylation marker is located.
[0021] In another aspect, the present invention provides a kit for detecting benign or malignant pulmonary nodules, the kit containing reagents for detecting the aforementioned methylation markers of benign or malignant pulmonary nodules.
[0022] In another aspect, the present invention provides the use of reagents for detecting the methylation levels of the aforementioned methylation markers of benign and malignant pulmonary nodules in the preparation of diagnostic kits for identifying benign and malignant pulmonary nodules in samples.
[0023] In another aspect, the present invention provides the application of detecting the methylation level of the aforementioned methylation markers of benign and malignant pulmonary nodules in the diagnosis of benign and malignant pulmonary nodules.
[0024] In one embodiment, the detection reagent further includes reagents used in any one or a combination of PCR amplification, quantitative real-time PCR, digital PCR, liquid phase-array assay, next-generation sequencing, third-generation sequencing, bisulfite sequencing, whole-genome methylation sequencing, and methylation array assay. For example, the detection reagent may be selected from the following: bisulfite and its derivatives, PCR buffer, polymerase, dNTPs, primers, probes, methylation-sensitive or insensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporter agents, exonucleases, alkaline phosphatase, internal standards, and controls.
[0025] In one embodiment, the biomarker is obtained from a liquid sample, said liquid sample being blood from mammals or other biological sources, such as peripheral blood, serum, plasma, ascites, urine, cerebrospinal fluid, sputum, saliva, etc. The mammals include rats, mice, and humans; preferably humans. In one embodiment, the sample is a fine-needle aspiration biopsy or plasma. The sample includes genomic DNA or cfDNA. In one embodiment, the sample preferably refers to plasma cfDNA. cfDNA refers to circulating cell-free DNA or cell-free DNA, degraded DNA fragments released into the plasma.
[0026] In one embodiment, the present invention provides biomarkers for screening benign and malignant pulmonary nodules based on cfDNA methylation.
[0027] In this invention, the terms "benign" and "malignant" refer to the nature of the pulmonary nodules. The pulmonary nodules include solid, partially solid, or ground-glass nodules, preferably partially solid or ground-glass nodules. "Malignant" pulmonary nodules generally refer to those exhibiting cancerous changes.
[0028] In another aspect, the present invention provides a method for constructing a predictive assessment model for benign and malignant pulmonary nodules, comprising the following steps:
[0029] (a) Collect benign and malignant lung nodule samples, and divide them into training and test sets;
[0030] (b) Extract cfDNA from the sample, construct a library, and sequence it;
[0031] (c) Perform methylation transformation on the sequence, compare the data, and calculate the AMF and MHF values of the sample;
[0032] (d) Construct a feature matrix from the sample data; build an algorithm model and screen out methylation markers of benign and malignant lung nodules based on the training set samples;
[0033] (e) Validate the model's performance using test set samples;
[0034] (f) Identify the methylation markers that will ultimately be used for the assessment of benign and malignant lung nodules.
[0035] In step (d), a logistic regression model is constructed for each feature in the training set to distinguish between malignant and benign lung nodules. The average AUC of the 3-fold cross-validation is calculated and sorted from high to low. The remaining features are added to the feature set in turn, and the logistic regression model is reconstructed. If the average AUC of the 5-fold cross-validation of the logistic regression model increases, the feature is retained; otherwise, it is removed. The optimal combination of biomarkers is obtained and modeling is performed using the optimal combination.
[0036] In one implementation, the algorithm model includes a machine learning model, which employs any one of principal component analysis, logistic regression, nearest neighbor analysis, support vector machine, and neural network models; preferably, it is a logistic regression model.
[0037] In another aspect, the present invention also provides a risk assessment model for benign or malignant pulmonary nodules, said risk assessment model being obtained using the aforementioned construction method, comprising:
[0038] The data acquisition module is used at least for data acquisition and obtaining sample datasets;
[0039] The sequencing module is used, at least, to obtain sequencing data based on the construction of sequencing libraries.
[0040] The data alignment module is at least used to align the sequencing data with a reference sequence and determine the methylation result of the markers in the sequencing data based on the alignment result;
[0041] The result determination module is at least used to calculate the predicted score threshold or the diagnostic score threshold through statistical model analysis to determine whether the sample to be tested is a benign or malignant pulmonary nodule; the statistical model is a logistic regression model.
[0042] The risk assessment model may also include a methylation module, at least for performing methylation to obtain methylated data.
[0043] This invention uses regression analysis to correlate the methylation levels of selected biomarkers with the benign or malignant nature of nodules, constructing a regression model to obtain a risk assessment or diagnostic model for benign or malignant pulmonary nodules. A logistic regression machine learning model constructed using the methylation levels of 35 methylation biomarkers from this invention can be used to differentiate between benign and malignant pulmonary nodules.
[0044] In one implementation, a new machine learning model is constructed using 13 methylation biomarkers: Seq ID NO:1, Seq ID NO:3, Seq ID NO:5, Seq ID NO:7, Seq ID NO:8, Seq ID NO:10, Seq ID NO:11, Seq ID NO:13, Seq ID NO:14, Seq ID NO:16, Seq ID NO:18, Seq ID NO:25, and Seq ID NO:27.
[0045] In another implementation, a machine learning model was constructed using 10 methylation markers: Seq ID NO:4, Seq ID NO:9, Seq ID NO:12, Seq ID NO:17, Seq ID NO:23, Seq ID NO:26, Seq ID NO:29, Seq ID NO:31, Seq ID NO:33, and Seq ID NO:35.
[0046] Furthermore, the present invention also provides an information data processing terminal, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs the following steps:
[0047] (a) Obtain the methylation level of at least one region of the Seq ID No:1 to Seq ID No:35 sequence in the sample to be tested; (b) Calculate the score by constructing a logistic regression diagnostic model; (c) Identify the benign or malignant nature of the pulmonary nodules based on the score.
[0048] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the following steps:
[0049] (a) Obtain the methylation level of at least one region of the Seq ID No:1 to Seq ID No:35 sequence in the sample to be tested; (b) Calculate the score by constructing a logistic regression diagnostic model; (c) Identify the benign or malignant nature of the pulmonary nodules based on the score.
[0050] A method for detecting benign or malignant pulmonary nodules for non-disease diagnostic purposes, comprising the following steps:
[0051] S1: Detect the methylation level of at least one region of the Seq ID No:1 to Seq ID No:35 sequences in the sample;
[0052] S2: The score is calculated based on the aforementioned risk assessment model for benign and malignant pulmonary nodules.
[0053] S3: Determine the benign or malignant nature of thyroid nodules based on the scoring.
[0054] Specifically, a nodule is identified as malignant when the methylation level of the target sequence in the sample meets a certain threshold. For example, for a sample to be tested, if the score is greater than the threshold, the result is positive, indicating a malignant nodule; otherwise, it is negative, indicating a benign nodule.
[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0056] This invention, for the first time, screened 35 methylation biomarkers with highly high correlation to the benign and malignant nature of pulmonary nodules based on high-throughput cfDNA methylation sequencing of plasma. These biomarkers can effectively distinguish between benign and malignant pulmonary nodules with high sensitivity and specificity. Verification has shown that the methylation levels of individual methylation biomarkers in this invention also have good classification effects. These biomarkers can be used to establish a risk prediction and assessment model for benign and malignant pulmonary nodules / lung cancer, for the purpose of differentiating between benign and malignant pulmonary nodules.
[0057] This invention also establishes a risk prediction and assessment model or diagnostic model based on the relationship between the methylation level of biomarkers and the benign or malignant nature of lung nodules. The model distinguishes between benign and malignant lung nodules by the prediction score. The model has the advantages of non-invasive detection, safe and convenient detection, high throughput, and high detection accuracy. Based on the methylation biomarkers obtained by this invention, detection costs can be effectively controlled while achieving good detection performance. Attached Figure Description
[0058] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0059] Figure 1 Screening process for methylation markers to identify benign and malignant pulmonary nodules;
[0060] Figure 2 Box plots are used to display the methylation levels of 35 methylation markers in the test and training sets;
[0061] Figure 3 The box plot shows the methylation level of Seq ID No:5 in the test and training sets;
[0062] Figure 4 Predict score distribution for the AllModel diagnostic model;
[0063] Figure 5The ROC curve is used to demonstrate the diagnostic performance of the AllModel model;
[0064] Figure 6 Predict the score distribution for the Sub1 diagnostic model;
[0065] Figure 7 The ROC curve is used to demonstrate the effectiveness of the Sub1 diagnostic model;
[0066] Figure 8 Predict the score distribution for the Sub2 diagnostic model;
[0067] Figure 9 The ROC curve is used to demonstrate the effectiveness of the Sub2 diagnostic model. Detailed Implementation
[0068] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] Unless otherwise specified in the examples, the procedures should be performed under standard conditions or conditions recommended by the manufacturer. Reagents or instruments whose manufacturers are not specified are all commercially available products.
[0070] Example 1: Methylation-targeted sequencing for screening benign and malignant methylation biomarkers in lung nodules
[0071] A total of 306 patients with pulmonary nodules were collected, and all enrolled patients signed informed consent forms. These samples were divided into training and test sets according to a certain ratio. The training set was used to build the machine learning model described below, and the test set was used to test the model's performance. Sample information is shown in Table 1 below.
[0072] Table 1. Statistics of Sample Information from Training and Test Sets
[0073]
[0074] This application obtains methylation sequencing data of plasma cfDNA from samples using the Methyl-Titan method (patent number: CN201910515830) and screens out methylation markers within it. The specific screening process is as follows:
[0075] 1. Extraction of plasma cfDNA samples
[0076] A 2 mL whole blood sample was collected from the patient using a streck blood collection tube. Plasma was separated by centrifugation within 3 days and then transferred to the laboratory. cfDNA was extracted using the QIAGEN QIAamp Circulating Nucleic Acid Kit according to the instructions.
[0077] 2. Sequencing and Data Preprocessing
[0078] The library was sequenced at 150bp paired ends using an Illumina Nextseq 500 sequencer, with a sequencing volume of at least 5M.
[0079] The Pear (v0.6.0) software merges paired-end sequencing data of the same fragment (150bp) from Illumina Hiseq X10 / Nextseq 500 / Novaseq sequencers into a single sequence with a minimum overlap length of 20bp and a minimum length of 30bp after merging.
[0080] The merged sequencing data were de-adapted using Trim_galore v 0.6.0 and cutadapt v1.8.1 software. The adapter sequence “AGATCGGAAGAGCAC” was removed from the 5' end of the sequence, and bases with sequencing quality values below 20 at both ends were removed.
[0081] 3. Sequencing data alignment
[0082] The reference genome data used in this article came from the UCSC database (UCSC:HG19, http: / / hgdownload.soe.ucsc.edu / goldenPath / hg19 / bigZips / hg19.fa.gz).
[0083] (a) First, HG19 was transformed into cytosine to thymine (CT) and adenine to guanine (GA) using Bismark software, and the transformed genomes were indexed using Bowtie2 software.
[0084] (b) The preprocessed data is also converted using CT and GA;
[0085] (c) Using Bowtie2 software, the transformed sequences were aligned to the transformed HG19 reference genome. The minimum seed sequence length was 20, and mismatches in the seed sequence were not allowed.
[0086] 4. Calculation of AMF and MHF for each sample
[0087] For each target methylation region HG19 CpG site, the methylation state corresponding to each site is obtained based on the above comparison results.
[0088] (a) Calculate the average methylation rate (AMF) value for the target methylation region. The formula for calculating AMF is as follows:
[0089]
[0090] Where M is the total number of CpG sites in the target methylation region, i is the number of CpG sites within the region, and N is the number of sites within the region. C,i N represents the number of reads sequenced as C at this CpG site (i.e., the number of methylated reads). T,i The number of reads sequenced as T for this CpG site (i.e., the number of unmethylated sequencing reads).
[0091] (b) Calculate the MHF value for the target methylation region. A target methylation region may have multiple methylation haplotypes. This value needs to be calculated for each methylation haplotype within the target region. An example of the MHF calculation formula is shown below:
[0092]
[0093] Where l represents the target methylation region, h represents the target methylation haplotype, Nl represents the number of reads located in the target methylation region, and Nl,h represents the number of reads containing the target methylation haplotype.
[0094] 5. Feature Matrix Construction
[0095] (a) Merge the AMF of each target methylation region of each sample in the training set and test set respectively. The MHF value is the feature matrix of the training set and test set, and target methylation regions with fewer than 100 reads are treated as missing values.
[0096] (b) Remove target methylation regions with a missing value ratio higher than 10%;
[0097] (c) Use the KNN algorithm to train a transformer for the training set matrix, and use the transformer to impute missing data in the feature matrices of the training and test sets.
[0098] 6. Identify methylation markers for benign and malignant pulmonary nodules based on the training set samples (see...). Figure 1 )
[0099] (a) In the training set, a logistic regression model is constructed for each feature to distinguish between malignant and benign pulmonary nodules. The average AUC of the 3-fold cross-validation is calculated and sorted from high to low.
[0100] (b) Add the remaining features to the feature set in sequence and rebuild the logistic regression model;
[0101] (c) If the average AUC of the 5-fold cross-validation of the logistic regression model increases, then the feature is retained; otherwise, it is removed.
[0102] (d) After traversing all features, the optimal combination of markers is obtained, and the optimal combination is used for modeling. Finally, the test set samples are used to verify the effect of the model.
[0103] 7. A total of 35 methylation markers for benign and malignant pulmonary nodules were screened out during the above process.
[0104] The above process screened 35 methylation biomarkers for benign and malignant lung nodules. Individual or combinations of these biomarkers can be used to differentiate between benign and malignant lung nodules. The methylation biomarker-associated gene refers to the gene corresponding to the nearest transcription start site (TSS) within 100 kb of the biomarker; specific associated genes are shown in Table 2. The genomic location of the methylation biomarker refers to its position within the UCSC.
[0105] (https: / / genome.ucsc.edu / cgi-bin / hgTracks?db=hg19) HG19 genome location.
[0106] This invention provides 35 methylation biomarkers. The methylation levels of these 35 methylation biomarkers in malignant and benign pulmonary nodules in the training and test sets are shown in Table 2 below. Figure 2 As shown. For benign and malignant pulmonary nodule samples in the training and test sets, the methylation levels of methylation markers were calculated separately. The median of each category was used as the methylation level for that category. The statistical significance of the difference in methylation between malignant and benign pulmonary nodules in the training and test sets was calculated using 'Wilcox.test'. If the p-value was <0.05, the methylation marker was considered to have a statistically significant difference in methylation between malignant and benign pulmonary nodule samples.
[0107] Of the 35 methylation biomarkers provided in this invention, all 35 showed significant methylation differences in the training set samples, and 30 showed significant methylation differences in the test set samples. These results indicate that the 35 methylation biomarkers we selected can also effectively distinguish between malignant and benign pulmonary nodule samples based on their methylation levels.
[0108] Using Seq ID No:5 as an example, this paper details the methylation levels of this methylation marker in malignant and benign pulmonary nodules in both the training and test sets. Figure 3 The methylation marker showed highly significant differences in methylation between malignant and benign pulmonary nodule samples in both the training and test sets, with a P-value of 5.8E-15 in the training set and 3.0E-05 in the test set.
[0109] Table 2. Statistical analysis and tests of methylation levels of 35 methylation biomarkers in the training and test sets.
[0110]
[0111]
[0112]
[0113] Example 2: AllModel - A Machine Learning Diagnostic Model for All Methylation Markers
[0114] In this embodiment, a logistic regression machine learning model is constructed using the methylation levels of 35 methylation markers from this invention to distinguish between benign and malignant pulmonary nodules.
[0115] The model is trained using samples from the training set in Example 1, and then the model's performance is tested using samples from the test set. The specific steps are as follows:
[0116] Using the logistic regression model from the sklearn (V1.0.1) package in Python (V3.9.7): AllModel = LogisticRegression();
[0117] Use the training set samples for training: AllModel.fit(Traindata,TrainPheno), where TrainData is the data in the training set, TrainPheno is the phenotype of the training set samples (1 for malignant nodules, 0 for benign nodules), and the relevant thresholds of the model are determined based on the training set samples.
[0118] Test using samples from the test set: TestPred = AllModel.predict_proba(TestData)[:,1], where TestData is the test set data and TestPred is the model's predicted score. Use this predicted score and the above threshold to determine whether the sample is a malignant lung nodule.
[0119] The model prediction scores for the training and test sets are shown below. Figure 4 The figure shows a significant difference in model scores between benign and malignant pulmonary nodule samples. The ROC curve is shown below. Figure 5 The AUC on the training set was 0.951, and the AUC on the test set was 0.904. A prediction score threshold of 0.662 was set; a score greater than this value was predicted as a malignant pulmonary nodule, and a score less than this value was predicted as a benign pulmonary nodule. At this threshold, the accuracy on the test set was 0.815, the specificity was 0.833, and the sensitivity was 0.806 (see Table 3).
[0120] Example 3: Machine Learning Diagnostic Model Sub1 for Random Methylation Marker Combination 1
[0121] To verify the effectiveness of the random methylation biomarker combination, this embodiment selected 13 methylation biomarkers (Seq ID No:1, Seq ID No:3, Seq ID No:5, Seq ID No:7, Seq ID No:8, Seq ID No:10, Seq ID No:11, Seq ID No:13, Seq ID No:14, Seq ID No:16, Seq ID No:18, Seq ID No:25, and Seq ID No:27) from 35 methylation biomarkers to construct a new machine learning model.
[0122] The method for constructing the machine learning model is consistent with Example 2. The features used only include the methylation levels of the aforementioned 13 methylation markers. The prediction scores of this model on the training and test sets are shown below. Figure 6 ROC curves are shown below. Figure 7 The model showed a significant difference in predicted scores between malignant and benign pulmonary nodules on both the training and test sets. The AUC on the training set was 0.933, and the AUC on the test set was 0.877. A prediction score threshold of 0.602 was set; scores above this threshold were predicted as malignant, and scores below this threshold were predicted as benign. At this threshold, the test set accuracy was 0.815, specificity was 0.733, and sensitivity was 0.855 (see Table 3).
[0123] Example 4: Machine Learning Diagnostic Model Sub2 for Random Methylation Marker Combinations 2
[0124] In this embodiment, an additional 10 methylation biomarkers were selected from the 35 methylation biomarkers in this invention to construct a machine learning model. The selected methylation biomarkers are: Seq ID No:4, Seq ID No:9, Seq ID No:12, Seq ID No:17, Seq ID No:23, Seq ID No:26, Seq ID No:29, Seq ID No:31, Seq ID No:33, and Seq ID No:35.
[0125] The method for constructing the machine learning model is consistent with Example 2. The features used only include the methylation levels of the aforementioned 10 methylation markers. The prediction scores of this model on the training and test sets are shown below. Figure 8 ROC curves are shown below. Figure 9The model showed a significant difference in predicted scores between malignant and benign pulmonary nodules on both the training and test sets. The AUC on the training set was 0.890, and the AUC on the test set was 0.889. A prediction score threshold of 0.610 was set; scores above this threshold were predicted as malignant, and scores below this threshold were predicted as benign. At this threshold, the accuracy on the test set was 0.815, the specificity was 0.700, and the sensitivity was 0.871 (see Table 3).
[0126] Table 3. Performance of Machine Learning Diagnostic Models
[0127]
[0128] Example 5: Effect of a single methylation marker model
[0129] This invention found that the methylation level of a single methylation biomarker among 35 methylation biomarkers also had a good classification effect. The effect of a single methylation biomarker in distinguishing between malignant and benign pulmonary nodules is shown in Table 4. A logistic regression diagnostic model was constructed using only Seq IDNo:5, with a training set AUC of 0.83, a test set AUC of 0.77, and a diagnostic score threshold of 0.68. The test set accuracy was 0.75, specificity was 0.7, and sensitivity was 0.77.
[0130] Table 4. Diagnostic effectiveness of single methylation biomarker models
[0131]
[0132]
[0133] This invention screened 35 methylation markers to differentiate between benign and malignant pulmonary nodules. Machine learning models constructed based on the methylation levels of different combinations or individual methylation markers of these markers can effectively distinguish between malignant and benign pulmonary nodule samples, providing an important reference for the diagnosis of benign and malignant pulmonary nodules.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention. SEQUENCE LISTING <110> Jiangsu Kunyuan Biotechnology Co., Ltd. <120> Methylation biomarkers for detecting benign and malignant pulmonary nodules and their applications <130> JH‐CNP220324 <160> 35 <170> PatentIn version 3.3 <210> 1 <211> 200 <212> DNA <213> Homo Sapiens <400> 1 ctccgaggct ggtcttgagg cacttctcta gtagcttctc caaaagactg agagtgccgg 60 cgtaggtatg acagtgaggg tacctcacag acccttctcc aaagtctggc gggccttggg 120 gttttcggg gccaccaggc tcggtggaat ttttgaaacg cttcgaaat acatagttc 180 ctctgtggag tgagtgccta 200 <210> 2 <211> 200 <212> DNA <213> Homo Sapiens <400> 2 gtccggaccc ccgtgaagag gactgaccgg caccggatac ttctatagca ttctcccaac 60 aaacgagatc taacgaaccc attggcaagg cggtcatccg gctgcactta aatgccgct 120 gcgtcctcgg tgatccattc cccaatctta ataaaacagc aattacctcg aggagcctgg 180 gatggaacat ctacacgccg 200 <210> 3 <211> 498 <212> DNA <213> Homo Sapiens <400> 3 ttttgactgt ttgtttgcct aagagatgac ttcccttgcc agaaaaaaaa aagtgttgta 60 aaaataaaag gaaacgggat tacgatgtaa aagacgaata gataaacccg gttccgcaga 120 tctgcggcgc gcgcgcctg cgacctcgga tacattcatt gaagttgccg cgcactcgta 180 cccgggttca cctcgcccct cgctcattcc tcccgaccaa ggcccatggt cagaggtgtc 240 ctcgccccgc ggccgtcaga gggcgcggcc tacactcgaa tgccggccga gccctccacg 300 cgctcggaac ttgggctttc cggtgcagcc tcccgcgat cgcaatgccc gctgccttc 360 ccgagcccag tccggaaccc gcctctcg gggaccttga cctcgcgcgg acctcgtcgc 420 cttgcttagc tctccctgac ctgctttga acccgctaat cccgtcccga acctcgccc 480 cgcggcaatc cctcgccc 498 <210> 4 <211> 220 <212> DNA <213> Homo Sapiens <400> 4 ggaatgtgtg agggtgggga gcgatcgtgc tggagaagac cgggcgcaaa cagtcgctgg 60 ggagattgca gcctttaagc ttttttctt attcgcatct tctggcttct ctcttccgtc 120 gaaccctttt ggcaaccgca gggaacacgc atcctcaggt tggcacggga ggcggcgagg 180 agctcccgga gccaccggct gctggatccc cctctccccc 220 <210> 5 <211> 382 <212> DNA <213> Homo Sapiens <400> 5 acgaagccgc ggccgccgcg cctgaggctg cagccgcggc cgccgccgcc gtgacgccca 60 cgtgcgggta gtagtgcagc ggcacgtgcg agtggaaggg gtagggcagg cttccggtgg 120 cggccgcgtg cgtcatcatg taggtgtaga agctggggtc ggctgggtgc ggccaggaca 180 tggccaggcg ctgccgcttg tccttcatgc gccggttctg gaaccacacc tgcggggaga 240 gacgcgccgc agcctgggtt agggagcgcc ccgtgttccc agctcctgtc ccaggacctc 300 tgccccttcc ggacctctga atggcttggt ctacttctct ccgaccaagc ccaaccccga 360 gtaccctgtg gtctcccagc tg 382 <210> 6 <211> 200 <212> DNA <213> Homo Sapiens <400> 6 acacagaagt ttcacggtgg gaggctgagt ggctttctcc cccggcgccg ttctcagggt 60 ctttctgcgg gtcgaagaag gacccgcggg agctgagagg cccaggtcgg aagcactccc 120 ggctggccca agagtagagg cgaagagcgt tgagtaggca tccatggact tttctttctg 180 ggacagattg tcaggctcac 200 <210> 7 <211> 397 <212> DNA <213> Homo Sapiens <400> 7 ccacagagtc ggggtcttca cggtaggttc tcgagcggga cgcgcgggtc cggaggctgc 60 ggttttccct gggtttgggg aatgggggta ggaactagga gggagctggg gccaaagagc 120 caagcgggct gggactggaa tgaaagcgct ctgggttgtg gagtgggtcg gggggcaagg 180 gtccgcgcta aggagccgaa aggggccggc cgcccccttc ccctatgcac cggcgcgcca 240 ctgcagatgg ctcaccctcc cccgccaaat cgctgctccc gccggagctg ccgtcgccat 300 gtcgttgaac ttgaatggtt cgtcctcccc ttttccctct tcggcggtca gcaggctcca 360 tactgggggc gccgggctgc ctgcacccca ccccgtg 397 <210> 8 <211> 376 <212> DNA <213> Homo Sapiens <400> 8 tccaggcggc gcgctgaggc cctcccttac cttccagcgg gaacccgcta cgcgggtagt 60 tctgccccgg gcccggccgc atcatcctgg gcacagcgcc ggccagcgtg gtcatcctgg 120 gggcagcttc gctcggaaat tatatccagg tgaaggcgaa acggaaaggc gagtgcggcg 180 cggatgaccc tcgggaacta tccggagcgt ggagagcccc tccccaaaac ggctggagag 240 agagggaggg acgcggggag gggggctgtc ggttcctagt ccagaggccg gagctggaac 300 ccgggaaagg ggaggacggg gaggccccgg agtccaggat cccgagccca gggcggaaaa 360 gtttggtacg agtctg 376 <210> 9 <211> 589 <212> DNA <213> Homo Sapiens <400> 9 ccactactca cgcgcgcact gcaggccttt gcgcacgacg ccccagatga agtcgccaca 60 gaggtcgcac cacgtgtgcg tggcgggccc cgcgggctgg aagcggtggc cacggccagg 120 gaccagctgc cgtgtggggt tgcacgcggt gcccgcgcg atgcgcagcg cgtggcacg 180 ctccagccgg gtgcggccct tcccagcgcg cccagcgggt gccagctccc gcagctcaat 240 gagctcaggc tcccccgaca tggcccggtt gggcccgtgc ttcgctggct ttgggcgcta 300 gcaagcgcgg gccgggcggg gccacagggc gggccccgac ttcagcgcct cccccaggat 360 ccagactggg cggcgggaag gagctgagga gagccgcgca atggaaacct gggtgcaggg 420 actgtggggc ccgaaggcgg ggctgggcgc gctctcgcag agcccccccc gccttgccct 480 tccttccctc cttcgtcccc tcctcacacc ccaccccgga cggccacaac gacggcgacc 540 gcaaagcacc acgcggagat acccgtgttt ctggaggcca gctttactg 589 <210> 10 <211> 351 <212> DNA <213> Homo Sapiens <400> 10 gttagtcttg ggggagcaga cgcaagagga ggcaagggcg ccgcgagctc cccggatgca 60 ctggtcccac aggccgtgcc cgagtggagc actgcgaatg gggccaagaa attttggcct 120 ttctcgccgg acctggctgc ctccgcgggc ctctccgcct accgcgctcc cgccgcggcc 180 cgactcccgc gggtctccgc gccgaaccca cctggctcct atcgcacggg acattcccga 240 cccacccacg ccgcgtcact gagcctctgt accgataccc ggcgcctccg ccagcagggc 300 ctggacgcac cgcctccttt gacctcgggc ttccccgcg ctccgctgct t 351 <210> 11 <211> 200 <212> DNA <213> Homo Sapiens <400> 11 taccttggct gatgctcgct cgccaggccg ggggctccg ccgcagcctt ttgacaggca 60 catgagccgc gagcttcga acctcgataa tatcatctcg agcgcgaaag tcaatacggt 120 gacagcgcgc ggccggatac aatccaatta cgctcggctg cccgggcgct cctggggctc 180 ggggtccggc ggccgagggt 200 <210> 12 <211> 214 <212> DNA <213> Homo Sapiens <400> 12 acgaggagat cgacctggaa agcatcgaca ttgacaagat cgacgagcac gatggcgacc 60 agagcaacga ggatgacgag gacaaggccg aggctccgca cgcgcccgca gccccttctg 120 ctcttgcccg ggaccaaggc tcgccgctgg cagcagccga cgttctcaag ccccaggact 180 cgcccttggg cctggcaaag gaggccccag agcc 214 <210> 13 <211> 200 <212> DNA <213> Homo Sapiens <400> 13 ggcttctctg cggtgtgagg tgtgaacgag gggctgaggc tgtggtggga agcgagaaag aggaggtggc tttggtctcc cggggagcc cctttacact tgggctccac ggactgcgtc 120 ctttgccctc aggcgcgcgc accgcgggag tccagagcaa attgccctta gatggccgcg 180 gccggggcagc ggggaggcag 200 <210> 14 <211> 231 <212> DNA <213> Homo Sapiens <400> 14 taccaacggc cctcggaggc aggacgaggc ggagagccac ccaagaaagg tggcggaggc ggggagaccc tgcgggcacg gctcacgcgc acatccccgg cttccccgggg ctccgcgcct 120 tcccaagagc cccgttgtct ccggcgtccc agggatcgcg tgggctccgc gcaatctctc 180 ccccacttta acggcgcgtt ttagccgccc ggcctaacgc ctctccccgc t 231 <210> 15 <211> 322 <212> DNA <213> Homo Sapiens <400> 15 60. gaggggcagt gtgtcacaga ctgcagccca ctccaacctc ggctcctgga gaaggggcgt cgaatctctc ttgggcatgg gagggaaaga cattccgagt tggctgggcg gagtggcagc 120 cttgagagtg acgagtgaca gcaaagcctc gtcctagcaa ggccttttac caacagcgcg 180 gcatgccctt tcgaggagag cgccaggccc tcgcactttg caagtcaaga gagcaaagaa 240 agcggggaca gggcgcgtaa tcgcaatgtc cggtcgcgcg tgtgcacgtg tctgtgtttg 300 catgtgtgcg tgagcatgtg ca 322 <210> 16 <211> 200 <212> DNA <213> Homo Sapiens <400> 16 cgggaagtga ggcggcggac agcggtgacg aggcgttgcc agatgagagc gcgccgccca 60 agatggggtg caccgggtgc acggagttgg ccgcgtgcgc ggggtggccg gccgagtggc 120 ccacggtccc gcagtgaaag gccgagtggt ggcccccata gatctcgcca accagcctct 180 tcatctcctc cagggagctg 200 <210> 17 <211> 218 <212> DNA <213> Homo Sapiens <400> 17 gagggaacgg tggaggccgc cgttctcttg ggcgcggcct ctgctggggg acggcgggga 60 tcgcagggcc gaggggccgg gcgcgcgcgg gggagggacc caggccaggg gccgctcgcc 120 tcggggcggg gtctctggag agcgcgcaga ggggacgctt cgtgaatgcc tgccggcctg 180 aaggatgtcg ctgatctctg cccccttcgc caaggccg 218 <210> 18 <211> 283 <212> DNA <213> Homo Sapiens <400> 18 gcgaagccgc ggggcagctc cgctcgcgct ccagtcgcag gatgtccttg accgagaagg 60 gggtggaggt gacggggctc agcagcatcc cgaaggcgga tggggcgggg ccgaggaggt 120 ccgggtgagg agcggcaccc tgaacttccc gtcttgtcgc tgcaggcccc gcagacagac 180 ccaagctctg ggacagacgc ccagcgtccc agacagcgcc ttcctctggg ccatgctggt 240 aggcccgggt ccagggccgg gtgacgagac cgtagccccc cat 283 <210> 19 <211> 234 <212> DNA <213> Homo Sapiens <400> 19 cccgccaggg ccgcatctgc ggcccgagac ttacttctgc tttgctttcc tgcacctggc 60 ccggggcagg ctttctagtt tctgcctcgg ccagaggagc cgacttccgc aggacgggct 120 tggggcaggg ggtgccggtg ggggcagcgg ggcgaggcga ggtcgcaggg agacgaccag 180 gcctgcttgg gcctgggccg ctccagccgc ctcgcctctt cagggacctc gccc 234 <210> 20 <211> 226 <212> DNA <213> Homo Sapiens <400> 20 cgggagggga ccgctggctg caggcgccct gctggagtcc ggaacgaggg tggtcagggt 60 cgcctcccca cggggtcgca gatgagggag tggggccgcg cagggaggct cctgcaggcc 120 gtgaccccca ggcctggccc gtagagagtg actgatggcg aggacccgct ggggacccga 180 gggcctcagc ctgcgcctgc gccggggctc cccgaccaag accttc 226 <210> 21 <211> 200 <212> DNA <213> Homo Sapiens <400> 21 tccagcgtcg tcccgggatt ctcggacacc acaaacgcca tcaaccacga gcaccggtgt 60 ccgtggctat tgccccgaat ggtccccatc cgcgtccccg ggaactccct cggcttttcg 120 cgcatccagg tccccagccc cagctactgg tgcgccccga gcccctaggt gccagagcgg 180 tggtcggccg ggctcctgcc 200 <210> 22 <211> 363 <212> DNA <213> Homo Sapiens <400> 22 ttcccatccc cccaccgaaa gcaaatcatt caacgacccc cgaccctccg acggcaggag ccccccgacc tcccaggcgg accgccctcc ctccccgcgc gcggggttccg ggcccggcga gagggcgcga gcacagccga ggccatggag gtgacggcgg accagccgcg ctgggtgagc 240. caccaccacc ccgccgtgct caacggggcag cacccggaca cgcaccaccc gggcctcagc cactcctaca tggacgcggc gcagtacccg ctgccggagg aggtggatgt gctttttaac 300 360. atcgacggtc aaggcaacca cgtcccgccc tactacgga actcggtcag ggccacggtg cag 363 <210> 23 <211> 241 <212> DNA <213> Homo Sapiens <400> 23 ccgaggtcac agcgcggaga aacctggcgg ggccccggac tccccggctt gggaaaagcg atgactgccc tgaactgctg gggcgttcga aatttccagg gtcccgaccc tccgtggggt 120 acgcgcgact tcggcgcaga tgtcagtccg ctgccttccg ggttgaggga gcgaggactc 180 cagacgaccc cagggccgct gtccaggccc agccccgcgt tctccacctc gccacctcgc 240 t 241 <210> 24 <211> 200 <212> DNA <213> Homo Sapiens <400> 24 cgtagttgtc tcctggctcc tggggtccgc ggagctctag atgtacctgc agctcctccc 60 gagtcctgca agccaccctt gtccctcttc tcccgctcac cccccggccc ccccatctct 120 tttgctattc cggggaaggc cacgcagggt gcaacccgga cgcgcccccg ggggaagccc 180 gcgacgcagc agccacaccc 200 <210> 25 <211> 200 <212> DNA <213> Homo Sapiens <400> 25 cctgggaggg aggttttgcc agataccagg tggactaggg tgagcgcccg agggccggga 60 cgcacgcacg ggccgggtag gatggcgctg gcgtcgatgc ccgcgcgctt cagggcctgg 120 tctggccgcc cctccatcct tgtcggtttc tcgggtcgcg gaccccgcgc ggcgccgggc 180 gatgctggcc tgcccgtggc 200 <210> 26 <211> 211 <212> DNA <213> Homo Sapiens <400> 26 tgaggtcccg acccaggcgg ctcggagtgc tccaggagcc acctgggtct gcgggcgcag 60 cgcggcgggg cgggagcggt ggcccgcagg ggccgcggcc tgcgatgaag gccggggggc 120 agcgctagca gcgaggtgcc acagtgggcc gaggagtctg ggctgtggcc cagggtagga 180 ccggctcaaa ctccagtgcc ctgattggag c 211 <210> 27 <211> 200 <212> DNA <213> Homo Sapiens <400> 27 gcaccgacac ctcatagacc ggccgcgtga agagcggcgc gttgtcgttc tcgtcgccca 60 cacgcaccgt gtagggccgc actgtgcgca gcgggggcgc gccgcgatcc tcggccacca 120 gcgtcaagtt gtactcggcg atgcgttcgc ggtccagcga cgccgcggtc accaccaggt 180 agctgcccgc gtaggccggc 200 <210> 28 <211> 206 <212> DNA <213> Homo Sapiens <400> 28 cgcagaacgc cgggacgcat ttccgagctc gggcccgacc caacccgctg ccattctagc 60 tacctagtgg agccgcggag agactcttga accgtgaggc tccttgatga gaaggtggag 120 aggctgccgg gctattctcc gagcatcttc cggccgagct ctgctcctct ccgtgcgctt 180 gcggcttctc cttgggctag cgcgtc 206 <210> 29 <211> 480 <212> DNA <213> Homo Sapiens <400> 29 aactgagatc gggagctgtc cccggcagag cgcactcacc tcggtcccag gtggactgaa 60 gtccagagcg gcgctgtgca gctggaaggg cgcgcgatag ctcaagttag aggcggcccc 120 ggggcgcggc gcaggacaca agacctcaaa ctggtacttg cacaggtagc cgttggcgcg 180 caggtggcat cgcatctcct tccagcctgc gggctcgacc ccaccggtgg cctggagtac 240 cgcgcatctc cgcgcggtgc aggagcgttg gggctcctcc acccactgca gcgtgtcgct 300 ttcgagaccg ccggggtcgg aggacagcca ggagaaaccc cgcaaaggct cgttctccag 360 ggtgcagtgg gaacgcctgc gctccagtgc gacccagaac agcaggtctt tggagccccc 420 tccgggccct gggcctgccc gcaggagcgc gagcacagcg cgcagctcgg cgcccgcacg 480 <210> 30 <211> 278 <212> DNA <213> Homo Sapiens <400> 30 ttcgtgcagt actgccccgg cacctggtgc tttatccaga tggtccacga ggagggctcg 60 ctgtcggtgc tggggtactc tgtgctctac tccagcctca tggcgctgct ggtcctcgcc 120 accgtgctgt gcaacctcgg cgccatgcgc aacctctatg cgatgcaccg gcggctgcag 180 cggcacccgc gctcctgcac cagggactgt gccgagccgc gcgcggacgg gagggaagcg 240 tcccctcagc ccctggagga gctggatcac ctcctgct 278 <210> 31 <211> 200 <212> DNA <213> Homo Sapiens <400> 31 ctgttcccag gtaactccga ctgggcactg gggagttaga aaagccagct ctttagccag 60 agcgcctagg gcgcggcgga gagcgggccg cccggcacca cgttccttct ggcagttccg 120 ccccagcctc ccagcgtctt gcgcctgtgg cggcggcagt acgggcctgg ggggtcaccc 180 acaggaaagg tgagattagc 200 <210> 32 <211> 359 <212> DNA <213> Homo Sapiens <400> 32 caataccagc agcaaaacgc aggtttttgg gggaactccc gccgcccgcc accaagggct 60 atctccagac gggcgccggg tgcagcgccg tgaccgggcg ccctggcgcc ggctcgggcg 120 cgaaattcag cggtggcaag cggagggtgg gcttggtaac cacccgcgcg cgcccgagcc 180 aagagtcgcg tactgtctgc ccgcggcaaa gttcgtcttt ctccgcttgg agggctgttc 240 ctacaccggt attaagaaac cgacttcgct agcgactgca agtgcttgcg attttgactt 300 tccgtccaca gttgagcgtc ttgcacttaa attcactgcg ccccgcatgc aacagtgcc 359 <210> 33 <211> 314 <212> DNA <213> Homo Sapiens <400> 33 ccccgtatct gccatgcaaa acgagggagc gttaggaagg aatccgtctt gtaaagccat 60 tggtcctggt catcagcctc tacccaatgc tttcgtgatg ctgctgctga tctatttggg 120 aagttggctg gctggcgagg cagagcctct cctcaaagcc tggctcccac ggaaaatatg 180 ctcagtgcag ccgcgtgcat gaatgaaaac gccgccgggc gcttctagtc ggacaaaatg 240 cagccgagaa ctccgctcgt tctgtgcgtt ctcctgtccc aggtagggaa gaggggctgc 300 cgggcgcgct ctgc 314 <210> 34 <211> 239 <212> DNA <213> Homo sapiens <400> 34 ccaacaggct gaaggtggag gcaaagcag cgctgtgctc cccacaggct agaagcctga 60 tgaagccaaa gataaccgga gggaccctt tttatggct tcaagatcg gaaagccaga 120 gatcacccca gggcaagta gcccacagca gccgcctacc cggtgcttt tctcagctca 180 gatgtatcca gataaaagtg cccagcggtg caccccaca cacctccgac aaatgaatg 239 <210> 35 <211> 336 <212> DNA <213> Homo sapiens <400> 35 ccgaattaat cagttctctc aacctgagtt actaagaag aaagtcctt ccaataaaa 60 ctgaaaatca ctgcgaatga caatactata ctacaagttc gttttggggc cggtgggtgg 120 gatgaggag aaagggcacg gatatcccg gaggggccgcg gagtgaggg gactatggtc 180 gcggtggaat ctctgttccg ctggcacatc cgcgcaggtg cggctctgag tgctgctcg 240 gggttacaga cctcggcatc cggctgcagg ggcagacaga gacctctct gctagggcgt 300 gcggtaggca tcgtatggag cccagagact gccgag 336
Claims
1. Methylation markers for detecting lung nodule malignancy, characterized in that, The methylation markers are sequence combinations shown in Seq ID No:1 to Seq ID No:35, or sequence combinations shown in Seq ID No:1, Seq ID No:3, Seq ID No:5, Seq ID No:7, Seq ID No:8, Seq ID No:10, Seq ID No:11, Seq ID No:13, Seq ID No:14, Seq ID No:16, Seq ID No:18, Seq ID No:25 and Seq ID No:27, or sequence combinations shown in Seq ID No:4, Seq ID No:9, Seq ID No:12, Seq ID No:17, Seq ID No:23, Seq ID No:26, Seq ID No:29, Seq ID No:31, Seq ID No:33 and Seq ID No:
35.
2. The use of the reagent for detecting the methylation level of the methylation markers of benign and malignant pulmonary nodules as described in claim 1 in the preparation of a diagnostic kit for identifying benign and malignant pulmonary nodules in a sample.
3. Use according to claim 2, characterized in that, The detection reagents also include reagents used in any one or a combination of PCR amplification, liquid phase chip method, second-generation sequencing, third-generation sequencing, bisulfite sequencing, whole genome methylation sequencing, and methylation chip method.
4. Use according to claim 3, characterized in that, The PCR amplification method includes quantitative real-time PCR or digital PCR.
5. A lung nodule benign-malignant risk assessment model, characterized in that, The risk assessment model includes: a data acquisition module, used at least to acquire a sample dataset; and a sequencing module, used at least to obtain sequencing data, wherein the sequencing data is the methylation level of the sequence regions shown in Seq ID No:1 to Seq ID No:35 in the sample, or the methylation level of the sequence regions shown in Seq ID No:1, Seq ID No:3, Seq ID No:5, Seq ID No:7, Seq ID No:8, Seq ID No:10, Seq ID No:11, Seq ID No:13, Seq ID No:14, Seq ID No:16, Seq ID No:18, Seq ID No:25 and Seq ID No:27, or the methylation level of the sequence regions shown in Seq ID No:4, Seq ID No:9, Seq ID No:12, Seq ID No:17, Seq ID No:23, Seq ID No:26, Seq ID No:29, Seq ID No:31, Seq ID No:33 and Seq ID No:
27. The methylation level of the sequence region shown in No:35; the data comparison module, which is at least used to compare the sequencing data with the reference sequence and determine the methylation result of the marker in the sequencing data based on the comparison result; the result determination module, which is at least used to calculate the prediction score threshold through statistical model analysis and determine whether the sample to be tested is a benign or malignant lung nodule.
6. An information data processing terminal comprising a memory, a processor and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it performs the following steps: (a) Obtain the methylation levels of the sequence regions shown in Seq ID No:1 to Seq ID No:35 in the sample to be tested, or the methylation levels of the sequence regions shown in Seq ID No:1, Seq ID No:3, Seq ID No:5, Seq ID No:7, Seq ID No:8, Seq ID No:10, Seq ID No:11, Seq ID No:13, Seq ID No:14, Seq ID No:16, Seq ID No:18, Seq ID No:25 and Seq ID No:27, or the methylation levels of the sequence regions shown in Seq ID No:4, Seq ID No:9, Seq ID No:12, Seq ID No:17, Seq ID No:23, Seq ID No:26, Seq ID No:29, Seq ID No:31, Seq ID No:33 and Seq ID No:
27. (a) Methylation level of the sequence region shown in No:35; (b) Score calculated by constructing a logistic regression diagnostic model; (c) Identification of benign or malignant pulmonary nodules based on the score.
7. A computer-readable storage medium having stored thereon a computer program, characterized in that When the computer program is executed by the processor, the following steps are performed: (a) Obtain the methylation levels of the sequence regions shown in Seq ID No:1 to Seq ID No:35 in the sample to be tested, or the methylation levels of the sequence regions shown in Seq ID No:1, Seq ID No:3, Seq ID No:5, Seq ID No:7, Seq ID No:8, Seq ID No:10, Seq ID No:11, Seq ID No:13, Seq ID No:14, Seq ID No:16, Seq ID No:18, Seq ID No:25 and Seq ID No:27, or the methylation levels of the sequence regions shown in Seq ID No:4, Seq ID No:9, Seq ID No:12, Seq ID No:17, Seq ID No:23, Seq ID No:26, Seq ID No:29, Seq ID No:31, Seq ID No:33 and Seq ID No:35; (b) Calculate the score by constructing a logistic regression diagnostic model; (c) Identify the benign or malignant nature of the pulmonary nodules based on the score.