Methylation Biomarkers for Distinguishing Benign and Malignant Pulmonary Nodules and Their Applications

By performing DNA differential methylation analysis on lung nodules tissue samples, 20 methylated biomarkers were selected to construct a predictive model, and a random forest algorithm was used to solve the problem of high false positives in early screening of lung cancer in the prior art, achieving high sensitivity and specific early diagnosis of lung cancer.

CN119464493BActive Publication Date: 2025-07-18SHANDONG PROVINCIAL HOSPITAL AFFILIATED TO SHANDONG FIRST MEDICAL UNIVERSITY (SHANDONG PROVINCIAL HOSPITAL) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411531338.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-07-18
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

The prior art has high sensitivity and low specificity problems in early screening of lung cancer, resulting in a high false positive rate, lack of sensitivity of existing tumor markers, and it is difficult to effectively distinguish between benign and malignant pulmonary nodules.

Method used

By performing DNA differential methylation analysis on lung nodules tissue samples, the optimal 20 methylated biomarkers were selected to construct a prediction model, and using a random forest algorithm for prediction, combined with sulfite sequencing and DNA methylation sequencing technology, a random forest classification model was constructed to distinguish benign and malignant pulmonary nodules.

Benefits of technology

The model has 100% sensitivity and specificity in the training center, 100% sensitivity and 91% specificity in the test center, which can effectively distinguish lung malignant tumors from benign lung nodules and has important early screening value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119464493B_ABST
    Figure CN119464493B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical fields of disease diagnosis and molecular biology, and particularly relates to methylation biomarkers for differentiating benign and malignant pulmonary nodules and their applications. The present invention performs DNA differential methylation analysis on tissue samples of pulmonary nodules, and constructs a prediction model by selecting the optimal 20 methylation biomarkers according to sensitivity. This prediction model has high sensitivity and specificity, can distinguish pulmonary malignant nodules (i.e., lung tumors) from pulmonary benign nodules based on tissue samples, predict the risk of lung cancer occurrence, has high model accuracy and good sensitivity, and is of great value for early screening of lung cancer patients, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of disease diagnosis and molecular biology, and particularly relates to methylation biomarkers for differentiating benign and malignant lung nodules and their applications. Background Art

[0002] Disclosing the information of this background art section is only intended to enhance the overall understanding of the present invention and does not necessarily constitute an admission or imply in any form that this information constitutes the prior art already known to those of ordinary skill in the art.

[0003] Lung cancer, as one of the most common cancers in the world, is a major cause of cancer-related deaths worldwide. Since early-stage lung cancer has basically no typical symptoms, most patients are diagnosed at an advanced stage, and the survival rate is significantly reduced. Research shows that the 5-year survival rate of stage I lung cancer is greater than 80%, while the 5-year survival rate of stage IV lung cancer is almost 0. Therefore, for the prognosis and treatment of lung cancer patients, early lung cancer screening and early diagnosis tools with high sensitivity and specificity are crucial.

[0004] Currently, low-dose spiral CT (LDCT) screening is widely used in the early detection of lung cancer and has become an effective method for early lung cancer screening. It can detect potential lung nodules at an early stage and effectively reduce the lung cancer mortality rate. However, due to only considering nodule size and density and ignoring other risk factors such as the basic medical history of the lungs, LDCT screening has the characteristics of high sensitivity and low specificity, and its false positive rate is as high as 96.4%, being able to detect many non-tumorous lung nodules. False positive results will cause unnecessary interventions and excessive anxiety, thus causing potential harm to patients. To solve the current dilemmas, people have tried to use common tumor markers for diagnosis, such as squamous cell carcinoma antigen (SCC-Ag), carcinoembryonic antigen (CEA), carbohydrate antigen 125 (CA125), and cytokeratin 19 fragment (CYFRA21-1), but they lack the corresponding sensitivity, and early lung cancer screening remains a huge challenge.

[0005] With the continuous deepening of research, abnormal DNA methylation is considered an important cause of cancer development. DNA methylation is an epigenetic modification of genes in which a methyl group is covalently added to the carbon 5 position of cytosine bases by DNA methyltransferases, mainly occurring in CpG islands rich in CpG dinucleotides. Most tumors are characterized by global DNA hypomethylation and local specific hypermethylation, which was found by Suzuki et al. to be particularly prominent in the early stages of non-small cell lung cancer. Hypermethylation in the promoter region usually silences tumor suppressor genes, while DNA hypomethylation activates proto-oncogenes, leading to genetic instability and thus promoting the growth of tumor cells. Therefore, early detection of DNA methylation has the potential to predict the occurrence of cancer. However, there are still few reports on the study of lung cancer screening based on DNA methylation at present. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention provides methylation biomarkers for differentiating benign and malignant lung nodules and their applications. The present invention performs DNA differential methylation analysis on tissue samples of lung nodules, and selects the optimal 20 methylation biomarkers according to sensitivity to construct a prediction model. This prediction model has high sensitivity and specificity, can distinguish lung malignant nodules (i.e., lung tumors) from lung benign nodules based on tissue samples, predict the risk of lung cancer occurrence, and has high model accuracy and good sensitivity. Based on the above research results, the present invention is thus completed.

[0007] Specifically, the present invention relates to the following technical solutions:

[0008] In the first aspect of the present invention, there is provided a methylation biomarker for differentiating benign and malignant lung nodules, and the methylation biomarker is selected from one or more of the following gene fragments:

[0009] chr5:93905357-93905359, chr5:134259480-134259482, chr5:134259489-134259491, chr5:93905256-93905258, chr11:67247741-67247743, chr1:243645964-243645966, chr5:134259522-134259524, chr15:45928449-45928451, chr5:154026439-154026441, chr5:154026479-154026481, chr8:42398664-42398666, chr1:47899804-47899806,

[0010] chr2:85362588-85362590, chr2:130875117-130875119, chr11:71189489-71189491, chr11:76887102-76887104, chr5:134259500-134259502, chr5:134259527-134259529, chr11:18808624-18808626, and chr15:25093827-25093829.

[0011] Furthermore, the methylation biomarker is a group composed of all the above gene fragments.

[0012] The present invention discovers through research that the sensitivity of a single gene may not be significant. However, by selecting the 20 genes with the highest sensitivity on the premise of relatively high specificity and combining them, the prediction model for benign and malignant pulmonary nodules of the present invention has extremely high sensitivity and specificity, thus greatly improving the prediction ability of the model.

[0013] In a second aspect of the present invention, there is provided the use of a reagent for detecting the methylation level of the above methylation biomarker in the preparation of a product for distinguishing between benign and malignant pulmonary nodules.

[0014] In a third aspect of the present invention, there is provided a system for distinguishing between benign and malignant pulmonary nodules, the system comprising:

[0015] i) An analysis module, the analysis module comprising: a detection reagent for determining the methylation biomarker selected from the above in a test sample of a subject;

[0016] ii) An evaluation module, the evaluation module comprising: differentiating and judging the benignity and malignancy of the test sample according to the expression level of the methylation biomarker determined in i).

[0017] More specifically, the random forest model is a random forest classification model constructed by RandomForestClassifier, and its parameter settings are: criterion = 'entropy', max_depth = 5, n_estimators = 10, oob_score = True, random_state = 124. The malignant prediction probability threshold is 0.7. For each sample, based on the DNA methylation level of its biomarker combination, the malignant prediction probability is calculated using the constructed random forest regression model. If the malignant prediction probability is greater than the threshold, it is judged as a malignant nodule; otherwise, it is judged as a benign nodule.

[0018] The beneficial technical effects of the above one or more technical solutions:

[0019] The above technical solution analyzes DNA methylation of actual collected lung nodule tissue samples through bisulfite sequencing. The optimal markers are screened according to the sensitivity and specificity of differentially methylated genes in patients with benign nodules and malignant nodules, and a prediction model is constructed and verified using the random forest algorithm. Finally, enrichment analysis of differentially methylated genes is performed using the GO and KEGG databases. This model has high accuracy and good sensitivity. The sensitivity of this model in the training set is 100%, the specificity is 100%, the AUC is 1.000, the sensitivity in the test set is 100%, the specificity is 91%, and the AUC is 0.99. It can distinguish lung malignancies from lung benign nodules based on tissue samples, which has important value for the early screening of lung cancer patients, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0021] Figure 1 It is the basic analysis of methylation sequencing data of tissue samples of lung nodule patients in the embodiments of the present invention. (A) Data quality control from the perspective of the proportion of valid data. (B) Data quality control from the perspective of sequencing coverage depth.

[0022] Figure 2 It is the receiver operating characteristic curve and the area under the curve of the training set and the test set under different numbers of markers in the embodiments of the present invention. Sorted in descending order according to sensitivity, the optimal n (10 - 100, with every 10 as a gradient) markers are selected, and a binary classification algorithm is trained in the training set (A) and the model is evaluated in the test set (B) using the random forest algorithm.

[0023] Figure 3 It is the heat map of methylation levels of the top 100 most significant markers in the training set in all samples (training set & test set) in the embodiments of the present invention. The colored bars represent the methylation levels of the corresponding markers. The bluer the color, the lower the methylation level, and the redder the color, the higher the methylation level.

[0024] Figure 4 It is the model performance detection under the top 20 markers in the embodiments of the present invention. (A) Heat map of methylation levels of the top 20 most significant markers in all samples (training set & test set) in the training set. The colored bars represent the methylation levels of the corresponding markers. The bluer the color, the lower the methylation level, and the redder the color, the higher the methylation level. (B) Receiver operating characteristic curve and the area under the curve of the training set and the test set under the top 20 markers. (C and D) Sensitivity and specificity of the training set (C) and the test set (D).

[0025] Figure 5GO functional enrichment analysis of differentially methylated genes in the embodiments of the present invention. (A and B) Cellular components. GO enrichment bubble chart (A) and directed acyclic graph (DAG) (B). (C and D) Molecular functions. GO enrichment bubble chart (C) and directed acyclic graph (D). (E and F) Biological processes. GO enrichment bubble chart (E) and directed acyclic graph (F). The GO enrichment bubble chart of differentially methylated genes reflects the number distribution of related genes on the GO terms enriched in biological processes and the enrichment significance p-value; the directed acyclic graph, where the branches represent inclusion relationships, and the defined function range becomes smaller from top to bottom, and the depth of color represents the enrichment degree.

[0026] Figure 6 KEGG pathway enrichment analysis of differentially methylated genes in the embodiments of the present invention. Detailed implementation manners

[0027] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0028] It should be noted that the terms used herein are only for describing the specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof. If the experimental methods in the following specific implementation manners are not specified with specific conditions, they are generally carried out according to the conventional methods and conditions in the field of molecular biology, and such techniques and conditions are fully explained in the literature. See, for example, the techniques and conditions described in Sambrook et al., "Molecular Cloning: A Laboratory Manual", or according to the conditions recommended by the manufacturer.

[0029] The present invention will be further described in conjunction with specific examples. The following examples are only for explaining the present invention and do not limit its content. If the specific experimental conditions are not indicated in the examples, they are generally carried out according to the conventional conditions or according to the conditions recommended by the sales company; the materials, reagents, etc. used in the examples, unless otherwise specified, can be obtained through commercial channels.

[0030] The term "biomarker" refers to "a characteristic that can be objectively detected and evaluated and can serve as an indicator of normal biological processes, pathological processes, or pharmacological responses to therapeutic interventions". For example, nucleic acid biomarkers (which can also be called gene biomarkers, such as DNA), protein biomarkers, cytokine markers, chemokine markers, carbohydrate biomarkers, antigen biomarkers, antibody biomarkers, etc. In the present invention, the biomarker is a methylation biomarker.

[0031] In the present invention, the terms "benign" and "malignant" indicate the nature of pulmonary nodules. Generally, benign nodules are characterized by a round or oval shape, good mobility, mostly without lobulation, and may have sharp corners or fibrous cords at the edges. The edges of inflammatory pulmonary nodules are mostly blurred, while the edges of benign non-inflammatory pulmonary nodules are mostly clear, regular, and even smooth. Malignant nodules show irregular shapes, mostly lobulated, or have spiculation (or spiny protrusions), pleural indentation signs, and vascular convergence signs, as well as uncontrollable growth, spread, and metastasis of malignant cells. The edges are mostly clear but not smooth, and the nodule-lung interface is rough or even spiculated. In a specific embodiment of the present invention, malignant pulmonary nodules are lung cancers.

[0032] In the present invention, differentiating between benign and malignant pulmonary nodules can include not only early screening for lung cancer patients but also late diagnosis of lung cancer, as well as risk assessment, prognosis, and disease identification of lung cancer.

[0033] The early screening refers to the possibility of detecting lung cancer before metastasis, especially before observable morphological changes in tissues or cells.

[0034] In a typical specific embodiment of the present invention, a methylation biomarker for differentiating between benign and malignant pulmonary nodules is provided. The methylation biomarker is selected from one or more of the following gene fragments: chr5:93905357-93905359, chr5:134259480-134259482, chr5:134259489-134259491, chr5:93905256-93905258, chr11:67247741-67247743, chr1:243645964-243645966, chr5:134259522-134259524, chr15:45928449-45928451, chr5:154026439-154026441, chr5:154026479-154026481, chr8:42398664-42398666, chr1:47899804-47899806,

[0035] chr2:85362588-85362590, chr2:130875117-130875119, chr11:71189489-71189491, chr11:76887102-76887104, chr5:134259500-134259502, chr5:134259527-134259529, chr11:18808624-18808626, and chr15:25093827-25093829.

[0036] Further, the methylation biomarker is a group composed of all the above gene fragments.

[0037] The present invention discovers through research that the sensitivity of a single gene may not be significant. However, by selecting the 20 genes with the highest sensitivity under the premise of relatively high specificity and combining them, the prediction model for benign and malignant pulmonary nodules of the present invention has extremely high sensitivity and specificity, thereby greatly improving the prediction ability of the model.

[0038] In the second aspect of the present invention, there is provided an application of a reagent for detecting the methylation level of the above methylation biomarker in the preparation of a product for distinguishing between benign and malignant pulmonary nodules.

[0039] Among them, the reagent for detecting the methylation level of the methylation biomarker includes reagents required for bisulfite conversion and DNA methylation sequencing.

[0040] The product for distinguishing between benign and malignant pulmonary nodules may specifically be a detection kit.

[0041] In the third aspect of the present invention, there is provided a system for distinguishing between benign and malignant pulmonary nodules, the system comprising:

[0042] i) An analysis module, the analysis module comprising: a detection reagent for determining the methylation biomarker selected from the above in a test sample of a subject;

[0043] ii) An evaluation module, the evaluation module comprising: differentiating and judging the benign or malignant nature of the test sample according to the methylation expression level of the methylation biomarker determined in i);

[0044] In the analysis module of i), the test sample is a pulmonary nodule tissue sample.

[0045] The specific evaluation process of the evaluation module in ii) includes: differentiating and judging the benign or malignant nature of the test sample based on the prediction model according to the methylation level of the methylation biomarker determined in i); the prediction model is a random forest model.

[0046] More specifically, the random forest model is a random forest classification model constructed by RandomForestClassifier, and its parameter settings are: criterion='entropy', max_depth=5, n_estimators=10, oob_score=True, random_state=124. The malignant prediction probability threshold is 0.7. For each sample, based on the DNA methylation level of its biomarker combination, the constructed random forest regression model is used to calculate the malignant prediction probability. If the malignant prediction probability is greater than the threshold, it is judged as a malignant nodule (i.e., lung cancer); otherwise, it is judged as a benign nodule.

[0047] The present invention will be further explained and illustrated by the following examples, but it does not constitute a limitation to the present invention. It should be understood that these examples are only used to illustrate the present invention and not to limit the scope of the present invention.

[0048] Example

[0049] 1. Test method:

[0050] 1.1 Research population

[0051] A total of 81 patients who underwent surgery at Shandong Provincial Hospital from January 2020 to June 2022 were selected, including 48 patients with lung cancer and 33 patients with benign pulmonary nodules. This study was conducted in accordance with the Helsinki Declaration. The patients included in this study were 18 years old and above, with positive pulmonary nodules (diameter < 3 cm) indicated by CT scan and subsequently underwent surgical resection. The enrolled patients denied the existence of other malignancies and had no previous history of cancer. None of the patients received any preoperative cancer treatment. The clinical information of the patients was collected, including the patient's age, gender, smoking history, nodule size, nodule location, pathological type and stage of the nodule. The pathological type was determined for all surgically resected tissue sections according to the 2015 WHO histological classification of lung cancer. All patients with pathologically confirmed malignancies were staged according to the eighth edition of the TNM guideline classification criteria. This study has obtained the approval of the Ethics Committee of Shandong Provincial Hospital Affiliated to Shandong First Medical University, and all participants have signed the informed consent form.

[0052] 1.2 Sample collection and tissue DNA isolation

[0053] Formalin-fixed paraffin-embedded (FFPE) tissue samples were sectioned from surgical resection. Tissue genomic DNA (GDNA) was extracted from FFPE tissue specimens using the Qiagen QIAamp DNA FFPE Tissue Kit (Qiagen, catalog number 56404). The gDNA was fragmented into 200 bp using an M220 Focused-ultrasonicator (Covaris, Inc.). According to the manufacturer's protocol, 50 ng of the gDNA fragments were used for bisulfite conversion.

[0054] 1.3 Bisulfite conversion and DNA methylation sequencing

[0055] Bisulfite conversion was performed using the EZ DNA Methylation-Gold TM Kit (Zymo Research, catalog number D5005) according to the manufacturer's protocol. Briefly, 130 μL of CT Conversion Reagent was added to 20 μL of the DNA sample, and the mixture was placed in a thermal cycler and incubated according to the following program: 10 minutes at 98 °C, 2.5 hours at 64 °C, and up to 20 hours at 4 °C. The bisulfite-converted DNA was mixed with 600 μL of M-Binding Buffer, and the mixed sample was added to a Zymo-Spin TM IC Column and centrifuged. 100 μL of M-Wash Buffer, 200 μL of M-Desulphonation Buffer, 200 μL of M-Wash Buffer, and 15 μL of LOW EDTA Buffer were added for washing and elution, respectively.

[0056] Construct a SWIFT library using the Accel-NGS Methyl-Seq DNA Library Kit (SWIFT, product number SW 30096). Briefly, denature the purified DNA at 95 °C for 2 minutes, and immediately cool it on ice after the incubation is completed. Ligate the adapters using Adaptase mix at 37 °C for 15 minutes and 95 °C for 2 minutes. Extend using Enzyme Y2 under the conditions of 98 °C for 1 minute + 62 °C for 2 minutes + 65 °C for 5 minutes, and then purify with AMPure XP Beads (Beckman Coulter, product number A63881) and elute in a volume of 15 μL. Ligate the adapters using Enzyme B3 at 25 °C for 15 minutes, and then purify with AMPure XP Beads (Beckman Coulter, product number A63881) and elute in a volume of 20 μL. Amplify the library using KAPAHiFi HotStart Uracil ReadyMix. The amplification conditions are: one cycle at 95 °C for 30 seconds and repeated cycles of 98 °C for 10 seconds + 60 °C for 30 seconds + 72 °C for 1 minute. Purify the amplified DNA with AMPure XP Beads and elute in a volume of 22 μL. Perform Qubit concentration measurement and use an Agilent 2100 Bioanalyzer for quality inspection. The pre-library containing more than 500 ng of DNA is considered to meet the subsequent targeted enrichment.

[0057] Next, perform targeted enrichment using the Twist Fast Hybridization and Wash Kit, 96 Reactions (Twist, product number 101175). Briefly, add 4 μL of Twist Methylation Custom Panel, 96 Reactions (Twist, product number 104408), 13 μL of Twist Universal Blockers, 96 Reactions (Twist, product number 100767), and 2 μL of Hybridization Enhancer (Twist, product number 100968) to the DNA pre-library, and heat and dry it at low temperature in a vacuum concentrator. Add 20 μL of Fast Hybridization Mix (Twist, product number 100968) and 30 μL of Hybridization Enhancer to the lyophilized sample from the previous step, and perform hybridization in a PCR instrument at 95 °C for 5 minutes and 60 °C for 15 minutes - 4 hours.

[0058] After hybridization, streptavidin magnetic beads (Twist, catalog number 100984) were used to pull down the DNA pre-library bound to the probe. 100 μL of streptavidin magnetic beads were used, added with 200 μL of Fast Binding Buffer (Twist, catalog number 100972) and washed three times, then resuspended in 200 μL of Fast Binding Buffer. All the hybridization solution was transferred to the equilibrated magnetic beads and thoroughly mixed on a rotator at room temperature for 30 minutes. 200 μL of pre-warmed Fast Wash Buffer I (Twist, catalog number 100972) was added, incubated at 65 °C for 5 minutes, placed on a magnetic stand for 1 minute, and the supernatant was removed, and repeated once. 200 μL of pre-warmed Wash buffer II (Twist, catalog number 100972) was added, incubated at 48 °C for 5 minutes, placed on a magnetic stand for 1 minute, and the supernatant was removed, and repeated twice. Finally, it was eluted in 45 μL of water and the solution was incubated on ice.

[0059] The library was further amplified using KAPA HiFi HotStart Ready Mix and P5 and P7 primers. The amplification program was as follows: 45 seconds at 98 °C for one cycle, 15 seconds at 98 °C + 30 seconds at 60 °C + 30 seconds at 72 °C for 15 cycles and 1 minute at 72 °C for one cycle. The amplified library was purified using DNA Purification Beads (Twist, catalog number 100984) and washed in 80% ethanol solution. Finally, the library was sequenced on the Illumina Nova 6000.

[0060] 1.4 Quality control

[0061] The sequencing data was pre-processed using fastp (v0.21.0) and trimmomatic (v0.39) to obtain clean data. Secondly, bismark (v0.23.0) was used to align the clean data to the hg19 reference genome and deduplicate the alignment results. Then, bismark_methylation_extractor was used to extract the methylation identification results. Finally, bamdst (v1.0.9) was used to perform statistics on the alignment results to obtain basic statistical information such as sequencing coverage.

[0062] 1.5 Differential methylation analysis

[0063] The samples were split into a training set and a test set in a ratio of 2:1. The R package DSS (version 2.42.0) was used to perform differential methylation analysis on the training set, obtaining the differential level diff and the differential significance p_value between the two groups for each CpG site. According to the methylation quantification of each sample between the two groups, the AUC and sensitivity (under the premise of a specificity of 0.9) were calculated. Differentially methylated sites were screened with the following criteria: p_value < 1e-2 & abs(diff) > 0.1 & sensitivity > 0.6.

[0064] 1.6 Biomarker Selection and Model Construction

[0065] Sorted in descending order according to sensitivity, the optimal n biomarkers were selected, and the random forest algorithm (R package randomForest_4.7 - 1.1) was used for binary classification algorithm training (training set) and model evaluation (test set). As the name implies, a random forest is a forest built in a random way, consisting of many decision trees, and there is no association between each decision tree in the random forest. After obtaining the forest, when a new input sample enters, each decision tree in the forest is used to make a judgment to see which class this sample belongs to (for classification algorithms), and then see which class is selected the most, and predict this sample as that class. Twenty biomarkers with the most significant differences were selected for model building, and the ROC method was used to evaluate the performance of this model in the training set and the test set.

[0066] Plasma samples from 48 patients with pulmonary malignant nodules and 33 patients with pulmonary benign nodules were collected and divided into a training set and a test set in a ratio of 2:1. The training set contained 32 pulmonary malignant nodule samples and 22 pulmonary benign nodule samples. The methylation levels of 20 differential intervals of the training set samples were calculated as molecular features, and the random forest machine learning classification algorithm, that is, the RandomForestClassifier function of sklearn.ensemble, was used to construct a random forest model to obtain a discriminant model for pulmonary benign and malignant nodules.

[0067] The specific construction method is as follows: Random Forest uses N (N = 10 in this embodiment) decision trees as the basic learners to form an ensemble learner. Randomness is introduced during the decision tree training process, including obtaining the samples for each decision tree by bootstrap sampling (i.e., the sampling process for the samples used by each decision tree belongs to random sampling with replacement), and randomly selecting the classification attributes for each decision tree (i.e., the methylation sites used for random forest decision classification are randomly selected from 20 differential intervals). The input training data is the label to which each sample belongs (label, i.e., positive or negative sample) and the methylation values of 20 differential intervals. The random forest algorithm is used to evaluate the importance degree of each methylation feature and the importance degree of each sample for model construction. The classification voting results of the N decision trees in the random forest are presented in the form of percentages, that is, the results of each sample for each label are represented by percentage probabilities.

[0068] When applying this model, the predict function in the Python language is used to discriminate the test set samples, and it can be determined whether the test samples are classified as malignant or benign. The predict_proba function can calculate the probability that the test sample is classified as a malignant nodule, which is specifically determined by the classification voting of the N decision trees in the random forest. The calculation formula is as follows:

[0069] Probability that the sample is classified as a malignant nodule: P1 = N1 / N;

[0070] Where N1 represents the number of decision trees that predict the sample as a malignant nodule, and N is the number of decision trees in the random forest.

[0071] When the probability P1 that the test sample is classified as a malignant nodule is ≥ 0.7, the test sample is judged as a positive sample (i.e., a malignant nodule); when the probability P1 that the test sample is classified as a malignant nodule is < 0.7, the test sample is judged as a negative sample (i.e., a benign nodule).

[0072] 1.7 Functional and pathway enrichment analysis of differentially methylated genes

[0073] The GO enrichment analysis is performed using the R package clusterProfiler (v4.2.2), and the KEGG pathway enrichment analysis is performed using WebGestalt (https: / / www.webgestalt.org / ).

[0074] 2 Experimental results

[0075] 2.1 Clinical cohort

[0076] From January 10, 2020 to June 27, 2022, all patients underwent pulmonary nodule resection surgery. A total of 81 patients' tissue samples (33 benign and 48 malignant) were used for DNA methylation analysis, model development and validation. The demographic and clinical characteristics of the patients are shown in Table 1. Specifically, stage I and stage II cancers accounted for 56.3% and 6.2% of the total cancer patients, respectively. The average size of benign nodules was 0.96 cm (0.50 - 1.65 cm), and the average size of malignant nodules was 0.94 cm (0.50 - 1.75 cm).

[0077] Table 1 Demographic and clinical characteristics of the patients

[0078]

[0079]

[0080] 2.2 Quality control of methylation sequencing data

[0081] The original sequencing data of the tissue samples of 81 patients with pulmonary nodules were preprocessed using fastp and trimmomatic software ( Figure 1 ). Figure A performs data quality control from the perspective of the proportion of valid data, and evaluates the data quality from the perspectives of raw_q20_rate (the proportion of bases with Q value greater than 20 in the original data), mapped_bases_rate (the proportion of successfully aligned bases), target_bases_rate (the proportion of bases falling in the panel target region), dedup_mapped_bases_rate (the proportion of successfully aligned bases after deduplication), and dedup_target_bases_rate (the proportion of bases falling in the panel target region after deduplication). Figure B performs data quality control from the perspective of sequencing coverage depth, and evaluates the data quality from the perspectives of target_average_depth (the average coverage depth of bases in the panel target region) and dedup_target_average_depth (the average coverage depth of bases in the panel target region after deduplication). The overall results show that the average values of raw_q20_rate in benign nodules and malignant nodules are 0.955 and 0.959 (>0.95), respectively, indicating that the tissue samples have good quality and a small probability of error, which increases the detection of false-positive variations and makes the diagnostic conclusion more accurate. mapped_bases_rate and target_bases_rate indicate good sequencing quality and higher capture efficiency, which can reduce the sequencing cost.

[0082] 2.3 Identification of differentially methylated regions

[0083] Eighty-one tissue samples were divided into a training set (22 benign and 32 malignant) and a test set (11 benign and 16 malignant) at a ratio of 2:1, such that the distribution of malignant tumors, age, and gender in the test set was balanced with that of the training set, as shown in Table 2. In both the training set and the test set, the percentages of benign and malignant nodules were 40.7% and 59.3%, respectively, and there were no significant differences in terms of gender and age (P>0.05).

[0084] Table 2 Comparison of clinical information and differences between the training set and the test set

[0085]

[0086] The differential levels diff and differential significance p_value between the two groups of 984,796 CpG sites were calculated using the DSS method. After screening based on the criteria: p_value<1e-2 & abs(diff)>0.1 & sensitivity>0.6, a total of 885 differentially methylated sites were finally obtained. The 885 differentially methylated sites that met the screening criteria were sorted in descending order of sensitivity, and the performance of the model under different numbers of markers was evaluated to select the optimal markers ( Figure 2 ). The top 10 markers in the model showed an AUC value of 1 in the training set and an AUC value of 0.97 in the test set. The receiver operating characteristic curve-AUC (ROC-AUC) containing 60 methylated target regions had reached 1 in the test set, indicating good model performance. According to the heatmap of methylation levels ( Figure 3 ), the methylation levels of these 100 markers were significantly different between the two groups. The methylation levels of most of the markers in the upper half were higher in malignant nodules than in benign nodules, and the methylation levels of most of the markers in the lower half were lower in malignant nodules than in benign nodules, which could clearly distinguish between benign and malignant nodules.

[0087] 2.4 Screening of diagnostic markers and construction of diagnostic models for benign and malignant nodules

[0088] Finally, the 20 most significantly different markers (Top20, Figure 4 A) were selected for model construction, and the specific information is shown in Table 3. To test the diagnostic ability of the model under the Top20 markers in distinguishing between lung malignant nodules and benign lung nodules, model performance evaluations were carried out in both the training set and the test set. The specificity of the model in the training set was 1 (0.846 - 1), and the sensitivity was 1 (0.891 - 1). In the test set, the specificity was 0.91 (0.587 - 0.998), and the sensitivity was 0.933 (0.794 - 1) ( Figure 4C and D). Its receiver operating characteristic curve - AUC (ROC - AUC) was 1.000 in the training set and reached 0.99 in the test set (Figures B and C), indicating that the model based on the Top20 markers had a great advantage in differentiating benign and malignant nodules, further demonstrating the good performance of the designed model.

[0089] Table 3 Top20 markers for the diagnosis of pulmonary benign and malignant nodules

[0090]

[0091]

[0092]

[0093] 2.5 Functional and pathway enrichment analysis of differentially methylated genes

[0094] To explore the biological pathways that differentially methylated genes might affect, gene enrichment analysis was performed on the Top100 markers. GO enrichment analysis showed that multiple cellular components ( Figure 5 A and B), molecular functions ( Figure 5 C and D), and biological processes ( Figure 5 E and F) were significantly enriched. In the cellular component category, focal adhesion, cell - substratum junction, cell front, etc. were enriched. In molecular functions, DNA - binding transcription factor, extracellular matrix structural constituent, and catalytic activity acting on glycoprotein were enriched. In the biological process category, embryonic epithelial morphogenesis, myotube differentiation, tubule formation, etc. were enriched. KEGG pathway analysis showed ( Figure 6 ) that multiple pathways might be involved: cancer - related pathways, signal transduction pathways, and infectious pathways. These pathways might be involved in the occurrence process of lung cancer.

[0095] In the construction of the prediction model of the present invention, the random forest algorithm was used to process the expression of differentially methylated genes, and the Top20 genes were selected as candidate markers for the diagnosis of benign and malignant pulmonary nodules, and the results showed superior diagnostic performance.

[0096] In this study, the methylation data of benign and malignant pulmonary nodules were systematically analyzed, and malignant tumors mainly focused on early lung cancer. In the screening of differentially methylated sites, the analysis was carried out based on the sensitivity and specificity of each gene, and then the optimal markers were selected. The sensitivity of a single gene may not be significant, but several genes with the highest sensitivity were selected for combination on the premise of relatively high specificity. The diagnostic model for benign and malignant pulmonary nodules had extremely high sensitivity and specificity, which greatly improved the predictive ability of the model. Among 885 differentially methylated sites meeting the screening criteria, the Top20 genes were selected as candidate genes. The sensitivity and specificity of the diagnostic model for benign and malignant pulmonary nodules constructed by the Top20 markers were 100% and 91% respectively. This method can be used as a supplementary means for LDCT screening and has important value for the early screening of lung cancer patients.

[0097] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. Use of a reagent for detecting the methylation level of a methylation biomarker in the preparation of a product for differentiating between benign and malignant lung nodules; The methylation biomarker is a group consisting of all of the following gene fragments: chr5:93905357-93905359, chr5:134259480-134259482, chr5:134259489-134259491, chr5:93905256-93905258, chr11:67247741-67247743, chr1:243645964-243645966, chr5:134259522-134259524, chr15:45928449-45928451, chr5:154026439-154026441, chr5:154026479-154026481, chr8:42398664-42398666, chr1:47899804-47899806, chr2:85362588-85362590, chr2:130875117-130875119, chr11:71189489-71189491, chr11:76887102-76887104, chr5:134259500-134259502, chr5:134259527-134259529, chr11:18808624-18808626 and chr15:25093827-25093829; The reference genome version is the hg19 reference genome.

2. The application according to claim 1, characterized in that, The reagent for detecting the methylation level of the methylation biomarker includes reagents required for bisulfite conversion and DNA methylation sequencing.

3. The application according to claim 1, characterized in that The product for differentiating between benign and malignant lung nodules is a detection kit.

4. A system for differentiating benign and malignant pulmonary nodules, characterized in that, The system includes: i) An analysis module, the analysis module comprising: a detection reagent for determining the methylation biomarker in a test sample of a subject; The methylation biomarker is a group consisting of all of the following gene fragments: chr5:93905357-93905359, chr5:134259480-134259482, chr5:134259489-134259491, chr5:93905256-93905258, chr11:67247741-67247743, chr1:243645964-243645966, chr5:134259522-134259524, chr15:45928449-45928451, chr5:154026439-154026441, chr5:154026479-154026481, chr8:42398664-42398666, chr1:47899804-47899806, chr2:85362588-85362590, chr2:130875117-130875119, chr11:71189489-71189491, chr11:76887102-76887104, chr5:134259500-134259502, chr5:134259527-134259529, chr11:18808624-18808626 and chr15:25093827-25093829; The reference genome version is the hg19 reference genome; ii) An evaluation module, the evaluation module includes: discriminating and judging the benign and malignant nature of the sample to be tested according to the methylation expression level of the methylation biomarker determined in i).

5. The system according to claim 4, wherein In the analysis module of i), the detection reagents include the reagents required for bisulfite conversion and DNA methylation sequencing.

6. The system according to claim 4, wherein In the analysis module of i), the sample to be tested is a lung nodule tissue sample.

Citation Information

Patent Citations

  • Methylation markers for identifying and diagnosing benign and malignant pulmonary nodules as well as screening method and application of methylation markers

    CN115976216A

  • Methylation markers for detection of benign / malignant pulmonary nodules or combination thereof, and application thereof

    WO2022161076A1