A training method for lung cancer prediction model, prediction device and application
By detecting the methylation levels of 127 markers in the blood and using a machine learning model to build a lung cancer prediction model, the problems of difficulty in early diagnosis, high false positive rate in imaging, and inaccurate screening of high-risk populations were solved, achieving high-sensitivity and high-specificity lung cancer detection.
Patent Information
- Application Number
- CN202211552486.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-12-05
AI Technical Summary
Existing lung cancer detection technologies have problems such as difficulty in early diagnosis, high false positive rate of imaging examinations, lack of highly accurate markers and screening standards for high-risk populations, resulting in poor lung cancer screening results.
By detecting the methylation levels of 127 markers in the blood, a lung cancer prediction model is constructed using a machine learning model, and combined with the methylation patterns of specific CpG sites, early diagnosis and risk assessment of lung cancer can be achieved.
It improves the sensitivity and specificity of lung cancer detection, reduces the trauma of imaging examinations, reduces the false positive rate, and provides a more accurate screening method for high-risk groups.
Smart Images

Figure CN115976209B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biological detection technology, and in particular to a training method for a lung cancer prediction model, a prediction device, and an application thereof. Background Art
[0002] Lung cancer has the highest morbidity and mortality rates worldwide. Clinical staging at diagnosis is crucial for the five-year survival rate of lung cancer patients, reaching 92% for early-stage lung cancer and only 5.8% for late-stage disease. Therefore, early diagnosis is crucial for improving the prognosis of lung cancer patients. However, lung cancer screening abroad is primarily based on chest imaging and molecular markers, which are less suitable for the Chinese population.
[0003] The difficulties in early diagnosis and treatment of lung cancer include the following aspects: First, early lung cancer has no characteristic imaging manifestations, and there is a lack of new imaging analysis technology for differential diagnosis of early lung cancer; second, patients with early lung cancer often have no characteristic clinical symptoms, and there is a lack of screening criteria and appropriate screening programs for high-risk groups; third, there is a lack of highly accurate markers for early diagnosis and treatment of lung cancer. Traditional molecules such as CEA have a sensitivity of less than 50% for diagnosing early lung cancer, and there is a lack of precise indicators to guide early diagnosis and treatment in clinical practice; fourth, small lung nodules are easily missed and their nature is difficult to determine. Manual labeling has a time-consuming and labor-intensive bottleneck, and artificial intelligence faces the problem of imbalance between "small data and data groups."
[0004] Currently, common lung cancer diagnostics primarily include enzyme-based assays. Plasma tumor marker testing is a common clinical test used for lung cancer screening and postoperative monitoring. Carcinoembryonic antigen (CEA) is a broad-spectrum tumor marker, with studies demonstrating a sensitivity and specificity of 69% and 68%, respectively, for diagnosing lung cancer. Other commonly used clinical markers for lung cancer include cytokeratin 19 fragment antigen (CYFRA21-1) and neuron-specific enolase (NSE), which are more effective in diagnosing squamous cell carcinoma and small cell carcinoma, respectively. However, due to tumor heterogeneity, tumor markers with sufficient specificity and sensitivity for diagnosing early-stage lung cancer have yet to be discovered. For example, alkaline phosphatase can be significantly elevated in patients with liver cancer and osteosarcoma. Glycoproteins, such as serum alpha-acid glycoprotein (α-AG) in lung cancer and CA19-9 in digestive system tumors, can be elevated. Tumor-associated antigens, such as carcinoembryonic antigen (CEA), can be elevated in gastrointestinal tumors, lung cancer, and breast cancer, and alpha-fetoprotein (AFP) can be elevated in liver cancer and malignant teratomas. At present, tumor markers lack specificity and are only of certain value in assisting diagnosis and judging prognosis.
[0005] Since the 1990s, with the development of low-dose computed tomography (LDCT) technology, lung cancer screening has entered the LDCT era. Clinical studies have shown that LDCT screening of high-risk individuals can reduce lung cancer mortality by 20% compared with chest X-rays. Lung cancer screening is effective in detecting stage I lung cancer and non-small cell lung cancer. However, while LDCT screening can detect malignant nodules, it also detects a large number of benign and indeterminate nodules, resulting in a high false-positive rate. Many false-positive nodules require further invasive testing, increasing patient anxiety and, in a small number of cases, developing complications from invasive testing. Overdiagnosis through LDCT can lead to false-positive results. While the existence of overdiagnosis in CT screening programs for lung cancer remains unclear, studies suggest that approximately 10%-12% of cancer cases identified through lung cancer screening are overdiagnosed.
[0006] In view of this, the present invention is proposed. Summary of the Invention
[0007] The purpose of the present invention is to provide a training method for a lung cancer prediction model, a prediction device and an application.
[0008] The present invention is achieved in that:
[0009] In a first aspect, embodiments of the present invention provide the use of a reagent for detecting marker methylation levels in the preparation of a product for predicting lung cancer, wherein the markers include at least 50 of markers 1 to 127; wherein the marker corresponding to each item in the following table includes a corresponding CpG site and / or a region covering the corresponding CpG site:
[0010] Table 1 Markers
[0011]
[0012]
[0013]
[0014]
[0015]
[0016] The hg19 reference genome sequence was used as the benchmark.
[0017] In a second aspect, an embodiment of the present invention further provides a kit for diagnosing or assisting in the diagnosis of lung cancer, which comprises the reagent for detecting the methylation level of the marker described in the aforementioned embodiment.
[0018] In a third aspect, an embodiment of the present invention provides a training method for a lung cancer prediction model, comprising: obtaining marker methylation results and labeling results of a training sample; wherein the marker is as described in the aforementioned embodiment, and the marker result is a label representing at least one of the risk of disease, disease progression, and prognostic risk of lung cancer in the sample; inputting the methylation results of the marker of the training sample into a pre-constructed prediction model to obtain a prediction result; the pre-constructed prediction model is a machine learning model that can predict at least one of the risk of disease, disease progression, and prognostic risk of lung cancer based on the methylation level of the marker; and updating the parameters of the pre-constructed prediction model based on the labeling results and the prediction results.
[0019] In a fourth aspect, embodiments of the present invention provide a lung cancer prediction device comprising an acquisition module and a prediction module. The acquisition module is configured to acquire the methylation level of a marker in a sample to be tested, where the marker is as described in the preceding embodiments; and the prediction module is configured to input the acquired marker methylation level into a prediction model trained using the training method described in the preceding embodiments to obtain a prediction result.
[0020] In a fifth aspect, an embodiment of the present invention provides an electronic device, comprising a processor and a memory, wherein the memory is used to store a program. When the program is executed by the processor, the processor implements the training method or lung cancer prediction method as described in the aforementioned embodiment. The steps of the prediction method include: obtaining the methylation level of a marker of a sample to be tested, the marker being as described in the aforementioned embodiment, inputting the obtained methylation level of the marker into a prediction model trained by the training method as described in the aforementioned embodiment to obtain a prediction result.
[0021] In a sixth aspect, an embodiment of the present invention provides a computer-readable medium having a computer program stored thereon, which, when processed and executed, implements the training method described in the aforementioned embodiment or the prediction method described in the aforementioned embodiment.
[0022] The present invention has the following beneficial effects:
[0023] (1) The present invention discovered a new lung cancer marker that has better sensitivity and specificity than traditional clinical detection methods and existing markers;
[0024] (2) Compared with clinical imaging detection methods, it is safer, non-invasive, and not affected by the physical condition of the person being tested;
[0025] (3) The present invention only requires the collection of a small amount of blood to complete, while imaging examinations are affected by the physiological activities of certain organs and cannot be performed on patients with certain special physical conditions, and certain radioactive substances may cause certain damage to the body. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 A technical roadmap for the data processing steps of the present invention;
[0028] Figure 2 The distribution differences of methylation abundance of 127 markers in different samples;
[0029] Figure 3 The prediction results of the prediction model constructed for 127 markers in samples of different malignancy levels;
[0030] Figure 4 ROC curve plot of the prediction model constructed for 127 markers;
[0031] Figure 5 ROC curve plot of the prediction model constructed for 50 markers. DETAILED DESCRIPTION
[0032] To make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer are used. Where the manufacturer of the reagents or instruments is not specified, all are conventional products that can be purchased commercially.
[0033] With the development of high-throughput sequencing technology and its expanding applications, liquid biopsies, particularly the detection of circulating tumor DNA (ctDNA), have become one of the most promising non-invasive diagnostic methods in precision oncology. Recent studies have shown that epigenetic modifications often occur in the early stages of tumor development. Whole-genome sequencing of cell-free DNA can extract more widespread epigenetic modification variants to improve diagnostic sensitivity and specificity. The most common of these is the detection of DNA methylation.
[0034] Existing methylation signal detection technologies are often limited to single sites or roughly measure the difference in methylation levels in a certain region between cancer and non-cancer groups, and cannot accurately locate epigenetic differences. CpG methylation in mammals is a relatively stable epigenetic modification that can be inherited through the action of certain enzymes during cell division. Because the local activity of these methylation-related enzymes is consistent, adjacent CpG sites on the same DNA molecule have similar methylation states. Therefore, CpG co-methylation analysis can be performed using linkage disequilibrium theoretical models established to simulate adjacent genetic variation on human chromosomes.
[0035] Based on this, the present invention uses unique detection and data analysis methods to quantify the methylation distribution pattern unique to lung cancer, making it a new tumor marker and applying it to the detection of lung cancer. The inventors of this application have discovered that by quantifying the differences in methylation levels in certain specific CpG regions, using machine learning methods to screen for CpG regions with different methylation levels between tumor cells and normal cells, and accurately locating the methylation-linked haplotype signals specific to lung cancer, a classification model with high accuracy can be constructed, thereby improving the accuracy of lung cancer detection.
[0036] The specific solutions provided in this application are as follows.
[0037] An embodiment of the present invention provides the use of a reagent for detecting marker methylation levels in the preparation of a product for predicting lung cancer, wherein the marker includes: the marker includes at least 50 of markers 1 to 127; wherein the marker corresponding to each item in Table 1 includes a corresponding CpG site and / or a region containing a corresponding CpG site.
[0038] The “region containing corresponding CpG sites” herein can be specifically understood as: the region between the two CpG sites with the longest distance between them on the genome among all the CpG sites corresponding to each marker.
[0039] In some embodiments, the markers include markers 1-50.
[0040] In some embodiments, the markers include markers 1-127.
[0041] In some embodiments, the lung cancer includes early stage lung cancer, mid stage lung cancer and late stage lung cancer.
[0042] In some embodiments, predicting lung cancer includes predicting at least one of the risk of lung cancer, disease progression, and prognostic risk.
[0043] In some embodiments, the reagents for detecting the methylation level of the marker include at least one of a methylation sequencing reagent, a methylation-specific PCR reagent, a methylation-sensitive single nucleotide primer extension reagent, a methylation-sensitive single-stranded conformation analysis reagent, and a methylation-sensitive denaturing gradient gel electrophoresis reagent. Optionally, the methylation sequencing reagent includes a bisulfite reagent, a sequencing library construction reagent, and a PCR amplification reagent. The reagents for detecting the methylation level of the marker can be obtained based on the above-mentioned marker design by combining conventional technical knowledge. The invention of this application is to propose a new marker for predicting lung cancer, rather than the detection method itself. The detection method for CpG site methylation can be obtained based on conventional technical knowledge in this field and will not be repeated here.
[0044] In some embodiments, the product comprises at least one of a reagent, a kit, and a predictive model.
[0045] An embodiment of the present invention further provides a kit for diagnosing or assisting in the diagnosis of lung cancer, which comprises the reagent for detecting the methylation level of the marker described in any of the aforementioned embodiments.
[0046] An embodiment of the present invention further provides a method for training a lung cancer prediction model, which includes:
[0047] Obtaining marker methylation results and labeling results for the training sample; wherein the marker is as described in any of the aforementioned embodiments, and the labeling result is a label representing at least one of the risk of lung cancer, disease progression, and prognostic risk of the sample;
[0048] Inputting the methylation results of the markers of the training samples into a pre-established prediction model to obtain a prediction result; the pre-established prediction model is a machine learning model that can predict at least one of the risk of lung cancer, disease progression, and prognosis risk based on the methylation levels of the markers;
[0049] Parameters of a pre-built prediction model are updated based on the annotation results and the prediction results.
[0050] In some embodiments, the tag may be a character or a string of characters.
[0051] In some embodiments, the prediction model includes any one of a random forest model, a support vector machine model, a gradient boosting model, and a logistic regression model. It is understood that, when the features or indicators for constructing the model are disclosed, each prediction model includes a variety of parameters (including general and adjustable parameters) that can be routinely adjusted and selected based on conventional technical knowledge in the field.
[0052] In some embodiments, when the prediction model is a random forest model, the formula of the random forest model is as follows:
[0053]
[0054] Where B represents the number of trees in the random forest, b represents the index of the tree, and f b represents the decision tree with index b, x' represents the methylation level input value of the marker of the sample to be tested (which can be 1 or 0), Represents the final predicted value of the random forest model for the test sample.
[0055] When the prediction model is a random forest model, the training parameters are as follows: the number of decision trees, n_estimators , is ≥ 50, and can be any one of 50, 100, 200, 300, 400, 500, and 600, or a range between any two. When generating a single decision tree, the number of features, max_features , is "log2 (logarithm of the total number of features)" or "sqrt." The tree depth, max_depth , is 1 to 10.
[0056] In some embodiments, when the prediction model is a random forest model, the setting parameters of the trained model include the following: the number of decision trees n_estimators is 500; the number of features max_features when generating a single decision tree is "log2"; and the depth of the tree max_depth is 3.
[0057] It should be noted that the markers of the present application all correspond to more than three CpG sites. When marking or calculating the "methylation result" or "methylation degree" of each marker, the methylation results of all CpG sites corresponding to the marker are considered. When all CpG sites corresponding to the marker are methylated, the marker is recorded as methylated; otherwise, it is recorded as unmethylated.
[0058] An embodiment of the present invention further provides a lung cancer prediction device, comprising:
[0059] an acquisition module, configured to acquire the methylation level of a marker of a sample to be tested, wherein the marker is as described in any of the aforementioned embodiments;
[0060] The prediction module is used to input the obtained methylation level of the marker into the prediction model trained by the training method described in any of the above embodiments to obtain a prediction result.
[0061] Optionally, the above modules can be stored in a memory in the form of software or firmware or fixed in the operating system (OS) of the electronic device provided in this application, and can be executed by a processor in the electronic device. At the same time, the data, program code, etc. required to execute the above modules can be stored in the memory.
[0062] In some embodiments, the training samples and the test samples can be independently blood samples or environmental samples containing blood samples. The blood samples can be whole blood samples or plasma samples.
[0063] An embodiment of the present invention also provides an electronic device, comprising a processor and a memory, wherein the memory is used to store a program. When the program is executed by the processor, the processor implements the training method or lung cancer prediction method as described in any of the foregoing embodiments. The steps of the prediction method include: obtaining the methylation level of a marker of a sample to be tested, the marker being as described in any of the foregoing embodiments, inputting the obtained methylation level of the marker into a prediction model trained by the training method as described in any of the foregoing embodiments, and obtaining a prediction result.
[0064] The electronic device may include a memory, a processor, a bus, and a communication interface. The memory, processor, and communication interface may be electrically connected to each other directly or indirectly to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more buses or signal lines. The processor may process information and / or data related to target recognition to perform one or more functions described in this application.
[0065] The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0066] A processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including a central processing unit (CPU) or a network processor (NP). It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0067] In actual applications, the electronic device can be a server, a cloud platform, a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device, a virtual reality device, etc. Therefore, the embodiments of the present application do not limit the type of electronic device.
[0068] In addition, an embodiment of the present invention further provides a computer-readable medium having a computer program stored thereon, wherein when the computer program is processed and executed, the training method as described in any of the foregoing embodiments or the prediction method as described in any of the foregoing embodiments is implemented.
[0069] The "computer-readable medium" in this article includes various media that can store program codes, such as USB flash drives, mobile hard drives, read-only memories, random access memories, magnetic disks or optical disks.
[0070] The features and performance of the present invention are further described in detail below with reference to the embodiments.
[0071] Example 1
[0072] A method for screening lung cancer markers and constructing a lung cancer prediction model specifically includes the following steps.
[0073] (1) Lung cancer patients, benign nodules, and healthy subjects were identified and cell-free DNA (cfDNA) was extracted from the plasma of these patients and healthy subjects and sequenced using bisulfite conversion sequencing.
[0074] Bisulfite conversion sequencing includes: methylating plasma free DNA (cfDNA) to convert unmethylated cytosine into thymine to obtain a methylated sample; constructing a sequencing library based on the methylated sample, and sequencing to obtain sequencing data.
[0075] The specific steps for methylation library construction and library sequencing include:
[0076] 1. 5-30 ng of cfDNA was used for the methylation library. 50 pg of internal reference DNA (166 bp) was spiked into the library. The library was then treated with the Enzymatics (USA) 5X ER / A-Tailing Enzyme Mix and WGS Ligase to ensure that the library was sequenceable on the Illumina NovaSeq 6000 sequencer.
[0077] 2. After ligation, the sample was treated with bisulfite and purified using the Lightning conversion reagent kit from Zymo Research.
[0078] 3. Amplify the recovered DNA using KAPA HiFi HS Uracil + ready Mix (KAPA);
[0079] 4. Purify using AMPure XP beads (Beckman) and elute the library using EB buffer (Qiagen);
[0080] 5. Take 500 ng of pre-library DNA, add blocking reagent (IDT) and probe (Twist) and evaporate to dryness at 60°C, then incubate with hybridization solution (IDT) for 16 hours;
[0081] 6. After hybridization, add streptavidin magnetic beads (IDT) for capture, and then use washing solution (IDT) for washing;
[0082] 7. The washed DNA was amplified using KAPA HiFi Hotstart ready Mix, and the amplified product was purified using AMPure XP beads (Beckman) to obtain the final library.
[0083] The final library was quantified using qPCR (KAPA SYBR Fast Kit, Roche) and then subjected to paired-end 150 bp sequencing on the Illumina NovaSeq 6000 sequencing platform.
[0084] (2) Quality control the sequencing data, align it to the reference genome, and obtain the methylation levels of all CpG sites on the genome;
[0085] The specific steps for quality control and alignment of sequencing data include:
[0086] 1. Use cutadapt software to filter the sequencing data, including filtering the sequencing adapter sequences, removing DNA fragments with sequencing read lengths less than 50bp, and removing DNA fragments with low average sequencing quality;
[0087] 2. Use BSMAP to align the filtered data with the lambda reference genome, calculate the ratio of methylated sequences to unmethylated sequences in the lambda internal reference, and perform quality control of the methylation conversion rate of the sequencing library;
[0088] 3. Use Bismark to align the filtered data with the hg19 reference genome (carrying the decoy bait sequence) to obtain the specific location information of each DNA fragment on the genome and the methylation status information of each CpG site;
[0089] 4. Use bamtools software to remove DNA fragments with low alignment quality, unaligned DNA fragments, and DNA fragments with incomplete paired-end reads;
[0090] 5. Sort the filtered DNA fragments according to the alignment position to facilitate subsequent analysis and processing.
[0091] (3) Use the processed sequencing data to quantify the differences in methylation signals of CpG sites across the genome and screen for regions with a high degree of linkage; use machine learning methods to screen for sites with a high weight in distinguishing cancer patients from healthy controls;
[0092] The specific steps for processing genome-wide CpG methylation differential region data include:
[0093] 1. Obtain high-quality CpG sites: Statistically capture the methylation levels of all sites within the region with a CG context, a sequencing depth greater than 10×, and coverage greater than 70% in lung cancer populations. Methylation level is defined as: methylation level = methylated C / (methylated C + unmethylated C);
[0094] 2. Obtain differential CpG sites: For the CpG sites obtained in step 1, only those with a fold difference of ≥1.4 between lung cancer and benign nodules, and a fold difference of ≥1.5 between lung cancer and healthy controls were retained. The fold difference was defined as: fold difference = mean methylation level of the CpG site in the positive population / mean methylation level of the CpG site in the negative population;
[0095] 3. Filtering based on methylation haplotypes: For each DMR region obtained in 3, at least five DNA fragments meeting any of the following conditions must be observed in the lung cancer plasma WGBS data;
[0096] (a) If the number of CpG sites in the DMR region is ≤ 6, at least 3 methylated Cs are required to be observed in the DNA fragment;
[0097] (b) If the number of CpG sites in the DMR region is greater than 6, at least 4 methylated Cs are required to be observed in the DNA fragment;
[0098] 4. Filtering based on methylation haplotype abundance: To ensure that the selected methylation haplotypes are lung cancer-specific biomarkers, the P-values of methylation haplotype abundance in the WGBS data of lung cancer tissue, colorectal cancer tissue, liver cancer tissue, and healthy tissue must be statistically significantly different (P ≤ 0.05), where the P-value is the result of a one-way analysis of variance test. The methylation haplotype abundance is defined as: methylation haplotype abundance = number of fragments meeting the three conditions / (number of fragments meeting the three conditions + number of fragments not meeting the three conditions);
[0099] 5. Use the least absolute shrinkage and selection operator (LASSO) regression algorithm, a machine learning method, to further reduce the dimensionality of the regions obtained in step 4. Regions with absolute weights greater than or equal to 0.001 are selected as the region combinations for model construction. After randomly selecting training set samples, repeat the above steps 100 times to obtain stable gene regions.
[0100] 6. Use the built-in Random Forest Importance (Random Forest Importance) machine learning method to further reduce the dimensionality of the region obtained in step 4. Based on the impurity, for each tree, the features are ranked by impurity (gini / entropy). The average of the entire forest is then taken, and the top 1000 features ranked from most important to least important are selected as potential candidate features. After randomly selecting training set samples, repeat the above steps 100 times to obtain stable gene regions.
[0101] 7. Take the intersection of the candidate features obtained in steps 5 and 6 as the final panel combination, namely the 127 methylation-linked haplotype region markers marked in Table 1.
[0102] (4) A machine learning classification model was constructed using Random Forest. The ROC curve was plotted and the optimal threshold was selected using Youden's index to evaluate the model performance, thereby reflecting the sensitivity and specificity of the method for tumor detection (see the technical roadmap for details). Figure 1 ).
[0103] The specific steps to build a machine learning classification model include:
[0104] Training set: 80 lung cancer patients and 92 healthy controls;
[0105] Validation set: 760 subjects, including 366 lung cancer patients (250 stage I lung cancer, 19 stage II lung cancer, 29 stage III lung cancer, 19 stage IV lung cancer, and 49 lung cancer with unknown stage information), 53 subjects with benign lung nodules, and 341 healthy subjects;
[0106] 1. Feature data extraction: Extract the methylation signal intensity of the markers screened by the above method from the sequencing data of each sample as input data. Specifically, for each marker (methylation haplotype), if all CpG sites on the same sequencing read (each marker or its corresponding CpG region) show methylation signals, it is marked as 1, otherwise it is marked as 0. Therefore, for each sample to be tested, a vector of length 127 is generated;
[0107] 2. Determining the optimal model parameters: Random Forest was used for model construction and iterative training. The training set samples were subjected to 10-fold cross-validation. Given a parameter space, the optimal parameter combination was searched. Through iterative training, the parameters that achieved optimal model performance were determined and recorded. The optimal sensitivity and specificity thresholds were found in the validation set samples.
[0108] 3. Model Performance Verification: The determined model's optimal parameters and thresholds were validated in an independent test set. The ROC curve was plotted and the AUC value was calculated. The performance on the final test set represented the overall performance of the model. The Youden's index was used to select the optimal threshold and evaluate model performance, thereby reflecting the sensitivity and specificity of the method for tumor detection.
[0109] The formula of the prediction model is as follows:
[0110]
[0111] Where B represents the number of trees in the random forest, b represents the index of the tree, and f brepresents the decision tree with index b, x' represents the methylation level input value of the marker of the sample to be tested (methylation is recorded as 1, otherwise it is recorded as 0), Represents the final predicted value of the random forest model for the test sample.
[0112] The optimal model parameters are: clf__oob_score (out-of-bag data): True; clf__bootstrap: True; clf__criterion: gini; clf__max_features: log2; clf__n_estimators (number of trees in the forest): 500; clf__criterion: gini. max_depth is 3.
[0113] Example 2
[0114] The 127 methylation abundance distribution differences in Table 1 were verified in 47 plasma samples of stage I lung cancer, 52 plasma samples of benign nodules, and 33 plasma samples of healthy individuals. It can be seen that lung cancer carries a higher intensity methylation signal, see Figure 2 .
[0115] Example 3
[0116] The malignancy degree predicted by the prediction model (Example 1) is distributed differently in 25 plasma samples of lung adenocarcinoma and 22 plasma samples of lung squamous cell carcinoma. It can be seen that as the malignancy degree of the tumor increases, the probability of malignancy predicted by the methylation model also increases, which is consistent with the pathogenesis of lung cancer. Figure 3 .
[0117] Example 4
[0118] A methylation prediction model was constructed using 80 lung cancer patients and 92 healthy controls. A lung cancer prediction model (Example 1) was constructed using all 127 methylation features (as shown in Table 1). 760 subjects were tested, including 366 lung cancer patients (250 patients with stage I lung cancer, 19 patients with stage II lung cancer, 29 patients with stage III lung cancer, 19 patients with stage IV lung cancer, and 49 patients with unknown stage lung cancer), 53 subjects with benign lung nodules, and 341 healthy subjects. The diagnostic performance is shown in Table 2. With a specificity of 92.89%, the detection rate of lung cancer reached 91.53%. Compared with existing technologies and serological markers, the detection results were significantly improved, with an overall AUC of 0.972 (as shown in Table 2). Figure 4 shown).
[0119] Table 2 Prediction results
[0120]
[0121] Example 5
[0122] A methylation prediction model was constructed using 80 lung cancer patients and 92 healthy controls (Example 1). A lung cancer prediction model was constructed using any 50 methylation haplotype region markers in Table 1 (see Table 3) (the construction method was the same as in Example 1, except for the number of markers). 760 subjects were tested, including 366 lung cancer patients (250 stage I lung cancer, 19 stage II lung cancer, 29 stage III lung cancer, 19 stage IV lung cancer, and 49 lung cancer with unknown staging information), 53 subjects with benign lung nodules, and 341 healthy subjects.
[0123] The diagnostic performance is shown in Table 4. With a specificity of 88.58%, the detection rate of lung cancer reached 84.15%, which is significantly improved compared with existing technologies and serological markers. The overall AUC is 0.924 (e.g. Figure 5 shown).
[0124] Table 3.50 lung cancer markers
[0125]
[0126]
[0127]
[0128] Table 4 Prediction results
[0129]
[0130] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. Use of a reagent for detecting marker methylation levels in the preparation of a product for predicting lung cancer, characterized in that: The markers include markers 1 to 127; wherein, each marker corresponding to an item in the following table includes a corresponding CpG site and / or a region containing a corresponding CpG site: The hg19 reference genome sequence was used as the benchmark.
2. The use according to claim 1, characterized in that The lung cancer includes early stage lung cancer, middle stage lung cancer and late stage lung cancer.
3. The use according to claim 1, characterized in that The prediction of lung cancer is the prediction of the risk of lung cancer.
4. The use according to any one of claims 1 to 3, characterized in that The reagents for detecting the methylation level of the marker include at least one of a methylation sequencing reagent, a methylation-specific PCR reagent, a methylation-sensitive single nucleotide primer extension reagent, a methylation-sensitive single-strand conformation analysis reagent, and a methylation-sensitive denaturing gradient gel electrophoresis reagent.
5. The use according to claim 4, characterized in that The methylation sequencing reagents include bisulfite reagents, sequencing library construction reagents and PCR amplification reagents.
6. The use according to claim 1, characterized in that The product includes at least one of a reagent, a kit, and a predictive model.
7. A kit for diagnosing or assisting in the diagnosis of lung cancer, characterized in that: It comprises the reagent for detecting the methylation level of a marker according to any one of claims 1 to 6.
8. A method for training a lung cancer prediction model, characterized in that: It includes: Obtaining marker methylation results and labeling results of the training sample; wherein the marker is as described in any one of claims 1 to 6, and the marker result is a label representing the risk of lung cancer in the sample; Inputting the methylation results of the markers of the training samples into a pre-built prediction model to obtain a prediction result; the pre-built prediction model is a machine learning model that can predict the risk of lung cancer based on the methylation levels of the markers; Parameters of a pre-built prediction model are updated based on the annotation results and the prediction results.
9. The training method according to claim 8, characterized in that The prediction model includes any one of a random forest model, a support vector machine model, a gradient boosting model and a logistic regression model.
10. The training method according to claim 9, characterized in that: When the prediction model is a random forest model, the parameter settings for model training include the following: the number of decision trees n_estimators is ≥ 50; the number of features max_features when generating a single decision tree is "log2" or "sqrt"; and the tree depth max_depth is 1 to 10.
11. The training method according to claim 10, characterized in that: The number of decision trees n_estimators is 50~600.
12. The training method according to claim 10, characterized in that: When the prediction model is a random forest model, the parameters of the trained model include the following: the number of decision trees, n_estimators, is 500; the number of features, max_features, when generating a single decision tree is "log2"; and the tree depth, max_depth, is 3.
13. A lung cancer prediction device, characterized in that: It includes: an acquisition module, configured to obtain the methylation level of a marker of a sample to be tested, wherein the marker is as described in any one of claims 1 to 6; The prediction module is used to input the obtained methylation level of the marker into the prediction model trained by the training method according to any one of claims 8 to 12 to obtain a prediction result.
14. An electronic device, characterized in that: The electronic device includes a processor and a memory, the memory is used to store a program. When the program is executed by the processor, the processor implements the training method or lung cancer prediction method according to any one of claims 8 to 12, wherein the steps of the prediction method include: obtaining the methylation level of a marker of a sample to be tested, wherein the marker is as described in any one of claims 1 to 6, inputting the obtained methylation level of the marker into a prediction model trained by the training method according to any one of claims 8 to 12, and obtaining a prediction result.
15. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is processed and executed, the training method according to any one of claims 8 to 12 or the prediction method according to claim 14 is implemented.
Citation Information
Patent Citations
Detection of lung neoplasia by analysis of methylated DNA
CN109563546A
Methylation gene related to lung cancer and detection kit of gene
CN110317875A
Lung cancer related methylation gene combination in plasma and application of lung cancer related methylation gene combination
CN112195245A
Training method of prediction model for non-small cell lung cancer immunotherapy curative effect and prediction device
CN114999653A