Methylation biomarkers for diagnosing lung cancer and use thereof

CN122168761APending Publication Date: 2026-06-09HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
Filing Date
2026-05-09
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing methods for diagnosing lung cancer, such as imaging examinations and tissue biopsies, are highly invasive and have a high false positive rate. Furthermore, the sensitivity and specificity of diagnosis based on serum tumor markers are difficult to meet clinical needs simultaneously, especially in the early diagnosis of lung cancer.

Method used

A new set of differentially methylated regions (chr1: 25328700-25329000, chr1: 56129400-56129700, chr10: 46957200-46957500, etc.) were developed as methylation biomarkers. Plasma cell-free DNA methylomics analysis was performed using cfMeDIP–seq technology to construct a diagnostic model to distinguish between lung cancer and non-tumor healthy controls, benign pulmonary nodules, and malignant pulmonary nodules.

Benefits of technology

It enables sensitive and specific in vitro diagnosis of lung cancer, significantly improving the diagnostic accuracy of early lung cancer and reducing the false positive rate of imaging examinations and the invasiveness of tissue biopsy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122168761A_ABST
    Figure CN122168761A_ABST
Patent Text Reader

Abstract

The application discloses a methylation biomarker for diagnosing lung cancer and application thereof, and belongs to the technical field of biomedical detection. The methylation biomarker is composed of 16 differential methylation regions (DMRs), and specific differential methylation regions are shown in Table 2. The application also provides products and devices for diagnosing lung cancer based on the 16 differential methylation regions, and the lung cancer diagnosis has the advantages of high sensitivity and strong specificity, and has important application value for clinical lung cancer diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of biomedical detection technology, specifically relating to methylation biomarkers for the diagnosis of lung cancer and their applications. Background Technology

[0002] Lung cancer is one of the leading causes of cancer-related deaths worldwide, and its prognosis is closely related to the clinical stage at diagnosis. Early-stage lung cancer often presents with no obvious specific symptoms, and most patients are diagnosed at an intermediate or advanced stage, resulting in a narrow treatment window and a significantly lower five-year survival rate. Therefore, developing efficient and accurate early diagnostic methods is crucial for improving the clinical outcomes of lung cancer patients.

[0003] Currently, the clinical diagnosis of lung cancer mainly relies on imaging examinations, histopathological biopsies, and serum tumor marker detection. While imaging methods such as low-dose spiral CT can detect lung lesions, they have limitations in differentiating between benign and malignant lesions and suffer from high false-positive rates and radiation exposure. Histopathological biopsy, as the gold standard for lung cancer diagnosis, is highly invasive and complex, making it unsuitable for large-scale screening and dynamic monitoring. In contrast, fluid-based molecular marker detection offers significant advantages such as non-invasiveness, repeatable sampling, and standardized testing procedures. It can reflect the occurrence and development of tumors at the molecular level, providing a highly promising technological approach for early screening and auxiliary diagnosis of lung cancer.

[0004] DNA methylation, as an important epigenetic modification, is considered an early event in tumorigenesis due to its abnormal changes. Therefore, differentially methylated regions (DMRs) have become a promising type of molecular biomarker in liquid biopsy. Although existing studies have identified and reported a large number of differentially methylated regions associated with lung cancer, and developed several diagnostic products and models based on them, the diagnostic sensitivity and specificity of a single or a few regions often fail to simultaneously meet clinical needs due to the complexity of lung cancer pathogenesis and the heterogeneity of population genetic backgrounds. To continuously improve diagnostic efficacy, especially the accurate identification of early-stage lung cancer, it is still necessary to continuously explore and develop new combinations of differentially methylated regions with higher discriminative performance to construct a more robust and accurate diagnostic biomarker system, which is of great significance for the clinical diagnosis of lung cancer. Summary of the Invention

[0005] In view of this, the primary objective of this application is to provide methylation biomarkers for the diagnosis of lung cancer. This application has developed and validated a new set of differentially methylated regions that can serve as methylation biomarkers for the diagnosis of lung cancer, and have the advantages of high sensitivity and high specificity, providing a new approach for the clinical diagnosis of lung cancer.

[0006] To achieve the above objectives, this application adopts the following technical solution: One aspect of this application discloses a methylation biomarker for diagnosing lung cancer, said methylation biomarker consisting of any one or more of the following differentially methylated regions: chr1: 25328700-25329000, chr1: 56129400-56129700, chr10: 46957200-46957500, chr11: 129433800-12943410 0, chr11: 55847100-55847400, chr15: 20546100-20546400, chr15: 20546700-20547000, chr16: 33294300-33294 600, chr18: 106800-107100, chr19: 50624100-50624400, chr2: 119060400-119060700, chr22: 19671600-196719 00, chr5: 17586000-17586300, chr8: 86742600-86742900, chr8: 86760300-86760600, chr9: 67103400-67103700; The reference genome version for the differentially methylated regions above is hg19.

[0007] Another aspect of this application discloses the use of reagents for detecting the methylated biomarkers described above in the preparation of products for diagnosing lung cancer.

[0008] Another aspect of this application discloses a kit for diagnosing lung cancer, the kit comprising reagents for detecting the methylation level of differentially methylated regions in a sample as defined above.

[0009] Another aspect of this application discloses a device for diagnosing lung cancer, comprising a data acquisition module and a data analysis module, wherein: The data acquisition module is configured to: obtain the methylation level of differentially methylated regions in the subject's sample as defined above; The data analysis module is configured to: based on the methylation level of the differentially methylated regions obtained by the data acquisition module, distinguish whether the subject has lung cancer through a constructed diagnostic model.

[0010] This application has at least the following beneficial effects: In this application, a genome-wide analysis of the plasma cell-free DNA methylation profile of lung cancer, healthy individuals without tumors, and patients with benign and malignant pulmonary nodules was performed using cell-free DNA immunoprecipitation high-throughput sequencing (cfMeDIP–seq). This analysis revealed differentially methylated regions closely related to the occurrence of lung cancer and identified abnormal changes in the methylation status of these regions in lung cancer.

[0011] Based on the findings of this application, the methylation status of methylation biomarkers can be analyzed using cfMeDIP–seq technology, enabling sensitive and specific in vitro diagnosis of lung cancer. This application demonstrates, through the detection of plasma cfDNA in lung cancer patients and healthy controls, that these differentially methylated regions can not only distinguish between lung cancer and non-tumor healthy controls, but also differentiate between benign and malignant pulmonary nodules. Therefore, this application provides a combination of differentially methylated regions that can be used for the in vitro diagnosis of lung cancer, and its application as a methylation biomarker for lung cancer diagnosis has significant clinical value. Attached Figure Description

[0012] Figure 1 This is a demonstration of the results of feature selection and model construction in this application, in which, Figure 1 A shows the number of feature genes selected based on LASSO and Boruta algorithms and the intersection results. Figure 1 The curve in Figure B is the ROC curve of a diagnostic lung cancer model constructed from 16 DMRs, used to distinguish between cancer patients and non-cancer controls in the test set.

[0013] Figure 2 For lung cancer model validation and independence analysis, among which, Figure 2 In the middle A, the ROC curve distinguishes between cancer patients and non-cancer controls in the validation set using the lung cancer model; Figure 2 In the B-line diagram, the ROC curve distinguishes between early and late-stage lung cancer patients and non-cancer controls in a lung cancer model. Figure 2 The middle C represents the ROC curve used to distinguish cancer patients from non-cancer controls based on lung cancer model, cfDNA concentration, and clinical characteristics. Figure 2 D represents the distribution of model risk scores among different clinical subgroups; Figure 2 E represents the model risk score distribution for normal controls, benign nodules, malignant nodules, and lung cancer groups; Figure 2 F represents the distribution of model risk scores among the non-tumor control group, early-stage lung cancer group, and late-stage lung cancer group.

[0014] Figure 3 To construct a model for benign and malignant pulmonary nodules, among which, Figure 3 In Figure A, the ROC curve of the multi-class lung cancer model distinguishes between normal controls, benign nodules, malignant nodules, and lung cancer groups in the test set. Figure 3The figure in B represents the ROC curve of the multi-class lung cancer model in the validation set, distinguishing between normal controls, benign nodules, malignant nodules, and lung cancer groups. Detailed Implementation

[0015] The embodiments of this application will be clearly and completely described below. The technical solutions in the embodiments described below are exemplary and only possible technical implementations of this application, not all possible implementations. Those skilled in the art can combine the embodiments of this application to obtain other embodiments without creative effort, and these embodiments are also within the protection scope of this application.

[0016] In this application, the term "lung cancer" refers to a malignant tumor originating from cells in the lungs, including squamous cell carcinoma, adenocarcinoma, and small cell lung cancer. Staging is typically performed using methods such as biopsy staining and / or imaging examinations, with clinical stages including Stage I, II, III, and IV. The term "lung tumor" refers to any abnormal mass of tissue occurring in the lung parenchyma or bronchi; it is a broad concept encompassing both benign and malignant tumors. "Benign lung tumors" refer to slow-growing, non-invasive, and generally non-life-threatening lung growths, such as hamartomas and inflammatory pseudotumors; while "malignant lung tumors" are lung cancer, which has the potential for invasive growth and distant metastasis.

[0017] In this application, the term "pulmonary nodule" refers to a focal, round or oval, high-density pulmonary shadow, typically not exceeding 3 cm in diameter, found on imaging examinations (such as CT scans). Pulmonary nodules can be classified into "benign pulmonary nodules" and "malignant pulmonary nodules" based on their pathological nature. Benign pulmonary nodules are usually caused by non-neoplastic lesions or benign tumors such as infectious granulomas, inflammation, and hamartomas, and their biological behavior is indolent, lacking the ability to invade or metastasize. Malignant pulmonary nodules, on the other hand, often represent early-stage or metastatic lung cancer lesions, exhibiting malignant biological characteristics such as continuous growth, invasion of surrounding tissues, and distant metastasis.

[0018] In this application, the term "diagnosis" includes the detection or identification of a subject's disease state or condition, determining the likelihood that a subject will develop a given disease or condition, determining the likelihood that a subject with a disease or condition will respond to treatment, distinguishing lesions of different natures (e.g., distinguishing between benign and malignant pulmonary nodules), determining the prognosis (or possible progression or regression) of a subject with a disease or condition, and determining the effect of treatment on a subject with a disease or condition. In some specific examples of this application, the diagnosis is intended to distinguish between lung cancer and healthy controls without tumors, or to distinguish between benign and malignant pulmonary nodules.

[0019] In this application, the term "sample" or "sample to be tested" refers to any substance that may contain the target area to be analyzed, including biological samples. Samples can be obtained directly from biological sources or processed samples. Samples include, but are not limited to, bodily fluids (e.g., whole blood, plasma, serum, urine, etc.), bronchoalveolar lavage fluid, and tissue samples. In some preferred embodiments, the sample to be tested is a plasma sample.

[0020] In this application, the term "subject" is appropriately used to refer to mammals. The mammals may be any kind of mammal, but are preferably humans (including non-tumor healthy controls and lung cancer patients).

[0021] In this application, the term "DNA methylation" or "methylation" refers to the process in which a methyl group is transferred to a specific base in vivo, catalyzed by DNA methyltransferases, using S-adenosylmethionine (SAM) as a methyl donor. In mammals, methylation primarily occurs at the C5 position of cytosine residues. In the human genome, a large amount of DNA methylation occurs at the cytosine in CpG dinucleotides, where C is cytosine, G is guanine, and p is a phosphate group.

[0022] In this application, the term "cfDNA (cell-free DNA)" refers to a fragment of free DNA that is not contained within intact cells in the corresponding bodily fluid sample from which a sample is obtained, but rather exists in the bodily fluid sample. In some embodiments of this application, the cfDNA is derived from a plasma sample of a subject.

[0023] In this application, the term "methylation biomarker" refers to an objectively detectable molecular indicator that, based on changes in the DNA methylation status or methylation level of a specific genomic region, can indicate a subject's biological state, pathological process, or response to therapeutic intervention. In this application, the methylation biomarker specifically refers to specific genomic regions (i.e., differentially methylated regions) exhibiting differential methylation patterns among subjects. Detecting the methylation levels of these regions enables in vitro diagnosis of lung cancer.

[0024] In this application, the term "differentially methylated region" (DMR) refers to a DNA region containing one or more differentially methylated sites. Under selected conditions of interest, such as compared to control conditions, a DMR with a significantly increased number or frequency of methylation modification sites can be termed a hypermethylated DMR, and a DMR with a significantly decreased number or frequency of methylation modification sites can be termed a hypomethylated DMR. A DMR serving as a methylation biomarker for lung cancer can be termed a lung cancer DMR.

[0025] In this application, the "methylation level" or "methylation profile" of a nucleic acid molecule refers to the presence or absence of one or more methylated nucleotide bases in the nucleic acid molecule. For example, a nucleic acid molecule containing methylated cytosine is considered methylated. The methylation state can be expressed as a methylation percentage. In some specific examples, the methylation percentage (%) refers to the percentage of methylated cytosine, which can be expressed as: methylation percentage (%) = number of methylated cytosines in the region / (number of methylated cytosines in the region + number of unmethylated cytosines in the region).

[0026] In some specific examples of this application, the methylation level of differentially methylated regions was determined and calculated using a method based on cfMeDIP-seq (cell-free Methylated DNA Immunoprecipitation Sequencing). cfMeDIP-seq is a technique that uses a 5-methylcytosine (5mC)-specific antibody to enrich methylated DNA fragments and then quantifies them using high-throughput sequencing. This technique is particularly suitable for samples with low cfDNA content, such as plasma, and has advantages such as being non-conversion dependent (no bisulfite treatment required), requiring low starting amounts, and having broad coverage.

[0027] In a specific example employing cfMeDIP-seq technology, the calculation of the methylation level may include the following steps: (1) Raw data processing and alignment: The raw reads obtained from sequencing are subjected to quality control and adapter removal. The high-quality reads are aligned to the reference genome (e.g., UCSC hg19 version) using alignment software (e.g., BWA, Bowtie2, etc.) to obtain the genome alignment file (e.g., BAM file) for each sample.

[0028] (2) Fragment count calculation: For each differentially methylated region defined in this application, the genome is cut into 300bp segments using the MEDIP package, and the number of fragments in each region is calculated.

[0029] (3) Normalization and expression of methylation levels: In order to eliminate technical bias caused by differences in total cfDNA amount, immunoprecipitation efficiency and sequencing depth among samples, it is necessary to normalize the above fragment numbers. As a specific example, methylation level can be expressed as "enrichment score" or "relative methylation level". A common calculation method is as follows: for each differentially methylated region, divide the number of fragments in the immunoprecipitation (IP) sample by the corresponding total number of fragments in the input sample, and then take the logarithm (e.g., log2 Fold Change). The resulting value represents the degree of methylation enrichment in that region.

[0030] (4) Alternative calculation methods: In some implementation schemes, an input control may be omitted, and a genome-wide standardization strategy may be adopted instead. For example, transcriptome expression quantification—CPM (Counts Per Million mapped reads)—can be used to correct for library size by counting reads within a region, and the corrected value is the quantitative indicator of the methylation level in that region. Alternatively, a relative methylation score may be used, the calculation method of which is well known to those skilled in the art.

[0031] It should be understood that regardless of the specific normalization algorithm used, as long as it can quantitatively reflect the degree of methylation modification of the differentially methylated regions defined in this application and be used for the construction of diagnostic models and sample discrimination, it falls within the protection scope of this application. In practical applications, those skilled in the art can choose an appropriate calculation method based on specific experimental design and clinical needs.

[0032] In this application, the performance of the diagnostic model can be measured by the area under the receiver operating characteristic (ROC) curve (AUC). AUC is a measure of diagnostic test accuracy; a larger area is better, with an optimal value of 1. Sensitivity and specificity are important indicators for evaluating the performance of diagnostic methods. Sensitivity is the proportion of true positive samples correctly identified, and specificity is the proportion of true negative samples correctly identified.

[0033] Furthermore, the technical solution of this application will be explained in detail.

[0034] The first aspect of this application provides methylation biomarkers for diagnosing lung cancer, said methylation biomarkers consisting of one or more differentially methylated regions. The physical locations of these regions are defined based on the human genome reference sequence GRCh37 / hg19 version. They are shown below in the format [chromosome: start physical location - end physical location]: chr1: 25328700-25329000, chr1: 56129400-56129700, chr10: 46957200-46957500, chr11: 129433800-12943410 0, chr11: 55847100-55847400, chr15: 20546100-20546400, chr15: 20546700-20547000, chr16: 33294300-33294 600, chr18: 106800-107100, chr19: 50624100-50624400, chr2: 119060400-119060700, chr22: 19671600-196719 00, chr5: 17586000-17586300, chr8: 86742600-86742900, chr8: 86760300-86760600, chr9: 67103400-67103700.

[0035] In some preferred examples, the methylation biomarker is a combination of all 16 differentially methylated regions described above.

[0036] The inventors of this application used cfMeDIP-seq technology to perform whole-genome methylation profiling analysis on plasma cfDNA samples from a large number of lung cancer patients and healthy non-tumor controls. Through feature screening, 16 differentially methylated regions were identified. These 16 specific genomic regions showed significant and consistent differences in methylation levels between lung cancer samples and non-tumor control samples (such as samples from healthy individuals or patients with benign pulmonary nodules). By detecting the methylation enrichment level of this specific set of regions, products, devices, or models with extremely high diagnostic efficacy can be constructed, thereby achieving precise diagnosis of lung cancer. The differentially methylated region combination provided in this application can comprehensively reflect the complex epigenetic changes in lung cancer, significantly improving the sensitivity and specificity of lung cancer diagnosis.

[0037] It should be understood that the genomic coordinates of the differentially methylated regions described in this application are based on a specific reference genome version (hg19). With the advancement of genomics research and updates to the reference genome version, the specific coordinate numbers may shift due to sequence corrections or changes in assembly methods. However, it is understood that the core of this application lies in the sequences of these regions themselves. In any version of the human genome reference sequence, any consecutive nucleic acid sequence that, through sequence alignment and identification, has at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the aforementioned regions falls within the scope of protection of this application. The scope of protection for the differentially methylated regions described in this application can be determined by comparing sequence alignment (such as BLAST, BLAT, etc.) with the reference genome hg19 version corresponding to the coordinates of the regions listed in the claims. Depending on the actual detection platform and bioinformatics analysis process, an offset range of ±500bp, preferably ±200bp, and more preferably ±100bp is allowed. Any consecutive nucleic acid sequence with at least 95%, at least 97%, or at least 99% sequence identity with the described region is within the scope of protection of this application. Regardless of subsequent iterations of sequencing technologies or analysis platforms, as long as the core target region detected by them exhibits the aforementioned sequence consistency with the region protected by this application, it constitutes an equivalent feature.

[0038] The second aspect of this application discloses the use of reagents for detecting methylated biomarkers in a test sample in the preparation of products for diagnosing lung cancer.

[0039] In this application, the reagent described is used to detect the methylation level of the differentially methylated regions mentioned above. Specifically, various methods known to those skilled in the art can be applied to detect the methylation level of genomic DNA or cfDNA in a sample. In some specific embodiments, applicable detection methods include, but are not limited to, cfMeDIP-seq technology, other immunoprecipitation methods based on antibodies or methylation-binding proteins, methylation-specific enzyme digestion methods, bisulfite sequencing (whole-genome bisulfite sequencing WGBS, simplified bisulfite sequencing RRBS, etc.), methylation-specific PCR, and mass spectrometry-based detection techniques. As a preferred example, the detection method is cfMeDIP-seq technology.

[0040] Therefore, the reagents for detecting methylation levels should be formulated according to the specific detection technology employed. As a specific example, when using cfMeDIP-seq technology, the reagents may include one or more of the following categories: (1) Sample processing reagents: reagents for extracting and purifying cell-free DNA (cfDNA) from plasma samples, such as DNA extraction kits containing lysis buffer, binding buffer, washing buffer and elution buffer.

[0041] (2) Immunoprecipitation reagents: These include antibodies or functional fragments thereof that specifically recognize and bind to 5-methylcytosine (5mC), and solid-phase carriers or magnetic beads (such as Protein G beads or streptavidin beads) for separating antibody-DNA complexes. Binding buffers and washing buffers required for the immunoprecipitation reaction are also included.

[0042] (3) Library construction and sequencing reagents: including enzymes and buffers for cfDNA end repair, A-tailing, adapter ligation, as well as DNA polymerase, dNTPs, universal primers and purification magnetic beads required for PCR amplification.

[0043] (4) Quantitative and quality control reagents: such as related detection reagents used to determine the initial amount of cfDNA, library concentration and fragment distribution.

[0044] It should be understood that, in practical applications, there can be a variety of choices regarding the specific composition and type of reagents used to detect methylation levels, as long as they can achieve in vitro qualitative or quantitative detection of the methylation status of the aforementioned specific differentially methylated regions. This application does not impose any specific limitations on this.

[0045] In some embodiments of this application, the use of the aforementioned differentially methylated regions for lung cancer diagnosis specifically refers to distinguishing between lung cancer and healthy controls without tumors, or distinguishing between benign and malignant pulmonary nodules.

[0046] In one application scenario, diagnosis refers to identifying subjects exhibiting pulmonary symptoms or having lung cancer risk factors (such as a long history of smoking, family history, etc.) to distinguish whether they are lung cancer patients, healthy individuals, or individuals with non-cancerous diseases (such as inflammation, tuberculosis, hamartomas, and other benign lung lesions). This distinction is of great significance in avoiding unnecessary invasive examinations and treatments, and reducing the physical and mental burden and economic costs on subjects.

[0047] In another application scenario, diagnosis refers to further differentiating whether a lung nodule is malignant (i.e., lung cancer) or benign in subjects found to have lung nodules through imaging (such as low-dose spiral CT-LDCT). Clinically, the vast majority of lung nodules detected by LDCT screening are benign, but their differential diagnosis remains a major clinical challenge. The differentially methylated region provided in this application can serve as a highly specific liquid biopsy method, accurately determining the benign or malignant nature of lung nodules at the molecular level, effectively compensating for the shortcomings of imaging methods and significantly reducing the risk of overdiagnosis.

[0048] Furthermore, in this application, the product may be any form that can be used for in vitro diagnostics, such as a formulation, reagent kit, gene chip, sequencing library, or diagnostic system, and this application does not make any specific limitation in this regard.

[0049] A third aspect of this application provides a kit for diagnosing lung cancer. The kit includes reagents for detecting the methylation level of differentially methylated regions in a test sample as defined in any of the preceding claims. Explanations and examples of the reagents have been detailed in the preceding section and will not be repeated here.

[0050] In some preferred embodiments, the test sample detected by the kit is plasma from the subject. Plasma samples contain cell-free DNA (cfDNA), which originates from various cells in the body, including tumor cells, and carries rich epigenetic information. Using plasma as a test sample only requires routine venous blood collection, offering significant advantages such as being minimally invasive, convenient, and repeatable, making it particularly suitable for early screening, dynamic monitoring, and health management of large populations for lung cancer.

[0051] The specific components in the kit can be configured appropriately based on different detection technologies, which is a capability possessed by those skilled in the art, and therefore will not be elaborated upon here. For example, the kit may employ a cfMeDIP-seq-based technology. In this case, the kit may contain separate sections for reagents used for cfDNA extraction, immunoprecipitation, library construction, and purification, and may optionally include an instruction manual guiding the implementation of the assay. The core components of the kit include a 5mC specific antibody and / or the corresponding immunoprecipitation reaction buffer system.

[0052] A fourth aspect of this application provides a device for diagnosing lung cancer. This device integrates or interconnects multiple functional modules to automate or semi-automatize the process from sample data acquisition to the final diagnostic result output.

[0053] In some specific embodiments, the device includes a data acquisition module and a data analysis module. Wherein: The data acquisition module is configured to obtain the methylation level of differentially methylated regions in the subject's sample as defined in any of the preceding descriptions. The term "obtain" here is broad, referring both to the module directly generating data through its integrated detection instrument and to the module receiving experimental data from an external, independent detection device via a communication interface. In the implementation using cfMeDIP-seq technology, the data obtained by the data acquisition module is the methylation enrichment fraction or other normalized methylation level value of the differentially methylated regions calculated after cfMeDIP-seq sequencing of the subject's plasma cfDNA and processing through a bioinformatics analysis workflow.

[0054] The data analysis module is configured to: based on the methylation level data of the differentially methylated regions obtained by the data acquisition module, use a pre-built diagnostic model to distinguish and determine whether the subject has lung cancer, and generate a diagnostic result.

[0055] In some specific examples, the data acquisition module further includes a sample acquisition unit and a detection unit. The sample acquisition unit is configured to collect or receive the subject's sample for testing; for example, it can be a syringe for loading plasma sample tubes. The detection unit integrates various functional components for implementing cfMeDIP-seq detection, such as a temperature control system, an automated liquid handling system, and a sequencer for performing immunoprecipitation reactions, library preparation, and sequencing.

[0056] In some specific examples, the data analysis module is essentially a computing device that includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to specifically implement the steps of building, validating, and applying the diagnostic model.

[0057] As a specific example, the computational steps performed by the processor may include: First, model building and validation are performed: The processor acquires methylation level data of differentially methylated regions from a given population of known samples (i.e., samples definitively diagnosed as lung cancer or not). The data is randomly divided into two groups: a training set and a test set. Then, using the training set data, a diagnostic model is built based on a machine learning algorithm. Various types of machine learning algorithms can be used; preferred examples include Random Forest, Ridge Regression, LASSO Regression, or Support Vector Machines (SVM). Finally, the diagnostic model is validated on the test set, evaluating its AUC, sensitivity, and specificity.

[0058] Secondly, clinical sample diagnosis is performed: For a new sample to be tested, after obtaining the methylation level data of its corresponding differentially methylated regions through the data acquisition module, the processor inputs this data into the constructed and validated diagnostic model. The model will calculate and output a discrimination score. The processor compares this score with a preset threshold (e.g., a cutoff value determined based on the Youden index) to ultimately determine whether the subject has lung cancer or not, or whether the subject's lung nodules are benign or malignant, and then outputs the results to the output device (such as a monitor or printer).

[0059] Among these advantages, using the random forest algorithm to construct the diagnostic model offers several benefits. Random forest is an ensemble learning algorithm that can effectively handle high-dimensional methylation data (i.e., multiple differentially methylated regions), is less prone to overfitting, and has good tolerance for noise and outliers. Furthermore, this algorithm can evaluate the diagnostic contribution of each biomarker, helping to understand the weight of different regions in lung cancer diagnosis. Of course, in practical applications, the data analysis module of this device can also choose other computer programs integrating different algorithms to achieve similar functions; this application does not limit this choice.

[0060] Furthermore, it is understood that in this application, the "threshold" or "cutoff value" refers to the discrimination threshold used in the diagnostic model to convert continuous model output scores or probabilities into binary diagnostic results (such as "having lung cancer" and "not having lung cancer"). When the score or probability calculated by the diagnostic model for the test sample is higher than this threshold, it is judged as positive (i.e., having lung cancer); when it is lower than or equal to this threshold, it is judged as negative (i.e., not having lung cancer). The setting of the threshold directly determines the sensitivity and specificity of the diagnostic method.

[0061] In some specific implementations, the determination of the threshold or cutoff value is based on receiver operating characteristic (ROC) curve analysis. Specifically, in the training set samples, the continuous predicted values ​​output by the diagnostic model are used as test variables, and the known clinical gold standard diagnostic results of the samples are used as state variables to plot ROC curves. Subsequently, the sensitivity and specificity corresponding to each possible critical point on the ROC curve are calculated, and the Youden index is further calculated, where the Youden index = sensitivity + specificity - 1. The critical point corresponding to the maximum value of the Youden index is determined as the optimal cutoff value. At this cutoff value, the diagnostic method achieves the best balance between sensitivity and specificity.

[0062] It should be understood that in actual clinical applications, the threshold setting can be flexibly adjusted according to specific application scenarios and clinical needs. For example, in lung cancer screening, to minimize missed diagnoses, the threshold can be appropriately lowered to improve sensitivity, allowing for a certain false positive rate. In auxiliary differential diagnosis, to avoid unnecessary invasive examinations, the threshold can be appropriately increased to improve specificity. Regardless of the threshold setting strategy used, as long as it utilizes the methylation level of the differentially methylated region defined in this application as the criterion, it falls within the protection scope of this application. The distinction between benign and malignant pulmonary nodules can be achieved using the above method, which will not be elaborated here.

[0063] The present application will be further illustrated below with reference to specific embodiments. It should be noted that the specific embodiments below are for illustrative purposes only and do not limit the scope of the present application in any way.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0065] In addition, unless otherwise specified, methods without detailed conditions or steps are conventional methods, and the reagents and materials used are commercially available.

[0066] Subject information involved in the following examples: A total of 182 participants were included in the study. All participants were recruited from the Second Affiliated Hospital of Anhui Medical University. Among them, 79 were lung cancer patients, and 103 were non-tumor healthy controls who underwent physical examinations during the same period. Furthermore, all participants with nodules underwent surgical resection. The benign or malignant nature of the nodules was determined based on histological diagnosis. Based on the size and malignancy of the lung nodules, all samples were divided into four groups: a normal control group without nodules (n=79), a benign nodule group (n=24) with nodules smaller than 3 cm, a malignant nodule group (n=46) with nodules larger than 3 cm, and a lung cancer group (n=33) with nodules larger than 3 cm. Specific information is shown in Table 1. Peripheral blood and clinical data, including age, sex, and stage, were collected from all lung cancer patients before treatment. Non-tumor samples were collected during routine physical examinations, excluding participants with a history of cancer treatment or recent medication use. All patients provided written informed consent. The study was approved by the Institutional Review Committee of the Ethics Committee of Anhui Medical University (YX2020-019) in accordance with all relevant ethical regulations.

[0067] Table 1. Statistical analysis of clinical information of 182 subjects

[0068] In Table 1, 1 The values ​​represent the mean (standard deviation SD) and n / N (%). The screening cohort is used to screen for differentially methylated regions, while the validation cohort serves as an independent clinical cohort to validate the diagnostic efficacy of the screened differentially methylated regions.

[0069] Example 1: Screening of differentially methylated regions 1.1 Subject Recruitment The subjects included in this study were the screening cohort shown in Table 1 (N=152).

[0070] 1.2 Plasma processing and cfDNA extraction and concentration detection 8 mL of peripheral blood was drawn from each subject using an EDTA anticoagulant tube. After collection, the blood was stored at 4°C, and plasma was extracted within 4 hours. The peripheral blood was centrifuged at 1600g for 10 minutes at 4°C, and the supernatant was transferred to a 2 mL EP tube. The EP tube was then centrifuged at 16000g for 10 minutes at 4°C, and the supernatant was again transferred to a 2 mL EP tube. Care was taken to avoid aspirating the leukocyte layer to prevent leukocyte DNA contamination of cfDNA. If cfDNA was not extracted immediately, the plasma should be stored at -80°C until cfDNA extraction.

[0071] Cell-free DNA (cfDNA) in plasma was extracted using the Qiagen Circulating Nucleic Acids Kit (Qiagen, Cat# 55114). The specific procedure is as follows: (1) Prepare the Carrier RNA solution. Add 310 μL of N-Nase-Free ddH2O to the Carrier RNA lyophilized powder provided with the kit, and shake until the Carrier RNA lyophilized powder is completely dissolved to prepare a 1 μg / μL Carrier RNA solution. Then, aliquot the Carrier RNA solution into 200 μL PCR tubes, 15 μL per tube, and store at -20℃. Each sample requires 1 μL of Carrier RNA solution. Each time the Carrier RNA solution is used, take the corresponding number of PCR tubes according to the number of samples, and discard them after use.

[0072] (2) Turn on the instruments and prepare the reagents. First, turn on the water bath and set the temperature to 60°C. Then turn on the constant temperature metal bath and set the temperature to 56°C. Next, turn on the ice maker. Finally, take out the plasma sample and carrier RNA solution and wait for the samples to equilibrate to room temperature before proceeding to the next step.

[0073] (3) Add plasma and 0.1 times the plasma volume of proteinase K solution to a 50 mL centrifuge tube, then add 0.8 times the plasma volume of ACL solution and 1 μL of carrier RNA solution. Cap the centrifuge tube and vortex for 30 seconds to mix thoroughly. Then immediately place it in a water bath and incubate at 60°C for 30 minutes.

[0074] (4) Take out a 50mL centrifuge tube, open the cap and add 1.8 times the plasma volume of ACB solution, vortex on a vortex mixer for 30 seconds, then close the cap and incubate on ice for 5 minutes.

[0075] (5) Assemble the vacuum filtration apparatus. Insert the QIAamp Mini column into the QIAvac 24 Plus apparatus, then insert a 20 mL extension tube into the QIAamp Mini column, and finally connect the QIAvac 24 Plus apparatus to the vacuum pump.

[0076] (6) Pour the mixture from step (4) into the extension tube. If the volume of the mixture is greater than 20 mL, it can be poured in several times. Turn on the vacuum pump, remove all the mixture, and then turn off the vacuum pump. After the vacuum pump pressure reaches 0, discard the 20 mL extension tube.

[0077] (7) Add 600 μL of ACW1 buffer to each QIAamp Mini column, turn on the vacuum pump, and after all the liquid has flowed out, turn off the vacuum pump and wait for the pressure to reach 0. Then add 750 μL of ACW2 buffer to the QIAamp Mini column, turn on the vacuum pump, and after all the liquid has flowed out, turn off the vacuum pump and wait for the pressure to reach 0. Next, add 750 μL of anhydrous ethanol to the QIAamp Mini column, turn on the vacuum pump, and after all the liquid has flowed out, turn off the vacuum pump and wait for the pressure to reach 0.

[0078] (8) Remove the QIAamp Mini column and place it into a new 2mL EP tube. Open the EP tube cap and incubate it in a 56°C constant temperature metal bath for 10 minutes to remove residual solvent and anhydrous ethanol.

[0079] (9) Place the QIAamp Mini column into a new 1.5 mL EP tube, and drop 30 μL of AVE buffer into the middle of the membrane of the QIAamp Mini column. Be careful not to let the pipette tip touch the membrane of the column to avoid damaging it. Cover the EP tube and incubate at room temperature for 3 minutes, then centrifuge at 13000 g for 1 minute.

[0080] (10) Add 30 μL of AVE buffer to the middle of the membrane of the QIAamp Mini column again, cover it, incubate at room temperature for 3 minutes, and then centrifuge at 13000g for 1 minute in a high-speed centrifuge. After centrifugation, discard the QIAamp Mini column. The cfDNA solution in the new EP tube is the final elution volume of 55 μL.

[0081] Subsequently, the concentration of cfDNA was detected using the Qubit dsDNA High Sensitivity Assay Kit (Thermo Fisher, Cat# Q33231) according to the instructions.

[0082] 1.3 cfMeDIP-seq library preparation and sequencing First, the cfDNA samples extracted in Section 1.2 were end-repaired and 3' A-tailed using the KAPA HyperPrep Library Construction Kit (KAPA Biosystems, Cat# KK8504). Then, the cfDNA samples were ligated to sequencing adapters using the NEBNext Adapter Kit (NEBNext Multiplex Oligos for Illumina, New England BioLabs, Cat# E7335L). The adapters were then cut using USER enzyme (New England BioLabs, Cat# E7335L), and the products were purified using the Qiagen MinElute PCR Purification Kit (Qiagen, Cat# 28004). Next, Medip libraries were prepared using the MagMeDIP DNA Methylation Immunoprecipitation Kit (Diagenode, C02010021), and the library products were purified using AMPureXP magnetic beads (Beckman Coulter, catalog number A63882).

[0083] The specific steps are as follows: (1) End repair and 3' end A addition. Take a PCR tube and transfer 50 μL of the cfDNA solution extracted in Section 1.2, 3 μL of KAPA End Repair A-Tailing Enzyme Mix, and 7 μL of KAPA End Repair A-Tailing Buffer to the PCR tube. Vortex to mix and briefly centrifuge. Then place the PCR tube in a PCR amplifier and set the program to 20℃ (30 min), 65℃ (30 min), and finally 4℃, with the hot cap temperature set to 75℃.

[0084] (2) Adapter ligation. Take the reaction product from the previous step, add 10 μL of KAPA DNA Ligase, 30 μL of KAPALigation Buffer and 10 μL of Adapter solution, vortex to mix and then centrifuge briefly. Then place it in a PCR amplification instrument, set the program to 20℃ (20 min), then maintain at 4℃, and set the hot lid temperature to none.

[0085] (3) USER enzyme cleavage of the adapter. Add 3 μL of USER enzyme to the reaction product from the previous step, vortex to mix and then centrifuge briefly. Then place it in a PCR amplifier, set the program to 37℃ (15 min), then maintain at 4℃, and set the hot cap temperature to 50℃.

[0086] (4) Purify the DNA mixture from the previous step using the Qiagen MinElute PCR Purification Kit. Add 5 volumes of PB Buffer to the reaction product from the previous step, vortex to mix, and briefly centrifuge. Then add to the purification column provided with the kit, centrifuge at 13300g for 1 min, and discard the filtered liquid. Next, add 750 μL of PE buffer to the purification column, centrifuge at 13300g for 1 min, and discard the filtered liquid. Finally, add 20 μL of EB Buffer to the center of the membrane of the purification column, let stand for 1 min, and centrifuge at 13300g for 1 min. The final elution volume is 20 μL.

[0087] (5) Prepare 1× MagBuffer A solution. Prepare 1× MagBuffer A solution by mixing 20 μL MagBuffer A (5×) and 80 μL ddH2O for each sample.

[0088] (6) Magnetic bead washing. Vortex the Magbeads until no precipitate remains at the bottom. Take 11 μL of Magbeads beads for each sample and add the corresponding volume to a 1.5 mL centrifuge tube. Place the tube on a magnetic rack for 1 min to allow complete adsorption. Use a pipette to aspirate and discard the supernatant. Remove the centrifuge tube and add 22 μL of 1× MagBuffer A solution for each sample. Mix well with a pipette, centrifuge lightly, and place on a magnetic rack for 1 min. Discard the supernatant again. Repeat the washing process once more, for a total of two washes. Finally, add 22 μL of 1× MagBuffer A solution for each sample to the centrifuge tube, mix well, and place on ice for later use.

[0089] (7) Prepare Mag master mix solution. Take 24 μL MagBuffer A 5×, 6 μL MagBuffer B and 3 μL ddH2O for each sample, mix well and place on ice for later use.

[0090] (8) Prepare antibody mix solution. Take 0.3 μL Antibody, 0.6 μL MagBuffer A 5×, 2.1 μL ddH2O and 2.0 μL MagBuffer C solution for each sample, shake to mix well, and place on ice for later use.

[0091] (9) Prepare hybridization reagents. Take 1.5 mL of EP tube for each sample, add 33 μL Mag master mix solution, 20 μL DNA purified in step (4) and 37 μL ddH2O, place the EP tube in a metal bath, heat at 95 °C for 3 min, and then insert it into ice for 2 min.

[0092] (10) Antibody hybridization. Take 79 μL of the mixed solution from the previous step and add it to an eight-tube, then add 5 μL of the antibody mix solution prepared in step (8) and 20 μL of Magbeads washed in step (6), mix well with a pipette, fix it on a vertical mixer, and mix by rotating at 4°C for more than 16 hours.

[0093] (11) Preparation of DNA elution reagents and instruments. First, turn on the ice maker and prepare the ice box. Place Mag Washbuffer 1, Mag Wash buffer 2, and a 200 μL magnetic rack on ice for pre-cooling. Then, turn on the two metal baths and set the temperatures to 55℃ and 100℃ respectively, ensuring that the metal baths reach the final temperature. Next, remove the Xp Beads from the freezer, vortex to mix until there is no precipitate at the bottom of the magnetic beads, and equilibrate to room temperature to avoid the low temperature causing a significant reduction in the ability of PEG to adsorb DNA. Finally, prepare the DIB mix solution. Prepare the corresponding volume of DIB mix solution according to the volume of 1 μL Proteinase K and 100 μL DIB solution for each sample. Vortex to mix thoroughly and place on ice for later use.

[0094] (12) Magnetic bead washing. Remove the sample from the vertical mixer, centrifuge briefly, place it on a magnetic rack and let it stand for 1 min, then discard the supernatant. Then add 100 μL MagWash Buffer-1 to the eight-tube bundle and place it on a vertical mixer at 4 °C and mix for 5 min.

[0095] (13) Repeat step (12) twice, and wash the magnetic beads three times with MagWash Buffer-1 solution.

[0096] (14) Remove the sample from the vertical mixer, centrifuge briefly, place it on a magnetic rack and let it stand for 1 min, then discard the supernatant. Add 100 μL of MagWash Buffer-2 to the eight-tube container, then place it on a vertical mixer at 4°C and mix by rotation for 5 min, then place it on a magnetic rack and let it stand for 1 min. Use a pipette to remove the supernatant, let it stand for another 1 min, and finally use a 10 μL pipette to remove the liquid at the bottom.

[0097] (15) Take 100 μL of the DIB mix solution from step (11) and add it to the washed magnetic beads containing DNA from the previous step. Shake well and transfer to a 1.5 mL centrifuge tube. Then place the centrifuge tube on a metal bath at 55°C and heat for 15 min to fully digest the 5 mC monoclonal antibody. Next, place the centrifuge tube on a metal bath at 100°C and heat for 15 min to fully inactivate Proteinase K (serine protease) and prevent it from affecting subsequent reactions. Finally, place the centrifuge tube on ice to cool. After the DIB mix solution has cooled, centrifuge briefly and place it on a magnetic rack to stand for 1 min. After the magnetic beads are adsorbed onto the tube wall, take the supernatant into a new 1.5 mL centrifuge tube.

[0098] (16) Purify DNA. Vortex the XP beads again to mix. Add 180 μL (1.8 times the volume) of XP beads to the product from the previous step. Vortex to mix, let stand for 5 min, then briefly centrifuge. Place on a magnetic rack and let stand for 3 min, then aspirate the supernatant. Remove the centrifuge tube, add 200 μL of 80% ethanol solution, vortex to mix, let stand for 5 min, then briefly centrifuge. Place on a magnetic rack and let stand for 3 min, then aspirate the supernatant. Wash again with 80% ethanol solution. Aspirate the supernatant and let stand for 1 min, then use a 10 μL pipette tip to aspirate the liquid from the bottom of the centrifuge tube. Place the magnetic beads on a magnetic rack and wait for them to dry.

[0099] (17) Add 23 μL of ddH2O to the centrifuge tube and mix it by pipetting. After standing for 5 min, centrifuge briefly, place it on a magnetic rack and stand for 3 min. Then, transfer 22 μL of the supernatant into a new eight-tube.

[0100] (18) PCR amplification. Add 1.5 μL i5 index, 1.5 μL i7 index and 25 μL Hot Start ReadMix solution to the eight-tube containing purified DNA from the previous step, and mix well with a pipette. Place the eight-tube into the PCR amplification instrument and set the program as follows: 95℃ (3 min); then 98℃ (20 s), 65℃ (15 s), 72℃ (30 s), for a total of 15 cycles; then 72℃ (1 min); and finally hold at 4℃.

[0101] (19) The amplified library was purified using AMPure XP magnetic beads (Beckman Coulter, Cat#A63882), following step (16). The library size distribution was analyzed using a Bio-Fragment Analyzer (BIOptic), and quantification was performed using the Qubit dsDNA High Sensitivity Assay Kit (Thermo Fisher, Cat#Q33231). The final library was sequenced at 150 bp at both ends using an Illumina Nova 6000 sequencer.

[0102] 1.4 Data Analysis 1.4.1 Sequencing data quality control and methylation level analysis The raw sequencing data were first quality controlled using FastQC (v0.11.7) software, and Trim Galore (v0.6.3) was used to remove non-compliant fragments and bases (including sequencing adapter sequences, bases with a quality lower than 20, and fragments with a length lower than 20 bp). Then, the sequencing data were aligned to the hg19 genome using bowtie2 software, and duplicate fragments were removed using samtools software. Finally, the MEDIP package in R was used to extract the number of sequencing fragments within each 300 bp sliding window of the genome. 10,318,911 sliding windows were obtained for each sample, and the value of each sliding window was the number of fragments falling within that sliding window.

[0103] 1.4.2 Standardization and DMR Analysis To screen for lung cancer-specific differential methylation regions (DMRs), standardization and differential analysis were performed on all lung tumor and normal samples. First, samples and fragments that did not meet the requirements were screened out based on several indicators. To minimize the impact of sequencing data quality, the total number of fragments after deduplication of all sample sequencing data was counted using bamdst (v1.0.9) software, and fragments with less than 10,000,000 fragments were removed. Since most methylation regions are enriched only in a few samples, methylation regions with more than 10 fragments in less than 10% of samples were removed to minimize their impact on subsequent data standardization. Fragments located in sex chromosome and mitochondrial genome regions were also removed to eliminate the influence of sex and mitochondrial genes. Next, methylation levels were standardized using the Voom function in the limma package. Finally, differential analysis was performed using the limma package in R software to screen for DMRs.

[0104] 1.4.3 Hold-out method and propensity score matching (PSM) Propensity score matching (PSM) is a method to reduce interference from other biases and confounding factors. The principle is to use propensity score to find one or more identical or similar samples as controls, thereby reducing interference from other unknown factors. To reduce interference from age, gender, and other possible confounding factors, a hold-out method was used for cross-validation, and the MatchIt package in R software was used to perform 1:1 propensity score matching between tumor and control groups. The scores were calculated using a logit regression model, and the obtained matching scores were then used as covariates in the limma difference analysis. Propensity score matching analysis excludes some samples. To include more samples in the difference analysis, a hold-out method was used to randomly select 80% of the samples from all screening cohorts as the training set, and the remainder as the test set. This was repeated 100 times, with propensity score matching and difference analysis performed each time to ensure that more samples appeared at least once in the difference analysis. The limma package was used to perform differential analysis on the training set samples for each cycle, and the top 300 DMRs for each cycle were extracted. Then, the frequency of all DMRs in each group was counted. After removing duplicate DMRs, the 50 DMRs with the highest frequency in each group were extracted as the specific methylation regions of that group. A total of 200 DMRs were obtained from the four groups.

[0105] 1.5 Feature Selection and Model Building for Machine Learning Algorithms Feature selection was performed using the LASSO and Boruta algorithms, and model building was performed using the random forest algorithm.

[0106] Specifically, the LASSO and Boruta algorithms were first used to perform feature filtering on the 200 DMRs obtained in Section 1.4.

[0107] Sixteen DMRs were selected, and their details are shown in Table 2. Among them, two were hypomethylated and 14 were hypermethylated. Figure 1 (A)

[0108] Table 2. 16 Differentially Methylated Regions (DMRs)

[0109] Further development of lung cancer models based on cfDNA-specific DMRs was conducted, and a random forest model was constructed using 16 screened DMRs. Based on the differential analysis grouping, 80% of the 152 samples from this example were used to train the models, and the remaining samples were used to test model power, resulting in a total of 100 models. The predicted values ​​of all test set samples were combined, and the AUC was calculated using the pROC package in R. The final AUC of the model distinguishing between lung cancer and non-tumor controls in the test set was 0.9854 (…). Figure 1(B) The results showed that the model could accurately distinguish between lung cancer and non-tumor controls.

[0110] In this embodiment, R (4.1.1) and Rstudio software were used for data analysis. Chi-square test was used for categorical variables, Wilcox test for two groups of continuous variables, and Kruskal test for multiple groups of continuous variables. A p-value < 0.05 was considered statistically significant. ROC curves were plotted using the pROC package in R software, and the Youden index was used to determine the optimal cutoff value.

[0111] Example 2: Validation of Model Diagnostic Efficacy The validation cohort in Table 1 was used as an independent clinical validation dataset. The methylation levels of the 16 DMRs shown in Table 2 were obtained according to the method in Example 1 and input into the lung cancer model in Example 1.

[0112] The results showed that the AUC of this lung cancer model in distinguishing between lung cancer and non-tumor controls was 0.9638 ( Figure 2 (A). Next, the samples in the screening and validation queues were merged for a more comprehensive analysis. This embodiment analyzed the power of the model to distinguish between cancer and non-cancer control groups in early and late stages. The results showed that the AUC value of the model in distinguishing between cancer and non-cancer control groups in early stages was 0.9683, and the AUC value in distinguishing between cancer and non-cancer control groups in late stages was 0.9983 (AUC value). Figure 2 (B). Next, this embodiment compared the ability of each clinical trait and model to distinguish between cancer patients and healthy control samples. The results showed that the methylation risk model had the highest power, with an AUC value of 0.9824, which was much higher than age (0.7684), sex (0.5262), and cfDNA concentration (0.7453). Figure 2 (C). The results show that the diagnostic model in this application can accurately distinguish between lung cancer and non-tumor controls at an early stage, and has higher accuracy than clinical characteristics.

[0113] Because there were significant differences in clinical traits between the cancer and non-cancer groups, this application further conducted a clinical subgroup analysis of the model risk score to analyze the relationship between the model and clinical traits. Age and cfDNA concentration were grouped according to the median, and then the differences in model risk scores between each group were statistically analyzed. The results showed that there were no differences in model risk values ​​among different age, sex, and cfDNA concentration groups. Figure 2 (D). However, the risk increases with disease progression and clinical stage; the risk is significantly higher in patients with benign nodules than in the normal control group, and the risk is significantly higher in the lung cancer group than in the malignant nodule group. Figure 2 In the middle stage (E), the risks in the early stages are far higher than in the later stages ( Figure 2(F). Therefore, this model is independent of age, sex, and cfDNA concentration, and can reflect disease progression.

[0114] Since the samples were divided into four groups based on nodule benignity / malignancy and size in this application, a multi-classification machine learning model was constructed based on these 16 DMRs using the training set described above to evaluate their ability to distinguish between benign and malignant lung nodules. The results show that the AUC of these 16 DMRs in distinguishing between benign and malignant lung nodules on the test set was 0.8899 (…). Figure 3 In the validation set, the AUC for distinguishing between benign and malignant pulmonary nodules was 0.8732 (A). Figure 3 (Middle B). The results showed that these 16 DMRs could not only distinguish between lung cancer and non-tumor controls, but also distinguish between benign and malignant lung nodules.

[0115] It should be noted that this application is not limited to the above-described embodiments. The above embodiments are merely examples, and any embodiments with the same structure and effect as the technical concept within the scope of this application are included in the technical scope of this application. Furthermore, various modifications that can be conceived by those skilled in the art to the embodiments, and other ways of constructing by combining some of the constituent elements of the embodiments, without departing from the spirit of this application, are also included in the scope of this application.

Claims

1. A methylation biomarker for diagnosing lung cancer, characterized in that, The methylation biomarker consists of one or more of the following differentially methylated regions. composition: chr1: 25328700-25329000, chr1: 56129400-56129700, chr10: 46957200-46957500, chr11: 129433800-12943410 0, chr11: 55847100-55847400, chr15: 20546100-20546400, chr15: 20546700-20547000, chr16: 33294300-33294 600, chr18: 106800-107100, chr19: 50624100-50624400, chr2: 119060400-119060700, chr22: 19671600-196719 00, chr5: 17586000-17586300, chr8: 86742600-86742900, chr8: 86760300-86760600, chr9: 67103400-67103700; The reference genome version for the differentially methylated regions above is hg19.

2. The use of the reagent for detecting the methylated biomarker described in claim 1 in the preparation of products for diagnosing lung cancer.

3. The application as described in claim 2, characterized in that, The reagent is used to detect the methylation level of the differentially methylated regions.

4. The application as described in claim 2, characterized in that, The diagnosis of lung cancer refers to distinguishing between lung cancer and healthy controls without tumors, or distinguishing between benign and malignant lung nodules.

5. A reagent kit for diagnosing lung cancer, characterized in that, The kit includes reagents for detecting the methylation level of differentially methylated regions in a sample as defined in claim 1.

6. The reagent kit as described in claim 5, characterized in that, The sample to be tested is the subject's plasma.

7. A device for diagnosing lung cancer, characterized in that, It includes a data acquisition module and a data analysis module, wherein: The data acquisition module is configured to: obtain the methylation level of the differentially methylated regions in the subject's sample as defined in claim 1; The data analysis module is configured to: based on the methylation level of the differentially methylated regions obtained by the data acquisition module, distinguish whether the subject has lung cancer through a constructed diagnostic model.

8. The apparatus as claimed in claim 7, characterized in that, The data acquisition module includes a sample acquisition unit and a detection unit, wherein the sample acquisition unit is configured to acquire a test sample from the subject, and the detection unit is configured to detect the methylation level of differentially methylated regions in the test sample.

9. The apparatus as claimed in claim 7, characterized in that, The data analysis module includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program stored in the memory to perform the following steps: The methylation levels of differentially methylated regions in known samples of a given population are randomly divided into two groups: a training set and a test set. Using the training set data, a diagnostic model is built based on machine learning algorithms, and the obtained diagnostic model is validated on the test set. The methylation level of the differentially methylated region of the sample to be tested is input into the diagnostic model to determine the diagnostic result.

10. The apparatus as claimed in claim 9, characterized in that, The machine learning algorithm is the random forest algorithm.