Lung cancer early screening model construction method based on lp-wgs and dna methylation and electronic device
By constructing a lung cancer early screening model based on LP-WGS and DNA methylation, using peripheral blood cfDNA for sequencing and feature extraction, and establishing a multi-feature cross-stacking machine learning model, the radiation exposure and low compliance rate problems of existing lung cancer screening methods are solved, and high-precision early lung cancer detection is achieved.
Patent Information
- Application Number
- CN202311311093.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-10
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2043-10-10
AI Technical Summary
Existing lung cancer screening methods, such as low-dose computed tomography, have problems with radiation exposure and low compliance rates. The biomarkers in liquid biopsies have low sensitivity and high false positive rates, making it difficult to achieve high-accuracy detection of early lung cancer.
A lung cancer early screening model based on low-depth whole-genome sequencing and DNA methylation was constructed. By collecting cfDNA from peripheral blood, sequencing and feature extraction were performed, and a multi-feature cross-stacking machine learning model was established for training and verification to improve detection accuracy.
It achieves non-invasive and accurate early prediction of lung cancer with high accuracy and sensitivity, can dynamically monitor lung cancer risks in real time, and improves the performance of lung cancer screening.
Smart Images

Figure CN117275585B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of biomedicine, and specifically to a method for constructing an early screening model for lung cancer based on LP-WGS and DNA methylation and an electronic device. Background Art
[0002] Lung cancer is the most common cancer in China. Most lung cancer patients do not experience specific symptoms in the early stages of the disease. Typically, 64.6% of cancer cases are diagnosed at an advanced stage (stage III / IV). Therefore, screening at a low tumor burden is crucial to reducing lung cancer-related mortality. Low-dose computed tomography (LDCT) is commonly used for lung cancer screening, but radiation exposure and ambiguous risk assessments result in a low compliance rate (35.6%) and may lead to overdiagnosis.
[0003] Liquid biopsies represent a promising approach for cancer screening, requiring only small amounts of biological fluid. Their advantages lie in their ease of access and cost-effectiveness, allowing for repeated sampling, which facilitates adherence to cancer screening. For years, blood biomarker tests, such as cytokeratin 19 fragment (CYFRA 21-1), neuron-specific enolase (NSE), and squamous cell carcinoma antigen (SCC-Ag), have been investigated as potential biomarkers for lung cancer. However, these biomarkers have demonstrated suboptimal performance in early diagnosis, with low sensitivity and high false-positive rates. Circulating free DNA (cfDNA) released from tumor tissue possesses unique genetic and epigenetic patterns similar to those of its source.
[0004] Compared to rare somatic mutations detectable in the blood, epigenetic dysregulation during tumorigenesis is an early event involving widespread changes in DNA methylation and chromatin structure across the genome. Methylation modifications often occur in specific genomic regions, such as CpG islands, which provides an opportunity to analyze the large number of changes generated during tumorigenesis through targeted sequencing. Recent studies have shown that epigenome-based models are superior to mutation-based models in the early detection of lung cancer. DNA methylation is closely related to the occurrence and progression of various tumors. Many studies have found that different diseases and even different stages of the same disease may have specific methylation patterns. The frequency of hypermethylation of CpG islands in tumor cells is much higher than that of gene mutations. Therefore, by detecting the methylation level of specific genes or the entire genome, the risk of lung cancer can be predicted.
[0005] In addition, studies using whole-genome sequencing (WGS) have found that cfDNA exhibits non-random fragmentation during tumorigenesis, and fragmentomic features such as fragmentation patterns and terminal motifs vary significantly across the course of the disease. Changes in fragmentation patterns reflect changes in chromatin structure prior to cfDNA release, while terminal motifs reflect changes in chromatin accessibility and nuclease activity. To date, several cfDNA fragmentation features have been used for lung cancer screening, including fragment size coverage, fragment size distribution, terminal motifs, breakpoint motifs, and copy number variation. Results from enzyme digestion experiments have shown that lower DNA methylation levels predict higher nucleosome accessibility and allow nucleases to cleave within nucleosomes to generate shortened DNA fragments, suggesting that DNA methylation may be an important regulator of cfDNA fragmentation.
[0006] Combining these and other promising liquid biopsy markers with current screening programs could significantly improve early screening and diagnosis of lung cancer. To our knowledge, although methylation and fragment signatures have been used separately for early lung cancer detection, few studies have integrated these two epigenetic signatures, which could potentially demonstrate high performance given their complementary contributions. With this in mind, this application is being proposed. Summary of the Invention
[0007] In order to solve one of the above technical defects, the embodiments of the present application provide a non-invasive method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation and an electronic device that can predict early lung cancer with high accuracy.
[0008] According to a first aspect of an embodiment of the present application, a method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation is provided, comprising the following steps:
[0009] S10, collect peripheral blood from lung cancer patients and healthy subjects, extract cfDNA from the peripheral blood, and establish a sample set;
[0010] S20, based on cfDNA, performs low-depth whole-genome and methylation-targeted sequencing to construct sequencing libraries;
[0011] S30, performing feature extraction on the low-depth whole-genome sequencing data and the methylation targeted sequencing data to obtain fragment features of the whole-genome sequencing data, methylation features of the methylation targeted sequencing data, and fragment group features of the methylation data;
[0012] S40: Based on the fragment features of whole-genome sequencing data and the methylation features of methylation targeted sequencing data, a multi-feature cross-stacked prediction model for early screening of lung cancer was constructed;
[0013] S50, training and verifying the lung cancer early screening prediction model through the sample set to obtain the final lung cancer early screening prediction model and prediction results.
[0014] Preferably, the S40, constructing a lung cancer early screening prediction model based on multi-feature cross-stack, comprises:
[0015] S401, establishing a single feature prediction model;
[0016] Including: using machine learning models to establish a classification model based on the characteristics of multiple sequencing data of samples;
[0017] S402, establishing an integrated model of a single feature;
[0018] This includes: concatenating multiple machine learning scores for each feature into a new feature vector, and then building an integrated model of the single feature based on logistic regression;
[0019] S403, establishing a multi-feature joint integration model;
[0020] It includes: splicing multiple feature scores of each sample into a new feature vector, and then establishing a multi-feature joint integration model based on logistic regression.
[0021] Preferably, the machine learning model includes at least one of a gradient boosting model, an XGBoost model, a random forest model, a logistic regression model and a multi-layer perceptron.
[0022] Preferably, the S30, performing feature extraction on the low-depth whole genome sequencing data and the methylation targeted sequencing data to obtain fragment features of the whole genome sequencing data, methylation features of the methylation targeted sequencing data, and fragment group features of the methylation data, includes:
[0023] S301, preprocessing low-depth whole-genome sequencing data and methylation targeted sequencing data respectively;
[0024] S302, performing feature extraction on the pre-processed low-depth whole genome sequencing data to obtain fragment features of the whole genome sequencing data;
[0025] S303 , performing feature extraction on the pre-processed methylation targeted sequencing data to obtain methylation features of the methylation targeted sequencing data and fragment group features of the methylation data.
[0026] Preferably, the pre-processed low-depth whole genome sequencing data is sequencing fragment information, including: the chromosome number of each fragment in the hg19 human reference genome, the fragment start position, the fragment end position, the fragment length, the GC content and the corrected weight value;
[0027] The fragment features of the whole-genome sequencing data include: whole-genome copy number variation, cfDNA long-short fragment ratio, cfDNA fragment size distribution, cfDNA nucleosome pattern and cfDNA 4bp motif end ratio.
[0028] Preferably, in S301, the methylation characteristics are: the proportion of abnormally high methylation fragments, the length ratio of methylated fragments, and the proportion of 4 bp motif ends of methylated fragments.
[0029] Preferably, in S301, the low-depth whole-genome sequencing data in the sequencing library is preprocessed, including:
[0030] S301-11, use fastp software to perform quality control and remove adapter sequences on the raw data after sequencing;
[0031] S301-12, using BWA-MEM, align the fastq data obtained in step S301-11 with the hg19 human reference genome to obtain a bam file after alignment;
[0032] S301-13, using public database information, removing reads from the hg19 genome blacklist regions, spacer regions, patch sequences, highly variable regions, and centromere-telomeric regions in the bam file obtained in step S301-12;
[0033] S301-14, using the aligned position information of the paired-end sequencing data, merging the sequencing data into a single cfDNA fragment, and removing fragments longer than 400 bases from the merged fragments;
[0034] S301-15, perform GC correction.
[0035] Preferably, in S301, preprocessing the methylation targeted sequencing data includes:
[0036] S301-21, use fastp software to perform quality control on the methylation targeted sequencing data and remove sequencing adapters to obtain fastq data;
[0037] S301-22, align the fastq data with the hg19 human reference genome to obtain the aligned bam file;
[0038] S301-23, remove PCR duplications in the bam file and create an index for the bam file after duplication removal;
[0039] S301-24, stitch and connect the CpG sites in the paired test sequences in the bam file to obtain the methylation status haplotype of each fragment.
[0040] Preferably, the step S303 of performing feature extraction on the pre-processed methylation targeted sequencing data to obtain methylation features of the methylation targeted sequencing data and fragment group features of the methylation data includes:
[0041] S303-1, determining tumor-specific methylation intervals by comparing differences in hypermethylated fragments; for the sample to be tested, within the tumor-specific methylation interval, counting fragments in which methylated C bases account for more than half of the total CpG sites or in which the total number of methylated C sites is greater than or equal to 5, and defining them as abnormal fragments; dividing the number of abnormal fragments by the total number of fragments covering the interval to obtain the proportion of hypermethylated abnormal fragments;
[0042] S303-2: Obtain cfDNA fragments within all panel intervals, extract 4 bp base sequences upstream and downstream of the start and end positions, calculate the fragment ratios of 256 4 bp base end sequences, and obtain the proportion of methylated fragments with 4 bp motif ends;
[0043] S303-3: Count the cfDNA fragments within each tumor-specific methylation interval, and based on the length information of the cfDNA fragments, obtain the number of short fragments with a length of 100-150 bp and the number of long fragments with a length of 151-210 bp, and obtain the length-to-short ratio of the methylation fragment group within the interval.
[0044] According to a second aspect of an embodiment of the present application, there is provided an electronic device, including:
[0045] memory; processor; and computer program;
[0046] The computer program is stored in the memory and configured to be executed by the processor to implement the method described above.
[0047] The method and electronic device for constructing a lung cancer early screening model based on LP-WGS and DNA methylation provided in the examples of this application have the following beneficial effects:
[0048] In this application, a sample set based on the fragment features of whole genome sequencing data and the methylation features of methylation targeted sequencing data was established. The lung cancer early screening prediction model based on multi-feature cross-stacking was trained and verified through the sample set to obtain the final lung cancer early screening prediction model and prediction results; the lung cancer early screening prediction model has the advantages of non-invasive detection and high prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0050] Figure 1 A schematic diagram of the process of constructing a lung cancer early screening model based on LP-WGS and DNA methylation provided in an embodiment of the present application;
[0051] Figure 2 Schematic diagram of the process of preprocessing and feature extraction of low-depth whole-genome sequencing data and methylation targeted sequencing data in the sequencing library in the embodiment of this application;
[0052] Figure 3 This is a coverage diagram of the CTCF transcription factor fragment in the examples of this application;
[0053] Figure 4 The copy number variation characteristics of each sample to be tested in the embodiment of this application;
[0054] Figure 5 This is a schematic diagram of the clinical sample characteristics in the examples of this application;
[0055] Figure 6 Schematic diagram of the prediction performance evaluation of the test set and the validation set in a specific embodiment. DETAILED DESCRIPTION
[0056] In order to make the technical solutions and advantages of the embodiments of the present application more clearly understood, the exemplary embodiments of the present application are further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, and are not an exhaustive list of all the embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other unless they conflict.
[0057] like Figures 1 to 4 As shown, the present invention provides a method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation, comprising the following steps:
[0058] S10, collect peripheral blood from lung cancer patients and healthy subjects, extract cfDNA from the peripheral blood, and establish a sample set;
[0059] S20, based on cfDNA, performs low-depth whole-genome and methylation-targeted sequencing to construct sequencing libraries;
[0060] S30, performing feature extraction on the low-depth whole-genome sequencing data and the methylation targeted sequencing data to obtain fragment features of the whole-genome sequencing data, methylation features of the methylation targeted sequencing data, and fragment group features of the methylation data;
[0061] S40: Based on the fragment features of whole-genome sequencing data and the methylation features of methylation targeted sequencing data, a multi-feature cross-stacked prediction model for early screening of lung cancer was constructed;
[0062] S50, training and verifying the lung cancer early screening prediction model through the sample set to obtain the final lung cancer early screening prediction model and prediction results.
[0063] In this embodiment, S10, collecting peripheral blood from lung cancer patients and healthy subjects to extract cfDNA from the peripheral blood, may include: separating plasma by a two-step centrifugation method, extracting cfDNA from the plasma using an Apostle automatic nucleic acid extractor; and detecting DNA concentration using a Qubit 4.0 fluorescence quantification instrument;
[0064] Among them, the blood collection volume of peripheral blood can be: 20mL / person.
[0065] In this embodiment, the S20 method performs low-depth whole-genome and methylation-targeted sequencing based on cfDNA to construct a sequencing library, which may include:
[0066] S201, for low-depth whole-genome sequencing, KAPA Hyper Plus Kit (Roche) can be used to prepare the sequencing library according to the standard operating procedures in the instructions, with an input of 10 ng of cfDNA for library construction; the library is sequenced using the Illumina Nova Seq 6000 sequencer for paired-end sequencing, with a sequencing length of 150 bp and a sequencing depth of approximately 3.3x.
[0067] S202: For methylation-targeted sequencing, the Next Enzymatic Methylation Library Construction Kit can be used and the pre-library prepared according to the standard protocol in the manufacturer's instructions. The cfDNA input for library construction is 10-50 ng. The library is captured using the TWIST custom probe Lucas22 and accompanying reagents according to the standard protocol. The captured library is then paired-end sequenced on an Illumina NovaSeq 6000 sequencer, with a sequencing length of 150 bp and a sequencing depth of approximately 1000x.
[0068] In this embodiment, a sample set based on the fragment features of whole genome sequencing data and the methylation features of methylation targeted sequencing data was established. The lung cancer early screening prediction model based on multi-feature cross-stacking was trained and verified through the sample set to obtain the final lung cancer early screening prediction model and prediction results; the lung cancer early screening prediction model has the advantages of non-invasive detection and high prediction accuracy; compared with traditional clinical detection methods, it has better detection sensitivity and specificity, can perform real-time dynamic monitoring, and has clinical application value.
[0069] In this embodiment, the S30 performs feature extraction on the low-depth whole genome sequencing data and the methylation targeted sequencing data to obtain fragment features of the whole genome sequencing data, methylation features of the methylation targeted sequencing data, and fragment group features of the methylation data; including:
[0070] S301, preprocessing low-depth whole-genome sequencing data and methylation targeted sequencing data respectively;
[0071] S302, performing feature extraction on the pre-processed low-depth whole genome sequencing data to obtain fragment features of the whole genome sequencing data;
[0072] S303 , performing feature extraction on the pre-processed methylation targeted sequencing data to obtain methylation features of the methylation targeted sequencing data and fragment group features of the methylation data.
[0073] Specifically, in S301, the methylation characteristics are: the proportion of abnormally high methylation fragments, the length ratio of methylated fragments, and the proportion of 4bp motif ends of methylated fragments.
[0074] Furthermore, in S301, the low-depth whole genome sequencing data is preprocessed, including:
[0075] S301-11, use fastp software to perform quality control and remove adapter sequences on the raw data after sequencing;
[0076] S301-12, use BWA-MEM to align the fastq data obtained in the above step S301-11 with the hg19 human reference genome to obtain the aligned bam file, use samblaster software to mark the repeated sequences in the bam file, and use samtool software to remove the repeated and low-quality reads in the bam file.
[0077] Since low-quality data can affect subsequent analysis results, removing low-quality reads can improve the accuracy of data analysis. In this embodiment, low-quality reads refer to sequencing reads with a quality value lower than 30 during alignment, which is a common industry standard.
[0078] S301-13, using public database information, removing reads from the hg19 genome blacklist regions, spacer regions, patch sequences, highly variable regions, and centromere-telomeric regions in the bam file obtained in step S301-12;
[0079] S301-14, using the aligned position information of the paired-end sequencing data, merging the sequencing data into a single cfDNA fragment, and removing fragments longer than 400 bases from the merged fragments;
[0080] Since different GC contents of fragments may lead to coverage deviation during the experimental PCR amplification process, it is necessary to perform GC correction on the fragment data obtained in step S301-14;
[0081] S301-15, GC correction: including:
[0082] S301-151, establish a baseline: select multiple (in a specific embodiment, 29) healthy individuals. For each healthy sample, select fragments between 50 and 400 bp in length from the files obtained in steps S301-11 to S301-14, and calculate the GC content of each fragment (GC content is the percentage of GC bases in the fragment). Stratify the GC content by 1% intervals, and count the number of fragments corresponding to each GC content stratification for each chromosome (excluding chromosomes X and Y).
[0083] S301-152, after obtaining the fragment counts of different chromosomes and different GC contents of 29 healthy samples through the above baseline construction process, the median of the 29 distributions was further calculated, i.e., the GC bias-corrected reference distribution;
[0084] S301-153, for the actual test sample, the file obtained after processing in steps S301-11 to S301-14, which includes: the chromosome number of each fragment in the hg19 human reference genome, the fragment start position, the fragment end position, the fragment length, and the GC content information;
[0085] According to the number of fragments (N_i) corresponding to each GC content of each test sample counted in S301-151 above, and based on the C bias correction reference distribution (N_n) of healthy samples obtained in S301-152 above, the weight value (G_i=N_n / N_i) at each GC content is calculated; therefore, corresponding weight values can be assigned to fragments with different GC contents, and the weight values are used to weight each fragment according to its GC content to complete GC correction.
[0086] The pre-processed low-depth whole-genome sequencing data is sequencing fragment information, including: the chromosome number of each fragment in the hg19 human reference genome, the fragment start position, the fragment end position, the fragment length, the GC content and the corrected weight value;
[0087] The fragment features of the whole-genome sequencing data include: whole-genome copy number variation, cfDNA long-short fragment ratio, cfDNA fragment size distribution, cfDNA nucleosome pattern and cfDNA 4bp motif end ratio.
[0088] Furthermore, the S302, performing feature extraction on the pre-processed low-depth whole genome sequencing data to obtain segment features of the whole genome sequencing data, includes:
[0089] S302-1, feature extraction of cfDNA fragment length ratio (Fragmentsizeratio, FSR); including:
[0090] The hg19 reference genome was split into 5Mb windows, the GC content of each window was calculated, and then windows with a GC content less than 0.3 and a sequence alignment ratio less than 0.8 were filtered out;
[0091] After the above steps, multiple retained windows are obtained; each window is represented by a chromosome number, a starting position, and an ending position;
[0092] According to the position information of the obtained test sample fragments and the position information of the multiple retained windows, the fragments located in each 5Mb window interval can be counted respectively, and based on the length information of these fragments, short fragments with a length of 100-150bp and long fragments with a length of 151-210bp can be obtained respectively;
[0093] The calculation expression for extracting the feature length-short segment ratio R is:
[0094]
[0095] in: For all long fragments of cfDNA with a length of 151-210bp, For all short fragments of cfDNA with a length of 100-150bp, is the weight value corresponding to the GC content of the chromosome where segment i is located.
[0096] S302-2, feature extraction of cfDNA fragment size distribution (FSD); including:
[0097] Screen the fragments in the range of 50-400 bp in each test sample and divide them into 70 length intervals (e.g., 50-55 bp, 56-60 bp, ..., 396-400 bp) according to the 5 bp interval.
[0098] The segments were stratified according to 70 length intervals within the autosomal long and short arm intervals, and the FSD value within each chromosome arm interval was calculated according to the following formula, ultimately obtaining multiple features:
[0099]
[0100] in, represents the FSD score of the i-th chromosome arm interval; represents the total number of segments covered in the i-th chromosome arm interval; j represents the j-th segment covered in the i-th interval; Indicates the weight value corresponding to the GC ratio of the jth fragment.
[0101] S302-3, feature extraction of the proportion of 4bp motif ends in cfDNA; including:
[0102] In each fragment of each sample, the 4bp base sequences upstream and downstream of the start and end positions were extracted respectively, and the proportion of 256 combinations of 4mer sequences at the end of the fragment after GC correction was counted as the corrected terminal motif feature.
[0103] S302-4, feature extraction of cfDNA nucleosome patterns (NPs); including:
[0104] First, transcription factor-related files were downloaded from the GTRD database and the CIS-BS database, and the common transcription factors present in both databases were retained; transcription factors with less than 10,000 binding sites on autosomes were removed to obtain multiple transcription factors for subsequent analysis.
[0105] Secondly, the central area of each binding site of each transcription factor on the genome is used as the window interval, and the coverage fragments of all relative sites relative to the center position are calculated. The GC weight values corresponding to the fragments are accumulated as the coverage of the relative site, and the coverage of the central site window is smoothed and denoised to finally obtain the coverage distribution waveform of each transcription factor fragment.
[0106] Next, extract three features from the coverage pattern curve and construct a feature matrix, including:
[0107] 1) Coverage of the central region of the transcription factor binding site, where the central region is the range from -30 bp to +30 bp of the transcription factor binding site;
[0108] 2) Average coverage of the region from -1kb to +1kb of the transcription factor binding site;
[0109] 3) Amplitude values calculated using Fast Fourier Transform.
[0110] S302-5, feature extraction of whole-genome copy number variation, including:
[0111] The autosomes of the reference genome are divided into non-overlapping windows of 5 Mb in length (in a specific embodiment, a total of 589 non-overlapping windows), the average sequencing depth of each window is calculated, the overall mean is used for normalization, and the sequencing depth is corrected according to the GC content of each window;
[0112] Construct a baseline, select multiple (in a specific embodiment, 20) healthy individuals, perform the analysis in the previous step on each of them, and further calculate the median of each window depth for multiple healthy samples to filter the window dispersion in the baseline samples;
[0113] For each test sample, the above analysis can obtain the GC-corrected depth value of each window; then, based on the logarithmic value of the sequencing depth of each window of the baseline sample, the logarithmic value of the copy number change of each window of each test sample is calculated, that is, log2 (sequencing depth of the test sample divided by the average sequencing depth of the baseline sample), thereby obtaining the copy number variation characteristics of each test sample; (in a specific embodiment, 485 copy number variation characteristics).
[0114] In this embodiment, in S301, the methylation targeted sequencing data is preprocessed, including:
[0115] S301-21, use fastp software to perform quality control on the methylation targeted sequencing data and remove sequencing adapters to obtain fastq data;
[0116] S301-22, align the fastq data with the hg19 human reference genome to obtain the aligned bam file;
[0117] S301-23, remove PCR duplications in the bam file and create an index for the bam file after duplication removal;
[0118] S301-24, stitch and connect the CpG sites in the paired test sequences in the bam file to obtain the methylation status haplotype of each fragment.
[0119] Stitching refers to merging paired sequenced sequences into one. During merging, the two sequences may have overlapping regions, no overlapping regions, or inconsistent overlapping regions. For overlapping regions, this embodiment handles the following: in uncovered regions, CpG sites are replaced with U, methylated sites are marked as 1, and unmethylated sites are marked as 0. Read1 and Read2 simultaneously cover the sites. If the bases are the same, the highest base and its quality value are retained. If the bases are different and both bases have a quality value ≥30, the base is recorded as N. If one base is ≥30, the base and its quality value ≥30 are retained. If both bases are <30, the base is recorded as N, and the haplotype methylation string is finally generated.
[0120] In this embodiment, the S303 of performing feature extraction on the pre-processed methylation targeted sequencing data to obtain methylation features of the methylation targeted sequencing data and fragment group features of the methylation data includes:
[0121] S303-1, determining tumor-specific methylation intervals by comparing differences in hypermethylated fragments; for the sample to be tested, within the tumor-specific methylation interval, counting fragments in which methylated C bases account for more than half of the total CpG sites or in which the total number of methylated C sites is greater than or equal to 5, and defining them as abnormal fragments; dividing the number of abnormal fragments by the total number of fragments covering the interval to obtain the proportion of hypermethylated abnormal fragments;
[0122] S303-2: Obtain cfDNA fragments within all panel intervals, extract 4 bp base sequences upstream and downstream of the start and end positions, calculate the fragment ratios of 256 4 bp base end sequences, and obtain the proportion of methylated fragments with 4 bp motif ends;
[0123] S303-3: Count the cfDNA fragments within each tumor-specific methylation interval, and based on the length information of the cfDNA fragments, obtain the number of short fragments with a length of 100-150 bp and the number of long fragments with a length of 151-210 bp, and obtain the length-to-short ratio of the methylation fragment group within the interval.
[0124] It should be noted that in this embodiment, based on the combination of multi-omics and multi-features, it is possible to fully utilize the tumor signals of various dimensions in cfDNA fragments, and has higher sensitivity and specificity than traditional clinical detection methods and single-omics models.
[0125] Example 2
[0126] like Figure 2 As shown, based on the first embodiment, in this embodiment, the S40 is to construct a lung cancer early screening prediction model based on multi-feature cross-stack; including:
[0127] S401, establishing a single feature prediction model;
[0128] Including: using machine learning models to establish a classification model based on the characteristics of multiple sequencing data of samples;
[0129] S402, establishing an integrated model of a single feature;
[0130] This includes: concatenating multiple machine learning scores for each feature into a new feature vector, and then building an integrated model of the single feature based on logistic regression;
[0131] S403, establishing a multi-feature joint integration model;
[0132] It includes: splicing multiple feature scores of each sample into a new feature vector, and then establishing a multi-feature joint integration model based on logistic regression.
[0133] Specifically, the machine learning model includes at least one of a gradient boosting machine (LGBM) model, an XGBoost model, a random forest model (RF), a logistic regression (LR) model, and a multi-layer perceptron (MLPC).
[0134] In this embodiment, the step S50 of training and validating the lung cancer early screening prediction model using the sample set to obtain the final lung cancer early screening prediction model and prediction results may include:
[0135] A 5-fold cross-validation validation strategy is adopted to train the single-feature prediction model using the divided training set to obtain a trained single-feature prediction model. In the trained single-feature prediction model, different types of features correspond to corresponding optimal machine learning models.
[0136] In this example, a lung cancer early screening prediction model based on multi-feature cross-stacking was constructed based on the fragment features of six types of whole-genome sequencing data in cfDNA of lung cancer and healthy subjects and the methylation features of methylation targeted sequencing data to improve the accuracy, compliance and accessibility of lung cancer screening.
[0137] An embodiment of the present application further provides an electronic device, including:
[0138] memory; processor; and computer program;
[0139] The computer program is stored in the memory and configured to be executed by the processor to implement the method described above.
[0140] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored; the computer program is executed by a processor to implement the method described above.
[0141] In the embodiments of the present application, the method, the electronic device and the computer-readable storage medium are based on the same inventive concept. Since the principles of solving problems by the method, the electronic device and the computer-readable storage medium are similar, the implementation of the method, the electronic device and the computer-readable storage medium can refer to each other, and the repeated parts will not be repeated.
[0142] like Figure 5 、 6 As shown, in a specific embodiment, the performance of the model is verified.
[0143] This study included 326 healthy and 326 tumor samples, of which 58.5% were stage I lung cancer and 6.6% were stage II lung cancer. The sample list is as follows: Figure 6 shown.
[0144] Using the model established in Example 1, the classification threshold was set based on the specificity of 0.95 in the test set, and the sensitivity and specificity in the validation set were calculated. In the validation set, the LP-WGS single-omics detection AUC was 0.9439 (95% CI: 0.9148-0.973), sensitivity was 0.687 (95% CI: 0.6076-0.7664), and specificity was 0.9847 (95% CI: 0.9637-1); the methylation detection AUC was 0.948 (95% CI: 0.9199-0.976), sensitivity was 0.7481 (95% CI: 0.6738-0.8224), and specificity was 0.9313 (95% CI: 0.888-0.9746); the performance of the combined model was AUC 0.9617 (95%CI:0.9376-0.9857);Sensitivity 0.7634(95%CI: 0.6906-0.8362);0.9542(95%CI: 0.9184-0.99) Compared with the single-omics performance, the performance is improved. For details, see Figure 6 result.
[0145] This shows that the model of this application has a good ability to predict early lung cancer.
[0146] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application may be implemented in various computer languages, such as C, VHDL, Verilog, object-oriented programming language Java, and interpreted scripting language JavaScript.
[0147] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0148] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0150] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0151] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation, characterized in that: The following steps are involved: S10, collect peripheral blood from lung cancer patients and healthy subjects, extract cfDNA from the peripheral blood, and establish a sample set; S20, based on cfDNA, performs low-depth whole-genome and methylation-targeted sequencing to construct sequencing libraries; S30, performing feature extraction on the low-depth whole-genome sequencing data and the methylation targeted sequencing data to obtain fragment features of the whole-genome sequencing data, methylation features of the methylation targeted sequencing data, and fragment group features of the methylation data; S40: Based on the fragment features of whole-genome sequencing data and the methylation features of methylation targeted sequencing data, a multi-feature cross-stacked prediction model for early screening of lung cancer was constructed; S50, training and validating the lung cancer early screening prediction model using the sample set to obtain the final lung cancer early screening prediction model and prediction results; The S30, performing feature extraction on the low-depth whole genome sequencing data and the methylation targeted sequencing data to obtain fragment features of the whole genome sequencing data, methylation features of the methylation targeted sequencing data, and fragment group features of the methylation data, includes: S301, preprocessing low-depth whole-genome sequencing data and methylation targeted sequencing data respectively; S302, performing feature extraction on the pre-processed low-depth whole genome sequencing data to obtain fragment features of the whole genome sequencing data; S303, performing feature extraction on the pre-processed methylation targeted sequencing data to obtain methylation features of the methylation targeted sequencing data and fragment group features of the methylation data; The pre-processed low-depth whole-genome sequencing data is sequencing fragment information, including: the chromosome number of each fragment in the hg19 human reference genome, the fragment start position, the fragment end position, the fragment length, the GC content and the corrected weight value; The fragment features of the whole-genome sequencing data include: whole-genome copy number variation, cfDNA long-short fragment ratio, cfDNA fragment size distribution, cfDNA nucleosome pattern and cfDNA 4bp motif end ratio.
2. The method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation according to claim 1, characterized in that: The S40 is to construct a lung cancer early screening prediction model based on multi-feature cross-stack; including: S401, establishing a single feature prediction model; Including: using machine learning models to establish a classification model based on the characteristics of multiple sequencing data of samples; S402, establishing an integrated model of a single feature; This includes: concatenating multiple machine learning scores for each feature into a new feature vector, and then building an integrated model of the single feature based on logistic regression; S403, establishing a multi-feature joint integration model; It includes: splicing multiple feature scores of each sample into a new feature vector, and then establishing a multi-feature joint integration model based on logistic regression.
3. The method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation according to claim 2, characterized in that: The machine learning model includes: at least one of a gradient boosting model, an XGBoost model, a random forest model, a logistic regression model and a multi-layer perceptron.
4. The method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation according to claim 3, characterized in that: In S301, the methylation characteristics are: the proportion of abnormally high methylation fragments, the length ratio of methylated fragments, and the proportion of 4 bp motif ends of methylated fragments.
5. The method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation according to claim 4, characterized in that: In S301, the low-depth whole-genome sequencing data in the sequencing library is preprocessed, including: S301-11, use fastp software to perform quality control and remove adapter sequences on the raw data after sequencing; S301-12, using BWA-MEM, align the fastq data obtained in step S301-11 with the hg19 human reference genome to obtain a bam file after alignment; S301-13, using public database information, removing reads from the hg19 genome blacklist regions, spacer regions, patch sequences, highly variable regions, and centromere-telomeric regions in the bam file obtained in step S301-12; S301-14, using the aligned position information of the paired-end sequencing data, merging the sequencing data into a single cfDNA fragment, and removing fragments longer than 400 bases from the merged fragments; S301-15, perform GC correction.
6. The method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation according to claim 4, characterized in that: In S301, preprocessing the methylation targeted sequencing data includes: S301-21, use fastp software to perform quality control on the methylation targeted sequencing data and remove sequencing adapters to obtain fastq data; S301-22, align the fastq data with the hg19 human reference genome to obtain the aligned bam file; S301-23, remove PCR duplications in the bam file and create an index for the bam file after duplication removal; S301-24, stitch and connect the CpG sites in the paired test sequences in the bam file to obtain the methylation status haplotype of each fragment.
7. The method for constructing a lung cancer early screening model based on LP-WGS and DNA methylation according to claim 4, characterized in that: The step S303 of extracting features from the pre-processed methylation targeted sequencing data to obtain methylation features of the methylation targeted sequencing data and fragment group features of the methylation data includes: S303-1, determining tumor-specific methylation intervals by comparing differences in hypermethylated fragments; for the sample to be tested, within the tumor-specific methylation interval, counting fragments in which methylated C bases account for more than half of the total CpG sites or in which the total number of methylated C sites is greater than or equal to 5, and defining them as abnormal fragments; dividing the number of abnormal fragments by the total number of fragments covering the interval to obtain the proportion of hypermethylated abnormal fragments; S303-2: Obtain cfDNA fragments within all panel intervals, extract 4 bp base sequences upstream and downstream of the start and end positions, calculate the fragment ratios of 256 4 bp base end sequences, and obtain the proportion of methylated fragments with 4 bp motif ends; S303-3: Count the cfDNA fragments within each tumor-specific methylation interval, and based on the length information of the cfDNA fragments, obtain the number of short fragments with a length of 100-150 bp and the number of long fragments with a length of 151-210 bp, and obtain the length-to-short ratio of the methylation fragment group within the interval.
8. An electronic device comprising: Memory; processor; and computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Application of plasma free DNA methylation marker in early screening of lung cancer and early screening device for lung cancer
CN114736968A
Construction of early liver cancer detection model containing multi-omics marker composition and kit
CN115851951A