Method for screening cancer specific cfDNA fragment feature combination in peripheral blood based on whole genome sequencing and application thereof
Through whole genome sequencing, screening multi-dimensional cfDNA fragment characteristics and combining machine learning algorithms, a cancer-specific cfDNA fragment feature combination was solved, solving the problems of high sequencing depth and single feature dimensions in the existing technology, and achieving high sensitivity and high accuracy early cancer diagnosis.
Patent Information
- Application Number
- CN202510249892.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-04
AI Technical Summary
In the early diagnosis of cancer based on cfDNA fragment characteristics, the prior art has the problem of high sequencing depth requirements, insufficient sensitivity and accuracy, and the existing feature dimensions are single, making it difficult to effectively identify early cancer signals.
Through whole genome sequencing, the characteristics of cfDNA fragments in multiple dimensions were screened, including the global length distribution index, the 5’-terminal 6bp motif frequency, the coverage and length distribution index of the genome 10kb window, and combined with machine learning algorithms such as random forest, support vector machine, gradient hoist and LASSO, a cancer-specific cfDNA fragment feature combination was constructed.
It improves the sensitivity and accuracy of early cancer screening, and provides a series of cancer-specific molecular characteristics combinations, suitable for early cancer risk assessment and prediction, with excellent predictive performance.
Smart Images

Figure CN120279994A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene diagnosis, and particularly, to a method for screening a combination of cancer-specific cfDNA fragment features in peripheral blood based on whole-genome sequencing and its application. Background Art
[0002] Cancer has gradually become one of the top killers threatening human health globally, with its incidence and mortality rates continuously rising. "Early detection and early treatment" is crucial for improving the survival rate of cancer patients. However, since many types of cancer show no obvious symptoms in the early stage and are difficult to detect, the development of efficient and accurate early diagnosis tools has become an urgent need.
[0003] With the development of genomics technology, new strategies for early cancer diagnosis have been provided by identifying trace tumor molecular signals in peripheral blood through high-throughput sequencing technology or signal capture amplification methods. Cell-free DNA (cfDNA) in peripheral blood is free DNA released into the blood from apoptotic or necrotic cells, containing characteristics of tissue-derived cells, including DNA methylation status, copy number variation, end sequencing pattern, coverage area, etc. In healthy individuals, the vast majority of cfDNA molecules are derived from blood cells. When an organ or tissue undergoes a lesion, it is usually accompanied by more active apoptosis and more cfDNA molecules released from tissue cells into peripheral blood. Existing studies have found that there are significant differences in the fragmentation characteristics of cfDNA between cancer patients and the normal population, so it can be used as an important cancer biomarker. By detecting specific molecular signals of peripheral blood cfDNA, such as tumor driver mutations, early cancer signals can be identified. The tumor early diagnosis technology based on cfDNA shows good application prospects.
[0004] Currently, the technical strategy based on cfDNA mutation detection is a cancer screening route with relatively high clinical acceptance. For example, Exact Science's CancerSEEK and Guardant Health's TEC-seq products identify cancer signals by detecting tumor-specific gene mutations in peripheral blood. However, since the level of ctDNA molecules derived from tumors in peripheral blood is extremely low, and not all ctDNA fragments carry mutations, this strategy often requires extremely high sequencing depths, posing a huge challenge to the screening sensitivity. In contrast, characteristics such as the fragment length, nucleosome footprint, and end motif pattern of cfDNA are highly correlated with tissue source types and pathological states, with higher signal abundances, enabling cancer screening based on low-depth sequencing. The fragmentomics characteristics of cfDNA are cancer molecular markers that have received much attention in recent years, including the length distribution of cfDNA fragments, end motif frequencies, fragment breakpoint coordinates, nucleosome footprints, and different topological structures. The DE LFI team first demonstrated the application prospect of cfDNA fragmentomics in early cancer screening in a Nature article published in 2019, achieving a detection sensitivity of 73% in 500 retrospective samples. However, in this study, DELF I only focused on the cfDNA fragment length distribution characteristics in a 5mb window on the genome, without incorporating information such as coverage and fragment end motif patterns, resulting in a relatively single feature dimension. Secondly, the 5mb genomic window has a large span, often covering the coding and non-coding regions of multiple genes, and there may be significant differences between different genes, transcriptional regulatory regions, and CDS regions. Using the average value of the entire 5mb window may lose specific tumor-related molecular features. Therefore, to improve the sensitivity and accuracy of early cancer screening, new cfDNA fragment features still need to be developed. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a method for screening a combination of cancer-specific cfDNA fragment features in peripheral blood based on whole-genome sequencing and its application.
[0006] One object of the present invention is to provide a method for screening a combination of cancer-specific cfDNA fragment features in peripheral blood based on whole-genome sequencing.
[0007] Another object of the present invention is to provide the application of the above method in the preparation of products for cancer prediction.
[0008] Still another object of the present invention is to provide a method for constructing a cancer prediction model.
[0009] To achieve the above objects, the present invention is realized through the following solutions:
[0010] The present invention conducts low-depth whole-genome sequencing on peripheral blood cfDNA, extracts and integrates cfDNA fragment eigenvalue in multiple dimensions, and uses a variety of machine learning algorithms including Random Forest, Support Vector Machine (SVM), Gradient Boosting Machine (GBM) and regression algorithm LASSO to screen cfDNA features highly related to cancer and construct models, providing a biomarker combination and method for early cancer screening and cancer prediction.
[0011] A method for screening a combination of cancer-specific cfDNA fragment features in peripheral blood based on whole-genome sequencing, comprising the following steps:
[0012] S1. Obtain paired-end sequencing data of peripheral blood cfDNA in a sample group, retain the paired-end reads aligned to the human reference genome, and obtain cfDNA fragment information based on the paired-end reads. Each pair of paired-end reads corresponds to a cfDNA fragment, and the cfDNA fragment information includes but is not limited to the position of the cfDNA in the genome, the nucleotide sequence of the cfDNA, and the length of the cfDNA; the sample group includes cancer patients and healthy individuals;
[0013] S2. Use the paired-end reads obtained in step S1 to obtain the eigenvalues of the following cfDNA fragment features:
[0014] The global length distribution index all-F-Index of the cfDNA fragment, the 5'-terminal 6bp motif frequency 6bp-f of the cfDNA fragment, the cfDNA fragment coverage Coverage_bin in a 10kb window of the genome, and the length distribution index F-Index_bin of the cfDNA fragment in a 10kb window of the genome;
[0015] Among them, the eigenvalue all-F-Index of the global length distribution index of the cfDNA fragment is calculated by formula (1);
[0016] Formula (1):
[0017] where i represents the length of the cfDNA fragment; P is a constant value, 160bp - 170bp; X i represents the proportion of cfDNA fragments with length i among all cfDNA fragments;
[0018] The 6bp motif frequency 6bp-f at the 5' end of the cfDNA fragment is the occurrence frequency of each 6bp motif in the terminal motif set; the terminal motif set includes set 1 and set 2, where set 1 consists of the 6bp motifs at the 5' end of the nucleotide sequence of each cfDNA fragment; set 2 consists of the reverse complementary sequences of the 6bp motifs at the 3' end of the nucleotide sequence of each cfDNA fragment; the nucleotide sequence of the cfDNA fragment is obtained by aligning the paired-end reads to the human reference genome;
[0019] The genomic 10kb window is obtained by the following method: the genomic region after removing the redundant regions is divided into non-overlapping segments at 10kb intervals to obtain M genomic 10kb windows; the redundant regions include the short arm regions, telomere regions, and centromere regions of 5 chromosomes, and the 5 chromosomes are chromosome 13, chromosome 14, chromosome 15, chromosome 21, and chromosome 22;
[0020] The characteristic value Coverage_bin of the cfDNA fragment coverage of the genomic 10kb window is calculated from the genomic 10kb window and formula (2);
[0021] Formula (2): Coverage_bin k = N k / N total ;
[0022] where Coverage_bin k represents the coverage of the cfDNA fragments in the k-th 10kb-bins, 1 ≤ k ≤ M; N k represents the number of cfDNA fragments aligned to the k-th 10kb-bins; N total represents the number of all cfDNA fragments aligned to the reference genome based on the paired-end reads;
[0023] The length distribution index F-Index_bin of the cfDNA fragments in the genomic 10kb window is calculated from the genomic 10kb window and formula (3);
[0024] Formula (3):
[0025] where i represents the length of the cfDNA fragment; P is a constant value, 160bp - 170bp; X' i represents the proportion of cfDNA fragments with length i among all cfDNA fragments in the k-th genomic 10kb window;
[0026] S3. Generate a feature matrix using the cfDNA fragment feature values obtained in step S2, and perform feature screening to obtain a feature combination; the feature screening is performed based on the accuracy of cancer prediction, and the cancer prediction includes predicting whether the sample is a cancer patient.
[0027] Preferably, in step S1, the number of people in the sample group is not less than 100.
[0028] More preferably, in step S1, among the sample group, there are not less than 50 cancer patients and not less than 50 healthy people.
[0029] Preferably, in step S1, the version of the human reference genome is hg38.
[0030] Preferably, in step S1, retain the paired-end reads mapped to chromosomes 1 to 22 of the human reference genome.
[0031] Preferably, in step S2, the feature values of the cfDNA fragment features further include: the 5'-end 6bp motif short-6bp-f of cfDNA fragments with a length less than 150bp, which is the occurrence frequency of each 6bp motif in the short-terminal motif set; the short-terminal motif set is obtained by retaining the subset corresponding to all cfDNA fragments with a length less than 150bp in the terminal motif set.
[0032] Preferably, in step S2, P is a constant value of 166bp.
[0033] Preferably, in step S3, the method of feature screening includes the following steps: using the feature matrix as input features and performing feature selection by the wrapper method.
[0034] More preferably, the feature selection method used by the wrapper method includes recursive feature elimination.
[0035] More preferably, the model construction methods used by the wrapper method include random forest, support vector machine, gradient boosting machine or LASSO regression.
[0036] Preferably, the cancer is lung cancer or liver cancer.
[0037] Use of any of the above methods in the preparation of a product for cancer prediction.
[0038] A system for screening cancer-specific cfDNA fragment feature combinations in peripheral blood based on whole-genome sequencing, including a data acquisition module, an analysis module and an output module;
[0039] The data acquisition module is used to obtain paired-end sequencing data of peripheral blood cfDNA of a sample group;
[0040] The analysis module is used to obtain a cancer-specific cfDNA fragment feature combination based on the paired-end sequencing data; the method for obtaining the cancer-specific cfDNA fragment feature combination includes the following steps:
[0041] S1. Retain the paired reads in the paired-end sequencing data that are aligned to the human reference genome, and each cfDNA fragment matches a pair of paired reads; the sample group includes cancer patients and healthy individuals;
[0042] S2. Use the paired reads obtained in step S1 to obtain the feature values of the following cfDNA fragment features:
[0043] The global length distribution index all-F-Index of the cfDNA fragment, the 5'-end 6bp motif frequency 6bp-f of the cfDNA fragment, the cfDNA fragment coverage Coverage_biin in the 10kb window of the genome, and the length distribution index F-Index_bin of the cfDNA fragment in the 10kb window of the genome;
[0044] Among them, the feature value all-F-Index of the global length distribution index of the cfDNA fragment is calculated by formula (1);
[0045] Formula (1):
[0046] Among them, i represents the length of the cfDNA fragment; P is a constant value, 160bp to 170bp; X i represents the proportion of cfDNA fragments with length i among all cfDNA fragments;
[0047] The 5'-end 6bp motif frequency 6bp-f of the cfDNA fragment is the occurrence frequency of each 6bp motif in the terminal motif set; the terminal motif set includes set 1 and set 2, and set 1 is composed of the 5'-end 6bp motifs of the nucleotide sequences of each cfDNA fragment; set 2 is composed of the reverse complementary sequences of the 3'-end 6bp motifs of the nucleotide sequences of each cfDNA fragment; the nucleotide sequence of the cfDNA fragment is obtained by aligning the paired reads to the human reference genome;
[0048] The 10kb window of the genome is obtained by the following method: The genomic region after removing the redundant regions is divided into non-overlapping segments according to 10kb intervals to obtain M 10kb windows of the genome; the redundant regions include the short arm regions, telomere regions, and centromere regions of 5 chromosomes, and the 5 chromosomes are chromosome 13, chromosome 14, chromosome 15, chromosome 21, and chromosome 22;
[0049] The eigenvalue Coverage_bin of the cfDNA fragment coverage in the 10kb window of the genome is calculated from the 10kb window of the genome and formula (2);
[0050] Formula (2): Coverage_bin k = N k / N total ;
[0051] Where Coverage_bin k represents the coverage of cfDNA fragments in the k-th 10kb-bins, 1 ≤ k ≤ M; N k represents the number of cfDNA fragments aligned to the k-th 10kb-bins; N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads;
[0052] The length distribution index F-Index_bin of the cfDNA fragments in the 10kb window of the genome is calculated from the 10kb window of the genome and formula (3);
[0053] Formula (3):
[0054] Where i represents the length of the cfDNA fragment; P is a constant value, 160bp - 170bp; X' i represents the proportion of cfDNA fragments of length i among all cfDNA fragments in the k-th 10kb window of the genome;
[0055] S3. Generate a feature matrix using the cfDNA fragment eigenvalues obtained in step S2, and perform feature screening to obtain a feature combination; the feature screening is performed based on the accuracy of cancer prediction, and the cancer prediction includes predicting whether a sample is a cancer patient;
[0056] The output module is used to output the cancer-specific cfDNA fragment feature combination obtained by the analysis module.
[0057] Preferably, the system further includes a storage module for storing the paired-end sequencing data of the peripheral blood cfDNA of the sample group obtained by the data acquisition module and the cancer-specific cfDNA fragment feature combination obtained by the analysis module.
[0058] Preferably, in step S1, the number of people in the sample group is not less than 100.
[0059] More preferably, in step S1, among the sample group, there are not less than 50 cancer patients and not less than 50 healthy people.
[0060] Preferably, in step S1, the version of the human reference genome is hg38.
[0061] Preferably, in step S1, the paired-end reads aligned to chromosomes 1 to 22 of the human reference genome are retained.
[0062] Preferably, in step S2, P is a constant value of 166 bp.
[0063] Preferably, in step S2, the eigenvalue of the cfDNA fragment feature further includes: the 5'-terminal 6-bp motif short-6bp-f of cfDNA fragments with a length less than 150 bp, which is the occurrence frequency of each 6-bp motif in the short-terminal motif set; the short-terminal motif set is obtained by retaining the subset corresponding to all cfDNA fragments with a length less than 150 bp in the terminal motif set.
[0064] Preferably, in step S3, the method for feature screening includes the following steps: using the feature matrix as the input feature and performing feature selection by the wrapper method.
[0065] More preferably, the feature selection method adopted by the wrapper method includes recursive feature elimination.
[0066] More preferably, the model construction methods adopted by the wrapper method include random forest, support vector machine, gradient boosting machine, or LASSO regression.
[0067] Preferably, the cancer is lung cancer or liver cancer.
[0068] A computer device includes a memory and a processor, and a computer program executable on the processor is stored on the memory; when the computer program is executed by the processor, the operations of the system for screening cancer-specific cfDNA fragment features in peripheral blood based on whole-genome sequencing are implemented.
[0069] A computer-readable storage medium stores a computer program capable of being executed by a processor, and when the computer program is executed by the processor, the operations of the system for screening cancer-specific cfDNA fragment features in peripheral blood based on whole-genome sequencing are implemented.
[0070] A method for constructing a cancer prediction model processes a training set sample by any of the above methods to obtain a cancer-specific cfDNA fragment feature combination; using the cancer-specific cfDNA fragment feature combination as the input feature and the cancer risk probability as the output result, and using AUC as the performance index, training a classification model to obtain the cancer prediction model.
[0071] Preferably, the classification model includes random forest, support vector machine, gradient boosting machine or LASSO regression.
[0072] Compared with the prior art, the present invention has the following beneficial effects:
[0073] The present invention combines cfDNA fragment information screening of multiple dimensions to screen characteristic indicators closely related to cancer, and provides a new index F-Index for evaluating the characteristic of fragment length distribution, which has a significant difference between cancer samples and healthy samples. In addition, the present invention also applies the motif characteristics of the 6 bases at the ends of cfDNA fragments and the distribution characteristics of cfDNA fragments within a 10kb window on the genome, and applies multiple machine learning algorithms to identify a series of cancer-specific molecular characteristic combinations, which are suitable for risk assessment and prediction of early cancer. Based on the characteristic combinations screened by the method of the present invention, excellent prediction performance is achieved in different classification models, and the determination accuracy is high. Description of the Drawings
[0074] Figure 1 It is a schematic flowchart of a method for screening a cancer-specific cfDNA fragment characteristic combination in peripheral blood based on whole genome sequencing.
[0075] Figure 2 It is a statistical chart of all-F-Index of the global length distribution index of cfDNA fragments of lung cancer samples and healthy samples in Application Example 1, where Controls represents healthy samples and Lung Cancers represents lung cancer samples.
[0076] Figure 3 It is a 6bp-f statistical chart of the 5'-terminal 6bp motifs of 11 cfDNA fragments of lung cancer samples and healthy samples in Application Example 1, where Controls represents healthy samples and Lung Cancers represents lung cancer samples.
[0077] Figure 4 It is a statistical chart of Coverage_bin of cfDNA fragment coverage of 4 genomic 10kb windows and F-Index_bin of length distribution index of cfDNA fragments of 5 genomic 10kb windows of lung cancer samples and healthy samples in Application Example 1, where Controls represents healthy samples and Lung Cancers represents lung cancer samples.
[0078] Figure 5 It is an AUC chart of a lung cancer prediction model constructed using 4 machine learning algorithms with validation set samples in Application Example 1.
[0079] Figure 6The distribution of the lung cancer risk probabilities obtained from the lung cancer prediction model constructed using the GBM algorithm in Application Example 1 in the training set samples and the validation set samples. Training represents the training set samples, Validation represents the validation set samples, Controls represents the healthy samples, and Cancers represents the lung cancer samples.
[0080] Figure 7 The AUC graphs of the lung cancer prediction models constructed using different input features with the validation set samples in Application Example 1. Integration represents the multi-dimensional model constructed with the lung cancer-specific cfDNA fragment feature combination as the input feature, Motif represents the single-dimensional model constructed with the 5'-terminal 6bp motifs of 11 cfDNA fragments as the input feature, and 10k-Bins represents the two-dimensional model constructed with the cfDNA fragment coverage of 4 genomic 10kb windows and the length distribution index of cfDNA fragments in 5 genomic 10kb windows as the input features.
[0081] Figure 8 The all-F-Index statistical graph of the global length distribution index of cfDNA fragments in liver cancer samples and healthy samples in Application Example 2. Controls represents the healthy samples, and Liver Cancers represents the liver cancer samples.
[0082] Figure 9 The Short_6bp-f statistical graph of the terminal 6bp motifs of 10 cfDNA fragments with lengths less than 150bp in liver cancer samples and healthy samples in Application Example 2. Controls represents the healthy samples, and Liver Cancer s represents the liver cancer samples.
[0083] Figure 10 The F-Index_bin statistical graph of the length distribution index of cfDNA fragments in 5 genomic 10kb windows in liver cancer samples and healthy samples in Application Example 2. Controls represents the healthy samples, and LiverCancers represents the liver cancer samples.
[0084] Figure 11 The AUC graphs of the liver cancer prediction models constructed using 4 machine learning algorithms with the validation set samples in Application Example 2.
[0085] Figure 12 The distribution of the liver cancer risk probabilities obtained from the liver cancer prediction model constructed using the randomForest algorithm in Application Example 2 in the training set samples and the validation set samples. Training represents the training set samples, Va lidation represents the validation set samples, Controls represents the healthy samples, and Cancers represents the liver cancer samples.
[0086] Figure 13 It is the AUC graph of the liver cancer prediction models constructed using different input features with the validation set samples in Application Example 2. Integration represents the multi-dimensional model constructed with the liver cancer-specific cfDNA fragment feature combination as the input feature. Motif represents the single-dimensional model constructed with the 6bp motif at the 5' end of 10 cfDNA fragments with a length less than 150bp as the input feature. 10k-Bins represents the single-dimensional model constructed with the length distribution index of cfDNA fragments in 5 genomic 10kb windows as the input feature. Detailed implementation manners
[0087] The present invention will be further elaborated in detail below in conjunction with the accompanying drawings of the specification and specific embodiments. The embodiments are only used to explain the present invention and are not used to limit the scope of the present invention. The test methods used in the following embodiments are all conventional methods unless otherwise specified; the materials, reagents, etc. used are all reagents and materials that can be obtained from commercial channels unless otherwise specified.
[0088] Example 1 A method for screening a cancer-specific cfDNA fragment feature combination in peripheral blood based on whole-genome sequencing
[0089] This example provides a method for screening a cancer-specific cfDNA fragment feature combination in peripheral blood based on whole-genome sequencing, and its process is as Figure 1 shown. The following provides a specific implementation manner.
[0090] 1. Acquisition and processing of plasma cfDNA whole-genome sequencing data
[0091] In the present invention, cfDNA is extracted, library constructed, and paired-end sequenced from the peripheral blood samples of the subjects. There is no particular limitation on the methods for cfDNA extraction, library construction, and sequencing, and they can be adjusted from existing technical methods. The following provides a specific method:
[0092] (1) Plasma separation
[0093] The fresh peripheral blood sample is centrifuged at 3000×g for 10 minutes at 4°C within 4 hours after blood collection, and the lower-layer blood cells are discarded. Subsequently, it is further centrifuged at 16000×g for 10 minutes, and the supernatant plasma is transferred to a cryotube to obtain plasma, which is stored at -80°C before nucleic acid extraction.
[0094] (2) cfDNA extraction
[0095] The thawed plasma samples were centrifuged at 16,000×g for 10 minutes, and cfDNA was extracted using the circulating Nucleic Acid Kit. The DNA concentration was detected using a Qubit nucleic acid quantifier, and the nucleic acid samples were stored at -20°C before library construction and sequencing.
[0096] (3) Library construction and paired-end sequencing
[0097] The QIAseq cfDNA All-in-One Kit (24) (Cat.No. / ID: 180023) was used to construct the cfDNA library without fragmenting cfDNA during the library construction process. The LabChip GX Touch HT Nucleic Acid Analyzer instrument was used for fragment detection. The Illumina Novaseq sequencing platform was used to perform paired-end sequencing on the cfDNA library with a sequencing read length of ~150 bp, a sequencing depth adaptable to 1×~20×, and an average raw data output of 3G~60G per sample. After sequencing, the raw fastq sequencing data was obtained.
[0098] (4) Data quality control processing and alignment analysis
[0099] The Fastp software was used to perform quality control processing on the raw fastq data based on default parameters to remove adapters, data with base quality values below Q30, sequences with a high N content (i.e., reads sequences with an N-Content greater than 2%), and PCR-introduced duplicate sequences, resulting in paired-end reads. Each cfDNA fragment matched a pair of paired-end reads (read1 and read2).
[0100] The bwa software was used to align the quality-controlled paired-end reads to the human reference genome hg38 (Genome Reference Consortium human genome build 38, GRCh38) to obtain the Bam file of the alignment results. The picard tool was used to preprocess the Bam file to remove multiply aligned paired-end reads, paired-end reads that did not align to the reference genome, and retain paired-end reads with a quality value greater than 30, resulting in a clean bam file.
[0101] 2. Analysis of cfDNA fragment characteristics and their characteristic values
[0102] This example involves a total of 4 types of cfDNA fragment characteristics: the global length distribution index of cfDNA fragments, the 6-bp motif at the 5' end of cfDNA fragments, the cfDNA fragment coverage in a 10-kb genomic window, and the length distribution index of cfDNA fragments in a 10-kb genomic window.
[0103] (1) Calculation of the global length distribution index eigenvalue all-F-Index of cfDNA fragments
[0104] The global length distribution index all-F-Index of cfDNA fragments reflects the enrichment degree of short fragments in cfDNA fragments. The larger the all-F-Index, the higher the proportion of short fragments, indicating that the overall cfDNA fragment distribution tends to be shorter in length and there is an enrichment phenomenon of short fragments.
[0105] The calculation method of all-F-Index is as follows:
[0106] Extract the alignment coordinates, directions, lengths, and base sequence information of each pair of paired-end reads from the clean bam file obtained in the previous step. In this analysis step, only retain the paired-end reads aligned to the regions of autosomes 1 to 22, and remove the paired-end reads aligned to the X and Y chromosomes, mitochondrial genome, or other supplementary contig regions to obtain a bed file of cfDNA fragment distribution on the genome (i.e., the genomic bed file). This file contains cfDNA fragment information in the genome, including: cfDNA fragment sequence, length, and position information on the chromosome.
[0107] Calculate all-F-Index according to the genomic bed file and formula (1).
[0108] Formula (1):
[0109] Among them, i represents the length of the cfDNA fragment; P is a constant value of 166bp, which is the general length of cfDNA fragments in the plasma of healthy people; X i represents the proportion of cfDNA fragments with length i among all cfDNA fragments.
[0110] (2) Calculation of the eigenvalue of the 5'-terminal 6bp motif of cfDNA fragments
[0111] cfDNA is derived from apoptosis or active release of tissues and blood cells and is affected by nuclease digestion.
[0112] The eigenvalue of the 5'-terminal 6bp motif of cfDNA fragments is the 5'-terminal 6bp motif frequency (6bp-f) of cfDNA fragments. This value reflects that the terminal motif pattern of cfDNA fragments is affected by nuclease activity and tissue type. Its statistical method is as follows:
[0113] For each cfDNA fragment, each pair of paired - end reads consists of read1 on the forward strand and read2 on the reverse strand. By aligning to the reference genome, the forward complete sequence of the cfDNA fragment (i.e., the nucleotide sequence from the 5'-end to the 3'-end) is obtained. The 6 - bp sequence at the 5'-end is defined as End_motif_F, and the 6 - bp sequence at the 3'-end is defined as End_motif_R. After that, the sequence of End_motif_F is not transformed, and all End_motif_F of cfDNA fragments form set 1; the sequence of End_motif_R is transformed by reverse complement of bases to obtain the reverse - complementary sequence of End_motif_R, and all reverse - complementary sequences of End_motif_R of cfDNA fragments form set 2. After merging set 1 and set 2, the set of 6 - bp motifs at the 5'-ends of all cfDNA fragments is obtained, denoted as the terminal motif set. Theoretically, there are at most 4096 different 6 - bp motifs, such as {AAAAAA,AAAAAT,...,ACGACG,CTTGTT,...,GGGGGG}. By counting the occurrence frequency of each 6 - bp motif in the terminal motif set, 6bp - f is obtained.
[0114] (3) Calculation of the coverage characteristic value of cfDNA fragments in 10 - kb windows of the genome
[0115] The coverage of cfDNA fragments in 10 - kb windows of the genome is affected by nucleosome distribution and structure and is related to tissue types. Generally, the higher the chromatin openness within the genomic window, the lower the coverage of cfDNA fragments. The calculation method of Coverage_bin is as follows:
[0116] For human chromosomes 13, 14, 15, 21, and 22, since the short - arm regions of these 5 chromosomes contain nucleolus - organizing regions and are rich in a large number of repetitive ribosomal RNA genes, it is difficult for genome - sequencing technologies to accurately analyze the complete sequences therein. Therefore, based on the bed file of the cfDNA fragment distribution on the genome obtained in this example, the short - arm regions of these 5 chromosomes, as well as the telomere and centromere regions, are removed. The remaining genomic regions are non - overlappingly segmented into M genomic 10 - kb windows (i.e., 10kb - bins) according to 10 - kb intervals (such as chr1:20000~chr1:29999). Theoretically, at most 287465 10kb - bins can be obtained, that is, the maximum value of M is 287465.
[0117] Coverage_bin is calculated according to the genomic bed file, 10kb - bins, and formula (2).
[0118] Formula (2): Coverage_bin k =Nk / N total ;
[0119] where Coverage_bin k represents the coverage of cfDNA fragments in the k-th 10kb-bins, where 1 ≤ k ≤ M; N k represents the number of cfDNA fragments aligned to the k-th 10kb-bins. Among them, reads spanning two adjacent 10kb-bins are regarded as being aligned to both 10kb-bins simultaneously. When calculating coverage, both 10kb-bins on both sides increase by one count. When calculating F-Index_bin k later, this read will also be included in the calculation; N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads.
[0120] (4) Calculation of the length distribution index F-Index_bin of cfDNA fragments in 10kb windows of the genome
[0121] The length distribution index F-Index_bin of cfDNA fragments in 10kb windows of the genome is affected by both nucleosome structure and histone modification and is related to tissue type and pathological state. The calculation method of F-Index_bin is as follows:
[0122] Calculate the F-Index_bin of cfDNA fragments in the k-th 10kb-bins according to the genome bed file, 10kb-bins, and formula (3).
[0123] Formula (3):
[0124] where i represents the length of the cfDNA fragment; P is a constant value of 166bp, which is the general length of cfDNA fragments in the plasma of healthy people; X’ i represents the proportion of cfDNA fragments with length i among all cfDNA fragments in the k-th 10kb-bins.
[0125] Through the calculation and statistics of this embodiment, for each sample, 1 all-F-Index value, and at most 4096 6bp-f values, 287465 Coverage_bin values, and 287465 F-Index_bin values are obtained.
[0126] 3. Determination of cancer-specific cfDNA fragment characteristics combination
[0127] Taking cancer patients and healthy individuals as the sample groups, based on the cfDNA fragment characteristics and their eigenvalues of each sample obtained in the previous step, a feature matrix is generated. Each row represents a sample, and each column represents the cfDNA fragment characteristics and their eigenvalues. All eigenvalues are normalized and transformed into a normal distribution.
[0128] Using the transformed feature matrix as the input features, feature selection is performed based on a machine learning model. Specifically, a model constructed by applying the Recursive Feature Elimination (RFE) algorithm combined with different machine learning algorithms (such as Random Forest, Support Vector Machine (SVM), Gradient Boosting Machine (GBM), or the LASSO regression algorithm) is used for feature screening. When constructing the model, the model parameters are set to 10-fold repeated cross-validation, and the number of repeated sampling iterations is 10. The accuracy of predicting whether a patient has cancer is used as the result evaluation parameter, and the optimal number of features and feature combination are obtained and returned, which is the cancer-specific cfDNA fragment feature combination.
[0129] 4. Construction and performance evaluation of the cancer prediction model
[0130] Based on the obtained cancer-specific cfDNA fragment feature combination, further combined with the corresponding machine learning algorithm used when screening this feature combination, a cancer prediction model is trained. The model parameters are set to 10-fold repeated cross-validation and 10 repeated iterations, and the prediction performance of the model is evaluated.
[0131] Example 2: A method for screening cancer-specific cfDNA fragment feature combinations in peripheral blood based on whole-genome sequencing
[0132] 1. Acquisition and processing of plasma cfDNA whole-genome sequencing data
[0133] The same as Example 1.
[0134] 2. Analysis of cfDNA fragment characteristics and their eigenvalues
[0135] Basically the same as Example 1, except that in addition to the 4 types of cfDNA fragment characteristics, it also involves the 5th type of cfDNA fragment characteristic: the 5'-terminal 6bp motif of cfDNA fragments with a length less than 150bp. This indicator reflects both short-fragment enrichment and nuclease cleavage bias at the same time, and its eigenvalue is the frequency of the 5'-terminal 6bp motif of cfDNA fragments with a length less than 150bp (short-6bp-f). The statistical method is as follows:
[0136] For each cfDNA fragment with a length less than 150 bp, the 6-bp sequence at its 5' end is defined as Short_End_motif_F, and the 6-bp sequence at its 3' end is defined as Short_End_motif_R. A subset corresponding to all cfDNA fragments with a length less than 150 bp is extracted from the set of end motifs, denoted as the short-end motif set, which is obtained by combining set 1' composed of Short_End_motif_F and set 2' composed of the reverse complementary sequences of Short_End_motif_R. The occurrence frequency of each 6-bp motif in the short-end motif set is counted to obtain short-6bp-f.
[0137] Through the calculation and statistics of this embodiment, for each sample, 1 all-F-Index, up to 4096 6bp-fs, 4096 short-6bp-fs, 287465 Coverage_bins, and 287465 F-Index_bins are obtained.
[0138] 3. Determination of cancer-specific cfDNA fragment characteristics combination
[0139] The same as in Example 1.
[0140] 4. Construction and performance evaluation of cancer prediction model
[0141] The same as in Example 1.
[0142] Application Example 1 Lung cancer prediction method based on the characteristics combination of plasma cell-free DNA fragments
[0143] 1. Peripheral blood samples of subjects
[0144] A total of 398 peripheral blood samples of lung cancer patients (i.e., lung cancer samples) and 168 peripheral blood samples of healthy individuals (i.e., healthy samples) are included in this application example. Each peripheral blood sample comes from a different individual and is sourced from the sample library of Shenzhen Hypereal Biotechnology Co., Ltd. Informed consent from the subjects has been obtained before sample testing.
[0145] 2. Determination of lung cancer-specific cfDNA fragment characteristics combination
[0146] Taking all the peripheral blood samples of the subjects as the sample group, processing them according to the method in Example 1 to generate a genomic bed file as shown in Table 1, and then continuing to analyze according to the method in Example 1 to obtain all the cfDNA fragment characteristics and their characteristic values of each sample and generate a characteristic matrix. Normalization and normalization conversions are performed on all the characteristic values respectively. The all-F-Indices of the global length distribution indices of the cfDNA fragments of the lung cancer samples and the healthy samples are respectively summarized and compared and analyzed, asFigure 2 As shown, compared with healthy samples, the all-F-Index of lung cancer samples is significantly increased, indicating that this cfDNA fragment feature is significantly associated with lung cancer and can be used as a potential molecular diagnostic marker for lung cancer.
[0147] Table 1 Genomic bed file format (partial) generated by cfDNA sequencing data alignment
[0148]
[0149]
[0150] Subsequently, the subject samples were randomly divided into a training set of samples and a validation set of samples. The training set of samples included 239 lung cancer samples and 101 healthy samples, and the validation set of samples included the remaining 159 lung cancer samples and 67 healthy samples.
[0151] For the training set of samples, according to the method of Example 1, the RFE algorithm and each machine learning algorithm (RandomForest, GBM, SVM or LASSO) were used to construct models respectively, and feature screening was carried out to return the optimal number of features and feature combinations. Among them, the top 20 feature combinations with relative weights obtained by each machine learning algorithm are shown in Tables 2-5, which are the lung cancer-specific cfDNA fragment feature combinations.
[0152] The 6bp motif {AAGTGC} at the 5' end of 1 cfDNA fragment appears in the lung cancer-specific cfDNA fragment feature combinations obtained by 4 machine learning algorithms. The 6bp motifs {AAGGGT, ACCTCT, TTTAGG} at the 5' end of 3 cfDNA fragments appear in the lung cancer-specific cfDNA fragment feature combinations obtained by 3 machine learning algorithms (RandomForest, SVM and GBM).
[0153] After summary analysis, taking the lung cancer-specific cfDNA fragment feature combination obtained based on GBM as an example, as Figure 3 and Figure 4 shown, the 6bp motifs at the 5' end of 11 cfDNA fragments, the cfDNA fragment coverage of 4 genomic 10kb windows and the length distribution index of the cfDNA fragments of 5 genomic 10kb windows are significantly different between lung cancer patients and healthy samples. Similar differences also exist in the lung cancer-specific cfDNA fragment feature combinations obtained based on the other 3 machine learning algorithms between lung cancer patients and healthy samples. The above results indicate that the lung cancer-specific cfDNA fragment feature combinations obtained by the method provided by the present invention are potential molecular diagnostic markers for lung cancer.
[0154] Table 2 Top 20 Feature Combinations with Relative Weights Obtained Based on RandomForest
[0155]
[0156]
[0157] Table 3 Top 20 Feature Combinations with Relative Weights Obtained Based on GBM
[0158]
[0159] Table 4 Top 20 Feature Combinations with Relative Weights Obtained Based on SVM
[0160]
[0161]
[0162] Table 5 Top 20 Feature Combinations with Relative Weights Obtained Based on LASSO
[0163]
[0164] 3. Construction and Performance Evaluation of Lung Cancer Prediction Model
[0165] (1) Model Construction
[0166] For the training set samples of this application example, the lung cancer-specific cfDNA fragment feature combinations of each sample in the training set as shown in Table 2 are respectively obtained as the input features of the randomForest model, and the lung cancer risk probability is used as the output result. The output result range is [0, 1]. The greater the lung cancer risk probability, the higher the possibility that the sample is a lung cancer sample. Combining the actual status (lung cancer and non-lung cancer) of each sample in the training set, 10-fold repeated cross-validation and 10 repeated iterations are performed, and the optimal training model is output based on the AUC index performance. During the 10-fold cross-validation process, the cutoff value is 0.5, and the judgment criterion is: if the lung cancer risk probability ≥ 0.5, the sample is judged as positive (i.e., a lung cancer sample); if the lung cancer risk probability < 0.5, the sample is judged as negative (i.e., a non-lung cancer sample).
[0167] For the training set of this application example, the lung cancer-specific cfDNA fragment feature combinations of each sample in the training set as shown in Table 3 are respectively obtained as the input features of the GBM model, and the lung cancer prediction model is constructed according to the same method as above.
[0168] For the training set of this application example, the lung cancer-specific cfDNA fragment feature combinations of each sample in the training set as shown in Table 4 are respectively obtained as the input features of the svmLinear model, and a lung cancer prediction model is constructed according to the same method as above.
[0169] For the training set of this application example, the lung cancer-specific cfDNA fragment feature combinations of each sample in the training set as shown in Table 5 are respectively obtained as the input features of the LASSO model. According to the same method as above, and with alpha set to seq(0, 1, by = 0.05) and lambda set to 10^seq(-2, 2, length = 100), a lung cancer prediction model is constructed with AUC as the performance index.
[0170] (2) Performance evaluation
[0171] Use the validation set samples of this application example to evaluate the performance of the optimal training models respectively constructed by the four machine learning algorithms in step (1) of this application example. Calculate the prediction sensitivity Sensitivity, specificity Specificity, positive predictive value PPV (Positive Predictive Value), negative predictive value NPV (Negative Predictive Value), and comprehensive accuracy Accuracy index according to formulas (4) - (8).
[0172] Formula (4): Sensitivity = TP / (TP + FN);
[0173] Formula (5): Specificity = TN / (TN + FP);
[0174] Formula (6): PPV = TP / (TP + FP);
[0175] Formula (7): NPV = TN / (TN + FN);
[0176] Formula (8): Accuracy = (TP + TN) / (TP + FN + TN + FP).
[0177] Table 6 Performance evaluation results of the lung cancer prediction model
[0178] Model randomForest gbm svmLinear lasso Sensitivity 0.874 0.874 0.849 0.893 Specificity 0.896 0.94 0.91 0.91 PPV 0.952 0.972 0.957 0.959 NPV 0.75 0.759 0.718 0.782 AUC 0.953 0.961 0.948 0.959 Accuracy 0.881 0.894 0.867 0.782
[0179] As Figure 5As shown in Table 6, the lung cancer prediction models constructed by the four machine learning algorithms all have excellent classification performance and can accurately determine whether a sample is a lung cancer sample. Among them, the lung cancer prediction model constructed based on the GBM algorithm is the best, and the best AUC in 10-fold repeated cross-validation reaches 0.961. Further summarize the lung cancer risk probabilities obtained from the training set samples and validation set samples of this model. As Figure 6 shown, it can be seen that the obtained judgment results show significant discrimination in both the training set samples and the validation set samples.
[0180] In addition, use the 6bp motif at the 5' end of 11 cfDNA fragments in the lung cancer-specific cfDNA fragment feature combination obtained based on GBM as input features, or use the cfDNA fragment coverage of 4 genomic 10kb windows and the length distribution index of cfDNA fragments in 5 genomic 10kb windows in the lung cancer-specific cfDNA fragment feature combination as input features; use the lung cancer risk probability as the output result, and construct a lung cancer prediction model based on the training set samples and the gradient boosting machine GBM according to the same method, and use the validation set samples for performance evaluation. As Figure 7 shown in Table 7, the AUC of the single-dimensional model constructed based on the 6bp motif at the 5' end of cfDNA fragments is 0.94, and the AUC of the two-dimensional model constructed based on the cfDNA fragment coverage of genomic 10kb windows and the length distribution index of cfDNA fragments in genomic 10kb windows reaches 0.92. Although the prediction performance is slightly lower than that of the multi-dimensional model constructed using all lung cancer-specific cfDNA fragment features, it still shows an accuracy of more than 85% and can accurately determine lung cancer samples.
[0181] Table 7 Performance evaluation results of models with different dimensions
[0182]
[0183] Application Example 2 Hepatocellular carcinoma prediction method based on plasma-free DNA fragment feature combination
[0184] 1. Peripheral blood samples of subjects
[0185] A total of 121 peripheral blood samples of hepatocellular carcinoma patients (i.e., hepatocellular carcinoma samples) and 213 peripheral blood samples of healthy people (i.e., healthy samples) were included in this application example. Each peripheral blood sample was from a different individual and all were from the sample bank of Shenzhen Hypatos Biotechnology Co., Ltd. Informed consent from the subjects was obtained before sample testing.
[0186] 2. Determination of hepatocellular carcinoma-specific cfDNA fragment feature combination
[0187] Using the peripheral blood samples of all subjects as a sample group, processing them according to the method in Example 2 to generate the genomic bed file as shown in Table 8, and then continuing to analyze according to the method in Example 2 to obtain all cfDNA fragment features and their characteristic values of each sample and generate a feature matrix, and performing normalization and normal transformation on all characteristic values. Respectively summarize the all-F-Index of the global length distribution index of cfDNA fragments in liver cancer samples and healthy samples and conduct a comparative analysis. As Figure 8 shown, compared with healthy samples, the all-F-Index of liver cancer samples is significantly increased, indicating that this cfDNA fragment feature is significantly associated with liver cancer and can be used as a potential molecular diagnostic marker for liver cancer.
[0188] Table 8 Genomic bed file format generated by aligning partial cfDNA sequencing data
[0189] #chr Start End Name score strand End_motif_F End_motif_R gc length Chr1 60196 60567 A00251:497:H7TFGDSXC:3:1576:21233:32784 43 - ACATTG ATCATC 0.337 371 Chr1 785481 785871 A00251:497:H7TFGDSXC:3:2165:11894:19805 58 + CCACGC GAGTGA 0.413 390 Chr3 59890 60267 A00251:497:H7TFGDSXC:4:2344:8205:5462 60 - CACAAGT TATAGA 0.438 377
[0190] After that, the subject samples are randomly divided into a training set sample and a validation set sample. The training set sample includes 73 liver cancer samples and 128 healthy samples, and the validation set sample includes the remaining 48 lung cancer samples and 85 healthy samples.
[0191] Apply the RFE algorithm and machine learning algorithms (Random Forest, GBM, SVM or LASSO) to the training set samples according to the method in Example 2 to construct a model and perform feature screening, and return the optimal number of features and feature combinations. Among them, the top 15 feature combinations with relative weights obtained by each machine learning algorithm are shown in Tables 9-12, which are the liver cancer-specific cfDNA fragment feature combinations.
[0192] The 6bp motif {TGCTTC} at the 5' end of 1 cfDNA fragment appears in the liver cancer-specific cfDNA fragment feature combinations obtained by 3 machine learning algorithms (SVM, GBM and LASSO).
[0193] After summary analysis, taking the liver cancer-specific cfDNA fragment feature combination obtained based on RandomForest as an example, as Figure 9 and Figure 10 shown, the 6bp motif at the end of 10 cfDNA fragments with a length less than 150bp and the length distribution index of cfDNA fragments in 5 genomic 10kb windows are significantly different between liver cancer patients and healthy samples. Similar differences also exist in the lung cancer patients and healthy samples for the liver cancer-specific cfDNA fragment feature combinations obtained based on the other 3 machine learning algorithms. The above results indicate that the lung cancer-specific cfDNA fragment feature combination obtained by the method provided by the present invention is a potential molecular diagnostic marker for liver cancer.
[0194] Table 9 Feature combinations ranked top 15 in relative weight obtained based on RandomForest
[0195]
[0196]
[0197] Table 10 Feature combinations ranked top 15 in relative weight obtained based on GBM
[0198]
[0199] Table 11 Feature combinations ranked top 15 in relative weight obtained based on SVM
[0200]
[0201]
[0202] Table 12 Feature combinations ranked top 15 in relative weight obtained based on LASSO
[0203]
[0204] 4. Construction and performance evaluation of liver cancer prediction model
[0205] (1) Model construction
[0206] For the training set samples of this application example, the liver cancer-specific cfDNA fragment feature combinations of each sample in the training set as shown in Table 9 are respectively obtained as the input features of the randomForest model, and the liver cancer risk probability is used as the output result. The output result range is [0, 1]. The greater the liver cancer risk probability, the higher the possibility that the sample is a liver cancer sample. Combining the actual status (liver cancer and non-liver cancer) of each sample in the training set, 10-fold repeated cross-validation and 10 repeated iterations are performed, and the optimal training model is output based on the AUC index performance. During the 10-fold cross-validation process, the cutoff value is 0.5, and the judgment criterion is: if the liver cancer risk probability ≥ 0.5, the sample is judged as positive (i.e., a liver cancer sample); if the liver cancer risk probability < 0.5, the sample is judged as negative (i.e., a non-liver cancer sample).
[0207] For the training set of this application example, the liver cancer-specific cfDNA fragment feature combinations of each sample in the training set as shown in Table 10 are respectively obtained as the input features of the GBM model, and the lung cancer prediction model is constructed according to the same method as above.
[0208] For the training set of this application example, the characteristic combinations of liver cancer-specific cfDNA fragments of each sample in the training set are obtained as shown in Table 11, and used as the input features of the svmLinear model. According to the same method above, a lung cancer prediction model is constructed.
[0209] For the training set of this application example, the characteristic combinations of liver cancer-specific cfDNA fragments of each sample in the training set are obtained as shown in Table 12, and used as the input features of the LASSO model. According to the same method above, and with alpha set to seq(0, 1, by = 0.05) and lambda set to 10^seq(-2, 2, length = 100), taking AUC as the performance index, a lung cancer prediction model is constructed.
[0210] (2) Performance evaluation
[0211] Use the validation set samples of this application example to evaluate the performance of the optimal training models constructed by the four machine learning algorithms in step (1) of this application example. Calculate the prediction sensitivity, specificity, positive predictive value PPV, negative predictive value NPV, and comprehensive accuracy Accuracy indicators according to formulas (4) to (8).
[0212] Table 13 Performance evaluation results of the liver cancer prediction model
[0213] Model randomForest gbm svmLinear lasso Sensitivity 0.958 0.979 0.938 0.979 Specificity 0.941 0.918 0.929 0.882 PPV 0.902 0.87 0.882 0.825 NPV 0.976 0.987 0.963 0.987 AUC 0.984 0.972 0.976 0.978 Accuracy 0.947 0.94 0.932 0.917
[0214] As Figure 11 and shown in Table 13, the liver cancer prediction models constructed by the 4 machine learning algorithms all have excellent classification performance and can accurately determine whether the sample is a liver cancer sample. Among them, the liver cancer prediction model constructed based on the randomForest algorithm is the best in the 10-fold repeated cross-validation, with the AUC reaching as high as 0.984. Further summarize the liver cancer risk probabilities obtained from the training set samples and validation set samples of this model. As Figure 12 shown, it can be seen that the obtained determination results show significant discrimination in both the training set samples and the validation set samples.
[0215] In addition, use the terminal 6bp motifs of 10 cfDNA fragments with lengths less than 150bp in the liver cancer-specific cfDNA fragment characteristic combination obtained based on RandomForest as the input features, or use the length distribution index of 5 cfDNA fragments in the genomic 10kb window of the liver cancer-specific cfDNA fragment characteristic combination as the input features; take the liver cancer risk probability as the output result, and construct a liver cancer prediction model based on the training set samples and randomForest according to the same method, and use the validation set samples for performance evaluation. As Figure 13As shown in Table 14, the AUC of the single-dimensional model constructed based on the 6-bp motif at the 5' end of cfDNA fragments with a length less than 150 bp is 0.963, and the AUC of the single-dimensional model constructed based on the length distribution index of cfDNA fragments in the 10-kb window of the genome reaches 0.961. Although the prediction performance is lower than that of the multi-dimensional model constructed using all hepatocellular carcinoma-specific cfDNA fragment features, it still shows an accuracy of more than 95% and can accurately determine hepatocellular carcinoma samples.
[0216] Table 14 Performance evaluation results of different-dimensional models
[0217]
[0218] Example 3 A system for screening cancer-specific cfDNA fragment features in peripheral blood based on whole-genome sequencing
[0219] This example provides a system for screening cancer-specific cfDNA fragment features in peripheral blood based on whole-genome sequencing, including a data acquisition module, an analysis module, a storage module, and an output module. The data acquisition module is used to obtain the paired-end sequencing data of peripheral blood cfDNA of a sample group. The analysis module is used to obtain the cancer-specific cfDNA fragment features combination according to the paired-end sequencing data in combination with the method of Example 1 or Example 2. The storage module is used to store the paired-end sequencing data of peripheral blood cfDNA of the sample group obtained by the data acquisition module and the cancer-specific cfDNA fragment features combination obtained by the analysis module. The output module is used to output the cancer-specific cfDNA fragment features combination obtained by the analysis module.
[0220] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present invention and not to limit the protection scope of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description and ideas. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for screening a combination of cancer-specific cfDNA fragment features in peripheral blood based on whole-genome sequencing, characterized in that, Including the following steps: S1. Obtain the paired-end sequencing data of peripheral blood cfDNA of a sample group, retain the paired-end reads aligned to the human reference genome, and obtain information on cfDNA fragments based on the paired-end reads; the sample group includes cancer patients and healthy individuals; S2. Use the paired-end reads obtained in step S1 to obtain the eigenvalue of the following cfDNA fragment characteristics: The global length distribution index all-F-Index of cfDNA fragments, the 5'-end 6bp motif frequency 6bp-f of cfDNA fragments, the cfDNA fragment coverage Coverage_bin of a 10kb window of the genome, and the length distribution index F-Index_bin of cfDNA fragments in a 10kb window of the genome; Among them, the eigenvalue all-F-Index of the global length distribution index of cfDNA fragments is calculated by formula (1); Formula (1): wherein, i represents the length of the cfDNA fragment; P is a constant value, 160bp to 170bp; X i represents the proportion of cfDNA fragments with length i among all cfDNA fragments; The 5'-end 6bp motif frequency 6bp-f of cfDNA fragments is the occurrence frequency of each 6bp motif in the terminal motif set; the terminal motif set includes set 1 and set 2, and set 1 is composed of the 5'-end 6bp motifs of the nucleotide sequences of each cfDNA fragment; set 2 is composed of the reverse complementary sequences of the 3'-end 6bp motifs of the nucleotide sequences of each cfDNA fragment; The 10kb window of the genome is obtained by the following method: divide the genomic region after removing redundant regions into non-overlapping segments according to 10kb intervals to obtain M 10kb windows of the genome; the redundant regions include the short arm regions, telomere regions, and centromere regions of 5 chromosomes, and the 5 chromosomes are chromosome 13, chromosome 14, chromosome 15, chromosome 21, and chromosome 22; The eigenvalue Coverage_bin of the cfDNA fragment coverage of the 10kb window of the genome is calculated by the 10kb window of the genome and formula (2); Formula (2): Coverage_bin k = N k / N total ; Among them, Coverage_bin k represents the coverage of cfDNA fragments in the k-th 10kb-bins, where 1 ≤ k ≤ M; N k represents the number of cfDNA fragments aligned to the k-th 10kb-bins; N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads; The length distribution index F-Index_bin of cfDNA fragments in the 10kb window of the genome is calculated by the 10kb window of the genome and formula (3); Formula (3): wherein, i represents the length of the cfDNA fragment; P is a constant value, 160bp to 170bp; X’ i represents the proportion of cfDNA fragments with length i among all cfDNA fragments in the k-th genomic 10kb window; S3. Generate a feature matrix with the cfDNA fragment characteristic values obtained in step S2, and perform feature screening to obtain a feature combination; the feature screening is performed based on the accuracy of cancer prediction, and the cancer prediction includes predicting whether the sample is a cancer patient.
2. The method according to claim 1, wherein In step S2, the eigenvalue of the cfDNA fragment characteristic further includes: the 5'-end 6bp motif short-6bp-f of cfDNA fragments with a length less than 150bp, which is the occurrence frequency of each 6bp motif in the short-terminal motif set; the short-terminal motif set is obtained by retaining the subset corresponding to all cfDNA fragments with a length less than 150bp in the terminal motif set.
3. The method according to claim 1, wherein In step S1, retain the paired-end reads of chromosomes 1 to 22 aligned to the human reference genome.
4. The method according to claim 1, characterized in that, In step S3, the method for feature screening includes the following steps: using the feature matrix as input features and adopting the wrapper method for feature selection.
5. The method according to claim 4, characterized in that, The feature selection method adopted by the wrapper method includes recursive feature elimination.
6. The method according to claim 4, wherein The model construction methods adopted by the wrapper method include random forest, support vector machine, gradient boosting machine or LASSO regression.
7. The method according to claim 1, wherein The cancer is lung cancer or liver cancer.
8. Use of the method according to any one of claims 1 to 7 in the preparation of a product for cancer prediction.
9. A method for constructing a cancer prediction model, characterized in that, Processing the training set samples by the method according to any one of claims 1 to 7 to obtain a cancer-specific cfDNA fragment feature combination; using the cancer-specific cfDNA fragment feature combination as input features, using the probability of cancer risk as the output result, and using AUC as the performance index to train the classification model to obtain the cancer prediction model.
10. The construction method according to claim 9, wherein ,, the classification model includes random forest, support vector machine, gradient boosting machine or LASSO regression.
Citation Information
Patent Citations
Cancer-related biomarker based on cfDNA sequencing and data analysis as well as application of cancer-related biomarker based on cfDNA sequencing and data analysis in cfDNA sample classification
CN111254194A
Early cancer prediction method based on low-depth WGS sequencing end features
CN115910349A
Methods and systems for free DNA fragment size density to assess cancer
CN116157868A
Cancer non-invasive early screening method and system based on cfDNA terminal sequence characteristics
CN117316280A
DNA methylation and gene expression as determinants of genome-wide cell-free DNA fragmentation
WO2024263526A2
Cited By
CFTR gene transcription start site region feature extraction method based on cfDNA and application thereof
CN122471019A