A method for screening a cancer-specific cfDNA fragment combination in peripheral blood based on whole genome sequencing and application thereof
By combining whole-genome sequencing and multi-dimensional cfDNA fragment features with machine learning algorithms, cancer-specific cfDNA fragment feature combinations are screened out, solving the problems of high sequencing depth and single feature dimension in existing technologies, and achieving higher sensitivity and accuracy in early cancer screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2026-03-27
AI Technical Summary
Current technologies for early cancer diagnosis based on cfDNA fragment features require high sequencing depth, lack sufficient sensitivity and accuracy, and have limited feature dimensions, resulting in poor early cancer screening performance.
By screening multi-dimensional cfDNA fragment features through whole-genome sequencing and combining them with machine learning algorithms such as random forest, support vector machine, gradient boosting machine and LASSO regression, cancer-specific cfDNA fragment feature combinations can be screened for early cancer screening and prediction.
It improves the sensitivity and accuracy of early cancer screening, providing higher accuracy in diagnosis and the ability to assess early cancer risk.
Smart Images

Figure CN120279994B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of genetic diagnosis, in particular, to a method for screening a cancer-specific cfDNA fragment combination in peripheral blood based on whole genome sequencing and application thereof. BACKGROUND
[0002] Cancer has gradually become one of the top killers threatening human health worldwide, and its incidence and mortality continue to rise. "Early detection and early treatment" is crucial to improve the survival rate of cancer patients. However, due to the fact that many types of cancer have no obvious symptoms at an early stage, it is difficult to be detected, so it is an urgent need to develop an efficient and accurate early diagnosis tool.
[0003] With the development of genomics technology, through high-throughput sequencing technology or signal capture amplification method, the microtumor molecular signal in peripheral blood is identified, which provides a new strategy for early diagnosis of cancer. Cell-free DNA (cfDNA) in peripheral blood is free DNA released into blood by apoptotic or necrotic cells, which contains the characteristics of tissue-derived cells, including DNA methylation status, copy number variation, end sequencing pattern, coverage area, etc. In healthy individuals, most cfDNA molecules are derived from blood cells. When an organ or tissue is diseased, it is usually accompanied by more active apoptosis and more cfDNA molecules released from tissue cells into peripheral blood. Previous studies have found that the fragmentation characteristics of cfDNA in cancer patients are significantly different from those in normal population, so it can be used as an important cancer marker. The specific molecular signal of cfDNA in peripheral blood, such as tumor driver mutation, can be detected to identify early signals of cancer. The tumor early diagnosis technology based on cfDNA shows good application prospect.
[0004] Currently, the technical strategy based on cfDNA mutation detection is the clinically acceptable cancer screening route, such as Exact Science's CancerSEEK and Guardant Health's TEC-seq product, which identifies cancer signals by detecting tumor-specific gene mutations in peripheral blood. However, due to the extremely low level of ctDNA molecules derived from tumors in peripheral blood, and not all ctDNA fragments carry mutations, this strategy often requires extremely high sequencing depth, which poses a great challenge to the sensitivity of screening. In comparison, the fragment length, nucleosome imprinting, and end motif pattern of cfDNA are highly related to the tissue source type and pathological state, and have higher signal abundance, which can realize cancer screening based on low-depth sequencing. The fragmentomic features of cfDNA are the cancer molecular markers that have attracted much attention in recent years, including the length distribution of cfDNA fragments, end motif frequency, fragment breakpoint coordinates, nucleosome imprinting, and different topological structures. The DELFI team first proved the application prospect of cfDNA fragmentomics for tumor early screening in a Nature article published in 2019, achieving a detection sensitivity of 73% in 500 retrospective samples. However, DELFI only focuses on the length distribution characteristics of cfDNA fragments in a 5mb window of the genome in this study, and does not include information such as coverage and fragment end motif pattern, and the feature dimension is relatively single. Secondly, the 5mb genomic window has a large span, often covering multiple gene coding and non-coding regions, and there may be significant differences between different genes, transcriptional regulatory regions, and CDS regions. Using the average value of the entire 5mb window may lose specific tumor-related molecular features. Therefore, in order to improve the sensitivity and accuracy of early cancer screening, new cfDNA fragment features still need to be developed. SUMMARY
[0005] In order to solve the above problems existing in the prior art, the present application provides a method for screening a combination of cancer-specific cfDNA fragment features in peripheral blood based on whole genome sequencing and its application.
[0006] An object of the present application is to provide a method for screening a combination of cancer-specific cfDNA fragment features in peripheral blood based on whole genome sequencing.
[0007] Another object of the present application is to provide the use of the method in the preparation of a cancer prediction product.
[0008] Still another object of the present application is to provide a method for constructing a cancer prediction model.
[0009] In order to achieve the above objects, the present application is realized by the following scheme:
[0010] The application extracts and integrates multiple dimension cfDNA fragment characteristic values by low-depth whole genome sequencing of peripheral blood cfDNA, uses various machine learning algorithms including random forest, support vector machine, gradient boosting machine and regression algorithm LASSO to screen cfDNA characteristics highly related to cancer and model construction, and provides a marker combination and method which can be used for early screening and cancer prediction.
[0011] A method for screening cancer-specific cfDNA fragment characteristic combination in peripheral blood based on whole genome sequencing, comprising the following steps:
[0012] S1. Obtain double-end sequencing data of peripheral blood cfDNA of a sample group, retain double-end reads aligned to a human reference genome, obtain cfDNA fragment information based on the double-end reads, each pair of double-end reads corresponds to a cfDNA fragment, and the cfDNA fragment information includes but is not limited to the position of cfDNA in the genome, the nucleotide sequence of cfDNA, and the length of cfDNA; the sample group includes cancer patients and healthy people;
[0013] S2. Obtain the characteristic values of the following cfDNA fragment characteristics using the double-end reads obtained in step S1:
[0014] Global length distribution index all-F-Index of cfDNA fragments, 5' end 6bp motif frequency 6bp-f of cfDNA fragments, cfDNA fragment coverage Coverage_bin of a 10kb window of the genome, and length distribution index F-Index_bin of cfDNA fragments of a 10kb window of the genome;
[0015] Wherein, the characteristic value of the global length distribution index of the cfDNA fragments all-F-Index is calculated by formula (1);
[0016] Formula (1):
[0017] Wherein, i represents the length of the cfDNA fragment; P is a constant value, 160bp-170bp; X i represents the proportion of cfDNA fragments with length i in all cfDNA fragments;
[0018] The 6bp motif frequency 6bp-f of the 5' end of the cfDNA fragment is the occurrence frequency of each 6bp motif in the end motif set; the end motif set includes set 1 and set 2, the set 1 consists of the 5' end 6bp motif of the nucleotide sequence of each cfDNA fragment; the set 2 consists of the reverse complement sequence of the 3' end 6bp motif of the nucleotide sequence of each cfDNA fragment; the nucleotide sequence of the cfDNA fragment is obtained by aligning the paired-end reads to the human reference genome;
[0019] The genomic 10kb window is obtained by the following method: dividing the genomic region excluding the redundant region into non-overlapping segments at intervals of 10kb, to obtain M genomic 10kb windows; the redundant region includes the short arm region of 5 chromosomes, the telomere region and the centromere region, the 5 chromosomes are chromosome 13, chromosome 14, chromosome 15, chromosome 21 and chromosome 22;
[0020] The characteristic value Coverage_bin of the cfDNA fragment coverage of the genomic 10kb window is calculated by the genomic 10kb window and formula (2);
[0021] Formula (2): Coverage_bin k = N k / N total ;
[0022] Wherein, Coverage_bin k represents the coverage of cfDNA fragments of the kth 10kb-bins, 1≤k≤M; N k represents the number of cfDNA fragments aligned to the kth 10kb-bins; N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads;
[0023] The length distribution index F-Index_bin of the cfDNA fragments of the genomic 10kb window is calculated by the genomic 10kb window and formula (3);
[0024] Formula (3):
[0025] Wherein, i represents the length of the cfDNA fragment; P is a constant value, 160bp-170bp; X' i represents the proportion of cfDNA fragments with length i in all cfDNA fragments in the kth genomic 10kb window;
[0026] S3. generating a feature matrix with the feature values of cfDNA fragments obtained in step S2, and performing feature screening to obtain a feature combination; the feature screening is screening based on the accuracy of cancer prediction, and the cancer prediction includes predicting whether the sample is a cancer patient.
[0027] Preferably, in step S1, the number of people in the sample group is not less than 100.
[0028] More preferably, in step S1, the number of cancer patients in the sample group is not less than 50, and the number of healthy people is not less than 50.
[0029] Preferably, in step S1, the version of the human reference genome is hg38.
[0030] Preferably, in step S1, the paired-end reads aligned to chromosomes 1-22 of the human reference genome are retained.
[0031] Preferably, in step S2, the feature values of the cfDNA fragment features further include: short-6bp-f, a 6bp motif at the 5' end of cfDNA fragments with a length of less than 150bp, which is the frequency of each 6bp motif in the short-end motif set corresponding to the subset of all cfDNA fragments with a length of less than 150bp.
[0032] Preferably, in step S2, P is a constant value of 166bp.
[0033] Preferably, in step S3, the feature screening method includes the following steps: using the feature matrix as input features, and using the wrapping method to perform feature selection.
[0034] More preferably, the feature selection method used by the wrapping method includes recursive feature elimination.
[0035] More preferably, the model construction method used by the wrapping method includes random forest, support vector machine, gradient boosting machine or LASSO regression.
[0036] Preferably, the cancer is lung cancer or liver cancer.
[0037] Use of any of the methods in the preparation of a product for predicting cancer.
[0038] A system for screening cancer-specific cfDNA fragment feature combinations in peripheral blood based on whole genome sequencing, comprising a data acquisition module, an analysis module and an output module;
[0039] The data acquisition module is used to acquire paired-end sequencing data of peripheral blood cfDNA of a sample group.
[0040] The analysis module is configured to obtain a cancer-specific cfDNA fragment feature combination according to the double-end sequencing data; and a method for obtaining the cancer-specific cfDNA fragment feature combination comprises the following steps:
[0041] S1. retaining double-end reads matched to a human reference genome in the double-end sequencing data, each cfDNA fragment matching a pair of double-end reads; the sample group comprising cancer patients and healthy people;
[0042] S2. obtaining feature values of the following cfDNA fragment features by using the double-end reads obtained in step S1:
[0043] a global length distribution index all-F-Index of cfDNA fragments, a 6bp motif frequency 6bp-f of 5' ends of cfDNA fragments, a cfDNA fragment coverage Coverage_biin of a 10kb window of a genome, and a length distribution index F-Index_bin of cfDNA fragments of the 10kb window of the genome;
[0044] wherein the feature value all-F-Index of the global length distribution index of cfDNA fragments is calculated by formula (1);
[0045] formula (1):
[0046] wherein i represents the length of a cfDNA fragment; P is a constant value, 160bp-170bp; X i represents the proportion of cfDNA fragments with a length of i in all cfDNA fragments;
[0047] the 6bp motif frequency 6bp-f of 5' ends of cfDNA fragments is the occurrence frequency of each 6bp motif in an end motif set; the end motif set comprises set 1 and set 2, the set 1 consisting of 6bp motifs at 5' ends of nucleotide sequences of each cfDNA fragment; the set 2 consisting of reverse complementary sequences of 6bp motifs at 3' ends of nucleotide sequences of each cfDNA fragment; the nucleotide sequence of the cfDNA fragment being obtained by aligning the double-end reads to the human reference genome;
[0048] the 10kb window of the genome is obtained by the following method: dividing a genomic region excluding a redundant region into non-overlapping intervals at intervals of 10kb, to obtain M 10kb windows of the genome; the redundant region comprising short arm regions of 5 chromosomes, telomere regions and centromere regions, the 5 chromosomes being chromosome 13, chromosome 14, chromosome 15, chromosome 21 and chromosome 22;
[0049] a feature value Coverage_bin of the cfDNA fragment coverage of the genomic 10 kb window, which is calculated by the genomic 10 kb window and formula (2);
[0050] Formula (2): Coverage_bin k = N k / N total ;
[0051] wherein, Coverage_bin k represents the cfDNA fragment coverage of the kth 10 kb-bins, 1≤k≤M; N k represents the number of cfDNA fragments aligned to the kth 10 kb-bins; N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads;
[0052] a length distribution index F-Index_bin of the cfDNA fragments of the genomic 10 kb window, which is calculated by the genomic 10 kb window and formula (3);
[0053] Formula (3):
[0054] wherein, i represents the length of the cfDNA fragment; P is a constant value, 160 bp-170 bp; X i represents the proportion of cfDNA fragments with a length of i in all cfDNA fragments in the kth genomic 10 kb window;
[0055] S3. Generating a feature matrix with the feature values of the cfDNA fragments obtained in step S2, and performing feature screening to obtain a feature combination; the feature screening is screening based on the accuracy of cancer prediction, and the cancer prediction includes predicting whether the sample is a cancer patient;
[0056] The output module is configured to output the cancer-specific cfDNA fragment feature combination obtained by the analysis module.
[0057] Preferably, the system further comprises a storage module configured to store the paired-end sequencing data of the peripheral blood cfDNA of the sample group obtained by the data acquisition module and the cancer-specific cfDNA fragment feature combination obtained by the analysis module.
[0058] Preferably, in step S1, the number of people in the sample group is not less than 100.
[0059] More preferably, in step S1, the number of cancer patients in the sample group is not less than 50, and the number of healthy people is not less than 50.
[0060] Preferably, in step S1, the version of the human reference genome is hg38.
[0061] Preferably, in step S1, the paired-end reads aligned to chromosomes 1-22 of the human reference genome are retained.
[0062] Preferably, in step S2, P is a constant value of 166 bp.
[0063] Preferably, in step S2, the feature values of the cfDNA fragment features further comprise: a 6bp motif short-6bp-f at the 5' end of cfDNA fragments with a length less than 150 bp, which is the frequency of each 6bp motif in the short-end motif set corresponding to the subset of all cfDNA fragments with a length less than 150 bp.
[0064] Preferably, in step S3, the method of feature screening comprises the following steps: using the feature matrix as input features, and performing feature selection using a wrapper method.
[0065] More preferably, the feature selection method used by the wrapper method comprises recursive feature elimination.
[0066] More preferably, the model construction method used by the wrapper method comprises random forest, support vector machine, gradient boosting machine or LASSO regression.
[0067] Preferably, the cancer is lung cancer or liver cancer.
[0068] A computer device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor; when the computer program is executed by the processor, the operation of the system for screening cancer-specific cfDNA fragment feature combinations in peripheral blood based on whole genome sequencing is realized.
[0069] A computer-readable storage medium storing a computer program executable by a processor, wherein when the computer program is executed by the processor, the operation of the system for screening cancer-specific cfDNA fragment feature combinations in peripheral blood based on whole genome sequencing is realized.
[0070] A method for constructing a cancer prediction model, wherein a training set sample is processed by any of the methods to obtain a cancer-specific cfDNA fragment feature combination; the cancer-specific cfDNA fragment feature combination is used as input features, the risk probability of cancer is used as output results, and the AUC is used as a performance indicator to train a classification model to obtain the cancer prediction model.
[0071] Preferably, the classification model comprises a random forest, a support vector machine, a gradient boosting machine or a LASSO regression.
[0072] Compared with the prior art, the application has the following beneficial effects:
[0073] The application combines multiple dimensions of cfDNA fragment information screening and cancer-related feature indicators, and provides a new index F-Index for evaluating fragment length distribution characteristics, which has significant differences between cancer samples and healthy samples. In addition, the application also applies the motif characteristics of the 6 bases at the end of the cfDNA fragment and the cfDNA fragment distribution characteristics within the 10 kb window of the genome. A series of cancer-specific molecular feature combinations are identified by applying multiple machine learning algorithms, which are suitable for early cancer risk assessment and prediction. The feature combinations screened by the method of the application have excellent prediction performance in different classification models, and the determination accuracy is high. BRIEF DESCRIPTION OF DRAWINGS
[0074] Figure 1 Flowchart of the method for screening cancer-specific cfDNA fragment feature combinations in peripheral blood based on whole genome sequencing.
[0075] Figure 2 Statistical chart of all-F-Index of global length distribution of cfDNA fragments of lung cancer samples and healthy samples in application example 1, Controls represents healthy samples, Lung Cancers represents lung cancer samples.
[0076] Figure 3 Statistical chart of 6bp-f of 11 cfDNA fragments of 6bp motif at the 5' end of lung cancer samples and healthy samples in application example 1, Controls represents healthy samples, Lung Cancers represents lung cancer samples.
[0077] Figure 4 Statistical chart of Coverage_bin of cfDNA fragment coverage of 4 genomic 10kb windows and F-Index_bin of length distribution index of cfDNA fragments of 5 genomic 10kb windows of lung cancer samples and healthy samples in application example 1, Controls represents healthy samples, Lung Cancers represents lung cancer samples.
[0078] Figure 5 AUC chart of lung cancer prediction model constructed by using 4 machine learning algorithms with validation set samples in application example 1.
[0079] Figure 6Distribution of lung cancer risk probability obtained from lung cancer prediction model built with GBM algorithm in application example 1 in training set samples and validation set samples, Training represents training set samples, Validation represents validation set samples, Controls represents healthy samples, Cancers represents lung cancer samples.
[0080] Figure 7 AUC plot of liver cancer prediction models built with different input features using validation set samples in application example 2, Integration represents multi-dimension model built with lung cancer specific cfDNA fragment feature combination as input feature, Motif represents single dimension model built with 11 cfDNA fragment 5’ end 6bp motif as input feature, 10k-Bins represents two-dimension model built with cfDNA fragment coverage of 4 genomic 10kb windows and length distribution index of cfDNA fragments of 5 genomic 10kb windows as input features.
[0081] Figure 8 all-F-Index statistical plot of global length distribution index of cfDNA fragments of liver cancer samples and healthy samples in application example 2, Controls represents healthy samples, Liver Cancers represents liver cancer samples.
[0082] Figure 9 Short_6bp-f statistical plot of 10 cfDNA fragments with length less than 150bp 5’ end 6bp motif of liver cancer samples and healthy samples in application example 2, Controls represents healthy samples, Liver Cancers represents liver cancer samples.
[0083] Figure 10 F-Index_bin statistical plot of length distribution index of cfDNA fragments of 5 genomic 10kb windows of liver cancer samples and healthy samples in application example 2, Controls represents healthy samples, Liver Cancers represents liver cancer samples.
[0084] Figure 11 AUC plot of liver cancer prediction models built with four machine learning algorithms using validation set samples in application example 2.
[0085] Figure 12 Distribution of lung cancer risk probability obtained from lung cancer prediction model built with randomForest algorithm in application example 2 in training set samples and validation set samples, Training represents training set samples, Validation represents validation set samples, Controls represents healthy samples, Cancers represents liver cancer samples.
[0086] Figure 13 AUC plot of liver cancer prediction models constructed using different input features using validation set samples in application example 2, Integration represents a multi-dimensional model constructed using liver cancer specific cfDNA fragment feature combinations as input features, Motif represents a single-dimensional model constructed using 5' end 6bp motifs of 10 cfDNA fragments with length less than 150bp as input features, and 10k-Bins represents a single-dimensional model constructed using length distribution index of cfDNA fragments in 5 genomic 10kb windows as input features. DETAILED DESCRIPTION
[0087] The application will be further described below in conjunction with the accompanying drawings and specific embodiments, which are only used to explain the application and are not used to limit the scope of the application. The test methods used in the following examples are conventional methods unless otherwise specified. The materials, reagents, etc. used are commercially available unless otherwise specified.
[0088] Example 1: A method for screening cancer specific cfDNA fragment feature combinations in peripheral blood based on whole genome sequencing
[0089] This embodiment provides a method for screening cancer specific cfDNA fragment feature combinations in peripheral blood based on whole genome sequencing, the process of which is shown in Figure 1 The following provides a specific implementation.
[0090] 1. Acquisition and processing of plasma cfDNA whole genome sequencing data
[0091] The present application extracts, library constructs and double-end sequences cfDNA from peripheral blood samples of subjects. The method for extracting, library constructing and sequencing cfDNA is not particularly limited and can be adjusted from existing technical methods. The following provides a specific method.
[0092] (1) Plasma separation
[0093] Fresh peripheral blood samples are centrifuged at 3000xg for 10 minutes at 4°C within 4 hours after blood collection, and the lower blood cells are discarded. Then further centrifuged at 16000xg for 10 minutes, and the supernatant plasma is transferred to a cryogenic tube, i.e. the plasma is obtained, which is stored at -80°C before nucleic acid extraction.
[0094] (2) cfDNA extraction
[0095] The thawed plasma sample was centrifuged at 16000xg for 10 minutes, and cfDNA was extracted using the circulating Nucleic Acid Kit. The DNA concentration was detected using a Qubit nucleic acid quantifier. The nucleic acid sample was stored at -20°C before library construction and sequencing.
[0096] (3) Library construction and paired-end sequencing
[0097] The cfDNA library was constructed using the library construction kit QIAseq cfDNA All-in-One Kit (24) (Cat.No. / ID: 180023), and the cfDNA was not interrupted during library construction. Fragment detection was performed using the LabChip GX Touch HT Nucleic Acid Analyzer instrument. The cfDNA library was sequenced using the Illumina Novaseq sequencing platform, with a read length of ~ 150 bp and a sequencing depth of 1x-20x. The average output of each sample was 3G-60G raw data. After sequencing, raw fastq sequencing data was obtained.
[0098] (4) Data quality control processing and alignment analysis
[0099] The raw fastq data was processed using Fastp software based on the default parameters, and the adapter, data with a base quality value less than Q30, sequences containing a high N content (i.e., removing reads with N-Content greater than 2%), and PCR-induced repetitive sequences were removed to obtain paired-end reads. Each cfDNA fragment was matched to a pair of paired-end reads (read1 and read2).
[0100] The bwa software was used to align the quality-controlled paired-end reads to the human reference genome hg38 (Genome Reference Consortium human genome build 38, GRCh38) to obtain the Bam file of the alignment results. The picard tool was used to pretreat the bam file, remove the multiple alignment paired-end reads, remove the paired-end reads that did not align to the reference genome, and retain the paired-end reads with a quality value greater than 30 to obtain the clean bam file.
[0101] 2. Analysis of cfDNA fragment characteristics and characteristic values
[0102] This example involves four types of cfDNA fragment characteristics: global length distribution index of cfDNA fragments, 5' end 6bp motif of cfDNA fragments, cfDNA fragment coverage of the 10kb window of the genome, and length distribution index of cfDNA fragments of the 10kb window of the genome.
[0103] (1) Calculation of the global length distribution index characteristic value all-F-Index of cfDNA fragments
[0104] The global length distribution index all-F-Index of cfDNA fragments reflects the enrichment degree of short fragments in cfDNA fragments. The greater the all-F-Index, the higher the proportion of short fragments, indicating that the cfDNA fragments as a whole are biased towards shorter length distribution, and there is an enrichment of short fragments.
[0105] The calculation method of all-F-Index is as follows:
[0106] Extract the alignment coordinates, direction, length and base sequence information of each pair of paired-end reads from the clean bam file obtained in the previous step. In this analysis step, only paired-end reads aligned to the 1-22 autosomal regions are retained, and paired-end reads aligned to the X, Y chromosomes, mitochondrial genome or other supplementary contig regions are removed to obtain a bed file of cfDNA fragment distribution on the genome (i.e. a genomic bed file). The file contains cfDNA fragment information on the genome, including cfDNA fragment sequence, length and location information on the chromosome.
[0107] According to the genomic bed file and formula (1), all-F-Index is calculated.
[0108] Formula (1):
[0109] Where i represents the length of the cfDNA fragment; P is a constant value of 166 bp, which is the common length of cfDNA fragments in healthy human plasma; X i represents the proportion of cfDNA fragments with length i in all cfDNA fragments.
[0110] (2) Calculation of the 5' end 6bp motif characteristic value of cfDNA fragments
[0111] cfDNA is derived from the apoptosis or active release of tissues and blood cells, which is affected by nuclease cleavage.
[0112] The characteristic value of the 5' end 6bp motif of cfDNA fragments is the 5' end 6bp motif frequency (6bp-f) of cfDNA fragments, which reflects the influence of nuclease activity and tissue type on the end motif pattern of cfDNA fragments, and its statistical method is as follows:
[0113] For each cfDNA fragment, each pair of paired-end reads consists of read 1 of the forward strand and read 2 of the reverse strand, by aligning to the reference genome, the forward complete sequence of the cfDNA fragment (i.e. the nucleotide sequence from 5' end to 3' end) is obtained, and the 6bp sequence of its 5' end is defined as End_motif_F, and the 6bp sequence of its 3' end is defined as End_motif_R. The sequence of End_motif_F is not converted, and the End_motif_F of all cfDNA fragments form set 1; the sequence of End_motif_R is converted to its reverse complement to obtain the reverse complement sequence of End_motif_R, and the reverse complement sequence of End_motif_R of all cfDNA fragments form set 2. After merging set 1 and set 2, the set of 5' end 6bp motifs of all cfDNA fragments is obtained, denoted as End_motif_set. There are theoretically at most 4096 different 6bp motifs, such as {AAAAAA, AAAAAT,..., ACGACG, CTTGTT,..., GGGGGG}, and the frequency of each 6bp motif in the End_motif_set is counted, i.e. 6bp-f is obtained.
[0114] (3) Calculation of the feature value of cfDNA fragment coverage in the genomic 10kb window
[0115] The coverage of cfDNA fragments in the genomic 10kb window is affected by the distribution and structure of nucleosomes, and is related to the type of tissue. Generally, the higher the degree of chromatin openness in the genomic window, the lower the coverage of cfDNA fragments. The calculation method of Coverage_bin is as follows:
[0116] Human chromosomes 13, 14, 15, 21 and 22, because the short arm region of these five chromosomes contains the nucleolus formation region, which is rich in a large number of repetitive ribosomal RNA genes, making it difficult for genomic sequencing technology to accurately analyze the complete sequence in these regions, therefore, based on the bed file of the distribution of cfDNA fragments on the genome obtained in this embodiment, the short arm region of these five chromosomes and the telomere and centromere region are removed, and the remaining genomic region is divided into non-overlapping intervals of 10kb (such as chr1:20000-chr1:29999) to obtain M genomic 10kb windows (i.e. 10kb-bins), and theoretically at most 287465 10kb-bins can be obtained, i.e. Mmax is 287465.
[0117] Coverage_bin is calculated according to the genomic bed file, 10kb-bins and formula (2).
[0118] Formula (2): Coverage_bin k = Nk / N total ;
[0119] Coverage_bin k represents the coverage of cfDNA fragments in the kth 10kb-bins, 1≤k≤M;N k represents the number of cfDNA fragments aligned to the kth 10kb-bins, wherein reads spanning two adjacent 10kb-bins are considered to be aligned to both 10kb-bins, and both 10kb-bins are increased by one count when calculating coverage, and the subsequent calculation of F-Index_bin k when the read is also included in the calculation;N total represents the number of all cfDNA fragments aligned to the reference genome based on paired-end reads.
[0120] (4) Calculation of length distribution index F-Index_bin of cfDNA fragments in genomic 10kb window
[0121] The length distribution index F-Index_bin of cfDNA fragments in genomic 10kb window is affected by both nucleosome structure and histone modification, and is related to tissue type and pathological state, and the calculation method of F-Index_bin is as follows:
[0122] According to the genomic bed file, 10kb-bins and formula (3), the F-Index_bin of cfDNA fragments in the kth 10kb-bins is calculated.
[0123] Formula (3):
[0124] wherein i represents the length of cfDNA fragments; P is a constant value of 166bp, which is the common length of cfDNA fragments in healthy human plasma; X’ i represents the proportion of cfDNA fragments with length i in all cfDNA fragments in the kth 10kb-bins.
[0125] Through the calculation and statistics of the embodiment, for each sample, one all-F-Index value, at most 4096 6bp-f values, 287465 Coverage_bin values and 287465 F-Index_bin values are obtained.
[0126] 3. Determination of cancer-specific cfDNA fragment feature combination
[0127] Taking cancer patients and healthy people as sample groups, the total cfDNA fragment features and feature values of each sample are obtained according to the above step, and a feature matrix is generated, each row representing a sample, and each column representing the cfDNA fragment features and feature values. All feature values are normalized and normalized.
[0128] The converted feature matrix is taken as the input feature, and feature selection is performed based on a machine learning model. Specifically, a model constructed by combining a recursive feature elimination (RFE) algorithm with different machine learning algorithms (such as random forest, support vector machine, gradient boosting machine, or regression algorithm LASSO) is used for feature screening. In the construction of the model, the model parameters are selected by 10-fold repeated cross-validation, and the number of repeated sampling iterations is 10. The accuracy of predicting whether a patient is a cancer patient is used as the result evaluation parameter, and the optimal feature quantity and feature combination are returned, which is the cancer-specific cfDNA fragment feature combination.
[0129] 4. Cancer prediction model construction and performance evaluation
[0130] Based on the obtained cancer-specific cfDNA fragment feature combination, further combined with the corresponding machine learning algorithm used in the screening of the feature combination, a cancer prediction model is trained, and the model parameters are selected by 10-fold repeated cross-validation, 10 repeated iterations, and the prediction performance of the model is evaluated.
[0131] Example 2: A method for screening cancer-specific cfDNA fragment feature combinations in peripheral blood based on whole genome sequencing
[0132] 1. Acquisition and processing of plasma cfDNA whole genome sequencing data
[0133] The same as example 1.
[0134] 2. Analysis of cfDNA fragment features and their feature values
[0135] Basically the same as example 1, the difference is that in addition to the 4 types of cfDNA fragment features, the 5th type of cfDNA fragment feature is also involved: the 6bp motif at the 5' end of cfDNA fragments with a length of less than 150bp, which reflects both short fragment enrichment and nuclease cutting bias. The feature value is the frequency of the 6bp motif at the 5' end of cfDNA fragments with a length of less than 150bp (short-6bp-f), and the statistical method is as follows:
[0136] For each cfDNA fragment with length below 150 bp, the 6 bp sequence at its 5' end is defined as Short_End_motif_F and the 6 bp sequence at its 3' end is defined as Short_End_motif_R. The subset of the End_motif_set corresponding to all cfDNA fragments with length below 150 bp is extracted and denoted as short-End_motif_set, which is the union of Set 1' consisting of Short_End_motif_F and Set 2' consisting of the reverse complement of Short_End_motif_R. The frequency of each 6 bp motif in short-End_motif_set is calculated, i.e. short-6bp-f is obtained.
[0137] Through the calculation and statistics of this embodiment, for each sample, one all-F-Index, at most 4096 6bp-f, 4096 short-6bp-f, 287465 Coverage_bin and 287465 F-Index_bin are obtained.
[0138] 3. Determination of lung cancer specific cfDNA fragment feature combination
[0139] The same as Embodiment 1.
[0140] 4. Construction and performance evaluation of lung cancer prediction model
[0141] The same as Embodiment 1.
[0142] Application Example 1: Lung cancer prediction method based on plasma free DNA fragment feature combination
[0143] 1. Peripheral blood samples of subjects
[0144] In this application example, 398 peripheral blood samples of lung cancer patients (i.e. lung cancer samples) and 168 peripheral blood samples of healthy people (i.e. healthy samples) were included, each peripheral blood sample was derived from a different individual, and all samples were derived from the sample library of Shenzhen Hyplons Biotechnology Co., Ltd. and the informed consent of the subjects was obtained before sample detection.
[0145] 2. Determination of lung cancer specific cfDNA fragment feature combination
[0146] All peripheral blood samples of subjects were used as sample group, and the method in Embodiment 1 was used for processing to generate the genomic bed file shown in Table 1, and then the method in Embodiment 1 was used for analysis to obtain the cfDNA fragment features and feature values of each sample and generate the feature matrix, and the normalization and normalization conversion were performed on all feature values. The all-F-Index of the global length distribution index of cfDNA fragments of lung cancer samples and healthy samples were summarized and compared, as shown in Table 2.Figure 2 As shown, the all-F-Index of lung cancer samples is significantly higher than that of healthy samples, indicating that the cfDNA fragment features have a significant correlation with lung cancer and can be used as potential molecular diagnostic markers for lung cancer.
[0147] Table 1: Format of genomic bed file generated by cfDNA sequencing data alignment (partial)
[0148]
[0149]
[0150] Then, the subject samples were randomly divided into training set samples and validation set samples, wherein the training set samples contained 239 lung cancer samples and 101 healthy samples, and the validation set samples contained the remaining 159 lung cancer samples and 67 healthy samples.
[0151] The training set samples were subjected to the RFE algorithm and each machine learning algorithm (RandomForest, GBM, SVM or LASSO) according to the method of Example 1 to construct models and perform feature screening, and the optimal number of features and feature combination were returned. Among them, the top 20 feature combinations according to the relative weight obtained by each machine learning algorithm are shown in Tables 2-5, which are lung cancer-specific cfDNA fragment feature combinations.
[0152] The 5' end 6bp motif {AAGTGC} of 1 cfDNA fragment appeared in the lung cancer-specific cfDNA fragment feature combinations obtained by the four machine learning algorithms. The 5' end 6bp motifs {AAGGGT, ACCTCT, TTTAGG} of 3 cfDNA fragments appeared in the lung cancer-specific cfDNA fragment feature combinations obtained by three machine learning algorithms (RandomForest, SVM and GBM).
[0153] Based on the lung cancer-specific cfDNA fragment feature combination obtained by GBM, as shown in Figure 3 and Figure 4 As shown, the 5' end 6bp motifs of 11 cfDNA fragments, the cfDNA fragment coverage of 4 genomic 10kb windows and the length distribution index of 5 genomic 10kb windows of cfDNA fragments have obvious differences between lung cancer patients and healthy samples. The lung cancer-specific cfDNA fragment feature combinations obtained based on the remaining three machine learning algorithms also have similar differences between lung cancer patients and healthy samples. The above results show that the lung cancer-specific cfDNA fragment feature combinations obtained by the method provided by the present application are potential molecular diagnostic markers for lung cancer.
[0154] Table 2 Top 20 feature combinations ranked by relative weight based on RandomForest
[0155]
[0156]
[0157] Table 3 Top 20 feature combinations ranked by relative weight based on GBM
[0158]
[0159] Table 4 Top 20 feature combinations ranked by relative weight based on SVM
[0160]
[0161]
[0162] Table 5 Top 20 feature combinations ranked by relative weight based on LASSO
[0163]
[0164] 3. Construction and performance evaluation of lung cancer prediction model
[0165] (1) Model construction
[0166] For the training set samples of the present application example, the lung cancer specific cfDNA fragment feature combinations of each sample in the training set as shown in Table 2 were obtained as input features of the randomForest model, and the lung cancer risk probability was taken as the output result, the output result range was [0, 1], the greater the lung cancer risk probability, the higher the possibility of the sample being a lung cancer sample. Combined with the actual state (lung cancer and non-lung cancer) of each sample in the training set, 10-fold repeated cross-validation was performed, 10 repeated iterations were performed, and the optimal training model was output based on the AUC index performance. In the 10-fold cross-validation process, the cutoff value was 0.5, and the determination criteria were: lung cancer risk probability ≥ 0.5, the sample was determined as positive (i.e. lung cancer sample); lung cancer risk probability < 0.5, the sample was determined as negative (i.e. non-lung cancer sample).
[0167] For the training set of the present application example, the lung cancer specific cfDNA fragment feature combinations of each sample in the training set as shown in Table 3 were obtained as input features of the GBM model, and the lung cancer prediction model was constructed according to the same method as described above.
[0168] For the training set of the present application example, the lung cancer specific cfDNA fragment feature combinations of each sample of the training set as shown in Table 4 were obtained, as the input features of the svmLinear model, and the lung cancer prediction model was constructed according to the same method as described above.
[0169] For the training set of the present application example, the lung cancer specific cfDNA fragment feature combinations of each sample of the training set as shown in Table 5 were obtained, as the input features of the LASSO model, and the lung cancer prediction model was constructed according to the same method as described above, with alpha set to seq(0, 1, by = 0.05) and lambda set to 10^seq(-2, 2, length = 100), and AUC as the performance indicator.
[0170] (2) Performance evaluation
[0171] The performance of the optimal training model constructed by the four machine learning algorithms in step (1) of the present application example was evaluated using the validation set samples of the present application example, and the prediction sensitivity Sensitivity, specificity Specificity, positive predictive value PPV (Positive Predictive Value), negative predictive value NPV (Negative Predictive Value) and comprehensive accuracy Accuracy indicators were calculated according to formulas (4)-(8).
[0172] Formula (4): Sensitivity = TP / (TP+FN);
[0173] Formula (5): Specificity = TN / (TN+FP);
[0174] Formula (6): PPV = TP / (TP+FP);
[0175] Formula (7): NPV = TN / (TN+FN);
[0176] Formula (8): Accuracy = (TP+TN) / (TP+FN+TN+FN).
[0177] Table 6 Performance evaluation results of lung cancer prediction model
[0178] Model randomForest gbm svmLinear lasso Sensitivity 0.874 0.874 0.849 0.893 Specificity 0.896 0.94 0.91 0.91 PPV 0.952 0.972 0.957 0.959 NPV 0.75 0.759 0.718 0.782 AUC 0.953 0.961 0.948 0.959 Accuracy 0.881 0.894 0.867 0.782
[0179] As Figure 5As shown in Table 6, the lung cancer prediction models constructed by the four machine learning algorithms all have excellent classification performance and can accurately determine whether the sample is a lung cancer sample. Among them, the lung cancer prediction model based on the GBM algorithm is optimal, with the best AUC reaching 0.961 in 10-fold repeated cross-validation. Further, the lung cancer risk probability obtained by the model in the training set sample and the validation set sample is shown in Table 6. Figure 6 As shown in Table 6, the lung cancer prediction models constructed by the four machine learning algorithms all have excellent classification performance and can accurately determine whether the sample is a lung cancer sample. Among them, the lung cancer prediction model based on the GBM algorithm is optimal, with the best AUC reaching 0.961 in 10-fold repeated cross-validation. Further, the lung cancer risk probability obtained by the model in the training set sample and the validation set sample is shown in Table 6.
[0180] In addition, 11 cfDNA fragments in the lung cancer-specific cfDNA fragment feature combination obtained based on the GBM were used as input features, or 4 genomic 10 kb window cfDNA fragment coverages and 5 genomic 10 kb window cfDNA fragment length distribution indexes in the lung cancer-specific cfDNA fragment feature combination were used as input features; the lung cancer risk probability was used as the output result, and the lung cancer prediction model was constructed based on the training set sample and the gradient boosting machine GBM by referring to the same method, and the performance was evaluated by the validation set sample. As shown in Table 6, Figure 7 As shown in Table 7, the AUC of the single-dimensional model constructed based on the 5' end 6bp motif of the cfDNA fragment is 0.94, and the AUC of the two-dimensional model constructed based on the genomic 10 kb window cfDNA fragment coverage and the genomic 10 kb window cfDNA fragment length distribution index reaches 0.92. Although the prediction performance is lower than that of the multi-dimensional model constructed by using all lung cancer-specific cfDNA fragment features, it still shows an accuracy of more than 85% and can realize accurate determination of lung cancer samples.
[0181] Table 7 Performance evaluation results of different dimensional models
[0182]
[0183] Application Example 2: Liver Cancer Prediction Method Based on Plasma Free DNA Fragment Feature Combination
[0184] 1. Peripheral blood samples of subjects
[0185] In this application example, 121 peripheral blood samples (i.e. liver cancer samples) of liver cancer patients and 213 peripheral blood samples (i.e. healthy samples) of healthy people were included, each peripheral blood sample was derived from a different individual, and all samples were derived from the sample library of Shenzhen Hailuo Si Biotechnology Co., Ltd. The informed consent of the subjects was obtained before sample detection.
[0186] 2. Determination of liver cancer-specific cfDNA fragment feature combination
[0187] With all the peripheral blood samples of the subjects as the sample group, the method in Example 2 was used for processing, and the genomic bed file shown in Table 8 was generated, and then the method in Example 2 was used for analysis to obtain the characteristics and characteristic values of all cfDNA fragments of each sample and generate a characteristic matrix, and all the characteristic values were normalized and normalized. Figure 8 As shown in Table 8, the all-F-Index of the global length distribution index of cfDNA fragments of the liver cancer samples and the healthy samples was compared and analyzed, and the all-F-Index of the liver cancer samples was significantly higher than that of the healthy samples, indicating that the cfDNA fragment characteristics were significantly associated with liver cancer and could be used as a potential molecular diagnostic marker for liver cancer.
[0188] Table 8: Part of the genomic bed file format generated by cfDNA sequencing data alignment
[0189] #chr Start End Name score strand End_motif_F End_motif_R gc length Chr1 60196 60567 A00251:497:H7TFGDSXC:3:1576:21233:32784 43 - ACATTG ATCATC 0.337 371 Chr1 785481 785871 A00251:497:H7TFGDSXC:3:2165:11894:19805 58 + CCACGC GAGTGA 0.413 390 Chr3 59890 60267 A00251:497:H7TFGDSXC:4:2344:8205:5462 60 - CACAAGT TATAGA 0.438 377
[0190] Then the samples of the subjects were randomly divided into training set samples and validation set samples, wherein the training set samples included 73 liver cancer samples and 128 healthy samples, and the validation set samples included the remaining 48 lung cancer samples and 85 healthy samples.
[0191] The RFE algorithm and machine learning algorithms (Random Forest, GBM, SVM or LASSO) were applied to the training set samples according to the method of Example 2 to construct a model and feature screening, and the optimal number of features and feature combination were returned, wherein the top 15 features ranked by the relative weight of each machine learning algorithm were shown in Tables 9-12, which were the liver cancer-specific cfDNA fragment feature combinations.
[0192] There was a 5' end 6bp motif {TGCTTC} of 1 cfDNA fragment in the liver cancer-specific cfDNA fragment feature combinations obtained by 3 machine learning algorithms (SVM, GBM and LASSO).
[0193] Based on the liver cancer-specific cfDNA fragment feature combination obtained by Random Forest, as shown in Tables 9-12, the 6bp motif at the end of 10 cfDNA fragments with a length less than 150bp and the length distribution index of 5 cfDNA fragments in the 10kb genomic window had significant differences between the liver cancer patients and the healthy samples. Figure 9 Figure 10 As shown in Tables 9-12, the 6bp motif at the end of 10 cfDNA fragments with a length less than 150bp and the length distribution index of 5 cfDNA fragments in the 10kb genomic window had significant differences between the liver cancer patients and the healthy samples.
[0194] Table 9 Top 15 feature combinations ranked by relative weight based on RandomForest
[0195]
[0196]
[0197] Table 10 Top 15 feature combinations ranked by relative weight based on GBM
[0198]
[0199] Table 11 Top 15 feature combinations ranked by relative weight based on SVM
[0200]
[0201]
[0202] Table 12 Top 15 feature combinations ranked by relative weight based on LASSO
[0203]
[0204] 4. Construction and performance evaluation of liver cancer prediction model
[0205] (1) Model construction
[0206] For the training set samples of the present application example, the liver cancer specific cfDNA fragment feature combinations shown in Table 9 were obtained for each sample in the training set as input features of the randomForest model, and the liver cancer risk probability was taken as the output result, the output result range was [0, 1], the greater the liver cancer risk probability, the higher the possibility of the sample being a liver cancer sample. Combined with the actual state (liver cancer and non-liver cancer) of each sample in the training set, 10-fold repeated cross-validation was performed, 10 repeated iterations were performed, and the optimal training model was output based on the AUC index performance. In the 10-fold cross-validation process, the cutoff value was 0.5, and the determination criteria were: liver cancer risk probability ≥ 0.5, the sample was determined to be positive (i.e. liver cancer sample); liver cancer risk probability < 0.5, the sample was determined to be negative (i.e. non-liver cancer sample).
[0207] For the training set of the present application example, the liver cancer specific cfDNA fragment feature combinations shown in Table 10 were obtained for each sample in the training set as input features of the GBM model, and the lung cancer prediction model was constructed according to the same method as described above.
[0208] For the training set of this application example, the liver cancer-specific cf DNA fragment feature combinations of each sample in the training set as shown in Table 11 were obtained and used as input features of the svmLinear model. The lung cancer prediction model was constructed in the same way as above.
[0209] For the training set of this application example, the liver cancer-specific cf DNA fragment feature combinations of each sample in the training set as shown in Table 12 were obtained and used as the input features of the LASSO model. Following the same method as above, with alpha set to seq(0,1,by=0.05) and lambda set to 10^seq(-2,2,length=100), and AUC as the performance index, a lung cancer prediction model was constructed.
[0210] (2) Performance Evaluation
[0211] The performance of the optimal training models constructed by the four machine learning algorithms in step (1) of this application example is evaluated using the validation set samples of this application example. The prediction sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and overall accuracy index are calculated according to formulas (4) to (8).
[0212] Table 13 Performance evaluation results of the liver cancer prediction model
[0213] Model randomForest gbm svmLinear lasso Sensitivity 0.958 0.979 0.938 0.979 Specificity 0.941 0.918 0.929 0.882 PPV 0.902 0.87 0.882 0.825 NPV 0.976 0.987 0.963 0.987 AUC 0.984 0.972 0.976 0.978 Accuracy 0.947 0.94 0.932 0.917
[0214] like Figure 11 As shown in Table 13, the liver cancer prediction models constructed using the four machine learning algorithms all exhibit excellent classification performance, accurately determining whether a sample is a liver cancer sample. Among them, the liver cancer prediction model based on the randomForest algorithm achieves the best result in 10-fold repeated cross-validation, with an AUC as high as 0.984. Further summarizing the liver cancer risk probabilities obtained by this model from the training set and validation set samples, as shown... Figure 12 As shown, the obtained judgment results exhibit significant discriminative power in both the training set and the validation set.
[0215] In addition, the last 6 bp motifs of 10 cfDNA fragments shorter than 150 bp from the liver cancer-specific cfDNA fragment feature combination obtained based on RandomForest are used as input features, or the length distribution index of cfDNA fragments in 5 genomic 10kb windows from the liver cancer-specific cfDNA fragment feature combination is used as input features; the liver cancer risk probability is used as the output result. Following the same method, a liver cancer prediction model is constructed based on the training set samples and randomForest, and its performance is evaluated using the validation set samples. Figure 13As shown in Table 14, the AUC of the single dimension model constructed based on the 6bp motif of the 5' end of the cfDNA fragments with length less than 150bp is 0.963, and the AUC of the single dimension model constructed based on the length distribution index of the cfDNA fragments in the 10kb window of the genome is 0.961. Although the prediction performance is reduced compared with the multi-dimension model constructed using all the liver cancer specific cfDNA fragment features, the accuracy is still more than 95%, which can realize the accurate determination of the liver cancer samples.
[0216] Table 14 Performance evaluation results of different dimension models
[0217]
[0218] Embodiment 3 A system for screening cancer specific cfDNA fragment feature combinations in peripheral blood based on whole genome sequencing
[0219] The embodiment provides a system for screening cancer specific cfDNA fragment feature combinations in peripheral blood based on whole genome sequencing, which comprises a data acquisition module, an analysis module, a storage module and an output module. The data acquisition module is used to acquire paired-end sequencing data of peripheral blood cfDNA of a sample group. The analysis module is used to obtain cancer specific cfDNA fragment feature combinations according to the paired-end sequencing data and the method of Embodiment 1 or Embodiment 2. The storage module is used to store the paired-end sequencing data of peripheral blood cfDNA of the sample group obtained by the data acquisition module and the cancer specific cfDNA fragment feature combinations obtained by the analysis module. The output module is used to output the cancer specific cfDNA fragment feature combinations obtained by the analysis module.
[0220] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application rather than limit the protection scope of the present application. For those skilled in the art, based on the above description and ideas, other different forms of changes or modifications can be made, which do not need to be or cannot be exhausted all the embodiments. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the claims of the present application.
Claims
1. A method for screening cancer-specific cfDNA fragment feature combinations in peripheral blood based on whole-genome sequencing, characterized in that, Includes the following steps: S1. Obtain paired-end sequencing data of peripheral blood cfDNA from the sample population, retain paired-end reads aligned to the human reference genome, and obtain cfDNA fragment information based on the paired-end reads; the sample population includes cancer patients and healthy individuals; S2. Using the paired-end reads obtained in step S1, the characteristic values of the following cfDNA fragments are obtained: The global length distribution index all-F-Index of cfDNA fragments, the 6bp frequency of the 5' end 6bp motif of cfDNA fragments (6bp-f), the coverage of cfDNA fragments in the 10kb window of the genome (Coverage_bin), and the length distribution index F-Index_bin of cfDNA fragments in the 10kb window of the genome. The characteristic value all-F-Index of the global length distribution index of the cfDNA fragment is calculated by formula (1); Formula (1): all - F - Index = ; Where i represents the length of the cfDNA fragment; P is a constant value, 160 bp to 170 bp; X i This represents the proportion of cfDNA fragments of length i among all cfDNA fragments; The 5' end 6bp motif frequency 6bp-f of the cfDNA fragment is the frequency of occurrence of each 6bp motif in the terminal motif set; the terminal motif set includes set 1 and set 2, set 1 consists of the 5' end 6bp motif of the nucleotide sequence of each cfDNA fragment; set 2 consists of the reverse complementary sequence of the 3' end 6bp motif of the nucleotide sequence of each cfDNA fragment. The 10kb genomic window is obtained by the following method: the genomic region with redundant regions removed is divided into non-overlapping 10kb intervals to obtain M 10kb genomic windows; the redundant regions include the short arm regions, telomere regions and centromere regions of 5 chromosomes, the 5 chromosomes being chromosomes 13, 14, 15, 21 and 22. The characteristic value Coverage_bin of the cfDNA fragment coverage of the 10kb window of the genome is calculated by the 10kb window of the genome and formula (2); Official (2): Coverage_bin k = ; Among them, Coverage_bin k This represents the coverage of the k-th 10kb-bin cfDNA fragment, where 1≤k≤M; N k This indicates the number of cfDNA fragments that were aligned to the kth 10kb-bin. N total This indicates the number of all cfDNA fragments aligned to the reference genome based on paired-end reads; The length distribution index F-Index_bin of the cfDNA fragment in the 10kb window of the genome is calculated by the 10kb window of the genome and formula (3); Official (3): F-Index_bin k = ; Where i represents the length of the cfDNA fragment; P is a constant value, 160 bp to 170 bp; X’ i This represents the proportion of cfDNA fragments of length i within the k-th 10kb window of the genome; S3. Generate a feature matrix using the cfDNA fragment feature values obtained in step S2, and perform feature screening to obtain feature combinations; the feature screening is based on the accuracy of cancer prediction, and the cancer prediction includes predicting whether the sample is a cancer patient.
2. The method according to claim 1, characterized in that, In step S2, the feature values of the cfDNA fragment feature further include: the 5' end 6bp motif short-6bp-f of cfDNA fragments with a length of less than 150bp, which is the frequency of occurrence of each 6bp motif in the short-end motif set; the short-end motif set is obtained by retaining the subset corresponding to all cfDNA fragments with a length of less than 150bp in the end motif set.
3. The method according to claim 1, characterized in that, In step S1, paired-end reads of chromosomes 1–22 aligned to the human reference genome are preserved.
4. The method according to claim 1, characterized in that, In step S3, the feature selection method includes the following steps: using the feature matrix as input features, and employing a wrapping method for feature selection.
5. The method according to claim 4, characterized in that, The feature selection method used in the packaging method includes recursive feature elimination.
6. The method according to claim 4, characterized in that, The model building methods used in the packaging method include random forest, support vector machine, gradient boosting machine or LASSO regression.
7. The method according to claim 1, characterized in that, The cancer in question is either lung cancer or liver cancer.
8. The use of the method according to any one of claims 1 to 7 in the preparation of cancer prediction products.
9. A method for constructing a cancer prediction model, characterized in that, The training set samples are processed using the method described in any one of claims 1 to 7 to obtain cancer-specific cfDNA fragment feature combinations; the cancer-specific cfDNA fragment feature combinations are used as input features, the cancer risk probability is used as the output result, and AUC is used as the performance index to train the classification model to obtain the cancer prediction model.
10. The construction method according to claim 9, characterized in that, The classification models include random forest, support vector machine, gradient boosting machine, or LASSO regression.
Citation Information
Patent Citations
Early cancer prediction method based on low-depth WGS sequencing end features
CN115910349A
Methods and systems for free DNA fragment size density to assess cancer
CN116157868A