Method for identifying open chromatin regions based on cfDNA whole genome sequencing data, cancer prediction model and system
By identifying open regions of chromatin and constructing a cancer prediction model through fragment dispersion analysis, this approach addresses the insufficient sensitivity of existing cfDNA fragment patterns in early cancer diagnosis, achieving high-precision early cancer diagnosis.
Patent Information
- Application Number
- CN202310639854.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-05-31
AI Technical Summary
Existing cfDNA fragment patterns lack sensitivity in early cancer diagnosis and cannot accurately depict chromatin accessibility in specific regions of the genome, thus affecting biological interpretability and medical application prospects.
We employed fragment dispersion analysis to identify open chromatin regions by calculating fragment dispersion in cfDNA whole-genome sequencing data, and constructed a cancer prediction model based on fragment dispersion. We then used classifiers such as support vector machines to predict cancer.
It improves the accuracy and precision of early cancer diagnosis, especially demonstrating high predictive accuracy and low cost in pan-cancer and single cancer prediction, and is suitable for low-cost sequencing of blood and urine samples.
Smart Images

Figure CN116665784B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, specifically to a method for analyzing cfDNA fragment patterns, a method for identifying open chromatin regions based on cfDNA whole-genome sequencing data, a cancer prediction model based on fragment dispersion and its construction method, and a cancer prediction system. Background Technology
[0002] Early diagnosis of cancer is crucial for its prognosis, but the discovery of non-invasive biomarkers with ultra-high specificity (>99%) and high sensitivity remains a significant challenge. Liquid biopsy technology has attracted widespread attention due to its low invasiveness and ability to provide comprehensive information. Cell-free DNA (cfDNA) profiling has been established for detecting tissue rejection after solid organ transplantation, non-invasive prenatal testing for fetal aneuploidy during pregnancy, and non-invasive tumor diagnosis. While using cfDNA mutation information or cfDNA methylation information for early cancer diagnosis is currently the mainstream method, its main bottleneck lies in its insufficient sensitivity in early-stage cancer patients.
[0003] In recent years, cfDNA fragment patterning has received increasing attention as an emerging technology. Circulating cell-free DNA (cfDNA) molecules in plasma are largely produced during cell death in various tissues throughout the body. These fragments carry a wealth of information about the human body, particularly information about nucleosome location, and are closely related to open chromatin and even gene expression. Whole-genome sequencing can obtain an individual's cfDNA fragment profile. Some studies have identified specific open chromatin fragment patterns, characterized by lower sequencing coverage and the absence of nucleosomes near transcription start sites (TSS).
[0004] Currently, there are several large-scale cfDNA fragment patterns. For example, the DELFI method proposed by Stephen Cristiano et al. in 2019 measures the proportion of short fragments within every 5MB of the genome. This method has achieved good classification results in seven types of cancer and is widely considered a classic method. However, this method does not pinpoint fragment pattern information to specific regulatory regions, which greatly affects its biological interpretability and medical application prospects. Another type is fragment patterns located in specific regions. In 2016, Snyder et al. proposed a WPS (Windowed Protected Score) fragment pattern, which is the ratio of the number of endpoints within a region to the number of fragment centers. The lower the score, the less protected the region is. This is a classic method in regional fragmentation patterns, but its cancer diagnostic performance has not been validated, and accurate nucleosome inference largely depends on deep sequencing. In 2019, Sun Kun et al. proposed an OCF fragmentation pattern, which calculates the difference between the start and end points of a fragment at a specific location within a region. A larger OCF value indicates a more open region. This method relies on known open chromatin regions. In 2022, Mohammad Shahrokh Esfahani et al. proposed a PFE fragmentation pattern, where the PFE value of a region is the Shannon entropy of that region's fragment length, used to measure the expression of known genes, called promoter fragment entropy. Its drawback is that it relies on known genes and ignores information about potential unknown regulatory regions. In 2022, Zhou Xiong Hui et al. proposed an IFS fragmentation pattern, whose signal mainly comes from low-coverage regions. A lower IFS value indicates a more open region, and this work used IFS values to identify open regions of the genome. However, low sequencing depth regions may originate from sequencing bias, meaning they still cannot accurately describe the chromatin openness of specific regions of the genome.
[0005] A new cfDNA fragment pattern is needed to accurately depict the openness of chromatin regions and effectively mine open regions in both cancer patients and healthy individuals, enabling precision medicine and early diagnosis of cancer. Summary of the Invention
[0006] To address the problems existing in the background technology, this invention provides a novel method for analyzing cfDNA fragment patterns—fragment dispersion—and a method for identifying open chromatin regions based on cfDNA whole-genome sequencing data, a cancer prediction model based on fragment dispersion and its construction method, and a cancer prediction system. Based on fragment dispersion, it is possible to better mine open chromatin regions, using this feature for early cancer diagnosis with high prediction accuracy.
[0007] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0008] In a first aspect, the present invention provides a method for analyzing cfDNA fragment patterns, comprising the following steps:
[0009] Based on the location of the sample cfDNA whole-genome sequencing data on the human reference genome, and combined with the interval division, the fragment dispersion (FDI) of that interval is calculated. The method for calculating the fragment dispersion is as follows:
[0010]
[0011] Where n is the total number of cfDNA endpoints in the interval, and for the j-th endpoint, the total number of endpoints within a 20bp radius around it is m. j Std(coverage) is the standard deviation of segment coverage in this interval.
[0012] Secondly, this invention provides a method for identifying open chromatin regions based on cfDNA whole-genome sequencing data, comprising the following steps:
[0013] (1) Match the whole genome sequencing data of sample cfDNA to the human reference genome to obtain the location information of cfDNA on the genome;
[0014] (2) Based on the interval division, calculate the fragment dispersion (FDI) of each interval. The method for calculating the fragment dispersion is as follows:
[0015]
[0016] Where n is the total number of cfDNA endpoints in the interval, and for the j-th endpoint, the total number of endpoints within a 20bp radius around it is m. j Std(coverage) is the standard deviation of segment coverage in this interval;
[0017] (3) Select regions in the genome with a fragment dispersion higher than a certain threshold as discrete regions, i.e. open chromatin regions.
[0018] Furthermore, in step (2), the whole genome is scanned with intervals of 120-400 bp and 15-25 bp, and the fragment dispersion of all intervals is calculated.
[0019] Furthermore, step (3) uses the beta distribution test to screen and obtain discrete regions.
[0020] Further, the specific steps of step (3) are as follows: fit the fragment dispersion data of all intervals with beta distribution, filter the regions with high fragment dispersion in the current chromosome and the local 2000-10000bp region with p value and FDR threshold, merge the overlapping regions, and filter to obtain the final discrete region, namely the chromatin open region.
[0021] Thirdly, the present invention also provides a method for constructing a cancer prediction model based on fragment discreteness, comprising the following steps:
[0022] (1) The steps to obtain the training set include obtaining whole-genome sequencing data of cfDNA from cancer patients and healthy patients, and using the above-mentioned method of identifying open chromatin regions based on whole-genome sequencing data of cfDNA to determine all discrete regions and fragment dispersion of each sample.
[0023] (2) Construction steps, including using a classifier to build a cancer prediction model based on the fragment dispersion and class label of the discrete region of each sample.
[0024] Furthermore, it also includes the process of performance testing of the aforementioned cancer prediction model using a validation set.
[0025] Furthermore, the aforementioned cancer prediction model is either a pan-cancer prediction model or a single cancer prediction model, and the aforementioned cancer population includes patients with any one or more of the following: breast cancer, liver cancer, bile duct cancer, colorectal cancer, stomach cancer, lung cancer, ovarian cancer, and pancreatic cancer.
[0026] Fourthly, the present invention also provides a cancer prediction model based on fragment discreteness constructed by the above construction method.
[0027] Fifthly, the present invention provides a cancer prediction system, comprising:
[0028] The alignment module is used to match the cfDNA genome sequencing data of healthy human samples to the human reference genome to obtain the location information of cfDNA in the genome; it is also used to match the cfDNA genome sequencing data of cancer human samples to the human reference genome to obtain the location information of cfDNA in the genome.
[0029] The calculation module is used to calculate the fragment dispersion of discrete regions in healthy individuals and cancer patients respectively, based on the above-mentioned method for identifying open chromatin regions based on cfDNA whole-genome sequencing data.
[0030] The modeling module is used to build a cancer prediction model using a classifier based on the fragment dispersion and class label of each sample's discrete region.
[0031] The diagnostic prediction module is used to predict whether a sample has cancer based on a cancer prediction model.
[0032] The beneficial effects of this invention are:
[0033] 1. This invention defines a novel cfDNA fragment pattern—fragment dispersion—to describe the openness of genomic regions. A statistical model is used to identify regions with high fragment dispersion as discrete regions, i.e., open chromatin regions. Validation was performed using conserved open regions common to B cells, T cells, and monocytes in plasma. Conserved open chromatin regions exhibit high fragment dispersion, while surrounding non-open chromatin regions show lower fragment dispersion. Compared to six other existing fragment patterns, fragment dispersion shows the highest correlation with gene expression. Furthermore, the open chromatin regions identified using this method are significantly enriched with known open chromatin regions, indicating that fragment dispersion can reflect chromatin openness in the genome, and compared to other cfDNA fragment patterns, it best reflects the chromatin openness of specific genomic regions.
[0034] 2. This invention presents a novel cancer prediction model based on fragment dispersion. This model demonstrates high prediction accuracy in both pan-cancer and single-cancer prediction, while also being low-cost and easy to use. This cancer prediction model has broad application prospects in the early prediction and diagnosis of cancer. Attached Figure Description
[0035] Figure 1 This is a fragment dispersion distribution diagram of the BH01 sample in the open region and its surrounding + / -2000bp region in Embodiment 2 of the present invention;
[0036] Figure 2 The correlation coefficients between fragment dispersion and other fragment patterns (OCF, number of fragments (sum(reads)), IFS, fragment length Shannon entropy (PFE), WPS) and gene expression in Example 2 of this invention;
[0037] Figure 3 This refers to the enrichment results of open regions and regulatory elements mined from the fragment dispersion in the healthy sample BH01 in Embodiment 2 of the present invention, wherein... Figure 3 a, 3b, 3c, and 3d are the enrichment results of the mined open regions and TSS, CTCF, TTS, and random regions, respectively;
[0038] Figure 4 This is the cross-validation result of the pan-cancer prediction model in Embodiment 3 of the present invention;
[0039] Figure 5 This is the validation result of the breast cancer prediction model in Embodiment 3 of the present invention, wherein Figure 5 'a' represents the result of cross-validation. Figure 5 b represents the result of independent verification;
[0040] Figure 6 This is the validation result of the liver cancer prediction model in Embodiment 3 of the present invention, wherein Figure 6 'a' represents the result of cross-validation. Figure 6 b represents the result of independent verification. Detailed Implementation
[0041] The principles and features of the present invention are described below with reference to the accompanying drawings and specific embodiments. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0042] Because open chromatin lacks the protection of nucleosomes, fragments are more likely to be randomly cleaved, resulting in a more even distribution of fragments in open regions. Based on this, the inventors proposed a new cfDNA fragment pattern—fragment dispersion—to describe the degree of openness of a region. Using this pattern, they identified open regions in chromatin and, consequently, proposed a cancer diagnostic model.
[0043] Example 1: A Novel cfDNA Pattern—Fragment Dispersion
[0044] The whole genome sequencing data of cfDNA from the sample was obtained and matched to the human reference genome to obtain the location information of cfDNA in the genome.
[0045] cfDNA whole-genome sequencing data consists of cell-free DNA sequencing data from bodily fluid samples, such as blood and urine. These two fluids are readily available and have a high content of cell-free DNA. Sequencing data can be generated using any sequencing platform that meets cfDNA quality requirements, but paired-end sequencing is required. cfDNA whole-genome sequencing data can come from public databases or from laboratory sequencing data. The cfDNA sequencing data undergoes impurity removal and quality control based on the original test data. This includes removing adapter and end sequences from the original sequencing data, labeling repetitive reads during sequencing, and using these to calculate alignment rate, repetition rate, and genome coverage to filter out reads that did not align, were of low quality, or were labeled with repetitions.
[0046] The matching method can be either BWA or Bowtie, with BWA being preferred. The human reference genome can be g19, hg38, GRCH37, GRCH38, etc., and the embodiments in this invention are based on GRCh37.
[0047] For cfDNA sequencing data, based on the interval division, the Fragmentation Dispersed Index (FDI) of that interval is calculated. The definition of Fragmentation Dispersed Index is:
[0048]
[0049] n is the total number of cfDNA endpoints in this interval. For the j-th endpoint, the total number of endpoints within a 20bp radius around it is m.j Std(coverage) is the standard deviation of segment coverage in this interval.
[0050] Based on the defined interval, the standard deviation is used to calculate the change in coverage of all base points (bp) within that interval, considering the distribution of the center point within that interval.
[0051]
[0052] N is the interval length, x i Let be the coverage of the i-th bp in the interval, and u be the mean coverage of 200bp in the interval.
[0053] Based on the structure of nucleosomes, interval lengths N of 120-400 bp can effectively describe the openness of the intervals. During whole-genome scanning, the interval spacing can be set to any value between 15-25 bp. In this embodiment of the invention, whole-genome scanning is preferably performed with intervals of 200 bp and intervals of 20 bp.
[0054] Example 2: Identification of Open Chromatin Regions Based on cfDNA Whole Genome Sequencing Data
[0055] Following the steps in Example 1, the fragment dispersion of all intervals of the sample cfDNA was obtained. After z-score normalization, the fragment dispersion of each sample's discrete region was obtained. The beta distribution test was used to screen regions with high fragment dispersion as discrete regions using p-value and FDR threshold (using p-value of 0.05 for screening, then FDR correction of p-value, and screening with threshold of 0.05). To reduce false positive results, regions that were significantly higher than the entire chromosome and the local 2000-10000 bp of the region were selected (using p-value of 0.05 for screening). Regions that were difficult to match (fragments with a mean mappability value below 0.9), regions without fragment coverage, and regions in the dark region were excluded. Overlapping regions were merged to obtain the final discrete region, i.e., the open chromatin region.
[0056] This embodiment uses the above method to mine open regions of chromatin in healthy plasma samples.
[0057] The cfDNA sequencing data used in this embodiment comes from the following sources:
[0058] Download high-depth cfDNA sequencing data (BH01) of plasma samples provided in Synder's (Synder et al. Cell 2016) paper from NCBI SRA.
[0059] Download RNA-seq sequencing data of healthy plasma samples described by Esfahani (Esfahani et al. Nature Biotechnology 2022) from NCBI GEO.
[0060] Download open region data for B cells, T cells, and monocytes from the RoadMap database.
[0061] Download the genome annotation information (TSS, TTS, etc.) for GRCh37 from GeneCode (https: / / www.gencodegenes.org / human / release_30lift37.html).
[0062] CTCF location information of GRCh37 was obtained from the Suresh paper (Suresh et al. Genome Research 2009).
[0063] (1) Calculate the fragment dispersion near open regions of chromatin.
[0064] Open regions common to B cells, T cells, and monocytes were defined as conserved open regions in plasma data. The method of this invention was used to calculate the dispersion of cfDNA fragments in these regions and the surrounding + / - 2000 bp in BH01 samples. The results are as follows: Figure 1 As shown in the figure, conserved open chromatin regions have high fragment dispersion, while other regions have low fragment dispersion. This indicates that the fragment dispersion of cfDNA can reflect the openness of chromatin in the genome and can be used to mine open regions.
[0065] (2) The correlation between the fragment dispersion of gene transcription sites and the gene expression of the corresponding genes.
[0066] Fragment dispersion (FDI) was calculated near the transcription start site (TSS) ([50bp, 150bp]) in the BH01 sample, and Spearman correlation coefficients were calculated with the corresponding gene expression. Correlation coefficients between gene expression and six other fragment patterns (OCF, number of fragments (sum(reads)), IFS, fragment length Shannon entropy (PFE), WPS) were also calculated. The results are as follows: Figure 2 As shown, the correlation between fragment dispersion and gene expression is the highest, indicating that it best reflects the chromatin accessibility of specific regions of the genome.
[0067] (3) Discovering open regions in the genome
[0068] The open chromatin regions of the BH01 sample were obtained using the method described in this embodiment, and then enriched with known active regulatory elements (transcription start site (TSS), CCTCC binding factor (CTCF)) and inactive regions (transcription termination region (TTS), random region) in hg19. The results are as follows: Figure 3 As shown, the open regions mined from healthy plasma samples were significantly enriched with TSS and CTCF, but not with TTS and random regions. The results indicate that the fragment dispersion of the present invention can reflect the chromatin openness in the genome, and the mined open regions are indeed significantly enriched with known open regions.
[0069] Example 3: Constructing a Cancer Prediction Model Based on Fragment Discreteness
[0070] cfDNA sequencing data from cancer patients and healthy individuals were collected to obtain a training set. This included obtaining whole-genome sequencing data of cfDNA from cancer patients and healthy individuals. The cfDNA sequencing data from cancer patients were merged as cancer cfDNA sequencing samples, and the cfDNA sequencing data from healthy individuals were merged as healthy cfDNA sequencing samples. The discrete regions and fragment dispersion of each sample were determined using the method in Example 2. The discrete regions from the two types of samples were merged to calculate the union of the discrete regions. The fragment dispersion (FDI) of each sample in these discrete regions was used as a feature. Based on the fragment dispersion and class label of the discrete regions of each sample, a cancer prediction model was constructed using a classifier. The classifier model can be constructed using mainstream classifiers. In this example, a support vector machine with a linear kernel function is preferred.
[0071] The aforementioned cancer prediction models can be pan-cancer prediction models or single-cancer prediction models. Cancers include, but are not limited to, breast cancer, lung cancer, liver cancer, colorectal cancer, gastric cancer, bile duct cancer, pancreatic cancer, and ovarian cancer. When the data used includes multiple cancers and controls, the union of samples from multiple cancer patients is used to identify the pan-cancer open regions of cancer patients and the open regions of controls. The constructed cancer prediction model can be used for pan-cancer prediction and diagnosis. When the data used consists of sequencing data from specific cancer patients and controls, the constructed cancer prediction model is a prediction model for that specific cancer and is used for the prediction and diagnosis of that specific cancer.
[0072] (1) Construction of pan-cancer prediction model
[0073] From finaledb( http: / / finaledb.research.cchmc.org / Download plasma cfDNA sequencing data from 208 cancer samples (54 breast cancers, 26 bile duct cancers, 27 colorectal cancers, 27 gastric cancers, 12 lung cancers, 28 ovarian cancers, and 34 pancreatic cancers) and 215 control samples described in Cristiano's paper (Cristiano et al. Nature 2019).
[0074] Twenty cancer patients (two lung cancer samples, and three samples each of breast cancer, cholangiocarcinoma, colorectal cancer, gastric cancer, ovarian cancer, and pancreatic cancer) were merged into one pan-cancer cfDNA sequencing sample, and 20 healthy samples were merged into one healthy cfDNA sequencing sample. The discrete regions (open chromatin regions) and fragment dispersion of each sample were determined using the method in Example 2. The union of the discrete regions in the cancer samples and the healthy samples was calculated. The fragment dispersion of the remaining samples in these discrete regions (open chromatin regions) was calculated and used as features. Based on the feature matrix and class labels, a pan-cancer prediction model was constructed using a support vector machine. The diagnostic performance of this pan-cancer prediction model was evaluated using a 10-fold 10-fold cross-validation model, and the sensitivity of the pan-cancer prediction model under high specificity was also calculated.
[0075] The ROC curves for the cross-validation results are shown below. Figure 4 As shown in the figure, the AUC of this pan-cancer prediction model is 0.9447 (95% CI: 0.90-0.98), while the sensitivity at 100% specificity is 72%, indicating that the model has high diagnostic accuracy in pan-cancer diagnosis.
[0076] (2) Construction of breast cancer prediction model
[0077] From finaledb( http: / / finaledb.research.cchmc.org / Download the plasma cfDNA sequencing data of 54 breast cancer samples and 215 control samples described in Cristiano's (Cristiano et al., Nature 2019) paper, as dataset 1. Obtain the cfDNA sequencing data of 25 breast cancer samples and 25 control samples from Zhou's (Zhou et al., Genome Medicine 2022) article (Zenodo.org.DOI:10.5281 / zenodo.6914806), as dataset 2.
[0078] Ten-fold cross-validation was performed using 269 samples from Dataset 1. In each fold, 20 breast cancer samples and 20 control samples were selected to calculate discrete regions (open chromatin regions). The fragment dispersion of discrete regions (open chromatin regions) for each sample in the training set was calculated to construct a breast cancer prediction model, which was then validated on the test set. Simultaneously, a cancer prediction model was constructed across Dataset 1 and independently validated on Dataset 2.
[0079] The ROC curve of the 10-fold cross-validation results is as follows: Figure 5 As shown in figure a, the AUC = 0.9617, the sensitivity at 100% specificity is 77%, and the ROC curve for independent validation results is shown in figure a. Figure 5 As shown in b, the AUC = 0.9040 and the sensitivity at 100% specificity is 76%, indicating that the breast cancer prediction model has high diagnostic performance on both local and independent datasets.
[0080] (3) Construction of a liver cancer prediction model
[0081] Download the cfDNA sequencing data of 90 liver cancer samples and 30 controls described in Jiang's (Jianget.al.PNAS2015) paper from finaledb (http: / / finaledb.research.cchmc.org / ), as dataset 1. Obtain the cfDNA sequencing data of 8 liver cancer samples and 8 control samples from Zhou's (Zhou et al.Genome Medicine 2022) article (Zenodo.org.DOI:10.5281 / zenodo.6914806), as dataset 2.
[0082] Ten-fold cross-validation was performed using 120 samples from Dataset 1. In each fold, 20 liver cancer samples and 20 control samples were selected to calculate discrete regions (open chromatin regions). The fragment dispersion of the open chromatin regions for each sample in the training set was calculated to construct a liver cancer prediction model, which was then validated on the test set. Simultaneously, a liver cancer prediction model was constructed across Dataset 1 and independently validated on Dataset 2.
[0083] The ROC curve of the 10-fold cross-validation results is as follows: Figure 6 As shown in figure a, the AUC = 0.9657, the sensitivity at 100% specificity is 96%, and the ROC curve for independent validation results is shown in figure a. Figure 6 As shown in b, AUC = 1.0000, and the sensitivity at 100% specificity is 100%, indicating that the liver cancer prediction model has high diagnostic performance on both local and independent datasets.
[0084] In the cancer prediction model of this invention, the cfDNA sequencing data comes from bodily fluid samples such as blood and urine. These bodily fluid samples are easy to obtain, and the cfDNA sequencing depth is low (<2X), resulting in lower costs. It has high sensitivity with high specificity in pan-cancer and single cancer prediction and diagnosis, and has good prospects in the field of early cancer diagnosis.
[0085] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for identifying open chromatin regions based on cfDNA whole genome sequencing data, characterized in that, The method comprises the following steps: (1) matching sample cfDNA whole genome sequencing data to a human reference genome to obtain location information of cfDNA on the genome; (2) calculating fragment dispersion index (FDI) of each interval according to the location of sample cfDNA whole genome sequencing data on the human reference genome and interval division, and the calculation method of the fragment dispersion index is as follows: FDI = wherein n is the total number of cfDNA end points in the interval, for the jth end point, the total number of end points within 20 bp around it is , and Std(coverage) is the standard deviation of the fragment coverage in the interval. (3) screening regions with fragment dispersion index higher than a certain threshold in the genome as discrete regions, i.e. chromatin open regions.
2. The method of identifying open chromatin regions based on cfDNA whole genome sequencing data according to claim 1, wherein, In step (2), the whole genome is scanned with an interval of 120-400 bp and a spacing of 15-25 bp, and the fragment dispersion index of all intervals is calculated.
3. The method for identifying open chromatin regions based on cfDNA whole-genome sequencing data according to claim 1, characterized in that, In step (3), the discrete regions are screened by beta distribution test.
4. The method for identifying open chromatin regions based on cfDNA whole genome sequencing data according to claim 3, characterized in that, The specific steps of step (3) are as follows: fitting the fragment dispersion index data of all intervals with a beta distribution, screening regions with high fragment dispersion index in the current chromosome and the local 2000-10000 bp of the region with p value and FDR threshold, merging overlapping regions, and screening the final discrete regions, i.e. chromatin open regions.
5. A method for constructing a cancer prediction model based on fragment diversity, characterized by, The method comprises the following steps: (1) obtaining training set, including obtaining cfDNA whole genome sequencing data of cancer population samples and healthy population samples, and determining all discrete regions and fragment dispersion index of each sample by the method of any one of claims 1-4; (2) constructing step, including using a classifier to construct a cancer prediction model according to the fragment dispersion index of the discrete regions of each sample and the class label. 6.The method of claim 5, wherein, The process of measuring the performance of the cancer prediction model by a validation set is also included.
7. The method of constructing a fragment dispersion-based cancer prediction model according to claim 5 or 6, wherein, The cancer prediction model is a pan-cancer prediction model or a single cancer prediction model, and the cancer population suffers from any one or more of breast cancer, liver cancer, cholangiocarcinoma, colorectal cancer, gastric cancer, lung cancer, ovarian cancer and pancreatic cancer.
8. A cancer prediction model based on fragment dispersity, characterized in that, The method is constructed by any one of claims 5-7.
9. A cancer prediction system characterized by, It comprises: An alignment module for matching cfDNA genome sequencing data of healthy population samples to a human reference genome to obtain location information of cfDNA on the genome; for matching cfDNA genome sequencing data of cancer population samples to a human reference genome to obtain location information of cfDNA on the genome; A calculation module for calculating the fragment dispersion index of the discrete regions of the healthy population and the cancer population respectively according to the method of any one of claims 1-4; A modeling module for constructing a cancer prediction model using a classifier according to the fragment dispersion index of the discrete regions of each sample and the class label; A diagnosis prediction module for predicting cancer of a to-be-tested sample according to the cancer prediction model and determining whether the sample has cancer.