Method for constructing early cancer prediction model based on open genomic regions and model

By constructing an early cancer prediction model based on open regions of the genome, and utilizing the number of cfDNA sequencing fragment endpoints and key community mining technology, the problems of insufficient signal specificity and low detection sensitivity in existing technologies are solved, and high-precision early cancer diagnosis is achieved.

CN115762746BActive Publication Date: 2026-02-10SHUANGLIANYUN (WUHAN) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210191818.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-02
Filing Date
2022-02-28
Publication Date
2026-02-10
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

Existing technologies using cfDNA for early cancer diagnosis suffer from insufficient signal specificity, significant DNA molecule damage caused by measurement methods, low detection sensitivity, and high fragment pattern noise, resulting in unsatisfactory diagnostic results.

Method used

We constructed an early cancer prediction model based on open regions of the genome. By collecting the number of endpoints of cfDNA sequencing fragments as features, we used the Poisson distribution test and convolutional neural network to identify open regions. We then combined mutual information and HotNet2 to mine key communities and built an integrated tumor risk prediction model.

Benefits of technology

It improves the predictive accuracy and sensitivity of early cancer diagnosis, reduces costs, and has broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762746B_ABST
    Figure CN115762746B_ABST
Patent Text Reader

Abstract

The application relates to a cancer early prediction model based on a genome open region, comprising the following steps: 1) obtaining all potential open region data in a genome; 2) collecting cfDNA sequencing data of plasma or urine of cancer patients and healthy people, and matching cfDNA sequencing fragments to the genome; 3) calculating the cfDNA end point number of each open region according to the open region data and the cfDNA sequencing fragments in the sample; 4) constructing a cancer early diagnosis model; and 5) diagnosing the open region according to the sequencing fragment end point number of the open region by using the early diagnosis model. The application uses the cfDNA sequencing fragment end point number in a body fluid as a feature to construct a tumor risk prediction model for cancer early diagnosis, the cfDNA sequencing data is easy to obtain, the cost required for sequencing is low, the prediction accuracy is high, and the application has a wide application prospect in the field of cancer early diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the medical field and relates to an early cancer prediction model based on open regions of the genome. Background Technology

[0002] Early cancer diagnosis is a diagnostic and treatment method specifically for patients in the early stages of cancer. Achieving early cancer diagnosis can alleviate patient suffering, psychological and economic burdens, and significantly improve prognosis. Liquid biopsy technology has attracted widespread attention due to its advantages such as low invasiveness and the presence of information from different tissue sources. Cell-free DNA (cfDNA) refers to DNA released into body fluids (leaving the cell) during apoptosis. Because the patterns of DNA released into body fluids (currently, most research focuses on blood) vary under different physiological conditions, it has very high scientific and clinical research value and is a popular research direction in the field of early cancer diagnosis.

[0003] Using cfDNA methylation for clinical diagnosis is one of the most researched and commercially relevant methods. However, DNA methylation information is affected not only by disease states such as cancer, but also by many other factors such as age and smoking status, resulting in insufficient specificity of the signals obtained from the data. Furthermore, bisulfite sequencing, the method used to measure DNA methylation data points, can damage the input DNA molecules, and its conversion efficiency is not high enough, further reducing the sensitivity of the data. Currently, methods using cfDNA methylation information for cancer diagnosis still have unsatisfactory diagnostic rates for early-stage cancer patients.

[0004] Using mutation information carried by cfDNA for corresponding clinical diagnosis is also a method that has attracted much attention. The main bottleneck of this type of method is that early-stage cancer releases relatively few DNA molecules into body fluids, and the proportion of tumor-specific mutations is also very small, making it difficult to detect tumor-specific mutations from DNA in body fluids. Current methods attempt to improve the sensitivity of early cancer diagnosis through deep coverage sequencing and simultaneous detection of multiple mutant molecules, but the results are still unsatisfactory. Earlier this year, Chinese scientists conducted a relatively systematic study on 10,000 Chinese patients and found that mutations could only be detected in 57.9% of stage I-III cancer patients.

[0005] Unlike cfDNA, which carries a low proportion of tumor-specific mutations, bodily fluids such as blood and urine contain a high number of cfDNA fragments. Furthermore, the fragment patterns (fragment length, fragment distribution across the genome, etc.) of cfDNA in samples from different pathological states vary, providing a basis for using cfDNA fragment patterns as features for cancer diagnosis. Researchers at Johns Hopkins School of Medicine divided the genome into adjacent 5MB gene regions and then used the proportion of short fragments in each region as features (the proportion of fragments <150bp; the fragment pattern in this method can be considered a variant of fragment length) for cancer diagnosis (this method is called DELFI). The combined predictive AUC reached 0.94 in seven cancer types, with a sensitivity of 0.73 at 98% specificity. However, the formation mechanism of cfDNA fragment patterns during apoptosis is highly complex and remains unclear. This introduces significant noise and uncertainty into the use of cfDNA for cancer screening. Moreover, this method cannot pinpoint specific regulatory regions to changes in fragment patterns, greatly affecting its biological interpretability and medical application prospects. Professor YMDennis Lo's research group at the Chinese University of Hong Kong discovered that the endpoints (end or start positions of fragments) of cfDNA in liver cancer patients tend to match at certain specific base positions, and this information can be used to predict liver cancer. However, the predictive performance of these methods still cannot meet practical needs.

[0006] Meanwhile, open regions of the genome, lacking protein protection, are more easily cleaved and digested as cfDNA sequencing fragments. Therefore, fragment endpoint information in open regions may be effective cancer diagnostic indicators, and using them as features for early cancer diagnosis could be a promising solution. Summary of the Invention

[0007] To address the aforementioned technical problems in the background art, this invention provides a method and prediction model for constructing an early cancer prediction model based on open regions of the genome. This invention uses the number of endpoints of cfDNA sequencing fragments in body fluids as a feature to construct a tumor risk prediction model for early cancer diagnosis. The cfDNA sequencing data of this invention is easy to obtain, the sequencing cost is low, and the prediction accuracy is high, thus having broad application prospects in the field of early cancer diagnosis.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] A method for constructing an early cancer prediction model based on open regions of the genome, comprising the following detailed steps:

[0010] 1) Based on the DNase peak data of white blood cells, immune cells and tissues in the database, obtain all potential open regions in the genome, and identify the open regions to obtain the identified open region data;

[0011] 2) Collect cfDNA sequencing data from plasma or urine of cancer patients and healthy individuals, and match the cfDNA sequencing data to the genome to form cfDNA sequencing fragments;

[0012] 3) Based on the identified open region data obtained in step 1) and the cfDNA sequencing fragments in the sample obtained in step 2), calculate the number of cfDNA endpoints in each open region, and mine key open regions in the cancer pathogenesis process based on the number of cfDNA endpoints.

[0013] 4) Construct an early cancer diagnosis model based on the number of cfDNA endpoints in key open regions during cancer development and the class label for each sample;

[0014] 5) For the sample to be diagnosed, the early diagnostic model from step 4) is used to make a diagnosis based on the number of sequencing fragment endpoints in its open region.

[0015] As a preferred embodiment, the specific implementation of step 1) of the present invention is as follows: according to the database, search for open region data of all cells and tissues that release a large amount of cfDNA into the blood, wherein the open region data are DNase peaks and ATAC-seq peaks; the database is Roadmap, encode and NCBI GEO; the open region data includes 1,381,219 open intervals.

[0016] Preferably, the specific implementation method for identifying the open area data in step 1) of the present invention is as follows:

[0017] 1.1) Based on the Poisson distribution test, all potential open region data in the genome are screened to obtain the screened open region data;

[0018] 1.2) The open region data obtained in step 1.1) is trained and validated based on the convolutional neural network to obtain the identified open region data.

[0019] As a preferred embodiment, step 2) of the present invention is specifically implemented as follows:

[0020] 2.1) Obtain cfDNA sequencing data from plasma of different cancer patients and multiple healthy individuals from the research group that published the paper in Nature;

[0021] 2.2) The pipeline is used to perform adapter trimming, genome matching and quality filtering on the cfDNA sequencing data, and finally form cfDNA sequencing fragments.

[0022] As a preferred embodiment, the specific implementation of step 3) in this invention is as follows:

[0023] For each sample, the number of cfDNA endpoints in each open region is counted, and the number of endpoints in all regions is standardized using z-score to obtain the number of open region fragment endpoints for each sample; the number of cfDNA endpoints is a 100bp interval upstream and downstream of the midpoint in the open region.

[0024] Preferably, the specific implementation of step 3) of this invention, which involves mining key open regions in the cancer pathogenesis process based on the number of cfDNA endpoints, is as follows:

[0025] 3.1) Predict open regions of cancer patients and healthy samples based on the number of cfDNA endpoints, merge all open regions to form candidate open regions, and the number of candidate open regions is at least greater than 2;

[0026] 3.2) Calculate the mutual information value between every two candidate intervals using a mutual information-based open network, and determine the correlation between the two candidate intervals based on the mutual information value;

[0027] 3.3) Based on the key community theory of HotNet2, the correlation of candidate intervals is analyzed to identify communities closely related to cancer occurrence. These communities are key open regions in the cancer pathogenesis process.

[0028] As a preferred embodiment, step 4) of the present invention is specifically implemented as follows:

[0029] For the key open regions in the cancer pathogenesis process obtained in step 3), multiple tumor risk prediction models are constructed using the bootstrapping strategy. The tumor risk prediction models with better performance are screened and integrated to form an early cancer diagnosis model.

[0030] An early cancer prediction model based on open regions of the genome, constructed using the method described above, is an early cancer prediction model based on open regions of the genome.

[0031] The application of early cancer prediction models based on open regions of the genome in early cancer prediction, as previously described.

[0032] Compared with existing technologies, the advantages of this invention are: it provides an early cancer prediction model based on open regions of the genome, which is easy to use, low in cost, and has high prediction accuracy. The prediction model of this invention has broad application prospects in the field of early cancer diagnosis. Attached Figure Description

[0033] Figure 1 This is a flowchart of an early cancer prediction model based on open regions of the genome, as described in this invention.

[0034] Figure 2 This is the ROC plot for cancer diagnosis in Embodiment 1 of the present invention.

[0035] Figure 3 The diagram shows the diagnostic results (sensitivity results) of cancer patients at different stages in Example 1 of the present invention.

[0036] Figure 4 This is the ROC plot for liver cancer diagnosis in Embodiment 2 of the present invention. Detailed Implementation

[0037] To better illustrate the purpose, technical solution, and advantages of the present invention, the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0038] See Figure 1 This invention provides a method for constructing an early cancer prediction model based on open regions of the genome. The specific method for constructing this model is as follows:

[0039] 1) Based on DNase peak data of white blood cells, immune cells, and tissues from databases such as Roadmap, all potential open regions in the genome were obtained. Genome open region identification was based on Poisson distribution test and convolutional neural network.

[0040] 1.1) Hotspot region screening based on Poisson distribution test. In MACS, a tool used to identify transcription factor binding sites in the genome, the Poisson distribution test is used to identify peaks in the genome. The main steps are as follows:

[0041] For each scanning window in the genome, the IFS (Integrated Fragmentation Score) is defined to describe the sequencing depth (number of fragments) and fragment length of that region, and its calculation formula is as follows:

[0042]

[0043] Where n is the number of fragments in the interval, leni is the length of the i-th sequencing fragment, Slen is the sum of the lengths of sequencing fragments in the entire genome, and Scoverage is the number of sequencing fragments in the entire genome. α is a value from 0 to 1. That is, IFS is a weighted value of the number of fragments in the interval and the fragment length (according to the definition of fragment length, it is a ratio based on the sum of the fragment lengths in the interval to the overall sequencing length distribution).

[0044] 1.2) Open Region Prediction Based on Convolutional Neural Network (CNN). Feature Extraction: For each hotspot region, extend a 200bp interval to the left and right from its center. Then, use its sequence information and the IFS value at each position (the IFS value of each bp is the sum of the IFS values ​​of all segments mapped from the center point to that position) as features, described as an 8×200 matrix. Each column represents one bp of information. For each bp, use an 8×1 vector to describe the sequence information and IFS information (the first four rows describe the sequence, and the last four rows describe the IFS values).

[0045] Convolutional Neural Network Architecture: A convolutional neural network with three hidden layers is proposed. The first and second layers are convolutional layers, and the third layer is a fully connected layer. Training and Validation: Training is performed using hotspot regions from single-numbered chromosomes, and testing is performed using double-numbered chromosomes. The evaluation metrics are AUC and Accuracy.

[0046] Identifying open regions in cancerous and normal samples is crucial for constructing diagnostic models and elucidating the pathogenesis of bladder cancer. Based on the biological hypothesis that open genomic regions, lacking histone protection, have more easily cleaved cfDNA fragments, a method was developed to identify cfDNA fragment omics models of genomic hotspot regions for predicting tumor risk. This method combines cfDNA fragment length and fragment quantity information of genomic regions. It was confirmed that these open hotspot regions contain both open regions and repeats, and these regions were used to construct cancer diagnostic models. Therefore, a key focus of this study is to identify corresponding genomic hotspot regions in urine samples based on cfDNA fragment patterns using the Poisson distribution test. Simultaneously, based on the differences in fragment patterns between open and repeat regions and genomic sequence information, a suitable convolutional neural network model was established to further distinguish between repeats and open regions.

[0047] 2) Based on cfDNA sequencing data from plasma / urine of cancer patients / healthy individuals, match cfDNA sequencing fragments to the genome.

[0048] Processing of cfDNA sequencing data. This project plans to use a mature pipeline to perform adapter trimming, genome matching, and quality filtering on the sequencing data. The main steps are: ① Trimmomatic (v0.36) is used to trim the adapters of the reads. ② BWA-MEM 0.7.15 is used to match the reads to the human reference genome. ③ Samblaster (v0.1.24) is used to mark the Polymerase Chain Reaction (PCR) repeats, and then samtools is used to delete reads and PCR repeats with low matching quality.

[0049] 3) Based on the open region data obtained in step 1) and the cfDNA sequencing fragments in the sample obtained in step 2), calculate the number of cfDNA endpoints in each open region. After identifying the open regions in the sample genome, key open regions in the cancer pathogenesis process are then identified as features for subsequent early cancer diagnosis models. Mutual information is widely used to calculate the correlation between discrete variables. Based on mutual information theory, the relationships between open regions are calculated to construct a co-open network; a biological network is established, combining the topological structure of network nodes (open regions) while considering the relationship between open regions and phenotypes, to identify high-weight key regions.

[0050] 3.1) Construction of a co-open network based on mutual information. For cfDNA sequencing data of healthy samples with high sequencing depth and cfDNA sequencing data of bladder cancer samples, the open regions of bladder cancer patients and healthy samples were predicted using the methods described above. All open regions were merged as candidate open regions.

[0051] For every two candidate intervals, their mutual information value is calculated as their correlation. Mutual information (MI) is a concept in information theory. Mutual information I(X, Y) represents the mutual information between variables X and Y, defined as follows:

[0052]

[0053] 3.2) Key Community Mining Based on HotNet2. A community refers to a dense subnetwork of vertices in a network that are closely connected to each other but have fewer connections to other vertices. In biological networks, a community is generally considered to interact and become a basic functional unit. If communities closely related to cancer development can be identified, these communities can serve as features in cancer diagnostic models, and the open regions contained within them may be key open regions in the cancer development process.

[0054] 4) Based on the number of cfDNA endpoints in open regions of urinary cfDNA samples and the class label for each sample, an early cancer prediction model was constructed. Using a bootstrapping strategy to build multiple tumor risk prediction models and integrating those with better performance into a diagnostic model will improve the performance and robustness of the diagnostic model. One aspect of this study is to construct a cfDNA-based cancer risk prediction model with high classification accuracy and generalization based on the characteristics of open region data from urinary cfDNA using a bootstrapping strategy.

[0055] Nodes in a dense subnetwork may work together to participate in a biological process, and multiple subnetworks may comprehensively describe the biological mechanisms involved in disease pathogenesis. Based on this biological hypothesis, this project aims to use a bootstrapping strategy to construct an ensemble tumor risk prediction model as a cancer diagnostic model. For each key community, samples are sampled from the training set using the bootstrapping strategy. Then, the IFS (z-score normalized) of the key open intervals in that community is used as a feature to construct a decision tree. This process is repeated M times (e.g., 10 times). Based on its predictive performance in the remaining samples, a weak tumor risk prediction model with some predictive ability is saved.

[0056] After the process of building weak tumor risk prediction models for each community is completed, all weak tumor risk prediction models are combined into the final integrated tumor risk prediction model according to the majority voting strategy.

[0057] 5) For the sample to be tested for risk prediction, risk prediction is made based on the number of sequencing fragment endpoints in its open region using the early risk prediction model in step (4).

[0058] Furthermore, this invention also provides an early cancer prediction model based on open regions of the genome, constructed using the prediction method of this model, and the application of this early cancer prediction model based on open regions of the genome in early cancer prediction.

[0059] The flowchart for early cancer prediction described in this invention is as follows: Figure 1 As shown. Unless otherwise specified, the experimental methods used in the embodiments are conventional methods, and the materials and reagents used are commercially available unless otherwise specified.

[0060] Example 1: Cancer diagnosis using the method of the present invention on a pan-cancer dataset

[0061] I. Collecting potential open regions of the genome

[0062] Find open region data (DNase peaks and ATAC-seq peaks) of all cells and tissues (all types of white blood cells, immune cells, tissues, etc.) that release a large amount of cfDNA into the bloodstream (including Roadmap). https: / / egg2.wustl.edu / roadmap / data / byFileType / peaks / consolidated / ;Blueprint : http: / / dcc.blueprint-epigenome.eu / # / files; encode: https: / / www.encodeproject.org / experiments NCBI GEO: https: / / www.ncbi.nlm.nih.gov / geo (Access Number: GSE118189, GSE74912, GSM2400294)) A total of 423 datasets were obtained. The open intervals in these 423 datasets were merged (adjacent open intervals were merged into one), resulting in a total of 1,381,219 open intervals.

[0063] II. Collect cfDNA sequencing data and class label information from pan-cancer samples and healthy samples.

[0064] The research group that published their paper in Nature in 2019 (Nature 2019, 570:385–389) obtained cfDNA sequencing data from plasma of 208 cancer patients (breast cancer, bile duct cancer, lung cancer, pancreatic cancer, colon cancer, ovarian cancer, and gastric cancer) and 215 healthy individuals (Illumina HiSeq 2000 / 2500, Pair-end). For each sample, its sequencing data was matched to the genome (hg19) using BWA.

[0065] III. Calculate the number of open region endpoints in cancer / healthy samples

[0066] For each sample, the number of cfDNA endpoints in each open region (a 100bp interval upstream and downstream of the midpoint of the open region) is counted, and the number of endpoints in all regions is standardized (z-score is used in this invention) to obtain the number of open region fragment endpoints for each sample.

[0067] IV. Constructing Cancer Diagnostic Models

[0068] All samples are divided into training and test sets (this invention uses 10x cross-validation). In the training set, a cancer diagnosis model is constructed using a support vector machine (default parameters, linear kernel).

[0069] V. Predictive Performance Evaluation

[0070] Based on 10-fold cross-validation, the AUC (area under the ROC curve) was used to evaluate the predictive performance of the cancer diagnostic model, while the sensitivity of the tumor risk prediction model at high specificity was calculated. The ROC curves for the prediction results in Example 1 are attached. Figure 2As shown in the figure, the AUC of this pan-cancer (including 7 types of cancer) diagnostic model is 0.9648 (95% CI: 0.9472-0.9824). Meanwhile, at 100% specificity, the model's sensitivity reaches 0.8181 (95% CI: 0.7376-0.8986); and at a specificity of 0.95, its specificity reaches 0.8895 (95% CI: 0.8091-0.9700). To further evaluate the performance of this predictive model in patients with early-stage cancer, the sensitivity (100% specificity) results for patients at different stages are shown in the appendix. Figure 3 The results show that even in stage I patients, the sensitivity of this model can still reach 0.8352 (95CI: 0.7007-0.9697).

[0071] Example 2: Early diagnosis of liver cancer using the method of the present invention

[0072] The steps in this embodiment are the same as in Embodiment 1, and the other steps are as follows:

[0073] II. Collect cfDNA sequencing data and class label information from liver cancer samples / healthy samples

[0074] The group that published their paper in PNAS in 2015 (Proc. Natl. Acad. Sci. 2015, 112: E1317–E1325) obtained cfDNA sequencing data from plasma of 90 liver cancer patients and 32 healthy individuals (Illumina HiSeq 2000, Pair-end). For each sample, the sequencing data was matched to the genome (hg19) using BWA.

[0075] III. Calculate the number of open region endpoints in cancer / healthy samples

[0076] Step three in this embodiment is the same as in embodiment 1.

[0077] IV. Constructing Cancer Diagnostic Models

[0078] Step four in this embodiment is the same as in embodiment 1.

[0079] V. Predictive Performance Evaluation

[0080] Based on 10-fold cross-validation, the AUC (area under the ROC curve) was used to evaluate the predictive performance of the cancer diagnostic model, while the sensitivity of the tumor risk prediction model at high specificity was also calculated. The ROC curves for the prediction results in Example 2 are attached. Figure 4As shown in the figure, the AUC of this diagnostic model is 0.9704 (95% CI: 0.9401-1.0000). Meanwhile, at 100% specificity, the sensitivity of this model can reach 0.9667 (95% CI: 0.9334-0.9999).

[0081] Table 1 Comparison of data from this method in hepatocellular carcinoma and pan-cancer.

[0082]

[0083] Table 2 shows the analysis results of the data in the paper by Cristiano S. et al.

[0084]

Claims

1. A method for constructing an early cancer prediction model based on open regions of the genome, characterized in that: The method includes the following steps: 1) Based on DNase peak data from leukocytes, immune cells, and tissues in the database, obtain all potential open regions in the genome, and identify these open regions to obtain identified open region data. Specifically, this involves searching the database for open regions of cells and tissues with a high amount of cfDNA released into the bloodstream. These open region data are DNase peaks and ATAC-seq peaks. The database includes Roadmap, Encode, and NCBI GEO. The open region data includes 1,381,219 open regions. The specific implementation method for identifying these open region data is as follows: 1.1) Based on the Poisson distribution test, all potential open region data in the genome are screened to obtain the screened open region data; 1.2) The open region data obtained in step 1.1) is trained and validated based on a convolutional neural network to obtain the identified open region data; 2) Collect cfDNA sequencing data from plasma or urine of cancer patients and healthy individuals, and match the cfDNA sequencing data to the genome to form cfDNA sequencing fragments; 3) Based on the identified open region data obtained in step 1) and the cfDNA sequencing fragments in the sample obtained in step 2), calculate the number of cfDNA endpoints in each open region, and mine key open regions in the cancer pathogenesis process based on the number of cfDNA endpoints; specifically: for each sample, count the number of cfDNA endpoints in each open region, and use z-score to standardize the number of endpoints in all regions to obtain the number of open region fragment endpoints for each sample; the number of cfDNA endpoints is a 100bp interval upstream and downstream of the midpoint of the open region; the specific implementation method for mining key open regions in the cancer pathogenesis process based on the number of cfDNA endpoints is as follows: 3.1) Predict open regions of cancer patients and healthy samples based on the number of cfDNA endpoints, merge all open regions to form candidate open regions, and the number of candidate open regions is at least greater than 2; 3.2) Calculate the mutual information value between every two candidate intervals using a mutual information-based open network, and determine the correlation between the two candidate intervals based on the mutual information value; 3.3) Based on the key community theory of HotNet2, the correlation of candidate intervals is analyzed to identify communities closely related to cancer occurrence. These communities are key open regions in the cancer pathogenesis process. 4) Construct an early cancer diagnosis model based on the number of cfDNA endpoints in key open regions during cancer development and the class label for each sample; 5) For the sample to be diagnosed, the early diagnostic model in step 4) is used to make a diagnosis based on the number of sequencing fragment endpoints in its open region.

2. The method for constructing an early cancer prediction model based on open regions of the genome according to claim 1, characterized in that: The specific implementation method of step 2) is as follows: 2.1) Obtain cfDNA sequencing data from plasma of different cancer patients and multiple healthy individuals from the research group that published the paper in Nature; 2.2) The pipeline is used to perform adapter trimming, genome matching and quality filtering on the cfDNA sequencing data, and finally form cfDNA sequencing fragments.

3. The method for constructing an early cancer prediction model based on open regions of the genome according to claim 2, characterized in that: The specific implementation method of step 4) is as follows: For the key open regions in the cancer pathogenesis process obtained in step 3), multiple tumor risk prediction models are constructed using the bootstrapping strategy. The tumor risk prediction models with better performance are screened and integrated to form an early cancer diagnosis model.

4. A method for constructing an early cancer prediction model based on open regions of the genome as described in any one of claims 1-3, resulting in an early cancer prediction model based on open regions of the genome.

5. The application of the early cancer prediction model based on open regions of the genome as described in claim 4 in early cancer prediction.

Citation Information

Patent Citations

  • Variant based disease diagnostics and tracking

    CN108603234A

  • Method and device for identifying chromatin open region based on sequencing data

    CN111724860A