Cancer risk assessment model based on low-pass cfdna sequencing data and use thereof

By filtering and processing cell-free DNA sequencing data, specific cancer risk prediction biomarkers are obtained and assessment models are constructed, which solves the problem of insufficient low-depth sequencing cancer risk assessment methods in existing technologies and realizes low-cost and accurate cancer risk assessment and diagnosis.

WO2026153310A1PCT designated stage Publication Date: 2026-07-23OMIXSCIENCE (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
OMIXSCIENCE (SHENZHEN) CO LTD
Filing Date
2026-01-13
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

The existing technologies for cancer based on low-depth cell-free DNA sequencing have limited effectiveness in addressing the challenges of pan-cancer risk assessment. Furthermore, there are still few efficient and low-cost cancer risk assessment methods available, and there is significant room for improvement in the diagnostic performance of low-depth sequencing.

Method used

A method for obtaining cancer risk prediction biomarkers is proposed. By filtering the cell-free DNA sequencing data of subject samples, gene mutation map data, the proportion of mutation sites in regulators, the length distribution data of cell-free DNA, the base distribution data of broken ends, and the distribution data of ends in nucleosomes are obtained as cancer risk prediction biomarkers to construct a widely applicable cancer risk assessment model.

Benefits of technology

It achieves low-cost, high-accuracy pan-cancer diagnosis, can distinguish cancer patients from patients with similar conditions, reduces testing costs, simplifies experimental procedures, has strong applicability, and is suitable for risk assessment of various types of cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2026072242_23072026_PF_FP_ABST
    Figure CN2026072242_23072026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a cancer risk assessment model based on low-pass cfDNA sequencing data and the use thereof. A method for acquiring a marker of the cancer risk assessment model comprises: filtering cfDNA sequencing data from a subject sample to obtain cfDNA gene mutation data, then, on the basis of the cfDNA gene mutation data, obtaining at least one of gene mutation profile data, data on the fraction of mutation sites in regulons, cfDNA length distribution data, cfDNA terminal base distribution data, and data on the fraction of cfDNA termini in nucleosomes of the subject, and using the data as the cancer risk prediction marker. The risk assessment model constructed by using the cancer risk prediction marker has a good diagnosis accuracy and wide applicability, and can be used for the risk assessment of a variety of cancers.
Need to check novelty before this filing date? Find Prior Art

Description

A Cancer Risk Assessment Model Based on Low-Depth Cell-Free DNA Sequencing Data and Its Application Technical Field

[0001] This invention relates to the field of biomedical technology, and in particular to a cancer risk assessment model based on low-depth cell-free DNA sequencing data and its application. Background Technology

[0002] Circulating cell-free DNA (CFDNA) is a fragment of DNA that is naturally degraded after cell death. Studies have shown that in healthy individuals, CFDNA mainly originates from blood cells and the liver, while in cancer patients, dead tumor cells also release CFDNA into the peripheral blood. However, CFDNA fragments from tumors differ significantly from those from non-tumor cells, exhibiting variations such as gene mutations, chromosomal copy number variations, and abnormal methylation levels. Therefore, cancer can be diagnosed by detecting tumor-derived CFDNA signals in peripheral blood, such as by identifying specific biomarkers of tumor-derived CFDNA.

[0003] In the field of related technologies, numerous cancer diagnostic techniques have been developed based on cell-free tumor DNA signals in peripheral blood, such as detection kits for high-frequency gene mutations and detection kits for tumor-specific methylation markers. However, currently, there are still relatively few validated, pan-cancer, and universally applicable cell-free DNA-based cancer risk assessment or diagnostic methods. Furthermore, research indicates that gene mutation is a hallmark feature of tumors; however, the frequency of tumor-derived gene mutations in cell-free DNA is typically low (mostly below 5%). Clonal hematopoiesis can also induce gene mutations, and various errors exist in molecular experiments and sequencing. These factors collectively make it difficult to identify tumor gene mutations from cell-free DNA. This usually requires very high sequencing depths (e.g., 200 times the coverage of the human genome) and background filtering using the genotypes of normal cells (e.g., white blood cells), resulting in high detection costs and making widespread adoption in cancer screening or diagnosis almost impossible. Moreover, current diagnostic technologies based on low-depth cell-free DNA sequencing (e.g., 2-3 times the coverage of the human genome) are mostly based on fragment omics analysis, and their diagnostic performance still has significant room for improvement.

[0004] Therefore, there is an urgent need to find a low-cost and accurate cancer risk assessment method based on low-depth cell-free DNA sequencing. Summary of the Invention

[0005] This invention aims to address at least one of the technical problems existing in the prior art. To this end, this invention proposes a machine-executed method for obtaining cancer risk prediction biomarkers, which does not require paired normal cell (such as white blood cell) genotypes during sequencing, is simple to operate, has low requirements for initial materials, and the obtained cancer risk prediction biomarkers can help achieve pan-cancer diagnosis.

[0006] This invention also proposes biomarkers for predicting cancer risk.

[0007] This invention also proposes a method for constructing a cancer risk assessment model, which has wide applicability and can be used to construct risk assessment models for various types of cancer. Moreover, the model has good predictive accuracy and low cost.

[0008] This invention also proposes a cancer risk assessment model.

[0009] The present invention also proposes a computer-readable storage medium for predicting cancer risk.

[0010] The present invention also proposes a device for predicting cancer risk.

[0011] The present invention also proposes the application of the above-mentioned method for obtaining cancer risk prediction biomarkers, method for constructing cancer risk assessment models, cancer risk assessment models, and computer-readable storage media or devices for predicting cancer risk in the development or screening of cancer-related drugs or products.

[0012] A first aspect of the present invention provides a machine-executed method for obtaining cancer risk prediction biomarkers, comprising:

[0013] S1. Filter the cell-free DNA sequencing data of the subject samples to obtain cell-free DNA gene mutation data;

[0014] S2. Based on the free DNA gene mutation data, obtain at least one of the following data from the subject: gene mutation map data, percentage data of mutation sites in regulators, length distribution data of free DNA, base distribution data of free DNA ends, and distribution data of free DNA ends in nucleosomes, and use it as a cancer risk prediction biomarker.

[0015] The method for obtaining cancer risk prediction biomarkers according to embodiments of the present invention has at least the following beneficial effects:

[0016] The cancer risk prediction biomarkers of this invention are obtained based on low-depth cell-free DNA sequencing data. They do not require paired normal cell (e.g., white blood cell) genotypes during sequencing, making the experimental operation simple, with low requirements for initial materials. Furthermore, the obtained cancer risk prediction biomarkers contribute to pan-cancer diagnosis. In addition, the method for obtaining cancer risk prediction biomarkers of this invention has no special requirements for the chemical reagents and sequencing platforms used in the experiment, making it highly versatile.

[0017] In some embodiments of the present invention, the subject sample includes a blood sample, a urine sample, a tear sample, a sweat sample, a semen sample, or a saliva sample.

[0018] In some embodiments of the present invention, the sequencing depth of the free DNA sequencing data is less than 10 times the human genome coverage;

[0019] Preferably, the sequencing depth of the free DNA sequencing data is 2 to 5 times the coverage of the human genome.

[0020] In some embodiments of the present invention, the filtering process includes discarding SNP site data in the SNP database with an abundance of more than a first threshold.

[0021] In some embodiments of the present invention, the SNP database is a large population SNP database, including but not limited to dbSNP, NyuWa, ChinaMap, Korea1K, ToMMo, and gnomAD databases.

[0022] In some embodiments of the present invention, the first threshold value ranges from 0.05% to 1%. SNP sites with abundance above the first threshold are relatively common gene diversity sites in the population. They may be congenital germline variations rather than somatic mutations originating from tumors, and therefore generally do not carry cancer-related information. Filtering them helps improve the reliability of the analysis results. In some embodiments of the present invention, the first threshold value includes, but is not limited to, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, and 1%.

[0023] In some alternative embodiments of the invention, the filtering process further includes discarding data located in the “blacklist” area defined by the ENCODE project.

[0024] It is understandable that the "blacklist" defined by the ENCODE project refers to areas that are prone to causing technical errors during data analysis. For details, please refer to: Amemiya, HM, Kundaje, A., and Boyle, AP (2019). The ENCODE Blacklist: Identification of Problematic Regions of the Genome. Sci Rep 9,9354.10.1038 / s41598-019-45839-z.

[0025] In some alternative embodiments of the present invention, the filtering process further includes discarding mutation data within the second threshold range with the highest coverage.

[0026] In some embodiments of the present invention, the value of the second threshold ranges from 0.1% to 5%. In some embodiments of the present invention, the value of the second threshold includes, but is not limited to, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, and 5%.

[0027] In some preferred embodiments of the present invention, the filtering process includes discarding SNP site data with an abundance of 0.1% or higher in the SNP database.

[0028] In some preferred alternative embodiments of the invention, the filtering process further includes discarding the "blacklist" region defined by the ENCODE project and the 1% mutation data with the highest coverage.

[0029] In some embodiments of the present invention, the filtering process further includes preliminary detection of gene mutations, discarding data whose reading alignment mass number is below a third threshold and whose base mass number is below a fourth threshold;

[0030] Specifically, the third threshold is taken from any value between 20 and 80, and the fourth threshold is taken from any value between 20 and 50.

[0031] In some embodiments of the present invention, the value of the third threshold includes, but is not limited to, 20, 30, 40, 50, 60, 70, and 80.

[0032] In some embodiments of the present invention, the value of the fourth threshold includes, but is not limited to, 20, 30, 40, and 50.

[0033] In some embodiments of the present invention, the filtering process further includes: using BWA software for comparison, using BCFtools software for preliminary detection of gene mutations, and discarding data with a reading comparison quality number below 60 and a base quality number below 30.

[0034] In some embodiments of the present invention, the method for obtaining the gene mutation map data includes: obtaining mutation pattern data based on the free DNA gene mutation data, then obtaining the proportion data of each mutation pattern, and generating gene mutation map data.

[0035] In some embodiments of the present invention, the mutation pattern data includes 96 types.

[0036] Specifically, the method for obtaining the gene mutation map data includes: converting the "A" and "G" mutation sites in the free DNA gene mutation data to "T" and "C" bases respectively, and recording one base before and after the mutation site to obtain 96 mutation pattern data. Then, normalization processing is performed to obtain the proportion data of each mutation pattern and generate gene mutation map data.

[0037] In some embodiments of the present invention, the 96 mutation pattern data include 6 basic mutation types, namely C>A, C>G, C>T, T>A, T>C, and T>G.

[0038] In some embodiments of the present invention, the method for obtaining the percentage data of the mutation site in the regulator includes: obtaining the number of sites in the cell-free DNA gene mutation data that overlap with the blood cell regulator, and calculating the proportion of these overlapping sites to all mutation sites, thereby obtaining the percentage data of the mutation site in the regulator.

[0039] The regulator can be one or more of the following: a strengthener, a promoter, an insulator, etc.

[0040] In some embodiments of the present invention, the regulator is an enhancer, and the method for obtaining the proportion of mutation sites in the enhancer includes: obtaining the location of enhancers in blood cells, which can be obtained by performing ChIP-seq experiments on H3K27ac histone modification; taking out sites that overlap with enhancers from all mutation sites, calculating their number, and dividing them by the total number of mutation sites to obtain the proportion of mutation sites in the enhancer.

[0041] In some embodiments of the present invention, the method for obtaining the length distribution data of the free DNA includes: dividing the free DNA gene mutation data into Mut-DNA fragments carrying mutated genotypes and Wt-DNA fragments not carrying mutated genotypes, and calculating the proportion of base fragments with a length below the fifth threshold in the Mut-DNA fragment or the Wt-DNA fragment respectively, and then obtaining the difference between the two proportions (that is, calculating the proportion of base fragments with a length below the fifth threshold in the Mut-DNA fragment and the proportion of base fragments with a length below the fifth threshold in the Wt-DNA fragment respectively, and then obtaining the difference between the two proportions), thereby obtaining the length distribution data of the free DNA, denoted as Diff-size.

[0042] In some embodiments of the present invention, the fifth threshold value ranges from 50 to 160, and the unit is bp.

[0043] In some embodiments of the present invention, the value of the fifth threshold includes, but is not limited to, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, and 160, in units of bp.

[0044] In some preferred embodiments of the present invention, the fifth threshold value ranges from 100 to 150 as any positive integer, with the unit being bp.

[0045] In some embodiments of the present invention, the length calculation method of the Mut-DNA fragment or the Wt-DNA fragment adopts conventional calculation methods in the art, such as obtaining the locus information of the two ends of each free DNA in the genome and calculating the distance between them, thereby obtaining the length information of the free DNA.

[0046] In some embodiments of the present invention, the method for obtaining the base distribution data of the free DNA fragment ends includes: calculating the proportion of fragments in the Mut-DNA fragment or the Wt-DNA fragment where the two bases upstream and downstream of the 5' break point are "CTCC", and then obtaining the difference between the proportions to obtain the base distribution data of the free DNA fragment ends, denoted as Diff-CTCC.

[0047] The terminology for the breakpoint bases of the cell-free DNA refers to the sequence consisting of four bases: the locus information at the 5' end of each cell-free DNA fragment (Mut-DNA fragment or Wt-DNA fragment) in the genome, and the two bases upstream and two bases downstream of the locus. This sequence represents the breakpoint bases of the cell-free DNA. Furthermore, for the cell-free DNA set under consideration (e.g., Mut-DNA), the breakpoint base sequence of each cell-free DNA fragment is calculated, and the number of each sequence is counted. This count is then divided by the total number of DNA fragments in the set to obtain the proportion of that sequence in the set, thus providing the distribution data of the breakpoint base sequences in the cell-free DNA set.

[0048] In some embodiments of the present invention, the method for obtaining the distribution data of the ends of the free DNA in the nucleosome includes: calculating the proportion of the Mut-DNA fragment whose ends are located within the sixth threshold of the nucleosome center and the proportion of the Wt-DNA fragment whose ends are located within the seventh threshold of the nucleosome center, and then obtaining the difference between the proportions, which is the distribution data of the ends of the free DNA in the nucleosome, denoted as Diff-nucleosome.

[0049] In some embodiments of the present invention, the sixth threshold and the seventh threshold are independently taken from any positive integer from 50 to 160, and the unit is bp.

[0050] In some embodiments of the present invention, the value of the sixth threshold includes, but is not limited to, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, and 160, in units of bp.

[0051] In some embodiments of the present invention, the value of the seventh threshold includes, but is not limited to, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, and 160, in units of bp.

[0052] In some preferred embodiments of the present invention, the sixth threshold and the seventh threshold are independently taken from any positive integer from 50 to 100, and the unit is bp.

[0053] Taking blood cell samples as an example, the location of nucleosome centers in blood cells can be obtained using micrococcal nuclease sequencing (Mnase-seq) technology. For each cell-free DNA molecule, its two ends (the smaller locus is called the upstream end, or U end, and the larger locus is called the downstream end, or D end) are analyzed separately. For each site, the nearest nucleosome center is located, and the distance is calculated; if the end is upstream of the nucleosome center, it is recorded as a negative value; otherwise, it is recorded as a positive value. For the set of cell-free DNA samples under consideration, the distances between the U and D ends of each cell-free DNA molecule and the nearest nucleosome center are calculated, and the frequency of different distances is counted. This is then divided by the number of DNA molecules to obtain the distribution data of the ends of cell-free DNA within the nucleosomes.

[0054] In some embodiments of the present invention, step S2 further includes acquiring terminal base distribution data of free DNA and / or high-contribution tumor mutation imprinting data, wherein:

[0055] The method for obtaining the terminal base distribution data of the free DNA includes: calculating the proportion of fragments with "CCCA" bases at the 5' end in the Mut-DNA fragment or the Wt-DNA fragment respectively, and then obtaining the difference between the proportions to obtain the terminal base distribution data of the free DNA, denoted as Diff-CCCA.

[0056] The method for obtaining high-contribution tumor mutation imprint data includes: obtaining background mutation imprints based on cell-free DNA gene mutation data of non-cancer patients, then merging tumor-specific gene mutation imprints in the tumor database, and deconvolving the gene mutation map obtained based on the cell-free DNA gene mutation data of the subjects to obtain the data.

[0057] A second aspect of the present invention provides cancer risk prediction biomarkers, including at least one of: gene mutation map data, data on the proportion of mutation sites in regulators, data on the length distribution of cell-free DNA, data on the terminal base distribution of cell-free DNA, data on the broken base distribution of cell-free DNA, and data on the distribution of the ends of cell-free DNA in nucleosomes.

[0058] In some embodiments of the present invention, the method for obtaining cancer risk prediction biomarkers employs the machine-executed method for obtaining cancer risk prediction biomarkers as described in any of the first aspects.

[0059] A third aspect of the present invention provides a method for constructing a cancer risk assessment model, comprising:

[0060] Cell-free DNA sequencing data were obtained from cancer patient samples and non-cancer patient samples, and corresponding cell-free DNA gene mutation data were obtained through filtering.

[0061] Based on the free DNA gene mutation data, at least one of the following data is obtained from the cancer patient samples and the non-cancer patient samples: gene mutation map data, percentage data of mutation sites in regulators, length distribution data of free DNA, base distribution data of free DNA ends, and distribution data of free DNA ends in nucleosomes. This data is then used as feature values ​​for the machine learning algorithm to generate a cancer risk assessment model.

[0062] The construction method according to embodiments of the present invention has at least the following beneficial effects:

[0063] (1) The cancer risk assessment model construction method of the present invention is based on low-depth cfDNA sequencing data and does not require paired normal cell (such as white blood cell) genotypes, which helps to reduce detection costs. In addition, the cancer risk assessment model construction method of the present invention has wide applicability and can be used to construct various types of cancer risk assessment models, and the prediction accuracy based on the model is good.

[0064] (2) The cancer risk assessment model of the present invention is simple to construct, and the sampling is convenient and safe. The assessment model of the present invention has good versatility, does not have specific requirements for the chemical reagents and sequencing platforms used in the experiment, and compared with traditional tissue biopsy or imaging examination, the cancer risk assessment model of the present invention adopts non-invasive blood sample collection in the construction process, which has the advantages of being simple, fast and non-invasive, and helps to reduce patient suffering and medical risks.

[0065] (3) The prediction model obtained by the cancer risk assessment model construction method based on low-depth cell-free DNA sequencing data of the present invention has good prediction accuracy. Compared with the conventional cancer risk assessment model, which can only be used to distinguish between cancer patients and normal people, the model constructed by the present invention can accurately distinguish between cancer patients and patients with similar diseases (such as liver cancer and non-tumor liver diseases, or liver cancer and hepatitis B carriers).

[0066] In some embodiments of the present invention, the cancers include lung cancer, breast cancer, liver cancer, bladder cancer, acute myeloid leukemia, adrenocortical carcinoma, bladder urothelial carcinoma, mammary ductal carcinoma, mammary lobular carcinoma, cervical cancer, bile duct cancer, colorectal cancer, esophageal cancer, gastric cancer, glioblastoma multiforme, head and neck squamous cell carcinoma, renal chromophobe carcinoma, renal clear cell carcinoma, renal papillary cell carcinoma, low-grade glioma, mesothelioma, lung adenocarcinoma, lung squamous cell carcinoma, ovarian serous cystadenocarcinoma, pancreatic cancer, pheochromocytoma and paraganglioma, prostate cancer, sarcoma, cutaneous melanoma, testicular cancer, thymic carcinoma, thyroid cancer, uterine sarcoma, endometrial cancer, and uveal melanoma.

[0067] In some embodiments of the present invention, the sample includes a blood sample, a urine sample, a tear sample, a sweat sample, a semen sample, or a saliva sample.

[0068] In some embodiments of the present invention, the methods for obtaining the gene mutation map data, the proportion data of mutation sites in regulators, the length distribution data of free DNA, the base distribution data of the broken ends of free DNA, and the distribution data of the ends of free DNA in nucleosomes are the same as those in the first aspect.

[0069] In some embodiments of the present invention, the feature value further includes terminal base distribution data of cell-free DNA. The method for obtaining the terminal base distribution data of cell-free DNA includes: dividing the cell-free DNA gene mutation data into Mut-DNA fragments carrying the mutant genotype and Wt-DNA fragments not carrying the mutant genotype, and calculating the proportion of fragments with the "CCCA" base at the 5' end in the Mut-DNA fragment or the Wt-DNA fragment respectively, and then obtaining the difference between the proportions, which is the terminal base distribution data of cell-free DNA, denoted as Diff-CCCA.

[0070] In some embodiments of the present invention, the feature value further includes high-contribution tumor mutation imprinting data, and the method for obtaining the high-contribution tumor mutation imprinting data includes:

[0071] Background mutation imprints are obtained from cell-free DNA gene mutation data of non-cancer patients, then tumor-specific gene mutation imprints from the tumor database are merged, and the gene mutation map obtained from the cell-free DNA gene mutation data of the subjects is deconvolved to obtain the final result.

[0072] In some embodiments of the present invention, the machine learning algorithm includes the Gradient Boosting Machines (GBM) algorithm.

[0073] The method for validating the cancer risk assessment model includes using nested cross-validation. Specifically, this involves dividing the dataset into training and validation sets multiple times, training the model on the training set, testing the results on the validation set, and then selecting the optimal validation model.

[0074] In a fourth aspect, the present invention provides a cancer risk assessment model, which is obtained using the cancer risk assessment model construction method described in any one of the third aspects.

[0075] The cancer risk assessment model of this invention exhibits excellent detection specificity and sensitivity, enabling early detection and prevention through personalized risk prediction. Furthermore, risk assessment using this model is cost-effective, requiring only low-depth cell-free DNA sequencing (sequencing depth of 2-3 times the human genome coverage), significantly reducing assessment costs compared to conventional sequencing depths of 200 times the human genome coverage.

[0076] A fifth aspect of the present invention provides a computer-readable storage medium for predicting cancer risk, the computer-readable storage medium storing a computer program that causes a computer to perform a method for obtaining cancer risk prediction biomarkers as described in any one of the first aspects, and / or a method for constructing a cancer risk assessment model as described in the third aspect.

[0077] In some embodiments of the present invention, the steps performed by the computer include:

[0078] Obtain cell-free DNA gene mutation data from the sample to be tested;

[0079] Based on the cell-free DNA gene mutation data of the test sample, at least one of the following data is obtained: gene mutation map data, percentage data of mutation sites in regulators, length distribution data of cell-free DNA, base distribution data of cell-free DNA ends, and distribution data of cell-free DNA ends in nucleosomes. These data are then used as feature values ​​to make predictions using the cancer risk assessment model.

[0080] A sixth aspect of the present invention provides an apparatus for predicting cancer risk, the apparatus comprising the following modules:

[0081] The acquisition module is used to acquire cell-free DNA sequencing data from subject samples and obtain cell-free DNA gene mutation data through filtering.

[0082] The calculation module is used to obtain at least one of the following data based on the free DNA gene mutation data: gene mutation map data of the subject, percentage data of mutation sites in regulators, length distribution data of free DNA, base distribution data of free DNA ends, and distribution data of free DNA ends in nucleosomes, and use it as the cancer risk prediction biomarker.

[0083] In some embodiments of the present invention, the computing module further includes performing the acquisition of high-contribution tumor mutation imprinting data, wherein the method for acquiring the high-contribution tumor mutation imprinting data includes:

[0084] Background mutation imprints are obtained from cell-free DNA gene mutation data of non-cancer patients, then tumor-specific gene mutation imprints from the tumor database are merged, and the gene mutation map obtained from the cell-free DNA gene mutation data of the subjects is deconvolved to obtain the final result.

[0085] In some embodiments of the present invention, the device further includes an analysis module for comparing the differences in cancer risk prediction biomarkers between the subject and non-cancer patients to determine whether there is a risk of developing cancer.

[0086] In some embodiments of the present invention, the method for acquiring the subject's gene mutation map data, the proportion of mutation sites in regulators, the length distribution data of free DNA, the base distribution data of the broken ends of free DNA, and the distribution data of the ends of free DNA in nucleosomes in the computing module of the device for predicting cancer risk is specifically referred to in the first aspect of the present invention.

[0087] A seventh aspect of the invention provides the use of the machine-executed method for acquiring cancer risk prediction biomarkers as described in any one of the first aspects, the cancer risk prediction biomarkers as described in any one of the second aspects, the method for constructing a cancer risk assessment model as described in any one of the third aspects, the cancer risk assessment model as described in the fourth aspect, the computer-readable storage medium as described in any one of the fifth aspects, or the apparatus for predicting cancer risk as described in any one of the sixth aspects, in any of the following:

[0088] A) Applications in the development or screening of cancer treatment drugs;

[0089] B) Applications in the development or screening of cancer prevention drugs;

[0090] C) Application in the preparation of products for predicting cancer prognosis.

[0091] Other features and advantages of the present invention will be set forth in the following description. Attached Figure Description

[0092] The present invention will be further described below with reference to the accompanying drawings and embodiments, wherein:

[0093] Figure 1 illustrates the method for constructing a cancer risk assessment model based on low-depth cell-free DNA sequencing according to the present invention.

[0094] Figure 2 is a comparison diagram of gene mutation patterns in the liver cancer group and the control group of this invention;

[0095] Figure 3 shows the unsupervised clustering results of the liver cancer dataset in this invention;

[0096] Figures 4a and 4b show the statistical and analytical results of the proportion of gene mutations located in enhancers in liver cancer patients and control groups of the present invention. Figure 4a shows the statistical results of the proportion of gene mutations in enhancers, and Figure 4b shows the ROC analysis results.

[0097] Figures 5a-5c show the deconvolution analysis results of this invention, where Figure 5a shows the total contribution rate statistics of tumor mutation imprints from COSMIC, Figure 5b shows the ROC analysis results, and Figure 5c shows the correlation analysis results with the abundance of tumor-derived DNA.

[0098] Figures 6a-6d show the detection results of the distribution pattern of free DNA length related to gene mutation. Figure 6a shows the length distribution statistics of Mut-DNA fragments and Wt-DNA fragments in liver cancer patients, Figure 6b shows the proportion of short fragments in the control group, Figure 6c shows the proportion of short fragments in liver cancer patients, and Figure 6d shows the Diff-size statistics of liver cancer patients and the control group.

[0099] Figures 7a-7f show the detection results of cell-free DNA terminal sequence usage patterns related to gene mutations. Figure 7a shows the usage rate of CCCA terminal sequences in the control group, Figure 7b shows the usage rate of CCCA terminal sequences in liver cancer patients, Figure 7c shows the statistical results of Diff-CCCA in liver cancer patients and the control group, Figure 7d shows the usage rate of CTCC breakpoint sequences in the control group, Figure 7e shows the usage rate of CTCC breakpoint sequences in liver cancer patients, and Figure 7f shows the statistical results of Diff-CTCC in liver cancer patients and the control group.

[0100] Figures 8a-8d show the detection results of the distribution pattern of cell-free DNA ends related to gene mutations in nucleosomes. Figure 8a shows the statistical results of the distribution of cell-free DNA ends in nucleosomes, Figure 8b shows the proportion of cell-free DNA ends inside nucleosomes in the control group, Figure 8c shows the proportion of cell-free DNA ends inside nucleosomes in liver cancer patients, and Figure 8d shows the statistical results of Diff-nucleosomes in the control group and liver cancer patients.

[0101] Figure 9 shows the ROC analysis results of this invention using Diff-size, Diff-CTCC, or Diff-nucleosome as detection markers, respectively.

[0102] Figures 10a and 10b show the validation results of the cancer risk assessment model constructed based on the present invention, where Figure 10a shows the statistical results of the prediction scores and Figure 10b shows the ROC analysis results.

[0103] Figures 11a-11c show the validation results using cell-free DNA length as a detection marker. Figure 11a shows the detection results based on Diff-size, Figure 11b shows the detection results based on commonly used cell-free DNA length, and Figure 11c shows the ROC analysis results. Detailed Implementation

[0104] The technical concept and effects of the present invention will be clearly and completely described below with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.

[0105] The terms "preferred," "more preferably," etc., as used in this invention refer to embodiments of the invention that provide certain beneficial effects under certain circumstances. However, other embodiments may also be preferred under the same or other circumstances. Furthermore, the description of one or more preferred embodiments does not imply that other embodiments are unavailable, nor is it intended to exclude other embodiments from the scope of protection of this invention.

[0106] When a range of values ​​is disclosed herein, the range shall be considered continuous and include the minimum and maximum values ​​of the range, as well as every value between the minimum and maximum values. Furthermore, when the range refers to integers, it includes every integer between the minimum and maximum values ​​within the range. Additionally, when multiple ranges are provided to describe a feature or characteristic, these ranges may be combined. In other words, unless otherwise specified, all ranges disclosed herein shall be understood to include any and all subranges to which they are incorporated.

[0107] In the description of this invention, the reference term "," includes all and any combination of one or more of the associated listed items.

[0108] In the description of this invention, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0109] Unless otherwise specified in the examples, the procedures should be performed under standard conditions or conditions recommended by the manufacturer. Reagents or instruments whose manufacturers are not specified are all commercially available products.

[0110] In peripheral blood cell-free DNA sequencing data, most gene mutation signals originate from factors unrelated to tumors, such as experimental or sequencing errors and clonal hematopoiesis. However, in cancer patients, mutation information carried by cell-free DNA released from tumors remains. The inventors, through previous research, discovered that tumor mutations are not random but follow specific patterns, unlike experimental errors or clonal hematopoiesis. Therefore, in low-depth sequencing data, it may be possible to predict the presence of tumors by specifically identifying these tumor-specific mutation patterns, thereby aiding in cancer diagnosis. Furthermore, further research revealed that the fragmented omics and epigenomic information of cell-free DNA from tumors differs from that of normal cells. Therefore, in cancer patients, the presence of this type of DNA alters the signals associated with cell-free DNA carrying gene mutations, potentially leading to the development of novel tumor biomarkers.

[0111] In a first aspect, the present invention provides a machine-executed method for obtaining cancer risk prediction biomarkers, comprising:

[0112] S1. Obtain cell-free DNA sequencing data from subject samples and obtain gene mutation data of cell-free DNA through filtering;

[0113] S2. Based on the free DNA gene mutation data, obtain at least one of the following data from the subject: gene mutation map data, percentage data of mutation sites in regulators, length distribution data of free DNA, base distribution data of free DNA ends, and distribution data of free DNA ends in nucleosomes, and use it as the cancer risk prediction biomarker.

[0114] The subjects' samples include blood samples, urine samples, tear samples, sweat samples, semen samples, or saliva samples; the sequencing depth of the cell-free DNA sequencing data is less than 10 times the human genome coverage, preferably 2 to 5 times the human genome coverage.

[0115] In some implementations, the filtering process includes: discarding SNP locus data with abundance above a first threshold in the SNP database, and discarding data located in the "blacklist" region defined by the ENCODE project and / or mutation data within the second threshold range with the highest coverage (optional), wherein the first threshold ranges from 0.05% to 1%, and the second threshold ranges from 0.1% to 5%. Specifically, the filtering process includes: discarding SNP locus data with abundance above 0.1% in the SNP database, and discarding mutation data located in the "blacklist" region defined by the ENCODE project and the 1% with the highest coverage (optional), wherein the SNP database is a large-scale population SNP database, including but not limited to dbSNP, NyuWa, ChinaMap, Korea1K, ToMMo, gnomAD, etc.

[0116] In addition, the filtering process includes preliminary detection of gene mutations, discarding data with a reading alignment quality number below the third threshold and a base quality number below the fourth threshold; the third threshold is any value between 20 and 80, and the fourth threshold is any value between 20 and 50. Specifically, the filtering process includes alignment using BWA software and preliminary detection of gene mutations using BCFtools software, discarding data with a reading alignment quality number below 60 and a base quality number below 30.

[0117] In some embodiments, the method for obtaining gene mutation map data includes: converting the "A" and "G" mutation sites in the cell-free DNA gene mutation data to "T" and "C" bases, respectively, and recording one base before and after each mutation site to obtain 96 mutation pattern data. Then, normalization processing is performed to obtain the proportion of each mutation pattern and generate gene mutation map data. The 96 mutation pattern data include six basic mutation types: C>A, C>G, C>T, T>A, T>C, and T>G.

[0118] In some embodiments, the method for obtaining the proportion data of the mutation site in the regulator includes: obtaining the number of sites in the cell-free DNA gene mutation data that overlap with the regulator, and calculating the proportion of these overlapping sites to all mutation sites, thereby obtaining the proportion data of the mutation site in the regulator. The regulator can be one or more of the following: enhancer, promoter, insulator, etc.

[0119] In some embodiments, the regulator is an enhancer, and the method for obtaining the percentage data of the mutation site in the enhancer includes: obtaining the location of the enhancer in blood cells, which can be obtained by performing a ChIP-seq experiment or an ATAC-seq experiment on H3K27ac histone modification; taking out the sites that overlap with the enhancer among all mutation sites, calculating their number, and dividing them by the total number of mutation sites to obtain the percentage data of the mutation site in the enhancer.

[0120] In some embodiments, the method for obtaining the length distribution data of cell-free DNA includes: dividing the cell-free DNA gene mutation data into Mut-DNA fragments carrying the mutated genotype and Wt-DNA fragments not carrying the mutated genotype, and calculating the proportion of base fragments with a length below a fifth threshold in each of the Mut-DNA fragments or Wt-DNA fragments, and then obtaining the difference between the two proportions, which is the length distribution data of the cell-free DNA, denoted as Diff-size. The fifth threshold is taken from any positive integer between 100 and 150, and the unit is bp.

[0121] The length of the Mut-DNA fragment or Wt-DNA fragment is calculated using conventional methods in the field. For example, the length of the free DNA can be obtained by acquiring the locus information of the two ends of each free DNA in the genome and calculating the distance between the two ends.

[0122] In some embodiments, the method for obtaining the base distribution data of the broken ends of free DNA includes: calculating the proportion of fragments with the "CTCC" base at the 5' break point in the Mut-DNA fragment or the Wt-DNA fragment respectively, and then obtaining the difference between the proportions to obtain the base distribution data of the broken ends of free DNA, denoted as Diff-CTCC.

[0123] In this context, the breakpoint bases of cell-free DNA refer to the sequence formed by extracting the locus information of the 5' end of each cell-free DNA fragment (Mut-DNA fragment or Wt-DNA fragment) in the genome, and extracting the sequence consisting of 4 bases: 2 bases upstream and 2 bases downstream of the locus. This sequence represents the breakpoint bases corresponding to the cell-free DNA fragment. Furthermore, for a given cell-free DNA set (such as Mut-DNA), the breakpoint base sequence of each cell-free DNA fragment is calculated, and the number of each sequence is counted. This count is then divided by the total number of DNA fragments in the set to obtain the proportion of that sequence in the set, thus yielding the distribution data of the breakpoint base sequences in the cell-free DNA set.

[0124] In some embodiments of the present invention, the method for obtaining the distribution data of the ends of cell-free DNA in nucleosomes includes: calculating the proportion of Mut-DNA fragments whose ends are located within a sixth threshold at the center of the nucleosome, and the proportion of Wt-DNA fragments whose ends are located within a seventh threshold at the center of the nucleosome, and then obtaining the difference between these proportions, which is the distribution data of the ends of cell-free DNA in nucleosomes, denoted as Diff-nucleosome. The sixth and seventh thresholds are independently taken from any positive integer between 10 and 100, and the unit is bp.

[0125] Taking blood cell samples as an example, the location of nucleosome centers in blood cells can be obtained using micrococcal nuclease sequencing (Mnase-seq) technology. For each cell-free DNA molecule, its two ends (the smaller locus is called the upstream end, or U end, and the larger locus is called the downstream end, or D end) are analyzed separately. For each site, the nearest nucleosome center is located, and the distance is calculated; if the end is upstream of the nucleosome center, it is recorded as a negative value; otherwise, it is recorded as a positive value. For the set of cell-free DNA samples under consideration, the distances between the U and D ends of each cell-free DNA molecule and the nearest nucleosome center are calculated, and the frequency of different distances is counted. This is then divided by the number of DNA molecules to obtain the distribution data of the ends of cell-free DNA within the nucleosomes.

[0126] In some embodiments, step S2 further includes acquiring terminal base distribution data of cell-free DNA and / or high-contribution tumor mutation imprinting data, wherein:

[0127] The method for obtaining the terminal base distribution data of cell-free DNA includes: dividing the cell-free DNA gene mutation data into Mut-DNA fragments carrying the mutant genotype and Wt-DNA fragments not carrying the mutant genotype, and calculating the proportion of fragments with the "CCCA" base at the 5' end in the Mut-DNA fragment or the Wt-DNA fragment respectively, and then obtaining the difference between the proportions, which is the terminal base distribution data of cell-free DNA, denoted as Diff-CCCA;

[0128] The method for obtaining high-contribution tumor mutation imprint data includes: obtaining background mutation imprints based on cell-free DNA gene mutation data from non-cancer patients, then merging tumor-specific gene mutation imprints from the tumor database, and deconvolving the gene mutation map obtained based on the cell-free DNA gene mutation data of the subjects to obtain the final result.

[0129] Secondly, as shown in Figure 1, this invention provides a method for constructing a cancer risk assessment model based on low-depth cell-free DNA sequencing data, specifically including the following steps:

[0130] (1) Collection and processing of peripheral blood:

[0131] Peripheral blood was collected from veins of both cancer and non-cancer patients using blood collection tubes containing anticoagulants. The blood was centrifuged at low temperature and medium speed for a period of time, then the supernatant was aspirated and transferred to a new centrifuge tube. The blood was then centrifuged again at low temperature and high speed, and the supernatant (i.e., plasma) was aspirated and transferred to a new centrifuge tube or cryopreservation tube. The plasma can be stored in an ultra-low temperature freezer until use.

[0132] (2) Extraction, sequencing, and analysis of cell-free DNA:

[0133] Cell-free DNA (cfDNA) was extracted from the plasma using commercially available kits and equipment, a sequencing library was constructed, and high-throughput sequencing was performed. Using a paired-end 100-base mode, the sequencing data averaged 60 million reads (approximately three times the coverage of the human genome). The raw sequencing data was first preprocessed (e.g., using Ktrim software) to remove sequencing adapters and low-quality cycles (optional step), and to remove repetitive sequences (optional step). Then, it was aligned to the human reference genome using BWA software, and repetitive sequences and low-quality alignment results were removed again (optional step) to obtain the final alignment results. Considering the influence of sex on the genome, data from the X and Y chromosomes can be removed before subsequent analysis.

[0134] (3) Obtaining gene mutation information in cell-free DNA:

[0135] For the whole genome sequencing information of the above cell-free DNA, BCFtools software was used to perform preliminary detection of gene mutations. DNA readings with an alignment quality number below 60 were discarded, and data with a sequencing quality number below 30 from the remaining DNA readings were also discarded.

[0136] The processed data is then filtered, discarding mutations located in the "blacklist" region defined by the ENCODE project and those with the highest coverage of 1% (generally issues in the data analysis process, and the results are unreliable). Common SNP sites with an abundance of 0.1% or higher in large population SNP databases (including dbSNP, NyuWa, ChinaMap, Korea1K, and ToMMo) are also discarded, while mutations recorded in the COSMIC database are retained. This yields the gene mutation data of peripheral blood cell-free DNA.

[0137] (4) Obtaining a gene mutation map:

[0138] Based on the gene mutation data of cell-free DNA obtained after the above filtering, the proportion of each mutation pattern was analyzed. Specifically, this included: using the recording method in the COSMIC database, sites with genomic A and G values ​​were converted into their complementary base pairs, C and T, for recording (C>A, C>G, C>T, T>A, T>C, T>G, a total of 6 basic types); simultaneously, one base before and after the mutation site was recorded (i.e., 4×4=16 types), obtaining 96 mutation patterns. For each sample, a gene mutation map was drawn based on the abundance of these 96 mutation patterns.

[0139] The above methods were used to obtain gene mutation profiles of cancer patients and non-cancer patients, and unsupervised cluster analysis was performed to determine whether there were significant differences.

[0140] (5) Deconvolution analysis:

[0141] Considering issues such as experimental errors and clonal hematopoiesis, a subset of samples from non-cancer patients were selected, and their cell-free DNA mutation information was used as a "background signature" to quantitatively simulate these signals. This "background signature" was combined with tumor-specific gene mutational signatures recorded in the COSMIC database to form a reference mutational signature set. Then, a computational method was used to deconvolve the mutation maps of cancer patient samples (generally using a non-negative matrix factorization algorithm) to obtain the contribution rate of each gene mutational signature.

[0142] Gene mutation imprints from COSMIC are extracted and used as tumor-related gene mutation information for further interpretation, such as determining whether the overall proportion differs between cancer patients and controls, and whether the dominant gene mutation imprint can predict the origin of cancer tissue.

[0143] (6) Obtain the proportion of mutation sites in enhancers:

[0144] Based on the gene mutation data of the free DNA obtained after the above filtering, the proportion of mutation sites in the known enhancer regions (obtained from the ENCODE project) was calculated.

[0145] Specifically, this includes downloading binding peak site information (entry name ENCSR222QLW) from the ENCODE project, obtained from ChIP-seq experiments targeting H3K27ac in T cells, which represents the enhancer region. For each sample, the proportion of mutation sites located within the enhancer region among all its mutation sites is calculated.

[0146] (7) Obtaining fragmented omics information:

[0147] For each sample, cell-free DNA reads covering the mutation sites are extracted and categorized into two groups based on their genotype at the mutation sites: one group carries the mutant genotype, termed Mut-DNA; the other group is wild-type, termed Wt-DNA. Then, fragment omics analysis is performed on Mut-DNA and Wt-DNA separately, specifically including the following steps:

[0148] Step 1: Based on the above Mut-DNA and Wt-DNA, plot their length distribution, and use the proportion of readings of 150 bases or less (i.e., shorter DNA) as the quantitative method of length distribution (Diff-size), and use 150 bases as the length characterization mode, which can be 100, 110, 120, 130, 140, etc.

[0149] Step 2: Calculate the distribution of the four base sequences starting at the 5' end, and use the proportion of CCCA as the quantitative method for the distribution of terminal bases (Diff-CCCA);

[0150] Step 3: Calculate the distribution of the 4 bases at the 5' breakpoint (i.e., 2 bases located upstream of the 5' of the free DNA and 2 bases at the 5' start), and use the proportion of CTCC as the quantitative method for the distribution of breakpoint bases (Diff-CTCC).

[0151] Step 4: Obtain the relative positions of the free DNA ends within the nucleosome, and use the proportion of ends within 50 bases of the nucleosome center as the quantitative method for the distribution of ends in the nucleosome (Diff-nucleosome). In addition, when calculating the distribution of ends in the nucleosome, other distance thresholds such as 50, 60, or 70 bases upstream or downstream can also be used, and U (upstream) and D (downstream) ends can also be calculated separately.

[0152] Step 5: Based on the above quantitative method, compare whether there is a statistical difference between Mut-DNA and Wt-DNA.

[0153] (8) Artificial intelligence-driven diagnostic assessment:

[0154] Gradient decision boosting trees (or other artificial intelligence techniques) are used to integrate (partially or entirely) the 96 gene mutation profiles obtained from cell-free DNA, the proportion of mutation sites in regulators, and fragment omics quantitative results (such as cell-free DNA length distribution data, cell-free DNA end base distribution data, and cell-free DNA end distribution data in nucleosomes), and nested cross-validation is used to develop and evaluate cancer diagnostic models.

[0155] Thirdly, this invention proposes a cancer risk assessment model, which is constructed using the cancer risk assessment model construction method described in the second aspect above.

[0156] The types of cancer included, depending on actual needs, are not limited to: lung cancer, breast cancer, liver cancer, bladder cancer, acute myeloid leukemia, adrenocortical carcinoma, urothelial carcinoma of the bladder, ductal carcinoma of the breast, lobular carcinoma of the breast, cervical cancer, bile duct cancer, colorectal cancer, esophageal cancer, gastric cancer, glioblastoma multiforme, squamous cell carcinoma of the head and neck, chromophobe renal carcinoma, clear cell renal carcinoma, papillary renal carcinoma, low-grade glioma, mesothelioma, lung adenocarcinoma, lung squamous cell carcinoma, ovarian serous cystadenocarcinoma, pancreatic cancer, pheochromocytoma and paraganglioma, prostate cancer, sarcoma, melanoma of the skin, testicular cancer, thymic cancer, thyroid cancer, uterine sarcoma, endometrial cancer, uveal melanoma, etc.

[0157] Fourthly, the present invention proposes a computer-readable storage medium for predicting cancer risk based on low-depth cell-free DNA sequencing data, the computer-readable storage medium storing a computer program that causes a computer to perform the cancer risk prediction biomarker acquisition method as described in any one of the first aspects, and / or the cancer risk assessment model construction method as described in the second aspect.

[0158] Fifthly, the present invention provides a device for predicting cancer risk, the device comprising the following modules:

[0159] The acquisition module is used to acquire cell-free DNA sequencing data from subject samples and obtain cell-free DNA gene mutation data through filtering.

[0160] The calculation module is used to obtain at least one of the following data based on the free DNA gene mutation data: gene mutation map data of the subject, percentage data of mutation sites in regulators, length distribution data of free DNA, base distribution data of free DNA ends, and distribution data of free DNA ends in nucleosomes, and use it as the cancer risk prediction biomarker.

[0161] In some instances, the device also includes an analysis module for comparing the differences in cancer risk prediction biomarkers between the subject and non-cancer patients to determine whether there is a risk of developing cancer, or for determining whether the subject has a risk of developing cancer based on the cancer risk assessment model constructed above.

[0162] The device implementation described above is merely illustrative. The modules described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0163] It is understood that all or some of the steps disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). It is understood that computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer.

[0164] Furthermore, it is understood that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0165] The following section provides a detailed explanation of the method for predicting cancer risk using gene mutation information from low-depth cell-free DNA sequencing data, using specific examples.

[0166] Example 1: Construction of a Cancer Risk Assessment Model Based on Low-Depth Cell-Free DNA Sequencing Data

[0167] The dataset used in this embodiment comes from the collaborating institution, Shenzhen Third People's Hospital, and includes 56 liver cancer patients and 24 controls (all hepatitis B carriers or non-cancerous liver diseases). All recruited patients provided informed consent, and the process was approved by the ethics committee. The method for constructing the cancer risk assessment model based on low-depth cell-free DNA sequencing data in this embodiment is as follows.

[0168] (1) Collection and processing of peripheral blood:

[0169] Approximately 6 mL of peripheral blood was collected from a vein using a blood collection tube containing an anticoagulant (EDTA in this example). The blood was centrifuged at 4°C and 1600 g / L for 15 minutes. The supernatant was then aspirated and transferred to a new centrifuge tube. The plasma was then centrifuged again at 4°C and 16000 g / L for 15 minutes. The supernatant (i.e., plasma) was then aspirated and transferred to a new centrifuge tube or cryopreservation tube. The plasma was stored at -80°C for later use.

[0170] (2) Extraction, sequencing, and analysis of cell-free DNA:

[0171] Cell-free DNA was extracted from the plasma using commercially available kits and equipment, a sequencing library was constructed, and high-throughput sequencing was performed at a depth of approximately 2-5x. The raw sequencing data was first preprocessed (using Ktrim software) to remove sequencing adapters and low-quality cycles, and to remove repetitive sequences. Then, it was compared with a human reference genome, and repetitive sequences and low-quality alignment results were removed to obtain the final alignment results. Whole-genome sequencing was performed using BWA software. Considering the influence of sex on the genome, data from the X and Y chromosomes were removed.

[0172] (3) Obtaining gene mutation information in cell-free DNA:

[0173] Preliminary gene mutation detection was performed using BCFtools software, with data having an alignment quality number below 60 and a base quality number below 30 discarded. The following filtering process was then performed:

[0174] Discard the "blacklist" regions defined by the ENCODE project and the mutations with the highest coverage of 1%; discard common SNP sites with an abundance of more than 0.1% in large population SNP databases (including dbSNP, NyuWa, ChinaMap, Korea1K, and ToMMo), but retain mutations recorded in the COSMIC database, thus obtaining the gene mutation data of peripheral blood cell-free DNA.

[0175] (4) Acquisition and analysis of gene mutation mapping data:

[0176] Based on the free DNA gene mutation data obtained after the above filtering, the proportion of each mutation pattern is analyzed. Specifically, this includes the following steps:

[0177] Using the recording method in the COSMIC database, sites in the genome designated as A and G were converted into their complementary base pairs, C and T, for recording (including six basic types: C>A, C>G, C>T, T>A, T>C, and T>G). Simultaneously, one base before and after the mutation site was recorded (i.e., 4×4=16 types), resulting in 96 mutation patterns. For each sample, a gene mutation map was constructed based on the abundance of these 96 mutation patterns. The differences in gene mutation patterns between the liver cancer patient group and the control group were then compared.

[0178] Figure 2 shows the gene mutation pattern atlas results in a control sample and a liver cancer sample. Further, applying an unsupervised clustering algorithm to this gene mutation atlas revealed that the detected samples could be clustered into two categories (as shown in Figure 3). One category (dark gray) contained significantly more cancer samples than the other (light gray) (P = 0.0016, chi-square test). This indicates that in cancer patients, gene mutation patterns (the distribution of the 96 mutation patterns) do indeed carry tumor-related information and have the potential to distinguish them from other patients, with a specificity of 79.2% and a sensitivity of 62.5%.

[0179] (5) Analysis of the proportion of mutation sites in enhancers:

[0180] Based on the gene mutation information of the free DNA obtained after the above filtering, known enhancer regions are obtained from the ENCODE project, and the proportion of mutation sites in the known enhancer regions is calculated.

[0181] The results, shown in Figures 4a and 4b, indicate that the proportion of gene mutations in the liver cancer patient samples was significantly higher than that in the control group (Figure 4a), with P = 0.00037 (Mann-Whitney rank-sum test). ROC analysis based on the above dataset showed that this parameter could distinguish between the two groups, with an AUC of 0.78 (Figure 4b), P < 0.00001 (Z-test).

[0182] (6) Deconvolution analysis:

[0183] Eight samples were randomly selected from the control group to form a "background mutation imprint," and deconvolution analysis was performed using 67 tumor mutation imprints from the COSMIC database (note that imprints annotated as sequencing error-related in this example were excluded). Details of the 67 tumor mutation imprints can be found at https: / / cancer.sanger.ac.uk / signatures / sbs / , including:

[0184] SBS1, SBS2, SBS3, SBS4, SBS5, SBS6, SBS7a, SBS7b, SBS7c, SBS7d, SBS8, SBS9, SBS10a, SBS10b, SBS10c, SBS10d, SBS11, SBS 12. SBS13, SBS14, SBS15, SBS16, SBS17a, SBS17b, SBS18, SBS19, SBS20, SBS21, SBS22a, SBS22b, SBS23, SBS24, SBS25, SBS2 6. SBS28, SBS29, SBS30, SBS31, SBS32, SBS33, SBS34, SBS35, SBS36, SBS37, SBS38, SBS39, SBS40a, SBS40b, SBS40c, SBS41, SBS42, SBS44, SBS84, SBS85, SBS86, SBS87, SBS88, SBS89, SBS90, SBS91, SBS92, SBS93, SBS94, SBS96, SBS97, SBS98, SBS99.

[0185] The results, shown in Figures 5a-5c, indicate that the vast majority of contributions in all samples came from "background mutation imprinting." However, the total contribution rate of tumor mutation imprinting from COSMIC was significantly higher in hepatocellular carcinoma patients than in the control group (Figure 5a; P = 0.0042, Mann-Whitney rank-sum test). ROC analysis confirmed that this parameter could distinguish between the two, with an AUC of 0.79 (Figure 5b; P < 0.00001, Z-test). Furthermore, this parameter was significantly positively correlated with the abundance of tumor-derived DNA (assessed using the ichorCNA algorithm) (Figure 5c; P = 0.00068, Spearman correlation).

[0186] (7) Fragmented omics analysis:

[0187] For each sample, since some sites are covered by multiple DNA molecules, cell-free DNA reads are extracted from the sites covering the mutations. These cell-free DNA reads are then divided into two categories based on their genotype at the mutation sites: one category carries the mutant genotype, called Mut-DNA; the other category is wild-type, called Wt-DNA. Fragmentation omics analysis is then performed on Mut-DNA and Wt-DNA separately, specifically including the following steps:

[0188] Step 1: Based on the above free DNA readings, plot their length distribution, and use the proportion of readings of 150 bases or less (i.e., shorter DNA) as the quantitative method of length distribution.

[0189] Step 2: Calculate the distribution of the four base sequences starting at the 5' end, and use the proportion of CCCA as a quantitative method for the distribution of terminal bases;

[0190] Step 3: Calculate the sequence distribution of the four bases at the 5' breakpoint (i.e., two bases are located upstream of the 5' of the free DNA, and the two bases at the 5' start), and use the proportion of CTCC as the quantitative method for the base distribution at the breakpoint;

[0191] Step 4: Obtain the relative position of the ends of free DNA within the nucleosome, and use the proportion of ends within 50 bases of the nucleosome center as the quantitative method for the distribution of ends in the nucleosome.

[0192] Step 5: Based on the above quantitative method, compare whether there is a statistical difference between Mut-DNA and Wt-DNA.

[0193] The results of the mutation-related cell-free DNA length distribution pattern are shown in Figures 6a-6d. They reveal that in both the control group and hepatocellular carcinoma (HCC) patients, Mut-DNA is consistently shorter than Wt-DNA (Figure 6a) and contains more fragments of 150 bases or shorter (Figures 6b and 6c). Furthermore, analysis based on the difference between the proportion of short fragments in Mut-DNA (here, the proportion of DNA fragments shorter than 150 bases) and the proportion of short fragments in Wt-DNA in each sample (Diff-size) shows that Diff-size is significantly higher in HCC patients than in the control group (Figure 6d).

[0194] The usage patterns of mutation-related cell-free DNA terminal sequences are shown in Figures 7a-7f, revealing a lower CCCA terminal sequence usage rate in Mut-DNA compared to Wt-DNA during sequence analysis (Figures 7a-7b, representing the control and hepatocellular carcinoma groups, respectively). Furthermore, using Diff-CCCA to characterize the CCCA terminal sequence usage rate in Mut-DNA and Wt-DNA revealed a decreasing trend in HCC patients compared to the control group (Figure 7c).

[0195] Meanwhile, in liver cancer patients, Mut-DNA had a lower CTCC breakpoint sequence usage rate than Wt-DNA, while there was no significant difference between the two in the control group (as shown in Figures 7d-7e, representing the control group and liver cancer group, respectively); in addition, based on the CTCC breakpoint sequence usage rate (Diff-CTCC) in Mut-DNA and Wt-DNA, it was found that it was significantly lower in liver cancer patients compared to the control group (as shown in Figure 7f).

[0196] The distribution patterns of mutation-related cell-free DNA ends in nucleosomes are shown in Figures 8a-8d. These results indicate that in liver cancer patients, Mut-DNA has a higher proportion of ends located inside nucleosomes compared to Wt-DNA (Figures 8a-8c). Furthermore, based on the difference between the proportion of ends located inside nucleosomes in Mut-DNA and Wt-DNA (Diff-nucleosome), this value was found to be significantly higher in liver cancer samples than in the control group (Figure 8d).

[0197] The data above indicate that Diff-size, Diff-CTCC, and Diff-nucleosome all show significant differences between liver cancer patients and the control group. Based on the above dataset, ROC analysis was performed using Diff-size, Diff-CTCC, or Diff-nucleosome as cancer risk assessment indicators. The results are shown in Figure 9, which shows that the AUC values ​​of Diff-size, Diff-CTCC, and Diff-nucleosome are 0.68, 0.64, and 0.68, respectively, indicating their potential as cancer diagnostic biomarkers.

[0198] (8) Cancer risk model construction:

[0199] Using the original proportions of the above 96 gene mutation types, the proportion of mutation sites in enhancers, and fragment omics (Diff-size, Diff-CCCA, Diff-CTCC, and Diff-nucleosome) as features, a cancer risk assessment model was developed using the GBM algorithm. The model was evaluated using 10-level nested cross-validation, and the data from the test set in the cross-validation was used as the prediction results (for details of the method, see Cristiano et al. Nature 2019; 570:385-389).

[0200] The results, shown in Figures 10a-10b, indicate that the predictive value for liver cancer patients was significantly higher than that for the control group, with an AUC value reaching 0.85. This demonstrates its good ability to assess cancer risk.

[0201] Example 2: Validation of a cancer risk assessment model based on low-depth cell-free DNA sequencing data

[0202] To test whether the cancer risk assessment model based on low-depth cell-free DNA sequencing data is effective in other cancers, this embodiment collected the Liang et al. dataset (Liang et al. Whole-genome sequencing of cell-free DNA yields genome-wide read distribution patterns to track tissue of origin in cancer patients. Clin Transl Med. 2020 Oct 11; 10(6):e177. doi:10.1002 / ctm2.177), which included 8 pre-treatment non-small cell lung cancer samples and 10 control samples.

[0203] Using the cancer risk assessment model construction method based on low-depth cell-free DNA sequencing data from Example 1 above, Diff-size in fragmented omics data was used as a predictive biomarker. The results are shown in Figures 11a-11c, which show that Diff-size has a significant difference between the two groups and the AUC value for diagnosis reaches 1, which can distinguish lung cancer patients with 100% accuracy.

[0204] In contrast, using the commonly used cell-free DNA length parameter (i.e., the proportion of cell-free DNA of 150 bases or less), the results are shown in Figure 11b. While some trends are visible, it is difficult to distinguish between lung cancer patients and the control group. This demonstrates that fragment omics Diff-size analysis based on gene mutation information from low-depth cell-free DNA sequencing data can be used to predict cancer risk.

[0205] In summary, this invention provides a method, computer-readable storage medium, and apparatus for predicting cancer risk based on low-depth cell-free DNA sequencing data. Starting from gene mutation data in low-depth cell-free DNA sequencing data, which is often overlooked in mainstream research, this invention identifies several novel cancer diagnostic biomarkers (such as gene mutation maps, the proportion of mutation sites in enhancers, and fragment omics information). These biomarkers can be used for various cancer types, exhibiting high universality and wide applicability.

[0206] Furthermore, the cancer risk assessment model constructed in this invention has low data requirements, only requiring conventional whole-genome sequencing. It has no special requirements for experimental protocols or sequencing platforms, resulting in low operational complexity and low cost, making it suitable for clinical application. Simultaneously, the biomarkers proposed in this invention have the potential to serve as partial feature sets, used in conjunction with biomarkers in mainstream research to develop large-scale, ultra-high-accuracy cancer diagnostic models using artificial intelligence technology. Moreover, in some cases, the Diff-size of this invention, compared to directly using length parameters, can better distinguish cancer patients from control groups.

[0207] The embodiments of the present invention have been described in detail above. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention. Furthermore, the embodiments of the present invention and the features thereof can be combined with each other unless otherwise specified.

Claims

1. A machine-executed method for acquiring cancer risk prediction biomarkers, characterized in that, include: S1. Filter the cell-free DNA sequencing data of the subject samples to obtain cell-free DNA gene mutation data; S2. Based on the free DNA gene mutation data, obtain at least one of the following data from the subject: gene mutation map data, percentage data of mutation sites in regulators, length distribution data of free DNA, base distribution data of free DNA ends, and distribution data of free DNA ends in nucleosomes, and use it as a cancer risk prediction biomarker.

2. The method for obtaining cancer risk prediction biomarkers according to claim 1, characterized in that, The sequencing depth of the cell-free DNA sequencing data is less than 10 times the coverage of the human genome. Preferably, the sequencing depth of the free DNA sequencing data is 2 to 5 times the coverage of the human genome.

3. The method for obtaining cancer risk prediction biomarkers according to claim 1, characterized in that, The filtering process includes discarding SNP site data in the SNP database with an abundance above a first threshold.

4. The method for obtaining cancer risk prediction biomarkers according to claim 3, characterized in that, The filtering process also includes discarding data located in the "blacklist" area defined by the ENCODE project; And / or, discard mutation data within the second threshold range with the highest coverage.

5. The method for obtaining cancer risk prediction biomarkers according to any one of claims 1 to 4, characterized in that, The method for obtaining the gene mutation map data includes: obtaining mutation pattern data based on the free DNA gene mutation data, then obtaining the proportion data of each mutation pattern and generating gene mutation map data; And / or, the method for obtaining the percentage data of the mutation site in the regulator includes: obtaining the number of sites in the cell-free DNA gene mutation data where the mutation site overlaps with the regulator, and calculating the percentage data of the site in the mutation site; And / or, the method for obtaining the length distribution data of the free DNA includes: dividing the free DNA gene mutation data into Mut-DNA fragments carrying the mutant genotype and Wt-DNA fragments not carrying the mutant genotype, and calculating the proportion of base fragments with a length below the fifth threshold in the Mut-DNA fragment or the Wt-DNA fragment respectively, and then obtaining the difference between the two proportions, which is the length distribution data of the free DNA, denoted as Diff-size; And / or, the method for obtaining the cell-free DNA fragment base distribution data includes: dividing the cell-free DNA gene mutation data into Mut-DNA fragments carrying the mutant genotype and Wt-DNA fragments not carrying the mutant genotype, and calculating the proportion of fragments in the Mut-DNA fragment or the Wt-DNA fragment with two bases upstream and downstream of the 5' break position as "CTCC", and then obtaining the difference between the two proportions to obtain the cell-free DNA fragment base distribution data, denoted as Diff-CTCC; And / or, the method for obtaining the distribution data of the ends of the free DNA in nucleosomes includes: dividing the free DNA gene mutation data into Mut-DNA fragments carrying the mutant genotype and Wt-DNA fragments not carrying the mutant genotype, and calculating the proportion of the Mut-DNA fragments whose ends are located within the sixth threshold of the nucleosome center and the proportion of the Wt-DNA fragments whose ends are located within the seventh threshold of the nucleosome center, and then obtaining the difference between the two proportions, which is the base distribution data of the free DNA ends, denoted as Diff-nucleosome; And / or, the mutation pattern data includes 96 mutation pattern data.

6. The method for obtaining cancer risk prediction biomarkers according to claim 5, characterized in that, The value of the fifth threshold is any positive integer between 50 and 160, and the unit is bp; And / or, the sixth threshold and the seventh threshold are independently taken from any positive integer between 50 and 160, in units of bp.

7. The method for obtaining cancer risk prediction biomarkers according to claim 6, characterized in that, Step S2 further includes acquiring terminal base distribution data of free DNA and / or high-contribution tumor mutation imprinting data, wherein: The method for obtaining the terminal base distribution data of the free DNA includes: dividing the free DNA gene mutation data into Mut-DNA fragments carrying the mutant genotype and Wt-DNA fragments not carrying the mutant genotype, and calculating the proportion of fragments with the "CCCA" base at the 5' end in the Mut-DNA fragment or the Wt-DNA fragment respectively, and then obtaining the difference between the two proportions, which is the terminal base distribution data of the free DNA, denoted as Diff-CCCA; The method for obtaining high-contribution tumor mutation imprint data includes: obtaining background mutation imprints based on cell-free DNA gene mutation data of non-cancer patients, then merging tumor-specific gene mutation imprints in the tumor database, and deconvolving the gene mutation map obtained based on the cell-free DNA gene mutation data of the subjects to obtain the data.

8. Cancer risk predictive biomarkers, characterized in that, include: At least one of the following: gene mutation map data, percentage data of mutation sites in regulators, length distribution data of cell-free DNA, terminal base distribution data of cell-free DNA, broken base distribution data of cell-free DNA, and distribution data of the ends of cell-free DNA in nucleosomes.

9. The cancer risk predictive biomarker according to claim 8, characterized in that, The cancer risk predictive biomarkers are obtained using the machine-executed cancer risk predictive biomarker acquisition method as described in any one of claims 1 to 7.

10. A method for constructing a cancer risk assessment model, characterized in that, include: The cell-free DNA sequencing data of cancer patient samples and non-cancer patient samples were filtered to obtain the corresponding cell-free DNA gene mutation data; Based on the free DNA gene mutation data, at least one of the following data is obtained from the cancer patient samples and the non-cancer patient samples: gene mutation map data, percentage data of mutation sites in regulators, length distribution data of free DNA, terminal base distribution data of free DNA, broken base distribution data of free DNA, and distribution data of the ends of free DNA in nucleosomes. This data is then used as feature values ​​for the machine learning algorithm to generate a cancer risk assessment model.

11. A cancer risk assessment model, characterized in that, The cancer risk assessment model was obtained using the method described in claim 10.

12. A computer-readable storage medium for predicting cancer risk, the computer-readable storage medium storing a computer program, characterized in that, The computer program causes the computer to perform the machine-executed method for obtaining cancer risk prediction biomarkers as described in any one of claims 1 to 7, and / or the method for constructing a cancer risk assessment model as described in claim 10.

13. A device for predicting cancer risk, characterized in that, The device includes the following modules: The filtering module is used to filter the cell-free DNA sequencing data of the subject samples to obtain cell-free DNA gene mutation data; The calculation module is used to obtain at least one of the following data based on the free DNA gene mutation data: gene mutation map data of the subject, percentage data of mutation sites in regulators, length distribution data of free DNA, base distribution data of free DNA ends, and distribution data of free DNA ends in nucleosomes, and use it as the cancer risk prediction biomarker.

14. The application of the machine-executed method for acquiring cancer risk prediction biomarkers according to any one of claims 1 to 7, the cancer risk prediction biomarkers according to any one of claims 8 to 9, the method for constructing a cancer risk assessment model according to claim 10, the cancer risk assessment model according to claim 11, the computer-readable storage medium for predicting cancer risk according to claim 12, or the apparatus for predicting cancer risk according to claim 13, in any of the following: A) Applications in the development or screening of cancer treatment drugs; B) Applications in the development or screening of cancer prevention drugs; C) Application in the preparation of products for predicting cancer prognosis.