Cancer non-invasive early screening method based on cfDNA sequencing coverage depth features near tss
By analyzing the sequencing coverage depth characteristics of cfDNA near the TSS through low-depth whole-genome sequencing and machine learning algorithms, the specificity and false positive problems of early screening for biliary and pancreatic malignancies were solved, enabling early and sensitive tumor detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 3D BIOMEDICINE SCI & TECH CO LTD
- Filing Date
- 2022-06-21
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies for early screening of biliary and pancreatic malignant tumors have low specificity and high false positive rates, and cfDNA mutation detection methods are not sensitive enough in early cancer diagnosis.
We analyzed the sequencing coverage depth characteristics of cfDNA near the transcription start site (TSS) using low-depth whole-genome sequencing technology, established an early screening model for biliary and pancreatic malignancies using machine learning algorithms, and optimized the model parameters using the LinearSVC algorithm and 5-fold cross-validation.
It enables early screening of biliary and pancreatic malignancies, reduces testing costs, improves sensitivity and specificity, and can detect changes in tumor characteristics at an earlier stage, providing more reliable clinical data support.
Smart Images

Figure CN117316281B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical testing technology, specifically relating to a non-invasive early cancer screening method based on the sequencing coverage depth characteristics of cfDNA near the TSS. Background Technology
[0002] In recent years, liquid biopsy technology has been widely used in clinical practice, especially in assisting in the diagnosis, treatment, and postoperative monitoring of cancer patients. Compared to traditional intraoperative sampling, liquid biopsy obtains samples through blood collection. Cell-free DNA (cfDNA) exists in blood plasma. In healthy individuals, cfDNA mainly originates from the spontaneous apoptosis of lymphocytes in the blood. After a series of digestive processes, the DNA molecules within the cell nucleus are fragmented and released into blood plasma and other bodily fluids. When a tumor develops, a large number of fragmented nucleic acid molecules from specific tumor cells are released into the blood plasma. Currently, the standard method for researching liquid biopsy and early cancer screening is to identify cfDNA released by tumors by detecting mutations in cancer-specific oncogenes or tumor suppressor genes. Whole-genome sequencing (WGS) of cfDNA can identify chromosomal abnormalities in cancer patients, but because the number of abnormal chromosomal changes in tumor-derived cfDNA is very small, especially in the early stages of cancer, detecting these changes can be challenging. A common limitation of using cfDNA mutation detection is the requirement to detect genomic-level mutational differences in cfDNA, such as in non-invasive prenatal diagnosis between fetuses and mothers, or in tumor diagnosis between tumor patients and healthy individuals. Diseases such as myocardial infarction, stroke, and autoimmune diseases are associated with elevated cfDNA levels, which may be a result of tissue damage. However, due to the lack of such differences in DNA mutational alterations, cfDNA cannot be specifically monitored, even though these alterations are remarkably similar to early stages of tumors. Furthermore, not all ctDNA derived from cancer cells carries mutational information, necessitating a new and more sensitive method to modify current conventional methods for cfDNA detection.
[0003] The chromatin state within cells from different tissue origins is not entirely consistent. Open chromatin regions are characterized by loosely connected nucleosomes, facilitating the binding and function of transposases and other cellular regulatory factors. Different cell populations have different open chromatin regions due to varying functional requirements. When tumor cells mutate, their function changes, and the open chromatin regions also alter compared to normal cells. In fact, research on cfDNA reflecting nucleosome footprints was reported as early as 2016. Based on these theoretical foundations, the field of cfDNA-based liquid biopsy for cancer has seen some new and significant breakthroughs. The literature Matthew W. Snyder, Martin Kircher, Andrew J. Hill, et al. Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. 2016, 164(1-2):57-68. Matthew et al. obtained a genome-wide nucleosome occupancy map by isolating cfDNA from circulating plasma. They found that the distribution pattern of cfDNA is closely related to tissue location. By studying cfDNA, they were able to predict the distribution pattern of nucleosomes and thus determine the specific source of cfDNA. This could be used for non-invasive detection in clinical situations. However, this was limited to the theoretical level and did not involve specific applications. It also lacked a comprehensive multi-omics evaluation of patient cfDNA.
[0004] cfDNA is DNA released into bodily fluids such as blood after apoptosis, where it is degraded by digestive enzymes. Open chromatin regions, lacking the protection of nucleosomes, are more easily digested into small fragments, resulting in small and shallow insertions of cfDNA in genome sequencing data. In actively transcribed genes, the promoter region approximately 150 bp upstream of the TSS (transcription start site) is a nucleosome-depleted region (NDR), an open chromatin region that facilitates the binding of complexes such as transcription factors. The TSS is flanked by well-positioned nucleosome arrays. In contrast, inactive promoters show neither significant nucleosome depletion nor strong nucleosome localization and phasing. Therefore, differences in the openness and nucleosome localization of the TSS site and surrounding chromosome regions lead to differences in the coverage depth pattern of whole-genome DNA sequencing. Thus, differences in WGS sequencing data coverage near the TSS of cfDNA can predict gene expression, and the expression of tissue-specific genes can infer the tissue origin of cfDNA. Further research in the literature Ulz, P., et al., Inferring expressed genes by whole-genome sequencing of plasma DNA. Nature Genetics, 2016, 48(10): p.1273-1278, revealed that this method based on the difference in cfDNA sequencing depth in the TSS region may provide a cheaper way to find signs of cancer in the blood. The authors found that cfDNA in healthy individuals mainly originates from apoptotic leukocytes in the blood, and the cfDNA sequencing depth coverage pattern reflects the gene expression characteristics of leukocytes. This suggests that the difference in sequencing coverage depth of cfDNA near the TSS region between healthy controls and cancer patients can be used as a characteristic identification method for ctDNA released by cancer cells. This literature only statistically analyzes the correlation between cfDNA sequencing coverage distribution and gene expression, aiming to infer the expression of related genes using cfDNA sequencing coverage characteristics, and does not explore the use of sequencing data coverage characteristics near TSS sites for early screening and clinical diagnosis of cancer patients.
[0005] Currently, there are no reported studies on the use of cfDNA sequencing depth coverage differences near the TSS in liquid biopsy for early screening of biliary and pancreatic malignancies. One study most relevant to this invention (PMID:33589745) analyzed plasma cfDNA low-depth WGS data from 2250 patients with cirrhosis, 508 patients with hepatocellular carcinoma, and 476 healthy controls. A machine learning model was built using the differences in sequencing depth near the TSS in the nucleosome localization information of cfDNA reactions. This model effectively screened for hepatocellular carcinoma patients from healthy individuals and patients with cirrhosis, achieving an area under the receiver operating characteristic (ROC) curve (AUC) of 0.97. However, this study is only a preliminary investigation of cfDNA for early screening of hepatocellular carcinoma; further research on other cancers, especially biliary and pancreatic malignancies, requires substantial clinical data. Summary of the Invention
[0006] The purpose of this invention is to provide a non-invasive early cancer screening method based on the sequencing coverage depth characteristics of cfDNA near the TSS. It primarily addresses the technical problems of low specificity and high false positive rate in the early screening of biliary and pancreatic malignancies in existing technologies.
[0007] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0008] A non-invasive early cancer screening method based on the sequencing coverage depth characteristics of cfDNA near the TSS (Transmission Site Side) is proposed. This method involves statistically analyzing the differences in sequencing data coverage depth patterns near the TSS between tumor-derived cfDNA and cfDNA from healthy individuals using low-depth whole-genome sequencing. This analysis establishes an early cancer screening model, enabling non-invasive early cancer screening. The low-depth whole-genome sequencing refers to a sequencing depth of 2X to 4X.
[0009] As a preferred embodiment, the parameter calculation method for the sequencing coverage depth of cfDNA near TSS is as follows: the region 500bp upstream and downstream of TSS [-250bp, 250bp] is defined as the central region, and the region upstream of TSS [-2000bp, -1000bp] and downstream of TSS [1000bp, 2000bp] is defined as the surrounding region; the NF value of the gene is the average coverage of the central region divided by the average coverage of the surrounding region.
[0010] As a preferred implementation scheme, the LinearSVC algorithm is used with the NF value of cfDNA near the TSS of 21334 genes as the feature variable, and the model coefficients are obtained by using 30-fold cross-validation with 30 repetitions to establish an early cancer screening model.
[0011] In a preferred embodiment, the cancer is a malignant tumor of the biliary and pancreatic system. This malignant tumor includes pancreatic cancer, gallbladder cancer, and bile duct cancer.
[0012] This invention also provides a non-invasive early cancer screening system based on the depth characteristics of cfDNA sequencing near the TSS, the system comprising:
[0013] The TSS data feature extraction module is used to obtain sequencing coverage depth feature data of cfDNA in the sample near the TSS.
[0014] The machine learning classification model building module is used to build an early cancer screening model based on the statistical differences in the coverage depth patterns of sequencing data near the TSS between tumor-derived cfDNA and healthy individual-derived cfDNA.
[0015] The Independent Validation Queue Evaluation Module is used to validate the predictive performance of the established machine learning classification model through an independent validation queue.
[0016] As a preferred embodiment, the TSS data feature extraction module includes:
[0017] The sequencing data alignment unit is used to align the sequencing data to the human reference genome hg19 after removing the sequencing adapters.
[0018] The reads filtering unit is used to filter and screen sequencing data. The filtering criteria are: only reads aligned to chromosomes 1-22 are considered; the quality score is greater than 20; the insert length is between 150 and 600; the paired ends must be proper; and the reference region of the read does not contain degenerate bases.
[0019] The gene screening and TSS determination unit is used to determine the TSS of gene transcription based on transcript annotation of the UCSChg19 genome; for genes with multiple TSS, only those with a difference of less than 50 bp are retained, and their average value is used as the TSS of the gene, while only genes on autosomes are considered.
[0020] The NF value calculation unit is used to obtain the sequencing coverage depth parameter of cfDNA near TSS. The calculation method is as follows: the region 500bp upstream and downstream of TSS [-250bp, 250bp] is defined as the central region; the region 2000bp upstream [-2000bp, -1000bp] and downstream [1000bp, 2000bp] of TSS is defined as the surrounding region; the NF value of a gene is the average coverage of the central region divided by the average coverage of the surrounding region.
[0021] As a preferred embodiment, the machine learning classification model building module includes:
[0022] The sample data classification unit is used to divide the samples into training and test sets in a 4:1 ratio, and to ensure that the distribution ratio of healthy controls and various cancer samples is consistent in the two sets.
[0023] The model parameter acquisition unit is used to process the sample data in the training set. In the training queue, genes are screened by the standard deviation of NF, and only genes with a standard deviation between [0,2] are retained. Then, the model parameters are obtained by using 30-fold cross-validation with 30 repetitions.
[0024] The model performance evaluation unit is used to plot the receiver operating characteristic curve of the training queue based on the model prediction value and pathological test results of each sample in the training queue.
[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0026] Currently, the imaging and serum marker diagnostic methods widely used in clinical practice for differentiating malignant tumors of the biliary and pancreatic systems have poor specificity, leading to false positives, and the diagnostic value of cfDNA-related mutation information tends to be late-stage. This invention, focusing on cfDNA fragmentation distribution as a cancer early diagnosis system, has the following characteristics: a. It employs low-depth whole-genome sequencing, significantly reducing sequencing costs compared to ultra-high-depth or high-depth target sequencing; b. It can detect abnormal changes in fragments earlier than cfDNA mutations at an earlier stage of cancer, making the method more sensitive than detecting cfDNA mutation information; c. The patients used are all at an earlier stage, allowing for earlier detection of related characteristic changes than patients used in the most recent studies, and the results have been validated through systematic and scientific analysis, demonstrating better diagnostic efficacy than existing classification models.
[0027] This invention employs low-depth (2X–4X) whole-genome sequencing of plasma cfDNA from 60 clinically detected cases of biliary and pancreatic malignancies and 31 healthy controls. Based on the differences in open genes in different tissues and the varying sequencing depth coverage patterns near and around the TSS in different regions of the whole genome between healthy individuals and those from different cancer sources, these features are used to establish an early screening model for biliary and pancreatic malignancies in the studied cohort, and the model's efficacy is evaluated. More importantly, this invention independently validates the early screening model in 47 patients with biliary and pancreatic tumors and 20 healthy individuals. This method uses a more precise analysis of tumor DNA in the blood to find clues for early tumor screening, providing more robust and reliable data support for precise clinical applications. Attached Figure Description
[0028] Figure 1 This is the ROC curve of the training set in Embodiment 1 of the present invention.
[0029] Figure 2 This is the ROC curve of the test set in Embodiment 1 of the present invention.
[0030] Figure 3 This is the ROC curve of the independent validation cohort subjects in Embodiment 1 of the present invention. Detailed Implementation
[0031] The technical solution of the present invention will be described in detail below with reference to the embodiments. Unless otherwise specified, all reagents and biological materials used below are commercial products.
[0032] Example 1
[0033] (1) Study cohort and clinical information
[0034] This study included 107 patients diagnosed with pancreatic and biliary tumors (pancreatic cancer, gallbladder cancer, and bile duct cancer) based on tumor markers, imaging examinations (such as ultrasound and abdominal CT scans), and pathological examinations, as well as 51 healthy individuals. Blood samples were collected from both patients and healthy individuals before surgery. Each enrolled patient received an accurate diagnosis based on postoperative pathological examination results.
[0035] (2) Blood collection, separation and storage
[0036] Whole blood samples from preoperative cancer patients and healthy controls were collected in 10 ml free nucleic acid preservation tubes (REF43803, BD, USA) and transported at room temperature. The received whole blood samples were centrifuged using a two-step centrifugation method to separate plasma. First, plasma and cellular components were separated by centrifugation at 1600g for 10 minutes at 4°C. The supernatant was carefully aspirated, taking care not to aspirate the leukocyte layer, and the hemolysis grade of the plasma was recorded. Samples with a hemolysis grade ≥5 were excluded from subsequent studies. Second, the plasma was centrifuged again at 16,000g for 15 minutes at 4°C to remove any residual cells or cell debris. The supernatant was transferred to centrifuge tubes and aliquoted into 1 ml tubes. The separated plasma samples were stored at -80°C.
[0037] (3) Extraction of cfDNA
[0038] Plasma samples were removed from the -80℃ freezer and placed in a water bath for static incubation at 37℃ for approximately 5 minutes. The plasma was then transferred to a low-temperature refrigerated centrifuge and centrifuged at 4℃, 1600g for 10 minutes. The supernatant was carefully aspirated into centrifuge tubes. Plasma cfDNA extraction was performed using the QIAamp Circulating Nucleic Acid Kit (55114, Qiagen, Shanghai, China) to extract cfDNA from 1 ml of plasma. Specific procedures were followed according to the product instructions. Finally, 30 μl of EB was used to elute the cfDNA. The total amount of extracted cfDNA was quantified using a Qubit real-time analyzer and the corresponding reagents (Q32854, Thermo Fisher, USA). The distribution of cfDNA fragments was detected using an Agilent 2100 bioanalyzer and the corresponding Agilent High Sensitivity DNA Kit & Reagents (5067-4626, Agilent, USA).
[0039] (4) cfDNA library preparation and WGS sequencing
[0040] Samples that passed cfDNA quality control were used for cfDNA library construction and WGS sequencing. Library preparation used the KAPADNA Hyper Prep kit (KK8504, KAPA, USA), with detailed procedures following the product instructions. Each cfDNA sample input was 10 ng. After end-closing and adding A-tails, adapters were ligated, purified, and amplified by PCR for seven cycles to enrich the library. After purification, DNA was eluted with 25 μl of elution buffer. Plasma cfDNA library concentration was measured using Qubit, and fragment distribution was determined at 4150 nm. Quality-controlled libraries were used for whole-genome sequencing on the NovoSeq 6000 platform, with a sequencing strategy of 2 x 150 bp and a sequencing depth of ~10 G (~3 × 10⁻⁶).
[0041] (5) Extraction of deep pattern features from sequencing data near TSS
[0042] DNA sequencing technology was used to obtain cfDNA from the patient's peripheral blood. The analysis workflow for the sequencing data is as follows:
[0043] 1) Sequencing data alignment. After removing the sequencing adapters from the DNA sequencing data, the sequencing data was aligned to the human reference genome hg19 using BWA software (version: 0.7.17-r1188) (genome download link: ftp: / / ftp-trace.ncbi.nih.gov / 1000genomes / ftp / technical / reference / human_g1k_v37.fasta.gz).
[0044] 2) Read filtering. Only reads aligned to chromosomes 1-22 are considered; the quality score is greater than 20; the insertion size is between 150 and 600; the ends must be proper pairs; and the reference region of the read does not contain degenerate bases.
[0045] 3) Gene screening and TSS determination. Transcription start sites (TSS) were determined based on transcript annotations of the UCSC hg19 genome. For genes with multiple TSSs, only those with a difference of less than 50 bp were retained, and their average value was used as the gene's TSS. Furthermore, only genes on autosomes were considered.
[0046] 4) Nucleosome footprint (NF) calculation. The central region is defined as the 500bp area upstream and downstream of the TSS [-250bp, 250bp]. The surrounding region is defined as the 2000bp area upstream [-2000bp, -1000bp] and downstream [1000bp, 2000bp] of the TSS. The NF value of a gene is calculated as: the average coverage of the central region divided by the average coverage of the surrounding region.
[0047] (6) Establishment of machine learning classification models
[0048] 1) Divide the samples into training and test sets. Divide all samples into training and test sets in a 4:1 ratio, ensuring that the distribution ratio of healthy controls and samples of various cancer types remains consistent in both sets.
[0049] 2) Processing the sample data in the training set. In the training queue, genes were screened using the standard deviation of NF, retaining only those with a standard deviation between [0,2], totaling 21334 genes. Then, 30-fold 5-fold cross-validation was used to obtain the model parameters.
[0050] 3) Evaluate model efficacy. Based on the model's predicted values and pathological test results for each sample in the training cohort, plot the receiver operating characteristic (ROC) curve for the training cohort. Using the predicted values as the standard, establish a series of thresholds to divide the training cohort into healthy individuals and cancer patients. Then, using the pathological test results as the true values, evaluate the model's predictive efficacy. The model's predictive efficacy evaluation methods include the area under the ROC curve (AUC, ranging from 0 to 1), positive predictive value (PPV, ranging from 0 to 1), specificity (ranging from 0 to 1), accuracy (ranging from 0 to 1), and sensitivity (ranging from 0 to 1). Higher values indicate better efficacy.
[0051] (7) Validation of the predictive power of the rating classification model
[0052] In the independent validation queue, the classification performance of the model is validated based on the classification model and predicted values determined in the training queue. The process is as follows:
[0053] 1) Variable identification. In the independent validation cohort, the NF values of 21,334 genes were used as variables.
[0054] 2) Model performance validation. Based on the expression levels of molecular markers and pathological examination results for each sample in the test set, ROC curves were plotted for the test set. Using the predicted values as the standard, the validation cohort was divided into a healthy population (same as the training cohort) and a cancer group, and the model's predictive performance was evaluated, including specificity, sensitivity, and accuracy; higher values indicate better performance.
[0055] Example 2
[0056] (1) Study cohort and clinical information
[0057] This study included two cohorts of 107 patients diagnosed with malignant biliary and pancreatic tumors based on tumor markers, imaging examinations (such as ultrasound and abdominal CT scans), and pathological examinations, as well as 51 healthy individuals. Blood samples were collected from both patients and healthy controls preoperatively. Each enrolled patient received an accurate diagnosis postoperatively based on pathological examination results.
[0058] The training and test sets included a total of 60 patients (29 with pancreatic cancer, 15 with gallbladder cancer, and 16 with bile duct cancer) and 31 healthy individuals (Table 1). All samples in the training cohort were divided into the training and test sets at a 4:1 ratio, ensuring consistent distribution of healthy controls and samples of each cancer type in both sets. Table 1 shows the grouping information for healthy controls and patients in the training and test sets. Analysis results indicated no significant differences in the gender ratio and the distribution of healthy controls and cancer patients between the training and test sets.
[0059] Table 1: Training and Test Set Information
[0060]
[0061] The independent validation cohort included 47 patients with biliary and pancreatic tumors (19 with pancreatic cancer, 8 with gallbladder cancer, and 20 with bile duct cancer) and 20 healthy individuals (Table 2). Table 2 shows the participant information in the training set and the independent validation cohort. The analysis results showed that there were no significant differences in the gender ratio and the distribution of healthy controls and cancer patients between the training set and the independent validation cohort samples.
[0062] Table 2: Training Set and Independent Validation Queue Information
[0063]
[0064]
[0065] (2) Health and Cancer Classification Scoring Model
[0066] Using a training queue and pathological test results, a scoring model for cancer patients and healthy individuals was constructed using the LinearSVC algorithm. The model consists of three parts: variables, model formulas, and predicted values. The process is as follows:
[0067] ① Model variables and parameters. In the training queue, the model uses the NF values of 21,334 genes as feature variables (model input variables are shown in Table 3). The model coefficients are obtained using 30-fold repeated 5-fold cross-validation.
[0068] Table 3
[0069] SEQ ID NO Model variable NF value 1 A1BG NF1 2 A1CF NF2 … …… …… 21333 ZYX NF21333 21334 ZZEF1 NF21334
[0070] ② Scoring Model. The scoring model formula is as follows:
[0071]
[0072] Where x i The formulas for calculating the model parameters w and b, which are the input variables, are as follows:
[0073]
[0074] Where λ is the penalty parameter, n is the number of samples, and y i These are the true values of the sample, where 1 represents cancer and -1 represents healthy.
[0075] Using a classification model and the NF value for each sample, the class prediction result for each sample can be obtained.
[0076] ③ Model performance evaluation.
[0077] To construct a diagnostic classification model for cancer patients, predicted values were used to divide the training cohort samples into healthy and cancer patients. Using pathological examination results as the ground truth, ROC curves were plotted based on the predicted values and pathological results in the cohort. The training set AUC value reached as high as 1. (See also...) Figure 1 The table shows the ROC curve for the training set. The PPV (accuracy), specificity, and sensitivity of the trained model were 100%, 100%, and 100%, respectively (Table 4). The results indicate that this risk prediction model has high sensitivity and NPV in the training set, and its predictive power for early cancer diagnosis is superior.
[0078] ④ Validation of the predictive performance of the discriminant model.
[0079] To verify the effectiveness of the discriminant model, a threshold was set based on the predicted values. Participants in the test set were used, with pathological test results as the true value. The model's effectiveness was verified based on the classification model and predicted values determined in the training queue. An ROC curve was plotted for the test set, and the AUC value reached 0.88. See also... Figure 2 The table shows the ROC curve for the test set. The model's predictive performance was evaluated, including accuracy, specificity, and sensitivity, which were 100%, 100%, and 66.7%, respectively (Table 4). The results indicate that this risk prediction model also exhibits high specificity, sensitivity, and accuracy on the test set, meaning its predictive performance is superior.
[0080] Table 4: Performance Evaluation of the Model with 21,334 Variables
[0081]
[0082]
[0083] (3) Evaluation of Independent Validation Queue
[0084] To further validate the effectiveness of the discriminative model, patients in the independent validation cohort were divided into healthy and cancer patient groups (same as the training and test sets) using a threshold set at 0.366 based on the predicted values. Using pathological examination results as the ground truth, the model's effectiveness was validated based on the classification model and predicted values determined in the training cohort. The ROC curve for the validation cohort was plotted, with an AUC value of 0.90. See also... Figure 3The ROC curves for the independent validation cohort are shown in Table 5. The model's predictive power, including accuracy, specificity, and sensitivity, was evaluated, yielding 92.5%, 85%, and 78.7%, respectively (Table 5). The results indicate that this risk prediction model also exhibits high specificity, sensitivity, and accuracy in the validation cohort, meaning its predictive power is superior.
[0085] Table 5: Performance evaluation of the model with 21,334 variables in the independent validation cohort.
[0086]
[0087] The above are merely some preferred embodiments of the present invention, and the present invention is not limited to the contents of these embodiments. For those skilled in the art, various changes and modifications can be made within the scope of the present invention's technical solutions, and any such changes and modifications are within the protection scope of the present invention.
Claims
1. A non-invasive early cancer screening method based on the sequencing coverage depth characteristics of cfDNA near the TSS (Transmission Sequence Side), the method comprising: statistically analyzing the differences in sequencing coverage depth patterns of tumor-derived cfDNA and healthy-derived cfDNA near the TSS using low-depth whole-genome sequencing, establishing an early cancer screening model to achieve non-invasive early cancer screening; the parameter calculation method for the sequencing coverage depth of cfDNA near the TSS is as follows: defining the region 500bp upstream and downstream of the TSS [-250bp, 250bp] as the central region, and the regions upstream [-2000bp, -1000bp] and downstream [1000bp, 2000bp] as the surrounding regions; the NF value of the gene is the average coverage of the central region divided by the average coverage of the surrounding regions; the cancers are pancreatic cancer, gallbladder cancer, and bile duct cancer.
2. The non-invasive early cancer screening method based on the depth characteristics of cfDNA sequencing near the TSS according to claim 1, characterized in that: The LinearSVC algorithm was used with the NF values of cfDNA near the TSS of 21334 genes as feature variables. The model coefficients were obtained using 30-fold cross-validation with 30 replicates, and an early cancer screening model was established.
3. A non-invasive early cancer screening system based on the depth characteristics of cfDNA sequencing near the TSS, characterized in that, The system includes: The TSS data feature extraction module is used to obtain sequencing coverage depth feature data of cfDNA in the sample near the TSS. The machine learning classification model building module is used to build an early cancer screening model based on the statistical differences in the coverage depth patterns of sequencing data near the TSS between tumor-derived cfDNA and healthy individual-derived cfDNA. The independent validation queue evaluation module is used to validate the predictive performance of the established machine learning classification model through an independent validation queue. The TSS data feature extraction module includes: The sequencing data alignment unit is used to align the sequencing data to the human reference genome hg19 after removing the sequencing adapters. The reads filtering unit is used to filter and screen sequencing data. The filtering criteria are: only reads aligned to chromosomes 1-22 are considered; the quality score is greater than 20; the insert length is between 150 and 600; the paired ends must be proper pairs; and the reference region of the read does not contain degenerate bases. The gene screening and TSS determination unit is used to determine the TSS of gene transcription based on transcript annotation of the UCSC hg19 genome; for genes with multiple TSS, only those with a difference of less than 50 bp are retained, and their average value is used as the TSS of the gene, while only genes on autosomes are considered. The NF value calculation unit is used to obtain the sequencing coverage depth parameter of cfDNA near TSS. The calculation method is as follows: the region 500bp upstream and downstream of TSS [-250bp, 250bp] is defined as the central region; the region 2000bp upstream [-2000bp, -1000bp] and downstream [1000bp, 2000bp] of TSS is defined as the surrounding region; the NF value of a gene is the average coverage of the central region divided by the average coverage of the surrounding region.
4. The non-invasive early cancer screening system based on the depth of cfDNA sequencing near the TSS according to claim 3, characterized in that, The machine learning classification model building module includes: The sample data classification unit is used to divide the samples into training and test sets in a 4:1 ratio, and to ensure that the distribution ratio of healthy controls and various cancer samples is consistent in the two sets. The model parameter acquisition unit is used to process the sample data in the training set. In the training queue, genes are screened by the standard deviation of NF, and only genes with a standard deviation between [0,2] are retained. Then, the model parameters are obtained by using 30-fold cross-validation with 30 repetitions. The model performance evaluation unit is used to plot the receiver operating characteristic curve of the training queue based on the model prediction value and pathological test results of each sample in the training queue.
Citation Information
Patent Citations
Non-invasive cancer early screening system based on cfDNA omics characteristics
CN113160889A
Disease Detection in Liquid Biopsies
US20230042332A1