Cancer early screening method and system based on sequencing coverage depth and fragment size characteristics near cfDNA transcription start site

By performing low-depth whole-genome sequencing on plasma cfDNA from patients with malignant biliary and pancreatic tumors, extracting DNA sequencing coverage depth and fragment size features near the TSS, and combining this with a machine learning model, the problem of insufficient sensitivity and accuracy in existing early cancer screening technologies has been solved, achieving low-cost, non-invasive early cancer screening.

CN121592773APending Publication Date: 2026-03-033D BIOMEDICINE SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411127926.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies lack sensitivity and accuracy in early cancer screening, especially non-invasive detection methods for biliary and pancreatic malignancies, which lack specificity and convenience. High-depth sequencing is costly and leads to a high number of false positives. Traditional biopsy methods are not suitable for early screening.

Method used

By performing low-depth whole-genome sequencing on plasma cfDNA from patients with malignant tumors, the DNA sequencing coverage depth pattern and fragment size characteristics near the transcription start site (TSS) are extracted. Combined with machine learning models, an early screening model is established, and cancer is identified using nucleosome imprinting (NF) values ​​and fragment ratio (TF) values.

Benefits of technology

It enables low-cost, non-invasive early cancer screening, improves the accuracy and sensitivity of screening for biliary and pancreatic malignancies, reduces the false positive rate, and provides an efficient non-invasive liquid biopsy method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004997003760000121
    Figure BDA0004997003760000121
  • Figure BDA0004997003760000131
    Figure BDA0004997003760000131
  • Figure BDA0004997003760000132
    Figure BDA0004997003760000132
Patent Text Reader

Abstract

The invention relates to a cancer non-invasive early screening method and system based on sequencing coverage depth and fragment size characteristics near a cfDNA transcription start site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of gene detection technology, specifically relating to a non-invasive early cancer screening method and system based on sequencing coverage depth and fragment size characteristics near the cfDNA transcription start site. Background Technology

[0002] In recent years, liquid biopsy technology has been widely used in clinical practice, especially in assisting in the diagnosis, treatment, and postoperative monitoring of cancer patients. Compared to traditional intraoperative sampling, liquid biopsy obtains samples through blood collection. Cell-free DNA (cfDNA) exists in blood plasma. In healthy individuals, cfDNA mainly originates from the spontaneous apoptosis of lymphocytes in the blood. After a series of digestive processes, the DNA molecules within the cell nucleus are fragmented and released into blood plasma and other bodily fluids. When a tumor develops, a large number of fragmented nucleic acid molecules from specific tumor cells are released into the blood plasma. Currently, the standard method for researching liquid biopsy and early cancer screening is to identify cfDNA released by tumors by detecting mutations in cancer-specific oncogenes or tumor suppressor genes. Whole-genome sequencing (WGS) of cfDNA can identify chromosomal abnormalities in cancer patients, but because the number of abnormal chromosomal changes in tumor-derived cfDNA is very small, especially in the early stages of cancer, detecting these changes can be challenging. A common limitation of using cfDNA mutation detection is the requirement to detect genomic-level mutational differences in cfDNA, such as in non-invasive prenatal diagnosis between fetuses and mothers, or in tumor diagnosis between tumor patients and healthy individuals. Diseases such as myocardial infarction, stroke, and autoimmune diseases are associated with elevated cfDNA levels, which may be a result of tissue damage. However, due to the lack of such differences in DNA mutational alterations, cfDNA cannot be specifically monitored, even though these alterations are remarkably similar to early stages of tumors. Furthermore, not all ctDNA derived from cancer cells carries mutational information, necessitating a new and more sensitive method to modify current conventional methods for cfDNA detection.

[0003] The chromatin state within cells from different tissue origins is not entirely consistent. Open chromatin regions are characterized by loosely connected nucleosomes, facilitating the binding and function of transposases and other cellular regulatory factors. Different cell populations have different open chromatin regions due to varying functional requirements. When tumor cells mutate, their function changes, and the open chromatin regions also alter compared to normal cells. In fact, research on cfDNA reflecting nucleosome footprints was reported as early as 2016. Based on these theoretical foundations, the field of cfDNA-based liquid biopsy for cancer has seen some new and significant breakthroughs. The literature Matthew W. Snyder, Martin Kircher, Andrew J. Hill, et al. Cell-free DNA Comprises an In Vivo Nucleosome Footprint that Informs Its Tissues-Of-Origin. 2016, 164(1-2):57-68. Matthew et al. obtained a genome-wide nucleosome occupancy map by isolating cfDNA from circulating plasma. They found that the distribution pattern of cfDNA is closely related to tissue location. By studying cfDNA, they were able to predict the distribution pattern of nucleosomes and thus determine the specific source of cfDNA. This could be used for non-invasive detection in clinical situations. However, this was limited to the theoretical level and did not involve specific applications. It also lacked a comprehensive multi-omics evaluation of patient cfDNA.

[0004] cfDNA is primarily formed after apoptosis, when DNA is degraded by digestive enzymes and released into bodily fluids such as blood. Open chromatin regions, lacking the protection of nucleosomes, are more easily digested into small fragments, resulting in small and shallow insertions of cfDNA in genome sequencing data. In actively transcribed genes, the promoter region (approximately 150 bp upstream of the TSS) is a nucleosome-depleted region (NDR), an open chromatin region that facilitates the binding of complexes such as transcription factors. The TSS is flanked by well-positioned nucleosome arrays. In contrast, inactive promoters show neither significant nucleosome depletion nor strong nucleosome localization and phasing. Therefore, differences in the openness and nucleosome localization of the TSS site and surrounding chromosome regions lead to differences in the coverage depth pattern of whole-genome DNA sequencing. Thus, differences in WGS sequencing data coverage near the TSS of cfDNA can predict gene expression, and the expression of tissue-specific genes can infer the tissue origin of cfDNA. Further research in the literature Ulz, P., et al., Inferring expressed genes by whole-genome sequencing of plasma DNA. Nature Genetics, 2016, 48(10): p.1273-1278, suggests that this method based on the difference in cfDNA sequencing depth in the TSS region may provide a cheaper way to find signs of cancer in the blood. The authors found that cfDNA in healthy individuals mainly originates from apoptotic leukocytes in the blood, and the cfDNA sequencing depth coverage pattern reflects the gene expression characteristics of leukocytes. This indicates that the difference in sequencing coverage depth of cfDNA near the TSS region between healthy controls and cancer patients can be used as a characteristic identification method for ctDNA released by cancer cells. However, this literature only statistically analyzes the correlation between cfDNA sequencing coverage distribution and gene expression, aiming to infer the expression of related genes using cfDNA sequencing coverage characteristics, and does not explore the use of sequencing data coverage characteristics near TSS sites for early screening and clinical diagnosis of cancer patients.

[0005] Current research on liquid biopsy and early cancer screening conventionally relies on detecting mutations in oncogenes or tumor suppressor genes to identify tumor-released circulating DNA (cfDNA). However, in the blood of early-stage cancer patients, the abundance of mutated circulating tumor DNA (ctDNA) is insufficient for effective detection. Furthermore, many non-tumor-derived cfDNA molecules in the blood exhibit false positives due to clonal hematopoiesis and epithelial clonal amplification. The sensitivity of circulating tumor DNA molecules carrying mutation information in the blood is significantly reduced in the diagnosis of cancer patients, especially in the early stages of cancer. Current methods based on cfDNA mutations require high sequencing depth and computational methods, and high-depth target sequencing is extremely expensive, hindering widespread clinical application. The most widely used clinical imaging examinations and tumor marker detection lack specificity for early screening of biliary and pancreatic malignancies, leading to frequent false positives. Currently, the most traditional clinical method for early screening of biliary and pancreatic malignancies is puncture sampling, followed by pathological section staining for differentiation. This method cannot effectively screen patients with early-stage biliary and pancreatic malignancies, and its accuracy, specificity, convenience, and sensitivity need to be improved due to factors such as puncture site and experience.

[0006] Therefore, there is an urgent need in this field for a non-invasive cancer screening method that can predict early-stage cancer patients using low-depth whole-genome sequencing, and can significantly reduce the cost of early cancer screening and improve screening accuracy. Summary of the Invention

[0007] The inventors of this application discovered that by analyzing fragment omics characteristics extracted from low-depth whole-genome sequencing of plasma cfDNA from patients with malignant tumors (taking biliary and pancreatic malignant tumors as an example), they found that the DNA sequencing coverage depth pattern and fragment size distribution characteristics near the transcription start site (TSS) of plasma cfDNA can be combined as a non-invasive liquid biopsy method to effectively screen for early-stage malignant tumor patients from healthy individuals. Furthermore, based on the combination of DNA sequencing coverage depth pattern and fragment size distribution characteristics near the transcription start site (TSS) of plasma cfDNA, the inventors constructed an early malignant tumor identification model, clarifying and validating its sensitivity and specificity.

[0008] Therefore, as demonstrated in this application, the combination of DNA sequencing coverage depth pattern features and fragment size distribution features near the plasma cfDNA transcription start site (TSS) can identify patients with malignant tumors, such as biliary and pancreatic malignant tumors. This overcomes the shortcomings of existing technologies and has great potential for development and translational application.

[0009] In a first aspect, this disclosure provides a non-invasive early cancer screening system based on sequencing coverage depth and fragment size characteristics near the cfDNA transcription start site (TSS), the system comprising:

[0010] Sequencing coverage depth pattern determination module is used to obtain sequencing coverage depth data near the TSS of cfDNA in the sample;

[0011] The cfDNA fragment feature extraction module is used to obtain cfDNA fragment size feature data in the sample;

[0012] The machine learning classification model building module is used to statistically analyze the differences in sequencing coverage depth and fragment size characteristics near the transcription start site between tumor-derived cfDNA and cfDNA from healthy individuals obtained through low-depth whole-genome sequencing, and to build an early cancer screening model.

[0013] The Independent Validation Queue Evaluation Module is used to validate the predictive performance of the established machine learning classification model through an independent validation queue.

[0014] In some implementations, the nucleosome imprinting (NF) value is calculated by dividing the average coverage of the central region by the average coverage of the surrounding region, thereby determining the sequencing coverage depth pattern. The central region is defined as a region approximately 500 bp upstream and downstream of the TSS [-250 bp, 250 bp], and the surrounding region is defined as approximately 2000 bp upstream of the TSS [-2000 bp, -1000 bp] and downstream of the TSS [1000 bp, 2000 bp].

[0015] In some implementations, reads are extracted from the region approximately [-1000bp, 1000bp] upstream and downstream of the TSS, and the number of short fragments with a length between approximately 150bp and 200bp and the number of long fragments with a length between approximately 220bp and 340bp are calculated. The TF value is then calculated by dividing the number of long fragments by the number of short fragments, thereby determining the cfDNA fragment characteristics.

[0016] In some implementations, the LinearSVC algorithm is used with 10-fold cross-validation to obtain model parameters and reference values, thereby establishing an early cancer screening model.

[0017] In some implementations, the cancer is a malignant tumor of the biliary and pancreatic tract.

[0018] In some implementations, the biliary and pancreatic malignancies include pancreatic cancer, gallbladder cancer, and bile duct cancer.

[0019] In some implementations, the system further includes a sequencing module, the sequencing module comprising:

[0020] The sequencing data alignment unit is used to align the sequencing data to the human reference genome hg19 after removing the sequencing adapters.

[0021] In some implementations, the sequencing module further includes:

[0022] The read filtering unit filters and retains reads that meet the following criteria: quality score greater than 20; insertion length between 150 and 600; proper end-to-end pairing; and reference region of the read does not contain degenerate bases.

[0023] In some implementations, the system further includes a gene screening and TSS determination module, wherein the transcription start site (TSS) of a gene is determined based on transcript annotation of the UCSChg19 genome, and for genes with multiple TSSs, only genes with a difference of less than 50 bp between different TSSs are retained, and their average value is used as the TSS of the gene.

[0024] In some implementations, the machine learning classification model building module includes:

[0025] The sample data classification unit is used to divide the samples into a training queue and a validation queue.

[0026] The model and parameter acquisition unit is used to process the sample data in the training set; in the training queue, 10-fold cross-validation is used to obtain model parameters and reference values.

[0027] The model performance evaluation unit is used to plot the receiver operating characteristic curve of the training queue based on the model prediction value and pathological test results of each sample in the training queue.

[0028] In some implementations, the model includes input variables, model formulas, model parameters, and reference values.

[0029] In a second aspect, this disclosure provides a non-invasive early cancer screening method based on sequencing coverage depth and fragment size characteristics near the transcription start site (TSS) of cfDNA. The method includes: statistically analyzing the differences in sequencing coverage depth and fragment size characteristics near the transcription start site of tumor-derived cfDNA and healthy individual-derived cfDNA using low-depth whole-genome sequencing to establish an early cancer screening model.

[0030] In some implementations, the nucleosome imprinting (NF) value is calculated by dividing the average coverage of the central region by the average coverage of the surrounding region, thereby determining the sequencing coverage depth pattern. The central region is defined as a region approximately 500 bp upstream and downstream of the TSS [-250 bp, 250 bp], and the surrounding region is defined as approximately 2000 bp upstream of the TSS [-2000 bp, -1000 bp] and downstream of the TSS [1000 bp, 2000 bp].

[0031] In some implementations, reads are extracted from the region approximately [-1000bp, 1000bp] upstream and downstream of the TSS, and the number of short fragments with a length between approximately 150bp and 200bp and the number of long fragments with a length between approximately 220bp and 340bp are calculated. The TF value is then calculated by dividing the number of long fragments by the number of short fragments, thereby determining the cfDNA fragment characteristics.

[0032] In some implementations, the LinearSVC algorithm is used with 10-fold cross-validation to obtain model coefficients and reference values, thereby establishing an early cancer screening model.

[0033] In some implementations, the cancer is a malignant tumor of the biliary and pancreatic tract.

[0034] In some implementations, the biliary and pancreatic malignancies include pancreatic cancer, gallbladder cancer, and bile duct cancer.

[0035] The following description and examples illustrate embodiments of the present invention in detail. It should be understood that the present invention is not limited to the specific embodiments described herein and therefore can be modified. Those skilled in the art will recognize that many variations and modifications exist in the present invention, all of which are included within its scope. Attached Figure Description

[0036] Figure 1 The ROC curve is based on the training queue in Example 2.

[0037] Figure 2 The ROC curve is based on the verification queue of Example 2. Detailed Implementation

[0038] The inventors of this application have discovered that by extracting fragment omics features from low-depth whole-genome sequencing of plasma cfDNA of malignant tumor patients (taking biliary and pancreatic malignant tumors as an example, with a sequencing depth of about 3×, and the median sequencing depth of this method is 2.9×), they found that the DNA sequencing coverage depth pattern and fragment size distribution features near the transcription start site (TSS) of plasma cfDNA can be combined as a non-invasive liquid biopsy method to effectively screen out early-stage tumor patients from healthy individuals. This overcomes the shortcomings of existing technologies and has great application value.

[0039] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” used herein are also intended to include the plural forms. Furthermore, the open-ended expressions “comprising” and “including” are to be interpreted as potentially containing structural components or method steps not mentioned, but it should be noted that these open-ended expressions also cover situations where the invention consists only of the stated components and method steps (i.e., they cover the closed-ended expressions “consisting of…”).

[0040] As used throughout, a range is used as a shorthand to describe each and all values ​​within that range. Any value within the range, such as an integer value, a value incremented by one-tenth (when the range ends with one decimal place), or a value incremented by one-hundredth (when the range ends with two decimal places), can be chosen as the end of the range. For example, the range 0.1–10 is used to describe all values ​​within that range, such as 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8…9.5, 9.6, 9.7, 9.8, 9.9, and 10 (in increments of one-tenth), and includes all subranges, such as 0.1–1.0, 2.0–3.0, 4.0–5.0, 6.0–7.0, 8.0–9.0, etc.

[0041] All scientific and technical terms mentioned in this specification have the same meaning as commonly understood by those skilled in the art, and in case of conflict, the definitions in this specification shall prevail. To make the description of this invention easier to understand, some terms are explained below.

[0042] The term "library construction" as used herein refers to the process of repairing cfDNA present in samples such as blood, bodily fluids, or feces and ligating it to a known DNA fragment, i.e., an adapter sequence (also called a linker), so that it can be used for high-throughput DNA sequencing on Illumina equipment. In this invention, "library construction" refers specifically to library construction for high-throughput sequencing.

[0043] The term "high-throughput sequencing" used in this article, also known as Next Generation Sequencing (NGS) or Massively Parallel Sequencing (MPS), refers to a sequencing technology that uses the principle of "sequencing while synthesis" to simultaneously perform parallel sequencing reactions on hundreds of thousands to millions of DNA molecules. The raw image data or electrochemical signals obtained are then analyzed through bioinformatics to ultimately obtain information such as the nucleic acid sequence or copy number of the sample. It is also called high-throughput sequencing, deep sequencing, or second-generation sequencing. The basic procedure of high-throughput sequencing involves randomly fragmenting the DNA to be tested into small fragments, constructing a library through steps such as end repair, ligation of adapter sequences, and PCR, and finally sequencing using sequencers such as Illumina or Ion Torrent.

[0044] The term "cfDNA," as used in this article, also known as cell-free DNA, refers to nucleic acid fragments, approximately 160-180 bp, that exist in a free state outside cells, such as in plasma, serum, or cerebrospinal fluid. These fragments are products of cellular DNA under physiological or pathological conditions. cfDNA can be released into circulation through secretion or cell death processes, such as cell necrosis or apoptosis. Some cfDNA is ctDNA.

[0045] The term "circulating tumor DNA (ctDNA)" as used in this article refers to cell-free DNA (cfDNA) fractions originating from tumors.

[0046] While various embodiments of the invention have been described above, it should be understood that they are provided by way of example only and not as limitations. Many changes to the disclosed embodiments may be made in accordance with the disclosure herein without departing from the spirit or scope of the invention. Therefore, the breadth and scope of the invention should not be limited by any of the embodiments described above.

[0047] All references mentioned herein are incorporated herein by reference. All publications and patent documents cited in this application are incorporated herein by reference for all purposes, and are cited as if they were individually cited.

[0048] Example

[0049] Unless otherwise stated, all materials used in the embodiments herein are commercially available, and all specific experimental methods used to conduct the experiments are conventional experimental methods in the art or are performed according to the steps and conditions recommended by the manufacturer, and can be conventionally determined by those skilled in the art as needed.

[0050] Example 1

[0051] Study cohorts and clinical information

[0052] This study included 55 patients with suspected biliary and pancreatic malignancies based on tumor markers and imaging examinations, 30 patients with gallbladder cancer, 62 patients with pancreatic cancer, and 71 healthy individuals. Blood samples were collected from all patients preoperatively. Each enrolled patient received an accurate pathological diagnosis postoperatively. The training cohort consisted of 58 patients with biliary and pancreatic cancer and 31 healthy controls, while the validation cohort consisted of 89 patients with biliary and pancreatic cancer and 40 healthy controls.

[0053] Blood collection, separation and storage

[0054] We collected whole blood from preoperative cancer patients and healthy controls in 10 ml free nucleic acid preservation tubes (43803, BEAVER, China) and transported them at room temperature. The received whole blood samples were separated into plasma using a two-step centrifugation method. First, the plasma and cellular components were separated by centrifugation at 16,000 g for 10 minutes at 4°C. The supernatant was carefully aspirated, taking care not to aspirate the leukocyte layer, and the hemolysis grade of the plasma was recorded. Samples with a hemolysis grade ≥5 were excluded from subsequent studies. Second, the plasma was centrifuged again at 16,000 g for 15 minutes at 4°C to remove any residual cells or cell debris. The supernatant was transferred to centrifuge tubes and aliquoted into 1 ml tubes. The separated plasma samples were stored at -80°C.

[0055] extraction of cfDNA

[0056] Plasma samples were removed from the -80℃ freezer and placed in a water bath for static incubation at 37℃ for approximately 5 minutes. The plasma was then transferred to a low-temperature refrigerated centrifuge and centrifuged at 16000g for 10 minutes at 4℃. The supernatant was carefully aspirated into centrifuge tubes. cfDNA extraction from plasma was performed using the QIAamp Circulating Nucleic Acid Kit (55114, Qiagen, Shanghai, China) to extract cfDNA from 1 ml of plasma. Specific procedures were followed according to the product instructions. Finally, 30 μl of EB was used to elute the cfDNA. The total amount of extracted cfDNA was quantified using a Qubit real-time analyzer and the corresponding reagents (Q32854, Thermo Fisher, USA). The distribution of cfDNA fragments was detected using an Agilent 2100 bioanalyzer and the corresponding Agilent High Sensitivity DNA Kit & Reagents (5067-4626, Agilent, USA).

[0057] cfDNA library preparation and WGS sequencing

[0058] Samples that passed cfDNA quality control were used for cfDNA library construction and WGS sequencing. Library preparation used the KAPADNA Hyper Prep kit (KK8504, KAPA, USA), with detailed procedures following the product instructions. Each cfDNA sample input was 10 ng. After end-closing and adding A-tails, adapters were ligated, purified, and amplified by PCR for seven cycles to enrich the library. After purification, DNA was eluted with 25 μl of elution buffer. The plasma cfDNA library concentration was determined using a Qubit assay, and fragment distribution was determined using a 4150 TapeStation automated electrophoresis system. Quality-controlled libraries were used for whole-genome sequencing on the NovoSeq 6000 platform, with a sequencing strategy of 2 x 150 bp and a sequencing depth of ~10 G (~3 × 10⁻⁶).

[0059] Extraction of depth coverage pattern and fragment size features from sequencing data near TSS

[0060] cfDNA from the patient's peripheral blood was obtained using DNA sequencing technology. The analysis workflow for the sequencing data is as follows:

[0061] 1) Sequencing data alignment. After removing sequencing adapters from the DNA sequencing data, the sequencing data was aligned to the human reference genome hg19 (genome download link: ftp: / / ftp-) using BWA software (version: 0.7.17-r1188).

[0062] trace.ncbi.nih.gov / 1000genomes / ftp / technical / reference / human_g1k

[0063] (v37.fasta.gz).

[0064] 2) Read filtering. Only reads aligned to chromosomes 1-22 are considered. Reads that retain the following sequences have a sequencing quality score greater than 20;

[0065] The insertion size is between 150 and 600; the ends must be properly paired; the reference region of the read does not contain degenerate bases.

[0066] 3) Gene screening and TSS determination. Transcription start sites (TSS) were determined based on transcript annotations of the UCSC hg19 genome. For genes with multiple TSSs, only those with a difference of less than 50 bp were retained, and their average value was used as the gene's TSS. Furthermore, only genes on autosomes were considered.

[0067] 4) Calculation of nucleosome footprint (NF) values ​​based on sequencing data depth coverage patterns. The central region is defined as the 500bp area upstream and downstream of the TSS [-250bp, 250bp]. The surrounding region is defined as the 2000bp area upstream [-2000bp, -1000bp] and downstream [1000bp, 2000bp]. The NF of a gene is calculated as the average coverage of the central region divided by the average coverage of the surrounding region.

[0068] 5) Calculation of TSS fragment ratio (TF) based on fragment size distribution. Reads from the upstream and downstream regions [-1000bp, 1000bp] of the gene's TSS are extracted from the BAM file, and their insert sizes are calculated. Then, the number of short fragments (150bp-200bp) and the number of long fragments (220bp-340bp) are calculated. The TF value is the ratio of the number of long fragments to the number of short fragments.

[0069] Machine learning classification model building

[0070] 1) Processing the sample data in the training set. In the training queue, for NF values, genes were selected based on the standard deviation of NF, retaining only genes with a standard deviation between [0,2], totaling 21333 genes with NF values. For TF values, genes with a TF value of NaN (missing data or unable to be calculated) in any sample were first removed. Then, genes were selected based on the standard deviation of TF, retaining only genes with a standard deviation between [0,0.5], totaling 18390 genes with TF values. The NF and TF values ​​were then combined to obtain 39723 features as training features. Finally, 10-fold cross-validation was used to obtain model parameters and reference values.

[0071] 2) Evaluate the model's effectiveness. Based on the predicted values ​​from 10-fold cross-validation and pathological test results for each sample in the training cohort, plot the receiver operating characteristic (ROC) curve for the training cohort. The model's predictive effectiveness is evaluated using methods including the area under the ROC curve (AUC, ranging from 0 to 1), and the positive predictive value (PPV, ranging from 0 to 1), specificity (range 0 to 1), accuracy (range 0 to 1), and sensitivity (range 0 to 1) calculated based on the reference value. Higher values ​​indicate better performance.

[0072] Validation of the predictive performance of the rating classification model

[0073] In the validation queue, the classification performance of the model is validated based on the classification model and predicted values ​​determined in the training queue. The process is as follows:

[0074] 1) Variable identification. In the validation queue, 39,723 combinations of 21,333 NF values ​​and 18,390 TF values ​​are identified as variables.

[0075] 2) Model Performance Validation. Based on the model predictions and pathological test results for each sample in the validation cohort, an ROC curve was plotted for the validation cohort. Using the reference value, the validation cohort was divided into a healthy population (same as the training cohort) and a cancer group, and the model's predictive performance was evaluated, including specificity, sensitivity, and accuracy; higher values ​​indicate better performance.

[0076] Example 2

[0077] Study cohorts and clinical information

[0078] This study included 147 patients with biliary and pancreatic malignancies and 71 healthy controls, divided into two independent study cohorts. The training cohort consisted of 31 healthy individuals, 28 patients with pancreatic ductal adenocarcinoma, 14 patients with gallbladder cancer, and 16 patients with bile duct cancer. The validation cohort consisted of 40 healthy individuals, 34 patients with pancreatic ductal adenocarcinoma, 16 patients with gallbladder cancer, and 39 patients with bile duct cancer. Clinical characteristics such as gender and TNM tumor stage are shown in Table 1.

[0079] Table 1. Clinical information of patients in the training and validation cohorts

[0080]

[0081]

[0082] Health and Cancer Classification Scoring Model

[0083] Using a training queue and pathological test results, a scoring model for cancer patients and healthy individuals was constructed using the LinearSVC algorithm. The model consists of three parts: variables, model formulas, and reference values. The process is as follows:

[0084] 1) Model variables and parameters. In the training queue, the model used the combination of NF and TF values ​​of 39,723 genes as feature variables (Table 2). A 10-fold cross-validation method was used to obtain the model parameters and reference values. The final parameters of the model were determined by training with samples from all training queues.

[0085] Table 2. Examples of Model Input Variables

[0086]

[0087] 2) Scoring Model. The scoring model formula is as follows:

[0088]

[0089] Where x i The formulas for calculating the model parameters w and b, which are the input variables, are as follows:

[0090]

[0091] Where λ is the penalty parameter, n is the number of samples, and y i It is the true value of the sample, y i The value can be 1 or -1, where 1 represents cancer and -1 represents health.

[0092] By using a classification model and a combination of the NF and TF values ​​for each sample, the class prediction result for each sample can be obtained.

[0093] 3) Evaluation of the effectiveness of the scoring model.

[0094] In the training cohort, based on the predicted values ​​obtained from 10-fold cross-validation and combined with the true values ​​from pathological examinations, a reference value of 0.281 was determined (the reference value ranges from -1 to 1). When the predicted risk value is less than the reference value, the sample is predicted as healthy; otherwise, it is predicted as biliary or pancreatic malignancy. The ROC curve of the training cohort is then plotted using the predicted and true values. Figure 1 The model's predictive power, AUC, specificity, sensitivity, and PPV were 0.954, 83.9%, 94.8%, and 91.7%, respectively (Table 3). The results indicate that:

[0095] In the training cohort, this risk classification model exhibits high AUC, specificity, sensitivity, and PPV, demonstrating superior predictive performance.

[0096] Independent Validation Queue Evaluation

[0097] To validate the model's efficacy in predicting biliary and pancreatic malignancies, an independent cohort was selected as the validation cohort. The model's efficacy was validated based on the risk classification model and reference value determined in the training cohort. Using a reference value of 0.281, patients in the validation cohort were divided into healthy individuals (same as the training cohort) and patients with biliary and pancreatic malignancies based on the model's predicted values. Then, using the pathological examination results of each patient and healthy individual as the true values, and combining them with the predicted values, ROC curves for the validation cohort were plotted. Figure 2 The AUC, specificity, sensitivity, and PPV of the model for predictive power were 0.954, 80.0%, 94.4%, and 91.3%, respectively (Table 3). The results indicate that this risk prediction model also exhibits high AUC, specificity, sensitivity, and PPV in the validation cohort, meaning it demonstrates superior predictive power.

[0098] Table 3. Model Performance Evaluation

[0099]

[0100]

[0101] While various embodiments of the invention have been described above, it should be understood that they are provided by way of example only and not as limitations. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications will fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A non-invasive early cancer screening system based on sequencing coverage depth and fragment size characteristics near the cfDNA transcription start site (TSS), the system comprising: Sequencing coverage depth pattern determination module is used to obtain sequencing coverage depth data near the TSS of cfDNA in the sample; The cfDNA fragment feature extraction module is used to obtain cfDNA fragment size feature data in the sample; The machine learning classification model building module is used to statistically analyze the differences in sequencing coverage depth and fragment size characteristics near the transcription start site between tumor-derived cfDNA and cfDNA from healthy individuals obtained through low-depth whole-genome sequencing, and to build an early cancer screening model. The Independent Validation Queue Evaluation Module is used to validate the predictive performance of the established machine learning classification model through an independent validation queue.

2. The system according to claim 1, wherein the nucleosome imprinting (NF) value is calculated by dividing the average coverage of the central region by the average coverage of the surrounding region to determine the sequencing coverage depth pattern, wherein the region 500 bp upstream and downstream of the TSS [-250 bp, 250 bp] is defined as the central region, and the region 2000 bp upstream of the TSS [-2000 bp, -1000 bp] and downstream of the TSS [1000 bp, 2000 bp] is defined as the surrounding region.

3. The system according to claim 1, wherein by extracting reads from the upstream and downstream regions of the TSS [-1000bp, 1000bp], the number of short fragments with a length between 150bp and 200bp and the number of long fragments with a length between 220bp and 340bp are calculated; and the TF value is calculated by dividing the number of long fragments by the number of short fragments, thereby determining the cfDNA fragment characteristics.

4. The system according to claim 1, wherein the LinearSVC algorithm is used, and 10-fold cross-validation is employed to obtain model parameters and reference values, thereby establishing an early cancer screening model.

5. The system according to any one of claims 1-4, wherein the cancer is a malignant tumor of the biliary and pancreatic tract.

6. The system according to claim 5, wherein the malignant biliary and pancreatic tumor includes pancreatic cancer, gallbladder cancer, and bile duct cancer.

7. The system according to any one of claims 1-4, further comprising a sequencing module, the sequencing module comprising: The sequencing data alignment unit is used to align the sequencing data to the human reference genome hg19 after removing the sequencing adapters.

8. The system of claim 7, wherein the sequencing module further comprises: The read filtering unit filters and retains reads that meet the following criteria: quality score greater than 20; insertion length between 150 and 600; proper end-to-end pairing; and reference region of the read does not contain degenerate bases.

9. The system according to any one of claims 1-4, the system further comprising a gene screening and TSS determination module, wherein the transcription start site (TSS) of the gene is determined based on transcript annotation of the UCSC hg19 genome, and for genes with multiple TSSs, only genes with a difference of less than 50 bp between different TSSs are retained, and their average value is used as the TSS of the gene.

10. The system according to any one of claims 1-4, wherein the machine learning classification model building module comprises: The sample data classification unit is used to divide the samples into a training queue and a validation queue. The model and parameter acquisition unit is used to process the sample data in the training set; In the training queue, 10-fold cross-validation is used to obtain model parameters and reference values; The model performance evaluation unit is used to plot the receiver operating characteristic curve of the training queue based on the model prediction value and pathological test results of each sample in the training queue.