A non-invasive early screening method and system for cancer based on cfDNA fragment length distribution characteristics

Through low-deep whole-gene sequencing and machine learning algorithms, the length distribution characteristics of cfDNA fragments are distinguished from cancer patients and healthy individuals, and the problems of low specificity and high false positives in liquid biopsy are solved, achieving efficient and low-cost early screening of biliary and pancreatic malignant tumors.

CN117316278BActive Publication Date: 2025-08-193D BIOMEDICINE SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210704961.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-08-19
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

In the prior art, the method of liquid biopsy for early screening of cancer has low specificity and high false positive rate, especially in early screening of bile-pancreatic malignant tumors, and high-deep sequencing is expensive and cannot be widely used.

Method used

Through low-deep whole-gene sequencing, the differences in fragment length distribution characteristics of tumor-derived cfDNA and healthy individuals were counted. The LinearSVC algorithm was used to establish an early screening model, and the cfDNA fragment length distribution characteristics were used to distinguish cancer patients and healthy individuals, and a non-invasive early screening method based on the length distribution characteristics of cfDNA fragment length distribution was established.

Benefits of technology

High specific screening of early bile-pancreatic malignant tumors is achieved, which reduces detection costs and improves the accuracy of screening, eliminates interference from clonal hematopoietic mutations, and provides more reliable early diagnosis data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117316278B_ABST
    Figure CN117316278B_ABST
Patent Text Reader

Abstract

The present invention discloses a non-invasive early screening method and system for cancer based on the length distribution characteristics of cfDNA fragments. This method uses a relatively low-depth whole-genome sequencing method to statistically analyze the differences in the fragment length distribution characteristics of tumor-derived cfDNA and healthy individual-derived cfDNA, establish an early screening model for cancer, and achieve non-invasive early screening for cancer. The scheme of the present invention focuses on the characteristic of blood cfDNA fragment size to distinguish ctDNA from non-tumor-derived cfDNA. It does not rely on the mutation detection of oncogenes or tumor suppressor genes, and eliminates the interference caused by clonal hematopoietic mutations. Secondly, the data of the embodiments of the present invention show that the size characteristics of blood cfDNA fragments can be used to distinguish between healthy people and early-stage tumor patients. Finally, due to the use of low-depth whole-genome sequencing technology, the detection costs involved in the scheme of the present invention are greatly reduced, which is conducive to future applications in the field of early screening of malignant tumors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical detection technology, and specifically relates to a non-invasive early screening method and system for cancer based on the cfDNA fragment length distribution characteristics. Background Art

[0002] The majority of human cancer mortality worldwide is due to late diagnosis, which results in poor therapeutic intervention. Therefore, early diagnosis of cancer is crucial. Traditional biomarker imaging technologies play an important role in cancer diagnosis; however, the specificity of traditional serum biomarkers is unsatisfactory for treatment guidance. Furthermore, due to radiation exposure and economic issues, imaging technologies cannot be used for real-time monitoring. From the 2016 FDA approval of the first plasma cfDNA (cell-free DNA) liquid biopsy product based on EGFR gene mutations to the demonstration that bTMB (blood tumor mutation burden) can predict the response to immunotherapy, liquid biopsies have made significant progress in the field of cancer treatment. Liquid biopsies are increasingly being considered for early cancer diagnosis, treatment guidance, and recurrence monitoring. They can provide information about the tumor. Furthermore, liquid biopsies offer a non-invasive alternative to traditional solid biopsies, which cannot be performed consistently in certain situations or in real time. Despite these numerous advantages, limitations remain, such as a lack of consensus on testing methods, the difficulty of analyzing large amounts of sequencing information, and insufficient evidence-based medicine.

[0003] Currently, the conventional approach for liquid biopsy and early cancer screening is to identify tumor-released cfDNA (cfDNA) by detecting mutations in oncogenes or tumor suppressor genes. Unfortunately, the abundance of circulating tumor DNA (ctDNA) in the blood is typically far lower than that of non-cancer-related DNA fragments, making its detection difficult, especially in the early stages of cancer. Previous studies have found that accurate tumor profiles can only be obtained when the abundance of ctDNA in cfDNA is 10% or greater. However, with the exception of some advanced tumors that release large amounts of ctDNA, the abundance of ctDNA in most cancer patients does not meet this standard. Currently, the main approach to improving the sensitivity and accuracy of ctDNA detection is to increase sequencing depth. However, this increase in sequencing depth can lead to false positives, as non-tumor DNA can also harbor various tumor-related mutations. High-depth sequencing is also extremely expensive, making it impractical for widespread clinical application. Moreover, mutational studies are limited to patients with advanced metastatic disease. In the literature Razavi, P., Li, BT, Brown, DNet al. High intensity sequencing reveals the sources of plasma circulating cell-free DNA variants. Nat Med 25, 1928–1937 (2019). Pedram et al. conducted an ultra-deep detection of 508 specific genes in the peripheral blood of 124 metastatic cancer patients and 47 healthy people with more than 60,000 times, and found a large number of cfDNA mutations, 81.6% from healthy people and 53.2% from cancer patients, all of which were caused by clonal hematopoiesis of white blood cells, not tumor cells. This will result in a high misjudgment rate when identifying cancer patients through mutations in specific target genes. Moreover, due to the great differences among cancer patients, after limiting specific target genes, some special cancer patients may be missed. Therefore, there are certain limitations to the detection of specific panel mutations in cfDNA. These problems have always restricted the application of liquid biopsy.

[0004] Another discovery by scientists may address this issue. Previous studies have shown that the length of cfDNA released into the blood by different cells varies. For example, many cfDNA fragments are 167 base pairs (similar to the length of a nucleosome), potentially related to caspase-dependent DNA cleavage during apoptosis. Infants' cfDNA fragments are significantly shorter than their mothers', a characteristic that has been exploited for prenatal diagnosis. These studies suggest that cfDNA from different cell sources exhibits unique length patterns, potentially serving as a signal to distinguish cfDNA origin. For various reasons, cancer genomes are disorganized in their packaging. This means that when cancer cells die, they release DNA into the blood in a chaotic manner. This difference in cfDNA fragment length distribution is likely present in circulating tumor cell DNA and non-tumor cell DNA. Studies have also reported that cfDNA fragment length, as a distinguishing characteristic between tumor and non-tumor cells, could mitigate the low sensitivity and false positives associated with detecting mutations in cfDNA. However, the few previous studies investigating this hypothesis have yielded conflicting results, hindering progress in this area. Therefore, scientific and systematic research on cfDNA fragmentation urgently needs to be improved and optimized.

[0005] Currently, there are a few reports on the use of cfDNA fragment size in liquid biopsy and early cancer screening. The two studies most similar to the present study (PMID: 30404863 and PMID: 31142840) are:

[0006] In the literature Mouliere, F., et al., Enhanced detection of circulating tumor DNA by fragment size analysis. Sci Transl Med, 2018.10(466). Mouliere et al. used a low-depth (0.4×) whole-genome sequencing method to detect the fragment length characteristics of cfDNA extracted from 344 plasma samples collected from 200 cancer patients (including 18 different types of cancer) and 65 plasma samples from healthy people. The analysis found that ctDNA fragments with cancer mutations are generally 20-40bp shorter than nucleosomal DNA fragments (167bp) and are enriched in the 90-150bp range. The researchers then used a method of selective sequencing to enrich short fragments to increase the abundance of ctDNA. The ratio of fragments in different length ranges was used as a feature to distinguish tumor blood samples from healthy blood samples through a machine learning algorithm. However, this method selectively enriches fragments, and fragments of other lengths also contain effective information such as tumor mutations. Taking only 90-150bp fragments will lose some information. In addition, the study used a training model combining the total proportion of fragment distribution and mutation characteristics, and lacked systematic consideration of the fragment size ratio information of specific functional sites in the entire genome. Therefore, this method still has much room for improvement and enhancement of cfDNA fragment characteristics as an early diagnosis indicator for cancer.

[0007] In the literature Cristiano, S., et al., Genome-wide cell-free DNA fragmentation in patients with cancer. Nature, 2019. 570(7761): p. 385-389. Cristiano et al. used whole-genome sequencing to detect blood samples from 208 patients with different stages of breast cancer, colorectal cancer, lung cancer, ovarian cancer, pancreatic cancer, gastric cancer, and bile duct cancer, and 215 healthy people. Circulating tumor DNA was found in 57% to more than 99% of the patients. This method is based on the detection of cfDNA using a lower-coverage whole-genome sequencing method. The number of short and long fragment cfDNA sequences mapped to different regions of the genome was analyzed in non-overlapping windows covering the genome to establish a model for predicting early-stage cancer. Moreover, the patients enrolled in the above-mentioned research institutes were mostly advanced cancer patients with a limited number of cancer types, and the fragment size statistics adopted a "one-size-fits-all" standard for multiple cancer types (short fragments were defined as 100 to 150 bp, and long fragments were defined as 151 to 220 bp). There was a lack of more in-depth large-sample statistics and research on determining the cancer-specific fragment optimization parameters for a single cancer type.

[0008] Based on this, it is very necessary for those skilled in the art to design a non-invasive early cancer screening method that can predict early cancer patients through low-depth whole-genome sequencing, and can significantly reduce the cost of early cancer screening and improve screening accuracy. Summary of the Invention

[0009] The purpose of this invention is to provide a non-invasive early cancer screening method based on the length distribution characteristics of cfDNA fragments. This method primarily addresses the technical issues of low specificity and high false positive rate in the early screening of biliary and pancreatic malignancies in existing technologies.

[0010] The technical solutions adopted by the present invention to solve the above technical problems are as follows:

[0011] A non-invasive early cancer screening method based on the fragment length distribution characteristics of cfDNA (cfDNA). This method uses relatively low-depth whole-genome sequencing to statistically analyze the differences in fragment length distribution characteristics between tumor-derived cfDNA and cfDNA from healthy individuals, thereby establishing an early cancer screening model and achieving non-invasive early cancer screening. The relatively low-depth whole-genome sequencing method refers to a sequencing depth of 2X to 4X.

[0012] As a preferred embodiment, the normalized z-score of the number of short and total cfDNA fragments within the 504 5Mb regions of the genome is calculated as the feature input value for model training. The total number of fragments is the sum of the defined number of long and short fragments.

[0013] The cfDNA fragments include short fragments and long fragments, wherein the length range of the cfDNA short fragments is [130,177] bp, and the length range of the long fragments is [177,237] bp.

[0014] As a preferred embodiment, the LinearSVC algorithm is adopted, and the 5-fold cross validation method is repeated 30 times to obtain the model coefficients and establish an early screening model for cancer.

[0015] As a preferred embodiment, the cancer is a pancreatic or biliary malignancy, including pancreatic cancer, gallbladder cancer, and bile duct cancer.

[0016] The present invention also provides a non-invasive early cancer screening system based on the cfDNA fragment length distribution characteristics, the system comprising:

[0017] cfDNA fragment feature extraction module, used to obtain cfDNA fragment size feature data in samples;

[0018] A machine learning classification model building module is used to establish an early cancer screening model based on the statistical differences in the fragment size characteristics of tumor-derived cfDNA and cfDNA from healthy individuals;

[0019] The independent validation cohort evaluation module is used to verify the predictive effectiveness of the established machine learning classification model through an independent validation cohort.

[0020] As a preferred embodiment, the cfDNA fragment feature extraction module includes:

[0021] Sequencing data alignment unit, used to remove sequencing data sequencing adapters and align the sequencing data to the human reference genome hg19;

[0022] The cfDNA fragment statistics unit is used to count cfDNA fragment length data information. The hg19 autosome is divided into 504 adjacent, non-overlapping window segments, each with a length of 5 Mb. Within each window region, the ratio of the number of cfDNA fragments greater than 130 bp and less than 177 bp to the number of cfDNA fragments greater than 177 bp and less than 237 bp is counted. Finally, the number of long and short cfDNA fragments within each 5 Mb interval is obtained.

[0023] The cfDNA fragment feature determination unit is used to determine the interval with the largest difference in fragment distribution between cancer patients and healthy controls based on the difference distribution of fragments between cancer patients and healthy controls; define the short fragment range [130,177] and the long fragment range [177,237], and then calculate the standardized z-score of the short fragment cfDNA and the total number of fragments in each of the 504 windows as the feature input value for model training.

[0024] As a preferred embodiment, the machine learning classification model building module includes:

[0025] The sample data classification unit is used to divide the samples into training and test sets in a ratio of 4:1, and to ensure that the distribution ratios of healthy controls and various cancer samples in the two sets are consistent;

[0026] The model parameter acquisition unit is used to process the sample data in the training set; in the training queue, the model parameters are obtained using 30 repeated 5-fold cross validation methods;

[0027] The model effectiveness evaluation unit is used to draw the receiver operating characteristic curve of the training cohort based on the model prediction value and pathological test results of each sample in the training cohort.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] The scheme of the present invention first focuses on the characteristic of blood cfDNA fragment size to distinguish ctDNA from cfDNA of non-tumor origin, and does not rely on mutation detection of oncogenes or tumor suppressor genes, which eliminates the interference caused by clonal hematopoietic mutations; secondly, the data of the embodiment of the present invention show that the use of blood cfDNA fragment size characteristics can distinguish healthy people and early tumor patients, and the low ctDNA content in early tumor patients does not lead to the inability to detect ctDNA signals in tumor patients; finally, due to the use of low-depth whole-genome sequencing technology, the detection cost involved in the scheme of the present invention is greatly reduced compared with other full-depth or ultra-depth NGS detection methods. These advantages are conducive to the future application of this scheme in the field of early screening of malignant tumors.

[0030] The present study used plasma samples from 60 cases of pancreatic and biliary tumors detected in the clinic and 31 healthy controls to perform low-depth (2X to 4X) whole-genome testing of cfDNA. Considering the positional distribution of cfDNA fragment size across the genome, an analysis system was established to distinguish between tumor cell and non-tumor cell DNA. Through systematic statistical analysis of the size and number of DNA fragments covering different regions of the whole genome in the blood, the cfDNA fragment size characteristics were trained and tested in the study cohort to establish a biomarker diagnostic model for early screening of pancreatic and biliary tumors. Furthermore, the independent validation cohort of this study included 94 pancreatic and biliary tumor patients and 40 healthy subjects, and successfully verified the efficacy of the diagnostic model based on the distribution characteristics of free DNA fragment length in the blood. This method uses a more accurate analysis of the length of cfDNA in the blood to find clues for early tumor screening, providing more solid and reliable data support for clinical precision applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is the fragment length distribution map at the whole genome level of cfDNA in Example 1 of the present invention.

[0032] Figure 2 This is a differential distribution diagram of cfDNA fragments between cancer patients and healthy individuals in Example 1 of the present invention.

[0033] Figure 3 This is the ROC curve diagram of the training set in Example 2 of the present invention.

[0034] Figure 4 This is the ROC curve diagram of the test set in Example 2 of the present invention.

[0035] Figure 5 This is the ROC curve of the independent validation cohort subjects in Example 2 of the present invention. DETAILED DESCRIPTION

[0036] The technical solution of the present invention is described in detail below with reference to the examples. Unless otherwise specified, the reagents and biological materials used below are all commercial products.

[0037] Example 1

[0038] (1) Study cohort and clinical information

[0039] The study included 154 patients diagnosed with pancreatic and biliary tumors (pancreatic cancer, gallbladder cancer, and bile duct cancer) based on tumor markers, imaging tests (such as ultrasound and abdominal CT scans), and pathology, as well as 71 healthy controls. Blood samples were collected from the patients and healthy controls before surgery. Each enrolled patient received an accurate diagnosis based on postoperative pathology.

[0040] (2) Blood collection, separation and storage

[0041] Preoperative whole blood samples from cancer patients and healthy controls were collected in 10 ml cell-free nucleic acid storage tubes (REF43803, BD, USA) and shipped at room temperature. The received whole blood samples were separated into plasma using a two-step centrifugation method. First, the plasma and cellular components were separated by centrifugation at 1600 g for 10 minutes at 4°C. The supernatant was carefully aspirated, taking care not to aspirate the white blood cell layer. The hemolysis grade of the plasma was recorded. Samples with a hemolysis grade ≥5 were not included in the subsequent study. Secondly, the plasma was centrifuged again at 16,000 g for 15 minutes at 4°C to remove any residual cells or cell debris. The supernatant was transferred to centrifuge tubes and aliquoted into 1 ml tubes. The separated plasma samples were stored in a refrigerator at -80°C.

[0042] (3) cfDNA extraction

[0043] Plasma samples were removed from a -80°C freezer and placed in a water bath. Incubated statically at 37°C for approximately 5 minutes, the plasma was transferred to a low-temperature refrigerated centrifuge and centrifuged at 1600 g for 10 minutes at 4°C. The supernatant was carefully aspirated into a centrifuge tube. Plasma cfDNA was extracted from 1 ml of plasma using the QIAamp Circulating Nucleic Acid Kit (55114, Qiagen, Shanghai, China). Specific procedures were described in the product manual. cfDNA was finally eluted using 30 μl of EB. The total amount of extracted cfDNA was quantified using a Qubit fluorometer and accompanying reagents (Q32854, Thermo Fisher, USA). cfDNA fragment distribution was assessed using an Agilent 2100 Bioanalyzer and the accompanying Agilent High Sensitivity DNA Kit & Reagents (5067-4626, Agilent, USA).

[0044] (4) cfDNA library construction and WGS sequencing

[0045] cfDNA samples that passed quality control were used for cfDNA library construction and WGS sequencing. The library was prepared using the KAPADNA Hyper Prep kit (KK8504, KAPA, USA). The detailed protocol is described in the product manual. The input volume for each cfDNA sample was 10 ng. The base ends were then blunted and A-tailed. The library was then enriched through seven cycles of adapter ligation, adapter purification, and PCR amplification. After purification, the DNA was eluted with 25 μl of elution buffer. The plasma cfDNA library concentration was determined using a Qubit assay, and the fragment distribution of the plasma cfDNA library was determined using a 4150 assay. Libraries that passed quality control were sequenced using the NovoSeq 6000 platform for whole-genome sequencing, using a 2x150bp sequencing strategy and a sequencing volume of ~10 Gb (~3×).

[0046] (5) cfDNA fragment size feature extraction

[0047] Using low-pass Whole Genome Sequencing (LP-WGS) technology, we obtained sequence information for cfDNA in the plasma of patients and healthy controls. The analysis process for sequencing data is as follows:

[0048] 1) Sequencing Data Alignment: After removing adapters from the raw fastq data obtained by LP-WGS sequencing (fastq), the sequencing data were aligned to the human reference genome hg19 (genome download link: ftp: / / ftp-trace.ncbi.nih.gov / 1000genomes / ftp / technical / reference / human_g1k_v37.fasta.gz) using BWA software (version: 0.7.12-r1039). Low-quality sequences and duplicate sequences were removed from the resulting BAM file.

[0049] 2) Statistics of cfDNA fragment length.

[0050] 3) Whole-genome fragment size distribution map. Exclude low-coverage regions of the hg19 reference genome and Duke black box regions; then divide the hg19 autosomes into 504 adjacent, non-overlapping window segments, each 5 Mb in length; within each window region, calculate the ratio of the number of cfDNA fragments greater than 130 bp and less than 177 bp to the number of cfDNA fragments greater than 177 bp and less than 237 bp; finally, calculate the number of cfDNA long and short fragments within each 5 Mb interval, and finally use the ratio to visualize the cfDNA fragment size map for the entire genome, see [1]. Figure 1 , which is the fragment length distribution map of cfDNA at the whole genome level.

[0051] (6) Establishment of machine learning classification model

[0052] 1) Determination of size characteristics. Based on the difference distribution of fragments between cancer patients and healthy controls, the interval with the largest difference in fragment distribution between cancer patients and healthy controls was determined. Figure 2 , is the difference distribution diagram of cfDNA fragments between cancer patients and healthy individuals, Figure 2 The vertical axis represents the difference in cfDNA fragment frequency between cancer patients and healthy individuals. We defined the short fragment range [130, 177] and the long fragment range [177, 237], and then calculated the normalized z-score of the number of short cfDNA fragments and the total number of fragments in each of the 504 windows, which served as feature input for model training.

[0053] 2) Divide the samples into training and test sets. All samples were divided into training and test sets in a ratio of 4:1, and the distribution ratio of healthy controls and various cancer samples in the two sets was kept consistent.

[0054] 3) Process the sample data in the training set. In the training cohort, use 30 repetitions of the 5-fold cross-validation method to obtain the model coefficients.

[0055] 4) Evaluate the effectiveness of the model. Based on the model's predicted value and pathological test results for each sample in the training set, draw the receiver operating characteristic curve (ROC curve) of the training set. Based on the predicted value, a series of thresholds are set to divide the training set into healthy people and cancer patients. The pathological test results are then used as the true value to evaluate the model's predictive effectiveness. The model's predictive effectiveness evaluation methods include the area under the ROC curve (AUC, Area UnderCurve, range 0-1), positive predictive value (PPV, Positive Predictive Value, range 0-1), specificity (range 0-1), accuracy (range 0-1), and sensitivity (range 0-1). The higher the value, the better the effect.

[0056] (7) Verification of the predictive effectiveness of the classification model

[0057] In an independent validation cohort, the model's ability to predict classification was verified based on the classification model and prediction values determined in the training cohort. The process was as follows:

[0058] 1) Confirmation of variables: In an independent validation cohort, the standard z-score of the number of short and total cfDNA fragments in 504 windows across the genome was used as the variable.

[0059] 2) Model performance validation. Based on the molecular marker expression levels and pathological test results of each sample in the test set, a receiver operating characteristic (ROC) curve was plotted for the test set. Based on the predicted values, the independent validation cohort was divided into a healthy population (using the same training and test sets) and a cancer group. The model's predictive performance, including specificity, sensitivity, and accuracy, was evaluated. Higher values indicate better performance.

[0060] Example 2

[0061] (1) Study cohort and clinical information

[0062] This study included two study cohorts totaling 154 patients diagnosed with pancreatic and biliary malignancies by tumor markers, imaging examinations (such as ultrasound, abdominal CT scan, etc.) and pathological examinations, as well as 71 healthy subjects. Blood samples were collected from patients before surgery, and blood samples from healthy controls were also collected. A total of 60 patients (29 with pancreatic cancer, 15 with gallbladder cancer, and 16 with bile duct cancer) and 31 healthy subjects were included in the training and test sets (Table 1). The samples were divided into training and test sets at a ratio of 4:1, and the distribution ratios of healthy controls and samples of various cancer types in the two sets were kept consistent. Table 1 shows the grouping information of healthy controls and patients in the training and test sets. The analysis results showed that there was no significant difference in the gender ratio and the distribution ratio of healthy controls and cancer patients between the training and test sets.

[0063] Table 1: Training set and test set information

[0064]

[0065]

[0066] The independent validation cohort included 94 patients with pancreatic and biliary tumors (37 pancreatic cancer, 17 gallbladder cancer, and 40 bile duct cancer) and 40 healthy controls. Table 2 shows participant information for the training set and independent validation cohort. Analysis results showed no significant differences in the gender ratio, healthy controls, or cancer patient distribution between the training and independent validation cohorts.

[0067] Table 2: Training set and independent validation cohort information

[0068]

[0069] (2) Health and cancer classification scoring model

[0070] Using the training set and pathology test results, we constructed a scoring model for cancer patients and healthy individuals using the LinearSVC algorithm. The model consists of three parts: variables, model formula, and predicted values. The process is as follows:

[0071] ① Model variables and parameters. In the training cohort, the model coefficients were obtained using 30 replicates of 5-fold cross-validation, using 1008 feature variables (Table 3) based on the standardized z-scores of the number of short fragments and total fragments within 504 5-Mb regions.

[0072] Table 3: Example of model input variables

[0073] Serial number Model variables Number of fragments 1 Bin1 short fragment N1 2 Bin1 all fragments N2 …… …… …… 1007 Bin504 short clip N1007 1008 Bin504 all clips N1008

[0074] ② Scoring model. The scoring model formula is as follows:

[0075]

[0076] where x i The calculation formulas of model parameters w and b are as follows:

[0077]

[0078] Where λ is the penalty parameter, n is the number of samples, y i is the true value of the sample, 1 for cancer and -1 for health.

[0079] Using the classification model and the fragment distribution in different regions of each sample at the whole genome level, the category prediction results of each sample can be obtained.

[0080] ③Model effectiveness evaluation.

[0081] In order to build an early screening classification model to distinguish cancer patients from healthy individuals, the training set samples were divided into healthy individuals and cancer patients using the predicted values. The pathology test results were used as the true values, and the ROC curve of the training set was drawn based on the predicted values and pathology results in the cohort. The AUC value of the training set reached 1, see Figure 3 , which is the receiver operating characteristic (ROC) curve for the training set. The trained model achieved PPV (accuracy), specificity, and sensitivity of 100%, 100%, and 100%, respectively (Table 4). These results demonstrate that this risk prediction model has high sensitivity and NPV in the training set, demonstrating superior predictive efficacy for early cancer diagnosis.

[0082] ④Verification of the predictive effectiveness of the discriminant model.

[0083] To verify the effectiveness of the discriminant model, the test set patients were divided into healthy and cancer patient groups (the same as the training set) using the threshold set by the predicted value. The pathological test results were used as the true value, and the effectiveness of the model was verified based on the classification model and predicted value determined in the training cohort. The ROC curve of the test set was drawn, and its AUC value reached 1. Figure 4 , which is the ROC curve for the test set. The model's predictive performance, including accuracy, specificity, and sensitivity, was evaluated, which were 100%, 100%, and 92%, respectively (Table 4). The results demonstrate that, in the test set, this risk prediction model also possesses high specificity, sensitivity, and accuracy, indicating excellent predictive performance.

[0084] Table 4

[0085]

[0086] (3) Independent verification cohort verification

[0087] To further validate the discriminant model's effectiveness, patients in the independent validation cohort were divided into healthy and cancer groups using a threshold set by the predicted value. Using the pathology test results as the true value, the model's effectiveness was verified based on the classification model and predicted values determined in the training and test sets. A receiver operating characteristic (ROC) curve for the independent validation cohort was plotted, and the area under the curve (AUC) for the independent validation cohort reached a high of 0.94. Figure 5 , and the ROC curve for the independent validation cohort. The model's predictive performance, including accuracy, specificity, and sensitivity, was assessed and found to be 90.1%, 77.5%, and 87.2%, respectively (Table 5). These results demonstrate that this risk prediction model also exhibited high specificity, sensitivity, and accuracy in the independent validation cohort, indicating superior predictive performance.

[0088] Table 5: Model effectiveness verification of independent validation cohort

[0089]

[0090]

[0091] The above are only some preferred embodiments of the present invention, and the present invention is not limited to the contents of the embodiments. For those skilled in the art, various changes and modifications can be made within the scope of the technical solution of the present invention, and any changes and modifications made are within the scope of protection of the present invention.

Claims

1. A method for constructing a noninvasive early cancer screening model based on the length distribution characteristics of cfDNA (cfDNA). The method comprises: using relatively low-depth whole-genome sequencing to statistically analyze the differences in the length distribution characteristics of tumor-derived cfDNA and cfDNA from healthy individuals, establishing an early cancer screening model and achieving noninvasive early cancer screening. The method for establishing the early cancer screening model comprises: calculating the normalized z-score of the number of short and total cfDNA fragments within 504 5-Mb regions of the genome as feature input values for model training. The cfDNA fragments include short and long fragments, with the length range of short cfDNA fragments being [130, 177] bp and the length range of long fragments being [177, 237] bp. The LinearSVC algorithm is used to construct scoring models for cancer patients and healthy individuals. A 5-fold cross-validation method with 30 repetitions is used to obtain model coefficients to establish an early cancer screening model. The cancers described are pancreatic cancer, gallbladder cancer, and bile duct cancer. The scoring model formula is as follows: where x i The calculation formulas of model parameters w and b are as follows: Where λ is the penalty parameter, n is the number of samples, y i is the true value of the sample, 1 for cancer and -1 for health.

2. A non-invasive early cancer screening system based on the cfDNA fragment length distribution characteristics, characterized by: The system comprises: cfDNA fragment feature extraction module, used to obtain cfDNA fragment size feature data in samples; A machine learning classification model building module is used to construct scoring models for cancer patients and healthy individuals based on the statistical differences in the fragment size characteristics of cfDNA derived from tumors and healthy individuals using the LinearSVC algorithm. A 30-fold repeated 5-fold cross-validation method is used to establish an early cancer screening model. The scoring model formula is as follows: where x i The calculation formulas of model parameters w and b are as follows: Where λ is the penalty parameter, n is the number of samples, y i is the true value of the sample, 1 for cancer and -1 for health; An independent validation cohort evaluation module is used to verify the predictive efficacy of the established machine learning classification model through an independent validation cohort; The cfDNA fragment feature extraction module includes: Sequencing data alignment unit, used to remove sequencing data sequencing adapters and align the sequencing data to the human reference genome hg19; The cfDNA fragment statistics unit is used to count cfDNA fragment length data information. The hg19 autosome is divided into 504 adjacent, non-overlapping window segments, each with a length of 5 Mb. Within each window region, the ratio of the number of cfDNA fragments greater than 130 bp and less than 177 bp to the number of cfDNA fragments greater than 177 bp and less than 237 bp is counted. Finally, the number of long and short cfDNA fragments within each 5 Mb interval is obtained. The cfDNA fragment feature determination unit is used to determine the interval with the largest difference in fragment distribution between cancer patients and healthy controls based on the difference distribution of fragments between cancer patients and healthy controls; define the short fragment range [130, 177] and the long fragment range [177, 237], and then calculate the standardized z-score of the number of short cfDNA fragments and total fragments in each of the 504 windows as the feature input value for model training; The cancers are pancreatic cancer, gallbladder cancer and bile duct cancer.

3. The non-invasive early cancer screening system based on cfDNA fragment length distribution characteristics according to claim 2, characterized in that: The machine learning classification model building module includes: The sample data classification unit is used to divide the samples into training and test sets in a ratio of 4:1, and to ensure that the distribution ratios of healthy controls and various cancer samples in the two sets are consistent; The model parameter acquisition unit is used to process the sample data in the training set; in the training queue, the model parameters are obtained using 30 repeated 5-fold cross validation methods; The model effectiveness evaluation unit is used to draw the receiver operating characteristic curve of the training cohort based on the model prediction value and pathological test results of each sample in the training cohort.

Citation Information

Patent Citations

  • Non-invasive cancer early screening system based on cfDNA omics characteristics

    CN113160889A

  • Analysis method and system based on ctDNA length

    CN113903401A