A method and device for constructing a cancer early screening model based on cfDNA difference statistics
By analyzing the cfDNA fragment omics characteristics of patients with biliary and pancreatic malignancies and healthy individuals through low-depth whole-genome sequencing and combining it with a machine learning model, the problem of insufficient sensitivity and specificity in the early diagnosis of biliary and pancreatic malignancies in existing technologies was solved, and minimally invasive, early and highly sensitive diagnosis and dynamic monitoring of biliary and pancreatic malignancies was achieved.
Patent Information
- Application Number
- CN202311693405.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-12-11
AI Technical Summary
Existing technologies make it difficult to effectively detect biliary and pancreatic malignancies through low-depth cfDNA whole-genome sequencing, especially in the early stages, where the sensitivity and specificity are insufficient, and imaging and serum marker detection have limitations.
Low-depth whole-genome sequencing was used to analyze the cfDNA fragment omics characteristics of patients with biliary and pancreatic malignancies and healthy individuals, including cfDNA fragment size, 5' end Motif characteristics, and depth patterns of sequencing data near the transcription start site (TSS). An early screening model was constructed by combining machine learning models, and a highly specific and sensitive diagnostic model was established using the cfDNA fragment omics characteristics and serum CA19-9 expression levels.
It achieves minimally invasive, early and highly sensitive diagnosis of malignant tumors of the gallbladder and pancreas, can dynamically monitor tumor changes, avoids information lag or loss caused by tumor heterogeneity, and improves the accuracy and reliability of diagnosis.
Smart Images

Figure CN119741967B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biotechnology and relates to a method and device for constructing a cancer early screening model based on cfDNA difference statistics. Background Art
[0002] Biliary tract cancer is the fifth most common gastrointestinal malignancy, mainly including gallbladder cancer and bile duct cancer. Pancreatic cancer is the most invasive and lethal digestive tract malignancy, and the gallbladder, bile duct and pancreas are anatomically located in close proximity. Early lesions of bile duct and pancreatic malignancies are usually asymptomatic, and the disease only becomes apparent after the tumor invades surrounding tissues or metastasizes to distant organs. Patients are usually in the late stage of the tumor when diagnosed, and late-stage patients often have a poor prognosis due to untimely treatment intervention. The five-year survival rate of pancreatic cancer patients is only 5%, and the five-year survival rate of bile duct cancer patients is less than 5%. Even if some patients can be treated with surgical intervention, the tumor is likely to recur after resection. Therefore, finding reliable early diagnostic biomarkers is more clinically valuable for the diagnosis and treatment of bile duct and pancreatic malignancies.
[0003] Currently, imaging, biopsy, cytology, and serum tumor marker testing are the main diagnostic methods for pancreatic and biliary tumors. Imaging methods primarily include ultrasound, MRI (magnetic resonance imaging), CT (computed tomography), and endoscopic retrograde cholangiopancreatography (ERCP). Ultrasound is the preferred screening method for pancreatic and biliary tumors, while MRI and CT are the primary modalities for diagnosis and staging. MRI better visualizes the anatomy of the liver, bile duct, and pancreas, as well as the extent of tumor involvement, but has poor sensitivity for tumors less than 1 cm in diameter. Brush cytology from biopsy and ERCP is the gold standard for diagnosing pancreatic and biliary tumors, but brush cytology has low sensitivity, while biopsy is prone to intraperitoneal infection and tumor metastasis. CA19-9 is an FDA-approved unique prognostic biomarker for resectable pancreatic and biliary tumors, but its sensitivity and specificity are limited and it is susceptible to interference from factors such as chronic pancreatitis, cholestasis, or cholangitis.
[0004] Liquid biopsies, due to their minimally or non-invasive sampling methods, cost-effectiveness, and convenience, are widely used for early diagnosis of cancer, postoperative monitoring, and companion diagnostics. When an organ or tissue experiences pathological conditions, resulting in increased cell death, more cell-free DNA (cfDNA) circulates in the blood. Therefore, plasma cfDNA carries genetic and epigenetic information about the pathological cells and their tissue of origin. However, not all patients can detect genetic variants due to low concentrations of tumor-associated cfDNA in plasma, particularly in the early stages of cancer. Plasma cfDNA fragmentomic analysis is an emerging field encompassing fragment size, end points, nucleosome footprints, and transcription start sites in plasma cfDNA. Studies have shown that genome-wide cfDNA fragmentomic profiles differ between cancer patients and non-cancer patients, and that cfDNA fragment size and end point sequence characteristics can be used for early diagnosis of cancer. A study by Cristiano et al. demonstrated that genome-wide cfDNA fragment size profiling has an 88% sensitivity for diagnosing cholangiocarcinoma (PMID: 31142840). However, studies using cfDNA fragmentomics from low-coverage whole-genome sequencing (LP-WGS) data to establish a detection model for anatomically proximal biliary and pancreatic malignancies are still lacking.
[0005] Therefore, there is an urgent need for research data from low-depth cfDNA whole-genome sequencing to assist in the clinical transformation of biomarkers in early diagnosis, postoperative detection, and companion diagnosis of biliary and pancreatic malignancies. Summary of the Invention
[0006] In response to the deficiencies in the existing technology and actual needs, the present invention provides a method and device for constructing a cancer early screening model based on cfDNA difference statistics. By performing whole-genome sequencing on the blood cfDNA of patients with biliary and pancreatic malignancies (biliary tract cancer and pancreatic cancer), and using healthy people as a control group, the omics characteristics of plasma cfDNA fragments are analyzed to discover and verify liquid biopsy biomarkers that can be used for the diagnosis of biliary and pancreatic malignancies, thereby facilitating the clinical transformation of biomarkers for the early diagnosis of biliary and pancreatic malignancies, prolonging the survival time of patients with biliary and pancreatic malignancies, and improving their quality of life. With the advantages of minimally invasive, high sensitivity for early diagnosis, and dynamic monitoring, the method overcomes the limitations of cfDNA gene mutation detection in the early stages of tumors and avoids information lag or loss caused by tumor heterogeneity.
[0007] In order to achieve the purpose of the invention, the present invention adopts the following technical solutions:
[0008] In the first aspect, the present invention provides a method for constructing a cancer early screening model based on cfDNA difference statistics, the construction method comprising: performing difference statistics on tumor-derived cfDNA and cfDNA from healthy individuals through a relatively low-depth whole-genome sequencing method, establishing an early cancer screening model, and realizing non-invasive early screening of cancer; the difference statistics include cfDNA fragment size feature difference statistics, cfDNA three consecutive groups of 3bp base combination category distribution frequency change trends upstream and downstream of the break position, and cfDNA sequencing data coverage depth pattern difference statistics near TSS.
[0009] The method of the present invention overcomes the limitations of cfDNA gene mutation detection in the early stages of tumors with its advantages of minimal invasiveness, high sensitivity for early diagnosis and dynamic monitoring, and avoids information lag or loss caused by tumor heterogeneity.
[0010] Preferably, the cfDNA fragments include short cfDNA fragments and long cfDNA fragments, the length of the short cfDNA fragments is [130, 177] bp, the length of the long cfDNA fragments is [177, 237] bp, and the ratio of the short cfDNA fragments to the long cfDNA fragments is the characteristic difference value of the cfDNA fragment size;
[0011] The statistical method for calculating the distribution frequency of three consecutive 3bp base combination categories upstream and downstream of the break site of the cfDNA includes: counting 9 bases starting from the sixth base at the 5' end of the cfDNA and grouping three consecutive bases into three groups of 3bp motifs; and counting the occurrence frequency of the three groups of 3bp motifs for each sample;
[0012] The parameter calculation method for the sequencing coverage depth of cfDNA near the TSS is as follows: the region 500 bp upstream and downstream of the TSS [-250 bp, 250 bp] is defined as the central region, and the upstream [-2000 bp, -1000 bp] and downstream [1000 bp, 2000 bp] of the TSS are defined as the surrounding regions; the NF value of the gene is the average coverage of the central region divided by the average coverage of the surrounding regions.
[0013] Preferably, the early cancer screening method further comprises: performing statistical analysis on the differences between tumor-derived cfDNA and cfDNA from healthy individuals, detecting the serum CA19-9 expression level, and then constructing a model.
[0014] Preferably, the cancer comprises pancreatic and biliary malignancies.
[0015] Preferably, the pancreatic and biliary malignancies include any one of pancreatic cancer, gallbladder cancer or bile duct cancer.
[0016] In a second aspect, the present invention provides a non-invasive early cancer screening device, comprising: a feature extraction module, a machine learning classification model building module, and an independent verification cohort evaluation module;
[0017] The feature extraction module is used to perform cfDNA fragment feature extraction, terminal Motif feature extraction and TSS data feature extraction;
[0018] The independent validation cohort evaluation module is used to perform validation of the predictive efficacy of the established machine learning classification model through an independent validation cohort.
[0019] Preferably, the cfDNA fragment feature extraction includes: a sequencing data comparison unit, a cfDNA fragment statistics unit and a cfDNA fragment feature determination unit;
[0020] The sequencing data alignment unit is used to perform the following steps: after removing the sequencing data sequencing adapters, aligning the sequencing data to the human reference genome hg19;
[0021] The cfDNA fragment statistics unit is used to perform the following steps: counting cfDNA fragment length data information; dividing the hg19 autosome into 504 adjacent, non-intersecting window segments, each with a length of 5 Mb; counting the ratio of the number of cfDNA segments with a length greater than 130 bp and less than 177 bp to the number of cfDNA segments with a length greater than 177 bp and less than 237 bp in each window region; and finally obtaining the number of long and short cfDNA segments in each 5 Mb interval.
[0022] The cfDNA fragment feature determination unit is configured to perform the following steps: determining the interval with the largest difference in fragment distribution between cancer patients and healthy controls based on the difference distribution of fragments between cancer patients and healthy controls; defining a short fragment range of [130, 177] and a long fragment range of [177, 237], and then calculating the normalized z-score of the number of short fragment cfDNA and the total number of fragments in each of the 504 windows as feature input values for model training;
[0023] The terminal motif feature extraction includes: a sequencing data alignment unit, a reads filtering unit, a combination motif calculation unit and a combination motif statistics unit;
[0024] The sequencing data alignment unit is used to perform the following steps: after removing the sequencing data sequencing adapters, aligning the sequencing data to the human reference genome hg19;
[0025] The read filtering unit is used to perform the following operations: filtering the sequencing data, with the following filtering criteria: only considering reads aligned to autosomes 1-22, with a quality score greater than 20; insert length between 150-600, with both ends properly paired; and the reference region of the read does not contain degenerate bases;
[0026] The combination motif calculation unit is used to perform the following steps: counting 9 bases from the 6th base at the 5' end and dividing them into three groups as its combination motif; for the 3' end read, counting 9 bases from the 6th base at the 3' end and dividing them into three groups in the order of 3' to 5', and then performing reverse complementation as its combination motif;
[0027] The combined motif statistics unit is used to count the frequencies of three groups of 3bp motifs for each sample. Each group of 3bp motifs has 64 different combinations, with a total of 192 frequencies as its training variables;
[0028] The TSS data feature extraction includes a sequencing data alignment unit, a read filtering unit, a gene screening and TSS determination unit, and an NF value calculation unit;
[0029] The sequencing data alignment unit is used to perform the following steps: after removing the sequencing data sequencing adapters, aligning the sequencing data to the human reference genome hg19;
[0030] The read filtering unit is used to perform the following operations: filtering the sequencing data, wherein the filtering criteria are: only reads aligned to autosomes 1-22 are considered, the quality score is greater than 20, the insert length is between 150-600, and the double ends must be properly paired; the reference region of the read does not contain degenerate bases;
[0031] The gene screening and TSS determination unit is used to perform the following steps: determining the TSS of gene transcripts based on the transcript annotation of the UCSC hg19 genome; for genes with multiple TSSs, only those with a difference of less than 50 bp are retained, and their average value is used as the TSS of the gene, and only genes on autosomes are considered;
[0032] The NF value calculation unit is used to perform the following steps: obtaining parameters for the sequencing coverage depth of cfDNA near the TSS, and the calculation method is as follows: defining the area 500bp upstream and downstream of the TSS [-250bp, 250bp] as the central area; defining 2000bp upstream [-2000bp, -1000bp] and downstream [1000bp, 2000bp] of the TSS as the surrounding area; the NF value of a gene is the average coverage of the central area divided by the average coverage of the surrounding area.
[0033] Preferably, the machine learning classification model building module includes: a model parameter acquisition unit and a model effectiveness evaluation unit;
[0034] The model parameter acquisition unit is used to perform the following steps: using a machine algorithm in a training set and a 5-fold or 10-fold cross validation method to obtain model parameters;
[0035] The model effectiveness evaluation unit is used to perform the following steps: drawing a receiver operating characteristic curve of the training set according to the model prediction value and pathological detection result of each sample in the training set.
[0036] In a third aspect, the present invention provides a cancer early screening model, wherein the input of the cancer early screening model is the score value of the risk scoring model and Log2(1+CA19-9); the output is the model score.
[0037] Preferably, the score value of the risk scoring model is the score value of the test set.
[0038] Preferably, when the model score is lower than 0, it is judged as healthy, and when it is higher than or equal to 0, it is judged as cancer.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] (1) The present invention has developed a blood cfDNA biomarker liquid biopsy technology that can be used for the early diagnosis of biliary and pancreatic malignancies. The technology uses cfDNA fragment size characteristics (independently defined size fragment range), 5' end Motif characteristics (independently developed 3x3bp combined Motif characteristics) and deep pattern characteristics of sequencing data near the transcription start site (TSS) to construct a risk prediction model with high specificity and high sensitivity. The technology also combines blood cfDNA fragment biomarkers and serum CA19-9 expression levels to establish a more effective diagnostic model.
[0041] (2) The present invention overcomes the limitations of cfDNA gene mutation detection in the early stages of tumors with its advantages of minimal invasiveness, high sensitivity for early diagnosis, and dynamic monitoring, and avoids information lag or loss caused by tumor heterogeneity. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1A This is the result of blood cfDNA fragment size characteristic analysis;
[0043] Figure 1B This is the result of the analysis of the 5' end Motif characteristics of blood cfDNA;
[0044] Figure 1C This is the result of deep pattern feature analysis of sequencing data near the transcription start site (TSS) of blood cfDNA;
[0045] Figure 2 This is the ROC curve diagram for the training cohort;
[0046] Figure 3 ROC curve diagram for the validation cohort;
[0047] Figure 4 This is the ROC curve of the blood cfDNA fragment omics combined with CA19-9 diagnostic model. DETAILED DESCRIPTION
[0048] To further illustrate the technical means and effects of the present invention, the present invention is further described below with reference to the embodiments and drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention.
[0049] If no specific techniques or conditions are specified in the examples, the experiments were carried out according to the techniques or conditions described in the literature in the field or according to the product instructions. If no manufacturer is specified for the reagents or instruments used, they are all conventional products that can be purchased through regular channels.
[0050] Example 1
[0051] This embodiment provides a method and model for detecting malignant tumors of the gallbladder and pancreas.
[0052] (1) Study cohort
[0053] The study included 147 patients with suspected pancreatic and biliary malignancies (55 with cholangiocarcinoma, 30 with gallbladder cancer, and 62 with pancreatic cancer) based on tumor marker and imaging findings, as well as 71 healthy controls. Blood samples were collected from all patients before surgery. Each enrolled patient received an accurate pathological diagnosis after surgery.
[0054] (2) Processing of blood samples from patients with pancreatic and biliary malignancies and healthy individuals
[0055] Blood samples from patients with biliary and pancreatic cancer and healthy individuals were collected in 10 mL cell-free DNA tubes (BEAVERCell Free DNATubes, 43803, China) and shipped at room temperature. Upon receipt, the blood was centrifuged at 1,600 g for 10 minutes at 4°C to remove cells and debris. The supernatant was transferred to a fresh 1.5 mL centrifuge tube and centrifuged again at 16,000 g for 15 minutes at 4°C. Finally, the supernatant was transferred and aliquoted into 1.5 mL centrifuge tubes at 1 mL per tube for storage at −80°C.
[0056] (3) Plasma cfDNA extraction from patients with pancreatic and biliary malignancies and healthy subjects
[0057] After removing plasma samples from -80°C, they were thawed in a 37°C water bath and centrifuged at 16,000 g for 10 minutes at 4°C. 0.99 mL of the supernatant was transferred to a fresh 1.5 mL centrifuge tube. CFDNA was subsequently extracted using the QIAamp Circulating Nucleic Acid Kit (55114, Qiagen, Shanghai, China). After sample lysis, protein digestion, column chromatography, washing, and elution, plasma cfDNA was collected and subsequently stored in a -20°C or -80°C freezer for long-term storage.
[0058] (4) Plasma cfDNA concentration and fragment size detection
[0059] Plasma cfDNA concentration was measured using a Qubit 3.0 fluorometer and accompanying reagents (Q32854, Thermo Fisher, USA). The fragment distribution of extracted plasma cfDNA was analyzed using a 2100 analyzer with accompanying chips and reagents (5067-4626, Agilent, USA). The specific detection procedures are described in the kit instructions.
[0060] (5) Plasma cfDNA micro-library construction and LP-WGS sequencing
[0061] Plasma cfDNA libraries were constructed using the KAPA DNA HyperPrep kit (KK8504, KAPA, USA). Each cfDNA sample was loaded with 3-10 ng of DNA. The final library was obtained through end-repair and A-tailing, adapter ligation and purification, and library amplification and purification. For detailed protocols, refer to the product manual. After passing quality control, the constructed library was sequenced using the Illumina NovaSeq 6000, generating 10 GB of sequencing data per sample.
[0062] (6) Plasma cfDNA fragment genomic feature extraction
[0063] 1) Plasma cfDNA fragment size feature extraction
[0064] 1. Sequencing Data Alignment. After removing adapters (fastq), the raw fastq data obtained by LP-WGS sequencing of plasma cfDNA from patients with pancreatic and biliary malignancies and healthy controls were aligned to the human reference genome hg19 (genome download link: ftp: / / ftp-trace.ncbi.nih.gov / 1000genomes / ftp / technical / reference / human_g1k_v37.fas ta.gz) using BWA software (version 0.7.12-r1039). Low-quality sequences and duplicate sequences were removed from the resulting BAM files.
[0065] 2. cfDNA fragment length statistics;
[0066] 3. Whole-genome fragment size distribution map. Regions with low coverage of the hg19 reference genome and the Duke black box regions were excluded. Next, the hg19 autosomes were divided into 504 adjacent, non-overlapping window segments, each 5 Mb in length. Within each window region, the ratio of the number of cfDNA fragments greater than 130 bp and less than 177 bp to the number of cfDNA fragments greater than 177 bp and less than 237 bp was calculated. Finally, the number of long and short cfDNA fragments within each 5 Mb interval was obtained, and the ratio was used to visualize the cfDNA fragment size map for the entire genome.
[0067] 2) 5' end motif feature extraction
[0068] 1. Sequencing Data Alignment. After removing sequencing adapters, the plasma cfDNA sequencing data from patients with pancreatic and biliary malignancies and healthy controls were aligned to the human reference genome hg19 using BWA software (version 0.7.17-r1188) (genome download link: ftp: / / ftp-trace.ncbi.nih.gov / 1000genomes / ftp / technical / reference / human_g1k_v37.fas ta.gz).
[0069] 2. Read filtering: Only reads that map to autosomes 1-22 are considered; the quality score is greater than 20; the insertion size is between 150 and 600 bp; the paired ends must be properly paired; and the reference region of the read does not contain degenerate bases.
[0070] 3. Combination motif calculation. For reads at the 5' end, count 9 bases starting from the 5' end and group three consecutive bases into three groups to form the combination motif. Since sequencing data is added with UMIs, each read is preceded by a UMI sequence. The actual calculation is to count 9 bases starting from the 6th base (inclusive) at the 5' end and group them into three groups as the combination motif. 4. Combination motif statistics. Each group of 3bp motifs has 64 different combinations. For each sample, the frequencies of its three groups of 3bp motifs are counted, for a total of 192 frequencies as its training variables.
[0071] 3) Deep pattern feature extraction of sequencing data near the transcription start site (TSS)
[0072] 1. Sequencing Data Alignment. After removing sequencing adapters, the plasma cfDNA sequencing data from patients with pancreatic and biliary malignancies and healthy controls were aligned to the human reference genome hg19 (genome download link: ftp: / / ftp-trace.ncbi.nih.gov / 1000genomes / ftp / technical / reference / human_g1k_v37.fas ta.gz) using BWA software (version 0.7.17-r1188).
[0073] 2. Read filtering: Only reads that map to autosomes 1-22 are considered; the quality score is greater than 20; the insertion size is between 150 and 600; the paired ends must be properly paired; and the reference region of the read does not contain degenerate bases.
[0074] 3. Gene screening and TSS determination. Gene transcription start sites (TSSs) were determined based on transcript annotation from the UCSC hg19 genome. For genes with multiple TSSs, only those with a difference of less than 50 bp were retained, and the average of these TSSs was used as the gene's TSS. Also, only genes on autosomes were considered.
[0075] 4. Nucleosome footprint (NF) calculation. The central region was defined as the region 500 bp upstream and downstream of the TSS (-250 bp, 250 bp). The surrounding region was defined as the region 2000 bp upstream (-2000 bp, -1000 bp) and downstream (1000 bp, 2000 bp). The NF of a gene was calculated as the average coverage of the central region divided by the average coverage of the surrounding region.
[0076] (7) Establishment and evaluation of a risk scoring model for pancreatic and biliary tumors
[0077] 1) Model Building: Utilizing the training cohort and pathology test results, a stacking algorithm was used to construct an integrated scoring model for distinguishing cancer patients from healthy controls, using three types of features as input. The model consists of three components: variables, primary model formulas, secondary model formulas, and predicted values. The process is as follows:
[0078] 1. Primary model variables and parameters of three types of features:
[0079] a.Fragment prediction value:
[0080] a) Model variables and parameters: A LinearSVC model was used, with 504 short fragments within a 5-Mb region and 1008 normalized z-scores of the total number of fragments as feature variables (Table 1). Model parameters were obtained using 10-fold cross-validation.
[0081] Table 1
[0082]
[0083]
[0084] b) Scoring model. The scoring model formula is as follows:
[0085] Score=w T x i -b
[0086] where x i As input variables, the calculation formulas of model coefficients w and b are as follows:
[0087]
[0088] Where λ is the penalty parameter, n is the number of samples, y i is the true value of the sample, 1 for cancer and -1 for health.
[0089] Using the classification model and the fragment distribution of different regions at the whole genome level of each sample, the category prediction result (Score) of each sample can be obtained.
[0090] b. Motif prediction value:
[0091] a) Model variables and parameters: A RandomForest model was used, with the frequencies of 192 3-bp motifs as feature variables (Table 2), and 10-fold cross-validation was used to obtain model parameters.
[0092] Table 2
[0093] Serial number Model variables 3bp motif frequency 1 AAA Freq1 2 AAT Freq2 3 …… …… 191 CCG Freq191 192 CCC Freq192
[0094] Scoring model. The scoring model formula is as follows:
[0095]
[0096] where x i is the input variable, P k (x i ) is the probability that the k-th decision tree predicts that sample i is cancer. Each decision tree is constructed based on Giniimpurity. The value I G The calculation formula is:
[0097] I G =2p c (1-p c )
[0098] where p c is the proportion of cancer samples in a node after each branch of the decision tree.
[0099] Using the classification model and the combined motif frequency of each sample, the category prediction result (Score) of each sample can be obtained.
[0100] c.NF predicted value:
[0101] a) Model variables and parameters: The LinearSVC model was used, with the NF values of 21,334 genes as feature variables (Table 3), and the model parameters were obtained using a 10-fold cross-validation method.
[0102] Table 3
[0103] Serial number Model variables NF value 1 A1BG NF1 2 A1CF NF2 … …… …… 21333 ZYX NF21333 21334 ZZEF1 NF21334
[0104] Scoring model. The scoring model formula and coefficient calculation formula are the same as those of primary model a. Using the classification model and the NF value of each sample, the category prediction result of each sample can be obtained.
[0105] 2. Secondary model variables and parameters. The input variables of the secondary model are the 5-fold cross-prediction values of the three primary models (Table 4). Then, a 10-fold cross-validation method is used to obtain the parameters and thresholds of the secondary model.
[0106] Table 4
[0107] Serial number Model variables Predicted value 1 Fragment prediction value Prediction value 1 2 Motif prediction value Prediction value 2 3 NF predicted value Prediction value 3
[0108] 3. Secondary model formula. The secondary model uses LogisticRegression, and the model formula is as follows:
[0109] Score=w T x i +b
[0110] where x i As input variables, the calculation formulas of model coefficients w and b are as follows:
[0111]
[0112]
[0113] Where yi is the true value of the sample, 1 for cancer and 0 for health. The threshold is 0.104, that is, the model scores below 0.104 are healthy, and above or equal to 0.104 are cancer.
[0114] 2) Model Performance Evaluation: Using a reference value of 0.104 as the standard, the training cohort samples were divided into healthy individuals (predicted to be healthy) and patients with pancreatic and biliary malignancies (predicted to have pancreatic and biliary malignancies). The model's predictive performance was evaluated using the pathological test results as the true value. Model predictive performance was evaluated using the area under the receiver operating characteristic (ROC) curve (AUC, range 0–1), positive predictive value (PPV, range 0–1), specificity (range 0–1), and sensitivity (range 0–1). Higher values indicate better performance.
[0115] (8) Validation of the risk scoring model for pancreatic and biliary tumors
[0116] In the validation cohort, the effectiveness of the model prediction is verified based on the risk score model and reference values determined in the training cohort. The process is as follows:
[0117] 1) Risk value. In the validation cohort, a risk value is calculated for each sample.
[0118] 2) Model performance validation. Using a reference value of 0.104, the validation cohort was divided into healthy individuals (the same as the training cohort) and patients with pancreatic and biliary tumors. Using the pathological test results as the true value, receiver operating characteristic (ROC) curves were plotted to evaluate the model's predictive performance, including area under the curve (AUC), predictive value perfusion (PPV), specificity, and sensitivity. Higher values indicate superior performance.
[0119] (9) Optimization of the risk scoring model for pancreatic and biliary tumors
[0120] Combining the fragmentomics diagnostic model with serum CA19-9 levels can further improve the accuracy of the model. Combined classification model based on CA19-9 levels:
[0121] 1) Model Building: Using the training cohort, combined with the risk score model scores and CA19-9 levels, a LinearSVC model was used to construct a joint classification model for cancer patients and healthy subjects. The model consists of three parts: variables, model formulas, and predicted values. The process is as follows:
[0122] 1. Model Variables and Parameters. The model input variables consist of the risk score model score and the log2(1 + CA19-9) level. The scores in the training set are the 10-fold cross-predicted values of the risk score model, and the scores in the test set are the predicted values of the final risk score model. Model parameters and thresholds are all default values. The variables and parameters of the joint classification model are shown in Table 5.
[0123] Table 5
[0124]
[0125] 2. Scoring model. The scoring model formula is similar to the formula for primary model a in the risk scoring model. The model threshold is 0, meaning that in the joint classification model, scores below 0 are considered healthy, and scores above or equal to 0 are considered cancerous.
[0126] 2) Model Performance Evaluation: Using a reference value of 0 as the standard, the training cohort samples were divided into healthy individuals (i.e., predicted to be healthy) and patients with pancreatic and biliary malignancies (i.e., predicted to have pancreatic and biliary malignancies). The pathological test results were used as the true value to evaluate the model's predictive performance. Model predictive performance was evaluated using the area under the receiver operating characteristic (ROC) curve (AUC, range 0–1), positive predictive value (PPV, range 0–1), specificity (range 0–1), and sensitivity (range 0–1). Higher values indicate better performance.
[0127] 3) Model Performance Verification: In the validation cohort, a risk value was calculated for each sample. Using a reference value of 0, the validation cohort patients were divided into healthy individuals (same as the training cohort) and patients with pancreatic and biliary tumors. Using the pathology test results as the true value, receiver operating characteristic (ROC) curves were plotted to evaluate the model's predictive performance, including area under the curve (AUC), predictive value perfusion (PPV), specificity, and sensitivity. Higher values indicate better performance.
[0128] (10) Application of risk scoring model for pancreatic and biliary tumors
[0129] 1) Peripheral blood was collected from patients suspected of having pancreatic and biliary malignancies based on imaging tests to obtain plasma cfDNA, and LP-WGS sequencing was used to determine the expression of biomarkers in the plasma cfDNA fragments.
[0130] 2) Using the risk scoring model, the cfDNA fragment omics biomarkers were used as input to obtain the model score. The score and serum CA19-9 expression levels were used as input to obtain the risk value for each patient calculated by the joint classification model;
[0131] 3) Compare the risk value with the model reference value to give a predicted result of each patient's risk of developing biliary and pancreatic malignancies.
[0132] Example 2
[0133] This example verifies the effectiveness of the risk scoring model for pancreatic and biliary malignancies.
[0134] (1) Study cohort and clinical information
[0135] This study enrolled 147 patients with pancreatic and biliary malignancies and 71 healthy controls, divided into two independent cohorts. The training cohort included 31 healthy subjects, 28 patients with pancreatic ductal adenocarcinoma, 14 patients with gallbladder cancer, and 16 patients with cholangiocarcinoma. The validation cohort included 40 healthy subjects, 34 patients with pancreatic ductal adenocarcinoma, 16 patients with gallbladder cancer, and 39 patients with cholangiocarcinoma. Clinical characteristics, including sex, TNM tumor stage, and serum tumor marker concentrations, are shown in Table 6.
[0136] Table 6
[0137]
[0138]
[0139] (2) Omic characteristics of blood cfDNA fragments
[0140] The blood cfDNA fragment size characteristics, 5' end Motif characteristics and deep pattern characteristics of sequencing data near the transcription start site (TSS) were extracted. The results showed that there were differences in blood cfDNA fragment size, 5' end Motif and deep pattern of sequencing data near the transcription start site (TSS) between healthy people and patients with pancreatic and biliary malignancies, which can be used to distinguish healthy people from patients with pancreatic and biliary malignancies ( Figure 1A-Figure 1C ).
[0141] (3) Establishment and evaluation of a risk scoring model for pancreatic and biliary tumors
[0142] In order to construct a risk classification model for pancreatic and biliary malignancies, the training cohort was used, combined with the pathological test results, and the Stacking algorithm was used to construct an integrated scoring model that uses three types of features as input to distinguish between cancer patients and healthy people. The model consists of three parts: variables, primary model formulas, secondary model formulas, and predicted values. In the training cohort, based on the cfDNAFragment prediction value, Motif prediction value, and NF prediction value of the plasma of each patient and healthy person in the primary model, combined with 10-fold cross-validation, the reference value was determined to be 0.104. When the risk value is less than the reference value, the sample is predicted to be healthy; otherwise, it is predicted to be pancreatic and biliary malignancies. Based on the plasma cfDNA fragment omics biomarkers and pathological test results of each patient in the training cohort, the ROC curve of the training cohort was drawn ( Figure 2 The model's predictive efficiency was 0.978%, 96.8%, 91.4%, and 98.1%, respectively (Table 7). These results demonstrate that this risk classification model has high AUC, specificity, sensitivity, and PPV in the training cohort, demonstrating superior predictive efficiency.
[0143] (4) Validation of the risk scoring model for pancreatic and biliary tumors
[0144] In order to verify the effectiveness of the risk scoring model in predicting pancreatic and biliary malignancies, another independent cohort was selected as the validation cohort. The model was validated based on the risk classification model and reference value determined in the training cohort. Based on the reference value of 0.104, the patients in the validation cohort were divided into healthy people (same as the training cohort) and patients with pancreatic and biliary malignancies. Based on the plasma cfDNA fragment biomarkers and pathological test results of each patient in the validation cohort, the ROC curve of the validation cohort was drawn ( Figure 3 The model's predictive performance was assessed using the AUC, specificity, sensitivity, and PPV, which were 0.941, 80.0%, 84.3%, and 90.4%, respectively (Table 7). The results demonstrated that this risk prediction model also had high AUC, specificity, sensitivity, and PPV in the validation cohort, indicating superior predictive performance.
[0145] Table 7
[0146]
[0147] (5) Optimization of the risk scoring model for pancreatic and biliary tumors
[0148] Combining the fragmentomics diagnostic model with the serum CA19-9 level can further improve the accuracy of the model. Combined classification model was established by combining the CA19-9 level ( Figure 4 ), when the specificity of the training cohort remained unchanged, the sensitivity of the training cohort could be increased to 93.1%, and the sensitivity and specificity of the validation cohort could be increased to 98.9% and 85.0%, respectively. The positive predictive values of the training cohort and validation cohort were 98.2% and 93.6%, respectively (Table 8). The results showed that the plasma cfDNA fragment omics combined with the serum CA19-9 expression level model can further improve the sensitivity of prediction.
[0149] Table 8
[0150]
[0151] In summary, the method of the present invention overcomes the limitations of cfDNA gene mutation detection in the early stages of tumors with its advantages of minimal invasiveness, high sensitivity for early diagnosis, and dynamic monitoring, and avoids information lag or loss caused by tumor heterogeneity.
[0152] The applicant states that the present invention is intended to illustrate the detailed methods of the present invention through the above-described embodiments, but the present invention is not limited to the above-described detailed methods, that is, it does not mean that the present invention must rely on the above-described detailed methods in order to be implemented. Those skilled in the art should understand that any improvements to the present invention, equivalent substitutions for various raw materials in the products of the present invention, addition of auxiliary ingredients, and selection of specific methods, etc., are all within the scope of protection and disclosure of the present invention.
Claims
1. A method for constructing a cancer early screening model based on cfDNA difference statistics, characterized in that: The construction method includes: performing statistical analysis of differences between tumor-derived cfDNA and cfDNA from healthy individuals through relatively low-depth whole-genome sequencing, establishing an early cancer screening model, and achieving non-invasive early cancer screening; The difference statistics include the difference statistics of cfDNA fragment size characteristics, the difference statistics of the distribution frequency change trend of three consecutive groups of 3bp base combination categories upstream and downstream of the cfDNA break position, and the difference statistics of the sequencing data coverage depth pattern of cfDNA near the TSS; The cfDNA fragments include short cfDNA fragments and long cfDNA fragments, the length of the short cfDNA fragments is [130, 177] bp, the length of the long cfDNA fragments is [177, 237] bp, and the ratio of the short cfDNA fragments to the long cfDNA fragments is the characteristic difference value of the cfDNA fragment size; The statistical method for the distribution frequency of three consecutive groups of 3 bp base combination categories upstream and downstream of the break position of cfDNA includes: counting 9 bases from the 5' end of cfDNA and grouping three consecutive bases into three groups of 3 bp motifs; and counting the occurrence frequency of the three groups of 3 bp motifs for each sample; The parameter calculation method for the sequencing coverage depth of cfDNA near the TSS is as follows: the region 500 bp upstream and downstream of the TSS [-250 bp, 250 bp] is defined as the central region, and the upstream [-2000 bp, -1000 bp] and downstream [1000 bp, 2000 bp] of the TSS are defined as the surrounding region; the NF value of the gene is the average coverage of the central region divided by the average coverage of the surrounding region; The cancer includes pancreatic and biliary malignancies, which include any one of pancreatic cancer, gallbladder cancer, and bile duct cancer.
2. The construction method according to claim 1, characterized in that The construction method also includes: performing statistical analysis on the differences between tumor-derived cfDNA and cfDNA from healthy individuals, detecting the serum CA19-9 expression level, and then constructing the model.
3. A non-invasive early screening device for cancer, characterized in that: The non-invasive early cancer screening device includes: a feature extraction module, a machine learning classification model building module and an independent verification cohort evaluation module; The feature extraction module is used to perform cfDNA fragment feature extraction, terminal Motif feature extraction and TSS data feature extraction; The independent validation cohort evaluation module is used to perform validation of the predictive efficacy of the established machine learning classification model through an independent validation cohort; The cfDNA fragment feature extraction includes: a sequencing data comparison unit, a cfDNA fragment statistics unit and a cfDNA fragment feature determination unit; The sequencing data alignment unit is used to perform the following steps: after removing the sequencing data sequencing adapters, aligning the sequencing data to the human reference genome hg19; The cfDNA fragment statistics unit is configured to perform the following steps: counting cfDNA fragment length data information; dividing the hg19 autosome into 504 adjacent, non-intersecting window segments, each with a length of 5 Mb; counting the ratio of the number of cfDNA segments with a length greater than 130 bp and less than 177 bp to the number of cfDNA segments with a length greater than 177 bp and less than 237 bp in each window region; and finally obtaining the number of long and short cfDNA segments within each 5 Mb interval. The cfDNA fragment feature determination unit is configured to perform the following steps: determining the interval with the largest difference in fragment distribution between cancer patients and healthy controls based on the difference distribution of fragments between cancer patients and healthy controls; defining a short fragment range of [130, 177] and a long fragment range of [177, 237], and then calculating the normalized z-score of the number of short cfDNA fragments and the total number of fragments in each of the 504 windows as feature input values for model training; The terminal motif feature extraction includes: a sequencing data alignment unit, a reads filtering unit, a combination motif calculation unit and a combination motif statistics unit; The sequencing data alignment unit is used to perform the following steps: after removing the sequencing data sequencing adapters, aligning the sequencing data to the human reference genome hg19; The read filtering unit is used to perform the following operations: filtering the sequencing data, with the following filtering criteria: only considering reads aligned to autosomes 1-22, with a quality score greater than 20; insert length between 150-600, with both ends properly paired; and the reference region of the read does not contain degenerate bases; The combination motif calculation unit is used to perform the following steps: counting 9 bases from the 5' end and dividing them into three groups as the combination motif; The combined motif statistics unit is used to count the frequencies of three groups of 3 bp motifs for each sample. Each group of 3 bp motifs has 64 different combinations, with a total of 192 frequencies as its training variables. The TSS data feature extraction includes a sequencing data alignment unit, a read filtering unit, a gene screening and TSS determination unit, and an NF value calculation unit; The sequencing data alignment unit is used to perform the following steps: after removing the sequencing data sequencing adapters, aligning the sequencing data to the human reference genome hg19; The read filtering unit is used to perform the following steps: filtering the sequencing data, with the following filtering criteria: only considering reads aligned to autosomes 1-22, with a quality score greater than 20, an insert length between 150 and 600, and both ends must be properly paired; and the reference region of the read does not contain degenerate bases; The gene screening and TSS determination unit is used to perform the following steps: determining the TSS of gene transcripts based on the transcript annotation of the UCSC hg19 genome; for genes with multiple TSSs, only those with a difference of less than 50 bp are retained, and their average value is used as the TSS of the gene, and only genes on autosomes are considered; The NF value calculation unit is used to perform the following steps: obtaining parameters for the sequencing coverage depth of cfDNA near the TSS, and the calculation method is as follows: defining the region 500 bp upstream and downstream of the TSS [-250 bp, 250 bp] as the central region; defining 2000 bp upstream [-2000 bp, -1000 bp] and downstream [1000 bp, 2000 bp] of the TSS as the surrounding region; and the NF value of a gene is the average coverage of the central region divided by the average coverage of the surrounding region.
4. The non-invasive early screening device for cancer according to claim 3, characterized in that: The machine learning classification model building module includes: a sample data classification unit, a model parameter acquisition unit and a model effectiveness evaluation unit; The sample data classification unit is used to perform the following steps: dividing the samples into a training set and a test set in a ratio of 4:1, and making the distribution ratios of healthy controls and samples of various cancer types in the two sets consistent; The model parameter acquisition unit is used to perform the following steps: using a machine algorithm in a training set and a 5-fold or 10-fold cross validation method to obtain model parameters; The model effectiveness evaluation unit is used to perform the following steps: drawing a receiver operating characteristic curve of the training set according to the model prediction value and pathological detection result of each sample in the training set.
5. A cancer early screening model product, characterized in that: The cancer early screening model product is a cancer early screening model product constructed by the construction method of claim 1 or 2, which is jointly established with the level of blood CA19-9 and the cancer early screening model based on cfDNA difference statistics. The input of the cancer early screening model product is the score value of the risk scoring model and Log2(1+CA19-9); The output is the model score.
6. The cancer early screening model product according to claim 5, characterized in that: The scoring value of the risk scoring model is the scoring value of the test set.
7. The cancer early screening model product according to claim 5 or 6, characterized in that: When the score of the cancer early screening model product is lower than 0, it is judged as healthy; when it is higher than or equal to 0, it is judged as cancer.
Citation Information
Patent Citations
Cancer detection model and construction method and kit thereof
CN113838533A
Multi-cancer early screening model construction method and detection device
CN114927213A