A colorectal cancer screening and subtype identification method based on circulating cell-free DNA fragment characteristics of genomic functional regions and application
Patent Information
- Application Number
- CN202610926736.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-15
Smart Images

Figure CN122762279A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical detection technology, specifically to a method, system, and application for colorectal cancer screening and subtype identification based on the characteristics of circulating cell-free DNA fragments in genomic functional regions. Background Technology
[0002] Colorectal cancer (CRC) is one of the most prevalent and deadliest malignant tumors worldwide, and early screening and accurate subtyping are crucial for its prevention and treatment. Circulating cell-free DNA (cfDNA) fragment characteristics refer to the characteristic signals exhibited by cfDNA fragments circulating in bodily fluids such as blood in terms of length distribution, terminal sequence motifs, breakpoint preferences, nucleosome localization, and coverage patterns. cfDNA fragment omics characteristics contain tumor-related genetic and epigenetic information, providing new research ideas for non-invasive tumor screening. Genomic functional regions refer to areas of the genome with specific biological functions or structural characteristics. When analyzing cfDNA, it is specifically meant that the breakage, distribution, and terminal sequences of cfDNA fragments in blood are not random, but precisely reflect the state of chromatin in these functional regions of the cells from which they originate (such as normal cells or tumor cells). However, traditional cfDNA analysis often ignores the functional differences between different regions of the genome, resulting in weak signals and insufficient specificity, making it difficult to meet the stringent clinical requirements for early screening sensitivity and subtype accuracy.
[0003] Based on current liquid biopsy techniques, colorectal cancer screening methods based on cfDNA have the following shortcomings that urgently need to be addressed: 1. Crude analytical methods and severe signal dilution: Current mainstream fragmentomics technologies (such as DELFI) typically divide the entire genome into fixed-size continuous regions for analysis. This method is essentially a "one-size-fits-all" approach, failing to consider the functional heterogeneity of the genome. The drawback is that this division artificially fragments complete biological functional units (such as an enhancer-promoter pair), causing weak tumor signals originating from the same regulatory region to be dispersed into multiple adjacent analytical regions, thus being severely diluted by significant background noise, significantly reducing the signal-to-noise ratio and sensitivity of the detection.
[0004] 2. Introducing significant irrelevant noise and limiting specificity: Under a fixed-region partitioning strategy, a large number of non-functional, non-informative genomic sequences (such as indolent chromatin regions) are included in the analysis. These regions themselves do not participate in tumorigenesis, and their cfDNA fragment characteristics show minimal variation, mainly reflecting individual differences and batch effects, constituting the main background noise. Analyzing these regions not only wastes computational resources but also severely interferes with the discrimination of true tumor signals, leading to a bottleneck in improving model specificity and an increased risk of false positives.
[0005] 3. Poor biological interpretability and weak clinical decision support: Because the analytical unit (fixed interval) does not directly correspond to any known biological function, existing technical models are often a "black box." While they can provide a "high risk" or "low risk" score, they cannot explain the biological source of the risk signal (e.g., which oncogenic pathway or regulatory element is abnormal). This makes it difficult for clinicians to develop further, personalized examination or intervention plans based on the results, limiting the technical value to the "screening" level and hindering its extension to "precision diagnosis" and "subtype guidance."
[0006] 4. Insufficient detection capability for early-stage cancer and precancerous lesions: The aforementioned signal dilution and noise interference issues are particularly prominent in early-stage cancers (stage I / II) and precancerous lesions (such as advanced adenomas) with extremely low tumor burden. Current technologies struggle to effectively capture and enrich the extremely weak cfDNA abnormal fragment signals released by these early lesions, thus their sensitivity for early and precancerous lesions is generally low, limiting their true cancer prevention value.
[0007] Therefore, there is an urgent need for a cfDNA analysis method that can effectively enrich tumor-specific signals, provide biological interpretations, and achieve accurate typing, in order to overcome the current technological bottlenecks. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a method for colorectal cancer screening and subtype identification based on the characteristics of circulating cell-free DNA fragments in specific genomic functional regions. This solves the problems of existing circulating cell-free DNA fragment omics detection methods, such as fixed window division failing to reflect genomic functional differences, tumor-related signals being easily diluted by background noise, insufficient sensitivity in early colorectal cancer detection, and lack of subtype identification capabilities.
[0009] The technical solution of the present invention: A method for colorectal cancer screening and subtype identification based on the characteristics of circulating cell-free DNA fragments in genomic functional regions includes the following steps: S1. Collect peripheral blood plasma samples from the subjects; S2. Extract circulating cell-free DNA from plasma samples, construct whole-genome libraries, and perform low-depth whole-genome sequencing on the constructed libraries to obtain circulating cell-free DNA fragment omics data. S3. Based on pre-constructed genomic functional regions, extract multi-dimensional fragment omics features from the circulating cell-free DNA fragment omics data; S4. Construct a sample feature dataset based on multi-dimensional fragment omics features, and construct an ensemble machine learning model based on the sample feature dataset, and train and optimize the ensemble machine learning model; S5. Input the test sample into the trained ensemble machine learning model and output the colorectal cancer risk score and subtype classification results.
[0010] Application of a colorectal cancer screening and subtype identification method based on the characteristics of circulating cell-free DNA fragments in genomic functional regions in early colorectal cancer screening, precancerous lesion risk assessment, auxiliary diagnosis of tumor clinical staging, and development of products for primary localization of left and right hemicolon or liquid biopsy detection.
[0011] Further, step S2 includes: S2-1. Extracting circulating cell-free DNA from plasma samples; 1) Take 400-600 µL of plasma sample, add 10-30 µL of proteinase K solution, 15-35 µL of magnetic beads, and 450-650 µL of lysis buffer, mix well, incubate at room temperature for 10-20 minutes, centrifuge, let stand, and remove the supernatant. 2) Add 600-800 µL of first wash buffer, mix well, centrifuge, let stand, and remove the supernatant; 3) Add 600-800 µL of second wash buffer, mix well, remove liquid, wash 1-3 times, centrifuge, let stand, remove liquid, and dry for 1-5 minutes; 4) Add 30-50 µL of DNA elution buffer, mix well, let stand, and collect the supernatant for later use; S2-2, End repair of circulating cell-free DNA; 1) Prepare the reaction system, which includes T4 DNA ligase buffer, ATP, dNTP mixture, polyethylene glycol, T4 polynucleotide kinase and Taq DNA polymerase; 2) Mix the reaction system with the product obtained in step S2-1 and perform PCR reaction. The reaction conditions are: 10-15℃ for 10-20 minutes, 35-40℃ for 10-20 minutes, and 55-60℃ for 40-50 minutes. S2-3, Connect the sequencing adapter containing a unique molecular identifier; 1) Prepare the reaction system, which includes T4 DNA ligase buffer and T4 DNA ligase; 2) Mix the reaction system with the product obtained in step S2-2 and perform PCR reaction. The reaction conditions are 20-30℃ incubation for 10-20 minutes. S2-4. Purify to obtain sequencing libraries; 1) Add 1.0-1.5 times the amount of magnetic beads to the product obtained in step S2-2, mix well, let stand for 10-20 minutes, centrifuge, let stand, and discard the supernatant; 2) Add 200-300 µL of 80% ethanol, let stand, remove the liquid, centrifuge, let stand, remove the liquid again, and dry. 3) Add 18-25 µL of elution buffer, mix well, let stand, centrifuge, let stand, and obtain the purified product; S2-5, Perform PCR amplification; 1) Configure the PCR amplification reaction system, which includes hot-start premixed enzyme, paired-end indexing primers, and template DNA; 2) Mix the reaction system with the products obtained in steps S2-4 and perform PCR reaction. The reaction steps include: a) reacting at 95-96℃ for 30-60 seconds; b) reacting at 98-99℃ for 10-30 seconds; c) reacting at 60-65℃ for 30-60 seconds; d) reacting at 68-70℃ for 30-60 seconds; repeat steps b) to d) at least 11 times; e) reacting at 72-75℃ for 5-7 minutes; f) finally incubating at 4-10℃. S2-6. Purify again to obtain the final sequencing library. 1) Add 1-1.2 times the amount of magnetic beads to the product obtained in steps S2-5, mix well, let stand for 10-20 minutes, centrifuge, let stand, and remove liquid; 2) Add 200-300 µL of 80% ethanol, let stand, remove the liquid, centrifuge, let stand, remove the liquid again, and dry. 3) Add 30-50 µL of elution buffer, mix well, let stand, centrifuge, let stand, and obtain the purified final sequencing library; S2-7. Take the purified final sequencing library obtained in step S2-6 and use a qubit 4 Fluorometer and AcuQ 1X dsDNA HS Assay Kit to detect the library concentration. S2-8. Perform high-throughput sequencing on the sequencing library to obtain raw sequencing data; S2-9. Identify and remove duplicate sequencing sequences generated by PCR amplification based on unique molecular identifiers.
[0012] 1) Perform quality assessment on the raw sequencing data; 2) Extract unique molecular identifier sequence information from sequencing reads; 3) Perform data quality cleaning and align the sequences to the reference genome; 4) Identify and mark repetitive reads based on unique molecular identifiers and alignment position information, and remove repetitive sequences generated by PCR amplification.
[0013] Furthermore, the genomic functional regions described in step S3 include enhancer regions, superenhancer regions, silencer regions, and DNA methylation regions.
[0014] Furthermore, the multidimensional fragment omics features mentioned in step S3 include the proportion of short fragments, terminal sequences, nucleosome footprints, and copy number variations.
[0015] Furthermore, the short fragment proportion feature was obtained by statistically analyzing the proportion of short fragments (100-150bp / 151-220bp) in different genomic regions (per 5Mb); the terminal sequence feature was obtained by analyzing the base preference of the 5′ end of the circulating free DNA fragment; the nucleosome footprint feature was obtained by analyzing the central coverage, average coverage, and amplitude of 270 transcription factor binding sites to infer nucleosome localization and occupation patterns; the copy number variation feature was obtained by statistically analyzing and correcting the sequencing depth of the whole genome window (1Mb bins) to obtain its relative copy number abundance, and then using a hidden Markov model (HMM) in conjunction with plasma tumor purity estimation to inversely deduce the true copy number variation status of the tumor.
[0016] Furthermore, the integrated machine learning model in step S4 includes multiple base learners and meta-learners; the base learners include at least one of the XGBoost model, LightGBM model, random forest model, and logistic regression model; the meta-learner is used to fuse the output results of each base learner and generate a final diagnostic score.
[0017] Furthermore, in step S4, cross-validation and hyperparameter optimization methods are used to train and optimize the ensemble machine learning model.
[0018] Furthermore, in step S4, the dataset is divided into training, testing, and validation sets in a 6-8:1-2:1-2 ratio to ensure the model's generalization ability.
[0019] Furthermore, the subtype classification results described in step S5 include the clinical stage of colorectal cancer and at least one origin of left-sided or right-sided colorectal cancer. This provides multi-dimensional decision support information for clinical practice.
[0020] The present invention also provides the application of the above method in early screening of colorectal cancer, risk assessment of precancerous lesions, auxiliary diagnosis of clinical staging of tumors, primary localization of left and right hemicolon, or development of liquid biopsy detection products.
[0021] Compared with existing technologies, the present invention has the following advantages: 1) Signal enrichment and high sensitivity: By targeting preset genomic functional regions, it focuses on analyzing high-information regulatory regions, effectively filtering out non-functional noise, and greatly improving the signal-to-noise ratio of weak tumor signals, making it especially suitable for screening early cancer and precancerous lesions.
[0022] 2) Multi-dimensional feature integration with strong discriminative power: By integrating features from multiple dimensions such as fragment size, terminal motifs, nucleosome footprints, and copy number variations, the model can systematically characterize tumor-related chromatin abnormalities, significantly improving the specificity and accuracy of the model.
[0023] 3) Strong biological interpretability: Features are extracted from clearly defined regulatory regions, and their changes can be associated with specific epigenetic events (such as nuclease expression imbalance, enhancer activation, etc.), which makes the risk score and subtype identification results have a clear mechanism explanation, thus contributing to precision medicine.
[0024] 4) Stable and reliable ensemble learning: The stacked ensemble strategy is used to fuse multiple heterogeneous models, which can capture complex nonlinear relationships, avoid overfitting of a single model, and ensure that the model can maintain robust performance in different cohorts.
[0025] 5) Low cost and high throughput: It adopts low-depth whole-genome sequencing, requiring only 6G of data per sample. Combined with functional region targeted analysis strategy, it significantly reduces detection costs while ensuring performance, making it suitable for large-scale population screening.
[0026] 6) Abundant information output: It can not only output disease risk scores, but also provide subtype information such as tumor stage and left and right hemicolon location, providing direct reference for clinical decisions such as surgical planning and targeted therapy. Attached Figure Description
[0027] Figure 1 This is the overall technical roadmap of the present invention.
[0028] Figure 2 An integrated multi-omics atlas of colorectal cancer (CRC). Among them, Figure 2 a provides an overview of the integrated research framework, demonstrating the correspondence between plasma cfDNA and tissue multi-omics molecular characteristics in colorectal cancer patients and healthy controls, as well as the scale of the construction of multi-omics datasets for plasma and tissue samples; Figure 2 b shows the whole genome cfDNA fragmentation profiles of healthy controls and colorectal cancer patients, indicating that the overall cfDNA in cancer samples is shifted toward shorter fragments; Figure 2 c represents the degree of cfDNA fragmentation and the results of circulating tumor DNA load analysis.
[0029] Figure 3 This is a graph showing the results of the correlation analysis between circulating cell-free DNA (cfDNA) and matched tumor tissue in colorectal cancer. Among them, Figure 3 a shows an overview of the epigenome maps of normal and tumor tissues; Figure 3 b shows the genome-wide distribution of the H3K27ac signal; Figure 3 c shows the comparative genomic distribution of H3K27ac in the promoter region; Figure 3d shows the cluster analysis results of the H3K27ac peak with other omics data; Figure 3 e indicates enhanced regions of activity that are increased or decreased in the tumor; Figure 3 f shows the sample stratification results based on tumor-specific enhancer features; Figure 3 g shows the consistency between tissue enhancer activity and circulating DNA characteristics; Figure 3 h shows the chromatin background of the CRC enhancer; Figure 3 i illustrates the enhancer-promoter interaction network; Figure 3 j shows the results of pathway enrichment analysis of enhancer-related target genes.
[0030] Figure 4 This figure shows the analysis results of key transcription factors and the HNF4A-FOXQ1 enhancer regulatory axis in CRC. Figure 4 a shows the results of screening enhancer-driven oncogenes by integrating tumor upregulated genes with enhancer-promoter associated target genes; Figure 4 b shows the expression patterns of 41 enhancer-related genes in normal and tumor tissues; Figure 4 c shows the prognostic correlation analysis results of the major enhancer-associated genes; Figure 4 d shows the results of the transcription factor motif enrichment analysis within the tumor-acquired enhancer region; Figure 4 e shows the association map between enriched transcription factors and prognostic target genes; Figure 4 f shows the enrichment of HNF4A binding sites in tumor-specific enhancer regions associated with prognostic genes; Figure 4 g shows the regulatory architecture of the FOXQ1 locus; Figure 4 h shows the results of the clinical relevance analysis of HNF4A; Figure 4 i shows the results of pathway enrichment analysis in tumors with high FOXQ1 expression; Figure 4 j shows the results of the association analysis between FOXQ1 expression and clinicopathological features and patient prognosis in an independent cohort; Figure 4 k to p show the results of single-cell transcriptome analysis, showing the expression patterns of FOXQ1 in specific epithelial subsets; Figure 4 q and r show the immunohistochemical staining and quantitative analysis results of FOXQ1 and HNF4A in normal controls and CRC tumors.
[0031] Figure 5 This is a graph showing the analysis results of super-enhancer remodeling and DNA methylation in colorectal cancer. Among them: Figure 5 a shows the difference in super-enhancer activity between colorectal cancer and normal tissue; Figure 5 b shows the results of aggregated enrichment analysis of active chromatin signals in tumor-specific and normal-specific super-enhancing subregions; Figure 5c shows the results of identifying super-enhancer-regulated genes by integrating regulatory interactions and transcriptional information; Figure 5 d shows the expression patterns of representative super-enhancer-related genes in normal and tumor tissues; Figure 5 e shows the prognostic analysis results of super-enhancer driver genes; Figure 5 f and g show a representative genomic view of the interaction between super-enhancer activation and long-range regulation at the NTMT1 and SOX9 gene loci; Figure 5 h shows the validation results of protein levels of super-enhancer-driven gene activation; Figure 5 i shows the genomic distribution of circulating DNA methylation signals in different functional regions; Figure 5 j shows the characteristics of tumor-associated DNA methylation changes in plasma; Figure 5 k shows the diagnostic performance of the DNA methylation-based classifier in distinguishing colorectal cancer from the control group; Figure 5 l shows the identification results of genes that have enhancer connections but lack transcriptional activation; Figure 5 m shows the epigenetic relationship between promoter hypermethylation and enhancer-related activation at the MTOR locus; Figure 5 n indicates genes exhibiting a synergistic effect of promoter hypermethylation and transcriptional repression in colorectal cancer; Figure 5 o shows the functional pathways enriched in methylation-silenced genes; Figure 5 p shows the results of the association analysis between methylation-mediated gene silencing and clinical prognosis; Figure 5 q shows the genome-wide distribution of tumor-associated DNA methylation signals.
[0032] Figure 6 This is a graph showing the results of the H3K27me3 landscape remodeling analysis in colorectal cancer. Among them: Figure 6 a and Figure 6 b shows the genome-wide distribution of H3K27me3 signaling in tumor tissue and normal tissue; Figure 6 c shows the tumor-associated regions of H3K27me3 gain and loss differences; Figure 6 d illustrates the identification process for distal silencing-promoter regulatory interactions; Figure 6 e shows the classification results of tumor samples and normal samples based on the H3K27me3 signal spectrum; Figure 6 f shows the genomic annotations of tumor-specific H3K27me3 gain and loss regions; Figure 6 g indicates the proportion of the dynamic H3K27me3 region involved in the silencing-promoter interaction; Figure 6 h illustrates the relationship between dynamic changes in H3K27me3 and gene transcription; Figure 6 i shows the results of functional pathway enrichment analysis related to genes regulated by dynamic H3K27me3; Figure 6j and k show the results of the clinical relevance analysis of genes regulated by silencers; Figure 6 l shows the epigenetic and chromatin interaction map of the VWA2 locus; Figure 6 m shows the results of the association analysis between VWA2 expression and patient survival; Figure 6 n shows a comparison of VWA2 protein levels in normal tissues and tumor tissues.
[0033] Figure 7 The figure shows the structural and diagnostic performance analysis results of the ERRFPM multimodal fragment omics model. Among them: Figure 7 a shows an overview of the ERRFPM multimodal stacking integration framework; Figure 7 b shows the overall changes in the terminal motifs of fragments in cfDNA derived from colorectal cancer; Figure 7 c shows the changes in tumor-associated nuclease expression in independent cohorts; Figure 7 d and Figure 7 e shows the analysis results of cfDNA fragmentation heterogeneity in CRC patients; Figure 7 f shows the difference in motif frequencies at the ends of 3-mer fragments; Figure 7 g shows a comparison of the diagnostic performance of single cfDNA feature classes and integrated models; Figure 7 h shows the diagnostic performance results of the model on the training set, validation set, and independent test queue; Figure 7 i shows the diagnostic performance results stratified by disease stage. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the following description will be provided in conjunction with the appendix. Figures 1 to 7 The present invention will be further described in detail with reference to specific embodiments. Example
[0035] A method for colorectal cancer screening and subtype identification based on the characteristics of circulating cell-free DNA fragments in genomic functional regions includes the following steps.
[0036] S1. Collect peripheral blood plasma samples from the subjects.
[0037] Peripheral blood plasma samples were collected from colorectal cancer patients and healthy controls to establish a case-control cohort. The fundamental differences in the content and fragment characteristics of circulating cell-free DNA (cfDNA) in the bloodstream between cancer patients and healthy individuals were utilized to provide labeled training data for subsequent modeling.
[0038] Specifically, peripheral blood samples were collected from 223 untreated CRC patients and 200 healthy individuals. All participants signed informed consent forms. A printed label was affixed to the side of the blood collection tube, and a nurse collected 5 mL of peripheral blood using a tube containing EDTA anticoagulant. The blood was gently mixed with the anticoagulant and sent to the laboratory for plasma separation within 1-3 hours. The collected blood was then centrifuged at 2500g for 10 minutes within 1-3 hours. The upper plasma layer and the middle leukocyte layer were collected. The plasma was centrifuged again at 2500g for 10 minutes, and the upper plasma layer was collected again. The collected plasma was aliquoted and stored with the leukocytes at -80°C.
[0039] S2. Extract circulating cell-free DNA from plasma samples, construct whole-genome libraries, and perform low-depth whole-genome sequencing on the constructed libraries to obtain circulating cell-free DNA fragment omics data.
[0040] S2-1. Extract circulating cell-free DNA from plasma samples.
[0041] Before extracting cfDNA from plasma using the BGI MGI Easy Cell-Free DNA Extraction Kit, dissolve the proteinase K powder (stored at 4°C) in dissolve buffer and mix thoroughly. Simultaneously, add 200 mL of anhydrous ethanol to Wash Buffer 2 and mix well. Frozen plasma taken from -80°C should be thawed slowly on ice. Take a 1.5 mL sterile centrifuge tube, add 20 µL of Protease K solution and 25 µL of MGI PureParticleG magnetic beads, then add 300 µL of thawed plasma sample. Add 550 µL of lysis buffer, vortex for 10 seconds, and incubate at room temperature for 15 minutes, inverting to mix every 5 minutes. Afterward, centrifuge briefly, place on a magnetic rack and let stand for 3 minutes. Discard the supernatant. Next, add 700µL of Wash Buffer 1, vortex for 15 seconds and briefly centrifuge. After incubating on a magnetic rack for 1 minute, remove the supernatant. Then add 700µL of Wash Buffer 2, vortex to mix, and remove the liquid, keeping the magnetic beads. Repeat once. Remove the EP tube, briefly centrifuge to collect the liquid on the tube wall, incubate on a magnetic rack for 1 minute, aspirate the liquid, and then dry for 2 minutes. Finally, add 30µL of DNA Elution Buffer, vortex to mix, and incubate at room temperature for 10 minutes, gently shaking 1-2 times during this period to accelerate DNA dissolution. Then incubate on a magnetic rack for 3 minutes. Transfer the supernatant to a new 1.5mL DNA LoBind® tube and store at -20°C for subsequent experiments. This entire process ensures efficient recovery of cfDNA and reduces the risk of degradation.
[0042] S2-2, End repair of circulating cell-free DNA.
[0043] 1) Prepare the reaction system according to Table 1 below, mix thoroughly to obtain a mixture, add 10 µL of the mixture to the product obtained by S2-1 with a total volume of 30 µL, mix thoroughly again, and centrifuge at 2500 g for 5 s.
[0044]
[0045] 2) Perform the PCR reaction. The reaction procedure is shown in Table 2 below.
[0046]
[0047] S2-3, Connect the sequencing adapter containing a unique molecular identifier.
[0048] 1) Configure the reaction system (single sample) according to Table 3 below.
[0049]
[0050] 2) Mix the mixture in Table 3 thoroughly. Add 65 µL of the product sample obtained in step S2-2 to the mixture, mix thoroughly again, and centrifuge at 2500 g for 5 s.
[0051] 3) Add 5 µL of the mixture in Table 3, mix thoroughly, and centrifuge at 2500 g for 5 s.
[0052] 4) Perform PCR, and the reaction conditions are shown in Table 4 below.
[0053]
[0054] S2-4. Purify to obtain sequencing libraries.
[0055] 1) Equilibrate the magnetic beads to room temperature for 30 min in advance. Add 1.2 times the amount of beads to the product sample obtained in each step S2-3, mix well and let stand for 15 min, inverting once every 5 min.
[0056] 2) Briefly centrifuge, let stand on a magnetic rack for 10 minutes to remove liquid.
[0057] 3) Add 200 µL of 80% ethanol, let stand on a magnetic rack for 1 min, then remove the liquid. Repeat this step.
[0058] 4) Centrifuge, collect the liquid from the tube wall at the bottom of the centrifuge tube, then place it on a magnetic rack and let it stand for 1 minute before discarding the liquid. Continue drying for 2 minutes, for a total of 3 minutes.
[0059] 5) After removing the centrifuge tube, add 18 µL of elution buffer, mix well, and let stand for 5 minutes, shaking once during this time.
[0060] 6) After centrifugation, place the tube on a magnetic rack for 10 min, and then transfer the liquid into a new 0.2 mL PCR tube.
[0061] S2-5, Perform PCR amplification.
[0062] 1) Configure the PCR amplification reaction system according to Table 5 below.
[0063]
[0064] 2) Mix the PCR amplification reaction system with the products obtained in steps S2-4, and centrifuge at 2500 g for 5 seconds. Perform the PCR reaction as follows: a) react at 95℃ for 60 seconds; b) react at 98℃ for 10 seconds; c) react at 60℃ for 30 seconds; d) react at 68℃ for 30 seconds; repeat steps b) to d) 11 times; react at 72℃ for 5 minutes, and finally incubate at 4℃.
[0065] S2-6, Purify again to obtain the final library. 1) Add 1.2 times the amount of magnetic beads to the product obtained in step S2-3, mix well, let stand for 10-20 minutes, centrifuge, let stand, and discard the liquid.
[0066] 2) Add 200 µL of 80% ethanol, let stand, remove the liquid, centrifuge, let stand, remove the liquid again, and dry.
[0067] 3) Add 30 µL of elution buffer, mix well, let stand, centrifuge, let stand, and obtain the final library.
[0068] S2-7. Take 1µL of the purified product and use a qubit 4 Fluorometer and an AcuQ 1X dsDNA HSAssay Kit to detect the library concentration.
[0069] S2-6. Perform high-throughput sequencing on the sequencing library to obtain raw sequencing data.
[0070] DNA library quality control was performed. 1 µL of library sample was used for concentration quantification using the Qubit 4.0 nucleic acid sequencer: the concentration of extracted cfDNA was determined using the Vazyme Equalbit 1× dsDNA HS Assay Kit. Before the experiment, all components of the kit were allowed to equilibrate at room temperature for at least 30 min. The entire detection process was conducted in the dark to avoid degradation of the fluorescent dye and its impact on the results. Separate standard and sample reaction systems were prepared: Standard 1 consisted of 190 μL Equalbit 1× dsDNA HS Working Solution mixed with 10 μL Standard 1; Standard 2 consisted of 190 μL Equalbit 1× dsDNA HS Working Solution mixed with 10 μL Standard 2; the sample reaction system consisted of 199 μL Equalbit 1× dsDNA HS Working Solution mixed with 1 μL cfDNA sample. All reactions were performed in 0.5 mL centrifuge tubes, and sample information was labeled on the tube caps. After mixing, the liquid was collected by brief centrifugation and incubated at room temperature for 2 min in the dark. Subsequently, the concentrations of standard 1, standard 2, and each cfDNA sample were measured sequentially according to the instrument's requirements using a Qubit 4.0 quantitative PCR instrument. The sample DNA concentration was automatically calculated based on the standard curve, and the results were recorded for subsequent library construction and data analysis. An additional 1 µL was analyzed using a TapeStation system to determine the electrophoretic pattern and average DNA fragment size of the library.
[0071] S2-7. Identify and remove duplicate sequencing sequences generated by PCR amplification based on unique molecular identifiers (UMIs). Removing duplicates by identifying UMIs effectively corrects random errors in the PCR amplification process, significantly improving the authenticity and accuracy of the data.
[0072] We first used FastQC software to perform quality control on the raw sequencing data. Both library construction and sequencing involve DNA polymerase and amplification steps, which can introduce biases that can lead to false positives. Therefore, we introduced a unique molecular tag (UMI) during low-depth WGS library construction. This is a completely random 2 bp or 4 bp nucleotide chain (such as TA or GTGA) to correct for sequencing errors or errors introduced during PCR amplification. The UMI was extracted and written into the read header using FastP (v 0.23.1) software. Simultaneously, data quality cleaning was performed, and the cleaned sequence files were aligned to the reference genome hg38 using bwa (0.7.18-r1243-dirty). Finally, duplicate reads were identified using GenCore based on the UMI and alignment markers.
[0073] S3. Based on pre-constructed genomic functional regions, extract multi-dimensional fragment omics features (such as...) from the circulating cell-free DNA fragment omics data. Figures 2-6 (As shown).
[0074] S3-1, Pre-construction of genomic functional regions. The construction process is as follows: Differential enhancers: H3K27ac cut & run sequencing was performed on 223 CRC tissues and paired normal tissues. Figure 3 (b-3d), peak identification was performed using MACS2 (q<0.01), and DiffBind was used to filter differential enhancers with |log2FC|>1 and FDR<0.05. For example... Figure 3 b. Genome-wide distribution of H3K27ac signaling in normal and tumor tissues, showing the redistribution of acetylation modifications in tumors to distant intergenic regions and exons. For example... Figure 3 Comparative genomic distribution of c,H3K27ac showed that its enrichment was higher in promoter regions in colorectal cancer patients. Figure 3 d. Cluster analysis was performed on all H3K27ac peaks in colorectal cancer tissues, along with cut & run data of open chromatin signals (ATAC-seq), epigenetic markers and transcription factor binding signals (H3K4me1, H3K27me3, and P300), CpG methylation data, and whole-genome sequencing (WGS) data of blood cfDNA. The clustering results showed that six clusters were the most ideal. Figure 3 e. H3K27ac-labeled enhancer regions with increased or decreased activity compared to normal tissue in colorectal cancer tumors. For example... Figure 3 f, Sample stratification based on tumor-specific enhancer features showed a clear separation between tumor and normal tissue, consistent with results from colorectal cancer cell lines. For example... Figure 3 g. Consistency between tissue enhancer activity and circulating DNA signatures: active tumor enhancers correspond to phased cfDNA fragmentation and enriched cfDNA methylation signals, while the repressed state lacks these plasma-related patterns. For example... Figure 3 h, chromatin background of CRC enhancers, shown in tumor upregulated and downregulated regions, with chromatin accessibility and H3K4me1 signal exhibiting coordinated gain or loss, respectively. Figure 3 i. A high-confidence enhancer-promoter interaction network links tumor-activated enhancers to downstream target genes. For example... Figure 3 j. Pathway enrichment analysis of enhancer-related target genes highlighted signaling pathways associated with CRC, including the PI3K-AKT and Hippo signaling pathways.
[0075] Super enhancers: Using the ROSE algorithm, the H3K27ac signal intensity was sorted on the genome to identify super enhancer regions that exceeded the background threshold. Figure 5 c-5e showcases the genes regulated by super enhancers and their prognostic analysis; Figure 5 f-5g represents representative super-enhancer-related gene loci.
[0076] Silent sub-region: Perform H3K27me3 CUT&RUN on CRC-organized and normal regions ( Figure 6 (a-6c) identifies regions of difference where the signal is gained or lost. Figure 6 d illustrates the process for identifying remote silencing-promoter interactions; Figure 6 e shows the sample classification results based on the H3K27me3 signal; Figure 6 f-6n presents the results of dynamic H3K27me3 region, related functional pathways, and VWA2 locus regulation analysis.
[0077] Differential DNA methylation regions: Based on cfMeDIP-seq data, regions with significant differences in methylation levels between CRC and the control group were identified (|Δβ|>0.2, FDR<0.05). Figure 5 l-5q presents the distribution of circulating DNA methylation signals, the performance of the methylation classifier, and the results of methylation-mediated gene silencing analysis.
[0078] The final result is a set of genomic functional regions including differential enhancers, super enhancers, H3K27me3 remodeling regions, and differentially methylated regions.
[0079] S3-2, Multi-dimensional Fragment Omics Feature Extraction For each of the above functional regions, the following four dimensions of features are extracted from the cfDNA sequencing data: (1) Short fragment ratio: Obtained by statistically analyzing the proportion of short fragments (100-150bp / 151-220bp) in different genomic regions (per 5Mb). The proportion of short fragments with a length ≤150 bp in each functional region to the total number of reads (Short Fragment Fraction, SFF) was calculated. This feature reflects the overall trend of shortening tumor-derived cfDNA fragments. Short fragment to long fragment ratio = (number of short fragments) / (number of long fragments, i.e., the number of fragments with a length >150 bp).
[0080] (2) Terminal sequence: obtained by the base preference of the 5′ end of the circulating free DNA fragment (Motif analysis), such as the frequency of occurrence of trinucleotide motifs.
[0081] like Figure 7 As shown in b, C-terminal motifs (such as CCC) are significantly reduced and A-terminal motifs (such as ACA and ATT) are significantly enriched in the cfDNA of CRC patients. Figure 7 c shows the association between changes in nuclease expression (downregulation of DNASE1L3, upregulation of DNASE1 and DFFB) and changes in terminal motifs. Figure 7 d-7e shows the results of cfDNA fragmentation heterogeneity analysis (Gini diversity index and FrEIA score). Figure 7 f shows the specific data on the difference in motif frequencies at the ends of 3-mer fragments.
[0082] (3) Nucleosome footprint: obtained by analyzing the central coverage, average coverage and amplitude of binding sites of 270 transcription factors to infer the nucleosome localization and occupation pattern.
[0083] Within a 5 kb radius on either side of the center point of each functional region (TSS for promoters, peak center for enhancers), the coverage depth is statistically analyzed every 50 bp window. Fourier transform is used to extract the periodic patterns of coverage fluctuations (period approximately 200 bp), and the nucleosome occupancy index (NOI = peak depth / trough depth) is calculated to quantify changes in nucleosome arrangement. This feature infers nucleosome localization and occupancy patterns based on fragment coverage depth.
[0084] (4) Copy number variation: The relative copy number abundance of the whole genome window (1Mb bins) is obtained by statistically analyzing and correcting the sequencing depth. Then, the true copy number variation status of the tumor is deduced by using the Hidden Markov Model (HMM) in combination with plasma tumor purity estimation.
[0085] Using CNVkit software and paired leukocyte DNA as a reference, the log2 relative copy number ratio (log2 ratio) of each functional region was calculated. For amplified or deleted regions greater than 2 Mb, segmentation and smoothing were performed, and the smoothed log2 ratio value for each functional region was extracted. This feature was calculated based on circulating cell-free DNA whole-genome sequencing data.
[0086] Each sample ultimately yields a feature vector, with dimension = number of functional regions × feature categories. To reduce dimensionality redundancy, variance filtering (removing features with variance below the 10th percentile) and correlation filtering (removing one of the feature pairs with a Pearson correlation coefficient > 0.95) are used for subsequent modeling.
[0087] S4. Construct a sample feature dataset based on multi-dimensional fragment omics features, and construct an ensemble machine learning model based on the sample feature dataset, and train and optimize the ensemble machine learning model; S4-1, Sample Feature Dataset Construction: Integrate all sample feature vectors (423 samples) extracted in step S3 and their corresponding clinical labels (whether it is CRC, clinical stage, left and right hemicolon position) to construct the sample feature dataset.
[0088] S4-2, Dataset Splitting: All 423 samples were randomly divided into a training set (254 cases, including 134 CRC cases and 120 control cases) in a 6:2:2 ratio, a validation set (85 cases, including 45 CRC cases and 40 control cases), and a test set (84 cases, including 44 CRC cases and 40 control cases). The distribution of each stage and the left and right hemicolons remained proportional across all sets. Figure 7 h-7i).
[0089] S4-3, Ensemble Machine Learning Model Construction: A stacking strategy is adopted. The model structure includes multiple base learners and one meta-learner.
[0090] The prediction model is constructed using a combination of hyperparameter tuning, cross-validation, and stacking ensemble learning. Rigorous partitioning ensures unbiased model evaluation. The stacking algorithm, through a meta-learner, integrates multiple base learners, capturing complex nonlinear relationships and significantly improving the model's prediction accuracy and generalization ability.
[0091] Five-fold cross-validation was performed on the training set, training four base learners: XGBoost, LightGBM, Random Forest, and Logistic Regression. Each base learner was hyperparameter-tuned using grid search. Then, using the cancer probabilities output by the four base learners during cross-validation as new features, a meta-learner was trained to learn the optimal fusion weights, forming a... Figure 7 The final stacked integration model shown in figure a (the model is named ERRFPM (Epigenetic Regulatory Regionsrelated Fragmentomics Prediction Model)).
[0092] S4-4, Model Training and Optimization: Cross-validation and hyperparameter optimization methods are used to train and optimize the ensemble machine learning model. Specific process: Five-fold cross-validation is performed on the training set to ensure the robustness of the model parameters.
[0093] A systematic search is performed in the hyperparameter space using grid search, with the validation set AUC as the optimization objective.
[0094] After each set of hyperparameter combinations has been cross-validated, the model performance is evaluated on the validation set. If the AUC is below 0.90, the base learner parameter combination is adjusted back or a feature selection step is added.
[0095] In this embodiment, the model's average AUC on the training set with 5-fold cross-validation is 0.986, and its AUC on the validation set is 0.978. The model's final parameters are locked after successful validation.
[0096] like Figure 7 As shown in h, the performance results of the ERRFPM model on the training set, validation set, and independent test set are as follows: AUC=0.986 for the training set, AUC=0.978 for the validation set, and AUC=0.969 for the test set, showing that the model has good generalization ability and no obvious overfitting.
[0097] S4-5, Model Performance Evaluation: On the independent test set (84 cases), the ERRFPM model achieved the following performance metrics ( Figure 7 h): Training set AUC = 0.986 The validation set AUC is 0.978. Test set AUC = 0.969 Test set sensitivity = 92.9% (41 / 44) Test set specificity = 90.0% (36 / 40) Test set accuracy = 91.7% (77 / 84) like Figure 7 As shown in i, the AUCs stratified by disease stage are as follows: Stage I AUC = 0.871, Stage II AUC = 0.875, Stage III AUC = 0.883, and Stage IV AUC = 0.924. The sensitivities for each stage are: Stage I 87.5% (7 / 8), Stage II 90.9% (10 / 11), Stage III 94.7% (18 / 19), and Stage IV 100% (6 / 6). These data confirm that the model also has a high detection capability for early-stage CRC, and the AUC increases with stage progression, reaching 0.924 for Stage IV.
[0098] S6. Input the test samples into the trained ensemble machine learning model, and output the colorectal cancer risk score and subtype classification results. S6-1, Colorectal Cancer Risk Score: For any plasma sample to be tested, after obtaining the feature vector according to steps S2-S4, it is input into the trained ERRFPM model. The model first outputs a probability value (0-1) from each of the four base learners, then inputs these four probability values into the meta-learner, and finally outputs a fused diagnostic score. Threshold setting: Based on the principle of maximizing the Youden index in the validation set (Youden index = sensitivity + specificity - 1), the optimal threshold is determined to be 0.52.
[0099] Judgment criteria: When the score is ≥ 0.52, it is considered "high risk" (high probability of colorectal cancer). When the score is less than 0.52, it is considered "low risk" (high probability of being healthy).
[0100] S6-2, Subtype Classification Results: The model of this invention can not only output a binary risk score, but also simultaneously provide the following subtype classifications (based on the same feature vector, achieved through an additional multi-classifier): (a) Clinical staging of colorectal cancer: During the training phase, a four-class logistic regression model is constructed (with CRC samples as input and labels as I / II / III / IV), also using a stacked ensemble structure.
[0101] (b) Identification of the origin of left-sided and right-sided colorectal cancer: CRC samples were categorized into left (distal to the splenic flexure, including the descending colon, sigmoid colon, and rectum) and right (ascending colon and transverse colon) based on the location of the primary lesion. An independent logistic regression model was trained using these binary labels (again, using stacked ensemble). On the test set, the results can provide a direct reference for clinical selection of chemotherapy regimens (e.g., the anti-EGFR monoclonal antibody cetuximab is only effective for left-sided CRC).
[0102] The above descriptions are merely embodiments of the present invention, and common knowledge such as specific technical solutions and / or characteristics are not described in detail here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the technical solutions of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A method for colorectal cancer screening and subtype identification based on circulating cell-free DNA fragment features of genomic functional regions, characterized in that, Includes the following steps: S1. Collect peripheral blood plasma samples from the subjects; S2. Extract circulating cell-free DNA from plasma samples, construct whole-genome libraries, and perform low-depth whole-genome sequencing on the constructed libraries to obtain circulating cell-free DNA fragment omics data. S3. Based on pre-constructed genomic functional regions, extract multi-dimensional fragment omics features from the circulating cell-free DNA fragment omics data; S4. Construct a sample feature dataset based on multi-dimensional fragment omics features, and construct an ensemble machine learning model based on the sample feature dataset, and train and optimize the ensemble machine learning model; S5. Input the test sample into the trained ensemble machine learning model and output the colorectal cancer risk score and subtype classification results.
2. The method of claim 1, wherein, Step S2 includes: S2-1. Extracting circulating cell-free DNA from plasma samples; 1) Take 400-600 µL of plasma sample, add 10-30 µL of proteinase K solution, 15-35 µL of magnetic beads, and 450-650 µL of lysis buffer, mix well, incubate at room temperature for 10-20 minutes, centrifuge, let stand, and remove the supernatant. 2) Add 600-800 µL of first wash buffer, mix well, centrifuge, let stand, and remove the supernatant; 3) Add 600-800 µL of second wash buffer, mix well, remove liquid, wash 1-3 times, centrifuge, let stand, remove liquid, and dry for 1-5 minutes; 4) Add 30-50 µL of DNA elution buffer, mix well, let stand, and collect the supernatant for later use; S2-2, End repair of circulating cell-free DNA; 1) Prepare the reaction system, which includes T4 DNA ligase buffer, ATP, dNTP mixture, polyethylene glycol, T4 polynucleotide kinase and Taq DNA polymerase; 2) Mix the reaction system with the product obtained in step S2-1 and perform PCR reaction. The reaction conditions are: 10-15℃ for 10-20 minutes, 35-40℃ for 10-20 minutes, and 55-60℃ for 40-50 minutes. S2-3, Connect the sequencing adapter containing a unique molecular identifier; 1) Prepare the reaction system, which includes T4 DNA ligase buffer and T4 DNA ligase; 2) Mix the reaction system with the product obtained in step S2-2 and perform PCR reaction. The reaction conditions are 20-30℃ incubation for 10-20 minutes. S2-4, magnetic bead purification; 1) Add 1.0-1.5 times the amount of magnetic beads to the product obtained in step S2-3, mix well, let stand for 10-20 minutes, centrifuge, let stand, and discard the liquid; 2) Add 200-300 µL of 80% ethanol, let stand, remove the liquid, centrifuge, let stand, remove the liquid again, and dry. 3) Add 18-25 µL of elution buffer, mix well, let stand, centrifuge, let stand, and obtain the purified product; S2-5, Perform PCR amplification; 1) Configure the PCR amplification reaction system, which includes hot-start premixed enzyme, paired-end indexing primers, and template DNA; 2) Mix the reaction system with the products obtained in steps S2-4 and perform PCR reaction. The reaction steps include: a) reacting at 95-96℃ for 30-60 seconds; b) reacting at 98-99℃ for 10-30 seconds; c) reacting at 60-65℃ for 30-60 seconds; d) reacting at 68-70℃ for 30-60 seconds; repeat steps b) to d) at least 11 times; e) reacting at 72-75℃ for 5-7 minutes; f) finally incubating at 4-10℃. S2-6. Purify again to obtain the final sequencing library. 1) Add 1-1.2 times the amount of magnetic beads to the product obtained in steps S2-5, mix well, let stand for 10-20 minutes, centrifuge, let stand, and remove liquid; 2) Add 200-300 µL of 80% ethanol, let stand, remove the liquid, centrifuge, let stand, remove the liquid again, and dry. 3) Add 30-50 µL of elution buffer, mix well, let stand, centrifuge, let stand, and obtain the purified final sequencing library; S2-7. Take the purified final sequencing library obtained in step S2-6 and detect the library concentration; S2-8. Perform high-throughput sequencing on the sequencing library to obtain raw sequencing data; S2-9. Identify and remove duplicate sequencing sequences generated by PCR amplification based on unique molecular identifiers. 1) Perform quality assessment on the raw sequencing data; 2) Extract unique molecular identifier sequence information from sequencing reads; 3) Perform data quality cleaning and align the sequences to the reference genome; 4) Identify and mark repetitive reads based on unique molecular identifiers and alignment position information, and remove repetitive sequences generated by PCR amplification.
3. The method of claim 1, wherein, The genomic functional regions mentioned in step S3 include enhancer regions, superenhancer regions, silencer regions, and DNA methylation regions.
4. The method according to claim 1, characterized in that, The multidimensional fragment omics features mentioned in step S3 include the proportion of short fragments, terminal sequences, nucleosome footprints, and copy number variations.
5. The method according to claim 4, characterized in that, The short fragment proportion feature was obtained by statistically analyzing the proportion of short fragments in different genomic regions; the terminal sequence feature was obtained by analyzing the base preference of the 5′ end of circulating free DNA fragments; the nucleosome footprint feature was obtained by analyzing the central coverage, average coverage, and amplitude of binding sites of 270 transcription factors to infer nucleosome localization and occupation patterns; the copy number variation feature was obtained by statistically analyzing and correcting the sequencing depth of the whole genome window to obtain its relative copy number abundance, and then using a hidden Markov model, combined with plasma tumor purity estimation, to reversely deduce the true copy number variation status of the tumor.
6. The method according to claim 1, characterized in that, The integrated machine learning model described in step S4 includes multiple base learners and meta-learners; the base learners include at least one of the XGBoost model, LightGBM model, random forest model, and logistic regression model; the meta-learner is used to fuse the output results of each base learner and generate a final diagnostic score.
7. The method according to claim 1, characterized in that, In step S4, cross-validation and hyperparameter optimization methods are used to train and optimize the ensemble machine learning model.
8. The method according to claim 1, characterized in that, In step S4, the dataset is divided into training set, test set and validation set according to the ratio of 6-8:1-2:1-2.
9. The method according to claim 1, characterized in that, The subtype classification results in step S5 include the clinical stage of colorectal cancer and at least one of the origins of left-sided and right-sided colorectal cancer.
10. The application of the method according to any one of claims 1 to 9 in the development of products for early screening of colorectal cancer, risk assessment of precancerous lesions, auxiliary diagnosis of clinical staging of tumors, primary localization of left and right hemicolon, or liquid biopsy detection.