A multi-omics sequencing-based early screening method for colorectal cancer
Patent Information
- Application Number
- CN202411393743.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-08
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2044-10-08
AI Technical Summary
然而,这些技术同样存在成本高、设备要求高、对患者的辐射暴露等问题
[0026]本发明提供了一种基于多组学测序的结直肠癌早期筛查方法,通过构建结直肠肿瘤早期筛查模型,解决了现有技术准确率较低、成本高、存在辐射等缺陷,实现了一种无创、成本效益高的结直肠癌检测方法。
Smart Images

Figure CN119274655B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biomedical technology, and in particular to an early screening method for colorectal cancer based on multi-omics sequencing. Background Technology
[0002] Colorectal cancer is one of the most common malignant tumors worldwide, with high incidence and mortality rates. According to the World Health Organization, colorectal cancer causes millions of deaths each year. Because early-stage colorectal cancer often lacks obvious symptoms, many patients are diagnosed at an advanced stage, which significantly impacts treatment outcomes. Therefore, early detection and intervention for colorectal cancer are crucial for improving patient survival rates.
[0003] Traditional colorectal cancer screening methods mainly include fecal occult blood tests and colonoscopy. While fecal occult blood tests are simple to perform, their sensitivity and specificity are relatively low, leading to missed or misdiagnosed cases. Colonoscopy, while highly accurate, is invasive, resulting in low patient acceptance and high cost. Advances in medical imaging technologies, such as CT colonography and magnetic resonance colonography, have provided non-invasive methods for colorectal cancer detection. However, these technologies also suffer from high costs, demanding equipment requirements, and radiation exposure for patients. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method for early screening of colorectal cancer based on multi-omics sequencing. By constructing an early screening model for colorectal tumors, a non-invasive and cost-effective method for detecting colorectal cancer is achieved, reducing the economic cost of sequencing and improving the accuracy and sensitivity of detection.
[0005] To achieve the above objectives, the present invention provides the following solution:
[0006] A method for early screening of colorectal cancer based on multi-omics sequencing includes:
[0007] Collect plasma samples for testing;
[0008] DNA methylation sequencing was performed on the plasma sample to be tested to obtain the sequencing data.
[0009] The sequencing data to be tested is input into the trained early screening model for colorectal tumors for classification to obtain the target screening results.
[0010] The training process of the colorectal cancer early screening model includes:
[0011] Colorectal cancer tissue and adjacent normal tissue samples were collected from colorectal cancer patients. Genomic DNA was extracted and DNA methylation sequencing was performed on the colorectal cancer tissue and adjacent normal tissue samples to obtain raw sequencing data.
[0012] The raw sequencing data were sequentially subjected to quality control and aligned with the human reference genome to obtain sequence correspondences;
[0013] The human reference genome was divided into several CpG island regions, and the methylation degree of different haplotypes in each CpG island region was analyzed to obtain reference methylation data;
[0014] Based on the sequence correspondence, the methylation degree difference analysis was performed using the p-values after Benjamini and Hochberg multiple test correction of the reference methylation data and the original sequencing data to obtain tumor-specific regions.
[0015] Plasma samples were collected from colorectal cancer patients and healthy individuals, and cfDNA was extracted and methylated from the plasma samples to obtain plasma cfDNA sequencing data.
[0016] The plasma cfDNA sequencing data were sequentially subjected to quality control and alignment with the human reference genome to obtain coordinate-determined sequencing data.
[0017] The coordinates are statistically analyzed to determine the degree of methylation of sequencing data in the tumor-specific region, thus obtaining a set of methylation features.
[0018] The coordinates are obtained to determine the coordinates of all reads of sequencing data on the reference genome. The p bases at the 3' end of the read data on the human reference genome are taken as the first base fragment, and the proportion of the first base fragment is calculated to obtain the terminal motif characteristics.
[0019] The 3' end of the read data is taken as the second base fragment by q bases upstream and q bases downstream of the human reference genome, and the proportion of the second base fragment is calculated to obtain the breakpoint motif feature;
[0020] The methylation feature set, the terminal motif feature, and the breakpoint motif feature are combined into an initial feature value;
[0021] Using the initial feature values as input data and the disease status as output, a pre-designed integrated machine learning classifier is trained to obtain the early screening model for colorectal tumors.
[0022] Preferably, the haploid includes any one or more of MM, MHL, CHALM, PDR, and Entropy.
[0023] Preferably, p is any integer between 4 and 10; q is any integer between 2 and 5.
[0024] Preferably, the integrated machine learning classifier includes: a first layer and a second layer; the first layer includes: support vector machine, random forest, XGBoost and CatBoost; the second layer includes: logistic regression network.
[0025] The present invention discloses the following technical effects:
[0026] This invention provides a method for early screening of colorectal cancer based on multi-omics sequencing. By constructing an early screening model for colorectal tumors, it solves the defects of existing technologies such as low accuracy, high cost, and radiation exposure, and realizes a non-invasive and cost-effective method for colorectal cancer detection. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a schematic diagram of a multi-omics sequencing-based early screening process for colorectal cancer provided in an embodiment of the present invention. Detailed Implementation
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] The purpose of this invention is to provide a method for early screening of colorectal cancer based on multi-omics sequencing. By constructing an early screening model for colorectal tumors, a non-invasive and cost-effective method for detecting colorectal cancer is achieved, reducing the economic cost of sequencing and improving the accuracy and sensitivity of detection.
[0031] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] Figure 1 This is a schematic diagram of the early colorectal cancer screening process based on multi-omics sequencing provided in this embodiment. Figure 1As shown, this embodiment provides a method for early screening of colorectal cancer based on multi-omics sequencing, including:
[0033] Step 100: Collect the plasma sample to be tested;
[0034] Step 200: Perform DNA methylation sequencing on the plasma sample to be tested to obtain the sequencing data.
[0035] Step 300: Input the sequencing data to be tested into the trained early screening model for colorectal tumors for classification to obtain the target screening results.
[0036] The training process for an early colorectal cancer screening model includes:
[0037] We collected colorectal cancer tissue and adjacent normal tissue samples from colorectal cancer patients, and performed genomic DNA extraction and DNA methylation sequencing on the colorectal cancer tissue and adjacent normal tissue samples to obtain raw sequencing data.
[0038] The raw sequencing data were sequentially subjected to quality control and aligned with the human reference genome to obtain sequence correspondences.
[0039] The human reference genome was divided into several CpG island regions, and the methylation degree of different haplotypes in each CpG island region was analyzed to obtain reference methylation data;
[0040] Based on the sequence correspondence, the difference in methylation degree was analyzed using the p-values after Benjamini and Hochberg multiple test correction of the reference methylation data and the original sequencing data to obtain the tumor-specific regions.
[0041] Plasma samples were collected from colorectal cancer patients and healthy individuals, and cfDNA was extracted and methylated from the plasma samples to obtain plasma cfDNA sequencing data.
[0042] The plasma cfDNA sequencing data were sequentially subjected to quality control and aligned with the human reference genome to obtain coordinate-determined sequencing data.
[0043] Statistical coordinates are used to determine the degree of methylation in tumor-specific regions of sequencing data, resulting in a set of methylation features.
[0044] The coordinates of all sequencing reads on the reference genome are determined by obtaining the coordinates. The p bases at the 3' end of the read on the human reference genome are taken as the first base fragment, and the proportion of the first base fragment is calculated to obtain the terminal motif characteristics.
[0045] The q bases upstream and q bases downstream of the 3' end of the read data in the human reference genome are taken as the second base fragment, and the proportion of the second base fragment is calculated to obtain the breakpoint motif characteristics.
[0046] The methylation feature set, terminal motif features, and breakpoint motif features are combined into an initial feature value;
[0047] Using initial feature values as input data and disease status as output, a pre-designed ensemble machine learning classifier is trained to obtain an early screening model for colorectal tumors.
[0048] Specifically, plasma samples from healthy individuals were used as negative controls for training the model, and plasma samples from colorectal cancer patients were processed in the same way.
[0049] Optionally, haploids include any one or more of MM, MHL, CHALM, PDR, and Entropy.
[0050] Preferably, p is any integer between 4 and 10; q is any integer between 2 and 5.
[0051] Specifically, the integrated machine learning classifier includes: a first layer and a second layer; the first layer includes: support vector machine, random forest, XGBoost and CatBoost; the second layer includes: logistic regression network.
[0052] Furthermore, a device for the early detection of colorectal cancer, the device comprising the following components:
[0053] Tissue sample sequencing module: Processes collected cancer tissue and adjacent normal tissue samples, extracts genomic DNA, constructs libraries and performs methylation sequencing to generate raw sequencing output data;
[0054] Blood sample sequencing module: processes collected blood samples, extracts cfDNA, constructs libraries and performs methylation sequencing, and also generates raw sequencing data;
[0055] Methylation feature set acquisition module: performs quality control on raw sequencing data and compares it with the human reference genome to identify the methylation status of cancer-specific regions under different haplotypes and form a methylation feature set;
[0056] Terminal motif feature acquisition module: Analyzes the p bases at the 3' end of each sequencing fragment to form a base fragment set, and calculates the distribution ratio of these fragments in the entire sample to determine the terminal motif features;
[0057] Breakpoint motif acquisition module: Analyzes q base pairs upstream and downstream of the 3' end of each sequencing fragment to form a base fragment set, and calculates the distribution ratio of these fragments to determine the breakpoint motif characteristics;
[0058] Model building module: The extracted methylation feature set, terminal motif features, and breakpoint motif features are used as inputs to the first layer of a composite machine learning classification framework, and the cancer diagnosis results are used as the output of the second layer to train the model and build an early screening model.
[0059] Optionally, a computer-readable medium stores a software program capable of performing the above-described method for constructing an early colorectal cancer screening model.
[0060] Specifically, the datasets used in the model building process are shown in Table 1.
[0061] Table 1
[0062]
[0063]
[0064] Preferably, the method for extracting and sequencing tissue samples is as follows:
[0065] Tissue samples were collected from the subjects, and genomic DNA was extracted using a genomic DNA extraction kit provided by TIANGEN, following the product instructions. A library was constructed from the extracted genomic DNA, and RRBS sequencing was then performed.
[0066] After sequencing is completed, the raw data undergoes quality control to select high-quality sequencing data, which are then compared with the human reference genome for further analysis.
[0067] Furthermore, the methods for extracting and sequencing plasma cfDNA samples:
[0068] A 10ml whole blood sample was collected from the subject using a cell-free DNA collection tube. After separating the plasma, cfDNA was extracted using a plasma DNA extraction kit provided by GENFINE, following the product instructions. The extracted cfDNA was used to construct a library using the aforementioned patented technology, followed by RRBS sequencing. After sequencing, the raw data underwent quality control, and high-quality sequencing data were selected and compared with the human reference genome for in-depth analysis.
[0069] Specifically, data processing. The early screening method in this embodiment covers the construction of multiple feature sets, including haploid methylation features, terminal motif features, and breakpoint motif features:
[0070] 1) Identification of haploid methylation feature sets:
[0071] This embodiment identifies multiple haploid states, including MM, MHL, CHALM, PDR, and Entropy. Each state reflects the degree of methylation within the CpG island region using a specific algorithm. By comparing the haploid methylation levels of cancerous tissue and adjacent normal tissue samples, significant differences in the CpG island region are revealed. These differences can serve as a set of methylation features for distinguishing cancer.
[0072] 2) Analysis of terminal motif characteristics:
[0073] In tumor tissue, tumor-specific DNA fragments are generated due to alterations in chromatin state and abnormal nuclease activity. Differences in the breakpoint locations of these fragments lead to variations in the proportion of terminal motif sequences. By extracting the six-base sequence from the 3' end of each read in cfDNA high-throughput sequencing data, we can calculate the distribution characteristics of the terminal motifs. Considering the permutations of the four bases (A, T, C, G), this embodiment can identify up to 4096 different terminal motif distributions.
[0074] 3) Breakpoint motif analysis:
[0075] Similar to terminal motif features, this feature focuses on the sequence adjacent to the break site, specifically the sequence of three base pairs upstream and downstream of the 3' end of each read. This sequence selection reveals a different feature distribution than that of the terminal motif. Considering the permutations of the four bases (A, T, C, G), this embodiment can identify up to 4096 different break site motif distributions.
[0076] Furthermore, based on feature extraction and the training set samples, this embodiment further constructs an ensemble machine learning classification model. This model consists of two layers: the first layer integrates four different machine learning algorithms, each trained independently on the training set samples to identify the probability of cancer. The second layer integrates the output of the first layer through a logistic regression model to generate the final classification decision. The algorithms involved and their brief descriptions are as follows:
[0077] 1) Logistic Regression Model:
[0078] Logistic regression is a widely used supervised learning model suitable for handling classification problems. It is known for its computational efficiency and intuitiveness by establishing a cost function and iteratively solving for the optimal model parameters using optimization methods. This method is suitable for binary classification problems, is fast, and easy to understand.
[0079] 2) Support Vector Machine:
[0080] Support Vector Machines (SVMs) are powerful supervised learning models used for classification and regression analysis. They maximize the margin between different classes by mapping data to a high-dimensional space and finding an optimal hyperplane. When training the model, the following parameters were set: kernel="linear", scale=T, probability=TRUE, type="C-classification".
[0081] 3) Random Forest:
[0082] Random forest is an ensemble learning method that improves prediction accuracy by constructing multiple decision trees and voting on or averaging their predictions. When training the model, the following parameters were set: method="rf", prox=TRUE, ntree=500, metric="Accuracy".
[0083] 4) XGBoost:
[0084] XGBoost is a high-efficiency gradient boosting decision tree framework that improves the performance of base learners through algorithm optimization, making it particularly suitable for handling large-scale datasets. When training the model, the following parameters were set: nrounds = 1000, params = list(booster = "gbtree", objective = "binary:logistic", eval_metric = "logloss")
[0085] 5) CatBoost:
[0086] CatBoost is an advanced gradient boosting decision tree algorithm specifically designed to handle categorical features and reduce model bias. It improves prediction accuracy through iterative optimization, making it particularly suitable for processing categorical features. When training the model, the following parameters are set: params = list(loss_function = 'Logloss', iterations = 1000).
[0087] Specifically, the process of establishing the colorectal cancer early screening detection model is as follows: In this embodiment, plasma samples are randomly divided into two parts, a training set and a validation set, in a 1:1 ratio. After quality control, 1863 samples are randomly divided into a training set (931 samples) and a test set (932 samples). The training set is used for model training, and the test set is used for model evaluation. The training set is used for model construction, while the validation set is used to evaluate the model's performance. This embodiment accurately identifies CpG island regions with specific methylation patterns through in-depth analysis of tumor tissue and corresponding normal tissue samples. These patterns include MM, MHL, CHALM, PDR, and Entropy haploid states. Furthermore, this embodiment combines haploid methylation characteristics in blood, motif characteristics at sequence ends, and motif characteristics near breakpoints to form a comprehensive set of biomarkers.
[0088] Furthermore, in the feature selection stage, this embodiment implements a p-value selection strategy based on Benjaminiand-Hochberg multiple test correction, using a p-value less than 0.01 as the threshold to ensure the statistical significance of the selection results. The features selected in this embodiment—MM, MHL, Entropy, CHALM, PDR, BPM_raw_4bp, and mer6_raw_4bp—are as follows:
[0089] The characteristics of MM, a total of 251 CpG islands, are shown in Table 2:
[0090] Table 2
[0091]
[0092]
[0093]
[0094]
[0095] The characteristics of MHL, totaling 255 CpG islands, are shown in Table 3:
[0096] Table 3
[0097]
[0098]
[0099]
[0100]
[0101] Entropy is characterized by a total of 182 CpG islands, as shown in Table 4:
[0102] Table 4
[0103]
[0104]
[0105]
[0106]
[0107] CHALM's characteristics, consisting of a total of 214 CpG islands, are shown in Table 5:
[0108] Table 5
[0109]
[0110]
[0111]
[0112] The characteristics of PDR, totaling 146 CpG islands, are shown in Table 6:
[0113] Table 6
[0114]
[0115]
[0116]
[0117] The BPM_raw_4bp feature contains a total of 321 motifs, as shown in Table 7:
[0118] Table 7
[0119]
[0120]
[0121]
[0122] The mer6_raw_4bp feature contains a total of 271 motifs, as shown in Table 8:
[0123] Table 8
[0124]
[0125]
[0126] Specifically, after feature selection and integration, this embodiment constructs an early colorectal cancer screening model containing seven features, which integrate the aforementioned five haploid states, terminal motifs, and breakpoint motifs. Using data from the training set, this embodiment inputs the samples into four different machine learning algorithms, each using the probability of cancer occurrence as its prediction output, forming 28 basic models that constitute the first layer of the model. Subsequently, the prediction results of these basic models are further input into the second-layer logistic regression model to integrate the predictions of each model and optimize the final screening results. This method used in this embodiment is called stack ensemble learning, which optimizes the combination of prediction results from different basic learners through meta-learning algorithms. The advantage of this method is that it can fully leverage the synergistic effect of multiple high-performance models to achieve more accurate classification or regression predictions.
[0127] Furthermore, model evaluation. Referring to 9, the model exhibits high specificity (training set: 97.50%, test set: 97.35%), sensitivity (training set: 91.72%, test set: 93.10%), and accuracy (training set: 95.7%, test set: 96.03%) on both the training and test sets, indicating that the model has good generalization ability and low overfitting risk.
[0128] Table 9
[0129]
[0130] Referring to Table 10, the sensitivity of the model at different stages is as follows: the sensitivity is relatively low in stages I and II, and relatively high in stages III and IV. However, even in the early stage of cancer (stage I), the sensitivity remains high (91%).
[0131] Table 10
[0132] training set 91.42% 91.50% 91.36% 100.00% test set 90.48% 94.29% 92.42% 100.00%
[0133] The beneficial effects of this invention are as follows:
[0134] This invention establishes a non-invasive and cost-effective method for detecting colorectal cancer by constructing an early screening model for colorectal tumors. This method reduces the economic cost of sequencing and significantly improves the accuracy and sensitivity of detection.
[0135] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0136] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for early screening of colorectal cancer based on multi-omics sequencing, characterized in that, include: Collect plasma samples for testing; DNA methylation sequencing was performed on the plasma sample to be tested to obtain the sequencing data. The sequencing data to be tested is input into the trained early screening model for colorectal tumors for classification to obtain the target screening results; The training process of the colorectal cancer early screening model includes: Colorectal cancer tissue and adjacent normal tissue samples were collected from colorectal cancer patients. Genomic DNA was extracted and DNA methylation sequencing was performed on the colorectal cancer tissue and adjacent normal tissue samples to obtain raw sequencing data. The raw sequencing data were sequentially subjected to quality control and aligned with the human reference genome to obtain sequence correspondences; The human reference genome was divided into several CpG island regions, and the methylation degree of different haplotypes in each CpG island region was analyzed to obtain reference methylation data; Based on the sequence correspondence, the methylation degree difference analysis was performed using the p-values after Benjamini and Hochberg multiple test correction of the reference methylation data and the original sequencing data to obtain tumor-specific regions. Plasma samples were collected from colorectal cancer patients and healthy individuals, and cfDNA was extracted and methylated from the plasma samples to obtain plasma cfDNA sequencing data. The plasma cfDNA sequencing data were sequentially subjected to quality control and alignment with the human reference genome to obtain coordinate-determined sequencing data. The coordinates are statistically analyzed to determine the degree of methylation of sequencing data in the tumor-specific region, thus obtaining a set of methylation features. The coordinates are obtained to determine the coordinates of all reads of sequencing data on the reference genome. The p bases at the 3' end of the read data on the human reference genome are taken as the first base fragment, and the proportion of the first base fragment is calculated to obtain the terminal motif characteristics. The 3' end of the read data is taken as the second base fragment by q bases upstream and q bases downstream of the human reference genome, and the proportion of the second base fragment is calculated to obtain the breakpoint motif feature; The methylation feature set, the terminal motif feature, and the breakpoint motif feature are combined into an initial feature value; Using the initial feature values as input data and the disease status as output, a pre-designed integrated machine learning classifier is trained to obtain the early screening model for colorectal tumors.
2. The method for early screening of colorectal cancer based on multi-omics sequencing according to claim 1, characterized in that, The haploids include any one or more of MM, MHL, CHALM, PDR, and Entropy.
3. The method for early screening of colorectal cancer based on multi-omics sequencing according to claim 1, characterized in that, p is any integer between 4 and 10; q is any integer between 2 and 5.
4. The method for early screening of colorectal cancer based on multi-omics sequencing according to claim 1, characterized in that, The ensemble machine learning classifier includes: a first layer and a second layer; the first layer includes: support vector machine, random forest, XGBoost and CatBoost; the second layer includes: logistic regression network.
Citation Information
Patent Citations
LP-WGS and DNA methylation-based lung cancer early screening model construction method and electronic equipment
CN117275585A
Multi-cancer-species detection and traceability system based on DNA high-throughput methylation sequencing
CN118298916A