Colorectal cancer early screening marker, detection method, detection device and computer readable medium

By performing high-resolution analysis of plasma sample cfDNA and integrating multiple algorithms, the non-invasive and economic problems of early screening for colorectal cancer are solved, and efficient early diagnosis of colorectal cancer and advanced intestinal adenoma detection are achieved.

CN113903398BActive Publication Date: 2025-06-06GENESEEQ TECH INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111053742.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-08
Publication Date
2025-06-06
Estimated Expiration
2041-09-08

AI Technical Summary

Technical Problem

There is a lack of non-invasive, economical and effective premature screening and testing methods for colorectal cancer suitable for a wide range of populations in the prior art. In particular, non-invasive testing methods are difficult to replace invasive colonoscopy, and the cost of fecal occult blood test is high and has not been included in medical insurance.

Method used

By WGS sequencing the plasma sample cfDNA, using high-resolution length distribution, sequence read proportion at the breakpoint of the 5' end of the read segment and copy number change analysis, combined with gradient lifter, random forest and deep network learning, a multi-feature multi-algorithm integrated model is constructed to achieve non-invasive and accurate diagnosis of colorectal cancer.

Benefits of technology

It provides high sensitivity and specific non-invasive colorectal cancer early screening, which can diagnose early bowel cancer and advanced bowel adenoma, has low detection throughput and is suitable for large-scale promotion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113903398B_ABST
    Figure CN113903398B_ABST
Patent Text Reader

Abstract

The present invention provides a marker for early screening of colorectal cancer, a detection method, a detection device and a computer-readable medium. The present invention performs high-resolution length distribution of differential DNA fragments between colorectal cancer and healthy people on the results of high-throughput sequencing, analyzes the read proportion of sequence reads at the 5-end breakpoint of the read segment and the copy number change of 1MB window, uses gradient boosting machine, random forest and deep network learning to perform training and modeling respectively, and finally constructs a multi-feature and multi-algorithm integrated model through a generalized linear model, thereby achieving the purpose of non-invasive and accurate diagnosis of colorectal cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an early screening method for colorectal cancer (CRC), belonging to the technical field of molecular biomedicine. Background Art

[0002] Colorectal cancer is a common malignant tumor. According to the White Paper on Colorectal Cancer and Precancerous Lesions in Chinese Physical Examination Population, the five-year survival rate of colorectal cancer in the Chinese population is 90.1%, 72.6%, 53.8% and 10.4% in stages I, II, III and IV respectively; the five-year survival rate of patients with local metastatic cancer was 89% between 2009 and 2015, much higher than the 21% of patients with distant metastatic cancer. Early detection and diagnosis of tumors are crucial to improving the survival rate of patients with colorectal cancer, and can also reduce the burden of medical expenses.

[0003] Colorectal endoscopy is the "gold standard" for the diagnosis of colorectal cancer, with a detection rate of up to 95%. However, colonoscopy is an invasive screening procedure, which is painful, requires a high physical condition of the patient, and involves surgical risks and complication risks, so patient compliance is low. At the same time, limitations such as insufficient colonoscopy resources in my country have led to a low screening penetration rate. In the short term, colonoscopy as a screening method cannot be promoted on a large scale, and there is an urgent need for support and supplementation from non-invasive detection methods.

[0004] The main non-invasive examination methods in my country are fecal occult blood tests, including guaiac fecal occult blood test (gFOBT) and immunochemical fecal occult blood test (FIT). They are highly sensitive and easy to collect and preserve specimens, but their prices are generally high and they are not included in medical insurance. Therefore, my country urgently needs to develop an effective, economical and practical screening method suitable for a wide range of people. Summary of the invention

[0005] The present invention provides a method for performing WGS sequencing on cfDNA of a plasma sample, performing high resolution fragmentation size distribution analysis on the high-throughput sequencing results for differential DNA fragments between colorectal cancer and healthy subjects, analyzing the motif break point 8mer ratio and 1Mb-bin copy number variant analysis, using a gradient boosting machine (GBM), random forest and deep learning to perform training and modeling respectively, and finally constructing a multi-feature and multi-algorithm integrated model through a generalized linear model (GLM), thereby achieving the purpose of non-invasive and accurate diagnosis of colorectal cancer.

[0006] A method for constructing a colorectal cancer early screening model comprises the following steps:

[0007] Step 1: extract and sequence cfDNA from samples of the positive group and the control group to obtain read data;

[0008] Step 2, aligning the read data results to the reference genome, and obtaining the number of reads in different length intervals within different window ranges on the reference genome as the first feature value;

[0009] Step 3, aligning the read data result to the reference genome to obtain the position of the 5' end of the read segment on the reference genome; obtaining the sequence data of m bp bases upstream and downstream of the position as a base fragment set; taking the proportion of each obtained base fragment in all fragments as the second characteristic value;

[0010] Step 4, dividing the reference genome into multiple windows, and obtaining the copy number data within each window as the third eigenvalue;

[0011] Step 5, taking the first, second and third eigenvalues ​​together as initial eigenvalues, screening out eigenvalues ​​with significant differences between the samples in the positive group and the control group from the initial eigenvalues ​​as model eigenvectors;

[0012] Step 6, input the model feature vectors of the samples in the positive group and the control group into the model, and use the probability of colorectal cancer as the model output value to train the model to obtain an early screening model.

[0013] The step 2 includes:

[0014] Step 2-1, divide the reference genome into multiple windows, and obtain the total number of reads, the number of short reads, and the number of ultra-long reads within each window respectively;

[0015] Step 2-2, taking the long arm and short arm of each chromosome as the region range, and obtaining the number of reads in different length gradient intervals within each range;

[0016] Step 2-3, taking the data obtained in steps 2-1 and 2-2 together as the first eigenvalue.

[0017] The short read segment refers to a length of 40-80bp, the ultra-long read segment is 200-300bp; all read segments refer to a length in the range of 40-300bp.

[0018] The size range of the window in step 2-1 is 2-7Mb.

[0019] The different length gradient intervals in step 2-2 refer to different length gradient ranges obtained by increasing the step length of 8-12 bp within the range of 40-300 bp.

[0020] The read numbers stated are normalized.

[0021] The m is any integer between 2 and 5.

[0022] In the step 5, the steps include: using the first, second and third eigenvalues ​​as input values ​​of the gradient boosting algorithm model, the random forest model and the deep network learning model respectively, training the samples with whether the patient has colon cancer as the output value, and obtaining eigenvectors with significant differences in each model.

[0023] In step 6, the steps include: using the characteristic vector of significant difference as the input value of the classifier model, using the probability of having colorectal cancer as the output value, and training the model with sample data from the positive group and the control group to obtain an early screening model.

[0024] The classifier model is a linear model, and the variables contained in the model are obtained by inputting the eigenvectors with significant differences among the first, second and third eigenvalues ​​into the gradient boosting algorithm model, random forest model and deep network learning model trained in step 5 respectively.

[0025] A colorectal cancer early screening detection device, comprising:

[0026] The sequencing module is used to extract and sequence cfDNA from samples in the positive group and the control group to obtain read data; the comparison module is used to compare the read data results to the reference genome;

[0027] A first feature value acquisition module is used to obtain the number of read segments in different length intervals within different window ranges on the reference genome as a first feature value;

[0028] The second feature value acquisition module is used to obtain the position of the 5' end of the read segment on the reference genome; obtain the sequence data of m bp bases upstream and downstream of the position as a base fragment set; and use the proportion of each obtained base fragment in all fragments as the second feature value;

[0029] A third eigenvalue acquisition module is used to divide the reference genome into multiple windows, and obtain copy number data within each window range as a third eigenvalue;

[0030] A screening module is used to use the first, second and third eigenvalues ​​as initial eigenvalues, and screen out eigenvalues ​​with significant differences between samples in the positive group and the control group from the initial eigenvalues ​​as model eigenvectors;

[0031] The model building module inputs the model feature vectors of the samples in the positive group and the control group into the model, and uses the probability of colorectal cancer as the model output value to train the model to obtain an early screening model.

[0032] The first feature value acquisition module includes:

[0033] A first read number statistics module is used to divide the reference genome into multiple windows, and obtain the total number of reads, the number of short reads, and the number of ultra-long reads within each window;

[0034] The second read number counting module is used to use the long arm and the short arm of each chromosome as the region range, and obtain the number of reads in different length gradient intervals within each range;

[0035] The merging module is used to use the data obtained from the first read number counting module and the second read number counting module as the first feature value.

[0036] A computer-readable medium having a computer program recorded thereon that can execute the method for constructing the above-mentioned colorectal cancer early screening model.

[0037] The above model can also use its sub-models separately:

[0038] A method for constructing a colorectal cancer early screening model comprises the following steps:

[0039] Step 1: extract and sequence cfDNA from samples of the positive group and the control group to obtain read data;

[0040] Step 2, aligning the read data results to the reference genome, obtaining the number of reads in different length intervals within different window ranges on the reference genome as the initial feature value;

[0041] Step 3, select the eigenvalues ​​with significant differences between the samples in the positive group and the control group from the initial eigenvalues ​​as the model eigenvectors;

[0042] Step 4: Input the model feature vectors of the samples in the positive group and the control group into the model, and use the probability of colorectal cancer as the model output value to train the model to obtain an early screening model.

[0043] The step 2 includes:

[0044] Step 2-1, divide the reference genome into multiple windows, and obtain the total number of reads, the number of short reads, and the number of ultra-long reads within each window respectively;

[0045] Step 2-2, taking the long arm and short arm of each chromosome as the region range, and obtaining the number of reads in different length gradient intervals within each range;

[0046] Step 2-3, the data obtained in steps 2-1 and 2-2 are taken together as initial eigenvalues.

[0047] The step 3 comprises: taking the initial eigenvalue as the input value of the model, classifying the samples based on whether they have colorectal cancer as the output value, sorting the samples according to the contribution value of each eigenvector, and obtaining the eigenvectors with significant differences in each model;

[0048] The model is selected from a gradient boosting algorithm model, a random forest model or a deep network learning model.

[0049] A method for constructing a colorectal cancer early screening model comprises the following steps:

[0050] Step 1: extract and sequence cfDNA from samples of the positive group and the control group to obtain read data;

[0051] Step 2, aligning the read data results to the reference genome to obtain the position of the 5' end of the read segment on the reference genome; obtaining the sequence data of m bp bases upstream and downstream of the position as a base fragment set; taking the proportion of each obtained base fragment in all fragments as the initial feature value;

[0052] Step 3, select the eigenvalues ​​with significant differences between the samples in the positive group and the control group from the initial eigenvalues ​​as the model eigenvectors;

[0053] Step 4: Input the model feature vectors of the samples in the positive group and the control group into the model, and use the probability of colorectal cancer as the model output value to train the model to obtain an early screening model.

[0054] The step 3 comprises: taking the initial eigenvalue as the input value of the model, classifying the samples based on whether they have colorectal cancer as the output value, sorting the samples according to the contribution value of each eigenvector, and obtaining the eigenvectors with significant differences in each model;

[0055] The model is selected from a gradient boosting algorithm model, a random forest model or a deep network learning model.

[0056] A method for constructing a colorectal cancer early screening model comprises the following steps:

[0057] A method for constructing a colorectal cancer early screening model comprises the following steps:

[0058] Step 1: extract and sequence cfDNA from samples of the positive group and the control group to obtain read data;

[0059] Step 2, dividing the reference genome into multiple windows, and obtaining the copy number data within each window as the initial feature value;

[0060] Step 3, select the eigenvalues ​​with significant differences between the samples in the positive group and the control group from the initial eigenvalues ​​as the model eigenvectors;

[0061] Step 4: Input the model feature vectors of the samples in the positive group and the control group into the model, and use the probability of colorectal cancer as the model output value to train the model to obtain an early screening model.

[0062] The step 3 comprises: taking the initial eigenvalue as the input value of the model, classifying the samples based on whether they have colorectal cancer as the output value, sorting the samples according to the contribution value of each eigenvector, and obtaining the eigenvectors with significant differences in each model;

[0063] The model is selected from a gradient boosting algorithm model, a random forest model or a deep network learning model.

[0064] Beneficial Effects

[0065] The WGScfDNA read length distribution, breakpoint sequence proportion and regional copy number changes of 115 healthy people and 195 patients with colorectal cancer / advanced intestinal adenoma were statistically analyzed, and three different training learning algorithms were used to build models, and all models were trained in a secondary set to improve the model's predictive performance for healthy and cancer groups. This invention provides a multi-molecular feature multi-training algorithm secondary integrated diagnostic model based on high-throughput and low-depth sequencing of plasma cfDNA for the first time. This model can diagnose not only early colorectal cancer but also advanced intestinal adenoma, and has the advantages of non-invasive detection, low throughput, high detection specificity and sensitivity. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a schematic diagram of the model building process;

[0067] Figure 2 It is a schematic diagram of the construction process of the quadratic set model;

[0068] Figure 3 It is a heat map showing the difference in the distribution characteristics of the length proportions of the top 50 high-resolution DNA fragments between the cancer group and the healthy group;

[0069] Figure 4 This is the heat map of the difference in the sequence proportion characteristics of the 5' end breakpoints of the top 50 reads between the cancer group and the healthy group

[0070] Figure 5 It is the difference heat map of the copy number change characteristics of the first 501Mb window between the cancer group and the healthy group;

[0071] Figure 6 It is the prediction AUC curve of the classifier with different training algorithms based on the distribution characteristics of high-resolution DNA fragment length ratio on the validation set and the test set;

[0072] Figure 7 It is the predicted AUC curve of the classifier with different training algorithms based on the sequence proportion characteristics at the 5' end breakpoint of the reads on the validation set and the test set;

[0073] Figure 8 It is the prediction AUC curve of the classifiers with different training algorithms for copy number variation features on the validation set and the test set;

[0074] Fig. 9 It is the AUC curve predicted by the classifier trained by the second set of different features of the classifier on the validation set;

[0075] Fig.10 It is a graph of the prediction results of the classifier trained with a secondary set of different features of the classifier on the validation set;

[0076] Fig.11 It is the AUC curve of different groups of predictions of the classifier after all models are combined on the validation set and the test set;

[0077] Fig.12 It is the AUC curve of different groups of predictions of the classifier after all models are combined on the validation set and the test set;

[0078] Fig.13 It is a graph of the prediction results of the classifier after all models are combined on the validation set and the test set;

[0079] Fig.14 It is the prediction result graph of the classifier after all models are combined on the validation set;

[0080] Fig.15 It is the prediction result graph of the classifier after all models are combined on the test set; DETAILED DESCRIPTION

[0081] The calculation method in the present invention is described in detail as follows:

[0082] The present invention first requires the steps of extracting, building a library, and sequencing cfDNA from a blood sample. The extraction and library building methods here are not particularly limited and can be adjusted from the extraction methods in the prior art. In the sequencing process here, the sequencing technology in the prior art can be used to obtain the base information of cfDNA.

[0083] The data sets used in the model building process of the present invention are as follows:

[0084]

[0085] Methods for extraction and sequencing of plasma cfDNA samples

[0086] A purple blood collection tube (EDTA anticoagulant tube) was used to collect 8 ml of whole blood samples from patients, and the plasma was centrifuged promptly (within 2 hours). After being transported to the laboratory, the plasma samples were extracted for ctDNA using the QIAGEN plasma DNA extraction kit according to the instructions. After the collected cfDNA samples were library-built, WGS-2 multiplication sequencing was performed. After obtaining the offline data, the data was aligned to the human reference genome to obtain the base data information of the corresponding read segments.

[0087] Data processing

[0088] The marker data in the present invention mainly utilizes three molecular features:

[0089] 1. High Resolution Fragmentation Size Distribution (HRFSD) For DNA fragment size distribution, it reflects the distribution characteristics of the length of cfDNA reads. High resolution DNA fragment size distribution (HRFSD) is used for machine learning to establish a prediction model to distinguish non-liver cancer patients (healthy people) from colorectal cancer patients (advanced intestinal adenoma CRA and colorectal cancer CRC). For DNA fragment size distribution, it reflects the distribution characteristics of the length of cfDNA reads. By comparing the lengths of cfDNA reads from 115 healthy people and 195 patients with colorectal cancer / advanced intestinal adenoma, it was found that the number of fragments between 40-80bp and 200-300bp was different between the two groups, which can be used as a distinguishing feature.

[0090] The cfDNA read length data was obtained by the following method: in the aligned bam, the quality, length and alignment position information of each read were recorded. The human reference genome was selected from the hg19 sequence provided by the University of California, Santa Cruz (UCSC). The human reference genome was cut into 572 windows according to the length of 5Mb, and the total number of reads (40-300bp), short reads (40-80bp) and ultra-long reads (200-300bp) in each window were counted. According to the statistical results of the number of various reads in all windows, the number of each read was standardized, that is, the standardized value = (original value-average value) / standard deviation. Thus, 572 sets of reads of different lengths were obtained.

[0091] At the same time, in order to obtain high-resolution read results, 41 regions of the long and short arms of each chromosome of the human reference genome are used as windows, as shown below:

[0092] chr1_p chr4_q chr8_p chr11_q chr16_q chr20_p chr1_q chr5_p chr8_q chr12_p chr17_p chr20_q chr2_9 chr5_q chr9_p chr12_q chr17_q chr21_q chr2_q chr6_p chr9_q chr13_q chr18_p chr22_q chr3_p chr6_q chr10_p chr14_q chr18_q chrX_p chr3_q chr7_p chr10_q chr15_q chr19_p chrX_q chr4_p chr7_q chr11_p chr16_p chr19_q

[0093] The 40-300 bp fragments were divided into 27 length gradients with an increment of 10 bp (e.g., 40-49 bp, 50-59 bp on the 1q arm of chr1...), and the number of fragments in each long and short arm window of each length gradient was counted and standardized to obtain a high-resolution DNA fragment size distribution result with a total of 2823 feature results (2823=572 all-read segment standardization results+572 short-read segment standardization results+572 extra-long single segment standardization results+41*27 length gradient standardization results).

[0094] 2. The proportion of sequence reads at the 5' end breakpoint of the read segment (Motif Breakpoint 8Mer, MTBK)

[0095] The human reference genome is a DNA double helix structure, which relies on base complementary pairing and hydrogen bond separation to link; during normal aging and cancer progression, the pH of the cell environment changes, thereby destroying the base complementary hydrogen bonds and causing breaks; due to the different base sequences at the breaks, the proportion of sequences containing information at different breakpoints will also be different. Collection method: After alignment, the basic information and aligned position of each read segment are recorded in the bam, and the 4bp sequence on the left and right of the human reference genome sequence coordinates where the 5' end of each read segment is located is confirmed. The number of read segments of each breakpoint sequence (a total of 4**8=65536 types) is counted, and the proportion of sequence read segments at 65536 breakpoints is calculated, for example, the proportion of AAAAAAAA read segments = the number of AAAAAAAA read segments / the total number of sequence read segments at all breakpoints.

[0096] 3.1Mb-Bin Copy Number Variation (CNV)

[0097] Copy number changes are highly correlated with individual cancers. Although some cancer-related genes or specific genomic intervals can be detected for copy number changes to distinguish them, there are still other rare or unknown genes or intervals that can provide potential copy number change information. Collection method: First, WGS data of 30 healthy people were collected, and the reference gene chromosomes 1-22 were divided into windows with a length of 1Mb without overlap. The read depth in each window was calculated for each sample using bedtools coverage, and corrected according to the GC content and average alignment ability record of each window (UCSC BigWig file). The median depth of 30 healthy people in each window was taken as a representative to obtain a population control baseline of 2475 window read depths; for each sample to be tested, the individual read depth information of 2475 windows was also obtained, and the hidden Markov model (HMM) and the depth of the population control baseline of each window were used to construct the logarithm of the copy number change in each window, that is, log2 (corrected normalized depth of the sample to be tested / corrected normalized depth of the population baseline), so as to obtain the copy number change information of each sample to be tested.

[0098] Through the above data acquisition, the initial data vectors of these three types of data can be obtained respectively. Next, the corresponding calculation method is designed:

[0099] The marker data in this invention mainly utilizes three single-feature machine learning algorithms:

[0100] 1. Gradient Boosting Machine (GBM)

[0101] The gradient boosting algorithm is a common algorithm in machine learning. Its basic principle is to train the newly added weak classifiers according to the negative gradient information of the current model loss function, and then combine the trained weak classifiers into the existing model in an accumulated form to obtain the optimal model. This model has the advantages of good training effect and low overfitting. To prevent GBM from overfitting or underfitting during the learning process, the GBM parameters are set as follows: ntrees = 200, max_depth = 9, learning_rate = 0.01, subample = 0.8. Crossvalidation = 10.

[0102] 2. Random Forest (RF)

[0103] Random forest is a powerful classification and regression tool. When a set of data is provided, random forest can randomly extract part of the information to generate a set of decision trees that help classification or regression, split the attributes of the nodes, and repeat the random extraction until it can no longer split; finally, combine all the split attribute results to obtain the final prediction result. To prevent RF from overfitting or underfitting during the learning process, set the RF parameters as follows: ntrees = 200, max_depth = 9, Crossvalidation = 10.

[0104] 3. Deep Learning (DL)

[0105] Deep learning is based on a multi-layer feedforward artificial neural network that is trained with stochastic gradient descent using back-propagation. The network can contain a large number of hidden layers consisting of neurons with hyperbolic tangent, rectification, and maximum power activation functions. Advanced features such as adaptive learning rate, rate annealing, momentum training, dropout, L1 or L2 regularization, checkpoints, and grid search enable high prediction accuracy. During training, each computing node trains a copy of the global model parameters on its local data using multithreading (asynchronous) and regularly contributes to the global model through model averaging on the network. The feedforward artificial neural network (ANN) model, also known as a deep neural network (DNN) or multi-layer perceptron (MLP), is the most common type of deep neural network. The main principle is to design multiple perceptrons with multiple inputs and multiple outputs to establish an appropriate number of neuron computing nodes and a multi-layer operation hierarchy, select appropriate input and output layers, and establish a functional relationship from input to output through network learning and tuning, which can be as close to the real association relationship as possible. To prevent over- or under-training during DL learning, the DL parameters are set as follows: epoch=300, hidden={100, 100, 100}, input_dropout_ratios=0.05, rho=0.95, mini_batch_size=10, Crossvalidation=10.

[0106] After obtaining the above three types of initial data information of 115 healthy people and 195 patients with colorectal cancer / advanced intestinal adenoma, the statistical results of high-resolution DNA fragment size distribution were used as input values ​​(the input vector of each sample contained 2823 characteristic values ​​composed of read segment ratios), and the test samples and normal samples were classified by three classification models respectively; similarly, after collecting the ratio of the number of read segments of the 5' end breakpoint sequence of the DNA fragments of patients and healthy people, the ratio of the sequence at the 5' end breakpoint of the DNA fragment (65536 types) was used as the input value, and the test samples and normal samples were classified by three classification models; similarly, the copy number information of 2475 windows was used as the input vector for sample classification through three classification models. Through the above calculation process, three types of data were substituted into three models for classification, and a total of 3×3=9 model calculation results were obtained. In each calculation result, the contribution value of each feature vector to the classification result can be obtained. After collecting the feature columns whose contribution values ​​of each molecular feature under different training algorithms are not 0, 1368 columns of HRFSD significant difference sets, 958 columns of MTBK significant difference sets, and 1073 columns of CNV significant difference sets were finally obtained. The contribution value of each molecular feature was ranked in the top 50 feature columns for differential analysis. As shown in the heatmap, the top 50 feature columns of each feature have differential signals in both the cancer and health groups.

[0107] For the MTBK dataset, the top 50 contribution sequences and contribution values ​​calculated under the three models are as follows:

[0108] GMB Model Data:

[0109]

[0110] RF model data:

[0111]

[0112]

[0113] DL model data:

[0114]

[0115]

[0116] In order to further improve the prediction performance of the classifier, the above 9 training model results are trained again (stacking). Stacking is an ensemble learning technology that combines multiple underlying weak classifiers (1 st -level basemodel) for meta-learning (2 nd-level meta-learning), collect the characteristics of each underlying classifier, find the optimal integration method, and thus improve the model prediction performance. The training algorithm used in this patent Stacking is the generalized linear model (GLM). The relationship between the mathematical expectation of the response variable and the linear combination of the predictor variables is established through the link function, and the 9 training models are converted into the final linear equation: ALL Stacked = Intercept + A*HRFSD_GMB + B*HRFSD_RF + C*HRFSD_DL + D*MTBK_GBM + E*MTBK_RF + F*MTBK_DL + G*CNV_GBM + H*CNV_RF + I*CNV_DL, where Intercept and AI are both linear equation parameters. HRFSD_GMB and others refer to the output values ​​(probability of disease) obtained by the model after obtaining the input data.

[0117] The specific coefficients are as follows:

[0118] name Corresponding coefficient Intercept -0.95688 A(HRFSD_GBM) 0.004297 B(HRFSD_RF) 0.139366 C(HRFSD_DL) 0.733057 D(MTBK_GBM) 0.788211 E(MTBK_RF) -0.08808 F(MTBK_DL) 0.944454 G(CNV_GBM) 0.337852 H(CNV_RF) -0.02318 I(CNV_DL) 0.612503

[0119] The model prediction performance is shown below with different features and input vectors to the training algorithm:

[0120]

[0121]

[0122] Among them, the HRFSD Stacked model refers to a linear equation model composed of the HRFSD GBM model, the HRFSD RF model, and the HRFSD DL model. The MTBK Stacked model refers to a linear equation model composed of the MTBKGBM model, the MTBKDL model, and the MTBKDL model. The CNV Stacked model refers to a linear equation model composed of the CNV GBM model, the CNV RF model, and the CNVDL model. The results are shown in the figure and the table below:

[0123]

[0124] Among them, the stacked model of HRFSD and MTBK refers to a linear model composed of three models of HRFSD and three models of MTBK; the stacked model of HRFSD and CNV refers to a linear model composed of three models of HRFSD and three models of CNV; the stacked model of MTBK and CNV refers to a linear model composed of three models of MTBK and three models of CNV;

[0125]

[0126] Each feature has a certain predictive effect under different training algorithms, and the prediction effect of a single feature in the secondary ensemble training is improved. The above 9 models were trained in the secondary ensemble, and the classifier prediction effect was the best, with the highest AUC reaching 0.988. At the same time, it was found that the ensemble model can effectively distinguish between healthy people and colorectal cancer, and healthy people and advanced intestinal adenomas, but because the molecular characteristics of colorectal cancer and advanced intestinal adenomas are similar, it is difficult to distinguish them (AUC = 0.594).

[0127] The results obtained from the above final ensemble model are shown in the following table:

[0128]

[0129]

[0130] Judging from the prediction results, the multi-feature set classifier results can correct the misjudgment of the single feature set classifier, with a sensitivity of 97.44% under a specificity of 94.83% in the validation set and test set. Sequence Listing <110> Nanjing Shihe Gene Biotechnology Co., Ltd. Nanjing Shihe Medical Equipment Co., Ltd. <120> Colorectal cancer early screening marker, detection method, detection device and computer readable medium <130> none <160> 150 <170> SIPOSequenceListing 1.0 <210> 1 <211> 8 <212> DNA <213> Artificial Sequence <400> 1 ccccattg 8 <210> 2 <211> 8 <212> DNA <213> Artificial Sequence <400> 2 cttaatag 8 <210> 3 <211> 8 <212> DNA <213> Artificial Sequence <400> 3 gtcccagt 8 <210> 4 <211> 8 <212> DNA <213> Artificial Sequence <400> 4 accccgtg 8 <210> 5 <211> 8 <212> DNA <213> Artificial Sequence <400> 5 ccgatttg 8 <210> 6 <211> 8 <212> DNA <213> Artificial Sequence <400> 6 tgcggtgc 8 <210> 7 <211> 8 <212> DNA <213> Artificial Sequence <400> 7 tacggtga 8 <210> 8 <211> 8 <212> DNA <213> Artificial Sequence <400> 8 gcgggttg 8 <210> 9 <211> 8 <212> DNA <213> Artificial Sequence <400> 9 atcgcgtg 8 <210> 10 <211> 8 <212> DNA <213> Artificial Sequence <400> 10 gcgattcg 8 <210> 11 <211> 8 <212> DNA <213> Artificial Sequence <400> 11 tgaaaccg 8 <210> 12 <211> 8 <212> DNA <213> Artificial Sequence <400> 12 cccattca 8 <210> 13 <211> 8 <212> DNA <213> Artificial Sequence <400> 13 gttcgttt 8 <210> 14 <211> 8 <212> DNA <213> Artificial Sequence <400> 14 ccctgtgt 8 <210> 15 <211> 8 <212> DNA <213> Artificial Sequence <400> 15 gccgatcc 8 <210> 16 <211> 8 <212> DNA <213> Artificial Sequence <400> 16 gcacagtt 8 <210> 17 <211> 8 <212> DNA <213> Artificial Sequence <400> 17 atagtgcg 8 <210> 18 <211> 8 <212> DNA <213> Artificial Sequence <400> 18 cccagtac 8 <210> 19 <211> 8 <212> DNA <213> Artificial Sequence <400> 19 gcccaatg 8 <210> 20 <211> 8 <212> DNA <213> Artificial Sequence <400> 20 gggtttca 8 <210> twenty one <211> 8 <212> DNA <213> Artificial Sequence <400> twenty one ccctcgaa8 <210> twenty two <211> 8 <212> DNA <213> Artificial Sequence <400> twenty two gcctagtc 8 <210> twenty three <211> 8 <212> DNA <213> Artificial Sequence <400> twenty three gattctca 8 <210> twenty four <211> 8 <212> DNA <213> Artificial Sequence <400> twenty four cggccgta 8 <210> 25 <211> 8 <212> DNA <213> Artificial Sequence <400> 25 aattcgct 8 <210> 26 <211> 8 <212> DNA <213> Artificial Sequence <400> 26 gaatggat 8 <210> 27 <211> 8 <212> DNA <213> Artificial Sequence <400> 27 acagtgtt 8 <210> 28 <211> 8 <212> DNA <213> Artificial Sequence <400> 28 tctcacgt 8 <210> 29 <211> 8 <212> DNA <213> Artificial Sequence <400> 29 cttggaaa 8 <210> 30 <211> 8 <212> DNA <213> Artificial Sequence <400> 30 atcacgct 8 <210> 31 <211> 8 <212> DNA <213> Artificial Sequence <400> 31 aacttcgg 8 <210> 32 <211> 8 <212> DNA <213> Artificial Sequence <400> 32 ctttcgtg 8 <210> 33 <211> 8 <212> DNA <213> Artificial Sequence <400> 33 attaatgt 8 <210> 34 <211> 8 <212> DNA <213> Artificial Sequence <400> 34 gctgatct 8 <210> 35 <211> 8 <212> DNA <213> Artificial Sequence <400> 35 gtaggacc 8 <210> 36 <211> 8 <212> DNA <213> Artificial Sequence <400> 36 cggtacgc 8 <210> 37 <211> 8 <212> DNA <213> Artificial Sequence <400> 37 tcaattcg 8 <210> 38 <211> 8 <212> DNA <213> Artificial Sequence <400> 38 ccgccgta 8 <210> 39 <211> 8 <212> DNA <213> Artificial Sequence <400> 39 catagaaa 8 <210> 40 <211> 8 <212> DNA <213> Artificial Sequence <400> 40 gcgtacaa 8 <210> 41 <211> 8 <212> DNA <213> Artificial Sequence <400> 41 aggcataa 8 <210> 42 <211> 8 <212> DNA <213> Artificial Sequence <400> 42 gcagcgaa 8 <210> 43 <211> 8 <212> DNA <213> Artificial Sequence <400> 43 caagcgta 8 <210> 44 <211> 8 <212> DNA <213> Artificial Sequence <400> 44 cacgacgc 8 <210> 45 <211> 8 <212> DNA <213> Artificial Sequence <400> 45 acaagaag 8 <210> 46 <211> 8 <212> DNA <213> Artificial Sequence <400> 46 acccggct 8 <210> 47 <211> 8 <212> DNA <213> Artificial Sequence <400> 47 ttgtatac 8 <210> 48 <211> 8 <212> DNA <213> Artificial Sequence <400> 48 gcgcgaaa 8 <210> 49 <211> 8 <212> DNA <213> Artificial Sequence <400> 49 tatagccg 8 <210> 50 <211> 8 <212> DNA <213> Artificial Sequence <400> 50 tcacaccc 8 <210> 51 <211> 8 <212> DNA <213> Artificial Sequence <400> 51 atgaattc 8 <210> 52 <211> 8 <212> DNA <213> Artificial Sequence <400> 52 agtactag 8 <210> 53 <211> 8 <212> DNA <213> Artificial Sequence <400> 53 cattctct 8 <210> 54 <211> 8 <212> DNA <213> Artificial Sequence <400> 54 agctgaac 8 <210> 55 <211> 8 <212> DNA <213> Artificial Sequence <400> 55 gagactcc 8 <210> 56 <211> 8 <212> DNA <213> Artificial Sequence <400> 56 cgcggtgt 8 <210> 57 <211> 8 <212> DNA <213> Artificial Sequence <400> 57 cut 8 <210> 58 <211> 8 <212> DNA <213> Artificial Sequence <400> 58 gttaatga 8 <210> 59 <211> 8 <212> DNA <213> Artificial Sequence <400> 59 tctaatga 8 <210> 60 <211> 8 <212> DNA <213> Artificial Sequence <400> 60 8 <210> 61 <211> 8 <212> DNA <213> Artificial Sequence <400> 61 cgcagcag 8 <210> 62 <211> 8 <212> DNA <213> Artificial Sequence <400> 62 tataatcg 8 <210> 63 <211> 8 <212> DNA <213> Artificial Sequence <400> 63 8 <210> 64 <211> 8 <212> DNA <213> Artificial Sequence <400> 64 gcccatta 8 <210> 65 <211> 8 <212> DNA <213> Artificial Sequence <400> 65 atttgtaa 8 <210> 66 <211> 8 <212> DNA <213> Artificial Sequence <400> 66 catttagg 8 <210> 67 <211> 8 <212> DNA <213> Artificial Sequence <400> 67 aacagcac 8 <210> 68 <211> 8 <212> DNA <213> Artificial Sequence <400> 68 ttcccagc 8 <210> 69 <211> 8 <212> DNA <213> Artificial Sequence <400> 69 atgaatac 8 <210> 70 <211> 8 <212> DNA <213> Artificial Sequence <400> 70 tacttccg 8 <210> 71 <211> 8 <212> DNA <213> Artificial Sequence <400> 71 accactgc 8 <210> 72 <211> 8 <212> DNA <213> Artificial Sequence <400> 72 agaagcag 8 <210> 73 <211> 8 <212> DNA <213> Artificial Sequence <400> 73 atcggcag 8 <210> 74 <211> 8 <212> DNA <213> Artificial Sequence <400> 74 cgtactca 8 <210> 75 <211> 8 <212> DNA <213> Artificial Sequence <400> 75 gcctgcac 8 <210> 76 <211> 8 <212> DNA <213> Artificial Sequence <400> 76 agtgctct 8 <210> 77 <211> 8 <212> DNA <213> Artificial Sequence <400> 77 cccactac 8 <210> 78 <211> 8 <212> DNA <213> Artificial Sequence <400> 78 tctgatct 8 <210> 79 <211> 8 <212> DNA <213> Artificial Sequence <400> 79 ggagcgta 8 <210> 80 <211> 8 <212> DNA <213> Artificial Sequence <400> 80 gaggcgtc 8 <210> 81 <211> 8 <212> DNA <213> Artificial Sequence <400> 81 ttgagcaa 8 <210> 82 <211> 8 <212> DNA <213> Artificial Sequence <400> 82 cataatgt 8 <210> 83 <211> 8 <212> DNA <213> Artificial Sequence <400> 83 cccagcac 8 <210> 84 <211> 8 <212> DNA <213> Artificial Sequence <400> 84 ttgggcag 8 <210> 85 <211> 8 <212> DNA <213> Artificial Sequence <400> 85 aaaagccg 8 <210> 86 <211> 8 <212> DNA <213> Artificial Sequence <400> 86 aacggtgc 8 <210> 87 <211> 8 <212> DNA <213> Artificial Sequence <400> 87 cggaatct 8 <210> 88 <211> 8 <212> DNA <213> Artificial Sequence <400> 88 ttggcgta 8 <210> 89 <211> 8 <212> DNA <213> Artificial Sequence <400> 89 gcttatgg 8 <210> 90 <211> 8 <212> DNA <213> Artificial Sequence <400> 90 gggtcaga 8 <210> 91 <211> 8 <212> DNA <213> Artificial Sequence <400> 91 ggcaatga 8 <210> 92 <211> 8 <212> DNA <213> Artificial Sequence <400> 92 cccccgta 8 <210> 93 <211> 8 <212> DNA <213> Artificial Sequence <400> 93 tgccccgtg 8 <210> 94 <211> 8 <212> DNA <213> Artificial Sequence <400> 94 ataagtat 8 <210> 95 <211> 8 <212> DNA <213> Artificial Sequence <400> 95 ttcagcac 8 <210> 96 <211> 8 <212> DNA <213> Artificial Sequence <400> 96 tccggcaa 8 <210> 97 <211> 8 <212> DNA <213> Artificial Sequence <400> 97 cattgcag 8 <210> 98 <211> 8 <212> DNA <213> Artificial Sequence <400> 98 tagagcac 8 <210> 99 <211> 8 <212> DNA <213> Artificial Sequence <400> 99 tctagtaa 8 <210> 100 <211> 8 <212> DNA <213> Artificial Sequence <400> 100 acaaattc 8 <210> 101 <211> 8 <212> DNA <213> Artificial Sequence <400> 101 cacggtga 8 <210> 102 <211> 8 <212> DNA <213> Artificial Sequence <400> 102 tcggacgt 8 <210> 103 <211> 8 <212> DNA <213> Artificial Sequence <400> 103 ttcggtgt 8 <210> 104 <211> 8 <212> DNA <213> Artificial Sequence <400> 104 tttcgtgg 8 <210> 105 <211> 8 <212> DNA <213> Artificial Sequence <400> 105 attcgttc 8 <210> 106 <211> 8 <212> DNA <213> Artificial Sequence <400> 106 acgcacca 8 <210> 107 <211> 8 <212> DNA <213> Artificial Sequence <400> 107 ccccgtat 8 <210> 108 <211> 8 <212> DNA <213> Artificial Sequence <400> 108 agcggtgc 8 <210> 109 <211> 8 <212> DNA <213> Artificial Sequence <400> 109 ggcggtac 8 <210> 110 <211> 8 <212> DNA <213> Artificial Sequence <400> 110 ttcaacgc 8 <210> 111 <211> 8 <212> DNA <213> Artificial Sequence <400> 111 gccggtcg 8 <210> 112 <211> 8 <212> DNA <213> Artificial Sequence <400> 112 ACTCGACC 8 <210> 113 <211> 8 <212> DNA <213> Artificial Sequence <400> 113 ctcacgca 8 <210> 114 <211> 8 <212> DNA <213> Artificial Sequence <400> 114 cctagtaa 8 <210> 115 <211> 8 <212> DNA <213> Artificial Sequence <400> 115 atggatcg 8 <210> 116 <211> 8 <212> DNA <213> Artificial Sequence <400> 116 ccgaatcc 8 <210> 117 <211> 8 <212> DNA <213> Artificial Sequence <400> 117 cggaacga 8 <210> 118 <211> 8 <212> DNA <213> Artificial Sequence <400> 118 tccgttct 8 <210> 119 <211> 8 <212> DNA <213> Artificial Sequence <400> 119 aggtacgg 8 <210> 120 <211> 8 <212> DNA <213> Artificial Sequence <400> 120 tcgcggga 8 <210> 121 <211> 8 <212> DNA <213> Artificial Sequence <400> 121 cggcgtgc 8 <210> 122 <211> 8 <212> DNA <213> Artificial Sequence <400> 122 acgtatac 8 <210> 123 <211> 8 <212> DNA <213> Artificial Sequence <400> 123 ccccgaac 8 <210> 124 <211> 8 <212> DNA <213> Artificial Sequence <400> 124 acctggag 8 <210> 125 <211> 8 <212> DNA <213> Artificial Sequence <400> 125 tggaggac 8 <210> 126 <211> 8 <212> DNA <213> Artificial Sequence <400> 126 gaccaaag 8 <210> 127 <211> 8 <212> DNA <213> Artificial Sequence <400> 127 ccctaagt 8 <210> 128 <211> 8 <212> DNA <213> Artificial Sequence <400> 128 atcggtag 8 <210> 129 <211> 8 <212> DNA <213> Artificial Sequence <400> 129 accattcc 8 <210> 130 <211> 8 <212> DNA <213> Artificial Sequence <400> 130 cccggatt 8 <210> 131 <211> 8 <212> DNA <213> Artificial Sequence <400> 131 tcaggact 8 <210> 132 <211> 8 <212> DNA <213> Artificial Sequence <400> 132 acggatcg 8 <210> 133 <211> 8 <212> DNA <213> Artificial Sequence <400> 133 atcggtcg 8 <210> 134 <211> 8 <212> DNA <213> Artificial Sequence <400> 134 tcctcggg 8 <210> 135 <211> 8 <212> DNA <213> Artificial Sequence <400> 135 tgtcgtag 8 <210> 136 <211> 8 <212> DNA <213> Artificial Sequence <400> 136 acgggcgg 8 <210> 137 <211> 8 <212> DNA <213> Artificial Sequence <400> 137 caagcgaa 8 <210> 138 <211> 8 <212> DNA <213> Artificial Sequence <400> 138 gccccgtgt 8 <210> 139 <211> 8 <212> DNA <213> Artificial Sequence <400> 139 ctatatca 8 <210> 140 <211> 8 <212> DNA <213> Artificial Sequence <400> 140 aggagttt 8 <210> 141 <211> 8 <212> DNA <213> Artificial Sequence <400> 141 cgctgtgt 8 <210> 142 <211> 8 <212> DNA <213> Artificial Sequence <400> 142 cccgatgt 8 <210> 143 <211> 8 <212> DNA <213> Artificial Sequence <400> 143 agccgtgc 8 <210> 144 <211> 8 <212> DNA <213> Artificial Sequence <400> 144 atatacgg 8 <210> 145 <211> 8 <212> DNA <213> Artificial Sequence <400> 145 caaggtga 8 <210> 146 <211> 8 <212> DNA <213> Artificial Sequence <400> 146 ttctagtt 8 <210> 147 <211> 8 <212> DNA <213> Artificial Sequence <400> 147 actacgga 8 <210> 148 <211> 8 <212> DNA <213> Artificial Sequence <400> 148 cacgggac 8 <210> 149 <211> 8 <212> DNA <213> Artificial Sequence <400> 149 gcgtgata 8 <210> 150 <211> 8 <212> DNA <213> Artificial Sequence <400> 150 ttagatca 8

Claims

1. A method for constructing a colorectal cancer early screening model. It is characterized in that The steps include: Step 1: extract and sequence cfDNA from samples of the positive group and the control group to obtain read data; Step 2, aligning the read data results to the reference genome, and obtaining the number of reads in different length intervals within different window ranges on the reference genome as the first feature value; Step 3, aligning the read data result to the reference genome to obtain the position of the 5' end of the read segment on the reference genome; obtaining the sequence data of m bp bases upstream and downstream of the position as a base fragment set; The proportion of each obtained base fragment in all fragments is taken as the second characteristic value; Step 4, dividing the reference genome into multiple windows, and obtaining the copy number data within each window as the third eigenvalue; Step 5, taking the first, second and third eigenvalues ​​together as initial eigenvalues, screening out eigenvectors in the initial eigenvalues ​​that have significant differences between the samples in the positive group and the control group as model eigenvectors; Step 6, taking the characteristic vector of significant difference between the samples of the positive group and the control group as the input value of the classifier model, and taking the probability of colorectal cancer as the model output value, the model is trained to obtain an early screening model; the classifier model is a linear model, and the variables contained in the model are obtained by inputting the characteristic vectors with significant differences among the first, second and third eigenvalues ​​into the gradient boosting algorithm model, random forest model and deep network learning model trained in step 5, respectively, and the classifier model uses a generalized linear model to perform secondary set training on the results of the 9 training models, and converts the 9 training models into final linear equations; The step 2 comprises: step 2-1, dividing the reference genome into multiple windows, and obtaining the total number of reads, the number of short reads, and the number of ultra-long reads within each window respectively; Step 2-2, taking the long arm and short arm of each chromosome as the region range, and obtaining the number of reads in different length gradient intervals within each range; the long arm and short arm refer to: chr1_p, chr4_q, chr8_p, chr11_q, chr16_q, chr20_p, chr1_q, chr5_p, chr8_q, chr12_p, chr17_p, chr20_q, chr2_p, chr5_q, chr9_p, chr12_q, chr 17_q, chr21_q, chr2_q, chr6_p, chr9_q, chr13_q, chr18_p, chr22_q, chr3_p, chr6_q, chr10_p, chr14_q, chr18 _q, chrX_p, chr3_q, chr7_p, chr10_q, chr15_q, chr19_p, chrX_q, chr4_p, chr7_q, chr11_p, chr16_p, chr19_q; Step 2-3, taking the data obtained in steps 2-1 and 2-2 as the first eigenvalue; The short read segment refers to a length of 40-80 bp, the ultra-long read segment is 200-300 bp; all read segments refer to a length in the range of 40-300 bp; The size range of the window in step 2-1 is 5Mb; The different length gradient intervals in step 2-2 refer to different length gradient ranges obtained by increasing the step length of 10 bp within the range of 40-300 bp; m is 4.

2. The method for constructing a colorectal cancer early screening model according to claim 1, It is characterized in that The read numbers stated are normalized.

Citation Information

Patent Citations

  • Method for predicting tumor risk value based on plasma multi-omics multi-dimensional characteristics and artificial intelligence

    CN112397143A

  • Lung cancer early screening marker, model construction method, detection device and computer readable medium

    CN113355421A

  • Ultra-sensitive detection of circulating tumor DNA through genome-wide integration

    US20210043275A1