Colorectal cancer prediction method and system based on metabolism-related DNA single-stranded fracture characteristics

By using the metabolism and DNA single-strand break characteristics of peripheral blood samples in colorectal cancer screening, machine learning models are constructed, solving the problem of insufficient sensitivity and specificity in the prior art, and achieving high accuracy and low cost colorectal cancer prediction.

CN119943160AActive Publication Date: 2025-05-06THE FIRST AFFILIATED HOSPITAL OF XIAMEN UNIV

Patent Information

Application Number
CN202510440364.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art has problems of insufficient sensitivity and specificity in colorectal cancer screening. Traditional methods are highly invasive, limited diagnostic accuracy of biomarkers, and lack of multimodal models and high-cost data analysis techniques, which limit their application in clinical and large-scale populations.

Method used

By collecting peripheral blood samples and clinical data, metabolic characteristics and DNA single-strand break characteristics were obtained, statistical preliminary screening and correlation analysis were performed, metabolic-related DNA single-strand break characteristics were screened out, and machine learning models were constructed for colorectal cancer prediction.

Benefits of technology

It significantly improves the prediction accuracy of colorectal cancer and precancerous lesions, with AUC reaching 0.866, which has early warning value, reduces detection costs, provides non-invasive detection methods, and is suitable for large-scale screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943160A_ABST
    Figure CN119943160A_ABST
Patent Text Reader

Abstract

The invention discloses a colorectal cancer prediction method and system based on metabolism-related DNA single-stranded fracture characteristics. The method comprises the following steps: acquiring metabolism characteristics and DNA single-stranded fracture characteristics; carrying out statistical preliminary screening on the DNA single-stranded fracture characteristics to screen out a first DNA single-stranded fracture characteristic; performing correlation analysis on the first DNA single-stranded fracture characteristics and the metabolism characteristics to obtain metabolism-related DNA single-stranded fracture characteristics; a Lasso regression model is combined to analyze the metabolism-related DNA single-stranded fracture characteristics, and a final metabolism-related DNA single-stranded fracture (MSSB) characteristic set is obtained; a prediction model is constructed, and training is carried out based on the final MSSB feature set; and performing colorectal cancer risk prediction by using the trained prediction model. The colorectal cancer related lesion screening model based on the multiple omics characteristics is high in generalization ability and accuracy, the cost is relatively low, the crowd acceptability is high, and important clinical application value is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to biomedical detection technology, and in particular to a method and system for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics. Background Art

[0002] Colorectal cancer (CRC) is one of the malignant tumors with the highest morbidity and mortality rates worldwide. Traditional colorectal cancer screening methods such as stool testing and colonoscopy have low sensitivity for detecting precancerous lesions (PL) and require invasive procedures. Although non-invasive biomarkers, such as carcinoembryonic antigen (CEA) and Septin 9 gene methylation detection, provide alternatives, their diagnostic accuracy and sensitivity are still limited. Current detection methods have many shortcomings: (1) Limitations of traditional screening methods: Stool testing and colonoscopy are standard screening methods, but fecal occult blood tests have low sensitivity for early precancerous lesions (PL), and colonoscopy is invasive and has the risk of missed detection, which brings discomfort and potential risks to patients. (2) Poor performance of existing biomarkers: Non-invasive biomarkers such as carcinoembryonic antigen (CEA) and septin 9 gene promoter methylation have limited diagnostic accuracy. CEA has a sensitivity of only 30-40% for early CRC, and septin 9 methylation detection has low specificity (AUC = 0.65-0.70). (3) Lack of multi-omics integration and lack of multimodal models: Specific biomarkers can only reflect biological changes at a certain level and lack a comprehensive assessment of the overall physiological state. Different biomarkers may have complex relationships that are independent or interactive with each other, and a single type of marker cannot fully reflect complex physiological and pathological processes. In the scenario of this patent, it is reflected in the fact that a single omics model cannot capture the metabolic-genomic interaction in the occurrence of CRC. (4) Complex data analysis and high cost: The various biomarker detection methods and analysis processes in the existing technology are complex and expensive, which limits their application in clinical and large-scale populations. In particular, some high-throughput omics technologies, such as epigenomics and transcriptomics, although they can provide rich information, their high cost and technical barriers hinder their widespread application. (5) Simple prediction models and poor generalization ability: Many colorectal cancer risk prediction models are currently based on simple linear regression. Whether it is a statistical method or a machine learning method, these models often have significant errors when processing high-dimensional omics data. First, simple linear regression models perform poorly when the relationship between features and disease diagnosis is complex. Omics data usually contain a large number of features, and linear regression models may not be able to capture the nonlinear relationships and interactions between these features. Simple linear regression is sensitive to the collinearity of features, which often leads to model instability and increased prediction errors in omics data. It may not provide sufficiently accurate prediction results when dealing with complex data sets. Summary of the invention

[0003] The purpose of the present invention is to solve the problems in the prior art.

[0004] The technical solution adopted by the present invention to solve the technical problem is: to provide a method for predicting colorectal cancer based on the characteristics of metabolism-related DNA single-strand breaks, comprising the following steps:

[0005] Ethylenediaminetetraacetate-anticoagulated peripheral blood samples and clinical data were collected to obtain the metabolic characteristics of peripheral blood plasma and the DNA single-strand break characteristics of peripheral blood cells;

[0006] Performing statistical preliminary screening on the DNA single-strand break characteristics, and obtaining the first DNA single-strand break characteristics;

[0007] Performing a correlation analysis on the first DNA single-strand break feature and the metabolic feature to obtain the metabolism-related DNA single-strand break feature; performing feature screening on the metabolism-related DNA single-strand break feature to obtain a final MSSB feature set for training and developing a machine learning model;

[0008] Use machine learning algorithms to build prediction models and train them based on the final MSSB feature set;

[0009] Use the trained prediction model to predict colorectal cancer.

[0010] Preferably, the statistical preliminary screening of the DNA single-strand break characteristics to screen out the first DNA single-strand break characteristics comprises the following steps:

[0011] The LPKM-SSB value of each gene in peripheral blood samples was calculated as the DNA single-strand break feature and expressed as:

[0012] LPKM-SSBi = (Ri * 10^5) / (total number of SSB positions in the sample);

[0013] Among them, LPKM-SSBi represents the LPKM - SSB value of SSB of gene i in each sample; Ri represents the number of SSB positions in the entire region of gene i;

[0014] The F1 score of DNA single-strand break feature for diagnosing colorectal cancer was calculated, and DNA single-strand break features with F1 score ≥ 0.85 were selected for one-way ANOVA. DNA single-strand break features with p ≤ 0.05 were selected as the first DNA single-strand break feature. The p value indicated the direct difference between normal and precancerous lesion or tumor groups in one-way ANOVA.

[0015] Preferably, the correlation analysis of the first DNA single-strand break feature and the metabolic feature to obtain the metabolism-related DNA single-strand break feature is specifically: using the O2PLS method to construct a correlation model between the metabolic feature and the first DNA break feature, and selecting several DNA single-strand break features with the highest correlation as the metabolism-related DNA single-strand break features.

[0016] Preferably, the metabolism-related DNA single-strand break features are subjected to feature screening, and a LASSO regression model with a 10-fold cross-validation-optimized LASSO regularization parameter λ is used for feature screening, or a random forest importance scoring algorithm is used for feature screening.

[0017] Preferably, the use of a machine learning algorithm to construct a prediction model and training based on the final MSSB feature set comprises the following steps:

[0018] Build multiple predictive models using a variety of machine learning algorithms;

[0019] Based on the final MSSB feature set, the optimal MSSB hyperparameters of each model were determined through 10-fold internal cross-validation;

[0020] Each model is retrained using the final MSSB feature set and the MSSB optimal hyperparameters; all trained models are evaluated, and the model with the best prediction effect is selected as the prediction model.

[0021] Preferably, the multiple machine learning algorithms include: logistic regression, K-nearest neighbor, Extra Trees classifier, random forest, extreme gradient boosting, light gradient boosting machine, support vector machine and neural network.

[0022] Preferably, the indicators used in the evaluation include: ROC curve, accuracy, Kappa coefficient, precision, and negative predictive value.

[0023] Preferably, the SHAP method is also used to rank the importance of input features for the trained model to explain the results of the prediction model.

[0024] Preferably, the trained prediction model is an MSSB-KNN model constructed using a K-nearest neighbor algorithm, and the MSSB-KNN model is predicted based on the LPKM-SSB values ​​of eight genes, wherein the eight genes include ABCB5, AC093642.5, AF146191.4, ANO3, CTD.2231H16.1, PRKG1, RASA3 and VIPR2; the prediction process of the MSSB-KNN model includes the following steps:

[0025] Calculate the LPKM-SSB values ​​of the eight genes in the peripheral blood samples of the subjects and input them into the MSSB - KNN model;

[0026] The MSSB-KNN model outputs the risk probability of colorectal cancer or colorectal precancerous lesions.

[0027] The present invention also provides a colorectal cancer prediction system based on metabolism-related DNA single-strand break characteristics, which is used to implement any of the above-mentioned colorectal cancer prediction methods based on metabolism-related DNA single-strand break characteristics, comprising:

[0028] The collection module collects EDTA-anticoagulated peripheral blood samples and clinical data to obtain the metabolic characteristics of peripheral blood plasma and the DNA single-strand break characteristics of peripheral blood cells;

[0029] The first screening module performs a statistical preliminary screening of the DNA single-strand break characteristics to obtain the first DNA single-strand break characteristics;

[0030] The second screening module performs a correlation analysis between the first DNA single-strand break feature and the metabolic feature to obtain the metabolism-related DNA single-strand break feature; performs feature screening on the metabolism-related DNA single-strand break feature to obtain a final MSSB feature set for training and developing a machine learning model;

[0031] The training module uses machine learning algorithms to build prediction models and performs training based on the final MSSB feature set;

[0032] The prediction module uses the trained prediction model to predict colorectal cancer.

[0033] The present invention has the following beneficial effects:

[0034] (1) Breakthrough in diagnostic performance: In multicenter validation, the model's AUC for colorectal cancer reached 0.866 (95% CI: 0.797- 0.934), significantly better than the traditional CEA and septin 9 methylation combined test (AUC: 0.708). In particular, in the detection of precancerous lesions, the AUCs of two independent cohorts were 0.873 and 0.916, respectively, which has early warning value and improves the prediction accuracy of colorectal cancer and precancerous lesions.

[0035] (2) Enhanced clinical applicability: Only 1 mL of peripheral blood is needed to complete the test, providing a non-invasive test method, avoiding invasive operations, reducing the pain of patients, and being suitable for large-scale screening. In addition, the test has demonstrated excellent diagnostic performance in two tertiary hospitals, and has a good prospect for promotion and application.

[0036] (3) Biomarker advantages: The discovered metabolism-related DNA breakage biomarkers have high sensitivity and specificity for CRC and PL, providing more effective non-invasive indicators for disease detection and more accurately reflecting the disease status.

[0037] (4) Accurate prediction by machine learning model: Through the optimization of machine learning models, more accurate prediction results can be provided, which has higher accuracy and predictive ability than existing biomarker detection methods.

[0038] (5) During the model training process, 8 patients with other cancer types were added to the control group, making the MSSB_KNN model have a certain degree of cancer type specificity. Further training of the model using relevant data can support the adaptation of multiple cancer types such as breast cancer and gastric cancer.

[0039] (6) Cost reduction: The model ultimately used for clinical applications only requires detection of SSB signals, and the cost of testing a single sample is less than RMB 500 (Septin9 fee, CEA fee).

[0040] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments, but the present invention is not limited to the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A method step diagram of an embodiment of the present invention;

[0042] Figure 2 It is a schematic diagram of a flow chart of an embodiment of the present invention;

[0043] Figure 3 This is a metabolite analysis diagram of an embodiment of the present invention;

[0044] Figure 4 This is a process diagram of feature selection of a method for predicting colorectal cancer and pre-lesions based on metabolism-related DNA breaks according to an embodiment of the present invention;

[0045] Figure 5 This is a prediction nomogram using three models or clinical features of MSSB_KNN, MSSB_LR and MSSB_NN in a validation cohort according to an embodiment of the present invention.

[0046] Figure 6 It is a comparison diagram of ROC curves of an embodiment of the present invention (including CEA, septin 9 methylation and MSSB_KNN model);

[0047] Figure 7 This is a web calculator interface according to an embodiment of the present invention.

[0048] Figure 8 4 is a system structure diagram of an embodiment of the present invention. DETAILED DESCRIPTION

[0049] See also Figure 1 and Figure 2 The method step diagram and flow chart of an embodiment of the present invention are shown, which include the following steps:

[0050] S101, collect peripheral blood samples anticoagulated with EDTA and clinical data to obtain metabolic characteristics of peripheral blood plasma and DNA single-strand break characteristics of peripheral blood cells;

[0051] S102, performing preliminary statistical screening on the DNA single-strand break characteristics to obtain a first DNA single-strand break characteristic;

[0052] S103, performing correlation analysis on the first DNA single-strand break feature and the metabolic feature to obtain metabolism-related DNA single-strand break features; performing feature screening on the metabolism-related DNA single-strand break features to obtain a final MSSB feature set for training and developing a machine learning model;

[0053] S104, constructing a prediction model using a machine learning algorithm and training it based on the final MSSB feature set;

[0054] S105, using the trained prediction model to predict colorectal cancer.

[0055] Specifically, the S101 includes:

[0056] S1011, Collection of clinical data: Clinical data were extracted from the participants' inpatient electronic medical records (EMRs), covering baseline patient information, including age, sex, laboratory test results (Septin 9 gene methylation status, CEA level) within one month before endoscopy or surgery, colonoscopy and pathological diagnosis, as well as TNM stage (T stage, N stage, M stage) and clinical classification.

[0057] At S1012, peripheral blood samples were collected, plasma was extracted for metabolomics analysis, and DNA from peripheral blood cells was extracted for fragmentation analysis.

[0058] S1013, Metabolomics Analysis: Plasma samples were obtained from EDTA-treated blood for metabolomics analysis. The specific steps were as follows: Plasma samples were thawed, vortexed, and aliquoted into 1.5 mL centrifuge tubes. Each sample was divided into six aliquots for multichannel analysis (30 µL per channel). When preparing mixed samples, 75 µL of serum from each sample was combined and vortexed to form a reference sample. To precipitate proteins, 90 µL of pre-cooled mass spectrometry-grade methanol was added to each sample, vortexed, centrifuged at 12,000 rpm for 10 minutes at 4°C, and the supernatant was transferred to a new tube and dried with nitrogen. For carboxyl metabolome analysis, 25 µL of a 3:1 volume ratio of acetonitrile / water was added to the sample, followed by 10 µL of the reaction catalyst and 25 µL of carbon-12 or carbon-13 labeling reagent, and incubated at 80°C for 60 minutes. After incubation, 40 µL of quenching reagent was added and incubated at the same temperature for another 30 minutes. Ultra-high performance liquid chromatography quadrupole time-of-flight mass spectrometry (UHPLC Q-TOF / MS) was used for analysis. Chromatographic separation was performed on an Agilent Eclipse Plus reversed-phase C18 column (150×2.1 mm, 1.8 µm). The mobile phases were 0.1% formic acid in water (A) and 0.1% formic acid in acetonitrile (B). The gradient program was 25% B at the start, increasing to 99% B within 10 minutes and maintaining for 5 minutes, returning to 25% B at 15.1 minutes and maintaining for 18 minutes. The column temperature was 40°C and the flow rate was 400 µL / min. The ion source parameters were: gas temperature 325°C; drying gas 8 L / min; nebulizer pressure 35 psi; sheath gas temperature 400°C; sheath gas flow rate 12 L / min; MS scan range m / z220-1000.

[0059] S1014, DNA break detection: Whole-genome DNA break analysis was performed on 200 μL of peripheral blood cell DNA using the single-strand break mapping at the nucleotide genome level (SSiNGLe) method. DNA was extracted according to the instructions provided by the manufacturer of the DNA extraction kit, and 100 ng of DNA was directly used for SSiNGLe library construction. Sequencing of all DNA break profile experiments was performed by sequencing companies on the Illumina / BGI platform using a paired-end 150 bp strategy with a sequencing depth of 2GB-5GB.

[0060] S1015, Feature calculation: Calculate the LPKM - SSB value representing the SSB of gene i in each sample using the formula LPKM - SSB = (Ri * 10^5) / (total number of SSB positions in the sample), where Ri is the number of SSB positions located in the entire region of gene i. Calculate the relative abundance (RA) of SSB in each sample using the formula RA = (SSB count in the specified region of the current sample) / ((total SSB count of all samples in the current cohort) / (total number of samples)).

[0061] Specifically, in S1014, the SSiNGLe method library building includes the following steps:

[0062] S201, preliminary preparation, take out the magnetic beads and place them at room temperature for equilibrium.

[0063] S202, PolyA tailing, including:

[0064] S2021, denaturation; if the DNA quality is sufficient, add the sample according to the following system: add 2 μL TdT10x buffer, 2 μL 2.5 mM CoCl2, 10 μL DNA (100 ng) and 5 μL water to the reaction system, with a total volume of 19 μL; if the DNA quality is insufficient, no water is added. After the addition is completed, the reaction system is placed in a C1000 touch PCR instrument for reaction, and the reaction conditions are 95℃ denaturation for 5 min, and then the reaction system is quickly placed on ice.

[0065] S2022, tailing; when configuring the tailing system, please note: prepare an extra system for every 3 samples in case there is not enough sample. The tailing system contains 1 μL of NEB terminal transferase diluted to 4U / μL (1:4 dilution) with 1x buffer and 2 μL of 10 mM dATP. After configuring the system, vortex mix for 10 s, then centrifuge for 10 s, and then dispense 3 μL of the system into the 19 μL denaturation system in the previous step, with a total volume of 22 μL. Place the mixed system in a C1000 touch PCR instrument for reaction, and the reaction conditions are incubation at 37°C for 30 min, and then temporarily stored at 4°C.

[0066] S2023, blocking; add 2 μL 10 mM ddNTP to the tailed system to a final volume of 24 μL. Place the system in a C1000 touch PCR instrument for reaction under the following conditions: incubate at 37°C for 30 min, react at 70°C for 10 min, and then temporarily store at 4°C.

[0067] S2024, magnetic bead purification: add twice the final volume (i.e. 48 μL) of magnetic beads to the above blocked system for purification. Use BGI MGISP-960 library construction instrument for purification, use freshly prepared 80% alcohol during purification, and finally elute with 16 μL ultrapure water.

[0068] S203, amplification reaction (using hot-start Taq enzyme and DNA-RNA oligo-d(T)50-r(T)3 oligonucleotide primer); the reaction system contained 0.4 μL rTaq (1U, Tiangen), 1 μL 10 μM oligo-d(T)50-r(T)3 primer, 2 μL 10 x Buffer, 1.6 μL 2.5 mM dNTP mix and 15 μL DNA purified in the previous step, with a total volume of 20 μL. The reaction system was placed in a C1000 touch PCR instrument for reaction. The reaction conditions included: (1) 94℃ 30 sec; (2) 94℃ 1 min; (3) 55℃ 30 sec; (4) 72℃ 30 sec; (2) to (4) were repeated for a total of 10 cycles; the reaction system was immediately placed on ice after the reaction was completed. For magnetic bead purification, use 2 volumes (i.e. 40 μL) of magnetic beads for purification and elute with 16 μL of ultrapure water.

[0069] S204, Poly C tailing; including:

[0070] S2041, denaturation; the reaction system contains 2 μL TdT 10x buffer, 2 μL 2.5 mM CoCl2 and 15 μL DNA eluted in the previous step, with a total volume of 19 μL. The reaction system was placed in a C1000 touch PCR instrument for reaction, and the reaction conditions were 95℃ denaturation for 5 min, and then the reaction system was immediately placed on ice.

[0071] S2042, tailing; add 1 μL of NEB TdT diluted to 5U / μL (1:4 dilution) with 1x buffer and 2 μL of 10 mM dCTP to the above denatured system, and the total volume reaches 22 μL. The system is placed in a C1000 touch PCR instrument for reaction, and the reaction conditions are incubation at 37℃ for 30 min, and then stored at 4℃.

[0072] S2043, magnetic bead purification; add 2 times the volume (i.e. 44 μL) of magnetic beads and use the BGI MGISP-960 library preparation instrument for purification. Use freshly prepared 80% alcohol for purification and finally elute with 15 μL ultrapure water.

[0073] S205, amplification reaction (using P5G10 and P7T12 primers); in the reaction system, 0.4 μL (1U, Tiangen) of rTaq was added alone, and it also contained 1 μL 10 μM P5G10 primer, 1 μL 10 μM P7T12 primer, 1.6 μL 2.5mM dNTP mix, 2 μL 10 x Buffer and 14 μL DNA eluted in the previous step, with a total volume of 20 μL. The reaction system was placed in a C1000 touch PCR instrument for reaction. The reaction conditions included: (1) 94℃ for 3 min; (2) 94℃ for 30 sec; (3) 55℃ for 1 min; (4) 72℃ for 1 min; (5) 94℃ for 30 sec; (6) 37℃ for 1 min; (7) 37℃ + 2℃ / 2 min (cycle); (8) repeat step (7) for a total of 17 cycles; (9) 72℃ for 2 min; (10) 94℃ for 30 sec; (11) 60℃ for 30 sec, for a total of 30 cycles; (12) 72℃ for 30 sec; (13) 72℃ for 7 min; (14) 4℃ hold. For magnetic bead purification, add 2 times the volume (i.e. 40 μL) of magnetic beads and use the BGI MGISP-960 library preparation instrument for purification. Use freshly prepared 80% alcohol for purification and finally elute with 14 μL ultrapure water.

[0074] S206, amplification reaction (using Illumina_P5 and Illumina_P7 primers); the total volume of the reaction system was 20 μL, including 0.4 μL rTaq, 1 μL 10 μM Illumina_P5 primer, 1 μL 10 μM Illumina_P7primer, 1.6 μL 2.5 mM dNTP mix, 2 μL 10 x Buffer and 14 μL DNA. The reaction conditions were 94℃ 3min, 94℃ 30 sec, 55℃ 30 sec (6 cycles for cell samples; 10 cycles for blood samples; 12 cycles for tissue samples), 72℃ 30 sec, 72℃ 7 min, 4℃ hold. Magnetic bead purification was performed; 2 times the volume (i.e. 40 μL) of magnetic beads were added, and the library was purified using the BGI MGISP-960 library preparation instrument. Freshly prepared 80% alcohol was used for purification, and finally eluted with 22 μL ultrapure water.

[0075] S207, concentration determination and storage; take 1 μL of the eluted product and use Qubit to determine the concentration, and freeze the remaining product at -20°C for library quality inspection and sequencing.

[0076] Specifically, in the embodiment of the present invention, the clinical data used by S1011 are derived from the participants' inpatient electronic medical records (EMR), covering the baseline information of the patients. Specifically, the variables collected include age, gender, laboratory test results (septin 9 gene methylation status, CEA level) one month before endoscopy or surgery, colonoscopy and pathological diagnosis results, as well as TNM stage (T stage, N stage, M stage) and clinical grade; and the clinical characteristics of different research cohorts are compared and statistically analyzed, including key variables such as age, gender distribution, disease stage, and biomarker levels, to evaluate the consistency and difference between cohorts. For metabolomics analysis, residual plasma samples collected from EDTA-treated blood in routine clinical tests were used, and DNA fragmentation analysis of peripheral blood cells was performed.

[0077] Specifically, in the embodiment of the present invention, S1013 also performs metabolomics mass spectrometry difference analysis and differential metabolite enrichment analysis, and uses multivariate statistical analysis (OPLS-DA) to analyze the plasma metabolomics mass spectrometry results. In addition, univariate statistical analysis (T test) was performed to identify differential metabolites between the NC, PL and CRC groups, and the screening criteria were VIP score ≥ 1 and p value < 0.05. KEGG pathway enrichment analysis was further performed to explore the regulatory pathways related to the differential metabolites. By obtaining the differential metabolites between the groups, alternative metabolites were obtained for the subsequent joint analysis of metabolism and SSB. For the analysis results, see Figure 3 As shown, through the Venn diagram ( Figure 3 (A) and (D) in the figure show the up-regulated and down-regulated metabolites in normal control (NC) compared with precancerous lesions (PL), normal control (NC) and colorectal cancer (CRC); the violin plots ( Figure 3 (B) and (C) show representative upregulation, as shown by violin plots ( Figure 3 (E) and (F) in Figure 3 show down-regulated metabolites. Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis of differential metabolites between normal control (NC) and colorectal cancer (CRC), normal control (NC) and precancerous lesions (PL), see Figure 3 (G) and (H) in the figure. NC stands for non-CRC / PL Control; PL stands for precancerous lesions; CRC stands for colorectal cancer.

[0078] Specifically, the S1015 analyzes the SSiNGLe data, calculates the Ri value of each gene in each sample, and thus calculates LPKM-SSB (representing the SSBs of gene i in each sample), and the calculation formula is as follows: LPKM-SSB=(Ri*10^5) / (the total number of positions of all SSBs in the sample);

[0079] Among them, Ri is the number of SSB positions in the entire region of gene i. The relative abundance (RA) of SSB in each sample is calculated as follows: RA=(SSB count in the specified region of the current sample) / ((Total number of SSBs in all samples in the current cohort) / (Total number of samples)). By calculating RA, we mainly look at the distribution of DNA breakage in different subgroups, reflecting the biological characteristics of DNA breakage.

[0080] For details, see Figure 4 As shown, in S102, the F1_Score function in R was used to calculate and select SSB genes with F1 scores greater than or equal to 0.85. The CreateTableOne function was further used to perform one-way analysis of variance (ANOVA) and select SSB genes with p values ​​less than or equal to 0.05. The Lasso regression model was constructed using the glmnet function, and the cv.glmnet function was used to perform 10-fold cross validation (CV) to select the optimal λ, thereby obtaining the final SSB gene set for model construction. See Figure 4 As shown, (A) and (B) show the association between LPKM-SSB values ​​of DNA single-strand break (SSB) genes and metabolite quantitative values, analyzed using the Two-Way Orthogonal Partial Least Squares (O2PLS) model. The loading plot shows the strength of the association, and coordinates with larger absolute values ​​represent stronger associations. The top 10 ranked elements are highlighted in red. (C) shows the loading matrix of SSB genes and metabolites, and elements close to the outer circle have higher correlations between the two omics data. (D) shows the change of regression coefficients in LASSO regression with Log(λ), and the importance of each variable is determined by the size of its coefficient. (E) shows the mean squared error (MSE) of 10-fold cross validation, indicating the optimal model by showing the change of MSE with Log(λ). The red lambda.min line indicates the value corresponding to the minimum MSE, and the gray lambda.1se line indicates a simpler model. (F) Highlights the feature importance assessed by LASSO regression, and the formula shows how the LPKM-SSB value and weight of each gene affect the final LASSO score.

[0081] Specifically, in S103, the O2PLS (bidirectional orthogonal partial least squares) method is applied to simultaneously model and predict paired samples of plasma metabolites and peripheral blood cell SSB in the training cohort, revealing the relationship between metabolomics and DNA fragmentation data sets, quantifying the correlation between the two omics, and identifying key SSB genes (MSSB) associated therewith. The top 25 MSSB genes associated with metabolites are selected, and a one-way ANOVA is performed using the CreateTableOne function. Genes with a p value ≤ 0.05 are selected, and a Lasso regression model is constructed using the glmnet function, followed by a 10-fold cross validation using the cv.glmnet function to determine the optimal λ, and the final MSSB gene set for model development is obtained. The results of the top 25 metabolism-related SSB genes obtained in the embodiment of the present invention are shown in the following table, where n in the table represents the number of samples in each group.

[0082]

[0083] For details, see Figure 2 As shown, the S104 includes:

[0084] S1041, Model construction: Eight machine learning (ML) algorithms including logistic regression (LR), K-nearest neighbor (KNN), Extra Trees classifier (ExtraTrees), random forest (RF), extreme gradient boosting (XGB), light gradient boosting machine (LGBM), support vector machine (SVM) and neural network (NN) were used to predict the likelihood of PL or CRC. Based on the best feature subset, the optimal hyperparameters of each algorithm were determined by 10 rounds of 10-fold cross validation and grid search using the default hyperparameters of the "caret" package. Finally, the model was retrained using the selected feature subset and the final hyperparameters derived from 10 rounds of 10-fold internal cross validation to obtain the MSSB-KNN model based on eight genes, including: ABCB5, AC093642.5, AF146191.4, ANO3, CTD.2231H16.1, PRKG1, RASA3, VIPR2.

[0085] S1042, Model Validation and Comparison: Use a variety of standard performance indicators to evaluate the effectiveness of the model, such as the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, F1 score, etc. Apply a calibration curve to reflect the consistency between the predicted probability and the observed results, and use decision curve analysis (DCA) to evaluate the net benefit of the model at various thresholds. Select the optimal prediction model based on the evaluation indicators in the training cohort and the validation cohort. Use the DeLong test to determine whether there is a significant difference in the AUC values ​​of different models, calculate the log loss (Log-loss) to measure the difference between the actual label and the predicted probability, and evaluate the prediction accuracy.

[0086] S1043, Model Interpretation: The game theory-based SHapley Additive exPlanation (SHAP) method is used to interpret the machine learning model, rank the importance of input features and clarify the results of the prediction model. By calculating the contribution of each feature to the prediction results, local and global explanations are provided to enhance the transparency and interpretability of the model and determine the practical significance of the model in predicting CRC / PL.

[0087] Specifically, in S1042, statistical calculations of diagnostic performance indicators are performed for all machine learning (ML) models, CEA, septin 9 methylation, and CEA_plus_septin 9 methylation (where the diagnostic criteria for CEA_plus_septin 9 methylation are defined as positive if any indicator is positive). In particular, when evaluating diagnostic performance, clinical data sets can be stratified according to cofactors such as age and gender to reflect the impact of different cofactors on diagnostic performance.

[0088] See also Figure 5 Shown are the prediction nomograms (Nomogram) using the three models or clinical characteristics of MSSB_KNN, MSSB_LR and MSSB_NN in the validation cohort, (A) is the nomogram of NC vs. CRC / PL in validation cohort 1; (B) is the nomogram of NC vs. PL in validation cohort 1; (C) is the nomogram of NC vs. CRC in validation cohort 1; (D) is the nomogram of NC vs. PL in validation cohort 2, where MSSB_KNN represents the metabolism-related DNA Single-strand Break model using K-Nearest Neighbors. Figure 5 The smaller the z statistic / coefficient significance p value, the better the classification effect of the corresponding model on the target group ( Figure 5In the table, “*” indicates that the p value is less than 0.05, “**” indicates that the p value is less than 0.01, and “***” indicates that the p value is less than 0.001). Figure 5 It can be seen that the classification effect of MSSB_KNN is the best in each cohort. In particular, the Delong test is used to determine whether the AUC values ​​of different models are significantly different: the smaller the p-value of the Delong test results of the AUC differences between different machine learning models, CEA, septin 9 methylation and CEA_plus_septin 9 methylation prediction methods, the more significant the AUC difference between the two methods. In addition, the log-loss value is used to evaluate the difference between the actual label (whether or not the patient actually has colorectal cancer or related precancerous lesions) and the predicted probability, thereby measuring the prediction accuracy of the model. The lower the log-loss value, the better the performance.

[0089] Specifically, in S1042, the primary outcome is to predict the diagnosis of colorectal precancerous lesions (PL) or colorectal cancer (CRC). Model performance is evaluated by indicators such as AUC, accuracy, sensitivity and specificity, and colonoscopy results are used as the "gold standard" reference; secondary outcomes: compare the diagnostic accuracy, sensitivity and specificity of the model in identifying CRC / PL, and compare with basic tests such as CEA and septin 9 methylation.

[0090] See also Figure 6As shown, it is a comparison diagram of the ROC curves of the embodiment of the present invention and other predictive indicators. The line graphs (A) to (D) show the AUC for distinguishing different groups in the training cohort, validation cohort 1 and validation cohort 2, wherein (A) shows NC vs. CRC / PL, (B) shows NC vs. PL, (C) shows NC vs. CRC, and (D) shows PL vs. CRC. The box plots (E) to (G) show the predicted probabilities of CRC or PL samples diagnosed by MSSB-KNN in the training cohort, validation cohort 1 and validation cohort 2, respectively. (H) to (I) show that the MSSB-KNN model can correctly detect CRC patients who may be misclassified by CEA or septin 9 methylation. CEA+ or MSSB+ represents participants detected by CEA (cutoff value ≥ 5ng / mL) or MSSB (cutoff value of predicted probability ≥ 0.5). Grouping: triangles represent CRC / PL, circles represent NC; septin 9 methylation positive (cutoff value Ct≤41) is red, negative is blue. Abbreviations: MSSB = Metabolism-related DNA Single-strand Break, KNN = K-Nearest Neighbors, septin 9 = septin 9 methylation, CEA_plus_septin 9 = defined as positive if either of the two indicators is positive, train = training cohort, test1 = validation cohort 1, test2 = validation cohort 2. Figure 6 It shows that using colonoscopy results as the gold standard, the MSSB_KNN model performs better than CEA and septin 9 methylation in distinguishing non-CRC / PL and CRC / PL patients.

[0091] Specifically, the S105 is implemented by using a model network calculator: Figure 7 The interface of the web calculator shown allows the user to input relevant feature values ​​(LPKM-SSB) in the interface, and the background automatically calculates and displays the prediction results and radar chart of colorectal cancer-related risks; the radar chart shows the performance of multiple indicators of the model in the comparison of different groups in the validation cohort, including AUC, balanced accuracy, F1 score, specificity, sensitivity, Kappa coefficient and accuracy, and the colored lines in the chart represent different comparison groups.

[0092] See also Figure 8 FIG. 1 is a system structure diagram of an embodiment of the present invention, including:

[0093] The collection module 801 collects peripheral blood samples and clinical data to obtain metabolic characteristics and DNA single-strand break characteristics;

[0094] The first screening module 802 performs a statistical preliminary screening of the DNA single-strand break features to screen out the first DNA single-strand break features;

[0095] The second screening module 803 performs a correlation analysis between the first DNA single-strand break feature and the metabolic feature to obtain the metabolism-related DNA single-strand break feature; combines the Lasso regression model and 10-fold cross validation to analyze the metabolism-related DNA single-strand break feature to obtain the final MSSB feature set for developing the machine learning model;

[0096] A training module 804 uses a machine learning algorithm to build a prediction model and performs training based on the final MSSB feature set;

[0097] The prediction module 805 uses the trained prediction model to predict colorectal cancer.

[0098] It can be seen that the present invention constructs a colorectal cancer-related lesion screening model based on multi-omics features with strong generalization ability and high accuracy. It has relatively low cost and high population acceptance, and has important clinical application value.

[0099] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for predicting colorectal cancer based on the characteristics of metabolism-related DNA single-strand breaks, characterized in that: The following steps are involved: Ethylenediaminetetraacetate-anticoagulated peripheral blood samples and clinical data were collected to obtain the metabolic characteristics of peripheral blood plasma and the DNA single-strand break characteristics of peripheral blood cells; Performing statistical preliminary screening on the DNA single-strand break characteristics, and obtaining the first DNA single-strand break characteristics; Performing a correlation analysis on the first DNA single-strand break feature and the metabolic feature to obtain the metabolism-related DNA single-strand break feature; performing feature screening on the metabolism-related DNA single-strand break feature to obtain a final MSSB feature set for training and developing a machine learning model; Use machine learning algorithms to build prediction models and train them based on the final MSSB feature set; Use the trained prediction model to predict colorectal cancer.

2. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: The statistical preliminary screening of the DNA single-strand break characteristics to screen out the first DNA single-strand break characteristics comprises the following steps: The LPKM-SSB value of each gene in peripheral blood samples was calculated as the DNA single-strand break feature and expressed as: LPKM-SSBi = (Ri * 10^5) / (total number of SSB positions in the sample); Among them, LPKM-SSBi represents the LPKM - SSB value of SSB of gene i in each sample; Ri represents the number of SSB positions in the entire region of gene i; The F1 score of DNA single-strand break feature for diagnosing colorectal cancer was calculated, and DNA single-strand break features with F1 score ≥ 0.85 were selected for one-way ANOVA. DNA single-strand break features with p ≤ 0.05 were selected as the first DNA single-strand break feature. The p value indicated the direct difference between normal and precancerous lesion or tumor groups in one-way ANOVA.

3. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: The correlation analysis of the first DNA single-strand break feature and the metabolic feature is performed to obtain the metabolism-related DNA single-strand break feature, specifically: using the O2PLS method to construct a correlation model between the metabolic feature and the first DNA break feature, and selecting several DNA single-strand break features with the highest correlation as the metabolism-related DNA single-strand break features.

4. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: The metabolism-related DNA single-strand break features are screened by using a LASSO regression model with a 10-fold cross-validation optimized LASSO regularization parameter λ, or by using a random forest importance scoring algorithm.

5. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: The method of using a machine learning algorithm to construct a prediction model and training it based on the final MSSB feature set includes the following steps: Build multiple predictive models using a variety of machine learning algorithms; Based on the final MSSB feature set, the optimal MSSB hyperparameters of each model were determined through 10-fold internal cross-validation; Each model is retrained using the final MSSB feature set and the MSSB optimal hyperparameters; all trained models are evaluated, and the model with the best prediction effect is selected as the prediction model.

6. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 5, characterized in that: The various machine learning algorithms include: logistic regression, K-nearest neighbor, Extra Trees classifier, random forest, extreme gradient boosting, light gradient boosting machine, support vector machine and neural network.

7. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 5, characterized in that: The evaluation indicators include: ROC curve, accuracy, Kappa coefficient, precision, and negative prediction value.

8. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: For the trained model, the SHAP method is also used to rank the importance of input features and explain the results of the prediction model.

9. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: The trained prediction model is an MSSB-KNN model constructed using a K-nearest neighbor algorithm. The MSSB-KNN model is predicted based on the LPKM-SSB values ​​of eight genes, including ABCB5, AC093642.5, AF146191.4, ANO3, CTD.2231H16.1, PRKG1, RASA3 and VIPR2. The prediction process of the MSSB-KNN model includes the following steps: Calculate the LPKM-SSB values ​​of the eight genes in the peripheral blood samples of the subjects and input them into the MSSB - KNN model; The MSSB-KNN model outputs the risk probability of colorectal cancer or colorectal precancerous lesions.

10. A colorectal cancer prediction system based on metabolism-related DNA single-strand break characteristics, characterized in that: The method for predicting colorectal cancer based on the characteristics of metabolism-related DNA single-strand breaks according to any one of claims 1 to 9 comprises: The collection module collects EDTA-anticoagulated peripheral blood samples and clinical data to obtain the metabolic characteristics of peripheral blood plasma and the DNA single-strand break characteristics of peripheral blood cells; The first screening module performs a statistical preliminary screening of the DNA single-strand break characteristics to obtain the first DNA single-strand break characteristics; The second screening module performs a correlation analysis between the first DNA single-strand break feature and the metabolic feature to obtain the metabolism-related DNA single-strand break feature; performs feature screening on the metabolism-related DNA single-strand break feature to obtain a final MSSB feature set for training and developing a machine learning model; The training module uses machine learning algorithms to build prediction models and performs training based on the final MSSB feature set; The prediction module uses the trained prediction model to predict colorectal cancer.

Citation Information

Patent Citations

  • Application of reagent for detecting expression level of 40 biomarkers in sample in preparation of kit for evaluating colorectal cancer risk

    CN115074446A

  • Colorectal cancer prediction system and method

    CN119339947A

  • Plasma proteome-based colorectal cancer precancerous lesion diagnosis model construction method

    CN119724556A

  • MALMPS model for identifying II / III stage colorectal cancer patients with high-risk recurrence and application

    CN119763844A

  • Colorectal cancer risk assessment

    WO2024138243A1

Cited By

  • Parkinson's disease screening model construction method based on expiration metabolites, screening method and related equipment

    CN122050857A