Colorectal cancer prediction method and system based on metabolism-related DNA single-strand break characteristics

By collecting the metabolism and DNA single-strand break characteristics of peripheral blood samples, a machine learning model was constructed to predict colorectal cancer, solving the problems of low sensitivity, high invasiveness and high cost in the existing technology, and achieving high accuracy and low cost colorectal cancer screening.

CN119943160BActive Publication Date: 2025-08-12THE FIRST AFFILIATED HOSPITAL OF XIAMEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510440364.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-08-12
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing colorectal cancer screening methods have problems such as low sensitivity, high invasiveness, lack of multiomic integration, complex data analysis and high cost, and poor generalization ability of predictive models. Traditional methods cannot fully reflect complex physiological and pathological processes.

Method used

By collecting peripheral blood samples with ethylenediaminetetraacetate anticoagulation, metabolic characteristics and DNA single-strand break characteristics were obtained, statistical preliminary screening and correlation analysis were performed, metabolic-related DNA single-strand break characteristics were screened out, machine learning models were constructed for colorectal cancer prediction, and a variety of algorithms were used for feature screening and model training, and finally prediction was used using the K-nearest neighbor algorithm.

Benefits of technology

It improves the prediction accuracy of colorectal cancer and precancerous lesions, provides non-invasive detection methods, reduces detection costs, enhances the generalization ability and clinical applicability of the model, and is suitable for large-scale screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943160B_ABST
    Figure CN119943160B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics. The method includes: obtaining metabolic characteristics and DNA single-strand break characteristics; performing statistical preliminary screening on the DNA single-strand break characteristics to screen out the first DNA single-strand break characteristics; performing correlation analysis on the first DNA single-strand break characteristics and the metabolic characteristics to obtain metabolism-related DNA single-strand break characteristics; analyzing the metabolism-related DNA single-strand break characteristics in combination with the Lasso regression model to obtain a final metabolism-related DNA single-strand break (MSSB) feature set; constructing a prediction model and training it based on the final MSSB feature set; and using the trained prediction model to predict colorectal cancer risk. The present invention constructs a colorectal cancer-related lesion screening model based on multi-omics features with strong generalization ability and high accuracy. The cost is relatively low and the population acceptance is high, which has important clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to biomedical detection technology, and in particular to a method and system for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics. Background Art

[0002] Colorectal cancer (CRC) is one of the malignant tumors with the highest morbidity and mortality rates worldwide. Traditional colorectal cancer screening methods such as stool testing and colonoscopy have low sensitivity for detecting precancerous lesions (PL) and require invasive procedures. Non-invasive biomarkers, such as carcinoembryonic antigen (CEA) and Septin 9 gene methylation testing, provide alternatives, but their diagnostic accuracy and sensitivity are still limited. Current detection methods have many shortcomings: (1) Limitations of traditional screening methods: Stool testing and colonoscopy are standard screening methods, but fecal occult blood tests have low sensitivity for early precancerous lesions (PL), and colonoscopy is invasive and has the risk of missed detection, causing discomfort and potential risks to patients. (2) Poor performance of existing biomarkers: Non-invasive biomarkers such as carcinoembryonic antigen (CEA) and septin 9 gene promoter methylation have limited diagnostic accuracy. CEA has a sensitivity of only 30-40% for early CRC, and septin 9 methylation detection has low specificity (AUC = 0.65-0.70). (3) Lack of multi-omics integration and lack of multimodal models: Specific biomarkers can only reflect biological changes at a certain level and lack a comprehensive assessment of the overall physiological state. Different biomarkers may have complex relationships that are independent or interactive with each other, and a single type of marker cannot fully reflect complex physiological and pathological processes. In the scenario of this patent, this is reflected in the fact that a single omics model cannot capture the metabolic-genomic interaction in the occurrence of CRC. (4) Complex data analysis and high cost: The various biomarker detection methods and analysis processes in the existing technology are complex and expensive, which limits their application in clinical and large-scale populations. In particular, some high-throughput omics technologies, such as epigenetics and transcriptomics, although they can provide rich information, their high cost and technical barriers hinder their widespread application. (5) Simple prediction models and poor generalization ability: Currently, many colorectal cancer risk prediction models are based on simple linear regression. Regardless of whether they are statistical methods or machine learning methods, these models often have significant errors when processing high-dimensional omics data. First, simple linear regression models perform poorly when the relationship between features and disease diagnosis is complex. Omics data often contain a large number of features, and linear regression models may not be able to capture the nonlinear relationships and interactions between these features. Simple linear regression is sensitive to feature collinearity, which often leads to model instability and increased prediction errors in omics data. It may not provide sufficiently accurate predictions when dealing with complex datasets. Summary of the Invention

[0003] The purpose of the present invention is to solve the problems in the prior art.

[0004] The technical solution adopted by the present invention to solve the technical problem is to provide a method for predicting colorectal cancer based on the characteristics of metabolism-related DNA single-strand breaks, comprising the following steps:

[0005] Ethylenediaminetetraacetic acid (EDTA)-anticoagulated peripheral blood samples and clinical data were collected to obtain metabolic characteristics of peripheral blood plasma and DNA single-strand break characteristics of peripheral blood cells;

[0006] Perform statistical preliminary screening on DNA single-strand break characteristics to obtain the first DNA single-strand break characteristics;

[0007] performing a correlation analysis on the first DNA single-strand break feature and the metabolic feature to obtain a metabolism-related DNA single-strand break feature; performing feature screening on the metabolism-related DNA single-strand break feature to obtain a final MSSB feature set for training and developing a machine learning model;

[0008] Use machine learning algorithms to build prediction models and train them based on the final MSSB feature set;

[0009] Use the trained prediction model to predict colorectal cancer.

[0010] Preferably, the statistical preliminary screening of DNA single-strand break characteristics to screen out the first DNA single-strand break characteristics comprises the following steps:

[0011] The LPKM-SSB value of each gene in peripheral blood samples was calculated as the DNA single-strand break feature and expressed as:

[0012] LPKM-SSBi=(Ri * 10^5) / (total number of SSB positions in the sample);

[0013] Where LPKM-SSBi represents the LPKM - SSB value of SSB of gene i in each sample; Ri represents the number of SSB positions in the entire region of gene i;

[0014] The F1 score of DNA single-strand break features for diagnosing colorectal cancer was calculated. DNA single-strand break features with an F1 score ≥ 0.85 were selected for one-way analysis of variance. DNA single-strand break features with p ≤ 0.05 were selected as the first DNA single-strand break feature. The p value represents the direct difference between the normal and precancerous lesion or tumor groups in the one-way analysis of variance.

[0015] Preferably, the correlation analysis of the first DNA single-strand break feature and the metabolic feature to obtain the metabolism-related DNA single-strand break feature is specifically: using the O2PLS method to construct a correlation model between the metabolic feature and the first DNA break feature, and selecting several DNA single-strand break features with the highest correlation as the metabolism-related DNA single-strand break features.

[0016] Preferably, the metabolism-related DNA single-strand break features are subjected to feature screening, and a LASSO regression model with a 10-fold cross-validation-optimized LASSO regularization parameter λ is used for feature screening, or a random forest importance scoring algorithm is used for feature screening.

[0017] Preferably, the method of constructing a prediction model using a machine learning algorithm and training the model based on the final MSSB feature set comprises the following steps:

[0018] Build multiple predictive models using a variety of machine learning algorithms;

[0019] Based on the final MSSB feature set, the optimal MSSB hyperparameters of each model were determined through 10-fold internal cross-validation;

[0020] Each model is retrained using the final MSSB feature set and the MSSB optimal hyperparameters; all trained models are evaluated, and the model with the best prediction effect is selected as the prediction model.

[0021] Preferably, the multiple machine learning algorithms include: logistic regression, K-nearest neighbor, Extra Trees classifier, random forest, extreme gradient boosting, light gradient boosting machine, support vector machine and neural network.

[0022] Preferably, the indicators used in the evaluation include: ROC curve, accuracy, Kappa coefficient, precision, and negative predictive value.

[0023] Preferably, the SHAP method is also used to rank the importance of input features for the trained model to explain the results of the prediction model.

[0024] Preferably, the trained prediction model is an MSSB-KNN model constructed using a K-nearest neighbor algorithm. The MSSB-KNN model performs prediction based on the LPKM-SSB values of eight genes, including ABCB5, AC093642.5, AF146191.4, ANO3, CTD.2231H16.1, PRKG1, RASA3, and VIPR2. The prediction process of the MSSB-KNN model includes the following steps:

[0025] Calculate the LPKM-SSB values of the eight genes in the peripheral blood samples of the subjects and input them into the MSSB-KNN model;

[0026] The MSSB-KNN model outputs the risk probability of colorectal cancer or colorectal precancerous lesions.

[0027] The present invention also provides a colorectal cancer prediction system based on metabolism-related DNA single-strand break characteristics, which is used to implement any of the above-mentioned colorectal cancer prediction methods based on metabolism-related DNA single-strand break characteristics, comprising:

[0028] The collection module collects EDTA-anticoagulated peripheral blood samples and clinical data to obtain the metabolic characteristics of peripheral blood plasma and the DNA single-strand break characteristics of peripheral blood cells;

[0029] The first screening module performs statistical preliminary screening on the DNA single-strand break characteristics to obtain the first DNA single-strand break characteristics;

[0030] The second screening module performs a correlation analysis between the first DNA single-strand break feature and the metabolic feature to obtain metabolism-related DNA single-strand break features; and performs feature screening on the metabolism-related DNA single-strand break features to obtain a final MSSB feature set for training and developing a machine learning model;

[0031] The training module uses machine learning algorithms to build prediction models and trains them based on the final MSSB feature set;

[0032] The prediction module uses the trained prediction model to predict colorectal cancer.

[0033] The present invention has the following beneficial effects:

[0034] (1) Breakthrough in diagnostic performance: In multicenter validation, the model achieved an AUC of 0.866 (95% CI: 0.797-0.934) for colorectal cancer, significantly outperforming the traditional combined detection of CEA and septin 9 methylation (AUC: 0.708). In particular, in the detection of precancerous lesions, the AUCs of two independent cohorts reached 0.873 and 0.916, respectively, demonstrating its early warning value and improving the accuracy of prediction for colorectal cancer and precancerous lesions.

[0035] (2) Enhanced clinical applicability: Only 1 mL of peripheral blood is required to complete the test, providing a non-invasive detection method, avoiding invasive procedures, reducing patient pain, and being suitable for large-scale screening. Furthermore, the test demonstrated excellent diagnostic performance in two tertiary hospitals, indicating a promising prospect for promotion and application.

[0036] (3) Biomarker advantages: The discovered metabolism-related DNA breakage biomarkers have high sensitivity and specificity for CRC and PL, providing more effective non-invasive indicators for disease detection and more accurately reflecting the disease status.

[0037] (4) Accurate prediction by machine learning model: Through the optimization of machine learning model, more accurate prediction results can be provided, which has higher accuracy and predictive ability than existing biomarker detection methods.

[0038] (5) During the model training process, eight patients with other cancer types were added to the control group, making the MSSB_KNN model have a certain degree of cancer type specificity. Further training of the model using relevant data can support the adaptation of multiple cancer types such as breast cancer and gastric cancer.

[0039] (6) Cost reduction: The model ultimately used for clinical application only needs to detect SSB signals, and the cost of a single sample test is less than 500 yuan (Septin9 charges, CEA charges).

[0040] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments, but the present invention is not limited to the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A diagram showing the steps of a method according to an embodiment of the present invention;

[0042] Figure 2 A schematic diagram of a flow chart of an embodiment of the present invention;

[0043] Figure 3 This is a metabolite analysis diagram of an embodiment of the present invention;

[0044] Figure 4 This is a process diagram of feature selection for a method for predicting colorectal cancer and early stage lesions based on metabolism-related DNA breaks according to an embodiment of the present invention;

[0045] Figure 5 This is a prediction nomogram using three models or clinical features, MSSB_KNN, MSSB_LR, and MSSB_NN, in a validation cohort according to an embodiment of the present invention.

[0046] Figure 6 2. ROC curve comparison diagram of the embodiment of the present invention (including CEA, septin 9 methylation and MSSB_KNN model);

[0047] Figure 7 This is a web calculator interface according to an embodiment of the present invention.

[0048] Figure 8 2 is a system structure diagram of an embodiment of the present invention. DETAILED DESCRIPTION

[0049] See also Figure 1 and Figure 2 FIG. 1 is a diagram showing steps and a flow chart of a method according to an embodiment of the present invention, which includes the following steps:

[0050] S101, collect EDTA-anticoagulated peripheral blood samples and clinical data to obtain metabolic characteristics of peripheral blood plasma and DNA single-strand break characteristics of peripheral blood cells;

[0051] S102, performing preliminary statistical screening on the DNA single-strand break characteristics to obtain a first DNA single-strand break characteristic;

[0052] S103, performing a correlation analysis between the first DNA single-strand break feature and the metabolic feature to obtain a metabolism-related DNA single-strand break feature; performing feature screening on the metabolism-related DNA single-strand break feature to obtain a final MSSB feature set for training and developing a machine learning model;

[0053] S104, constructing a prediction model using a machine learning algorithm and training it based on the final MSSB feature set;

[0054] S105, using the trained prediction model to predict colorectal cancer.

[0055] Specifically, the S101 includes:

[0056] S1011, Collection of clinical data: Clinical data were extracted from the participants' inpatient electronic medical records (EMRs), covering baseline patient information, including age, sex, laboratory test results (Septin 9 gene methylation status, CEA level) within one month before endoscopy or surgery, colonoscopy and pathological diagnosis, as well as TNM stage (T stage, N stage, M stage) and clinical classification.

[0057] At S1012, peripheral blood samples were collected, plasma was extracted for metabolomics analysis, and DNA from peripheral blood cells was extracted for fragmentation analysis.

[0058] S1013, Metabolomics Analysis: Plasma samples were obtained from EDTA-treated blood for metabolomics analysis. The following steps were performed: Plasma samples were thawed, vortexed, and aliquoted into 1.5 mL centrifuge tubes. Each sample was aliquoted into six aliquots (30 µL per channel) for multichannel analysis. To prepare a pooled sample, 75 µL of serum from each sample was combined and vortexed to form a reference sample. To precipitate proteins, 90 µL of pre-chilled mass spectrometry-grade methanol was added to each sample. After vortexing, the sample was centrifuged at 12,000 rpm for 10 minutes at 4°C. The supernatant was transferred to a fresh tube and dried under nitrogen. For carboxyl metabolome analysis, 25 µL of a 3:1 volume ratio of acetonitrile / water was added to the sample, followed by 10 µL of a reaction catalyst and 25 µL of a carbon-12 or carbon-13 labeling reagent. The sample was incubated at 80°C for 60 minutes. Following incubation, 40 µL of a quenching reagent was added, and the mixture was incubated for an additional 30 minutes at the same temperature. Ultra-high-performance liquid chromatography-quadrupole time-of-flight mass spectrometry (UHPLC Q-TOF / MS) was used for chromatographic separation on an Agilent Eclipse Plus reversed-phase C18 column (150 × 2.1 mm, 1.8 µm). The mobile phase consisted of 0.1% formic acid in water (A) and 0.1% formic acid in acetonitrile (B). The gradient program started at 25% B, increased to 99% B over 10 minutes and held for 5 minutes, and then returned to 25% B at 15.1 minutes and held for 18 minutes. The column temperature was 40°C, and the flow rate was 400 µL / min. The ion source parameters were: gas temperature 325°C; drying gas 8 L / min; nebulizer pressure 35 psi; sheath gas temperature 400°C; sheath gas flow rate 12 L / min; and the MS scan range was m / z 220–1000.

[0059] S1014, DNA breakage profiling: Whole-genome DNA fragmentation analysis was performed on 200 μL of peripheral blood cell DNA using the single-strand break mapping at the nucleotide level (SSiNGLe) method. DNA was extracted according to the manufacturer's instructions for the DNA extraction kit, and 100 ng of DNA was used directly for SSiNGLe library preparation. All DNA fragmentation profiling experiments were performed by sequencing companies on the Illumina / BGI platform using a paired-end 150 bp strategy with a sequencing depth of 2-5 GB.

[0060] S1015, Feature Calculation: Calculate the LPKM - SSB value representing the SSBs for gene i in each sample using the formula LPKM - SSB = (Ri * 10^5) / (total number of SSB positions in the sample), where Ri is the number of SSB positions located in the entire region of gene i. Calculate the relative abundance (RA) of SSBs in each sample using the formula RA = (SSB count in the specified region of the current sample) / ((total SSB count for all samples in the current cohort) / (total number of samples)).

[0061] Specifically, in S1014, the library construction using the SSiNGLe method includes the following steps:

[0062] S201, preliminary preparation, take out the magnetic beads and place them at room temperature for equilibrium.

[0063] S202, PolyA tailing, including:

[0064] S2021, denature. If the DNA quality is sufficient, add the following system: add 2 μL of TdT10x buffer, 2 μL of 2.5 mM CoCl2, 10 μL of DNA (100 ng), and 5 μL of water to the reaction system for a total volume of 19 μL. If the DNA quality is insufficient, omit the water. After addition, place the reaction system in a C1000 touch PCR instrument and denature at 95°C for 5 minutes. Immediately place the reaction system on ice.

[0065] S2022, tailing; When preparing the tailing system, please note: prepare an additional aliquot for every three samples to prevent insufficient sample. The tailing system contains 1 μL of NEB terminal transferase diluted to 4 U / μL (1:4 dilution) in 1x buffer and 2 μL of 10 mM dATP. After preparing the system, vortex and mix for 10 seconds, then centrifuge for 10 seconds. Then, aliquot 3 μL of this system into the 19 μL denaturing system in the previous step, for a total volume of 22 μL. Place the mixed system in a C1000 touch PCR instrument and incubate at 37°C for 30 minutes, then store at 4°C.

[0066] S2023, blocking; add 2 μL of 10 mM ddNTPs to the tailed system above to a final volume of 24 μL. This system was placed in a C1000 touch PCR instrument and reacted under the following conditions: incubation at 37°C for 30 min, reaction at 70°C for 10 min, and storage at 4°C.

[0067] S2024, magnetic bead purification: Add twice the final volume (i.e., 48 μL) of magnetic beads to the blocked system for purification. Purification was performed using the BGI MGISP-960 library preparation instrument. Freshly prepared 80% ethanol was used during purification, and the final elution was performed with 16 μL of ultrapure water.

[0068] S203, amplification reaction (using hot-start Taq enzyme and DNA-RNA oligo-d(T)50-r(T)3 oligonucleotide primer); the reaction system contains 0.4 μL rTaq (1U, Tiangen), 1 μL 10 μM oligo-d(T)50-r(T)3 primer, 2 μL 10 x Buffer, 1.6 μL 2.5 mM dNTP mix, and 15 μL of DNA purified in the previous step, with a total volume of 20 μL. The reaction system was placed in a C1000 touch PCR instrument for reaction. The reaction conditions included: (1) 94℃ for 30 seconds; (2) 94℃ for 1 minute; (3) 55℃ for 30 seconds; (4) 72℃ for 30 seconds; repeat (2) to (4) for a total of 10 cycles; after the reaction is completed, the reaction system was immediately placed on ice. Magnetic bead purification was performed using 2 volumes (i.e., 40 μL) of magnetic beads and eluted with 16 μL of ultrapure water.

[0069] S204, Poly C tailing; includes:

[0070] S2041: Denaturation; The reaction system contains 2 μL of TdT 10x buffer, 2 μL of 2.5 mM CoCl2, and 15 μL of DNA eluted from the previous step, for a total volume of 19 μL. The reaction system is placed in a C1000 touch PCR instrument and denatured at 95°C for 5 minutes. The reaction system is then immediately placed on ice.

[0071] S2042, tailing; add 1 μL of NEB TdT diluted to 5 U / μL (1:4 dilution) in 1x buffer and 2 μL of 10 mM dCTP to the denatured system above, bringing the total volume to 22 μL. The system was placed in a C1000 touch PCR instrument and incubated at 37°C for 30 minutes, followed by storage at 4°C.

[0072] S2043, magnetic bead purification; add 2 volumes (i.e., 44 μL) of magnetic beads and use the BGI MGISP-960 library preparation instrument for purification. Use freshly prepared 80% alcohol for purification and finally elute with 15 μL of ultrapure water.

[0073] S205, amplification reaction (using P5G10 and P7T12 primers); the reaction system contained 0.4 μL of rTaq alone (1 U, Tiangen), 1 μL of 10 μM P5G10 primer, 1 μL of 10 μM P7T12 primer, 1.6 μL of 2.5 mM dNTP mix, 2 μL of 10 x Buffer, and 14 μL of DNA eluted in the previous step, for a total volume of 20 μL. The reaction system was placed in a C1000 touch PCR instrument for reaction. The reaction conditions included: (1) 94℃ for 3 min; (2) 94℃ for 30 sec; (3) 55℃ for 1 min; (4) 72℃ for 1 min; (5) 94℃ for 30 sec; (6) 37℃ for 1 min; (7) 37℃ + 2℃ / 2 min (cycle); (8) repeat step (7) for a total of 17 cycles; (9) 72℃ for 2 min; (10) 94℃ for 30 sec; (11) 60℃ for 30 sec, for a total of 30 cycles; (12) 72℃ for 30 sec; (13) 72℃ for 7 min; and (14) 4℃ hold. For magnetic bead purification, add 2 times the volume (i.e., 40 μL) of magnetic beads and use the BGI MGISP-960 library preparation instrument for purification. Use freshly prepared 80% alcohol for purification and finally elute with 14 μL of ultrapure water.

[0074] S206, amplification reaction (using Illumina_P5 and Illumina_P7 primers); the total reaction volume was 20 μL, containing 0.4 μL rTaq, 1 μL 10 μM Illumina_P5 primer, 1 μL 10 μM Illumina_P7 primer, 1.6 μL 2.5 mM dNTP mix, 2 μL 10x buffer, and 14 μL DNA. Reaction conditions were 94°C for 3 min, 94°C for 30 sec, 55°C for 30 sec (6 cycles for cell samples; 10 cycles for blood samples; 12 cycles for tissue samples), 72°C for 30 sec, 72°C for 7 min, and a 4°C hold. Magnetic bead purification was performed by adding 2 volumes (i.e., 40 μL) of magnetic beads and using a BGI MGISP-960 library preparation instrument. Purification was performed using freshly prepared 80% ethanol and eluted with 22 μL of ultrapure water.

[0075] S207, concentration determination and storage; take 1 μL of the eluted product and use Qubit to determine the concentration, and freeze the remaining product at -20°C for library quality control and sequencing.

[0076] Specifically, in the present embodiment, the clinical data used by S1011 were derived from the participants' inpatient electronic medical records (EMRs), covering the patients' baseline information. Specifically, variables collected included age, sex, laboratory test results (septin 9 gene methylation status, CEA levels) from the month prior to endoscopic or surgical procedures, colonoscopy and pathological diagnosis results, TNM staging (T stage, N stage, M stage), and clinical grade. Clinical characteristics were compared statistically across study cohorts, including key variables such as age, sex distribution, disease stage, and biomarker levels, to assess consistency and variability between cohorts. For metabolomics analysis, residual plasma samples collected from EDTA-treated blood during routine clinical testing were used, and DNA fragmentation analysis was performed on peripheral blood cells.

[0077] Specifically, in the embodiment of the present invention, S1013 also performed metabolomics mass spectrometry difference analysis and differential metabolite enrichment analysis, and used the multivariate statistical analysis (OPLS-DA) method to analyze the plasma metabolomics mass spectrometry results. In addition, univariate statistical analysis (T test) was performed to identify differential metabolites between NC, PL and CRC groups, and the screening criteria were VIP score ≥ 1 and p value < 0.05. KEGG pathway enrichment analysis was further performed to explore the regulatory pathways related to differential metabolites. By obtaining differential metabolites between groups, alternative metabolites were obtained for subsequent joint analysis of metabolism and SSB. The analysis results are shown in Figure 3 As shown in the Venn diagram ( Figure 3 (A) and (D) in Figure 3 show the up-regulated and down-regulated metabolites in normal control (NC) compared with precancerous lesions (PL), and normal control (NC) compared with colorectal cancer (CRC); the violin plots ( Figure 3 (B) and (C) show representative upregulation, as shown by violin plots ( Figure 3 (E) and (F) in Figure 3 show down-regulated metabolites. Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis of differential metabolites between normal control (NC) and colorectal cancer (CRC), and normal control (NC) and precancerous lesions (PL) was performed. Figure 3 (G) and (H) in Figure 1. NC denotes non-CRC / PL control; PL denotes precancerous lesions; and CRC denotes colorectal cancer.

[0078] Specifically, the S1015 analyzes the SSiNGLe data and calculates the Ri value of each gene in each sample, thereby calculating LPKM-SSB (representing the number of SSBs of gene i in each sample) using the following formula: LPKM-SSB=(Ri*10^5) / (the total number of positions of all SSBs in the sample);

[0079] Where Ri is the number of SSB positions within the entire region of gene i. The relative abundance (RA) of SSBs in each sample is calculated as follows: RA = (SSB count in a given region of the current sample) / ((Total number of SSBs across all samples in the current cohort) / (Total number of samples)). Calculating RA primarily examines the distribution of DNA breakage severity across different subgroups, reflecting the biological characteristics of DNA breakage.

[0080] For details, see Figure 4 As shown, in S102, the F1_Score function in R was used to calculate and select SSB genes with F1 scores greater than or equal to 0.85. A one-way analysis of variance (ANOVA) was further performed using the CreateTableOne function to select SSB genes with p-values less than or equal to 0.05. The Lasso regression model was constructed using the glmnet function, and a 10-fold cross validation (CV) was performed using the cv.glmnet function to select the optimal λ, thereby obtaining the final SSB gene set for model construction. Figure 4 (A) and (B) show the association between LPKM-SSB values of DNA single-strand break (SSB) genes and metabolite quantification values, analyzed using a two-way orthogonal partial least squares (O2PLS) model. The loading plots show the strength of the association, with coordinates with larger absolute values representing stronger associations. The top 10 ranked elements are highlighted in red. (C) shows the loading matrix of SSB genes and metabolites. Elements closer to the outer circle have higher correlations between the two omics data sets. (D) shows the variation of the regression coefficients in the LASSO regression as a function of Log(λ). The importance of each variable is determined by the size of its coefficient. (E) shows the mean squared error (MSE) from 10-fold cross-validation, indicating the optimal model by plotting the MSE as a function of Log(λ). The red lambda.min line indicates the value corresponding to the minimum MSE, and the gray lambda.1se line indicates a simpler model. (F) Highlights the feature importance assessed by LASSO regression, and the formula shows how the LPKM-SSB value and weight of each gene affect the final LASSO score.

[0081] Specifically, in S103, the O2PLS (two-way orthogonal partial least squares) method was applied to simultaneously model and predict paired samples of plasma metabolites and peripheral blood cell SSB in the training cohort, revealing the relationship between the metabolomics and DNA fragmentation datasets, quantifying the correlation between the two omics, and identifying key SSB genes (MSSB) associated with them. The top 25 MSSB genes associated with metabolites were selected, and a one-way analysis of variance was performed using the CreateTableOne function. Genes with a p-value ≤ 0.05 were selected, and a Lasso regression model was constructed using the glmnet function. Subsequently, a 10-fold cross-validation was performed using the cv.glmnet function to determine the optimal λ, resulting in the final MSSB gene set for model development. The results of the top 25 metabolism-related SSB genes obtained in the examples of the present invention are shown in the following table, where n in the table represents the number of samples in each group.

[0082]

[0083] For details, see Figure 2 As shown, the S104 includes:

[0084] S1041, Model Construction: Eight machine learning (ML) algorithms, including logistic regression (LR), K-nearest neighbor (KNN), ExtraTrees classifier (ExtraTrees), random forest (RF), extreme gradient boosting (XGB), light gradient boosting machine (LGBM), support vector machine (SVM), and neural network (NN), were used to predict the likelihood of PL or CRC. Based on the optimal feature subset, the optimal hyperparameters for each algorithm were determined through 10 rounds of 10-fold cross-validation and a grid search using the default hyperparameters of the "caret" package. Finally, the model was retrained using the selected feature subset and the final hyperparameters derived from 10 rounds of 10-fold internal cross-validation. This resulted in an MSSB-KNN model based on eight genes: ABCB5, AC093642.5, AF146191.4, ANO3, CTD.2231H16.1, PRKG1, RASA3, and VIPR2.

[0085] S1042, Model Validation and Comparison: Model effectiveness was assessed using a variety of standard performance metrics, such as area under the receiver operating characteristic curve (AUC), sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and F1 score. Calibration curves were used to reflect the agreement between predicted probabilities and observed outcomes, and decision curve analysis (DCA) was used to assess the net benefit of the model at various thresholds. The optimal prediction model was selected based on the evaluation metrics in the training and validation cohorts. The DeLong test was used to determine whether the AUC values of different models differed significantly. Log-loss was calculated to measure the difference between the actual label and the predicted probability to assess predictive accuracy.

[0086] S1043, Model Interpretation: Use the game theory-based SHapley Additive exPlanation (SHAP) method to interpret machine learning models, rank the importance of input features and clarify the results of the prediction model. By calculating the contribution of each feature to the prediction result, local and global explanations are provided to enhance the transparency and interpretability of the model and determine the practical significance of the model in predicting CRC / PL.

[0087] Specifically, in S1042, statistical calculation of diagnostic performance indicators is performed for all machine learning (ML) models, CEA, septin 9 methylation, and CEA_plus_septin 9 methylation (where the diagnostic criteria for CEA_plus_septin 9 methylation is defined as a positive diagnosis if any of the indicators is positive). In particular, when evaluating diagnostic performance, the clinical dataset can be stratified by cofactors such as age and gender to reflect the impact of different cofactors on diagnostic performance.

[0088] See also Figure 5 Shown are the prediction nomograms (Nomograms) in the validation cohort using the three models or clinical characteristics of MSSB_KNN, MSSB_LR, and MSSB_NN. (A) is the nomogram of NC vs. CRC / PL in validation cohort 1; (B) is the nomogram of NC vs. PL in validation cohort 1; (C) is the nomogram of NC vs. CRC in validation cohort 1; (D) is the nomogram of NC vs. PL in validation cohort 2. Among them, MSSB_KNN represents the metabolism-related DNA single-strand break model using K-Nearest Neighbors method. Figure 5 The smaller the z statistic / coefficient significance p value is, the better the classification effect of the corresponding model on the target group is ( Figure 5In the table, “*” indicates that the p-value is less than 0.05, “**” indicates that the p-value is less than 0.01, and “***” indicates that the p-value is less than 0.001). Figure 5 It can be seen that the MSSB_KNN classification performance was the best in all cohorts. Specifically, the Delong test was used to determine whether the AUC values of different models differed significantly: the smaller the p-value of the Delong test results for the difference in AUC between different machine learning models, CEA, septin 9 methylation, and CEA_plus_septin 9 methylation prediction methods, the more significant the difference in AUC between the two methods. Furthermore, the logarithmic loss (Log-loss) value was used to evaluate the difference between the actual label (whether or not a patient actually has colorectal cancer or related precancerous lesions) and the predicted probability, thereby measuring the model's predictive accuracy. The lower the Log-loss value, the better the performance.

[0089] Specifically, in S1042, the primary outcome was to predict the diagnosis of colorectal precancerous lesions (PL) or colorectal cancer (CRC). Model performance was evaluated using metrics such as AUC, accuracy, sensitivity, and specificity, with colonoscopy results serving as the gold standard. Secondary outcomes included comparing the model's diagnostic accuracy, sensitivity, and specificity in identifying CRC / PL, and comparing it with baseline tests such as CEA and septin 9 methylation.

[0090] See also Figure 6Figures 1 and 2 show receiver operating characteristic (ROC) curves comparing the present invention's examples with other predictive indicators. (A) to (D) line graphs show the area under the curve (AUC) for differentiating different groups in the training cohort, validation cohort 1, and validation cohort 2. (A) shows NC vs. CRC / PL, (B) shows NC vs. PL, (C) shows NC vs. CRC, and (D) shows PL vs. CRC. (E) to (G) boxplots show the predicted probabilities of CRC or PL samples diagnosed by MSSB-KNN in the training cohort, validation cohort 1, and validation cohort 2, respectively. (H) to (I) demonstrate that the MSSB-KNN model can correctly detect CRC patients who might have been misclassified by CEA or septin 9 methylation. CEA+ or MSSB+ represent participants detected by CEA (cutoff ≥ 5 ng / mL) or MSSB (prediction probability cutoff ≥ 0.5). Grouping: Triangles represent CRC / PL, circles represent NC; septin 9 methylation-positive (cutoff Ct ≤ 41) is red, and negative is blue. Abbreviations: MSSB = Metabolism-related DNA Single-strand Break, KNN = K-Nearest Neighbors, septin 9 = septin 9 methylation, CEA_plus_septin 9 = defined as positive if either of the two indicators is positive, train = training cohort, test1 = validation cohort 1, test2 = validation cohort 2. Figure 6 It shows that using colonoscopy results as the gold standard, the MSSB_KNN model performs better than CEA and septin 9 methylation in distinguishing non-CRC / PL and CRC / PL patients.

[0091] Specifically, the S105 is implemented using a model network calculator: see Figure 7 The web calculator interface shown allows the user to input relevant feature values (LPKM-SSB) in the interface, and the background automatically calculates and displays the prediction results and radar chart of colorectal cancer-related risks; the radar chart shows the performance of multiple indicators of the model in comparing different groups in the validation cohort, including AUC, balanced accuracy, F1 score, specificity, sensitivity, Kappa coefficient and accuracy, and the colored lines in the chart represent different comparison groups.

[0092] See also Figure 8 FIG. 1 is a system structure diagram of an embodiment of the present invention, including:

[0093] The collection module 801 collects peripheral blood samples and clinical data to obtain metabolic characteristics and DNA single-strand break characteristics;

[0094] The first screening module 802 performs a statistical preliminary screening of DNA single-strand break features to screen out the first DNA single-strand break feature;

[0095] The second screening module 803 performs a correlation analysis between the first DNA single-strand break feature and the metabolic feature to obtain the metabolism-related DNA single-strand break feature; combines the Lasso regression model and 10-fold cross-validation to analyze the metabolism-related DNA single-strand break feature to obtain the final MSSB feature set for developing the machine learning model;

[0096] The training module 804 uses a machine learning algorithm to build a prediction model and performs training based on the final MSSB feature set;

[0097] The prediction module 805 uses the trained prediction model to predict colorectal cancer.

[0098] It can be seen that the present invention constructs a colorectal cancer-related lesion screening model based on multi-omics features with strong generalization ability and high accuracy. It has relatively low cost and high population acceptance, and has important clinical application value.

[0099] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for predicting colorectal cancer based on the characteristics of metabolism-related DNA single-strand breaks, characterized in that: The following steps are involved: Ethylenediaminetetraacetic acid (EDTA)-anticoagulated peripheral blood samples and clinical data were collected to obtain metabolic characteristics of peripheral blood plasma and DNA single-strand break characteristics of peripheral blood cells; The LPKM-SSB value of each gene in peripheral blood samples was calculated as the DNA single-strand break feature and expressed as: LPKM-SSBi=(Ri*10^5) / (total number of SSB positions in the sample); Where LPKM-SSBi represents the LPKM-SSB value of SSB of gene i in each sample; Ri represents the number of SSB positions in the entire region of gene i; Performing statistical preliminary screening on DNA single-strand break features to obtain the first DNA single-strand break feature; performing a correlation analysis on the first DNA single-strand break feature and the metabolic feature to obtain a metabolism-related DNA single-strand break feature; performing feature screening on the metabolism-related DNA single-strand break feature to obtain a final MSSB feature set for training and developing a machine learning model; Use machine learning algorithms to build prediction models and train them based on the final MSSB feature set; Use the trained prediction model to predict colorectal cancer; The correlation analysis of the first DNA single-strand break feature and the metabolic feature is performed to obtain the metabolism-related DNA single-strand break feature, specifically: performing metabolomics mass spectrometry difference analysis and differential metabolite enrichment analysis on the metabolic feature to obtain alternative metabolic features; using the O2PLS method to construct a correlation model between the alternative metabolic features and the first DNA break feature, and selecting several DNA single-strand break features with the highest correlation as the metabolism-related DNA single-strand break features; The metabolism-related DNA single-strand break features are screened by using a LASSO regression model with a 10-fold cross-validation-optimized LASSO regularization parameter λ, or a random forest importance scoring algorithm is used for feature screening; The trained prediction model makes predictions based on the LPKM-SSB values of eight genes, including ABCB5, AC093642.5, AF146191.4, ANO3, CTD-2231H16.1, PRKG1, RASA3, and VIPR2.

2. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: The statistical preliminary screening of the DNA single-strand break characteristics to screen out the first DNA single-strand break characteristics comprises the following steps: The LPKM-SSB value of each gene in peripheral blood samples was calculated as a DNA single-strand break signature; The F1 score of DNA single-strand break features for diagnosing colorectal cancer was calculated. DNA single-strand break features with an F1 score ≥ 0.85 were selected for one-way analysis of variance. DNA single-strand break features with p ≤ 0.05 were selected as the first DNA single-strand break feature. The p value represents the direct difference between the normal and precancerous lesion or tumor groups in the one-way analysis of variance.

3. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: The method of using a machine learning algorithm to build a prediction model and training it based on the final MSSB feature set includes the following steps: Build multiple predictive models using a variety of machine learning algorithms; Based on the final MSSB feature set, the optimal MSSB hyperparameters of each model were determined through 10-fold internal cross-validation; Each model is retrained using the final MSSB feature set and the MSSB optimal hyperparameters; all trained models are evaluated, and the model with the best prediction effect is selected as the prediction model.

4. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 3, characterized in that: The various machine learning algorithms include: logistic regression, K-nearest neighbors, Extra Trees classifier, random forest, extreme gradient boosting, light gradient boosting machine, support vector machine and neural network.

5. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 3, characterized in that: The evaluation indicators include: ROC curve, accuracy, Kappa coefficient, precision and negative predictive value.

6. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: For the trained model, the SHAP method is also used to rank the importance of input features and explain the results of the prediction model.

7. The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to claim 1, characterized in that: The trained prediction model is an MSSB-KNN model constructed using the K-nearest neighbor algorithm. The prediction process of the MSSB-KNN model includes the following steps: Calculate the LPKM-SSB values of the eight genes in the peripheral blood samples of the subjects and input them into the MSSB-KNN model; The MSSB-KNN model outputs the risk probability of colorectal cancer or colorectal precancerous lesions.

8. A colorectal cancer prediction system based on metabolism-related DNA single-strand break characteristics, characterized by: The method for predicting colorectal cancer based on metabolism-related DNA single-strand break characteristics according to any one of claims 1 to 7 comprises: The collection module collects EDTA-anticoagulated peripheral blood samples and clinical data to obtain the metabolic characteristics of peripheral blood plasma and the DNA single-strand break characteristics of peripheral blood cells; The first screening module performs a statistical preliminary screening of DNA single-strand break features to obtain the first DNA single-strand break feature; The second screening module performs a correlation analysis between the first DNA single-strand break feature and the metabolic feature to obtain metabolism-related DNA single-strand break features; and performs feature screening on the metabolism-related DNA single-strand break features to obtain a final MSSB feature set for training and developing a machine learning model; The training module uses machine learning algorithms to build prediction models and trains them based on the final MSSB feature set; The prediction module uses the trained prediction model to predict colorectal cancer.

Citation Information

Patent Citations

  • Colorectal cancer prediction system and method

    CN119339947A