Method for predicting peripheral neurotoxicity induced by treatment of colorectal cancer through oxaliplatin

By acquiring patient information and clinical characteristic data, and employing multi-algorithm screening and gradient boosting decision tree models, a personalized peripheral neurotoxicity prediction system was constructed. This system solved the problem of accurately predicting peripheral neurotoxicity in oxaliplatin-treated colorectal cancer, and improved the identification accuracy and treatment efficacy of high-risk patients.

CN120824019APending Publication Date: 2025-10-21AFFILIATED HOSPITAL OF JIANGNAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511257539.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Current technology cannot accurately predict peripheral neurotoxicity caused by oxaliplatin treatment for colorectal cancer, leading to treatment interruption or dosage adjustment and affecting treatment efficacy.

Method used

By acquiring patient information and clinical characteristic data, a core feature dataset is determined through multi-algorithm collaborative screening. A gradient boosting decision tree model is used for prediction, and a personalized peripheral neurotoxicity prediction system is constructed by combining baseline feature analysis and risk correction.

Benefits of technology

It enables accurate prediction of peripheral neurotoxicity in patients treated with oxaliplatin for colorectal cancer, improves the accuracy of identifying high-risk patients, supports personalized treatment plans and early intervention measures, and enhances the reliability and effectiveness of treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120824019A_ABST
    Figure CN120824019A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting peripheral neurotoxicity induced by treatment of colorectal cancer through oxaliplatin, and relates to the technical field of drug application, and the method comprises the following steps: obtaining pre-drug patient information before treatment of a patient with oxaliplatin and clinical feature data after use; processing and feature analysis are carried out based on pre-drug patient information and clinical feature data, and a core feature data set related to neurotoxin is determined; and based on the core feature data set, predicting the oxaliplatin-induced peripheral neurotoxicity of the patient by adopting a pre-trained gradient boosting decision tree model, and outputting a prediction risk result. According to the method, the occurrence risk of peripheral neurotoxicity can be accurately predicted according to clinical indexes so as to guide clinical accurate medication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of drug application technology, and in particular to a method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer. Background Art

[0002] Oxaliplatin is a first-line chemotherapy drug for colorectal cancer and plays an important role in tumor treatment. However, its treatment is often associated with significant adverse reactions, particularly oxaliplatin-induced peripheral neurotoxicity (OIPN), a common adverse reaction manifesting as numbness, tingling, paresthesia, and functional impairment in the limbs. OIPN can severely impact patients' quality of life, treatment tolerance, and long-term prognosis. Because its occurrence is closely related to drug dosage, duration of use, and individual variability, it often necessitates treatment interruption or dose adjustment, potentially compromising tumor treatment efficacy.

[0003] Clinical management of OIPN mainly relies on empirical judgment or symptomatic treatment after toxicity occurs, which has poor effects on one side and cannot achieve early prevention and precise intervention.

[0004] Therefore, how to provide a predictive technology that can accurately predict the risk of peripheral neurotoxicity (OIPN) based on clinical indicators to guide clinical precision drug use has outstanding economic value and significance. Summary of the Invention

[0005] In order to provide a method that can accurately predict the risk of peripheral neurotoxicity (OIPN) based on clinical indicators to guide clinical precision medication, this application provides a method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer.

[0006] In the first aspect, the invention objectives of this application are achieved by adopting the following technical solutions: Methods for predicting oxaliplatin-induced peripheral neurotoxicity in colorectal cancer include: Obtain pre-drug patient information before oxaliplatin treatment and clinical characteristics data after treatment; Performing processing and feature analysis based on the pre-drug patient information and the clinical feature data to determine a core feature data set related to neurotoxins; Based on the core feature data set, a pre-trained gradient boosting decision tree model is used to predict the occurrence of oxaliplatin-induced peripheral neurotoxicity in patients and output a predicted risk result.

[0007] By adopting the above-mentioned technical solution, data preprocessing includes missing value imputation and variable transformation. This application provides a prediction technology that can effectively and accurately predict the risk of peripheral neurotoxicity in patients with colorectal cancer receiving oxaliplatin treatment, helping to guide clinical selection of personalized treatment plans and the implementation of early intervention measures for neurotoxicity. Specifically, data preprocessing improves data integrity and quality, and comparative analysis of pre-treatment patient information and post-oxaliplatin clinical characteristics ensures the reliability of trial results. The core feature dataset is a multi-dimensional core feature set obtained through redundant variable removal, feature importance ranking, and feature screening to identify high-risk factors for oxaliplatin-induced peripheral neurotoxicity in colorectal cancer. It can cover potential risk factors for OIPN, allowing for accurate identification of potential risk factors in the gradient boosting decision tree model. In practical application, the pre-trained gradient boosting decision tree model (GBDT model) achieved an AUC (95% CI) of 0.908 on the validation set, with a sensitivity of 0.832 and a specificity of 0.803, achieving early risk warning and strong clinical applicability. Furthermore, during model training, the gradient boosting decision tree model performed superiorly in cross-validation. Pre-trained gradient boosting decision tree models can support real-time evaluation in clinical applications, forming a closed-loop management system of "prediction-intervention-reassessment."

[0008] In a preferred example of the present application, the method for determining the core feature dataset includes: Through LASSO regression, the initial feature variable set is preliminarily screened out from the preprocessed original data set; The Boruta algorithm is used to confirm the initial key feature variable set in the preprocessed raw data set; Based on RFECV recursive feature elimination, a set of key feature variables is selected from the preprocessed original data set; Use the pre-trained gradient boosting decision tree model to rank the variables in the pre-processed original dataset to select the top K key risk features; The final core feature variables and the corresponding core feature data set are obtained by taking an intersection based on the initial feature variable set, the initial key feature variable set and the key feature variable set.

[0009] By adopting the above technical solution, after collaborative screening with multiple algorithms, the final core feature data is determined by taking the intersection based on the screening results of each algorithm. This can effectively reduce redundant features and improve the ability to resist overfitting. The multi-dimensional core feature extraction method adopted in this application has a sensitivity improvement of 12.7% compared to the single variable dimension model of trauma.

[0010] By adopting the above technical solutions, using 5-fold cross-validation and SHAP summary plots, we ensured the model's generalization and interpretability. The model's high AUC, sensitivity, and specificity ensured its effectiveness in actual clinical applications, providing reliable decision support for doctors.

[0011] In a preferred example, the present application further includes, after data preprocessing: The pre-treatment clinical characteristic data were divided into a peripheral neurotoxicity group and a non-peripheral neurotoxicity group. The distribution of baseline characteristics of each group was calculated based on the median and quartiles. The baseline characteristics included age, red blood cell count, mean corpuscular volume, total bilirubin, indirect bilirubin, and creatine kinase; The Mann-Whitney U test was used to perform hypothesis testing on the baseline characteristics of the two groups, and a P value was generated to assess the significance of the difference. If the P value was less than 0.05, it was considered a significant difference and entered into the LASSO regression for initial screening; If there are significant differences in features, the inverse probability weighting method is used to correct data bias.

[0012] By adopting the above technical solution, based on detailed analysis and correction of baseline characteristics, potential data bias was eliminated, the quality of model training data was ensured, and the robustness and fairness of the model were improved, enabling it to maintain high predictive accuracy in different patient groups.

[0013] In a preferred example, the present application further includes an inverse relationship correction step: Negatively correlated features that were negatively correlated with neurotoxicity were extracted from the obtained baseline analysis results; Determining the negative correlation strength of the negative correlation feature according to the SHAP dependency graph, and determining the high-risk inverse feature; The high-risk inverse feature is assigned a risk coefficient increment ΔR, calculated as follows: ΔR = 0.3 × (1 – actual value / lower limit of reference range); The predicted risk result is risk-corrected based on the risk coefficient increment ΔR to obtain a final risk value.

[0014] By adopting the above technical solution and introducing a risk correction mechanism for negatively correlated features, the prediction accuracy of the model is further refined, the complex interactions between different features are taken into account, and the authenticity and credibility of the prediction results are improved.

[0015] In a preferred embodiment of the present application, after obtaining the pre-drug patient information and post-treatment clinical characteristic data of the patient using oxaliplatin, the application further includes: Individualized clinical and genetic data of colorectal cancer patients receiving oxaliplatin chemotherapy were selected and classified and integrated into a multidimensional feature dataset for neurotoxicity risk assessment; the multidimensional feature dataset included a susceptibility-related factor set and a vulnerability-related factor set; Taking each patient as the basic analysis unit, the peripheral neurotoxicity susceptibility assessment is performed according to the preset first evaluation factor to obtain the corresponding neurotoxicity susceptibility score, and the neurotoxicity risk level is calculated based on the susceptibility score; Taking each patient as the basic analysis unit, the patient's nervous system vulnerability is assessed according to the preset second evaluation factor to obtain the corresponding neurological function vulnerability data; Based on each patient's neurotoxicity risk level and nervous system vulnerability data, an individualized peripheral neurotoxicity comprehensive risk prediction result is generated.

[0016] By adopting the above technical solution, peripheral neurotoxicity susceptibility assessment refers to a patient's innate tendency to develop neurotoxicity due to drug exposure and genetic background. The neurotoxicity risk level is derived from susceptibility, representing the likelihood of neurotoxicity induction, reflected in the graded neurotoxicity risk level. Vulnerability refers to the patient's inherent ability to resist damage. This application uses a multidimensional peripheral neurotoxicity prediction method framework to integrate patients' pre-drug information and post-treatment clinical data to construct a multidimensional feature dataset containing both "susceptibility" and "vulnerability" dimensions. Based on this, a personalized comprehensive risk prediction result is generated, achieving a shift from "single exposure risk" to "individualized comprehensive risk". Traditional predictions often rely on single factors such as chemotherapy dose or cycle. This invention distinguishes between "susceptibility-related factors" (drug exposure + genetic background) and "vulnerability-related factors" (neurological function status). It also addresses the problem that existing methods lack the ability to assess patients' intrinsic vulnerability. By introducing the assessment dimension of the basic state of the nervous system, the model focuses not only on "whether the drug may cause toxicity" but also on "whether the patient is susceptible to toxic effects", thereby improving the accuracy of identifying high-risk populations.

[0017] In a preferred example of the present application, the peripheral neurotoxicity susceptibility assessment is performed on each patient as a basic analysis unit according to a preset first evaluation factor to obtain a corresponding neurotoxicity susceptibility score, and the neurotoxicity risk level is calculated based on the susceptibility score, specifically including: The first evaluation factors include cumulative oxaliplatin dose, single administration dose, infusion rate, number of chemotherapy cycles, age, history of diabetes, and GSTP1 and ERCC1 gene mutation status; A multivariate logistic regression model was used to train a susceptibility scoring model based on historical cohort data to output the neurotoxicity susceptibility probability value for each patient; The neurotoxicity susceptibility probability value is converted into a susceptibility score, and the hazard levels of different hazard levels are divided in combination with a preset score threshold to generate neurotoxicity hazard data.

[0018] By adopting the above technical solution, key pharmacokinetic parameters such as oxaliplatin cumulative dose, single dose, infusion rate, and number of chemotherapy cycles are combined with age, history of diabetes, and genetic polymorphisms such as GSTP1 and ERCC1 that are known to affect platinum drug metabolism and DNA repair ability. This forms a set of biologically reasonable susceptibility-related factors to improve the suitability and quantification of neurotoxicity susceptibility assessment. By adopting a multivariate logistic regression model and training a scoring model based on historical cohort data, it is possible to quantify the independent contribution of each factor to neurotoxicity and output a continuous susceptibility probability value, avoiding the bias caused by subjective experience judgment.

[0019] In a preferred embodiment of the present application, the patient's nervous system vulnerability assessment is performed based on a preset second evaluation factor with each patient as the basic analysis unit to obtain corresponding neurological vulnerability data, specifically including: The second evaluation factor includes baseline nerve conduction velocity, ankle reflex status, presence of peripheral numbness / tingling symptoms, concomitant use of neurotoxic drugs, vitamin B12 level, and SCN2A gene polymorphism; Standardize and assign values ​​to each patient's second vulnerability factor to construct a vulnerability feature vector; Performing a weighted summation of all second evaluation factor scores involved in a single patient according to the vulnerability characteristic vector to obtain a comprehensive neurological vulnerability score; Based on the comprehensive score of nervous system vulnerability, it is converted into a neurological function vulnerability index through piecewise linear mapping.

[0020] By adopting the above technical solution, objective neurological function indicators such as baseline nerve conduction velocity, ankle reflex, symptom status, and vitamin B12 level are combined with neuroexcitability-related gene polymorphisms such as SCN2A to construct a comprehensive indicator reflecting the "vulnerability" of the patient's nervous system, filling the gap in the assessment of host factors in existing predictive models; through standardized assignment and weighted summation, multi-source heterogeneous data such as continuous variables (such as nerve conduction velocity), categorical variables (such as ankle reflex), medication history, and genotype are fused into a unified "vulnerability feature vector". In practical applications, the comprehensive score can be converted into a vulnerability index with a score of 0-10 through piecewise linear mapping. The numerical values ​​are intuitive and have good clinical usability.

[0021] In the second aspect, the invention objective of this application is achieved by adopting the following technical solutions: A prediction system for peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer, comprising: A data acquisition module is used to obtain pre-drug patient information before oxaliplatin treatment and clinical characteristic data after treatment; Data preprocessing module, used to preprocess pre-drug patient information and clinical characteristic data; The core data screening module is used to analyze and screen data based on pre-processed pre-drug patient information and clinical characteristic data to determine the core feature data set; The model application and result output module is used to predict the occurrence of oxaliplatin-induced peripheral neurotoxicity in patients based on the core feature data set using a pre-trained gradient boosting decision tree model and output the predicted risk results. In the third aspect, the invention objectives of this application are achieved using the following technical solutions: A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer.

[0022] Fourthly, the invention objectives of this application are achieved by adopting the following technical solutions: A computer program product comprises a computer program / instruction, which, when executed by a processor, implements the steps of the method for predicting peripheral neurotoxicity induced by oxaliplatin in treating colorectal cancer.

[0023] In summary, this application includes at least one of the following beneficial technical effects: 1. Through systematic data collection and analysis, it can accurately predict the risk of peripheral neurotoxicity in patients receiving oxaliplatin treatment, helping doctors identify high-risk patients at an early stage; 2. LASSO regression, Boruta algorithm, and RFECV recursive feature elimination were used to screen out key feature variable sets. The combination of multiple feature selection methods effectively screened out the feature variables most relevant to peripheral neurotoxicity and reduced unnecessary noise data. 3. A shift from "single exposure risk" to "individualized comprehensive risk" has been achieved: Traditional predictions mostly rely on single factors such as chemotherapy dose or cycle. This invention distinguishes between "susceptibility-related factors" (drug exposure + genetic background) and "vulnerability-related factors" (neurological function status); and solves the problem of insufficient assessment of patients' intrinsic vulnerability by existing methods: by introducing the assessment dimension of the basic state of the nervous system, the model not only focuses on "whether the drug may cause toxicity", but also on "whether the patient is susceptible to toxicity", thereby improving the accuracy of identifying high-risk groups. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1This is a flow chart of a method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer in one embodiment of the present application.

[0025] Figure 2 This is a schematic diagram of feature screening using the Boruta algorithm in a method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer in one embodiment of the present application.

[0026] Figure 3 This is a feature importance ranking diagram output by a gradient boosted decision tree (GBDT) model in a method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer in one embodiment of the present application.

[0027] Figure 4 It is an intersection graph of the results of the Boruta algorithm, the RFECV algorithm, and the GBDT model used in the method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer in one embodiment of the present application.

[0028] Figure 5 This is a SHAP summary diagram used in the method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer in one embodiment of the present application. DETAILED DESCRIPTION

[0029] The present application is further described in detail below with reference to the accompanying drawings.

[0030] In one embodiment, if Figure 1 As shown, the present application discloses a method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer, which specifically comprises the following steps: S1: Obtain pre-drug patient information before oxaliplatin treatment and clinical characteristics data after treatment.

[0031] In this embodiment, pre-drug patient information refers to information that already exists or is available before the patient starts oxaliplatin chemotherapy, including but not limited to: demographic information (such as age and gender), basic medical history (Past Medical History, PMH), tumor characteristics (such as colorectal cancer staging (CCS), differentiation grade (Grade)), physical status assessment (such as NRS2002 nutritional score), baseline laboratory test indicators (such as blood routine, liver and kidney function, coagulation function, tumor markers, etc.). Post-use clinical characteristic data refer to treatment and response-related data generated during or after oxaliplatin chemotherapy, usually including but not limited to: number of chemotherapy cycles (NOCC), single dose of oxaliplatin (OXA), total cumulative dose of oxaliplatin (Total_OXA), exposure to cold objects (ETCO), laboratory indicators dynamically monitored during chemotherapy (such as blood routine, changes in liver and kidney function), and most importantly, the occurrence of peripheral neurotoxicity (presence of neurotoxicity, group) and its severity (Grade).

[0032] Specifically, data on eligible colorectal cancer patients were collected from hospital information systems (HIS), laboratory information systems (LIS), electronic medical records (EMRs), or specialized research databases. Inclusion criteria typically included patients receiving oxaliplatin-containing chemotherapy. The collected data must encompass both pre-treatment patient information and post-treatment clinical characteristics as defined above.

[0033] S2: Processing and feature analysis are performed based on the pre-drug patient information and the clinical feature data to determine a core feature data set related to neurotoxins.

[0034] In this embodiment, processing includes data preprocessing. Analysis refers to preliminary statistical description and comparison of data to understand data distribution and identify potential risk factors or differences. This application utilizes baseline analysis. Data screening, or feature screening, refers to the process of selecting a subset of raw variables (features) that is most relevant and informative for the prediction target (OIPN occurrence). The core feature dataset is the data subset obtained after screening that contains the most critical features for predicting OIPN risk.

[0035] In this embodiment, data preprocessing includes: S201: Construct an original data set based on pre-drug patient information and clinical characteristic data. The original data set contains multiple variables, and variables with a missing rate greater than a preset missing rate threshold are eliminated.

[0036] In this embodiment, the preset missing rate threshold can be set according to data distribution and clinical importance, and is 30% in this embodiment.

[0037] Specifically, pre-treatment patient information included variables such as age, sex, BMI, underlying diseases (e.g., diabetes), and gene mutations (e.g., KRAS / NRAS status). Clinical characteristic data included variables such as the number of chemotherapy cycles, single oxaliplatin dose (OXA), cumulative dose (Total_OXA), liver function indicators (ALT, AST), renal function indicators (eGFR), and laboratory abnormalities (e.g., elevated D-dimer). The missing proportion of each variable was calculated, and variables such as KRAS, NRAS, BRAF, height, and weight were excluded.

[0038] S202: Obtain missing data, use random forest imputation method to capture nonlinear relationships between variables, and fill in missing data.

[0039] In this example, the random forest imputation method was used to fill missing data. Compared with traditional linear interpolation methods (such as mean imputation and linear regression imputation), the random forest imputation method can better capture nonlinear relationships and interactions between variables. The random forest imputation method uses the nonlinear modeling capabilities of the random forest algorithm to predict missing values ​​using known data. It is applicable to both numerical and categorical variables. Nonlinear inertia refers to complex interactions between variables, such as the combined effect of age and chemotherapy dose on liver function.

[0040] Specifically, when training the random forest imputation method, each variable with missing values ​​(such as ALT) is used as the target variable, and other variables are input into the random forest model as features. The parameters are set to 100 trees and a maximum depth of 5 to avoid overfitting.

[0041] S203: Perform normalized naming processing on the variables.

[0042] Specifically, standardized naming refers to converting variable names into a unified format to enhance readability and maintainability, for example, changing "ETCO" to "cold exposure history."

[0043] In this embodiment, naming rules and variable renaming are pre-established. For example, in the prefix rule, lab_ represents laboratory indicators, such as lab_ALT; demo_ represents demographic information, such as demo_age; and variable renaming converts the original variable name into a standardized name, such as ETCO into cold_exposure; UA into uric_acid, etc.

[0044] Furthermore, the preprocessed complete dataset is randomly divided into a training set and a test set according to a preset ratio (e.g., 7:3). The training set is used for model training and parameter tuning, while the test set is used for final evaluation of model performance. For example, if the total sample size is 829, the training set will have 580 cases and the test set will have 249 cases.

[0045] Specifically, the method for determining the core feature dataset in step S2 includes: S21: Use LASSO regression to preliminarily screen out the initial feature variable set from the preprocessed original data set.

[0046] In this example, LASSO (Least Absolute Shrinkage and Selection Operator) regression is a variable selection method for linear regression. It uses an L1 regularization penalty term to reduce the coefficients of unimportant variables to zero, thereby achieving feature selection. The initial feature variable set is the set of features with non-zero coefficients retained after LASSO regression selection.

[0047] Specifically, the preprocessed dataset was normalized (to a mean of 0 and a variance of 1) to eliminate dimensionality effects. The target variable was the presence or absence of OIPN (a binary variable). Regularization parameter tuning was performed using 10-fold cross-validation to select the optimal regularization strength λ, with the criterion being to minimize the cross-validation error. In this example, the minimum mean square error λ was 0.022.

[0048] For example, based on a total sample of 829 cases, a training set of 580 cases, and a test set of 249 cases, the LASSO regression algorithm screened out 26 features from 50+ variables in the original data set by retaining features with non-zero coefficients, including Sex, CCS (Colorectal Cancer Staging), ETCO (Exposure to cold objects, cold exposure history), NRS2002 (Nutritional Risk Screening 2002, 0-7 points, the higher the score, the greater the nutritional risk), BAS (Basophil, basophil count), T (Tumor Stage, primary tumor stage), Age (age at diagnosis), Total_OXA (cumulative dose of oxaliplatin), OXA (single dose of oxaliplatin), BI (Bleeding Index, bleeding index), D_Dimer (D-dimer), CEA (Carcinoembryonic Antigen, carcinoembryonic antigen), TBIL (Total Bilirubin, total bilirubin), DBIL (Direct Bilirubin (direct bilirubin), IBIL (indirect bilirubin), PA (prealbumin), ALP (alkaline phosphatase), CK (creatine kinase), LDL_C (low-density lipoprotein cholesterol), APOA_1 (apolipoprotein A1), CHE (cholinesterase), eGFR (estimated glomerular filtration rate), UA (uric acid), BMI (body mass index), PT (prothrombin time), and INR (international normalized ratio). Linear relationships were captured using LASSO regression to retain features significantly associated with OIPN, such as chemotherapy dose and liver function indicators.

[0049] S22: The Boruta algorithm is used to confirm the initial set of key feature variables in the preprocessed original data set.

[0050] In this embodiment, the Boruta algorithm is a feature selection method based on random forests. It creates "shadow" features (randomly perturbed versions) of the original features, compares the importance of the original features and the shadow features, and identifies the truly relevant features. Figure 2 As shown, by running multiple iterations, features are ultimately classified as "Confirmed" (confirmed to be important), "Rejected" (not important), or "Tentative" (pending). Confirmed features are retained. For example, (Document 1), Boruta confirmed 21 initial key features: Sex, Age, CCS, Total_OXA, OXA, ETCO, T, BMI, PT, CEA, TBIL, DBIL, IBIL, PA, ALP, CK, CHE, eGFR, UA, LDL_C, and APOA_1.

[0051] Specifically, we created randomly shuffled copies (shadow features) of each original feature, ran a random forest model (number of trees = 500), and calculated the importance scores (e.g., Gini impurity reduction) of all features (original + shadow), repeating this for 100 iterations. During the statistical testing phase, we calculated the Z score of each original feature's importance relative to the best shadow feature. If the Z score was significant (p < 0.05), it was marked as "Confirmed" (important feature); otherwise, it was marked as "Unconfirmed" (unimportant feature).

[0052] S23: Based on RFECV recursive feature elimination, a set of key feature variables is selected from the preprocessed original data set.

[0053] In this example, the RFECV (Recursive Feature Elimination with Cross-Validation) algorithm uses recursive feature elimination combined with cross-validation to select the optimal feature subset by gradually eliminating the least important features. The key feature variable set is the feature subset with the best average performance in cross-validation.

[0054] Specifically, a random forest was selected as the base model (with a default number of 100 trees). The recursive elimination process included step 1: training the model using all original features from the original dataset, calculating feature importance, and selecting the least important feature. This elimination process was repeated until only one feature remained. During the cross-validation evaluation phase, 5-fold cross-validation was performed for each feature subset size (e.g., from 26 to 1), and the average AUC was calculated. The feature set corresponding to the subset size with the highest average AUC was selected.

[0055] Finally, the RFECV algorithm selected 18 features: Sex, Age, CCS, Total_OXA, OXA, ETCO, BI, BAS, BMI, PT, CEA, DBIL, PA, ALP, CHE, eGFR, UA, and APOA_1.

[0056] S24: Use the pre-trained gradient boosting decision tree model to rank the variables of the preprocessed original dataset by feature importance and screen the top K key risk features.

[0057] In this embodiment, K is 15. The Gradient Boosted Decision Tree (GBDT) model calculates the importance score of each feature in the model (e.g., based on the number of times the feature is used for splitting or the total purity gain it brings), and then sorts the features in descending order according to the importance score. Figure 3 As shown in the figure, the top 15 important features ranked by the gradient boosted decision tree (GBDT) model are: Total_OXA, ETCO, BMI, CEA, Age, UA, TBIL, eGFR, CHE, APOA_1, Sex, ALP, OXA, CCS, and DBIL. Feature importance is ranked based on the number of times a feature is used for splitting in the GBDT model or the total information gain it brings.

[0058] Specifically, the importance is calculated by calculating the cumulative information gain (or number of splits) of each feature and normalizing it to obtain the importance score.

[0059] S25: Taking the intersection of the initial feature variable set, the initial key feature variable set and the key feature variable set to obtain the final core feature variables and the corresponding core feature data set.

[0060] In this embodiment, the intersection of the results of the three feature selection methods is taken, such as Figure 4 As shown in the figure, 14 core characteristic variables were finally determined: CEA, CHE, eGFR, BMI, DBIL, OXA, ALP, APOA_1, ETCO, Total_OXA, Sex, UA, CCS, and Age.

[0061] Specifically, set A is defined as the LASSO results of step S21: an initial set of 26 feature variables; set B is defined as the Boruta results of step S22: an initial set of 21 key feature variables; and set C is defined as the RFECV results of step S23: an initial set of 18 key feature variables. A∩B∩C is assumed. These 14 core features and OIPN labels are extracted from the original dataset to form the final modeling dataset (i.e., the core feature dataset). This multi-method cross-validation approach avoids single-method bias (e.g., LASSO may miss nonlinear features). The initial LASSO regression screening in step S21 achieves linear regression and retains strongly correlated features; the Boruta algorithm in step S22 identifies key features, captures nonlinear relationships, and statistically verifies their importance; and the RFECV recursive feature elimination in step S23 ensures generalization through cross-validation, balancing performance and complexity. GBDT importance ranking correlates the final prediction model and quantifies feature contributions. Dimensionality optimization balances predictive power and clinical operability.

[0062] In actual application, the GBDT model trained with 14 core features screened by S21-S25 in 829 patient data achieved a test set AUC of 0.892.

[0063] In this embodiment, after data preprocessing, the following steps are further included: S101: The pre-treatment clinical characteristic data were divided into a peripheral neurotoxicity group and a non-peripheral neurotoxicity group. The distribution of baseline characteristics of each group was calculated based on the median and quartiles. The baseline characteristics included age, red blood cell count, mean corpuscular volume, total bilirubin, indirect bilirubin, and creatine kinase.

[0064] In this example, using a preprocessed dataset (this example uses an actual test process with a total sample size of 829 cases, 580 training cases, and 249 test cases as an example), the patients were divided into a non-neurotoxicity group (group = 0) consisting of 238 patients and a neurotoxicity group (group = 1) consisting of 591 patients. The group variable (presence of neurotoxicity) for grouping was defined during the data preprocessing stage.

[0065] Specifically, for the specified baseline characteristics (age (Age), red blood cell count (RBC), mean corpuscular volume (MCV), total bilirubin (TBIL), indirect bilirubin (IBIL), and creatine kinase (CK)), the median (Median) and quartiles (Q1-Q3) of each group were calculated. Descriptive statistics (median and quartiles) were used to calculate the distribution of baseline characteristics because the data may not be normally distributed.

[0066] Example of feature distribution results: Age: ALL group (all): Median [Q1-Q3]=62.000 [56.000, 69.000] group=0: 60.000 [51.000, 66.000] group=1: 63.000 [56.000, 70.000].

[0067] RBC (Red Blood Cell Count): ALL group: 3.950 [3.620, 4.320] group=0: 4.135 [3.842, 4.460] group=1: 3.850 [3.550, 4.230].

[0068] MCV (Mean Corpuscular Volume): ALL group: 93.300 [87.900, 97.700] group=0: 90.050 [85.825, 94.400] group=1: 94.600 [89.500, 98.900].

[0069] TBIL (Total Bilirubin): ALL group: 10.700 [8.300, 14.400] group=0: 9.800 [7.725, 12.075] group=1: 11.200 [8.700, 14.900].

[0070] IBIL (Indirect Bilirubin): ALL group: 8.800 [6.700, 11.500] group=0: 8.000 [6.400, 9.875] group=1: 9.100 [7.000; 12.200].

[0071] CK (Creatine Kinase): ALL group: 56.000 [40.000, 76.000] group=0: 50.500 [37.000, 70.000] group=1: 58.000 [43.000, 80.000].

[0072] By quantifying the central tendency and dispersion of the characteristics of each group, potential risk factors were identified (e.g., the median Age and TBIL were higher in group = 1).

[0073] S102: The Mann-Whitney U test was used to perform hypothesis testing on the baseline characteristics of the two groups, and a P value was generated to assess the significance of the difference. If the P value was less than 0.05, it was considered a significant difference and entered into the LASSO regression for initial screening.

[0074] In this example, the Mann-Whitney U test (a nonparametric test) was used because baseline characteristics (such as age and RBC count) may not be normally distributed. The test hypothesis: The null hypothesis (H0) is that there is no significant difference in the distribution of baseline characteristics between the two groups (group = 0 and group = 1). The alternative hypothesis (H1) is that there is a significant difference in the distribution. The significance threshold was set at a P value of 0.05. A P value < 0.05 was considered significant.

[0075] The actual test results are: Age: P value = 0.001 (significant), indicating that there is a significant difference in age between the two groups (group = 1 is older).

[0076] RBC: P value < 0.001 (significant), indicating that there is a significant difference in red blood cell count between the two groups (group = 0 is higher).

[0077] MCV: P value < 0.001 (significant), indicating that the mean corpuscular volume was significantly different between the two groups (group = 1 was higher).

[0078] TBIL: P value < 0.001 (significant), indicating that total bilirubin was significantly different between the two groups (group = 1 was higher).

[0079] IBIL: P value < 0.001 (significant), indicating that there was a significant difference in indirect bilirubin between the two groups (group = 1 was higher).

[0080] CK: P value < 0.001 (significant), indicating that there was a significant difference in creatine kinase between the two groups (group = 1 was higher).

[0081] Output the test results, generate a P value table, and retain only the features with P < 0.05 as the initial screening variables for input into LASSO regression.

[0082] S103: If there are significant differences in features, the inverse probability weighting method is used to correct data bias.

[0083] In this example, based on the test results of step S102, all specified characteristics (Age, RBC, MCV, TBIL, IBIL, CK) showed significant differences (P < 0.05), indicating that the data may have selection bias (for example, patients in group 1 are older and have higher bilirubin levels). The source of bias may be that in observational studies, imbalances in baseline characteristics between groups may lead to model bias.

[0084] Specifically, using the inverse probability weighting (IPW) method, each patient's propensity score (P) is first calculated. Using a logistic regression model, a propensity score is predicted based on all baseline characteristics (including age, RBC count, etc.). Weights are calculated as: weight = 1 / Pensity Score (for group = 1) or 1 / (1-Pensity Score) (for group = 0). These weights are then applied to each sample in subsequent analyses (such as LASSO regression) to balance the distribution of characteristics between groups. For example, given the high proportion of elderly patients in group = 1, IPW weights mitigate this effect, ensuring unbiased feature comparisons.

[0085] Age is specified in the baseline characteristics (age, RBC, MCV, TBIL, IBIL, CK) because age is a potential risk factor for OIPN (baseline analysis showed that group = 1 had an older age group). RBC and MCV reflect anemia and may be associated with chemotherapy tolerance. TBIL and IBIL are indicators of liver function. Oxaliplatin metabolism depends on the liver, and abnormal values ​​may increase the risk of neurotoxicity. CK is a marker of muscle damage, and OIPN may contribute to functional impairment.

[0086] S3: Based on the core feature dataset, a pre-trained gradient boosting decision tree model is used to predict oxaliplatin-induced peripheral neurotoxicity in patients and output the predicted risk results.

[0087] In this embodiment, the pre-trained Gradient Boosting Decision Tree (GBDT) model refers to a machine learning model trained using the GBDT algorithm using training samples from a core feature dataset. The model is capable of learning the complex relationship between features and OIPN risk. The predicted risk outcome is the probability value (e.g., 0.85 indicates an 85% risk) or classification result (e.g., "high risk" / "low risk") of OIPN occurring in a newly input patient, output by the model after analyzing the core feature data.

[0088] During actual model training, a gradient boosting decision tree model was used to predict oxaliplatin-induced peripheral neurotoxicity in patients because the gradient boosting decision tree model was the optimal model. The above analysis showed that 14 features (CEA, CHE, eGFR, BMI, DBIL, OXA, ALP, APOA_1, ETCO, Total_OXA, Sex, UA, CCS, and Age) were associated with peripheral neurotoxicity. These variables were selected as independent factors to establish a peripheral neurotoxicity prediction model based on a machine learning algorithm. During the experiment, five machine learning algorithms, including XGBoost, RandomForest, AdaBoost, GBDT, and GNB, were compared. To avoid overfitting and select the optimal model, a 5-fold cross-validation was performed on the training set, and the average accuracy, sensitivity, specificity, positive predictive value, negative predictive value, F1 score, and AUC of the five machine learning models were obtained.

[0089] For the validation set, the AUC (95% CI), accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1 score of the GBDT model were 0.908 (0.854-0.962), 0.833 (0.811-0.855), 0.871 (0.854-0.888), 0.734 (0.690-0.778), 0.894 (0.879-0.910), 0.688 (0.647-0.729), and 0.882 (0.868-0.897), respectively. Compared with other machine learning models, the performance indicators of each model in the training set and validation set are shown in Tables 1 and 2 respectively: Classification Model AUC (95% CI) cutoff (95% CI) Accuracy (95% CI) Sensitivity (95% CI) Specificity (95% CI) Positive predictive value (95% CI) XGBoost 0.992 (0.985-0.999) 0.641(0.618-0.664) 0.964(0.954-0.974) 0.963(0.952-0.974) 0.968(0.958-0.978) 0.987(0.983-0.991) RandomForest 0.944 (0.922-0.965) 0.683(0.671-0.695) 0.876(0.872-0.880) 0.874(0.865-0.882) 0.881(0.851-0.912) 0.95(0.938-0.962) AdaBoost 0.946 (0.926-0.966) 0.515(0.511-0.518) 0.874(0.855-0.893) 0.86(0.825-0.895) 0.911(0.887-0.934) 0.962(0.953-0.970) GBDT 0.997 (0.994-1.000) 0.631(0.594-0.669) 0.978(0.975-0.982) 0.978(0.969-0.988) 0.978(0.965-0.992) 0.992(0.986-0.997) GNB 0.845 (0.807-0.883) 0.496(0.372-0.619) 0.788(0.786-0.790) 0.792(0.776-0.809) 0.778(0.737-0.818) 0.902(0.889-0.916) Table 1 Classification Model AUC (95% CI) cutoff (95% CI) Accuracy (95% CI) Sensitivity (95% CI) Specificity (95% CI) Positive predictive value (95% CI) XGBoost 0.894 (0.830-0.959) 0.641(0.618-0.664) 0.828(0.816-0.840) 0.868(0.838-0.899) 0.722(0.666-0.779) 0.891(0.873-0.908) RandomForest 0.874 (0.805-0.942) 0.683(0.671-0.695) 0.803(0.781-0.826) 0.813(0.763-0.864) 0.777(0.712-0.843) 0.906(0.885-0.927) AdaBoost 0.875 (0.804-0.945) 0.515(0.511-0.518) 0.814(0.788-0.840) 0.816(0.772-0.859) 0.809(0.687-0.932) 0.921(0.876-0.966) GBDT 0.908 (0.854-0.962) 0.631(0.594-0.669) 0.833(0.811-0.855) 0.871(0.854-0.888) 0.734(0.690-0.778) 0.894(0.879-0.910) GNB 0.828 (0.745-0.910) 0.496(0.372-0.619) 0.769(0.724-0.813) 0.78(0.701-0.859) 0.741(0.697-0.785) 0.887(0.879-0.895) Table 2 Based on the indicator data shown in Tables 1 and 2 above, the GBDT model is selected as the optimal risk prediction model.

[0090] Specifically, the method for constructing the gradient boosting decision tree model in step S3 includes: S301: Based on the core feature dataset, a gradient boosting decision tree model is trained using 5-fold cross validation.

[0091] In this embodiment, the core feature dataset contains 14 core features, namely, 14 key predictors identified through multiple rounds of feature screening (LASSO regression, Boruta algorithm, and RFECV recursive feature elimination): CEA (carcinoembryonic antigen), CHE (cholinesterase), eGFR (estimated glomerular filtration rate), BMI (body mass index), DBIL (direct bilirubin), OXA (single dose of oxaliplatin), ALP (alkaline phosphatase), APOA_1 (apolipoprotein A1), ETCO (history of cold exposure), Total_OXA (cumulative dose of oxaliplatin), Sex (sex), UA (uric acid), CCS (colorectal cancer stage), and Age (age).

[0092] Specifically, 5-fold cross validation was performed, and the model optimization hyperparameters included learning rate 0.1, maximum tree depth 5, and number of subtrees 100.

[0093] S302: During the model training process, the contribution of each target feature to the model output is calculated using a SHAP summary plot to determine the most relevant predictor variables for risk prediction.

[0094] In this example, SHAP (SHapley Additive exPlanations) quantifies the contribution of each feature to the model output based on game theory, with positive values ​​promoting neurotoxicity risk and negative values ​​reducing risk. Figure 5 As shown in the figure, key variables are predicted based on the top variables ranked by SHAP importance. In this example, the top three variables were selected. In actual measurements, using the SHAP summary chart, Total_OXA (cumulative dose), ETCO (cold exposure history), and BMI ranked in the top three, indicating that these factors have the greatest impact on predicting OIPN.

[0095] Specifically, as attached Figure 5 As shown in the figure, the contribution of each core feature was visualized through the SHAP summary diagram. The analysis results showed that high Total_OXA, ETCO=1 (history of cold exposure), and BMI>24 significantly increased the risk of OIPN (SHAP value>0).

[0096] S303: The performance indicators of the gradient boosting decision tree model in the validation set must meet AUC ≥ 0.892, sensitivity ≥ 0.832, and specificity ≥ 0.803.

[0097] In this example, the area under the curve (AUC) measures the model's ability to distinguish between positive and negative OIPN patients (0.892 indicates 89.2% accuracy). Sensitivity (recall) measures the model's ability to correctly identify patients with positive OIPN (83.2%). Specificity measures the model's ability to correctly identify patients with negative OIPN (80.3%). The threshold (95% CI) for the AUC metric is [0.849–0.934].

[0098] Specifically, the gradient boosting decision tree model can assist clinical decision-making and guide medication adjustments. For example, patients with high Total_OXA are advised to take the drug in divided doses; patients with ETCO=1 are advised to avoid cold stimulation.

[0099] In one embodiment, after data preprocessing, an inverse relationship correction step is further included: S1001: Extract negative correlation features that are negatively correlated with neurotoxicity from the obtained baseline analysis results.

[0100] In this embodiment, based on the baseline analysis results of step S103, features negatively correlated with neurotoxicity are extracted: Red blood cell count (RBC): median 4.135 in group 0 and 3.850 in group 1 (P < 0.001); Hemoglobin (HGB): group = 0 group 122.500; group = 1 group 120.000 (P = 0.144); Albumin (ALB): group = 0 group 38.300; group = 1 group 38.200 (P = 0.948, negative trend); Prealbumin (PA): group = 0 group 248.250; group = 1 group 235.900 (P = 0.007).

[0101] The results show that negatively correlated features include red blood cell count, hemoglobin, albumin, and prealbumin. Higher values ​​for these negatively correlated features indicate a lower risk of neurotoxicity. High values ​​for RBC or HGB reflect improved oxygen transport capacity, potentially reducing susceptibility to neurological damage. High values ​​for ALB or PA indicate good nutritional status, potentially enhancing chemotherapy tolerance.

[0102] Specifically, screening was based on statistical significance (P < 0.05) or negative trend. S1002: Determine the negative correlation strength of the negative correlation feature based on the SHAP dependency graph and determine the high-risk inverse feature.

[0103] In this embodiment, the SHAP dependency graph uses the actual value of the feature as the horizontal axis and the SHAP value (negative value indicates reduced risk) as the vertical axis. The high-risk inverse characteristic is that the SHAP value is sensitive to changes in the feature value (steep slope), and the actual value is often lower than the lower limit of the reference range.

[0104] For example, the SHAP dependency graph for RBC shows that when RBC < 3.8, the SHAP value drops sharply (risk increases sharply), thus being classified as a high-risk feature. The SHAP value for PA is significantly negative when it is < 200 mg / L, also being classified as a high-risk feature.

[0105] S1003: Assign a risk coefficient increment ΔR to the high-risk inverse feature, calculated as follows: ΔR=0.3×(1–actual value / lower limit of reference range).

[0106] In this embodiment, the actual value is the patient's current test value; the lower limit of the reference range of RBC for men is 4.0×10¹² / L; and the lower limit of the reference range of PA is 200 mg / L.

[0107] For example, if a male patient's RBC is 3.5 × 10¹² / L, the incremental risk factor assigned by the high-risk inverse signature for RBC is ΔR = 0.3 × (1–3.5 / 4.0) = 0.3 × 0.125 = 0.0375. For the same patient with a PA of 180 mg / L, the incremental risk factor assigned by the high-risk inverse signature for PA is ΔR = 0.3 × (1–180 / 200) = 0.3 × 0.1 = 0.03. Assuming the probability of the GBDT outputting a risk prediction value is 0.6, the adjusted value is 0.6 + 0.0375 + 0.03 = 0.6675.

[0108] S1004: Perform risk correction on the predicted risk result based on the risk coefficient increment ΔR to obtain a final risk value.

[0109] In this embodiment, the predicted risk result is the initial predicted probability of the GBDT model for the test set samples; the risk correction is the superimposed risk coefficient increment ΔR.

[0110] For example, patient B's initial predicted probability was 0.68 (intermediate risk), but because RBC=3.6 (ΔR=0.03), after correction it was 0.71, which upgraded the risk to high risk. It was recommended to adjust the oxaliplatin dose.

[0111] In another embodiment, after obtaining pre-drug patient information and post-treatment clinical characteristic data of the patient using oxaliplatin, the method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer further includes: S100: Individualized clinical and genetic data of colorectal cancer patients receiving oxaliplatin chemotherapy were selected and classified and integrated into a multidimensional feature dataset for neurotoxicity risk assessment; the multidimensional feature dataset included a susceptibility-related factor set and a vulnerability-related factor set.

[0112] In this embodiment, individualized clinical and genetic data refers to medical data related to the patient's oxaliplatin treatment and with individual differences, including pre-treatment baseline clinical indicators such as age and liver and kidney function, and dynamic clinical indicators during / after treatment such as the number of chemotherapy cycles and neurotoxic symptoms, as well as genetic variation data related to neurotoxicity such as gene polymorphisms such as GSTP1 and SCN2A.

[0113] Specifically, the susceptibility-related factor set refers to factors directly associated with the risk of neurotoxicity from oxaliplatin exposure, reflecting the probability of a patient developing toxicity due to drug exposure, such as cumulative and single oxaliplatin doses. The vulnerability-related factor set refers to factors related to the patient's nervous system's resistance to neurotoxicity, reflecting the vulnerability of the patient's own neurological function, such as baseline nerve conduction velocity and vitamin B12 levels. Genetic data can be obtained through genetic testing (such as Illumina arrays or next-generation sequencing), focusing on loci known to be associated with oxaliplatin neurotoxicity, including GSTP1 (rs1695), UGT1A1 (rs8175347), ERCC1 (rs11615), FANCC (rs144848), and SCN2A (rs17183814). Specific testing methods follow the "Regulations on the Management of Clinical Gene Amplification Testing Laboratories."

[0114] S200: Taking each patient as the basic analysis unit, perform peripheral neurotoxicity susceptibility assessment according to the preset first evaluation factor to obtain the corresponding neurotoxicity susceptibility score, and calculate the neurotoxicity risk level based on the susceptibility score.

[0115] In this embodiment, the first evaluation factor includes the cumulative dose of oxaliplatin, the single administration dose, the infusion rate, the number of chemotherapy cycles, age, the history of diabetes, and the GSTP1 and ERCC1 gene mutation status.

[0116] Specifically, the first evaluation factor refers to the factor that directly reflects the "risk of neurotoxicity caused by oxaliplatin exposure", including treatment-related factors, patient baseline factors and genetic factors. Treatment-related factors are the cumulative dose, single dose, infusion rate and number of chemotherapy cycles of oxaliplatin. Patient baseline factors include age and history of diabetes. Genetic factors include GSTP1 and ERCC1 gene mutation status.

[0117] In this embodiment, step S200 includes: using a multi-factor logistic regression model to train a susceptibility scoring model based on historical cohort data to output a neurotoxicity susceptibility probability value for each patient; converting the neurotoxicity susceptibility probability value into a susceptibility score, and dividing the risk levels into different levels of risk based on a preset score threshold to generate neurotoxicity risk data.

[0118] In this embodiment, the susceptibility score adopts a 0-100 point system, and different score thresholds are set, such as <30 points for low risk, 30-70 points for medium risk, and >70 points for high risk, thereby generating a risk level.

[0119] Specifically, a multivariate logistic regression model was trained using the LogisticRegression module of the Python scikit-learn library with a regularization parameter C=1.0 and fine-tuned using 5-fold cross-validation. The input of the multivariate logistic regression model was the standardized value of the first evaluation factor (Z-score standardization, mean=0, standard deviation=1), and the output was the probability of neurotoxicity, ranging from 0 to 1. The validation set performance indicators of the multivariate logistic regression model were required to meet the requirements of "AUC ≥ 0.892, sensitivity ≥ 0.832, and specificity ≥ 0.803." If these indicators were not met, the feature selection strategy was adjusted, such as adding an RFECV recursive elimination step.

[0120] The susceptibility score is converted into neurotoxicity risk data by converting the probability value p output by the model into a susceptibility score of 0-100: score = p×100; and based on the ROC curve of the validation set, the optimal cutoff value (such as the point with the largest Youden index) is determined, and the score threshold is set, such as: <30 points for low risk (probability of occurrence <30%), 30-70 points for moderate risk (30%≤probability of occurrence≤70%), and >70 points for high risk (probability of occurrence >70%).

[0121] S300: Taking each patient as a basic analysis unit, the patient's nervous system vulnerability is assessed according to a preset second evaluation factor to obtain corresponding neurological function vulnerability data.

[0122] In this embodiment, the second evaluation factor includes baseline nerve conduction velocity, ankle reflex status, the presence of peripheral numbness / tingling symptoms, concomitant use of neurotoxic drugs, vitamin B12 level, and SCN2A gene polymorphism. The second evaluation factor refers to a factor reflecting the "patient's nervous system's resistance to neurotoxicity" and includes neurological function indicators, concomitant medications, and genetic factors. Neurological function indicators include baseline nerve conduction velocity, ankle reflex status, peripheral numbness / tingling symptoms, concomitant medications mainly refer to whether other neurotoxic drugs other than oxaliplatin are used in combination, and genetic factors such as SCN2A gene polymorphism.

[0123] In this embodiment, step S300 includes: standardizing and assigning values ​​to each second vulnerability factor of each patient to construct a vulnerability feature vector; performing a weighted summation of all second evaluation factor scores involved in a single patient based on the vulnerability feature vector to obtain a comprehensive neurological vulnerability score; and converting the comprehensive neurological vulnerability score into a neurological function vulnerability index through piecewise linear mapping.

[0124] Specifically, if the baseline nerve conduction velocity (NCS) is lower than the lower limit of the normal value (such as <40m / s), it is scored as 3 points, moderate slowing (40-50m / s) is scored as 2 points, mild slowing (50-60m / s) is scored as 1 point, and normal (≥60m / s) is scored as 0 points. Neuroelectrophysiological testing (such as electromyography) is used to obtain the motor conduction velocity (MCV) and sensory conduction velocity (SCV) of the median nerve and common peroneal nerve; absent ankle reflex is scored as 2 points, and weakened ankle reflex is scored as 1 point. Normal is scored as 0 points and assessed by neurological examination; peripheral numbness or tingling symptoms are scored as 2 points, and no symptoms are scored as 0 points; current use of neurotoxic drugs such as paclitaxel and vincristine is scored as 2 points, and no use is scored as 0 points; vitamin B12 levels below the normal range (<200 pg / mL) are scored as 1 point, and normal is scored as 0 points; carrying the SCN2A (rs17183814) risk allele (such as AA or AG type) is scored as 2 points, and the wild type (GG) is scored as 0 points.

[0125] For example, the weights of each factor were determined by expert consensus (neurologists and statisticians). The corresponding weight vectors for NCS, ankle reflex, symptoms, concomitant medication, vitamin B12, and SCN2A were [0.2, 0.15, 0.15, 0.1, 0.1, 0.15]. The overall score = Σ(factor score × corresponding weight), ranging from 0 to 10. For example, if a patient scored 3 for NCS, 2 for ankle reflex, and 0 for all other factors, the overall score would be 3 × 0.2 + 2 × 0.15 = 0.9 + 0.3 = 1.2.

[0126] Specifically, the weighted composite scoring method uses weights to reflect the contribution of each secondary assessment factor to neurological vulnerability, and calculates a composite score. Neurological vulnerability data is scored on a 0-10 scale, with higher scores indicating poorer neurological function reserve.

[0127] To enhance clinical interpretability, the composite score was mapped to a neurological vulnerability index on a 0-10 scale: If the comprehensive score is ≤2 points, the neurological vulnerability index is 0-3 points, indicating low vulnerability; if the comprehensive score is 2 points < comprehensive score ≤5 points, the neurological vulnerability index is 4-6 points, indicating medium vulnerability; if the comprehensive score is >5 points, the neurological vulnerability index is 7-10 points, indicating high vulnerability.

[0128] S400: Based on each patient's neurotoxicity risk level and nervous system vulnerability data, an individualized peripheral neurotoxicity comprehensive risk prediction result is generated.

[0129] In this example, a two-dimensional risk matrix model is established, with the horizontal axis representing the neurotoxicity risk level (low, medium, high) and the vertical axis representing the nervous system vulnerability index (low, medium, high). Based on the relationship between the two, a comprehensive risk level is determined: low risk, early warning risk, or high risk. A comprehensive peripheral neurotoxicity risk prediction result is output, including the risk level and the ranking of the main driving factors.

[0130] Specifically, the low risk of the comprehensive risk level is: low risk + low vulnerability / medium risk + low vulnerability; the warning risk is: low risk + medium vulnerability / medium risk + medium vulnerability / high risk + low vulnerability; the high risk is: medium risk + high vulnerability / high risk + medium vulnerability / high risk + high vulnerability.

[0131] It should be understood that the serial numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0132] In one embodiment, a system for predicting peripheral neurotoxicity induced by oxaliplatin in treating colorectal cancer is provided. The system for predicting peripheral neurotoxicity induced by oxaliplatin in treating colorectal cancer corresponds to the method for predicting peripheral neurotoxicity induced by oxaliplatin in treating colorectal cancer in the above embodiment.

[0133] The system for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer includes a data acquisition module, a data preprocessing module, a core data screening module, and a model application and result output module. A detailed description of each functional module is as follows: A data acquisition module is used to obtain pre-drug patient information before oxaliplatin treatment and clinical characteristic data after treatment; Data preprocessing module, used to preprocess pre-drug patient information and clinical characteristic data; The core data screening module is used to analyze and screen data based on pre-processed pre-drug patient information and clinical characteristic data to determine the core feature data set; The model application and result output module is used to predict oxaliplatin-induced peripheral neurotoxicity in patients based on the core feature data set using a pre-trained gradient boosting decision tree model and output the predicted risk results.

[0134] For the specific limitations of the system for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer, please refer to the limitations of the method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer above, and will not be repeated here; the various modules in the above-mentioned system for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer can be implemented in whole or in part by software, hardware, and a combination thereof; the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0135] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: S1: Obtain pre-drug patient information before oxaliplatin treatment and clinical characteristics data after treatment; S2: Process and analyze pre-drug patient information and clinical characteristic data to determine the core feature dataset related to neurotoxins; S3: Based on the core feature dataset, a pre-trained gradient boosting decision tree model is used to predict oxaliplatin-induced peripheral neurotoxicity in patients and output the predicted risk results.

[0136] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0137] In one embodiment, particularly according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method for predicting oxaliplatin-induced peripheral neurotoxicity in the treatment of colorectal cancer, as described above. In such an embodiment, the computer program can be downloaded and installed from a network via a communication module and / or installed from a removable medium. When executed by a central processing unit (CPU), the computer program performs the various functions defined in the present invention.

[0138] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0139] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer, characterized in that: include: Obtain pre-drug patient information before oxaliplatin treatment and clinical characteristics data after treatment; Performing processing and feature analysis based on the pre-drug patient information and the clinical feature data to determine a core feature data set related to neurotoxins; Based on the core feature data set, a pre-trained gradient boosting decision tree model is used to predict the occurrence of oxaliplatin-induced peripheral neurotoxicity in patients and output a predicted risk result.

2. The method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer according to claim 1, characterized in that: The method for determining the core feature data set includes: Through LASSO regression, the initial feature variable set is preliminarily screened out from the preprocessed original data set; The Boruta algorithm is used to confirm the initial key feature variable set in the preprocessed raw data set; Based on RFECV recursive feature elimination, a set of key feature variables is selected from the preprocessed original data set; Use the pre-trained gradient boosting decision tree model to rank the variables in the pre-processed original dataset to select the top K key risk features; The final core feature variables and the corresponding core feature data set are obtained by taking an intersection based on the initial feature variable set, the initial key feature variable set and the key feature variable set.

3. The method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer according to claim 2, characterized in that: After data preprocessing, it also includes: The pre-treatment clinical characteristic data were divided into a peripheral neurotoxicity group and a non-peripheral neurotoxicity group. The distribution of baseline characteristics of each group was calculated based on the median and quartiles. The baseline characteristics included age, red blood cell count, mean corpuscular volume, total bilirubin, indirect bilirubin, and creatine kinase; The Mann-Whitney U test was used to perform hypothesis testing on the baseline characteristics of the two groups, and a P value was generated to assess the significance of the difference. If the P value was less than 0.05, it was considered a significant difference and entered into the LASSO regression for initial screening; If there are significant differences in features, the inverse probability weighting method is used to correct data bias.

4. The method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer according to claim 3, characterized in that: Also included is an inverse relationship correction step: Negatively correlated features that were negatively correlated with neurotoxicity were extracted from the obtained baseline analysis results; Determining the negative correlation strength of the negative correlation feature according to the SHAP dependency graph, and determining the high-risk inverse feature; The high-risk inverse feature is assigned a risk coefficient increment ΔR, calculated as follows: ΔR = 0.3 × (1 – actual value / lower limit of reference range); The predicted risk result is risk-corrected based on the risk coefficient increment ΔR to obtain a final risk value.

5. The method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer according to claim 1, characterized in that: After obtaining the patient's pre-drug information before oxaliplatin treatment and the clinical characteristics data after treatment, it also includes: Individualized clinical and genetic data of colorectal cancer patients receiving oxaliplatin chemotherapy were selected and classified and integrated into a multidimensional feature dataset for neurotoxicity risk assessment; the multidimensional feature dataset included a susceptibility-related factor set and a vulnerability-related factor set; Taking each patient as the basic analysis unit, the peripheral neurotoxicity susceptibility assessment is performed according to the preset first evaluation factor to obtain the corresponding neurotoxicity susceptibility score, and the neurotoxicity risk level is calculated based on the susceptibility score; Taking each patient as the basic analysis unit, the patient's nervous system vulnerability is assessed according to the preset second evaluation factor to obtain the corresponding neurological function vulnerability data; Based on each patient's neurotoxicity risk level and nervous system vulnerability data, an individualized peripheral neurotoxicity comprehensive risk prediction result is generated.

6. The method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer according to claim 5, characterized in that: The method uses each patient as a basic analysis unit, performs peripheral neurotoxicity susceptibility assessment according to a preset first evaluation factor, obtains a corresponding neurotoxicity susceptibility score, and calculates the neurotoxicity risk level based on the susceptibility score, specifically including: The first evaluation factors include cumulative oxaliplatin dose, single administration dose, infusion rate, number of chemotherapy cycles, age, history of diabetes, and GSTP1 and ERCC1 gene mutation status; A multivariate logistic regression model was used to train a susceptibility scoring model based on historical cohort data to output the neurotoxicity susceptibility probability value for each patient; The neurotoxicity susceptibility probability value is converted into a susceptibility score, and the hazard levels of different hazard levels are divided in combination with a preset score threshold to generate neurotoxicity hazard data.

7. The method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer according to claim 5, characterized in that: The method uses each patient as a basic analysis unit and performs a neurological vulnerability assessment on the patient according to a preset second evaluation factor to obtain corresponding neurological vulnerability data, specifically including: The second evaluation factor includes baseline nerve conduction velocity, ankle reflex status, presence of peripheral numbness / tingling symptoms, concomitant use of neurotoxic drugs, vitamin B12 level, and SCN2A gene polymorphism; Standardize and assign values ​​to each patient's second vulnerability factor to construct a vulnerability feature vector; Performing a weighted summation of all second evaluation factor scores involved in a single patient according to the vulnerability characteristic vector to obtain a comprehensive neurological vulnerability score; Based on the comprehensive score of nervous system vulnerability, it is converted into a neurological function vulnerability index through piecewise linear mapping.

8. A system for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer, characterized in that the system include: A data acquisition module is used to obtain pre-drug patient information before oxaliplatin treatment and clinical characteristic data after treatment; A data preprocessing module, configured to perform data preprocessing on the pre-drug patient information and the clinical characteristic data; A core data screening module, configured to perform analysis and data screening based on the pre-processed pre-drug patient information and the pre-processed clinical characteristic data to determine a core characteristic data set; The model application and result output module is used to predict the occurrence of oxaliplatin-induced peripheral neurotoxicity in patients based on the core feature data set using a pre-trained gradient boosting decision tree model and output the predicted risk results.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer as claimed in any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method for predicting peripheral neurotoxicity induced by oxaliplatin in the treatment of colorectal cancer as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Model for predicting severe SA-AKI risk of ICU sepsis patient based on clinical variables

    CN122291004A