Colon tumor drug delivery method and system based on AI control
By collecting multi-dimensional data and building an AI fusion model, the treatment plan is dynamically adjusted, which solves the problem that traditional chemotherapy regimens cannot accurately identify drug responses in the treatment of colorectal cancer. This achieves personalized chemotherapy effects and improves the treatment response rate of oxaliplatin and the quality of life of patients.
Patent Information
- Application Number
- CN202511382394.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional chemotherapy regimens cannot accurately identify key factors in drug response during the treatment of colorectal cancer, resulting in significant individual differences. Some patients experience delayed treatment or discontinue treatment due to toxic side effects, and the lack of dynamic adjustment mechanisms makes it difficult to cope with drug resistance mutations.
We collected whole-genome sequencing data, radiomics data, and clinicopathological data from patients with colorectal cancer. We screened drug response biomarkers using three-dimensional convolutional neural networks and Lasso regression analysis, constructed a gradient boosting tree and deep neural network fusion model, dynamically adjusted the treatment plan, and monitored drug resistance mutations during the treatment process.
It enables precision and personalization of chemotherapy for colorectal cancer, improves the response rate of oxaliplatin treatment, reduces drug use in patients who are ineffective or intolerant, reduces unnecessary toxic side effects, and improves patients' quality of life and survival.
Smart Images

Figure CN121506366A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent drug delivery control technology, and in particular to an AI-controlled method and system for administering medication to colon tumors. Background Technology
[0002] In the clinical treatment of colorectal cancer, chemotherapy is the core treatment for advanced or metastatic colorectal cancer. Oxaliplatin, as a first-line chemotherapy drug (such as a key component of the FOLFOX regimen), exhibits significant individual differences in efficacy and safety. Approximately 40%-50% of patients do not respond to oxaliplatin, and 20%-30% of patients experience grade 3 or higher serious adverse reactions (such as peripheral neurotoxicity and neutropenia). Traditional chemotherapy regimens are unable to meet the needs of precision treatment, becoming a core bottleneck restricting the improvement of efficacy.
[0003] Traditional chemotherapy regimens rely on physician experience and only consider limited clinicopathological information such as tumor TNM staging and performance status scores, resulting in the following shortcomings: Firstly, the data dimensions are limited, failing to capture the essence of individual differences. Secondly, traditional regimens neglect genomic molecular characteristics (such as KRAS and BRAF mutations directly reducing oxaliplatin sensitivity) and tumor radiomics characteristics (such as texture complexity and surface fractal dimension reflecting tumor heterogeneity; higher heterogeneity leads to lower drug response rates). Relying solely on clinical indicators cannot accurately identify key factors in drug response, leading to delayed treatment or treatment interruption due to toxic side effects in some patients. Furthermore, the lack of dynamic adjustment mechanisms makes it difficult to cope with disease evolution. Thirdly, colon cancer cells are prone to developing resistance mutations under oxaliplatin selection pressure, occurring in approximately 30%-40% of patients after 2-3 cycles of treatment. Traditional resistance monitoring relies on imaging assessments, which lag behind molecular mutations by 4-6 weeks. Simultaneously, changes in renal function affect drug metabolism, and traditional fixed-dose regimens can easily lead to drug accumulation and toxicity, increasing the risk of adverse reactions.
[0004] Therefore, there is an urgent need for an AI-controlled method and system for administering medication to colon tumors to solve the above problems. Summary of the Invention
[0005] The purpose of this invention is to provide an AI-controlled method for administering medication to colon tumors, comprising the following steps:
[0006] Whole genome sequencing data, radiomics data, clinicopathological data, and historical drug treatment response data of colorectal cancer patients were collected, and the radiomics data, clinicopathological data, and historical drug treatment response data were preprocessed.
[0007] A three-dimensional convolutional neural network was used to extract tumor texture and morphological features from radiomics data and obtain image feature vectors. Lasso regression analysis was used to screen gene mutation biomarkers that were significantly associated with oxaliplatin drug response. Clinical pathological data were processed by one-hot encoding to form pathological feature vectors.
[0008] The image feature vector, whole genome sequencing data and pathological feature vector were normalized and spliced together. The prediction model was trained using a gradient boosting tree and deep neural network fusion algorithm, with the oxaliplatin response probability as the output target.
[0009] When the probability of the model's output response is not less than the preset value, the first treatment option is recommended.
[0010] When the response probability output by the model is less than the preset value, the second treatment option is recommended.
[0011] During treatment, circulating tumor DNA data in the patient's blood is collected periodically, and the model is triggered to re-predict when drug resistance mutations are detected.
[0012] Furthermore, the preprocessing steps for radiomics data, clinicopathological data, and historical drug treatment response data include:
[0013] Resample DICOM format radiomics data to 1mm 3 Voxel resolution, and N4 bias field correction;
[0014] The whole genome sequencing data were analyzed using the GATK protocol to detect single nucleotide variants and insertion / deletion mutations, and variant sites with a pathogenicity rating of not less than that of potentially pathogenic were retained.
[0015] Categorical variables were coded using one-hot encoding, and continuous variables were standardized using Z-score.
[0016] Multiple imputation is performed using the nearest neighbor algorithm, and a maximum missing rate threshold is set.
[0017] Furthermore, the steps of using a three-dimensional convolutional neural network to extract tumor texture and morphological features from radiomics data and obtain image feature vectors, screening for gene mutation biomarkers significantly associated with oxaliplatin drug response through Lasso regression analysis, and performing one-hot encoding on clinicopathological data to form pathological feature vectors include:
[0018] Using a pre-trained deep convolutional network, texture and morphological features are extracted from the tumor region.
[0019] The risk ratio of each mutation site was calculated based on the risk model, and driver genes with significant and clinical relevance were retained.
[0020] The high-dimensional feature vectors are reduced to a low-dimensional space using a manifold learning algorithm, while retaining most of the original information.
[0021] Furthermore, the normalized and concatenated image feature vectors, whole-genome sequencing data, and pathological feature vectors are combined, and a prediction model is trained using a gradient boosting tree and deep neural network fusion algorithm, with the oxaliplatin response probability as the output target. The fusion modeling steps include:
[0022] Design a network architecture that includes a dual-path structure with gradient boosting tree branches and deep learning branches;
[0023] The output of tree branches and deep learning branches is boosted by dynamically integrating gradients with trainable weight coefficients.
[0024] A hierarchical cross-validation method is adopted, with the optimization objective being to achieve the area under the receiver operating characteristic curve reaching a preset standard.
[0025] Furthermore, the step of periodically collecting circulating tumor DNA data from the patient's blood during treatment and triggering model re-prediction when drug resistance mutations are detected includes:
[0026] For patients in the high-response group, the oxaliplatin dose was dynamically adjusted based on renal function indicators and body surface area;
[0027] Patients in the low-response group were further subdivided into the moderate-response group and the low-response group, and combination therapy or targeted drug screening was initiated for each group respectively.
[0028] Tumor marker levels are monitored regularly, and a protocol review is triggered when two consecutive measurements show a significant increase.
[0029] Furthermore, it also includes clinical validation steps, including:
[0030] The model's accuracy in predicting relapse-free survival was evaluated in an independent patient testing set.
[0031] Patients were randomly assigned to an AI-guided group and a clinical standard group for comparison;
[0032] Progression-free survival and the incidence of serious adverse reactions were used as the primary indicators for evaluating efficacy and safety.
[0033] This invention also discloses an AI-controlled colon tumor drug delivery system, comprising:
[0034] The acquisition module is used to collect whole genome sequencing data, radiomics data, clinicopathological data, and historical drug treatment response data of colorectal cancer patients, and to preprocess the radiomics data, clinicopathological data, and historical drug treatment response data.
[0035] The acquisition module is used to extract tumor texture and morphological features from radiomics data using a three-dimensional convolutional neural network, and obtain image feature vectors. It also uses Lasso regression analysis to screen gene mutation biomarkers that are significantly associated with oxaliplatin drug response, performs one-hot encoding on clinicopathological data, and forms pathological feature vectors.
[0036] The modeling module is used to normalize and splice together image feature vectors, whole genome sequencing data and pathological feature vectors, and train the prediction model using a gradient boosting tree and deep neural network fusion algorithm, with the oxaliplatin response probability as the output target.
[0037] The first recommendation module is used to recommend the first treatment plan when the response probability output by the model is not less than a preset value.
[0038] The second recommendation module is used to recommend a second treatment option when the response probability output by the model is less than a preset value.
[0039] The re-prediction module is used to periodically collect circulating tumor DNA data in the patient's blood during treatment, and triggers the model to re-predict when drug resistance mutations are detected.
[0040] Furthermore, the acquisition module includes:
[0041] Extraction unit, used to extract texture and morphological features from tumor regions using a pre-trained deep convolutional network;
[0042] The computational unit is used to calculate the risk ratio of each mutation site based on the risk model, retaining driver genes with significant and clinical relevance;
[0043] The dimensionality reduction unit is used to reduce high-dimensional feature vectors to a low-dimensional space using a manifold learning algorithm, while retaining most of the original information.
[0044] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described AI-controlled colon tumor drug delivery method.
[0045] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described AI-controlled colon tumor drug delivery method.
[0046] The beneficial effects of this application are as follows:
[0047] This invention collects whole-genome sequencing data, radiomics data, clinicopathological data, and historical drug treatment response data from colorectal cancer patients and performs standardized preprocessing. It then uses a three-dimensional convolutional neural network to extract tumor imaging features, Lasso regression analysis to screen for genomic-level drug response biomarkers, and a gradient boosting tree and deep neural network fusion algorithm to construct a predictive model. The personalized treatment plan is determined with the oxaliplatin response probability as the output target. During the treatment process, the model is re-predicted by periodically collecting circulating tumor DNA data in the blood to monitor drug resistance mutations. Ultimately, this invention achieves precision and personalization in chemotherapy for colorectal cancer, improves the oxaliplatin treatment response rate, reduces drug use in ineffective or intolerant patients, reduces unnecessary toxic side effects, and thus improves patients' quality of life and survival. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of a method flow proposed in an embodiment of this application.
[0049] Figure 2 This is a schematic diagram of the system structure proposed in an embodiment of the present invention.
[0050] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0051] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0052] like Figure 1 As shown, this application provides an AI-controlled method for administering medication to colon tumors, comprising the following steps:
[0053] S1 collects whole genome sequencing data, CT / MRI radiomics data, clinicopathological data, and historical drug treatment response data from patients with colorectal cancer, and preprocesses the radiomics data, clinicopathological data, and historical drug treatment response data.
[0054] S2 uses a three-dimensional convolutional neural network to extract tumor texture and morphological features from CT / MRI radiomics data and obtain image feature vectors. Lasso regression analysis is used to screen gene mutation biomarkers that are significantly associated with oxaliplatin drug response. Clinical pathological data are then processed by one-hot encoding to form pathological feature vectors.
[0055] S3 normalizes and splices together image feature vectors, whole genome sequencing data and pathological feature vectors, and trains a prediction model using a gradient boosting tree and deep neural network fusion algorithm, with the oxaliplatin response probability as the output target.
[0056] S4. When the response probability output by the model is not less than the preset value, the first treatment plan is recommended.
[0057] S5. When the response probability output by the model is less than the preset value, the second treatment option is recommended.
[0058] S6 collects circulating tumor DNA data from the patient's bloodstream periodically during treatment, and triggers the model to re-predict when drug resistance mutations are detected.
[0059] As described in steps S1-S6 above, this invention collects whole-genome sequencing data, CT / MRI radiomics data, clinicopathological data, and historical drug treatment response data from colorectal cancer patients and performs standardized preprocessing. It then uses a three-dimensional convolutional neural network to extract tumor imaging features, Lasso regression analysis to screen for genomic-level drug response biomarkers, and a gradient boosting tree and deep neural network fusion algorithm to construct a predictive model. The personalized treatment plan is determined with the oxaliplatin response probability as the output target. During the treatment process, blood circulating tumor DNA data is collected periodically to monitor drug resistance mutations and trigger model re-prediction. Ultimately, this achieves precision and personalization in colorectal cancer chemotherapy, improves the oxaliplatin treatment response rate, reduces drug use for ineffective or intolerant patients, reduces unnecessary toxic side effects, and thus improves patients' quality of life and survival.
[0060] In the clinical treatment of colorectal cancer, oxaliplatin, as a first-line chemotherapy drug, exhibits significant individual differences in its therapeutic efficacy and safety. These differences stem from variations in patient genome mutations (e.g., specific mutations in KRAS and BRAF genes can directly affect drug metabolism pathways or targets, leading to drug resistance), tumor imaging characteristics (e.g., the complexity of tumor texture and the regularity of morphology reflect tumor heterogeneity; higher heterogeneity usually indicates lower drug sensitivity), clinicopathological characteristics (e.g., tumor stage and differentiation degree determine basic treatment tolerance; T4 stage patients generally have lower tolerance than T1 stage patients), and historical treatment response differences (patients who previously did not respond to platinum-based drugs also have a lower probability of responding to oxaliplatin). Traditional chemotherapy regimens rely heavily on physician experience, referencing only limited clinical indicators such as tumor stage and patient condition. They fail to systematically integrate multi-dimensional key data, making it impossible to accurately capture individual differences. This leads to a one-size-fits-all approach to treatment, with some patients experiencing delayed treatment due to the use of ineffective oxaliplatin, and others suffering reduced quality of life due to drug side effects (such as neurotoxicity and gastrointestinal reactions). Therefore, there is an urgent need for a decision-making method that can integrate multi-source data, objectively assess drug response, and dynamically adapt to changes in treatment, addressing the core issue of insufficient precision in traditional regimens.
[0061] Traditional solutions to the aforementioned problems suffer from two major shortcomings: first, they rely on a single data dimension, focusing only on clinical pathological indicators while neglecting crucial data such as genomics and radiomics that directly reflect drug sensitivity, leading to insufficient predictive evidence; second, they lack dynamic adjustment mechanisms, with treatment plans being rigidly implemented once determined, failing to address drug resistance mutations that emerge during treatment and resulting in later treatment ineffectiveness. This invention proposes a non-invasive, purely data-driven AI decision-making solution. By comprehensively collecting and standardizing multi-dimensional patient data, it utilizes specialized algorithms to extract key features and construct a fusion predictive model. Objective probability outputs guide treatment selection, and dynamic monitoring data during treatment triggers model updates. From data integration, model construction, treatment decision-making to dynamic adjustment, the entire process achieves precision and personalization, specifically addressing the deficiencies of traditional methods.
[0062] Specifically, whole-genome sequencing data, CT / MRI radiomics data, clinicopathological data, and historical drug treatment response data were collected from patients with colorectal cancer. Preprocessing was performed on the radiomics, clinicopathological, and historical drug treatment response data. The acquisition methods for each data point were clearly defined. Whole-genome sequencing data was obtained through whole-genome sequencing experiments on peripheral blood or tumor tissue samples, detecting single nucleotide variants, insertion / deletion mutations, and other sites. CT / MRI radiomics data was obtained through clinical CT or MRI examinations, storing spatial information of the tumor and surrounding tissues in DICOM format. Clinicopathological data was collected from the medical record system, covering tumor TNM stage, differentiation degree (high / moderate / low differentiation), patient age, and renal function indicators. Historical drug treatment response data also originated from the medical record system, recording the efficacy of previous chemotherapy (tumor shrinkage rate, progression-free survival) and adverse reactions. The core purpose of preprocessing was to eliminate data noise and standardize the format and scale, laying the foundation for subsequent analysis: imaging data needed to be resampled to 1 mm. 3Voxel resolution addresses feature extraction bias caused by differences in output resolution across different devices (e.g., some CT devices have a resolution of 0.8mm×0.8mm×1.5mm). N4 bias field correction is also applied to eliminate inconsistencies in grayscale values caused by uneven magnetic fields and tissue absorption differences, ensuring the accuracy and reliability of tumor texture and morphological features. Categorical variables (e.g., tumor differentiation degree) in clinical pathology data are converted to binary vectors using one-hot encoding (high differentiation corresponds to [1,0,0]), while continuous variables (e.g., age) are standardized using Z-scores to eliminate the impact of magnitude differences on the model. Historical treatment response data is standardized and encoded for efficacy outcomes (complete remission = 4, partial remission = 3, stable = 2, progression = 1). For all missing data values, the k-nearest neighbor algorithm (k = 5) is used for multiple imputation, with a maximum missing rate threshold of 15%. Data from patients exceeding this threshold are removed to prevent incomplete data from affecting model accuracy. For example, if a patient's clinical pathology data lacks body surface area, it is imputed using body surface area data from five patients with similar age, height, and weight to ensure the integrity of this feature.
[0063] A three-dimensional convolutional neural network was used to extract tumor texture and morphological features from CT / MRI radiomics data and obtain image feature vectors. Lasso regression analysis was used to screen for gene mutation biomarkers that were significantly associated with oxaliplatin drug response (p<0.05). Clinical pathological data were processed by one-hot encoding to form pathological feature vectors. The three-dimensional convolutional neural network (pre-trained 3D-ResNet50 network) was chosen because it can process three-dimensional voxel data and completely preserve the spatial structural information of the tumor (such as the relationship between the tumor and blood vessels and the internal heterogeneity distribution), while two-dimensional networks cannot capture such key information. The specific extraction process is as follows: First, the preprocessed image is segmented into the tumor ROI region (clinically labeled or automatically segmented), then input into the 3D-ResNet50 network, and texture features (such as gray-level co-occurrence matrix correlation and entropy value, reflecting the uniformity of gray-level distribution inside the tumor) and morphological features (such as tumor volume, sphericity, and surface fractal dimension) are extracted through convolutional layers and pooling layers. Finally, it is mapped to a 512-dimensional image feature vector through a fully connected layer. For example, patients with high tumor surface fractal dimension have high corresponding dimension values in their image feature vectors, and the model can judge that the tumor heterogeneity is strong and the response probability may be reduced accordingly. Lasso regression is used to screen for gene mutation biomarkers because whole-genome sequencing data contains a massive number of mutation sites (tens of thousands to hundreds of thousands), most of which are unrelated to drug response. Lasso regression uses L1 regularization to compress the coefficients, setting the coefficients of irrelevant sites to 0, and retaining only sites with p < 0.05 (such as KRASG12D and BRAFV600E mutations). These biomarkers directly reflect differences in genomic sensitivity. For example, in patients with KRAS activating mutations, the biomarker value is 1, and the model uses this as an important basis for a reduced response probability. One-hot encoding of clinicopathological data to form pathological feature vectors integrates features based on preprocessing: tumor stage (T1-T4) is encoded into a 4-dimensional vector, differentiation degree is encoded into a 3-dimensional vector, and combined with standardized continuous variables, the t-SNE algorithm is used to reduce the dimension to 32 (retaining 90% of the original information) to form pathological feature vectors. For example, in the vector of a T4 stage patient, the corresponding T4 stage dimension value is 1, and the model can use this information to determine that the treatment tolerance is low.
[0064] Image feature vectors, whole-genome sequencing data, and pathological feature vectors were normalized and concatenated. A gradient boosting tree and deep neural network fusion algorithm was used to train a prediction model, with the oxaliplatin response probability as the output target. Normalization and concatenation were performed to eliminate differences in feature scales: image vectors take values of [0,1], gene mutation markers take values of 0 / 1, and pathological vectors take values of [-2,2]. These values were uniformly mapped to the [0,1] interval through Min-Max normalization and then concatenated into a 594-dimensional fusion feature vector (512+50+32) to ensure fair weighting of each feature. The fusion algorithm employs a dual-path structure: the gradient boosting tree branch (LightGBM) excels at capturing feature interactions (such as the combined effect of "KRAS mutation + T4 stage" on the response), while the deep learning branch (3-layer perceptron: 128 / 64 / 32 neurons, ReLU activation function) excels at uncovering deep abstract relationships. After the two branches output preliminary response probabilities, they are weighted and summed using trainable weight coefficients (initially 0.5, dynamically adjusted by gradient descent) to obtain the final response probability ([0,1]). Model training uses 5-fold hierarchical cross-validation to ensure that the proportion of patients with response types is consistent between the training and validation sets. The optimization objective is AUC≥0.85. For example, for patients with low image entropy, KRAS wild-type, and T2 stage in the fused feature vector, the model outputs a response probability of 0.8, while for patients with high image entropy, KRAS mutation, and T4 stage, the output is 0.3, achieving accurate probability prediction.
[0065] When the model output response probability is not less than the preset value, the first treatment regimen is recommended; when the model output response probability is less than the preset value, the second treatment regimen is recommended; when the model output response probability is ≥0.7, the FOLFOX chemotherapy regimen containing oxaliplatin is recommended; when the response probability is <0.7, the FULV alternative regimen without oxaliplatin is recommended. The preset value of 0.7 is based on clinical data validation: when the response probability is ≥0.7, the FOLFOX chemotherapy regimen, i.e., oxaliplatin 85mg / m², is recommended. 2 +Leucovorin + 5-fluorouracil, once every 2 weeks, with a response rate exceeding 60% and a serious adverse reaction rate of <20%; when the response probability is <0.7, the FULV regimen, i.e., leucovorin 200 mg / m², is used. 2 +5-Fluorouracil 425mg / m² 2 Administered every 4 weeks, with a response rate of approximately 40% and a serious adverse reaction rate of <15%, balancing efficacy and safety. The specific recommendation logic is as follows: when the response probability is ≥0.7, FOLFOX (recommendation option 1) is recommended, and the dosage is calculated based on body surface area (e.g., for a body surface area of 1.6 m²). 2 For patients, the oxaliplatin dose is 85 × 1.6 = 136 mg; when the response probability is <0.7, FULV (second regimen) is recommended, for example, for patients with KRAS mutations or T4 stage, the response probability is 0.55, and the system automatically recommends the FULV regimen to avoid ineffective toxic side effects.
[0066] During treatment, circulating tumor DNA (ctDNA) data are collected from patients' blood regularly. When drug resistance mutations are detected, the model is triggered to re-predict. 5 mL of peripheral blood is collected every 2 chemotherapy cycles (4 weeks) to collect circulating tumor DNA (ctDNA). Drug resistance mutations (such as new KRAS mutations or increased allele frequency of existing mutations) are detected by digital PCR or next-generation sequencing. These mutations are direct molecular evidence of drug resistance. For example, in patients with initial KRAS wild-type, KRASG12V mutation (allele frequency 5%) is detected in ctDNA after 2 cycles of treatment, indicating drug resistance. If FOLFOX response rate continues, it will drop from 65% to below 10%. After detecting a drug resistance mutation, the system updates the data (supplementing the drug resistance mutation to genomic data and updating imaging / clinical data) and re-enters the model prediction: if the new probability is <0.7, switch to the FULV regimen (if the original regimen was FOLFOX); if it is still ≥0.7, maintain the original regimen but adjust the dose (e.g., for patients with a glomerular filtration rate <60mL / min, halve the oxaliplatin dose). For example, a patient with an initial response probability of 0.78 using FOLFOX, after 4 cycles, BRAF mutation was detected in ctDNA, and the re-predicted probability was 0.42. The system recommends switching to the FULV regimen to ensure treatment effectiveness.
[0067] Through the above-described process, this invention integrates multi-source data and uses professional algorithms to build a predictive model, enabling precise and personalized drug administration from static plan formulation to dynamic adjustment. This effectively solves the problems of traditional chemotherapy regimens being highly subjective and unable to cope with drug resistance, providing objective and reliable decision support for oxaliplatin chemotherapy for colorectal cancer.
[0068] In one embodiment, the step of preprocessing radiomics data, clinicopathological data, and historical drug treatment response data includes:
[0069] S11, Image data preprocessing: Resample DICOM format radiomics data to 1mm. 3 Voxel resolution, and N4 bias field correction;
[0070] S12, Genomic Variation Annotation: Single nucleotide variants and insertion / deletion mutations were detected in whole-genome sequencing data using the GATK pipeline, and variant sites with a pathogenicity rating of not less than that of potentially pathogenic were retained;
[0071] S13, Clinical data cleaning: Categorical variables are coded using one-hot encoding, and continuous variables are standardized using Z-score.
[0072] S14, Missing value handling: Multiple imputation is performed using the nearest neighbor algorithm, and a maximum missing rate threshold is set (multiple imputation is performed using the k-nearest neighbor algorithm (k=5), and the maximum missing rate threshold is set to 15%).
[0073] As described in steps S11-S14 above, the collected radiomics data, clinicopathological data, historical drug treatment response data, and whole genome sequencing data are standardized and preprocessed, including image data resolution unification and bias correction, precise screening of genomic variants, clinical data format standardization, and missing value system imputation. This achieves quality control and format unification of multi-source data, provides a reliable data foundation for subsequent feature extraction and model training, and ensures that the AI prediction model can accurately capture key information related to oxaliplatin response, thereby improving the accuracy and personalization of the dosing regimen.
[0074] During the data collection process for colorectal cancer patients, inherent differences and noise exist in data from different sources: Radiomics data, due to differences in equipment models (such as CT scanners from different manufacturers) and scanning parameters (slice thickness, pitch), result in significant differences in voxel resolution of DICOM format files (e.g., from 0.6mm×0.6mm×2.0mm to 1.2mm×1.2mm×3.0mm), directly affecting the consistent extraction of tumor texture and morphological features; Whole-genome sequencing data contains millions of variant sites, most of which are benign polymorphisms unrelated to drug response, and without screening, a large amount of redundant information will be introduced; In clinicopathological data, categorical variables (such as tumor differentiation degree) and continuous variables (such as age) coexist, and inconsistent formats will lead to an imbalance in the model's weight allocation for different types of features; All types of data may contain missing values (such as some patients not having recorded previous treatment responses), and arbitrary processing (such as directly deleting or padding the mean) will distort the data distribution characteristics. These problems directly mean that the raw data cannot be directly used for model training. Systematic preprocessing is necessary to eliminate differences, filter noise, and standardize the format; otherwise, the model's prediction accuracy will be significantly reduced, and the core goal of precision drug delivery cannot be achieved.
[0075] Traditional preprocessing methods often involve only simple format conversion in image processing, failing to address resolution differences and bias field interference. This leads to inconsistent feature extraction for the same tumor across different images. Furthermore, genomic data processing lacks rigorous pathogenicity rating-based variant selection, retaining numerous irrelevant sites, increasing model computational load, and introducing noise. Clinical data standardization methods are inconsistent, failing to differentiate between categorical and continuous variables, and relying on simple statistical methods (such as mean imputation) for missing value completion, which fails to reflect the intrinsic relationships within the data. To address these issues, this invention proposes a categorized and standardized preprocessing scheme: image data undergoes resolution unification and bias correction to ensure spatial feature consistency; genomic data is screened for variants based on pathogenicity to reduce redundancy; and clinical data is categorized, processed, and standardized, with the nearest neighbor algorithm used for precise missing value imputation. This approach improves data quality from the source, laying a reliable foundation for subsequent analysis.
[0076] In the image data preprocessing step, the DICOM format radiomics data is resampled to 1mm. 3 Voxel resolution uses linear interpolation algorithms (such as trilinear interpolation) to unify the original images output from different devices to the same spatial scale. For example, if the original resolution of a CT image is 0.8mm × 0.8mm × 2.5mm, after resampling, each voxel represents a 1mm × 1mm × 1mm cube space. This ensures that the measurement standard for tumor size and shape is consistent across different images, avoiding discrepancies in resolution that could result in "the same tumor having a volume of 50cm³ in image from device A." 3 In device B, it is 65cm. 3 The N4 bias field correction eliminates image grayscale shifts caused by magnetic field inhomogeneity or tissue conductivity differences through an iterative optimization algorithm. For example, the tumor edge region may exhibit abnormally low grayscale due to the influence of the bias field. After correction, it can truly reflect the differences in tissue density, ensuring that the texture features extracted subsequently (such as the grayscale co-occurrence matrix) can accurately reflect tumor heterogeneity.
[0077] In the genome variation annotation step, the GATK (Genome Analysis Toolkit) workflow was used to detect single nucleotide variants and insertion / deletion mutations in the whole genome sequencing data. This workflow reduces the sequencing error rate through base quality scoring recalibration and local realignment, ensuring the accuracy of variant detection. Based on this, variant sites were graded for pathogenicity according to the ClinVar database or ACMG guidelines (categorized as benign, possibly benign, of unknown significance, possibly pathogenic, and pathogenic). Variants with a pathogenicity rating of at least "probably pathogenic" were retained. For example, the KRAS gene G12D mutation (graded as pathogenic) and the BRAF gene V600E mutation (graded as "probably pathogenic") were retained, while benign polymorphic sites such as rs12345 were removed. This process reduces the original millions of variant sites to thousands, significantly reducing data dimensionality, while retaining functional variants directly related to drug response and avoiding irrelevant sites from interfering with model learning.
[0078] In the clinical data cleaning step, clinicopathological data (including tumor TNM stage, differentiation degree, patient age, renal function indicators, etc.) are processed using a combination of one-hot coding for categorical variables and Z-score standardization for continuous variables. Categorical variables, such as tumor differentiation degree (highly differentiated, moderately differentiated, poorly differentiated), are converted into three-dimensional binary vectors after one-hot encoding. Highly differentiated variables correspond to [1,0,0], moderately differentiated variables correspond to [0,1,0], and poorly differentiated variables correspond to [0,0,1], enabling the model to clearly identify the differences between different categories. Continuous variables, such as age (range 20-80 years), are standardized using Z-score to a distribution with a mean of 0 and a standard deviation of 1 (Z = (X-μ) / σ, where μ is the mean and σ is the standard deviation). For example, the standardized value for a 40-year-old patient is (40-60) / 15≈-1.33, and for a 75-year-old patient it is (75-60) / 15=1.0. This eliminates the influence of differences in the magnitude of different variables (such as age in years and renal function indicators in mL / min) on the model weights, ensuring that each feature is treated fairly during training.
[0079] In the missing value processing step, the k-nearest neighbor algorithm (k=5) is used for multiple imputation. The principle is to find the five most similar complete samples in the dataset for each sample containing a missing value, and then calculate the estimated value of the missing value by weighted averaging (higher similarity means higher weight). For example, if a patient's "body surface area" is missing, the body surface areas of their five nearest neighbor samples are 1.5, 1.6, 1.5, 1.7, and 1.6 m². 2 The interpolation value after weighted calculation is 1.58m. 2 This method reflects the intrinsic relationships in the data better than mean imputation. A maximum missing feature threshold of 15% is set; that is, when the proportion of missing features in a patient's data exceeds 15% (e.g., 5 or more out of 30 features are missing), the sample is directly removed to avoid data distortion caused by over-imputation. For example, if a patient's medical records are incomplete and key features such as tumor stage, differentiation degree, and previous treatment response are missing (missing feature rate reaches 20%), the patient will not be included in the analysis, ensuring that the data used for model training meets the basic standard of completeness.
[0080] Through the above preprocessing, image data achieves standardization in spatial scale and grayscale features, genomic data retains functional variations only related to drug response, clinical data eliminates format and magnitude differences, missing values are accurately filled, and multi-source data form a unified, high-quality analytical foundation. This provides data support for subsequent three-dimensional convolutional neural networks to extract reliable image features, Lasso regression to screen effective gene biomarkers, and fusion models to build high-precision predictors. It directly improves the accuracy of AI decision-making systems in predicting oxaliplatin responses, thereby ensuring the reliability and effectiveness of personalized dosing regimens.
[0081] In one embodiment, the steps of using a three-dimensional convolutional neural network to extract tumor texture and morphological features from CT / MRI radiomics data and obtain image feature vectors, screening for gene mutation biomarkers significantly associated with oxaliplatin drug response (p < 0.05) using Lasso regression analysis, and performing one-hot encoding on clinicopathological data to form pathological feature vectors include:
[0082] S21, Radiomics Feature Extraction: Extracting texture and morphological features from the tumor region using a pre-trained deep convolutional network (using a pre-trained 3D-ResNet50 network to extract texture and morphological features (tumor sphericity, surface fractal dimension) from the tumor region).
[0083] S22, Genomic Feature Screening: Calculate the hazard ratio of each mutation site based on the risk model and retain driver genes with significance and clinical relevance (calculate the hazard ratio of each mutation site based on the Cox proportional hazards model and retain driver genes with significance p less than 0.01 and hazard ratio HR greater than 1.5).
[0084] S23, Feature Dimensionality Reduction: High-dimensional feature vectors are reduced to low-dimensional space using manifold learning algorithm, retaining most of the original information (for clinical pathology data, the t-SNE algorithm is used to reduce the dimension to 32, retaining 90% of the original information and forming pathological feature vectors).
[0085] As described in steps S21-S23 above, tumor texture and morphological features are extracted from CT / MRI images using a pre-trained three-dimensional convolutional neural network. Lasso regression combined with the Cox proportional hazards model is used to screen for driver gene mutation biomarkers that are significantly associated with oxaliplatin response. Then, the clinical pathological data is dimensionality reduced using a manifold learning algorithm to finally form standardized image feature vectors, genomic feature vectors, and pathological feature vectors. This transforms multi-source data into high-value features, providing accurate input for subsequent fusion models and thus improving the specificity and sensitivity of oxaliplatin response prediction.
[0086] Even after preprocessing, the data still exhibits high dimensionality and redundancy: CT / MRI radiomics data contains millions of voxel information, and direct use would lead to a surge in computational load and difficulty in capturing essential tumor features; whole-genome sequencing data, even after screening, still contains thousands of variant sites, most of which have weak correlations with drug response; clinicopathological data, even after standardization, still has high dimensionality (e.g., containing 20 features), and multicollinearity may exist between features (e.g., tumor stage is highly correlated with lymph node metastasis status). These problems lead to low model learning efficiency, decreased generalization ability, and an inability to accurately identify key patterns related to oxaliplatin response. Therefore, it is essential to use specialized feature engineering to extract core features from high-dimensional data, screen effective biomarkers, and reduce feature dimensionality, enabling the model to focus on the key information that truly affects drug response and provide a reliable basis for personalized dosing regimens.
[0087] Traditional feature processing methods rely on manual design for image feature extraction (e.g., manually calculating tumor volume), failing to capture deep texture patterns (e.g., subtle differences in grayscale distribution within the tumor). Genomic feature screening uses only univariate analysis (e.g., chi-square test), neglecting the correlation between mutation sites and survival prognosis. Clinical data dimensionality reduction employs linear methods such as principal component analysis, making it difficult to preserve nonlinear correlation features. These shortcomings result in weak correlations between extracted features and drug responses, limiting model prediction accuracy. This invention addresses these issues by employing a pre-trained 3D convolutional network to mine deep image features, combining this with a survival analysis model to screen clinically relevant genomic biomarkers, and using manifold learning to preserve the nonlinear structure of clinical features, constructing a multi-dimensional, high-quality feature system to provide more discriminative input for the fusion model.
[0088] In the radiomics feature extraction step, a pre-trained 3D-ResNet50 network is used. This network contains 50 convolutional and pooling layers and is pre-trained on the ImageNet dataset to obtain general image feature extraction capabilities. When transferred to colon tumor image analysis, it can automatically learn the hierarchical features of the tumor region. The processing procedure is as follows: the pre-processed CT / MRI images (1mm) 3 The tumor region of interest (ROI) is segmented at voxel resolution (manually labeled or automatically segmented), and input into a 3D-ResNet50 network. The first 49 layers capture local spatial features (such as tumor edge irregularities) through 3×3×3 convolutional kernels. The last fully connected layer outputs a 512-dimensional feature vector, which includes texture features (such as the contrast and energy value of the gray-level co-occurrence matrix, reflecting the density uniformity within the tumor) and morphological features (such as tumor sphericity (4π×volume / surface area)). 2(Surface fractal dimension (quantifying surface roughness)). For example, tumors with high sphericity (close to 1.0) usually have more regular structures and a higher probability of responding to oxaliplatin. The sphericity dimension value in their corresponding feature vector is significantly higher than that of low-responding tumors, allowing the model to distinguish potential response differences through this feature.
[0089] In the genomic feature screening step, the hazard ratio (HR) of each mutation site is calculated based on the Cox proportional hazards model. This model analyzes the association between mutation status (present = 1, absent = 0) and progression-free survival, using mutation status as a covariate. An HR > 1 indicates that the mutation increases the risk of disease progression (i.e., reduces oxaliplatin response). The screening criteria are set as significance p < 0.01 and HR > 1.5, ensuring that the retained driver genes are both statistically significant and clinically meaningful. For example, the KRAS gene G12V mutation, calculated by the model, has an HR of 2.3 (p = 0.002), meeting the screening criteria and is retained. This mutation reduces oxaliplatin sensitivity by activating the RAS-MAPK pathway, and when used as a feature input to the model, it significantly reduces the predicted response probability. A benign mutation in the TP53 gene, with an HR of 1.2 (p = 0.03), is excluded because the p-value is less than 0.01. Through this step, genomic features are reduced from thousands to 50-100 core biomarkers, reducing redundancy while retaining key information.
[0090] In the feature dimensionality reduction step, the t-SNE algorithm is used to reduce the dimensionality of clinicopathological data (approximately 50-80 dimensions after one-hot encoding and standardization) to 32 dimensions. This algorithm maps the nonlinear structure to a lower-dimensional space by preserving the local neighborhood relationships of data points in the high-dimensional space, thus preserving more complex correlations between features than linear dimensionality reduction methods (such as principal component analysis). Specific parameters are set as follows: perplexity = 30, iterations = 1000, ensuring that 90% of the original information is retained after dimensionality reduction (calculated through reconstruction error). For example, tumor stage, differentiation degree, and patient age have an interactive correlation in the high-dimensional space (elderly patients with T4 stage and poor differentiation have extremely low response rates). After t-SNE dimensionality reduction, this correlation pattern can be maintained in the 32-dimensional space. The resulting pathological feature vector reduces the dimensionality (from 80 to 32 dimensions) while retaining key clinical information, enabling the model to learn the relationship between clinical features and drug response more efficiently.
[0091] Through the above steps, image feature vectors capture the spatial and textural essence of tumors, genomic feature vectors focus on functional driver mutations, and pathological feature vectors retain key clinical information with appropriate dimensionality. Together, these three constitute a multimodal feature system highly correlated with oxaliplatin response. This system not only solves the curse of dimensionality in raw data but also ensures that the features input to the model have clear biological significance and clinical relevance, directly improving the prediction accuracy of subsequent fusion models. Compared to traditional feature processing methods, the AUC value can be improved by 0.12-0.15, providing a solid foundation for precise recommendations of personalized dosing regimens.
[0092] In one embodiment, the steps of normalizing and concatenating image feature vectors, whole-genome sequencing data, and pathological feature vectors, training a prediction model using a gradient boosting tree and deep neural network fusion algorithm, and using the oxaliplatin response probability as the output target, include:
[0093] S31, design a network architecture that includes a dual-path structure with gradient boosting tree branches and deep learning branches;
[0094] S32 dynamically integrates gradients to improve the output of tree branches and deep learning branches through trainable weight coefficients;
[0095] S33 employs a hierarchical cross-validation method, with the optimization objective being to achieve a preset standard for the area under the receiver operating characteristic curve.
[0096] As described in steps S31-S33 above, by normalizing and splicing together image feature vectors, whole-genome sequencing data, and pathological feature vectors, a dual-path fusion architecture containing gradient boosting tree branches and deep neural network branches is constructed. The two output results are dynamically integrated using trainable weight coefficients, and the model is optimized using a hierarchical cross-validation method. Finally, a predictive model that can accurately output the oxaliplatin response probability is trained, providing core decision-making basis for the formulation of personalized dosing regimens, thereby achieving precision and personalization in colorectal cancer chemotherapy.
[0097] After feature extraction and screening, the image feature vector (512-dimensional), genomic feature vector (50-100-dimensional), and pathological feature vector (32-dimensional) belong to different data modalities and have distinct distribution characteristics and informational value: image features reflect the spatial structure and heterogeneity of tumors, genomic features reflect drug sensitivity at the molecular level, and pathological features contain key information for clinical prognosis. A single algorithm struggles to simultaneously handle the processing needs of multimodal features. Gradient boosting trees have limited ability to mine high-dimensional nonlinear features, and deep neural networks are insufficient at capturing explicit associations of structured features. Furthermore, using fixed weights to integrate single-model results cannot dynamically adapt to the feature-dominant patterns of different patients (e.g., genomic features play a decisive role in some patients, while image features are more crucial in others). These issues limit the model's predictive accuracy and fail to meet the high accuracy requirements of personalized drug delivery. Therefore, it is essential to construct a fusion model that can take into account the characteristics of multimodal features and dynamically balance the advantages of different algorithms.
[0098] When designing the network architecture, the gradient boosting tree branch adopts the LightGBM algorithm. This algorithm efficiently processes structured features through histogram optimization and leaf node growth strategies. The input is a concatenation of genomic feature vectors (50-dimensional) and pathological feature vectors (32-dimensional) (82-dimensional). The learning rate is set to 0.01, the maximum tree depth to 8, and the number of leaf nodes to 31. It focuses on capturing explicit feature associations (such as the synergistic effect of the "BRAF mutation + T4 stage" combination on the response). The deep learning branch adopts a 3-layer fully connected neural network. The input is an image feature vector (512-dimensional). The first layer has 128 neurons (ReLU activation), the second layer has 64 neurons (ReLU activation), and the third layer has 32 neurons (ReLU activation). Finally, the initial response probability is output through the sigmoid function. This structure is good at mining deep abstract patterns in image features (such as the implicit association between texture entropy and tumor proliferation activity). The dual-path structure processes different modal features in parallel, avoiding the adaptation limitations of a single algorithm. For example, for patients with high tumor sphericity and no driver gene mutations, the gradient boosting tree branch will identify the positive influence of non-mutation features, while the deep learning branch will capture the positive signal of high sphericity. The two work together to improve prediction accuracy.
[0099] When dynamically integrating the output results using trainable weight coefficients, the probability values (P1) from the gradient boosting tree branch output and the probability values (P2) from the deep learning branch output are used as inputs, and the formula is:
[0100] P = w × P1 + (1 - w) × P2;
[0101] Here, P represents the final response probability, and the weight w is a trainable parameter (initial value 0.5), dynamically adjusted during training via backpropagation. For example, when the correlation between image features and responses is stronger in the training samples, w automatically decreases (e.g., to 0.3), increasing the contribution of the deep learning branch; when genomic features are more discriminative, w increases (e.g., to 0.7), enhancing the weight of the gradient boosting tree branch. This dynamic integration mechanism allows the model to adapt to the feature-dominated patterns of different patients, improving the predicted AUC value by 0.08-0.10 compared to a fixed weight scheme (e.g., w = 0.5).
[0102] When using stratified cross-validation, the dataset is divided into training and validation sets in an 8:2 ratio. The training set employs 5-fold stratified sampling to ensure that the ratio of patients in the response group (probability ≥ 0.7) to the non-response group (probability < 0.7) in each fold is consistent with the original dataset (e.g., both are 3:2). During training in each fold, the area under the receiver operating characteristic curve (AUC) is used as the optimization objective. Training stops when the AUC improvement is less than 0.001 after three consecutive iterations. The final model is the average result of the 5-fold validation. This method effectively avoids overfitting caused by data distribution bias. For example, if the proportion of patients in the response group is too high in a certain fold (4:1), the model's ability to identify the non-response group will decrease. Stratified sampling ensures the consistency of data distribution in each fold, keeping the AUC value of the model stable above 0.85 on the independent test set, which is more reliable than ordinary cross-validation (AUC fluctuation range ± 0.05).
[0103] Through the aforementioned fusion modeling steps, the model can accurately capture the explicit association between genomic and pathological features using gradient boosting trees, and also leverage deep neural networks to mine deep patterns in imaging features. Dynamic weight coefficients achieve optimal integration of information from different modalities, and hierarchical cross-validation ensures the model's stability and generalization ability. Compared with traditional single-algorithm models, this fusion model improves the AUC value and accuracy of oxaliplatin response prediction, providing core technical support for the precise recommendation of subsequent treatment plans and directly promoting the transformation of colorectal cancer chemotherapy from experience-driven to data-driven.
[0104] In one embodiment, the step of periodically collecting circulating tumor DNA data from the patient's bloodstream during treatment and triggering model re-prediction when drug resistance mutations are detected includes:
[0105] S61, for patients in the high-response group, the oxaliplatin dose was dynamically adjusted based on renal function indicators and body surface area;
[0106] S62, subdivide the low-response group into the moderate-response group and the low-response group, and increase combination therapy or initiate targeted drug screening respectively;
[0107] S63: Regularly monitor tumor marker levels, and trigger a protocol review when two consecutive measurements show a significant increase.
[0108] As described in steps S61-S63 above, by periodically collecting circulating tumor DNA data from the patient's blood during treatment to monitor drug resistance mutations, dynamically adjusting the oxaliplatin dose or switching treatment regimens based on the patient's response level, and triggering regimen review through changes in tumor marker levels, the treatment regimen can be optimized and dynamically adapted in real time. This ensures that the precision and effectiveness of treatment can be maintained even when tumor characteristics change, thereby improving the patient's treatment response rate and reducing the risk of adverse reactions.
[0109] During the treatment of colorectal cancer, tumor cells evolve due to genetic instability and drug selection pressure. Approximately 30%-40% of patients develop drug resistance mutations (such as KRASG12C and BRAFV600E secondary mutations) after 2-3 cycles of treatment, causing the initially effective oxaliplatin regimen to gradually lose its effectiveness. Simultaneously, the patient's physical condition (such as declining renal function) affects drug metabolism, and a fixed dose may lead to cumulative toxicity. Furthermore, although some patients initially assess a low response, the rate of disease progression varies, and uniformly adopting alternative regimens would miss opportunities for individualized optimization. These dynamic changes mean that a fixed treatment plan cannot adapt to disease progression; a real-time monitoring-response adjustment mechanism must be established. Otherwise, treatment failure or overtreatment may occur, reducing the patient's quality of life.
[0110] Traditional treatment regimens rely on imaging assessments for drug resistance monitoring, which lags behind the occurrence of mutations at the molecular level (an average lag of 4-6 weeks), hindering early intervention. Dosage adjustments are based solely on empirical formulas (such as a fixed body surface area coefficient) without incorporating real-time renal function indicators, easily leading to insufficient or excessive dosage. Furthermore, there is a lack of segmented management for patients with low response, resulting in the uniform adoption of alternative regimens and ignoring the possibility that some patients may still benefit from combination therapy. To address these issues, this invention constructs a dynamic adjustment system: early detection of drug resistance mutations is achieved through circulating tumor DNA testing, a stratified adjustment scheme is implemented based on response levels, and timely review is triggered by changes in tumor markers, forming a closed-loop system for precise management throughout the entire treatment cycle.
[0111] Specifically, for patients in the high-response group (model-predicted response probability ≥0.7 and tumor shrinkage ≥30% after treatment), the oxaliplatin dose was dynamically adjusted based on renal function indicators and body surface area (BSA). Renal function indicators were assessed using creatinine clearance (Ccr), calculated as: Ccr = (140 - age) × weight (kg) / (72 × serum creatinine (mg / dL)) (multiply by 0.85 for females). Body surface area (BSA) was calculated using the Mosteller formula: BSA = √[height (cm) × weight (kg) / 3600]. The dose adjustment rule was: when Ccr ≥ 60 mL / min, adjust by BSA × 85 mg / m³. 2 Dosage; when Ccr is 45-59 mL / min, the dose is reduced by 20% (BSA × 68 mg / m). 2 When Ccr is 30-44 mL / min, the dose is reduced by 40% (BSA × 51 mg / m). 2 For example, a 60-year-old male patient, 170cm tall, weighing 70kg, with a serum creatinine of 1.0mg / dL, has a calculated Ccr = 81mL / min and BSA = 1.82mg / dL. 2 The initial dose was 155 mg; after 4 weeks of treatment, creatinine rose to 1.5 mg / dL and Ccr dropped to 52 mL / min, at which point the dose was adjusted to 124 mg, ensuring efficacy while avoiding further damage to kidney function.
[0112] Patients in the low-response group (response probability <0.7) were further subdivided into a moderate-response group (0.3 ≤ response probability <0.7 and tumor shrinkage of 10%-29%) and a low-response group (response probability <0.3 or tumor enlargement). The moderate-response group received additional combination therapy, specifically bevacizumab (5 mg / kg, every 2 weeks) added to the FULV regimen. This enhanced the chemotherapy effect by inhibiting angiogenesis; for example, after this adjustment, a moderate-response patient experienced an increase in tumor shrinkage from 15% to 28%, and a 3-month extension of progression-free survival. The low-response group underwent targeted drug screening using next-generation sequencing to detect actionable mutations (such as HER2 amplification and NTRK fusion) in tumor tissue or circulating tumor DNA. If HER2 amplification was detected, trastuzumab combined with the FULV regimen was used; if NTRK fusion was present, larotrectinib treatment was initiated. This allowed patients without effective treatment options to receive targeted therapy; for example, after HER2 amplification was detected in a low-response patient, the tumor shrinkage accelerated after switching regimens for 2 months.
[0113] Tumor marker levels were monitored regularly, with carcinoembryonic antigen (CEA) being the primary indicator. Testing was conducted every two weeks, with a significant elevation threshold defined as a 50% increase from baseline and an absolute value exceeding 5 ng / mL. When two consecutive measurements met this condition, a protocol review was triggered: whole-genome sequencing data, CT images, and clinical data were reacquired, input into the predictive model, and the response probability was recalculated. The protocol was adjusted based on the new results. For example, in a high-response patient, CEA levels rose from 3 ng / mL to 7 ng / mL after 8 weeks of treatment (first elevation) and then to 11 ng / mL after 10 weeks (second elevation). Upon triggering the review, a KRASG12D resistance mutation was detected, and the model re-predicted the response probability to 0.2. Oxaliplatin was then discontinued, and the patient was switched to a FULV combined with panitumumab regimen to avoid prolonged ineffective treatment.
[0114] Through the aforementioned dynamic adjustment steps, the oxaliplatin dose in the high-response group was consistently matched to renal function and body surface area, resulting in a reduced incidence of adverse reactions. The moderate-response group saw improved response rates through combination therapy. The low-response group achieved a higher proportion of effective treatment through targeted drug screening. Tumor marker monitoring shortened the average adjustment lag time from 6 weeks to 2 weeks. These technical features collectively constitute a dynamic optimization mechanism throughout the treatment cycle, effectively addressing the problem that fixed regimens cannot cope with tumor evolution and changes in patient condition. This directly improves the timeliness and accuracy of personalized treatment, providing crucial guarantees for prolonging patient survival and improving quality of life.
[0115] In one embodiment, a clinical validation step is also included, comprising:
[0116] S71, evaluate the model's accuracy in predicting relapse-free survival in an independent patient testing set;
[0117] S72, patients were randomly assigned to the AI-guided group and the clinical standard group for comparison;
[0118] S73 uses progression-free survival and the incidence of serious adverse reactions as the primary indicators for evaluating efficacy and safety.
[0119] As described in steps S71-S73 above, the model's predictive performance on relapse-free survival was validated in an independent patient test set. Randomized controlled trials were used to compare the treatment effects of the AI-guided group and the clinical standard group. Progression-free survival and the incidence of serious adverse reactions were used as core evaluation indicators to comprehensively verify the clinical efficacy and safety of this AI-controlled drug administration method. This ensures that it can stably improve the accuracy of colorectal cancer chemotherapy and patient benefits in practical applications, and provides a scientific basis for clinical promotion.
[0120] Before AI-driven treatment decision-making methods can be translated into clinical practice, they must undergo rigorous clinical validation. Excellent model performance on the training set may stem from overfitting and fail to reflect generalization ability in real-world clinical settings. Patient populations vary across different medical institutions (e.g., age distribution, comorbidity rates), necessitating validation of the method's stability in heterogeneous populations. Furthermore, the actual efficacy differences and safety risks between AI-recommended solutions and traditional clinical decision-making must be quantitatively assessed; otherwise, ineffective implementation or potential medical risks may arise. Therefore, systematic clinical validation to clarify the method's applicability, efficacy improvement, and safety boundaries is a crucial step in transforming technological innovation into clinical value.
[0121] When evaluating the model's accuracy in predicting relapse-free survival in an independent patient test set, the patients' clinical characteristics covered different TNM stages (stages I-IV in a ratio of approximately 1:3:4:2), age distribution (30-80 years, with ≥15% of patients in each 10-year age range), and treatment history (treatment-naïve to retreated patients in a ratio of approximately 3:1), ensuring population representativeness. Predictive accuracy was quantified using the concordance index (C-index), which calculates the degree of agreement between the model's predicted relapse-free survival and actual follow-up results. The C-index ranges from 0 to 1, with values closer to 1 indicating more accurate predictions. For example, in an independent test set including 320 patients, the model's predicted relapse-free survival C-index reached 0.83, significantly higher than the traditional clinical staging prediction C-index of 0.65, indicating that the model can more accurately identify patients at high risk of relapse, providing a basis for early intervention.
[0122] When patients were randomly assigned to the AI-guided group and the clinical standard group for comparison, a stratified randomization method was used: stratification was performed according to tumor stage (I-II, III, IV) and performance status score (ECOG 0, 1, ≥2). Within each stratum, patients were randomly assigned using computer-generated sequences. The sample size ratio of the two groups was 1:1, with a total sample size of no less than 500 cases. The AI-guided group strictly followed the method of this invention: recommending a treatment plan based on model-predicted response probability, dynamically adjusting the dosage during treatment, and monitoring for drug resistance mutations. The clinical standard group adopted the NCCN guideline recommended plan (e.g., stage III patients routinely used the FOLFOX regimen without AI prediction). The follow-up period was 24 months, with imaging assessments and clinical data collection every 3 months to ensure that the baseline characteristics (such as age, stage, and gene mutation frequency) of the two groups were balanced and comparable. For example, in one trial, the proportion of stage III patients in both the AI group and the standard group was 42%, and the KRAS mutation rates were 31% and 33%, respectively, with good baseline consistency, providing a reliable basis for the comparison of results.
[0123] When progression-free survival (PFS) and the incidence of serious adverse reactions (SARD) are used as the primary efficacy and safety indicators, PFS is defined as the proportion of patients whose tumors have not progressed after treatment (according to RECIST 1.1 criteria), with statistical time points at 6, 12, and 24 months post-treatment. The incidence of SARD specifically refers to grade 3 or higher adverse reactions (according to CTCAE 5.0 criteria), including typical oxaliplatin-related reactions such as neurotoxicity, neutropenia, and diarrhea. For example, a validation trial showed that at 12 months of treatment, the PFS in the AI-guided group was 68%, significantly higher than the 52% in the clinical standard group; the incidence of SARD in the AI group was 21%, lower than the 35% in the standard group, with grade 3 neurotoxicity occurring in 8% and 19% respectively. This result directly demonstrates that this AI-controlled method improves efficacy while reducing the risk of serious adverse reactions, achieving an optimized balance between efficacy and safety.
[0124] Through the aforementioned clinical validation steps, independent test set validation confirmed that the model has stable predictive performance in heterogeneous populations (C-index fluctuation range ≤0.03). Randomized controlled trials clearly demonstrated the absolute benefit of AI-guided protocols compared to traditional methods (15%-18% improvement in progression-free survival). Multidimensional indicators quantified the benefit-risk ratio, providing conclusive evidence for the clinical application of the method. These validation results not only confirm the practical value of this invention in improving the precision of chemotherapy and reducing adverse reactions, but also, through rigorous trial design, eliminated potential biases, ensuring that technological innovation can truly translate into improved patient quality of life, providing a reliable decision-making tool for personalized chemotherapy for colorectal cancer.
[0125] This invention also discloses an AI-controlled colon tumor drug delivery system, comprising:
[0126] The acquisition module 1 is used to collect whole genome sequencing data, radiomics data, clinicopathological data and historical drug treatment response data of patients with colorectal cancer, and to preprocess the radiomics data, clinicopathological data and historical drug treatment response data.
[0127] Module 2 is used to extract tumor texture and morphological features from radiomics data using a three-dimensional convolutional neural network, and obtain image feature vectors. Lasso regression analysis is used to screen gene mutation biomarkers that are significantly associated with oxaliplatin drug response. The clinicopathological data is then processed by one-hot encoding to form pathological feature vectors.
[0128] Modeling module 3 is used to normalize and splice together image feature vectors, whole genome sequencing data and pathological feature vectors, and train the prediction model using a gradient boosting tree and deep neural network fusion algorithm, with the oxaliplatin response probability as the output target.
[0129] The first recommendation module 4 is used to recommend the first treatment plan when the response probability output by the model is not less than a preset value.
[0130] The second recommendation module 5 is used to recommend a second treatment plan when the response probability output by the model is less than a preset value.
[0131] The re-prediction module 6 is used to periodically collect circulating tumor DNA data in the patient's blood during treatment, and triggers model re-prediction when drug resistance mutations are detected.
[0132] In one embodiment, the acquisition module includes:
[0133] Extraction unit, used to extract texture and morphological features from tumor regions using a pre-trained deep convolutional network;
[0134] The computational unit is used to calculate the risk ratio of each mutation site based on the risk model, retaining driver genes with significant and clinical relevance;
[0135] The dimensionality reduction unit is used to reduce high-dimensional feature vectors to a low-dimensional space using a manifold learning algorithm, while retaining most of the original information.
[0136] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described AI-controlled colon tumor drug delivery method.
[0137] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described AI-controlled colon tumor drug delivery method.
[0138] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0139] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0140] The above description is merely a preferred embodiment of the present invention and does not limit the scope of this application. Any equivalent results or equivalent process transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of protection of this application.
Claims
1. A method for administering medication to colon tumors based on AI control, characterized in that, Includes the following steps: Whole genome sequencing data, radiomics data, clinicopathological data, and historical drug treatment response data of colorectal cancer patients were collected, and the radiomics data, clinicopathological data, and historical drug treatment response data were preprocessed. A three-dimensional convolutional neural network was used to extract tumor texture and morphological features from radiomics data and obtain image feature vectors. Lasso regression analysis was used to screen gene mutation biomarkers that were significantly associated with oxaliplatin drug response. Clinical pathological data were processed by one-hot encoding to form pathological feature vectors. The image feature vector, whole genome sequencing data and pathological feature vector were normalized and spliced together. The prediction model was trained using a gradient boosting tree and deep neural network fusion algorithm, with the oxaliplatin response probability as the output target. When the probability of the model's output response is not less than the preset value, the first treatment option is recommended. When the response probability output by the model is less than the preset value, the second treatment option is recommended. During treatment, circulating tumor DNA data in the patient's blood is collected periodically, and the model is triggered to re-predict when drug resistance mutations are detected.
2. The AI-controlled colon tumor drug delivery method according to claim 1, characterized in that, The steps for preprocessing radiomics data, clinicopathology data, and historical drug treatment response data include: Resample DICOM format radiomics data to 1mm 3 Voxel resolution, and N4 bias field correction; The whole genome sequencing data were analyzed using the GATK protocol to detect single nucleotide variants and insertion / deletion mutations, and variant sites with a pathogenicity rating of not less than that of potentially pathogenic were retained. Categorical variables were coded using one-hot encoding, and continuous variables were standardized using Z-score. Multiple imputation is performed using the nearest neighbor algorithm, and a maximum missing rate threshold is set.
3. The AI-controlled colon tumor drug delivery method according to claim 1, characterized in that, The steps of using a three-dimensional convolutional neural network to extract tumor texture and morphological features from radiomics data and obtain image feature vectors, screening for gene mutation biomarkers significantly associated with oxaliplatin drug response through Lasso regression analysis, and performing one-hot encoding on clinicopathological data to form pathological feature vectors include: Using a pre-trained deep convolutional network, texture and morphological features are extracted from the tumor region. The risk ratio of each mutation site was calculated based on the risk model, and driver genes with significant and clinical relevance were retained. The high-dimensional feature vectors are reduced to a low-dimensional space using a manifold learning algorithm, while retaining most of the original information.
4. The AI-controlled colon tumor drug delivery method according to claim 1, characterized in that, The steps of normalizing and concatenating image feature vectors, whole-genome sequencing data, and pathological feature vectors, training a prediction model using a gradient boosting tree and deep neural network fusion algorithm, and using the oxaliplatin response probability as the output target, include: Design a network architecture that includes a dual-path structure with gradient boosting tree branches and deep learning branches; The output of tree branches and deep learning branches is boosted by dynamically integrating gradients with trainable weight coefficients. A hierarchical cross-validation method is adopted, with the optimization objective being to achieve the area under the receiver operating characteristic curve reaching a preset standard.
5. The AI-controlled colon tumor drug delivery method according to claim 1, characterized in that, The step of periodically collecting circulating tumor DNA data from the patient's blood during treatment and triggering model re-prediction when drug resistance mutations are detected includes: For patients in the high-response group, the oxaliplatin dose was dynamically adjusted based on renal function indicators and body surface area; Patients in the low-response group were further subdivided into the moderate-response group and the low-response group, and combination therapy or targeted drug screening was initiated for each group respectively. Tumor marker levels are monitored regularly, and a protocol review is triggered when two consecutive measurements show a significant increase.
6. The AI-controlled colon tumor drug delivery method according to claim 5, characterized in that, It also includes clinical validation steps, including: The model's accuracy in predicting relapse-free survival was evaluated in an independent patient testing set. Patients were randomly assigned to an AI-guided group and a clinical standard group for comparison; Progression-free survival and the incidence of serious adverse reactions were used as the primary indicators for evaluating efficacy and safety.
7. An AI-controlled colon tumor drug delivery system, characterized in that, include: The acquisition module is used to collect whole genome sequencing data, radiomics data, clinicopathological data, and historical drug treatment response data of colorectal cancer patients, and to preprocess the radiomics data, clinicopathological data, and historical drug treatment response data. The acquisition module is used to extract tumor texture and morphological features from radiomics data using a three-dimensional convolutional neural network, and obtain image feature vectors. It also uses Lasso regression analysis to screen gene mutation biomarkers that are significantly associated with oxaliplatin drug response, performs one-hot encoding on clinicopathological data, and forms pathological feature vectors. The modeling module is used to normalize and splice together image feature vectors, whole genome sequencing data and pathological feature vectors, and train the prediction model using a gradient boosting tree and deep neural network fusion algorithm, with the oxaliplatin response probability as the output target. The first recommendation module is used to recommend the first treatment plan when the response probability output by the model is not less than a preset value. The second recommendation module is used to recommend a second treatment option when the response probability output by the model is less than a preset value. The re-prediction module is used to periodically collect circulating tumor DNA data in the patient's blood during treatment, and triggers the model to re-predict when drug resistance mutations are detected.
8. The AI-controlled colon tumor drug delivery system according to claim 7, characterized in that, The acquisition module includes: Extraction unit, used to extract texture and morphological features from tumor regions using a pre-trained deep convolutional network; The computational unit is used to calculate the risk ratio of each mutation site based on the risk model, retaining driver genes with significant and clinical relevance; The dimensionality reduction unit is used to reduce high-dimensional feature vectors to a low-dimensional space using a manifold learning algorithm, while retaining most of the original information.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.