Colon cancer lymph node metastasis risk assessment method and system
By performing multi-level screening and analysis of imaging and pathological data of colon cancer patients, combining ROC curves and Logistic regression, risk assessment is performed using Nomogram model, the problem of relying on a single indicator in the existing technology is solved, and a more accurate and comprehensive one-time assessment is achieved.
Patent Information
- Application Number
- CN202510184770.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing methods of colon cancer lymph node metastasis risk assessment are too dependent on a single evaluation indicator, ignoring the correlation between multiple clinical indicators and their comprehensive impact, resulting in incomplete and accurate assessment.
By obtaining the patient's imaging data and pathological characteristic data, standardized preprocessing and multi-level screening were performed, combined with ROC curve analysis and multivariate Logistic regression analysis, an abnormal screening index set was generated, and abnormal grading was performed through the Nomogram model to finally obtain the final abnormality evaluation information.
A more comprehensive and accurate assessment of the risk of lymph node metastasis in colon cancer has been achieved, improving the accuracy and reliability of the assessment, and helping to develop personalized treatment plans to avoid overtreatment.
Smart Images

Figure CN120015325A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of colon cancer lymph node detection, and in particular to a colon cancer lymph node metastasis risk assessment method and system. Background Art
[0002] Colon cancer is a common malignant tumor of the digestive tract, and its lymph node metastasis status is a key factor affecting the patient's prognosis and treatment options. With the continuous development of medical imaging technology and pathological detection methods, how to accurately assess the risk of lymph node metastasis in patients with colon cancer in order to develop individualized treatment strategies has become one of the current research focuses. Existing lymph node metastasis risk assessment methods often rely too much on a single evaluation indicator, such as only considering imaging features or pathological features, while ignoring the correlation between multiple clinical indicators and their combined impact. Summary of the invention
[0003] The main purpose of the present invention is to provide a method and system for assessing the risk of lymph node metastasis in colon cancer, which can more comprehensively assess the abnormal information of lymph node metastasis.
[0004] To achieve the above object, the present invention provides a method for assessing the risk of lymph node metastasis of colon cancer, comprising: Obtain the patient's imaging data and pathological characteristic data, and perform standardized preprocessing to obtain indicator data; Performing preliminary abnormal screening on the indicator data based on clinical standards to obtain a preliminary abnormal factor group; Performing ROC curve analysis on the preliminary abnormal factor group to obtain multimodal abnormal data; Performing multivariate logistic regression analysis and screening on the multimodal abnormal data to obtain an abnormal screening indicator set; Performing preliminary indicator prediction on the abnormal screening indicator set to obtain preliminary abnormal prediction data; Inputting the preliminary abnormal prediction data into a preset Nomogram model for abnormal analysis to obtain an initial abnormal index; The initial abnormality index is graded by curves, and an abnormality assessment is performed to obtain final abnormality assessment information.
[0005] Furthermore, the imaging data and pathological characteristic data of the patient are obtained, and standardized preprocessing is performed to obtain index data, including: Performing region segmentation processing on the imaging data to obtain target region image data; Extracting feature vectors from the target area image data to obtain initial image feature data; Digitally scanning the pathological characteristic data to obtain a digital pathological image; Performing image block processing according to the digital pathological image to obtain a local area image set; Performing color space conversion processing on the local area image set to obtain HSV space feature data; Perform morphological feature extraction according to the HSV spatial feature data to obtain a pathological morphological feature set; Performing feature fusion on the initial image feature data and the pathological morphology feature set to obtain fused feature data; Performing normalization processing on the fused feature data to obtain a standardized feature vector; A feature index table is established according to the standardized feature vector to obtain the indicator data.
[0006] Furthermore, the index data is preliminarily screened for abnormalities based on clinical criteria to obtain a preliminary abnormal factor group, including: Classifying the index data into clinical manifestation levels to obtain multi-dimensional clinical classification data; Constructing a correlation matrix of each dimension for the multidimensional clinical classification data to obtain an indicator correlation matrix; Performing hierarchical clustering on the indicator association matrix to obtain indicator group data; The intra-group heterogeneity of the index group data was calculated according to the preset Kruskal-Wallis test algorithm to obtain the inter-group difference coefficient; Performing threshold screening on the inter-group difference coefficient to obtain a key indicator set; Calculate the single factor abnormality ratio according to the key indicator set to obtain abnormal contribution data; The abnormal contribution data is subjected to abnormal grouping processing to obtain the preliminary abnormal factor group.
[0007] Furthermore, the ROC curve analysis is performed on the preliminary abnormal factor group to obtain multimodal abnormal data, including: Calculating the true positive rate and the false positive rate of the preliminary abnormal factor group to obtain initial evaluation coordinates; Perform scattered point distribution sampling according to the initial evaluation coordinates to obtain predicted sampling points; Performing a cubic spline curve interpolation operation on the predicted sampling points to obtain abnormal interpolation parameters; Constructing a piecewise function for the abnormal interpolation parameter to obtain a piecewise function set of ROC curves; Performing curve integration operation according to the ROC curve piecewise function set to obtain initial area under the curve data; Performing interval correction on the initial area under the curve data to obtain a corrected curve area value; Calculating the confidence interval range according to the corrected curve area value to obtain a prediction accuracy index; The prediction accuracy index is divided into segmented thresholds, and the sensitivity and specificity of each segment are calculated to obtain graded evaluation data; Abnormal quantitative scoring is performed according to the graded evaluation data to obtain multimodal abnormal data.
[0008] Furthermore, the multivariate logistic regression analysis and screening of the multimodal abnormal data is performed to obtain an abnormal screening index set, including: Performing a multicollinearity test on the multimodal abnormal data to obtain an independence evaluation value of the abnormal factors; Performing threshold screening on the independence evaluation values of the abnormal factors to obtain an independent prediction factor set; Performing multivariate logistic regression coefficient calculation on the independent prediction factor set to obtain an abnormal prediction weight coefficient; Optimizing and iterating the anomaly prediction weight coefficient to obtain an anomaly scoring parameter set; Performing prediction analysis and processing according to the abnormal scoring parameter set to obtain prediction accuracy data; Performing threshold evaluation analysis on the prediction accuracy data to obtain an evaluation reliability result; The independent prediction factor set is sorted by importance according to the reliability evaluation result to obtain the final abnormal screening indicator set.
[0009] Furthermore, the performing preliminary indicator prediction on the abnormal screening indicator set to obtain preliminary abnormal prediction data includes: Performing combined parameter optimization on the abnormal screening index set to obtain an optimized index set; Perform anomaly classification calculation according to the optimized indicator set to obtain pre-classified anomaly values; Performing cross-validation analysis on the pre-classified outliers to obtain a validation data set; Perform Bootstrap sampling calculation on the verification data set to obtain a resampled prediction set; Performing multi-algorithm fusion on the resampled prediction set to obtain an integrated probability value; Perform abnormal hierarchical division according to the integrated probability value to obtain multi-level abnormal data; Probability calibration is performed according to the multi-level abnormal data to obtain the preliminary abnormal prediction data.
[0010] Furthermore, the preliminary abnormal prediction data is input into a preset Nomogram model for abnormal analysis to obtain an initial abnormal index, including: Inputting the preliminary abnormality prediction data into the Nomogram model, performing standardized input processing on the preliminary abnormality prediction data through the data standardization layer of the Nomogram model to obtain standardized feature data; Performing interval division and basic score calculation on the standardized feature data through the variable scoring layer of the Nomogram model to obtain variable scoring data; Performing prediction contribution analysis and weight coefficient generation on the standardized feature data through the weight allocation layer of the Nomogram model to obtain feature weight data; The variable scoring data and the feature weight data are weighted and fused through the single scoring layer of the Nomogram model to obtain weighted scoring data; The weighted score data is scored and accumulated by the total score calculation layer of the Nomogram model to obtain abnormal total score data; Performing Logistic function probability mapping and abnormality calibration on the abnormal total score data through the probability conversion layer of the Nomogram model to obtain an initial abnormality probability; The initial abnormal probability is subjected to nonlinear correction and threshold analysis through the abnormal evaluation layer of the Nomogram model to obtain and output the initial abnormal index.
[0011] Furthermore, the initial abnormality index is graded by curves, and abnormality evaluation is performed to obtain final abnormality evaluation information, including: Performing probability interval division calculation on the initial abnormality index to obtain a transfer abnormality probability segment; Solving the density function according to the transfer abnormality probability section to obtain an abnormality density distribution curve; Performing numerical integration operation on the abnormal density distribution curve to obtain cumulative abnormal distribution data; Extract critical points according to the accumulated abnormal distribution data to obtain an abnormal threshold sequence; Performing interval segmentation processing on the abnormal threshold sequence to obtain an abnormal level determination point set; According to the abnormality level determination point set, the initial abnormality index is graded and mapped to obtain an initial abnormality level; Verifying the abnormal preliminary judgment level to obtain an abnormal verification result; A semantic description is generated according to the abnormality verification result to obtain the abnormality assessment information.
[0012] The present invention also provides a colon cancer lymph node metastasis risk assessment system, which is applied to any of the colon cancer lymph node metastasis risk assessment methods described above, comprising: An acquisition module, which is used to obtain the patient's imaging data and pathological characteristic data, and perform standardized preprocessing to obtain indicator data; An analysis module, the analysis module is used to perform preliminary abnormal screening on the indicator data based on clinical standards to obtain a preliminary abnormal factor group; A correlation module, wherein the correlation module is used to perform ROC curve analysis on the preliminary abnormal factor group to obtain multimodal abnormal data; A processing module, wherein the processing module is used to perform multivariate logistic regression analysis and screening on the multimodal abnormal data to obtain an abnormal screening indicator set; A control module, the control module is used to perform preliminary indicator prediction on the abnormal screening indicator set to obtain preliminary abnormal prediction data; An execution module, wherein the execution module is used to input the preliminary abnormality prediction data into a preset Nomogram model for abnormality analysis to obtain an initial abnormality index; A generation module is used to perform curve grading on the initial abnormality index and perform abnormality evaluation to obtain final abnormality evaluation information.
[0013] The present invention provides a method and system for assessing the risk of lymph node metastasis of colon cancer, which has the following beneficial effects: By comprehensively analyzing the patient's imaging data and pathological feature data and combining clinical standards for multi-level screening, the abnormal information of lymph node metastasis can be more comprehensively evaluated, thereby improving the accuracy of abnormal evaluation. By integrating multimodal abnormal data into ROC curve analysis and performing association analysis based on clinical indicators, a refined evaluation of lymph node metastasis is achieved, which helps to formulate personalized treatment plans and avoid overtreatment. Systematic analysis of abnormal screening indicators based on Logistic regression can ensure that the evaluation model has stable predictive efficiency in different clinical scenarios, reduce evaluation bias, and improve the overall accuracy of the model. By comprehensively analyzing multi-dimensional abnormal prediction data, using the Nomogram model for abnormal grading, and using the curve grading strategy to achieve accurate evaluation of abnormalities, the prediction effect is effectively improved and the misjudgment rate is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a flow chart of a method for assessing risk of lymph node metastasis of colon cancer provided by the present invention; Figure 2 This is a structural diagram of a colon cancer lymph node metastasis risk assessment system provided by the present invention.
[0015] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0017] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods.
[0018] Reference Figure 1 As shown, the present invention provides 1. A method for assessing the risk of lymph node metastasis of colon cancer, characterized by comprising: Step S1: Obtain the patient's imaging data and pathological characteristic data, and perform standardized preprocessing to obtain index data; Step S2: Perform preliminary abnormal screening on the indicator data based on clinical criteria to obtain a preliminary abnormal factor group; Step S3: Perform ROC curve analysis on the preliminary abnormal factor group to obtain multimodal abnormal data; Step S4: Perform multivariate logistic regression analysis and screening on the multimodal abnormal data to obtain an abnormal screening indicator set; Step S5: Preliminary indicator prediction is performed on the abnormal screening indicator set to obtain preliminary abnormal prediction data; Step S6: inputting the preliminary abnormal prediction data into the preset Nomogram model for abnormal analysis to obtain an initial abnormal index; Step S7: Curve grading of the initial abnormality index and abnormality evaluation are performed to obtain final abnormality evaluation information.
[0019] Based on the above steps, the detailed process is as follows: Step S1: Collect the original data of the patient's CT, MRI and other imaging examinations, including key features such as tumor size, location, depth of infiltration, and invasion of surrounding tissues. Synchronously collect the patient's pathological examination results, covering pathological features such as histological type, degree of differentiation, and vascular invasion of biopsy or surgical specimens. Perform standardization on all collected data to eliminate data differences between different examination equipment and different hospitals. Standardization includes normalization of numerical data, unifying data of different dimensions to the same scale; encoding conversion of categorized data, and conversion of text descriptions into computable numerical forms. The data cleaning stage removes outliers, processes missing values, and ensures the integrity and reliability of the data. The indicator data formed is a structured and standardized data set, which provides a basis for subsequent analysis.
[0020] Step S2: Perform preliminary abnormal screening of indicator data based on clinical standards to obtain a preliminary abnormal factor group. According to the existing clinical diagnosis and treatment guidelines, establish clinical screening criteria for lymph node metastasis risk assessment. The screening criteria include recognized high-risk factors such as tumor size (such as >5cm), invasion depth (such as T3 / T4 stage), and degree of differentiation (poor differentiation). According to these clinical standards, systematic screening of standardized indicator data is carried out to identify abnormal indicators related to lymph node metastasis. The screening process is combined with univariate analysis to calculate the correlation between each indicator and lymph node metastasis. Analyze the clinical correlation between indicators and eliminate duplicate or redundant information. This step narrows the scope of indicators of interest and establishes a preliminary abnormal factor group containing potential predictive factors, laying the foundation for subsequent statistical analysis.
[0021] Step S3: Perform receiver operating characteristic (ROC) curve analysis on the screened preliminary abnormal factor group. For each potential predictive factor, calculate the sensitivity and specificity at different critical values, draw the ROC curve, and calculate the area under the curve (AUC). ROC analysis determines the optimal critical value of each predictive factor to obtain ideal specificity while maintaining high sensitivity. By comparing the ROC curves and AUC values of each predictive factor, the predictive efficacy of each factor is evaluated. Integrate the characteristics of different modal data (imaging, pathology, etc.), classify and organize indicators with good predictive efficacy, and construct a multimodal abnormal data set. This data set contains the optimal critical value of the predictive factor and the prediction efficacy evaluation results, providing basic data support for multiple regression analysis.
[0022] Step S4: Multivariate Logistic regression analysis was performed on the obtained multimodal abnormal data. The lymph node metastasis status was used as the dependent variable (binary variable: 0 for no metastasis and 1 for metastasis), and all the selected predictive factors were used as independent variables to construct a multivariate Logistic regression model. In the specific modeling process, multicollinearity diagnosis was performed on all variables, variance inflation factor (VIF) was calculated, and variables with VIF>10 were eliminated. Subsequently, forward, backward or stepwise regression methods were used to screen out statistically significant variables with P<0.05 as the inclusion criteria and P>0.10 as the elimination criteria. The regression coefficient (β value), standard error (SE), Wald test statistic, odds ratio (OR value) and its 95% confidence interval were calculated for each selected variable. The interaction terms between the predictive factors were further analyzed, and the statistically significant interaction terms were included in the model. The overall goodness of fit of the model was tested, and indicators such as -2 times log-likelihood (-2LL), Cox&Snell R² and Nagelkerke R² were calculated. The Hosmer-Lemeshow test was performed to evaluate the calibration of the model. Finally, the factors with significant predictive value were combined to form an abnormal screening index set, which included not only the weight coefficient of each predictive factor, but also its standardized regression coefficient to reflect the relative contribution of each factor to the prediction of lymph node metastasis risk.
[0023] Step S5: Based on the multivariate logistic regression equation established in step S4, the prediction model was constructed and validated. All research subjects were randomly divided into a training set (70%) and a validation set (30%). In the training set, the actual index data of the patients were substituted into the regression equation to calculate the predicted probability of lymph node metastasis for each patient. The prediction results were internally validated, and the stability and reliability of the model were evaluated using the Bootstrap resampling technique (sampling times were not less than 1000 times). At the same time, the 10-fold cross-validation method was used to randomly divide the data into 10 parts, and 9 parts of the data were used in turn for modeling and 1 part of the data for validation to obtain the average predictive efficiency of the model. The diagnostic performance indicators such as accuracy, sensitivity, specificity, positive predictive value, and negative predictive value of the prediction model were calculated. The ROC curve was drawn and the AUC value was calculated to evaluate the discrimination of the model. By comparing the predicted value with the actual observed value, a calibration curve was drawn to evaluate the calibration degree of the prediction model. The above validation process was repeated in the validation set to evaluate the external efficacy of the model. Finally, a complete preliminary abnormal prediction data set including individual prediction probability, prediction accuracy index, Bootstrap validation results and cross-validation results is generated. This data set provides a reliable data basis for the subsequent Nomogram model construction.
[0024] Step S6: Based on the predicted data, a nomogram prediction model is constructed. The regression coefficient of each predictor in the multivariate logistic regression model is converted into a visual scoring scale of 0-100 points. In the nomogram, a corresponding score axis is set for each predictor, and the score size is proportional to the regression coefficient of the factor. A nomogram containing the scores of each predictor and its corresponding predicted probability is drawn, with the horizontal axis representing the score and the vertical axis representing each predictor and its value range. Different types of predictors (continuous variables, categorical variables) are standardized and converted into a unified scoring standard. A total score axis is set, and the individual scores of each predictor are accumulated to obtain a total score reflecting the overall risk level. The total score is converted into the predicted probability of lymph node metastasis through a preset conversion function. The model is internally validated, and the C-index (consistency index) is calculated to evaluate the discrimination of the model. A calibration curve is drawn to evaluate the prediction accuracy of the Nomogram model, with the horizontal axis representing the predicted probability and the vertical axis representing the actual probability. The final output is an initial abnormal index data set containing standardized scores, predicted probabilities, and C-index values.
[0025] Step S7: Risk stratify the initial abnormal index obtained. Based on ROC curve analysis, the Youden Index is calculated to determine the optimal critical value, and the patients are divided into high-risk group and low-risk group. The survival curve is drawn by the Kaplan-Meier method, and the prognostic differences of patients in different risk groups are compared by Log-rank test. The positive predictive value, negative predictive value, likelihood ratio and other diagnostic indicators of each risk level are calculated. The histogram and density curve of the risk prediction probability are drawn to show the distribution characteristics of different risk levels. Decision curve analysis (DCA) is performed to evaluate the clinical net benefit of the prediction model under different critical probabilities. A quantile-based risk stratification standard is established to divide patients into three groups: low risk (<25%), medium risk (25%-75%) and high risk (>75%). The stratification results are correlated with clinical pathological characteristics to verify the rationality of risk stratification. The final output includes complete abnormal evaluation information including risk level, prediction probability, survival analysis results, DCA curve and risk distribution characteristics.
[0026] The present invention provides a method for assessing the risk of lymph node metastasis in colon cancer. By comprehensively analyzing the patient's imaging data and pathological feature data and combining clinical standards for multi-level screening, it is possible to more comprehensively assess the abnormal information of lymph node metastasis, thereby improving the accuracy of abnormal assessment. By integrating multimodal abnormal data into ROC curve analysis and performing association analysis based on clinical indicators, a refined assessment of lymph node metastasis is achieved, which is helpful in formulating personalized treatment plans and avoiding overtreatment. Systematic analysis of abnormal screening indicators based on Logistic regression can ensure that the evaluation model has stable predictive efficiency in different clinical scenarios, reduce evaluation bias, and improve the overall accuracy of the model. By comprehensively analyzing multi-dimensional abnormal prediction data, using the Nomogram model for abnormal grading, and realizing accurate assessment of abnormalities through a curve grading strategy, the prediction effect is effectively improved and the misjudgment rate is reduced.
[0027] In one embodiment, imaging data and pathological characteristic data of the patient are obtained and standardized preprocessed to obtain index data, including: In the process of imaging data processing, the grayscale threshold range of the target area is set to 130-180, and the region growing algorithm is used to segment the imaging data, and the adjacent pixels that meet the threshold conditions are classified into the same area, thereby obtaining the target area image data. The deep convolutional neural network is used to extract the feature information of each dimension including texture features, edge features and shape features, and the feature vector matrix is constructed to form the initial image feature data.
[0028] Pathological feature data processing uses digital scanning under a 400x optical microscope with a scanning resolution of 0.25 μm / pixel to convert pathological sections into digital pathological images. The digital pathological images are divided into blocks of 256×256 pixels to generate multiple non-overlapping local area images to form a local area image set.
[0029] During the color space conversion process, the RGB color space is converted to the HSV color space. The conversion formula is the preset standard RGB to HSV conversion matrix to obtain the HSV space feature data. Based on the HSV space feature data, the morphological operators are used to extract the morphological features such as the area, perimeter, and roundness of the nucleus, and the pathological morphological feature set is constructed.
[0030] In the feature fusion stage, a weighted fusion strategy is used to fuse the initial image feature data with the pathological morphology feature set. The weight coefficient is determined by the cross-validation method to generate fused feature data. The fused feature data is normalized to the maximum and minimum values, and the normalization interval is set to [0,1] to obtain a standardized feature vector.
[0031] In the process of constructing the indicator data, a KD tree index structure is established based on the standardized feature vector. The index key value includes the values of each dimension of the feature vector. The corresponding sample identification and category information are stored in the index table, thereby obtaining complete indicator data. This indicator data will be used as input for the subsequent risk assessment model to assess and predict the risk of lymph node metastasis of colon cancer.
[0032] This embodiment uses a region growing algorithm combined with a preset grayscale threshold range to segment image data, thereby achieving accurate extraction of the target area and avoiding the subjectivity and uncertainty of traditional manual segmentation methods. The multi-dimensional feature extraction method based on a deep convolutional neural network ensures the comprehensiveness and representativeness of image features, providing a reliable data basis for subsequent risk assessment. By introducing HSV color space conversion and morphological feature extraction, the limitations of traditional RGB space feature extraction are broken, making the expression of pathological features richer and more accurate. The weighted fusion strategy is used to fuse image features and pathological features, making full use of the complementary advantages of the two types of data and improving the comprehensiveness of feature expression. The feature index structure design based on the KD tree optimizes the storage and retrieval efficiency of data and speeds up the processing speed of risk assessment. Through the maximum and minimum value normalization processing, the problem of inconsistent dimensions of different features is solved, and the accuracy and stability of the risk assessment model are improved. The design of the overall solution realizes the automation and standardization of the risk assessment of lymph node metastasis of colon cancer, providing an objective and reliable basis for clinical diagnosis.
[0033] In one embodiment, the indicator data is preliminarily screened for abnormalities based on clinical criteria to obtain a preliminary abnormal factor group, including: The collected indicator data are screened for abnormalities using preset clinical standards. In this method, the patient's clinical indicator data are divided into different hierarchical categories according to preset rules, such as age is divided into the youth group (18-44 years old), the middle-aged group (45-59 years old) and the elderly group (over 60 years old), and the tumor size is divided into T1 stage (≤2cm), T2 stage (2-5cm) and T3 stage (>5cm) and other multi-dimensional classification data.
[0034] Based on multidimensional clinical classification data, the correlation matrix between indicators was constructed, and the matrix data structure was formed by calculating the correlation coefficients between the indicators. The correlation matrix reflects the strength of the relationship between different clinical indicators. The correlation coefficient is calculated using the Pearson correlation coefficient method, and the value range is [-1,1]. The indicator correlation matrix formed was divided using a hierarchical clustering algorithm, and the clustering distance threshold was set to 0.6. The indicators with strong correlation were grouped together to obtain multiple indicator groups with intrinsic correlation.
[0035] The Kruskal-Wallis test was performed on each indicator group data to calculate the heterogeneity of each indicator within the group. The significance level α of the test was set to 0.05. When the p value was less than α, it indicated that there was a significant difference in the indicators within the group. The calculated inter-group difference coefficient was screened by the set threshold value, and the threshold value was 0.3. The indicators with a difference coefficient greater than the threshold were screened as the key indicator set.
[0036] For the selected key indicator set, calculate the abnormal ratio of each indicator. The calculation method of the abnormal ratio is the ratio of the number of abnormal samples of the indicator to the total number of samples, and the calculation result is standardized to the interval [0,1] to obtain the abnormal contribution data. The K-means clustering algorithm is used to group the abnormal contribution data, and the cluster number K value is set to 3. Indicators with similar abnormal contributions are divided into the same group to form a preliminary abnormal factor group.
[0037] This embodiment achieves a refined classification of clinical indicators of patients and improves the accuracy of data analysis by setting multi-dimensional clinical classification standards. The indicator association matrix and hierarchical clustering algorithm are used to effectively capture the intrinsic correlation characteristics between different clinical indicators, avoiding the limitations of the traditional single indicator evaluation method. The Kruskal-Wallis test is used to perform heterogeneity analysis on the indicator group, ensuring the scientific screening of key indicators in the evaluation process and reducing the interference of irrelevant indicators on the evaluation results. Through abnormal contribution calculation and group processing, a complete abnormal factor identification mechanism is established to improve the accuracy of risk assessment. This method organically combines a variety of statistical methods, while ensuring the reliability of the evaluation results, it also significantly improves the efficiency of risk assessment, providing clinicians with a more objective and scientific basis for decision-making.
[0038] In one embodiment, ROC curve analysis is performed on the preliminary abnormal factor group to obtain multimodal abnormal data, including: In the calculation of the true positive rate and false positive rate, each sample point in the preliminary abnormal factor group is classified and judged, and the judgment result is compared with the actual label. The true positive rate calculation formula is TPR=TP / (TP+FN), and the false positive rate calculation formula is FPR=FP / (FP+TN). The calculated (FPR, TPR) coordinate pair constitutes the initial evaluation coordinate set. The evaluation coordinates reflect the classification performance of the model under different decision thresholds.
[0039] The scattered point distribution sampling process adopts a stratified random sampling method, and dense sampling is performed on the basis of the initial evaluation coordinates. The sampling interval is [0,1]×[0,1], and the number of sampling points is set to 1000. The distribution characteristics of the sampling points are consistent with the original evaluation coordinates to ensure that the sampling results are representative. The predicted sampling points obtained by sampling are used for subsequent curve fitting.
[0040] The cubic spline interpolation operation uses natural boundary conditions to construct a cubic polynomial function between adjacent sampling points. The interpolation process satisfies the function value continuity, first-order derivative continuity, and second-order derivative continuity constraints. The interpolation parameters include the polynomial coefficient matrix and node position information, which are used to describe the local morphological characteristics of the curve.
[0041] The piecewise function construction is based on the abnormal interpolation parameters, dividing the entire ROC curve into multiple subintervals. The function expression in each subinterval is determined by the corresponding cubic polynomial. The selection of segmentation points follows the principle of maximum curvature change to ensure that the piecewise function can accurately describe the changing trend of the curve. All piecewise functions together constitute the ROC curve piecewise function set.
[0042] The curve integral operation uses the composite trapezoidal quadrature method to calculate the area of the region enclosed by the ROC curve and the coordinate axis. The integral interval is [0,1], and the integral step is set to 0.001. The integral result is the area under the initial curve data, which reflects the overall classification performance of the model.
[0043] The interval correction process introduces bootstrap resampling technology to generate 1000 resampled data sets. The area under the curve is calculated repeatedly for each resampled data set to obtain the area value distribution. By calculating the mean and standard deviation of the distribution, the initial area under the curve is corrected for deviation to obtain the corrected curve area value.
[0044] The confidence interval range is calculated using the normal distribution assumption, and the significance level α is set to 0.05. Based on the corrected curve area value and its standard deviation, the upper and lower limits of the (1-α) confidence interval are calculated. The confidence interval range is used to quantify the uncertainty of the prediction accuracy and constitute a prediction accuracy index.
[0045] The segmented threshold division adopts an equal spacing method to divide the predicted probability into three risk levels: high, medium, and low. The sensitivity and specificity are calculated separately within the threshold interval corresponding to each risk level. The sensitivity calculation formula is Se=TP / (TP+FN), and the specificity calculation formula is Sp=TN / (TN+FP). The calculation results constitute the graded evaluation data.
[0046] The abnormal quantitative scoring is based on the graded evaluation data and adopts the weighted summation method. The weight coefficient is determined by the expert scoring, and the scoring indicators include sensitivity, specificity, positive predictive value and negative predictive value. The scoring result is the multimodal abnormal data, which is used for subsequent risk classification and early warning.
[0047] This embodiment significantly improves the accuracy and reliability of the risk assessment of lymph node metastasis of colon cancer by performing ROC curve analysis on the preliminary abnormal factor group, integrating multiple modeling techniques such as true positive rate calculation, scattered distribution sampling and cubic spline curve interpolation. The piecewise function construction combined with the curve integral operation realizes the accurate quantification of the performance of the evaluation model and provides more stable decision support for clinical diagnosis. The introduction of interval correction and confidence interval calculation effectively reduces the uncertainty of the evaluation results and improves the robustness of the prediction model. The multi-level segmented threshold division and abnormal quantitative scoring mechanism make the risk assessment results more clinically practical, facilitate medical staff to quickly identify high-risk patients, and realize accurate treatment plan formulation. This method overcomes the limitations of strong subjectivity and single quantitative indicators in traditional evaluation methods, and significantly improves the scientificity and objectivity of the risk assessment of lymph node metastasis of colon cancer.
[0048] In one embodiment, multivariate logistic regression analysis is performed on the multimodal abnormal data to obtain an abnormal screening index set, including: The multimodal abnormality dataset contains indicator data in multiple dimensions, including age, gender, tumor size, tumor location, CEA level, CA199 level, gene mutation type, and CT image characteristics.
[0049] When performing multicollinearity test on multimodal abnormal data, the variance inflation factor (VIF) method was used to calculate the correlation between each abnormal factor. When the VIF value was greater than 10, multicollinearity was determined to exist. The VIF value of each abnormal factor was calculated to form an abnormal factor independence evaluation value matrix. After testing, the VIF value of age and gene mutation type was 2.3, the VIF value of tumor size and CT characteristics was 3.1, and the VIF value of CEA and CA199 levels was 8.5, all less than 10, indicating that there was no significant multicollinearity between the factors.
[0050] When performing threshold screening on the independence evaluation values of abnormal factors, the VIF threshold was set to 5, and factors with VIF values less than 5 were screened as independent predictors. After screening, six factors including age, gender, tumor size, tumor location, gene mutation type, and CT features were included in the set of independent predictors.
[0051] When calculating the multivariate logistic regression coefficient, the lymph node metastasis status was used as the dependent variable, and the factors in the independent predictor set were used as the independent variables. The regression coefficient of each factor was obtained by the maximum likelihood estimation method. The regression coefficient of age was calculated to be 0.023, the regression coefficient of tumor size was 0.456, and the regression coefficient of gene mutation was 0.378. These coefficients constituted the abnormal prediction weight coefficient set.
[0052] When optimizing the abnormal prediction weight coefficients iteratively, the cross-validation method was used to divide the data set into a training set and a validation set in a ratio of 7:3, and the weight coefficients were optimized through repeated iterative training. After 100 iterations of optimization, the final abnormal score parameter set was obtained: age 0.025, tumor size 0.468, and gene mutation 0.385.
[0053] In the prediction analysis and processing phase, the optimized abnormal scoring parameter set is applied to the validation set data to calculate the risk prediction score for each sample. When the prediction score is greater than or equal to 0.5, the sample is judged as positive (predicted to have lymph node metastasis), and the prediction score is less than 0.5. The sample is judged as negative (predicted to have no lymph node metastasis). The prediction results are compared with the actual lymph node metastasis in the validation set, and the number of cases with true positives (predicted to have metastasis and actually have metastasis), false positives (predicted to have metastasis but actually have no metastasis), true negatives (predicted to have no metastasis and actually have no metastasis), and false negatives (predicted to have no metastasis but actually have metastasis) is counted. Through ROC curve analysis, the prediction accuracy data was obtained: the sensitivity was 85.7%, the specificity was 82.3%, and the area under the curve (AUC) was 0.864.
[0054] When performing threshold evaluation analysis on the prediction accuracy data, various prediction performance indicators were calculated based on the above statistical data. After evaluation, the positive predictive value (the number of true positives divided by the total number of predicted positives) was 83.6%, the negative predictive value (the number of true negatives divided by the total number of predicted negatives) was 84.5%, and the overall accuracy (the number of correctly predicted cases divided by the total number of cases) was 84.0%, which constituted the evaluation reliability results.
[0055] According to the reliability evaluation results, the independent predictors were ranked in importance based on the regression coefficients of each predictor and their statistical significance. The ranking results showed that tumor size, gene mutation type, and age were the three indicators with the greatest predictive value, which constituted the final abnormal screening indicator set.
[0056] This embodiment collects multi-dimensional clinical indicator data and constructs a multimodal abnormal data set, combined with a multivariate logistic regression analysis method, to achieve a comprehensive assessment of the risk of lymph node metastasis in colon cancer. Through the variance inflation factor test and threshold screening, independent predictors are effectively identified, mutual interference between factors is avoided, and the accuracy of the evaluation results is improved. The prediction weight coefficients are adjusted by cross-validation and iterative optimization methods, so that the prediction performance of the model is significantly improved, and the prediction accuracy reaches 84.0%. By counting true positives, false positives and other indicators and performing ROC curve analysis, the prediction efficiency of the model is comprehensively evaluated, providing reliability guarantee for clinical application. The final screened abnormal indicator set highlights the importance of key predictive factors such as tumor size and gene mutation type, providing a simple and practical reference for clinicians to assess the risk of lymph node metastasis in patients.
[0057] In one embodiment, preliminary indicator prediction is performed on the abnormal screening indicator set to obtain preliminary abnormal prediction data, including: When optimizing the combination parameters, the grid search method is used to optimize the parameter combination of the selected indicator set. Specifically, the parameter search range is set and the grid points are divided. The performance indicators of the indicator combination are calculated at each grid point, including accuracy, specificity, and sensitivity. By comparing the performance indicators under different parameter combinations, the optimal parameter combination is selected to form an optimized indicator set. The optimized indicator set contains multi-dimensional data such as screened clinical indicators, imaging features, and laboratory test results.
[0058] The anomaly classification calculation uses a support vector machine (SVM) classifier to classify the optimized indicator set. The SVM classifier uses the radial basis kernel function to construct the optimal classification hyperplane in the feature space and divide the sample data into normal and abnormal groups. The classification result outputs the probability value of each sample belonging to the abnormal class to form a pre-classified abnormal value.
[0059] Cross-validation analysis uses the K-fold cross-validation method to randomly divide the data set into K mutually exclusive subsets. In each round of validation, one of the subsets is selected as the validation set, and the remaining K-1 subsets are used as the training set. This is repeated K times to obtain the validation data set. The K value is usually set to 5 or 10. In this embodiment, the K value is 5.
[0060] During the Bootstrap sampling calculation process, the validation data set is sampled with replacement, and the number of sampling is 100 times the original sample size. After each sampling, the random forest algorithm is used to model and predict the sampled data to obtain the predicted probability value. All sampling results are integrated to form a resampled prediction set.
[0061] The multi-algorithm fusion adopts the Stacking integrated learning framework, takes the prediction results of multiple base learners (including random forest, XGBoost, LightGBM, etc.) as feature input, uses logistic regression as a meta-learner for integration, and outputs the integrated probability value.
[0062] The abnormal stratification is based on the integrated probability value to set multiple risk thresholds. Those below 0.3 are classified as low-risk, those between 0.3-0.7 are classified as medium-risk, and those above 0.7 are classified as high-risk, forming multi-level abnormal data. The threshold setting is determined based on clinical practice experience and ROC curve analysis results.
[0063] Probability calibration uses the Platt calibration method to convert multi-level anomaly data into calibrated probability values. The relationship between the original prediction score and the true label is fitted by the logarithmic probability regression model to obtain the calibration parameters. The calibration parameters are used to map and transform the anomaly prediction probability and output the preliminary anomaly prediction data.
[0064] This embodiment optimizes parameters through a grid search method to ensure the optimal performance of the indicator combination and improve the prediction accuracy of the model. The SVM classifier is combined with cross-validation analysis to enhance the reliability and stability of the classification results. The application of Bootstrap sampling technology effectively overcomes the limitation of insufficient data, expands the sample size, and makes the prediction results more statistically significant. The multi-algorithm fusion strategy makes full use of the advantages of different algorithms and improves the overall prediction performance through the Stacking framework. The setting of risk stratification is based on clinical practice experience, making the prediction results more practical. The introduction of probability calibration ensures a good correspondence between the predicted probability and the actual risk, and improves the clinical application value of the evaluation results. This method significantly improves the accuracy and reliability of the risk assessment of lymph node metastasis in colon cancer, and provides strong support for clinical diagnosis and treatment decisions.
[0065] In one embodiment, the preliminary abnormal prediction data is input into a preset Nomogram model for abnormal analysis to obtain an initial abnormal index, including: In the data standardization layer, the input preliminary anomaly prediction data is standardized. For continuous variables, the Z-score standardization method is used to convert the values to a distribution with a mean of 0 and a standard deviation of 1; for categorical variables, one-hot encoding is used to convert them into numerical features. Standardization ensures that feature data of different dimensions are comparable.
[0066] In the variable scoring layer, the standardized characteristic data are divided into intervals based on the experience of clinical medical experts. Continuous variables are divided into cut-off points according to clinical grouping standards, and categorical variables maintain the original category division. Basic scores are assigned to each interval, ranging from 0 to 100 points, and the scoring standards are determined by referring to the statistical results of clinical research.
[0067] In the weight distribution layer, the LASSO regression method is used to analyze the contribution of each feature to the prediction result. The optimal regularization parameter is selected through cross-validation to obtain the feature importance ranking. The weight coefficient is calculated based on the feature importance, and the sum of the weight coefficients is 1, which reflects the relative importance of each feature in the prediction.
[0068] In the single-item scoring layer, the variable score is multiplied by the corresponding feature weight and summed to obtain the weighted score. The weighted score comprehensively considers the predictive ability and clinical importance of the feature, reflecting the differentiation of risk assessment.
[0069] In the total score calculation layer, the weighted scores of all features are summed to obtain the total score of the anomaly. The total score ranges from 0 to 1000, and the higher the score, the greater the risk.
[0070] In the probability conversion layer, the logistic function is used to map the total abnormal score to a probability value between 0 and 1. The probability mapping curve is fitted based on the training set data and calibrated by the Hosmer-Lemeshow test to ensure the accuracy of the predicted probability.
[0071] In the anomaly assessment layer, the optimal probability threshold is determined based on ROC curve analysis, which corresponds to the maximum point of the Youden index. The probability value is nonlinearly corrected by the spline function to obtain the corrected anomaly index. If the anomaly index is greater than the threshold, it is judged as high risk, and if it is less than the threshold, it is judged as low risk.
[0072] This embodiment uses a multi-level Nomogram model structure to accurately evaluate the risk of lymph node metastasis of colon cancer. The design of the data normalization layer and the variable scoring layer enables unified processing of different types of feature data, overcoming the evaluation bias caused by inconsistent feature dimensions in traditional methods. The weight allocation layer uses the LASSO regression method to determine the feature weights, avoiding the interference of subjective factors on the evaluation results and improving the objectivity of risk assessment. Through the design of the single scoring layer and the total score calculation layer, the clinical importance of the features is organically combined with the statistical significance, enhancing the interpretability of the evaluation results. The probability conversion layer introduces the Logistic function mapping and calibration mechanism to solve the problem of the nonlinear relationship between the score and the actual risk. The abnormal assessment layer optimizes the threshold selection based on the ROC curve to improve the accuracy of risk classification. The overall model design fully considers the needs of clinical practice, and has good clinical application value while ensuring the accuracy of the evaluation.
[0073] In one embodiment, the initial abnormality index is graded by curves and abnormality evaluation is performed to obtain final abnormality evaluation information, including: In the probability interval division stage, the initial anomaly index is divided into 100 small intervals evenly using the equal distance interval division method, and the anomaly index in the range of 0-1 is evenly divided into 100 small intervals, each with a length of 0.01. By counting the number of sample points in each interval, the transfer anomaly probability of each interval is calculated to form a transfer anomaly probability segment data set.
[0074] The density function solution process uses the kernel density estimation method, selects the Gaussian kernel function as the benchmark kernel function, and sets the bandwidth parameter to 0.05. The center point of each probability interval is used as the sampling point, and the kernel density estimation calculation is performed in combination with the interval probability value to generate a continuous anomaly density distribution curve. This curve reflects the probability density distribution characteristics corresponding to different anomaly index values.
[0075] The trapezoidal integration method is used in the numerical integration operation stage to perform numerical integration on the abnormal density distribution curve in the interval [0,1]. The integration step is set to 0.001. By accumulating the integral values of each sub-interval, the cumulative distribution function value of the abnormal index from 0 to any point x is obtained to form a cumulative abnormal distribution data sequence.
[0076] The critical point extraction process is based on the rate of change of the cumulative distribution function. The change rate threshold is set to 0.1, and when the first-order derivative of the cumulative distribution function is greater than the threshold, it is determined to be a critical point. The anomaly index values of all points that meet the conditions are extracted and sorted to obtain the anomaly threshold sequence. This sequence contains the key nodes where the degree of anomaly changes significantly.
[0077] The interval segmentation process defines the area between two adjacent critical points in the abnormal threshold sequence as an abnormal level interval. Combined with clinical experience, the abnormal level is divided into five levels: very low risk (0-0.2), low risk (0.2-0.4), medium risk (0.4-0.6), high risk (0.6-0.8) and very high risk (0.8-1.0). Based on these interval nodes, the abnormal level determination point set is constructed.
[0078] In the hierarchical mapping stage, the initial anomaly index is compared with the anomaly level judgment point set. By determining the interval range that the initial anomaly index falls into, it is mapped to the corresponding risk level, thereby obtaining the initial anomaly judgment level result.
[0079] The abnormality verification process introduces a clinical validation data set, which contains pathologically confirmed lymph node metastasis status information. The consistency indicators between the initial abnormality level and the actual metastasis situation are calculated, including accuracy, sensitivity, and specificity. The validation standard is set to an accuracy of not less than 85%, and a sensitivity and specificity of not less than 80%.
[0080] Generate a standardized assessment description based on the abnormal verification results. The assessment information includes a qualitative description of the risk level, quantitative indicator values, and credibility analysis. The qualitative description uses a unified semantic template and combines specific numerical indicators to form a complete abnormal assessment report. This report provides a reference for risk assessment for clinicians.
[0081] This embodiment introduces probability interval division and kernel density estimation methods to achieve accurate quantification of the abnormal index, making the risk assessment process of colon cancer lymph node metastasis more objective and accurate. The combined application of density function solution and numerical integration operation effectively captures the abnormal distribution characteristics and enhances the reliability of the assessment results. Based on the hierarchical mapping mechanism of critical point extraction, a scientific risk level classification system is established to improve the clinical practicality of the assessment results. The introduction of multi-dimensional evaluation indicators in the abnormal verification link ensures the accuracy and credibility of the assessment conclusions.
[0082] The present invention also provides a colon cancer lymph node metastasis risk assessment system, which is applied to any of the above colon cancer lymph node metastasis risk assessment methods, comprising: The acquisition module is used to obtain the patient's imaging data and pathological characteristic data, and perform standardized preprocessing to obtain indicator data; An analysis module is used to perform preliminary abnormal screening of indicator data based on clinical standards to obtain a preliminary abnormal factor group; The association module is used to perform ROC curve analysis on the preliminary abnormal factor group to obtain multi-modal abnormal data; A processing module is used to perform multivariate Logistic regression analysis and screening on multimodal abnormal data to obtain an abnormal screening indicator set; A control module, the control module is used to perform preliminary indicator prediction on the abnormal screening indicator set to obtain preliminary abnormal prediction data; An execution module, the execution module is used to input the preliminary abnormal prediction data into a preset Nomogram model for abnormal analysis to obtain an initial abnormal index; The generation module is used to curve grade the initial anomaly index and perform anomaly evaluation to obtain the final anomaly evaluation information.
[0083] The present invention provides a colon cancer lymph node metastasis risk assessment system, which can more comprehensively assess the abnormal information of lymph node metastasis by comprehensively analyzing the patient's imaging data and pathological feature data, and combining clinical standards for multi-level screening, thereby improving the accuracy of abnormal assessment. By integrating multimodal abnormal data into ROC curve analysis and performing association analysis based on clinical indicators, a refined assessment of lymph node metastasis is achieved, which helps to formulate personalized treatment plans and avoid overtreatment. Systematic analysis of abnormal screening indicators based on Logistic regression can ensure that the evaluation model has stable predictive efficiency in different clinical scenarios, reduce evaluation bias, and improve the overall accuracy of the model. By comprehensively analyzing multi-dimensional abnormal prediction data, using the Nomogram model for abnormal grading, and realizing accurate assessment of abnormalities through a curve grading strategy, the prediction effect is effectively improved and the misjudgment rate is reduced.
[0084] It should be noted that technicians in the relevant technical field can clearly understand that for the convenience and conciseness of description, the specific working process of the system and each module described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0085] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for assessing the risk of lymph node metastasis in colon cancer, characterized in that: include: Obtain the patient's imaging data and pathological characteristic data, and perform standardized preprocessing to obtain indicator data; Performing preliminary abnormal screening on the indicator data based on clinical standards to obtain a preliminary abnormal factor group; Performing ROC curve analysis on the preliminary abnormal factor group to obtain multimodal abnormal data; Performing multivariate logistic regression analysis and screening on the multimodal abnormal data to obtain an abnormal screening indicator set; Performing preliminary indicator prediction on the abnormal screening indicator set to obtain preliminary abnormal prediction data; Inputting the preliminary abnormal prediction data into a preset Nomogram model for abnormal analysis to obtain an initial abnormal index; The initial abnormality index is graded by curves, and an abnormality assessment is performed to obtain final abnormality assessment information.
2. The method for assessing the risk of lymph node metastasis of colon cancer according to claim 1, characterized in that: The imaging data and pathological characteristic data of the patient are obtained and standardized preprocessed to obtain index data, including: Performing region segmentation processing on the imaging data to obtain target region image data; Extracting feature vectors from the target area image data to obtain initial image feature data; Digitally scanning the pathological characteristic data to obtain a digital pathological image; Performing image block processing according to the digital pathological image to obtain a local area image set; Performing color space conversion processing on the local area image set to obtain HSV space feature data; Perform morphological feature extraction according to the HSV spatial feature data to obtain a pathological morphological feature set; Performing feature fusion on the initial image feature data and the pathological morphology feature set to obtain fused feature data; Performing normalization processing on the fused feature data to obtain a standardized feature vector; A feature index table is established according to the standardized feature vector to obtain the indicator data.
3. The method for assessing the risk of lymph node metastasis of colon cancer according to claim 1, characterized in that: The index data is preliminarily screened for abnormalities based on clinical criteria to obtain a preliminary abnormal factor group, including: Classifying the index data into clinical manifestation levels to obtain multi-dimensional clinical classification data; Constructing a correlation matrix of each dimension for the multidimensional clinical classification data to obtain an indicator correlation matrix; Performing hierarchical clustering on the indicator association matrix to obtain indicator group data; The intra-group heterogeneity of the index group data was calculated according to the preset Kruskal-Wallis test algorithm to obtain the inter-group difference coefficient; Performing threshold screening on the inter-group difference coefficient to obtain a key indicator set; Calculate the single factor abnormality ratio according to the key indicator set to obtain abnormal contribution data; The abnormal contribution data is subjected to abnormal grouping processing to obtain the preliminary abnormal factor group.
4. The method for assessing the risk of lymph node metastasis of colon cancer according to claim 1, characterized in that: The performing ROC curve analysis on the preliminary abnormal factor group to obtain multimodal abnormal data includes: Calculating the true positive rate and the false positive rate of the preliminary abnormal factor group to obtain initial evaluation coordinates; Perform scattered point distribution sampling according to the initial evaluation coordinates to obtain predicted sampling points; Performing a cubic spline curve interpolation operation on the predicted sampling points to obtain abnormal interpolation parameters; Constructing a piecewise function for the abnormal interpolation parameter to obtain a piecewise function set of ROC curves; Performing curve integration operation according to the ROC curve piecewise function set to obtain initial area under the curve data; Performing interval correction on the initial area under the curve data to obtain a corrected curve area value; Calculating the confidence interval range according to the corrected curve area value to obtain a prediction accuracy index; The prediction accuracy index is divided into segmented thresholds, and the sensitivity and specificity of each segment are calculated to obtain graded evaluation data; Abnormal quantitative scoring is performed according to the graded evaluation data to obtain multimodal abnormal data.
5. The method for assessing the risk of lymph node metastasis of colon cancer according to claim 1, characterized in that: The multimodal abnormal data is subjected to multivariate logistic regression analysis and screening to obtain an abnormal screening index set, including: Performing a multicollinearity test on the multimodal abnormal data to obtain an independence evaluation value of the abnormal factors; Performing threshold screening on the independence evaluation values of the abnormal factors to obtain an independent prediction factor set; Performing multivariate logistic regression coefficient calculation on the independent prediction factor set to obtain an abnormal prediction weight coefficient; Optimizing and iterating the anomaly prediction weight coefficient to obtain an anomaly scoring parameter set; Performing prediction analysis and processing according to the abnormal scoring parameter set to obtain prediction accuracy data; Performing threshold evaluation analysis on the prediction accuracy data to obtain an evaluation reliability result; The independent prediction factor set is sorted by importance according to the reliability evaluation result to obtain the final abnormal screening indicator set.
6. The method for assessing the risk of lymph node metastasis of colon cancer according to claim 1, characterized in that: The performing preliminary indicator prediction on the abnormal screening indicator set to obtain preliminary abnormal prediction data includes: Performing combined parameter optimization on the abnormal screening index set to obtain an optimized index set; Perform anomaly classification calculation according to the optimized indicator set to obtain pre-classified anomaly values; Performing cross-validation analysis on the pre-classified outliers to obtain a validation data set; Perform Bootstrap sampling calculation on the verification data set to obtain a resampled prediction set; Performing multi-algorithm fusion on the resampled prediction set to obtain an integrated probability value; Perform abnormal hierarchical division according to the integrated probability value to obtain multi-level abnormal data; Probability calibration is performed according to the multi-level abnormal data to obtain the preliminary abnormal prediction data.
7. The method for assessing the risk of lymph node metastasis of colon cancer according to claim 1, characterized in that: The initial abnormality prediction data is input into a preset Nomogram model for abnormality analysis to obtain an initial abnormality index, including: Inputting the preliminary abnormality prediction data into the Nomogram model, performing standardized input processing on the preliminary abnormality prediction data through the data standardization layer of the Nomogram model to obtain standardized feature data; Performing interval division and basic score calculation on the standardized feature data through the variable scoring layer of the Nomogram model to obtain variable scoring data; Performing prediction contribution analysis and weight coefficient generation on the standardized feature data through the weight allocation layer of the Nomogram model to obtain feature weight data; The variable scoring data and the feature weight data are weighted and fused through the single scoring layer of the Nomogram model to obtain weighted scoring data; The weighted score data is scored and accumulated by the total score calculation layer of the Nomogram model to obtain abnormal total score data; Performing Logistic function probability mapping and abnormality calibration on the abnormal total score data through the probability conversion layer of the Nomogram model to obtain an initial abnormality probability; The initial abnormal probability is subjected to nonlinear correction and threshold analysis through the abnormal evaluation layer of the Nomogram model to obtain and output the initial abnormal index.
8. The method for assessing the risk of lymph node metastasis of colon cancer according to claim 1, characterized in that: The initial abnormality index is subjected to curve grading and abnormality evaluation to obtain final abnormality evaluation information, including: Performing probability interval division calculation on the initial abnormality index to obtain a transfer abnormality probability segment; Solving the density function according to the transfer abnormality probability section to obtain an abnormality density distribution curve; Performing numerical integration operation on the abnormal density distribution curve to obtain cumulative abnormal distribution data; Extract critical points according to the accumulated abnormal distribution data to obtain an abnormal threshold sequence; Performing interval segmentation processing on the abnormal threshold sequence to obtain an abnormal level determination point set; According to the abnormality level determination point set, the initial abnormality index is graded and mapped to obtain an initial abnormality level; Verifying the abnormal preliminary judgment level to obtain an abnormal verification result; A semantic description is generated according to the abnormality verification result to obtain the abnormality assessment information.
9. A colon cancer lymph node metastasis risk assessment system, characterized in that: The method for assessing the risk of lymph node metastasis of colon cancer as described in any one of claims 1 to 8 above comprises: An acquisition module, which is used to obtain the patient's imaging data and pathological characteristic data, and perform standardized preprocessing to obtain indicator data; An analysis module, the analysis module is used to perform preliminary abnormal screening on the indicator data based on clinical standards to obtain a preliminary abnormal factor group; A correlation module, wherein the correlation module is used to perform ROC curve analysis on the preliminary abnormal factor group to obtain multimodal abnormal data; A processing module, wherein the processing module is used to perform multivariate logistic regression analysis and screening on the multimodal abnormal data to obtain an abnormal screening indicator set; A control module, the control module is used to perform preliminary indicator prediction on the abnormal screening indicator set to obtain preliminary abnormal prediction data; An execution module, wherein the execution module is used to input the preliminary abnormality prediction data into a preset Nomogram model for abnormality analysis to obtain an initial abnormality index; A generation module is used to perform curve grading on the initial abnormality index and perform abnormality evaluation to obtain final abnormality evaluation information.
Citation Information
Cited By
Clinical test data intelligent analysis method and system
CN120299591A
Individual dynamic metabolic abnormality early warning system based on function type data analysis
CN121545750A