Creatinine and CAR combination-based cancer malignant fluid prognosis composite index CCAR construction method

By constructing a CCAR index based on the creatinine-to-C-reactive protein ratio and using a random forest machine learning model for nonlinear mapping, the limitations of existing technologies in assessing the metabolic and immune status of cancer cachexia patients are overcome. This enables more accurate survival prediction and risk stratification, is applicable to various cancer types, and supports personalized treatment.

CN120809194AActive Publication Date: 2025-10-17BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV

Patent Information

Application Number
CN202510848966.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-17
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing inflammation-related prognostic indicators have significant limitations in their applicability across different cancer types and patient populations, making it difficult to simultaneously assess the metabolic, immune, and nutritional status of patients with cancer cachexia.

Method used

A composite prognostic index for cancer cachexia, CCAR, based on the creatinine-to-C-reactive protein ratio (CAR), was constructed. A random forest machine learning model was used for nonlinear mapping, and a network calculator system was combined to perform risk stratification and survival prediction.

Benefits of technology

It provides more accurate survival predictions for cancer cachexia patients, is applicable to multiple cancer types, improves predictive efficacy and clinical applicability, has high reliability and visualization effects, and supports individualized treatment decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809194A_ABST
    Figure CN120809194A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical systems, and particularly discloses a method for constructing a cancer malignant fluid prognosis composite index CCAR based on a creatinine and CAR combination, and the method comprises the following steps: S1, obtaining serum biomarker data of a cancer malignant fluid patient, including creatinine, C-reactive protein and albumin; s2, calculating the ratio CAR of the C reactive protein to the albumin, wherein CAR = CRP / Albumin; s3, taking the CAR and the Cr as input features, and constructing a nonlinear composite index CCAR through a random forest machine learning model; s4, developing a network calculator system; the CCAR index constructed by the invention can comprehensively reflect the inflammation burden, the metabolic state and the immune level of the cancer and malignant fluid patients, and compared with the traditional index, the CCAR index can provide more accurate survival prognosis prediction; creatinine, the ratio of C-reactive protein to albumin and other parameters contained in the CCAR index are all clinical common blood examination items, the obtaining approach is simple and convenient, and wide clinical application is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of medical systems, in particular to a construction method of a cancer cachexia prognosis composite index CCAR based on a combination of creatinine and CAR. BACKGROUND

[0002] Cancer, as a major public health problem with continuously rising incidence and mortality worldwide, its progression and prognosis prediction of patients has always been the core direction of medical research. Studies have shown that systemic inflammation is a key factor in the occurrence and development of cancer, and inflammatory markers such as CRP, albumin, and creatinine have a certain predictive value for cancer prognosis. However, the applicability of traditional inflammation-related prognostic indicators (such as NLR, PLR, CAR, etc.) in different cancer types and patient populations has significant limitations.

[0003] Cancer cachexia, as a common complication of cancer, seriously affects the prognosis of patients. Existing inflammation-related indicators focus on reflecting the inflammatory state of patients, but for cancer cachexia patients, how to simultaneously assess their metabolic, immune, and nutritional status remains a difficult problem to be solved. Existing research has confirmed that inflammation burden, metabolic abnormalities, and immune status are closely related to the occurrence and progression of cancer cachexia, so developing a prognostic indicator that can comprehensively reflect these factors has important clinical significance. SUMMARY

[0004] In view of the deficiencies of the prior art, the application provides a construction method of a cancer cachexia prognosis composite index CCAR based on a combination of creatinine and CAR, which solves the problems in the background art.

[0005] To achieve the above purpose, the following technical scheme is adopted: a construction method of a cancer cachexia prognosis composite index CCAR based on a combination of creatinine and CAR, comprising the following steps: S1, obtaining serum biomarker data of cancer cachexia patients, including creatinine (Cr), C-reactive protein (CRP), and albumin (Albumin); S2, calculating the C-reactive protein to albumin ratio CAR=CRP / Albumin; S3, taking CAR and creatinine (Cr) as input features, and constructing a nonlinear composite index CCAR through a random forest machine learning model; S4, developing a network calculator system: The front-end interface receives the Cr, CRP, and Albumin data input by the user; The back-end server calculates the CAR value and standardizes the CAR and Cr; The standardized data is input into the pre-trained random forest model to generate the CCAR value; The output CCAR value, risk stratification result, low risk: CCAR≤0.56, high risk: CCAR>0.56, and survival rate prediction, 1 year / 3 years / 5 years.

[0006] Preferably, the risk stratification threshold CCAR=0.56 is determined by the maximum selected rank statistic in survival analysis, used to distinguish the prognosis risk level of cancer patients with malignant liquid degeneration; The network calculator front end integrates a dynamic heat map visualization module to map the negative correlation between the CCAR value and the survival probability, where an increase in the CCAR value corresponds to a decrease in the survival probability; The heat map adopts a bivariate gradient coloring scheme, with the horizontal axis representing the CCAR value and the vertical axis representing the predicted survival time, and the color band gradually changes from green for high survival probability to red for low survival probability.

[0007] Preferably, the calculation formula of the CAR ratio is: ; Where the correction coefficient =0.1 is used to balance and dimensional differences, so that value is in a reasonable calculation interval.

[0008] Preferably, the unit of the is milligrams per liter, representing the concentration of C-reactive protein in serum; the unit of the is grams per deciliter, representing the concentration of albumin in serum, both of which are obtained through clinical blood tests.

[0009] Preferably, the CCAR index realizes the following nonlinear mapping relationship through a random forest model: ; Where is a prediction function constructed by the random forest algorithm, which realizes nonlinear feature combination by integrating multiple decision trees; and are the mean and standard deviation of the creatinine in the training set, respectively, and are the mean and standard deviation of the CAR value in the training set, respectively, used for standardization processing of input features; The standardized features satisfy the normal distribution with a mean of 0 and a standard deviation of 1, ensuring the stability and generalization ability of the random forest model training; The prognosis prediction performance of the CCAR index in the validation queue needs to meet: AUC≥0.75 (AUC of traditional indicators CAR / NLR / PLR increased by >0.10); KM survival curve log-rank test p<0.001 between high-risk group and low-risk group; The CCAR indicator is applicable to the following cancer cachexia types: pancreatic cancer, gastric cancer, colorectal cancer, lung cancer, and is verified by cohort in each cancer type (KM curve stratification significance <0.05).

[0010] Preferably, the CCAR value is combined with the TNM stage to generate a composite prognostic model, and the risk score formula is: ; Where the weight coefficient =0.7, =0.3 is determined by fitting the Cox proportional hazards model to maximize the prediction accuracy of the model for survival risk.

[0011] Preferably, the standardization in S3 uses the StandardScaler algorithm to make the input data satisfy the distribution with mean 0 and standard deviation 1; S3 also includes the following steps before constructing the nonlinear composite indicator CCAR through the random forest machine learning model: Rank the feature importance of cancer cachexia patients using LASSO regression, specifically: collect multi-dimensional clinical feature data of cancer cachexia patients including creatinine, hemoglobin, and platelets; Use the LASSO regression model to filter each feature and calculate the coefficient of each feature; Determine creatinine as the key prediction feature based on the feature coefficient; The training steps of the random forest model include: Divide the collected cancer cachexia patient data into a training set and a validation set; Standardize CAR and creatinine in the training set; Set the parameters of the random forest model, including the number of trees and the maximum depth; Train the model using the training set; Evaluate the trained model through the validation set, adjust the parameters to optimize the model performance; Save the optimized random forest model and standardization parameters.

[0012] Preferably, the backend server processing flow in S4 includes the following steps: Receive patient creatinine, C-reactive protein, and albumin data sent by the front end through POST requests; The input data is verified for validity, and if the data is outside the normal range, a prompt is given to re-enter; according to The CAR value is calculated according to the formula of 0.1; The mean and standard deviation obtained in the training stage are used to standardize the CAR and creatinine; The pre-trained random forest model is called to generate the CCAR value; According to the size relationship between the CCAR value and 0.56, the risk is stratified, and the survival rates of 1 year, 3 years and 5 years are calculated; The calculation results are returned to the front end in JSON format.

[0013] Preferably, the implementation of the dynamic heat map visualization module includes the following steps: Collect the CCAR values and corresponding survival data of historical cancer malnutrition patients; Construct a two-dimensional matrix of CCAR values and survival time, and calculate the average survival probability of different CCAR value intervals at each survival time point; Determine the horizontal axis of the heat map as the CCAR value, the vertical axis as the survival time, and the color band from green to red, corresponding to high survival probability to low survival probability; Based on the CCAR value of the current patient, highlight the corresponding survival probability area in the heat map; Show the heat map on the front-end interface to intuitively present the relationship between CCAR value and survival probability.

[0014] Preferably, the risk stratification threshold of CCAR=0.56 is determined as follows: Sort the patients in the training cohort by CCAR value from small to large; Iterate through all possible split points and calculate the statistic, formula: ; Where, is the actual number of deaths in the low-risk group, is the expected number of deaths in the low-risk group, is the actual number of deaths in the high-risk group, is the expected number of deaths in the high-risk group, and the expected number of deaths is calculated based on the overall cohort survival curve; Select the split point that maximizes statistic as the initial threshold; Verify the stability of the initial threshold by the method, and finally determine the risk stratification threshold as 0.56; The Maxstat statistic is used to evaluate the discrimination ability of the segmentation point on the survival risk, and the larger the statistic value indicates that the survival difference between the high-risk group and the low-risk group under the segmentation point is more significant.

[0015] The application provides a construction method of a cancer cachexia prognosis composite indicator CCAR based on a combination of creatinine and CAR. 1. Outstanding comprehensive evaluation ability: the CCAR indicator constructed by the application can comprehensively reflect the inflammatory burden, metabolic state and immune level of a cancer cachexia patient, and can provide more accurate survival prognosis prediction compared with traditional indicators.

[0016] 2. Strong clinical applicability: the CCAR indicator contains parameters such as creatinine, C-reactive protein and albumin ratio (CAR), which are common clinical blood test items, and the acquisition method is simple and convenient, which is convenient for clinical wide application.

[0017] 3. High reliability of the indicator: the selected inflammatory and metabolic parameters (creatinine, C-reactive protein and albumin) are clinically recognized markers that can effectively reflect the patient's systemic inflammatory burden and metabolic level, and have high reliability and universality.

[0018] 4. Accurate risk stratification: the CCAR model can more accurately predict the survival prognosis of a cancer cachexia patient, and realize risk stratification through the indicator, helping to formulate individualized treatment decisions in clinical practice.

[0019] 5. Convenient and efficient clinical application: the network calculator developed based on CCAR makes the model more convenient and efficient in clinical application, and can quickly provide individualized prognosis evaluation for the patient for the doctor, and provide strong support for treatment decision.

[0020] 6. Excellent prediction performance: the CCAR indicator has significant prognosis prediction performance in the verification queue, AUC≥0.75, and the AUC of the traditional indicator CAR / NLR / PLR is improved by more than 0.10, and the KM survival curve log-rank test p<0.001 of the high-risk group and the low-risk group.

[0021] 7. Wide application range: the CCAR indicator is applicable to various cancer cachexia types such as pancreatic cancer, gastric cancer, colorectal cancer and lung cancer, and has been verified in each type, and the KM curve stratification significance p<0.05.

[0022] 8. Good visualization effect: the dynamic heat map visualization module integrated in the front end of the network calculator can visually present the negative correlation between the CCAR value and the survival probability through color coding, and improve the clinical interpretability.

[0023] 9. Model optimization: LASSO regression is used to screen key features, combined with a random forest model to build CCAR, ensuring the accuracy and stability of the model, and the risk stratification threshold determination method is scientific. The reliability of the threshold is guaranteed by Maxstat statistics and Bootstrap method.

[0024] 10. More accurate composite model: The composite prognostic model generated by combining CCAR value with TNM staging determines the weight coefficient through Cox proportional hazards model fitting, further improving the accuracy of survival risk prediction. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 is a flowchart of a method for constructing a cancer cachexia prognostic composite indicator CCAR based on the combination of creatinine and CAR; Figure 2 is a graph of ranking the importance of cancer cachexia patient features by LASO regression; Figure 3 is an AUC curve graph comparing CCAR with other classic indicators in different cohorts; Figure 4 is a Kaplan-meier survival curve graph of CCAR in different cohorts; Figure 5 is a Kaplan-meier survival curve graph of CCAR in various tumors; Figure 6 is a flowchart of constructing the CCAR index and developing a network calculator; Figure 7 is the interface and heat map visualization of the CCAR calculator for survival prediction of cancer cachexia. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.

[0027] As shown in Figures 1-7 , the present application provides a technical solution: a method for constructing a cancer cachexia prognostic composite indicator CCAR based on the combination of creatinine and CAR, comprising the following steps: S1, obtaining serum biomarker data of cancer cachexia patients, including creatinine (Cr), C-reactive protein (CRP) and albumin (Albumin); S2, calculate the C-reactive protein to albumin ratio CAR = CRP / Albumin; S3, taking CAR and creatinine (Cr) as input features, constructing a nonlinear composite indicator CCAR through a random forest machine learning model; S4, developing a network calculator system: The front-end interface receives user input of Cr, CRP, and Albumin data; the back-end server calculates the CAR value and standardizes CAR and Cr; the standardized data is input into the pre-trained random forest model to generate the CCAR value; the CCAR value, risk stratification results (low risk: CCAR≤0.56, high risk: CCAR>0.56), and survival rate prediction (1 year / 3 years / 5 years) are output.

[0028] More specifically, the risk stratification threshold CCAR=0.56 is determined by the largest selected rank statistic in survival analysis, which is used to distinguish the prognosis risk level of cancer cachexia patients; the network calculator front-end integrates a dynamic heat map visualization module to map the negative correlation between CCAR value and survival probability with color coding, where an increase in CCAR value corresponds to a decrease in survival probability; the heat map uses a bivariate gradient coloring scheme, with the horizontal axis representing CCAR value and the vertical axis representing predicted survival time, and the color band gradually changes from green for high survival probability to red for low survival probability.

[0029] The implementation of the dynamic heat map visualization module includes the following steps: collecting CCAR values and corresponding survival data of historical cancer cachexia patients; constructing a two-dimensional matrix of CCAR values and survival time, calculating the average survival probability of different CCAR value intervals at each survival time point; determining the horizontal axis of the heat map as CCAR value and the vertical axis as survival time, with the color band gradually changing from green to red, corresponding to high survival probability to low survival probability; based on the CCAR value of the current patient, highlighting its corresponding survival probability area in the heat map; displaying the heat map on the front-end interface to intuitively present the relationship between CCAR value and survival probability.

[0030] The risk stratification threshold CCAR=0.56 is determined as follows: sort the patients in the training cohort by CCAR value from small to large; traverse all possible split points and calculate the statistic for each split point, with the formula: ; wherein is the actual number of deaths in the low-risk group, is the expected number of deaths in the low-risk group, is the actual number of deaths in the high-risk group, is the expected number of deaths in the high-risk group, which is calculated based on the overall cohort survival curve; select the split point that maximizes the statistic as the initial threshold; by The stability of the initial threshold value is verified, and the risk stratification threshold value is finally determined to be 0.56; The statistical quantity is used to evaluate the discrimination ability of the split point on the survival risk, and the larger the statistical quantity value indicates that the survival difference between the high-risk group and the low-risk group under the split point is more significant.

[0031] In this embodiment, the Bootstrap method is used to verify the stability of the initial threshold value, and the reliability of the threshold value is evaluated by multiple resampling. The risk stratification threshold value is finally determined to be CCAR=0.56, which is used to distinguish the prognosis risk level of cancer cachexia patients. Wherein, CCAR≤0.56 is determined as low risk, and CCAR>0.56 is determined as high risk.

[0032] The implementation of the dynamic heat map visualization module is as follows: Data collection and preprocessing: Collect the CCAR values and corresponding survival data of historical cancer cachexia patients, including the survival status at different time points, clean and standardize the data to ensure the consistency and availability of the data.

[0033] Two-dimensional matrix construction: A two-dimensional matrix is constructed with CCAR values and survival time as dimensions. The CCAR values are divided into multiple intervals, and the average survival probability of each CCAR value interval at different survival time points (such as 1 year, 3 years, 5 years) is calculated to form a probability distribution matrix.

[0034] Heat map parameter definition: The horizontal axis is defined as the CCAR value, covering the complete value range from low to high; the vertical axis is defined as the predicted survival time, including key time nodes such as 1 year, 3 years, and 5 years; the color band scheme adopts a two-variable gradient coloring scheme, with the color band gradually changing from green to red. Green corresponds to a high survival probability area, and red corresponds to a low survival probability area. The color depth is positively correlated with the survival probability; patient data highlighting is based on the current input patient CCAR value, which automatically locates and highlights the corresponding survival probability area in the heat map. It is highlighted through color deepening or border marking, etc., which is convenient for clinicians to quickly identify; the front-end display and interaction integrates the constructed heat map into the network calculator front-end interface, and uses HTML5, CSS3 and JavaScript technology to realize dynamic rendering, supports mouse hovering to view specific probability values, etc. Interactive functions, intuitively present the negative correlation between CCAR value and survival probability, that is, the trend of decreasing survival probability with increasing CCAR value.

[0035] More specifically, the calculation formula of the CAR ratio is: ; Where the correction coefficient =0.1 is used to balance the dimensional difference between and , so that The value is in a reasonable calculation interval.

[0036] The unit is milligrams per liter, representing the concentration of C-reactive protein in serum; The unit is grams per deciliter, representing the concentration of albumin in serum, both of which are obtained through clinical blood tests.

[0037] Biomarker data acquisition: The concentration of C-reactive protein in patient serum is obtained through clinical blood tests, and the detection method uses immunoturbidimetry or enzyme-linked immunosorbent assay (ELISA). The detection result is represented by milligrams per liter (mg / L). The specific operation steps are as follows: Collect 5mL of fasting venous blood from the patient and place it in an anticoagulant tube; centrifuge at 3000rpm for 10 minutes to separate the serum; use a fully automatic biochemical analyzer to determine the CRP concentration based on the immunoturbidimetry principle, and the detection linear range is 0.5-200mg / L.

[0038] Albumin concentration detection: The same blood sample is collected synchronously, and the bromocresol green (BCG) method is used to detect the serum albumin concentration, and the result is represented by grams per deciliter (g / dL). The specific implementation steps include: Take 100μL of serum, add 500μL of BCG reagent, mix well; incubate at 37°C for 5 minutes, and measure the absorbance at 630nm wavelength; calculate the albumin concentration through the standard curve, and the detection range is 20-60g / L (i.e. 2.0-6.0g / dL).

[0039] CAR ratio calculation process: dimensional balance processing, since the unit of CRP is mg / L and the unit of Albumin is g / dL, there is a difference in the dimension of the two Direct division will result in a large numerical value of the calculation result. By introducing a correction coefficient k=0.1 for dimensional standardization, the specific formula is: ; This coefficient converts the CAR value into a dimensionless index, with a value range controlled between 0.01-10.0, which meets the numerical readability requirements of clinical indicators.

[0040] Take a patient's test data as an example: CRP=71.4mg / L; Albumin=32.6g / dL; Substitute the formula: .

[0041] Computational logic verification: Through 1000 historical sample tests, the rationality of the correction coefficient k = 0.1 is verified. When not corrected (k = 1), the CAR value range is 0.05-25.0, and 90% of the samples are concentrated in 1.0-5.0, with a large numerical span. After using k = 0.1, the CAR value range is narrowed to 0.005-2.5, and 90% of the samples are concentrated in 0.1-0.5, which is more in line with the input feature distribution requirements of the machine learning model.

[0042] Clinical application scenarios: For cancer cachexia patients admitted to the emergency department, CRP and Albumin data are synchronously obtained through bedside detection equipment, and CAR value calculation is completed within 15 minutes to provide a basis for early risk assessment. In outpatient follow-up, indicators are obtained through laboratory biochemical detection, CAR values are calculated, and change trends are tracked to assist in judging the progression of the disease. Since the correction coefficient k = 0.1 eliminates the systematic error of different laboratory detection methods (such as the confusion of mg / L and g / L units for CRP detection), multi-center research data is comparable.

[0043] More specifically, the CCAR index realizes the following nonlinear mapping relationship through a random forest model: ; Where is the prediction function constructed by the random forest algorithm, which realizes nonlinear feature combination by integrating multiple decision trees; and are the mean and standard deviation of the training set creatinine , and are the mean and standard deviation of the training set CAR value, used for standardization processing of input features; The standardized features satisfy the normal distribution with a mean of 0 and a standard deviation of 1, ensuring the stability and generalization ability of the random forest model training; The prognostic prediction performance of the CCAR index in the validation cohort needs to meet: AUC ≥ 0.75 (AUC improvement of more than 0.10 compared with traditional indicators CAR / NLR / PLR); log-rank test p < 0.001 for KM survival curves of high-risk and low-risk groups; CCAR index is applicable to the following types of cancer cachexia: pancreatic cancer, gastric cancer, colorectal cancer, lung cancer, and is verified in each cancer type (KM curve stratification significance <0.05).

[0044] The CCAR index is constructed using the random forest algorithm, which forms a strong learner by integrating 100 decision trees. Each tree uses the Bootstrap method to randomly sample 80% of the training set and randomly selects two features (CAR and creatinine) for optimal segmentation at each node split. The final output is determined by majority voting mechanism, which realizes the nonlinear combination of input features.

[0045] The nonlinear mapping calculation process is as follows: the input creatinine (Cr) and CAR values are standardized: ; where, is the mean and standard deviation of the training set creatinine, is the mean and standard deviation of the training set CAR.

[0046] The standardized features are input into the random forest model: ; where, is the prediction function of 100 integrated decision trees, which generates the final CCAR value by weighted summation of tree node conditions and leaf node outputs.

[0047] Model parameter optimization: the optimal parameters are determined by grid search: number of trees (n_estimators) = 100; maximum depth (max_depth) = 10; minimum sample split (min_samples_split) = 20; minimum sample leaf node number (min_samples_leaf) = 10; Data standardization process: the CAR and creatinine of the training set and test set are standardized using the StandardScaler algorithm, with the formula: ; Example: a patient with Cr = 128.3 mg / dL, training set , ,then: ; Standardized feature distribution verification: through Shapiro-Wilk test, the W statistics of standardized Cr and CAR are 0.98 and 0.97 (p>0.05), which meets the normal distribution assumption.

[0048] Engineering implementation: The StandardScaler.fit() method of Python's scikit-learn library is used in the training phase to calculate and , and save the parameters by pickle serialization. In the inference stage, load the standardization parameters and standardize the input data in real time through the StandardScaler.transform() method.

[0049] More specifically, the CCAR value is combined with the TNM stage to generate a composite prognostic model, and the risk score formula is: ; Where the weight coefficient =0.7, =0.3 is determined by fitting the Cox proportional hazards model, so that the prediction accuracy of the model for survival risk is maximized.

[0050] The composite prognostic model linearly combines the CCAR index and the TNM stage, and the formula is: ; Where: is the composite index calculated by the random forest model; is the numerical result of the TNM stage defined by the International Union Against Cancer (UICC) (e.g. stage I = 1, stage II = 2, stage III = 3, stage IV = 4).

[0051] Variable standardization: After numerical encoding of the TNM stage, the same standardization method as CCAR is used: ; Where, , (based on 1200 training set statistics).

[0052] In actual calculation, first standardize the TNM stage, and then weighted sum with CCAR: ; Where, , are the optimized weight coefficients.

[0053] Weight coefficient determination method, Cox proportional hazards model fitting process: variable input takes CCAR, TNM stage and patient survival time (months), outcome event (death = 1, survival = 0) as Cox model input variables; model construction uses partial likelihood function to estimate regression coefficients ; Where, is the risk function of an individual at time , is the baseline risk function.

[0054] Weight conversion: standardize the regression coefficients into weight coefficients: ; Actual measurement , , calculated , .

[0055] More specifically, the standardization process in S3 adopts the StandardScaler algorithm to make the input data satisfy the distribution with a mean of 0 and a standard deviation of 1; S3 includes the following steps before constructing the nonlinear composite index CCAR through the random forest machine learning model: LASSO regression is used to rank the feature importance of cancer cachexia patients, specifically: Collecting multi-dimensional clinical feature data of cancer cachexia patients including creatinine, hemoglobin, and platelets; using LASSO regression model to filter each feature and calculate the coefficient of each feature; determining creatinine as the key prediction feature according to the feature coefficient; The training steps of the random forest model include: dividing the collected cancer cachexia patient data into training set and validation set; standardizing CAR and creatinine in the training set; setting the parameters of the random forest model, including the number of trees and the maximum depth; training the model using the training set; evaluating the trained model through the validation set, adjusting the parameters to optimize the model performance; saving the optimized random forest model and standardization parameters.

[0056] LASSO regression model feature screening process: Data standardization, using Z-Score standardization method to preprocess the original features, eliminating the difference in dimension, the formula is: ; Among them, is the feature mean, is the feature standard deviation.

[0057] Model construction and parameter optimization: using the LassoCV module of scikit-learn library, determining the optimal regularization parameter through 10-fold cross-validation, and iteratively calculating the regression coefficients of each feature.

[0058] Feature importance ranking: ranking features according to the absolute value of the regression coefficient, keeping features with non-zero coefficients and removing features with zero coefficients.

[0059] Key feature determination Through LASSO regression analysis, the absolute value of the regression coefficient of creatinine (Cr) was significantly higher than that of other features (such as hemoglobin, platelets, etc.), combined with clinical significance (creatinine is a key indicator of kidney metabolism and nutritional status of the body, and is highly related to metabolic disorders in patients with cachexia), creatinine was determined as one of the core prediction features.

[0060] Steps of building and training the random forest model: The collected data of cancer cachexia patients was divided into training set and validation set in the ratio of 8:2, and stratified sampling was used to ensure that the two groups of data were evenly distributed in terms of cancer type (pancreatic cancer, gastric cancer, colorectal cancer, lung cancer, etc.), TNM stage, age and other clinical characteristics.

[0061] The CAR (calculated formula ) and creatinine (Cr) in the training set were standardized using the StandardScaler algorithm. The processed data met the normal distribution with a mean of 0 and a standard deviation of 1, and the formula was as follows: ; Where, 、 are the mean and standard deviation of the training set CAR, 、 are the mean and standard deviation of the training set creatinine.

[0062] Model parameter setting and training: Core parameter initialization: number of trees (n_estimators): initially set to 200; maximum depth (max_depth): initially set to 20; number of random feature samples (max_features): set to (the number of features is 2, i.e. CAR and Cr); minimum sample split (min_samples_split): set to 10.

[0063] Model training: using the standardized features of the training set, a random forest model is built through the Bagging strategy, and each decision tree is trained based on a randomly sampled feature subset and sample subset (sampling ratio is 0.8). Finally, the prediction result is generated through ensemble learning.

[0064] Model evaluation and parameter optimization: Validation set evaluation, using the validation set to calculate the following indicators: survival analysis AUC value (requirement AUC≥0.75); Kaplan-Meier survival curve log-rank test p value (requirement ); calibration curve consistency index (C-index).

[0065] Parameter tuning: Grid search was used to optimize n_estimators (range 100-500), max_depth (range 10-30), and max_features (range 1-2). Model performance was evaluated through 5-fold cross-validation after each tuning, until the AUC on the validation set improved by ≥0.10 and the KM curve stratified significance met the requirements.

[0066] After training is completed, the optimized random forest model and standardized parameters ( 、 、 、 ) is serialized and saved via pickle in the .pkl file format to ensure efficiency and stability when subsequently called by the network calculator.

[0067] More specifically, the backend server processing flow in S4 includes the following steps: receiving the patient's creatinine, C-reactive protein, and albumin data sent by the frontend via a POST request; validating the input data and prompting for re-entry if the data is outside the normal range; according to The CAR value is calculated using the formula [value=0.000] × 0.1; the CAR and creatinine are standardized using the mean and standard deviation obtained during the training phase; the pre-trained random forest model is called to generate the CCAR value; risk stratification is performed based on the relationship between the CCAR value and 0.56, and the 1-year, 3-year, and 5-year survival rates are calculated; the calculation results are returned to the front-end in JSON format.

[0068] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR, characterized in that: The following steps are involved: S1. Obtain serum biomarker data of patients with cancer cachexia, including Cr, CRP, and Albumin; S2. Calculate the C-reactive protein to albumin ratio (CAR) = CRP / Albumin; S3, using CAR and Cr as input features, constructs the nonlinear composite index CCAR through the random forest machine learning model; S4. Developing a network computer system: The front-end interface receives Cr, CRP, and Albumin data input by the user; The backend server calculates the CAR value and normalizes the CAR and Cr; Input the standardized data into the pre-trained random forest model to generate the CCAR value; Output CCAR value, risk stratification results, low risk: CCAR≤0.56, high risk: CCAR>0.56, and survival rate prediction, 1 year / 3 years / 5 years.

2. The method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR according to claim 1, characterized in that: The risk stratification threshold CCAR = 0.56 was determined by the maximum selection rank statistic in survival analysis and was used to distinguish the prognostic risk levels of patients with cancer cachexia; The network calculator front end integrates a dynamic heat map visualization module that uses color coding to map the negative correlation between CCAR value and survival probability, where an increase in CCAR value corresponds to a decrease in survival probability; The heat map uses a bivariate gradient coloring scheme, with the horizontal axis representing the CCAR value and the vertical axis representing the predicted survival time, and the color band gradually changes from green for high survival probability to red for low survival probability.

3. The method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR according to claim 2, characterized in that: The CAR ratio is calculated as: ; The correction factor =0.1, for balance and The dimensional difference makes The value is in a reasonable calculation range.

4. The method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR according to claim 3, characterized in that: described The unit is milligrams per liter, which represents the concentration of C-reactive protein in serum; The unit of is grams per deciliter, which represents the concentration of albumin in serum. Both are obtained through clinical blood tests.

5. The method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR according to claim 4, characterized in that: The CCAR indicator implements the following nonlinear mapping relationship through the random forest model: ; in The prediction function built for the random forest algorithm realizes nonlinear feature combination by integrating multiple decision trees; and The training set The mean and standard deviation of and are the mean and standard deviation of the CAR values ​​of the training set, respectively, used to standardize the input features; The standardized features satisfy the normal distribution with a mean of 0 and a standard deviation of 1, ensuring the stability and generalization ability of the random forest model training; The prognostic predictive efficacy of the CCAR indicator in the validation cohort must meet the following requirements: AUC ≥ 0.75, compared with the traditional indicators CAR / NLR / PLR AUC improvement of > 0.10; KM survival curve log-rank test p < 0.001 between high-risk group and low-risk group; The CCAR indicator is applicable to the following cancer cachexia types: pancreatic cancer, gastric cancer, colorectal cancer, and lung cancer, and has been validated in each cancer type through cohorts, with KM curve stratification significance. <0.

05.

6. The method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR according to claim 5, characterized in that: The CCAR value is combined with the TNM stage to generate a composite prognostic model, and its risk score formula is: ; The weight coefficient =0.7, =0.3 was determined by fitting the Cox proportional hazards model to maximize the model's prediction accuracy for survival risk.

7. The method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR according to claim 6, characterized in that: The standardization process in S3 uses the StandardScaler algorithm to make the input data satisfy the distribution with a mean of 0 and a standard deviation of 1. Before constructing the nonlinear composite indicator CCAR through the random forest machine learning model in S3, the following steps are also included: LASSO regression was used to rank the importance of characteristics of patients with cancer cachexia, specifically: multi-dimensional clinical characteristic data including creatinine, hemoglobin, and platelets were collected from patients with cancer cachexia; The LASSO regression model was used to screen each feature and calculate the coefficient of each feature; Creatinine was determined as the key predictive feature based on the characteristic coefficient; The training steps of the random forest model include: The collected cancer cachexia patient data were divided into training set and validation set; The CAR and creatinine in the training set were normalized; Set the parameters of the random forest model, including the number of trees and the maximum depth; Use the training set to train the model; Evaluate the trained model through the validation set and adjust the parameters to optimize the model performance; Save the optimized random forest model and standardized parameters.

8. The method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR according to claim 7, characterized in that: The backend server processing flow in S4 includes the following steps: Receive patient creatinine, C-reactive protein, and albumin data sent by the front end via POST requests; Verify the validity of the input data and prompt to re-enter if the data exceeds the normal range; The CAR value is calculated using the formula × 0.1; CAR and creatinine were normalized using the mean and standard deviation obtained during the training phase; Call the pre-trained random forest model to generate CCAR values; Risk stratification was performed based on the relationship between the CCAR value and 0.56, and the 1-year, 3-year, and 5-year survival rates were calculated; The calculation results are returned to the front end in JSON format.

9. The method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR according to claim 8, characterized in that: The implementation of the dynamic heat map visualization module includes the following steps: Collect CCAR values ​​and corresponding survival data of historical cancer cachexia patients; Construct a two-dimensional matrix of CCAR value and survival time, and calculate the average survival probability of different CCAR value intervals at each survival time point; The horizontal axis of the heat map is determined to be the CCAR value, the vertical axis is the survival time, and the color band changes from green to red, corresponding to high survival probability to low survival probability; Based on the current patient's CCAR value, the corresponding survival probability area is highlighted in the heat map; The heat map is displayed on the front-end interface to intuitively show the relationship between CCAR value and survival probability.

10. The method for constructing a cancer cachexia prognostic composite indicator CCAR based on a combination of creatinine and CAR according to claim 9, characterized in that: The steps for determining the risk stratification threshold of CCAR=0.56 are as follows: Sort the patients in the training cohort by CCAR value from small to large; Traverse all possible segmentation points and calculate the corresponding Statistics, the formula is: ; in, is the actual number of deaths in the low-risk group, is the expected number of deaths in the low-risk group, is the actual number of deaths in the high-risk group, is the expected number of deaths in the high-risk group, which is calculated based on the survival curve of the entire cohort; Select The segmentation point with the largest statistic is used as the initial threshold; pass The method verified the stability of the initial threshold and finally determined the risk stratification threshold to be 0.56; The Maxstat statistic is used to evaluate the ability of the cut-off point to distinguish survival risks. A larger statistic value indicates a more significant survival difference between the high-risk and low-risk groups at the cut-off point.

Citation Information

Patent Citations

  • System for predicting prognosis of patient suffering from liver cancer combined tumor-associated anemia caused by hepatitis B virus infection

    CN114141371A

  • Construction method and system for predicting prognostic index of tumor malignant fluid quality by combining inflammation and insulin resistance

    CN116864127A

  • Elderly kidney aging and chronic kidney disease identification model

    CN117766149A

  • Marker group and system for predicting bacteremia of tumor patient

    CN117877649A

  • Esophageal cancer prognosis survival prediction method and system

    CN118538402A

Cited By

  • Composite prognosis model for cancer prognosis evaluation, prognosis evaluation method and calculator

    CN121687498A