A leukemia patient survival risk prediction method based on cross-domain representation migration and adaptive fusion
Patent Information
- Application Number
- CN202611054051.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-15
AI Technical Summary
[0006]本发明旨在提供一种基于跨域表征迁移与自适应融合的白血病患者生存风险预测方法,以解决现有白血病患者生存风险预测存在目标样本数量有限、多源医学数据存在分布偏移、传统方法难以充分挖掘医学变量非线性交互信息的技术问题
[0020] Introducing external auxiliary datasets into the modeling process alleviates the problem of insufficient information caused by insufficient target patient samples in a single center and improves the predictive stability of the model under small sample conditions;
Smart Images

Figure CN122762297A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence-assisted medical data analysis, specifically to a method for predicting the survival risk of leukemia patients based on cross-domain representation transfer and adaptive fusion. Background Technology
[0002] Leukemia is a highly heterogeneous malignant hematological disease, and patient survival outcomes are influenced by a variety of clinical factors. Accurate survival risk assessment is of great significance for patient prognostic stratification and follow-up management.
[0003] Traditional prognostic analyses often employ statistical models, which struggle to fully uncover the complex nonlinear interactions between clinical variables. While machine learning models improve data fitting capabilities, the limited number of follow-up leukemia cases collected from single centers makes modeling with small samples prone to instability.
[0004] To expand the sample size, existing studies have attempted to incorporate patient data from outside the target region to aid modeling. However, patient populations from different regions exhibit baseline differences, and directly merging data can easily lead to inter-regional distribution shifts, causing negative transfer and reducing predictive performance on the target population.
[0005] Therefore, there is an urgent need for a survival risk prediction scheme that is adapted to small-sample medical data scenarios, can achieve cross-domain adaptive alignment in high-dimensional feature space, and integrate multi-level risk features to complete the prognostic assessment of leukemia patients. Summary of the Invention
[0006] This invention aims to provide a method for predicting the survival risk of leukemia patients based on cross-domain representation transfer and adaptive fusion, in order to solve the technical problems of limited target sample size, distribution bias in multi-source medical data, and difficulty in fully mining nonlinear interaction information of medical variables in existing methods for predicting the survival risk of leukemia patients.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A survival risk prediction method for leukemia patients based on cross-domain representation transfer and adaptive fusion includes the following steps:
[0009] S1. Obtain medical data of leukemia patients and divide it into external auxiliary datasets and local target datasets based on the data source. External auxiliary datasets are used to provide migration assistance information, while local target datasets are used to build survival risk prediction models for the target population.
[0010] S2. Preprocess the medical data of the two types of datasets, and perform high-dimensional mapping of the original medical features through representation learning to obtain high-dimensional risk representation features of patients.
[0011] S3. Construct a cross-domain feature space based on high-dimensional risk representation features, calculate the distribution differences between datasets, and determine the migration weights corresponding to samples from outside the domain based on the distribution differences.
[0012] S4. Based on the aforementioned transfer weights, adjust the contribution of out-of-domain auxiliary samples in the model training process, and integrate the patient's original medical characteristics, high-dimensional risk representation characteristics, and risk prediction characteristics to form a comprehensive risk prediction feature;
[0013] S5. Input the comprehensive risk prediction features into the survival risk prediction model, and output the patient risk score by combining the patient's survival time and event status information.
[0014] S6. Based on the risk score, the patient's risk level is classified, and the survival risk prediction results of leukemia patients are obtained.
[0015] Furthermore, the representation learning employs a pre-trained tabular feature representation model to perform high-dimensional mapping of patient medical variables.
[0016] Furthermore, the inter-domain distribution adaptation is based on a high-dimensional risk representation space, and the kernel mean matching method is used to calculate the migration weight of the auxiliary samples outside the domain.
[0017] Furthermore, the comprehensive risk prediction features consist of the patient's original medical characteristics, high-dimensional risk representation features, and short-term survival event probability features corresponding to a preset time window.
[0018] Furthermore, the survival risk prediction model adopts a gradient boosting tree-based survival analysis model.
[0019] Compared with the prior art, the present invention has the following beneficial effects:
[0020] Introducing external auxiliary datasets into the modeling process alleviates the problem of insufficient information caused by insufficient target patient samples in a single center and improves the predictive stability of the model under small sample conditions;
[0021] Cross-domain distribution adaptation is implemented in a high-dimensional risk representation space, which differs from the original clinical variable space alignment method. It utilizes the potential nonlinear association information between variables to reduce the impact of inter-domain distribution offset on the prediction results.
[0022] By integrating multi-level risk characteristics to construct a comprehensive predictive input, and by utilizing explicit clinical information and implicit risk characterization information, it is helpful to uncover multi-dimensional risk information of patients.
[0023] We adopted a survival analysis model adapted to censored follow-up data to improve the model's adaptability to long-term prognostic data;
[0024] An integrated processing architecture for representation transfer, distribution correction, and survival prediction is formed, and patient risk stratification results are output, providing a reference for medical staff to conduct prognostic assessment and individualized follow-up management. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the overall process of the survival risk prediction method for leukemia patients based on cross-domain representation transfer and adaptive fusion in an embodiment of the present invention;
[0026] Figure 2 This is a schematic diagram of the cross-domain representation migration framework between external auxiliary data and local target data in an embodiment of the present invention;
[0027] Figure 3 This is a schematic diagram of the domain distribution adaptation and migration weight calculation process based on a high-dimensional risk representation space in an embodiment of the present invention;
[0028] Figure 4 This is a schematic diagram of the multi-layer risk feature fusion and survival risk prediction process in an embodiment of the present invention;
[0029] Figure 5 This is a schematic diagram of patient risk stratification based on risk scores in an embodiment of the present invention. Detailed Implementation
[0030] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited to the following embodiments.
[0031] Example 1: Patient Data Acquisition and Preprocessing
[0032] This embodiment addresses the task of predicting the survival risk of leukemia patients by constructing a cross-domain predictive data environment that includes medical information from patients from different sources.
[0033] First, medical data of leukemia patients were acquired and divided into external auxiliary datasets and local target datasets based on the data source.
[0034] in:
[0035] The extra-domain auxiliary dataset is used to provide cross-domain migration assistance information, including patient basic information, disease status information, risk factor information, and survival follow-up information;
[0036] The local target dataset is used to build a survival risk prediction model for the target patient group and for the final model performance evaluation.
[0037] This embodiment uses a publicly available dataset of leukemia patients for validation.
[0038] The external auxiliary data contains 5,676 patient samples, of which 600 were selected for model training after screening. The local target data contains 590 patient samples, which are divided according to the ratio of training set to test set, with 472 samples used as target domain training data and 118 samples used as independent test data.
[0039] Test data is used only for the final model performance evaluation and is not involved in model training, feature processing, transfer weight calculation, or parameter determination.
[0040] In this embodiment, the patient's medical variables input include:
[0041] Clinical risk factors include age, sex, genetic risk, environmental exposure factors, lifestyle factors, disease stage, and comprehensive disease score.
[0042] Because patient groups from different sources differ in terms of population structure, disease status, and risk factor distribution, there may be data distribution offsets between the external auxiliary dataset and the local target dataset.
[0043] Therefore, this embodiment first preprocesses the patient's medical data.
[0044] Specifically, it includes:
[0045] (1) Handling missing values;
[0046] Missing data in patient medical variables are identified and supplemented or processed according to the variable type to reduce the impact of missing information on the model training process.
[0047] (2) Outlier handling;
[0048] Anomaly detection is performed on continuous medical indicators to reduce the impact of extreme samples on the stability of model training.
[0049] (3) Data standardization processing;
[0050] To address the issue of differences in the dimensions and ranges of values among various medical variables, continuous variables are standardized to map different characteristics to a unified scale.
[0051] For the The original medical characteristics of the patients after pretreatment are represented as follows:
[0052] in:
[0053] Indicates the dimensions of patient medical variables;
[0054] Indicates the first Original medical characteristics of patients after pretreatment.
[0055] The above data acquisition and preprocessing process provides standardized input for subsequent cross-domain characterization transfer and survival risk prediction.
[0056] Example 2: Generation of High-Dimensional Risk Features Based on Representation Learning
[0057] Because of the complex nonlinear interactions among patients' medical variables, directly using the original medical variables for cross-domain transfer may not be sufficient to fully express the patient's potential risk patterns.
[0058] Therefore, this embodiment first performs high-dimensional representation learning on the patient's medical characteristics.
[0059] Specifically, the original medical features of patients in the external auxiliary dataset and the local target dataset are input into the pre-trained tabular feature representation model. The model learns the potential correlation between different medical variables and generates high-dimensional risk representation features of patients.
[0060] The high-dimensional risk characterization process is represented as follows:
[0061] in:
[0062] Indicates the patient's original medical characteristics;
[0063] Representation learning model;
[0064] This indicates the patient's high-dimensional risk profile.
[0065] This embodiment uses the TabPFN model as a pre-trained table feature representation model.
[0066] The TabPFN model, based on a tabular data pre-training mechanism, is able to learn the potential relationships between different medical variables under limited sample conditions, mapping the original medical variables to a high-dimensional feature space containing patient risk information.
[0067] In this embodiment, the TabPFN model outputs high-dimensional embedding features of patients, and further uses principal component analysis to perform dimensionality reduction, retaining 36 principal components as the final risk representation, so as to balance the ability to express risk information and the computational efficiency of the model.
[0068] Through the above processing, patient medical data is mapped from the original variable space to a high-dimensional risk representation space, providing a unified feature basis for subsequent inter-domain distribution adaptation.
[0069] Example 3: Domain Distribution Adaptation Based on High-Dimensional Risk Representation Space
[0070] Because the external auxiliary dataset and the local target dataset have different sources, they differ in the distribution of patient risk characteristics.
[0071] Directly matching samples within the original medical variable space may not fully reflect the complex nonlinear relationships between medical variables.
[0072] Therefore, this embodiment completes inter-domain distribution adaptation within the high-dimensional feature space based on the high-dimensional risk characterization features obtained in Embodiment 2.
[0073] Let the high-dimensional representation of the external auxiliary data be:
[0074] The local target data is represented in high dimension as follows:
[0075] in:
[0076] Represents the set of auxiliary samples from outside the domain;
[0077] This represents the local target sample set.
[0078] This embodiment uses the kernel mean matching method to calculate the distribution differences between patient data from different sources and generates out-of-domain sample migration weights.
[0079] Its optimization objective is expressed as:
[0080] in:
[0081] Indicates the first The migration weights corresponding to each out-of-domain auxiliary sample;
[0082] This represents a nonlinear feature mapping function.
[0083] This embodiment uses the RBF kernel function for distribution matching, and the kernel parameters are set as follows:
[0084] The migration weight constraint parameters are set as follows:
[0085] By solving the above optimization objective, the migration weights corresponding to each external auxiliary sample are obtained.
[0086] Samples with higher weights indicate that their high-dimensional risk feature distribution is closer to the local target patient group, and thus contribute more to the subsequent model training process.
[0087] Furthermore, to verify the domain distribution adaptation effect, this embodiment compares the distribution distance between domains before and after weighting.
[0088] The results show:
[0089] Before weighting: RBF-MMD = 0.00399; mean distance = 0.24066.
[0090] After weighting: RBF-MMD = 0.00239; mean distance = 0.06766.
[0091] The results indicate that the distributional differences between external auxiliary data and local target data are reduced after adjusting the migration weights in the high-dimensional space.
[0092] By using the above method, this embodiment avoids information bias caused by direct migration in the original medical variable space and improves the effectiveness of migrating external auxiliary patient data to the target patient survival risk prediction task.
[0093] Example 4: Multi-layer risk feature fusion
[0094] After completing the adaptation of the distribution of auxiliary samples outside the domain, this embodiment further integrates patient risk information at different levels to construct a comprehensive risk prediction feature.
[0095] Traditional medical risk prediction methods typically rely on patients' original clinical variables for modeling, making it difficult to simultaneously represent both the patient's explicit medical status and potential risk patterns. Therefore, this embodiment integrates the patient's original medical characteristics, high-dimensional risk representation features, and short-term risk probability features to form a comprehensive risk prediction input.
[0096] For the The patient's original medical characteristics are as follows:
[0097] in:
[0098] This represents the dimensions of the patient's original medical variables.
[0099] The high-dimensional risk representation vector obtained through representation learning is represented as follows:
[0100] in:
[0101] This represents a high-dimensional risk representation dimension.
[0102] In this embodiment, after principal component analysis:
[0103] That is, 36 principal components are retained as the final high-dimensional risk representation.
[0104] Furthermore, the probability of short-term survival events in patients is calculated based on a preset time window:
[0105] in:
[0106] Indicates the patient's survival time;
[0107] Indicates the preset prediction time window;
[0108] This indicates the probability that the target event will occur within the time window.
[0109] The above three types of risk information are integrated:
[0110] in:
[0111] This indicates the patient's comprehensive risk prediction characteristics.
[0112] Through the above-described fusion method, this embodiment simultaneously utilizes:
[0113] (1) Explicit risk information contained in the patient’s original medical characteristics;
[0114] (2) Information on potential risk patterns obtained through representation learning;
[0115] (3) Risk change information corresponding to the short-term time window.
[0116] Compared to prediction methods that use only a single type of feature, this embodiment can generate a more complete description of patient risk, providing comprehensive input for subsequent survival risk prediction.
[0117] Example 5: Construction of a Gradient Boosting Survival Risk Prediction Model Based on the Cox Objective Function
[0118] Because follow-up data for leukemia patients often includes cases where the target event did not occur in some patients during the observation period, i.e., right-censored data exists.
[0119] Therefore, this embodiment uses a gradient boosting tree-based survival analysis model suitable for censored follow-up data to predict patient survival risk.
[0120] Specifically, the comprehensive risk prediction features obtained in Example 4 Input the survival risk prediction model.
[0121] For the The survival information of the patients is represented as follows:
[0122] Among them:
[0123] Indicates the patient's observation period;
[0124] Indicates the event status.
[0125] Event state is defined as:
[0126]
[0127] This embodiment uses a gradient boosting tree-based survival model based on the Cox proportional hazards objective function for training.
[0128] The model learns the nonlinear relationship between a patient's overall risk characteristics and survival risk stepwise through multiple weak learners, which can be expressed as follows:
[0129] in:
[0130] Indicates the first A weak learner;
[0131] This indicates the number of weak learners.
[0132] The model outputs a relative survival risk score for patients:
[0133] in:
[0134] The larger the value, the higher the patient's corresponding survival risk.
[0135] Furthermore, based on the ideas of the Cox proportional hazards model:
[0136] in:
[0137] Indicates the patient's time Risk function under certain conditions;
[0138] This represents the baseline risk function.
[0139] This embodiment uses a gradient boosting tree structure to learn the complex nonlinear relationship between medical variables and survival risk, and combines patient survival time and event status information to achieve ranking and prediction of patients' relative risk.
[0140] Compared to prediction methods that only classify events, this embodiment can make full use of time information during the follow-up process, making it more suitable for long-term prognostic risk analysis scenarios for leukemia patients.
[0141] Example 6: Risk Stratification and Predictive Performance Evaluation
[0142] To verify the effectiveness of the method of the present invention in the task of predicting the survival risk of leukemia patients, this embodiment uses independent test data to evaluate the model performance.
[0143] Evaluation metrics include the Concordance Index (C-index), the area under the time-dependent receiver operating characteristic curve (AUC), and the survival differences in risk stratification.
[0144] 1. Consistency Index Evaluation
[0145] The consistency index is used to evaluate the degree of consistency between the model and the patient survival risk ranking.
[0146] The calculation formula is as follows:
[0147] in:
[0148] This represents the patient risk score predicted by the model;
[0149] This indicates the actual observation time for the patient.
[0150] A higher C-index indicates a stronger consistency between the model and the patient risk ranking.
[0151] 2. Short-term risk prediction and assessment
[0152] To evaluate the risk prediction capability of patients experiencing target events within a preset time window, this embodiment uses the area under the receiver operating characteristic curve (AUC) for evaluation.
[0153] Its calculation form is as follows:
[0154] in:
[0155] TPR stands for True Rate;
[0156] FPR represents the false positive rate.
[0157] 3. Risk stratification assessment
[0158] Based on the patient risk score output by the model, patients are classified into risk levels.
[0159] In this embodiment, the median risk score of the patients in the test set was used as the risk classification threshold to divide the 118 test patients into a high-risk group and a low-risk group, with 59 patients in each group.
[0160] Furthermore, the Kaplan-Meier method was used to estimate survival curves for patients in different risk groups:
[0161] in:
[0162] Indicates a point in time Number of patients who experienced the event;
[0163] This indicates the number of patients in the high-risk concentration at that point in time.
[0164] The Log-rank test was used to compare survival differences between different risk groups.
[0165] The results show:
[0166] The event incidence rate was 67.80% in the high-risk group and 30.51% in the low-risk group. The Log-rank test result was... .
[0167] This indicates that the risk score obtained based on the method of the present invention can effectively distinguish patients with different prognostic levels.
[0168] 4. Model Performance Comparison
[0169] To further verify the effectiveness of the method of the present invention, this embodiment sets up multiple comparison models under the same test set conditions.
[0170] in:
[0171] (1) Traditional Cox survival model;
[0172] (2) Only the TabPFN short-term risk prediction model is used;
[0173] (3) The method of directly merging out-of-domain samples without performing migration weight correction;
[0174] (4) Complete model of the present invention.
[0175] The experimental results are as follows:
[0176] Table 1 Comparison of prediction performance of different models
[0177] Traditional Cox Model 0.5863 0.5994 TabPFN Risk Prediction Model 0.5872 0.6023 Directly merge out-of-domain sample models 0.6113 0.6509 Complete model of the invention 0.7139 0.7585
[0178] The experimental results show that, compared with traditional survival analysis models and migration methods without domain adaptation, the method of this invention achieves higher risk ranking ability and short-term risk differentiation ability.
[0179] This demonstrates that high-dimensional representation transfer, domain distribution correction, and multi-level risk feature fusion can improve the prediction of survival risk for leukemia patients under small sample conditions.
[0180] Example 7: Module Validation and Multi-Module Collaborative Analysis
[0181] To verify the synergistic effect between the core technology modules of this invention, this embodiment further verifies the effectiveness of the modules.
[0182] While maintaining consistency in data partitioning methods, training strategies, and evaluation metrics, the key modules of this invention are removed one by one, and the model performance under different schemes is compared.
[0183] The following comparison scheme is set up:
[0184] Table 2 Module validity verification results
[0185] Control group 1 Complete model 0.7139 0.7585 Control group 2 Remove TabPFN high-dimensional representation module 0.5479 0.5667 Control group 3 Remove high-dimensional spatial domain distribution adaptation module 0.6113 0.6509 Control group 4 Remove short-term survival probability feature 0.6463 0.6801 Control group 5 Replace the XGB survival model with the traditional Cox model. 0.6244 0.6474 Control group 6 Training using only local target data 0.5720 0.5766
[0186] The above verification results show that there are clear data processing dependencies among the various technical modules of this invention.
[0187] Specifically:
[0188] First, the representation learning module maps patient medical data from different sources to a unified high-dimensional risk representation space, so that potential risk patterns can be effectively expressed.
[0189] Secondly, the distribution adaptation method is used to adjust the contribution of external auxiliary samples in the high-dimensional risk representation space to reduce the distribution differences between patient data from different sources.
[0190] Finally, the original medical characteristics, high-dimensional risk representations, and short-term risk probability characteristics are integrated and input into the survival risk prediction model to complete the risk score.
[0191] The model performance declined to varying degrees after the removal of different modules, indicating that the combined effect of the above modules can achieve a synergistic prediction effect.
[0192] Example 8: Model Parameter Setting and Standardization Implementation
[0193] To ensure the reproducibility of the method of the present invention, this embodiment describes the main parameters and implementation methods in the model training process.
[0194] It should be noted that the following parameters are only a preferred setting in this embodiment. They can be adjusted according to the actual scale of medical data and application scenarios without affecting the technical effect of the present invention.
[0195] (1) Representation learning parameter settings
[0196] This embodiment uses the TabPFN model to learn high-dimensional representations of patient medical variables.
[0197] The model input consists of preprocessed patient medical characteristics: Model outputs high-dimensional risk embeddings for patients:
[0198] Principal component analysis was further used to reduce the dimensionality of the high-dimensional embedding.
[0199] In this embodiment, 36 principal components are retained as the final risk characterization features to reduce feature redundancy and improve model computation efficiency.
[0200] (2) Domain distribution adaptation parameter settings
[0201] This embodiment uses the kernel mean matching method to calculate the migration weight of out-of-domain auxiliary samples.
[0202] in:
[0203] High-dimensional spatial distribution matching is performed using the RBF kernel function;
[0204] In this embodiment, the kernel parameters are set to .
[0205] The migration weight constraint parameters are set to
[0206] The above settings allow external auxiliary samples to receive different contribution weights based on the similarity of their characteristic distributions to the local target patient group.
[0207] (3) Parameter settings for survival risk prediction model
[0208] This embodiment uses a gradient boosting tree-based survival model based on the Cox proportional risk objective function for risk prediction.
[0209] The main parameters of the model are set as follows:
[0210] Tree depth: 3; learning rate: 0.11; number of base learners: 75; L2 regularization parameter: 0.10.
[0211] By setting the parameters mentioned above, the model can learn the nonlinear relationship between patient medical characteristics and survival risk while controlling model complexity.
[0212] (4) Model training and testing process settings
[0213] This embodiment strictly separates training data and test data.
[0214] in:
[0215] During the training phase, only the target training data and the out-of-domain auxiliary data adjusted for transfer weights are used to build the model.
[0216] The test set is only used for the final model performance evaluation.
[0217] During model training:
[0218] Data processing parameters are determined solely based on the training data;
[0219] The representation learning process is completed solely based on data from the training phase.
[0220] The domain distribution adaptation weights are calculated using only the training data;
[0221] Test data is not used in any model parameter updates or weight calculation processes.
[0222] The above process avoids the leakage of test information and ensures the independence and reliability of model evaluation results.
Claims
1. A method for predicting survival risk in leukemia patients based on cross-domain representation transfer and adaptive fusion, characterized in that, Includes the following steps: S1. Obtain medical data of leukemia patients and divide the patient data into an external auxiliary dataset and a local target dataset according to the source of the patient data. The external auxiliary dataset is used to provide migration assistance information, and the local target dataset is used to establish a survival risk prediction model for target patients. S2. Preprocess the patient medical data in the extra-domain auxiliary dataset and the local target dataset, and perform high-dimensional mapping of patient medical features through representation learning methods to obtain high-dimensional feature representations containing potential patient risk information; S3. Construct a cross-domain feature space based on the high-dimensional feature representation, calculate the distribution difference between the extra-domain auxiliary dataset and the local target dataset, and determine the migration weights corresponding to the extra-domain auxiliary samples according to the distribution difference to reduce the data distribution offset between patient data from different sources; S4. Based on the aforementioned transfer weights, adjust the contribution of out-of-domain auxiliary samples in the model training process, and integrate the patient's original medical characteristics, high-dimensional risk representation characteristics, and risk prediction characteristics to form a comprehensive risk prediction feature; S5. Input the comprehensive risk prediction features into the survival risk prediction model, and generate a patient risk score based on the patient's survival time information and event status information. S6. Based on the risk score, classify the patient into risk levels to obtain the survival risk prediction results for leukemia patients.
2. The method for predicting survival risk of leukemia patients based on cross-domain representation transfer and adaptive fusion according to claim 1, characterized in that, The external auxiliary dataset and the local target dataset come from different regions or medical systems, and both contain basic patient information, disease status, risk factors, and survival follow-up data.
3. The method according to claim 1, characterized in that, The representation learning described in step S2 uses a pre-trained tabular feature representation model to mine the nonlinear interaction relationships between various medical variables and map the original features to generate high-dimensional risk representation features of patients; the pre-trained tabular feature representation model is the TabPFN model.
4. The method according to claim 3, characterized in that, The inter-domain distribution differences are calculated based on high-dimensional risk representation features, rather than directly based on the original patient variables, in order to reduce the distribution bias between different data sources.
5. The method according to claim 4, characterized in that, In step S3, the kernel mean matching method is used to calculate the transfer weights corresponding to each external auxiliary sample. The weights are used to adjust the degree of influence of each external auxiliary sample when participating in model training.
6. The method according to claim 1, characterized in that, The comprehensive risk prediction feature described in step S4 is composed of three types of features: the patient's original medical characteristics, high-dimensional risk representation features, and short-term survival event probability features corresponding to the preset time window.
7. The method according to claim 1, characterized in that, The survival risk prediction model described in step S5 is a gradient boosting tree-based survival analysis model, which can simultaneously utilize patient observation duration and endpoint censoring markers to output a relative survival risk score for the patient.
8. The method according to claim 1, characterized in that, Step S6 divides patients into multiple risk groups based on their overall risk scores and uses a survival difference test to determine the stratification effect of different risk groups.