A prognosis prediction method based on multi-modal intermediate fusion

By integrating pathological images, radiological images, and clinical information, and utilizing deep convolutional neural networks and multi-instance learning methods, a multimodal prognostic prediction score is generated. This addresses the insufficient accuracy of existing models in predicting the risk of recurrence after renal clear cell carcinoma surgery, achieving a more efficient prediction effect.

CN119905237BActive Publication Date: 2026-01-06SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411734971.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2026-01-06
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

Existing recurrence risk prediction models for clear cell renal cell carcinoma rely on single-modality data, which cannot fully capture the intrinsic heterogeneity of the tumor, resulting in decreased prediction accuracy. Furthermore, single nucleotide polymorphism sequencing is a complex and costly process.

Method used

We employ residual deep convolutional neural networks and predefined omics to extract features from pathological and radiological images. By combining multi-instance learning and deep survival networks, we integrate pathological, imaging, and clinical information to generate multimodal prognostic prediction scores.

Benefits of technology

It improves the accuracy of predicting the risk of recurrence after clear cell renal cell carcinoma surgery, overcomes the problem of missing data, and significantly enhances the predictive performance of traditional models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119905237B_ABST
    Figure CN119905237B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of prognosis prediction method based on multi-modal intermediate fusion, which extracts pathological features from pathological images by deep convolutional neural network and pre-defined omics, and applies multi-instance learning to aggregate these features to form pathological representation, while using deep convolutional neural network and pre-defined omics to extract radiological features to form radiological representation, then using deep survival network to integrate pathological representation, radiological representation and clinical variables to generate multi-modal prognosis prediction score.Compared with prior art, the present application integrates the information of three modalities of pathology, radiology and clinic through deep survival network, significantly improves the prediction performance of traditional prognosis prediction indicators and single-modal model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tumor image analysis technology, and in particular to a prognostic prediction method based on multimodal intermediate fusion. Background Technology

[0002] Renal cell carcinoma is considered one of the most common malignant tumors of the urinary system. Clear cell renal cell carcinoma (ccRCC), as the major subtype of RCC, accounts for approximately 70% of renal cancer cases. While surgery is the primary initial treatment, it is noteworthy that approximately 20%–30% of ccRCC patients experience recurrence and metastasis after surgery, significantly impacting overall survival. Therefore, accurately identifying high-risk individuals for recurrence is crucial for clinicians to make better decisions and intervene promptly.

[0003] Currently, researchers have developed various tools based on clinical characteristics (including TNM staging and histopathological grade) to assess the risk of tumor recurrence after surgery for clear cell renal cell carcinoma. Among these, the Leibovich score and UISS score are the most commonly used assessment criteria. However, in the clinical decision-making process, physicians have observed significant differences in survival outcomes among patients within the same risk subgroup, indicating that traditional clinical characteristics may not be sufficient to fully capture the intrinsic heterogeneity of tumors. Furthermore, a prospective study of eight prognostic prediction models widely used in a clinical trial cohort showed that the predictive accuracy of all developed models significantly decreased compared to the initially reported results.

[0004] In recent years, some studies have begun to apply artificial intelligence for prognostic prediction. Most models use single-modal data (e.g., utilizing only pathological, imaging, or genetic information). In contrast, multimodal models can integrate comprehensive and complementary medical data, thereby improving the accuracy of predicting postoperative recurrence in patients. One study developed a multimodal recurrence model for ccRCC by integrating clinical features, single nucleotide polymorphisms (SNPs), and whole-section images of tumor tissue. However, SNP sequencing is complex and costly. Currently, a prognostic prediction model that can effectively integrate multimodal information is lacking. Summary of the Invention

[0005] The purpose of this invention is to provide a prognostic prediction method based on multimodal intermediate fusion. Pathological features are extracted from pathological images using residual deep convolutional neural networks and predefined omics, and these features are aggregated using multi-instance learning to form a pathological representation. Simultaneously, radiological features are extracted using residual deep convolutional neural networks and predefined omics to form an image representation. Subsequently, a deep survival network is used to integrate the pathological representation, image representation, and clinical information to generate a multimodal prognostic prediction score for accurate prognostic prediction.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] A prognostic prediction method based on multimodal intermediate fusion includes the following steps:

[0008] Step S1, Data Acquisition: Obtain whole slide images (WSI) stained with hematoxylin-eosin (HE) as pathological images, obtain arterial phase three-dimensional computed tomography (3D CT) sequences as radiological images, and obtain the patient's clinical variables;

[0009] Step S2, Data Preprocessing: The pathological images are segmented, stained, normalized, and filtered; the radiological images are converted to different data formats and the tumor regions are delineated; and the clinical variables are coded.

[0010] Step S3, Pathological Image Feature Representation: Extract manually defined features and semantic features based on deep convolutional neural networks from each pathological patch after data preprocessing, and aggregate the features of different pathological patches through a multi-instance learning method to obtain the feature representation of the pathological image;

[0011] Step S4, Radiographic Image Feature Representation: Based on the radiographic image and the outlined tumor contour, manually defined features and semantic features based on deep convolutional neural networks are extracted, and dimensionality reduction is performed through principal component analysis to obtain the feature representation of the radiographic image.

[0012] Step S5: Construction of multimodal prognostic prediction model: A deep survival network is used to integrate features of pathological images, radiological images and clinical variables to output a multimodal prognostic prediction score.

[0013] The preprocessing of the pathological images in step S2 specifically includes the following steps:

[0014] Segmentation: Dividing pathological images into small pathological blocks of a preset size;

[0015] Color normalization: Color normalization is performed on each pathological patch in RGB three-channel format;

[0016] Filtering: The OTSU method is used to calculate the proportion of tissue region in each pathological block, and pathological blocks with a tissue region proportion less than a preset value are filtered out.

[0017] The preprocessing of the radiological image in step S2 specifically includes the following steps:

[0018] Format conversion: Convert DICOM format radiographic images to NII format;

[0019] Resampling: Resampling radiographic images at predetermined slice intervals;

[0020] Gray-level normalization: Normalizes the gray levels of a radiographic image to a preset range;

[0021] Tumor region delineation: For each 2D scan section in the 3D sequence of radiological images, delineate the tumor outline;

[0022] Region of Interest Extraction: By delineating the tumor region, a 3D region of interest containing the tumor is cropped from the original radiological image;

[0023] Maximum tumor section selection: Based on the 3D region of interest containing the tumor and the tumor annotation obtained by cropping, the 2D image section containing the largest tumor is selected.

[0024] The tumor region delineation uses a pre-trained deep segmentation learning model nnU-Net in public data to automatically delineate kidney tumors, followed by manual review of the images to correct any incorrect tumor boundaries identified by the model.

[0025] The clinical variables include basic vital signs and tumor-related indicators. The basic vital signs include gender, age, and ECOG PS score. The tumor-related indicators include tumor size, pT stage, pN stage, TNM stage, necrosis status, and whether it is sarcomatoid renal cell carcinoma.

[0026] In step S2, the encoding and preprocessing of basic vital signs indicators specifically involves:

[0027] For the gender indicator, male is coded as 1 and female as 0;

[0028] For the age indicator, 1 is coded for ages exceeding a preset age value, and 0 is coded for ages not exceeding a preset age value.

[0029] For the ECOG PS score, which is used to assess the patient's physical condition, an ECOG PS score ≥ 1 is coded as 1, and an ECOG PS score < 1 is coded as 0.

[0030] In step S2, the encoding preprocessing of tumor-related indicators specifically involves:

[0031] For tumor size, a tumor with a maximum diameter greater than or equal to a preset value is coded as 1, and a tumor with a maximum diameter less than a preset value is coded as 0.

[0032] For pT staging, the pT staging is used to assess the size and invasion of the tumor, with T1 coded as 1, T2 coded as 2, and T3 coded as 3;

[0033] For pN staging, which is used to assess lymph node status, N0 is encoded as 0 and N1 is encoded as 1;

[0034] For TNM staging, stage I is coded as 1, stage II as 2, and stage III as 3;

[0035] In the case of necrosis, the presence of tumor necrosis is coded as 1, and the absence of necrosis is coded as 0;

[0036] For whether it is sarcomatoid renal cell carcinoma, 1 is the code for sarcomatoid renal cell carcinoma and 0 is the code for not being sarcomatoid renal cell carcinoma.

[0037] In step S3, the semantic features based on the deep convolutional neural network are obtained by extracting the image representation through the ResNet network; the specific steps for extracting manually defined features include:

[0038] Cell nucleus segmentation was performed using a pre-trained HoverNet model, and cell types were predicted.

[0039] Based on cell nucleus segmentation results and pathological images, features were extracted using the HistomicsTK package in Python.

[0040] For each pathological patch, the average value of the HistomicsTK features of the most frequently occurring cell type is calculated and concatenated with the one-hot encoding of that cell type to form the hand-defined features extracted for each pathological patch.

[0041] The multi-instance learning method used in step S3 to aggregate features of different pathological fragments employs a Transformer structure.

[0042] In step S4, the semantic features based on the deep convolutional neural network are extracted through the ResNet network; the manually defined features include shape features, first-order statistical features, texture features, filtering features, and wavelet features defined in the Pyradiomics package; when performing principal component analysis on shape, first-order statistical, texture, filtering, and wavelet features, only the principal components that cause the explanation ratio to exceed the preset value are retained.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] (1) Compared with existing prognostic prediction models that utilize pathological images or genetic information, this invention utilizes information from three modalities—pathology, imaging, and clinical—that are easier to obtain and more comprehensive in clinical practice. Information acquisition is easier and the available data is richer, which can improve the training accuracy of the prognostic prediction model.

[0045] (2) The present invention utilizes an intermediate fusion method to balance the situation of missing data and the performance of the model.

[0046] (3) This invention uses deep survival networks to integrate multimodal prognostic information, which significantly improves the performance of traditional prognostic prediction indicators and single-modal models in predicting recurrence risk. Attached Figure Description

[0047] Figure 1 This is a flowchart of the method of the present invention;

[0048] Figure 2 This is a flowchart of sample screening in an embodiment of the present invention;

[0049] Figure 3 This is a consistency correction curve between the nomogram prediction and the actual results established by the multimodal prognosis prediction model in this embodiment of the invention.

[0050] Figure 4 This is a schematic diagram illustrating the effect of the multimodal prognostic prediction model on patient risk stratification in an embodiment of the present invention. Part A represents the stratification results of all cohorts, part B represents the results of the training set, part C represents the results of the validation set, and part D represents the results of the test set. Detailed Implementation

[0051] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0052] This embodiment uses the prognostic prediction of patients with clear cell renal cell carcinoma as an example to provide a prognostic prediction method based on multimodal intermediate fusion, such as... Figure 1 As shown, the method includes the following steps:

[0053] Step S1: Data Acquisition: Obtain 40x magnified whole slide images (WSI) stained with hematoxylin-eosin (HE) from patients with clear cell renal cell carcinoma after surgery as pathological images. Obtain the patient's preoperative arterial phase three-dimensional computed tomography (3D CT) sequence from the database as radiological images. Obtain the patient's clinical variables.

[0054] The pathological images in this embodiment are derived from HE-stained images of postoperative samples from patients with clear cell renal cell carcinoma at multiple hospitals, as well as renal tumor diagnostic slide images from the TCGA database. The radiological images in this embodiment are derived from preoperative arterial phase 3D CT sequences of patients with clear cell renal cell carcinoma acquired internally by multiple hospitals.

[0055] Step S2, Data Preprocessing: The pathological images are segmented, stained, normalized, and filtered; the radiological images are converted to different data formats and the tumor regions are delineated; and the clinical variables are coded.

[0056] Step S21: Preprocessing of pathological images, specifically including the following steps:

[0057] Segmentation: The 40x magnified pathological image is segmented into small pathological blocks, each with a size of 768×768 pixels;

[0058] Color normalization: Color normalization is performed on each 768×768 pixel pathological patch in RGB three-channel format;

[0059] Filtering: The proportion of tissue region in each pathological block was calculated using the OTSU method, and pathological blocks with a tissue region proportion of less than 0.3 were filtered out.

[0060] Step S22: Preprocessing of radiological images, specifically including the following steps:

[0061] Format conversion: Convert DICOM format radiographic images to NII format;

[0062] Resampling: Resampling the radiographic image with a slice interval of 1 mm;

[0063] Gray-level normalization: Normalizes the gray level of the radiographic image to [-140, 260];

[0064] Tumor region delineation: For each 2D scan section in the 3D sequence of radiological images, delineate the tumor outline;

[0065] Region of Interest Extraction: By delineating the tumor region, a 3D region of interest containing the tumor is cropped from the original radiological image;

[0066] Maximum tumor section selection: Based on the 3D region of interest containing the tumor and the tumor annotation obtained by cropping, the 2D image section containing the largest tumor is selected.

[0067] In this embodiment, the tumor region delineation uses the nnU-Net deep segmentation learning model pre-trained in public data to automatically delineate the kidney tumor. Afterwards, manual review of the images is performed to confirm or modify the tumor boundaries that are incorrectly segmented by the model.

[0068] Step S23, preprocessing of clinical variables, specifically:

[0069] In this embodiment, clinical variables include basic vital signs and tumor-related indicators. Basic vital signs include gender, age, and ECOG PS score. Tumor-related indicators include tumor size, pT stage, pN stage, TNM stage, necrosis status, and whether it is sarcomatoid renal cell carcinoma.

[0070] For the gender indicator, male is coded as 1 and female as 0;

[0071] For the age indicator, those over 60 years old are coded as 1, and those under 60 years old are coded as 0;

[0072] For the ECOG PS score (used to assess a patient's performance status), ECOG PS ≥ 1 is coded as 1, and ECOG PS < 1 is coded as 0.

[0073] For tumor size, tumors with a maximum diameter ≥10cm are coded as 1, and tumors with a maximum diameter <10cm are coded as 0;

[0074] For pT staging (used to assess tumor size and invasiveness), T1 is coded as 1, T2 as 2, and T3 as 3;

[0075] For pN staging (used to assess lymph node status), N0 is encoded as 0 and N1 is encoded as 1;

[0076] For TNM staging, stage I is coded as 1, stage II as 2, and stage III as 3;

[0077] In the case of necrosis, the presence of tumor necrosis is coded as 1, and the absence of necrosis is coded as 0;

[0078] For whether it is sarcomatoid renal cell carcinoma, 1 is the code for sarcomatoid renal cell carcinoma and 0 is the code for not being sarcomatoid renal cell carcinoma.

[0079] Step S3, Pathological Image Feature Representation: For each pathological patch after data preprocessing, manually defined features and semantic features based on deep convolutional neural networks are extracted. The features of different pathological patches are aggregated through a multi-instance learning method to obtain the feature representation of the pathological image.

[0080] Specifically, the semantic features based on deep convolutional neural networks are obtained by extracting image representations through the ResNet network. In particular, a ResNet network pre-trained on the ImageNet dataset is used, and the corresponding image representation is obtained by removing its last classification layer.

[0081] The specific steps for extracting manually defined features include:

[0082] Cell nucleus segmentation was performed using the HoverNet model pre-trained on the PanNuke dataset, and the cell type (any one of tumor cells, inflammatory cells, junctional cells, dead cells, and normal epithelial cells) was predicted.

[0083] Based on cell nucleus segmentation results and pathological images, features were extracted using the HistomicsTK package in Python.

[0084] For each pathological patch, the average value of the HistomicsTK features of the most frequently occurring cell type is calculated and concatenated with the one-hot encoding of that cell type to form the hand-defined features extracted for each pathological patch.

[0085] After completing the feature representation of each pathological section, a Transformer-based multi-instance learning method is used to aggregate the features of different sections, thereby generating the feature representation of the whole pathological slide.

[0086] In this embodiment, a pathology-based prognostic prediction is obtained by connecting linear layers. This prediction is then compared with the patient's actual survival label, and the parameters of the multi-instance learning part are optimized by calculating negative log loss.

[0087] Step S4, Radiographic Image Feature Representation: Based on the radiographic image and the outlined tumor contour, manually defined features and semantic features based on deep convolutional neural networks are extracted, and dimensionality reduction is performed through principal component analysis to obtain the feature representation of the radiographic image.

[0088] The semantic features based on deep convolutional neural networks are extracted through a ResNet network. Specifically, the extracted semantic features take the largest tumor section obtained in the radiological image preprocessing stage as input, and use a ResNet model pre-trained on the ImageNet dataset to remove the last classification layer in order to obtain the semantic feature representation of the image.

[0089] Manually defined features are used to extract shape features, first-order statistical features, texture features, filtering features, and wavelet features defined in the Pyradiomics package, based on the 3D region of interest containing the tumor and tumor annotations obtained during the radiological image preprocessing stage. When performing principal component analysis on the shape, first-order statistical, texture, filtering, and wavelet features, only principal components that explain a proportion greater than 0.8 are retained.

[0090] The principal components extracted from each feature dimension are concatenated to generate a summary radiological feature. This summary feature is then input into a Cox proportional hazards model to obtain a prognostic prediction. This prediction is compared with the patient's actual survival label, and the model parameters for the radiological image feature extraction part are optimized.

[0091] Step S5: Construction of multimodal prognostic prediction model: A deep survival network is used to integrate features of pathological images, radiological images and clinical variables to output a multimodal prognostic prediction score.

[0092] In this embodiment, a deep survival network is used to extend the linear covariate modeling of the traditional proportional hazards model Cox to a nonlinear form of a neural network.

[0093] This embodiment uses the time-dependent area under the receiver operating characteristic (AUC) curve at different time points to evaluate model performance. Furthermore, for low-risk and high-risk groups, the Kaplan-Meier method is used for survival estimation, risk stratification is performed based on the median predicted scores of the training set, and the significance is evaluated using the log-rank test. Statistical analysis and hypothesis testing are performed using the survival, survcomp, survminer, timeROC, and rms software packages in R4.2.2.

[0094] To verify the accuracy of this invention, the following examples are provided for verification:

[0095] 1) Dataset

[0096] Data were collected from patients with clear cell renal cell carcinoma at six hospitals between 2014 and 2021. The discovery cohort included 788 patients from three centers, some of whom had low-quality CT images or tumor tissue pathology slides, resulting in a lack of complete multimodal information. From the cohort with complete multimodal information, 134 patients were randomly selected as the validation cohort, and the remaining 654 patients were assigned to the training cohort, ensuring that patients lacking multimodal information were only assigned to the training cohort. Furthermore, in constructing the pathological model, this invention also used diagnostic slide data from 503 patients with clear cell renal cell carcinoma in the Cancer Genome Atlas (TCGA) database. An independent external test cohort included 357 patients from the remaining three hospitals, such as... Figure 2 As shown.

[0097] Table 1 presents the basic information of all patients in the internal center. Baseline characteristics were uniformly distributed across the training, validation, and testing cohorts, with no statistically significant differences between groups (p>0.05). Overall, the majority of patients were male (67.4%), and more than half were over 60 years of age (50.5%). In terms of pathological features, most patients presented stage I (79.0%) and grade 2 pathology (63.5%). A minority of patients showed poor prognostic features, such as tumor necrosis (14.0%) and sarcomatoid differentiation (1.7%). In different risk stratification models, most patients were classified as low-risk or intermediate-risk. The median follow-up time was 49 months, during which 197 patients (17.2%) experienced recurrence of renal cell carcinoma. P-values ​​in the table were obtained using Pearson's Chi-squared test, Fisher's exact test, or Kruskal-Wallis rank sum test.

[0098] Table 1. Patient demographic data and baseline characteristics

[0099]

[0100] 2) Comparison of different prediction models

[0101] To verify the advantages of the proposed Multimodal Input (MPRS) model in predicting recurrence of clear cell renal cell carcinoma, this embodiment compares several unimodal models, including a pathology-based recurrence prediction model (HPRS), an imaging-based recurrence prediction model (RPRS), and a clinical data-based recurrence prediction model (CPRS). Furthermore, this invention also compares the Leibovich score and UISS score, commonly used in assessing the severity of renal cell carcinoma.

[0102] In all cohorts, the MPRS model outperformed traditional unimodal models (HPRS, RPRS, CPRS) and Leibovic and UISS scores in predicting recurrence of clear cell renal cell carcinoma (Table 2). Specifically, in the training cohort, the MPRS model had an AUC of 0.907 (95% CI, 0.844–0.969) for 3-year recurrence prediction and 0.835 (95% CI, 0.780–0.890) in the validation cohort. In 5-year survival prediction, MPRS continued to demonstrate the highest discriminative power, with an AUC of 0.897 (95% CI, 0.820–0.974) in the validation cohort and 0.829 (95% CI, 0.768–0.891) in the test cohort.

[0103] Table 2 Comparison of different prediction models

[0104]

[0105] Furthermore, the MPRS calibration curves for each cohort showed a close alignment with the 45° reference line, further demonstrating a significant agreement between MPRS predictions and actual observations. Figure 3 As shown in the figure. Through stratified analysis of high- and low-risk groups, the relapse-free survival between the high-risk and low-risk groups defined by MPRS showed significant differences in the training, validation, and testing cohorts (p<0.001, e.g., ...). Figure 4 As shown in the figure, this fully demonstrates the effectiveness of deep learning-based multimodal input in predicting recurrence of clear cell renal cell carcinoma.

[0106] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A prognosis prediction method based on multi-modal intermediate fusion, characterized in that, The method comprises the following steps: Step S1, data collection: obtaining a H-E stained whole section image as a pathology image, obtaining an arterial phase three-dimensional computed tomography sequence as a radiology image, and obtaining clinical variables of the patient; Step S2, data preprocessing: performing patching, staining normalization processing and filtering preprocessing on the pathology image, performing data format conversion and tumor region delineation preprocessing on the radiology image, and performing encoding preprocessing on the clinical variables; Step S3, pathology image feature representation: extracting hand-defined features and semantic features based on a deep convolutional neural network from each pathology patch after data preprocessing, and obtaining feature representation of the pathology image by aggregating features of different pathology patches through a multi-instance learning method; Step S4, radiology image feature representation: performing hand-defined feature extraction and semantic feature extraction based on a deep convolutional neural network on the radiology image and the delineated tumor contour, and obtaining feature representation of the radiology image through dimension reduction processing by principal component analysis; Step S5, multi-modal prognosis prediction model construction: integrating features of the pathology image, the radiology image and the clinical variables by using a deep survival network, and outputting a multi-modal prognosis prediction score; The semantic features based on the deep convolutional neural network in step S3 are obtained by a ResNet network; The specific steps of extracting the hand-defined features include: Using a pre-trained HoverNet model to perform nucleus segmentation and predict cell types; Using a HistomicsTK package in Python to extract features based on the nucleus segmentation result and the pathology image; For each pathology patch, calculating the average value of the HistomicsTK features of the cell type with the highest frequency, and splicing the average value with the one-hot encoding of the cell type to form the hand-defined features extracted from each pathology patch; The semantic features based on the deep convolutional neural network in step S4 are obtained by a ResNet network; the hand-defined features include shape features, first-order statistical features, texture features, filtering features and wavelet features defined in a Pyradiomics package; when performing principal component analysis on the shape, first-order statistical, texture, filtering and wavelet features, only the principal components whose explanation ratio exceeds a preset value are retained.

2. The prognosis prediction method based on multi-modal intermediate fusion according to claim 1, characterized in that, The preprocessing of the pathology image in step S2 specifically includes the following steps: Patching: dividing the pathology image into pathology patches of a preset size; Color normalization: performing color normalization on each pathology patch in RGB three-channel format; Filtering: calculating the proportion of the tissue region in each pathology patch by using an OTSU method, and filtering the pathology patches with a tissue region proportion less than a preset value.

3. The prognosis prediction method based on multi-modal intermediate fusion according to claim 1, characterized in that, The preprocessing of the radiology image in step S2 specifically includes the following steps: Format conversion: converting the radiology image in DICOM format into NII format; Resampling: resampling the radiology image at a set slice interval; Gray scale normalization: normalizing the gray scale of the radiology image to a preset interval; Tumor region delineation: delineating a tumor contour for each 2D scan slice in the 3D sequence of the radiology image; Region of interest extraction: 3D region of interest containing tumor is cropped from original radiology image by the drawn tumor region; Maximum tumor cross-section selection: 2D image cross-section containing maximum tumor is selected based on the cropped 3D region of interest containing tumor and tumor annotation.

4. The prognosis prediction method based on multi-modal intermediate fusion according to claim 3, characterized in that, The tumor region drawing uses a deep segmentation learning model nnU-Net pre-trained in public data to automatically draw kidney tumors, and then manual reading is performed to modify the tumor boundary that is incorrect in model segmentation.

5. The prognosis prediction method based on multi-modal intermediate fusion according to claim 1, characterized in that, The clinical variables include basic vital sign indicators and tumor-related indicators, wherein the basic vital sign indicators include gender, age and ECOG PS score, and the tumor-related indicators include tumor size, pT stage, pN stage, TNM stage, necrosis, and whether it is a sarcomatoid renal cell carcinoma.

6. The prognosis prediction method based on multi-modal intermediate fusion according to claim 5, characterized in that, In the step S2, the encoding preprocessing of the basic vital sign indicators is specifically: For the gender indicator, male is coded as 1 and female is coded as 0; For the age indicator, age greater than a preset age value is coded as 1 and age less than the preset age value is coded as 0; For the ECOG PS score, the ECOG PS score is used to evaluate the physical condition of the patient, and ECOG PS≥1 is coded as 1 and ECOG PS<1 is coded as 0.

7. The prognosis prediction method based on multi-modal intermediate fusion according to claim 5, characterized in that, In the step S2, the encoding preprocessing of the tumor-related indicators is specifically: For the tumor size, tumor maximum diameter greater than or equal to a preset value is coded as 1 and tumor maximum diameter less than the preset value is coded as 0; For the pT stage, the pT stage is used to evaluate the size and infiltration of the tumor, and T1 is coded as 1, T2 is coded as 2, and T3 is coded as 3; For the pN stage, the pN stage is used to evaluate the lymph node status, and N0 is coded as 0 and N1 is coded as 1; For the TNM stage, stage I is coded as 1, stage II is coded as 2, and stage III is coded as 3; For the necrosis, the presence of tumor necrosis is coded as 1 and the absence of necrosis is coded as 0; For whether it is a sarcomatoid renal cell carcinoma, sarcomatoid renal cell carcinoma is coded as 1 and non-sarcomatoid renal cell carcinoma is coded as 0.

8. The prognosis prediction method based on multi-modal intermediate fusion according to claim 1, characterized in that, The multi-instance learning used to aggregate features of different pathological pieces in the step S3 adopts a Transformer structure.