Tumor prognosis risk prediction method based on multi-modal medical image fusion

Through the tumor prognosis risk prediction method of multimodal medical image fusion, the MRI and WSI characteristics were extracted using ResNet34 and transformer encoder, combined with Grad-CAM analysis, the problem of CRPC risk prediction after androgen deprivation treatment in prostate cancer patients was solved, and more efficient and accurate risk assessment was achieved.

CN120299694APending Publication Date: 2025-07-11SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510189848.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the risk of prostate cancer patients developing into castration-resistant prostate cancer after androgen deprivation treatment, and the multimodal data characteristics fusion effect is poor, the calculation complexity is high, and the model is insufficient interpretability.

Method used

The tumor prognostic risk prediction method based on multimodal medical image fusion was adopted, and the MRI and WSI image features were extracted through the ResNet34 network, and the self-attention and cross-attention transformer encoder were used for feature fusion, and the model interpretability analysis was combined with Grad-CAM to output the risk probability of the patient's development into CRPC.

Benefits of technology

It improves the accuracy of multimodal data feature fusion, reduces the computational complexity, enhances the generalization ability and interpretability of the model, and can accurately analyze the lesion areas that the model focuses on, improving the accuracy of CRPC risk prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299694A_ABST
    Figure CN120299694A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical image processing, in particular to a tumor prognosis risk prediction method based on multi-modal medical image fusion, which comprises the following steps: firstly, acquiring patient data, and preprocessing the patient data; the patient data comprises an MRI image and a WSI image of the prostate cancer patient; then, MRI image features and WSI image features are respectively extracted from the preprocessed data; next, the MRI image features and the WSI image features are fused, and fused features are obtained; and finally, inputting the fusion features into a deep learning prediction model, and outputting a risk probability value that the patient develops into CRPC. The tumor prognosis risk prediction method based on multi-modal medical image fusion provided by the invention can effectively predict the risk of prostate cancer patients developing into castration-resistant prostate cancer after androgen deprivation treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of medical image processing, and particularly to a tumor prognosis risk prediction method based on multi-modal medical image fusion. Background Art

[0002] Patients with prostate cancer (PCa), as a common tumor of the male urogenital system, are an important issue affecting the health of middle-aged and elderly men. The standard clinical treatment for advanced prostate cancer is usually androgen deprivation therapy, which mainly includes castration therapy, anti-androgen drug therapy, and maximum androgen blockade therapy, namely castration combined with anti-androgen drug therapy. Androgen deprivation therapy (ADT) plays a crucial role in the comprehensive treatment of PCa. Although ADT can temporarily control the condition of most patients and achieve biochemical and clinical remission, some patients develop into castration-resistant prostate cancer (CRPC) due to androgen receptor antagonist resistance 1 to 2 years after treatment. The occurrence of CRPC marks the further deterioration of the condition and shows a significantly poor prognosis. Therefore, how to effectively predict the risk that PCa patients develop into CRPC after receiving treatment is a key challenge in the treatment of prostate cancer. Summary of the Invention

[0003] An embodiment of this application provides a tumor prognosis risk prediction method based on multi-modal medical image fusion, which can effectively predict the risk that prostate cancer patients develop into castration-resistant prostate cancer after androgen deprivation therapy.

[0004] To solve the above technical problems, an embodiment of this application provides a tumor prognosis risk prediction method based on multi-modal medical image fusion, including the following steps: First, obtain patient data and preprocess the patient data; the patient data includes MRI images and WSI images of prostate cancer patients; then, extract MRI image features and WSI image features from the preprocessed data respectively; next, fuse the MRI image features and WSI image features to obtain fused features; finally, input the fused features into a deep learning prediction model to output the risk probability value that the patient develops into CRPC.

[0005] In some exemplary embodiments, obtaining patient data and preprocessing the patient data includes: collecting prostate MRI images and WSI images of prostate cancer patients; respectively outlining the tumor lesion areas of the MRI images and WSI images and annotating the tumor cell areas.

[0006] In some exemplary embodiments, the tumor lesion region of the MRI image is outlined and the tumor cell region is labeled, including: for the MRI data, 5 slices are equally spaced and extracted for each of the T1 and T2 modalities. A single slice is compressed into a 224×224 pixel image, and each patient corresponds to an MRI data of 10×1×224×224.

[0007] In some exemplary embodiments, the tumor lesion region of the WSI image is outlined and the tumor cell region is labeled, including: for the WSI data, at a magnification of 20×10, it is cropped according to the pixel size of 1024×1024 to obtain n patches of the tumor lesion region with a size of 3×1024×1024, and the patches are compressed into a data dimension of 3×224×224. Each patient corresponds to a WSI data of 1×n×3×224×224.

[0008] In some exemplary embodiments, MRI image features are extracted from the preprocessed data, including: using a pre-trained ResNet34 network to extract MRI image features, and representing the MRI image features as a 1×512-dimensional MRI feature vector.

[0009] In some exemplary embodiments, WSI image features are extracted from the preprocessed data, including: using a pre-trained ResNet34 network to perform preliminary patch-level feature extraction on n WSI patches to obtain n 1×512-dimensional feature vectors; using a 1×512-dimensional token as the global feature representation of the WSI, and passing the n patch features and the token through a transformer network based on the self-attention mechanism to achieve WSI-level feature extraction, obtaining a feature sequence of m×512; where m = n + 1, and the token corresponds to the global preliminary feature of the WSI.

[0010] In some exemplary embodiments, the MRI image features and the WSI image features are fused to obtain fused features, including: adding the MRI features and the WSI features vectorially as the preliminary fused feature vector; using a cross-attention-based transformer encoder network to perform inter-modal feature interaction to obtain a 1×512-dimensional fused feature vector as the fused feature.

[0011] In some exemplary embodiments, after obtaining the fused features and before inputting the fused features into the deep learning prediction model, it further includes: inputting the fused features into a multi-layer perceptron, and using a multi-layer fully connected network to process the fused features to improve the prediction performance of the model.

[0012] In some exemplary embodiments, the deep learning prediction model is a fully connected neural network model; by inputting the fused features into the trained fully connected neural network model, the risk probability value of the patient developing into CRPC is output; the training process of the fully connected neural network model includes: training the built fully connected neural network model on a server with 4 A6000 graphics cards, using the Pytorch deep learning framework, and its training parameters are: learning rate 0.001, the optimizer adopts the SGD optimizer, the number of iteration rounds is 200 times, and an early stopping learning strategy is adopted to avoid overfitting.

[0013] In some exemplary embodiments, after inputting the fused features into the deep learning prediction model and outputting the risk probability value of the patient developing into CRPC, it further includes: grabbing the gradients in the gradient backward propagation process through backward_hook to generate the global heat map of the WSI, and generating the local heat map of a single patch through Grad-CAM to display the data attention focus of the model predicting the patient's survival probability, which has clinical significance and model interpretability.

[0014] The technical solution provided by the embodiments of the present application has at least the following advantages:

[0015] The embodiments of the present application provide a tumor prognosis risk prediction method based on multimodal medical image fusion, and the method includes the following steps: First, obtain patient data and preprocess the patient data; the patient data includes the MRI images and WSI images of prostate cancer patients; then, extract the MRI image features and WSI image features from the preprocessed data respectively; next, fuse the MRI image features and WSI image features to obtain fused features; finally, input the fused features into the deep learning prediction model to output the risk probability value of the patient developing into CRPC.

[0016] The present application provides a tumor prognosis risk prediction method based on multimodal medical image fusion. Aiming at the risk problem of prostate cancer patients developing into CRPC after ADT treatment, an end-to-end deep learning model with multiple attention mechanism fusions based on the patient's MRI data and pathological tissue H&E staining WSI data is proposed, which can effectively extract and fuse the MRI modality and WSI modality data for the survival analysis of patients developing into CRPC. Aiming at the black box problem of the deep learning model, the present application also uses technologies such as Grad-CAM to perform interpretability analysis on the model, and can accurately analyze the lesion areas concerned by the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Unless otherwise stated, the figures in the drawings do not constitute a scale limitation.

[0018] Figure 1 This is a flowchart of a tumor prognosis risk prediction method based on multimodal medical image fusion provided by an embodiment of the present application.

[0019] Figure 2 This is a specific flowchart of a tumor prognosis risk prediction method based on multimodal medical image fusion provided by an embodiment of the present application. Detailed implementation manners

[0020] As can be seen from the background art, how to effectively predict the risk that PCa patients develop into CRPC after receiving treatment is a key challenge in the treatment of prostate cancer. To address this problem, the present application designs an automatic detection method for predicting the risk that prostate cancer patients develop into castration-resistant prostate cancer after androgen deprivation therapy.

[0021] Currently, in the aspect of medical image feature extraction, common methods include traditional feature extraction methods, deep learning-based feature extraction methods such as convolutional neural networks (CNNs) and vision transformers (ViTs), etc. However, most traditional feature extraction methods rely on single-modal data, which is insufficient for the analysis efficiency and accuracy of complex lesions. Although the application of deep learning and neural network methods has been improved, how to efficiently fuse information between multiple data sources and improve the robustness and generalization ability of the model remains a difficult problem.

[0022] Research based on multimodal data provides new possibilities for the prognosis assessment of prostate cancer. Magnetic resonance imaging (MRI) and whole-slide imaging (WSI) of pathological sections are widely used to evaluate the morphology, structure, and microscopic tissue characteristics of tumors. As a non-invasive imaging technique, MRI can provide rich soft tissue contrast and spatial information of tumors, and plays an important role especially in the localization and staging of prostate tumors. The pathological section WSI can reveal the microscopic structure and heterogeneity of tumors by finely showing the cellular and molecular changes of prostate tissue. Therefore, combining the two-modal data of MRI and WSI, especially in predicting which patients may progress to CRPC, can provide more comprehensive information for the prognosis of prostate cancer patients.

[0023] The main disadvantages of the prior art include: (1) The heterogeneity of features in multiple modalities is high, and the current feature fusion methods for multimodal data have poor effects and it is difficult to fully utilize the information between different modalities. (2) In the process of feature extraction and fusion of the existing methods, the computational complexity is relatively high, and it is difficult to perform calculations in an end-to-end manner, which affects the processing efficiency. (3) The existing methods lack evidence in terms of model interpretability and it is difficult to explain which features or regions of the data play a key role in the prediction.

[0024] In view of these problems, the purpose of this application is to provide a multi-modal data feature extraction and fusion method based on deep learning. Regarding the risk problem that prostate cancer patients develop into CRPC after ADT treatment, it is an end-to-end deep learning model that fuses multiple attention mechanisms based on the patient's MRI data and pathological tissue H&E staining WSI data. It can effectively extract and fuse MRI modality and WSI modality data for the survival analysis of patients developing into CRPC. The method includes the following steps: First, obtain patient data and preprocess the patient data; the patient data includes MRI images and WSI images of prostate cancer patients; then, extract MRI image features and WSI image features from the preprocessed data respectively; next, fuse the MRI image features and WSI image features to obtain fused features; finally, input the fused features into a deep learning prediction model to output the risk probability value of the patient developing into CRPC. The tumor prognosis risk prediction method based on multi-modal medical image fusion provided by this application can effectively predict the risk that prostate cancer patients develop into castration-resistant prostate cancer after androgen deprivation therapy.

[0025] This application provides a tumor prognosis risk prediction method based on multi-modal medical image fusion. The single-modal features of WSI and MRI are extracted by a specific feature extractor, and then the extracted features are fused and interacted through a transformer encoder modality fusion network based on self-attention and cross-attention to obtain fused features, which are used to predict the risk probability of PCa patients developing into CRPC. In addition, in view of the black box problem of the deep learning model, this application also uses technologies such as Grad-CAM to perform interpretability analysis on the model, and can accurately analyze the lesion areas concerned by the model.

[0026] The following will elaborate on each embodiment of this application with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of this application, many technical details are presented to help readers better understand this application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can still be implemented.

[0027] See Figure 1 , the embodiment of this application provides a tumor prognosis risk prediction method based on multi-modal medical image fusion, including the following steps:

[0028] Step S101, obtain patient data and preprocess the patient data; the patient data includes MRI images and WSI images of prostate cancer patients.

[0029] Step S102: Extract MRI image features and WSI image features from the preprocessed data respectively.

[0030] Step S103: Fuse the MRI image features and the WSI image features to obtain fused features.

[0031] Step S104: Input the fused features into a deep learning prediction model to output the risk probability value of the patient developing into CRPC.

[0032] In view of the problem of how to effectively predict the risk that PCa patients develop into CRPC after receiving treatment and prognosis, this application provides a tumor prognosis risk prediction method based on multi-modal medical image fusion. On the one hand, this application proposes a fusion strategy for two modalities of data, namely MRI and WSI. On the other hand, this application also proposes a deep learning prediction model that deeply fuses two modalities of data, MRI and WSI, to predict the risk that prostate patients develop into castration-resistant prostate cancer after androgen deprivation therapy.

[0033] In some embodiments, in step S101, patient data is obtained and preprocessed for the patient data, including the following steps:

[0034] Step S1011: Collect prostate MRI images and WSI images of prostate cancer patients.

[0035] Step S1012: Outline the tumor lesion regions (Region of Interest, ROI) of the MRI images and the WSI images respectively, and label the tumor cell regions.

[0036] Specifically, in step S1011, collect the tumor MRI images of prostate cancer patients, the digital pathological images (Whole Slide Image, WSI) of the H&E staining of the patient's tumor tissue, and the clinical information of the patient.

[0037] In some embodiments, in step S1012, outlining the tumor lesion region of the MRI image and labeling the tumor cell region includes: for the MRI data, 5 slices are equally spaced and extracted for each of the T1 and T2 modalities. A single slice is compressed into a 224×224 pixel image, and each patient corresponds to a 10×1×224×224 MRI data.

[0038] In some embodiments, in step S1012, the tumor lesion area of the WSI image is outlined to label the tumor cell area, including: for the WSI data, at a magnification of 20×10, cropping is performed according to the pixel size of 1024×1024 to obtain n patches of size 3×1024×1024 for the ROI, and the patches are compressed into a data dimension of 3×224×224. Each patient corresponds to 1 WSI data of n×3×224×224.

[0039] In some embodiments, in step S102, extracting MRI image features from the preprocessed data includes: using a pre-trained ResNet34 network to extract MRI image features, and representing the MRI image features as a 1×512-dimensional MRI feature vector.

[0040] In some embodiments, in step S102, extracting WSI image features from the preprocessed data includes: using a pre-trained ResNet34 network to perform preliminary patch-level feature extraction on n WSI patches to obtain n 1×512-dimensional feature vectors; using a 1×512-dimensional token as the global feature representation of the WSI, and passing the n patch features and the token (class token) through a transformer network based on the self-attention mechanism to achieve WSI-level feature extraction, obtaining a feature sequence of m×512; where m = n + 1, and the class token corresponds to the global preliminary feature of the WSI.

[0041] In some embodiments, in step S103, fusing the MRI image features and the WSI image features to obtain fused features includes: performing vector addition on the MRI features and the WSI features as the preliminary fused feature vector; using a transformer encoder network based on cross attention to perform inter-modal feature interaction to obtain a 1×512-dimensional fused feature vector as the fused feature.

[0042] In some embodiments, after obtaining the fused features in step S103 and before inputting the fused features into the deep learning prediction model in step S104, it further includes:

[0043] Step S1031: Input the fused features into a multi-layer perceptron (MLP), and use a multi-layer fully connected network to process the fused features to improve the prediction performance of the model.

[0044] Finally, the fused feature vectors of the patient's MRI and WSI are input into the deep learning prediction model, and the risk probability value Risk (0 < Risk < 1) of the patient developing into CRPC is output.

[0045] In some embodiments, the deep learning prediction model in step S104 is a fully connected neural network model; by inputting the fused features into the trained fully connected neural network model, the risk probability value of the patient developing into CRPC is output; the training process of the fully connected neural network model includes: training the built fully connected neural network model on a server with 4 A6000 graphics cards, using the Pytorch deep learning framework, and its training parameters are: learning rate 0.001, the optimizer uses the SGD optimizer, the number of iteration rounds is 200 times, and an early stopping learning strategy is adopted to avoid overfitting.

[0046] In some embodiments, after inputting the fused features into the deep learning prediction model in step S104 and outputting the risk probability value of the patient developing into CRPC, it further includes:

[0047] Step S105, grab the gradients in the process of backward gradient transmission through backward_hook to generate the global heat map of WSI, and generate the local heat map of a single patch through Grad-CAM to display the data attention focus of the model predicting the patient's survival probability, which has clinical significance and model interpretability.

[0048] This application proposes a tumor prognosis risk prediction method based on multi-modal medical image fusion, which is used to predict the probability of PCa patients developing into CRPC after ADT treatment. The prediction method provided by this application will be introduced in detail through specific embodiments below.

[0049] First, data collection and preprocessing are carried out.

[0050] Collect the prostate MRI images of prostate cancer patients, and uniformly save the data format as a.nii format file; collect the digital pathological images of the H&E stained sections of the patients, and uniformly save the data format as a.svs format file.

[0051] The ROI regions in MRI and WSI are outlined respectively. For MRI data, 5 slices are evenly sampled for each of the T1 and T2 modalities, and each single slice is compressed into an image of 224×224 pixels. Each patient corresponds to an MRI data of 10×1×224×224 (since MRI is a grayscale image, its number of color channels is 1); for WSI data, at a magnification of 20×10, it is cropped according to the pixel size of 1024×1024 to obtain n patches of size 3×1024×1024 for the ROI (since WSI is an RGB image, its number of color channels is 3), and the patches are compressed into a data dimension of 3×224×224. Each patient corresponds to a WSI data of n×3×224×224.

[0052] Then, model construction is carried out; the model construction includes a feature extraction module, a feature fusion module, and a multi-layer perceptron processing module.

[0053] (1) Feature extraction module: The ResNet34 network can effectively avoid the problem of gradient disappearance, improve the depth and accuracy of feature extraction, and at the same time take into account a relatively low computational complexity. It is used in the model for feature extraction of MRI and WSI patches: a) MRI feature extraction: As Figure 2 shown, the pre-trained ResNet34 network is used to implement feature extraction of MRI, which is represented as a 1×512-dimensional MRI feature vector; b) WSI feature extraction: Continuing to refer to Figure 2 , the pre-trained ResNet34 is used to perform preliminary patch-level feature extraction on n WSI patches to obtain n 1×512-dimensional feature vectors. In addition, a class token is used as the global feature representation (1×512) of the WSI. Subsequently, the n patch features and the class token are passed through a self-attention-based transformer network to achieve WSI-level feature extraction, resulting in a feature sequence of (n + 1)×512, where the class token corresponds to the global preliminary feature of the WSI.

[0054] (2) Feature fusion: After feature extraction, the present application fuses the features of different modalities through an attention mechanism. Specifically, the MRI feature and the WSI feature are added as vectors to form a preliminary fusion feature vector, and further passed through a cross-attention-based transformer encoder network to assign different weights to each modality, so as to highlight the features most useful for disease diagnosis in each modality, perform feature interaction between modalities, and finally obtain a 1×512-dimensional fusion feature vector.

[0055] (3) Multi-layer perceptron processing: Input the fused features into a multi-layer perceptron, and further process the features through a multi-layer fully connected network to improve the prediction performance of the model.

[0056] Then, model training is carried out. The established deep learning model is trained on a server with 4 A6000 graphics cards, using the Pytorch (1.4.0) deep learning framework. Its training parameters are: learning rate 0.001, the optimizer uses the SGD optimizer, the number of iteration rounds is 200 times, and a learning strategy of early stopping is applied to avoid overfitting.

[0057] Finally, the gradients in the process of gradient backpropagation are captured through backward_hook to generate a global attention map of the WSI, and local heatmaps of single patches are generated through Grad-CAM, which can display the data attention focus of the model's predicted patient survival probability, having clinical significance and model interpretability.

[0058] Compared with the prior art, the tumor prognosis risk prediction method based on multi-modal medical image fusion provided by this application has the following advantages:

[0059] The prediction method of this application effectively improves the accuracy of multi-modal data feature fusion, and automatically weights the features of different modalities through the attention mechanism. This application optimizes the model computational complexity, reduces the processing time, adapts to the needs of large-scale data processing, and proposes an end-to-end trainable model that shows stronger generalization ability on different data sets, can effectively avoid overfitting, improves the stability of diagnosis, and can effectively predict the risk probability of prostate cancer patients developing into castration-resistant prostate cancer after androgen deprivation therapy.

[0060] This application has been verified through experiments. On the prostate cancer data set, the prediction method of this application has significantly improved in indicators such as accuracy and recall compared with traditional methods.

[0061] With the above technical solutions, the embodiments of this application provide a tumor prognosis risk prediction method based on multi-modal medical image fusion. The method includes the following steps: First, obtain patient data and preprocess the patient data; the patient data includes MRI images and WSI images of prostate cancer patients; then, extract MRI image features and WSI image features from the preprocessed data respectively; next, fuse the MRI image features and WSI image features to obtain fused features; finally, input the fused features into a deep learning prediction model to output the risk probability value of the patient developing into CRPC.

[0062] The present application provides a tumor prognosis risk prediction method based on multimodal medical image fusion. Aiming at the risk problem that prostate cancer patients develop into CRPC after ADT treatment, an end-to-end deep learning model integrating multiple attention mechanisms based on the patient's MRI data and pathological tissue H&E staining WSI data is proposed, which can effectively extract and fuse the MRI modality and WSI modality data for the survival analysis of patients developing into CRPC. Aiming at the black box problem of the deep learning model, the present application also uses technologies such as Grad-CAM to perform interpretability analysis on the model, and can accurately analyze the lesion areas concerned by the model.

[0063] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.

Claims

1. A tumor prognosis risk prediction method based on multimodal medical image fusion, characterized in that It includes the following steps: Obtain patient data and preprocess the patient data; the patient data includes MRI images and WSI images of prostate cancer patients; Extract MRI image features and WSI image features from the preprocessed data respectively; Fuse the MRI image features and the WSI image features to obtain fused features; Input the fused features into a deep learning prediction model to output the risk probability value of the patient developing into CRPC.

2. The tumor prognosis risk prediction method based on multimodal medical image fusion according to claim 1, wherein Obtain patient data and preprocess the patient data, including: Collect prostate MRI images and WSI images of prostate cancer patients; Outline the tumor lesion areas of the MRI images and the WSI images respectively and label the tumor cell areas.

3. The tumor prognosis risk prediction method based on multimodal medical image fusion according to claim 2, wherein Outline the tumor lesion area of the MRI image and label the tumor cell area, including: For MRI data, extract 5 slices at equal intervals for T1 and T2 modalities respectively. A single slice is compressed into a 224×224 pixel image, and each patient corresponds to a 10×1×224×224 MRI data.

4. The tumor prognosis risk prediction method based on multimodal medical image fusion according to claim 2, wherein Outline the tumor lesion area of the WSI image and label the tumor cell area, including: For WSI data, at a magnification of 20×10 times, crop according to the pixel size of 1024×1024 to obtain n patches of the tumor lesion area with a size of 3×1024×1024, and compress the patches into a data dimension of 3×224×224. Each patient corresponds to 1 n×3×224×224 WSI data.

5. The tumor prognosis risk prediction method based on multi-modal medical image fusion according to claim 1, characterized in that Extract MRI image features from the preprocessed data, including: Adopt a pre-trained ResNet34 network to extract MRI image features and represent the MRI image features as a 1×512-dimensional MRI feature vector.

6. The tumor prognosis risk prediction method based on multimodal medical image fusion according to claim 1, wherein Extract WSI image features from the preprocessed data, including: Adopt a pre-trained ResNet34 network to perform preliminary patch-level feature extraction on n WSI patches to obtain n 1×512-dimensional feature vectors; Use a 1×512-dimensional token as the global feature representation of the WSI, and pass the n patch features and the token through a transformer network based on the self-attention mechanism to achieve WSI-level feature extraction, obtaining a feature sequence of m×512; where m = n + 1, and the token corresponds to the global preliminary feature of the WSI.

7. The tumor prognosis risk prediction method based on multimodal medical image fusion according to claim 1, characterized in that Fuse the MRI image features and the WSI image features to obtain fused features, including: Perform vector addition on the MRI features and the WSI features as the preliminary fused feature vector; Adopt a transformer encoder network based on cross-attention to perform inter-modal feature interaction to obtain a 1×512-dimensional fused feature vector as the fused feature.

8. The tumor prognosis risk prediction method based on multimodal medical image fusion according to claim 1, wherein After obtaining the fused features and before inputting the fused features into the deep learning prediction model, it further includes: Input the fused features into a multi-layer perceptron, and use a multi-layer fully connected network to process the fused features to improve the prediction performance of the model.

9. The method for predicting tumor prognosis risk based on multimodal medical image fusion according to claim 1, wherein The deep learning prediction model is a fully connected neural network model; By inputting the fusion features into the trained fully-connected neural network model, the risk probability value of the patient developing into CRPC is output; The training process of the fully-connected neural network model includes: training the built fully-connected neural network model on a server with 4 A6000 graphics cards, using the Pytorch deep learning framework, with its training parameters being: learning rate 0.001, the optimizer using the SGD optimizer, the number of iteration rounds being 200 times, and adopting an early stopping learning strategy to avoid overfitting.

10. The tumor prognosis risk prediction method based on multimodal medical image fusion according to claim 1, characterized in that, After inputting the fusion features into the deep learning prediction model and outputting the risk probability value of the patient developing into CRPC, it further includes: Grabbing the gradients during the gradient backward transmission process through backward_hook to generate the global heat map of the WSI, and generating the local heat map of a single patch through Grad-CAM to display the data attention focus of the model predicting the patient's survival probability, which has clinical significance and model interpretability.

Citation Information

Cited By

  • Multi-modal deep learning fusion model construction method for breast cancer

    CN121439245A