A Thrombolysis Prediction Method Based on Deep Learning and Multimodal Fusion
The deep learning-based multi-modal fusion method addresses the inefficiencies of current brain image segmentation by using a gated parallel Mamba network to capture long-range dependencies, improving stroke lesion segmentation and treatment suitability assessments.
Patent Information
- Application Number
- CN202411560031.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-11-04
AI Technical Summary
The existing automatic segmentation method performs poorly in capturing long-distance dependence between image pixels and high demand for computing resources, resulting in inefficiency in processing large-scale stroke image data, hindering its widespread use in clinical environments.
Using a thrombolysis prediction method based on deep learning and multimodal fusion, we can realize lesion segmentation and thrombolysis prediction of ischemic stroke by constructing a gated parallel Mamba-based stroke automatic segmentation model, combining multimodal MRI images, clinical features and imagingomics features.
It improves the efficiency and accuracy of lesion identification, reduces the consumption of computing resources, provides scientific and automated decision-making support, and improves the quality and safety of medical decision-making.
Smart Images

Figure CN119400394B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of thrombolysis prediction, and particularly to a thrombolysis prediction method based on deep learning and multimodal fusion. Background Art
[0002] Stroke is an extremely common cerebrovascular disease globally and is also one of the main causes of death and disability, characterized by high incidence and high fatality rate. Taking the United States as an example, there are approximately 795,000 newly diagnosed or recurrent stroke patients each year, and the prevalence of adult stroke is as high as 3%.
[0003] Ischemic stroke is the main type of stroke, accounting for about 70 - 80% of all strokes. As one of the effective means for treating ischemic stroke, thrombolytic therapy can dissolve blood clots and restore blood flow to the brain, thereby saving the damaged brain tissue. Although the lesions revealed by magnetic resonance imaging (MRI) in brain imaging examinations are usually regarded as irreversible ischemic injuries (i.e., ischemic core), some patients are still suitable for thrombolytic therapy.
[0004] How to accurately predict which patients belong to the group suitable for thrombolytic therapy has always been a challenge for clinicians. In the current medical technology background, automatic segmentation of lesions is an indispensable basic link for predictive analysis. However, existing automatic segmentation methods (such as models based on convolutional neural networks (CNNs) or Transformers) have obvious deficiencies. These methods perform poorly in capturing long - range dependencies between image pixels and have high demands for computing resources, resulting in heavy computational burdens and efficiency bottlenecks in practical applications. This technical limitation seriously hinders their wide application in clinical settings, especially when dealing with large - scale stroke image data, the problem is particularly prominent.
[0005] To address this gap, the present invention proposes a thrombolysis prediction method based on deep learning and multimodal fusion, so as to provide efficient and accurate decision - making support for whether a patient is suitable for thrombolytic therapy. Summary of the Invention
[0006] The purpose of the present invention is to provide a thrombolysis prediction method based on deep learning and multimodal fusion to solve the problems in the background art.
[0007] To achieve the above purpose, the present invention provides a thrombolysis prediction method based on deep learning and multimodal fusion, which specifically includes the following steps:
[0008] S1. Construct a segmentation model: Collect MRI images of ischemic stroke patients, construct a stroke automatic segmentation model based on gated parallel Mamba, and obtain the segmentation model weights;
[0009] S2. Image registration: Align multiple unlabeled pre- and post-thrombolysis MRI images to be registered into the same spatial reference system for image registration;
[0010] S3. Lesion segmentation: Use the segmentation model weights obtained in S1 to perform inference on the MRI images after S2 image registration, and extract prediction labels;
[0011] S4. Thrombolysis prediction: Combine the prediction labels obtained in S3 to construct a multimodal stroke thrombolysis prediction model. Use the pre-thrombolysis three-dimensional MRI, clinical features, and radiomics features to be predicted as inputs for thrombolysis prediction to determine whether the patient is suitable for thrombolysis treatment.
[0012] Preferably, S1 specifically includes the following steps:
[0013] S11. Data acquisition: Collect a dataset of multimodal MRI images of multiple ischemic stroke patients. Each MRI image includes original images of four modalities, namely diffusion-weighted imaging, apparent diffusion coefficient, T2-weighted imaging, and T2*-weighted imaging, and a label image. The label image includes two label categories, namely the background area and the lesion area;
[0014] S12. Construct a stroke automatic segmentation model based on gated parallel Mamba: Based on the standard encoder-decoder architecture in nnUNet, set a gated parallel Mamba module in the Bottleneck layer;
[0015] S13. Model training and validation: Use the self-configuration method of nnUNet to automatically configure hyperparameters for the dataset, use the dataset as input for model training, and after training is completed, the model generates a set of segmentation model weights.
[0016] Preferably, in S12, the gated parallel Mamba module adjusts the information flow of the Mamba module through a gate node and two parallel branches to further compress the image features; the image features include batch size, number of channels, image height, image width, and image depth;
[0017] The specific process is as follows: After the image features pass through a residual block and then through a Flatten operation, a vector v with a size of (B, H×W×D, C) is obtained, where B, C, H, W, and D are the batch size, number of channels, image height, image width, and image depth in sequence;
[0018] The vector v is represented as (B, L, C). After layer normalization of the vector v, it enters the two branches. In one branch, after passing through the Mamba module and then through a gated unit, a weight ω, ω∈(0,1) is obtained. Multiply the obtained weight by the new vector obtained from the other branch to obtain the vector F:
[0019]
[0020] Wherein, v1 and v2 are the feature vectors input to the two branches respectively, σ(·) represents the process of generating weights through the gating unit, and M(·) is the calculation process of the Mamba module;
[0021] After the vector F is reshaped, it is converted into (B ′ , H ′ ×W ′ ×D ′ , C ′ ), and finally, after transposition, it outputs a vector of size (B ′ , C ′ , H ′ , W ′ , D ′ ).
[0022] Preferably, the S2 specifically includes the following steps:
[0023] S21. To avoid the convergence problem caused by too large deviation of the starting position in subsequent registration, the multi-instance unlabeled MRI images before and after thrombolysis to be registered are aligned by spatial centroid, expressed as:
[0024] T = C f - C m ;
[0025] Wherein, T is the translation distance of the image, C f is the centroid coordinate of the fixed image, and C m is the centroid coordinate of the moving image;
[0026] S22. Perform rigid body transformation on the MRI image processed in S21, expressed as:
[0027] T(x) = R·x + t;
[0028] Wherein, T(x) is the coordinate of the image to be registered after transformation, x is the coordinate of the image to be registered, R is the rotation matrix, and t is the translation vector;
[0029] S23. Perform Mattes mutual information measurement on the MRI image processed in S22, expressed as:
[0030] MI(F, M) = H(F) + H(M) - H(F, M);
[0031] Wherein, MI(F, M) is the mutual information measurement between images F and M, H(F) and H(M) are the entropies of images F and M respectively, and H(F, M) is the joint entropy of images F and M;
[0032] S24. Perform linear interpolation calculation on the MRI image processed in S23, expressed as:
[0033]
[0034] Wherein, I(x, y) is the pixel value of the new coordinate point, a and b are the offsets of the coordinates from the nearest neighbor points, and x0, y0, x1, and y1 are the integer coordinates of the adjacent points of the pixel to be interpolated in the original grid.
[0035] Preferably, step S3 specifically includes the following steps:
[0036] The MRI image after the image registration in S2 is inferred using the segmentation model weights obtained in S1 to obtain a corresponding label image, which is manually corrected by an expert. The volume difference of the lesion before and after thrombolysis is used as the prediction label, expressed as:
[0037]
[0038] Wherein, L(A v ,B v ) is the prediction label, A v is the volume of the lesion after thrombolysis in this case, and B v is the volume of the lesion before thrombolysis in this case.
[0039] Preferably, in S4, the input of the multimodal stroke thrombolysis prediction model includes three groups of features of different modalities, specifically three-dimensional MRI before thrombolysis, clinical features, and radiomics features. Each group of features is subjected to feature extraction through its respective processing module, and finally fused and spliced in the fully connected layer to obtain the final feature vector. The obtained feature vector is processed, and finally the prediction result for classification is output.
[0040] Preferably, the multimodal stroke thrombolysis prediction model includes a three-dimensional MRI processing branch, a clinical feature processing branch, and a radiomics feature processing branch. In the three-dimensional MRI data processing branch, a residual module and an attention module are introduced to extract features. The operation formula of the residual module is as follows:
[0041] y = F(x,{W i ) + x;
[0042] Wherein, x and y are the input and output respectively, and F(x,{W i ) represents the features extracted through the convolutional layer and the batch normalization layer.
[0043] Therefore, a thrombolysis prediction method based on deep learning and multimodal fusion of the present invention has the following beneficial effects:
[0044] (1) The present invention proposes an automatic segmentation model for stroke based on the gated parallel Mamba network. By introducing the Mamba structure and combining the local feature extraction ability of the CNN convolutional layer, it can effectively capture long-range dependencies in images and quickly and accurately segment ischemic stroke lesions, greatly improving the efficiency and accuracy of lesion recognition.
[0045] (2) The algorithm proposed by the present invention achieves a linear time complexity, significantly reducing the consumption of computing resources and breaking through the bottleneck of the existing technology.
[0046] (3) The stroke thrombolysis prediction model based on multi-modal data proposed by the present invention comprehensively utilizes the patient's MRI images, clinical data, and radiomics information to achieve accurate prediction of the recoverability of lesions, thereby providing scientific and automated decision support for whether the patient is suitable for thrombolysis treatment. At the same time, the model is flexibly designed, can properly handle various data combination inputs, effectively cope with the problem of missing modal information, enhance the comprehensiveness and reliability of the prediction, and plays an important role in clinical medicine.
[0047] (4) Compared with the traditional method that only uses the volume of ischemic stroke lesions as the decision basis, this technology effectively breaks through its limitations by comprehensively considering more dimensions of information, provides more scientific and comprehensive decision support for doctors, helps to more accurately judge whether the patient is suitable for thrombolysis treatment, thereby optimizing the formulation of treatment plans and improving the quality and safety of medical decisions.
[0048] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0049] Figure 1 is the overall process framework diagram of the method proposed by the present invention;
[0050] Figure 2 is the architecture diagram of the automatic stroke segmentation model of the embodiment of the present invention;
[0051] Figure 3 is the architecture diagram of the gated parallel Mamba module of the embodiment of the present invention;
[0052] Figure 4 is the architecture diagram of the multi-modal stroke thrombolysis prediction model of the embodiment of the present invention;
[0053] Figure 5 is the architecture diagram of the residual module of the embodiment of the present invention;
[0054] Figure 6 is the architecture diagram of the attention module of the embodiment of the present invention;
[0055] Figure 7Comparison chart of the segmentation results of the stroke automatic segmentation model according to the embodiments of the present invention with the true labels under some test samples;
[0056] Figure 8 ROC curve graph of the prediction model according to the embodiments of the present invention with only three-dimensional MRI data before thrombolysis input;
[0057] Figure 9 ROC curve graph of the prediction model according to the embodiments of the present invention with only clinical feature data input;
[0058] Figure 10 ROC curve graph of the prediction model according to the embodiments of the present invention with only radiomics feature data input;
[0059] Figure 11 ROC curve graph of the prediction model according to the embodiments of the present invention with three-dimensional MRI before thrombolysis and clinical feature data input;
[0060] Figure 12 ROC curve graph of the prediction model according to the embodiments of the present invention with three-dimensional MRI before thrombolysis and radiomics feature data input;
[0061] Figure 13 ROC curve graph of the prediction model according to the embodiments of the present invention with clinical features and radiomics feature data input;
[0062] Figure 14 ROC curve graph of the prediction model according to the embodiments of the present invention with three-dimensional MRI data before thrombolysis, clinical features and radiomics features input. Detailed implementation manners
[0063] The technical solutions of the present invention are further described below through the drawings and embodiments.
[0064] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0065] Embodiment
[0066] As Figure 1 shown, the present invention provides a thrombolysis prediction method based on deep learning and multimodal fusion. For this method, multimodal magnetic resonance images from the Department of Radiology of Tianjin Huanhu Hospital are used to test the influence effect of multimodal data fusion on the model performance. The steps are as follows:
[0067] S1. Construct a segmentation model: Collect MRI images of ischemic stroke patients, construct a stroke automatic segmentation model based on gated parallel Mamba, and obtain the segmentation model weights; specifically:
[0068] S11. Data acquisition: Collect 172 cases of multi-modal MRI images from the Department of Radiology of Huanhu Hospital in Tianjin. Each case includes original images of four modalities, namely diffusion-weighted imaging, apparent diffusion coefficient, T2-weighted imaging, and T2*-weighted imaging, and a label image. The label image includes two label categories: background area and lesion area.
[0069] S12. Construct an automatic stroke segmentation model based on gated parallel Mamba:
[0070] As Figure 2 shown, based on the standard encoder-decoder architecture in nnUNet, a gated parallel Mamba module is set and implemented in the Bottleneck layer. This module can capture long-range dependencies in the image while capturing local features, and at the same time compress the extracted information to increase the expression ability of the model.
[0071] The input of the model is a vector of 4×16×128×128, and the output of the segmentation model is two channels, 0 and 1, corresponding to the two label categories of the label image in S11.
[0072] As Figure 3 shown, the gated parallel Mamba module specifically includes the following steps:
[0073] Adjust the information flow of the Mamba module through a gate node and two parallel branches to further compress the image features and increase the expression ability of the model. The image features B, C, H, W, and D are the batch size, number of channels, image height, image width, and image depth, respectively.
[0074] The specific process is as follows: After passing through a residual block, the image features go through a Flatten operation to obtain a vector v with a size of (B, H×W×D, C), denoted as (B, L, C). The vector v is layer-normalized and then enters two branches. In one branch, after passing through the Mamba module, a weight ω, ω∈(0, 1), is obtained through a gated unit. The obtained weight is multiplied by a new vector obtained from the other branch to get a vector F:
[0075]
[0076] In the formula, v1 and v2 are the feature vectors input to the upper and lower branches respectively, σ(·) represents the process of generating the weight through the gated unit, and M(·) is the calculation process of the Mamba module;
[0077] The vector F is reshaped to (B ′ , H ′ ×W ′ ×D ′, C ′ ), and finally transposed and output as (B ′ , C ′ , H ′ , W ′ , D ′ ) sized vector;
[0078] S13. Model training and validation: The nnUNet self-configuration method is used to automatically configure hyperparameters for the dataset. The Batchsize is set to 8, the Epoch is 800, and the initial learning rate is 0.01. The 5-fold cross-validation method is used for validation;
[0079] As Figure 7 shown, the 2D and 3D segmentation effects of this method on some samples of the dataset used are presented. It can effectively capture the features and information of each modality image on lesions of different sizes and shapes, and achieve accurate segmentation of lesions.
[0080] S2. Image registration: Multiple unlabeled pre- and post-thrombolysis MRI images to be registered are aligned to the same spatial reference system for image registration. Another 690 cases of unlabeled pre- and post-thrombolysis MRI data provided by Huanhu Hospital are registered. Each case includes five modalities: diffusion-weighted imaging, apparent diffusion coefficient, T2-weighted imaging, T2*-weighted imaging, and fluid-attenuated inversion recovery. Specifically:
[0081] S21. To make the approximate positions of MRI images in physical space consistent and avoid convergence problems caused by excessive deviation of the starting position in subsequent registration, spatial centroid alignment is required, expressed as:
[0082] T = C f - C m ;
[0083] In the formula, T is the distance the image needs to be translated, C f is the centroid coordinate of the fixed image, and C m is the centroid coordinate of the moving image.
[0084] S22. Perform a rigid body transformation on the MRI images processed in S21, that is, the MRI images are registered by rotation and translation, without involving scaling or shearing. This transformation can keep the shape of the image structure unchanged and only adjust the position in space, expressed as:
[0085] T(x) = R · x + t;
[0086] In the formula, T(x) is the coordinate of the image to be registered after transformation, x is the coordinate of the image to be registered, R is the rotation matrix, and t is the translation vector;
[0087] S23. Perform Mattes mutual information measurement on the MRI image processed in S22. By maximizing the mutual information, the algorithm can find the transformation that maximizes the correlation after aligning the MRI images, which is expressed as:
[0088] MI(F,M) = H(F) + H(M) - H(F,M);
[0089] In the formula, MI(F,M) is the mutual information measurement between images F and M, H(F) and H(M) are the entropies of images F and M respectively, and H(F,M) is the joint entropy of images F and M.
[0090] S24. Linear interpolation: After S21, S22, and S23, the pixel coordinates of the MRI image may no longer be integer coordinates, and the image needs to be resampled. Therefore, the linear interpolation method is used to perform interpolation calculation on the MRI image processed in S23 to obtain new pixel values, which is expressed as:
[0091]
[0092] In the formula, I(x,y) is the pixel value of the new coordinate point, a and b are the offsets of the coordinate from the nearest neighbor points, and x0, y0, x1, and y1 are the integer coordinates of the adjacent points of the pixel to be interpolated in the original grid.
[0093] S3. Lesion segmentation: Use the segmentation model weights obtained in S1 to perform inference on the MRI image after S2 image registration, and extract the prediction label; specifically:
[0094] Perform inference on 690 cases of data after S2 image registration using the segmentation model weights obtained in S1 to obtain the corresponding label image, which is manually corrected by experts. The difference in the lesion volume before and after thrombolysis in the label image is used as the prediction label, which is expressed as:
[0095]
[0096] In the formula, L(A v ,B v ) is the prediction label, A v is the lesion volume after thrombolysis in this case, B v is the lesion volume before thrombolysis in this case, and the calculation result of 0 represents suitable for thrombolytic treatment, and 1 represents not suitable for thrombolytic treatment.
[0097] S4. Thrombolysis prediction: Combine the prediction label obtained in S3 to construct a multimodal stroke thrombolysis prediction model, and use the pre-thrombolysis three-dimensional MRI, clinical features, and radiomics features to be predicted as inputs for thrombolysis prediction to determine whether the patient is suitable for thrombolytic treatment. Specifically:
[0098] S41. Construct a multimodal stroke thrombolysis prediction model:
[0099] As shown Figure 4 in the figure, the input includes features of three different modalities: three-dimensional MRI before thrombolysis, clinical features, and radiomics features, corresponding to the three-dimensional MRI processing branch, clinical feature processing branch, and radiomics feature processing branch in sequence. Each feature is extracted through its respective processing module, and finally fused and spliced in the fully connected layer to obtain the final feature vector. The obtained feature vector is processed to output the prediction result for classification, and the output size is 2, corresponding to the two categories of whether thrombolytic therapy is suitable in S3.
[0100] The three-dimensional MRI processing branch before thrombolysis is processed through a series of 3D convolutions, residual modules, and attention modules. The input size is (C, H, W, D), where C is the number of channels. In this method, 5 modalities are used, so C is 5. H, W, and D represent the height, width, and depth of the 3D MRI respectively. In the first layer of image feature extraction, after the input image passes through a 3D convolutional layer with a convolutional kernel size of 3×3×3 and a padding of 1, the dimension becomes (32, 128, 128, 128);
[0101] As shown Figures 5 - 6 in the figure, next, the MRI data is processed through two residual blocks. By introducing residual connections, the residual blocks enable the network to retain the features of the lower layers even in the deep layers, thus alleviating the problem of gradient disappearance. Its core idea is to learn the residual function, expressed as:
[0102] y = F(x, {W i}) + x;
[0103] In the formula, x and y are the input and output respectively, and F(x, {W i}) represents the features extracted through the convolutional layer and batch normalization layer;
[0104] The residual blocks perform a series of convolution and batch normalization operations and adjust the number of channels, and the output dimension becomes (128, 128, 128, 128). Then, through the attention module, a weight mask is generated and the feature map at each position is weighted. The output dimension remains the same as the input, that is, (128, 128, 128, 128).
[0105] To reduce the spatial dimension of the feature map, the model uses adaptive average pooling operation to reduce the spatial dimension from (128, 128, 128) to (128, 4, 4, 4). The 3D feature map after pooling is flattened and converted into a one-dimensional vector, and its dimension is reduced to 256 through the fully connected layer.
[0106] The model can also process clinical features and radiomics features, and these two types of data are respectively dimensionally reduced by a linear layer. Clinical data includes gender and age. After the radiomics data is extracted by the Pyradiomics tool, the SelectKbest method is used to screen out the top 10 most valuable features as the input of the prediction model. The clinical data and radiomics information each pass through two fully connected networks, and the final output dimension is 64.
[0107] After processing the image features, clinical features, and radiomics features, the model concatenates these three parts of features and performs splicing on the feature dimension to form a final feature vector, and finally outputs the prediction result for classification.
[0108] This model can flexibly select different data combinations for input according to requirements. It can choose to use only image, clinical, or radiomics information, or combine multiple data sources. The corresponding feature extraction module is selectively executed according to the flag variable specified during input, and the outputs are feature-spliced.
[0109] If only image features are used, the size of the spliced feature vector is 256; if image, clinical, and radiomics features are used simultaneously, the size of the spliced feature vector is 384. The spliced feature vector is then processed through several fully connected networks, and finally the prediction result for classification is output. The final output size is 2, corresponding to the two categories of whether thrombolytic therapy is suitable in S3.
[0110] S42. Model training: The model adopts a 5-fold cross-validation method, uses the Adam optimizer, sets the learning rate to 0.01, the Batch size to 4, and trains for 50 epochs. The cross-entropy loss function is used during the training process, which is expressed as:
[0111]
[0112] In the formula, is the cross-entropy loss calculated from the true label encoding and the output probability, y is the encoding of the true label, is the output probability of the model, and C is the number of prediction categories.
[0113] Such as Figures 8 - 14As shown, in this embodiment, 690 cases of data were explored to study the impact of different combinations of multimodal data on the performance of the prediction model, and the ROC curve obtained after five-fold cross-validation was presented. The results indicate that the model integrating multimodal data shows a significant performance improvement compared to the single-modal model. The accuracy of the model using MRI data is 77%. After further combining MRI data with radiomics data, the accuracy of the model is increased to 90%. When combining MRI data, clinical data, and radiomics data, the prediction accuracy of the model reaches the highest 92%, demonstrating the significant enhancement effect of multimodal data fusion on the model performance.
[0114] By comprehensively using MRI data, clinical data, and radiomics data, it is possible to accurately evaluate whether a patient is suitable for thrombolytic therapy. In addition, when dealing with complex situations such as missing clinical data, the model also demonstrates good prediction stability and ability.
[0115] Therefore, for a thrombolysis prediction method based on deep learning and multimodal fusion of the present invention, an automatic segmentation model of stroke based on the gated parallel Mamba network is proposed. By introducing the Mamba structure and combining the local feature extraction ability of the CNN convolutional layer, it can effectively capture long-range dependencies in the image, quickly and accurately segment ischemic stroke lesions, and greatly improve the efficiency and accuracy of lesion recognition.
[0116] The algorithm proposed in the present invention achieves a linear time complexity, significantly reducing the consumption of computing resources and breaking through the bottleneck of the existing technology.
[0117] The stroke thrombolysis prediction model based on multimodal data proposed in the present invention comprehensively utilizes the patient's MRI images, clinical data, and radiomics information to achieve accurate prediction of the recoverability of lesions, thereby providing scientific and automated decision support for whether a patient is suitable for thrombolytic therapy. At the same time, the model is flexibly designed, can properly handle various combinations of data inputs, effectively cope with the problem of missing modal information, enhance the comprehensiveness and reliability of prediction, and plays an important role in clinical medicine.
[0118] Compared with the traditional method that only uses the volume of ischemic stroke lesions as the decision basis, this technology effectively breaks through its limitations by considering more dimensions of information, provides more scientific and comprehensive decision support for doctors, helps to more accurately judge whether a patient is suitable for thrombolytic therapy, thereby optimizing the formulation of treatment plans and improving the quality and safety of medical decisions.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A thrombolysis prediction method based on deep learning and multimodal fusion, characterized in that Specifically, it includes the following steps: S1. Build a segmentation model: Collect MRI images of ischemic stroke patients, build an automatic stroke segmentation model based on gated parallel Mamba, and obtain the segmentation model weights. In S1, the gated parallel Mamba module adjusts the information flow of the Mamba module through a gate node and two parallel branches to further compress the image features. The image features include batch size, number of channels, image height, image width, and image depth. The specific process is as follows: After passing through a residual block, the image features go through a Flatten operation to obtain a vector with a size of wherein wherein, , , , , are in sequence the batch size, the number of channels, the image height, the image width, and the image depth; vector , the vector is layer-normalized and enters two branches. In one branch, after passing through the Mamba module, the weights are obtained through the gating unit. The obtained weights are multiplied by the new vector obtained from the other branch to get the vector : ; wherein, , are the feature vectors input to the two branches respectively, represents the process of generating weights through the gating unit, is the calculation process of the Mamba module; Vector After Reshape, it is converted to , and finally output after transposition vector of size; S2. Image registration: Align multiple unlabeled pre- and post-thrombolysis MRI images to be registered into the same spatial reference system for image registration. S3. Lesion segmentation: Use the segmentation model weights obtained in S1 to perform inference on the MRI images after S2 image registration, and extract prediction labels. S4. Thrombolysis prediction: Combine the prediction labels obtained in S3 to build a multimodal stroke thrombolysis prediction model. Use the pre-thrombolysis three-dimensional MRI, clinical features, and radiomics features to be predicted as inputs for thrombolysis prediction to determine whether the patient is suitable for thrombolysis treatment.
2. The thrombolysis prediction method based on deep learning and multimodal fusion according to claim 1, wherein S1 specifically includes the following steps: S11. Data acquisition: Collect a dataset of multi-modal MRI images of multiple ischemic stroke patients. Each MRI image includes original images of four modalities, namely diffusion-weighted imaging, apparent diffusion coefficient, T2-weighted imaging, and weighted imaging, and a label image. The label image includes two label categories, namely the background area and the lesion area; S12. Build an automatic stroke segmentation model based on gated parallel Mamba: Based on the standard encoder-decoder architecture in nnUNet, set the gated parallel Mamba module in the Bottleneck layer. S13. Model training and validation: Use the nnUNet self-configuration method to automatically configure hyperparameters for the dataset, use the dataset as input for model training, and after training, the model generates a set of segmentation model weights.
3. The thrombolysis prediction method based on deep learning and multi-modal fusion according to claim 2, wherein, S2 specifically includes the following steps: S21. To avoid convergence problems caused by excessive deviation of the starting position in subsequent registration, perform spatial centroid alignment on multiple unlabeled pre- and post-thrombolysis MRI images to be registered, expressed as: ; In the formula, is the distance by which the image needs to be translated, is the centroid coordinate of the fixed image, is the centroid coordinate of the moving image; S22. Perform rigid body transformation on the MRI images processed in S21, expressed as: ; In the formula, is the coordinate of the image to be registered after transformation, is the coordinate of the image to be registered, is the rotation matrix, is the translation vector; S23. Perform Mattes mutual information measurement on the MRI images processed in S22, expressed as: ; wherein, is the mutual information measure between the images and ; and are the entropies of the images and respectively, and is the joint entropy of the images and ; S24. Perform linear interpolation calculation on the MRI images processed in S23, expressed as: ; wherein, is the pixel value of the new coordinate point, and are the offsets of the coordinates from the nearest neighbor points, 、 、 and are the integer coordinates of the adjacent points of the pixel to be interpolated in the original grid.
4. The thrombolysis prediction method based on deep learning and multimodal fusion according to claim 3, wherein S3 specifically includes the following steps: Use the segmentation model weights obtained in S1 to perform inference on the MRI images after S2 image registration to obtain the corresponding label image, which is manually corrected by experts. Use the volume difference of the lesions before and after thrombolysis as the prediction label, expressed as: ; In the formula, is the predicted label, is the volume of the lesion after thrombolysis for the corresponding case, is the volume of the lesion before thrombolysis for the corresponding case.
5. A thrombolysis prediction method based on deep learning and multimodal fusion according to claim 4, characterized in that: In S4, the inputs of the multimodal stroke thrombolysis prediction model include three groups of features of different modalities, specifically pre-thrombolysis three-dimensional MRI, clinical features, and radiomics features. Each group of features is subjected to feature extraction through its respective processing module before input, and finally fused and spliced in the fully connected layer to obtain the final feature vector. The obtained feature vector is processed, and finally the prediction result for classification is output.
6. The thrombolysis prediction method based on deep learning and multimodal fusion according to claim 5, characterized in that: The multimodal stroke thrombolysis prediction model includes a three-dimensional MRI processing branch, a clinical feature processing branch, and a radiomics feature processing branch. In the three-dimensional MRI data processing branch, a residual module and an attention module are introduced to extract features. The operation formula of the residual module is as follows: ; In the formula, and are the input and output respectively, represents the features extracted by the convolutional layer and the batch normalization layer.
Citation Information
Patent Citations
Breast cancer neoadjuvant chemotherapy curative effect prediction device based on multi-feature fusion
CN114974575A
Self-adaptive efficient image super-resolution method
CN118350995A