Medical image processing method based on bimodal magnetic resonance image feature fusion
By using a three-dimensional convolutional neural network based on R3D-18 and a cross-attention mechanism, the problem of multimodal MRI data fusion was solved, achieving more efficient image classification accuracy and stability, especially in the case of imbalanced data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
- Filing Date
- 2026-01-09
- Publication Date
- 2026-05-12
AI Technical Summary
Existing medical image processing methods struggle to effectively integrate multimodal MRI data, leading to difficulties in aligning cross-modal features, which affects the joint modeling of three-dimensional volume data. Furthermore, the models are prone to undergeneralization or decreased sensitivity to minority class samples during training.
Feature extraction is performed using a 3D convolutional neural network based on R3D-18, and a cross-attention mechanism is introduced for modal feature fusion. The correlation between different modal features is modeled at the voxel level through cross-attention correlation interaction modeling, and the model is optimized by combining Focal Loss and LDAM Loss.
It improves the discriminativeness and robustness of multimodal MRI data, enhances the accuracy and stability of image classification, and performs better, especially in cases of data class imbalance.
Smart Images

Figure CN122023282A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and medical image processing technology, and in particular relates to a medical image processing method based on the fusion of dual-modal magnetic resonance imaging features. Background Technology
[0002] With the development of medical imaging technology, magnetic resonance imaging (MRI) has been widely used in clinical image analysis due to its high resolution for soft tissues. For some genetic or developmental diseases, automatic segmentation, classification, and phenotypic quantification of target regions based on MRI images can help improve image interpretation efficiency and provide a reference for subsequent clinical analysis. Taking Klinefelter syndrome-related images as an example, its image representation often exhibits characteristics such as large individual differences and hidden features. Relying on a single imaging sequence is insufficient to fully characterize the structural and functional information of the target region. Therefore, in actual acquisition, different sequences / modalities (such as T2-weighted imaging, diffusion-weighted imaging, etc.) are often combined for comprehensive observation.
[0003] However, bimodal / multi-sequence MRI data differ in spatial resolution, imaging scale, and signal intensity distribution, making direct alignment of cross-modal features difficult and posing a challenge to the joint modeling of 3D volume data. Existing medical image processing methods often employ strategies such as feature stitching, weighted summation, or post-decision fusion to achieve multimodal fusion, but these often only superimpose information at the channel level, lacking explicit characterization of fine-grained spatial correspondences and correlations between different modalities. Furthermore, some methods use 2D slices as processing units, failing to fully utilize the spatial context information of 3D volume data, thus affecting the discriminativeness and stability of the fused features. In addition, real medical image samples often suffer from limited quantity, uneven class distribution, and acquisition noise, making models prone to undergeneralization or decreased sensitivity to minority class samples during training.
[0004] Therefore, medical image processing methods for bimodal MRI three-dimensional volume data can effectively model the correlation interaction of different modal features and generate more discriminative fusion representations while ensuring cross-modal spatial consistency, thereby improving the robustness and accuracy of three-dimensional medical image analysis tasks. Summary of the Invention
[0005] The purpose of this invention is to address the problem that existing multi-sequence magnetic resonance image fusion and classification methods are unable to fully model the complex relationships between different sequences in medical image processing, thus limiting the accuracy of target detection or classification. This invention proposes a medical image processing method based on dual-modal magnetic resonance image feature fusion to improve the accuracy and robustness of detection results.
[0006] This invention is achieved through the following technical solution: A medical image processing method based on dual-modal magnetic resonance imaging feature fusion includes the following steps: S1. Acquire at least two magnetic resonance imaging sequence image data of the target area as sequence inputs for two modes, and preprocess the image data to obtain standardized three-dimensional volume data. S2. Extract features from the three-dimensional volume data corresponding to different magnetic resonance imaging sequences to obtain the corresponding sequence features; S3. Map the 3D features of different modalities to a feature sequence of a unified dimension, and introduce a cross-attention correlation interaction modeling method to model the correlation between features of different modalities at the voxel level, so as to obtain the fused multimodal feature representation. S4. Based on the multimodal feature representation, perform feature aggregation and classification analysis, and output the probability distribution of the target region in different categories.
[0007] A storage device that stores instructions and data for implementing a medical image processing method based on dual-modal magnetic resonance imaging feature fusion.
[0008] A medical image processing device based on dual-modal magnetic resonance imaging feature fusion includes: a processor and a storage device; the processor loads and executes instructions and data in the storage device to implement a medical image processing method based on dual-modal magnetic resonance imaging feature fusion.
[0009] The present invention has the following beneficial effects: A multimodal deep learning framework is proposed that directly utilizes complete DWI and T2 sequences without explicit ROI segmentation. R3D-18 is used as the backbone network and modified for single-channel medical image input to preserve complete 3D volumetric information and avoid reliance on manual annotation. Furthermore, a mid-term feature fusion mechanism is designed, combined with a bidirectional cross-attention mechanism to enhance intermodal complementarity, thereby effectively modeling cross-modal interaction information while maintaining modal feature independence. This framework can simultaneously learn modality-specific and joint representations, thus improving the model's robustness on heterogeneous data.
[0010] Experimental results show that the proposed framework outperforms the single-modal baseline model in all metrics. In particular, multimodal fusion demonstrates better performance in AUC and result stability, further validating the effectiveness of the complementary information from DWI and T2. Furthermore, the introduction of Focal Loss and LDAM Loss effectively mitigates the impact of data class imbalance, making the model's classification performance more reliable and thus improving the overall accuracy of image classification. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of the medical image processing method based on dual-modal magnetic resonance imaging feature fusion according to the present invention; Figure 2 This is a schematic diagram of the training framework for unimodal and bimodal tasks; Figure 3 The results are based on the ROC curves of the fold-by-fold cross-validation of the R3D-18 deep learning model on DWI data. Figure 4 The results are based on the ROC curve of the fold-by-fold cross-validation of the R3D-18 deep learning model on T2 data. Figure 5 The results are based on the R3D-18 deep learning model and the fold-by-fold cross-validation ROC curves on dual-modal magnetic resonance data. Figure 6 This is a schematic diagram of the hardware device of the present invention. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Example 1 Please see Figure 1 As shown, this invention is a medical image processing method based on dual-modal magnetic resonance imaging feature fusion, comprising the following steps: S1. Acquire at least two magnetic resonance imaging sequence image data of the target area as sequence inputs for two modes, and preprocess the image data to obtain standardized three-dimensional volume data. It should be noted that the two magnetic resonance imaging sequences used in this invention are DWI and T2 sequences.
[0015] All subjects were scanned on a 3.0 T MRI scanner (MAGNETOM Skyra, Siemens Healthcare, Erlangen, Germany) using an 18-channel body coil and a 32-channel spine coil. The parameters were as follows: (1) Axial T2-weighted turbo spin echo sequence, TE / TR = 104 / 6500 ms, FOV = 180 x 180 mm², matrix = 384 x 320, slice thickness = 3 mm; voxel size = 0.5 x 0.5 x 3.0 mm; (2) Axial DWI, using single-shot echo planar imaging, TE / TR = 78 / 5300 ms, FOV = 220 x 176 mm², matrix = 90 x 90, slice thickness = 3 mm, voxel size = 1.2 x 1.2 x 3.0 mm, b value (number of excitations) = 0 (1), 50 (2), 100 (3), 500 (3), 1000 (5), 500 (5), 2000 (7), 2500 (7), 3000 (7) s / mm², total scan time = 7.5 minutes.
[0016] Regarding the target region, in the present invention, obstructive azoospermia (OA) and non-obstructive azoospermia (NOA) were taken as examples, and the regions of interest (ROIs) in the testis were selected. The ROI regions were carefully delineated by a board-certified radiologist with 11 years of professional experience in abdominal MRI interpretation. These ROIs were precisely outlined on the automatically generated apparent diffusion coefficient (ADC) parameter maps and the corresponding axial T2-weighted imaging sequences, and these images were acquired by a clinical 3.0-tesla MRI scanner. Using the ITK-SNAP software platform, the ROI delineation was performed by manually and precisely tracing layer by layer along the inner edge of the testicular tunica albuginea (testicular capsule).
[0017] It should be noted that step S1 is specifically as follows: S11. Unify the depth dimension D of the image sequence to a preset value S: When D > S, perform central cropping, and when D < S, perform symmetric zero-padding to unify the spatial scale; Specifically, considering that the number of slices D (i.e., dimension) in the T2 sequence varies in different cases, the present invention unified the depth of the three-dimensional volume data to 22 slices: When D > 22, perform central cropping, and when D < 22, perform symmetric zero-padding before and after to ensure that all samples can be stacked into a consistent volume dimension .
[0018] S12. Calculate the global mean μ and standard deviation σ of the non-zero voxels in the image sequence, and perform a normalization transformation on each volume data vol: in, This represents the data after standardization transformation.
[0019] S2. Extract features from the three-dimensional volume data corresponding to different magnetic resonance imaging sequences to obtain the corresponding sequence features; It should be noted that the feature extraction network described in step S2 adopts a three-dimensional convolutional neural network based on the R3D-18 structure, and has been improved to adapt to the characteristics of medical image data, specifically including: The number of input channels for the 3D convolution in the input layer is set to 1 to adapt to single-channel medical magnetic resonance imaging data; the weights of the model pre-trained on the ImageNet dataset are used as initialization parameters, and the weights are initialized by taking the mean value over the channel dimension; the global pooling layer and classification layer in the original network are removed, and the 3D convolutional features up to layer 4 are retained as the output of the feature extraction network.
[0020] Based on this, a new classification module is constructed, in which Dropout layer and fully connected layer are applied sequentially to the convolutional features. The inactivation probability of Dropout is set to 0.5, and the output dimension of the fully connected layer is set to 2 to output the binary classification result. The corresponding classification probability distribution is obtained through the Softmax layer.
[0021] Furthermore, during model training, the 3D convolutional neural network undergoes pre-training weight initialization before use, and data augmentation processing is applied to the standardized 3D volume data during the training phase to expand the training sample set. This data augmentation processing includes spatial augmentation methods such as random flipping and random discrete rotation, as well as intensity augmentation methods such as affine transformation and Gaussian noise perturbation.
[0022] In summary, during the network pre-training of this invention, a total of 30 individuals with Klinefelter syndrome (mean age: 29.4 years; age range: 24–39 years) and 112 individuals without Klinefelter syndrome (mean age: 31.0 years; age range: 23–40 years) were collected. Information on NOA individuals is as follows: 47, XXY karyotype (30 cases); non-Klinefelter syndrome: Y chromosome microdeletion affecting the AZFc or AZFb+c region (15 cases), history of cryptorchidism (4 cases), mumps and / orchitis (23 cases), history of radiotherapy (1 case), and idiopathic NOA (70 cases).
[0023] Based on the aforementioned data characteristics, pelvic MRI images of 142 individuals were further collected. Each individual included both DWI and T2 sequences. The DWI sequence consisted of 22 slices covering the entire pelvic volume, with each slice having a resolution of 180×144. The T2 sequence consisted of approximately 22 slices of varying numbers, with a corresponding resolution of 384×384. Therefore, each individual possessed three-dimensional volumetric data in both modalities. Among all individuals, 30 had Klinefelter syndrome, and the remainder did not. Given the imbalanced data distribution, 5-fold cross-validation was used for model training and performance evaluation. In each fold, the dataset was divided into a training set (n=90), a validation set (n=23), and a test set (n=29). The proportions of individuals with Klinefelter syndrome in the training, validation, and test sets remained approximately 60%, 20%, and 20%, respectively. During modeling, individuals with Klinefelter syndrome were labeled as 1, and those without Klinefelter syndrome were labeled as 0. The training set is used to train the model, the validation set is used to fine-tune the model and classifier, and the test set is used only for final performance evaluation.
[0024] It should be noted that during the training phase, a hybrid loss function consisting of a weighted sum of Focal Loss and LDAM Loss is used for optimization.
[0025] The hybrid loss function is defined as follows:
[0026] in, For Focal Loss, For LDAM Loss, and These are preset weighting coefficients; The Focal Loss is defined as:
[0027] in, The weight of category y, This is the focusing parameter, and its value is 2.5. This represents the model's predicted probability for the true class y. The LDAM Loss introduces a class-adaptive interval to the logits vector. Processing:
[0028] in, Adjusted according to the number of samples per category: , Let be the number of samples of class y in the training set. This is a hyperparameter used to control the overall margin range, where k represents the class index in the class samples; j Indicates excluding categories y In addition, for any category index in the category samples, the LDAM Loss is ultimately defined as the temperature-scaled cross-entropy loss, as follows:
[0029] in It is the category weight. It is cross-entropy loss. s This is the temperature scaling factor.
[0030] S3. Perform interactive processing and fusion on the different sequence features to obtain a fused feature representation; Specifically, step S3 is as follows: S31. Perform spatial resolution alignment processing on the first modality feature sequence and the second modality feature sequence; To achieve spatial consistency of features from different sequences, this invention employs trilinear interpolation to resample the three-dimensional features of one modality, adjusting its spatial dimensions to match those of another modality. Specifically:
[0031] S32. Flatten the aligned 3D features along the spatial dimensions to form a feature sequence; Subsequently, the three-dimensional features are unfolded into a sequence along the spatial dimension, denoted as... The flattening features corresponding to DWI and T2 are respectively .
[0032] S33. Input the feature sequence into the bidirectional cross-attention unit to calculate the enhanced first modality feature sequence and the second modality feature sequence respectively; Specifically, the channel is first projected and divided into blocks:
[0033] in Next, bidirectional cross-attention is applied to the dimensionality-reduced space to obtain cross-modal enhanced sequence representations. .
[0034] S34. Restore the enhanced first modal feature sequence and second modal feature sequence to their corresponding spatial feature forms, and concatenate them in the channel dimension to obtain the concatenated features; S35. The spliced features are subjected to channel compression processing through 1×1×1 three-dimensional convolution to obtain the fused dual-modal feature vector.
[0035] The enhanced features are then restored to their spatial form and concatenated along the channel dimension. They are then compressed using a 1×1 3D convolution (BN+ReLU) to obtain the fused features. .
[0036] S4. Perform global average pooling on the bimodal feature vector, and process it through a Dropout layer and a fully connected layer to output a classification probability value; the classification probability value represents the probability that the target region belongs to a category.
[0037] Finally, the fused features are processed by three-dimensional global average pooling, then Dropout (0.5) and a fully connected layer to output two logits, and finally the classification probability is obtained by Softmax.
[0038] Let the predicted logit vector of the input sample (x,y) be... Its softmax probability is
[0039] in This represents the total number of categories. In this invention, it refers to the binary classification of Klinefelter syndrome, so K=2. These are real category labels. This indicates that the prediction is the true category. The probability of.
[0040] Example 2 All experiments in this invention were implemented using the PyTorch framework and run on an NVIDIA GeForce RTX 3090 GPU. The maximum number of training epochs was set to 100, and an early stopping mechanism was introduced to terminate training prematurely if the validation set performance showed no improvement over a long period, thus preventing overfitting. The batch size was set to 4, and the initial learning rate was 5 × 10⁻⁶. -5 The weight decay is 5 × 10⁻⁶. -4 All experiments used a fixed random seed of 42 to ensure the reproducibility of results. The hyperparameters of the loss function were set as follows: =2.5, temperature scaling factor =30, Focal Loss weights = 0.5, LDAM Loss weight =0.6.
[0041] The input of this invention is volume data. This included DWI and T2 sequences. A classification model was built based on a 3D convolutional neural network ResNet-18 (R3D-18), and experiments were conducted on single-modal (DWI or T2) and dual-modal (DWI + T2) tasks. The training framework is as follows: Figure 2 As shown.
[0042] In single-modal experiments, the model receives input from only a single modality to evaluate its independent predictive performance.
[0043] The performance of the proposed process is evaluated using AUC and ROC curves. ROC curves are a commonly used visualization method in binary classification tasks, visually demonstrating the classifier's performance by plotting the true positive rate and false positive rate at different thresholds; AUC, on the other hand, summarizes the overall model performance with a single numerical value. Its significant advantage lies in maintaining high robustness even under imbalanced sample distribution.
[0044] Due to the relatively limited sample size of Klinefelter syndrome patients in this invention, AUC was considered the most suitable performance evaluation metric. Furthermore, the ROC curve can reflect the trade-off between specificity and sensitivity of the classifier. For both DWI and T2 unimodal inputs and their fused input, ROC curves were plotted for each fold, and the corresponding AUC values were calculated. The final results are summarized and reported as the mean ± standard deviation of the AUC for each fold and its 95% confidence interval (95% CI).
[0045] like Figure 3 As shown, the ROC curve results of the model in 5-fold cross-validation are displayed when only DWI data is used as input. The curves in each fold show high classification performance, with AUC values exceeding 0.87, indicating that DWI-based features can provide effective support for the prediction of Klinefelter syndrome.
[0046] like Figure 4 As shown, when only the T2 sequence is used as input, the ROC curve results of the R3D-18-based deep learning model in five-fold cross-validation are displayed below. It can be seen that the ROC curves at different folds are all above the diagonal (the baseline of the random classifier) in terms of overall trend, indicating that the model has a certain discriminative ability on the T2 data, but its overall performance exhibits some fluctuation.
[0047] like Figure 5As shown in the figure, when DWI and T2 sequences are input simultaneously and feature fusion is performed in the intermediate stage, the ROC curve results of the R3D-18-based deep learning model in five-fold cross-validation are as follows. The ROC curves at each fold are significantly better than the random classification baseline, indicating that multimodal feature fusion can effectively improve the model's discriminative ability. The overall performance is stable and generally higher than the single-modal experimental results.
[0048] This invention further analyzed the performance of the model under different modal inputs on the five-fold cross-validation and aggregated test set, and the results are shown in Table 1. As can be seen from the table, the dual-modal input (DWI+T2) achieved the highest AUC value on both the cross-validation and aggregated five-fold test set, outperforming single-modal DWI or T2. This indicates that multimodal feature fusion can effectively improve the model's discriminative performance and stability.
[0049] Table 1. Model performance on five-fold cross-validation and pooled test set under different modal inputs.
[0050] Example 3 Please see Figure 6 , Figure 6 This is a schematic diagram of the hardware device in operation according to an embodiment of the present invention. The hardware device specifically includes: a medical image processing device 401 based on dual-modal magnetic resonance image feature fusion, a processor 402, and a storage device 403.
[0051] A medical image processing device 401 based on dual-modal magnetic resonance image feature fusion: The medical image processing device 401 based on dual-modal magnetic resonance image feature fusion implements the medical image processing method based on dual-modal magnetic resonance image feature fusion.
[0052] Processor 402: The processor 402 loads and executes the instructions and data in the storage device 403 to implement the medical image processing method based on dual-modal magnetic resonance imaging feature fusion.
[0053] Storage device 403: The storage device 403 stores instructions and data; the storage device 403 is used to implement the medical image processing method based on dual-modal magnetic resonance imaging feature fusion.
[0054] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0055] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A medical image processing method based on dual-modal magnetic resonance imaging feature fusion, characterized in that: Includes the following steps: S1. Acquire two magnetic resonance imaging sequence image data of the target area as sequence inputs for two modes, and preprocess the image data to obtain standardized three-dimensional volume data. S2. Extract features from the three-dimensional volume data corresponding to different modal magnetic resonance imaging sequences to obtain the corresponding sequence features; S3. Map the 3D features of different modalities to a feature sequence of a unified dimension, and introduce a cross-attention correlation interaction modeling method to model the correlation between features of different modalities at the voxel level, so as to obtain the fused multimodal feature representation. S4. Based on the multimodal feature representation, perform feature aggregation and classification analysis, and output the probability distribution of the target region in different categories.
2. The medical image processing method based on dual-modal magnetic resonance imaging feature fusion according to claim 1, characterized in that, Step S1 is as follows: S11. Unify the depth dimension D of the image sequence to a preset value S: when D > S, perform center cropping on the depth dimension; when D < S, perform symmetrical zero padding on the depth dimension to achieve spatial scale unification. S12. Select non-zero voxels in the image sequence, calculate their global mean μ and standard deviation σ, and perform a standardization transformation on each volume data vol: in, This represents the volume data after standardization transformation.
3. The medical image processing method based on dual-modal magnetic resonance imaging feature fusion according to claim 2, characterized in that, The feature extraction network in step S2 is a three-dimensional convolutional neural network constructed based on the R3D-18 backbone structure; the three-dimensional convolutional neural network performs structural adaptation of the input layer and channel configuration for single-channel medical image volume data, and initializes the network parameters using pre-trained model weights; the three-dimensional convolutional neural network outputs the feature map of the preset intermediate layer as the sequence feature of the corresponding modality during the forward inference process.
4. The medical image processing method based on dual-modal magnetic resonance imaging feature fusion according to claim 3, characterized in that, Step S3 is as follows: S31. Spatial resolution alignment is performed on the sequence features of the first mode and the sequence features of the second mode to make them have the same size in spatial dimension; S32. Flatten the aligned first modality sequence features and second modality sequence features along the spatial dimension to form a feature sequence; S33. Input the feature sequence into the bidirectional cross-attention unit to calculate the enhanced first modality feature sequence and the second modality feature sequence respectively; S34. Reconstruct the enhanced first modal feature sequence and the second modal feature sequence into corresponding spatial feature forms, and splice them in the channel dimension to obtain spliced features; S35. The spliced features are subjected to channel compression processing through 1×1×1 three-dimensional convolution to obtain the fused dual-modal feature vector.
5. A medical image processing method based on dual-modal magnetic resonance imaging feature fusion according to claim 4, characterized in that, The 3D convolutional neural network is initialized using pre-trained model weights during model construction. During model training, data augmentation is performed on the standardized 3D volume data to expand the training sample set. The data augmentation includes spatial transformation augmentation and intensity augmentation, wherein the spatial transformation augmentation includes random flipping and random rotation operations, and the intensity augmentation includes affine transformation and noise perturbation.
6. The medical image processing method based on dual-modal magnetic resonance imaging feature fusion as described in claim 5, characterized in that: During model training, a loss function designed for class imbalance is used to optimize the model. This loss function includes a loss term to enhance the learning ability of minority class samples and a loss term to adjust the discriminative boundary between classes. The hybrid loss function is defined as follows: in, For Focal Loss, For LDAM Loss, and These are preset weighting coefficients; The Focal Loss is defined as: in, The weight of category y, To focus parameters, This represents the model's predicted probability for the true class y. The LDAM Loss introduces a class-adaptive interval to the logits vector. Processing: in, Adjusted according to the number of samples per category: , Let be the number of samples of class y in the training set. This is a hyperparameter used to control the overall margin range, where k represents the class index in the class samples; j Indicates excluding categories y In addition, for any category index in the category samples, the LDAM Loss is ultimately defined as the temperature-scaled cross-entropy loss, as follows: in It is the category weight. It is cross-entropy loss. s This is the temperature scaling factor.
7. A storage device, characterized in that, The storage device stores program instructions and data. When the program instructions are executed by the processor, they are used to implement the medical image processing method based on dual-modal magnetic resonance imaging feature fusion as described in any one of claims 1 to 6.
8. A medical image processing device based on dual-modal magnetic resonance imaging feature fusion, characterized in that, include: A processor and a storage device; the storage device stores program instructions and data, and the processor is used to load and execute the program instructions to implement the medical image processing method based on dual-modal magnetic resonance imaging feature fusion as described in any one of claims 1 to 6.