Hepatocellular carcinoma microvessel invasion prediction system based on preoperative multi-mode 3D MRI and application of hepatocellular carcinoma microvessel invasion prediction system

Through the deep learning prediction system of multimodal 3D MRI imaging, the prediction problem of preoperative microvascular invasion of hepatocellular carcinoma is solved, the prediction accuracy and reliability are improved, personalized treatment plans are helped to reduce the recurrence rate of hepatocellular carcinoma.

CN120259193APending Publication Date: 2025-07-04XUZHOU MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510266280.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

There is a lack of effective preoperative methods in the prior art to predict microvascular invasion of hepatocellular carcinoma, resulting in high recurrence and metastasis rates after surgery and the inability to provide active intervention in the early stage.

Method used

A hepatocellular carcinoma microvascular invasion prediction system based on multimodal 3D MRI imaging is adopted, including data acquisition, lesion outline, lesion area cropping, deep learning network feature extraction, mixed feature fusion and weight allocation and classification modules. Local and global features are extracted using CNN and Transformer branches, and feature weighting and loss supervision are performed through multi-layer perceptron and SoftMax activation functions.

Benefits of technology

It has achieved non-invasive and efficient prediction of microvascular invasion of hepatocellular carcinoma before surgery, improved the survival rate of patients and reduced the recurrence rate, and provided a basis for personalized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_5
    Figure SMS_5
  • Figure SMS_6
    Figure SMS_6
Patent Text Reader

Abstract

The invention discloses a hepatocellular carcinoma microvessel invasion prediction system based on a preoperative multi-mode 3D MRI image and application thereof, and the system comprises a data collection module, an image preprocessing module, a lesion sketching module, a lesion area cutting module, a deep learning network feature extraction module, a mixed feature fusion module, and a weight distribution and classification module. The weight distribution module comprises a feature weighted fusion module and a loss supervision module; the deep network feature extraction module extracts local features and global features of a 3D image by using a CNN branch and a Transform branch, and the features of the two branches are fused by using the mixed feature fusion module. According to the method, a new scheme is provided for non-invasive and efficient preoperative prediction of hepatocellular carcinoma microvascular invasion, a basis can be provided for preoperative treatment mode selection and postoperative evaluation of a hepatocellular carcinoma patient, and the method is high in prediction precision and high in reliability and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and particularly to a prediction system for microvascular invasion of hepatocellular carcinoma based on preoperative multimodal 3D MRI images and its application. Background Art

[0002] Hepatocellular carcinoma (HCC) is a primary malignant tumor of the liver, with high malignancy and relatively poor prognosis. Despite the continuous improvement of the comprehensive treatment plan combining surgical treatment, local ablation and targeted drugs, the mortality rate of liver cancer remains high. On the one hand, due to the lack of obvious early symptoms of liver cancer, most patients are already in the advanced stage at the time of diagnosis, and the possibility of radical surgery or transplantation is almost lost; on the other hand, the postoperative recurrence rate and metastasis rate of patients are extremely high. Microvascular invasion (MVI) is one of the most important factors for recurrence after surgical treatment of liver cancer, indicating the presence of highly invasive biological behavior of the tumor. MVI is a microscopic histopathological finding, which refers to the presence of cancer cell nests in the vascular lumen lined by endothelial cells under the microscope, mainly in the portal vein branches (including intra-capsular blood vessels). Pathological grading method: M0: no MVI is found; M1 (low-risk group): ≤ 5 MVIs, and occur in the liver tissue adjacent to the cancer; M2 (high-risk group): > 5 MVIs, or MVI occurs in the liver tissue far from the cancer. In the present invention, M0 is classified into the negative group of microvascular invasion, and M1 and M2 grades are classified into the positive group of microvascular invasion. At present, there is no effective means for preoperative evaluation of MVI invasion in HCC clinically, and only pathological diagnosis can be obtained by taking specimens after surgery. If the presence of MVI can be predicted before surgery and active intervention can be given at an early stage, the prognosis of patients can be significantly improved. Preoperative prediction of MVI is of great significance for improving the long-term survival rate of HCC patients and reducing the postoperative recurrence rate. Therefore, there is an urgent need in this field for a solution that can accurately and meticulously achieve preoperative prediction of MVI. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a prediction system for microvascular invasion of hepatocellular carcinoma based on preoperative multimodal MRI and its application in view of the above-mentioned deficiencies in the prior art.

[0004] To solve the above technical problem, the technical solution adopted by the present invention is: In the first aspect of the present invention, there is provided a prediction system for microvascular invasion of hepatocellular carcinoma based on preoperative multimodal 3D MRI, characterized in that the system includes:

[0005] A data acquisition module for acquiring multimodal MRI images, which include AP images, DP images, DWI images and T2WI-FS images;

[0006] A lesion delineation module for delineating lesions in AP images, DP images, DWI images, and T2WI-FS images;

[0007] A lesion area cropping module for cropping the delineated lesion areas in AP images, DP images, DWI images, and T2WI-FS images in 3D form;

[0008] A deep learning network feature extraction module that extracts local features and global features from the cropped images of each modality: AP images, DP images, DWI images, and T2WI-FS images using CNN branches and Transformer branches respectively;

[0009] A hybrid feature fusion module for fusing the local features and global features obtained from the CNN branches and Transformer branches within each modality to obtain the deep learning hybrid features of each modality;

[0010] A weight assignment and classification module, which includes a multi-modal feature weighted fusion classification module and a deep network loss supervision module;

[0011] The multi-modal weighted fusion classification module uses a multi-layer perceptron to generate the weights corresponding to the features of each modality, multiplies element-wise to obtain the weighted features, and then adds the weighted features of each modality as the fusion feature for final classification; then, the deep network loss supervision module generates the loss supervision weights of each modality for dynamically adjusting the feature extraction process of each modality.

[0012] Preferably, the multi-modal feature weighted fusion classification module uses a multi-layer perceptron to perform dimensional mapping on the high-dimensional features formed by concatenating the hybrid features of each modality, normalizes using the Sigmoid activation function, generates the corresponding feature weights of each modality, multiplies element-wise with the input of the hybrid features of each modality to form the weighted features of each modality, and adds the weighted features of each modality as the fusion feature to participate in the final classification;

[0013] The deep network loss supervision module sums the weighted features of each modality obtained in the multi-modal feature weighted fusion module respectively, and then passes them into the SoftMax activation function to obtain the loss supervision weights of each modality for adjusting the features of each modality obtained by feature extraction.

[0014] Preferably, the method for the deep learning network feature extraction module to extract features is specifically as follows:

[0015] 1) In the CNN branch, the extraction process of local features has a total of 4 stages, and there are specific local feature blocks in each stage. The calculation method is as follows:

[0016]

[0017] Among them, L i represents the output of the local feature block in the i-th stage, where i ∈ {1, 2, 3, 4}, and L i-1 represents the output of the local feature block in the (i - 1)-th stage. represents a 3D depth convolution with a convolution kernel size of 3, and C k=1 represents a 3D standard convolution with a convolution kernel size of 1. LN represents layer normalization, and Ge represents the GELU activation function. According to the order of the extraction process, the number of local feature blocks in each stage is (2, 2, 4, 2) in sequence;

[0018] 2) In the Transformer branch, the extraction process of global features also has a total of 4 stages. According to the order of the extraction process, the number of specific global feature blocks in each stage is (2, 2, 4, 2) in sequence; There is a separate depth convolution block in the global feature block: DWBlock. Suppose the input is X with a size of D×H×W and having C channels and and respectively represent the feature input of DWBlock and the feature matrix with a size of D×H×W and having C channels. D, H, and W represent depth, height, and width respectively. Then DWBlock is expressed as:

[0019]

[0020] That is, the calculation method of the global feature block is as follows:

[0021]

[0022] Among them, Pooling is a 3D average pooling layer, and MLP is a multi-layer perceptron. G i respectively represent the outputs of the pooling layer, DWBlock, and global feature block.

[0023] Preferably, the method for the hybrid feature fusion module to fuse the features extracted by the deep learning network feature extraction module is specifically as follows:

[0024] Feed the global features of each modality extracted by the deep learning network feature extraction module into the channel attention mechanism: CA mechanism. The CA mechanism utilizes the mutual dependence between channels to improve the feature representation of specific semantics; Feed the local features extracted by the deep learning network feature extraction module into the spatial attention mechanism: SA mechanism to enhance local details and suppress irrelevant regions; The outputs of the CA mechanism and the SA mechanism are fused in features and finally input into the IRMLP of the inverted residual structure. The corresponding calculation methods of the CA mechanism, the SA mechanism, and the IRMLP are as follows:

[0025] CA(X) = Sig(MLP(AvgPool(X)) + MLP(MaxPool(X)));

[0026] SA(X) = Sig(C k=7 (Concat[AvgPool(X), MaxPool(X)]));

[0027]

[0028] Where Sig represents the Sigmoid activation function, BN is batch normalization, AvgPool is the average pooling layer, MaxPool is the max pooling layer, and Concat is the concatenation operation; the hybrid feature fusion operation is as follows:

[0029]

[0030] Where, represents element-wise multiplication, is the global feature weighted by the CA mechanism, is the local feature weighted by the SA mechanism, F i4 is the final output of the hybrid feature fusion module.

[0031] Preferably, the specific method of the weight allocation and classification module is:

[0032] Input the output of the hybrid feature fusion module of each modality into the average pooling layer and perform a flattening operation to generate the feature F AP , F DP , F DWI , F T2WI-FS corresponding to each modality, and first concatenate these features to form a multi-modal feature tensor;

[0033] Among them, the multi-modal feature weighted fusion classification sub-module generates the feature weights corresponding to each modality through a multi-layer perceptron combined with the Sigmoid activation function, and multiplies them element-wise with F AP , F DP , F DWI , F T2WI-FS respectively to form the weighted features F′ AP , F′ DP , F′ DWI , F′ T2WI-FS , add the weighted features of each modality as the fusion feature, and use the fully connected layer to achieve the final classification;

[0034] In the deep network loss supervision submodule, the feature weights corresponding to each modality in the multimodal feature weighted fusion classification submodule are summed up respectively, and then passed into the SoftMax activation function to obtain the loss supervision weight α of each modality. AP ,α DP ,α DWI ,α T2WI-FS , which is used to supervise the loss of each modal feature for classification and dynamically adjust the features of each modality obtained during feature extraction; that is, the loss function of multimodal deep supervision is:

[0035] L DS =α AP *L AP +α DP *L DP +α DWI *L DWI +α T2WI-FS *L T2WI-FS ;

[0036] The loss of the fusion feature classification by adding the weighted features of each modality is recorded as L fusion , then the total loss function of the model can be expressed as:

[0037] L total =L fusion +γL DS ;

[0038] Where γ represents the balance parameter between the loss function of deep supervision and the loss of fusion features.

[0039] Preferably, the lesion segmentation module uses 3Dslicer software or ITK Snap software to manually delineate lesions on AP images, DP images, DWI images and T2WI-FS images.

[0040] Preferably, the system further comprises an image preprocessing module, which preprocesses the multimodal MRI image acquired by the data acquisition module and inputs the preprocessed image into the lesion delineation module.

[0041] Preferably, the image preprocessing module preprocesses the multimodal MRI images acquired by the data acquisition module through resampling and Z-score standardization.

[0042] A second aspect of the present invention provides an application of the system as described above, which is used for preoperative prediction of microvascular invasion of hepatocellular carcinoma based on multimodal MRI images.

[0043] Preferably, the method for preoperatively predicting microvascular invasion of hepatocellular carcinoma based on multimodal MRI images comprises the following steps:

[0044] S1. Data acquisition:

[0045] Obtain the multi-modal MRI images of the patient through the data acquisition module, and the multi-modal MRI images include AP images, DP images, DWI images, and T2WI-FS images;

[0046] S2. Image preprocessing:

[0047] Perform resampling and Z-value normalization preprocessing on the AP images, DP images, DWI images, and T2WI-FS images obtained by the data acquisition module through the image preprocessing module;

[0048] S3. Lesion delineation:

[0049] Perform lesion delineation on the preprocessed AP images, DP images, DWI images, and T2WI-FS images through the lesion delineation module;

[0050] S4. Lesion area cropping:

[0051] Perform 3D cropping on the delineated lesion areas of the AP images, DP images, DWI images, and T2WI-FS images through the lesion area cropping module;

[0052] S5. Deep learning network feature extraction:

[0053] Use the CNN branch and the Transformer branch of the deep learning network feature extraction module to extract local features and global features from the cropped AP images, DP images, DWI images, and T2WI-FS images respectively;

[0054] S6. Hybrid feature fusion:

[0055] Fuse the local features and global features obtained by the CNN branch and the Transformer branch within each modality through the hybrid feature fusion module to obtain the deep learning hybrid features of each modality;

[0056] S7. Weight assignment and classification:

[0057] Use the multi-modal weighted fusion classification module to generate the corresponding weights of each modality feature by using a multi-layer perceptron. After multiplying element by element to obtain the weighted features, the weighted features of each modality are added as the fusion feature for final classification; then generate the loss supervision weights of each modality through the network loss supervision module for dynamically adjusting the process of extracting each modality feature.

[0058] The beneficial effects of the present invention are:

[0059] The present invention provides a system for predicting microvascular invasion of hepatocellular carcinoma based on preoperative multimodal 3D MRI images and its application. The present invention provides a new solution for non-invasive and efficient prediction of microvascular invasion of liver cancer before surgery, which helps doctors formulate more personalized treatment plans, thereby improving the survival rate of patients and reducing the recurrence rate. The prediction of the present invention has high accuracy and strong reliability, and has broad application prospects. Description of the Drawings

[0060] Figure 1 Flow chart for predicting microvascular invasion of hepatocellular carcinoma by the system in Example 1;

[0061] Figure 2 Lesion delineation results in Example 1;

[0062] Figure 3 Model architecture for predicting microvascular invasion of hepatocellular carcinoma by the system of the present invention;

[0063] Figure 4 Details of the deep learning network feature extraction module and the hybrid feature fusion module in Example 1;

[0064] Figure 5 ROC curves and AUC values of the test set under single-modal and multi-modal fusion of the microvascular invasion prediction system of hepatocellular carcinoma in Example 1;

[0065] Figure 6 DCA curves of the test set under single-modal and multi-modal fusion of the microvascular invasion prediction system of hepatocellular carcinoma in Example 1.

[0066] Figure 7 Confusion matrices of the test set under single-modal and multi-modal fusion of the microvascular invasion prediction system of hepatocellular carcinoma in Example 1. Detailed Description of the Invention

[0067] The following further describes the present invention in detail with reference to the embodiments, so that those skilled in the art can implement it according to the description in the specification.

[0068] It should be understood that the terms such as "having", "including" and "comprising" used herein do not exclude the presence or addition of one or more other elements or their combinations.

[0069] Example 1

[0070] This example provides a system for predicting microvascular invasion of hepatocellular carcinoma based on preoperative multimodal 3D MRI. The system includes:

[0071] A data acquisition module, which is used to obtain preoperative multi-modal MRI images of treated patients. The multi-modal MRI images include DCE-MRI images (including arterial phase AP images and delayed phase DP images), DWI images, and T2WI-FS images;

[0072] An image preprocessing module, which is used to perform resampling and Z-value normalization preprocessing on the AP images, DP images, DWI images, and T2WI-FS images obtained by the data acquisition module;

[0073] A lesion delineation module, which is used to delineate lesions in the AP images, DP images, DWI images, and T2WI-FS images;

[0074] A lesion area cropping module, which is used to crop the delineated lesion areas in the AP images, DP images, DWI images, and T2WI-FS images in 3D form;

[0075] A deep learning network feature extraction module, which respectively extracts local features and global features from the cropped AP images, DP images, DWI images, and T2WI-FS images by using a CNN branch and a Transformer branch;

[0076] A hybrid feature fusion module, which is used to fuse the local features and global features obtained by the CNN branch and the Transformer branch within each modality to obtain the deep learning hybrid features of each modality;

[0077] A weight assignment and classification module, which includes a multi-modal feature weighted fusion classification module and a deep network loss supervision module;

[0078] The multi-modal feature weighted fusion classification module uses a multi-layer perceptron to perform dimensional mapping on the high-dimensional features formed by splicing the hybrid features of each modality, normalizes them by using a Sigmoid activation function, generates the corresponding feature weights of each modality, and then multiplies them element-wise with the hybrid features of each modality to form the weighted features of each modality. The weighted features of each modality are added together as the fusion features to participate in the final classification;

[0079] The network loss supervision module sums up the weighted features of each modality obtained in the multi-modal feature weighted fusion module, and then respectively passes them into a SoftMax activation function to obtain the loss supervision weights of each modality, which are used to adjust the features of each modality obtained by feature extraction.

[0080] Refer to Figure 1 , The method for the system to perform preoperative prediction of microvascular invasion of hepatocellular carcinoma based on multi-modal MRI images includes the following steps:

[0081] S1. Data acquisition:

[0082] The preoperative multimodal MRI images of the treated patients are obtained through the data acquisition module, and the multimodal MRI images include AP images, DP images, DWI images, and T2WI-FS images;

[0083] Specifically, in this example, the collected data includes training set data and test set data, and the data acquisition method is as follows:

[0084] Retrospectively collect the data from January 2019 to December 2022 in the Affiliated Hospital of Xuzhou Medical University as the experimental data set (n = 174), among which there are 47 cases with MVI (microvascular invasion) and 127 cases without MVI. The data set is randomly divided into a training set (n = 120) and a test set (n = 54) according to the ratio of 7:3. Among them, there are 32 cases with MVI and 88 cases without MVI in the training set; there are 15 cases with MVI and 39 cases without MVI in the test set; perform image transformation operations such as random translation, rotation, scaling, and Gaussian noise on the 3D MRI images of 47 patients with MVI to achieve the effect of data augmentation. The 80 augmented 3D MRI images are also randomly divided according to the ratio of 7:3 and are respectively added to the already divided training set (n = 176) and test set (n = 78). Finally, there are 88 cases with MVI and 88 cases without MVI in the training set; there are 39 cases with MVI and 39 cases without MVI in the test set;

[0085] S2. Image preprocessing:

[0086] Resample the AP images, DP images, DWI images, and T2WI-FS images obtained by the data acquisition module through the image preprocessing module, and sample the voxel block size to: 1mm × 1mm × 1mm; then use the Z-score function for standardization processing to reduce the differences brought by images collected by different devices;

[0087] S3. Lesion delineation:

[0088] Delineate the lesions on the preprocessed AP images, DP images, DWI images, and T2WI-FS images through the lesion delineation module. Specifically in this embodiment:

[0089] A radiologist with 10 years of clinical experience uses 3Dslicer software to manually delineate the tumor regions of all cases. Appropriately expand the delineation range for the uncertain tumor margins (expand outward by 1mm); if there are multiple tumors, select the larger tumor for delineation; avoid non-tumor tissues such as large blood vessels around during delineation. The delineation schematic diagram is as Figure 2 shown. All segmentation results are reviewed by experienced radiologists (with 15 years of clinical experience);

[0090] S4. Lesion area cropping:

[0091] The delineated lesion regions of the AP image, DP image, DWI image, and T2WI-FS image are cropped in 3D by the lesion region cropping module. During the cropping process, the cross-section, sagittal plane, and coronal plane of the lesion region are used as references, and the tumor lesion edge is used as the benchmark, extending 2 pixels outward to ensure the complete retention of the lesion region and its surrounding tissue information;

[0092] S5. Deep learning network feature extraction:

[0093] The cropped AP image, DP image, DWI image, and T2WI-FS image are respectively used by the deep learning network feature extraction module to extract local features and global features using the CNN branch and the Transformer branch. The specific process is as follows:

[0094] First, the cropped AP image, DP image, DWI image, and T2WI-FS image are scaled, where C, D, H, and W are (1, 64, 64, 64).

[0095] 1) In the CNN branch, the local feature extraction process has a total of 4 stages, and there are specific local feature blocks in each stage. The calculation method is as follows:

[0096]

[0097] Among them, L i represents the output of the local feature block in the i-th stage, i ∈ 1, 2, 3, 4, and L i-1 represents the output of the local feature block in the i - 1 stage, represents a 3D depth convolution with a convolution kernel size of 3, and C k=1 represents a 3D standard convolution with a convolution kernel size of 1. LN represents layer normalization, and Ge represents the GELU activation function. According to the order of the extraction process, the number of local feature blocks in each stage is (2, 2, 4, 2), and the C, D, H, and W of the output after each stage passing through the local feature block are (32, 32, 32, 32), (64, 16, 16, 16), (128, 8, 8, 8), and (256, 4, 4, 4) respectively.

[0098] 2) In the Transformer branch, the input image is first subjected to patch embedding operation. The patch size is set to 2, and the output C, D, H, W becomes (32, 32, 32, 32). The process of extracting global features also has a total of 4 stages, and the number of specific global feature blocks in each stage is (2, 2, 4, 2). In the latter three stages, patch merging is separately performed for downsampling. After passing through the global feature blocks in each stage, the output C, D, H, W are (32, 32, 32, 32), (64, 16, 16, 16), (128, 8, 8, 8), and (256, 4, 4, 4) respectively. A separate depthwise convolution block (DWBlock) is designed in the global feature block. Suppose the input is X and respectively represent the feature input of the DWBlock and the feature matrix with C channels and size D×H×W. D, H, and W represent depth, height, and width respectively. Then the DWBlock can be expressed as:

[0099]

[0100] That is, the calculation method of the global feature block is as follows:

[0101]

[0102] Among them, Pooling is a 3D average pooling layer, and MLP is a multi-layer perceptron. G i respectively represent the outputs of the pooling layer, DWBlock, and global feature block.

[0103] S6. Hybrid feature fusion:

[0104] The local features and global features obtained from the CNN branch and Transformer branch within each modality are fused through the hybrid feature fusion module to obtain the deep learning hybrid features of each modality. The specific process is as follows:

[0105] The global features within each modality finally extracted by the model are fed into the channel attention (CA) mechanism, which utilizes the mutual dependence between channels to improve the feature representation of specific semantics; the extracted local features are fed into the spatial attention (SA) mechanism to enhance local details and suppress irrelevant regions. The outputs of each attention mechanism path will be subjected to feature fusion and finally input into the MLP (IRMLP) of the inverted residual structure. The calculation methods of the corresponding CA mechanism, SA mechanism, and IRMLP are as follows:

[0106] CA(X) = Sig(MLP(AvgPool(X)) + MLP(MaxPool(X)));

[0107] SA(X) = Sig(C k=7 (Concat[AvgPool(X), MaxPool(X)]));

[0108]

[0109] Among them, Sig represents the Sigmoid activation function, BN is batch normalization, AvgPool is the average pooling layer, MaxPool is the max pooling layer, and Concat is the concatenation operation. The hybrid feature fusion operation can be expressed as follows:

[0110]

[0111] Among them, represents element-wise multiplication, is the global feature weighted by the CA mechanism, is the local feature weighted by the SA mechanism, F i4 is the final output of the hybrid feature fusion module;

[0112] S7. Weight Assignment and Classification:

[0113] Through the multi-modal weighted fusion classification module, the multi-layer perceptron is used to generate the weights corresponding to the features of each modality. After obtaining the weighted features by element-wise multiplication, the weighted features of each modality are added as the fusion feature for final classification; then, through the deep network loss supervision module, the loss supervision weights of each modality are generated for dynamically adjusting the feature extraction process of each modality. The specific process is as follows:

[0114] The outputs of the hybrid feature fusion modules of each modality are input into the average pooling layer and flattened to generate the corresponding features F AP , F DP , F DWI , F T2WI-FS of each modality. First, these features are concatenated to form a multi-modal feature tensor.

[0115] Among them, the multi-modal feature weighted fusion classification sub-module, through the multi-layer perceptron combined with the Sigmoid activation function, generates the feature weights corresponding to each modality, and multiplies them element-wise with F AP , F DP , F DWI , F T2WI-FS respectively to form the weighted features F' AP , F' DP , F′DWI , F' T2WI-FS , the weighted features of each modality are added as the fused feature, and the fully connected layer (FC) is used to achieve the final classification;

[0116] In the deep network loss supervision sub-module, the corresponding feature weights of each modality in the multi-modal feature weighted fusion classification sub-module are respectively summed, and then respectively passed into the SoftMax activation function to obtain the loss supervision weights α of each modality AP , α DP , α DWI , α T2WI-FS , which is used to supervise the loss of each modality feature for separate classification and dynamically adjust the modality features obtained during the feature extraction process. That is, the loss function of multi-modal deep supervision is:

[0117] L DS = α AP * L AP + α DP * L DP + α DWI * L DWI + α T2WI-FS * L T2WI-FS ;

[0118] The loss of the fused feature classification obtained by adding the weighted features of each modality is denoted as L fusion , then the total loss function of the model can be expressed as:

[0119] L total = L fusion + γL DS ;

[0120] Where γ represents the balance parameter between the loss function of deep supervision and the loss of the fused feature. This feature is learnable and is predefined as 1.0. The loss function used by the model is the cross-entropy loss (Cross Entropy Loss).

[0121] Table 1 gives the individual prediction results of microvascular invasion of hepatocellular carcinoma by 4-modal 3D MRI:

[0122] Table 1 Prediction results of microvascular invasion of hepatocellular carcinoma based on single modality

[0123]

[0124]

[0125] Table 2 gives the final prediction results of microvascular invasion of hepatocellular carcinoma based on 4-modal MRI and the results of its ablation experiments:

[0126] Table 2 Final prediction results and ablation results of microvascular invasion of hepatocellular carcinoma based on multimodal MRI

[0127]

[0128] As can be seen from Table 1 and Table 2, compared with the results of single modality, the deep learning model based on multimodal fusion studied in the present invention has obtained the best prediction performance, and with the addition of different modules, the results of the model in multiple indicators such as AUC, ACC, SEN, and SPE have been significantly improved.

[0129] Figure 5 、 Figure 6 The ROC curve and DCA curve of the test set under single modality and multimodal fusion are given. As Figure 5 can be seen, the AUC value of the prediction result of fusing multiple modalities is the highest, which is 0.9513. And Figure 6 it shows that fusing multiple modalities has more clinical decision-making benefits than single modality.

[0130] Figure 7 The confusion matrix of the test set under single modality and multimodal fusion is given. As Figure 7 can be seen, compared with (a), (b), (c), and (d) of AP, DP, DWI, and T2WI-FS modalities, (e) under multimodal fusion achieves the largest number of correctly predicted negative and positive samples.

[0131] Although the implementation embodiments of the present invention have been disclosed above, it is not limited to the applications listed in the specification and implementation manners. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details.

Claims

1. A prediction system for microvascular invasion of hepatocellular carcinoma based on preoperative multi-modal 3D MRI, characterized in that, The system includes: A data acquisition module for acquiring multi-modal MRI images, which include AP images, DP images, DWI images, and T2WI-FS images; A lesion delineation module for delineating lesions in AP images, DP images, DWI images, and T2WI-FS images; A lesion area cropping module for cropping the delineated lesion areas in AP images, DP images, DWI images, and T2WI-FS images in 3D form; A deep learning network feature extraction module that extracts local features and global features from the cropped images of each modality: AP images, DP images, DWI images, and T2WI-FS images using CNN branches and Transformer branches respectively; A hybrid feature fusion module for fusing the local features and global features obtained from the CNN branches and Transformer branches within each modality to obtain the deep learning hybrid features for each modality; A weight assignment and classification module, which includes a multi-modal feature weighted fusion classification module and a deep network loss supervision module; The multi-modal weighted fusion classification module generates weights corresponding to the features of each modality using a multi-layer perceptron, multiplies element-wise to obtain weighted features, and then adds the weighted features of each modality as the fusion feature for final classification; then, the deep network loss supervision module generates loss supervision weights for each modality to dynamically adjust the feature extraction process of each modality.

2. The hepatocellular carcinoma microvascular invasion prediction system based on preoperative multimodal MRI according to claim 1, wherein, The multi-modal feature weighted fusion classification module performs dimensionality mapping on the high-dimensional features formed by splicing the hybrid features of each modality using a multi-layer perceptron, normalizes using the Sigmoid activation function, generates the corresponding feature weights for each modality, multiplies them element-wise with the input of the hybrid features of each modality to form the weighted features of each modality, and adds the weighted features of each modality as the fusion feature to participate in the final classification; The deep network loss supervision module sums up the weighted features of each modality obtained in the multi-modal feature weighted fusion module, and then respectively inputs them into the SoftMax activation function to obtain the loss supervision weights for each modality, which are used to adjust the features of each modality obtained by feature extraction.

3. The hepatocellular carcinoma microvascular invasion prediction system based on preoperative multimodal MRI according to claim 2, wherein The method for the deep learning network feature extraction module to perform feature extraction is specifically as follows: 1) In the CNN branch, the local feature extraction process has a total of 4 stages, and there are specific local feature blocks in each stage. The calculation method is as follows: Among them, L i represents the output of the local feature block in the i-th stage, where i ∈ 1, 2, 3, 4, and L i-1 represents the output of the local feature block in the (i - 1)-th stage, represents a 3D depth convolution with a convolution kernel size of 3, and C k=1 represents a 3D standard convolution with a convolution kernel size of 1. LN represents layer normalization, and Ge represents the GELU activation function. According to the order of the extraction process, the number of local feature blocks in each stage is (2, 2, 4, 2) in sequence; 2) In the Transformer branch, the extraction process of global features also has 4 stages in total. According to the order of the extraction process, the specific number of global feature blocks in each stage is (2, 2, 4, 2) in sequence; there is a separate depth convolution block in the global feature block: DWBlock. Suppose the input is of size D×H×W with C channels X and respectively represent the feature input of DWBlock and the feature matrix of size D×H×W with C channels. D, H, and W represent depth, height, and width respectively. Then DWBlock is expressed as: That is, the calculation method of the global feature block is as follows: Among them, Pooling is a 3D average pooling layer, and MLP is a multi-layer perceptron. They respectively represent the outputs of the pooling layer, DWBlock, and global feature block.

4. The hepatocellular carcinoma microvascular invasion prediction system based on preoperative multimodal MRI according to claim 3, wherein The method for the hybrid feature fusion module to fuse the features extracted by the deep learning network feature extraction module is specifically as follows: Feed the global features of each modality extracted by the deep learning network feature extraction module into the channel attention mechanism: CA mechanism. The CA mechanism utilizes the mutual dependence between channels to improve the feature representation of specific semantics. Feed the local features extracted by the deep learning network feature extraction module into the spatial attention mechanism: SA mechanism to enhance local details and suppress irrelevant regions. The outputs of the CA mechanism and the SA mechanism are fused, and finally input into the IRMLP of the inverted residual structure. The calculation methods of the corresponding CA mechanism, SA mechanism, and IRMLP are as follows: CA(X) = Sig(MLP(AvgPool(X)) + MLP(MaxPool(X))); SA(X) = Sig(C k=7 (Concat[AvgPool(X), MaxPool(X)])); Among them, Sig represents the Sigmoid activation function, BN is batch normalization, AvgPool is the average pooling layer, MaxPool is the max pooling layer, and Concat is the concatenation operation. The hybrid feature fusion operation is expressed as follows: Among them, represents element-wise multiplication, is the global feature after being weighted by the CA mechanism, is the local feature after being weighted by the SA mechanism, F i4 is the final output of the hybrid feature fusion module.

5. The prediction system for microvascular invasion of hepatocellular carcinoma based on preoperative multimodal MRI according to claim 4, wherein The specific method of the weight assignment and classification module is as follows: Input the output of the hybrid feature fusion module of each modality into the average pooling layer and perform a flattening operation to generate the feature F corresponding to each modality AP ,F DP ,F DWI ,F T2WI-FS , and first splice these features to form a multi-modal feature tensor; Among them, the multi-modal feature weighted fusion classification sub-module uses a multi-layer perceptron with a Sigmoid activation function to generate the feature weights corresponding to each modality, and multiplies them element-wise with F AP , F DP , F DWI , F T2WI-FS respectively to form the weighted features F' AP , F' DP , F' DWI , F' T2WI-FS . The weighted features of each modality are added as the fusion feature, and the final classification is achieved using a fully connected layer; In the deep network loss supervision sub-module, the feature weights corresponding to each modality in the multi-modal feature weighted fusion classification sub-module are respectively summed, and then are respectively input into the SoftMax activation function to obtain the loss supervision weights α AP , α DP , α DWI , α T2WI-FS , which are used to supervise the losses of each modality feature for separate classification and dynamically adjust each modality feature obtained in the feature extraction process; that is, the loss function of multi-modal deep supervision is: L DS = α AP * L AP + α DP * L DP + α DWI * L DWI + α T2WI-FS * L T2WI-FS ; The loss of the classification of the fused features obtained by adding the weighted features of each modality is denoted as L fusion , then the total loss function of the model can be expressed as: L total = L fusion + γL DS ; Among them, γ represents the balance parameter between the loss function of deep supervision and the loss of the fused features.

6. The hepatocellular carcinoma microvascular invasion prediction system based on preoperative multimodal MRI according to claim 1, wherein The lesion segmentation module manually outlines the lesions in the AP image, DP image, DWI image, and T2WI-FS image using 3Dslicer software or ITK Snap software.

7. The hepatocellular carcinoma microvascular invasion prediction system based on preoperative multi-modal 3D MRI according to claim 1, wherein The system further includes an image preprocessing module, which preprocesses the multi-modal MRI images acquired by the data acquisition module and then inputs them into the lesion outlining module.

8. The prediction system for microvascular invasion of hepatocellular carcinoma based on preoperative multimodal MRI according to claim 7, wherein The image preprocessing module preprocesses the multi-modal MRI images acquired by the data acquisition module through resampling and Z-score normalization.

9. An application of the system according to any one of claims 1-8, which is used for preoperative prediction of microvascular invasion of hepatocellular carcinoma based on multi-modal MRI images.

10. The application according to claim 9, characterized in that, The method for the system to perform preoperative prediction of microvascular invasion of hepatocellular carcinoma based on multi-modal MRI images includes the following steps: S1. Data acquisition: Acquire the multi-modal MRI images of the patient through the data acquisition module. The multi-modal MRI images include AP images, DP images, DWI images, and T2WI-FS images; S2. Image preprocessing: Perform resampling and Z-value normalization preprocessing on the AP images, DP images, DWI images, and T2WI-FS images acquired by the data acquisition module through the image preprocessing module; S3. Lesion outlining: Outline the lesions in the preprocessed AP images, DP images, DWI images, and T2WI-FS images through the lesion outlining module; S4. Lesion area cropping: Perform 3D cropping on the outlined lesion areas of the AP images, DP images, DWI images, and T2WI-FS images through the lesion area cropping module; S5. Deep learning network feature extraction: Use the CNN branch and the Transformer branch of the deep learning network feature extraction module to extract local features and global features from the cropped AP images, DP images, DWI images, and T2WI-FS images respectively; S6. Hybrid feature fusion: The local features and global features obtained by the CNN branch and the Transformer branch within each modality are fused through the hybrid feature fusion module to obtain the deep learning hybrid features of each modality; S7. Weight assignment and classification: Through the multi-modal weighted fusion classification module, the weights corresponding to the features of each modality are generated by using a multi-layer perceptron. After obtaining the weighted features by element-wise multiplication, the weighted features of each modality are added as the fusion feature for final classification; then, the network loss supervision module generates the loss supervision weights of each modality, which are used to dynamically adjust the feature extraction process of each modality.