Liver cancer mvi tree classification model based on self-supervised learning fine-grained identification
Patent Information
- Application Number
- CN202311703231.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-12-12
AI Technical Summary
其次,该方法使用的深度学习方法为简单的卷积神经网络,对图像的特征提取能力也有限
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the following description, in conjunction with embodiments, further illustrates the invention. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention.
Smart Images

Figure CN117911742B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of medical imaging, deep learning, and computer vision, and particularly relates to the use of fine-grained classification and pre-trained models. Background Technology
[0002] Liver cancer is the sixth most common cancer worldwide, with high incidence and mortality rates. Microvascular invasion (MVI) around the tumor is an independent risk factor for postoperative recurrence in liver cancer. While MVI detection is usually performed through postoperative pathological examination, preoperative prediction of MVI has significant clinical value for patient treatment and postoperative assessment. MVI is classified into three grades: M0, M1, and M2. Past research has largely focused on the classification of MVI as positive or negative.
[0003] Currently, existing MVI classification methods have poor accuracy, mainly because the inter-class differences between M1 and M2 are small. In general scenarios, three-class classification models tend to concentrate the classification results on M0 and M2, failing to learn the class features of M1. On the other hand, binary classification models of M1 and M2 lack sufficiently powerful feature extraction capabilities to achieve adequate classification performance.
[0004] The closest existing technology is: Anjun Song, Yueyue Wang, Wentao Wang, et al., “Using deep learning to predict microvascular invasion in hepatocellular carcinoma based on dynamic contrast-enhanced MRI combined with clinical parameters,” Journal of Cancer Research and Clinical Oncology. The principle of this paper's "Deep learning-based prediction of microvascular invasion in hepatocellular carcinoma based on dynamic contrast-enhanced MRI combined with clinical parameters" is as follows: Features from eight types of MRI images are extracted using a convolutional neural network (CNN), and a classification vector is constructed by combining it with clinical features. A fully connected layer (FC) is then used for three-class classification, categorizing the microvascular invasion (MVI) into M0, M1, and M2. However, because this method only uses a simple three-class classification method for MVI grading, M1 is easily misclassified into the other two categories. Secondly, the deep learning method used is a simple convolutional neural network, which has limited feature extraction capabilities. Ultimately, the accuracy of its experimental results is also low. Summary of the Invention
[0005] This invention proposes a tree-like hierarchical model for liver cancer MVI based on self-supervised learning and fine-grained recognition.
[0006] This invention abandons the three-classification approach and instead splits the hierarchical task into two tree-like binary classification models to avoid the problem that M1 and M2 are too similar.
[0007] To improve the classification accuracy of severity classification, this invention integrates fine-grained classification and self-supervised learning methods and applies them to the MVI severity classification model.
[0008] The model of this invention uses contrastive learning to pre-train the feature extraction module, obtains self-attention maps in layers for the extracted features, and obtains classification features in fine-grained manner by combining the self-attention maps.
[0009] In summary, this significantly improves the accuracy and precision of MVI classification.
[0010] Technical solution
[0011] This invention discloses a tree-like grading model for hepatocellular carcinoma MVI based on self-supervised learning and fine-grained classification. The model's construction includes: a data preprocessing module, an MVI negative / positive classification model, and an MVI severity classification model.
[0012] The data preprocessing module acquires MRI images of the patient's liver to form a dataset and performs preprocessing.
[0013] The MVI negative / positive classification model takes two preprocessed modalities, T1 and T1D, as input. Features are extracted using two identical deep residual networks (ResNet18) as backbones, resulting in 3-dimensional features: number of channels, length, and width. After obtaining the features for both modalities, they are concatenated along the channel dimension. Then, two consecutively connected convolutional layers (Conv*2) fuse the multimodal information along the channel dimension. Finally, a fully connected layer (FC) is used to obtain the negative / positive classification result, where the MVI positive level includes M1 and M2.
[0014] The MVI severity classification model takes T1 images of MVI-positive patients as input and outputs the patient's MVI grading result (M1 or M2).
[0015] After the MRI images are processed by the data preprocessing module, they are first passed through a trained MVI negative-positive classification model. If the classification result is negative, the case grade is output as M0. If the classification result is positive, the image is then input into the trained MVI severity classification model to obtain the grade result of M1 or M2.
[0016] By implementing and applying the model of this invention, the MVI level of liver cancer patients can be accurately determined using small sample training, greatly improving the accuracy of MVI grading. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the complete structure of the model of the present invention.
[0018] Figure 2 This is a schematic diagram of the MVI negative-positive classification model structure in this invention.
[0019] Figure 3 The schematic diagram of SimCLR pre-training used in this invention and the structural schematic diagram of the MVI severity classification model in this invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the following description, in conjunction with embodiments, further illustrates the invention. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention.
[0021] This article presents a grading study of MVI. When applied in the future, the model of this invention can help liver cancer patients determine whether to adopt aggressive treatment methods.
[0022] In the application of this invention's research findings: It utilizes deep learning methods to classify patients' MVI (malignant visceral infarction) using preoperative magnetic resonance imaging (MRI) images, thereby accelerating MVI diagnosis and meeting the need for preoperative MVI prediction. This invention's technical solution itself is merely a tool for diagnosis and decision support, not a direct solution for treating and saving lives.
[0023] Experiments showed that when using deep learning methods for MVI grading, the classification accuracy for negative (M0) and positive (M1, M2) cases was high, while the accuracy was poor when performing fine-grained classification of positive patients. To address these challenges, this invention employs a tree structure for MVI case grading. For positive cases where grading is difficult to distinguish, this invention introduces a self-supervised learning method for model pre-training, along with fine-grained classification and self-attention methods to help the model learn deep features and fine-grained distinctions between different categories. Through the implementation and application of this invention's model, accurate MVI grading of liver cancer patients can be determined using small-sample training, significantly improving the accuracy of MVI grading.
[0024] This invention discloses a tree-like grading model for hepatocellular carcinoma MVI based on self-supervised learning and fine-grained classification. The model's construction includes: a data preprocessing module, an MVI negative / positive classification model, and an MVI severity classification model.
[0025] The data preprocessing module acquires MRI images of the patient's liver to form a dataset and performs preprocessing.
[0026] The MVI negative-positive classification model is as follows: Figure 2 :
[0027] The preprocessed images of two modalities, T1 and T1D, are used as input. Features are extracted using two identical deep residual networks (ResNet18) as backbones, resulting in 3-dimensional features: number of channels, length, and width. After obtaining the features of the two modalities, they are concatenated along the channel dimension. Then, two consecutively connected convolutional layers (Conv*2) are used to fuse the multimodal information along the channel dimension. Finally, a fully connected layer (FC) is used to obtain the negative / positive classification results, where the MVI positive level includes M1 and M2.
[0028] The MVI severity classification model is as follows: Figure 3 Input is the T1 image of a patient with MVI positivity, and output is the patient's MVI grading result (M1 or M2).
[0029] After the MRI images are processed by the data preprocessing module, they are first passed through a trained MVI negative-positive classification model. If the classification result is negative, the case grade is output as M0. If the classification result is positive, the image is then input into the trained MVI severity classification model to obtain the grade result of M1 or M2.
[0030] Details are as follows
[0031] The data preprocessing module performs image preprocessing on the dataset, specifically including:
[0032] S1.1, Read the T1 and T1D images of the MRI images in the dataset;
[0033] S1.2, The tumor region is segmented and labeled to obtain tumor segmentation labels. The label format is the coordinates of each vertex of the labeled polygon in the image.
[0034] S1.3, find the minimum and maximum x and y coordinates of all segmentation points to determine the minimum rectangular closure region that encloses the entire tumor region.
[0035] S1.4 expands the four boundaries of the closure region outward.
[0036] S1.5, truncate the expanded closure region and resample it to the same size 224*224.
[0037] The MVI severity classification model processing procedure includes:
[0038] S3.1, as Figure 3 As shown in (a): First, the feature extraction module (using ResNet34 encoder) is pre-trained using images. Then, the contrastive learning method SimCLR is used to train the T1 image x.i The enhanced image x is obtained through data augmentation. j Then x i and x j The input is fed into a two-channel ResNet34 encoder with shared parameters to obtain their feature representations h. i and h j Then, these two feature representations are input into two prediction heads g(.) with the same parameters to obtain the feature vector z. i and z j Finally, maximize the two eigenvectors (z) i and z j The similarity between the two is calculated, and the parameters are updated. The trained ResNet34 encoder is then used for... Figure 3 (b) Feature extraction backbone network.
[0039] (Note: The prediction head g(.) is a neural network composed of fully connected layers.)
[0040] S3.2 uses the ResNet34 encoder trained in S3.1 as the backbone network for feature extraction. The image is input into the ResNet34 encoder to obtain the feature maps (f2, f3) output from the last two convolutional layers.
[0041] S3.3, the feature maps (f2, f3) obtained from S3.2 are passed through a convolutional layer and a pooling layer (not shown in the figure) with non-shared parameters to obtain multi-scale attention maps (a2, a3), and then summed to obtain a multi-scale fused attention map (A). The dimension of attention map (A) is 32*7*7, where the three dimensions are the channel, length and width of the attention map, respectively.
[0042] In step S3.4, attention pooling is performed on the final feature map (f3) processed by the ResNet34 encoder obtained in step S3.2 and the multi-scale fused attention map (A) obtained in step S3.3. The 32 channels of the attention map (A) are treated as 32 two-dimensional attention maps, and each is multiplied by the feature map (f3) to obtain 32 attention-guided feature maps (PF). Global average pooling is performed on the 32 feature maps (PF) in both length and width dimensions to obtain a one-dimensional vector with the length of 32 channels. The vectors are concatenated to obtain the final classification feature (P). The obtained classification feature (P) is input into the fully connected layer to obtain the final classification result.
[0043] The MVI severity classification model described above, under contrastive learning pre-training, further employs fine-grained classification ideas to extract attention maps from feature maps in layers, and uses the attention maps to guide feature map generation to form classification feature vectors.
[0044] Specifically, a contrastive learning method is first used to pre-train a deep residual network (specifically, the ResNet34 encoder) to enable it to extract MVI features, thereby improving the convergence speed of model training. Then, attention maps are used to guide the generation of classification vectors from the feature maps, guiding the model to focus on fine-grained regions in the image during classification, thus improving the model's ability to learn and recognize fine-grained features.
[0045] Example
[0046] Data Preprocessing Module 1
[0047] In this module, the Region of Interest (ROI) is cropped based on the tumor segmentation labels of the tumor region. First, the smallest closure region X of the segmented region needs to be found. Then, the width W of X is calculated as one-third w. The four boundaries of the closure region are expanded outwards by the length w. The expanded position is determined using the following formula (taking the right boundary of the closure region as an example):
[0048]
[0049] Where L represents the x-coordinate of the right boundary of the expanded closure region, x represents the x-coordinate of the right boundary of the closure region before the transformation, and W img This represents the right boundary of the entire image. The above formula is mainly to avoid cropping the image beyond its boundaries. The purpose of expansion is that MVI needs to focus on the patient's tumor and the vascular invasion in the surrounding area. The cropped image is then enlarged to 224*224.
[0050] Building a hierarchical model
[0051] The overall hierarchical model is obtained by organizing two binary classification networks through a tree structure.
[0052] The network model can be referenced. Figure 1 .
[0053] In the first stage of the model, the model structure is as follows: Figure 2 (Right now Figure 1 (Top right section) This section uses an MVI negative / positive classification network to determine whether a patient has MVI symptoms. The network extracts multimodal features through two feature extraction branches. The feature extraction part uses ResNet18 pre-trained on ImageNet. After obtaining the bimodal features, they are concatenated along the channel dimension. Two convolutions are then used to halve the channel dimension of the fused feature map without changing the feature map size, thus achieving the purpose of fusing multimodal features. Finally, a fully connected layer is used to obtain the classification result.
[0054] In the second stage of the model, the model structure is as follows: Figure 3 (b)(i.e.) Figure 1(Lower right section) The MVI severity classification model is used to further classify patients who are positive in the first stage.
[0055] The feature extraction part of the model consists of a SimCLR pre-trained ResNet34, and the specific pre-trained model structure is as follows: Figure 3 (a) During the training phase, the last two feature maps of the feature extraction module, f2 and f3, are obtained and fed into a multi-scale attention extractor to obtain attention map A. At this point, the size of attention map A is 32*7*7, and the final feature map f3 after feature extraction is 512*7*7. Attention pooling is performed using this data, dividing the attention map into 32 7*7 two-dimensional attention maps. Each attention map is multiplied by the 512*7*7 feature map to obtain an attention-guided feature map PF of size 512*7*7. This PF is then pooled in the spatial dimension to obtain a 512*1*1 feature vector. The feature vectors obtained after processing the 32 attention maps are concatenated to obtain the final 16384*1 feature vector P used for classification. Finally, the model passes through a fully connected layer to obtain the classification result.
[0056] The MVI tree-type hierarchical composite model of the present invention was evaluated and tested using common classification indicators such as accuracy, precision, recall, F1 index, sensitivity, and specificity, as well as commonly used medical evaluation indicators.
[0057] All the above metrics indicate that higher values generally mean better model performance. Example: Model training and experimental verification:
[0058] Taking the actual collected dataset as an example, a total of 363 images were obtained. During the experiment, the training and test sets were divided in a 4:1 ratio. All experimental results presented in this invention are from five-fold cross-validation. All experiments in this invention used PyTorch as the algorithm building tool. The grading experiments were performed using a 1080 GPU, while the self-supervised learning SimCLR-related experiments were performed using a 3090 GPU. Images during training were recorded using TensorBoard. The training batch size was set to 32, using the Adamw optimizer with an initial learning rate of 1e-3 and an optimizer weight-decay of 1e-6. Cosine annealing was used as the learning rate update strategy. The cross-entropy loss function was used for the MVI severity classification task, and the focal loss function was used for the MVI negative / positive classification task, with alpha set to 0.4 and gamma set to 3. The final accuracy obtained by this invention was 0.7193, precision 0.7417, recall 0.7089, and F1 score 0.7076. We validated our work on all basic methods, such as ResNet18, and found that their accuracy was below 0.6. We validated our work on the baseline method of this invention, MA-Net, which achieved an accuracy of 0.6186, nearly 10% lower than our invention. These comparative validations demonstrate that our model can improve the accuracy of MVI classification.
[0059] convolution 0.5787 0.6179 0.5806 0.5740 ResNet18 0.5646 0.5368 0.5518 0.5306 ResNet34 0.5816 0.5832 0.5628 0.5526 ResNet34 (SimCLR) 0.5687 0.6111 0.5550 0.5417 MA-Net 0.6186 0.6755 0.6088 0.5908 This invention model 0.7193 0.7417 0.7089 0.7076
[0060] Table 1: Comparative Experimental Results. Except for the model of this invention, all experiments in the table treat the classification as a three-class classification problem. ResNet34 (SimCLR) refers to the results of experiments using ResNet34 pre-trained with SimCLR. MA-Net is the benchmark method compared during the research period of this invention.
Claims
1. A hepatocellular carcinoma MVI (Multi-VI) hierarchical classification method based on self-supervised learning and fine-grained classification, characterized in that, include: Data preprocessing module, MVI negative / positive classification model, and MVI severity classification model; The data preprocessing module acquires MRI images of the patient's liver to form a dataset and performs preprocessing. The MVI negative / positive classification model takes two preprocessed modal images, T1 and T1D, as input. Features are extracted using two identical deep residual networks as backbone networks, resulting in 3-dimensional features: number of channels, length, and width. After obtaining the features of the two modalities, the features are concatenated along the channel dimension. The multimodal information is then fused along the channel dimension using two consecutively connected convolutional layers. Finally, a fully connected layer is used to obtain the negative / positive classification result, where the MVI positive level includes M1 and M2. The MVI severity classification model takes T1 images of MVI-positive patients as input and outputs the patient's MVI grading result, i.e., M1 or M2. After the MRI images are processed by the data preprocessing module, they are first passed through the trained MVI negative and positive classification model. If the classification result is negative, the case grade result is output as M0. If the classification result is positive, the image is then input into the trained MVI severity classification model to obtain the grade result of M1 or M2. The MVI severity classification model processing procedure includes: S3.1 First, the feature extraction module is pre-trained using images. The contrastive learning method SimCLR is used to train the T1 image. The enhanced image is obtained through data augmentation. After that and The inputs are fed into a two-channel ResNet34 encoder with shared parameters to obtain their feature representations. and Then, these two feature representations are input into two prediction heads g(.) with the same parameters to obtain the feature vector. and Finally, the similarity between the two feature vectors is maximized, and the parameters are updated; the trained encoder ResNet34 is used as the feature extraction backbone network. S3.2, use the ResNet34 encoder trained in S3.1 as the backbone network for feature extraction; input the image into the ResNet34 encoder to obtain the feature maps f2 and f3 output by the last two convolutional layers; S3.3, The feature maps f2 and f3 obtained from S3.2 are passed through a convolutional layer and a pooling layer with non-shared parameters respectively to obtain multi-scale attention maps a2 and a3, and then summed to obtain a multi-scale fused attention map A. The dimensions of attention map A are 32*7*7, where the three dimensions are the channels, length and width of the attention map, respectively. In step S3.4, attention pooling is performed on the final feature map f3 obtained from the ResNet34 encoder in step S3.2 and the multi-scale fused attention map A obtained in step S3.
3. The 32 channels of attention map A are treated as 32 two-dimensional attention maps, and each is multiplied by feature map f3 to obtain 32 attention-guided feature maps PF. Global average pooling is performed on the 32 feature maps PF in both length and width dimensions to obtain a one-dimensional vector with the length of 32 channels. The vectors are concatenated to obtain the final classification feature P. The obtained classification feature P is input into the fully connected layer to obtain the final classification result.
2. The hepatocellular carcinoma MVI tree-like grading method based on self-supervised learning and fine-grained classification as described in claim 1, characterized in that, The data preprocessing module completes the image preprocessing process for the dataset, specifically including: S1.1, Read the T1 and T1D images of the MRI images in the dataset; S1.2, The tumor region is segmented and labeled to obtain tumor segmentation labels. The label format is the coordinates of each vertex of the labeled polygon in the image. S1.3, find the minimum and maximum x and y coordinates of all segmentation points to determine the minimum rectangular closure region that encloses the entire tumor region; S1.4 expands the four boundaries of the closure region outward; S1.5, truncate the expanded closure region and resample it to the same size 224*224.
3. The hepatocellular carcinoma MVI tree-like grading method based on self-supervised learning and fine-grained classification as described in claim 1, characterized in that, The batch size for training the self-supervised learning SimCLR was set to 32, using the adamw optimizer, with an initial learning rate of 1e-3 and an optimizer weight-decay of 1e-6. The learning rate update strategy uses cosine annealing; the cross-entropy loss function is used for the MVI severity classification task, and the focal loss function is used for the MVI negative-positive classification task, with alpha set to 0.4 and gamma set to 3.