Method for establishing liver cancer patient target immunity treatment effect evaluation model
Through the deep learning method based on MRI images, the 3D-Unet and 3D-CNN model combined with the Transformer module is used to achieve high accuracy assessment of the targeted immunotherapy effect of liver cancer patients, solving the problem of insufficient evaluation in the prior art, and is suitable for the formulation of personalized treatment plans.
Patent Information
- Application Number
- CN202510404818.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, when evaluating the effect of targeted immunotherapy in patients with liver cancer, the lack of comprehensive analysis of tumor heterogeneity and imaging characteristics, resulting in insufficient accuracy of efficacy prediction and limited scope of application of the method.
The deep learning method based on MRI images is adopted, and the liver and tumor regions are automatically segmented through the 3D-Unet network model, and the image features are extracted in combination with the 3D-CNN model and the clinical features are extracted, and a multimodal network model is constructed to evaluate the target treatment-free effect.
It has achieved high accuracy assessment of target treatment-free effect in liver cancer patients, reduced manual intervention, improved predictive sensitivity and specificity, and is suitable for the formulation of personalized treatment plans.
Smart Images

Figure CN120339214A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and relates to a method for establishing a target immune treatment effect evaluation model for liver cancer patients, in particular to a method for establishing a target immune treatment effect prediction and / or evaluation model for liver cancer patients based on MRI images. Background Art
[0002] Hepatocellular carcinoma (HCC) is the main type of liver malignant tumor, accounting for more than 90% of liver cancer cases. Due to the insidious onset and rapid progression of HCC, most patients are already in the advanced stage with metastasis at the time of diagnosis, missing the opportunity for surgical treatment. Currently, advanced HCC mainly relies on systemic treatment, such as targeted drugs (sorafenib, lenvatinib) and immune checkpoint inhibitors (such as PD1 monoclonal antibody). Although target immune combination therapy (targeted drug combined with immunotherapy) has shown potential in clinical trials, there are still significant problems of treatment resistance and disease progression in clinical practice, and some patients cannot benefit from it.
[0003] Traditional methods rely on clinical experience or single biomarkers to screen patients, lacking a comprehensive analysis of tumor heterogeneity and imaging features, resulting in insufficient accuracy in predicting treatment efficacy and prone to over-treatment or delayed treatment. In recent years, multiple methods have been developed in the medical field to predict or evaluate the efficacy of targeted immunotherapy for liver cancer patients. For example, the patent document with publication number CN114708977A discloses an imaging genomics feature acquisition method based on gadoxetic acid disodium-enhanced MRI for predicting the histological grade of hepatocellular carcinoma. Inclusion and exclusion criteria are set for patients with HCC pathological information on MRI, and the total number of patients eligible for imaging genomics feature acquisition is counted. Patients meeting the inclusion criteria undergo routine Gd-EOB-DTPA enhanced MRI examination; imaging genomics features are extracted from the obtained gadoxetic acid disodium-enhanced MRI. The patent document with publication number CN115240802A discloses a method for predicting the histological grade of hepatocellular carcinoma based on gadoxetic acid disodium-enhanced MRI imaging genomics. Gadoxetic acid disodium-enhanced MRI imaging genomics features are obtained; the inter-observer consistency of the extracted MRI imaging genomics features in each group is evaluated using ICC; imaging genomics features with an ICC value lower than 0.75 are excluded; the average value of the obtained consistencies is used as the final result; characteristic parameters are screened through a decision tree and a model for predicting the histological grade of hepatocellular carcinoma is constructed, including: serum AST level, 2 hepatobiliary phase imaging genomics features of AUClow and mode, and 4 arterial phase imaging genomics features of energy, inversevariance, IDMN, and maximum. The patent document with publication number CN116189761A discloses a method for accurately predicting the efficacy of DEB-TACE combined with PD-1 inhibitor for liver cancer based on multi-omics data. The patent document with publication number CN117911742A discloses a dendritic grading model for MVI in liver cancer based on self-supervised learning and fine-grained classification, characterized by including: a data preprocessing module, an MVI negative / positive classification model, and an MVI severity classification model; wherein, the data preprocessing module collects MRI images of the patient's liver to form a data set and performs preprocessing; after the MRI images are processed by the data preprocessing module, first through the trained MVI negative / positive classification model, if the classification result is negative, the grading result of this case is output as M0, if the classification result is positive, the image is continuously input into the trained MVI severity classification model to obtain a grading result of M1 or M2.The patent document with the publication number CN119170210A discloses a prediction model for predicting the combined treatment effect of liver cancer based on multi-modal MRI radiomics features and its construction method. Before combined treatment, multiple machine learning models are established through non-invasive multi-modal contrast-enhanced MRI radiomics features and clinical factors to predict the tumor treatment response of transarterial chemoembolization combined with molecular targeting and immunotherapy, and the applicability of the radiomics model is verified using an external dataset; in addition, combining this radiomics model with a clinical model has a more excellent prediction efficiency, and the radiomics features provide incremental prediction value for clinical factors. Summary of the Invention
[0004] However, these existing technical methods all require collecting specific variable factors and are only suitable for special application scenarios. The applicable range and even accuracy of predicting and evaluating the target immune treatment effect of liver cancer patients are very limited. We have made improvements in this regard and proposed a prediction and / or evaluation model for the target immune treatment effect of liver cancer patients based on MRI images and its establishment method. Specifically, the present invention provides the following technical solutions;
[0005] A method for establishing an evaluation model for the target immune treatment effect of liver cancer patients, comprising the following steps:
[0006] S1. Obtain the arterial-phase MRI images of HCC patients and perform preprocessing;
[0007] S2. Input the preprocessed MRI images into a 3D-Unet network model for training to construct an automatic liver tumor segmentation model;
[0008] S3. Input the MRI images of HCC patients receiving target immune treatment into the 3D-Unet network model to obtain liver tumor region images, and then input them into a 3D-CNN model to learn imaging features;
[0009] S4. Fuse the imaging features extracted by the 3D-CNN model from the MRI sequence data with the clinical features extracted by the Transformer module to construct a multi-modal network model, output a binary classification result of the target immune treatment efficacy, and the obtained model can be used to predict and / or evaluate the target immune treatment effect of HCC patients.
[0010] In one embodiment, the spacing of the MRI sequence is resampled to [0.83, 0.83, 3.60] and adjusted to a size of 256×256×48.
[0011] Preferably, the 3D-Unet network model is divided into 2 stages. The first stage performs liver segmentation, and the second stage performs liver tumor mass segmentation.
[0012] Furthermore, in the 3D-CNN model, the data initially needs to be processed in the convolutional layer, followed by BN, ReLU, and a max-pooling layer. Then, four residual blocks are selected in sequence and sent to the average pool. Subsequently, it is processed in the fully connected layer, and finally, the Softmax activation function is used for classification.
[0013] In one implementation, the Transformer module is used to extract features from clinical data, and the Transformer module consists of two layers of encoders.
[0014] Preferably, the MRI image is scaled using the SplineInterpolated Zoom algorithm to generate new pixel values through spline interpolation, and the liver tumor annotation is completed manually using the ITK-SNAP software.
[0015] In a specific implementation, the 3D-Unet network model consists of an encoder and a decoder. The encoder part of the 3D-Unet network model contains four basic blocks, and each block contains two convolutional layers, a normalization layer, and a LeakyReLU layer (α = 0.01). The decoder upsamples through a transposed convolutional layer (kernel 2×2×2, stride 2×2×2) and concatenates with the encoder features.
[0016] Furthermore, the residual block is connected to the global average pooling layer. The input size of the segmented tumor image is 128×128×64 and is normalized by the mean and standard deviation.
[0017] Preferably, in the two layers of encoders of the Transformer module, the hidden layer dimension is 256, the number of self-attention heads is 8, and the Dropout probability of each layer of encoder is set to 0.2, retaining 80% of the neuron outputs.
[0018] In a preferred implementation, the multi-modal network model training adopts 10-fold cross-validation and 5 rounds of iterative optimization. Each iteration randomly samples in a stratified manner. The loss function is binary cross-entropy, and the learning rate of the optimizer is set to 10 -4 , and the batch size is 8.
[0019] In the evaluation model for the efficacy of target immunotherapy for liver cancer patients established based on MRI images in the present invention, the automatic segmentation of the liver and tumor regions is achieved through a two-stage 3D-Unet network model. The segmentation results are accurate and smooth, effectively reducing the errors and time-consuming problems of manual delineation. By combining the MRI image features extracted by the 3D-CNN model with the clinical features extracted by the Transformer module, the biological characteristics of the tumor and the individual information of the patient are comprehensively captured, significantly improving the prediction sensitivity and specificity. The multi-modal network model based on deep learning can dynamically integrate imaging and clinical data, provide personalized prediction of the efficacy of target immunotherapy for patients, and assist doctors in formulating precise treatment strategies. The whole process from image preprocessing to segmentation, feature extraction, and classification is automated, reducing manual intervention and improving the efficiency of clinical decision-making. The robustness of the model is enhanced through techniques such as residual blocks, cross-validation, and Dropout, ensuring stable performance on different datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is an example diagram of the annotation analysis of liver tumors provided by the present invention;
[0021] Figure 2 It is a schematic diagram of the liver tumor segmentation pattern based on the 3D-Unet network model provided by the present invention;
[0022] Figure 3 It is an architecture diagram of the 3D-CNN model for tumor response prediction provided by the present invention;
[0023] Figure 4 It is an overall structure diagram of the multi-modal network for tumor response prediction provided by the present invention;
[0024] Figure 5 It is a partial flow schematic diagram of the method for predicting the efficacy of target immunotherapy for liver cancer patients based on MRI images provided by the present invention;
[0025] Figure 6 The formula of the Leaky ReLU activation function in the middle layer of the model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] There are many defects in the existing methods for predicting the efficacy of target immunotherapy for liver cancer patients based on MRI images. These methods are manual (similar to primitive) methods. Doctors need to delineate liver tumors by themselves and then extract traditional radiomics features through machines. To overcome these defects, we have introduced deep learning technology into the prediction of the efficacy of target immunotherapy for liver cancer patients based on MRI images and established a corresponding multi-modal deep learning model based on the convolutional neural network CNN. We use CNN to extract imaging features to replace the extraction of traditional radiomics features through machines in the existing technology.
[0027] As a branch of machine learning, deep learning does not require manual selection of relevant features. Instead, it can autonomously and adaptively identify and extract comprehensive and complex lesion information in sample images according to the learning objective. By comprehensively considering multi-layer data and using an integrated learning model, we can more accurately evaluate the biological characteristics of tumors and provide guidance for the personalized treatment of HCC. Therefore, fusing deep learning with HCC MRI images to predict the sensitivity differences of HCC patients in targeted immunotherapy combination treatment is a future research direction worthy of attention.
[0028] Therefore, based on the existing technology, we propose a method for establishing a prediction model for the efficacy of targeted immunotherapy in HCC patients, specifically a method for HCC MRI image recognition and segmentation and prediction of the efficacy of targeted immunotherapy based on deep learning, to individually judge whether HCC patients are suitable for receiving targeted immunotherapy. This method has the advantages of high regional segmentation accuracy, smooth segmentation, and high sensitivity and specificity of the prediction effect.
[0029] Compared with the existing technology, the present invention provides an end-to-end fully automated multi-modal method for predicting the efficacy of targeted immunotherapy in HCC patients, which is a true artificial intelligence and can liberate manual operations including steps such as delineating liver tumors. Briefly speaking, the method of the present invention has the following characteristics:
[0030] 1. A two-stage CNN segmentation model to complete the full-automatic segmentation of the tumor region (the existing technologies are all manual or semi-automatic methods);
[0031] 2. Based on the segmented tumor MRI images, extract the features of the images through CNN, rather than the traditionally defined radiomics features (the existing technologies are all pre-defined radiomics features);
[0032] 3. Fuse the image features extracted by CNN and the clinical features, and mine multi-modal features through a neural network.
[0033] The prediction / evaluation model and method for the efficacy of targeted immunotherapy in HCC patients constructed by the present invention are more targeted in use and are applicable to targeted combined immunotherapy, rather than being limited to special application scenarios of the existing technology such as being limited to "Tace + targeted + immunotherapy". The two are completely different application scenarios because the treatment administration methods of clinical TACE are different, and TACE treatment partially affects the effect of targeted combined immunotherapy. And our research excludes patients who have undergone TACE, and the applicable range is patients receiving targeted combined immunotherapy, with higher prediction accuracy.
[0034] The present invention proposes a method for the recognition and segmentation of liver cancer MRI images and the prediction of the efficacy of targeted immunotherapy based on deep learning, so as to individually determine whether HCC patients are suitable for receiving targeted immunotherapy. This technical solution has the advantages of high regional segmentation accuracy, smooth segmentation, high sensitivity and specificity of prediction effect, etc.
[0035] The technical solution of the present invention will be described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; and the structures shown in the drawings are only schematic and do not represent real objects.
[0036] It should be understood that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0037] Embodiment 1
[0038] Please refer to Figure 5 , the present invention provides a method for establishing an evaluation model for the efficacy of targeted immunotherapy for liver cancer patients, which is a method for predicting the efficacy of targeted immunotherapy for liver cancer patients based on MRI images, and includes the following steps:
[0039] S1. Obtain the arterial-phase MRI images of HCC patients and perform preprocessing. By preprocessing, the noise interference in the images can be removed, making the details of the images clearer. For example, features such as the boundaries and internal structures of tumors can be more clearly shown, which helps to more accurately identify and distinguish tumor tissues from normal tissues. At the same time, correcting image distortion can ensure the accuracy of the spatial information of the images, reduce the deviation of feature extraction caused by image deformation, and thus improve the accuracy of subsequent liver tumor segmentation and treatment effect evaluation;
[0040] S2. Input the preprocessed MRI images into a 3D-Unet network model for training to construct an automatic liver tumor segmentation model; compared with traditional manual segmentation methods, the automatic segmentation model can complete the segmentation tasks of a large number of images in a short time, saving a large amount of manpower and time costs. At the same time, through the training of a large amount of data, it can learn various features and patterns of liver tumors, reduce the subjectivity and inconsistency that may occur in the manual segmentation process, improve the reliability of the segmentation results. In addition, the automatic segmentation model can also provide more accurate tumor information for subsequent treatment plan formulation and efficacy evaluation, such as the size, location, volume, etc. of the tumor;
[0041] S3. Input the MRI images of HCC patients receiving target immunotherapy into the 3D-Unet network model to obtain liver tumor region images, and then input them into the 3D-CNN model to learn imaging features. The 3D-CNN model can analyze the liver tumor region images from multiple angles and levels, and extract various imaging features including the morphology, texture, density, etc. of the tumor. These features can reflect the biological characteristics and treatment responses of the tumor, helping doctors more accurately evaluate the treatment effect and predict the prognosis of patients. In addition, by learning a large number of imaging features, the 3D-CNN model can also discover some potential treatment effect evaluation indicators, providing a reference for the formulation of personalized treatment plans.
[0042] S4. Integrate the imaging features extracted by the 3D-CNN model from the MRI sequence data with the clinical features extracted by the Transformer module. The integration method can adopt the splicing method to construct a multimodal network model and output the binary classification result of the target immunotherapy efficacy to evaluate the target immunotherapy effect of HCC patients. Imaging features reflect the imaging manifestations of tumors, such as the size, morphology, density, etc. of the tumor; clinical features reflect the clinical information of patients, such as age, gender, liver function indicators, tumor markers, etc. Integrating the two can comprehensively evaluate the treatment effect from multiple angles and reduce the limitations of single-feature evaluation. The multimodal network model can capture the complex relationship between imaging features and clinical features, thus more accurately predicting the treatment effect and providing strong support for doctors to formulate personalized treatment plans.
[0043] In this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device / equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device / equipment. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical or equivalent elements in the process, method, article or device / equipment including the said element.
[0044] Further, the spacing of the MRI sequence is resampled to [0.83, 0.83, 3.60] and adjusted to a size of 256×256×48. Unify the format and size of the image data for easy processing and analysis. Resampling can make the MRI sequences of different patients have the same spacing, eliminate image differences caused by different scanning parameters, and improve the comparability of images. Adjusting the image size can ensure the consistency of all image input dimensions and reduce training difficulties and feature extraction biases caused by inconsistent sizes.
[0045] Furthermore, the 3D-Unet network model is divided into two stages. The first stage is for liver segmentation, and the second stage is for liver tumor mass segmentation. Step-by-step segmentation can significantly improve the accuracy and efficiency of segmentation. Performing liver segmentation first can narrow the scope of subsequent tumor segmentation and reduce unnecessary calculations and interference information. At the same time, the results of liver segmentation can provide more accurate contextual information for tumor segmentation, which is conducive to identifying the boundaries and locations of tumors. Staged processing can also reduce complexity, improve training stability and convergence speed, thereby speeding up the training and segmentation process.
[0046] Furthermore, in the 3D-CNN model, the data is initially processed in a convolutional layer, followed by a BN, ReLU and a maximum pooling layer, and then four residual blocks are selected in turn, and then sent to the average pool, and then processed in a fully connected layer, and finally classified using the Softmax activation function; it can effectively extract image features and accurately classify them. The convolutional layer can automatically extract local features in the image, and can capture feature information of different scales and directions through different convolution kernels. The BN layer can accelerate the convergence speed of the 3D-CNN model, improve the stability of the 3D-CNN model, and reduce the gradient during training. The ReLU layer introduces nonlinear factors to increase the expressive power of the 3D-CNN model, enabling the 3D-CNN model to learn more complex feature patterns. The maximum pooling layer can reduce the dimension of the features and reduce the amount of calculation while retaining important feature information. The residual block can solve the gradient disappearance problem in deep neural networks, allowing the 3D-CNN model to train deeper networks and extract more advanced features. The average pooling and fully connected layers can integrate and classify the extracted features, and finally output the classification probability through the Softmax activation function, which can accurately classify the treatment effect.
[0047] Furthermore, the Transformer module is used to extract features from the clinical data. The Transformer module consists of two layers of encoders. The self-attention mechanism of the Transformer module can globally model the input clinical data and automatically assign different weights to different features, thereby highlighting important feature information. The two-layer encoder can further explore the deep features in the clinical data and improve the efficiency and accuracy of feature extraction. By extracting key features from clinical data, the Transformer module can provide more valuable information for the multimodal network model, which helps to improve the accuracy of target immune treatment effect evaluation.
[0048] Furthermore, the scaling of MRI images was performed using the SplineInterpolated Zoom algorithm to generate new pixel values through spline interpolation, and liver tumor annotation was done manually using ITK-SNAP software;
[0049] SplineInterpolated Zoom, that is, spline interpolation and scaling transformation. In computer graphics and image processing, using spline interpolation and scaling transformation is an effective method to achieve image smoothing and detail enhancement. Especially in image processing, when we want to locally magnify an image, using spline interpolation can provide better visual effects than simple pixel interpolation (such as nearest-neighbor interpolation or bilinear interpolation).
[0050] Spline interpolation is a commonly used data interpolation method in mathematics and computer graphics. It approximates data by constructing a series of smooth curves (i.e., spline curves) between data points. In image processing, common spline interpolation methods include B-spline and B-spline interpolation.
[0051] When scaling an image, we can use spline interpolation to replace the traditional pixel interpolation method to obtain a smoother and more natural magnification effect.
[0052] Furthermore, the 3D-Unet network model consists of an encoder and a decoder. The encoder part of the 3D-Unet network model contains four basic blocks, and each basic block contains two convolutional layers, a normalization layer, and a Leaky ReLU layer (α = 0.01); the decoder upsamples through a transposed convolutional layer (kernel 2×2×2, stride 2×2×2) and concatenates with the encoder features; the convolutional layer can automatically extract local features of the image, the normalization layer can normalize the input data, making the data distribution more stable, thus accelerating the convergence of the 3D-Unet network model; the non-linear characteristic of the Leaky ReLU layer can introduce more feature combinations and improve the expression ability of the 3D-Unet network model; the transposed convolutional layer can implement the upsampling operation to restore the spatial information of the image; feature concatenation can combine high-level features and low-level features extracted by the encoder to improve the accuracy of segmentation.
[0053] Furthermore, the residual block is connected to the global average pooling layer. The input size of the segmented tumor image is 128×128×64 and is normalized by the mean and standard deviation; the global average pooling layer can convert the high-dimensional feature map into a low-dimensional vector, reducing the data dimension, reducing the computational amount, and improving the training efficiency; normalization can make the data have the same scale, reducing training difficulties and gradient vanishing problems caused by data scale differences.
[0054] Furthermore, in the two layers of the Transformer module encoder, the hidden layer dimension is 256, the number of self-attention heads is 8, and the Dropout probability of each layer of the encoder is set to 0.2, retaining 80% of the neuron outputs; a hidden layer dimension of 256 can provide sufficient capacity to learn complex features in clinical data, and 8 self-attention heads can focus on the input data from different perspectives, improving the efficiency and accuracy of feature extraction;
[0055] Furthermore, the training of the multi-modal network model uses 10-fold cross-validation and 5 rounds of iterative optimization. Each iteration uses random stratified sampling, the loss function is binary cross-entropy, and the learning rate of the optimizer is set to 10 -4 , and the batch size is 8; 10-fold cross-validation can make full use of the data for evaluation and training, reducing the risk of overfitting; by dividing the data into 10 subsets and using 9 of them for training and 1 for validation each time, more accurate evaluation results can be obtained; 5 rounds of iterative optimization can gradually converge to the optimal solution; random stratified sampling can ensure that the data distributions of each training set and validation set are similar, reducing the training bias caused by uneven data distribution; the binary cross-entropy loss function is suitable for binary classification problems and can effectively measure the difference between the prediction result and the true label.
[0056] Example 2
[0057] Please refer to Figures 1-4 , and the specific applications of the HCC patient target immune therapy effect evaluation model and the HCC patient target immune therapy effect prediction method based on MRI images provided by the present invention are as follows:
[0058] I. Preprocess the images and manually draw the outlines;
[0059] 1. Resample the spacing of the MRI sequence to [0.83, 0.83, 3.60]; this makes the spatial resolutions of MRI images from different sources reach a unified standard, greatly improving the comparability between images;
[0060] 2. Use the SplineInterpolated Zoom scaling method to adjust the MRI images to a size of 256×256×48, and use the spline interpolation algorithm to generate new pixel values when scaling the images;
[0061] 3. Manually use the ITK-SNAP software to complete the liver tumor contour annotation; relying on the medical knowledge and experience of professionals, ensure the accuracy and reliability of the liver tumor annotation, provide accurate annotation data for subsequent training, and improve the accuracy of liver tumor recognition and segmentation;
[0062] II. Construction and training of the MRI image recognition and segmentation model for HCC patients;
[0063] 1. Use the preprocessed MRI images as the dataset input to the 3D-Unet network model;
[0064] 2. The 3D-Unet network model consists of an encoder and a decoder;
[0065] 3. The encoder part of the 3D-Unet network model includes four consecutive basic blocks. Each basic block contains two convolutional layers, followed by a normalization layer and a Leaky ReLU layer respectively; the number of feature channels of the four basic blocks are set to 30, 60, 120, and 240 respectively; the formula of the Leaky ReLU activation function in the middle layer of the 3D-Unet network model is as Figure 6 shown, (where, is set to 0.01):
[0066] 4. The decoder part of the 3D-Unet network model consists of three transposed convolutional layers. The transposed convolutional layer expands the spatial size of the feature map through the sliding operation of the convolutional kernel; the convolutional kernel size is 2x2x2, the stride is 2x2x2, and then followed by a basic block; the feature map is upsampled by a factor of 2 through the transposed convolutional layer, and then connected to the corresponding feature map from the encoder, introducing low-level detail information into the high-level features, and then input into the basic block; the number of feature channels of the basic blocks in the decoder module are 120, 60, and 30 respectively;
[0067] 5. The loss function is a combination of the cross-entropy loss function and the Dice Loss, and the mixing of the two loss functions each accounts for a weight of 0.5; combining the cross-entropy loss function and the Dice Loss can measure the difference between the prediction result and the true label from different angles, improving the accuracy and robustness of the segmentation;
[0068] 6. Based on the above loss function, use the Adam optimizer to adjust the weights; the initial learning rate is set to 0.001, the training batch size is set to 2, and the number of training epochs for each stage is set to 100;
[0069] III. 3D-CNN predicts the target immune efficacy of HCC patients;
[0070] 1. Readjust the liver tumor region segmented from the above MRI into a 3D volume of 128×128×64, and perform image normalization processing through the mean and standard deviation of the input sequence;
[0071] 2. The 3D-CNN model contains four residual blocks. The first residual block consists of 64 convolutional filters, the second residual block consists of 128 convolutional filters, the third residual block consists of 256 convolutional filters, and the fourth residual block consists of 512 convolutional filters. After the four residual blocks, a global average pooling layer is connected. It calculates the average value of all elements within each specific 1×1×1 region of the input feature map and uses it as the value of the output feature map at the corresponding position. Its output is passed to the fully connected layer, and the binary classification task is completed through the fully connected layer and the Softmax function.
[0072] 3. The multi-modal network is trained and validated through a 10-fold cross-validation scheme. The entire dataset is randomly divided into 10 subsets containing 18 elements each. Each subset is used as a validation set once, while the other 9 subsets are used as training sets for training. After training, the entire dataset is randomly grouped again, and the above operations are repeated for a total of 5 iterations. The data for each fold is randomly stratified sampled.
[0073] 4. The binary cross-entropy loss function is selected as the loss function. The binary cross-entropy loss function is suitable for binary classification problems, can accurately measure the difference between the prediction results of the 3D-CNN model and the true labels, effectively guide the 3D-CNN model to adjust parameters, and improve the accuracy of predicting the target immune efficacy of HCC patients.
[0074] 5. The Adam optimization algorithm is used to optimize the network weights, and the learning rate is set to 10-4. During training, the batch size is fixed at 8, and the number of training epochs is 100.
[0075] IV. Multi-modal prediction of the target immune efficacy of HCC patients;
[0076] 1. Clinical features are extracted through the transformer module. The Transformer module contains two encoder layers, where the hidden layer size is 256 and the number of self-attention heads is 8. The Dropout probability in the encoder layer is set to 0.2, indicating that the outputs of 80% of the neurons are retained and 20% are discarded (set to 0).
[0077] 2. The 3D-CNN model and the Transformer module each output a 512-dimensional vector, which contain the information and features extracted from the MRI image sequence and clinical data respectively. The vectors output by the 3D-CNN model and the Transformer module are concatenated into a 1024-dimensional vector, and are fused and dimensionality-reduced through the fully connected layer, and converted into a 256-dimensional vector. The softmax output layer is used to predict the target immune efficacy of HCC patients.
[0078] 3. The 10-fold cross-validation and 5 rounds of iteration steps are as described above.
[0079] As described above, it is only the preferred specific implementation manner of the present invention; however, the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its improved concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A method for establishing a target immune therapy effect evaluation model for liver cancer patients, characterized in that, It includes the following steps: S1. Obtain the arterial-phase MRI images of HCC patients and perform preprocessing; S2. Input the preprocessed MRI images into the 3D-Unet network model for training to construct an automatic liver tumor segmentation model; S3. Input the MRI images of HCC patients receiving target immunotherapy into the 3D-Unet network model to obtain liver tumor region images, and then input them into the 3D-CNN model to learn imaging features; S4. Fuse the imaging features extracted by the 3D-CNN model from the MRI sequence data with the clinical features extracted by the Transformer module to construct a multimodal network model, output the binary classification result of the target immunotherapy efficacy, and obtain a model for predicting or evaluating the target immunotherapy effect of HCC patients.
2. The method according to claim 1, characterized in that The spacing of the MRI sequence is resampled to [0.83, 0.83, 3.60] and adjusted to a size of 256×256×48.
3. The method according to claim 1, characterized in that, The 3D-Unet network model is divided into two stages. The first stage performs liver segmentation, and the second stage performs liver tumor mass segmentation.
4. The method according to claim 1, characterized in that, In the 3D-CNN model, the data initially needs to be processed in the convolutional layer, followed by BN, ReLU, and a max-pooling layer. Then, four residual blocks are selected in sequence, and then sent to the average pool. Subsequently, it is processed in the fully connected layer, and finally, the Softmax activation function is used for classification.
5. The method according to claim 1, wherein The Transformer module is used to extract features from clinical data. The Transformer module consists of two layers of encoders.
6. The method according to claim 2, characterized in that, The scaling of the MRI images adopts the SplineInterpolated Zoom algorithm to generate new pixel values through spline interpolation, and the liver tumor annotation is completed manually using the ITK-SNAP software.
7. The method according to claim 3, characterized in that, The 3D-Unet network model consists of an encoder and a decoder. The encoder part of the 3D-Unet network model contains four basic blocks, and each block contains two convolutional layers, a normalization layer, and a LeakyReLU layer (α = 0.01); The decoder upsamples through a transposed convolutional layer (kernel 2×2×2, stride 2×2×2) and concatenates with the encoder features.
8. The method according to claim 4, characterized in that, The residual block is connected to the global average pooling layer. The input size of the segmented tumor image is 128×128×64, and it is normalized by the mean and standard deviation.
9. The method according to claim 5, characterized in that, In the two layers of encoders of the Transformer module, the hidden layer dimension is 256, the number of self-attention heads is 8, and the Dropout probability of each layer of the encoder is set to 0.2, retaining 80% of the neuron outputs.
10. The method according to claim 1, characterized in that, The training of the multimodal network model uses 10-fold cross-validation and 5 rounds of iterative optimization. Random stratified sampling is performed in each iteration. The loss function is binary cross-entropy, and the learning rate of the optimizer is set to 10 -4 , and the batch size is 8.
Citation Information
Patent Citations
Gadoxetate disodium enhanced MRI (Magnetic Resonance Imaging)-based radiomics characteristic acquisition method for predicting histological grade of hepatocellular carcinoma
CN114708977A
Method for predicting histological grade of hepatocellular carcinoma based on gadoxate disodium enhanced MRI (Magnetic Resonance Imaging) radiomics
CN115240802A
Method and device for accurately predicting curative effect of DEB-TACE combined PD-1 inhibitor of liver cancer based on multi-omics data
CN116189761A
Liver cancer MVI tree-shaped grading model based on self-supervised learning fine granularity recognition
CN117911742A
Prediction model for predicting liver cancer combined treatment effect based on multi-mode MRI radiomics characteristics and construction method and application thereof
CN119170210A