Liver cancer immunotherapy efficacy evaluation method fusing imageomics and deep learning
By constructing a multimodal image fusion network and a multi-source data coupling model, the problems of subjectivity and insufficient feature mining in the evaluation of liver cancer immunotherapy in existing technologies are solved. This enables objective quantitative evaluation of the liver cancer immune microenvironment and dynamic monitoring of efficacy, supporting individualized treatment decisions.
Patent Information
- Application Number
- CN202610444686.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-07
AI Technical Summary
Current technologies for evaluating liver cancer immunotherapy rely on manual image reading and have limited evaluation indicators, making it difficult to capture microscopic features during the immunotherapy process. Furthermore, the depth of radiomics information mining is insufficient, failing to accurately reflect the dynamic evolution of the tumor immune microenvironment.
A multimodal image fusion network and a multi-source data coupling model were constructed. Through multi-scale deep feature extraction, cross-modal attention mechanism and hierarchical graph neural network, an objective quantitative assessment of the immune microenvironment status of liver cancer was achieved. The results were then combined with imaging features and clinical information for comprehensive analysis.
It enables efficient and automated evaluation of the efficacy of immunotherapy for liver cancer, can identify the spatial distribution characteristics of immune cells and dynamically monitor the evolution of efficacy, and provides reliable support for individualized treatment decisions.
Smart Images

Figure CN122348049A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of medical image processing and artificial intelligence, specifically involving a method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning. Background Technology
[0002] With the synergistic evolution of precision medicine and medical imaging technology, medical imaging-assisted diagnosis plays an increasingly crucial role in the entire process of tumor diagnosis and treatment. Liver cancer, a highly prevalent malignant tumor worldwide, has seen its treatment shift from traditional radiotherapy and chemotherapy to personalized comprehensive treatment models represented by immunotherapy. In this process, accurately depicting the growth, evolution, tissue morphology, and metabolic characteristics of tumors through imaging techniques has become a core element in assisting clinicians in making treatment decisions, monitoring disease progression, and predicting patient prognosis.
[0003] Intelligent assessment technologies integrating radiomics and deep learning aim to reveal tumor heterogeneity information that is difficult to identify with the naked eye through in-depth mining and quantitative analysis of multidimensional medical images. This type of technology utilizes high-performance convolutional neural networks to automatically extract high-dimensional features from heterogeneous image data such as CT and MRI, and combines this with mathematical modeling to establish the correlation between image phenotypes and the tumor microenvironment. Its fundamental principle lies in leveraging the powerful feature generalization capabilities of artificial intelligence to transform complex visual information into quantifiable assessment indicators, providing more objective decision-making basis for clinicians.
[0004] Current technologies still face significant challenges in evaluating liver cancer immunotherapy. Traditional imaging assessment methods rely heavily on radiologists' manual interpretation of images, resulting in outcomes limited by expert subjective experience. Furthermore, evaluation indicators often focus on macroscopic morphological changes such as tumor diameter, failing to capture complex microscopic features like inflammatory responses and immune cell infiltration during immunotherapy. Simultaneously, conventional feature extraction algorithms lack sufficient depth in mining radiomics information, leading to underutilization of semantic complementarity between multimodal images and an inability to accurately reflect the dynamic evolution of the tumor immune microenvironment. Existing models often lack the ability to efficiently couple deep imaging features with multi-source clinical data, resulting in a single evaluation dimension and limited predictive accuracy, failing to meet the clinical needs for predicting the efficacy and optimizing treatment regimens in liver cancer immunotherapy.
[0005] Therefore, a scheme for evaluating the efficacy of liver cancer immunotherapy that integrates radiomics and deep learning is desired. Summary of the Invention
[0006] The purpose of this invention is to provide a method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning, which can solve the problems in the background art mentioned above.
[0007] To achieve the above objectives, the technical solution adopted by this invention is: a method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning, comprising the following specific steps: Step 1: Obtain multimodal medical imaging data and corresponding clinical information data of liver cancer patients. The multimodal medical imaging data includes at least enhanced computed tomography (CT) images and magnetic resonance imaging (MRI) images. The clinical information data includes the patient's basic physiological indicators, pathological staging information, and previous treatment records. Step 2: Perform standardized preprocessing on the multimodal medical image data, including spatial registration, grayscale normalization, and automatic lesion region segmentation, to generate a multimodal image volume dataset in a unified coordinate system; Step 3: Construct a multi-scale deep feature extraction network to independently extract high-dimensional image feature vectors from the multimodal image volume dataset, and perform semantic alignment and weight allocation on the feature vectors of different modalities through a cross-modal attention mechanism; Step 4: Fuse the semantically aligned multimodal image feature vectors to generate a comprehensive image representation vector, and construct a multi-source heterogeneous feature matrix by combining the clinical information data; Step 5: Based on the multi-source heterogeneous feature matrix, train and deploy a liver cancer immunotherapy efficacy prediction model. The model adopts a hierarchical structure design, with the bottom layer used to identify the state of the tumor immune microenvironment and the upper layer used to output efficacy evaluation results. Step 6: Input the multimodal medical imaging data and clinical information data of the patient to be evaluated into the efficacy prediction model to obtain quantitative evaluation results, including the probability of immune response, the trend of tumor burden change, and the treatment response level.
[0008] Preferably, the acquisition of multimodal medical image data in step 1 follows a unified time window specification to ensure the temporal consistency of each modality of image within the treatment cycle, and the image resolution meets the requirement that the internal texture details of the lesion can be identified.
[0009] Preferably, in step 2, the automatic segmentation of the lesion area adopts a cascaded three-dimensional fully convolutional neural network architecture. The first-level network is used to roughly locate the liver region, and the second-level network focuses on the fine delineation of the boundaries of the lesions in the liver. The segmentation results are post-processed morphologically to eliminate isolated noise points.
[0010] Preferably, the multi-scale deep feature extraction network in step 3 contains multiple parallel residual coding paths, each path corresponding to a different receptive field scale, used to capture multi-level image semantic information from local texture to global structure.
[0011] Preferably, the cross-modal attention mechanism dynamically adjusts the contribution weight of each modality in the fusion process by calculating the correlation scores between different modal feature channels, thereby enhancing the sensitivity to immune-related image phenotypes.
[0012] Preferably, in step 4, the construction of the multi-source heterogeneous feature matrix adopts a feature embedding mapping strategy, which transforms discrete clinical variables into continuous vector representations and concatenates them with continuous image feature vectors to form a unified input format.
[0013] Preferably, the underlying module of the liver cancer immunotherapy efficacy prediction model in step 5 adopts a graph neural network structure, which divides the tumor region into multiple sub-region nodes and models the spatial distribution pattern of immune cell infiltration through the connection relationship between nodes.
[0014] Preferably, the upper-level module adopts a gated loop unit structure, which integrates the evaluation results of historical treatment stages to achieve dynamic tracking and prediction of the evolution trend of efficacy.
[0015] Preferably, the treatment response level in the quantitative assessment results is divided into four categories: complete response, partial response, disease stability, and disease progression. The determination is based on the joint calibration of changes in imaging features and clinical endpoint events.
[0016] Preferably, the method further includes an online update mechanism for the efficacy prediction model. When the cumulative number of newly labeled samples reaches a predetermined number, incremental retraining of the model parameters is triggered to adapt to changes in new drug regimens or patient groups in clinical practice.
[0017] Compared with the prior art, the present invention has the following beneficial effects: This invention achieves an objective quantitative assessment of the immune microenvironment status of liver cancer by constructing a multimodal image fusion network and a multi-source data coupling model, overcoming the technical shortcomings of traditional manual image reading, which is highly subjective and lacks sufficient feature mining. The cross-modal attention mechanism adopted effectively improves the utilization efficiency of semantic complementarity between different image modalities and enhances the ability to identify immunotherapy-specific image phenotypes. This invention, by introducing a hierarchical modeling strategy using graph neural networks and gated recurrent units, not only characterizes the spatial distribution features of immune cells within tumors but also enables continuous monitoring of the dynamic evolution of therapeutic efficacy. The overall solution is highly automated and scalable, and can be seamlessly integrated into clinical diagnosis and treatment processes, providing reliable technical support for individualized decision-making in liver cancer immunotherapy. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention; Figure 2This is a schematic diagram illustrating the core principle framework of the multi-scale deep feature extraction and cross-modal attention mechanism in this invention; Figure 3 This is a flowchart illustrating the logical flow of multimodal image standardization preprocessing and cascaded three-dimensional fully convolutional neural network segmentation in this invention. Figure 4 This is a schematic diagram of the core principle framework of the hierarchical efficacy prediction model based on graph neural networks and gated recurrent units in this invention; Figure 5 This is a schematic diagram of the multi-level interaction relationship and data flow of the multi-source heterogeneous feature matrix construction and quantitative evaluation results output in this invention. Detailed Implementation
[0019] Example 1: Please refer to the appendix Figure 1 To be continued Figure 5 To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments.
[0020] In the aforementioned method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning, step 1 specifically involves acquiring multimodal medical imaging data and corresponding clinical information data of liver cancer patients. In practice, the multimodal medical imaging data includes at least enhanced computed tomography (CT) images and magnetic resonance imaging (MRI) images. Enhanced CT images cover volumetric data across multiple phases, including the arterial phase, portal venous phase, and equilibrium phase, clearly reflecting the tumor's blood supply characteristics and enhancement patterns. MRI images include, but are not limited to, T1-weighted imaging, T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient maps. These sequences provide rich soft tissue contrast, capable of characterizing cell density and water molecule movement restriction within the tumor.
[0021] Furthermore, the acquisition of the multimodal medical imaging data follows a unified time window standard. Specifically, for patients undergoing immunotherapy, the acquisition time difference between different modalities is controlled within 48 hours to ensure a high degree of temporal consistency in the physiological state reflected by the images. Regarding image resolution, the slice thickness for enhanced computed tomography is set between 1 mm and 2 mm, and the interslice spacing for magnetic resonance imaging is set between 3 mm and 5 mm, with an in-plane resolution of no less than 0.7 mm × 0.7 mm, ensuring that the fine textures, capsule integrity, and edge infiltration features within the lesions are clearly discernible in digital space.
[0022] The acquisition of the clinical information data involves in-depth analysis of the electronic medical record system. This clinical information data includes the patient's basic physiological indicators, such as age, sex, height, weight, and body mass index; pathological staging information includes tumor size, number, vascular invasion, lymph node metastasis status, and distant metastasis markers, specifically following the Barcelona hepatocellular carcinoma staging criteria or the Chinese hepatocellular carcinoma staging criteria. In addition, past treatment records cover surgical resection history, transcatheter arterial chemoembolization history, ablation therapy history, and the dosage and duration of targeted drug use. To ensure data completeness, laboratory test indicators also need to be collected, including alpha-fetoprotein levels, liver function classification parameters such as alanine aminotransferase (ALT), aspartate aminotransferase (AST), total bilirubin, and albumin levels, as well as immune-related peripheral blood indicators such as the neutrophil-to-lymphocyte ratio.
[0023] In the above method, step 2 specifically involves standardizing the multimodal medical image data. The first step in the preprocessing is spatial registration, which transforms images acquired by different devices and at different time points into a unified anatomical coordinate system. A rigid registration algorithm based on mutual information is used as the initial alignment method, followed by a non-rigid registration technique based on free shape deformation to correct liver morphological deformation caused by respiratory motion or patient position differences. In grayscale normalization processing, for computed tomography images, the original Henle unit values are linearly mapped to a preset grayscale range, and extreme values exceeding the enhancement range of normal liver and tumors are truncated; for magnetic resonance imaging images, a histogram-based matching method or a Z-score normalization method is used to eliminate signal intensity differences generated by different scanners by subtracting the mean and dividing by the standard deviation.
[0024] In step 2, the automatic segmentation of the lesion region employs a cascaded three-dimensional fully convolutional neural network architecture. This architecture comprises two logically related sub-networks. The first-level network is configured as a coarse localization network, taking downsampled whole abdominal body data as input. It utilizes deep convolutional layers to extract global spatial features, aiming to quickly locate the boundary contours of the liver from complex abdominal tissue and generate a mask for the liver region. The second-level network is a fine segmentation network, taking the liver region segments determined by the first-level network as input. This network adopts an encoder-decoder structure, introducing multi-scale jump connections to focus on the fine delineation of intrahepatic lesions. At the encoder end, the receptive field is progressively increased through continuous convolution and pooling operations; at the decoder end, resolution is restored through transposed convolution or upsampling operations, and fused with features from the corresponding level of the encoder to preserve local detail information. The segmentation results then enter the morphological post-processing stage, applying opening and closing operations to eliminate isolated holes or small noise points, and extracting the largest target region as the final lesion mask through connected component analysis. Finally, the system generates a multimodal image volume dataset in a unified coordinate system, where each voxel contains the feature response value of the corresponding modality.
[0025] In the above method, step 3 specifically involves constructing a multi-scale deep feature extraction network. This network comprises multiple parallel residual coding paths. Each path is designed with a specific receptive field scale. The first path uses a small 3×3×3 convolutional kernel to capture local lesion texture features, such as fine calcifications or micro-necrotic foci within the tumor; the second path uses a medium 5×5×5 convolutional kernel to perceive the integrity of the tissue structure; and the third path uses dilated convolutions to acquire large-scale global structural information, such as the spatial relationship between the tumor and surrounding large blood vessels. Within each path, residual connections are used to element-wise sum the input features with the nonlinearly transformed features to alleviate the gradient vanishing problem in deep networks.
[0026] Subsequently, a cross-modal attention mechanism is used to semantically align and weight the feature vectors of different modalities. The working principle of this cross-modal attention mechanism is as follows: First, the enhanced computed tomography (CT) feature map and the magnetic resonance imaging (MRI) feature map are mapped into query vector space, key vector space, and value vector space, respectively. A preliminary relevance score is obtained by calculating the dot product between the CT query vector and the MRI key vector, and then scaling it by the square root of the feature dimension. This relevance score is processed by a normalized exponential function to form an attention weight distribution map. This attention weight distribution map reflects the semantic relevance of different modalities at the same anatomical location. Finally, the attention weights are weighted and summed with the corresponding modal value vectors to achieve feature enhancement and alignment. This mechanism dynamically adjusts the contribution weights of each modality in the fusion process. If a region shows more discriminative immune cell infiltration features in the MRI image, the system will automatically strengthen the feature weight of that region under that modality, enhancing the sensitivity to immune-related image phenotypes.
[0027] In the above method, step 4 specifically involves fusing the semantically aligned multimodal image feature vectors to generate a comprehensive image representation vector. The fusion process employs a multi-level cascade strategy. First, the aligned modal features are concatenated along the channel dimension. Then, dimensionality compression and nonlinear interaction are performed through a 1×1×1 convolutional layer to extract common feature representations across modalities.
[0028] For the aforementioned clinical information data, step 4 employs a feature embedding mapping strategy. For discrete variables, such as gender or past treatment history, one-hot encoding or a pre-trained embedding layer is used to transform them into high-dimensional continuous vector representations. For continuous variables, such as age or laboratory indicators, after min-max standardization, they are directly used as vector components. Subsequently, the transformed clinical feature vectors are concatenated with the continuous comprehensive image feature vectors to form a multi-source heterogeneous feature matrix. During the concatenation process, a feature balancing factor is introduced. By multiplying the image features and clinical features by their respective scaling coefficients, the balance between the two types of features in terms of dimensions and numerical distribution is ensured, constructing a unified input format that can comprehensively characterize the patient's state.
[0029] In the above method, step 5 specifically involves training and deploying a liver cancer immunotherapy efficacy prediction model based on the multi-source heterogeneous feature matrix. This model employs a hierarchical structure, logically divided into a bottom-level microenvironment identification module and an upper-level efficacy output module. The bottom-level module uses a graph neural network structure, its core logic being to divide the tumor region into multiple sub-region nodes with anatomical significance or geometric features. Each node carries the local imaging features and clinical mapping features of that region. By calculating the Euclidean distance or cosine similarity between nodes, an adjacency matrix is constructed to model the connection relationships between nodes. In graph convolution operations, the features of each node are aggregated and updated according to the adjacency matrix and the features of its neighboring nodes. Specifically, the next layer feature of the current node is equal to the nonlinear transformation result of its own features and the features of its neighboring nodes after weighted averaging. Through multi-layer graph convolution, the model can effectively capture the spatial distribution patterns of immune cell infiltration within the tumor, such as marginal or diffuse infiltration.
[0030] The upper-level module employs a gated recurrent unit (ROU) structure designed to process efficacy evolution data with time-series attributes. This structure includes a reset gate and an update gate. The update gate controls the proportion of historical efficacy features retained from the previous time step, while the reset gate determines the fusion ratio between current and historical information. By integrating imaging assessment results and clinical feedback from historical treatment phases, the gated recurrent unit can learn the patterns of efficacy changes over time, enabling dynamic tracking and prediction of efficacy trends. During model training, a cross-entropy loss function is used as the optimization objective, and the network weights are adjusted through backpropagation to minimize the difference between predicted values and actual efficacy labels.
[0031] In the above method, step 6 specifically involves inputting the multimodal medical imaging data and clinical information data of the patient to be evaluated into the efficacy prediction model to obtain quantitative evaluation results. These quantitative evaluation results are generated through the model's Softmax output layer and include indicators across multiple dimensions. First, there is the probability of immune response, represented as a continuous value between 0 and 1; the closer the value is to 1, the higher the likelihood of a positive response from the patient to the current immunotherapy regimen. Second, there is the trend of tumor burden change, which outputs the expected percentage change in tumor volume or metabolic activity by comparing the characteristic differences between baseline and follow-up images.
[0032] Finally, the treatment response level is determined. Based on internationally accepted criteria for evaluating the efficacy of treatment in solid tumors, and after jointly calibrating changes in imaging features with clinical endpoint events, the response level is divided into four categories. Category 1 is complete response, defined as the disappearance of all target lesions with no new lesions appearing, and the duration meeting a preset threshold. Category 2 is partial response, defined as a reduction in the total diameter of target lesions of 30% or more. Category 3 is stable disease, defined as a reduction in target lesion diameter that does not reach the level of partial response, and an increase that does not reach the level of disease progression. Category 4 is disease progression, defined as an increase in the total diameter of target lesions exceeding 20%, or the appearance of one or more new lesions.
[0033] Furthermore, the method described in this embodiment also includes an online update mechanism for the efficacy prediction model. During clinical application, the system continuously collects newly labeled samples that have been validated against the gold standard. When the number of newly added samples accumulates to a predetermined threshold (e.g., 50 or 100 cases), the system triggers an incremental retraining process. During retraining, a learning rate decay strategy is adopted. While retaining the old model's ability to recognize general features, the model parameters are finely adjusted to adapt it to new clinical drug regimens, more advanced imaging equipment, or specific patient group characteristics.
[0034] Example 2: In another preferred embodiment, the present invention provides an implementation scheme for evaluating the efficacy of liver cancer immunotherapy in a distributed clinical environment. In this embodiment, the execution of each step has been deeply optimized in terms of hardware configuration and parallel computing logic.
[0035] For step 1, the data acquisition process involves standardized interface integration with the hospital's internal image archiving and communication system. The system automatically extracts image sequences that meet specific sequence requirements using the DICOM query retrieval protocol. Regarding clinical information retrieval, natural language processing technology is used to parse unstructured text from discharge summaries and pathology reports, extracting key named entities such as tumor differentiation degree and positive microvascular invasion, and transforming them into structured feature vectors.
[0036] For the preprocessing in step 2, a multi-threaded parallel acceleration strategy is adopted to improve processing efficiency. When segmenting lesions using a 3D fully convolutional neural network, the 3D volumetric data is divided into multiple overlapping sub-blocks. Each sub-block is independently input into the parallel computing unit of the graphics processing unit. After each sub-block completes the convolution operation, a linear weighted fusion algorithm is used to smooth the probability map of the overlapping region, eliminating the abrupt breaks at the sub-block boundaries. The morphological post-processing is executed in a multi-core environment of the central processing unit, and multi-level caching optimization reduces frequent memory swapping of large volumetric data.
[0037] For the multi-scale feature extraction in step 3, this embodiment introduces an attention feature pyramid structure. In the residual encoding path, multi-scale convolutions are performed not only at the same level but also lateral connections are established between feature maps at different depths. Strong semantic features at higher levels are fused with strong spatial features at lower levels through upsampling. In the cross-modal attention mechanism, a multi-head attention pattern is introduced. By dividing the feature channels into multiple subspaces and independently calculating relevance scores within each subspace, the model can capture complementary information between modalities from different attribute dimensions (such as brightness, texture, and contrast).
[0038] For feature fusion in step 4, this embodiment employs a nonlinear mapping technique based on manifold learning. Since multimodal imaging features and clinical features often reside in different high-dimensional manifold spaces, a local tangent space permutation algorithm is constructed to map features from different sources into a common low-dimensional manifold space. Within this space, the Euclidean distance between features accurately reflects the similarity of patients' immune status. The embedding process of clinical variables incorporates prior weights from a domain expert knowledge base; for clinically recognized strong predictors (such as PD-L1 expression), higher initial weight coefficients are assigned in the fusion matrix.
[0039] For the model construction in step 5, the graph neural network of the bottom-level module adopts a dynamic graph construction strategy. Instead of being limited to static anatomical connections, it dynamically updates the edge weights between nodes in each training iteration based on feature similarity. This approach can capture the heterogeneous changes in the immune microenvironment caused by tumor necrosis and hemorrhage. The gated recurrent unit of the upper-level module introduces a bidirectional recurrent mechanism, utilizing historical data to predict the future and also using subsequent follow-up data to retrospectively calibrate the previous evaluation results.
[0040] For the quantitative assessment in step 6, the system provides an interactive visualization interface. In addition to outputting the response probability and response level, the system can also mark high-risk areas of interest to the model on the 3D image in the form of a heatmap. For example, if the model determines that a certain area has a high trend of negative tumor growth, that area will be highlighted in red on the enhanced computed tomography image, along with a description of the imaging characteristics of the area, such as abnormally increased enhancement or blurred edges.
[0041] Example 3: In yet another preferred embodiment, the present invention focuses on improving the robustness of the model and adapting to multi-center data.
[0042] In step 1, considering the differences in scanning protocols among different medical institutions, the system introduces an automatic image quality evaluation module. For images with excessive noise, insufficient contrast, or severe motion artifacts, the system will automatically issue a warning and prompt image enhancement processing. The image enhancement process employs denoising and super-resolution reconstruction techniques based on generative adversarial networks to map low-quality images to high-quality standard images, reducing the impact of front-end data differences on subsequent feature extraction.
[0043] In the segmentation task of step 2, for some liver cancer lesions with highly irregular shapes and severe adhesion to surrounding normal tissue, this embodiment adds an active contour evolution mechanism to the cascaded network. Using the probability map generated by deep learning as the initial energy field, and through iteratively solving partial differential equations, the segmentation boundary evolves towards the position of maximum image gradient, achieving precise fitting of complex lesion edges.
[0044] In step 3, to address the computational overhead of cross-modal attention mechanisms, this embodiment employs a spatially sparse attention strategy. The system only performs correlation calculations on voxels within the liver and lesion regions, ignoring irrelevant background areas, thus reducing the computational complexity from the quadratic level of the input size to the linear level. Simultaneously, a feature alignment loss based on contrastive learning is introduced. By bringing closer the distance between different modal features of the same patient and widening the distance between features of different patients, the network is forced to learn robust features specific to each individual.
[0045] In step 4, a missing value compensation mechanism is introduced into the construction of the multi-source heterogeneous feature matrix. In clinical practice, some patients may lack certain laboratory indicators. The system utilizes a pre-trained variational autoencoder to infer and fill in the missing values based on the probability distribution of other existing clinical and imaging features, ensuring the integrity of the input matrix and avoiding model failure due to missing data.
[0046] In step 5, during model training, transfer learning and domain adaptation techniques were employed to address the risk of overfitting caused by small sample annotations. First, pre-training was performed on a large-scale general liver tumor image dataset to acquire basic visual feature perception capabilities. Subsequently, fine-tuning was performed on a specialized dataset for liver cancer immunotherapy. During fine-tuning, the maximum mean difference operator was used as a constraint to reduce the distributional differences between the source and target domain data in the feature space, thereby improving the model's generalization performance on data with different center points.
[0047] In the evaluation results output of step 6, the system adds an uncertainty estimation function. Monte Carlo random sampling technique is used to evaluate the stability of the model's output under the current input. If the model's prediction of a patient's response probability fluctuates significantly, the system will output a lower confidence score and suggest that clinicians combine it with other methods such as biopsy for comprehensive judgment.
[0048] Furthermore, the online update mechanism in this embodiment incorporates a federated incremental learning framework. When multiple collaborating medical institutions generate new data, there is no need to upload the original data to the central server; gradient calculations can be performed locally. The central server aggregates the gradient update values from each node, performs a weighted average, updates the global model parameters, and then distributes them to each node. This approach achieves continuous evolution of model performance while protecting patient privacy.
[0049] In the above embodiments, all numerical calculation logic involved has been described in detail in text form. For example, in the standardization process of feature vectors, the specific operation is to subtract the arithmetic mean of the feature in the sample set from the original value, and then divide the resulting difference by the standard deviation of the feature. In the application of activation functions in neural networks, for linear rectifier units, the logic is: when the input value is greater than 0, the output value is equal to the input value; when the input value is less than or equal to 0, the output value is equal to 0. For normalized exponential functions, the logic is: calculate the exponent of each element in the input vector to the base of the natural constant, and then divide each exponent by the sum of all exponents to obtain the probability distributions.
[0050] In multi-scale convolution, the so-called receptive field expansion is achieved by inserting a predetermined number of zeros between the kernel elements, thus increasing the sampling coverage without increasing the number of parameters. In the feature fusion concatenation operation, two or more one-dimensional feature vectors are concatenated end-to-end in a predetermined order to form a new vector with a length equal to the sum of the lengths of the individual sub-vectors. In the calculation of the loss function, cross-entropy logically manifests as: a weighted sum of the logarithms of the true labels and the predicted probabilities, with the negative of the sum, representing the degree of deviation between the predicted and true distributions.
[0051] Through the synergistic cooperation of the above steps, the liver cancer immunotherapy efficacy evaluation method integrating radiomics and deep learning achieved by this invention can mine deep-seated immune-related phenotypes from multi-dimensional imaging data, and combine them with multi-source clinical information to construct a complete technical closed loop from microscopic microenvironment identification to macroscopic efficacy prediction. The system exhibits a high degree of automation and objectivity when processing massive amounts of medical data, greatly improving the scientific rigor and accuracy of individualized liver cancer immunotherapy treatment plans.
[0052] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning, characterized in that, Includes the following steps: Acquire multimodal medical imaging data and corresponding clinical information data of liver cancer patients. The multimodal medical imaging data includes at least enhanced computed tomography (CT) images and magnetic resonance imaging (MRI) images. The clinical information data includes the patient's basic physiological indicators, pathological staging information, and previous treatment records. The multimodal medical image data is subjected to standardized preprocessing, including spatial registration, grayscale normalization, and automatic lesion region segmentation, to generate a multimodal image volume dataset in a unified coordinate system; A multi-scale deep feature extraction network is constructed to independently extract high-dimensional image feature vectors from the multimodal image volume dataset, and a cross-modal attention mechanism is used to perform semantic alignment and weight allocation on the feature vectors of different modalities. The semantically aligned multimodal image feature vectors are fused to generate a comprehensive image representation vector, and a multi-source heterogeneous feature matrix is constructed by combining the clinical information data. Based on the aforementioned multi-source heterogeneous feature matrix, a liver cancer immunotherapy efficacy prediction model is trained and deployed. The liver cancer immunotherapy efficacy prediction model adopts a hierarchical structure design, with the bottom layer used to identify the state of the tumor immune microenvironment and the upper layer used to output efficacy evaluation results. The multimodal medical imaging data and clinical information data of the patients to be evaluated are input into the liver cancer immunotherapy efficacy prediction model to obtain quantitative evaluation results, including the probability of immune response, the trend of tumor burden change, and the treatment response level.
2. The method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning according to claim 1, characterized in that, To acquire multimodal medical imaging data and corresponding clinical information data of liver cancer patients, the following operations are performed: the acquisition of multimodal medical imaging data follows a unified time window standard to ensure that the acquisition time difference between enhanced computed tomography images and magnetic resonance imaging images is controlled within 48 hours. The enhanced computed tomography images include volume data from three phases: the arterial phase, the portal venous phase, and the equilibrium phase, with the slice thickness set between 2 millimeters and 1 millimeter. The magnetic resonance imaging includes longitudinal relaxation-weighted imaging, transverse relaxation-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient maps, with the slice spacing set between 3 mm and 5 mm and the in-plane resolution not less than 0.7 mm × 0.7 mm. The clinical information data is obtained by parsing the electronic medical record system, including the patient's age, gender, height, weight, body mass index, tumor size, number, vascular invasion, lymph node metastasis status, distant metastasis markers, surgical resection history, transcatheter arterial chemoembolization history, ablation therapy history, targeted drug dosage and duration, alpha-fetoprotein level, alanine aminotransferase (ALT), aspartate aminotransferase (AST), total bilirubin, albumin level, and neutrophil-to-lymphocyte ratio.
3. The method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning according to claim 1, characterized in that, The spatial registration and grayscale normalization in the standardized preprocessing perform the following operations: a rigid registration algorithm based on mutual information is used as the initial alignment method, and then a non-rigid registration technique based on free shape deformation is applied to correct the liver morphological deformation caused by respiratory motion or patient position differences. For computed tomography images, the raw Henle unit values are linearly mapped to a preset grayscale range, and extreme values that exceed the range of normal liver and tumor enhancement are truncated. For magnetic resonance imaging images, histogram matching or standardization methods are used to eliminate signal intensity differences from different scanning devices by subtracting the mean and dividing by the standard deviation.
4. The method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning according to claim 1, characterized in that, The standardized preprocessing automatically segments the lesion area and performs the following operations: a cascaded three-dimensional fully convolutional neural network architecture is adopted, in which the first-level network is used as a coarse localization network, the input is downsampled whole abdominal body data, and the global spatial features are extracted using deep convolutional layers to generate a liver region mask. The second-level network serves as a fine-grained segmentation network. Its input is the liver region segments determined by the first-level network. This network employs an encoder and decoder structure and introduces multi-scale jump connections between corresponding levels. At the encoder end, the receptive field is increased through continuous convolution and pooling operations. At the decoder end, the resolution is restored through transposed convolution or upsampling operations and fused with encoder features to preserve local detail information. The segmentation results were post-processed morphologically, and isolated holes and noise points were eliminated by opening and closing operations. Connectivity analysis was then applied to extract the target region with the largest volume as the final lesion mask.
5. The method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning according to claim 1, characterized in that, The multi-scale deep feature extraction network performs the following operations: The multi-scale deep feature extraction network contains multiple parallel residual coding paths, each path is configured with a receptive field scale; the first path uses a 3×3×3 convolutional kernel to capture local lesion texture features. The second path uses a 5×5×5 convolutional kernel to perceive the integrity of the organizational structure; the third path uses dilated convolutions with dilation rates to obtain large-scale global structural information; within each path, residual connections are used to sum the input features and the features after nonlinear transformation element-wise to alleviate gradient vanishing in deep networks.
6. The method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning according to claim 1, characterized in that, The cross-modal attention mechanism performs semantic alignment and weight allocation on feature vectors of different modalities, and performs the following operations: mapping the enhanced computed tomography feature map and the magnetic resonance imaging feature map into query vector space, key vector space and value vector space, respectively; The dot product between the query vector from enhanced computed tomography (CT) and the key vector from magnetic resonance imaging (MRI) is calculated and then scaled by the square root of the feature dimension to obtain a relevance score. The relevance scores are processed by a normalized exponential function to generate an attention weight distribution map, which reflects the semantic relevance of different modalities at the same anatomical location. The attention weights are weighted and summed with the value vectors of the corresponding modalities to achieve feature enhancement and alignment, and the contribution weights of each modality in the fusion process are dynamically adjusted based on the relevance scores.
7. The method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning according to claim 1, characterized in that, The construction of the multi-source heterogeneous feature matrix involves the following operations: concatenating the aligned modal features along the channel dimension, performing dimensionality compression and nonlinear interaction through a 1×1×1 convolutional layer, and extracting cross-modal common feature representations. Discrete variables in clinical information data are transformed into high-dimensional continuous vector representations using one-hot encoding or pre-trained embedding layers; continuous variables in clinical information data are processed by min-max standardization and then used as vector components. The transformed clinical feature vector is concatenated with the continuous comprehensive image feature vector. A feature balancing factor is introduced during the concatenation process. By multiplying the image features and clinical features by the corresponding scaling factors, the two types of features are ensured to be balanced in terms of dimensionality and numerical distribution.
8. The method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning according to claim 1, characterized in that, The underlying layer of the liver cancer immunotherapy efficacy prediction model is used to identify the state of the tumor immune microenvironment and perform the following operations: using a graph neural network structure, the tumor region is divided into multiple sub-region nodes with anatomical significance or geometric features, and each node carries the local imaging features and clinical mapping features of the region. Calculate the Euclidean distance or cosine similarity between nodes to construct an adjacency matrix and model the connection relationships between nodes; In graph convolution operations, the features of each node are aggregated and updated according to the adjacency matrix and the features of its neighboring nodes, so that the features of the current node in the next layer are equal to the nonlinear transformation result of its own features and the features of its neighboring nodes, in order to capture the spatial distribution pattern of immune cell infiltration inside the tumor.
9. The method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning according to claim 1, characterized in that, The upper layer of the liver cancer immunotherapy efficacy prediction model is used to output efficacy evaluation results and performs the following operations: integrating the evaluation results of historical treatment stages using a gated loop unit structure, wherein the gated loop unit structure includes a reset gate and an update gate; By updating the gate, the proportion of historical therapeutic effects characteristics from the previous moment retained in the current state is controlled; by resetting the gate, the fusion ratio of current input information and historical information is determined, thereby achieving dynamic tracking of the therapeutic effect evolution trend. During the model training phase, the cross-entropy loss function is used as the optimization objective. The logic of the cross-entropy loss function is to take the weighted sum of the logarithms of the true label and the predicted probability and take the opposite number. The network weights are then adjusted through the backpropagation algorithm to minimize the difference between the predicted value and the actual efficacy label.
10. The method for evaluating the efficacy of liver cancer immunotherapy by integrating radiomics and deep learning according to claim 1, characterized in that, The quantitative assessment results and model update mechanism perform the following operations: the immune response probability is expressed as a continuous value between 0 and 1; the tumor burden change trend is output by comparing the characteristic differences between baseline images and follow-up images, and the expected percentage change in tumor volume or metabolic activity is output. The treatment response levels are divided into four categories: complete response is defined as the disappearance of all target lesions and the appearance of no new lesions; partial response is defined as a reduction in the total diameter of target lesions of 30% or more. Disease stability is defined as the reduction of target lesions not reaching a partial response and the increase not reaching the level of disease progression. Disease progression is defined as an increase of more than 20% in the total diameter of target lesions or the appearance of new lesions; Establish an online update mechanism that triggers incremental retraining of model parameters when the cumulative number of newly labeled samples reaches a predetermined threshold. The incremental retraining employs a learning rate decay strategy, adjusting model parameters to adapt to new clinical drug regimens, imaging equipment, or patient population characteristics.