A liver cancer pathological differentiation degree prediction system and method fusing multi-sequence magnetic resonance imaging and semantic information
By fusing multi-sequence magnetic resonance imaging with semantic information, the problems of data utilization being singular and feature extraction lacking specificity in the analysis of pathological differentiation degree of hepatocellular carcinoma were solved. This approach achieved deep complementarity between image and text features, improving prediction accuracy and clinical application efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF ELECTRONICS SCI & TECH OF CHINA
- Filing Date
- 2026-06-09
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies for image analysis of the pathological differentiation of hepatocellular carcinoma suffer from problems such as limited data utilization, lack of targeted feature extraction, low efficiency of multimodal fusion, and reliance on manual annotation, making it difficult to effectively capture fine-grained texture features and achieve deep information complementarity.
By employing a multi-sequence magnetic resonance imaging and semantic information fusion approach, and through lesion segmentation and region cropping, combined with structured diagnostic text generation and cross-modal attention mechanisms, we achieve deep complementarity between image and text features, construct an end-to-end analysis architecture, and reduce reliance on manual annotation.
It improves the predictive accuracy and interpretability of the pathological differentiation degree of hepatocellular carcinoma, reduces the time and errors of manual annotation, and enhances the robustness and clinical application efficiency of the model.
Smart Images

Figure CN122369896A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing, and specifically relates to a method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information. Background Technology
[0002] Hepatocellular carcinoma (HCC) is the most common primary malignant liver tumor, characterized by high mortality and morbidity. The histopathological features corresponding to the degree of tumor differentiation are crucial for clinical prognostic assessment. Compared to well-differentiated or moderately differentiated HCC, poorly differentiated HCC has a higher recurrence and metastasis rate, often exhibiting more significant tissue heterogeneity, atypical blood supply patterns, and invasiveness on imaging. According to the Clinical Guidelines for Liver Cancer, imaging examinations can be used for initial screening and localization of liver cancer, while histological assessment is the standard for diagnosis and grading. However, relying solely on radiologists' visual interpretation of routine imaging features is insufficient to fully quantify the heterogeneity characteristics within the tumor, and the predictive accuracy for different degrees of differentiation is limited. Therefore, developing a computational method that can effectively extract and quantitatively characterize the tissue heterogeneity features of HCC from multi-sequence magnetic resonance imaging (MRI) is of significant clinical value in assisting radiologists in making non-invasive and objective predictions of pathological differentiation.
[0003] Current technologies in liver image analysis are mainly developing in the following directions:
[0004] (I) Evolution of Feature Extraction Methods. From early texture analysis and traditional radiomics to deep learning models, feature extraction capabilities have continuously improved. Texture analysis based on enhanced magnetic resonance imaging, such as gray-level run length inhomogeneity and intensity statistics, has shown the ability to reflect changes in the internal microstructure of tumors. Traditional radiomics models, such as those based on handcrafted features and machine learning, have also demonstrated their effectiveness in image feature extraction. With the development of artificial intelligence technology, deep learning models have been widely applied to the interpretation of medical images using various imaging modalities. Many studies have utilized the ability of convolutional neural networks to efficiently extract medical image features to identify the imaging characteristics of hepatocellular carcinoma. Yang et al. proposed a multi-channel three-dimensional convolutional neural network; Zhou et al. introduced a model based on DenseNet and integrating squeezing and excitation modules, achieving an area under the receiver operating characteristic curve (AUC) of 0.83.
[0005] (II) Attempts at Multimodal Fusion. To further improve the accuracy of image analysis, some studies have adopted multimodal learning methods to fuse data from different image modalities in order to capture the image features of lesions more comprehensively. In terms of image-level fusion, Li et al. proposed BSAFusion, which uses bidirectional stepwise alignment and a multimodal fusion module. Zhou et al. introduced entropy-based feature metrics and novel fusion modules to improve tumor visualization. Text-image fusion further enriches feature representation, and common paradigms include linear combination, attention-based fusion, and decision-level fusion. Jia et al. further improved model performance through multimodal fusion and attention-based metric learning. Studies have shown that incorporating semantic knowledge can improve the interpretability of the model. Gao et al. used recurrent neural networks (RNNs) to embed semantic prior information for image classification of liver lesions. Inspired by visual language models, MedCLIP uses contrastive learning to align images with reports to achieve cross-modal retrieval. MedUnifie utilizes the Transformer architecture to unify heterogeneous modalities and promote cross-domain generalization.
[0006] (III) Common Deficiencies of Existing Solutions. Despite the progress made in the above studies, the following common problems still exist when applied to image analysis of the pathological differentiation degree of hepatocellular carcinoma: First, most models rely solely on image data, ignoring the reasoning process of clinical experts comprehensively utilizing multi-source information for interpretation in actual workflows; Second, many architectures focus on optimizing feature extraction algorithms, but neglect the inherent difficulties of complex anatomical structures and high heterogeneity of lesions in liver medical images, making it difficult to capture fine-grained texture features that are crucial for determining the pathological differentiation degree; Third, effectively fusing information from different imaging modalities remains challenging, and existing methods often lead to data redundancy, high computational costs, and limited representation quality, making it difficult to achieve deep complementarity of information from different modalities; Fourth, current methods usually require doctors to manually label lesion locations before analysis, a process that is both time-consuming and prone to labeling errors, limiting the application of the model in large-scale clinical scenarios.
[0007] Taking the LiAIDS system and the STIC model as examples, the former uses shallow physical stitching to fuse images and clinical data, ignoring differences in the feature space domain, and relies on element-wise multiplication to fuse multi-phase feature maps, lacking a decoupling mechanism; the latter relies entirely on RNN self-supervised mining of temporal patterns, without introducing medical prior knowledge, and directly stitches together image features and clinical numerical features. These typical solutions have all failed to effectively solve the above four common problems.
[0008] Therefore, how to fully integrate the features and semantic information of multi-sequence magnetic resonance images to improve the accuracy and interpretability of image analysis of the pathological differentiation degree of hepatocellular carcinoma is a technical problem that urgently needs to be solved. Summary of the Invention
[0009] To address the problems of singular data utilization, lack of targeted feature extraction, low efficiency of multimodal fusion, and reliance on manual annotation in the existing technologies, this invention provides a system and method for predicting the pathological differentiation degree of liver cancer by fusing multi-sequence magnetic resonance imaging and semantic information, in order to achieve the following objectives:
[0010] (1) Introduce the reasoning process of clinical experts in making comprehensive judgments by utilizing multi-source information in actual workflows, and simulate the decision-making logic of multi-source information comprehensive analysis;
[0011] (2) To address the inherent difficulties of complex anatomical structures and high heterogeneity of lesions in liver medical images, a targeted feature enhancement strategy was designed to improve the ability to capture fine-grained texture features.
[0012] (3) Construct a multimodal fusion module to achieve efficient complementarity between image features and semantic information, and avoid data redundancy and limited representation quality;
[0013] (4) Integrate an end-to-end analysis architecture to eliminate the reliance on manually labeled lesion locations.
[0014] To achieve the above objectives, the present invention adopts the following technical solution:
[0015] A method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information includes the following steps:
[0016] S1. Data Acquisition: Acquire multi-sequence magnetic resonance imaging (MRI) data of the target object, and perform lesion segmentation and region cropping on each sequence of MRI data to obtain multi-sequence lesion region images;
[0017] S2. Structured Diagnostic Text Generation: Design structured cue words, take structured cue words and multi-sequence lesion area images as input, and use a pre-trained large language model to generate structured diagnostic text;
[0018] S3. Feature Extraction and Fusion: Based on structured diagnostic text and multi-sequence lesion region images, a lesion differentiation prediction model is used to predict the degree of pathological differentiation of liver cancer; specifically, it includes the following sub-steps:
[0019] S31. Image Feature Representation Generation: For the multi-sequence lesion region images obtained in step S1, specific features and common structural features are extracted sequentially using a parallel dual-branch encoder. The specific features and common structural features are then fused using an enhanced lightweight residual fusion module to obtain the fused features of each sequence. The fused features of all sequences are then stitched together and processed by a channel compression layer to generate a unified image feature representation.
[0020] S32. Text Feature Generation: The structured diagnostic text generated in step S2 is segmented and encoded using a RoBERTa encoder to extract global embedding vectors representing different sequences; the cosine similarity between the global embedding vectors of each sequence is calculated and a similarity matrix is constructed. After attention normalization and weighted aggregation, text feature representation is generated.
[0021] S33. Cross-modal feature fusion: A bidirectional cross-modal attention mechanism is adopted. First, the first attention output is calculated with text features as the query and image features as the key and value. Then, the second attention output is calculated with image features as the query and text features updated by the first attention interaction as the key and value. The two outputs are fused to obtain the cross-modal fused feature.
[0022] S4. Differentiation Prediction: Based on cross-modal fusion features, quantitative indicators are generated to characterize the differentiation grade of hepatocellular carcinoma, thereby obtaining information on the pathological differentiation degree of liver cancer.
[0023] Furthermore, the multi-sequence MRI data acquired in step S1 includes arterial phase, portal venous phase, delayed phase, and T2-weighted images.
[0024] Furthermore, the process of segmenting lesions and cropping regions for each sequence in step S1 includes:
[0025] S11. Preliminary screening: Perform two-dimensional rapid detection based on MRI data to remove MRI data with no lesions or lesion areas with pixels smaller than 50% of the largest lesion pixel of the target object;
[0026] S12. Three-dimensional segmentation: Using the MRI data obtained in S11 as input, a deep learning model is used to perform three-dimensional segmentation of HCC lesions on the input data;
[0027] S13. Post-processing: Based on the segmentation results of step S12, perform connected component analysis and retain only the HCC connected region with the largest volume in three-dimensional space; based on the minimum bounding rectangle of the largest HCC connected region, crop and expand the original image to suppress the background, and adaptively weight and enhance the texture based on the contrast, energy, and correlation statistics of the gray-level co-occurrence matrix, and use the processed image patch as the final lesion area image data.
[0028] Furthermore, the deep learning model in step S12 is a 3D U-Net that integrates non-local attention and pixel attention mechanisms within the nnU-Net framework; it employs a foreground-aware sampling strategy to prioritize sampling image patches containing lesion voxels for training.
[0029] Furthermore, the method for designing structured prompt words in step S2 specifically includes:
[0030] S21. Domain Knowledge Extraction: Based on the LI-RADS standard, extract the semantic features required for the diagnosis of hepatocellular carcinoma; the semantic features include lesion diameter, signal intensity, signal distribution, lesion morphology, lesion edge and anatomical location;
[0031] S22. Structured cue word design: Based on semantic features, the image analysis task is uniformly defined as: identifying lesion features at anatomical locations during the image sequence.
[0032] Furthermore, the process of simultaneously extracting specific features and common structural features from the multi-sequence lesion region images obtained in step S1 using a parallel dual-branch encoder sequence by sequence includes:
[0033] Construct a parallel dual-branch encoder, which consists of a dedicated encoder and a shared encoder; wherein:
[0034] The dedicated encoder is a three-dimensional encoder built separately for each MRI lesion region image data sequence, used to capture the sequence-specific lesion diagnostic information reflected by the sequence under its imaging mechanism and obtain dedicated features;
[0035] The shared encoder is a single three-dimensional encoder used by all MRI lesion region image data sequences to extract common structural features from all sequences.
[0036] The MRI lesion region image data of each sequence are simultaneously sent to a dedicated encoder and a shared encoder, and the two are processed in parallel.
[0037] Furthermore, in step S31, the process of fusing specific features and common structural features through the enhanced lightweight residual fusion module to obtain the fused features of each sequence includes:
[0038] An enhanced lightweight residual fusion module is constructed, which consists of modal branches and shared branches;
[0039] In the modality branch: local spatial features are extracted through convolution operations. Subsequently, the extracted features are updated by weighting them using a channel attention mechanism to obtain weighted features and their corresponding attention weights. In the modality sharing branch: local spatial features are extracted through convolution operations. Then, the extracted features are updated by weighting them using a local contextual attention mechanism to obtain weighted shared features and their corresponding attention weights. .
[0040] The weighted features are added to the shared features according to the following formula to obtain the fused features. :
[0041] ;
[0042] in, This is represented as element-wise multiplication;
[0043] Fusion features The image is processed using depthwise separable convolutions, and global context enhancement and residual connections are introduced to fuse multi-scale information. For a total of N image sequences, the fusion features of all sequences are concatenated along the channel dimension to obtain the concatenated features, which are the fusion features of each sequence. The fusion features of all sequences are then compressed through a single channel compression layer to form a unified image representation. :
[0044] ;
[0045] in, This indicates a splicing operation along the channel dimension. Indicates channel compression mapping, This represents the fusion feature corresponding to the i-th image sequence.
[0046] Furthermore, the method for predicting the pathological differentiation degree of liver cancer also includes: using a total loss function. To update the processing parameters; among them, and These are weight parameters; For classification loss; For cross-modal contrast loss; This represents the distribution alignment loss.
[0047] A system for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information includes a lesion segmentation module, a semantic description generation module, and a lesion differentiation prediction module; the lesion differentiation prediction module includes a multimodal feature extraction unit, a cross-modal feature fusion unit, and a classification prediction unit; the modules work together to realize the method for predicting the pathological differentiation degree of liver cancer.
[0048] A computer-readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements the method for predicting the pathological differentiation degree of liver cancer.
[0049] By adopting the above technical solution, the present invention has the following advantages:
[0050] 1. This invention transforms the professional medical features contained in images that conform to the LI-RADS diagnostic criteria into structured diagnostic text information. Then, using lesion region images and structured diagnostic text as input, a bidirectional cross-modal attention mechanism is employed to align and complement the visual features in the images that are difficult to quantify with the semantic information in the text. This creates a synergistic effect between image and text information, enabling the model to capture deep pathological correlations that cannot be effectively expressed by a single image modality. This solves the problem of existing technologies that only analyze pixel-level image features and have a single information dimension, improving the accuracy and reliability of predictions and providing a more precise basis for clinical decision-making.
[0051] 2. This invention employs an attention-based 3D segmentation model to segment lesions and crop regions from raw multi-sequence MRI data, replacing the time-consuming and subjectively subjective manual layer-by-layer delineation process performed by traditional radiologists. Due to the algorithm-driven automated processing, the segmentation results are highly repeatable and objective, significantly improving clinical application efficiency and laying the foundation for subsequent structured text generation and pathological differentiation prediction. This automates the entire process from raw image preprocessing and lesion region extraction to subsequent prediction results, further enhancing clinical application efficiency while ensuring objectivity.
[0052] 3. This invention employs a multi-source information fusion strategy combining MRI images and textual information during the fusion process, and designs a shared + specific feature extraction architecture within the image processing channel. In clinical practice, factors such as patient respiratory movements may introduce artifacts and lead to a decrease in the quality of some image sequences. In a single-modality model, such artifacts may directly cause prediction failure; however, in this invention, the features of different image sequences and the semantic information of the text have a certain degree of redundancy and complementarity. Even if image artifacts exist, the model can still make a comprehensive judgment based on the image features of other high-quality sequences and the corresponding accurate textual descriptions. This mutual complementarity of multi-source information enhances the model's anti-interference ability, improves its tolerance to individual data quality issues, and demonstrates better stability in complex clinical scenarios, thereby improving the model's robustness and reducing its sensitivity to the quality of a single image sequence. Attached Figure Description
[0053] Figure 1 This is a flowchart of the method for predicting the degree of pathological differentiation of liver cancer according to the present invention;
[0054] Figure 2 This is a diagram of the lesion region image network architecture obtained by segmenting multi-sequence MRI in the embodiment;
[0055] Figure 3 This is a flowchart illustrating the structured diagnostic text generation process in an embodiment.
[0056] Figure 4This is a diagram of the network architecture for multimodal feature extraction fusion and pathological differentiation prediction in the embodiment. Detailed Implementation
[0057] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0058] like Figure 1 As shown in the figure, this embodiment provides a method for predicting the pathological differentiation degree of liver cancer by fusing multi-sequence magnetic resonance imaging and semantic information, including the following steps:
[0059] S1. Data Acquisition: Acquire multi-sequence magnetic resonance imaging (MRI) data of the target subject, and perform slice segmentation processing on each sequence of MRI data to obtain multi-sequence lesion region images. The target subject can be a healthy subject, a suspected hepatocellular carcinoma patient, or a confirmed hepatocellular carcinoma patient; the acquired multi-sequence MRI data includes arterial phase, portal venous phase, delayed phase, and T2-weighted images. The process of lesion segmentation and region cropping for each sequence in this embodiment includes the following sub-steps:
[0060] S11. Preliminary screening: Perform two-dimensional rapid detection based on MRI data to remove MRI data with no lesions or lesion areas having pixels smaller than 50% of the largest lesion pixel size of the target object.
[0061] S12. Three-dimensional segmentation: In this embodiment, the MRI data obtained in S11 is used as input, and the lesion segmentation module is used to perform three-dimensional segmentation.
[0062] (1) Automated segmentation model configuration:
[0063] The lesion segmentation module is a deep learning model based on a 3D U-Net architecture and enhanced by an attention mechanism. For example... Figure 2 As shown, the deep learning model employs an encoder-decoder structure, integrating a non-local attention mechanism in the encoder (downsampling) path to associate spatially distant but semantically related regions. Specifically:
[0064] Given an input feature map X, three 1×1×1 convolutions are used to obtain:
[0065] ;
[0066] in, , , It uses a learnable 1×1×1 convolutional kernel, which also serves to reduce dimensionality. The attention weights are calculated as follows:
[0067] ;
[0068] Where i represents the index of the current location to be enhanced, and j represents the index of the location participating in the association calculation. This represents the indexes of all positions traversed during the normalization summation process; This represents the query feature vector at position i. This represents the key feature vector at position j. f(i,j) represents the feature vector at position j; f(i,j) represents the normalized attention weight between positions i and j. This represents the output feature obtained at position i after nonlocal attention aggregation. This operation is performed once after each downsampling stage.
[0069] In the decoder (upsampling) path, transposed convolution is used for upsampling, and skip connections are used to preserve high-resolution spatial details. Unlike the non-local attention in the encoder, a 3D pixel attention mechanism is also set at the end of each decoding unit to calculate voxel-level weights, thereby highlighting lesion areas while suppressing background noise. This attention mechanism is calculated as follows:
[0070] ;
[0071] Where X represents the input feature map, This represents a convolution operation, used for linear mapping of the input feature map; This represents the Sigmoid activation function, used to map the output to the [0,1] interval, generating a pixel attention map. The importance weights of each voxel location are encoded; This represents the output feature map after pixel attention weighting; This indicates element-wise multiplication.
[0072] (2) Segmentation Implementation and Post-processing:
[0073] The segmentation model is implemented and trained within the nnU-Net framework. This embodiment designs the backbone network structure of this framework: based on the standard 3D U-Net backbone, a non-local attention module is added after each downsampling stage of the encoder to model the semantic relationships between far-distance voxels in space; and a 3D pixel attention module is added at the end of each upsampling unit of the decoder to weight and highlight lesion regions and suppress background noise at the voxel level. Apart from the aforementioned backbone network improvements, data preprocessing, training scheduling, sliding window inference, and post-processing all follow the default implementation of the nnU-Net framework; this embodiment does not modify the aforementioned automatic configuration. For training sampling, this embodiment enables the foreground oversampling mechanism built into the nnU-Net framework to alleviate the foreground-background class imbalance problem.
[0074] Refined Input Generation: Based on the largest connected regions obtained from the above segmentation, background suppression and texture enhancement steps are performed on the original image. Specifically:
[0075] Background suppression: The minimum bounding rectangle of the lesion region is determined based on the maximum connected region. The original image is cropped to remove redundant background regions outside the rectangle. The region of interest (ROI) obtained after cropping is extended outward by 50 pixels to form the lesion region.
[0076] Texture enhancement: The lesion region obtained after background suppression is subjected to texture enhancement operation, and finally a single image is generated in which the background region is suppressed and the texture of the lesion region is enhanced. This image is used as the final lesion region image and as the input for subsequent prediction tasks.
[0077] In some specific embodiments, the texture enhancement method is as follows:
[0078] Calculate the gray-level co-occurrence matrix of the lesion region obtained by background suppression: Definition This indicates that, under the condition of displacement vector Δ, the grayscale value is... and Pixel pairs in the displacement vector The probability of both appearing together is given by Δx, where Δy represents the displacement in the horizontal direction, i and j represent discrete gray level indices, corresponding to the gray values of the center pixel and its neighboring pixels, respectively; I(x,y) represents the gray value of the image at coordinates (x,y). Let represent the grayscale value of a neighboring pixel that is at a distance Δ from coordinates (x, y) by a displacement vector Δ; therefore, the grayscale co-occurrence matrix is defined as follows:
[0079] ;
[0080] Where N represents the total number of valid pixel pairs that satisfy the displacement relationship Δ; M represents the upper limit of the pixel range of the image participating in the statistics in the corresponding coordinate direction; This is an indicator function.
[0081] A series of texture features are extracted from the gray-level co-occurrence matrix, including:
[0082] 1. Contrast: Used to measure the intensity of grayscale differences in an image.
[0083] ;
[0084] 2. Energy: Used to reflect the uniformity or repeatability of the texture.
[0085] ;
[0086] 3. Correlation: Used to measure the degree of linear correlation between image textures.
[0087] ;
[0088] in, and These represent the mean gray values of the gray-level co-occurrence matrix in the row and column directions, respectively, and are used to characterize the center position of the gray-level distribution; and These represent the gray-level standard deviations of the gray-level co-occurrence matrix in the row and column directions, respectively, used to characterize the degree of dispersion of the gray-level distribution; This represents the joint deviation of gray levels relative to their respective means. The larger the correlation value, the stronger the linear dependence between the gray levels of adjacent pixels, and the more regular the texture structure.
[0089] The mean values of contrast, energy, and correlation are normalized and linearly stretched to map a weight, which is then multiplied by the pixel values of the original slice, thereby adaptively enhancing or suppressing the texture details of the image according to the complexity of the texture.
[0090] S2. Structured Diagnostic Text Generation. This step aims to transform the professional medical features contained in the images that conform to the LI-RADS diagnostic criteria into structured diagnostic text information. The implementation process is as follows: Figure 3 As shown:
[0091] S21. Domain Knowledge Extraction: Based on the diagnostic content in the LI-RADS standard, an analysis was conducted to sort out its assessment content and extract a set of semantic features required for the diagnosis of hepatocellular carcinoma. The diagnostic content includes indicators such as lesion size, arterial phase hyperenhance, enhancing capsule, non-peripheral clearance, and threshold increase; the semantic features include lesion diameter, signal intensity, signal distribution, lesion morphology, lesion margin, special features, and anatomical location.
[0092] S22. Structured Cue Word Design: To guide the large language model in standardized analysis, a cue word template is designed. This template defines the image analysis task as: identifying lesion features of anatomical locations during image sequences. It is used to generate structured and clinically meaningful descriptions for each input lesion region image, thereby enhancing the model's generalization ability when processing different text inputs.
[0093] S23. Using structured prompts and multi-sequence lesion region images as input, a pre-trained large-scale language model (LLM) is used to generate structured diagnostic text. The large-scale language model used in this embodiment is ChatGPT-4o; under the dual guidance of system prompts for setting the dialogue scenario and basic context, and user prompts for specifying the analysis task, the input multi-sequence lesion region images are analyzed, and a natural language text description is generated, thereby obtaining structured diagnostic text.
[0094] Regarding the image-to-text semantic generation part, its core function is to transform visual features into structured text descriptions. This function does not solely rely on the aforementioned large language model; it can be replaced by other large-scale language models with multimodal understanding and generation capabilities, such as Google's Gemini series, Anthropic's Claude series, or the open-source Llama series. Models specifically optimized for vision-language tasks can also be used. Semantic information can also be achieved through a non-generative approach based on discriminative models and template filling. This involves training an independent discriminative classification model for each key clinical image feature to determine whether the input lesion image possesses that feature.
[0095] S3. Feature Extraction and Fusion: This step uses a lesion differentiation prediction module to predict the pathological differentiation degree of liver cancer based on structured diagnostic text and multi-sequence lesion region images. During this process, image features and text semantic features are collaboratively processed and deeply fused, allowing the two heterogeneous information to complement each other and improve accuracy. Specifically... Figure 4 As shown:
[0096] S31. Image Feature Representation Generation:
[0097] (1) Parallel feature extraction with two branches: For the multi-sequence lesion region images obtained in step S1, the specific features and common structural features are extracted synchronously by a parallel dual-branch encoder for each sequence.
[0098] The dedicated branch is a 3D encoder set independently for each sequence, used to capture the unique diagnostic information of that sequence and obtain its unique features; the shared branch is a 3D encoder shared by all sequences, used to extract the common structural features of all sequences; the input of each sequence is processed in parallel by the corresponding dedicated encoder and the shared encoder.
[0099] (2) Fusion of specific features and shared features: The process of fusing specific features and shared structural features through an enhanced lightweight residual fusion module to obtain the fused features of each sequence includes:
[0100] An enhanced lightweight residual fusion module is constructed, which consists of modal branches and shared branches;
[0101] In the modality branch: local spatial features are extracted through convolution operations. Subsequently, the extracted features are updated using a channel attention mechanism to obtain weighted features and their corresponding attention weights. In the modality sharing branch: local spatial features are extracted through convolution operations. Then, the extracted features are updated by weighting them using a local contextual attention mechanism to obtain weighted shared features and their corresponding attention weights. .
[0102] The weighted features are added to the shared features according to the following formula to obtain the fused features. :
[0103] ;
[0104] in, This is represented as element-wise multiplication;
[0105] Fusion features Processing is performed using depthwise separable convolutions, and global context enhancement and residual connections are introduced to fuse multi-scale information;
[0106] For a total of N image sequences, the fusion features of all sequences are concatenated along the channel dimension to obtain the fusion features of each sequence; after compressing the fusion features of all sequences through a channel compression layer, a unified image feature representation is formed. :
[0107] ;
[0108] in, This indicates a splicing operation along the channel dimension. Indicates channel compression mapping, This represents the fusion feature corresponding to the i-th image sequence. S32. Text Feature Generation: The structured diagnostic text generated in step S2 is segmented and encoded using a RoBERTa encoder, outputting token-level hidden representations and global embedding vectors. The cosine similarity between the global embedding vectors of each sequence is calculated, and a similarity matrix is constructed. The elements in the i-th row of matrix S are summed to obtain the cumulative semantic similarity between the i-th sequence and the other sequences. Softmax normalization is performed along the sequence dimension to obtain the attention weights corresponding to the i-th sequence. Finally, unified text features are generated after attention normalization and weighted aggregation.
[0109] ;
[0110] Where N represents the total number of input MRI image sequences; ∈ℝ d d represents the global embedding vector obtained by RoBERTa encoding the structured diagnostic text corresponding to the i-th sequence, where d is the feature dimension output by RoBERTa. F represents the weight coefficient corresponding to the i-th sequence obtained after attention normalization of the cosine similarity matrix, and its value reflects the importance of the sequence in the multi-sequence semantic set; txt This indicates that the embedding vectors of each sequence are weighted. The unified text feature representation obtained after weighted summation is used for subsequent cross-modal fusion.
[0111] S33. Cross-modal feature fusion: This step adopts a bidirectional cross-modal attention mechanism to enhance the interaction between image feature representation and text feature representation. Specifically, the attention outputs in two directions are calculated separately: first, the first attention output is calculated with text features as query Q, image features as key K and value V, and then the second attention output is calculated with image features as query Q, and text features updated by the first attention interaction as key K and value V. The outputs in the two directions are fused to obtain the final cross-modal fused feature.
[0112] The core calculation formula of the above bidirectional cross-modal attention mechanism adopts standard scaled dot product attention:
[0113] ;
[0114] Where, d k The dimension of the key feature is used to scale the dot product result to control gradient stability.
[0115] After obtaining the attention outputs from both directions, the two are fused. The fusion method can be either concatenation or element-wise addition to obtain the final cross-modal fused feature. In this embodiment, element-wise addition is preferred to reduce the feature dimensionality while maintaining computational efficiency.
[0116] S4. Based on cross-modal fusion features, quantitative indicators are generated to characterize the differentiation grade of hepatocellular carcinoma, thereby obtaining information on the pathological differentiation degree of liver cancer for doctors to use as a reference for interpretation.
[0117] The aforementioned method for predicting the pathological differentiation degree of liver cancer further includes: using a total loss function to update processing parameters; the total loss function includes classification loss, cross-modal contrast loss, and distribution alignment loss, and its expression is:
[0118] ;
[0119] in, Indicates the total loss; Indicates classification loss; Indicates cross-modal contrast loss; This represents the distribution alignment loss; and These represent the weight parameters for cross-modal contrast loss and distribution alignment loss, respectively.
[0120] The classification loss mitigates the class imbalance problem through a dynamic scaling factor, allowing the model to focus more on learning samples that are difficult to classify. Its mathematical expression is as follows:
[0121]
[0122] Where B represents the number of samples in a training batch; C represents the number of categories; b represents the b-th sample; and c represents the c-th category. This represents the true label indication of the b-th sample in class c; This represents the probability that the model predicts the b-th sample belongs to the c-th class; This represents the class weight corresponding to class c, used to alleviate the class imbalance problem; This represents the focus modulation factor, used to reduce the contribution of easily classified samples to the loss, making the model pay more attention to difficult-to-classify samples.
[0123] The cross-modal contrastive loss is used to enhance the similarity of matched text-image pairs in the embedding space and suppress the similarity of mismatched text-image pairs. For the i-th text feature vector and the j-th image feature vector in a batch, their similarity is defined as:
[0124] ;
[0125] in, This represents the similarity between the i-th text feature and the j-th image feature; This represents the result of L2 normalization of the i-th text feature vector; This represents the result of L2 normalization of the j-th image feature vector; Indicates the transpose operation; This represents the temperature coefficient, used to adjust the smoothness of the similarity distribution.
[0126] Building upon this, a symmetric bidirectional contrastive loss function is employed to enforce semantic alignment between images and text within the embedding space. This loss function is learned by maximizing the similarity of matching image-text pairs and minimizing the similarity of non-matching pairs. Specifically, it is defined as follows:
[0127] ;
[0128] in, This represents the similarity normalization term between the i-th text feature and all image features in the batch; This represents the similarity normalization term between the i-th image feature and all text features in the batch.
[0129] The distribution alignment loss is used to measure and reduce the distribution differences of different image modal features in the latent space. In this embodiment, the distribution alignment loss uses Jensen-Shannon divergence. Let the shared feature corresponding to the m-th modality in the b-th sample be represented as... Then its probability distribution is expressed as:
[0130] ;
[0131] The average distribution of each modality is expressed as:
[0132] ;
[0133] Furthermore, the distribution alignment loss is defined as:
[0134] ;
[0135] Where M represents the number of input image modalities; This represents the Kullback-Leibler divergence, used to measure the difference between two probability distributions. It is achieved by minimizing... This can make the distribution of different image modalities more consistent in the shared feature space.
[0136] This embodiment also provides a liver cancer pathological differentiation degree prediction system that integrates multi-sequence magnetic resonance imaging and semantic information. The system consists of a lesion differentiation prediction model, which includes a lesion segmentation module, a semantic description generation module, and a lesion differentiation prediction module. The lesion differentiation prediction module includes a multimodal feature extraction unit, a cross-modal feature fusion unit, and a classification prediction unit. The modules work together to realize the liver cancer pathological differentiation degree information prediction method.
[0137] The above methods and systems were experimentally verified:
[0138] Data augmentation: During the training of the lesion differentiation prediction model, various data augmentation operations were applied to each input MRI image, including random segmentation, random translation, geometric transformation, and flipping.
[0139] Training configuration: Both the lesion segmentation module and the lesion differentiation prediction module underwent independent 5-fold cross-validation training. The segmentation model used the SGD optimizer; the lesion differentiation prediction module used the AdamW optimizer. Hyperparameters such as the learning rate were selected and adjusted based on the cross-validation results.
[0140] In summary, the method and system of this invention achieve automated prediction of the pathological differentiation degree of hepatocellular carcinoma by deeply integrating MRI images with semantic information. After acquiring the data, automated lesion segmentation is first used to identify the lesion region corresponding to hepatocellular carcinoma in the original MRI, achieving a significant improvement in efficiency. Based on this, structured cue words are designed according to the professional medical features of the LI-RADS diagnostic criteria. Then, using the lesion region image as input, a large language model is used to transform the unstructured image features into a standardized, physician-understandable medical semantic text description, resulting in structured diagnostic text. For both image and text modalities, image feature representations and text features are extracted separately. A parallel dual-channel cross-modal fusion mechanism is used to collaboratively process and deeply fuse the acquired image features and text semantic features, making the two heterogeneous information complementary.
Claims
1. A method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information, characterized in that, Includes the following steps: S1. Data Acquisition: Acquire multi-sequence MRI data of the target object, and perform lesion segmentation and region cropping on each sequence of MRI data to obtain multi-sequence lesion region images; S2. Structured Diagnostic Text Generation: Design structured cue words, take structured cue words and multi-sequence lesion area images as input, and use a pre-trained large language model to generate structured diagnostic text; S3. Feature Extraction and Fusion: Based on structured diagnostic text and multi-sequence lesion region images, a lesion differentiation prediction model is used to predict the pathological differentiation degree of liver cancer; specifically, it includes the following sub-steps: S31. Image Feature Representation Generation: For the multi-sequence lesion region images obtained in step S1, specific features and common structural features are extracted sequentially using a parallel dual-branch encoder. The specific features and common structural features are then fused using an enhanced lightweight residual fusion module to obtain the fused features of each sequence. The fused features of all sequences are then stitched together and processed by a channel compression layer to generate a unified image feature representation. S32. Text Feature Generation: The structured diagnostic text generated in step S2 is segmented and encoded using a RoBERTa encoder to extract global vectors representing different sequences; the cosine similarity between the global vectors of each sequence is calculated and a similarity matrix is constructed. After attention normalization and weighted aggregation, the text feature representation is generated. S33. Cross-modal feature fusion: A bidirectional cross-modal attention mechanism is adopted. First, the first attention output is calculated with text features as the query and image features as the key and value. Then, the second attention output is calculated with image features as the query and text features updated by the first attention interaction as the key and value. The two outputs are fused to obtain the cross-modal fused feature. S4. Differentiation Prediction: Based on cross-modal fusion features, quantitative indicators are generated to characterize the differentiation grade of hepatocellular carcinoma, thereby obtaining information on the pathological differentiation degree of liver cancer.
2. The method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information according to claim 1, characterized in that, The multi-sequence MRI data acquired in step S1 includes arterial phase, portal venous phase, delayed phase, and T2-weighted images.
3. The method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information according to claim 1, characterized in that, The process of segmenting lesions and cropping regions for each sequence in step S1 includes: S11. Preliminary screening: Perform two-dimensional rapid detection based on MRI data to remove MRI data with no lesions or lesion areas with pixels smaller than 50% of the largest lesion pixel of the target object; S12. Three-dimensional segmentation: Using the MRI data obtained in S11 as input, a deep learning model is used to perform three-dimensional segmentation of HCC lesions on the input data; S13. Post-processing: Based on the segmentation results of step S12, perform connected component analysis and retain only the HCC connected region with the largest volume in three-dimensional space; based on the minimum bounding rectangle of the largest HCC connected region, crop and expand the original image to suppress the background, and adaptively weight and enhance the texture based on the contrast, energy, and correlation statistics of the gray-level co-occurrence matrix, and use the processed image patch as the final lesion area image data.
4. The method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information according to claim 3, characterized in that, The deep learning model in step S12 is a 3D U-Net that integrates non-local attention and pixel attention mechanisms within the nnU-Net framework; it uses a foreground-aware sampling strategy to prioritize sampling image patches containing lesion voxels for training.
5. The method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information according to claim 1, characterized in that, The method for designing structured prompt words in step S2 specifically includes: S21. Domain Knowledge Extraction: Based on the LI-RADS standard, extract the semantic features required for the diagnosis of hepatocellular carcinoma; the semantic features include lesion diameter, signal intensity, signal distribution, lesion morphology, lesion edge and anatomical location; S22. Structured cue word design: Based on semantic features, the image analysis task is uniformly defined as: identifying lesion features at anatomical locations during image sequences.
6. The method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information according to claim 1, characterized in that, The process of extracting specific features and common structural features sequentially from the multi-sequence lesion region images obtained in step S1 using a parallel dual-branch encoder includes: A parallel dual-branch encoder is constructed, which consists of a dedicated encoder and a shared encoder. The dedicated encoder is a three-dimensional encoder constructed separately for each MRI lesion region image data sequence, used to capture the sequence-specific lesion diagnostic information reflected by the sequence under its imaging mechanism and obtain dedicated features. The shared encoder is the same three-dimensional encoder shared by all MRI lesion region image data sequences, which is used to extract the common structural features of all sequences. The MRI lesion region image data of each sequence are simultaneously fed into a dedicated encoder and a shared encoder, and the two are processed in parallel.
7. The method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information according to claim 6, characterized in that, In step S31, the process of fusing specific features and common structural features through the enhanced lightweight residual fusion module to obtain the fused features of each sequence includes: An enhanced lightweight residual fusion module is constructed, which consists of modal branches and shared branches; In the modality branch: local spatial features are extracted through convolution operations. Subsequently, the extracted features are updated using a channel attention mechanism to obtain weighted features and their corresponding attention weights. ; In the modality sharing branch: local spatial features are extracted through convolution operations. Then, the extracted features are updated by weighting them using a local context attention mechanism to obtain weighted shared features and their corresponding attention weights. ; The weighted information is added to the shared information according to the following formula to obtain the fused feature. : ; in, This is represented as element-wise multiplication; Fusion features The image is processed using depthwise separable convolutions, and global context enhancement and residual connections are introduced to fuse multi-scale information. For a total of N image sequences, the fusion features of all sequences are concatenated along the channel dimension to obtain the concatenated features, which are the fusion features of each sequence. The fusion features of all sequences are then compressed through a single channel compression layer to form a unified image representation. : ; in, This indicates a splicing operation along the channel dimension. Indicates channel compression mapping, This represents the fusion feature corresponding to the i-th image sequence.
8. The method for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information according to claim 1, characterized in that, The method for predicting the pathological differentiation degree of liver cancer also includes: using a total loss function. To update the processing parameters; among which, and For weight parameters, For classifying losses, For cross-modal contrast loss, This represents the distribution alignment loss.
9. A system for predicting the pathological differentiation degree of liver cancer by integrating multi-sequence magnetic resonance imaging and semantic information, characterized in that, It includes a lesion segmentation module, a semantic description generation module, and a lesion differentiation prediction module; the lesion differentiation prediction module includes a multimodal feature extraction unit, a cross-modal feature fusion unit, and a classification prediction unit; each module works together to implement the liver cancer pathological differentiation degree prediction method as described in any one of claims 1 to 8.