Liver cancer image automatic segmentation and recognition method fusing multi-modal images

By combining dynamic processing procedures and segmentation network models, the artifact interference characteristics of multimodal image data are resolved, thus solving the artifact interference problem in existing technologies for multimodal image data. This enables automated and accurate tumor boundary segmentation under different equipment conditions and image quality, improving segmentation efficiency and stability, and supporting the development of clinical treatment plans.

CN121033079BActive Publication Date: 2026-02-10THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511524994.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-02-10
Estimated Expiration
2045-10-24

AI Technical Summary

Technical Problem

Existing multimodal medical imaging data is susceptible to artifact interference in liver cancer segmentation, resulting in unstable segmentation performance, difficulty in adapting to image data of different qualities, increased manual operation costs, and reduced segmentation efficiency and accuracy.

Method used

A dynamic processing flow and segmentation network model are adopted. Depending on whether the image data contains artifact interference features, temporal registration or feature enhancement processing is performed to generate tumor boundary segmentation results. The model parameters of multiple medical institutions are coordinated through a federated learning framework to dynamically optimize the convolution kernel size and feature receptive field of the network model.

Benefits of technology

It enables automated and precise tumor boundary segmentation under different equipment conditions and image quality, reduces manual intervention, improves segmentation efficiency and stability, provides detailed tumor morphological information, and supports the development of clinical treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033079B_ABST
    Figure CN121033079B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical image processing, and discloses a liver cancer image automatic segmentation and identification method fusing multi-modal images. The method comprises the following steps: acquiring multi-modal medical image data in a first processing stage, and judging whether the data contains a pseudo-interference feature; if the data contains the pseudo-interference feature, triggering a first processing procedure, performing time sequence registration processing on the data, acquiring a first image feature set after registration in a second processing stage, generating a first tumor boundary segmentation result based on a first segmentation network model and the first image feature set; if the data does not contain the pseudo-interference feature, triggering a second processing procedure, performing feature enhancement processing on the data, acquiring an enhanced second image feature set in a third processing stage, and generating a second tumor boundary segmentation result based on a second segmentation network model and the second image feature set. The method can be used for targeted processing of multi-modal image data of different qualities, improves segmentation adaptability and reliability, and meets the requirements of clinical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, specifically to an automatic segmentation and recognition method for liver cancer images that integrates multimodal images. Background Technology

[0002] In the clinical diagnosis and treatment of liver cancer, medical image segmentation technology is a crucial means of achieving precise tumor localization, and its results directly affect the formulation and implementation of subsequent treatment plans. With the development of medical imaging technology, multimodal medical imaging data (such as CT, MRI, and ultrasound) have gradually become the main data source for liver cancer image analysis due to their ability to provide richer information about the tumor and surrounding tissues. However, current liver cancer segmentation methods based on multimodal imaging still face many challenges in practical applications.

[0003] Multimodal medical imaging data is susceptible to interference from various factors during acquisition, resulting in artifacts. For example, a patient's respiratory movements may cause motion artifacts in CT images, and electromagnetic interference from equipment may cause signal artifacts in MRI images. These artifacts can obscure or distort the true features of the tumor region, making it difficult to accurately match the spatial positions between different modalities, thereby affecting the accuracy of subsequent feature extraction and segmentation results.

[0004] Existing segmentation methods rely on relatively simple preprocessing approaches for image data, lacking targeted strategies. For multimodal image data free of artifact interference, conventional preprocessing procedures are insufficient to fully extract subtle tumor features, such as grayscale variations at tumor edges and textural differences between tumors and normal tissues. Furthermore, traditional segmentation network models are often designed for single-modal images or fixed types of multimodal images, failing to dynamically adjust processing procedures and model parameters based on the actual quality of the image data (whether artifacts are present). This results in insufficient segmentation performance stability when dealing with multimodal image data of varying quality, making it difficult to meet clinical demands for accuracy and reliability in liver cancer segmentation.

[0005] In clinical practice, the quality of multimodal imaging data varies significantly. Some primary healthcare institutions experience even higher rates of artifacts in their acquired images due to equipment limitations or inadequate adherence to operational procedures. Existing segmentation methods struggle to adapt to this diverse data quality, often requiring manual assessment of image quality before selecting appropriate processing methods. This not only increases manual costs but may also introduce additional interference due to human error, further reducing segmentation efficiency and accuracy. This hinders the widespread clinical adoption and application of automated liver cancer segmentation technology. Summary of the Invention

[0006] The purpose of this invention is to provide an automatic segmentation and recognition method for liver cancer images that integrates multimodal images, so as to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention provides an automatic segmentation and recognition method for liver cancer images by fusing multimodal imaging, the method comprising:

[0008] In the first processing stage, multimodal medical image data is acquired, and it is determined whether the multimodal medical image data contains artifact interference features.

[0009] If the multimodal medical image data contains artifact interference features, the first processing flow is triggered: the multimodal medical image data is subjected to temporal registration processing, the first image feature set after registration is obtained in the second processing stage, and the first tumor boundary segmentation result is generated based on the first segmentation network model and the first image feature set.

[0010] If the multimodal medical image data does not contain artifact interference features, the second processing flow is triggered: feature enhancement processing is performed on the multimodal medical image data, the enhanced second image feature set is obtained in the third processing stage, and the second tumor boundary segmentation result is generated based on the second segmentation network model and the second image feature set.

[0011] Preferably, the generation of the first tumor boundary segmentation result based on the first segmentation network model and the first image feature set includes:

[0012] The first image feature set is input into the first feature encoder to extract multi-scale tumor feature maps;

[0013] The probability distribution of tumor regions is calculated based on the multi-scale tumor feature map.

[0014] The first tumor boundary segmentation result is generated based on the probability distribution of the tumor region.

[0015] Preferably, the generation of the second tumor boundary segmentation result based on the second segmentation network model and the second image feature set includes:

[0016] The second image feature set is input into the second feature encoder to extract the spatial correlation feature map;

[0017] Calculate vascular invasion characteristic indicators based on the spatial correlation feature map;

[0018] A second tumor boundary segmentation result is generated based on the aforementioned vascular invasion feature indicators.

[0019] Preferably, the method further includes:

[0020] Obtain the tumor region features corresponding to the first tumor boundary segmentation result or the second tumor boundary segmentation result;

[0021] The tumor region features are input into the recognition network model to calculate the malignancy grade index.

[0022] The liver cancer subtype identification results are generated based on the malignancy grading index.

[0023] Preferably, the step of generating liver cancer subtype identification results based on the malignancy grading index includes:

[0024] Determine whether the high-risk threshold has been reached based on the aforementioned severity grading index;

[0025] If the high-risk threshold is reached, the first identification path is triggered: extracting microvascular invasion features and generating liver cancer subtype identification results based on the microvascular invasion features;

[0026] If the high-risk threshold is not reached, the second identification path is triggered: extracting tumor capsule integrity features and generating liver cancer subtype identification results based on the tumor capsule integrity features.

[0027] Preferably, the method further includes:

[0028] The parameters of local segmentation models from multiple medical institutions are coordinated based on a federated learning framework.

[0029] The weights of the first segmentation network model and the second segmentation network model are updated based on the coordinated model parameters;

[0030] The multimodal medical image data is reprocessed using the updated weights.

[0031] Preferably, the parameters of the local segmentation model coordinating multiple medical institutions based on the federated learning framework include:

[0032] Obtain the differences in the distribution of local model features among various medical institutions;

[0033] Calculate the model aggregation weights based on the differences in the feature distributions;

[0034] The model aggregates weights and fuses the local segmentation model parameters of the multiple medical institutions.

[0035] Preferably, the method further includes:

[0036] Dynamic optimization parameters are generated based on the first tumor boundary segmentation result or the second tumor boundary segmentation result.

[0037] The convolution kernel sizes of the first segmentation network model and the second segmentation network model are adjusted based on the dynamic optimization parameters.

[0038] The adjusted convolution kernel size is used to process subsequent input multimodal medical image data.

[0039] Preferably, adjusting the convolution kernel size of the first segmentation network model and the second segmentation network model based on the dynamic optimization parameters includes:

[0040] Calculate the feature receptive field adjustment based on the dynamic optimization parameters;

[0041] The convolutional layer parameters of the first segmentation network model and the second segmentation network model are updated based on the feature receptive field adjustment.

[0042] Preferably, the method further includes:

[0043] Obtain the pathological feature vector corresponding to the liver cancer subtype identification result;

[0044] The pathological feature vectors are fed back to the feature decoders of the first segmentation network model and the second segmentation network model;

[0045] The boundary smoothness parameters for subsequent tumor boundary segmentation are optimized based on the feedback results.

[0046] Compared with the prior art, the beneficial effects of the present invention are:

[0047] When multimodal medical image data contains artifact interference features, the temporal registration process in the triggered first processing flow can effectively eliminate the impact of artifacts on the spatial matching of different modal images, enabling precise alignment of each modal image in both time and space dimensions. This provides a more reliable foundation for the subsequent acquisition of the first image feature set. The generation of the first tumor boundary segmentation result based on the first segmentation network model and the first image feature set can fully utilize the complementary information of each registered modal image, reduce feature distortion caused by artifacts, more accurately capture the spatial location and morphological features of the tumor region, and avoid segmentation omissions or misjudgments caused by artifact interference.

[0048] When multimodal medical image data does not contain artifact interference features, the triggered second processing flow includes feature enhancement, which can specifically strengthen subtle tumor features in the image. This includes enhancing grayscale gradient changes at tumor edges and highlighting texture differences between tumors and normal tissues. This allows the second image feature set to more comprehensively and clearly reflect the true characteristics of the tumor. Combining the second image feature set with a second segmentation network model can fully uncover potential tumor features within the image, more accurately define tumor boundaries, restore the true morphology of the tumor, and avoid problems such as blurred or inaccurate segmentation boundaries due to insufficient feature extraction.

[0049] This method achieves adaptive adjustment of image data quality by dynamically selecting the processing flow and corresponding segmentation network model. It automates the entire process from data quality assessment to segmentation result generation without manual intervention, significantly reducing manual steps, lowering labor costs and mitigating the impact of human error, thus improving the efficiency and stability of liver cancer segmentation. In clinical applications, whether in large hospitals with superior equipment and fewer image data artifacts, or in primary healthcare institutions with limited equipment and a higher incidence of image data artifacts, this method can automatically match the optimal processing scheme based on the actual quality of the acquired image data, ensuring reliable tumor boundary segmentation results in different scenarios.

[0050] This method, through phased processing and targeted model application, can fully leverage the advantages of multimodal image data. After the adaptation process, the complementary nature of the features of different image modalities is more fully utilized. It can not only more accurately identify tumor areas, but also effectively distinguish tumors from surrounding normal tissues, blood vessels and other structures, providing clinicians with more detailed information on tumor morphology and location. This helps doctors to have a more comprehensive understanding of the tumor situation and provides more accurate imaging evidence for the formulation of treatment plans. At the same time, it creates favorable conditions for the promotion and application of automatic liver cancer segmentation technology in medical institutions at different levels, and better meets the diverse application needs of clinical practice. Attached Figure Description

[0051] Figure 1 This is a schematic diagram illustrating the working principle of the automatic segmentation and recognition method for liver cancer images fused with multimodal images as described in this invention.

[0052] Figure 2 The flowchart for processing the first segmentation network model;

[0053] Figure 3 A flowchart generated from the results of liver cancer subtype identification;

[0054] Figure 4 Flowchart for updating the federated learning model;

[0055] Figure 5 A flowchart for dynamically optimizing the convolution kernel size. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Please see Figure 1This invention provides an automatic segmentation and recognition method for liver cancer images based on the fusion of multimodal images, the method comprising:

[0058] In the first processing stage, multimodal medical image data is acquired, which typically includes, but is not limited to, image information from different modalities such as CT, MRI, and PET. Artifact interference feature analysis is performed on the multimodal medical image data. A pre-trained convolutional neural network is used to extract noise patterns, motion artifacts, or metallic artifact features from the images, and a classifier is used to determine whether significant artifact interference features exist. If the multimodal medical image data contains artifact interference features, the first processing flow is triggered: temporal registration processing is performed on the multimodal medical image data. An elastic registration algorithm based on mutual information is used to align image sequences at different time points or different modalities. In the second processing stage, the registered first image feature set is acquired. This feature set contains spatially aligned multimodal image patches. Based on the first segmentation network model and the first image feature set, a first tumor boundary segmentation result is generated. The first segmentation network model adopts an encoder-decoder structure. The encoder extracts hierarchical features, and the decoder gradually restores spatial details and outputs a segmentation mask. If the multimodal medical image data does not contain artifact interference features, the second processing flow is triggered: feature enhancement processing is performed on the multimodal medical image data, applying contrast-limited adaptive histogram equalization or deep learning-based super-resolution reconstruction to improve image clarity; in the third processing stage, the enhanced second image feature set is obtained, which includes image data with texture enhancement and edge sharpening; based on the second segmentation network model and the second image feature set, a second tumor boundary segmentation result is generated, and the second segmentation network model integrates an attention mechanism to focus on the tumor region.

[0059] Example 1: See Figure 2 In the multimodal medical image data acquisition phase, the system receives medical images from different imaging devices, including computed tomography (CT), magnetic resonance imaging (MRI), and positron emission tomography (PET). These images may be affected by various artifacts during acquisition, such as motion artifacts caused by patient movement, metal artifacts from implants, or noise artifacts caused by equipment calibration issues. To accurately determine the presence of these artifact interference features, the system uses a pre-trained deep convolutional network to analyze the input images. This network identifies regions where artifact interference may exist by analyzing the image's frequency features, texture patterns, and local consistency. The network outputs a probability value representing the degree to which the image is affected by artifacts. When this probability value exceeds a preset threshold, the system determines that the data contains artifact interference features.

[0060] When artifact interference is detected, the system initiates the first processing flow. Temporal registration is a crucial step in this flow, aiming to spatially align images acquired at different time points or in different modalities. This registration process is particularly important because patients may experience positional changes or organ movement during multiple scans. The system employs a mutual information-based elastic registration algorithm, which can handle grayscale differences between images from different modalities while adapting to local deformations. The registration process first extracts key feature points from the images, then establishes correspondences between these feature points, and finally aligns the moving image with the reference image using an elastic transformation model. After registration, the system obtains a spatially consistent first image feature set, which contains aligned image data from different modalities.

[0061] The first segmentation network model is responsible for processing the first image feature set after registration. This network adopts an encoder-decoder architecture. The encoder part consists of multiple convolutional and pooling layers, progressively extracting multi-scale features from the image. During the encoding process, the network generates tumor feature maps at different resolutions. These feature maps contain both low-level features rich in detail and high-level features with clear semantic information. The decoder part then uses upsampling and skip connection operations to progressively restore spatial details and fuse features at different scales. The network finally outputs a probability distribution map of each pixel belonging to the tumor region. This probability distribution reflects the model's confidence in the tumor location.

[0062] The process of generating the first tumor boundary segmentation result based on the probability distribution of tumor regions involves post-processing operations. The system first performs thresholding on the probability map, marking regions with probability values ​​higher than the threshold as tumor regions. Then, morphological operations are used to optimize the segmentation result, including filling small holes, smoothing boundaries, and removing isolated noise points. These operations help improve the accuracy and visual continuity of the segmentation result. When the input multimodal medical image data does not contain significant artifact interference features, the system triggers a second processing flow. This flow begins with feature enhancement processing, aiming to improve the visual quality and feature discriminability of the image. The system employs a contrast-limited adaptive histogram equalization technique, which can enhance local contrast while suppressing noise amplification. The processing divides the image into multiple small regions, performs histogram equalization independently within each region, and avoids over-enhancement through contrast limiting. For particularly blurry or low-contrast regions, the system may also employ a deep learning-based super-resolution method, using a trained neural network to recover image details.

[0063] The enhanced image forms a second image feature set, which is then input into a second segmentation network model. This network, based on the standard encoder-decoder structure, incorporates an attention mechanism and a specialized feature extraction module. The second feature encoder uses dilated convolutions to expand the receptive field, enabling it to capture a wider range of contextual information. This design allows the network to extract spatial correlation feature maps containing the relationship between the tumor and surrounding tissues. These feature maps not only include features of the tumor itself but also spatial relationship information between the tumor and surrounding structures such as blood vessels and liver parenchyma.

[0064] The process of calculating vascular invasion characteristic indicators based on spatial correlation feature maps involves a specially designed feature analysis module. This module analyzes the spatial relationship between the tumor margin region and vascular structures, quantifying the likelihood of vascular invasion by calculating directional gradient histograms and texture features. These characteristic indicators include the angle between the vascular orientation and the tumor boundary, the density changes of the vessels near the tumor, and the integrity characteristics of the vessel wall.

[0065] When generating the second tumor boundary segmentation result based on vascular invasion feature indicators, the system incorporates these indicators as additional constraints into the segmentation process. During network training, vascular invasion features are used as auxiliary supervisory signals to help the network learn more accurate boundary features. During inference, these feature indicators are used to adjust the precise location of the segmentation boundary, especially in the region where the tumor and blood vessels meet, making the segmentation result more consistent with anatomical facts.

[0066] After each processing stage, the system performs a quality assessment of the intermediate results. If abnormal or low-quality results are found, corresponding reprocessing or error correction procedures are triggered. This design ensures that the system can produce reliable segmentation results when faced with input data of varying quality. The network implementation adopts a modular design, allowing each functional module to be independently updated and optimized. The encoder uses a pre-trained convolutional neural network as the backbone, while the decoder employs custom upsampling and feature fusion modules. A multi-task learning strategy is used during training to simultaneously optimize segmentation accuracy and auxiliary task performance. This design enables the network to learn richer and more robust feature representations. In practical deployment, the system also considers computational efficiency and real-time requirements. Through model quantization and inference optimization techniques, the system ensures efficient operation on standard medical computing equipment. Furthermore, the system provides adjustable parameter settings, allowing users to adjust them according to specific application scenarios and accuracy requirements.

[0067] Example 2: See Figure 3Before initiating the subtype identification process, the system first acquires tumor boundary segmentation results from the first or second processing flow. These segmentation results identify tumor regions in liver CT or MRI images using binary masks. Based on these segmentation results, the system extracts corresponding tumor region features from the original multimodal images. The extraction process employs a multi-scale analysis method to capture the morphological features of the tumor at different resolutions. These features include overall tumor shape descriptors, such as geometric parameters like roundness, concavity / convexity, and aspect ratio; texture features are calculated using the gray-level co-occurrence matrix, including quantitative indicators such as contrast, correlation, and entropy; and intensity features statistically analyze the distribution characteristics of pixel values ​​within the tumor region, including statistics such as mean, variance, and skewness.

[0068] When tumor region features are input into the recognition network model, the system employs a deep neural network architecture for processing. The network contains multiple fully connected layers, with non-linear activation functions used for transformation between each layer. The network design considers the relevance and importance of medical features, assigning different weights to different types of features through an attention mechanism. For example, texture features with high diagnostic value are automatically assigned higher attention weights. During the forward propagation of the network, the feature vector undergoes layer-by-layer transformations, ultimately outputting a malignancy grading index. This index is a continuous numerical value reflecting the potential invasiveness of the tumor, and its calculation is based on weight parameters trained using a large amount of clinical data.

[0069] The process of generating liver cancer subtype identification results based on malignancy grading indicators employs a hierarchical decision-making mechanism. The system sets a high-risk threshold, determined through statistical analysis of a large amount of clinical case data. When the malignancy grading index reaches or exceeds this threshold, it indicates that the tumor has a high malignancy potential, triggering the first identification path. In this path, the focus is on analyzing microvascular invasion characteristics, a crucial indicator for assessing liver cancer invasiveness. The system extracts microvascular patterns from the tumor margin region from enhanced image data and uses a high-precision segmentation algorithm to identify smaller-diameter vascular structures. By analyzing the distribution density, orientation, and relationship of blood vessels to the tumor boundary, the system quantifies the severity of microvascular invasion. These features are input into a specially trained subtype classifier, which can distinguish different types of vascular invasion patterns and generate specific liver cancer subtype identification results accordingly.

[0070] When the malignancy grading index does not reach the high-risk threshold, the system activates a second identification path, which focuses on analyzing the integrity characteristics of the tumor capsule. The tumor capsule is an important indicator in liver cancer pathology, and its integrity is closely related to the biological behavior of the tumor. The system uses an edge enhancement algorithm to strengthen the display of the capsule region, and then uses a dynamic programming algorithm to track the continuity of the capsule. The feature extraction process includes capsule thickness uniformity analysis, integrity scoring, and the location and quantification of breakpoints. The system also evaluates the interface characteristics between the capsule and surrounding liver tissue, including parameters such as clarity and regularity. All these feature parameters are integrated into a comprehensive evaluation model, which learns the correlation between capsule features and liver cancer subtypes through machine learning algorithms, and finally outputs subtype identification results based on capsule integrity characteristics.

[0071] Throughout the identification process, the system employs a multi-dimensional feature fusion strategy. In addition to key features such as microvascular invasion or capsule integrity, the system integrates auxiliary features including tumor size, location, and multifocality. These features are fused using a feature weighting mechanism, with more important features assigned higher weights. The identification network is trained end-to-end, learning the mapping relationship from raw features to subtype classification through extensive labeled data. The network's output layer uses a softmax function to generate the probability distribution for each subtype, ultimately selecting the category with the highest probability as the identification result.

[0072] The system also incorporates result verification and confidence assessment mechanisms. For each generated subtype identification result, the system calculates a confidence score based on feature consistency and the distribution of classification probabilities. When the confidence score falls below a predetermined threshold, the system initiates a review process, improving the reliability of the results through additional feature analysis or manual review. Furthermore, the system establishes a feedback learning mechanism, feeding the final pathologically confirmed results back into the identification network for continuous model optimization and improvement. In clinical practice, this implementation considers differences in equipment and imaging protocols across different medical institutions. The system exhibits strong adaptability, capable of handling image data with varying resolutions and contrast agent enhancement schemes. The feature extraction process employs standardized processing to ensure consistent feature representations across data from different sources. The identification network is also trained on multi-center data, demonstrating good generalization ability. The system also provides interpretable output, offering not only the subtype identification result but also key feature evidence supporting that conclusion. For example, when identified as a specific subtype, the system highlights the most contributing feature indicators, such as vascular patterns or capsule features, which helps clinicians understand and verify the automated identification results. This interpretable design enhances the system's usability and credibility in clinical practice.

[0073] Through feature pre-computation and model optimization, the entire process from segmentation to identification can be completed within a reasonable timeframe, meeting real-time clinical needs. The system adopts a modular design, allowing each processing stage to be independently optimized and updated, maintaining system scalability and maintainability. It achieves a complete automated workflow from raw medical images to liver cancer subtype identification, providing crucial auxiliary information for clinical decision-making. By integrating multimodal image features and deep learning technology, the system can handle complex medical image analysis tasks, demonstrating its application value in the field of precision liver cancer diagnosis.

[0074] Example 3: See Figure 4 In the initialization phase of the federated learning framework, the central server first defines a unified model architecture and initial parameters. The specific architectures of the first and second segmentation network models are standardized, including hyperparameters such as the number of encoder layers, the number of convolutional kernels, and the upsampling method of the decoder. Each participating medical institution deploys these models locally and trains them using its own multimodal medical image data. Each institution maintains an independent database containing image data from different sources, such as CT and MRI, along with their corresponding annotation information. During local training, each institution uses stochastic gradient descent to update model parameters, and the loss function incorporates metrics such as segmentation accuracy and boundary smoothness.

[0075] Obtaining the differences in feature distribution among local models from various medical institutions is a crucial step in federated learning. The central server evaluates the variability in feature distribution by analyzing the model parameters uploaded by each institution. Specifically, the server calculates the feature output of each local model on a standard validation set and compares the statistical properties of these outputs. The difference assessment includes the bias of feature means, the correlation of feature dimensions, and the coverage of the feature space. For example, for the feature encoder output of the first segmentation network model, the server calculates the distribution differences in the channel dimension of the feature maps generated by the models from each institution. These differences are quantified as distance metrics for subsequent weight calculations.

[0076] The process of calculating the model aggregation weights based on differences in feature distributions employs an adaptive weighting strategy. The weight calculation considers not only the model's performance metrics but also the uniqueness and diversity of the institution's data distribution. The server calculates the aggregation weight for each institution using the following formula:

[0077]

[0078] in: This represents the aggregate weight of the k-th medical institution. It involves adjusting parameters to control the smoothness of the weight distribution. Measuring the characteristic distribution of the k-th institution and the global characteristic distribution The difference in distance between them This represents the total number of institutions participating in federal learning. Distance difference The calculation is based on the KL divergence or Wasserstein distance of the characteristic distribution, which can effectively capture the essential differences between distributions.

[0079] When fusing local segmentation model parameters from multiple medical institutions based on model aggregation weights, the central server performs a weighted averaging operation. For each parameter matrix in the model, the server calculates its weighted average:

[0080]

[0081] in: Represents global model parameters. These are the local model parameters of the k-th institution. This aggregation method ensures that institutions with more representative and higher-quality feature distributions contribute more to the global model.

[0082] After updating the weights of the first and second segmentation network models based on the coordinated model parameters, the central server distributes the updated global model to each participating institution. Each institution initializes its local model using the global model parameters and then continues fine-tuning training using local data. This process is repeated periodically, forming multiple rounds of federated learning. In each round, the server re-evaluates the differences in feature distribution among the institutions and dynamically adjusts the aggregation weights to adapt to changes in data distribution.

[0083] When reprocessing multimodal medical imaging data using the updated weights, each institution first processes its local images to be analyzed using the latest global model. The processing includes forward propagation computation and post-processing. For the first segmentation network model, the focus is on the registered multimodal data; the model leverages the enhanced generalization ability obtained through federated learning to better handle differences in imaging equipment across different institutions. The second segmentation network model focuses on the feature-enhanced data, improving the segmentation accuracy of fine structures through shared knowledge.

[0084] The federated learning framework also includes privacy protection mechanisms. Institutions upload only model parameters, not the original data. Differential privacy technology adds appropriate noise to the parameters to prevent the original data from being inferred from the model parameters. Simultaneously, encrypted transmission is used during communication to ensure the security of parameter exchange. In practical implementation, the system also incorporates an anomaly detection mechanism. The server monitors the model parameters uploaded by each institution, detecting outliers or malicious attacks. For parameters deviating from the normal range, the system automatically reduces their weight or temporarily excludes the institution from aggregation, maintaining the stability of federated learning. The central server maintains the version history of the global model, recording the participating institutions, aggregation weights, and performance metrics for each round of federated learning. This allows for rollback to previous model versions when needed and facilitates tracking the model's evolution.

[0085] This federated learning framework also supports incremental learning capabilities. When a new medical institution wants to join, it can directly use the current global model as initial parameters and then participate in subsequent federated learning rounds. This design allows the system to continuously expand its knowledge base while maintaining model compatibility and consistency. Through this distributed learning approach, medical institutions can collaboratively contribute to improving the liver cancer segmentation model while protecting data privacy. The data characteristics of different institutions, such as different population distributions, different equipment parameters, and different acquisition protocols, can all be integrated and utilized through federated learning, resulting in a more generalizable model. This collaborative model is particularly suitable for the medical field, respecting data privacy regulations while pooling resources to advance the development of medical artificial intelligence.

[0086] Example 4: This example details the process of dynamically optimizing network model parameters based on tumor boundary segmentation results. This implementation analyzes historical segmentation performance metrics and automatically adjusts key model parameters, enabling the segmentation network to adaptively handle medical image data with different characteristics and improving the model's adaptability and stability in various clinical scenarios. Before starting the dynamic optimization process, the system first collects relevant performance metrics from either the first or second tumor boundary segmentation results. These metrics are derived from the analysis of processed cases, including segmentation accuracy assessment, boundary sharpness measurement, and tumor morphological feature statistics for each case. Segmentation accuracy assessment is achieved by calculating the overlap between the predicted segmentation region and expert annotations; these metrics reflect the model's performance in specific cases. Boundary sharpness measurement analyzes the sharpness and continuity of the segmentation boundaries, quantified by edge gradient strength and boundary smoothness parameters. Tumor morphological feature statistics include geometric features such as tumor size distribution, shape complexity, and location information. All this data is organized into a structured set of optimization parameters, serving as the basis for dynamic adjustment.

[0087] The process of generating dynamic optimization parameters employs time-series analysis. The system maintains a sliding window of historical performance data, continuously monitoring and recording the segmentation results of several recently processed cases. Analyzing this data, the system identifies performance trends and patterns, such as persistently low tumor segmentation accuracy within a specific size range or insufficient tumor boundary clarity at certain locations. Based on these analyses, the system generates a set of numerical optimization parameters that encode the direction and extent to which the model needs adjustment. These optimization parameters include numerical values ​​such as convolution kernel size adjustment, receptive field correction coefficients, and feature extraction intensity.

[0088] When adjusting the convolutional kernel sizes of the first and second segmentation network models based on dynamically optimized parameters, the system implements a parameter mapping mechanism. Each optimized parameter is mapped to a specific network layer and convolutional kernel adjustment instruction. For example, when the system detects a decline in segmentation performance for small tumors, it generates optimized parameters to reduce the convolutional kernel size, enabling the network to capture finer feature details. Conversely, for large tumors or diffuse lesions, the system may suggest increasing the convolutional kernel size to obtain broader contextual information. The adjustment process is not a simple replacement of the convolutional kernels, but rather a smooth transition achieved through reparameterization, avoiding drastic fluctuations in model performance.

[0089] When processing subsequent input multimodal medical image data using the adjusted convolutional kernel size, the system employs a progressive tuning strategy. New convolutional kernel parameters are first tested on a validation set to evaluate their adaptability to different data types. They are then gradually applied to actual case processing while performance changes are continuously monitored. This progressive application approach ensures the stability and reliability of the model tuning. Table 1 shows some key parameter tuning examples recorded by the system during dynamic optimization:

[0090] Table 1: Record of dynamic adjustment of convolution kernel size.

[0091]

[0092] When calculating the receptive field adjustment based on dynamic optimization parameters, the system considers the combined influence of multiple factors. Besides tumor size, these include the complexity of the tumor boundary, the contrast difference with surrounding tissues, and the image quality score. The system establishes a multidimensional optimization function that maps these factors to specific receptive field adjustment suggestions. Receptive field adjustment involves not only the convolution kernel size but also adjustments to the dilation rate and modifications to the pooling operation.

[0093] When updating the convolutional layer parameters of the first and second segmentation network models based on feature receptive field adjustments, the system employs structural reparameterization techniques. For convolutional kernels requiring enlargement, the system achieves an equivalent large receptive field by combining multiple small convolutional kernels and dilated convolutions, while maintaining computational efficiency. For convolutional kernels requiring reduction, the system uses techniques such as depthwise separable convolution and grouped convolution to reduce kernel size while maintaining feature extraction capabilities. The update process ensures network structure compatibility; all adjustments are made within the existing architectural framework without altering the overall network topology. A rollback mechanism is also implemented. After each parameter adjustment, the system retains the model state before the adjustment for a period. If performance degradation or anomalies are detected during subsequent monitoring, the system can quickly revert to its previous stable state. This design ensures the reliability and safety of clinical applications.

[0094] The dynamic optimization process also includes a cross-model coordination mechanism. When simultaneously adjusting the first and second segmentation network models, the system ensures that the parameter adjustments between the two models remain consistent. For example, when increasing the convolution kernel size of the first segmentation network model, the relevant parameters of the second segmentation network model are adjusted accordingly to maintain consistency between the two models in handling similar features. The system also implements a case-type-based differentiated adjustment strategy. For different types of liver cancer cases, such as nodular, massive, or diffuse types, the system employs different optimization parameter mapping rules. These rules are based on statistical analysis of a large amount of clinical data to ensure that the adjustment strategy conforms to the characteristics of different pathological types.

[0095] All adjustments are performed automatically in the background without manual intervention. Clinicians only need to focus on the quality of the final segmentation results, without needing to concern themselves with the details of the underlying model adjustments. The system provides adjustment logs and performance reports, facilitating technical personnel to monitor the optimization process and maintain system operation. Through this dynamic optimization mechanism, the segmentation network model can continuously adapt to changing clinical needs and data characteristics, maintaining the stability and advancement of segmentation performance. This adaptive design enables the system to maintain reliable performance in diverse clinical environments, providing continuous technical support for liver cancer diagnosis.

[0096] Example 5: In the stage of obtaining the pathological feature vector corresponding to the liver cancer subtype identification results, the system extracts multi-dimensional pathological semantic features from the subtype classification model. These features include subtype classification probability distribution, representing the probability of belonging to various liver cancer subtypes; malignancy score, quantifying the invasiveness level of the tumor; morphological descriptors, encoding the tumor's growth pattern and structural characteristics; and molecular feature simulation vectors, reflecting potential gene expression patterns. The system fuses these heterogeneous information into a fixed-dimensional pathological feature vector through a feature encoding layer. This vector carries the biological behavioral characteristics of the tumor. For example, for highly invasive subtypes, the vector will strengthen the encoding of invasive growth patterns; for subtypes with expansive growth, the feature weight of capsule integrity will be emphasized.

[0097] When feeding pathological feature vectors back to the feature decoders of the first and second segmentation network models, the system employs a dedicated feature fusion interface. At each upsampling stage of the decoder, the pathological feature vector is transformed through a fully connected layer to match the number of channels in the current feature map, and then a feature concatenation operation is performed. The concatenated feature map is then processed by convolutional layers to achieve cross-modal feature interaction, enabling deep fusion of pathological semantic information and visual features. Particularly in the second segmentation network model, the system employs an attention modulation mechanism: the pathological feature vector generates a spatial attention weight map, dynamically adjusting the feature response intensity at different locations in the decoder. For vascular invasive hepatocellular carcinoma, this mechanism enhances feature activation in the tumor margin region; for fibrolamellar hepatocellular carcinoma, it focuses on feature extraction from the internal septa.

[0098] When optimizing the boundary smoothness parameters for subsequent tumor boundary segmentation based on feedback results, the system reconstructs the post-segmentation processing flow. Traditional conditional random fields only consider the grayscale consistency and spatial continuity of the image itself, while this system introduces a pathological feature constraint term. This constraint term defines different boundary regularization intensities based on subtype characteristics: for invasive hepatocellular carcinoma with blurred boundaries, the smoothing constraint intensity is reduced to preserve irregular edges; for expansive hepatocellular carcinoma with clear boundaries, the smoothing constraint is enhanced to eliminate jagged edges caused by noise. In specific implementation, the system constructs a triplet structure for the energy function: unary potential energy combines original image features, binary potential energy maintains spatial continuity, and the newly added ternary potential energy encodes the association rules between pathological features and boundary morphology. This design ensures that the segmentation boundary conforms to both local image features and subtype pathological characteristics.

[0099] When processing a new case, the system retrieves cases with similar pathological features from the historical database and analyzes their optimal boundary smoothness parameter configurations. Using a feature matching algorithm, it identifies several historical cases whose pathological feature vectors are closest to the current case, and takes the average of their boundary smoothness parameters as the initial setting. During processing, the system monitors boundary quality indicators in real time, such as the degree of boundary fragmentation and the frequency of curvature abrupt changes, and fine-tunes the smoothness parameters based on the monitoring results. This transfer learning strategy based on similar cases significantly improves the accuracy and adaptability of parameter settings.

[0100] When the first segmentation network model processes the registered multi-phase data, the system transmits the vascular feature knowledge learned by the second segmentation network model through pathological feature vectors. Specifically, vascular invasion feature indicators are encoded as sub-components of the pathological feature vectors. When the decoder of the first segmentation network model receives this component, it activates the vascular perception convolutional kernel, enhancing the ability to analyze multi-phase enhancement patterns. This design achieves synergistic enhancement between the two segmentation network models. After each pathological feature vector feedback, the system automatically evaluates the impact of this operation on the segmentation results. Evaluation methods include comparing the boundary differences before and after feedback, calculating the rate of change of the feature activation map, and other indicators. If the evaluation results show that the feedback has no significant impact, the system will initiate a feature recalibration process: adjusting the encoding weights of the pathological feature vectors or enhancing the expressive power of the feature fusion layer to ensure the effective transmission of the feedback signal.

[0101] Over long-term operation, the system develops a knowledge base mapping pathological features to segmentation parameters. The processing record for each successful case is stored, including the pathological feature vector, the boundary smoothness parameter configuration used, and the final segmentation quality score. By periodically analyzing this data, the system automatically discovers hidden correlations between different subtypes and optimal parameter settings, and updates its parameter recommendation strategy accordingly. This self-evolutionary mechanism enables the system to continuously accumulate clinical experience and constantly improve the accuracy of boundary optimization.

[0102] The entire implementation process adopts a modular design, with the feedback path integrated as an independent plug-in into the segmentation network. This architecture allows for flexible adjustments to the feedback strategy and parameter optimization algorithm without altering the core segmentation model. Simultaneously, the system provides visualization tools to demonstrate how pathological features influence the specific boundary formation process; for example, highlighting the boundary segments most affected by pathological features, providing doctors with intuitive interpretive analysis. Through this closed-loop feedback mechanism, the segmentation network model can transform high-level pathological diagnostic knowledge into low-level image processing constraints, ensuring that segmentation boundaries not only meet visual accuracy requirements but also align with the biological characteristics of tumors. This segmentation optimization strategy, which integrates pathological semantics, enhances the clinical relevance of the segmentation results while maintaining technical rigor.

[0103] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0104] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for automatic segmentation and recognition of liver cancer images by fusing multimodal imaging, characterized in that, include: In the first processing stage, multimodal medical image data is acquired, and it is determined whether the multimodal medical image data contains artifact interference features. If the multimodal medical image data contains artifact interference features, the first processing flow is triggered: the multimodal medical image data is subjected to temporal registration processing, the first image feature set after registration is obtained in the second processing stage, and the first tumor boundary segmentation result is generated based on the first segmentation network model and the first image feature set. If the multimodal medical image data does not contain artifact interference features, then the second processing flow is triggered: the multimodal medical image data is subjected to feature enhancement processing, the enhanced second image feature set is obtained in the third processing stage, and the second tumor boundary segmentation result is generated based on the second segmentation network model and the second image feature set. Also includes: Obtain the tumor region features corresponding to the first tumor boundary segmentation result or the second tumor boundary segmentation result; The tumor region features are input into the recognition network model to calculate the malignancy grade index. Based on the aforementioned malignancy grading indicators, liver cancer subtype identification results are generated; The process of generating liver cancer subtype identification results based on the malignancy grading index includes: Determine whether the high-risk threshold has been reached based on the aforementioned severity grading index; If the high-risk threshold is reached, the first identification path is triggered: extracting microvascular invasion features and generating liver cancer subtype identification results based on the microvascular invasion features; If the high-risk threshold is not reached, the second identification path is triggered: extracting tumor capsule integrity features and generating liver cancer subtype identification results based on the tumor capsule integrity features.

2. The automatic segmentation and recognition method for liver cancer images based on fused multimodal images according to claim 1, characterized in that, The generation of the first tumor boundary segmentation result based on the first segmentation network model and the first image feature set includes: The first image feature set is input into the first feature encoder to extract multi-scale tumor feature maps; The probability distribution of tumor regions is calculated based on the multi-scale tumor feature map. The first tumor boundary segmentation result is generated based on the probability distribution of the tumor region.

3. The automatic segmentation and recognition method for liver cancer images based on fused multimodal images according to claim 1, characterized in that, The generation of the second tumor boundary segmentation result based on the second segmentation network model and the second image feature set includes: The second image feature set is input into the second feature encoder to extract the spatial correlation feature map; Calculate vascular invasion characteristic indicators based on the spatial correlation feature map; A second tumor boundary segmentation result is generated based on the aforementioned vascular invasion feature indicators.

4. The automatic segmentation and recognition method for liver cancer images based on fused multimodal images according to claim 1, characterized in that, Also includes: The parameters of local segmentation models from multiple medical institutions are coordinated based on a federated learning framework. The weights of the first segmentation network model and the second segmentation network model are updated based on the coordinated model parameters; The multimodal medical image data is reprocessed using the updated weights.

5. The automatic segmentation and recognition method for liver cancer images based on fused multimodal images according to claim 4, characterized in that, The parameters of the local segmentation model that coordinates multiple medical institutions based on the federated learning framework include: Obtain the differences in the distribution of local model features among various medical institutions; Calculate the model aggregation weights based on the differences in the feature distributions; The model aggregates weights and fuses the local segmentation model parameters of the multiple medical institutions.

6. The automatic segmentation and recognition method for liver cancer images based on fused multimodal images according to claim 1, characterized in that, Also includes: Dynamic optimization parameters are generated based on the first tumor boundary segmentation result or the second tumor boundary segmentation result. The convolution kernel sizes of the first segmentation network model and the second segmentation network model are adjusted based on the dynamic optimization parameters. The adjusted convolution kernel size is used to process subsequent input multimodal medical image data.

7. The automatic segmentation and recognition method for liver cancer images based on fused multimodal images according to claim 6, characterized in that, The step of adjusting the convolution kernel size of the first segmentation network model and the second segmentation network model based on the dynamic optimization parameters includes: Calculate the feature receptive field adjustment based on the dynamic optimization parameters; The convolutional layer parameters of the first segmentation network model and the second segmentation network model are updated based on the feature receptive field adjustment.

8. The automatic segmentation and recognition method for liver cancer images based on fused multimodal images according to claim 1, characterized in that, Also includes: Obtain the pathological feature vector corresponding to the liver cancer subtype identification result; The pathological feature vectors are fed back to the feature decoders of the first segmentation network model and the second segmentation network model; The boundary smoothness parameters for subsequent tumor boundary segmentation are optimized based on the feedback results.

Citation Information

Patent Citations

  • CT metal artifact correction and super-resolution method for unsupervised deep dictionary learning

    CN117726706A