Medical image segmentation method and system based on parallel coding, variation fusion and uncertainty optimization of visual basic model

By combining the parallel coding architecture and the uncertainty optimization module, the problem of insufficient global structure modeling in coronary CT angiography image segmentation is solved, high-precision and high-robust coronary artery segmentation is achieved, and the segmentation effect is improved.

CN120707578APending Publication Date: 2025-09-26SECOND AFFILIATED HOSPITAL OF COLLEGE OF MEDICINEOF XIAN JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510818741.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies in coronary CT angiography image segmentation have problems such as insufficient global structure modeling, limited feature fusion depth, patch strategy that destroys continuity, and lack of uncertainty optimization, resulting in defects such as fragmentation, anatomical inconsistency, and omission of fine blood vessels in the segmentation results, making it difficult to meet the needs of high-precision and high-robustness segmentation.

Method used

It adopts a parallel coding architecture, combines the visual basic model ViT encoder and CNN encoder, and realizes adaptive fusion and refined processing of global and local features through a cross-branch variational fusion module and an uncertainty optimization module based on evidence learning, thereby improving segmentation performance.

Benefits of technology

The accuracy and robustness of coronary artery segmentation have been significantly improved, especially in the segmentation effects of small vessels, low contrast and complex morphological areas, and the model's overall perception and fine segmentation capabilities of complex vascular structures have been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707578A_ABST
    Figure CN120707578A_ABST
Patent Text Reader

Abstract

The invention relates to the field of medical image segmentation, and provides a parallel coding, variation fusion and uncertainty optimization segmentation method and system based on a visual basic model, and the method comprises the steps: 1, taking a coronary artery CT image as a research object, building a parallel coding architecture containing a visual basic model ViT encoder and a CNN encoder, and carrying out the parallel coding architecture; respectively extracting global and local features; 2, introducing a cross-branch variational fusion module, and generating a high-quality feature map by modeling potential distribution of two types of features and adopting a variational attention mechanism to realize adaptive feature fusion; and step 3, in the decoding process, each layer of output of the feature decoder is sent to an uncertainty optimization module based on evidence learning, and fine processing is carried out on an uncertain area. According to the method, the advantages of parallel coding, variation fusion and uncertainty optimization strategies based on the visual basic model are fully played, the challenges of complex structure, low contrast and the like in coronary artery segmentation are effectively solved, and the segmentation performance is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and in particular to a medical image segmentation method and system based on parallel encoding, variational fusion and uncertainty optimization of a visual basic model. Background Art

[0002] In the field of medical image analysis, especially in coronary computed tomography angiography (CCTA) image processing, accurate image segmentation techniques are crucial for clinical diagnosis, disease assessment, and treatment planning. CCTA has become the standard noninvasive method for assessing coronary artery anatomy and pathology (Marano, R., Rovere, G., Savino, G., Flammia, FC, Carafa, MRP, Steri, L., Merlino, B., Natale, L.: Ccta in the diagnosis of coronary artery disease. La radiologia medica 125, 1102–1113 (2020)). High-precision segmentation of coronary arteries in CCTA images is crucial for assessing stenosis severity, plaque morphology, and guiding clinical decision-making in CAD management. Despite advances in imaging technology, accurate coronary artery segmentation in CCTA images remains challenging due to several inherent factors that complicate the delineation of vascular structures.

[0003] Deep learning has shown significant potential in coronary artery segmentation, with good scalability and higher segmentation accuracy. U-Net and its variants remain the core architecture of many current advanced models (Song, A., Xu, L., Wang,L., Wang, B., Yang, X., Xu, B., Yang, B., Greenwald, SE: Automatic coronaryartery segmentation of ccta images with an efficient feature-fusion-and-rectification 3d-unet. IEEE Journal of Biomedical and Health Informatics 26(8), 4044–4055 (2022)), for example, 3D-FFR-UNet (Song, A., Xu, L., Wang, L., Wang, B.,Yang, X., Xu, B., Yang, B., Greenwald, SE: Automatic coronary artery segmentation of ccta images with an efficient feature-fusion-and-rectification 3d-unet. IEEE Journal of Biomedical and Health Informatics 26(8), 4044–4055 (2022) enhanced feature fusion capabilities by introducing dense convolution modules, while Dong et al. (Dong, C., Xu, S., Dai, D., Zhang, Y., Zhang, C., Li, Z.: A novel multi-attention, multi-scale 3D deep network for coronary artery segmentation. Medical Image Analysis 85, 102745 (2023)) adopted a multi-scale attention mechanism to capture more detailed vascular structural features. However, although convolutional neural network (CNN)-based methods perform well in extracting local features, they still have shortcomings in maintaining the continuity of vascular anatomical structures, often resulting in segmentation fragmentation and anatomical inconsistencies in complex vascular regions.

[0004] In contrast, the visual Transformer (ViT)-based method (Zhou, HY, Guo, J., Zhang, Y., Han, X., Yu, L., Wang, L., Yu, Y.: nnformer: Volumetric medical imagesegmentation via a 3d transformer. IEEE Transactions on Image Processing (2023)) has obvious advantages in modeling global features, but its performance in preserving fine-grained details is limited, making it difficult to accurately depict slender and curved vascular structures, which are crucial for clinical diagnosis.

[0005] To overcome these limitations, hybrid architectures combining the strengths of CNNs and ViTs have been proposed in recent years and have shown promising application prospects. For example, Pan et al. (Pan, C., Qi, B., Zhao, G., Liu, J., Fang, C., Zhang, D., Li, J.: Deep 3D vessel segmentation based on cross transformer network. In: 2022 IEEE international conference on bioinformatics and biomedicine (BIBM). pp. 1115–1120. IEEE (2022)) proposed a cross-Transformer network that integrates a UNet for local features and Transformers for long-range dependencies. Similarly, Ensembled-SAMs integrate nnU-Net (Isensee, F., Jaeger, PF, Kohl, SA, Petersen, J., Maier-Hein, KH: nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18(2), 203–211 (2021)) with SAMs (Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, AC, Lo, WY, et al.: Segment anything. In: Proceedings of the IEEE / CVF International Conference on Computer Vision. pp.4015–4026 (2023) achieved performance improvement, but its processing method relied on separate processing of 2D slices and merging of the results, failed to achieve feature-level fusion in 3D space, and ignored the continuity information between different slices.

[0006] CN119991723A discloses a medical image segmentation method combining a convolutional neural network with a graph network, comprising the following steps: collecting coronary artery CT angiography three-dimensional image data through a computed tomography (CT) device to form a data set CoronarySet, randomly dividing the data set CoronarySet into a training set CoronarySet1 and a test set CoronarySet2; segmenting the data of the training set CoronarySet1 into multiple patches; inputting the patches into a dual-branch parallel encoder to obtain a multi-layer coronary artery texture feature map and a multi-layer coronary artery topology feature map; constructing a feature fusion module, and integrating the features of the coronary artery into the coronary artery texture feature map based on the attention mechanism module. The texture feature maps and topological feature maps of different layers are fused to obtain fused vascular features; the shallow features of the 3D convolution module and the fused vascular features are input into each layer of the segmentation decoder in turn, and the result of the last layer of the output decoder is obtained to obtain the coronary artery segmentation result consistent with the input 3D image size, so as to train the 3D visual image network medical segmentation network; after the 3D visual image network medical segmentation network is trained, the image to be segmented in the test set CoronarySet2 is divided into patches, and each patch is input into the segmentation network to obtain the 3D vascular segmentation result of the patches. Finally, the 3D vascular segmentation results of multiple patches are spliced ​​to obtain a 3D vascular segmentation medical image of a single coronary artery CT angiography image.

[0007] The medical image segmentation method combining convolutional neural networks and graph networks proposed in the above method attempts to improve the segmentation effect in the segmentation task of coronary CT angiography (CCTA) images through a dual-branch encoder and feature fusion module. However, due to insufficient global structural modeling, limited feature fusion depth, patch strategy that destroys continuity, and lack of uncertainty optimization, it is difficult to effectively cope with challenges such as small blood vessels, low contrast, and complex morphology. As a result, the segmentation results have defects such as fragmentation, anatomical inconsistency, and omission of fine blood vessels, which makes it difficult to meet the clinical needs for high-precision and high-robustness segmentation. Summary of the Invention

[0008] In response to the problems in the prior art, the present invention provides a medical image segmentation method and system based on parallel coding, variational fusion and uncertainty optimization of a visual basic model.

[0009] The present invention is achieved through the following technical solutions: A medical image segmentation method and system based on parallel coding, variational fusion and uncertainty optimization of a visual basic model, comprising the following steps: Step 1: Using coronary artery CT images as the research object, a parallel encoding architecture is established, which includes a ViT encoder based on the visual foundation model and a CNN encoder. The ViT encoder from the visual foundation model obtains global features, and the CNN encoder extracts local features. Step 2: Design a cross-branch variational fusion module to model the potential distribution of global features and local features, and then introduce a variational attention mechanism to adaptively fuse global features and local features to obtain a feature map. ; Step 3: The feature map The input feature decoder is decoded. During the decoding process, the feature maps output by each layer of the decoder are input into the uncertainty optimization module based on evidence learning for refined processing.

[0010] Preferably, in step 2, the process of modeling the potential distribution of global structures and local features extracted from the visual base model ViT branch and the CNN branch is: The cross-branch variational fusion module is equipped with independent encoders for the ViT encoder from the visual base model to obtain global features and the CNN encoder to extract local features. The encoder uses a multi-layer perceptron to parameterize the potential distribution of global features and local features, and establishes a corresponding Gaussian distribution model. The reparameterization technique is used to train the model through the mean and standard deviation to obtain global latent variables and local latent variables.

[0011] Preferably, a variational attention mechanism is introduced to adaptively fuse global features and local features to obtain a fused feature map The process is: First, the global latent variables and local latent variables are input into the corresponding encoders based on multi-layer perceptrons. and encoder , generate the intermediate potential distribution; then pass Function to obtain the corresponding fusion weight and ; The weighted combination of the last two latent variables is the final fusion feature map .

[0012] Preferably, in step 3, the uncertainty optimization module obtains optimized features through evidence uncertainty estimation, multi-scale feature fusion and uncertainty-guided optimization strategy.

[0013] Preferably, the process of estimating the uncertainty of evidence is: First, Dirichlet distribution is used to model uncertainty, and the non-negative activation function Softplus From the feature map of the last layer of the decoder F Evidence graph , ,in, F Represents the feature map of the last layer output of the decoder; constructs the Dirichlet distribution parameters , calculate the Dirichlet intensity S , estimated uncertainty U , U The smaller the value, the more reliable the prediction for that area; U The larger the value, the higher the uncertainty.

[0014] Preferably, in step 3, the process of multi-scale feature fusion includes fusing the feature maps output by each layer of feature decoder: first, the low-resolution feature map is aligned with the high-resolution feature map by upsampling, and then these feature maps are gradually spliced ​​and fused by channel, and then the spatial attention block is introduced to obtain the fused feature. .

[0015] Preferably, in step 3, the uncertainty-guided optimization strategy is implemented by integrating the initial predictions P , uncertainty U and fusion features The optimization process is: Initial Forecast P : Feature map of the last layer of the decoder F ,go through Function, get the initial prediction P ;

[0016] uncertainty U : Modeling the uncertainty in segmentation results by adopting Dirichlet distribution; First, through the initial prediction P , uncertainty U and fusion features Constructing a reliability mask To suppress areas with higher abundance:

[0017] in, Exponential decay is applied to high uncertainty regions; Then, the attention mechanism is introduced to adaptively highlight important spatial regions by generating dynamic weights:

[0018] The final optimized feature is expressed as:

[0019] Among them, the weight Automatically balance the contribution between the initial prediction and the fused features.

[0020] A medical image segmentation method and system based on parallel coding, variational fusion and uncertainty optimization of a visual basic model, including a feature extraction module, a cross-branch variational fusion module, a decoder and an uncertainty optimization module based on evidence learning; The feature extraction module uses coronary artery CT images as the research object and establishes a parallel encoding architecture consisting of a ViT encoder based on the visual foundation model and a CNN encoder. The ViT encoder from the visual foundation model obtains global features, and the CNN encoder extracts local features. The cross-branch variational fusion module is used to model the potential distribution of global features and local features respectively, and then introduces the variational attention mechanism to adaptively fuse global features and local features to obtain feature maps. ; The decoder is used to decode the feature map During decoding, the feature maps output by each layer of the decoder are input into the uncertainty optimization module based on evidence learning to refine the uncertain areas and improve the segmentation performance.

[0021] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method when executing the computer program.

[0022] A storage medium stores a computer program, which implements the steps of the method when executed by a processor.

[0023] Compared with the prior art, the present invention has the following beneficial effects: This paper presents a medical image segmentation method based on parallel coding, variational fusion, and uncertainty optimization using a visual base model. This method uses a parallel coding architecture to synergistically fuse global and local feature representations. Specifically, the method first activates the last two modules in the encoder (ViT) derived from the visual base model, combined with an attention-guided enhancement (AGE) module to extract global image features. Simultaneously, a CNN encoder is introduced to efficiently extract local features, thereby forming a complementary representation of the input medical image.

[0024] To further effectively fuse global and local information, the present invention designs a cross-branch variational fusion module (CVF). This module models the potential distribution of global features and local features and introduces a variational attention mechanism to adaptively fuse the global features obtained from the ViT encoder of the visual base model and the local features extracted by the CNN encoder, thereby achieving more accurate context perception and feature complementarity.

[0025] To further enhance the robustness and reliability of segmentation results, this paper proposes an Evidential-learning Uncertainty Refinement (EUR) module. This module effectively guides decision-making during segmentation by incorporating evidence uncertainty estimates. Combining multi-scale feature aggregation with a spatial attention mechanism, it further enhances the model's spatial localization capabilities. Furthermore, an uncertainty-guided optimization strategy refines regions of high uncertainty, significantly improving the model's segmentation performance and robustness in complex scenarios such as low contrast and blurred boundaries.

[0026] The present invention discloses a medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of a visual base model. This method fully utilizes the advantages of the ViT encoder from the visual base model in global feature extraction and the ability of the convolutional neural network (CNN) encoder in local detail modeling. By constructing a parallel coding architecture, the complementary fusion of the two feature representations is achieved, thereby significantly improving the accuracy of coronary artery segmentation.

[0027] The cross-branch variational fusion module (CVF) introduced in this paper achieves an adaptive balance between the contributions of global and local features through potential distribution learning and variational attention mechanism, thereby enhancing the model's ability to model macrostructures and microstructures (i.e., large-scale vascular trunks and small branches), and improving its overall perception and fine segmentation of complex vascular structures.

[0028] The Evidential-learning Uncertainty Refinement Module (EUR) proposed in this paper improves the segmentation accuracy of the model in blurred and low-contrast areas through evidential uncertainty estimation, multi-scale feature fusion, and uncertainty-guided optimization strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a flowchart of a medical image segmentation method proposed by the present invention based on parallel encoding, variational fusion and uncertainty optimization of a visual basic model.

[0030] Figure 2 This is a processing diagram of a medical image segmentation method proposed in the present invention based on parallel encoding, variational fusion and uncertainty optimization of a visual basic model.

[0031] Figure 3The structural diagram of the cross-branch variational fusion module (CVF) proposed in this invention is shown.

[0032] Figure 4 The structural diagram of the Evidential-learning Uncertainty Refinement Module (EUR) proposed in this invention is shown.

[0033] Figure 5 This figure demonstrates the visual effects of a medical image segmentation method proposed in this paper, which combines parallel encoding, variational fusion, and uncertainty optimization based on a visual foundation model. In the figure, specific regions of interest are marked with cyan, yellow, and green dashed circles to facilitate visual comparison. These marked areas highlight the advantages of the new method in detail processing and accuracy. DETAILED DESCRIPTION

[0034] The present invention will be further described in detail below with reference to specific embodiments, which are intended to explain the present invention rather than to limit it.

[0035] The present invention discloses a medical image segmentation method and system based on parallel coding, variational fusion and uncertainty optimization of a visual basic model. Figure 1 、 2 , including the following steps: Step 1: Using coronary CT images as the research object, a parallel encoding architecture is established, consisting of a ViT encoder based on the visual base model and a CNN encoder. The ViT encoder from the visual base model is combined with an attention-guided enhancement module to capture global image features, while the CNN encoder focuses on extracting rich local features.

[0036] Step 2: Design a cross-branch variational fusion module (CVF) to model the potential distribution of global features and local features, and then introduce a variational attention mechanism to adaptively fuse global features and local features to obtain a feature map. .

[0037] Among them, the CVF module is designed to fuse the global and local features extracted from the ViT branch and the CNN branch based on the visual basic model. This module contains two core components: potential distribution learning and variational attention fusion. Its specific structure is as follows Figure 3 As shown in Figure 2, this module fuses global features with local features through latent distribution learning and variational attention fusion.

[0038] The process of modeling the latent distribution of global and local features using Latent Distribution Learning is: The cross-branch variational fusion module is equipped with independent encoders ( and ) is used to capture the inherent differences and complementarities between global features and local features. The encoder uses multi-layer perceptrons (MLPs) to and local features The latent distribution of is parameterized, and the corresponding Gaussian distribution model is established. The reparameterization trick is used to train the model through the mean and standard deviation to obtain global latent variables and local latent variables.

[0039] The mean and standard deviation are:

[0040] in, For MLPs from The latent variable mean predicted in is used to model the potential distribution center of the ViT branch; For the same MLPs from The predicted standard deviation indicates the distribution range of the ViT feature; For MLPs from The mean of the latent variables predicted in is used to model the potential distribution center of the CNN branch; For the same MLPs from The predicted standard deviation indicates the distribution range of CNN features.

[0041] and Sampling is performed according to normal distribution, that is, and .

[0042] The reparameterization technique ensures differentiability during training, allowing the CVF module to learn robust feature representations that take into account both deterministic and stochastic variations. As a result, the latent variables encapsulate richer contextual information, which is crucial for subsequent tasks such as segmentation.

[0043]

[0044]

[0045]

[0046] Where, and From the standard normal distribution A random variable sampled from .

[0047] This mechanism enables the CVF module to learn robust feature representations that account for both deterministic and stochastic variations. Consequently, the latent variables encapsulate richer contextual information, which is crucial for downstream tasks such as segmentation. By introducing this uncertainty-aware fusion strategy, the model demonstrates enhanced adaptability and stability in the face of complex vascular structures and imaging noise.

[0048] The variational attention fusion mechanism is introduced to fuse global features and local features. The process of obtaining the fused feature map is as follows: First, the global latent variables and local latent variables are input into the corresponding encoders based on multi-layer perceptrons. and encoder , generating an intermediate latent distribution:

[0049]

[0050] in, and They represent the intermediate potential distributions generated after being processed by their respective MLP encoders. Specifically, and Represented by global features Through the encoder After processing, the mean and variance of the Gaussian distribution are generated; similarly, It is composed of local features Through the encoder The processed output, and Represented by local features Through the encoder After processing, the mean and variance of the Gaussian distribution are generated; finally, and They are sampled from the corresponding Gaussian distribution.

[0051] Similar to and ,These encoders also ensure the consistency and robustness of the latent feature representation.

[0052] Afterwards, Function to obtain the corresponding fusion weight and :

[0053] The weighted combination of the last two latent variables is the final fusion feature representation :

[0054] in, It is a learnable weight matrix that is automatically learned through training and is used to achieve optimal feature transformation.

[0055] This mechanism can adaptively balance the contributions of global and local features, thereby enhancing the model's ability to model macrostructures and microstructures (i.e., large-scale vascular trunks and small branches), and improving its overall perception and fine segmentation of complex vascular structures.

[0056] Step 3: The feature map Input feature decoder. During the decoding process, the feature maps output by each layer of the decoder are input into the Evidential-learning Uncertainty Refinement Module (EUR) for refinement processing, which can improve the segmentation accuracy and obtain an optimized segmented image.

[0057] The uncertainty optimization module obtains the optimized feature map through evidence uncertainty estimation, multi-scale feature fusion and uncertainty-guided optimization strategy, which enhances the segmentation robustness of the model in fuzzy and low-contrast areas. Its specific structure is as follows: Figure 4 shown.

[0058] Deep learning models often exhibit overconfidence in medical image segmentation, which reduces the reliability of predictions. To address this issue, the EUR module adopts an evidence-based learning paradigm based on Subjective Logic (Jsang, A.: Subjective Logic: A formalism for reasoning under uncertainty. Springer Publishing Company, Incorporated (2018)), modeling uncertainty through "evidence" rather than direct probability.

[0059] The process of estimating the uncertainty of evidence is: First, Dirichlet Distribution (Xu, Y., Tang, J., Men,A., Chen, Q.: Eviprompt: A training-free evidential prompt generation methodfor adapting segment anything model in medical images. IEEE Transactions on Image Processing (2024)) is used to model voxel-level uncertainty. Specifically, through the non-negative activation function Softplus From the feature map of the last layer of the decoder F Evidence graph , ;in, F Represents the feature map output by the last layer of the decoder.

[0060] Then, construct the Dirichlet distribution parameters : ,in , K is the number of categories.

[0061] Then, the Dirichlet intensity is calculated S : ; Finally, the uncertainty estimate : The formula is: .

[0062] This method is able to effectively identify high-uncertainty regions, such as blood vessel boundaries or low-contrast areas, thereby guiding more accurate segmentation decisions.

[0063] To enhance the network's ability to perceive contextual information, the EUR module fuses multi-scale features from different decoder stages. The multi-scale feature fusion process involves fusing the feature maps output by each layer of feature decoders: First, the low-resolution feature map is aligned with the high-resolution feature map by upsampling, and then gradually concatenated and fused by channel:

[0064]

[0065] in: Indicates the i Characteristics of the layer; Indicates channel alignment convolution; Indicates the magnification ratio is The upsampling operation. represents the intermediate features after fusion; t represents the intermediate variable.

[0066] The fused features are spliced ​​by channel as follows: ,in Represents a splicing operation.

[0067] Subsequently, the Spatial Attention Block (SAB) was introduced (Liao, M., Zou, Z., Wan, Z., Yao, C., Bai, X.: Real-time scene text detection with differentiable binarization and adaptive scale fusion. IEEE transactions on pattern analysis and machine intelligence 45(1), 919–931 (2022)) to further improve the spatial positioning ability and obtain fusion features. : , thereby enhancing cross-scale information interaction and spatial sensitivity, ensuring more robust feature representation.

[0068] The EUR module integrates the initial forecast P , uncertainty U and fusion features The segmentation results are refined and optimized. The process is as follows: Initial Forecast P : Feature map of the last layer of the decoder F ,go through Function, get the initial prediction P ;

[0069] uncertainty U : Modeling the uncertainty in segmentation results by adopting Dirichlet distribution; First, through the initial prediction P , uncertainty U and fusion features Constructing a reliability mask to suppress areas of high uncertainty:

[0070] in, Exponential decay is applied to high uncertainty regions; Then, the attention mechanism is introduced to adaptively highlight important spatial regions by generating dynamic weights:

[0071] The final optimized feature is expressed as:

[0072] Among them, the weight Automatically balance the contribution between the initial prediction and the fused features to improve segmentation accuracy in uncertain or complex areas.

[0073] The present invention discloses a medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of a visual basic model. Taking coronary artery CT images as the research object, a parallel coding architecture including a ViT encoder based on a visual basic model and a CNN encoder is established to more comprehensively extract global structure and local features. A cross-branch variational fusion module (CVF) is used to model the potential distribution of global and local features, and a variational attention mechanism is introduced to adaptively fuse global and local features. Combined with an uncertainty optimization module (EUR) based on evidence learning, high-uncertainty areas are refined and optimized, enhancing the robustness and segmentation accuracy of the model in complex scenarios such as low contrast and fuzzy boundaries. The overall solution has stronger generalization ability and clinical application value.

[0074] The proposed method is evaluated on three datasets (CCTA119, MICCAI 2020 ASOCA Challenge dataset, and ICAS-100 dataset). The CCTA119 dataset is a self-built dataset containing 119 coronary CT angiography (CCTA) volume data from a tertiary-level medical institution with a resolution of , the matrix size is All cases were annotated by three radiologists with at least five years of experience.

[0075] MICCAI 2020 ASOCA Challenge dataset (Gharleghi, R., Adikari, D., Ellenberger, K., Webster, M., Ellis, C., Sowmya, A., Ooi, S., Beier, S.: Annotated computed tomography coronary angiogram images and associated dataof normal and diseased arteries. Scientific Data 10(1), 128 (2023)): Contains 40 CCTA scans.

[0076] The ICAS-100 dataset is a subset of the ImageCAS dataset (Zeng, A., Wu, C., Lin, G., Xie, W., Hong, J., Huang, M., Zhuang, J., Bi, S., Pan, D., Ullah, N., et al.: Imagecas: A large-scale dataset and benchmark for coronary artery segmentation based on computed tomography angiography images. ComputerizedMedical Imaging and Graphics 109, 102287 (2023)), which contains a total of 100 CCTA scan data.

[0077] Evaluation metrics: Dice Similarity Coefficient (DSC) and Average Symmetric Surface Distance (ASSD).

[0078] Implementation details Experiments were conducted using five-fold cross validation on the CCTA119, ASOCA, and ICAS-100 datasets. The specific divisions are as follows: CCTA119: 95 training cases, 24 testing cases; ASOCA: 32 cases for training and 8 cases for testing; ICAS-100: 80 cases for training and 20 cases for testing.

[0079] All experiments were performed on a NVIDIA 3090 GPU based on the PyTorch framework. The network was trained using the Adam optimizer with an initial learning rate of , a total of 600 epochs were trained with a batch size of 2. During the training phase, the full body data was randomly cropped to The sub-volume blocks are used for training; in the testing phase, a sliding window strategy is adopted to perform inference with the same sub-volume size, moving half the window length each time to cover the entire volume data to ensure the integrity and continuity of the segmentation results.

[0080] Part IV: Comparison with State-of-the-Art Methods A medical image segmentation method based on parallel encoding, variational fusion, and uncertainty optimization of a visual basis model is compared with nine state-of-the-art segmentation methods, including: Convolutional neural network (CNN)-based methods: 3D-Unet (Çiçek, Ö., Abdulkadir, A., Lienkamp, ​​SS, Brox, T., Ronneberger, O.: 3d u-net: learning densevolumetric segmentation from sparse annotation. In: International conferenceon medical image computing and computer-assisted intervention. pp. 424–432. Springer (2016)), S 2 CA-Net (Zhou, L., Jiang, Y., Li, W., Hu, J., Zheng, S.: Shape-scale co-awareness network for 3d brain tumor segmentation. IEEETransactions on Medical Imaging (2024)), I 2U-Net (Dai, D., Dong, C., Yan, Q., Sun, Y., Zhang, C., Li, Z., Xu, S.: I2u-net: A dual-path u-net with rich information interaction for medical image segmentation. Medical Image Analysis p. 103241 (2024)); Transformer-based methods: UNETR (Hatamizadeh, A., Tang, Y., Nath, V., Yang, D., Myronenko, A., Landman, B., Roth, H.R., Xu, D.: Unetr: Transformers for 3d medical image segmentation. In: Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. pp. 574–584 (2022)), TransUNet (Chen, J., Mei, J., Li, X., Lu, Y., Yu, Q., Wei, Q., Luo, X., Xie, Y., Adeli, E., Wang, Y., et al.: Transunet: Rethinking the u-net architecture design for medical image segmentation through the lens of transformers. Medical Image Analysis 97, 103280 (2024)), nnFormer (Zhou, H.Y., Guo, J., Zhang, Y., Han, X., Yu, L., Wang, L., Yu, Y.: nnformer: Volumetric medical image segmentation via a 3d transformer. IEEE Transactions on Image Processing (2023)); Methods specific to vascular segmentation: CS 2Net (Mou, L., Zhao, Y., Fu, H., Liu, Y., Cheng, J., Zheng, Y., Su, P., Yang, J., Chen, L., Frangi, AF, et al.: Cs 2 -net:Deep learning segmentation of curvilinear structures in medical imaging. Medical image analysis 67, 101874 (2021), 3D-FFR-Unet (Dong, C., Xu, S., Li,Z.: A novel end-to-end deep learning solution for coronary artery segmentation from ccta. Medical Physics (2022)), VSNet (Xu, J., Dong, A., Yang, Y., Jin, S., Zeng, J., Xu, Z., Jiang, W., Zhang, L., Dong, J., Wang, B.: Vsnet: Vessel structure-aware network for hepatic and portal veinsegmentation. Medical Image Analysis 101, 103458 (2025)).

[0081] All comparative experiments use publicly available code to ensure fairness.

[0082] As shown in Table 1, extensive comparative experiments are conducted on three datasets: CCTA119, ASOCA, and ICAS-100, and the results are shown in Table 1.

[0083] Table 1 Comparative experimental results of the medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization proposed in this invention and nine currently most advanced segmentation methods

[0084] On the CCTA119 dataset, our proposed method performs best among all compared models, improving DSC by 5.66% and reducing ASSD by 0.91mm compared to 3D-UNet. Compared to the top-performing CNN model, I2U-Net, our method also achieves a 2.75% improvement in DSC. Furthermore, our method outperforms the strongest Transformer-based and specialized methods for vessel segmentation, achieving a 2.70% improvement in DSC over nnFormer and a 2.24% improvement over VSNet, respectively.

[0085] On the ASOCA dataset, the proposed method also achieved the best performance, with DSC exceeding VSNet by 2.11% and ASSD reduced by 0.13 mm.

[0086] On the ICAS-100 dataset, although the dataset may have labeling errors resulting in relatively low overall indicators, the method proposed in this invention still maintains the best performance among all compared methods.

[0087] In summary, a medical image segmentation method based on parallel encoding, variational fusion and uncertainty optimization of a visual basic model has shown significant advantages both in terms of quantitative indicators and generalization ability on different datasets, especially in dealing with challenging medical image segmentation tasks such as small blood vessels, low contrast and complex structures.

[0088] Figure 5 Qualitative comparison results are shown. It can be seen that existing methods generally suffer from over-segmentation, under-segmentation, or both, resulting in less than ideal segmentation results. In the figure, specific regions of interest are marked with cyan, yellow, and green dotted circles to facilitate more intuitive visual comparison. These marked areas highlight the advantages of the new method in detail processing and accuracy. In contrast, the segmentation results of the method proposed in this invention are highly consistent with the true label (ground truth), especially in complex vascular structures. The performance is better, further verifying the effectiveness of the proposed method.

[0089] Table 2 Cross-validation experimental results of the medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization proposed in this paper

[0090] The model was trained on the CCTA119 dataset and tested on two independent datasets, ASOCA and ICAS-100, to assess its generalization and robustness. The proposed method performed well in cross-validation. As shown in Table 2, when trained on the CCTA119 dataset and tested on the ASOCA dataset, the proposed method achieved a DSC of 85.26%, significantly outperforming 3D-UNet (79.14%, a 6.12% improvement) and VSNet (82.67%, a 2.59% improvement). In terms of the ASSD metric, the proposed method achieved 0.83 mm, also outperforming 3D-UNet (1.64 mm, a 0.81 mm improvement) and VSNet (1.02 mm, a 0.19 mm improvement).

[0091] In addition, under the cross-dataset test setting of CCTA119→ICAS-100, the proposed method also maintains consistent excellent performance, further verifying its strong generalization ability.

[0092] Ablation experiments are used to evaluate the contributions of key components in the proposed method, including: Enhanced-ViT (a ViT encoder enhanced by activating the last two ViT modules and introducing an AGE module); CVF module (cross-branch variational fusion module); EUR module (Uncertainty Optimization module based on evidence learning).

[0093] The experiment starts with a basic encoder-decoder network Net1 (Zhang, Z., Liu, Q., Wang, Y.:Road extraction by deep residual u-net. IEEE Geoscience and Remote SensingLetters 15(5), 749–753 (2018)), and gradually adds the above components: Net2: Add Enhanced-ViT on the basis of Net1, and use the summation method to perform feature fusion; Net3: Replace the summation fusion in Net2 with the CVF module; Net4: Add the EUR module separately to Net1; The present invention: Integrate all the proposed components.

[0094] Table 3 Ablation experiment results of the medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of the visual basic model proposed in this paper on the CCTA119 dataset

[0095] As shown in Table 3, Net2 achieved a 1.25% improvement in DSC over Net1; Net3 achieved a further 1.21% improvement over Net2; and Net4 achieved a 1.19% improvement over Net1. The proposed method, which integrates all components, achieved a significant 4.39% DSC improvement over Net1. This result fully validates the effectiveness of each component and their synergistic effect, demonstrating that Enhanced-ViT, CVF, and EUR collectively significantly enhance the segmentation performance of the overall framework.

[0096] This paper proposes a medical image segmentation method based on parallel coding, variational fusion, and uncertainty optimization of the Vision Foundation Model. By designing an efficient parallel coding architecture, this method fully exploits the potential of the Vision Foundation Model (VFM) in medical image segmentation and achieves high-precision segmentation of coronary artery structures. The system mainly includes the following four key components: The encoder from the visual base model is used to extract global high-level semantic features of the image, enhancing the model's ability to model the overall structure of blood vessels; CNN encoder: responsible for capturing local low-level detail features, making up for the shortcomings of Transformer-type models in perceiving fine-grained information; CVF module (Cross-Branch Variational Fusion Module): Through latent distribution modeling and variational attention mechanism, it realizes the adaptive fusion of ViT and CNN branch features from the visual base model, improving context consistency and feature complementarity; The EUR module (Evidential-learning Uncertainty Refinement Module) effectively guides segmentation decisions by introducing evidence uncertainty estimation. Combining multi-scale feature fusion with a spatial attention mechanism further enhances the model's spatial localization capabilities. Furthermore, an uncertainty-guided optimization strategy refines regions of high uncertainty, significantly improving the model's segmentation performance and robustness in complex scenarios such as low contrast and blurred boundaries.

[0097] Extensive experiments on the self-built CCTA119 dataset and two publicly available medical image datasets demonstrate that this method outperforms current leading segmentation models in segmentation performance, demonstrating excellent effectiveness, robustness, and strong generalization. These results suggest that this method has significant application prospects and potential for supporting intelligent diagnosis and clinical decision support for coronary artery disease (CAD).

[0098] The present invention also discloses a medical image segmentation system based on parallel coding, variational fusion and uncertainty optimization of a visual base model, including a feature extraction module, a cross-branch variational fusion module, a decoder and an uncertainty optimization module based on evidence learning; the feature extraction module takes coronary artery CT images as the research object, and establishes a parallel coding architecture including a ViT encoder based on a visual base model and a CNN encoder. The ViT encoder from the visual base model is combined with an attention-guided enhancement module to obtain global features, and the CNN encoder extracts local features; the cross-branch variational fusion module is used to model the potential distribution of global and local features, and then introduces a variational attention mechanism to adaptively fuse the global and local features of the blood vessels to obtain a feature map ; The decoder is used to decode the feature map During decoding, the feature maps output by each layer of the decoder are input into the uncertainty optimization module based on evidence learning to refine the uncertain areas and improve the segmentation performance.

[0099] like Figure 2 As shown, the system consists of three main components: (1) Parallel encoding architecture: Combining an encoder based on a visual base model (e.g., based on a pre-trained model such as SAM-Med3D (Wang, H., Guo, S., Ye, J., Deng, Z., Cheng, J., Li, T., Chen, J.,Su, Y., Huang, Z., Shen, Y., et al.: Sam-med3d: towards general-purpose segmentation models for volumetric medical images. arXiv preprint (2023))) and an encoder based on a 3D UNet shape, the 3D voxel feature representations are extracted from both global and local perspectives. The two encoders work in parallel to extract complementary feature representations from the input 3D medical image: the former is good at capturing global contextual information, while the latter focuses on extracting local structural details, thereby achieving a comprehensive modeling of the complex morphology of the coronary arteries. To further enhance the model's ability to understand global features, the last two modules are activated in the encoder of the visual base model (ViT), and an attention-guided enhancement module (AGE) is introduced, such as Figure 2(b) This module effectively highlights the continuity and key morphological features of blood vessels and improves the overall structural perception ability by combining the attention mechanism (Dong, C., Xu, S., Dai, D., Zhang, Y., Zhang, C., Li, Z.: A novel multi-attention, multi-scale 3D deep network for coronary artery segmentation. Medical Image Analysis 85, 102745 (2023)) with a multi-layer fusion strategy.

[0100] (2) Cross-branch Variational Fusion Module (CVF): This module adaptively fuses the latent feature representations from the two encoder outputs by modeling the potential distribution of global and local features and applying a variational attention mechanism. By modeling the latent distribution and introducing a variational attention mechanism, this module achieves adaptive weighted fusion of features from different sources, enhancing the model's expressiveness in complex structural areas.

[0101] (3) Uncertainty Optimization Module (EUR) based on Evidence Learning: This module is used to optimize the uncertainty regions in the segmentation results. This module introduces evidence uncertainty estimation, combines multi-scale feature aggregation with spatial attention mechanism, and adopts uncertainty-guided optimization strategy to make fine corrections to the uncertainty regions, thereby improving the robustness of the segmentation results.

[0102] The present invention also discloses an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method when executing the computer program.

[0103] The present invention also discloses a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described are implemented.

[0104] Experiments on an internal dataset (CCTA119) and two publicly available CCTA datasets (ASOCA and ICAS-100) demonstrate that the proposed segmentation framework outperforms state-of-the-art segmentation models in terms of both the Dice Similarity Coefficient (DSC) and Average Symmetric Surface Distance (ASSD), demonstrating strong generalization performance. This technology is expected to be widely used in automated CAD diagnosis, preoperative planning, and postoperative follow-up, providing an efficient and accurate decision-making tool for clinicians.

[0105] The above description is merely a preferred embodiment of the present invention and is not intended to impose any limitation on the technical solution of the present invention. Those skilled in the art should understand that, without departing from the spirit and principles of the present invention, the technical solution can also be subjected to several simple modifications and replacements, and these modifications and replacements are also within the scope of protection covered by the claims.

Claims

1. A medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of a visual basic model, characterized by: The following steps are involved: Step 1: Using coronary artery CT images as the research object, a parallel encoding architecture is established, which includes a ViT encoder based on the visual foundation model and a CNN encoder. The ViT encoder from the visual foundation model obtains global features, and the CNN encoder extracts local features. Step 2: Design a cross-branch variational fusion module to model the potential distribution of global features and local features, and then introduce a variational attention mechanism to adaptively fuse global features and local features to obtain a feature map. ; Step 3: The feature map The input feature decoder is decoded. During the decoding process, the feature maps output by each layer of the decoder are input into the uncertainty optimization module based on evidence learning for refined processing.

2. The medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of a visual base model according to claim 1 is characterized in that: In step 2, the process of modeling the potential distribution of global features and local features is: The cross-branch variational fusion module is equipped with independent encoders for the ViT encoder from the visual base model to obtain global features and the CNN encoder to extract local features. The encoder uses a multi-layer perceptron to parameterize the potential distribution of global features and local features, and establishes a corresponding Gaussian distribution model. The reparameterization technique is used to train the model through the mean and standard deviation to obtain global latent variables and local latent variables.

3. The medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of a visual base model according to claim 1, characterized in that: Introducing the variational attention mechanism to adaptively fuse global features and local features to obtain a fused feature map The process is: First, the global latent variables and local latent variables are input into the corresponding encoders based on multi-layer perceptrons. and encoder , generate the intermediate potential distribution; then pass Function to obtain the corresponding fusion weight and ; The weighted combination of the last two latent variables is the final fusion feature map .

4. The medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of a visual base model according to claim 1, characterized in that: In step 3, the uncertainty optimization module obtains the optimized features through evidence uncertainty estimation, multi-scale feature fusion, and uncertainty-guided optimization strategy.

5. The medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of a visual base model according to claim 4 is characterized in that: In step 3, the process of estimating the uncertainty of evidence is: First, Dirichlet distribution is used to model uncertainty, and the non-negative activation function From the feature map of the last layer of the decoder F Evidence graph , ,in, F Represents the feature map of the last layer output of the decoder; constructs the Dirichlet distribution parameters α , calculate the Dirichlet intensity S , estimated uncertainty U , U The smaller the value, the more reliable the prediction for that area; U The larger the value, the higher the uncertainty.

6. The medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of a visual base model according to claim 5, characterized in that: In step 3, the multi-scale feature fusion process includes fusing the feature maps output by each layer of feature decoder: first, the low-resolution feature map is aligned with the high-resolution feature map by upsampling, and then these feature maps are gradually spliced ​​and fused by channel, and then the spatial attention block is introduced to obtain the fused feature. .

7. The medical image segmentation method based on parallel coding, variational fusion and uncertainty optimization of a visual base model according to claim 6, characterized in that: In step 3, the uncertainty-guided optimization strategy integrates the initial predictions P , uncertainty U and fusion features The optimization process is: Initial Forecast P : Feature map of the last layer of the decoder F ,go through Function, get the initial prediction P ; uncertainty U : Modeling the uncertainty in segmentation results by adopting Dirichlet distribution; First, through the initial prediction P , uncertainty U and fusion features Constructing a reliability mask To suppress areas with higher abundance: in, Exponential decay is applied to high uncertainty regions; Then, the attention mechanism is introduced to adaptively highlight important spatial regions by generating dynamic weights: The final optimized feature is expressed as: ; Among them, the weight Automatically balance the contribution between the initial prediction and the fused features.

8. A medical image segmentation system based on parallel coding, variational fusion and uncertainty optimization of a visual basic model, characterized by: It includes feature extraction module, cross-branch variational fusion module, decoder and uncertainty optimization module based on evidence learning; The feature extraction module uses coronary artery CT images as the research object and establishes a parallel encoding architecture based on the visual foundation model ViT encoder and CNN encoder. The ViT encoder from the visual foundation model obtains global features, and the CNN encoder extracts local features. The cross-branch variational fusion module is used to model the potential distribution of global features and local features respectively, and then introduces the variational attention mechanism to adaptively fuse global features and local features to obtain feature maps. ; The decoder is used to decode the feature map During the decoding process, the feature maps output by each layer of the decoder are input into the uncertainty optimization module based on evidence learning to perform fine processing on the uncertain areas to improve the segmentation performance.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Convolutional neural network and graph network combined medical image segmentation method

    CN119991723A