A coronary angiography image segmentation method and system based on semantic guidance
Patent Information
- Application Number
- CN202610911496.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-06-24
AI Technical Summary
[0004]本发明针对现有冠脉造影图像分割方法中多尺度特征融合过程存在语义不一致、血管结构易断裂及细小分支易丢失的问题,提供一种基于语义引导的冠脉造影图像分割方法及系统,旨在通过构建统一语义基底并进行多尺度语义重建,以引入全局语义信息对特征恢复过程进行引导
(1)本发明将预训练的医学视觉基础模型的多尺度原始全局语义特征通过语义对齐、语义基底构建、语义门控机制和多尺度语义重建,捕获血管分割导向的多尺度语义重建特征,并将语义特征适配到解码器中。该方法使不同尺度特征在生成过程中保持语义一致性,解码过程持续受到统一语义基底的约束,分割结果的全局连贯性得到显著增强,血管主干连续完整。
Smart Images

Figure CN122454196B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing technology, specifically relating to a semantically guided coronary angiography image segmentation method and system, which is particularly suitable for automatic segmentation of blood vessels in X-ray coronary angiography images. Background Technology
[0002] Cardiovascular disease is one of the major global health burdens, and coronary heart disease (CHD), as the most common cardiovascular disease, is crucial for reducing mortality risk through early and accurate diagnosis. X-ray coronary angiography (XCA) is widely considered the gold standard for diagnosing CHD. By segmenting blood vessels in XCA images, vascular structural information can be automatically extracted, helping to detect stenosis or occlusion, assess the extent of lesions, and assist in clinical diagnosis.
[0003] However, XCA images often suffer from uneven illumination, low contrast, blurred boundaries between blood vessels and the background, and strong interference from catheters, making it difficult for traditional image segmentation methods (such as thresholding, region growing, and edge detection) to achieve ideal results. In recent years, deep learning-based methods, especially U-Net and its variants, have made significant progress in medical image segmentation. However, existing methods still face the following challenges when processing slender, overlapping, and low-contrast coronary artery structures: (1) Repeated downsampling and upsampling processes weaken global blood vessel connectivity, resulting in structural breaks and topological inconsistencies in the segmentation results; (2) In the encoder-decoder structure, local details and global semantics are difficult to be effectively integrated, there is semantic inconsistency between multi-scale features, and small blood vessel branches are easily lost in low-contrast areas; (3) Existing methods mostly introduce semantic information through multi-scale feature fusion or attention mechanism. In essence, they are still feature-level weighted fusion, lacking unified expression and structured modeling of semantic information, and making it difficult to constrain the semantic consistency between features of different scales as a whole. (4) Some methods introduce semantic knowledge or cross-modal alignment mechanisms to use semantic information for feature enhancement or result interpretation. However, these methods mainly rely on semantic information to weight or fuse existing features. They do not involve generating features from semantic information or participating in the structural control of the segmentation process, making it difficult to provide continuous and effective guidance for the feature recovery process. (5) Existing semantic guidance methods usually only introduce semantic information in the encoding stage or feature fusion stage, and lack a mechanism to continuously guide feature recovery during the decoding process, which causes the semantic information to gradually decay during feature reconstruction, thereby affecting the integrity and connectivity of the small blood vessel structure; Therefore, how to construct a unified semantic expression and continuously exert semantic constraints during the feature recovery process to effectively guide the global consistency modeling and decoding process of multi-scale features, and improve the structural integrity and accuracy of coronary artery segmentation without significantly increasing model complexity, has become an urgent technical problem to be solved. Summary of the Invention
[0004] This invention addresses the problems of semantic inconsistency, susceptibility to vascular structure fragmentation, and loss of fine branches in the multi-scale feature fusion process of existing coronary angiography image segmentation methods. It provides a semantically guided coronary angiography image segmentation method and system, aiming to guide the feature recovery process by constructing a unified semantic basis and performing multi-scale semantic reconstruction. This global semantic information is not only used for feature representation enhancement but also participates in the feature reconstruction process as a structural constraint, thereby continuously guiding feature recovery during the decoding stage, improving the connectivity and integrity of vascular structures, and enhancing adaptability to low-contrast and complex background interference scenarios.
[0005] To address the above problems, the present invention provides the following technical solution: A semantically guided coronary angiography image segmentation method includes the following steps: Step S1: Obtain the coronary angiography image to be segmented and perform preprocessing; Step S2: Perform dual-stream parallel feature extraction on the preprocessed coronary angiography images: Step S21: Perform hierarchical downsampling on the preprocessed coronary angiography image using the trunk encoder to extract multi-scale local structural features; Step S22: Extract multi-scale original global semantic features from the pre-processed coronary angiography image using the image encoder of the pre-trained medical vision basic model; Step S3: Perform semantic reconstruction on the original global semantic features at multiple scales: Step S31: Perform semantic alignment on the original multi-scale global semantic features to obtain aligned multi-scale semantic features; Step S32: Concatenate the aligned multi-scale semantic features along the channel dimension to obtain concatenated features; in the unified semantic space, perform gated fusion on the concatenated features to generate a spatial adaptive weight map; use the spatial adaptive weight map to adaptively fuse the aligned multi-scale semantic features to construct a unified semantic base with global consistency. Step S33: Based on the unified semantic base, multi-scale semantic reconstruction features that match the spatial resolution and number of channels of each layer of the decoder are generated in reverse through different scale mapping functions. The multi-scale semantic reconstruction features include features of at least two different spatial scales. Step S4: At each layer of the decoder, the multi-scale semantic reconstruction features of the corresponding scale and the multi-scale local structural features of the backbone encoder are injected into the decoder layer by layer to guide the decoder to recover the feature map layer by layer. Step S5: Activate and binarize the feature map output by the decoder to output the blood vessel segmentation result.
[0006] Specifically, the multi-scale original global semantic features include features at three different scales: the first original global semantic feature F 32 Second original global semantic feature F 64 and the third original global semantic feature F 128 The semantic alignment specifically involves: aligning the third original global semantic feature F... 128 Second original global semantic feature F 64 Upsampled to the first original global semantic feature F using bilinear interpolation. 32 Using the same spatial resolution, the number of channels is then uniformly mapped to the same value through three independent 1×1 convolutions to obtain the first alignment feature after alignment. Second alignment feature and third alignment feature .
[0007] Specifically, step S32 includes: The second alignment feature Third alignment feature and the first alignment feature By concatenating along the channel dimension, the concatenated features are obtained. ; A 1×1 convolution is applied to the splicing feature. Mapping is performed to generate a single-channel feature map, which is then activated by the Sigmoid function to obtain a spatially adaptive weight map G. This spatially adaptive weight map G is then used to adaptively fuse high-level semantic features and local detail features to obtain a unified semantic basis. : ; Where ⊙ denotes element-wise multiplication. The unified semantic basis is not only used for semantic expression alignment, but also serves as a generation source, generating multi-scale semantic reconstruction features in reverse through a cross-scale mapping function, thereby achieving structured control over the decoding process.
[0008] Specifically, the multi-scale semantic reconstruction features generated in reverse in step S33 include four features with different spatial resolutions and channel numbers: the first semantic reconstruction feature Second semantic reconstruction features Third semantic reconstruction features and fourth semantic reconstruction features ; Represents the feature dimension.
[0009] Specifically, in step S4, the decoder performs layer-by-layer calculations according to resolution: ; Where R represents the spatial resolution of the current feature map, This represents the output characteristics of the decoder at a resolution of R. This indicates the decoding features of the layer above the decoder (lower resolution). This represents the local structural features derived from the backbone encoder at a resolution of R, used to provide local structural information. This represents the semantic reconstruction features at resolution R, primarily providing global semantic constraints. This indicates an upsampling operation, used to map low-resolution features to the current resolution R.
[0010] Specifically, the upsampling operation UP(·) is implemented using bilinear interpolation or transposed convolution.
[0011] Specifically, the pre-trained medical vision basic model is the MedSAM2 model, and the backbone encoder and the decoder adopt a U-Net structure or a variant thereof.
[0012] The present invention provides a semantically guided coronary angiography image segmentation system, comprising: Image preprocessing module: used to acquire the coronary angiography image to be segmented and perform preprocessing; The dual-stream feature extraction module includes a backbone encoder unit and an image encoder unit of a pre-trained medical vision basic model. The backbone encoder unit is used to perform hierarchical downsampling on the pre-processed coronary angiography image to extract multi-scale local structural features. The image encoder unit of the pre-trained medical vision basic model is used to extract multi-scale original global semantic features of the pre-processed coronary angiography image. The semantic reconstruction module includes a semantic alignment unit, a semantic basis construction unit, and a multi-scale semantic reconstruction unit. The semantic alignment unit performs semantic alignment on the original multi-scale global semantic features to obtain aligned multi-scale semantic features. The semantic basis construction unit concatenates the aligned multi-scale semantic features along the channel dimension to obtain concatenated features. It then performs gated fusion on these concatenated features in a unified semantic space to generate a spatially adaptive weight map. This spatially adaptive weight map is used to adaptively fuse the aligned multi-scale semantic features to construct a unified semantic basis. The multi-scale semantic reconstruction unit, based on the unified semantic basis, generates multi-scale semantic reconstruction features that match the spatial resolution and number of channels of each layer of the decoder through different scale mapping functions. These multi-scale semantic reconstruction features include features at least two different spatial scales. Decoding module: includes a decoder unit and a semantic injection unit; the decoder unit is used to recover the feature map layer by layer; the semantic injection unit is used to inject the multi-scale semantic reconstruction features of the corresponding scale and the multi-scale local structural features extracted by the backbone encoder unit into the decoder unit at each layer of the decoder, so as to guide the decoder unit to recover the feature map layer by layer; Output module: used to activate and binarize the feature map output by the decoder unit to obtain the blood vessel segmentation result.
[0013] The present invention also provides an electronic device, comprising: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor. When the computer program is executed by the at least one processor, it causes the at least one processor to execute the semantically guided coronary angiography image segmentation method.
[0014] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the semantically guided coronary angiography image segmentation method described above.
[0015] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention captures multi-scale semantic reconstruction features guided by blood vessel segmentation by using multi-scale original global semantic features of a pre-trained medical vision basic model through semantic alignment, semantic basis construction, semantic gating mechanism and multi-scale semantic reconstruction, and adapts the semantic features into the decoder. This method ensures that features of different scales maintain semantic consistency during the generation process, and the decoding process is continuously constrained by a unified semantic basis. The global coherence of the segmentation results is significantly enhanced, and the main blood vessel trunk is continuous and complete.
[0016] (2) By constructing a unified semantic base, the present invention enables different spatial locations to share consistent global semantic information. During the decoding process, global constraints are applied to the local structure recovery, which effectively reduces the structural breakage caused by repeated downsampling and upsampling, and the blood vessel topology in the segmentation result remains intact.
[0017] (3) In this invention, the semantic features at each scale are generated in reverse from a unified semantic base, which avoids the semantic shift problem caused by scale differences in traditional methods, so that small blood vessel branches can still be accurately identified and preserved in low contrast areas, and the blood vessel tree structure in the segmentation result is more complete.
[0018] (4) This invention learns an adaptive weight map of the learning space to dynamically balance the contributions of high-level semantic features and local detail features in different vascular regions (main trunk and small branches). The gating operation is applied to a unified semantic space, which can suppress semantic conflict regions, strengthen structurally consistent regions, filter background noise and interference from non-target structures such as catheters, resulting in a cleaner background in the segmentation results and significantly improved anti-interference ability.
[0019] (5) This invention derives multi-scale semantic reconstruction features from a unified semantic base that precisely match the spatial resolution and number of channels of each layer of the decoder, and injects them into the decoder layer by layer, avoiding the problem of increased parameter quantity and feature redundancy caused by traditional splicing, and achieving non-destructive semantic guidance.
[0020] (6) This invention can effectively alleviate the structural discontinuity problem in blood vessel segmentation by semantic guidance alone, without the need to introduce a complex local structure modeling module. It improves segmentation performance while maintaining low computational complexity, and is especially suitable for coronary angiography images with low contrast and strong background interference. It has good clinical application prospects. Attached Figure Description
[0021] Figure 1 This is an overall flowchart of the method of the present invention; Figure 2 This is a flowchart for semantic reconstruction of multi-scale original global semantic features; Figure 3 This is a visual comparison of the segmentation results of the method of the present invention and the comparative method. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.
[0023] Example 1
[0024] This embodiment proposes a semantically guided coronary angiography image segmentation method. Its core lies in constructing a basic model-guided semantic reconstruction framework (FSR-Framework) to address the semantic inconsistencies, structural breaks, and loss of small vessel branches encountered in existing methods during multi-scale feature fusion. Unlike existing methods that directly stitch together or weightedly fuse multi-scale features, this embodiment employs a mechanism of first unifying semantic representation before multi-scale reconstruction. Specifically, semantic information from different scales is mapped to a unified semantic space to construct a semantic base. This semantic base is then used to generate vessel segmentation-oriented multi-scale semantic features to guide the decoding process. This mechanism ensures semantic consistency across different scales during feature generation, thereby improving the connectivity and integrity of the vascular structure.
[0025] like Figure 1 As shown in this embodiment, a semantically guided coronary angiography image segmentation method includes the following steps: Step S1: Obtain the coronary angiography image to be segmented and perform preprocessing; Step S2: Perform dual-stream parallel feature extraction on the preprocessed coronary angiography images: Step S21: Perform hierarchical downsampling on the preprocessed coronary angiography image using the trunk encoder to extract multi-scale local structural features; Step S22: Extract multi-scale original global semantic features from the pre-processed coronary angiography image using the image encoder of the pre-trained medical vision basic model; Step S3: Perform semantic reconstruction on the original global semantic features at multiple scales: Step S31: Perform semantic alignment on the original multi-scale global semantic features to obtain aligned multi-scale semantic features; Step S32: Concatenate the aligned multi-scale semantic features along the channel dimension to obtain concatenated features; in the unified semantic space, perform gated fusion on the concatenated features to generate a spatial adaptive weight map; use the spatial adaptive weight map to adaptively fuse the aligned multi-scale semantic features to construct a unified semantic base with global consistency. Step S33: Based on the unified semantic base, multi-scale semantic reconstruction features that match the spatial resolution and number of channels of each layer of the decoder are generated in reverse through different scale mapping functions. The multi-scale semantic reconstruction features include features of at least two different spatial scales. Step S4: At each layer of the decoder, the multi-scale semantic reconstruction features of the corresponding scale and the multi-scale local structural features of the backbone encoder are injected into the decoder layer by layer to guide the decoder to recover the feature map layer by layer. Step S5: Activate and binarize the feature map output by the decoder to output the blood vessel segmentation result.
[0026] The following provides a more detailed explanation of each step.
[0027] In step S1, the coronary angiography images are preprocessed, including converting them to grayscale images, standardizing intensity using two-dimensional minimax normalization, and then applying constrained contrast adaptive histogram equalization (CLAHE). It should be noted that during the training phase, to expand the sample data of coronary angiography images, image enhancement is required before preprocessing. Image enhancement includes: random rotation (…). The geometric transformations include: 10° rotation, horizontal flipping, random scaling (0.8~1.25), random cropping (aspect ratio 0.7~1.3), and random brightness adjustment (0.7~1.3). All geometric transformations are applied synchronously to the corresponding real-world annotations to ensure consistency between the image and the label.
[0028] Specifically, in step S2, the preprocessed coronary angiography image , Here, H represents the feature dimension, and W represents the height. In this embodiment, H=256 and W=256. The backbone encoder consists of four downsampling stages, each halving the spatial resolution of the feature map and doubling the number of channels. The output feature resolution of each level of the backbone encoder is as follows: Trunk encoder layer 1:
[0029] Main encoder layer 2:
[0030] Main encoder layer 3:
[0031] Main encoder layer 4 (bottleneck layer):
[0032] The backbone encoder ultimately obtains multi-scale local structural features {E1, E2, E3, E4}, where the first local structural feature is... Resolution 64×64, number of channels 96; second local structural features Resolution 32×32, number of channels 192; third local structural features Resolution 16×16, number of channels 384; fourth local structural feature , 8×8 resolution, 768 channels.
[0033] In this embodiment, the pre-trained MedSAM2 model is selected as the basic model for medical vision. Its image encoder receives the pre-processed coronary angiography image as input and outputs three multi-scale original global semantic features: First original global semantic features (Top-level global features, spatial resolution 32×32); Second original global semantic features (High-resolution local features, spatial resolution 64×64); Third original global semantic features (Higher resolution local features, spatial resolution 128×128).
[0034] Reference Figure 2 The specific process of semantic reconstruction of the multi-scale original global semantic features in step S3 of this embodiment is as follows: Step S31: The multi-scale original global semantic features { , , The data is uniformly aligned to a spatial resolution of 32×32, and channel projection is performed using 1×1 convolution to achieve semantic consistency. ; ; ; in, This indicates that the bilinear interpolation is upsampled to 32×32; This represents a 1×1 convolutional channel projection, mapping the three features to a unified semantic space (with a unified number of channels of 256), resulting in aligned multi-scale semantic features. }, For the first alignment feature, For the second alignment feature , This is the third alignment feature.
[0035] Step S32: Concatenate the aligned multi-scale semantic features along the channel dimension to obtain the concatenated features. , Indicates a splicing operation; In a unified semantic space, a semantic gating mechanism is introduced to perform gated fusion of the concatenated features, and a spatial adaptive weight map G is generated through convolutional mapping and Sigmoid. ; in, This is a 1×1 convolution operation with 768 input channels, 1 output channel, 1 stride, 0 padding, and no bias. The sigmoid activation function compresses the output to (0,1); For spatially adaptive weighted graphs, when G(i,j) is close to 1, this position mainly draws from high-level semantic features. When G(i,j) is close to 0, the information is mainly derived from local detail features at that location. and Information acquisition. Gating operations operate on a unified semantic space, rather than the original multi-scale features. Their function is to suppress semantically conflicting regions, strengthen structurally consistent regions, and filter out background interference.
[0036] A unified semantic basis is constructed using the multi-scale semantic features after adaptive fusion and alignment of the spatial adaptive weight graph G. ; Among them, the unified semantic base , where ⊙ represents element-wise multiplication. It acts on areas with large blood vessels. and It acts on the branching areas of small blood vessels.
[0037] Step S33: Based on the unified semantic base Multi-scale semantic reconstruction features that match the spatial resolution and number of channels of each layer of the decoder are generated in reverse by different scale mapping functions. The multi-scale semantic reconstruction features include features of at least two different spatial scales. ; in, This represents the semantic reconstruction features of the k-th layer. This represents the mapping function of the k-th layer. Then, derived from the same semantic base to ensure cross-scale semantic consistency, it outputs four multi-scale semantic reconstruction features guided by vessel segmentation: the first semantic reconstruction feature... Second semantic reconstruction features Third semantic reconstruction features and fourth semantic reconstruction features .
[0038] Specifically, in step S4, the decoder performs semantic-guided decoding, calculating layer by layer from the deepest layer to the shallowest layer according to resolution: ; A spatial resolution-based representation is used, where R represents the spatial resolution of the current feature map. This represents the output characteristics of the decoder at a resolution of R. This indicates the decoding features of the layer above the decoder (lower resolution). This represents the local structural features derived from the backbone encoder at a resolution of R, used to provide local structural information. This represents the semantic reconstruction features at resolution R, primarily providing global semantic constraints. This indicates an upsampling operation, used to map low-resolution features to the current resolution R.
[0039] The resolution of each layer of the decoder is symmetrical to that of the backbone encoder. The resolution of the decoder is: Decoder layer 1 (deepest layer):
[0040] Decoder Layer 2:
[0041] Decoder Layer 3:
[0042] Decoder Layer 4 (Output Layer): .
[0043] To verify the effectiveness of this invention, this embodiment uses a publicly available coronary angiography dataset, sourced from the "Automatic Regional Coronary Artery Disease Diagnosis Based on X-ray Angiography Images (ARCADE)" challenge hosted by the 26th International Conference on Medical Image Computing and Computer-Aided Intervention (MICCAI). Each patient's image consists of 0 to 12 frames. The frames with the best contrast were selected and detailed annotations were performed using the Syntax Score method, covering 26 different regions, totaling 1500 images. All images are accompanied by pixel-level vessel annotations. The dataset is divided into a training set (80%, 1200 images) and a test set (20%, 300 images), with 200 images from the training set used as the validation set. All images are accompanied by pixel-level vessel annotations.
[0044] This experiment was conducted on a workstation equipped with two Intel Xeon Silver 4310 2.10 GHz CPUs, one NVIDIA GeForce RTX 4090 24G GPU, and 128GB of RAM. The deep learning framework used was PyTorch 2.2.1 (CUDA 11.8). The batch size was set to 16, and training lasted for 150 epochs. The optimizer used was AdamW (with an initial learning rate of...). (Weight decay is 0.01), and the learning rate scheduler used is CosineAnnealingLR.
[0045] In coronary artery segmentation tasks, due to the small size of the vessel region and the large size of the background region, traditional loss functions (such as binary cross-entropy loss) often result in high accuracy for background prediction but poor foreground prediction. To address the imbalance between foreground and background pixels, Dice loss is introduced. This loss focuses on the overlap between the predicted result and the ground truth label, thus providing a more accurate evaluation of model performance. Therefore, this embodiment combines binary cross-entropy loss and Dice loss as the model's loss function, with the total loss Loss = Loss BCE +LossDice Loss BCE Represents the binary cross-entropy loss, Loss Dice This indicates Dice's loss.
[0046] To evaluate the segmentation performance of the model, the following five metrics were used: Intersection over Union (IoU), F1 score (F1), accuracy (Acc), specificity (Spe), and sensitivity (Sen).
[0047] To verify the role of the semantic basis reconstruction mechanism proposed in this invention in multi-scale semantic modeling, the following comparative experiments were designed. Each comparative method corresponds to a different feature utilization method in the prior art, in order to analyze the improvement effect of this invention relative to different technical paths.
[0048] Base Model (UNet): The standard encoder-decoder architecture (UNet) is used as a baseline model to evaluate segmentation performance without introducing any additional semantic modeling mechanisms.
[0049] UNet + Concat (Model 1) directly concatenates multi-scale semantic features to the decoding stage to simulate common multi-scale feature fusion strategies in existing technologies. This method is used to verify whether simple fusion can effectively improve structural consistency.
[0050] UNet + Attention (Model 2) introduces an attention mechanism to weight features during the decoding stage. This method is used to verify whether the role of traditional attention mechanisms in semantic modeling is sufficient to replace the method of this invention.
[0051] The method of this invention achieves semantic consistency modeling through semantic alignment, semantic basis construction, and multi-scale semantic reconstruction. This method is used to verify whether multi-scale reconstruction using a unified semantic basis can more effectively improve segmentation performance compared to direct fusion or attention weighting. Table 1 presents the quantitative results of different methods on the test set.
[0052] Table 1. Comparison of Experimental Results for Different Feature Usage Methods
[0053] As can be seen from the results in Table 1, the method of this invention outperforms the comparative method in all indicators, especially showing significant improvements in F1, Sen, and IoU. Further analysis reveals that: Compared to the UNet model, the performance improvement indicates that by introducing a unified semantic basis, global structural information can continue to play a constraining role during the decoding process, thereby reducing the phenomenon of broken blood vessels.
[0054] Compared to the UNet+Concat model 1, the method of this invention has better performance, indicating that simple concatenation cannot eliminate semantic differences between multi-scale features, while this invention achieves unified expression through semantic basis, effectively improving semantic consistency.
[0055] Compared to the UNet+Attention model2, the method of this invention still performs better, indicating that the attention mechanism only adjusts local weights, while the method of this invention reconstructs in a unified semantic space, which is more conducive to structure preservation in terms of mechanism.
[0056] From the visualization results Figure 3 It can be observed that the UNet model has obvious breaks in the low-contrast region; although the UNet+Concat model 1 enhances the response in some regions, it has the problem of noise diffusion; the UNet+Attention model 2 improves at the boundary, but the overall structural consistency is still insufficient; in contrast, the method of the present invention shows a more continuous vascular trunk structure, more complete small branches and a cleaner background region. The reason is that the method of the present invention uniformly constrains the features at each scale through semantic basis, so that the segmentation results have higher spatial consistency.
[0057] Therefore, this invention constructs a unified semantic base, enabling different spatial locations to share consistent global semantic information. This imposes global constraints on local structure recovery during decoding, effectively reducing vascular rupture. Through multi-scale semantic reconstruction, features at each scale are generated from the same semantic source, avoiding the semantic shift problem caused by scale differences in traditional methods, thereby improving the recognition ability of small vascular branches. The semantic gating mechanism filters features in a unified semantic space, suppressing background noise and interference from non-target structures such as catheters, improving segmentation stability.
[0058] To further verify the performance advantages and technological improvements of the method of this invention compared to existing mainstream segmentation methods, representative models from different technical approaches were selected for comparative analysis. The selected methods cover typical approaches in the current field of medical image segmentation, including methods based on convolutional structure optimization and global modeling methods based on Transformer. Specifically, the following models were selected as comparison objects: U-Net++, a model based on convolutional structure improvement, and Swin-UNet, a model based on Transformer structure. These methods represent the current mainstream feature fusion optimization path and global modeling enhancement path, respectively. Specific experimental results are shown in the table below: Table 2. Experimental results compared with mainstream methods
[0059] The results show that the method of this invention outperforms the comparative methods in all evaluation metrics, especially in terms of F1 score and IoU, where it demonstrates greater stability. A comparison with the aforementioned methods verifies the advantages of the method of this invention in unified semantic modeling and multi-scale semantic consistency reconstruction.
[0060] U-Net++ improves segmentation performance by enhancing information transfer between features at different levels through improved skip connection structures. However, feature fusion in this type of method still relies on local convolution operations and lacks explicit modeling of global semantic consistency. In coronary angiography images, due to the slender and continuous nature of blood vessel structures, high requirements are placed on cross-regional structural consistency. Relying solely on local feature fusion can easily lead to breaks in the main vessel trunk and difficulty in fully restoring small branches. In contrast, the method of this invention constructs a unified semantic basis, maps features at different scales to a consistent semantic space, and uses this semantic representation to generate multi-scale features in reverse, thus ensuring cross-scale semantic consistency from a mechanistic perspective and effectively improving the connectivity of blood vessel structures.
[0061] Swin-UNet, by introducing a self-attention mechanism, models global information, thus mitigating the limited receptive field of convolutional networks to some extent. However, the global modeling of such methods is mainly reflected in the feature extraction stage, and the features still propagate independently in a hierarchical manner during subsequent decoding, lacking a unified semantic constraint mechanism. Therefore, inconsistencies in the semantics of features at different scales and deviations between detail recovery and the global structure may still exist during multi-scale restoration. The method of this invention, through a semantic basis reconstruction mechanism, continuously introduces unified semantic constraints during the decoding stage, enabling features at all scales to share a consistent semantic source during generation, thereby improving detail recovery capabilities while maintaining global structural consistency.
[0062] In summary, although existing mainstream methods have certain advantages in feature representation or global modeling, they still have limitations in complex vascular structure segmentation tasks due to the lack of a unified semantic basis and multi-scale semantic reconstruction mechanism.
[0063] The method of this invention achieves consistent modeling of multi-scale features in a unified semantic space through a semantic basis reconstruction mechanism. This mechanism is different from existing methods based on feature fusion or attention modeling. It effectively solves the problem of semantic inconsistency of multi-scale features and achieves significant improvements in both structural integrity and segmentation accuracy. This verifies the effectiveness and advancement of the method in the field of medical image segmentation.
[0064] In summary, the semantically guided coronary angiography image segmentation method proposed in this invention effectively solves the problem of inconsistent multi-scale feature semantics in existing methods by constructing a unified semantic representation and performing multi-scale semantic reconstruction. Experimental results show that this method can stably improve segmentation performance under different contrast strategies, especially in terms of vessel connectivity, preservation of small branches, and anti-interference ability. The above results verify the effectiveness of the structural design of the method and its application value in medical image segmentation tasks.
[0065] Example 2
[0066] This embodiment provides a semantically guided coronary angiography image segmentation system for performing the method described in Embodiment 1. The system includes the following modules: Image preprocessing module: Used to acquire and preprocess the coronary angiography images to be segmented. The preprocessing includes: converting the coronary angiography images to grayscale images, performing intensity normalization using two-dimensional maximum-minimum normalization, and enhancing contrast using constrained contrast adaptive histogram equalization. During the training phase, this module can also perform data augmentation on the training images, including random rotation, horizontal flipping, random scaling, random cropping, and random brightness adjustment; all geometric transformations are simultaneously applied to the corresponding ground truth annotations.
[0067] Dual-stream feature extraction module: includes a backbone encoder unit and an image encoder unit of a pre-trained medical vision basic model.
[0068] The trunk encoder unit is used to perform hierarchical downsampling on the preprocessed coronary angiography images to extract multi-scale local structural features. Specifically, the trunk encoder unit includes four downsampling stages, outputting multi-scale local structural features at four different spatial resolutions.
[0069] The image encoder unit of the pre-trained medical vision foundation model is used to extract multi-scale raw global semantic features from the preprocessed coronary angiography image. In this embodiment, the medical vision foundation model is the MedSAM2 model, whose image encoder outputs raw global semantic features at three different scales, including the highest-level global features, high-resolution local features, and higher-resolution local features.
[0070] The semantic reconstruction module includes a semantic alignment unit, a semantic basis construction unit, and a multi-scale semantic reconstruction unit.
[0071] The semantic alignment unit is used to semantically align the original multi-scale global semantic features to obtain aligned multi-scale semantic features. Specifically, the semantic alignment unit upsamples the high-resolution local features and even higher-resolution local features to the same spatial resolution as the highest-level global features through bilinear interpolation, and then maps the number of channels of the three features to the same value through three independent 1×1 convolutions, outputting the first alignment feature, the second alignment feature, and the third alignment feature.
[0072] The semantic base construction unit is used to concatenate aligned multi-scale semantic features along the channel dimension to obtain concatenated features; to perform gated fusion on the concatenated features in a unified semantic space to generate a spatially adaptive weight map; and to construct a unified semantic base by adaptively fusing the aligned multi-scale semantic features using the spatially adaptive weight map. Specifically, the semantic base construction unit includes a concatenation subunit, a gated weight generation subunit, and a fusion subunit: the concatenation subunit is used to concatenate the first aligned feature, the second aligned feature, and the third aligned feature along the channel dimension; the gated weight generation subunit is used to map the concatenated features using a 1×1 convolution to generate a single-channel feature map, and then obtains a spatially adaptive weight map through a Sigmoid activation function; the fusion subunit is used to calculate the unified semantic base using the spatially adaptive weight map according to a preset adaptive fusion formula.
[0073] The multi-scale semantic reconstruction unit is used to generate multi-scale semantic reconstruction features that match the spatial resolution and number of channels of each layer of the decoder, based on the unified semantic basis and through different scale mapping functions. The multi-scale semantic reconstruction features include features at least two different spatial scales. In this embodiment, four semantic reconstruction features with different spatial resolutions and number of channels are output, corresponding to the deepest layer, the second layer, the third layer, and the output layer of the decoder, respectively.
[0074] Decoding module: includes decoder unit and semantic injection unit.
[0075] The decoder unit is used to recover the feature map layer by layer. The decoder unit is symmetrical to the trunk encoder unit and contains four upsampling stages.
[0076] The semantic injection unit is used to inject the multi-scale semantic reconstruction features of the corresponding scale and the multi-scale local structural features extracted by the backbone encoder unit into the decoder unit at each layer of the decoder, guiding the decoder unit to recover the feature map layer by layer. Specifically, the decoder unit calculates layer by layer according to resolution: the output feature map of the current layer is equal to the upsampled feature map of the previous layer plus the local structural features of the backbone encoder of the corresponding scale plus the semantic reconstruction features of the corresponding scale.
[0077] Output module: This module activates and binarizes the feature map output by the decoder unit to obtain the blood vessel segmentation result. Specifically, the output module maps the decoder output layer features to a single-channel feature map through a 1×1 convolution, then uses a Sigmoid activation function to obtain a blood vessel segmentation probability map, and finally binarizes it with a threshold of 0.5 to output the blood vessel segmentation mask.
[0078] The modules described above work together to execute the segmentation method described in Example 1. In this embodiment, the system can use the loss function and training hyperparameters described in Example 1 to train the model during the training phase, and directly load pre-trained weights for segmentation during the inference phase.
[0079] Example 3
[0080] This embodiment provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the at least one processor to perform the semantically guided coronary angiography image segmentation method described in any one of Embodiments 1.
[0081] The processor may be a central processing unit, a graphics processing unit, a digital signal processor, an application-specific integrated circuit (ASIC), or other programmable logic devices. The memory may be random access memory, read-only memory, flash memory, hard disk, or other non-volatile storage media. The electronic device may be a server, workstation, desktop computer, laptop computer, embedded device, or mobile terminal.
[0082] Example 4
[0083] This embodiment provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the semantically guided coronary angiography image segmentation method described in any one of Embodiment 1.
[0084] The computer-readable storage medium includes, but is not limited to, magnetic storage media, optical storage media, semiconductor storage media, or other non-volatile storage media. The computer program may be written in a high-level programming language or a low-level instruction set, and may be compiled or interpreted and executed on the target platform.
[0085] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for semantic guidance based coronary angiogram segmentation, the method comprising: Includes the following steps: Step S1: Obtain the coronary angiography image to be segmented and perform preprocessing; Step S2: Perform dual-stream parallel feature extraction on the preprocessed coronary angiography images: Step S21: Perform hierarchical downsampling on the preprocessed coronary angiography image using the trunk encoder to extract multi-scale local structural features; Step S22: Extract multi-scale original global semantic features from the pre-processed coronary angiography image using the image encoder of the pre-trained medical vision basic model; Step S3: Semantic reconstruction of the original global semantic features at multiple scales: Step S31: performing semantic alignment on the multi-scale original global semantic features to obtain aligned multi-scale semantic features; the multi-scale original global semantic features include three features of different scales: a first original global semantic feature F 32 , a second original global semantic feature F 64 , and a third original global semantic feature F 128 ; the semantic alignment specifically comprises: up-sampling the third original global semantic feature F 128 and the second original global semantic feature F 64 to the same spatial resolution as the first original global semantic feature F 32 by bilinear interpolation, and then mapping the channel numbers to the same number of values by three independent 1×1 convolutions to obtain an aligned first aligned feature F , an aligned second aligned feature F , and an aligned third aligned feature F . Step S32: Align the second feature Third alignment feature and the first alignment feature By concatenating along the channel dimension, the concatenated features are obtained. In a unified semantic space, the splicing features are... Gated fusion is performed using a 1×1 convolution on the spliced features. The mapping is performed to generate a single-channel feature map, which is then passed through the Sigmoid activation function to obtain a spatial adaptive weight map G. By utilizing the multi-scale semantic features obtained through adaptive fusion and alignment of the spatially adaptive weighted graph, a unified semantic basis with global consistency is constructed. : ; Where ⊙ represents element-wise multiplication; Step S33: Based on the unified semantic base, multi-scale semantic reconstruction features that match the spatial resolution and number of channels of each layer of the decoder are generated in reverse through different scale mapping functions. The multi-scale semantic reconstruction features include features of at least two different spatial scales. Step S4: At each layer of the decoder, the multi-scale semantic reconstruction features of the corresponding scale and the multi-scale local structural features of the backbone encoder are injected into the decoder layer by layer to guide the decoder to recover the feature map layer by layer. Step S5: Activate and binarize the feature map output by the decoder to output the blood vessel segmentation result.
2. The coronary angiography image segmentation method according to claim 1, characterized in that, The multi-scale semantic reconstruction features generated in step S33 include four features with different spatial resolutions and channel numbers: the first semantic reconstruction feature Second semantic reconstruction features Third semantic reconstruction features and fourth semantic reconstruction features ; Represents the feature dimension.
3. The method according to claim 2, characterized in that, In step S4, the decoder performs layer-by-layer calculations according to resolution: ; Where R represents the spatial resolution of the current feature map, This represents the output characteristics of the decoder at a resolution of R. This indicates the decoding features of the layer above the decoder. This represents the local structural features derived from the backbone encoder at a resolution of R. This represents the semantic reconstruction features at a resolution of R, and UP(·) represents the upsampling operation.
4. The method according to claim 3, characterized in that, The upsampling operation UP(·) is implemented using bilinear interpolation or transposed convolution.
5. The method according to claim 1, characterized in that, The pre-trained medical vision baseline model is the MedSAM2 model, and the backbone encoder and the decoder adopt a U-Net structure or a variant thereof.
6. A semantically guided coronary angiography image segmentation system, characterized in that, include: Image preprocessing module: used to acquire the coronary angiography image to be segmented and perform preprocessing; The dual-stream feature extraction module includes a backbone encoder unit and an image encoder unit of a pre-trained medical vision basic model. The backbone encoder unit is used to perform hierarchical downsampling on the pre-processed coronary angiography image to extract multi-scale local structural features. The image encoder unit of the pre-trained medical vision basic model is used to extract multi-scale original global semantic features of the pre-processed coronary angiography image. The semantic reconstruction module includes a semantic alignment unit, a semantic basis construction unit, and a multi-scale semantic reconstruction unit. The semantic alignment unit is used to semantically align the multi-scale original global semantic features to obtain aligned multi-scale semantic features; the multi-scale original global semantic features include features at three different scales: the first original global semantic feature F 32 Second original global semantic feature F 64 and the third original global semantic feature F 128 The semantic alignment specifically involves: aligning the third original global semantic feature F... 128 Second original global semantic feature F 64 Upsampled to the first original global semantic feature F using bilinear interpolation. 32 Using the same spatial resolution, the number of channels is then uniformly mapped to the same value through three independent 1×1 convolutions to obtain the first alignment feature after alignment. Second alignment feature and third alignment feature ; The semantic base construction unit is used to align the second feature Third alignment feature and the first alignment feature By concatenating along the channel dimension, the concatenated features are obtained. In a unified semantic space, the splicing features are... Gated fusion is performed using a 1×1 convolution on the spliced features. The mapping is performed to generate a single-channel feature map, which is then passed through the Sigmoid activation function to obtain a spatial adaptive weight map G. By utilizing the multi-scale semantic features obtained through adaptive fusion and alignment of the spatially adaptive weighted graph, a unified semantic basis with global consistency is constructed. : ; Where ⊙ represents element-wise multiplication; The multi-scale semantic reconstruction unit is used to generate multi-scale semantic reconstruction features that match the spatial resolution and number of channels of each layer of the decoder based on the unified semantic base and through different scale mapping functions. The multi-scale semantic reconstruction features include features at least two different spatial scales. Decoding module: includes a decoder unit and a semantic injection unit; the decoder unit is used to recover the feature map layer by layer; the semantic injection unit is used to inject the multi-scale semantic reconstruction features of the corresponding scale and the multi-scale local structural features extracted by the backbone encoder unit into the decoder unit at each layer of the decoder, so as to guide the decoder unit to recover the feature map layer by layer; Output module: used to activate and binarize the feature map output by the decoder unit to obtain the blood vessel segmentation result.
7. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor. When the computer program is executed by the at least one processor, it causes the at least one processor to perform the semantically guided coronary angiography image segmentation method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The system contains a computer program that, when executed by a processor, implements the semantically guided coronary angiography image segmentation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Refined target detection method based on superpixel guidance and double-flow feature fusion
CN121330460A
Cascade global-to-local cross-scale fusion colonoscope polyp segmentation method
CN122115847A