Coronary vessel segmentation system and method fusing high and low frequency characteristics

By combining cascaded multi-scale strip convolution and an adaptive module for separating high- and low-frequency features, the problem of overly smooth segmentation results and local-to-global feature imbalance in existing coronary artery segmentation algorithms is solved, achieving higher accuracy and more stable coronary artery segmentation.

CN120931928APending Publication Date: 2025-11-11SOUTHWEAT UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511057588.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing deep learning-based coronary artery segmentation algorithms suffer from problems such as overly smooth segmentation results, imbalance between local and global features, and insufficient anti-interference ability when processing coronary angiography images. They are unable to accurately segment the fine branches and global structure of blood vessels in complex backgrounds.

Method used

The traditional downsampling module is replaced by a cascaded multi-scale strip convolution (CMSC) module and a high-low frequency feature separation adaptive module (HLFAE). By extracting features through cascaded strip convolution kernels of different scales and adaptively enhancing the high- and low-frequency components, efficient and accurate coronary artery segmentation is achieved.

Benefits of technology

It significantly improves segmentation accuracy and robustness, can capture vascular details and global structure more clearly, effectively avoids over- or under-segmentation, enhances the model's resistance to complex backgrounds and noise, and improves segmentation accuracy and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931928A_ABST
    Figure CN120931928A_ABST
Patent Text Reader

Abstract

The invention provides a coronary blood vessel segmentation system and method fusing high and low frequency features. The method comprises the steps that a cascade multi-scale stripe convolution module is cascaded with stripe convolution operations of different scales, and features are extracted in a multi-level mode from the global mode to the local mode; the high and low frequency feature separation adaptive module separates an input feature map into high-frequency and low-frequency components, adaptive enhancement is carried out on the high-frequency and low-frequency components respectively, and finally the enhanced high-frequency and low-frequency components are fused; the multi-scale stripe high and low frequency enhancement module firstly inputs the input feature map into the cascade multi-scale stripe convolution module to extract features, then inputs the output feature map into the high and low frequency feature separation adaptive module to carry out high and low frequency separation and adaptive enhancement, and finally outputs the obtained feature map. Based on the technical scheme of the invention, the segmentation precision and the detail capture capability can be improved; local details and a global structure can be clearly distinguished, and the problem of excessive segmentation or insufficient segmentation is effectively avoided; the robustness of the model is enhanced, and the blood vessel can be stably and accurately segmented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision, image recognition and other technologies, and in particular to a coronary artery segmentation system and method that integrates high and low frequency features. Background Technology

[0002] Coronary artery disease (CAD) is one of the most common cardiovascular diseases worldwide and a leading cause of death. With an aging population and the prevalence of unhealthy lifestyles, its prevalence continues to rise, posing a serious challenge to global healthcare systems. Extracting complete and clear vascular tree structures from complex coronary angiography images has become a key scientific problem for improving diagnostic accuracy and clinical application value. Deep learning-based medical image segmentation methods have become a popular approach.

[0003] In the field of coronary angiography image segmentation, deep learning technology is gradually replacing traditional methods. Traditional vessel segmentation methods, such as threshold-based segmentation and region growing algorithms, while simple and efficient, perform poorly when faced with complex backgrounds, noise interference, and individual differences in coronary angiography images. They struggle to accurately distinguish vessels from surrounding tissues, especially when dealing with small vessel branches and low-contrast regions, easily leading to missed or incorrect segmentation.

[0004] Traditional machine learning methods, such as support vector machines and random forests, while offering improvements over conventional methods by using manually designed features for classification, are still limited by their sensitivity to noise and the constraints of manually designed features. In coronary angiography images, due to the complex and diverse vascular structures and the presence of numerous non-target regions, these methods struggle to extract sufficiently effective features, resulting in segmentation accuracy that fails to meet clinical needs.

[0005] In recent years, significant progress has been made in vessel segmentation methods based on convolutional neural networks (CNNs). Fully convolutional networks (FCNs) have achieved end-to-end pixel-level prediction for the first time, but they lack full utilization of contextual information. When dealing with the complex morphology of coronary vessels and background noise, they cannot effectively capture the global structure and local details of the vessels, resulting in low segmentation accuracy.

[0006] U-Net improves local detail preservation through its symmetric encoder-decoder structure and skip connections, but it suffers from high computational complexity and is prone to overfitting during training, limiting its generalization ability. Furthermore, it struggles to balance local and global feature information, often resulting in oversegmentation or undersegmentation when segmenting coronary arteries, failing to accurately and completely capture the overall topological structure of the vessels.

[0007] To optimize performance, various improvement schemes have emerged. For example, ResUNet introduces residual connections to alleviate gradient vanishing, Attention U-Net incorporates an attention mechanism to enhance the ability to focus on important features, U-Net++ optimizes feature fusion through nested dense skip connections, and U-Net3+ introduces full-scale skip connections and multi-level feature fusion. However, these methods still have certain limitations when dealing with problems such as low contrast, complex backgrounds, complex vascular structures with small branches, blurred boundaries, and low signal-to-noise ratio in coronary angiography images. They struggle to effectively separate and enhance high-frequency and low-frequency features in images, cannot accurately segment small branches of blood vessels in complex backgrounds, and need improvement in balancing local and global feature information.

[0008] 1. Overly Smooth Segmentation Results: Existing deep learning-based coronary artery segmentation algorithms often exhibit overly smoothed segmentation results when processing coronary angiography images. This results in the inability to clearly present the edges of blood vessels and some fine structures, losing crucial details and affecting doctors' accurate judgment of vascular lesions, such as making it difficult to identify early, small vascular stenosis or lesion sites.

[0009] 2. Imbalance between Local and Global Features: Current algorithms struggle to balance local and global feature information. During segmentation, they either overemphasize local details, neglecting the overall topological structure of the blood vessel, resulting in oversegmentation and incorrectly dividing local regions of the vessel into multiple parts; or they overemphasize the global structure, ignoring local details such as small vascular branches, leading to undersegmentation and an inability to fully represent the entire blood vessel, hindering comprehensive and accurate diagnosis of coronary artery diseases. Traditional square convolutional kernels may fail to fully capture the elongated features of blood vessels when dealing with complex vascular morphologies.

[0010] 3. Complex Background and Noise Interference: Coronary angiography images suffer from low contrast, complex backgrounds, intricate vascular structures with fine branches, blurred boundaries, and low signal-to-noise ratios. Existing algorithms are insufficiently resistant to interference in these complex situations, easily affected by non-target tissue structures and motion artifacts, making it difficult to accurately extract vascular features, leading to decreased segmentation accuracy and impacting the accuracy and reliability of diagnosis. Summary of the Invention

[0011] To address the problems in the existing technologies, this application proposes a coronary artery segmentation system and method that integrates high- and low-frequency features. This system is a coronary artery segmentation network, FS-UNet3+, that combines high- and low-frequency features. The core technologies include a cascaded multi-scale strip convolution (CMSC) module and a high- and low-frequency feature separation adaptive module (HLFAE). The CMSC module replaces square convolution kernels with strip convolution kernels and cascades them, enabling multi-level feature extraction from global to local perspectives, effectively capturing vessel morphology. The HLFAE module first separates the high- and low-frequency components of the input features, then adaptively enhances them separately. High-frequency enhancement focuses on details, while low-frequency enhancement preserves the global structure, and finally, the components are fused. Based on the U-Net3+ architecture, a multi-scale strip high- and low-frequency enhancement module (MSFE) that integrates these two modules replaces the traditional downsampling module, thereby achieving efficient and accurate coronary artery segmentation.

[0012] 1) Cascaded Multi-Scale Strip Convolution (CMSC) Module: The traditional square convolution kernel is replaced with a strip convolution kernel, and a cascaded approach is used to refine the feature map at different scales, thereby achieving multi-level feature extraction and effectively capturing vascular morphological features.

[0013] 2) High- and low-frequency feature separation adaptive module (HLFAE): The input features are separated into high and low frequencies, and the high-frequency and low-frequency components are adaptively enhanced respectively. The high-frequency enhancement focuses on the details of blood vessels, while the low-frequency enhancement preserves the global structure of blood vessels, and finally achieves effective fusion.

[0014] 3) Network architecture optimization: Based on the U-Net3+ architecture, the traditional downsampling module is replaced with a multi-scale strip high and low frequency enhancement module (MSFE) that integrates the CMSC module and HLFAE module to achieve efficient and accurate coronary artery segmentation.

[0015] A coronary artery segmentation system integrating high- and low-frequency features according to the present invention, in one embodiment, includes:

[0016] The cascaded multi-scale strip convolution module cascades strip convolution operations at different scales, maintaining attention to global and local information through hierarchical feature extraction and feature feedback mechanisms.

[0017] The high- and low-frequency feature separation adaptive module separates the input feature map into high-frequency components and low-frequency components, performs adaptive enhancement on the high- and low-frequency components respectively, and finally fuses the enhanced high- and low-frequency components.

[0018] The multi-scale strip high- and low-frequency enhancement module first inputs the input feature map into the cascaded multi-scale strip convolution module for feature extraction, then inputs the output feature map of the cascaded multi-scale strip convolution module into the high- and low-frequency feature separation adaptive module for high- and low-frequency separation and adaptive enhancement, and finally outputs the obtained feature map.

[0019] In one implementation, the cascaded multi-scale strip convolution module first performs strip convolution operations at different scales on the input feature map, then concatenates the convolution results at different scales, and then performs feature fusion through a 1×1 convolution kernel to obtain the final output feature map; the different scales include, but are not limited to, using 1×n and n×1 strip convolution kernels, where n is a different integer.

[0020] In one implementation, the adaptive enhancement in the high- and low-frequency feature separation adaptive module uses a combination of convolutional layers and activation functions to adaptively adjust the degree of enhancement based on the statistical information of the features. The high- and low-frequency feature separation adaptive module fuses the enhanced high-frequency and low-frequency components by concatenating the enhanced high-frequency and low-frequency components, and then performs feature fusion through a 1×1 convolutional kernel to obtain the final output feature map.

[0021] In one implementation, the high- and low-frequency feature separation adaptive module separates the input feature map into high-frequency components and low-frequency components, including: input feature map X∈R B×C×H×W Where B is the batch size, C is the number of channels, H and W are the spatial dimensions of the feature map; X is the feature map; X low Low-frequency characteristics; X high For high-frequency features; R is the set of real numbers; the input feature map first extracts low-frequency information through pooling, and then calculates high-frequency information using the difference. The formula for separating high- and low-frequency features is:

[0022] X low =AvgPool(X)

[0023] X high =X-Upsample(X low )

[0024] AvgPool(X) performs 2×2 average pooling to extract low-frequency information; Upsample(X) low The original size is restored by bilinear interpolation, and then subtracted from X to obtain high-frequency information.

[0025] In one embodiment, the high- and low-frequency feature separation adaptive module adaptively enhances the high- and low-frequency components respectively, including:

[0026] Dilated convolution feature extraction is performed using feature_extraction(X) to form high-frequency enhanced features; the formula for high-frequency feature extraction is:

[0027] X' high ==feature_extraction(X high )

[0028]

[0029] The feature_extraction(X) function consists of four convolutional layers with different dilation rates, where d represents the dilation rate of the convolution; || represents the channel concatenation operation; and X' high The high-frequency enhancement feature is the enhanced high-frequency characteristic; Represents the high-frequency feature X high Perform a 3x3 convolution operation with a kernel size of 4. This results in a feature map with four times the number of channels. Then, perform dimensionality reduction using 1x1 and 3x3 convolutions, as shown in the formula:

[0030] X′ high =Conv 1×1 (Conv 3×3 (Concat(X′ high ))

[0031] Low-frequency features are extracted through pooling operations, and weights are calculated using channel attention (SEAttention) to enhance them. The formula for enhancing low-frequency features is as follows:

[0032] S low =σ(W2·ReLU(W1·AvgPool(X) low ))+

[0033] W2·ReLU(W1·MaxPool(X low )))

[0034] X′ low ==X low ·S low

[0035] SEAttention calculates global pooling and uses a multilayer perceptron to perform non-linear modeling on the globally pooled features to generate channel attention weights S. low W1 and W2 are the parameters of the fully connected layer, and σ is the sigmoid activation function; calculate the weighted low-frequency feature X. low MaxPool is the maximum pooling method; X′ low is the enhanced low-frequency feature; ReLU is the activation function.

[0036] In one implementation, the high- and low-frequency feature separation adaptive module finally fuses the enhanced high- and low-frequency components, including: introducing an adaptive mechanism to assign weights to different channels according to the importance of the features; the high- and low-frequency feature fusion formula and the self-attention module formula are as follows:

[0037] X fusion =CBR(Concat(X′) high ,X′low ))

[0038] X final ==SEAttention(X) fusion )

[0039] Y = X final +x

[0040] Among them, X fusion This is a feature obtained by fusing high and low frequency characteristics; X final X is the final enhanced feature map after adjustment by the SE attention module; X is the original input, and Y is the final output feature map; `conact` is the concatenation, which concatenates two feature maps along the channel dimension; first, high-frequency and low-frequency features are concatenated, and then fused using the CBR process. The CBR process includes extracting local spatial features from the concatenated features through a 3×3 convolutional layer, performing batch normalization, and introducing a non-linear combination operation through the ReLU activation function; finally, it is adjusted again by the SE attention module; the final enhanced feature X is... final The input X is added back as a residual connection; the adaptive mechanism includes, but is not limited to, the SE attention mechanism.

[0041] In one embodiment, a coronary artery segmentation method that integrates high- and low-frequency features includes the following steps:

[0042] Step S1: Input the input feature map into the cascaded multi-scale strip convolution step for feature extraction, and obtain the feature map output by the cascaded multi-scale strip convolution step.

[0043] Step S2: Input the output feature map of the cascaded multi-scale strip convolution step into the high-low frequency feature separation adaptive step for high-low frequency separation and adaptive enhancement to obtain the final output feature map;

[0044] The cascaded multi-scale strip convolution step involves cascading strip convolution operations of different scales to extract features at multiple levels, from global to local.

[0045] The high- and low-frequency feature separation adaptive step separates the input feature map into high-frequency components and low-frequency components, adaptively enhances the high- and low-frequency components respectively, and finally fuses the enhanced high- and low-frequency components.

[0046] In one embodiment, the cascaded multi-scale strip convolution step in step S1 includes: first, performing strip convolution operations of different scales on the input feature map; then, concatenating the convolution results of different scales; and then performing feature fusion through a 1×1 convolution kernel to obtain the final output feature map; the different scales include, but are not limited to, using 1×n and n×1 strip convolution kernels, where n is a different integer.

[0047] In one embodiment, the high- and low-frequency feature separation adaptive step in step S2 includes:

[0048] Step S21: Separate high and low frequencies from the input features. The input feature map X∈R B×C×H×W Where B is the batch size, C is the number of channels, H and W are the spatial dimensions of the feature map; X is the feature map; X low Low-frequency characteristics; X high For high-frequency features; R is the set of real numbers; the input feature map first extracts low-frequency information through pooling, and then calculates high-frequency information using the difference. The formula for separating high- and low-frequency features is:

[0049] X low =AvgPool(X)

[0050] X high =X-Upsample(X low )

[0051] AvgPool(X) performs 2×2 average pooling to extract low-frequency information; Upsample(X) low The original size is recovered by bilinear interpolation, and then subtracted from X to obtain high-frequency information;

[0052] Step S22: Adaptively enhance the high and low frequency components by using feature_extraction(X) for dilated convolution feature extraction to form high-frequency enhanced features; the formula for high-frequency feature extraction is:

[0053] X' high ==feature_extraction(X high )

[0054]

[0055] The feature_extraction(X) function consists of four convolutional layers with different dilation rates, where d represents the dilation rate of the convolution; || represents the channel concatenation operation; and X' high The high-frequency enhancement feature is the enhanced high-frequency characteristic; Represents the high-frequency feature X high Perform a 3x3 convolution operation with a kernel size of 4; the final result is a feature map with four times the number of channels, which is then subjected to 1x1 and 3x3 convolutions for dimensionality reduction, as shown in the formula:

[0056] X′ high =Conv 1×1 (Conv 3×3 (Concat(X′ high ))

[0057] Low-frequency features are extracted through pooling operations, and weights are calculated using channel attention (SEAttention) to enhance them. The formula for enhancing low-frequency features is as follows:

[0058] S low =σ(W2·ReLU(W1·AvgPool(X) low ))+

[0059] W2·ReLU(W1·MaxPool(X low )))

[0060] X′ low ==X low ·S low

[0061] SEAttention calculates global pooling and uses a multilayer perceptron to perform non-linear modeling on the globally pooled features to generate channel attention weights S. low W1 and W2 are the parameters of the fully connected layer, and σ is the sigmoid activation function; calculate the weighted low-frequency feature X. low MaxPool is the maximum pooling method; X′ low The signal is the enhanced low-frequency feature; ReLU is the activation function.

[0062] Step S23: The enhanced high- and low-frequency components are fused, and an adaptive mechanism is introduced to assign weights to different channels according to the importance of the features; the high- and low-frequency feature fusion formula and the self-attention module formula are as follows:

[0063] X fusion =CBR(Concat(X′) high ,X′ ow ))

[0064] X final ==SEAttention(X) fusion )

[0065] Y = X final +X

[0066] Among them, X fusion This is a feature obtained by fusing high and low frequency characteristics; X finalX is the final enhanced feature map after adjustment by the SE attention module; X is the original input, and Y is the final output feature map; `conact` is the concatenation, which concatenates two feature maps along the channel dimension; first, high-frequency and low-frequency features are concatenated, and then fused using the CBR process. The CBR process includes extracting local spatial features from the concatenated features through a 3×3 convolutional layer, performing batch normalization, and introducing a non-linear combination operation through the ReLU activation function; finally, it is adjusted again by the SE attention module; the final enhanced feature X is... final The input X is added back as a residual connection; the adaptive mechanism includes, but is not limited to, the SE attention mechanism.

[0067] The above-mentioned technical features can be combined in various suitable ways or replaced by equivalent technical features, as long as the purpose of the present invention can be achieved.

[0068] This invention provides a coronary artery segmentation system and method that integrates high- and low-frequency features. Compared with existing technologies, it has at least the following advantages: ① Improved segmentation accuracy and detail capture capability: By using a cascaded multi-scale strip convolution (CMSC) module, the traditional square convolution kernel is replaced with a strip convolution kernel, and a cascaded strategy is adopted. This allows for progressive refinement of feature maps at different scales, fusing features at multiple levels from global to local. This significantly enhances the model's ability to express vascular morphological features and greatly improves the segmentation accuracy of small vascular branches.

[0069] ② Optimizing the balance between global and local features: The High-Low Frequency Feature Separation Adaptive Module (HLFAE) performs high- and low-frequency separation and adaptive enhancement on the input features. Enhancement of high-frequency components helps capture detailed information such as vessel edges and textures, while enhancement of low-frequency components preserves global structural information such as the overall direction and distribution of blood vessels. This complementary approach enables the model to clearly distinguish between local details and global structure when dealing with complex backgrounds and multi-scale features, effectively avoiding over-segmentation or under-segmentation.

[0070] ③. Enhanced Model Robustness: The strip convolution kernel design of the CMSC module and the processing method of high and low frequency features by the HLFAE module significantly improve the model's resistance to complex backgrounds and noise. When faced with interference such as low contrast, complex backgrounds, and motion artifacts in coronary angiography images, it can effectively suppress background interference and accurately focus on blood vessels and their course. This invention has significant advantages in handling background interference and edge details, and can stably and accurately segment blood vessels in coronary angiography images with different imaging conditions and significant individual differences. Attached Figure Description

[0071] The invention will now be described in more detail with reference to embodiments and the accompanying drawings.

[0072] Figure 1A schematic diagram of the cascaded multi-scale strip convolution CMSC module of the present invention is shown;

[0073] Figure 2 A schematic diagram of the adaptive HLFAE module for high and low frequency feature separation of the present invention is shown;

[0074] Figure 3 A schematic diagram of the feature extraction module of the present invention is shown;

[0075] Figure 4 A schematic diagram of the channel attention module of the present invention is shown;

[0076] Figure 5 A schematic diagram of the FS-UNet3+ structure of the present invention is shown. Detailed Implementation

[0077] The invention will now be further described with reference to the accompanying drawings.

[0078] This invention provides a coronary artery segmentation system and method that integrates high- and low-frequency features.

[0079] In one embodiment, the proposed coronary artery segmentation network FS-UNet3+ combines high- and low-frequency features, and is an improvement upon the U-Net3+ architecture. Its core lies in introducing a cascaded multi-scale strip convolutional (CMSC) module and a high- and low-frequency feature separation adaptive module (HLFAE), and fusing them into a multi-scale strip high- and low-frequency enhancement module (MSFE) to replace the traditional downsampling module, thereby achieving more accurate coronary artery segmentation. This includes:

[0080] (1) Concatenated Multiscale Strip Convolution (CMSC) module

[0081] Principle: The CMSC module replaces square convolution kernels with strip convolution kernels, which are better suited to the slender structure of blood vessels. By cascading strip convolution operations of different scales, features can be extracted at multiple levels from global to local, enhancing the model's ability to express vascular morphological features.

[0082] Implementation details: In practical applications, the CMSC module first performs strip convolution operations at different scales on the input feature map, such as using 1×n and n×1 strip convolution kernels (where n is a different integer). Then, the convolution results at different scales are concatenated, and then feature fusion is performed through a 1×1 convolution kernel to obtain the final output feature map.

[0083] (2) High- and low-frequency feature separation adaptive module (HLFAE)

[0084] Principle: To better balance local details and global structural information, the HLFAE module separates the input feature map into high-frequency and low-frequency components. The high-frequency components contain detailed information such as the edges and textures of blood vessels, while the low-frequency components contain global structural information such as the overall direction and distribution of blood vessels. Adaptive enhancement is applied to both high- and low-frequency components separately, and finally, the enhanced high- and low-frequency components are fused.

[0085] Specific implementation: High-frequency separation: Use a learnable filter to separate the input feature map into high-frequency components and low-frequency components.

[0086] Adaptive Enhancement: High-frequency and low-frequency components are enhanced separately using an adaptive enhancement submodule. This submodule can employ a combination of convolutional layers and activation functions, adaptively adjusting the degree of enhancement based on the statistical information of the features.

[0087] Fusion: The enhanced high-frequency components and low-frequency components are concatenated and then fused using a 1×1 convolutional kernel to obtain the final output feature map.

[0088] The module first separates the input features into high and low frequencies. Input feature map: X∈R B×C×H×W Where B is the batch size, C is the number of channels, H and W are the spatial dimensions of the feature map; X is the feature map; X low Low-frequency characteristics; X high For high-frequency features, R is the set of real numbers, where R represents the fact that each element of the feature map is a real value. The input feature map is first processed by pooling to extract low-frequency information, and then the high-frequency information is calculated using the difference. The formula for separating high and low frequency features is as follows:

[0089] X low =AvgPool(X)

[0090] X high =X-Upsample(X low )

[0091] AvgPool(X) performs 2×2 average pooling to extract low-frequency information. Upsample(X) low The original size is restored by bilinear interpolation, and then subtracted from X to obtain high-frequency information.

[0092] To enhance the capture of contextual information at different scales, various dilation rates (e.g., 1, 3, 5, 7) were used for convolution. This multi-scale processing can capture the morphological features of blood vessels at different scales, making it particularly suitable for fine-grained feature learning against complex backgrounds. Figure 3As shown, these convolutions can capture detailed information, especially in image edges or textured regions. Dilated convolution feature extraction is performed using feature_extraction(X) to form high-frequency enhanced features. The formula for high-frequency feature extraction is as follows:

[0093] X' high ==feature_extraction(X high )

[0094]

[0095] The feature_extraction(X) function consists of four convolutional layers with different dilation rates, where d represents the dilation rate of the convolution; || represents the channel concatenation operation; and X' high The high-frequency enhancement feature is the enhanced high-frequency characteristic; Represents the high-frequency feature X high A 3x3 convolution operation is performed, where 'd' represents the dilation rate. This results in a feature map with four times the number of channels. This is then further reduced in dimensionality using 1x1 and 3x3 convolutions, as shown in the following formula:

[0096] X′ high =Conv 1×1 (Conv 3×3 (Concat(X′ high ))

[0097] Low-frequency features are extracted through pooling operations, such as... Figure 4 As shown, low-frequency features represent the macroscopic structure of the image. Weights are calculated using channel attention (SEAttention) to enhance low-frequency features. The formula for low-frequency feature enhancement is as follows:

[0098] S low =σ(W2·ReLU(W1·AvgPool(X) low ))+

[0099] W2·ReLU(W1·MaxPool(X low )))

[0100] X′ low ==X low ·S low

[0101] SEAttention is calculated using global pooling + MLP to generate channel attention weights S. low W1 and W2 are the parameters of the fully connected layer, and σ is the sigmoid activation function; calculate the weighted low-frequency feature X. lowMaxPool is a max pooling method that uses downsampling to extract low-frequency features; X′ low is the enhanced low-frequency feature; ReLU is the activation function; MLP is a multilayer perceptron that performs non-linear modeling on the globally pooled features to generate importance weights for each channel.

[0102] Finally, an adaptive mechanism (such as the SE attention mechanism) is introduced, which can assign weights to different channels based on the importance of features. This means that the network can dynamically enhance the representation capabilities of high-frequency and low-frequency features, further improving the perception of image details and global structure. The high- and low-frequency feature fusion formula and the self-attention module formula are as follows:

[0103] X fusion =CBR(Concat(X′) high ,X′ low ))

[0104] X final ==SEAttention(X) fusion )

[0105] Y = X final +X

[0106] First, high-frequency and low-frequency features are concatenated, then fused using CBR (Convolutional Normalization), where CBR(X) = convolution (3x3) + batch normalization + ReLU activation; CBR is conv + BatchNorm + ReLU. The CBR fusion module performs feature processing: a 3×3 convolutional layer is used to extract local spatial features, batch normalization is used to accelerate training convergence and improve stability, and the ReLU activation function introduces non-linearity to enhance the model's expressive power. In other words, the CBR process involves extracting local spatial features from the concatenated features using a 3×3 convolutional layer, performing batch normalization, and introducing a non-linear combination operation through the ReLU activation function. This CBR combination operation effectively fuses the input high-frequency and low-frequency features, improving the network's ability to perceive image details and structural information. Finally, the SE attention module is used for further adjustments. The final enhanced feature X is then... final Adding back the input X as a residual connection ensures gradient stability and preserves the original information; X fusion This is a feature obtained by fusing high and low frequency characteristics; X final X is the final enhanced feature map after adjustment by the SE attention module; X is the original input, Y is the final output feature map (with residual connections); and contact is the concatenation, which concatenates the two feature maps along the channel dimension.

[0107] (3) Multi-scale strip high and low frequency enhancement module (MSFE)

[0108] Principle: The MSFE module integrates the CMSC module and the HLFAE module, giving full play to the advantages of both to achieve multi-scale and multi-level extraction and enhancement of vascular features.

[0109] Specific implementation: The input feature map is first input into the CMSC module for feature extraction, and then the output feature map of the CMSC module is input into the HLFAE module for high and low frequency separation and adaptive enhancement. The final feature map is the output of the MSFE module.

[0110] Based on the advantages of U-Net3+, this invention proposes an improved network structure, High-Low Frequency Separation UNet3+ (HLFS-UNet3+), which replaces the traditional downsampling module in U-Net3+ with a Hybrid Downsampling Module (MSFE). Figure 5 As shown, this approach further improves the accuracy and robustness of coronary artery segmentation. Specifically, the MSFE module first replaces the traditional square convolution kernel (e.g., 3×3) with strip convolutions (e.g., 3×1, 5×1) to more efficiently extract the slender structural features of blood vessels and enhance sensitivity to details. Building upon this, a cascaded convolution (CMSC) structure is adopted. Small-scale convolutions (e.g., 3×1) are used to capture local detail features, followed by larger-scale convolutions (e.g., 5×1 or 7×1) to extract global contextual information, thus achieving multi-level feature learning from local to global perspectives. Furthermore, the MSFE module introduces an HLFAE module to separate and adaptively enhance high-frequency and low-frequency features during downsampling. The HLFAE module first extracts and enhances low-frequency features, then combines the enhanced low-frequency features with high-frequency features, fully utilizing multi-scale information in the image. Through this hybrid design, the MSFE module can not only effectively handle the complex structure of blood vessels and background noise but also significantly improve the model's ability to preserve detail information, thus providing a superior feature extraction solution for coronary artery segmentation tasks.

[0111] In one embodiment, the network architecture design of this invention proposes a coronary artery segmentation network FS-UNet3+ that combines high- and low-frequency features. Based on the U-Net3+ architecture, it replaces the traditional downsampling module with a multi-scale strip high- and low-frequency enhancement module (MSFE) that integrates cascaded multi-scale strip convolutional (CMSC) modules and high- and low-frequency feature adaptive enhancement modules (HLFAE), thereby improving the accuracy and robustness of coronary artery segmentation. This architecture design can effectively handle complex vascular structures and background noise, taking into account both detailed and global features to achieve more accurate segmentation, which is the core architectural innovation that distinguishes it from existing technologies.

[0112] In one embodiment, the Cascaded Multi-Scale Strip Convolution (CMSC) module of the present invention replaces the traditional square convolution kernel with strip convolution kernels (such as 7×1, 5×1, 3×1) and adopts a cascaded convolution strategy. First, global features are extracted through larger-scale strip convolution, and then passed to smaller-scale strip convolution to extract local detail information. This achieves multi-level feature fusion from global to local, effectively improving the feature extraction capability for slender vascular structures, reducing computational complexity, and improving segmentation accuracy, especially showing significant advantages in the segmentation of small vascular branches.

[0113] In one embodiment, the High-Low Frequency Feature Separation Adaptive Module (HLFAE) of the present invention performs high- and low-frequency separation on the input features, adaptively enhancing the high-frequency (detail information) and low-frequency (global structural information) components respectively. High-frequency features are processed through multi-scale convolution, low-frequency features are enhanced using a channel attention mechanism, and finally, high- and low-frequency features are fused and an adaptive mechanism is introduced to adjust the weights, improving the model's ability to perceive features at different scales, better balancing local details and global structure, and accurately segmenting blood vessels in complex backgrounds.

[0114] In one embodiment, the present invention employs a modular collaborative optimization approach: the fusion of the CMSC and HLFAE modules, which work together to provide higher-quality multi-scale feature inputs to the HLFAE module, which in turn further optimizes these features, thereby improving model performance from different perspectives. This collaborative optimization design between modules significantly enhances the model's ability to identify and segment both the overall structure of blood vessels and small branches, achieving significant improvements across multiple evaluation metrics and proving crucial for achieving high-precision coronary artery segmentation.

[0115] In one embodiment, this invention improves segmentation accuracy and detail capture capability: by using a cascaded multi-scale strip convolution (CMSC) module, the traditional square convolution kernel is replaced with a strip convolution kernel, and a cascading strategy is adopted. This allows for progressive refinement of feature maps at different scales, fusing features at multiple levels from global to local. This significantly enhances the model's ability to express vascular morphological features, greatly improving the segmentation accuracy of small vascular branches. Compared with existing technologies, this invention shows a significant improvement in the mPrecision metric. For example, in experiments on the XCAD dataset, it improved by 2.79% compared to the baseline model (from 90.33% to 93.12%), accurately delineating the subtle edges and branches of blood vessels, allowing doctors to more clearly observe the morphological structure of blood vessels and helping to detect early, small lesions.

[0116] This invention optimizes the balance between global and local features: the High-Low Frequency Feature Separation Adaptive Module (HLFAE) performs high- and low-frequency separation and adaptive enhancement on the input features. High-frequency component enhancement helps capture detailed information such as vessel edges and textures, while low-frequency component enhancement preserves global structural information such as the overall direction and distribution of vessels. This complementary approach enables the model to clearly distinguish between local details and global structure when dealing with complex backgrounds and multi-scale features, effectively avoiding over-segmentation or under-segmentation. In experiments, this invention improves the average Dice similarity coefficient (mDice) by 10.51% (reaching 90.78%) and the average intersection-over-union ratio (mIoU) by 10.39% (reaching 83.67%) compared to the best-performing models in the prior art, more accurately presenting the overall topological structure of vessels and providing doctors with more complete and accurate vascular information to assist in the comprehensive assessment of coronary artery disease.

[0117] This invention enhances model robustness: the strip convolution kernel design of the CMSC module and the processing method of high and low frequency features by the HLFAE module significantly improve the model's resistance to complex backgrounds and noise. When faced with interference such as low contrast, complex backgrounds, and motion artifacts in coronary angiography images, it can effectively suppress background interference and accurately focus on blood vessels and their course. Heatmaps generated using Class Activation Mapping (CAM) visualization technology show that the model of this invention has more concentrated and significant red activation areas in the blood vessel region and maintains high attention in the vessel edge region, while other models have more blue areas. This indicates that this invention has significant advantages in handling background interference and edge details, and can stably and accurately segment blood vessels in coronary angiography images with different imaging conditions and significant individual differences.

[0118] While the invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features described herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other described embodiments.

Claims

1. A coronary artery segmentation system integrating high- and low-frequency features, characterized in that, include: The cascaded multi-scale strip convolution module combines strip convolution operations of different scales to maintain attention to both global and local information through hierarchical feature extraction. The high- and low-frequency feature separation adaptive module separates the input feature map into high-frequency components and low-frequency components, performs adaptive enhancement on the high- and low-frequency components respectively, and finally fuses the enhanced high- and low-frequency components. The multi-scale strip high- and low-frequency enhancement module first inputs the input feature map into the cascaded multi-scale strip convolution module for feature extraction, then inputs the output feature map of the cascaded multi-scale strip convolution module into the high- and low-frequency feature separation adaptive module for high- and low-frequency separation and adaptive enhancement, and finally outputs the obtained feature map.

2. The coronary artery segmentation system integrating high and low frequency features according to claim 1, characterized in that, The cascaded multi-scale strip convolution module first performs strip convolution operations at different scales on the input feature map, then concatenates the convolution results at different scales, and then performs feature fusion through a 1×1 convolution kernel to obtain the final output feature map; the different scales include, but are not limited to, using 1×n and n×1 strip convolution kernels, where n is a different integer.

3. The coronary artery segmentation system integrating high and low frequency features according to claim 1, characterized in that, The adaptive enhancement in the high- and low-frequency feature separation adaptive module uses a combination of convolutional layers and activation functions to adaptively adjust the degree of enhancement based on the statistical information of the features. The high- and low-frequency feature separation adaptive module fuses the enhanced high-frequency and low-frequency components by concatenating the enhanced high-frequency and low-frequency components, and then performs feature fusion through a 1×1 convolutional kernel to obtain the final output feature map.

4. The coronary artery segmentation system integrating high and low frequency features according to claim 1, characterized in that, The high- and low-frequency feature separation adaptive module separates the input feature map into high-frequency components and low-frequency components, including: input feature map X∈R B×C×H×W Where B is the batch size, C is the number of channels, H and W are the spatial dimensions of the feature map; X is the feature map; X low Low-frequency characteristics; X high For high-frequency features; R is the set of real numbers; the input feature map first extracts low-frequency information through pooling, and then calculates high-frequency information using the difference. The formula for separating high- and low-frequency features is: X low =AvgPool(X) X high =X-Upsample(X low ) AvgPool(X) performs 2×2 average pooling to extract low-frequency information; Upsample(X) low The original size is restored by bilinear interpolation, and then subtracted from X to obtain high-frequency information.

5. The coronary artery segmentation system integrating high and low frequency features according to claim 4, characterized in that, The adaptive module for high- and low-frequency feature separation performs adaptive enhancement on the high- and low-frequency components respectively, including: Dilated convolution feature extraction is performed using feature_extraction(X) to form high-frequency enhanced features; the formula for high-frequency feature extraction is: X‘ high ==feature_extraction(X high ) The feature_extraction(X) function consists of four convolutional layers with different dilation rates, where d represents the dilation rate of the convolution. Expansion ratio; || represents channel splicing operation; X' high The high-frequency enhancement feature is the enhanced high-frequency characteristic; Represents the high-frequency feature X high Perform a 3x3 convolution operation with a kernel size of 4. This results in a feature map with four times the number of channels. Then, perform dimensionality reduction using 1x1 and 3x3 convolutions, as shown in the formula: X′ high =Conv 1×1 (Conv 3×3 (Concat(X′ high )) Low-frequency features are extracted through pooling operations, and weights are calculated using channel attention (SEAttention) to enhance them. The formula for enhancing low-frequency features is as follows: S low =σ(W2·ReLU(W1·AvgPool(X low ))+ W2·ReLU(W1·MaxPool(X low ))) X′ low ==X low ·S low SEAttention calculates global pooling and uses a multilayer perceptron to perform non-linear modeling on the globally pooled features to generate channel attention weights S. low W1 and W2 are the parameters of the fully connected layer, and σ is the sigmoid activation function; calculate the weighted low-frequency feature X. low MaxPool is the maximum pooling method; X′ low is the enhanced low-frequency feature; ReLU is the activation function.

6. The coronary artery segmentation system integrating high and low frequency features according to claim 4, characterized in that, The high- and low-frequency feature separation adaptive module finally fuses the enhanced high- and low-frequency components, including: introducing an adaptive mechanism to assign weights to different channels according to the importance of the features; the high- and low-frequency feature fusion formula and the self-attention module formula are as follows: X fusion =CBR(Concat(X′ high ,X′ low )) X final ==SEAttention(X fusion ) Y=X final +X Among them, X fusion This is a feature obtained by fusing high and low frequency characteristics; X final X is the final enhanced feature map after adjustment by the SE attention module; X is the original input, and Y is the final output feature map; `conact` is the concatenation, which concatenates two feature maps along the channel dimension; first, high-frequency and low-frequency features are concatenated, and then fused using the CBR process. The CBR process includes extracting local spatial features from the concatenated features through a 3×3 convolutional layer, performing batch normalization, and introducing a non-linear combination operation through the ReLU activation function; finally, it is adjusted again by the SE attention module; the final enhanced feature X is... final The input X is added back as a residual connection; the adaptive mechanism includes, but is not limited to, the SE attention mechanism.

7. A method for coronary artery segmentation that integrates high- and low-frequency features, characterized in that, Includes the following steps: Step S1: Input the input feature map into the cascaded multi-scale strip convolution step for feature extraction, and obtain the feature map output by the cascaded multi-scale strip convolution step. Step S2: Input the output feature map of the cascaded multi-scale strip convolution step into the high-low frequency feature separation adaptive step for high-low frequency separation and adaptive enhancement to obtain the final output feature map; The cascaded multi-scale strip convolution step involves cascading strip convolution operations of different scales to extract features at multiple levels, from global to local. The high- and low-frequency feature separation adaptive step separates the input feature map into high-frequency components and low-frequency components, adaptively enhances the high- and low-frequency components respectively, and finally fuses the enhanced high- and low-frequency components.

8. The coronary artery segmentation method integrating high- and low-frequency features according to claim 7, characterized in that, The cascaded multi-scale strip convolution step in step S1 includes: first, performing strip convolution operations at different scales on the input feature map; then, concatenating the convolution results at different scales; and finally, performing feature fusion through a 1×1 convolution kernel to obtain the final output feature map; the different scales include, but are not limited to, using 1×n and n×1 strip convolution kernels, where n is a different integer.

9. The coronary artery segmentation method integrating high- and low-frequency features according to claim 7, characterized in that, The adaptive step for high- and low-frequency feature separation in step S2 includes: Step S21: Separate high and low frequencies from the input features. The input feature map X∈R B×C×H×W Where B is the batch size, C is the number of channels, H and W are the spatial dimensions of the feature map; X is the feature map; X low Low-frequency characteristics; X high For high-frequency features; R is the set of real numbers; the input feature map first extracts low-frequency information through pooling, and then calculates high-frequency information using the difference. The formula for separating high- and low-frequency features is: X low =AvgPool(X) X high =X-Upsample(X low ) AvgPool(X) performs 2×2 average pooling to extract low-frequency information; Upsample(X) low The original size is recovered by bilinear interpolation, and then subtracted from X to obtain high-frequency information; Step S22: Adaptively enhance the high and low frequency components by using feature_extraction(X) for dilated convolution feature extraction to form high-frequency enhanced features; the formula for high-frequency feature extraction is: X‘ high ==feature_extraction(X high ) The feature_extraction(X) function consists of four convolutional layers with different dilation rates, where d represents the dilation rate of the convolution; || represents the channel concatenation operation; and X' high The high-frequency enhancement feature is the enhanced high-frequency characteristic; Represents the high-frequency feature X high Perform a 3x3 convolution operation with a kernel size of 4; the final result is a feature map with four times the number of channels, which is then subjected to 1x1 and 3x3 convolutions for dimensionality reduction, as shown in the formula: X′ high =Conv 1×1 (Conv 3×3 (Concat(X′ high )) Low-frequency features are extracted through pooling operations, and weights are calculated using channel attention (SEAttention) to enhance them. The formula for enhancing low-frequency features is as follows: S low =σ(W2·ReLU(W1·AvgPool(X low ))+ W2·ReLU(W1·MaxPool(X low ))) X‘ low ==X low ·S low SEAttention calculates global pooling and uses a multilayer perceptron to perform non-linear modeling on the globally pooled features to generate channel attention weights S. low W1 and W2 are the parameters of the fully connected layer, and σ is the sigmoid activation function; calculate the weighted low-frequency feature X. low MaxPool is the maximum pooling method; X′ low The signal is the enhanced low-frequency feature; ReLU is the activation function. Step S23: The enhanced high- and low-frequency components are fused, and an adaptive mechanism is introduced to assign weights to different channels according to the importance of the features; the high- and low-frequency feature fusion formula and the self-attention module formula are as follows: X fusion =CBR(Concat(X′ high ,X′ low )) X final ==SEAttention(X fusion ) Y=X final +X Among them, X fusion This is a feature obtained by fusing high and low frequency characteristics; X final X is the final enhanced feature map after adjustment by the SE attention module; X is the original input, and Y is the final output feature map; `conact` is the concatenation, which concatenates two feature maps along the channel dimension; first, high-frequency and low-frequency features are concatenated, and then fused using the CBR process. The CBR process includes extracting local spatial features from the concatenated features through a 3×3 convolutional layer, performing batch normalization, and introducing a non-linear combination operation through the ReLU activation function; finally, it is adjusted again by the SE attention module; the final enhanced feature X is... final The input X is added back as a residual connection; the adaptive mechanism includes, but is not limited to, the SE attention mechanism.