Fully-supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelet frequency domain
By employing a direction-aware dual-tree complex wavelet-based fully supervised and semi-supervised coronary artery segmentation method, the problems of low accuracy of distal vessels and scarcity of labeled data in coronary artery segmentation are solved, achieving efficient and accurate automated coronary artery segmentation and improving the robustness and diagnostic efficiency of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHEAST UNIV
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-21
AI Technical Summary
Existing coronary artery segmentation technologies suffer from low precision in distal vessel segmentation, insufficient utilization of multi-scale frequency information and directional cues, and a scarcity of high-quality labeled data. This makes traditional convolutional neural networks prone to breakage when segmenting oblique vessels, hindering efficient and accurate automated coronary artery segmentation.
We employ a fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelet transform. Through 28-direction high-frequency subband decomposition and multi-axis/single-axis dedicated branch architecture using 3D dual-tree complex wavelet transform, combined with the collaborative perception of frequency domain single-axis branch and multi-axis branch, we achieve explicit geometric modeling of the complex 3D spatial topology of coronary arteries and learn CCTA image features using a small amount of labeled data.
It significantly improves the voxel accuracy and topological continuity of coronary artery segmentation, reduces the dependence on high-quality labeled data, enhances the model's generalization ability and robustness, maintains the stability of segmentation results under different scanning devices and individual patient differences, reduces the burden of manual segmentation for doctors, and improves diagnostic efficiency.
Smart Images

Figure CN121904360A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a unified coronary artery segmentation method based on direction-aware dual-tree complex wavelet frequency domain, which is both fully supervised and semi-supervised, and belongs to the field of medical image processing and artificial intelligence technology. Background Technology
[0002] Currently, the application of artificial intelligence (AI) technology in the medical field has been widely explored and has achieved remarkable results. Computer vision, as an important branch of AI, occupies a key position in medical image processing. With the help of technologies such as deep learning, computers can automatically classify, identify, and analyze structures and lesions in medical images, providing auxiliary support for clinicians' diagnosis and treatment, and effectively improving the efficiency and accuracy of medical image diagnosis.
[0003] Cardiovascular disease is a prevalent and serious health problem worldwide. Coronary artery disease (CAD), as the most common and extremely dangerous type of cardiovascular disease, is experiencing an increasing prevalence due to population aging and lifestyle changes, placing enormous pressure on healthcare systems and the socio-economic sphere. CAD, or coronary atherosclerotic heart disease, is caused by atherosclerosis of the coronary arteries, leading to narrowing or even blockage of the blood vessels. This results in myocardial ischemia and hypoxia, and in severe cases, myocardial necrosis, ultimately developing into heart disease. CAD is one of the leading causes of myocardial infarction and sudden death; therefore, early and accurate diagnosis of CAD is crucial for slowing disease progression and ensuring patient safety.
[0004] Computed Tomography Angiography (CTA), as a mature imaging technology, is gradually becoming the main means of clinical diagnosis of coronary heart disease due to its significant advantages of speed, non-invasiveness, and high resolution. In the traditional cardiovascular disease diagnosis process, there is a high reliance on the physician's personal experience. Although coronary CTA can provide relatively rich imaging information, due to the limitations of visual information characteristics, it is difficult for physicians to intuitively and quickly determine whether there is stenosis in the coronary arteries. Precise vessel segmentation results can provide physicians with quantitative analysis information, helping them to more accurately assess the patient's condition and promote a more refined analysis of the coronary artery lesion status. With the ever-increasing volume of medical imaging data and the growing need for automated processing, developing automated, efficient, and accurate algorithms to assist in the analysis and diagnosis of medical images has become a crucial issue that urgently needs to be addressed in the medical field. In recent years, the National Health Commission has clearly stated that it will continue to promote the application of AI-assisted diagnostic technologies at the national level. The rapid development of deep learning (DL) technology, especially convolutional neural networks (CNNs), has brought about a breakthrough transformation in the field of medical image processing and analysis. Deep learning has demonstrated powerful feature learning and recognition capabilities in medical image segmentation, but its performance is often limited by the availability of large-scale labeled data. This contradiction is particularly prominent in coronary CT angiography (CCTA) segmentation tasks: on the one hand, coronary arteries have complex and tortuous branches with small diameters, exhibiting high geometric and topological heterogeneity in three-dimensional space. Traditional convolutional neural networks have significant limitations in representing their subtle orientations and multi-directional structures, especially prone to segmenting fragmented obliquely running vessels; on the other hand, high-quality CCTA labeled data is scarce and costly to acquire, severely restricting the training effect and generalization ability of fully supervised models. To address these issues, this invention proposes a breakthrough design at the anatomical accuracy level: through 28-directional high-frequency subband decomposition using three-dimensional dual-tree complex wavelet transform and a multi-axis / single-axis dedicated branch architecture, it achieves explicit geometric modeling of the complex three-dimensional topology of coronary arteries for the first time. This method fundamentally solves the key deficiency of traditional discrete wavelet transform, which only supports a limited number of directions (such as 0°, 45°, and 90°) and cannot effectively capture oblique vascular structures. For example, in common 30°-60° bifurcation regions such as the diagonal branch of the left anterior descending artery, traditional methods are prone to microvascular rupture due to insufficient directional representation. Through the collaborative sensing of single-axis and multi-axis branches in the frequency domain, the system simultaneously captures the main extension trend of the blood vessel and the details of the cross-sectional boundary, thereby achieving significant improvements in voxel-level accuracy and three-dimensional topological continuity, especially showing a breakthrough improvement in the segmentation of slender distal blood vessels.This method not only improves the accuracy of anatomical structure recognition in models with limited labeled data, but also provides a feasible technical path for the automated segmentation of coronary arteries from the laboratory to real-time clinical applications, which has important theoretical value and clinical translation prospects. Summary of the Invention
[0005] This invention addresses the problems existing in the prior art by providing a fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets. It aims to solve the problems of low segmentation accuracy of distal vessels and insufficient utilization of multi-scale frequency information and directional cues in existing coronary artery segmentation techniques. The method guides the network to fully utilize unlabeled data to learn CCTA image features with a small number of labels, thereby obtaining accurate coronary artery segmentation results and providing technical support for clinicians to carry out diagnostic work.
[0006] To address the aforementioned technical problems, this invention provides a fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets, comprising the following steps:
[0007] (1) Data Acquisition and Multidimensional Preprocessing: CCTA Imaging Data Acquisition, Storage, and Preprocessing: Acquire the patient's CCTA image DICOM data and store it in NIFTI format using 3D Slicer software; and perform the following preprocessing operations: Set appropriate window width and window level according to the coronary artery CT value range (usually 300-700 HU), and resample all images to a uniform spatial resolution (e.g., 0.4×0.4×0.4mm). 3 The image intensity is normalized to the range of [0,1]. The training data is augmented by random rotation (±15°), random scaling (0.9-1.1 times), and random translation (±10 pixels) to increase data diversity.
[0008] (2) Coronary Artery Annotation and Dataset Construction: Using the Materialise interactive medical image control system (MIMICS) interactive medical image processing software, based on prior knowledge of cardiac anatomy, regions of interest (ROIs) were extracted from the preprocessed CCTA images in step 1 to eliminate interference from irrelevant tissue areas. Experienced radiologists manually and accurately annotated the coronary artery structures within the extracted ROIs, including the left anterior descending artery, left circumflex artery, right coronary artery, and their main branches. Subsequently, a second specialist cross-validated and corrected the manual annotation results to ensure accuracy and reliability. The corrected final annotation results were stored in NIFTI format.
[0009] (3) Training of the deep learning XNet-DT segmentation model: The coronary artery dataset constructed in step (2) is randomly divided into four equal sub-datasets, and the model is trained and evaluated using a four-fold cross-validation method. During the training process, one sub-dataset is selected as the validation set, and the remaining three sub-datasets are used as the training set and input into the direction-aware dual-tree complex wavelet network (XNet-DT). The training process is repeated four times to ensure that each sub-dataset can be used as a validation set to participate in the model evaluation, thus ensuring the objectivity and accuracy of the model evaluation results. In each fold of the training process, only the coronary artery labels corresponding to 25% of the CCTA image data in the training set are input into the network for fully supervised training; the remaining 75% of the CCTA image data in the training set are input into the deep learning segmentation model in an unlabeled form, and semi-supervised training is performed by combining the wavelet domain perturbation consistency regularization strategy to fully explore the effective information in the unlabeled data and improve the segmentation performance of the model.
[0010] (4) Coronary artery segmentation prediction: The direction-aware dual-tree complex wavelet segmentation model trained in step (3) and whose performance was confirmed to meet the standards through four-fold cross-validation was applied to newly acquired CCTA image data; the coronary artery structure in the new CCTA image was automatically segmented through model inference, and accurate coronary artery segmentation results were output. The segmentation results can be directly used for subsequent clinical coronary artery stenosis detection and personalized treatment plan formulation. Specifically, the original CCTA image was decomposed into multiple scales and directions through dual-tree complex wavelet transform. The frequency domain features and spatial domain features obtained from the decomposition were respectively input into three functionally specialized branches for processing, so as to achieve balanced utilization of low-frequency / high-frequency information and direction-aware feature extraction. In the three-branch encoder-decoder architecture, the encoders of the frequency domain single-axis branch (U branch), frequency domain multi-axis branch (M branch), and spatial domain branch (S branch) all consist of 5 convolutional blocks and 4 pooling layers. The input features (the U / M branches are the dynamic fusion of the low-frequency subband after dual-tree complex wavelet transform (DT-CWT) decomposition and the corresponding high-frequency subband; the S branch is the original CCTA image) are extracted by convolutional blocks (Conv3d+InstanceNorm3d+ReLU) and downsampled by pooling layers (MaxPool3d) to obtain encoder feature maps at different scales. The encoder and decoder enhance the global feature capture capability through skip connections and feature concatenation. In each branch, asymmetric hierarchical embedding of spatial-frequency cross-domain fusion modules (SFFM) integrates the concatenated features to obtain cross-domain fused features, achieving cross-branch feature complementarity while avoiding gradient loss. The decoder part corresponds to the three branches. Each branch has a decoder, consisting of four convolutional blocks and four trilinear interpolation upsampling layers. Each decoder receives features from the corresponding branch encoder processed by the multi-scale global feature perception module, concatenates them with the fusion features passed from SFFM, and then progressively upsamples and integrates them. Finally, the output convolutional layer converts these features into three prediction results focusing on different domains (single-axis frequency domain, multi-axis frequency domain, and spatial domain). Labeled and unlabeled data input to the network are used to generate prediction results. The prediction results from labeled data are compared with the ground truth labels using a softmax function to calculate the supervised loss (Dice loss + cross-entropy loss + centerline boundary Dice loss). The prediction results from unlabeled data obtained through the S-branch and U / M-branch are used to generate unsupervised probabilities and pseudo-labels using softmax and argmax functions respectively, which are then used to calculate the cross-loss function and jointly optimize the model parameters.
[0011] The three-branch heterogeneous input architecture of the XNet-DT model in step (3) is characterized by: multi-scale and multi-directional decomposition of the original CCTA image through dual-tree complex wavelet transform, and inputting the frequency domain features and spatial domain features obtained from the decomposition into three functionally specialized branches for processing, thereby achieving balanced utilization of low-frequency / high-frequency information and direction-aware feature extraction. The multi-axis frequency domain branch takes the low-frequency features of dual-tree complex wavelet transform (DT-CWT) combined with multi-axis high-frequency features as input, and extracts global structure and multi-directional vascular details through an encoder; the single-axis frequency domain branch takes the low-frequency features of DT-CWT combined with single-axis high-frequency features as input, and focuses on details and edges in a single main direction; the spatial domain branch directly processes the original image, providing semantic context and large-scale structural features. Each branch encoder consists of 5 convolutional blocks and 4 pooling layers. A spatial-frequency cross-domain fusion module (SFFM) achieves cross-branch feature complementarity—the spatial domain branch and the multi-axis frequency domain branch are fused at deep features (layers 3-4), while the spatial domain branch and the single-axis frequency domain branch are fused at shallow features (layers 1-2). The fused features are then passed to each branch decoder via skip connections. Each decoder consists of 4 convolutional blocks and 4 upsampling layers, outputting the prediction result for the corresponding branch. The model is optimized using both supervised loss based on labeled data and cross-branch consistency loss based on unlabeled data, balancing high- and low-frequency information, enhancing direction awareness, and preserving feature differences between branches.
[0012] In step (3), the encoder inputs data to the network and first extracts feature map F1 from the data using a single convolutional block. The dimension of feature map F1 is defined as follows: W, H, and D represent the width, height, and depth of the image space, respectively, and C represents the number of channels. W = H = 112, D = 80, and C = 16. Subsequently, feature maps F2, F3, and F4 are extracted sequentially through a combination of convolutional blocks and pooling layers. With each feature extraction, the spatial dimension is reduced to half the size of the previous layer's feature map, and the channel dimension is expanded to four times the number of channels in the previous layer's feature map. Each convolutional block consists of two sets of convolutional layers with identical structures. Each set of convolutional layers sequentially includes a 3D convolution (Conv3d) with a kernel of (3,3,1), a 3D instance normalization (InstanceNorm3d), and an activation function (ReLU). The pooling layer uses a 3D max pooling (MaxPool3d) with a kernel of (2,2,2).
[0013] The dual-tree complex wavelet transform module in step (3) performs a first-level dual-tree complex wavelet decomposition on the original CCTA image I(x,y,z):
[0014]
[0015] Where C LLL This is a low-frequency approximate sub-band, representing the global structural information of the image. These are high-frequency detail subbands in six directions, corresponding to direction angles θ. i ∈{±15°,±45°,±75°}. Let the low-pass filter be h0[n] and the high-pass filter be h1[n]. After parallel filtering by tree A and tree B, the complex coefficients are obtained:
[0016]
[0017] Similarly, coefficients for other directions can be obtained. The first-level decomposition of the two-dimensional image produces eight sub-bands, but two of these sub-bands are merged due to symmetry, ultimately resulting in one low-frequency sub-band L and six high-frequency sub-bands. Based on the anatomical characteristics of the coronary arteries, the high-frequency sub-bands are further divided into uniaxial high-frequency sub-bands H. U and multi-axis high-frequency subband H M .
[0018] F S =conv 1·1·1 (I(x,y,z))
[0019] F U =ReLU(α·conv) 1·1·1 (H U )+AugPool(L))
[0020] F M =ReLU(β·conv) 1·1·1 (H M )+AugPool(L))
[0021] The feature processing flow of the Spatial Frequency Cross-Domain Fusion (SFFM) module in step (3) is as follows:
[0022]
[0023] Asymmetric fusion design is used in the spatial frequency cross-domain fusion module (SFFM) in each layer branch. The U branch and S branch are fused at the deep layer (layers 3 and 4) encoder features, and the M branch and S branch are fused at the shallow layer (layers 1 and 2) encoder features. During the fusion process, the branch features of the nth layer of the received encoder are used. First, the frequency domain features are converted into corresponding channel dimensions through 1×1×1 convolution. After ReLU activation, they are concatenated with the spatial domain convolution features to form fused features. Then, the concatenated features are integrated through a convolution-3×3×3 operation to obtain cross-domain fused features. This achieves cross-branch feature complementarity and avoids gradient loss.
[0024] The loss function in step (3) optimizes the model by minimizing the supervised loss on labeled images while simultaneously minimizing the triple output consistency loss on unlabeled images. Therefore, the total loss function is the same regardless of whether training is fully supervised or semi-supervised. The definition is as follows:
[0025]
[0026] in This indicates a loss of oversight. This represents the unsupervised loss. The parameter λ is a weight used to balance... and The contributions. Among them λ increases linearly with the number of training iterations. max Set to 0.5, where t represents the number of iterations in the current model training. max This represents the total number of training iterations for the model.
[0027]
[0028] in and Let represent the U, S, M branch predictions for the i-th image, and y i This represents the true label. Composite loss function. The specific formula is as follows:
[0029]
[0030] in and Let be the Dice loss for the k-branch and the Binary Cross-Entropy (CE) loss, respectively, where k∈{U,S,M}. α and β are hyperparameters for balancing the losses, both set to 0.5. The Centerline Boundary Dice (CBDice) loss takes into account the slender structure of blood vessels. Specifically designed for vessel segmentation, CBDice weights small structures and boundary regions to help the model focus on these important areas. These three loss functions work together to optimize the performance of XNet-DT. The Dice and Centerline Boundary Dice losses ensure stable global training of the model, while CBDice provides additional training guidance for the details of the coronary arteries.
[0031] Unsupervised loss Defined as:
[0032]
[0033] In the formula and This is achieved through cross-supervision loss:
[0034]
[0035] in and They represent respectively by and Generated pseudo-labels, Dice loss The loss is used to optimize the overlap between model predictions and ground truth labels, and its standard definition is the negative Dice similarity coefficient in binary segmentation tasks. For the predicted probability map p and the ground truth label y, this loss can be expressed as:
[0036]
[0037] Where V is the total number of voxels in a single image, p v ∈[0,1] represents the probability that the v-th v-th void is predicted to be a coronary artery. v ∈[0,1] is the true label of the vth v element (1 represents the coronary artery, 0 represents the background), and ∈>0 is a smoothing constant used to ensure numerical stability.
[0038] Compared to existing technologies, the advantages of this invention are as follows: This invention proposes a fully supervised-semi-supervised fusion segmentation scheme based on direction-aware dual-tree complex wavelet, effectively addressing the industry pain point of high cost and limited quantity of accurate CCTA data annotation in the medical imaging field. This invention creatively proposes a "fully supervised-semi-supervised" hybrid training framework and a "frequency domain perturbation consistency" regularization strategy. This scheme only requires supervised learning using 25% of the accurately labeled data in the training set to anchor the correct anatomical structures, while guiding the model to apply wavelet domain perturbations to the remaining 75% of unlabeled data, and forcing consistent predictions under different perturbations. This mechanism drives the model to mine essential vascular feature representations across the frequency domain from massive amounts of unlabeled data, rather than memorizing surface patterns from labeled data, thereby reducing the model's dependence on expensive expert annotations by approximately 75%, greatly alleviating the core bottleneck of scarce high-quality annotations in the medical imaging field, and significantly improving the model's generalization ability and feature robustness. Based on 28-directional high-frequency subband decomposition and multi-axis / single-axis dedicated branch design using 3D dual-tree complex wavelet transform, this method achieves explicit geometric modeling of the complex 3D spatial topology of coronary arteries for the first time. This design not only solves the fundamental problem of microvascular rupture caused by the inability of traditional three-directional discrete wavelet transform (DWT) to effectively characterize oblique vessels (such as diagonal branches bifurcating at 30°-60°), but also accurately captures the main extension direction and cross-sectional boundary details of the vessels through the synergy of frequency domain single-axis and multi-axis branches. This significantly improves voxel accuracy and topological continuity, especially for the segmentation of delicate distal vessels. Regarding model robustness and practicality, thanks to the inherent approximate translation invariance and superior direction selectivity of dual-tree complex wavelet transform, this method exhibits strong inherent resistance to common interference sources in CCTA images, such as noise, partial volume effects, and motion artifacts. The three-branch fusion architecture ensures the complementarity and balance of global semantics, multi-directional geometry, and single-axis detail information, maintaining stable and reliable segmentation results despite image heterogeneity caused by different scanning devices, protocols, and individual patient differences. Ultimately, this invention provides a highly standardized, end-to-end automated segmentation process. The trained model can quickly and accurately process new input CCTA images, and its output high-quality coronary artery 3D model can be directly used for clinical stenosis quantification, plaque nature assessment, and personalized interventional surgery planning. This significantly reduces the burden of repetitive manual segmentation for doctors, improves diagnostic consistency and efficiency, and has clear clinical translational value and broad prospects for industrial application. Attached Figure Description
[0039] Figure 1 This is a schematic diagram of the overall process of the present invention;
[0040] Figure 2 This is a schematic diagram of the deep learning segmentation network architecture of the present invention;
[0041] Figure 3A schematic diagram of the spatial-frequency domain cross-domain feature fusion module structure in a deep learning segmentation network;
[0042] Figure 4 A schematic diagram of the dual-tree complex wavelet transform frequency domain information extraction module in a deep learning segmentation network;
[0043] Figure 5 This is a schematic diagram of the segmentation result of the present invention. Detailed Implementation
[0044] The embodiments of the present invention will now be described in further detail with reference to the accompanying drawings.
[0045] Example 1: As Figure 1 The method for coronary artery segmentation based on direction-aware dual-tree complex wavelets, as shown, includes the following steps:
[0046] (1) Data acquisition and multi-dimensional preprocessing: DICOM format data of coronary CT angiography (CCTA) images of patients are acquired, and 3D Slicer software is used to convert the DICOM format data and store it in NIFTI format to complete the standardized storage of data.
[0047] The following preprocessing operations were performed on the converted NIFTI format CCTA images: ① Window width and level adjustment: Based on the CT value characteristics of the coronary arteries, the window width and level were set to parameters suitable for the coronary artery CT value range (300-700 HU) to enhance the contrast between blood vessels and surrounding tissues; ② Spatial resolution unification: All CCTA images were resampled to a preset uniform spatial resolution (e.g., 0.4×0.4×0.4mm) using a resampling algorithm. 3 ) Eliminate resolution differences caused by different acquisition devices; ③ Intensity normalization: Map the preprocessed image intensity values to the [0,1] interval to reduce the interference of intensity fluctuations on subsequent model training; ④ Data augmentation: Perform random data augmentation operations on the CCTA image data used for training, specifically including random rotation (rotation angle range of ±15°), random scaling (scaling factor range of 0.9-1.1 times), and random translation (translation pixel range of ±10 pixels), to improve the generalization ability of the model by expanding data diversity.
[0048] (2) Coronary Artery Annotation and Dataset Construction: Using the Materialise interactive medical image control system (MIMICS) interactive medical image processing software, based on prior knowledge of cardiac anatomy, regions of interest (ROIs) were extracted from the preprocessed CCTA images in step 1 to eliminate interference from irrelevant tissue areas. Experienced radiologists manually and accurately annotated the coronary artery structures within the extracted ROIs, including the left anterior descending artery, left circumflex artery, right coronary artery, and their main branches. Subsequently, a second specialist cross-validated and corrected the manual annotation results to ensure accuracy and reliability. The corrected final annotation results were stored in NIFTI format.
[0049] (3) Training of the deep learning XNet-DT segmentation model: The coronary artery dataset constructed in step (2) is randomly divided into four equal sub-datasets, and the model is trained and evaluated using a four-fold cross-validation method. During the training process, one sub-dataset is selected as the validation set, and the remaining three sub-datasets are used as the training set and input into the direction-aware dual-tree complex wavelet network (XNet-DT). The training process is repeated four times to ensure that each sub-dataset can be used as a validation set to participate in the model evaluation, thus ensuring the objectivity and accuracy of the model evaluation results. In each fold of the training process, only the coronary artery labels corresponding to 25% of the CCTA image data in the training set are input into the network for fully supervised training; the remaining 75% of the CCTA image data in the training set are input into the deep learning segmentation model in an unlabeled form, and semi-supervised training is performed by combining the wavelet domain perturbation consistency regularization strategy to fully explore the effective information in the unlabeled data and improve the segmentation performance of the model.
[0050] (4) Coronary artery segmentation prediction: such as Figure 2As shown, the deep learning segmentation model in step (3) of the direction-aware dual-tree complex wavelet-based fully supervised and semi-supervised coronary artery segmentation method has been trained and its performance has been confirmed by four-fold cross-validation. The direction-aware dual-tree complex wavelet segmentation model (XNet-DT) performs segmentation prediction on newly acquired CCTA image data: First, the original CCTA image is decomposed into multiple scales and directions through dual-tree complex wavelet transform (DT-CWT). The frequency domain features and spatial domain features obtained from the decomposition are respectively input into three functionally specialized branches (as shown in the figure, namely the frequency domain single-axis branch (U branch), the frequency domain multi-axis branch (M branch), and the spatial domain branch (S branch)) for processing, so as to achieve balanced utilization of low-frequency / high-frequency information and direction-aware feature extraction. In this three-branch encoder-decoder architecture, the encoder structures of the three branches are identical, each consisting of 5 convolutional blocks and 4 pooling layers. The input features differ between branches: the U / M branch inputs a dynamic fusion of the low-frequency subband after DT-CWT decomposition and the corresponding high-frequency subband; the S branch directly inputs the original CCTA image. The input features are extracted by convolutional blocks (Conv3d+InstanceNorm3d+ReLU) and downsampled by pooling layers (MaxPool3d) to obtain encoder feature maps at different scales. The encoder and decoder concatenate features through skip connections to enhance global feature capture capabilities. Simultaneously, each branch asymmetrically embeds a spatial frequency cross-domain fusion module (SFFM, as shown in the figure) to integrate the concatenated features, obtaining cross-domain fused features. This achieves cross-branch feature complementarity while avoiding gradient loss. The decoder section has one decoder for each of the three branches. Each decoder consists of four convolutional blocks and four trilinear interpolation upsampling layers. It receives features from the encoder of the corresponding branch after they have been processed by the multi-scale global feature perception module. These features are then concatenated with the fusion features transmitted by SFFM and gradually upsampled and integrated. Finally, the output convolutional layer converts the features into prediction results for three different domains of interest (single-axis frequency domain, multi-axis frequency domain, and spatial domain).
[0051] It should be noted that during each training iteration, only the coronary artery labels corresponding to 25% of the CCTA image data in the training set are input into the network for fully supervised training; the remaining 75% of the CCTA image data in the training set are input into the deep learning segmentation model in unlabeled form, and semi-supervised training is performed using a wavelet domain perturbation consistency regularization strategy to fully explore the effective information in the unlabeled data and improve the model's segmentation performance. The specific model optimization logic is as follows: the labeled and unlabeled data input into the network are used to predict the results. The prediction results of the labeled data are used to generate supervised probabilities through the softmax function, and the supervised loss (Dice loss + cross-entropy loss + centerline boundary Dice loss) is calculated by comparing them with the true labels; the prediction results of the unlabeled data are obtained through the S branch and U / M branch, and the unsupervised probabilities and pseudo-labels are generated by the softmax function and argmax function, respectively, and the cross-loss function is calculated. The two types of losses are used together to optimize the model parameters.
[0052] like Figure 2 As shown, the core advantage of the three-branch heterogeneous input architecture in the direction-aware dual-tree complex wavelet-based fully supervised and semi-supervised coronary artery segmentation method lies in the following: the multi-axis frequency domain branch uses low-frequency features of DT-CWT combined with multi-axis high-frequency features as input, extracting global structure and multi-directional vascular details through the encoder; the single-axis frequency domain branch uses low-frequency features of DT-CWT combined with single-axis high-frequency features as input, focusing on details and edges in a single principal direction; the spatial domain branch directly processes the original image, providing semantic context and large-scale structural features. Differentiated cross-branch fusion is achieved through SFFM—the spatial domain branch and the multi-axis frequency domain branch fuse at deep features (layers 3-4), and the spatial domain branch fuses at shallow features (layers 1-2). The fused features are passed to the decoders of each branch via skip connections, balancing high and low frequency information, enhancing direction awareness, and maintaining feature differences between branches.
[0053] Regarding the specific implementation details of the encoder: The input network data is first used to extract an initial feature map by a convolutional block. The dimension of this feature map is defined as C×W×H×D (where W, H, and D represent the width, height, and depth of the image space dimension, respectively, and C represents the number of channels). Subsequently, multi-scale feature maps are extracted sequentially through a combination of "convolutional block + pooling layer". After each feature extraction, the spatial dimension is reduced to 1 / 2 of the previous layer's feature map, and the channel dimension is expanded to 4 times the number of channels of the previous layer's feature map. Each convolutional block consists of two sets of convolutional layers with the same structure. Each set of convolutional layers sequentially includes a 3D convolution (Conv3d) with a kernel of (3,3,1), a 3D instance normalization (InstanceNorm3d), and an activation function (ReLU). The pooling layer is implemented using a 3D max pooling (MaxPool3d) with a kernel of (2,2,2). The segmentation model trained in step (3) is used to segment the coronary arteries of the CCTA image data.
[0054] like Figure 3 As shown, the dual-tree complex wavelet transform module in the direction-aware dual-tree complex wavelet-based fully supervised and semi-supervised coronary artery segmentation method performs first-level dual-tree complex wavelet decomposition on the original CCTA image I(x,y,z):
[0055]
[0056] Where C LLL This is a low-frequency approximate sub-band, representing the global structural information of the image. These are high-frequency detail subbands in six directions, corresponding to direction angles θ. i ∈{±15°,±45°,±75°}. Let the low-pass filter be h0[n] and the high-pass filter be h1[n]. After parallel filtering by tree A and tree B, the complex coefficients are obtained:
[0057]
[0058] Similarly, coefficients for other directions can be obtained. The first-level decomposition of the two-dimensional image produces eight sub-bands, but two of these sub-bands are merged due to symmetry, ultimately resulting in one low-frequency sub-band L and six high-frequency sub-bands. Based on the anatomical characteristics of the coronary arteries, the high-frequency sub-bands are further divided into uniaxial high-frequency sub-bands H. U and multi-axis high-frequency subband H M .
[0059] F S =conv 1·1·1 (I(x,y,z))
[0060] F U =ReLU(α·conv) 1·1·1 (H U )+AvgPool(L))
[0061] F M =ReLU(β·conv) 1·1·1 (H M )+AugPool(L))
[0062] like Figure 4 As shown, the feature processing flow of the spatial frequency cross-domain fusion module (SFFM) in the direction-aware dual-tree complex wavelet-based fully supervised and semi-supervised coronary artery segmentation method is as follows:
[0063]
[0064] Asymmetric fusion design is used in the spatial frequency cross-domain fusion module (SFFM) in each layer branch. The U branch and S branch are fused at the deep layer (layers 3 and 4) encoder features, and the M branch and S branch are fused at the shallow layer (layers 1 and 2) encoder features. During the fusion process, the branch features of the nth layer of the received encoder are used. First, the frequency domain features are converted into corresponding channel dimensions through 1×1×1 convolution. After ReLU activation, they are concatenated with the spatial domain convolution features to form fused features. The concatenated features are then integrated through a convolution-3×3×3 operation to obtain cross-domain fused features. This achieves cross-branch feature complementarity and avoids gradient loss.
[0065] A fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets is characterized by the following: the loss function of the model is optimized by minimizing the supervised loss on labeled images and simultaneously minimizing the triple output consistency loss on unlabeled images. Therefore, the total loss function is optimized in both fully supervised and semi-supervised training. The definition is as follows:
[0066]
[0067] in This indicates a loss of oversight. This represents the unsupervised loss. The parameter λ is a weight used to balance... and The contributions. Among them λ increases linearly with the number of training iterations. max Set to 0.5, where t represents the number of iterations in the current model training. max This represents the total number of training iterations for the model.
[0068] Monitoring losses Defined as:
[0069]
[0070] in and Let represent the U, S, M branch predictions for the i-th image, and y i This represents the true label. Composite loss function. The specific formula is as follows:
[0071]
[0072] in and Let be the Dice loss for the k-branch and the Binary Cross-Entropy (CE) loss, respectively, where k∈{U,S,M}. α and β are hyperparameters for balancing the losses, both set to 0.5. The Centerline Boundary Dice (CBDice) loss takes into account the slender structure of blood vessels. Specifically designed for vessel segmentation, CBDice weights small structures and boundary regions to help the model focus on these important areas. These three loss functions work together to optimize the performance of XNet-DT. The Dice and Centerline Boundary Dice losses ensure stable global training of the model, while CBDice provides additional training guidance for the details of the coronary arteries.
[0073] Unsupervised loss Defined as:
[0074]
[0075] In the formula and This is achieved through cross-supervision loss:
[0076]
[0077] in and They represent respectively by and Generated pseudo-labels, Dice loss The loss is used to optimize the overlap between model predictions and ground truth labels, and its standard definition is the negative Dice similarity coefficient in binary segmentation tasks. For the predicted probability map p and the ground truth label y, this loss can be expressed as:
[0078]
[0079] Where V is the total number of voxels in a single image, p v ∈[0,1] represents the probability that the v-th v-th void is predicted to be a coronary artery. v∈[0,1] is the true label of the vth v element (1 represents the coronary artery, 0 represents the background), and ∈>0 is a smoothing constant used to ensure numerical stability.
[0080] Input the CCTA image data for testing, and use the trained deep learning segmentation model as described in step (3) to segment the coronary arteries in the CCTA image data. The generated coronary artery segmentation results are as follows: Figure 5 As shown.
[0081] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention. Equivalent transformations or substitutions made based on the above technical solutions all fall within the scope of protection of the claims of the present invention.
Claims
1. A fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets, characterized in that, include: (1) CCTA image data acquisition, storage and preprocessing: Acquire the patient's CCTA image DICOM data and use 3DSlicer software to store the data in NIFTI format; and perform the following preprocessing operations: set appropriate window width and window level according to the range of coronary artery CT values (usually 300-700HU), resample all images to a uniform spatial resolution, normalize the image intensity to the range of [0,1], and perform random rotation, random scaling and random translation operations on the training data to enhance the data and increase the data diversity; (2) Coronary artery annotation and dataset construction: The Materialise interactive medical image control system (MIMICS) interactive medical image processing software was used to extract the region of interest from the preprocessed CCTA images based on prior knowledge of the cardiac region. The coronary artery structures, including the left anterior descending artery, the left circumflex artery, the right coronary artery and its main branches, were manually annotated by radiologists. The annotation results were verified and corrected by a second expert. The manual annotation results were stored in NIFTI format. The CCTA image data and coronary artery labels were combined to form the coronary artery dataset. (3) Training of deep learning XNet-DT segmentation model: The dataset constructed in step (2) is randomly divided into 4 parts and trained using four-fold cross-validation. Each part of the data is used as the validation set, and the remaining three parts of the data are used as the training set and input into the network. The network is trained four times so that each part of the data is validated, thereby ensuring the accuracy of the model evaluation. In the training of each fold, only 25% of the coronary artery labels corresponding to the CCTA image data are set to be input into the network, and the remaining data are input into the deep learning segmentation model in an unlabeled form for training. (4) Coronary artery segmentation prediction: Use the segmentation model trained in step (3) to perform coronary artery segmentation inference on new CCTA image data.
2. The fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets according to claim 1, characterized in that, The XNet-DT segmentation model in step (3) includes a dual-tree complex wavelet transform module, a three-branch encoder-decoder architecture, and a spatial-frequency cross-domain fusion module. The specific structure and functions are as follows: The original CCTA image is decomposed into multiple scales and directions using a dual-tree complex wavelet transform. The frequency domain features and spatial domain features obtained from the decomposition are input into three specialized branches for processing, achieving balanced utilization of low-frequency / high-frequency information and direction-aware feature extraction. In the three-branch encoder-decoder architecture, the encoders of the frequency domain single-axis branch (U branch), frequency domain multi-axis branch (M branch), and spatial domain branch (S branch) all consist of 5 convolutional blocks and 4 pooling layers. The input features ( The U / M branch is a dynamic fusion of the low-frequency subband after dual-tree complex wavelet transform (DT-CWT) decomposition and the corresponding high-frequency subband. The S branch is the original CCTA image. After feature extraction by convolutional blocks (Conv3d+InstanceNorm3d+ReLU) and downsampling by pooling layers (MaxPool3d), encoder feature maps of different scales are obtained. The encoder and decoder enhance the global feature capture capability by splicing features through skip connections. In each branch, asymmetric hierarchical embedding of spatial frequency cross-domain fusion modules (SFFM) integrates information of spliced features to obtain cross-domain fused features, which not only achieves cross-branch feature complementarity but also avoids gradient loss. The decoder section has one decoder for each of the three branches. Each decoder consists of four convolutional blocks and four trilinear interpolation upsampling layers. It receives features from the encoder of the corresponding branch after processing by the multi-scale global feature perception module, concatenates them with the fusion features passed by SFFM, and then gradually upsamples and integrates them. Finally, it is converted into prediction results for three different domains of interest through the output convolutional layer: single-axis frequency domain, multi-axis frequency domain, and spatial domain. The labeled and unlabeled data input to the network are used to predict the results. The prediction results obtained from the labeled data are compared with the real labels by generating supervision probabilities through the softmax function to calculate the supervision loss (Dice loss + cross-entropy loss + centerline boundary Dice loss). The prediction results obtained from the unlabeled data through the S branch and U / M branch are used to generate unsupervised probabilities and pseudo-labels through the softmax function and argmax function respectively to calculate the cross loss function and jointly optimize the model parameters.
3. The fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets according to claim 2, characterized in that: The XNet-DT model employs a three-branch heterogeneous input architecture. It performs multi-scale, multi-directional decomposition of the original CCTA image using dual-tree complex wavelet transform (DT-CWT). The resulting frequency domain and spatial domain features are then input into three functionally specialized branches for processing, achieving balanced utilization of low-frequency / high-frequency information and direction-aware feature extraction. The multi-axis frequency domain branch uses low-frequency features from DT-CWT combined with multi-axis high-frequency features as input, extracting global structure and multi-directional vascular details through an encoder. The single-axis frequency domain branch uses low-frequency features from DT-CWT combined with single-axis high-frequency features... The input focuses on details and edges in a single main direction; the spatial domain branch directly processes the original image, providing semantic context and large-scale structural features. Each branch encoder consists of 5 convolutional blocks and 4 pooling layers. Cross-branch feature complementarity is achieved through the spatial-frequency cross-domain fusion module (SFFM)—the spatial domain branch and the multi-axis frequency domain branch are fused in deep features (3-4 layers), and the spatial domain branch and the single-axis frequency domain branch are fused in shallow features (1-2 layers). The fused features are passed to each branch decoder via skip connections; each decoder consists of 4 convolutional blocks and 4 upsampling layers, and outputs the prediction results of the corresponding branch.
4. The fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets according to claim 3, characterized in that: In the encoder, the input network data is first processed by a convolutional block to extract feature map F1. The dimension of feature map F1 is defined as follows: W, H, and D represent the width, height, and depth of the image space dimension, respectively, and C represents the number of channels. W = H = 112, D = 80, and C = 16. Subsequently, feature maps F2, F3, and F4 are extracted sequentially through a combination of convolutional blocks and pooling layers. With each feature extraction, the spatial dimension is reduced to half of the previous layer's feature map, and the channel dimension is expanded to four times the number of channels in the previous layer's feature map. Each convolutional block consists of two sets of convolutional layers with the same structure. Each set of convolutional layers sequentially includes a three-dimensional convolution (Conv3d) with a kernel of (3,3,1), a three-dimensional instance normalization (InstanceNorm3d), and an activation function (ReLU). The pooling layer uses a three-dimensional max pooling (MaxPool3d) with a kernel of (2,2,2).
5. The fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets according to claim 4, characterized in that: The dual-tree complex wavelet transform module performs a first-level dual-tree complex wavelet decomposition on the original CCTA image I(x,y,z): Where C LLL This is a low-frequency approximate sub-band, representing the global structural information of the image. These are high-frequency detail subbands in six directions, corresponding to direction angles θ. i For the region ∈{±15°,±45°,±75°}, let the low-pass filter be h0[n] and the high-pass filter be h1[n]. After parallel filtering by tree A and tree B, we obtain the complex coefficients: The first-level decomposition of the two-dimensional image generated a total of 8 subbands. Two of these subbands were merged due to symmetry, resulting in one low-frequency subband L and six high-frequency directional subbands. Based on the anatomical characteristics of the coronary arteries, the high-frequency subband was further divided into a uniaxial high-frequency subband H. U and multi-axis high-frequency subband H M , F S =conv 1·1·1 (I(x,y,z)) F U =ReLU(α·conv 1·1·1 (H U )+AvgPool(L)) F M =ReLU(β·conv 1·1·1 (H M )+AvgPool(L))。 6. The fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets according to claim 5, characterized in that: The spatial frequency cross-domain fusion module (SFFM) processes features as follows: Asymmetric fusion design is used in the spatial frequency cross-domain fusion module (SFFM) in each layer branch. The U branch and S branch are fused at the deep layer (layers 3 and 4) encoder features, and the M branch and S branch are fused at the shallow layer (layers 1 and 2) encoder features. During the fusion process, the branch features of the nth layer of the encoder are received. First, the frequency domain features are converted into corresponding channel dimensions through 1×1×1 convolution. After ReLU activation, they are concatenated with the spatial domain convolution features to form fused features. Then, the concatenated features are integrated through a convolution-3×3×3 operation to obtain cross-domain fused features. This achieves cross-branch feature complementarity and avoids gradient loss.
7. The fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets according to claim 2, characterized in that: The loss function, which the model optimizes by minimizing the supervised loss on labeled images and simultaneously minimizing the triple output consistency loss on unlabeled images, is the total loss function in fully supervised or semi-supervised training. The definition is as follows: in This indicates a loss of oversight. This represents the unsupervised loss, where the parameter λ is a weight used to balance... and The contribution of λ increases linearly with the number of training iterations, and the supervised loss... Defined as: in and Let represent the U, S, M branch predictions for the i-th image, and y i The composite loss function represents the true label. The specific formula is as follows: in and Let be the Dice loss for the k-branch and the Binary Cross-Entropy (CE) loss, respectively, where k∈{U,S,M}, and α and β are hyperparameters of the balancing loss, both set to 0.
5. Unsupervised loss Defined as: In the formula and This is achieved through cross-supervision loss: in and They represent respectively by and Generated pseudo-labels, Dice loss Used to optimize the overlap between model predictions and ground truth labels, its standard definition is the negative Dice similarity coefficient in binary segmentation tasks. For the predicted probability map p and the ground truth label y, this loss is expressed as: Where V is the total number of voxels in a single image, p v ∈[0,1] represents the probability that the v-th v-th void is predicted to be a coronary artery, y v ∈[0,1] is the true label of the vth v element (1 represents the coronary artery, 0 represents the background), and ∈>0 is a smoothing constant used to ensure numerical stability.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelet as described in any one of claims 1 to 7.
9. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by the processor, the computer instructions implement the fully supervised and semi-supervised coronary artery segmentation method based on direction-aware dual-tree complex wavelets as described in any one of claims 1-7.