Three-dimensional coronary artery image segmentation method and system

Through the multi-stage segmentation method, the DS-Unet model and the technology of extracting image blocks in the vascular centerline is solved, and the problem of insufficient accuracy and computational complexity when processing complex coronary images in the prior art is achieved, and efficient and accurate three-dimensional coronary image segmentation is achieved.

CN119963576APending Publication Date: 2025-05-09FUZHOU UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510051410.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing deep learning-based medical image segmentation method is difficult to achieve ideal accuracy when processing tiny blood vessels such as complex coronary arteries, and due to hardware resource limitations, it is impossible to directly process complete three-dimensional coronary artery image data.

Method used

A multi-stage segmentation method is adopted, first pre-processing and coarse segmentation of the three-dimensional coronary CT images, and the blood vessel area is initially segmented through the DS-Unet model, and then the blood vessel center line is extracted, and the image blocks are extracted along the center line for subdivision, and finally a final segmented image matching the original input size is combined to generate.

Benefits of technology

It improves the efficiency and accuracy of three-dimensional coronary artery image segmentation, can capture blood vessel details more accurately, reduces computational complexity and memory overhead, and ensures the continuity and accuracy of blood vessel segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963576A_ABST
    Figure CN119963576A_ABST
Patent Text Reader

Abstract

The invention relates to a three-dimensional coronary artery image segmentation method and system, and the method comprises the steps: S1, carrying out the preprocessing of a three-dimensional coronary artery CT image, including window width and window level adjustment, edge detection and bilateral filtering, so as to reduce the noise and enhance the image contrast; s2, a DS-Unet model is adopted to carry out coarse segmentation on the preprocessed image, a blood vessel region is preliminarily segmented, the DS-Unet model is based on a UNet model and integrates a dense residual block and a sliding window self-attention mechanism, the overall structure and boundary of a blood vessel can be effectively captured, and it is ensured that the blood vessel is accurately recognized; s3, extracting a blood vessel center line from the blood vessel region subjected to coarse segmentation, and extracting an image block with a set size along the blood vessel center line to form an image block data set; and S4, performing fine segmentation on the newly extracted image block data set by adopting another DS-Unet model, and combining segmented results to generate a final segmented image matched with the original input size. The method and the system are beneficial to improving the efficiency and the accuracy of three-dimensional coronary artery image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and in particular to a three-dimensional coronary artery image segmentation method and system. Background Art

[0002] Coronary artery disease has always been a major threat to global health. As the most common form of coronary artery disease, accurate coronary artery segmentation and 3D reconstruction are essential for the diagnosis and treatment of cardiovascular disease. Computed tomography coronary angiography (CTCA) plays a key role in clinical applications and research. However, coronary CT angiography (CCTA) images often face common medical imaging challenges, including unbalanced foreground and background distribution, small target area, and inconsistent image quality, which makes computer-aided analysis of coronary artery images very important.

[0003] Although deep learning-based image segmentation methods have achieved remarkable results in medical image processing, they still face many challenges when dealing with complex structures, especially the segmentation of small structures such as blood vessels. In recent years, the rapid development of deep learning technology has opened up new horizons for medical image processing, especially in medical image segmentation. Compared with traditional methods, deep learning methods have significantly improved segmentation accuracy. Convolutional neural networks (CNNs) have been widely used in medical imaging, and classic networks such as U-Net have set new standards for biomedical image segmentation. In order to cope with the complexity of three-dimensional medical images, researchers have proposed volume-based network architectures such as 3D U-Net and V-Net, which make full use of the richness of volume data and inter-layer information. In addition, networks based on 3D U-Net have been improved in many aspects, adding residual connections, attention mechanisms, and dense connections to further improve segmentation accuracy.

[0004] Although CNN-based networks have shown good performance in medical image segmentation, it is difficult to capture enough detailed information when dealing with complex small blood vessels such as coronary arteries. In order to overcome these challenges, Transformer-based models have begun to attract widespread attention. Its powerful global modeling capabilities and parallel computing advantages have also shown great potential in medical image segmentation. VisionTransformer uses self-attention mechanisms and fully connected layers to process image data, which can model global information in the image without relying on traditional convolution operations, achieving excellent performance.

[0005] In order to solve the computational complexity problem when processing large-size images, Swin Transformer, as an improved version of Transformer, adopts a layered window mechanism to reduce computational and memory overhead, making it more suitable for processing large-size images. At the same time, the combination of Transformer and CNN has become a current research hotspot. By applying Transformer to the encoder part for global modeling and using CNN for local feature recovery in the decoder part, this hybrid model gives full play to the advantages of both and significantly improves segmentation accuracy and efficiency.

[0006] However, in practical applications, due to hardware resource limitations (such as insufficient graphics card performance), it is impossible to directly process the complete three-dimensional coronary artery image data. In this case, it is particularly important to adopt a multi-stage segmentation method based on the coronary centerline. This method can effectively reduce the memory overhead required for each processing by gradually segmenting the blood vessels in stages, thereby avoiding the computational bottleneck caused by directly processing high-dimensional three-dimensional data. At the same time, the centerline-based segmentation strategy can maintain the continuity of the vascular structure and ensure the integrity and accuracy of the vascular segmentation.

[0007] In the field of coronary artery image segmentation, although many deep learning-based methods have made some progress, they still face many challenges, especially when dealing with complex vascular structures. Although the existing deep learning-based image segmentation methods can handle most cases, it is still difficult to achieve ideal accuracy when dealing with small blood vessels or coronary arteries. Therefore, studying how to design an efficient and accurate coronary artery medical image segmentation system has become a technical problem that needs to be solved urgently. Summary of the invention

[0008] The object of the present invention is to provide a three-dimensional coronary artery image segmentation method and system, which are conducive to improving the efficiency and accuracy of three-dimensional coronary artery image segmentation.

[0009] In order to achieve the above object, the technical solution adopted by the present invention is: a three-dimensional coronary artery image segmentation method, comprising the following steps:

[0010] Step S1: preprocessing the three-dimensional coronary artery CT image, including window width and window position adjustment, edge detection and bilateral filtering, to reduce noise and enhance image contrast;

[0011] Step S2: The preprocessed image is roughly segmented using the DS-Unet model to preliminarily segment the blood vessel area. The DS-Unet model is based on the UNet model and integrates a dense residual block and a sliding window self-attention mechanism, which can effectively capture the overall structure and boundaries of the blood vessels and ensure that the blood vessels are accurately identified;

[0012] Step S3: extracting the blood vessel centerline from the roughly segmented blood vessel region, and extracting an image block of a set size along the blood vessel centerline to form an image block data set;

[0013] Step S4: Use another DS-Unet model to finely segment the newly extracted image block dataset, and combine the segmented results to generate a final segmented image that matches the original input size.

[0014] Furthermore, the step S1 includes the following steps:

[0015] Step S11: Load the three-dimensional coronary artery CT image and adjust the window width and window position of the image; by setting the window width and window position, adjust the pixel value of the image to a range suitable for subsequent processing:

[0016]

[0017] Where T(l) is the adjusted pixel value, l is the original image pixel value, and W level is the window level, W width is the window width;

[0018] Step S12: Use the Sobel filter method to calculate the gradient of each slice image, so as to highlight the edge of the image and provide a basis for subsequent denoising processing; extract the gradient amplitude of each pixel in the original three-dimensional data, and extract the boundary area according to the size of the gradient amplitude:

[0019]

[0020] Among them, edges(x,y) is the edge strength at the pixel (x,y) in the image, gradient x (x,y), gradient y (x, y) are the horizontal and vertical gradient values ​​at the pixel (x, y) respectively; I(x+i, y+j) is the grayscale value of the pixel (x+i, y+j) in the image, where i = -1, 0, 1, j = -1, 0, 1, G x (i+1,j+1),G y (i+1, j+1) is the convolution kernel of the Sobel filter in the horizontal and vertical directions at the pixel (x+i, y+j);

[0021] Step S13: Use bilateral filtering to remove noise inside the extracted boundary, taking into account spatial similarity and grayscale similarity, while removing noise and retaining the edge of the image:

[0022]

[0023] Among them, I(x,y) is the gray value of the pixel (x,y) in the image before filtering, I new (x, y) is the gray value of pixel (x, y) in the filtered image, σ s is the standard deviation in the spatial domain, σ r is the standard deviation of pixel values;

[0024] Step S14: normalize the denoised image to ensure that the pixel values ​​are in the range of [0,1], and adjust the size to the same size to ensure that the image data has a uniform scale, thereby improving the calculation efficiency and accuracy:

[0025]

[0026] Among them, I(x,y) is the gray value of the pixel (x,y) in the image before normalization, I noemalized (x,y) is the grayscale value of pixel (x,y) in the normalized image.

[0027] Furthermore, based on the UNet model, the DS-Unet model improves the convolution before downsampling into a dense residual convolution, and adds a sliding window self-attention mechanism as a jump connection in the downsampled and upsampled data to supplement the detail features. The overall structure of the DS-Unet model consists of an encoder and a decoder; in the encoder part, the spatial dimension of the feature map is gradually reduced through multiple downsamplings, while the number of channels is increased; before each downsampling, a dense residual block is used to process the data, and the dense residual block realizes feature reuse and gradient flow by connecting the outputs of all previous layers; in the decoder part, the spatial dimension of the feature map is gradually restored through multiple upsamplings, and the size of the feature map is restored to the same as the input; the jump connection in the decoder splices the feature map of the corresponding layer of the encoder with the feature map of the current layer of the decoder to retain the detail information.

[0028] Furthermore, the dense residual block comprises:

[0029] First convolutional layer: In the first part of the dense residual block, the input image is processed by a 1×1×1 convolution to reduce the number of channels of the input image, and is normalized by a BN layer to ensure the stability of the values ​​during the convolution process. After normalization, the ReLU activation function is applied to increase the nonlinear transformation. The output of this process is the first intermediate feature map, which is further used in subsequent convolution operations.

[0030] The second convolutional layer receives the output from the first convolutional layer and processes it through a 3×3×3 convolution. After the convolution, the BN layer is also applied to normalize the output features, and a nonlinear transformation is introduced through the ReLU activation function to further extract the deep features of the image and ensure that the spatial dimension is not lost.

[0031] A residual connection is performed based on the output result of the second convolutional layer, that is, the original input is adjusted through a 1×1 convolution to make it consistent with the number of channels of the second output image and then added to avoid the gradient vanishing problem;

[0032] The third convolutional layer: receives the image after residual connection, and continues to perform 3×3×3 convolution processing on the features through convolution to further enhance the feature expression ability; finally, by adding the input image, the output of the first convolutional layer and the output of the third convolutional layer, a deep feature map integrating multi-level features is obtained;

[0033] The operation of the dense residual block is expressed as follows:

[0034] Out(x)=H2(x+H2(H1(x)))+x+H2(x)+H2(H1(X))

[0035] Among them, x represents the input feature map, H1(x) represents the BN-ReLU-Conv1×1×1 operation, and H2(x) represents the BN-ReLU-Conv3×3×3 operation.

[0036] Furthermore, the implementation method of the skip connection between downsampling and upsampling is as follows: the deep feature map after dense residual convolution is divided into blocks through the Patch Embedding layer and the Swin Transformer encoder is used to extract features for each image block, and the feature information in each image block is fused through the self-attention mechanism to adaptively capture context information. The result extracted by the Swin Transformer encoder is spliced ​​with the upsampling result to realize the skip connection;

[0037] The implementation method of the Patch Embedding layer operation is:

[0038] First, the input high-level semantic feature map is cut into several small blocks of preset size, each of which is mapped to a vector space through a convolution operation to generate the initial patch embedding matrix X∈R (N2)×d , where N represents the number of small blocks after the image is cut, and d represents the feature vector dimension of each small block; specifically, through a 3D convolution layer, the convolution kernel size and stride are set to the small block size, so as to extract the features of each small block and map them to a new vector space, and the number of channels of the output feature map is d, that is, the feature dimension of each small block; in order to improve the stability of the model, the feature map after the convolution operation is normalized; the LN layer is used to standardize the convolution result, enhance the numerical stability during the training process, and accelerate the convergence of the model; finally, a processed feature map is obtained;

[0039] The Swin Transformer encoder appears in a two-stage serial structure; in the first stage, window-based multi-head self-attention (W-MSA) is used, and then in the second stage, shifted window-based multi-head self-attention (SW-MSA) is used; the self-attention calculation method is selected according to whether the current Swin Transformer encoder is odd or even:

[0040]

[0041] Where W-MSA and SW-MSA represent conventional multi-head self-attention modules and window-based multi-head self-attention blocks, respectively; z' and z'+1 represent the outputs of W-MSA and SW-MSA, respectively; LN and MLP represent layer normalization and multi-layer perceptron, respectively;

[0042] In order to efficiently calculate the shift window mechanism, a three-dimensional cyclic shift is adopted; based on the following self-attention formula:

[0043]

[0044] Among them, Q, K, V represent query, key and value respectively; d represents the size of query and key; B is the relative position encoding; then, the features processed by the attention mechanism are nonlinearly transformed and enhanced through a multi-layer perceptron (MLP); in the entire network process, dense residual connections and self-attention masks are combined to ensure the stability and effectiveness of the model.

[0045] Furthermore, in step S2, the implementation method of the coarse segmentation is as follows: extracting features from the input feature map through convolution layers in multiple dense residual blocks, and cutting and mapping the feature map to the vector space through Patch Embedding after each convolution layer; then, capturing the spatial dependency through the self-attention operation of the Swin Transformer module; in the downsampling process, the feature map is gradually reduced in dimension through maximum pooling, and then the feature map is restored using convolution upsampling, and spliced ​​with deeper features through jump connections to fuse information of different scales; finally, a coarse segmentation result is obtained.

[0046] Furthermore, step S3 includes the following steps:

[0047] Step S31: After the rough segmentation, a binary dilation algorithm is applied to dilate the blood vessels to enhance the connectivity in the segmentation results:

[0048]

[0049] Among them, A is the original binary image, represents the dilation operation, which is defined as performing the local maximization of the structure element for each pixel in the image, where B is the structure element;

[0050] Step S32: After dilating the blood vessels, the two largest connected domains are extracted as the left and right coronary arteries using the connected domain analysis, and the skeleton of the blood vessels is extracted as the centerline using the surface refinement algorithm:

[0051]

[0052] Among them, A is the input blood vessel image, B is the structural element, represents the erosion operation, which shrinks each target area in the image; n represents the number of iterations, and the skeleton of the image is gradually extracted through multiple erosion operations;

[0053] Step S33: based on the extracted blood vessel centerline, a fixed-size image block is cut out around the blood vessel centerline; the center of each image block is a point on the blood vessel centerline; for each point on the centerline, the boundary of the image block is determined by the size of the image block.

[0054] Furthermore, in step S4, the implementation method of the fine segmentation is: save each cropped image block and record the cropping position, and the center line provides key information of the cropping area; input the image block into another DS-Unet model for fine segmentation to obtain a fine segmentation result, extract the image block from the stored small block file by reading the recorded position of each small block, and merge them into the original image according to the recorded coordinate position; each small block is merged with the corresponding area of ​​the original image by taking the maximum value to ensure that the content of the small block covers the original area; finally, the image is binarized to generate the final segmentation result.

[0055] Furthermore, in the deep learning training process, firstly, three-dimensional coronary artery CT image data for model training is obtained, and labels are annotated on the images, that is, blood vessels are segmented in the images, and then the original images and the labeled images are preprocessed to form a training data set; then, a DS-Unet model for coarse segmentation is trained with the training data set, and after the training is completed, the DS-Unet model can be used for coarse segmentation; then, the images in the training data set are input into the DS-Unet model for coarse segmentation, and then, through the processing of step S3, image blocks centered on the centerline of the blood vessels are cropped from the images, and then an image block data set consisting of unlabeled image blocks and labeled image blocks is constructed; and then another DS-Unet model for fine segmentation is trained with the obtained image block data set.

[0056] The present invention also provides a three-dimensional coronary artery image segmentation system, comprising a memory, a processor, and computer program instructions stored in the memory and executable by the processor. When the processor executes the computer program instructions, the above method can be implemented.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] (1) The dense residual module is used to extract features from the original image. The dense residual module enhances the network's ability to recognize small blood vessels and complex structures by effectively integrating feature information at different levels. This module can accurately capture the details of blood vessels through multi-level feature extraction and information fusion, especially when dealing with complex anatomical structures. The dense residual structure improves the network's learning ability, optimizes the network's training process, and helps the network converge faster, thereby ensuring improved segmentation accuracy.

[0059] (2) The SWIN Transformer module is used to capture long-range dependencies and multi-scale contextual information in images. The SWIN Transformer reduces computational complexity through the local window self-attention mechanism, making it more efficient when processing large-size images. This module not only improves segmentation accuracy, but also further optimizes computational efficiency, especially in high-resolution images and detailed blood vessel segmentation.

[0060] (3) By extracting the coronary artery centerline and segmenting the image based on the centerline, the image segmentation task is decomposed into smaller regions for processing, thereby avoiding the memory consumption and computational complexity caused by directly processing the entire three-dimensional image. This staged segmentation method effectively reduces the computational burden while ensuring the continuity and accuracy of vascular segmentation, and is particularly suitable for the segmentation of complex and small blood vessels. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 A flowchart of a method implementation of an embodiment of the present invention;

[0062] Figure 2 It is a cross-sectional diagram of data before and after preprocessing in an embodiment of the present invention;

[0063] Figure 3 A network structure diagram of a DS-Unet model in an embodiment of the present invention;

[0064] Figure 4 is a structural diagram of a dense residual block in an embodiment of the present invention;

[0065] Figure 5 This is a result diagram of segmenting the same case using different methods in an embodiment of the present invention. DETAILED DESCRIPTION

[0066] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0067] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0068] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0069] like Figure 1 As shown, this embodiment provides a three-dimensional coronary artery image segmentation method, comprising the following steps:

[0070] Step S1: Preprocess the three-dimensional coronary artery CT image, including window width and window position adjustment, edge detection and bilateral filtering, to reduce noise and enhance image contrast. Specifically, the following steps are included:

[0071] Step S11: Load the three-dimensional coronary artery CT image and adjust the window width and window position of the image; by setting the window width and window position, adjust the pixel value of the image to a range suitable for subsequent processing:

[0072]

[0073] Where T(l) is the adjusted pixel value, l is the original image pixel value, and W level is the window level, W width It is the window width.

[0074] Step S12: Use the Sobel filter method to calculate the gradient of each slice image, so as to highlight the edge of the image and provide a basis for subsequent denoising processing; extract the gradient amplitude of each pixel in the original three-dimensional data, and extract the boundary area according to the size of the gradient amplitude:

[0075]

[0076] Among them, edges(x,y) is the edge strength at the pixel (x,y) in the image, gradient x (x,y), gradient y (x, y) are the horizontal and vertical gradient values ​​at the pixel (x, y) respectively; I(x+i, y+j) is the grayscale value of the pixel (x+i, y+j) in the image, where i = -1, 0, 1, j = -1, 0, 1, G x (i+1,j+1),Gy (i+1, j+1) is the horizontal and vertical convolution kernel of the Sobel filter at pixel (x+i, y+j).

[0077] Step S13: Use bilateral filtering to remove noise inside the extracted boundary, taking into account spatial similarity and grayscale similarity, while removing noise and retaining the edge of the image:

[0078]

[0079] Among them, I(x,y) is the gray value of the pixel (x,y) in the image before filtering, I new (x, y) is the gray value of pixel (x, y) in the filtered image, σ s is the standard deviation in the spatial domain, σ r is the standard deviation of the pixel values.

[0080] Step S14: normalize the denoised image to ensure that the pixel values ​​are in the range of [0,1] to facilitate the input of subsequent algorithms, and adjust the size to the same size to ensure that the image data has a uniform scale, thereby improving calculation efficiency and accuracy:

[0081]

[0082] Among them, I(x,y) is the gray value of the pixel (x,y) in the image before normalization, I noemalized (x,y) is the grayscale value of pixel (x,y) in the normalized image.

[0083] Step S2: The DS-Unet model is used to roughly segment the preprocessed image and preliminarily segment the blood vessel area. The DS-Unet model is based on the UNet model and integrates a dense residual block (Dense Residual Block) and a sliding window self-attention mechanism (Swin Transformer Block), which can effectively capture the overall structure and boundaries of the blood vessels and ensure that the blood vessels are accurately identified.

[0084] The DS-Unet model is based on the UNet model. The convolution before downsampling is improved into dense residual convolution, and the sliding window self-attention mechanism is added to the downsampled and upsampled data as a jump connection to supplement the detail features. The overall structure of the DS-Unet model consists of an encoder and a decoder. In the encoder part, the spatial dimension of the feature map is gradually reduced through multiple downsampling, while the number of channels is increased. Before each downsampling, a dense residual block is used to process the data. The dense residual block realizes feature reuse and gradient flow by connecting the outputs of all previous layers. In the decoder part, the spatial dimension of the feature map is gradually restored through multiple upsampling, and the size of the feature map is restored to the same as the input. The jump connection in the decoder splices the feature map of the corresponding layer of the encoder with the feature map of the current layer of the decoder to retain the detail information. In the downsampling and upsampling structures, the sliding window self-attention mechanism is combined to enhance the feature expression ability of the jump connection. Specifically, the convolution feature map after the dense residual block is input into the Patch Embedding layer for block processing, and the SwinTransformer encoder is used to extract features for each image block. The results extracted by the Swin Transformer encoder are concatenated with the upsampled feature map to further supplement and enhance the detail information, thereby improving the segmentation accuracy.

[0085] Specifically, the dense residual block includes:

[0086] First convolutional layer: In the first part of the dense residual block, the input image is processed by a 1×1×1 convolution to reduce the number of channels of the input image, and is normalized by a BN layer to ensure the stability of the values ​​during the convolution process. After normalization, the ReLU activation function is applied to increase the nonlinear transformation; the output of this process is the first intermediate feature map, which is further used in subsequent convolution operations.

[0087] The second convolutional layer receives the output from the first convolutional layer and processes it through a 3×3×3 convolution. After the convolution, the BN layer is also applied to normalize the output features, and the nonlinear transformation is introduced through the ReLU activation function to further extract the deep features of the image and ensure that the spatial dimension is not lost.

[0088] A residual connection is performed based on the output result of the second convolutional layer, that is, the original input is adjusted through a 1×1 convolution to make it consistent with the number of channels of the second output image and then added to avoid the gradient vanishing problem.

[0089] The third convolutional layer: receives the image after residual connection, and continues to perform 3×3×3 convolution processing on the features through convolution to further enhance the feature expression ability; finally, by adding the input image, the output of the first convolutional layer and the output of the third convolutional layer, a deep feature map that integrates multi-level features is obtained.

[0090] The operation of the dense residual block is expressed as follows:

[0091] Out(x)=H2(x+H2(H1(x)))+x+H2(x)+H2(H1(X))

[0092] Among them, x represents the input feature map, H1(x) represents the BN-ReLU-Conv1×1×1 operation, and H2(x) represents the BN-ReLU-Conv3×3×3 operation.

[0093] The implementation method of the jump connection between downsampling and upsampling is as follows: for the deep feature map after dense residual convolution, the Patch Embedding layer is used to divide the image block into blocks and the Swin Transformer encoder is used to extract features for each image block. The feature information in each image block is fused through the self-attention mechanism to adaptively capture the context information. The result extracted by the Swin Transformer encoder is spliced ​​with the upsampling result to realize the jump connection.

[0094] The implementation method of the Patch Embedding layer operation is:

[0095] First, the input high-level semantic feature map is cut into several small blocks of preset size, each of which is mapped to a vector space through a convolution operation to generate the initial patch embedding matrix X∈R (N2)×d , where N represents the number of small blocks after the image is cut, and d represents the feature vector dimension of each small block; specifically, through a 3D convolution layer, the convolution kernel size and stride are set to the small block size, so as to extract the features of each small block and map them to a new vector space, and the number of channels of the output feature map is d, that is, the feature dimension of each small block; in order to improve the stability of the model, the feature map after the convolution operation is normalized; the LN layer is used to standardize the convolution result, enhance the numerical stability during the training process, and accelerate the convergence of the model; finally, a processed feature map is obtained. This feature map contains the local information of the input image and provides a basis for the subsequent self-attention operation.

[0096] The Swin Transformer encoder appears in a two-stage serial structure; in the first stage, window-based multi-head self-attention (W-MSA) is used, and then in the second stage, shifted window-based multi-head self-attention (SW-MSA) is used; the self-attention calculation method is selected according to whether the current Swin Transformer encoder is odd or even:

[0097]

[0098] Among them, W-MSA and SW-MSA represent the conventional multi-head self-attention module and the window-based multi-head self-attention block respectively; z' and z'+1 represent the outputs of W-MSA and SW-MSA respectively; LN and MLP represent layer normalization and multi-layer perceptron respectively.

[0099] In order to efficiently calculate the shift window mechanism, a three-dimensional cyclic shift is adopted; based on the following self-attention formula:

[0100]

[0101] Among them, Q, K, V represent query, key and value respectively; d represents the size of query and key; B is the relative position encoding; then, the features processed by the attention mechanism are nonlinearly transformed and enhanced through a multi-layer perceptron (MLP); in the entire network process, dense residual connections and self-attention masks are combined to ensure the stability and effectiveness of the model.

[0102] In step S2, the implementation method of the coarse segmentation is as follows: extracting features from the input feature map through convolution layers in multiple dense residual blocks, and cutting and mapping the feature map to the vector space through Patch Embedding after each convolution layer; then, capturing the spatial dependency through the self-attention operation of the Swin Transformer module; in the downsampling process, the feature map is gradually reduced in dimension through maximum pooling, and then the feature map is restored by convolution upsampling, and spliced ​​with deeper features through jump connections to fuse information of different scales; finally, a coarse segmentation result is obtained.

[0103] Step S3: extracting the blood vessel centerline from the roughly segmented blood vessel region, and extracting image blocks of a set size along the blood vessel centerline to form an image block data set. Specifically, the following steps are included:

[0104] Step S31: After the rough segmentation, a binary dilation algorithm is applied to dilate the blood vessels to enhance the connectivity in the segmentation results:

[0105]

[0106] Where A is the original binary image (e.g., the labeled image of the segmented blood vessel area), represents the dilation operation, defined as performing a local maximization of the structure element on each pixel in the image, and B is the structure element.

[0107] Step S32: After dilating the blood vessels, the two largest connected domains are extracted as the left and right coronary arteries using the connected domain analysis, and the skeleton of the blood vessels is extracted as the centerline using the surface refinement algorithm:

[0108]

[0109] Among them, A is the input blood vessel image, B is the structural element, represents the erosion operation, which shrinks each target area in the image; n represents the number of iterations, and the skeleton of the image is gradually extracted through multiple erosion operations.

[0110] Step S33: based on the extracted blood vessel centerline, a fixed-size image block is cut out around the blood vessel centerline; the center of each image block is a point on the blood vessel centerline; for each point on the centerline, the boundary of the image block is determined by the size of the image block.

[0111] Step S4: Use another DS-Unet model to finely segment the newly extracted image block dataset, and combine the segmented results to generate a final segmented image that matches the original input size.

[0112] In step S4, the implementation method of the fine segmentation is as follows: save each cropped image block and record the cropping position, and the center line provides key information of the cropping area; input the image block into another DS-Unet model for fine segmentation to obtain a fine segmentation result, extract the image block from the stored small block file by reading the recorded position of each small block, and merge them into the original image according to the recorded coordinate position; each small block is merged with the corresponding area of ​​the original image by taking the maximum value to ensure that the content of the small block covers the original area; finally, the image is binarized to generate a final segmentation result.

[0113] In the deep learning training process, firstly, three-dimensional coronary artery CT image data for model training is obtained, and labels are annotated on the images, that is, blood vessels are segmented in the images, and then the original images and the labeled images are preprocessed to form a training data set; then, a DS-Unet model for coarse segmentation is trained with the training data set, and after the training is completed, the DS-Unet model can be used for coarse segmentation; then, the images in the training data set are input into the DS-Unet model for coarse segmentation, and then through the processing of step S3, an image block centered on the blood vessel centerline is cropped from the image, and then an image block data set consisting of unlabeled image blocks and labeled image blocks is constructed; and then another DS-Unet model for fine segmentation is trained with the obtained image block data set.

[0114] The following is further described in detail through specific embodiments.

[0115] Image preprocessing is crucial in image segmentation. It can improve algorithm performance, enhance model robustness, and reduce noise interference. Low-quality images may lead to inaccurate segmentation results, while preprocessing can effectively eliminate noise, blur, and artifacts, improve image quality, and thus improve segmentation accuracy. At the same time, preprocessing can also highlight important features in the image, such as edges, colors, and textures, to help the model better distinguish different areas.

[0116] S1. Load the 3D coronary artery CT image and adjust the window width and window position of the image. In this step, the pixel value of the image is adjusted to more clearly display the coronary artery area by setting the window width to 540 and the window position to 200:

[0117]

[0118] Among them, l is the original image pixel value, W level is the window level, W width It is the window width.

[0119] The Sobel filter method is used to calculate the gradient of each slice of the image, thereby highlighting the edge of the image and providing a basis for subsequent denoising. Sobel edge extraction detects edges by calculating the gradient of each pixel in the image. A 3×3 convolution kernel is used to extract gradient information in the horizontal and vertical directions, respectively, to highlight the horizontal and vertical edges. By calculating the amplitude of the horizontal and vertical gradients, the edge intensity map of the image can be obtained. Furthermore, pixels with a gradient amplitude greater than the image average are marked as edges, thereby generating a boundary mask to highlight the true edge area.

[0120] The generated boundary mask is used to determine the area within the boundary, and bilateral filtering is used to denoise the interior of the extracted boundary, taking into account spatial similarity and grayscale similarity to remove noise while retaining the edge of the image.

[0121]

[0122] Where I(x,y) is the grayscale value of the image pixel, σ s is the standard deviation in the spatial domain, σ r is the standard deviation of the pixel values.

[0123] Normalization and image resizing play an important role in image segmentation. Normalization can standardize the image pixel values ​​to a uniform range, which can eliminate the differences in pixel value distribution between different images and help improve the training efficiency and accuracy of the model. Scaling can adjust the image to a uniform size, ensuring that different images have the same scale when input, avoiding model processing problems caused by size differences. We normalize the denoised image and then resize the image to (128, 128, 128). The preprocessed data slices are as follows: Figure 2 shown.

[0124] S2. Use the preprocessed data set as the input of the neural network for rough segmentation.

[0125] Network architecture such as Figure 3 As shown, DR is a dense residual block and ST is a jump connection module including Patch Embedding and SwinTransformer self-attention mechanism:

[0126] The network structure of the present invention is based on an encoder-decoder architecture, and its main goal is to achieve end-to-end mapping from input images to segmentation results, without manually designing features or inserting intermediate steps, ensuring the efficiency and simplicity of the training process. In the encoder stage, the network uses dense residual module stacking, extracts multi-scale features through convolution, and gradually reduces the spatial resolution through downsampling operations while increasing the number of feature channels. The dense residual module can enhance the ability of feature extraction, especially when capturing complex vascular structures and small blood vessels. At the same time, in order to model global context information, the network introduces Swin Transformer modules at multiple key positions, and uses the self-attention mechanism to capture long-distance dependencies and global features, thereby effectively supplementing the local feature extraction capability of convolution. The deep features extracted by the encoder and the shallow features are transmitted to the decoder part through jump connections, realizing the fusion of multi-level features and improving the detail expression of segmentation. In the decoder stage, upsampling is first performed through deconvolution, and the spatial resolution is gradually restored while reducing the number of channels. Each decoding block consists of two layers of convolution, and the spliced ​​features are further processed. The Batch Norm layer and the ReLU layer are used for feature normalization and increasing nonlinear expression capabilities. Through skip connections, shallow features are fused with the progressively upsampled features in the decoder in the channel dimension, which can retain more boundary details and texture information while restoring the resolution.

[0127] Among them, the dense residual module consists of three convolutional layers and two residual connections, such as Figure 4 As shown:

[0128] First convolutional layer: In the first part of the dense residual block, the input image is processed through a 1×1×1 convolution to reduce the number of channels of the input image, and normalized through the BN layer to ensure the stability of the values ​​during the convolution process. After normalization, the ReLU activation function is applied to increase the nonlinear transformation. The output of this process will become the first intermediate feature map, which will be further used in subsequent convolution operations.

[0129] The second convolutional layer receives the output from the first convolutional layer and processes it through a 3×3×3 convolution. After the convolution, BN is also applied to normalize the output features, and a nonlinear transformation is introduced through the ReLU activation function to further extract the deep features of the image and ensure that the spatial dimension is not lost. A residual connection is performed based on the output result of the second convolutional layer, that is, the original input is adjusted through a 1×1 convolution to make it consistent with the number of channels of the second output image and then added to avoid the gradient disappearance problem.

[0130] The third convolutional layer receives the image after residual connection and continues to perform 3×3×3 convolution on the features through convolution to further enhance the feature expression ability. Finally, by adding the input image, the output of the first convolutional layer and the output of the third convolutional layer, a deep feature map integrating multi-level features is obtained.

[0131] The dense residual operation can be expressed as:

[0132] Out(x)=H2(x+H2(H1(x)))+x+H2(x)+H2(H1(X))

[0133] Where x represents the input feature map, H1(x) represents the BN-ReLU-Conv1×1×1 operation, and H2(x) represents the BN-ReLU-Conv3×3×3 operation.

[0134] The encoder part of the network extracts image features layer by layer through multiple dense residual blocks. The dense residual module can efficiently integrate feature information at different levels. The encoder outputs three different levels of feature maps: (16, 128, 128, 128), (32, 64, 64, 64) and (64, 32, 32, 32), representing multi-scale features of step-by-step downsampling. Figure 3 Middle ST module: The specific process of the jump connection between downsampling and upsampling is to divide the deep feature map after dense residual convolution into blocks through the PatchEmbedding layer and use the Swin Transformer self-attention mechanism to extract features from each image block. The result is spliced ​​with the upsampling result to realize the jump connection.

[0135] After the output of each level of dense residual block, the feature map is processed by the Patch Embedding layer to cut the feature map into several small blocks of preset size. For example, in the first level feature map, the size of the cut small block is 8×8×8, while in the second and third level feature maps, the size of the cut small block is 4×4×4 and 2×2×2 respectively. These small blocks are mapped to the vector space through a 3D convolution operation (the convolution kernel and stride are equal to the small block size) to generate the feature matrix X∈R (N2)×d, where N is the number of cut patches and d is the dimension of the feature vector. To ensure the stability of the model, the Patch Embedding output is normalized layer by layer to further enhance the numerical stability during training and accelerate the convergence of the model.

[0136] The feature map after Patch Embedding will be sent to the Swin Transformer module for global feature extraction. The depth of each Swin Transformer module is 2, using window-based multi-head self-attention (W-MSA) and shifted window multi-head self-attention (SW-MSA) respectively. These modules effectively improve the accuracy of segmentation by capturing spatial dependencies and cross-region context information. Specifically, the module calculates the attention weight through query (Q), key (K) and value (V), and combines the multi-layer perceptron (MLP) to perform nonlinear transformation and enhancement on the features after attention processing. At the same time, residual connections and attention masks are introduced to ensure the stability and effectiveness of the model.

[0137]

[0138] Where W-MSA and SW-MSA represent the conventional multi-head self-attention module and the window-based multi-head self-attention block, respectively. z' and z'+1 represent the outputs of W-MSA and SW-MSA, respectively. LN and MLP represent layer normalization and multi-layer perceptron, respectively.

[0139] After the rough segmentation, the binary dilation algorithm is first applied to dilate the blood vessels to enhance the connectivity in the segmentation results. The dilation operation uses a small ball with a radius of 4 as the structural element and performs local maximization on each pixel in the original binary image to obtain the dilated blood vessel area. Its expression is:

[0140]

[0141] Where A is the original binary image (e.g., the labeled image of the segmented blood vessel area), represents the dilation operation, which is defined as performing a local maximization of the structure element on each pixel in the image, and B is the structure element.

[0142] After dilating the blood vessels, the connected domain analysis is used to extract the two largest connected domains, representing the left and right coronary arteries, and the skeleton of the blood vessels is extracted through the surface refinement algorithm to obtain the center line. The skeleton extraction process is expressed as:

[0143]

[0144] Where A is the input blood vessel image, B is the structural element, represents the erosion operation, which shrinks each target area in the image. n represents the number of iterations, and the skeleton of the image is gradually extracted through multiple erosion operations.

[0145] According to the extracted vascular centerline, the original image and the corresponding label image are segmented into image blocks of size 64×64×64 according to the centerline, so as to serve as the input for fine segmentation. The center position of each cropped block corresponds to a point on the vascular centerline, and the boundary of the cropped block is determined by the preset block size. Each cropped image block and label block will be saved, and the cropping position will be recorded. The saved image blocks and label blocks will be used as input data for model training. The centerline provides key information of the cropped area, so that the training process can focus on the vascular area.

[0146] The image block and label block data are input into the network to obtain the fine segmentation result. By reading the recorded position of each small block, the image blocks are extracted from the stored small block file and merged into the original image according to the recorded coordinate position. Each small block is merged with the corresponding area of ​​the original image by taking the maximum value to ensure that the content of the small block covers the original area. Finally, the restored image will be binarized to generate the final segmentation result.

[0147] Experimental Results

[0148] All experiments were conducted using PyTorch 2.0.0+cu118 on a GeForce RTX3080Ti GPU. The proposed networks and the compared methods all follow the same training method. Unless otherwise stated, all networks use Dice coefficient loss as the loss function. The size of each data box is 512x512xD (where D ~ 240), and the size of the cut block is 64x64x64. The output segmentation confidence map is binarized with a threshold of 0.5 to generate the final segmentation result. Due to the large amount of data, 4-fold cross validation is used to divide the training and test sets. For both coarse and fine segmentation processes, the number of training iterations for all network models is set to 50, using the Adam optimizer, a training batch size of 4, and an initial learning rate of 0.001. If the validation loss does not improve after 10 iterations, the learning rate is halved.

[0149] We applied four commonly used evaluation metrics to assess the effectiveness of different methods: Dice Similarity Coefficient (DSC), Backtracking, Precision, and Hausdorff Distance (HD).

[0150]

[0151] HD(A,B)=max{max a∈A min b∈B d(a,b),max b∈B minα∈A d(a,b)

[0152] Among them, TP represents samples correctly identified as coronary arteries; FN represents samples predicted as background but actually belong to coronary arteries; FP represents samples predicted as coronary arteries but actually belong to background. S(TP+FN) represents the surface voxel set of actual coronary arteries, while S(TP+FP) represents the surface voxel set of predicted coronary arteries. d(a,b) represents the contour area difference.

[0153] To verify the effectiveness of the network model in the proposed method, we compare some common models with the same dataset. As can be seen in Table 1. The Swin Transformer Block integrated in the skip connection enables the model to effectively focus on global and local context. The focus on multi-scale attention is particularly beneficial for distinguishing coronary arteries from surrounding tissues of similar density, which is reflected in the superior recall rate (0.821) of DS-UNet, which is 5.1% higher than 3D-UNet. Although Swin UNETR also adopts a Transformer-based method, its lower DSC (0.804) and recall rate (0.773) indicate that the lack of dense connections may limit its ability to fully utilize spatial information, especially when dealing with smaller and complex vascular structures. Although AttenUNet performs strongly in reducing segmentation boundary errors (HD is 25.12), DS-UNet consistently surpasses it in precision and recall. This shows that although the attention mechanism helps to improve boundary accuracy, the combination of dense residual connections and attention mechanism in DS-UNet provides a more comprehensive solution that balances fine-grained segmentation with global feature extraction. Table 2 shows the performance comparison of several different models in coronary artery image segmentation. 3D-UNet, as a baseline model, has average precision and recall, and there are large errors in the segmentation boundaries. DRB-Unet, which introduces dense residual blocks (DRBs) to replace traditional convolutional layers, significantly improves precision and recall, and reduces the Hausdorff distance, indicating that the segmentation is more accurate. STB-UNet, with Swin Transformer blocks (STBs) as skip connections, performs best in recall and can capture more true positive areas, but has slightly lower precision and rougher segmentation boundaries. SwinUNETR improves precision through a global self-attention mechanism, but has a lower recall, resulting in inferior segmentation performance to STB-UNet. Finally, DS-UNet combines DRB and STB to achieve the best overall performance, with significantly improved precision and recall, and smoother segmentation boundaries, indicating that the combination of these two modules provides the best balance between fine-grained segmentation and global feature extraction.

[0154] Table 1 Comparison of results of various network models

[0155]

[0156] Table 2 Comparison of the effects of each module in the network

[0157]

[0158] To evaluate the effectiveness of our multi-stage segmentation approach that combines coarse and patch-based segmentation, we compared it with the traditional patch-based segmentation approach. The comparison aims to highlight the advantage of multi-stage segmentation in accurately capturing the complex structure of coronary arteries. In the patch-based approach, the training set is skeletonized using the ground truth labels. Patches are extracted around the skeleton points, and labeled and unlabeled regions are randomly selected, maintaining a 1:1 ratio to ensure a balanced representation. As shown in Table 3, the segmentation results show that the pure patch-based segmentation method has a significantly lower DSC than the multi-stage approach. The Hausdorff distance (HD) of the patch-based approach is also significantly higher, indicating that there are a large number of mis-segmented small vessels in the periphery of the coronary arteries. These results indicate that the multi-stage segmentation approach provides a more robust segmentation that is effective in reducing segmentation errors and improving the overall integrity of the coronary arteries.

[0159] Table 3 Results of two segmentation methods

[0160]

[0161] Figure 5 The segmentation results of the two methods are visually compared. Pure patch-based segmentation methods tend to misclassify some non-coronary vessels as coronary arteries despite the use of post-processing techniques such as connected domain analysis. This misclassification is mainly due to the poor connectivity of the segmented regions and the failure to segment small and complex parts of the coronary arteries, resulting in low overall accuracy. In contrast, the multi-stage method alleviates these problems by combining coarse segmentation (capturing the global anatomical structure) and fine patch-based segmentation, improving vessel coherence and reducing the inclusion of irrelevant vascular structures.

[0162] According to the experimental results, the system can effectively complete the preprocessing of the data set and use a neural network model that integrates multi-level feature information and self-attention mechanism for multi-stage processing. This not only enhances the recognition ability of small blood vessels and complex anatomical structures, thereby ensuring the details and accuracy of the segmentation results, but also improves the processing efficiency of large-size images by reducing computational complexity. In particular, it shows significant advantages in the segmentation of high-resolution images and complex structures, avoiding the memory consumption and computational complexity problems that may occur when directly processing the entire three-dimensional image, and ensuring the continuity and accuracy of blood vessel segmentation. This method provides strong technical support for the accurate diagnosis of coronary artery disease, and also lays a solid foundation for the widespread application of medical image segmentation tasks in the future, and has broad application prospects.

[0163] This embodiment also provides a three-dimensional coronary artery image segmentation system, including a memory, a processor, and computer program instructions stored in the memory and executable by the processor. When the processor executes the computer program instructions, the above method can be implemented.

[0164] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0165] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0166] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.

[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0168] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. A three-dimensional coronary artery image segmentation method, characterized in that: The following steps are involved: Step S1: preprocessing the three-dimensional coronary artery CT image, including window width and window position adjustment, edge detection and bilateral filtering, to reduce noise and enhance image contrast; Step S2: The preprocessed image is roughly segmented using the DS-Unet model to preliminarily segment the blood vessel area. The DS-Unet model is based on the UNet model and integrates a dense residual block and a sliding window self-attention mechanism, which can effectively capture the overall structure and boundaries of the blood vessels and ensure that the blood vessels are accurately identified; Step S3: extracting a blood vessel centerline from the roughly segmented blood vessel region, and extracting an image block of a set size along the blood vessel centerline to form an image block data set; Step S4: Use another DS-Unet model to finely segment the newly extracted image block dataset, and combine the segmented results to generate a final segmented image that matches the original input size.

2. A three-dimensional coronary artery image segmentation method according to claim 1, characterized in that: The step S1 comprises the following steps: Step S11: Load the three-dimensional coronary artery CT image and adjust the window width and window position of the image; by setting the window width and window position, adjust the pixel value of the image to a range suitable for subsequent processing: Where T(l) is the adjusted pixel value, l is the original image pixel value, and W level is the window level, W width is the window width; Step S12: Use the Sobel filter method to calculate the gradient of each slice image, so as to highlight the edge of the image and provide a basis for subsequent denoising processing; extract the gradient amplitude of each pixel in the original three-dimensional data, and extract the boundary area according to the size of the gradient amplitude: Among them, edges(x,y) is the edge strength at the pixel (x,y) in the image, gradient x (x,y), gradient y (x, y) are the horizontal and vertical gradient values ​​at the pixel (x, y) respectively; I(x+i, y+j) is the grayscale value of the pixel (x+i, y+j) in the image, where i = -1, 0, 1, j = -1, 0, 1, G x (i+1,j+1),G y (i+1, j+1) is the convolution kernel of the Sobel filter in the horizontal and vertical directions at the pixel (x+i, y+j); Step S13: Use bilateral filtering to remove noise inside the extracted boundary, taking into account spatial similarity and grayscale similarity, while removing noise and retaining the edge of the image: Among them, I(x,y) is the gray value of the pixel (x,y) in the image before filtering, I new (x, y) is the gray value of pixel (x, y) in the filtered image, σ s is the standard deviation in the spatial domain, σ r is the standard deviation of pixel values; Step S14: normalize the denoised image to ensure that the pixel values ​​are in the range of [0,1], and adjust the size to the same size to ensure that the image data has a uniform scale, thereby improving the calculation efficiency and accuracy: Among them, I(x,y) is the gray value of the pixel (x,y) in the image before normalization, I noemalized (x,y) is the grayscale value of pixel (x,y) in the normalized image.

3. A three-dimensional coronary artery image segmentation method according to claim 1, characterized in that: The DS-Unet model improves the convolution before downsampling into dense residual convolution based on the UNet model, and adds a sliding window self-attention mechanism as a jump connection in the downsampled and upsampled data to supplement the detail features. The overall structure of the DS-Unet model consists of an encoder and a decoder. In the encoder part, the spatial dimension of the feature map is gradually reduced through multiple downsampling, while the number of channels is increased. Before each downsampling, a dense residual block is used to process the data. The dense residual block realizes feature reuse and gradient flow by connecting the outputs of all previous layers. In the decoder part, the spatial dimension of the feature map is gradually restored through multiple upsampling, and the size of the feature map is restored to the same as the input. The jump connection in the decoder splices the feature map of the corresponding layer of the encoder with the feature map of the current layer of the decoder to retain the detail information.

4. A three-dimensional coronary artery image segmentation method according to claim 3, characterized in that: The dense residual block comprises: First convolutional layer: In the first part of the dense residual block, the input image is processed by a 1×1×1 convolution to reduce the number of channels of the input image, and is normalized by a BN layer to ensure the stability of the values ​​during the convolution process. After normalization, the ReLU activation function is applied to increase the nonlinear transformation. The output of this process is the first intermediate feature map, which is further used in subsequent convolution operations. The second convolutional layer receives the output from the first convolutional layer and processes it through a 3×3×3 convolution. After the convolution, the BN layer is also applied to normalize the output features, and a nonlinear transformation is introduced through the ReLU activation function to further extract the deep features of the image and ensure that the spatial dimension is not lost. A residual connection is performed based on the output result of the second convolutional layer, that is, the original input is adjusted through a 1×1 convolution to make it consistent with the number of channels of the second output image and then added to avoid the gradient vanishing problem; The third convolutional layer: receives the image after residual connection, and continues to perform 3×3×3 convolution processing on the features through convolution to further enhance the feature expression ability; finally, by adding the input image, the output of the first convolutional layer and the output of the third convolutional layer, a deep feature map integrating multi-level features is obtained; The operation of the dense residual block is expressed as follows: Out(x)=H2(x+H2(H1(x)))+x+H2(x)+H2(H1(X)) Among them, x represents the input feature map, H1(x) represents the BN-ReLU-Conv1×1×1 operation, and H2(x) represents the BN-ReLU-Conv3×3×3 operation.

5. A three-dimensional coronary artery image segmentation method according to claim 4, characterized in that: The implementation method of the skip connection between downsampling and upsampling is as follows: the deep feature map after dense residual convolution is divided into blocks through the PatchEmbedding layer and the Swin Transformer encoder is used to extract features for each image block. The feature information in each image block is fused through the self-attention mechanism to adaptively capture context information. The result extracted by the Swin Transformer encoder is spliced ​​with the upsampling result to realize the skip connection; The implementation method of the Patch Embedding layer operation is: First, the input high-level semantic feature map is cut into several small blocks of preset size, each of which is mapped to a vector space through a convolution operation to generate the initial patch embedding matrix X∈R (N2)×d , where N represents the number of small blocks after the image is cut, and d represents the feature vector dimension of each small block; specifically, through a 3D convolution layer, the convolution kernel size and stride are set to the small block size, so as to extract the features of each small block and map them to a new vector space, and the number of channels of the output feature map is d, that is, the feature dimension of each small block; in order to improve the stability of the model, the feature map after the convolution operation is normalized; the LN layer is used to standardize the convolution result, enhance the numerical stability during the training process, and accelerate the convergence of the model; finally, a processed feature map is obtained; The Swin Transformer encoder appears in a two-stage serial structure; in the first stage, window-based multi-head self-attention (W-MSA) is used, and then in the second stage, shifted window-based multi-head self-attention (SW-MSA) is used; the self-attention calculation method is selected according to whether the current Swin Transformer encoder is odd or even: Where W-MSA and SW-MSA represent conventional multi-head self-attention modules and window-based multi-head self-attention blocks, respectively; z' and z'+1 represent the outputs of W-MSA and SW-MSA, respectively; LN and MLP represent layer normalization and multi-layer perceptron, respectively; In order to efficiently calculate the shift window mechanism, a three-dimensional cyclic shift is adopted; based on the following self-attention formula: Among them, Q, K, V represent query, key and value respectively; d represents the size of query and key; B is the relative position encoding; then, the features processed by the attention mechanism are nonlinearly transformed and enhanced through a multi-layer perceptron (MLP); in the entire network process, dense residual connections and self-attention masks are combined to ensure the stability and effectiveness of the model.

6. A three-dimensional coronary artery image segmentation method according to claim 1, characterized in that: In step S2, the implementation method of the coarse segmentation is as follows: extracting features from the input feature map through convolution layers in multiple dense residual blocks, and cutting and mapping the feature map to the vector space through Patch Embedding after each convolution layer; then, capturing the spatial dependency through the self-attention operation of the SwinTransformer module; in the downsampling process, the feature map is gradually reduced in dimension through maximum pooling, and then the feature map is restored by convolution upsampling, and spliced ​​with deeper features through jump connections to fuse information of different scales; finally, a coarse segmentation result is obtained.

7. A three-dimensional coronary artery image segmentation method according to claim 1, characterized in that: The step S3 comprises the following steps: Step S31: After the rough segmentation, a binary dilation algorithm is applied to dilate the blood vessels to enhance the connectivity in the segmentation results: D=A⊕B Where A is the original binary image, ⊕ represents the dilation operation, which is defined as performing the local maximization of the structure element for each pixel in the image, and B is the structure element; Step S32: After dilating the blood vessels, the two largest connected domains are extracted as the left and right coronary arteries using the connected domain analysis, and the skeleton of the blood vessels is extracted as the centerline using the surface refinement algorithm: Among them, A is the input blood vessel image, B is the structural element, represents the erosion operation, which shrinks each target area in the image; n represents the number of iterations, and the skeleton of the image is gradually extracted through multiple erosion operations; Step S33: based on the extracted blood vessel centerline, a fixed-size image block is cut out around the blood vessel centerline; the center of each image block is a point on the blood vessel centerline; for each point on the centerline, the boundary of the image block is determined by the size of the image block.

8. A three-dimensional coronary artery image segmentation method according to claim 7, characterized in that: In step S4, the fine segmentation is implemented by: saving each cropped image block and recording the cropping position, and the center line provides key information of the cropping area; The image blocks are input into another DS-Unet model for fine segmentation to obtain the fine segmentation results. The image blocks are extracted from the stored small block files by reading the recorded positions of each small block, and they are merged into the original image according to the recorded coordinate positions; each small block is merged with the corresponding area of ​​the original image by taking the maximum value to ensure that the content of the small block covers the original area; Finally, the image is binarized to generate the final segmentation result.

9. A three-dimensional coronary artery image segmentation method according to claim 1, characterized in that: In the deep learning training process, firstly, three-dimensional coronary artery CT image data for model training is obtained, and labels are annotated on the images, that is, blood vessels are segmented in the images, and then the original images and the labeled images are preprocessed to form a training data set; then, a DS-Unet model for coarse segmentation is trained with the training data set, and after the training is completed, the DS-Unet model can be used for coarse segmentation; then, the images in the training data set are input into the DS-Unet model for coarse segmentation, and then through the processing of step S3, an image block centered on the blood vessel centerline is cropped from the image, and then an image block data set consisting of unlabeled image blocks and labeled image blocks is constructed; and then another DS-Unet model for fine segmentation is trained with the obtained image block data set.

10. A three-dimensional coronary artery image segmentation system, characterized in that: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method according to any one of claims 1 to 9 can be implemented.

Citation Information

Cited By

  • Medical image segmentation method and system based on heterogeneous feature collaborative extraction, medium and equipment

    CN120451573A

  • Medical image segmentation method and system for true cavity, false cavity and false cavity thrombus of aortic dissection

    CN120876852A

  • Medical image optimization method and system based on vascular branch selective blurring

    CN121437478A