Automatic Segmentation Method of Coronary Angiography Images Based on SFAG-DeepLabv3+
By improving the ASPP module of the DeepLabv3+ network as an ADP module, combined with GCS module and Bicubic upsampling, the problems of high background noise and low contrast of coronary angiography images are solved, segmentation accuracy and efficiency are improved, and more accurate and efficient detection of coronary artery and stenosis plaques are achieved.
Patent Information
- Application Number
- CN202410449402.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-04-15
AI Technical Summary
Coronary angiography images have high background noise, low contrast, and local blurred blood vessel boundaries. The existing segmentation methods are difficult to effectively deal with linear vascular structures and lesions, resulting in limited accuracy and stability of segmentation results.
The automatic segmentation method of coronary angiography image based on SFAG-DeepLabv3+ is adopted. By improving the ASPP module of the DeepLabv3+ network as an ADP module, an adaptive hybrid expansion convolution and dual pooling layer are introduced, combined with GCS module and Bicubic upsampling, the feature extraction and information fusion capabilities are enhanced, information loss is reduced, and segmentation continuity is improved.
It improves the segmentation accuracy and efficiency of coronary angiography images, reduces the false positive rate, enhances the stability of feature extraction and segmentation results of linear blood vessels, and achieves more accurate and efficient detection of coronary artery and stenosis plaques.
Smart Images

Figure CN118447241B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image segmentation, and relates to a method for segmenting coronary angiography images, in particular to an automatic segmentation method for coronary angiography images based on SFAG-DeepLabv3+. Background Art
[0002] Cardiovascular diseases (CVDs) are mostly caused by coronary artery diseases (CADs). The coronary circulation system is the main blood supply source for the myocardium. The deposition of lipids, cholesterol, calcium salts and other substances on the inner wall of the coronary artery can be transformed into plaques. The plaques will gradually increase and harden over time, and the arterial lumen will be narrowed due to the accumulation of atherosclerotic plaques, leading to coronary heart disease. Severe stenosis will cause blood flow obstruction and endanger life. Therefore, the quantification of coronary artery stenosis and plaques is crucial for diagnosing coronary artery diseases and assessing the risks of patients.
[0003] Coronary angiography images are the gold standard for diagnosing coronary artery diseases. Doctors observe the severity of vascular stenosis or narrowing by flushing the coronary vessels with a radioactive opaque dye. However, making a diagnosis based on this requires doctors to analyze hundreds of image frames to evaluate the morphology of plaques and the degree of vascular stenosis, which is highly subjective and time-consuming, and may even lead to misdiagnosis and omission. Moreover, due to the influence of contrast agents, projection angles and irradiation intensities during the imaging process of coronary angiography images, not all of the presented images are coronary angiography images. Coronary angiography images themselves also have the characteristics of unbalanced foreground-background distribution, low contrast and high background noise. Directly segmenting coronary angiography images will cause problems of over-segmentation or under-segmentation, affecting doctors' diagnosis of coronary artery stenosis. Therefore, a high-performance medical image diagnosis assistance system is needed to automatically segment the coronary artery vascular structure to achieve accurate detection of coronary arteries and stenotic plaques.
[0004] Traditional coronary artery segmentation relies on manual intervention and requires a large amount of computing time for a large number of coronary angiography images. In addition, traditional coronary artery segmentation methods usually have difficulty dealing with complex vascular structures, lesions and vascular branches, etc., resulting in limitations in the accuracy and stability of segmentation results, thereby reducing their efficiency and reliability in clinical applications.
[0005] With the rapid development of deep learning technology, CNN plays an increasingly crucial role in the field of medical image processing. Especially in the field of medical image segmentation, CNN demonstrates better performance and effectiveness than traditional segmentation methods, providing a more accurate and efficient image analysis tool for medical research and clinical practice. As a semantic segmentation algorithm, each generation of the DeepLab series algorithms has good segmentation performance. In 2018, the Google team proposed the DeepLabv3+ convolutional network, which is one of the semantic segmentation models with the best current performance. DeepLabv3+ introduced an encoder-decoder structure on the basis of v3 to further improve the segmentation accuracy. The encoder part adopts the structure of DeepLabv3, while the decoder part restores the spatial information through upsampling and feature fusion. This structure enables DeepLabv3+ to generate a more refined segmentation boundary while maintaining a high segmentation accuracy. However, for coronary angiography images, the ASPP module of DeepLabv3+ uses parallel dilated convolutions. Since the calculation method of dilated convolutions is similar to a checkerboard form, the receptive field is increased by inserting zero values into the standard convolution kernel. The convolution results obtained in each layer of the parallel dilated convolutions come from independent sets of dilated convolutions in that layer, and there is no continuity and correlation between the convolution results, thus resulting in the loss of local information. On the other hand, due to the linear structure of coronary arteries, the square convolution kernel of dilated convolutions has insufficient feature extraction for linear blood vessels and it is difficult to establish the correlation between different vascular branches.
[0006] In addition, when current scholars conduct automatic segmentation research on coronary angiography images using deep learning convolutional neural networks, although the end-to-end network structure is popular due to its high performance, there will be a certain amount of information loss during the compression and decompression process from the encoder to the decoder, thus causing the phenomenon of fractures when the model segments blood vessels. Summary of the Invention
[0007] Aiming at the problems of large background noise, low contrast, locally blurred blood vessel boundaries in coronary angiography images, and the insufficient extraction and fusion of linear blood vessel features in existing segmentation methods for coronary angiography images, the present invention provides an automatic segmentation method for coronary angiography images based on SFAG-DeepLabv3+.
[0008] The object of the present invention is achieved through the following technical solutions:
[0009] An automatic segmentation method for coronary angiography images based on SFAG-DeepLabv3+ includes the following steps:
[0010] Step 1, construct a CAG coronary angiography image dataset:
[0011] Step 101: Obtain the original coronary angiography image data;
[0012] Step 102: Process the original coronary angiography images using the Swin Transformer classification network to eliminate non-coronary angiography images;
[0013] Step 103: Select some coronary angiography images from the data processed in Step 102 to construct a CAG dataset;
[0014] Step 104: Use FSE data augmentation to augment the CAG dataset;
[0015] Step 105: Use Labelme to annotate the CAG dataset to obtain a labeled dataset, and randomly divide this dataset into a training set and a test set;
[0016] Step 2: Construct a coronary angiography image segmentation network based on AG-DeepLabv3+;
[0017] Step 201: Improve the ASPP module in the encoder part of the DeepLabv3+ network to obtain an ADP module. The ADP module includes a hybrid dilated convolutional layer and a double pooling layer with a serial structure having adaptive feature recalibration, where:
[0018] The convolutional layer part combines dilated convolution, enhances information integration in both spatial and channel dimensions, and forms a hybrid dilated convolution with a serial structure of adaptive feature recalibration. The specific steps are as follows: First, to enable each pixel of the high-level feature map to utilize the square region of the low-level feature layer without zero convolution or missing edges, three hybrid dilated convolutions with serial structures and dilation rates of 1, 2, and 5 are adopted. Then, the feature map after the hybrid dilated convolution is enhanced by information integration in both spatial and channel dimensions. In the channel dimension, the feature map after the dilated convolution is changed from [H, W, C] to [1, 1, C] through the global average pooling layer, and two 1×1×1 convolutions are used for information processing to obtain a C-dimensional vector. Then, the sigmoid function is used for normalization to obtain the corresponding mask, and through channel-wise multiplication, the feature map calibrated by channel information is obtained. In the spatial dimension, a convolution with 1 channel and a kernel size of 1×1 is used to change the feature map extracted by the hybrid dilated convolution from [H, W, C] to [H, W, 1], where H, W, and C are the height, length, and number of channels of the feature map respectively. The sigmoid activation is used to obtain the spatial attention map, which is finally directly applied to the feature map after the dilated convolution to complete the spatial information calibration of the hybrid dilated convolution feature extraction, forming a hybrid dilated convolution layer with a serial structure of adaptive feature recalibration, thereby promoting the acquisition of multi-scale information and local feature information of coronary angiography images during network training, and at the same time reducing the interference of non-coronary artery feature information in the angiography images;
[0019] The dual pooling layer combines a bar pooling kernel and a square pooling kernel of global average pooling and is in a parallel structure with the convolutional layer for feature fusion. Let x ∈ R C×H×W , where R is the input tensor. x is input into two parallel paths, each path containing a horizontal or vertical bar pooling layer, followed by a one-dimensional convolutional layer with a kernel size of 3 for modulating the current position and its neighboring features. To obtain the output z ∈ R C×H×W , y h and y υ are combined to obtain y c,i,j ∈ R C×H×W . Then, after a 1×1 convolution and a Sigmoid layer, element-wise multiplication is performed with the original input to obtain the final output z. For the global information of the coronary angiography image, the value of each pixel is averaged in its corresponding channel and each pixel point is replaced with the average value of its channel. After this process, each channel obtains an average value, and these average values form a new feature map. The new feature map is upsampled using a 2×2 convolution to obtain its output value;
[0020] Step 202: By introducing a GCS module between the encoder and the decoder, the information loss caused during the compression and decompression of information from the encoder to the decoder is solved. Specifically: the first stage of the GCS module is based on Gaussian excitation, and globally contextual features of low-level semantic information and high-level semantic information are integrated and enhanced at the channel level; the second stage uses spatial excitation to enhance the feature expression ability of high-level semantic information and low-level semantic information; the two stages are carried out simultaneously to enhance the network's feature extraction ability for global contextual information in terms of space and channels.
[0021] Step 203: In the decoder part of the DeepLabv3+ network, Bicubic upsampling is adopted. By utilizing more adjacent pixel information, the continuity of the model during blood vessel segmentation is improved.
[0022] Step 3: Use the training set obtained in Step 1 to train the coronary angiography image segmentation network based on AG-DeepLabV3+ constructed in Step 2, and input the test set into the trained coronary angiography image segmentation network based on AG-DeepLabV3+ to obtain the segmentation result.
[0023] Compared with the prior art, the present invention has the following advantages:
[0024] 1. The present invention uses a two-dimensional Swin Tansformer classification network to screen coronary angiography images to filter out non-coronary artery images, and then enhances the image quality through the proposed Filtering Smoothing Equalization (FSE) image enhancement method to improve the segmentation efficiency and effectiveness of coronary arteries.
[0025] 2. Based on the DeepLabv3+ backbone network, the present invention proposes the AG-DeepLabv3+ network. In the encoder part, based on the ASPP module, an Adaptive Dual Dilated Pooling (ADP) module is proposed to construct a strengthened feature extraction network. This module solves the problem of insufficient local feature extraction of linear blood vessels by a single dilated convolution through an adaptive feature recalibration hybrid dilated convolution layer and a dual pooling layer, strengthens the network's feature fusion ability for local and global contexts of coronary angiography images, realizes the model's feature extraction of multi-scale local information and global context information of linear coronary arteries, and maximally avoids the problem of missing small blood vessels during the segmentation process. A Gaussian Context Space (GCS) module is introduced between the encoder and the decoder. By adaptively reorganizing features in the channel dimension and the spatial dimension, the information loss caused during the compression and decompression of information from the encoder to the decoder is reduced, and the problem of blood vessel breakage during blood vessel segmentation caused by information loss during the compression and decompression of information from the encoder to the decoder is solved. In the decoder part, bicubic interpolation upsampling is adopted. By utilizing more adjacent pixel information, the continuity of blood vessel segmentation is improved. Brief Description of the Drawings
[0026] Figure 1 It is a flowchart of an automatic segmentation method for coronary angiography images based on SFAG-DeepLabv3+;
[0027] Figure 2 It is a diagram of the ADP module;
[0028] Figure 3 It is a diagram of the GCS module;
[0029] Figure 4 It is a flowchart of the overall implementation. Detailed Implementation Manner
[0030] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings, but are not limited thereto. Any modification or equivalent replacement of the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered by the protection scope of the present invention.
[0031] The present invention provides an automatic segmentation method for coronary angiography images based on SFAG-DeepLabv3+. First, the original coronary angiography is sent into a two-dimensional Swin Tansformer classification network to screen the coronary angiography images and eliminate non-coronary arteries; key frames at appropriate viewpoints in the end-diastolic phase of the heart are selected from the screened coronary angiography images to establish a CAG data set; to obtain high-quality coronary angiography images, a filtering smoothing equalization image enhancement technology (FSE) is proposed to preprocess the images; the enhanced data set is sent into the network for training, effectively alleviating the phenomenon of high false positive rate caused by low-quality coronary angiography images. As Figure 1 shown, the specific steps are as follows:
[0032] Step A1: Construct a CAG coronary angiography image data set and send the data set into a coronary angiography image segmentation network for training.
[0033] The specific steps are as follows:
[0034] Step B1: Obtain the original coronary angiography image data.
[0035] In the present invention, the original coronary angiography image data is provided by the Department of Cardiology, the Second Affiliated Hospital of Harbin Medical University, and all data are real data.
[0036] Step B2: Use the Swin Transformer classification network to process the original coronary artery angiography images and eliminate non-coronary angiography images.
[0037] Since the original coronary angiography image data contains coronary angiography images and non-coronary angiography images, directly segmenting the original data will increase the false positive rate of network segmentation. In this step, before segmentation, the Swin Transformer classification network is used to screen the original coronary angiography images to remove non-coronary angiography images.
[0038] Step B3: Select some coronary angiography images from the data processed in Step B2 to construct the CAG dataset.
[0039] In this step, coronary angiography images at the end-diastolic phase of the heart are selected from the screened coronary angiography images to construct the CAG dataset.
[0040] Step B4: Use FSE data augmentation to augment the CAG dataset.
[0041] The present invention adopts FSE image processing technology to redistribute the gray levels of the image by weighted averaging each pixel point to achieve the purpose of image enhancement. The specific steps are as follows:
[0042] Step B41: Use a Gaussian function as a smoothing filter to filter the image, and slide a 5×5 window on the coronary angiography image. Let the center point coordinates of the smoothing filter be (x, y), then the weight w of each position (i, j) in the smoothing filter is calculated as shown in Equation (1):
[0043]
[0044] Among them, σ is the standard deviation, the weight of the pixel point at the center of the window is the largest, and the weights of the pixel points farther away from the center gradually decrease.
[0045] By normalizing all the weights, that is, ensuring that their sum is 1, the normalized weight k can be obtained, as shown in Equation (2):
[0046]
[0047] Among them, m and n traverse each position in the filter, that is, from -2 to 2.
[0048] In the coronary angiography image, align the center point of the smoothing filter with the current image pixel position, extract the pixel values within the image area covered by the smoothing filter, multiply the extracted pixel values by the corresponding Gaussian weights, and add all the products to obtain the filtered pixel value; assign the filtered pixel value to the corresponding position of the output image. By this operation, the noise in the image is reduced to achieve image smoothing.
[0049] Step B42: Calculate the grayscale histogram of the filtered coronary angiography image to enhance the image contrast. The specific steps are as follows:
[0050] Step B421: Assume that the image grayscale range is [0, L - 1] (L is the number of grayscale levels), count the number of occurrences of each pixel value and normalize it to obtain the frequency of each pixel value.
[0051] Step B422: Calculate the cumulative distribution function (CDF): The CDF is the integral of the frequency distribution function (PDF), which represents the probability of each pixel value appearing in the original image. For a grayscale value a, where a = 0, 1, 2,..., L - 1, the calculation formula of the CDF is shown in Equation (3):
[0052]
[0053] where P(b) represents the frequency of pixels with grayscale value b appearing in the image.
[0054] Step B423: Calculate the equalized pixel value: Map each pixel value in the original image to a new pixel value so that the histogram after equalization approximates a uniformly distributed histogram. The calculation formula of this mapping function is shown in Equation (4):
[0055]
[0056] where H(a) represents the mapped pixel value, M and N represent the width and height of the image respectively, and min represents the minimum pixel value in the original image. Remap each pixel in the original image according to the new mapping function, thereby realizing the redistribution of the image grayscale levels and enhancing the image contrast.
[0057] In this step, while globally denoising the coronary angiography image through the FSE image enhancement technology, the image contrast is enhanced, achieving the effect of simultaneously realizing multiple objectives such as smoothing the image, enhancing image details, and removing noise. In addition, this processing method better preserves the overall structure and details of the image, improves the image quality, and helps with the subsequent coronary artery segmentation task.
[0058] Step B5: Label the CAG dataset and perform random partitioning.
[0059] In this step, use Labelme to label the obtained CAG dataset to obtain a labeled dataset.
[0060] Step A2: Construct a coronary angiography image segmentation network based on AG-DeepLabv3+.
[0061] The coronary angiography image segmentation network based on AG-DeepLabv3+ uses the DeepLabv3+ network as the basic model. The ASPP module is improved in its encoder structure to propose the ADP module, the GCS module is proposed between its encoder and decoder, and Bicubic interpolation is used in the decoder upsampling part. The specific construction steps are as follows:
[0062] Step B6: Improve the ASPP module in the encoder part of the DeepLabv3+ network to obtain the ADP module.
[0063] (1) Adaptive Feature Re-calibration Hybrid Dilated Convolution
[0064] For coronary angiography images, the ASPP module in the DeepLabv3+ network is prone to the problem of local information loss caused by zero convolution during the operation process, and it cannot fully capture the correlation information between different branch vessels. To capture the vascular information of different scales while retaining the local feature information of the vessels and the correlation information between different branch vessels, the convolutional layer of the ADP module uses a hybrid dilated convolution with a serial structure of adaptive feature re-calibration. For the selection of the dilation rate of the dilated convolution, assume that N dilated convolutions are continuously stacked, the convolution kernel size is K, and the dilation coefficients of each dilated convolution correspond to [r 1 , r 2 ,...... r n . The maximum distance between two non-zero elements is shown in Equation (5):
[0065]
[0066] where M i is the maximum distance between two non-zero elements in the i-th layer, and r i is the dilation coefficient of the i-th layer. For the last dilated convolution, its maximum distance is M n = r n , that is, the maximum distance is the dilation rate of this layer, and the purpose is to make M 2 ≤K, that is, the maximum distance between two non-zero elements in the second layer is less than or equal to the size of the convolution kernel of this layer. In the present invention, a 3×3 convolution kernel is selected, so three consecutive dilated convolutions are used. To enable each pixel of the high-level feature map to utilize the square area of the low-level feature layer without zero convolution or missing edges, let M 1 = 1. From Equation (5), M 1 = max[positive, negative, r 1 , and r 1= 1, so when selecting the dilation rate of dilated convolution, to avoid the loss of underlying feature information, the first dilation rate is set to 1. According to this design criterion, when r = [1, 2, 5] and r = [1, 3, 7], the existence of gaps between non-zero elements can be avoided. After comparative experiments, the present invention finally adopts the dilation rate of r = [1, 2, 5]. The feature map after hybrid dilated convolution is enhanced by integrating information in both the spatial and channel dimensions to form a hybrid dilated convolution with an adaptive feature recalibration serial structure. Specifically, in the channel dimension, the feature map after dilated convolution is changed from [H, W, C] to [1, 1, C] through a global average pooling layer, and two 1×1×1 convolutions are used for information processing to obtain a C-dimensional vector. Then, the sigmoid function is used for normalization to obtain the corresponding mask. Through channel-wise multiplication, the feature map calibrated by channel information is obtained; in the spatial dimension, a convolution with a channel number of 1 and a kernel size of 1×1 is used to change the feature map extracted by hybrid dilated convolution from [H, W, C] to [H, W, 1], where H, W, and C are the height, length, and number of channels of the feature map respectively. The sigmoid activation is used to obtain the spatial attention map, which is finally directly applied to the feature map after dilated convolution to complete the spatial information calibration of hybrid dilated convolution feature extraction, forming a convolutional layer with adaptive feature recalibration. Thereby, it promotes the acquisition of multi-scale information and local feature information of coronary angiography images during network training, while reducing the interference of non-coronary artery feature information in the angiography images. As Figure 2 (a) shows.
[0067] (2) Dual pooling layer
[0068] In the convolutional layer of the ADP module, dilated convolution is used to extract multi-scale features of blood vessels. However, due to the special linear structure of coronary arteries, although dilated convolution expands the receptive field by adjusting the dilation rate, this expansion is limited. Even after dilation, the effective size of the convolutional kernel is still physically limited, which results in limited perception ability of dilated convolution for linear structures. Based on this, the ADP module introduces a new dual pooling layer. This pooling structure combines a bar-shaped pooling kernel and a square pooling kernel of global average pooling and performs feature fusion in a parallel structure with the convolutional layer to make up for the problem of weak feature extraction of dilated convolution for the linear structure of blood vessels, while improving the global information extraction ability of the model for coronary angiography images.
[0069] For the linear structure of blood vessels, the present invention uses horizontal and vertical strip pooling operations to capture blood vessel feature information from different spatial dimensions, such as Figure 2(b) as shown above. It can establish long-range dependencies between widely distributed vascular regions, encode regions in the shape of stripes, while the narrow kernel shape in another dimension helps to focus on capturing local vascular detail information. Let \(x\in\mathbb{R}\) C×H×W , where \(\mathbb{R}\) is the input tensor. Input \(x\) into two parallel paths, each path contains a horizontal or vertical strip pooling layer, followed by a one-dimensional convolutional layer with a kernel size of 3, which is used to modulate the current position and its neighboring features. Given \(y\) h \(\in\mathbb{R}\) C×H and \(y\) υ \(\in\mathbb{R}\) C×W . To obtain the output \(z\in\mathbb{R}\) C×H×W that contains more useful global priors, combine \(y\) h and \(y\) υ together to get \(y\) c,i,j \(\in\mathbb{R}\) C×H×W , as shown in Equation (6):
[0070]
[0071] Then, after passing through a \(1\times1\) convolution and a Sigmoid layer, multiply element-wise with the original input to obtain the final output \(z\), and its formula is shown in Equation (7):
[0072] \(z = S(x,\sigma(f(y)))\) (7)
[0073] where \(S\) is element-wise multiplication, \(\sigma\) is the Sigmoid function, and \(f\) is the \(1\times1\) convolution.
[0074] In the above process, each position in the output tensor is allowed to establish relationships with various positions in the input tensor. For example, in Figure 2 (b), the square bounded by the black box in the output tensor is connected to all positions (the rectangles where the red and purple boxes are located) that have the same horizontal or vertical coordinates as it. Therefore, by repeating the above aggregation process multiple times, dependencies between different vascular branches can be constructed in the entire coronary angiogram image.
[0075] In addition, for the global information of the coronary angiogram image, the value of each pixel is averaged in its corresponding channel and each pixel is replaced with the average value of its channel. After this process, each channel obtains an average value, and these average values form a new feature map. Compared with the original feature map, the size of this new feature map will be reduced to \(1 / Q\) of the original, where \(Q\) is the pooling coefficient of 3. The new feature map is upsampled using a \(2\times2\) convolution to obtain its output value.
[0076] Step B7: By introducing a GCS module between the encoder and the decoder, the information loss caused during the compression and decompression process from the encoder to the decoder is solved.
[0077] The structure of the GCS module is as Figure 3 shown. First, the first stage of this module is based on Gaussian excitation to integrate and enhance the global context features of the underlying semantic information and the high-level semantic information at the channel level. Secondly, inspired by the spatial attention mechanism, the second branch of this module uses spatial excitation to enhance the feature expression ability of the high-level semantic information and the low-level semantic information. The two stages are carried out simultaneously to enhance the network's feature extraction ability of the global context information in terms of space and channels.
[0078] In the first stage, the GCS module extracts the context information of the low-level semantic information and the high-level semantic information from the encoder by considering the relationship between pixels in the channel dimension, which helps the decoder better reconstruct the image. This stage is divided into three parts, namely global context aggregation (GCA), normalization, and Gaussian context excitation (GCE). The purpose of the GCA operation is to obtain channel statistics by performing GPA spatial aggregation on the global context information in the feature map, so as to help the network capture long-range dependencies. The purpose of the normalization operation is to stabilize the distribution of the global context of the feature information from the encoder. The mean shift is defined as z - μ, where represents the average value of the global context z. This mean shift measures the deviation between z and μ. In the channel dimension, the present invention uses a Gaussian function with a negative correlation hypothesis to implement the transformation and activation operations. However, directly setting the mean deviation as the input will make the Gaussian function unstable because different input samples will result in inconsistent mean shift distributions. Based on this, a factor σ of a specific instance is introduced to stabilize the distribution, ensuring a zero mean and a unit variance. At this time, the global context is as shown in Equation (8):
[0079]
[0080] where σ represents the standard deviation of the global context, and the calculation formula of σ is as shown in Equation (9):
[0081]
[0082] In the formula, ε is a small constant. This way of obtaining is consistent with the result of normalizing z, and it is simplified to: In the Gaussian context excitation stage, the normalized global context is excited through a Gaussian function (including two steps of transformation and activation) to obtain the attention activation value (attention map), which is then converted into key information that can be used by the decoder. The Gaussian function formula is shown in Equation (10):
[0083]
[0084] where is the amplitude of the Gaussian function, b represents the mean of the Gaussian function, and c represents the standard deviation of the Gaussian function, which is used to control the diversity of the channel attention map. The larger the standard deviation, the smaller the diversity of the activation values between channels. To satisfy a negatively correlated relationship for channel attention learning, the attention activation value g is shown in Equation (11):
[0085]
[0086] In the second stage, two-stage operations including channel squeeze and spatial excitation, namely Squeeze and Excitation, are performed on the low-level semantic information and high-level semantic information from the encoder. First, in the Squeeze stage, a convolutional layer with an output channel of 1 and a convolutional kernel size of 1×1 is used to perform a convolutional operation on the input feature map, compressing the size of the feature map from [H, W, C] to [H, W, 1] to obtain a weight matrix with a dimension of [H, W, 1], where C is the number of channels, and H and W are the spatial dimensions. Second, in the Excitation stage, the Sigmoid function is used to normalize the above weight matrix to obtain the final weight matrix. The weight matrix is multiplied by the original input feature map to obtain the weighted feature map. By correcting the information of the low-level semantic information and high-level semantic information from the encoder in the spatial dimension, it helps to extract the information of the vascular global context in the spatial dimension.
[0087] B8. In the decoder part of the DeepLabv3+ network, bicubic upsampling is adopted to improve the continuity of the model during vascular segmentation by utilizing more neighboring pixel information.
[0088] For a given point, the output pixel in the target image is the weighted average of the nearest 4×4 neighborhood pixels. Assume the size of the source image A is m×n, and the size of the target image G after scaling by K times is M×N, that is Each pixel of A is known. To find the value of a certain pixel (X, Y) in the target image G, it is necessary to first find the corresponding pixel (x, y) of the pixel (X, Y) in the original image, and then use the 16 pixels closest to the pixel (x, y) in the original image as the parameters for calculating the pixel value at the target image, and use the Bicubic basis function to find the weights of the 16 pixels. The Bicubic function is shown in Equation (12):
[0089]
[0090] For the pixel point (x, y) to be interpolated, take its nearby 4×4 neighborhood points (x i , y j ), where i, j = 0, 1, 2, 3. Perform interpolation calculation according to Equation (13), where d refers to the weighting coefficient, then the target pixel value is:
[0091]
[0092] Step A3: Input the test set into the coronary angiography image segmentation network based on AG-DeepLabv3+, and use the obtained pre-trained weights to obtain the segmentation result.
[0093] As Figure 4 shown, integrating the foregoing image segmentation method, the overall implementation process can be specifically divided into the following steps:
[0094] First, input the training set with enhanced input data;
[0095] Secondly, use the proposed ADP module to replace the ASPP block in the encoder stage to obtain the improved AG-DeepLabv3+ network;
[0096] Secondly, propose the GCS module between the encoder and the decoder to extract global context features from the features from the encoder;
[0097] Secondly, use Bicubic upsampling in the decoder stage;
[0098] Secondly, obtain the pre-trained model;
[0099] Finally, use the test set for network evaluation.
[0100] Example:
[0101] In this example, the coronary angiography image samples are obtained from the Department of Cardiology, the Second Affiliated Hospital of Harbin Medical University. Coronary angiography images of 58 coronary angiography patients are obtained and a database is established. The acquisition process of all data is carried out in accordance with the approved rules and regulations.
[0102] The coronary angiography image data format used in this embodiment is Dicom, with angiographic views and video information. The imaging rate of coronary angiography is 15 frames per second, the bit depth is 24 per pixel, the window width is 256, the window level is 128, and the pixels of each frame of angiography image are 512×512. In the coronary angiography image, the coronary artery runs on the surface of the heart and distributes around the heart, mainly having three major branches: the left anterior descending branch (LAD), the left circumflex branch (LCX), and the right coronary artery (RCA). In addition, under these three major branches, there are also derivative branches such as Diag (diagonal branch), S (septal branch), AVG (atrioventricular groove branch), CB (conus branch), RV (right ventricular branch), OM (obtuse marginal branch), AM (acute marginal branch), PD (posterior descending branch), PL (left ventricular posterior branch), etc. Since coronary angiography shows a two-dimensional image, there may be problems such as vascular image overlap and vascular eccentric lesions, and different projection angles affect the determination of vascular inner diameter and lesion length. Therefore, multiple positions and multiple angles of projection are required. The projection positions for collecting the left coronary angiography images of the CAG dataset in the present invention are LAO+CRA (left anterior oblique plus cranial position), LAO+CAU (left anterior oblique plus caudal position), AP+CAU (anteroposterior plus caudal position), and the projection positions for the right coronary angiography images are AP+CRA (anteroposterior plus cranial position), RAO+CRA (right anterior oblique plus cranial position), RAO+CAU (right anterior oblique plus caudal position).
[0103] To reduce the influence of non-coronary angiography images on the accuracy of coronary artery segmentation, before segmenting the coronary angiography images, the present invention uses a Swin-Transformer classification network to screen the coronary angiography images. The classification network uses 3280 angiography images, and the training set and the test set are divided according to a ratio of 2:1.
[0104] When training the segmentation network, the present invention selects the key frames at an appropriate viewing point at the end-diastolic phase of the heart to establish the CAG dataset, because there is enough contrast agent during this time period to make the entire blood vessel visible and there is no unrecognizable coronary artery structure. Among the 58 cases of patient data finally used for training the coronary angiography image segmentation network, 30 cases of patients are in the training set and 28 cases of patients are in the test set, and each case of patient contains 10 - 51 pictures.
[0105] Evaluation indicators:
[0106] To evaluate the effectiveness of network segmentation, the present invention uses Dice (Dice similarity coefficient), IOU (Intersection over Union), Recall, and Accuracy as the evaluation indicators for this experiment. The formulas for each indicator are shown in Formulas (14), (15), (16), and (17):
[0107]
[0108]
[0109]
[0110]
[0111] Among them, TP refers to the number of coronary artery pixels in the predicted image; TN refers to the number of background pixels in the predicted image; FP refers to false positives, that is, the number of background pixels misidentified as blood vessel pixels; FN refers to false negatives, that is, the number of blood vessel pixels misidentified as background pixels. Dice is used to evaluate the similarity between the blood vessel segmentation result and the ground truth; IOU is used to evaluate the overlap degree between the segmentation result and the ground truth; Recall represents the number of pixels correctly predicted as coronary arteries among all coronary artery pixels; Accuracy is the proportion of samples predicted as blood vessels among all correctly predicted samples.
[0112] Experimental Results and Analysis:
[0113] Comparative Experiment:
[0114] To verify the effectiveness of the SFAG-DeepLabv3+ segmentation method in coronary angiography image segmentation, the present invention reproduced several mainstream methods in this field, namely U-Net, PSPNet, AngioNet, HRNetv2, Segformer, DeepLabv3+, and conducted a comparative experiment with the method proposed by the present invention. The experimental results are shown in Table 1. Since the datasets used by most coronary artery segmentation methods are of private nature, there is currently no suitable public coronary artery dataset available for testing. Under the same conditions, the present invention used the self-created CAG dataset for the comparative experiment, and the final value of the evaluation index of the model on the test set is the average result of multiple runs.
[0115] As can be seen from the data in Table 1, compared with the original backbone network DeepLabv3+, the SFAG-DeepLabv3+ model has improved in terms of Dice coefficient, intersection over union (IOU), pixel accuracy (PA), and accuracy (Accuracy) in coronary artery segmentation, and its parameter quantity only increases by 2.012M. Among them, the average Dice coefficient of the model on the test set increases from 0.9001 to 0.9285, the intersection over union increases from 0.8184 to 0.8666, the recall rate increases from 0.8756 to 0.9249, and the accuracy increases from 0.9686 to 0.9770. In addition, the method proposed in the present invention is superior to other methods in all evaluation indicators. Among them, the PSPnet network has the least parameter quantity, but SFAG-DeepLabv3+ is far higher than the PSPNet segmentation network in all four evaluation indicators. Especially in the IOU evaluation indicator, it has increased by 16.34% compared with PSPnet, and there is a higher similarity between the vascular segmentation result and the ground truth, showing better network performance in coronary angiography image segmentation.
[0116] Table 1 Comparison of results of different segmentation networks
[0117]
[0118] Ablation experiment:
[0119] To further explore the influence of each part in SFAG-DeepLabv3+ on the coronary artery segmentation performance, the present invention conducts a large number of ablation experimental studies on FSE, ADP, GCS, and Bicubic based on the constructed CAG coronary angiography image dataset. The experimental results are shown in Table 2.
[0120] Table 2 Ablation experiment results
[0121]
[0122]
[0123] The ablation study results in Table 2 show that the individual FSE, ADP, GCS, and Bicubic are effective in improving the segmentation performance, and the network shows the best segmentation performance when the four act simultaneously. Among them, the recall rate improves the most when the four act simultaneously. Compared with the original DeepLabv3+ backbone network, its recall rate increases by 4.93%. Since coronary artery segmentation with a higher recall rate usually has greater clinical value in diagnosing coronary artery stenosis and other coronary artery diseases, among all the network structures compared in the ablation study, SFAG-DeepLabv3+ is the best choice for coronary artery segmentation.
Claims
1. An automatic segmentation method for coronary angiography images based on SFAG-DeepLabv3+, characterized in that The method comprises the following steps: Step 1: Construct a CAG coronary angiography image dataset: Step 101, obtaining original coronary angiography image data; Step 102, using the Swin Transformer classification network to process the original coronary angiography image and remove non-coronary angiography images; Step 103, selecting part of the coronary angiography images from the data processed in step 102 to construct a CAG data set; Step 104: Use FSE data enhancement to perform data enhancement on the CAG dataset; Step 105: Use Labelme to label the CAG dataset to obtain a labeled dataset, and randomly divide the dataset into a training set and a test set; Step 2: Build a coronary angiography image segmentation network based on AG-DeepLabv3+: Step 201, improve the ASPP module of the DeepLabv3+ network encoder part to obtain an ADP module, wherein the ADP module includes a hybrid dilated convolution layer and a double pooling layer with a serial structure of adaptive feature recalibration, wherein: The convolution layer is combined with hybrid dilated convolution, and information integration and enhancement are performed in the spatial and channel dimensions at the same time to form a hybrid dilated convolution with a serial structure and adaptive feature recalibration. The specific steps are as follows: First, three hybrid dilated convolutions with serial structures and dilation rates of 1, 2, and 5 are used; then, the feature map after hybrid dilated convolution is simultaneously integrated and enhanced in the spatial and channel dimensions. In the channel dimension, the feature map after dilated convolution is changed from [H, W, C] to [1, 1, C] through the global average pooling layer, and two 1×1×1 convolutions are used for information processing to obtain a C-dimensional vector, which is then normalized using the sigmoid function to obtain the corresponding mask, and channel-wise multiplication is performed to obtain a feature map calibrated with channel information; in the spatial dimension, a convolution with a channel number of 1 and a convolution kernel size of 1×1 is used to change the feature map extracted by the hybrid dilated convolution from [H, W, C] to [H, W, 1], where H, W, and C are the height, length, and number of channels of the feature map, respectively, and sigmoid activation is used to obtain spatial attention. map, and finally directly applied to the feature map after the dilated convolution to complete the spatial information calibration of the hybrid dilated convolution feature extraction, forming a hybrid dilated convolution layer with a serial structure with adaptive feature recalibration; The double pooling layer combines the strip pooling kernel and the square pooling kernel of the global average pooling with the convolution layer into a parallel structure and then performs feature fusion. Let x∈R C×H×W , R is the input tensor, x is input into two parallel paths, each path contains a horizontal or vertical strip pooling layer, followed by a one-dimensional convolutional layer with a kernel size of 3, which is used to modulate the current position and its neighboring features, in order to obtain an output z∈R containing more useful global priors C×H×W , y h and υ Put together, we get y c,i,j ∈R C ×H×W Then, after a 1×1 convolution and a Sigmoid layer, the corresponding elements are multiplied with the original input to obtain the final output z; for the global information of the coronary angiography image, the value of each pixel is averaged in its channel and each pixel point is replaced by the average value of its channel. After this process, each channel obtains an average value, and these average values constitute a new feature map. The new feature map is upsampled by 2×2 convolution to obtain its output value; Step 202, by introducing a GCS module between the encoder and decoder, wherein: the first stage of the GCS module is based on Gaussian excitation, and the global context feature integration and enhancement of the bottom semantic information and the high-level semantic information at the channel level; the second stage adopts spatial excitation to enhance the feature expression ability of the high-level semantic information and the low-level semantic information; the two stages are performed simultaneously; Step 203: Bicubic upsampling is used in the DeepLabv3+ network decoder part; Step 3: Use the training set obtained in step 1 to train the coronary angiography image segmentation network based on AG-DeepLabV3+ constructed in step 2, and input the test set into the trained coronary angiography image segmentation network based on AG-DeepLabV3+ to obtain the segmentation result.
2. The automatic segmentation method of coronary angiography images based on SFAG-DeepLabv3+ according to claim 1, characterized in that In step 103, a coronary angiography image at the end of diastole is selected from the screened coronary angiography images to construct a CAG data set.
3. The method for automatic segmentation of coronary angiography images based on SFAG-DeepLabv3+ according to claim 1, characterized in that The specific steps of step 104 are as follows: (1) Using a Gaussian function as a smoothing filter to filter the image, a 5×5 window is used to slide on the coronary angiography image, the center point of the smoothing filter is aligned with the current image pixel position, and the pixel values in the image area covered by the smoothing filter are extracted. The extracted pixel values are multiplied by the corresponding Gaussian weights, and all the products are added to obtain the filtered pixel values; Assign the filtered pixel value to the corresponding position of the output image; (2) Calculate the grayscale histogram of the filtered coronary angiography image.
4. The method for automatic segmentation of coronary angiography images based on SFAG-DeepLabv3+ according to claim 3, characterized in that The specific steps of (2) are as follows: Assuming that the grayscale range of the image is [0, L-1], L is the number of grayscale levels, count the number of times each pixel value appears and normalize it to get the frequency of each pixel value; Calculate the cumulative distribution function CDF: For a gray value a, where a=0,1,2,...,L-1, the calculation formula of CDF is as follows: Among them, P(b) represents the frequency of pixels with gray value b appearing in the image; Calculate the equalized pixel value: map each pixel value in the original image to a new pixel value so that the equalized histogram is approximately a uniformly distributed histogram, and remap each pixel in the original image according to the new mapping function.
5. The method for automatic segmentation of coronary angiography images based on SFAG-DeepLabv3+ according to claim 4, characterized in that The calculation formula of the mapping function is as follows: Among them, H(a) represents the mapped pixel value, M and N represent the width and height of the image respectively, and min represents the minimum pixel value in the original image.
6. The method for automatic segmentation of coronary angiography images based on SFAG-DeepLabv3+ according to claim 1, characterized in that In the first stage of step 202, the GCS module extracts low-level semantic information and high-level semantic information context information from the encoder by considering the relationship between pixels in the channel dimension. This stage is divided into three parts, namely global context aggregation GCA, normalization and Gaussian context excitation GCE, where: The purpose of the GCA operation is to obtain channel statistics by performing GPA spatial aggregation on the global context information in the feature map; The purpose of the normalization operation is to stabilize the distribution of the global context from the encoder feature information, and the mean shift is defined as z-μ, where Represents the average value of the global context z. In the channel dimension, the Gaussian function with negative correlation assumption is used to implement the conversion and activation operations. σ is introduced to stabilize the distribution and ensure 0 mean and 1 variance. At this time, the global context for: Where σ represents the standard deviation of the global context; this method obtains The method is consistent with the result of normalizing z, which can be simplified to: In the Gaussian context excitation stage, the normalized global context is excited by a Gaussian function to obtain the attention activation value and convert it into key information that can be used by the decoder.
7. The method for automatic segmentation of coronary angiography images based on SFAG-DeepLabv3+ according to claim 1, characterized in that In the second stage of step 202, the low-level semantic information and high-level semantic information from the encoder are subjected to a two-stage operation including channel squeezing and spatial excitation, namely, Squeeze and Excitation. First, in the Squeeze stage, a convolution operation is performed on the input feature map using a convolution layer with an output channel of 1 and a convolution kernel size of 1×1, and the size of the feature map is compressed from [H, W, C] to [H, W, 1] to obtain a weight matrix with a dimension of [H, W, 1], where C is the number of channels, and H and W are the spatial dimensions; secondly, in the Excitation stage, the weight matrix is normalized using a Sigmoid function to obtain a final weight matrix; the weight matrix is multiplied by the original input feature map to obtain a weighted feature map.
8. The method for automatic segmentation of coronary angiography images based on SFAG-DeepLabv3+ according to claim 1, characterized in that In step 203, for a given point, the output pixel in the target image is the weighted average of the nearest 4×4 neighboring pixels. Assuming that the size of the source image A is m×n, the size of the target image G after scaling K times is M×N, that is, Each pixel of A is known. To find the value of a pixel (X, Y) in the target image G, we must first find the pixel (x, y) corresponding to the pixel (X, Y) in the original image. Then, we use the 16 pixels closest to the pixel (x, y) in the original image as the parameters for calculating the pixel value in the target image. We use the Bicubic basis function to find the weights of the 16 pixels. For the pixel (x, y) to be interpolated, we take the 4×4 neighboring points (x i ,y j ), i, j, = 0, 1, 2, 3, the interpolation calculation is performed as follows, where d refers to the weighting coefficient, and the target pixel value is:
Citation Information
Patent Citations
Improved remote sensing image segmentation method based on DeepLabV3
CN114387518A
Triangular mesh curved surface generation method and device, equipment and storage medium
CN115482358A