Coronary artery image segmentation method and system based on three-dimensional deep learning network

By adopting a three-dimensional deep learning network in coronary image segmentation, combining deconfusion, frequency cross attention and frequency enhancement modules, the problems of boundary noise, low contrast and false positive points in coronary segmentation are solved, and higher segmentation accuracy and recall are achieved.

CN119941758APending Publication Date: 2025-05-06NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510041050.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing coronary artery segmentation methods are prone to noise when boundary segmentation, resulting in unsmooth boundary and low contrast between the coronary artery and background, which can easily lead to undersegment and false positive points.

Method used

The coronary image segmentation method based on a three-dimensional deep learning network is adopted to suppress high-frequency components by adding a deconfusion module before each downsampling stage of the encoder; a frequency cross attention module is used instead of a jump connection to reduce the semantic gap between the encoder and the decoder; a frequency enhancement module is used in the decoder to enhance the coronary-related frequency and weaken the irrelevant frequency.

Benefits of technology

Effectively inhibit high-frequency components, reduce boundary segmentation errors, improve the accuracy and recall of coronary artery segmentation, and better capture the details and structural characteristics of the coronary artery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941758A_ABST
    Figure CN119941758A_ABST
Patent Text Reader

Abstract

The invention discloses a coronary artery image segmentation method and system based on a three-dimensional deep learning network, and relates to the technical field of medical image segmentation. The method is based on a novel multi-space and multi-frequency three-dimensional deep learning network, large-scale changes of coronary arteries can be processed, and representative features of the coronary arteries in complex anatomical structures and forms can be extracted; according to the invention, a de-confusion module is designed for weakening high-frequency components which may cause confusion and reducing the influence of a confusion phenomenon on a network; a frequency cross attention module is adopted, so that the semantic gap between an encoder and a decoder is reduced, and the accuracy of coronary artery segmentation is improved; hidden multi-scale context information is effectively extracted; furthermore, a frequency enhancement module is applied to enhance the coronary artery related frequency and weaken the irrelevant frequency, and global information and local information are concerned, so that the network can better capture the detail and structural features of the coronary artery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and specifically relates to a coronary artery image segmentation method and system based on a three-dimensional deep learning network. Background Art

[0002] Cardiovascular disease is the leading cause of death in recent years, mainly including coronary artery disease (CAD), which has a high morbidity and mortality rate, causing one-quarter of deaths each year. Coronary computed tomography angiography (CCTA) is a method of visualizing coronary arteries using contrast agents. Due to its non-invasiveness and high sensitivity, it is widely used in the clinical diagnosis of cardiovascular diseases. In clinical practice, coronary artery segmentation (CAS) is an important auxiliary diagnostic method. However, the large amount of CCTA data and the complex vascular structure of the coronary arteries make manual processing very time-consuming. Therefore, accurate and automatic coronary artery segmentation is an important solution that can improve diagnostic efficiency and reduce the burden on doctors.

[0003] At present, neural network models have achieved great success in medical image segmentation (MIS) tasks. At present, medical image segmentation tasks are mainly based on U-Net and its deformations, such as 2D U-Net, 3D U-Net, Attention U-Net and V-Net. Early medical image segmentation methods mainly used 2D methods. With the development of technology, 3D methods gradually emerged and achieved remarkable success in many medical segmentation tasks. Since 3D methods can extract depth information, their performance is usually better than 2D methods. Although U-Net has achieved remarkable performance in medical image segmentation, as a general model, it is still not enough to solve the challenge of coronary artery segmentation. Some improvements made for coronary artery segmentation are briefly listed. Mirunalini et al. use CNN to determine whether there is a coronary artery in the slice, and then use ResUNet for segmentation. Wang et al. add a Local Contextual Transformer module at the jump connection of U-Net to obtain local contextual information. Dong et al use the attention-guided feature fusion module at the jump connection to fuse the channel information of the decoder and encoder. However, these coronary artery segmentation methods all calculate images in the spatial domain, while ignoring the use of frequency domain information. We can transform images from the spatial domain to the frequency domain through Fourier transform. In the spatial domain, the image variables are spatial coordinates and the values ​​are grayscale values. In the frequency domain, the image variables are frequency components and the values ​​are amplitudes. Each frequency amplitude in the frequency domain is obtained by weighting all spatial domain grayscales, and has a global receptive field.

[0004] There are some challenges in coronary artery segmentation: First, downsampling will cause some noise to appear at the boundary (unsmooth jagged edges), which can easily lead to boundary segmentation errors. Second, the contrast between the coronary artery and the background is low, which can easily misjudge the coronary artery as the background, resulting in under-segmentation. Finally, the coronary artery has complex structural details, and the decoder does not extract enough detail information, resulting in false positive points. Summary of the invention

[0005] In view of the shortcomings of the existing technology, a coronary artery image segmentation method and system based on a three-dimensional deep learning network are provided, which effectively suppress the high-frequency components that cause confusion and reduce the semantic gap between the U-shaped network encoder and decoder. It is used to enhance the frequencies related to the coronary arteries and weaken the irrelevant frequencies. At the same time, it pays attention to global information and local information, and achieves good performance results in the coronary artery image segmentation task.

[0006] In a first aspect, the present invention provides a coronary artery image segmentation method based on a three-dimensional deep learning network, comprising the following specific steps:

[0007] Step 1: Acquire a coronary artery CT image dataset; the coronary artery CT image dataset includes a plurality of coronary artery CT images and their corresponding coronary artery labels;

[0008] Step 2: preprocessing the acquired coronary artery CT image dataset to obtain a preprocessed coronary artery CT image dataset;

[0009] The preprocessing is to truncate the grayscale interval of the coronary artery CT image and retain the set grayscale range;

[0010] Step 3: Construct a three-dimensional deep learning network FU-Net;

[0011] The three-dimensional deep learning network adopts a U-Net encoder-decoder architecture, in which a de-obfuscation module is added before each downsampling stage of the encoder, a frequency enhancement module is applied to the decoder to enhance the image through frequency features, and a frequency cross-attention module is used to replace the skip connection between the encoder and the decoder;

[0012] The de-obfuscation module is a dynamic learnable mask used to reduce frequency components higher than the Nyquist frequency in the coronary artery CT image; the Nyquist frequency is half of the sampling frequency;

[0013] The dynamic learnable mask is expressed as:

[0014]

[0015] Among them, M(u,v) is the mask value at the frequency component (u,v), u is the frequency value of the first component in the frequency domain, v is the frequency value of the second component in the frequency domain, α is the learnable weight, and R is the range of frequency components less than the Nyquist frequency NF:

[0016] R={(u,v)|u≤NF and v≤NF} (3)

[0017] The frequency cross attention module first obtains the frequency representations G and X corresponding to the high-level feature g and the low-level feature x output at a certain stage in the decoder after fast Fourier transformation; then obtains the query matrix by linear projection of the frequency representation G corresponding to the high-level feature g, and obtains the key matrix and the value matrix by linear projection of the frequency representation X corresponding to the low-level feature x; multiplies the query matrix Q with the key matrix K by Hadamard product to obtain the attention matrix A; multiplies the attention matrix A with the value matrix V by Hadamard product again to obtain the weight sum Y, and finally converts the weight sum Y to the spatial domain by inverse Fourier transform to obtain the result y in the spatial domain;

[0018] The frequency enhancement module adopts a two-branch structure, including a frequency domain branch and a space domain branch, for simultaneously extracting features in the frequency domain and the space domain, and then converting the frequency domain features f fre and the spatial domain feature f spa By adding element by element, the frequency-enhanced output is obtained;

[0019] The frequency domain branch receives the input f obtained by concatenating the output of the previous stage of the decoder and the spatial domain result y obtained by the corresponding frequency cross-attention module in And through Fourier transform, we can get the amplitude F of its frequency representation F amp and phase angle F phi , which represents the amplitude of F with respect to frequency amp and phase angle F phi Linear networks are used to extract features, and then an inverse Fourier transform iFFT is used to obtain the frequency domain features f after filtering irrelevant frequencies. fre ;

[0020] The spatial domain branch receives the input f obtained by concatenating the output of the previous stage of the decoder and the spatial domain result y obtained by the corresponding frequency cross attention module in , and the concatenated input f in Perform two convolution operations to extract the relevant coronary artery spatial features and obtain the spatial domain feature f spa ;

[0021] Step 4: Using the preprocessed coronary artery CT image dataset to train the three-dimensional deep learning network to obtain a trained three-dimensional deep learning network;

[0022] Step 5: Obtain the coronary artery CT image to be segmented and input it into the trained three-dimensional deep learning network to obtain the segmentation result.

[0023] In a second aspect, the present invention provides a coronary artery image segmentation system based on a three-dimensional deep learning network, which is used to implement a coronary artery image segmentation method based on a three-dimensional deep learning network, including a data acquisition module and a three-dimensional deep learning network FU-Net;

[0024] The data acquisition module is used to acquire the coronary artery CT image to be segmented as the input of the three-dimensional deep learning network FU-Net;

[0025] The three-dimensional deep learning network FU-Net is used to perform image segmentation on the coronary artery CT image to be segmented to obtain a segmentation result;

[0026] In a third aspect, the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program as described above for the coronary artery image segmentation method based on a three-dimensional deep learning network;

[0027] In a fourth aspect, the present invention provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the coronary artery image segmentation method based on a three-dimensional deep learning network described in the present invention is implemented.

[0028] Compared with the prior art, the present invention has at least the following beneficial effects:

[0029] The method described in the present invention is based on a novel multi-space, multi-frequency three-dimensional deep learning network (FU-Net). The method proposed in the present invention can handle large-scale changes in coronary arteries and extract representative features under their complex anatomical structures and morphologies; the present invention designs a de-obfuscation module (DAM) to weaken high-frequency components that may cause confusion and reduce the impact of the confusion phenomenon on the network; adopts a frequency cross attention (FCA) module to reduce the semantic gap between the encoder and the decoder, thereby improving the accuracy of coronary artery segmentation; effectively extracts hidden multi-scale contextual information; further applies a frequency enhancement module (FEM) module to enhance frequencies related to the coronary arteries and weaken irrelevant frequencies, while paying attention to global information and local information, so that the network can better capture the details and structural features of the coronary arteries. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1It is a structural block diagram of a three-dimensional deep learning network in an embodiment of the present invention;

[0031] Figure 2 This is a schematic diagram of the structure of a deobfuscation module in an embodiment of the present invention;

[0032] Figure 3 It is a structural schematic diagram of frequency cross attention in an embodiment of the present invention;

[0033] Figure 4 This is a schematic diagram of the structure of a frequency enhancement module in an embodiment of the present invention;

[0034] Figure 5 This is a result diagram comparing different methods in the embodiments of the present invention;

[0035] Among them, (a) is the result diagram of U-Net; (b) is the result diagram of Attention U-Net; (c) is the result diagram of DenseUNet; (d) is the result diagram of V-Net; (e) is the result diagram of UNETR; (f) is the result diagram of GFUNet; (g) is the result diagram of ResUNet*; (h) is the result diagram of LCTDRNet*; (i) is the result diagram of CASNet*; (j) is the result diagram of FU-Net; (k) is the result diagram of Ground truth. DETAILED DESCRIPTION

[0036] The specific embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings and Examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present invention.

[0037] The present invention proposes a multi-space, multi-frequency three-dimensional deep learning network, FU-Net, which comprehensively responds to the challenge of coronary artery segmentation. It is based on the classic encoder-decoder structure and includes three core modules. First, the present invention replaces the traditional skip connection with a frequency cross-attention (FCA) module based on frequency information. The FCA module can suppress irrelevant background noise while enhancing relevant coronary artery features. By using the FCA module, the fusion of encoding and decoding features can be effectively guided, and the semantic gap between encoder and decoder features can be reduced, thereby separating the coronary artery from irrelevant coronary veins and background noise. Secondly, the present invention proposes a de-aliasing (DAM) module, which is added before the downsampling operation to effectively suppress high-frequency components and reduce the confusion caused by them. Specifically, the present invention first designs a dynamic mask to suppress high-frequency components that may cause confusion through weights, thereby improving the recall rate of coronary artery segmentation. Then, by calculating the importance of different scales in space, these hierarchical features are dynamically adjusted to adapt to blood vessels of different scales. In addition, the present invention also designs a frequency enhancement module (FEM) for enhancing the frequencies related to the coronary arteries and weakening the irrelevant frequencies, while focusing on global information and local information, so that the network can better capture the details and structural features of the coronary arteries. Experiments on the coronary artery CT image dataset collected by the present invention show that the method described in the present invention achieves good performance and is superior to other advanced methods.

[0038] In this embodiment, a coronary artery image segmentation method based on a three-dimensional deep learning network includes the following specific steps:

[0039] Step 1: Acquire a coronary artery CT image dataset; the coronary artery CT image dataset includes a plurality of coronary artery CT images and their corresponding coronary artery labels;

[0040] In this embodiment, the coronary artery label is to label the pixels of the coronary arteries in the coronary artery CT image as white, and the pixels irrelevant to the background as black;

[0041] Step 2: preprocessing the acquired coronary artery CT image dataset to obtain a preprocessed coronary artery CT image dataset;

[0042] In this embodiment, the preprocessing is to truncate the grayscale interval of the coronary artery CT image, and retain the grayscale range related to the coronary artery: [-260,760];

[0043] Step 3: Construct a three-dimensional deep learning network FU-Net;

[0044] In order to achieve the accuracy and robustness of blood vessel segmentation, the present invention proposes a multi-space and multi-frequency three-dimensional deep learning network (FU-Net), such as Figure 1As shown in the figure, a U-Net encoder-decoder architecture is adopted, in which a de-obfuscation module (DAM) is added before each downsampling stage of the encoder, a frequency enhancement module (FEM) is applied to the decoder to enhance the image through frequency features, and a frequency cross attention (FCA) module is used to replace the skip connection between the encoder and decoder;

[0045] The present invention integrates the three proposed modules FCA, DAM and FEM into a U-shaped encoder-decoder architecture. In this embodiment, the encoder includes five stages E1, E2, E3, E4 and E5, which gradually extracts the coding feature maps of each scale and gradually extracts the spatial features to more advanced semantic features. The decoder includes four stages D1, D2, D3 and D4;

[0046] When downsampling is performed in U-Net, it usually causes confusion. According to the Nyquist-Shannon Sampling Theorem, frequency components exceeding half of the sampling rate will be identified as low-frequency components. For example, the jagged edges at the coronary artery boundary will cause boundary segmentation errors. This phenomenon has little effect on larger segmentation targets, but due to the small scale of the coronary artery, boundary segmentation errors will have a greater impact on the accuracy. Therefore, a frequency domain-based de-obfuscation module DAM is designed. Figure 2 As shown, de-aliasing must be done before sampling, because the aliasing phenomenon is a sampling problem, so it cannot be solved after sampling. Therefore, before each downsampling, DAM uses a mask to reduce the frequency components above the Nyquist frequency (NF) in the coronary CT image, thereby reducing the aliasing phenomenon. DAM is a dynamic learnable mask, in which the low-frequency components in the central area represent the background information with smoother grayscale changes, and the weight is set to 1; the high-frequency components in the boundary area represent the boundary information with more drastic grayscale changes, and the weight is set to a learnable parameter α, α∈[0,1];

[0047] The sampling rate (SR) is:

[0048]

[0049] Where SR is the sampling rate, C out is the number of channels of the output image, C in is the number of channels of the input image, H out is the height of the output image, H in is the height of the input image, W out is the width of the output image, W in is the width of the input image;

[0050] The Nyquist frequency is half of the sampling rate. A dynamic learnable mask is designed. The mask value of the frequency component less than the Nyquist frequency (NF) in the coronary artery CT image is set to 1, and the mask value of the frequency component greater than the Nyquist frequency (NF) is set to α (ranging between 0 and 1).

[0051]

[0052] Where M(u,v) is the mask value at the frequency component (u,v), u is the frequency value of the first component in the frequency domain, v is the frequency value of the second component in the frequency domain, and R is the range of frequency components less than the Nyquist frequency:

[0053] R={(u,v)|u≤NF and v≤NF} (3)

[0054] The frequency cross attention module FCA, such as Figure 3 As shown, firstly, the high-level feature g outputted at a certain stage in the decoder and the low-level feature x outputted at a certain stage in the encoder are subjected to fast Fourier transform to obtain the frequency representations G and X corresponding to the high-level feature g and the low-level feature x; then the frequency representation G corresponding to the high-level feature g is subjected to linear projection (Linear) to obtain the query matrix (Query Q), and at the same time, the frequency representation X corresponding to the low-level feature x is subjected to linear projection (Linear) to obtain the key matrix (Key K) and the value matrix (Value V); inspired by the convolution theorem, it can be seen that the product in the frequency domain is equivalent to the convolution in the spatial domain. Based on this principle, the Hadamard product is used to replace the two matrix multiplications in the cross attention. Specifically, the Hadamard product is used to multiply the query matrix Q with the key matrix K to obtain the attention matrix A, which is equivalent to the global circular convolution of the query matrix q in the spatial domain and the key matrix k in the spatial domain; the Hadamard product is used again to multiply the attention matrix A with the value matrix V to obtain the weight sum Y, and finally the inverse Fourier transform is used to convert the weight sum Y to the spatial domain to obtain the result y in the spatial domain;

[0055] A=Q⊙K (4)

[0056] y=iFFT(A⊙V) (5)

[0057] Where, ⊙ represents Hadamard product, iFFT represents inverse Fourier transform;

[0058] In order to solve the challenge of the complex structure and details of the coronary artery, it is hoped to enhance the image by frequency, retain the important frequencies and remove the irrelevant frequencies. The frequency enhancement module is as follows Figure 4As shown, a two-branch structure is adopted, including a frequency domain branch and a spatial domain branch, which are used to simultaneously extract features in the frequency domain and the spatial domain. The frequency domain branch is used to extract global information, and the spatial domain branch is used to extract local information. Specifically, the frequency domain branch receives the input f obtained by splicing the output of the previous stage of the decoder and the spatial domain result y obtained by the corresponding frequency cross attention module. in And through Fourier transform, we can get the amplitude F of its frequency representation F amp and phase angle F phi , which represents the amplitude of F with respect to frequency amp and phase angle F phi A linear network (MLP) is used for feature extraction to capture useful amplitudes and phase angles, remove irrelevant amplitudes and phase angles, and then an inverse Fourier transform iFFT is used to obtain the frequency domain feature f after filtering irrelevant frequencies. fre ; The spatial domain branch concatenates the input f in , perform two convolution (Conv) operations to extract the relevant coronary artery spatial features and obtain the spatial domain feature f spa , then the frequency domain feature f fre and the spatial domain feature f spa By adding element by element, the output of the FEM module after frequency enhancement is obtained;

[0059] Specifically, the frequency branch extracts the frequency domain feature f fre :

[0060] f fre =iFFT(MLP(FFT(f in ))) (6)

[0061] Among them, MLP represents linear network, FFT represents Fourier transform;

[0062] The spatial branch extracts the spatial domain feature f spa :

[0063] f spa =Conv(Conv(f in )) (7)

[0064] Among them, Conv represents the convolution operation;

[0065] Step 4: Using the preprocessed coronary artery CT image dataset to train the three-dimensional deep learning network to obtain a trained three-dimensional deep learning network;

[0066] Step 5: Obtain the coronary artery CT image to be segmented and input it into the trained three-dimensional deep learning network to obtain the segmentation result;

[0067] In specific implementation, the present invention first uses five commonly used evaluation indicators, namely Dice similarity coefficient (DSC), recall rate (Recall), precision rate (Precision), average symmetric surface distance (ASSD) and Hausdorff distance (HD) for quantitative evaluation. The DSC score represents the overlap rate between actual blood vessels and those identified as blood vessels; the recall rate (also known as sensitivity) is the proportion of actual blood vessels correctly identified as blood vessels, and the precision rate is the proportion of actual blood vessels correctly identified as blood vessels; HD is the maximum distance from one set to the nearest point in another set, describing the similarity of the two sets. The smaller the HD, the higher the similarity of the two sets; ASSD measures the average symmetric surface distance between the predicted segmentation boundary and the true segmentation boundary. The smaller the ASSD, the better the segmentation result. They are calculated by the following formulas:

[0068]

[0069] Among them, TP stands for true positive, indicating that the real coronary artery is predicted; FP stands for false positive, indicating that the area that is not a coronary artery is predicted as a real coronary artery; FN stands for false negative, indicating that the real coronary artery is not predicted. Among them, S(TP+FN) and S(TP+FP) are the voxel sets of the real coronary artery and the voxel set of the predicted coronary artery. d[a,S(TP+FN)] is the shortest distance from voxel a to the set S(TP+FN); d[b,S(TP+FP)] is the shortest distance from voxel b to the set S(TP+FP).

[0070] CCTA data needs to be preprocessed before being fed into the segmentation network because it includes a lot of tissue under radiodensity. Intercepting the HU values ​​of the coronary arteries can improve the segmentation performance. In our dataset, HU values ​​in the range of [-260,760] were used. The preprocessing operation effectively removes irrelevant areas and noise.

[0071] The proposed FU-Net was implemented on two GeForce RTX 3090 GPUs using python3.8 and torch1.12.1. We set the batch size to 3, the input block size to 1×16×512×512, trained for 200 epochs, and set the initial learning rate to 10 -4 , which decays to 10 after the 100th epoch. -5 , which decays to 10 after the 160th epoch. -6 . Comparison with the state-of-the-art methods.

[0072] In order to further analyze the effectiveness of the proposed method, the method of the present invention is compared with nine medical image segmentation methods, four of which are classic general segmentation networks, including 3D U-Net, Attention U-Net, DenseUNet, V-Net, UNETR, one frequency-based general segmentation network GFUNet, and the other three are the latest methods for coronary artery segmentation, including: CASNet, ResUNet and LCTDRNet.

[0073] Table 1 quantitatively shows the performance of FU-Net and 9 comparison methods on CCTA coronary artery datasets, including 4 general segmentation networks, 1 frequency-based general segmentation network, and 3 networks designed specifically for vascular segmentation. It can be seen that the FU-Net of the present invention has the best performance. It can be seen that FU-Net has achieved 83.72%, 84.13% and 83.96% in DSC, Recall and Precision, respectively, which is better than other state-of-the-art methods. Compared with the traditional 3D U-Net, DSC is improved from 80.94% to 93.72%. The recall rate and precision are also improved from 81.76% and 80.70% to 84.13 and 83.96%, respectively, verifying the effectiveness of the FU-Net proposed in the present invention. On the other hand, it can be concluded from the HD95 metric in Table 1 that the similarity of the background truth value of the predicted CA between UNETR and the present invention method is the highest, because the precision rates are only 88.86% and 83.96%, respectively. However, the smaller recall obtained by UNETR indicates that more labeled coronary arteries are not segmented into vessels. Overall, this method has better reliability for coronary artery segmentation in terms of DSC, Recall, Precision, HD95, ASSD, and Params.

[0074] Table 1 Performance of FU-Net and 9 comparison methods on CCTA coronary artery dataset

[0075] Method Dice↑ SE↑ PC↑ HD↓ ASSD↓ Params↓ 3DU-Net 0.8094 0.8176 0.8070 11.1209 0.7520 10.10M AttentionU-Net 0.8145 0.8101 0.8252 11.6268 0.8594 10.15M DenseUNet 0.8231 0.8090 0.8435 11.1018 0.8173 16.99M V-Net 0.8255 0.8070 0.8520 9.5826 0.6521 65.24M UNETR 0.8215 0.7698 0.8886 8.2533 0.6720 92.58M GFUNet 0.8148 0.8282 0.8085 8.0977 0.5761 28.95M CASNet* 0.8261 0.7923 0.8717 7.4781 0.5590 5.35M ResUNet* 0.8128 0.8284 0.8029 11.3792 0.8184 10.11M LCTDRNet* 0.8250 0.7796 0.8838 8.5253 0.6306 23.52M FU-Net(Ours) 0.8372 0.8413 0.8396 6.1826 0.5094 10.26M

[0076] Table 1 compares the method of the present invention with other methods on the coronary artery data set quantitatively, and the best results are shown in bold. "↑" and "↓" indicate that the larger the value, the better the performance, and "↓" indicates that the smaller the value, the better the performance. Figure 5 Comparison of the results of different methods. The yellow and white boxes in the second and third rows of the background ground truth respectively indicate that the compared method may have problems of under-segmentation and over-segmentation.

[0077] In this embodiment, a coronary artery image segmentation system based on a three-dimensional deep learning network is used to implement a coronary artery image segmentation method based on a three-dimensional deep learning network, including a data acquisition module and a three-dimensional deep learning network FU-Net;

[0078] The data acquisition module is used to acquire the coronary artery CT image to be segmented as the input of the three-dimensional deep learning network FU-Net;

[0079] The three-dimensional deep learning network FU-Net is used to perform image segmentation on the coronary artery CT image to be segmented to obtain a segmentation result;

[0080] In this embodiment, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program as the coronary artery image segmentation method based on a three-dimensional deep learning network according to any one of the first aspects;

[0081] The computer device may be a laptop computer, a desktop computer or a workstation.

[0082] The processor may be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or an off-the-shelf field programmable gate array (FPGA).

[0083] The memory described in the present invention may be an internal storage unit of a notebook computer, a desktop computer or a workstation, such as a memory or a hard disk; or an external storage unit, such as a mobile hard disk or a flash memory card.

[0084] In this embodiment, a computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the coronary artery image segmentation method based on a three-dimensional deep learning network described in the present invention can be implemented.

[0085] Computer-readable storage media may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer-readable storage media may include: read-only memory (ROM), random access memory (RAM), solid-state drive (SSD) or optical disk, etc. Among them, random access memory may include resistance random access memory (ReRAM) and dynamic random access memory (DRAM).

[0086] In summary, the present invention proposes a three-dimensional deep learning network FU-Net, which can generate attention features of adaptive frequency features. The network can efficiently select relevant information of coronary arteries from multi-space and multi-frequency features, enhance the fusion of frequency and spatial features at different levels, and obtain effective semantic representation. Three core modules are proposed: FCA module, FEM module and DAM module. The FCA module is designed to fuse the features of the encoder and decoder while suppressing irrelevant background noise. The DAM module is set before downsampling in the network, and effectively suppresses high-frequency components that may cause confusion by implicitly and dynamically adjusting the weights of the feature map. In addition, the FEM module is used to learn more frequency representations to learn the complex vascular structure of the coronary artery. Compared with other state-of-the-art methods, the method proposed in the present invention uses the collected CCTA dataset and achieves good performance results in the coronary artery image segmentation task. A large number of experimental results show that the network has good potential in solving the coronary artery image segmentation task.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A coronary artery image segmentation method based on a three-dimensional deep learning network, characterized in that: The specific steps include: Step 1: Acquire a coronary artery CT image dataset; the coronary artery CT image dataset includes a plurality of coronary artery CT images and their corresponding coronary artery labels; Step 2: preprocessing the acquired coronary artery CT image dataset to obtain a preprocessed coronary artery CT image dataset; Step 3: Construct a three-dimensional deep learning network FU-Net; Step 4: Using the preprocessed coronary artery CT image dataset to train the three-dimensional deep learning network to obtain a trained three-dimensional deep learning network; Step 5: Obtain the coronary artery CT image to be segmented and input it into the trained three-dimensional deep learning network to obtain the segmentation result.

2. The coronary artery image segmentation method based on a three-dimensional deep learning network according to claim 1, characterized in that: The preprocessing in step 2 is to truncate the grayscale interval of the coronary artery CT image and retain the set grayscale range.

3. The coronary artery image segmentation method based on a three-dimensional deep learning network according to claim 1, characterized in that: The three-dimensional deep learning network adopts a U-Net encoder-decoder architecture, in which a de-obfuscation module is added before each downsampling stage of the encoder, a frequency enhancement module is applied to the decoder to enhance the image through frequency features, and a frequency cross-attention module is used to replace the skip connection between the encoder and the decoder; The de-obfuscation module is a dynamic learnable mask used to reduce frequency components higher than the Nyquist frequency in the coronary artery CT image; the Nyquist frequency is half of the sampling frequency; The frequency cross attention module first obtains the frequency representations G and X corresponding to the high-level feature g and the low-level feature x output at a certain stage in the decoder after fast Fourier transformation; then obtains the query matrix by linear projection of the frequency representation G corresponding to the high-level feature g, and obtains the key matrix and the value matrix by linear projection of the frequency representation X corresponding to the low-level feature x; multiplies the query matrix Q with the key matrix K by Hadamard product to obtain the attention matrix A; multiplies the attention matrix A with the value matrix V by Hadamard product again to obtain the weight sum Y, and finally converts the weight sum Y to the spatial domain by inverse Fourier transform to obtain the result y in the spatial domain; The frequency enhancement module adopts a two-branch structure, including a frequency domain branch and a space domain branch, for simultaneously extracting features in the frequency domain and the space domain, and then converting the frequency domain features f fre and the spatial domain feature f spa By adding element by element, the frequency-enhanced output is obtained.

4. The coronary artery image segmentation method based on a three-dimensional deep learning network according to claim 3, characterized in that: The frequency domain branch receives the input f obtained by concatenating the output of the previous stage of the decoder and the spatial domain result y obtained by the corresponding frequency cross-attention module in And through Fourier transform, we can get the amplitude F of its frequency representation F amp and phase angle F phi , which represents the amplitude of F with respect to frequency amp and phase angle F phi Linear networks are used to extract features, and then an inverse Fourier transform iFFT is used to obtain the frequency domain features f after filtering irrelevant frequencies. fre .

5. The coronary artery image segmentation method based on three-dimensional deep learning network according to claim 3, characterized in that: The spatial domain branch receives the input f obtained by concatenating the output of the previous stage of the decoder and the spatial domain result y obtained by the corresponding frequency cross attention module in , and the concatenated input f in Perform two convolution operations to extract the relevant coronary artery spatial features and obtain the spatial domain feature f spa .

6. The coronary artery image segmentation method based on three-dimensional deep learning network according to claim 3, characterized in that: The dynamic learnable mask is expressed as: Among them, M(u,v) is the mask value at the frequency component (u,v), u is the frequency value of the first component in the frequency domain, v is the frequency value of the second component in the frequency domain, α is the learnable weight, and R is the range of frequency components less than the Nyquist frequency NF: R={(u,v)|u≤NF and v≤NF} (3).

7. A coronary artery image segmentation system based on a three-dimensional deep learning network, used to implement the coronary artery image segmentation method based on a three-dimensional deep learning network according to any one of claims 1 to 6, characterized in that: Includes data acquisition module and 3D deep learning network FU-Net; The data acquisition module is used to acquire the coronary artery CT image to be segmented as the input of the three-dimensional deep learning network FU-Net; The three-dimensional deep learning network FU-Net is used to perform image segmentation on the coronary artery CT image to be segmented to obtain a segmentation result.

8. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program according to any one of claims 1 to 6, which is a coronary artery image segmentation method based on a three-dimensional deep learning network.

9. A computer-readable storage medium, characterized in that: A computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the coronary artery image segmentation method based on a three-dimensional deep learning network described in any one of claims 1 to 6 is implemented.