Coronary CTA segmentation method and system based on multi-scale feature learning network

By adopting a deep learning network based on the U-Net architecture in coronary CTA image segmentation, the problem of difficulty in extracting multi-scale structural features in the prior art is solved, and fast and accurate coronary CTA image segmentation is achieved.

CN113920132BActive Publication Date: 2025-06-06HANGLOK-TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111117885.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-23
Publication Date
2025-06-06
Estimated Expiration
2041-09-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract multi-scale structural features in coronary CTA image segmentation, resulting in inaccurate vascular segmentation.

Method used

A deep learning network based on U-Net architecture is adopted to extract multi-scale features through encoding modules and decoding modules, and the segmentation accuracy is improved by using data expansion and Dice loss functions.

Benefits of technology

Fast and accurate coronary CTA image segmentation is achieved, and the segmentation performance of blood vessels at different scales is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113920132B_ABST
    Figure CN113920132B_ABST
Patent Text Reader

Abstract

The present invention relates to a technical solution of a coronary CTA segmentation method and system based on a multi-scale feature learning network, comprising: acquiring a coronary CTA image, extracting a sub-block image of a set size from the coronary CTA image, and performing data expansion on the sub-block image to obtain a training set; training the training set through a deep learning network based on multi-scale feature extraction of a U-Net architecture to obtain a segmentation model for coronary CTA; performing segmentation processing on the acquired coronary CTA image through the segmentation model to predict a segmented coronary CTA image. The beneficial effect of the present invention is: achieving fast and accurate coronary CTA segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer and medical treatment, and in particular to a coronary CTA segmentation method and system based on a multi-scale feature learning network. Background Art

[0002] Cardiovascular disease has become one of the leading causes of death in the world. Thanks to recent advances in multi-detector CT technology, 3D CT angiography (CTA) has become the standard examination method for this disease. Extracting the vascular structure of tubular arteries is an important step in detecting and analyzing vascular abnormalities and lesions such as aneurysms, stenoses, and plaques. In addition, accurate and complete vascular structures can provide an important basis for hemodynamic analysis, functional assessment, and interventional surgery planning.

[0003] With the growing clinical needs, manual extraction of coronary artery structure is a tedious and time-consuming process, which is not feasible in clinical practice. Therefore, computer-assisted semi-automatic or automatic segmentation methods have become the main way of vascular segmentation. Despite a lot of research in the past, computer-assisted segmentation of coronary artery structure remains a difficult task. This is usually due to the low image resolution, the presence of (motion) artifacts in the data, the presence of adjacent structures with similar intensity to the blood vessels, and lesions such as stenosis and calcification. At the same time, in CTA images, the coronary arteries have the characteristics of large scale changes and complex topology, which increases the difficulty of coronary artery segmentation in CTA.

[0004] Existing coronary artery segmentation algorithms can be divided into traditional vascular segmentation algorithms and deep learning-based segmentation methods. Traditional algorithms mainly include: threshold-based methods, deformation model-based methods and statistical model-based methods. The main idea is to artificially formulate some rules that conform to vascular characteristics, but artificial rules may not be applicable to all complex vascular structures. The main idea of ​​the deep learning-based algorithm is to design a deep neural network model to train on labeled vascular data, automatically extract vascular features and achieve vascular segmentation. Although deep learning-based methods have shown good results in the field of vascular segmentation, it is still difficult to effectively extract multi-scale structural features, resulting in inaccurate vascular segmentation. Summary of the invention

[0005] The object of the present invention is to solve at least one of the technical problems existing in the prior art, and provides a coronary CTA segmentation method based on a multi-scale feature learning network, characterized in that the method comprises: S100, acquiring a coronary CTA image, extracting a sub-block image of a set size from the coronary CTA image, and performing data expansion on the sub-block image to obtain a training set; S200, training the training set through a deep learning network based on multi-scale feature extraction of a U-Net architecture to obtain a segmentation model for coronary CTA; S300, performing segmentation processing on the acquired coronary CTA image through the segmentation model to obtain a segmented coronary CTA image.

[0006] According to the harmful website detection method based on generative adversarial network and deep learning, S100 includes: extracting sub-block images of 64x64x64 cubic pixels from the acquired coronary CTA images, increasing data under different transformations by data enhancement, and improving the generalization ability of the network.

[0007] According to the detection method of harmful websites based on generative adversarial networks and deep learning, the deep learning network for multi-scale feature extraction based on the U-Net architecture includes an encoding module and a decoding module, the encoding module includes a first encoding module and a second encoding module, the first encoding module is used to process the training set through batch normalization and RELU activation function to obtain an output feature map, and the spatial size of the output feature map is similar to the input feature map. Figure 1 The second encoding module receives the feature map output by the first encoding module, transmits it in parallel to several branches with different convolution kernel sizes, and connects them in series at the end of training; the decoding module is a cascaded dilated convolution, which obtains receptive fields of different sizes through dilated convolution and learns the contextual features of the CTA image.

[0008] According to the detection method of harmful websites based on generative adversarial networks and deep learning, the first encoding module is configured as: a U-Net-based encoding module, including a 3x3x3 convolution, the 3x3x3 convolution of the first encoding module has a hole rate of 1, a step size of 1, and a 3x3x3 convolution using batch normalization and RELU activation function to obtain an output feature map, and the output feature map space size is the same as the input feature map. Figure 1 To.

[0009] According to the detection method of harmful websites based on generative adversarial networks and deep learning, the first encoding module includes: an Inception-like sparsely connected architecture, the first encoding module is composed of a first branch, a second branch and a third branch; the first branch is composed of a maximum pooling layer and a 1x1x1 convolution in series, and the 1x1x1 convolution of the first branch has a hole rate of 1, and a step size of 2; the second branch is composed of a 1x1x1 convolution and a 3x3x3 convolution in series, the 1x1x1 convolution of the second branch has a hole rate of 1, and a step size of 1, and the hole rate of the 3x3x3 convolution of the second branch is 1, and the step size is 2; the third branch is composed of a 1x1x1 convolution, a 1x1x5 convolution, and a 1x The first encoding module performs downsampling processing by convolution or maximum pooling with a step size of 2 through the first branch, the second branch, and the third branch to obtain a feature map with the same scale; the feature maps output by the first branch, the second branch, and the third branch are combined through a series operation, and finally the output feature map is obtained by batch normalization and RELU activation function.

[0010] According to the detection method of harmful websites based on generative adversarial networks and deep learning, the decoding module includes: adopting an Inception-like sparsely connected architecture, the decoding module includes a fourth branch, a fifth branch, a sixth branch and a seventh branch, the fourth branch, the fifth branch, the sixth branch and the seventh branch are convolution branches with different receptive fields; the fourth branch is a 3x3x3 convolution, the fourth branch 3x3x3 convolution has a hole rate = 1, and a step size = 1; the fifth branch is composed of a 3x3x3 convolution and a 1x1x1 convolution in series, the 3x3x3 convolution of the fifth branch has a hole rate = 3, a step size = 1, and the 1x1x1 convolution of the fifth branch has a hole rate = 1, and a step size = 1; the The sixth branch is composed of 3x3x3 convolution, 3x3x3 convolution and 1x1x1 convolution in series, wherein the hole rate of the 3x3x3 convolution of the sixth branch is 1 or 3, and the step size is 1, and the hole rate of the fifth branch is 1, and the step size is 1; the seventh branch is composed of 3x3x3 convolution, 3x3x3 convolution, 3x3x3 convolution and 1x1x1 convolution in series, wherein the hole rate of the seventh branch is 1 or 3 or 5, and the step size is 1, wherein the hole rate of the 1x1x1 convolution of the seventh branch is 1, and the step size is 1; the decoding module uses residual connection to prevent gradient disappearance, and integrates the input feature map and the output feature maps of the four branches in series, and finally obtains the output feature map using batch normalization, RELU activation function and upsampling.

[0011] According to the harmful website detection method based on generative adversarial network and deep learning, a 1x1x1 convolution layer is further provided at the tail of the deep learning network for multi-scale feature extraction based on the U-Net architecture, and the 1x1x1 convolution layer is processed by the Sigmoid activation function to obtain a blood vessel segmentation prediction result.

[0012] According to the harmful website detection method based on generative adversarial networks and deep learning, the training process of the deep learning network based on multi-scale feature extraction of the U-Net architecture uses the Adam optimizer and the Dice loss function to perform corresponding learning processing.

[0013] The technical solution of the present invention also includes a coronary CTA segmentation system based on a multi-scale feature learning network, wherein the system includes: an image pre-acquisition module, used to acquire coronary CTA images, extract sub-block images of a set size from the coronary CTA images, and perform data expansion on the sub-block images to obtain a training set; a training module, used to train the training set through a deep learning network based on multi-scale feature extraction of a U-Net architecture to obtain a segmentation model for coronary CTA; a prediction module, used to perform segmentation processing on the acquired coronary CTA image through the segmentation model to obtain a segmented coronary CTA image.

[0014] The beneficial effect of the present invention is: achieving rapid and accurate coronary CTA segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The present invention is further described below in conjunction with the accompanying drawings and embodiments;

[0016] Figure 1 Shown is a flow chart of a method according to an embodiment of the present invention;

[0017] Figure 2 Shown is a deep learning network diagram according to an embodiment of the present invention;

[0018] Figure 3 Shown is a U0 module network diagram according to an embodiment of the present invention;

[0019] Figure 4 Shown is a network diagram of a U1 module according to an embodiment of the present invention;

[0020] Figure 5 Shown is a U2 module network diagram according to an embodiment of the present invention;

[0021] Figure 6 Shown is a system block diagram according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] This section will describe in detail the specific embodiments of the present invention. The preferred embodiments of the present invention are shown in the accompanying drawings. The purpose of the accompanying drawings is to supplement the description of the text part of the specification with graphics, so that people can intuitively and vividly understand each technical feature and the overall technical solution of the present invention, but it cannot be understood as a limitation on the scope of protection of the present invention.

[0023] In the description of the present invention, “several” means one or more, “more” means more than two, “greater than”, “less than”, “exceed”, etc. are understood as not including the number itself, and “above”, “below”, “within”, etc. are understood as including the number itself.

[0024] In the description of the present invention, the consecutive numbering of the method steps is for the convenience of review and understanding. Combined with the overall technical scheme of the present invention and the logical relationship between the various steps, adjusting the implementation order between the steps will not affect the technical effect achieved by the technical scheme of the present invention.

[0025] In the description of the present invention, unless otherwise clearly defined, words such as "setting" should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meaning of the above words in the present invention in combination with the specific content of the technical solution.

[0026] Glossary:

[0027] conv, convolution;

[0028] rate, void rate;

[0029] stride

[0030] Concatenation, cascade;

[0031] ReLU, a nonlinear function;

[0032] Upsample, upsample.

[0033] Figure 1 The figure shows a flow chart of a method according to an embodiment of the present invention, which includes: acquiring a coronary CTA image, extracting a sub-block image of a set size from the coronary CTA image, and performing data expansion on the sub-block image to obtain a training set; training the training set through a deep learning network based on multi-scale feature extraction of a U-Net architecture to obtain a segmentation model for coronary CTA; performing segmentation processing on the acquired coronary CTA image through the segmentation model to predict a segmented coronary CTA image.

[0034] Figure 2 The figure shows a deep learning network diagram according to an embodiment of the present invention. In this embodiment, a sub-block image of 64x64x64 cubic pixels is first extracted from the acquired coronary CTA image, and then data under different transformations is increased by data enhancement, thereby improving the generalization ability of the network. The training data is then trained on the constructed network model. The training process uses the Adam optimizer and the Dice loss function, and the optimal model is saved during the training. Finally, the training model is applied to the test data to obtain the final segmentation result.

[0035] The network is based on U-Net architecture, and two Inception-like modules are designed as encoding modules and decoding modules respectively embedded in the U-Net network, extracting multi-scale semantic features from different receptive fields, capturing more high-level semantic features in the encoder-decoder and retaining more spatial information in the decoder, thereby improving the network's segmentation performance for blood vessels of different scales. The architecture of the two Inception-like modules uses branches with different convolution kernel sizes in parallel, obtains rich high-level contextual information through a deeper and wider network, and retains detailed spatial information, achieving accurate segmentation of coronary arteries in CTA images.

[0036] Figure 3 The figure shows the network diagram of the U0 module according to an embodiment of the present invention. U0 is the original encoding module in U-Net, which has only one 3x3x3 convolution (void rate = 1; step size = 1). The output feature map is obtained by batch normalization and RELU activation function. The spatial size of the output feature map is similar to that of the input feature map. Figure 1 To.

[0037] Figure 4 FIG. 1 is a network diagram of a U1 module according to an embodiment of the present invention. The U1 module adopts an Inception-like sparsely connected architecture. The input feature map is transmitted in parallel to several branches with different convolution kernel sizes, and then connected in series to serve as an encoder module. Figure 4 As shown. The U1 module includes three branches, branch one is composed of a maximum pooling layer and a 1x1x1 convolution (void rate = 1; step size = 2) in series, branch two is composed of a 1x1x1 convolution (void rate = 1; step size = 1) and a 3x3x3 convolution (void rate = 1; step size = 2) in series, and branch three is composed of a 1x1x1 convolution (void rate = 1; step size = 2), a 1x1x5 convolution (void rate = 1; step size = 1), a 1x5x1 convolution (void rate = 1; step size = 1) and a 5x1x1 convolution (void rate = 1; step size = 1) in series. For the large convolution kernel of this module, the present invention uses asymmetric convolutions such as 1x1x5, 1x5x1 and 5x1x1 to replace the symmetric convolution of 5x5x5, so that the depth of the network can be increased while reducing the parameters of the network. The three branches use convolution or maximum pooling with a step size of 2 to achieve downsampling, so that they have feature maps of the same size. Since the three branches have convolution kernels of different sizes, effective extraction of multi-scale features is achieved. The feature maps output by the three branches are combined using a series operation, and finally the output feature map is obtained using batch normalization and RELU activation function.

[0038] Figure 5 FIG. 1 is a network diagram of a U2 module according to an embodiment of the present invention. Similarly, the U2 module also adopts an Inception-like sparsely connected architecture for a decoder module, such as Figure 5As shown in Figure 2, unlike the U1 module which uses a large convolution kernel, the U2 module uses cascaded dilated convolutions to obtain receptive fields of different sizes, thereby learning contextual features of different scales. The U2 module contains four convolution branches with different receptive fields. Branch one is a 3x3x3 convolution (voiding rate = 1; stride = 1), branch two is composed of 3x3x3 convolution (voiding rate = 3; stride = 1) and 1x1x1 convolution (voiding rate = 1; stride = 1) in series, branch three is composed of 3x3x3 convolution (voiding rate = 1; stride = 1), 3x3x3 convolution (voiding rate = 3; stride = 1) and 1x1x1 convolution (voiding rate = 1; stride = 1) in series, and branch four is composed of 3x3x3 convolution (voiding rate = 1; stride = 1), 3x3x3 convolution (voiding rate = 3; stride = 1), 3x3x3 convolution (voiding rate = 5; stride = 1) and 1x1x1 convolution (voiding rate = 1; stride = 1) in series. In addition, in this module, residual connections are used to prevent gradient disappearance, and the input feature map and the output feature maps of the four branches are connected in series. Finally, batch normalization, RELU activation function and upsampling are used to obtain the output feature map.

[0039] After passing the last 1x1x1 convolutional layer of the encoder-decoder, the Sigmoid activation function is used to obtain the vessel segmentation prediction result.

[0040] Figure 6 The system block diagram according to an embodiment of the present invention is shown. Figure 6 As shown, the image acquisition module is used to acquire coronary CTA images, extract sub-block images of a set size from the coronary CTA images, and perform data expansion on the sub-block images to obtain a training set; the training module is used to train the training set through a deep learning network based on multi-scale feature extraction of the U-Net architecture to obtain a segmentation model for coronary CTA; the prediction module is used to perform segmentation processing on the acquired coronary CTA images through the segmentation model to obtain a segmented coronary CTA image.

[0041] It should be appreciated that the method steps in the embodiments of the present invention can be implemented or implemented by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level process or object-oriented programming language to communicate with a computer system. However, if necessary, the program can be implemented in an assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, the program can be run on a programmed ASIC for this purpose.

[0042] Furthermore, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that may be executed by one or more processors.

[0043] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, a RAM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.

[0044] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents physical and tangible objects, including specific visual depictions of physical and tangible objects produced on the display.

[0045] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in the technical field without departing from the purpose of the present invention.

Claims

1. A coronary CTA segmentation method based on multi-scale feature learning network, It is characterized in that The method includes: S100, acquiring a coronary CTA image, extracting a sub-block image of a set size from the coronary CTA image, and performing data expansion on the sub-block image to obtain a training set; S200, training the training set by a deep learning network based on multi-scale feature extraction of a U-Net architecture to obtain a segmentation model for coronary CTA; S300, performing segmentation processing on the acquired coronary CTA image through the segmentation model to obtain a segmented coronary CTA image; The deep learning network for multi-scale feature extraction based on the U-Net architecture includes an encoding module and a decoding module, wherein the encoding module includes a first encoding module and a second encoding module, wherein the first encoding module is used to process the training set through batch normalization and RELU activation function to obtain an output feature map, and the spatial size of the output feature map is consistent with the input feature map; the second encoding module receives the feature map output by the first encoding module, transmits it in parallel to several branches with different convolution kernel sizes, and connects them in series at the end of training; the decoding module is a cascaded dilated convolution, which obtains receptive fields of different sizes through dilated convolution and learns contextual features of CTA images; The first encoding module is configured as follows: A U-Net-based encoding module includes a 3x3x3 convolution, wherein the 3x3x3 convolution of the first encoding module has a hole rate of 1 and a step size of 1, and the 3x3x3 convolution uses batch normalization and a RELU activation function to obtain an output feature map, and the spatial size of the output feature map is consistent with the input feature map; The first encoding module includes: An Inception-like sparsely connected architecture is adopted, wherein the first encoding module is composed of a first branch, a second branch, and a third branch; The first branch is composed of a maximum pooling layer and a 1x1x1 convolution in series, and the 1x1x1 convolution of the first branch has a dilation rate of 1 and a stride of 2; The second branch is composed of a 1x1x1 convolution and a 3x3x3 convolution connected in series, the 1x1x1 convolution of the second branch has a dilation rate of 1 and a stride of 1, and the 3x3x3 convolution of the second branch has a dilation rate of 1 and a stride of 2; The third branch is composed of 1x1x1 convolution, 1x1x5 convolution, 1x5x1 convolution and 5x1x1 convolution in series, the 1x1x1 convolution of the third branch has a hole rate of 1, a step size of 1, a hole rate of 1x1x5 convolution of the third branch, a step size of 2, a hole rate of 1x5x1 convolution of the third branch, a step size of 1, and a hole rate of 1x5x1 convolution of the third branch. The first encoding module performs downsampling processing by convolution or maximum pooling with a step size of 2 on the first branch, the second branch, and the third branch to obtain feature maps with the same scale; The feature maps output by the first branch, the second branch, and the third branch are combined through a series operation, and finally the output feature map is obtained by using batch normalization and a RELU activation function.

2. The coronary CTA segmentation method based on a multi-scale feature learning network according to claim 1, It is characterized in that The S100 includes: The collected coronary CTA images are used to extract sub-block images with a size of 64x64x64 cubic pixels, and data enhancement is used to increase the data under different transformations to improve the generalization ability of the network.

3. The coronary CTA segmentation method based on a multi-scale feature learning network according to claim 1, It is characterized in that The decoding module comprises: An Inception-like sparsely connected architecture is adopted, and the decoding module includes a fourth branch, a fifth branch, a sixth branch, and a seventh branch, and the fourth branch, the fifth branch, the sixth branch, and the seventh branch are convolution branches with different receptive fields; The fourth branch is a 3x3x3 convolution, and the fourth branch 3x3x3 convolution has a dilation rate of 1 and a step size of 1; The fifth branch is composed of a 3x3x3 convolution and a 1x1x1 convolution connected in series, the 3x3x3 convolution of the fifth branch has a dilution rate of 3 and a step size of 1, and the 1x1x1 convolution of the fifth branch has a dilution rate of 1 and a step size of 1; The sixth branch is composed of 3x3x3 convolution, 3x3x3 convolution and 1x1x1 convolution in series, wherein the sixth branch 3x3x3 convolution has a dilation rate of 1 or 3 and a stride of 1, and the sixth branch 1x1x1 convolution has a dilation rate of 1 and a stride of 1; The seventh branch is composed of 3x3x3 convolution, 3x3x3 convolution, 3x3x3 convolution and 1x1x1 convolution in series, wherein the void rate of the 3x3x3 convolution of the seventh branch is 1 or 3 or 5, and the step size is 1, wherein the void rate of the 1x1x1 convolution of the seventh branch is 1, and the step size is 1; The decoding module uses residual connection to prevent gradient disappearance, and integrates the input feature map and the output feature maps of the four branches in series, and finally obtains the output feature map using batch normalization, RELU activation function and upsampling.

4. The coronary CTA segmentation method based on a multi-scale feature learning network according to claim 1, It is characterized in that A 1x1x1 convolution layer is also provided at the tail of the deep learning network for multi-scale feature extraction based on the U-Net architecture. The 1x1x1 convolution layer is processed by a Sigmoid activation function to obtain a blood vessel segmentation prediction result.

5. The coronary CTA segmentation method based on a multi-scale feature learning network according to claim 1, It is characterized in that The training process of the deep learning network for multi-scale feature extraction based on the U-Net architecture uses the Adam optimizer and the Dice loss function to perform corresponding learning processing.

6. A coronary CTA segmentation system based on a multi-scale feature learning network, wherein the system include: An image acquisition module, used for acquiring a coronary CTA image, extracting a sub-block image of a set size from the coronary CTA image, and performing data expansion on the sub-block image to obtain a training set; A training module, used for training the training set through a deep learning network based on multi-scale feature extraction of a U-Net architecture to obtain a segmentation model for coronary CTA; A prediction module, configured to perform segmentation processing on the acquired coronary CTA image through the segmentation model to obtain a segmented coronary CTA image; The coronary CTA segmentation system based on the multi-scale feature learning network also includes an encoding module and a decoding module, wherein the encoding module includes a first encoding module and a second encoding module, wherein the first encoding module is used to process the training set through batch normalization and RELU activation function to obtain an output feature map, and the spatial size of the output feature map is consistent with the input feature map; the second encoding module receives the feature map output by the first encoding module, transmits it in parallel to several branches with different convolution kernel sizes, and connects them in series at the end of training; the decoding module is a cascaded dilated convolution, which obtains receptive fields of different sizes through dilated convolution and learns the contextual features of the CTA image; The first encoding module is configured as follows: A U-Net-based encoding module includes a 3x3x3 convolution, wherein the 3x3x3 convolution of the first encoding module has a hole rate of 1 and a step size of 1, and the 3x3x3 convolution uses batch normalization and a RELU activation function to obtain an output feature map, and the spatial size of the output feature map is consistent with the input feature map; The first encoding module includes: An Inception-like sparsely connected architecture is adopted, wherein the first encoding module is composed of a first branch, a second branch, and a third branch; The first branch is composed of a maximum pooling layer and a 1x1x1 convolution in series, and the 1x1x1 convolution of the first branch has a dilation rate of 1 and a stride of 2; The second branch is composed of a 1x1x1 convolution and a 3x3x3 convolution connected in series, the 1x1x1 convolution of the second branch has a dilation rate of 1 and a stride of 1, and the 3x3x3 convolution of the second branch has a dilation rate of 1 and a stride of 2; The third branch is composed of 1x1x1 convolution, 1x1x5 convolution, 1x5x1 convolution and 5x1x1 convolution in series, the 1x1x1 convolution of the third branch has a hole rate of 1, a step size of 1, a hole rate of 1x1x5 convolution of the third branch, a step size of 2, a hole rate of 1x5x1 convolution of the third branch, a step size of 1, and a hole rate of 1x5x1 convolution of the third branch. The first encoding module performs downsampling processing by convolution or maximum pooling with a step size of 2 on the first branch, the second branch, and the third branch to obtain feature maps with the same scale; The feature maps output by the first branch, the second branch, and the third branch are combined through a series operation, and finally the output feature map is obtained by using batch normalization and a RELU activation function.

Citation Information

Patent Citations

  • U-shaped cavity full-convolution integral segmentation network identification model based on remote sensing image

    CN111160276A

  • Skin ultrasonic image segmentation method based on improved UNet network model

    CN112132813A