3D Blood Vessel and Trachea Segmentation Method and System

Through the method of combining multimodal data generation and dual model structure, using technical means such as three-dimensional attention layer and residual connection, the problems of data imbalance and noise interference in 3D blood vessel and tracheal segmentation are solved, achieving higher segmentation accuracy and robustness.

CN115564782BActive Publication Date: 2025-06-10CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211253872.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-13
Publication Date
2025-06-10
Estimated Expiration
2042-10-13

AI Technical Summary

Technical Problem

The prior art has problems of data imbalance, noise interference and convergence difficulties caused by large data volume in 3D blood vessel and tracheal segmentation, resulting in low segmentation accuracy.

Method used

Using a method of combining multimodal data generation and dual model structure (coarse model and fine model), multimodal data generation does not intercept the original data, a training model including encoder, decoder and three-dimensional attention layer is constructed, and important spatial and channel features are extracted using the three-dimensional attention layer, and segmentation accuracy is improved through technical means such as residual connection and sliding window reasoning.

Benefits of technology

Effectively alleviate the impact of data imbalance, improve the accuracy of 3D vascular or tracheal segmentation, and improve the accuracy and robustness of segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115564782B_ABST
    Figure CN115564782B_ABST
Patent Text Reader

Abstract

The present invention proposes a 3D blood vessel and trachea segmentation method and system. The method is as follows: obtaining 3D blood vessel or trachea data samples; generating multi-modal data from the 3D blood vessel or trachea data samples; the training model includes a coarse model and a fine model; scaling the multi-modal data, and then training the scaled data and labels in the coarse model to segment the target region, restoring the coordinates of the target region to obtain the region of interest of the original image data and its coordinates; cropping the tissue to be segmented on the original image data according to the coordinates of the region of interest to obtain a voxel block corresponding to the region of interest, performing voxel expansion in the six face directions of the voxel block to obtain a training voxel block, generating multi-modal data from the training voxel block, and then training in the fine model. This 3D blood vessel and trachea segmentation method can mitigate the impact of data imbalance during training and effectively improve the accuracy of blood vessel or trachea segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 3D segmentation, and in particular to the segmentation of 3D tubular tissues, and specifically to a method and system for 3D blood vessel and trachea segmentation. Background Art

[0002] There are a large number of tubular structures distributed in the organ systems of the human body, such as pulmonary arteries, pulmonary tracheas, etc. These tubular structure tissues often have fractal characteristics and a tree-like or reticular topological structure. With the development of related imaging devices, the images of tubular tissues (such as blood vessels and tracheas) collected can reach the level of several pixels. Therefore, retaining their fine branch topologies has become a key issue in extracting their structures. Manual extraction of their structures often has the problems of time-consuming and large subjective differences, and the research on automatic analysis methods has become a hot topic. The structure can be extracted through the Hessian matrix, but its defect is that it is easy to miss detections at bifurcations and the operation is complex. When using the gradient vector flow method for extraction, it is insensitive to weak edges. Tubular structures are usually anisotropic, and the level set algorithm based on the active contour model will reduce the evolution speed of the evolution curve due to its inhibitory effect on high curvatures. In recent years, deep learning methods, such as U-Net, have achieved good results in medical image segmentation. However, directly applying relevant networks for segmentation cannot obtain good results because in actual situations, the data has: (1) unbalanced foreground and background samples, (2) noise interference, and (3) convergence difficulties caused by a large amount of data. Summary of the Invention

[0003] In order to overcome the defects existing in the above-mentioned prior art, the purpose of the present invention is to provide a method and system for 3D blood vessel and trachea segmentation.

[0004] In order to achieve the above object of the present invention, the present invention provides a method for 3D blood vessel and trachea segmentation, including the following steps:

[0005] Obtain 3D blood vessel or trachea data samples;

[0006] Generate multi-modal data for the 3D blood vessel or trachea data samples without intercepting the data when generating the multi-modal data;

[0007] Construct a training model, where the training model includes a rough model and a fine model;

[0008] Scale the multi-modal data, and then train the scaled data and labels in the rough model to segment the target area, and restore the coordinates of the target area to obtain the region of interest of the original image data and its coordinates;

[0009] Crop the tissue to be segmented on the original image data according to the coordinates of the region of interest to obtain a voxel block corresponding to the region of interest, perform voxel expansion in the six face directions of the voxel block to obtain a training voxel block, generate multi-modal data from the training voxel block, and then train it in a refined model;

[0010] Perform 3D blood vessel or trachea segmentation on the 3D blood vessel or trachea data to be segmented in the trained training model.

[0011] This 3D blood vessel and trachea segmentation method can mitigate the impact of data imbalance during training and effectively improve the accuracy of blood vessel or trachea segmentation compared to traditional segmentation methods.

[0012] A preferred embodiment of this 3D blood vessel and trachea segmentation method: When scaling the multi-modal data, use trilinear interpolation to scale the multi-modal data and use nearest neighbor interpolation to scale the labels.

[0013] A preferred embodiment of this 3D blood vessel and trachea segmentation method: Adopt sliding window inference during voxel expansion, and use a Gaussian kernel to generate an importance map during the inference process to smooth the inference result. This preferred embodiment can reduce stitching artifacts.

[0014] A preferred embodiment of this 3D blood vessel and trachea segmentation method: During the training of the refined model, select one direction from three directions of the training voxel block with equal probability for training. The three directions are the directions corresponding to the coronal plane, horizontal plane, and sagittal plane of the training voxel block respectively.

[0015] Perform weighted averaging on the feature maps generated by training in the three directions and then activate.

[0016] Multi-angle data training and inference result in higher segmentation accuracy for this preferred embodiment.

[0017] A preferred embodiment of this 3D blood vessel and trachea segmentation method: When training the refined model, enhance the training voxel block. The enhancement methods include at least two types, and each enhancement method corresponds to a probability. Whether to execute the enhancement method corresponding to the probability is controlled according to the probability, and then multi-modal data is generated from the enhanced training voxel block. This increases the diversity of training data.

[0018] A preferred embodiment of this 3D blood vessel and trachea segmentation method: Both the coarse model and the refined model include an encoder, a decoder, and a three-dimensional attention layer;

[0019] There is a corresponding three-dimensional attention layer after each layer of the encoder and decoder to extract important spatial and channel features;

[0020] The three-dimensional attention layer includes a three-dimensional spatial attention module and a three-dimensional channel attention module;

[0021] The three-dimensional channel attention module enables the network to focus on important feature channels while suppressing feature channels irrelevant to the current task;

[0022] The three-dimensional spatial attention module enables the network to focus on the feature maps at important positions while suppressing features at positions irrelevant to the current task;

[0023] The feature map results of the three-dimensional channel attention module and the three-dimensional spatial attention module are added together to obtain the fused feature map.

[0024] A preferred solution of the 3D blood vessel and trachea segmentation method: Add a residual connection after each layer of the encoder, so that the output of each layer of the encoder is y = H(x) + x, where y represents the output feature map of the network layer, x represents the feature map input to the network layer, and H(x) represents the result of the linear transformation of the feature map x input to the network layer. The existence of x preserves the shallow-layer features of small blood vessels or tracheas in the network.

[0025] A preferred solution of the 3D blood vessel and trachea segmentation method: The upper part of the decoder uses a transposed convolution with kernel = 2×2×2 and stride = 2×2×2;

[0026] The decoder part consists of two anisotropic convolutions with kernel sizes of 3×3×1 and 1×1×3 respectively.

[0027] In this preferred solution, the transposed convolution is used to mitigate the impact of the checkerboard effect, and the anisotropic convolution is used to adapt to the three-dimensional geometric deformation of blood vessels or tracheas.

[0028] A preferred solution of the 3D blood vessel and trachea segmentation method: The input channels of each layer of the three-dimensional attention layer are the same as the output channels of each layer of the encoder / decoder;

[0029] After passing through the three-dimensional attention layer, the output feature map is dimensionally reduced. For binary classification tasks, sigmoid is used to activate the network output, and for multi-classification tasks, softmax is used for activation. Then, the class with the highest probability is taken as the segmentation result.

[0030] The present invention also provides a 3D blood vessel and trachea segmentation system, including a processor and a memory. The processor and the memory are communicatively connected. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the 3D blood vessel and trachea segmentation method as described above. This 3D blood vessel and trachea segmentation system has all the advantages of the above 3D blood vessel and trachea segmentation method.

[0031] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present invention. Description of the Drawings

[0032] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0033] Figure 1 is a schematic flowchart of a 3D blood vessel and trachea segmentation method;

[0034] Figure 2 is a schematic diagram of a 3D segmentation network based on a three-dimensional attention mechanism;

[0035] Figure 3 is a schematic diagram of a three-dimensional attention layer network;

[0036] Figure 4 is the segmentation result of the pulmonary artery;

[0037] Figure 5 is the segmentation result of the pulmonary trachea;

[0038] Figure 6 is the training pulmonary artery loss decline graph. Detailed implementation manners

[0039] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0040] In the description of the present invention, unless otherwise specified and defined, it should be noted that the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it may be a mechanical connection or an electrical connection, or it may be the communication inside two elements. It may be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.

[0041] As Figure 1 shown, the present invention provides an embodiment of a 3D blood vessel and trachea segmentation method, and this embodiment includes the following steps:

[0042] Obtain 3D blood vessel or trachea data samples;

[0043] Perform multi-modal data generation on the 3D blood vessel or trachea data samples;

[0044] When performing multi-modal generation, in order not to lose the information of the original data, no truncation is performed when generating the corresponding modal data. Therefore, the value of each pixel of different modal data is determined by the formula, The meaning of this formula is a method for intensity transformation of images (especially CT images), that is, the way to convert the Intensity of a certain point into Pixel. Among them, Intensity represents the intensity value of the CT image, subscript max and subscript min represent the maximum and minimum intensities in the data, Pixel is the result after conversion, and Intensity max and Intensity min are set according to the intensity distributions of different tissues.

[0045] Construct a training model, and the training model includes a coarse model and a fine model.

[0046] In this embodiment, both the coarse model and the fine model include an encoder, a decoder, and a three-dimensional attention layer. In this embodiment, as Figure 2 shown, there are 4 layers in total for the encoder and the decoder. After each layer of the encoder or the decoder, there is a corresponding three-dimensional attention layer to extract important spatial and channel features. Residual connections are added after each layer of the encoder, so that the output of each layer of the encoder is y = H(x) + x, where y represents the output feature map of a certain network layer, x represents the feature map input to the network layer, and H(x) represents the result of linear transformation of the feature map x input to the network layer. Due to the existence of x, the shallow features of small blood vessels or tracheas are retained.

[0047] The difference between the upper part of the decoder and the traditional U-Net is that a transposed convolution with kernel = 2×2×2 and stride = 2×2×2 is used to mitigate the impact brought by the checkerboard effect; the decoder part is two anisotropic convolutions with kernel sizes of 3×3×1 and 1×1×3 respectively to adapt to the three-dimensional geometric deformation of blood vessels or tracheas.

[0048] The three-dimensional attention layer includes a three-dimensional spatial attention module and a three-dimensional channel attention module; the three-dimensional channel attention module is used to make the network focus on important feature channels and at the same time suppress the feature channels irrelevant to the current task; the three-dimensional spatial attention module is used to make the network focus on the feature maps at important positions and at the same time suppress the features at positions irrelevant to the current task.

[0049] After the three-dimensional convolutional layer in the network, activation is performed in the way of InstanceNorm + ReLu. Because anisotropic convolutions are used, the convolutional methods of each layer of the encoder and the decoder are slightly different. The types of three-dimensional convolutional layers of the specific encoder and decoder are shown in Table 1.

[0050] Table 1 Types of three-dimensional convolutional layers

[0051]

[0052] Training the training model includes training the rough model and the refined model.

[0053] When training the rough model, the generated multi-modal data is scaled. To ensure the correct annotation of the scaled image, trilinear interpolation is used to scale the multi-modal data, and nearest-neighbor interpolation is used to scale the labels. During training, the scaled data and labels are combined into 5D data ([B, C, H, W, D]) as a whole according to corresponding rules for training, the target area is segmented, and the coordinates of the target area are restored to obtain the region of interest of the original image data and its coordinates. Here, the corresponding rules refer to: after the scaling process, the multi-modal data and its labels are discrete and unordered. Therefore, we need to organize them in the form of a matrix, and splice the 4D three-dimensional data in the B dimension, that is, the batch dimension, similar to queuing. This process does not change the content of the three-dimensional data, but organizes them in an orderly manner to form five-dimensional data ([B, C, H, W, D]). Among them, B is the batch dimension, indicating how many three-dimensional data are in the matrix; C is the channel of the three-dimensional image, H is the height of the three-dimensional image, W is the width of the three-dimensional image, and D is the depth of the three-dimensional image.

[0054] The substances contained in the medical image data before and after scaling are the same. According to the formula the region of interest on the original image data can be obtained. In this formula, Original Coordinate represents the original voxel coordinates, Resized Coordinate represents the scaled voxel coordinates, Resized Spacing represents the voxel spacing after scaling, Original Spacing represents the original voxel spacing. The meaning of the formula is that the scaled voxel coordinates can be converted to the original voxel coordinates through the voxel spacing of the scaled image and the original voxel spacing.

[0055] When training the refined model, the training data is generated according to the labels, that is, the tissue to be segmented is cropped on the original image data according to the coordinates of the region of interest corresponding to the labels to obtain the voxel blocks corresponding to the region of interest. For the robustness of the algorithm, voxel expansion is performed in the 6 face directions of the voxel blocks according to different tasks to obtain the training voxel blocks. For example, in the task of segmenting small targets, such as aneurysms, expand the range by 5 voxels, that is: expand 5 voxels in the height direction, expand 5 voxels in the width direction, and expand 5 voxels in the depth direction; in the task of segmenting large targets, such as the liver, it may expand the range by 10 voxels, that is: expand 10 voxels in the height direction, expand 10 voxels in the width direction, and expand 10 voxels in the depth direction. The specific parameters of the expansion are set according to different segmentation tasks.

[0056] In this embodiment, sliding window inference is adopted during voxel expansion, and the sliding windows overlap according to a set ratio. The default overlap ratio in the system of this embodiment is preferably but not limited to 0.25. The feature results of the overlapping regions are averaged according to the number of overlaps, that is Among them, the parameter Overlap feature is the feature map of the overlapping features after weighting; feature i is the feature map i of the overlapping part, and j represents that there are j overlapping feature maps in total. This formula means that during sliding window inference, the values of the feature maps in the overlapping part are obtained by adding and averaging the feature maps of each sliding window in the overlapping region.

[0057] During the inference process, a Gaussian kernel is used to generate an importance map to smooth the inference result. The probability density function of the Gaussian distribution is σ is the standard deviation, and σ in the formula is taken according to the size of the voxel block. σ = 0.125 * patchsize, μ = 0, where patchsize is the size of the sliding block of the sliding window. This size is the information of [height, width, depth], and each dimension has its corresponding σ. During smoothing, smoothing is performed separately from these three dimensions. Therefore, x in this formula takes values within the range determined by the dimension value of the height or width or depth of the feature map. For example, if the dimension size of the height of the feature map is 192, then the value range of x is [-96, 96]. Since in this embodiment, a patch-based training strategy is adopted during training, during inference, in order to ensure the continuity and consistency of the inference result, here sliding window inference is performed, and the segmentation result is smoothed by a Gaussian kernel to reduce stitching artifacts and make it friendly for clinical diagnosis. According to the above method, a set number of 3D vascular or tracheal data samples are selected. The set number of 3D vascular or tracheal data samples can be set according to the specific task. If the 3D vascular or tracheal data samples are relatively large, a smaller number is set. If the 3D vascular or tracheal data samples are relatively small, a larger number is set. The purpose of this is to make full use of computing resources. Then, a specified number of training voxel blocks of the same size are cropped from each selected sample. The number of training voxel blocks that can be cropped from each voxel block is determined by the size of the cropped voxel block. Patch num =(Volume H -Patch H )×(Volume W -Patch W )×(Volume D -Patch D ), where Patch num is the number of voxel blocks, Volume H is the height of the original voxel data, Patch His the height of the cropped voxel block, Volume W is the width of the original voxel data, Patch W is the width of the cropped voxel block, Volume D is the depth of the original voxel data, Patch D is the depth of the cropped voxel block. The meaning of this formula is that the number of cropped voxel blocks is determined by the height, width, and depth of the cropped voxel block and the height, width, and depth of the original voxel data. The cropped voxel blocks control the attributes of their positive and negative samples according to a certain ratio according to different tasks. Here, the ratio setting is determined according to the task because it involves the balance of accuracy and generalization. Specifically, when implementing, choose one of the ratios according to the task requirements: positive sample: negative sample > 1:1, the recall rate and accuracy are generally higher, but there may be false positives; positive sample: negative sample = 1:1 is the default, which is a relatively balanced setting; positive sample: negative sample < 1:1, there are fewer false positives and better generalization, but the recall rate may decrease because it is difficult to learn the distribution of positive samples. Then, the cropped training voxel blocks are used to generate multi-modal data and combined into 5D data ([B, C, H, W, D]) for training in the fine model. When training the fine model, one direction is selected from the 3 directions of the training voxel block with equal probability for training. These 3 directions are the directions corresponding to the coronal plane, horizontal plane, and sagittal plane of the voxel block (H, W, D), (D, H, W), (D, W, H) respectively. All three directions will be inferred during inference to generate corresponding feature maps, and the feature maps generated by training in the 3 directions are weighted and averaged and then activated. The value of the finally activated feature map final feature = active func ((feature 1 + feature 2 + feature 3 ) / 3), where active func is the activation function, such as softmax or sigmoid, feature 1 + feature 2 + feature 3 are the feature maps of the coronal plane, horizontal plane, and sagittal plane. The meaning of the formula is that the activated feature of a voxel block is activated after taking the average of the values of its coronal plane, horizontal plane, and sagittal plane and adding them together.

[0058] During training, the training data is augmented. The augmentation methods include random scaling, random Gaussian noise, random Gaussian smoothing, random intensity transformation, etc. To increase the diversity of the data, in this embodiment, the augmentation methods include at least two. The strategy for the training data is not to execute all the augmentation methods sequentially, but each augmentation method has its own independent probability to control whether to execute the method, and the probability is default set to 0.15 in the system.

[0059] During the training of the coarse model and the fine model, dice loss is used for optimization. Among them, TP: True Positive, the prediction is 1 and the label is 1; FP: False Positive, the prediction is 1 and the label is 0; FN: False Negative, the prediction is 0 and the label is 1. For different tasks, whether binary classification or multi-classification, in order to speed up the training, the loss value of the background is not calculated.

[0060] As Figure 3 shown, the input channels of each three-dimensional attention layer are the same as the output channels of each encoder / decoder layer. For the three-dimensional channel attention layer, the execution process of the algorithm is shown in Table 2.

[0061] Table 2 Channel Attention Execution Process

[0062]

[0063]

[0064] For the three-dimensional spatial attention layer, the execution process of the algorithm is shown in Table 3.

[0065] Table 3 Spatial Attention Execution Process

[0066]

[0067] Finally, the results of the two feature maps are added together to obtain the fused feature map.

[0068] After passing through the three-dimensional attention layer, the output feature map is dimensionally reduced. For binary classification tasks, sigmoid is used to activate the network output, and its mathematical expression is Put the values of the feature map into the formula to get the activated feature, which is used for the subsequent binarization operation. The threshold for binarization is 0.5, that is, those greater than or equal to 0.5 are set to 1, and those less than 0.5 are set to 0; for multi-classification, softmax is used for activation, and a probability value is assigned to each channel, and its mathematical expression is Among them, n is the number of channels, and the probability of a certain channel is obtained through this fraction, x iIt is the feature map of this channel, and then the category with the highest probability is taken as the segmentation result. The binary classification task here means that there are two categories in the classification task. For example, segmenting blood vessels and background from a CT image is a binary classification task; the multi-classification task means that there are multiple categories in the classification task. For example, segmenting the liver, kidney, spleen, and pancreas from a CT image is a multi-classification task.

[0069] Train the training model according to the above method. After training is completed, perform 3D blood vessel or trachea segmentation on the 3D blood vessel or trachea data to be segmented in the trained training model.

[0070] This embodiment gives the result diagrams of segmenting the pulmonary artery and pulmonary trachea, as Figure 4 and Figure 5 shown. At the same time, it also gives the training pulmonary artery loss reduction diagram, as Figure 6 shown.

[0071] This application also proposes an embodiment of a 3D blood vessel and trachea segmentation system. In this embodiment, the 3D blood vessel and trachea segmentation system includes a processor and a memory. The processor and the memory are communicatively connected. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the 3D blood vessel and trachea segmentation method as described above. Among them, the 3D blood vessel or trachea data sample can be stored in the memory, or can be acquired by an image acquisition module and stored in the memory.

[0072] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0073] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A 3D blood vessel and trachea segmentation method, characterized in that, it includes the following steps: Obtain 3D blood vessel or trachea data samples; Generate multi-modal data for the 3D blood vessel or trachea data samples without intercepting the data when generating the multi-modal data; Construct a training model, the training model includes a coarse model and a fine model; Both the coarse model and the fine model include an encoder, a decoder and a three-dimensional attention layer; After each layer of encoder and decoder, there is a corresponding three-dimensional attention layer to extract important spatial and channel features; The three-dimensional attention layer includes a three-dimensional spatial attention module and a three-dimensional channel attention module; The three-dimensional channel attention module includes an adaptive three-dimensional pooling layer, a first three-dimensional convolutional layer, a LeakyReLU activation layer, a second three-dimensional convolutional layer and a first sigmoid activation layer that sequentially process the input feature map. The feature output by the first sigmoid activation layer is multiplied by the input feature map by channel and then the feature map result of the three-dimensional channel attention module is output; The three-dimensional spatial attention module includes a three-dimensional convolutional layer and a second sigmoid activation layer that sequentially process the input feature map. The feature output by the second sigmoid activation layer is multiplied by the input feature map and then the feature map result of the three-dimensional spatial attention module is output; The feature map results output by the three-dimensional channel attention module and the three-dimensional spatial attention module are added to obtain a fused feature map; Scale the multi-modal data, and then train the scaled data and labels in the coarse model to segment the target area, restore the coordinates of the target area, and obtain the region of interest of the original image data and its coordinates; Crop the tissue to be segmented on the original image data according to the coordinates of the region of interest to obtain a voxel block corresponding to the region of interest. Perform voxel expansion in the six face directions of the voxel block to obtain a training voxel block. Generate multi-modal data for the training voxel block and then train it in the fine model; When training the fine model, select one direction from the three directions of the training voxel block with equal probability. The three directions are the directions corresponding to the coronal plane, the horizontal plane, and the sagittal plane of the training voxel block respectively, Perform weighted averaging on the feature maps generated by training in the three directions and then activate; Segment the 3D blood vessel or trachea data to be segmented in the trained training model.

2. The 3D blood vessel and trachea segmentation method according to claim 1, characterized in that, When scaling the multi-modal data, perform trilinear interpolation on the multi-modal data for scaling and perform nearest neighbor interpolation on the label for scaling.

3. The 3D blood vessel and trachea segmentation method according to claim 1, characterized in that, During voxel expansion, sliding window inference is adopted, and during the inference process, a Gaussian kernel is used to generate an importance mapping graph to smooth the inference result.

4. The 3D blood vessel and trachea segmentation method according to claim 1, characterized in that, When training a refined model, the training voxel blocks are enhanced. The enhancement methods include at least two kinds, and each enhancement method corresponds to a probability. Whether to execute the enhancement method corresponding to the probability is controlled according to the probability. Then, the enhanced training voxel blocks are used for multi-modal data generation.

5. The 3D blood vessel and trachea segmentation method according to claim 1, wherein, a residual connection is added after each layer of the encoder, so that the output of each layer of the encoder is y = H(x) + x, where y represents the output feature map of the network layer, x represents the feature map input to the network layer, and H(x) represents the result of the linear transformation of the feature map x input to the network layer.

6. The 3D blood vessel and trachea segmentation method according to claim 1, wherein, the upper part of the decoder uses a transposed convolution with kernel = 2×2×2 and stride = 2×2×2; the decoder part is two anisotropic convolutions, and the kernel sizes are 3×3×1 and 1×1×3 respectively.

7. The 3D blood vessel and trachea segmentation method according to claim 1, wherein, the input channels of each layer of the three-dimensional attention layer are the same as the output channels of each layer of the encoder / decoder; after passing through the three-dimensional attention layer, the output feature map is dimensionally reduced. For a binary classification task, sigmoid is used to activate the network output, and for a multi-classification task, softmax is used for activation. Then, the class with the highest probability is taken as the segmentation result.

8. A 3D blood vessel and trachea segmentation system, wherein, it includes a processor and a memory. The processor and the memory are communicatively connected. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the 3D blood vessel and trachea segmentation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Blood vessel image segmentation method and device based on deep learning, equipment and medium

    CN113205537A

  • Blood vessel image segmentation method and device based on CRDNet

    CN113205538A