A method and system for cardiac vessel image segmentation

By combining convolutional coding networks and capsule coding networks with Transformer encoders, the problem of discontinuous segmentation results in cardiac and vascular image segmentation is solved, achieving more efficient and accurate cardiac and vascular segmentation.

CN117058175BActive Publication Date: 2025-11-25GUANGDONG UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311018267.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2025-11-25
Estimated Expiration
2043-08-11

AI Technical Summary

Technical Problem

Existing technologies neglect spatial and semantic information between features in cardiac vascular image segmentation, resulting in discontinuous segmentation results, undersegmentation, and unclear edge segmentation.

Method used

Convolutional coding networks are used to extract feature information at different scales and levels. Capsule coding networks are combined to perform multiple feature extractions and dimensional transformations. The high-level semantic feature map is cut into image patches through the Patch Embedding layer. The Transformer encoder is used to perform self-attention mechanism fusion and multiple skip connections are made to obtain contextual information, finally obtaining the segmentation result.

Benefits of technology

It improves the accuracy and continuity of cardiac and vascular image segmentation, reduces computational load, lowers training costs, and enhances segmentation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058175B_ABST
    Figure CN117058175B_ABST
Patent Text Reader

Abstract

The application discloses a kind of cardiovascular image segmentation method and system, method includes: S1, to heart image is normalized, and obtain preprocessed image;S2, by convolutional coding network is convolved to preprocessed image to extract different scale and level feature information, obtain multi-scale feature map;S3, by several layers of capsule coding network is carried out multiple feature extraction and multiple dimension transformation to multi-scale feature map, obtain high-level semantic feature map;S4, high-level semantic feature map is cut into several image blocks, and using encoder carries out feature extraction, by self-attention mechanism fusion feature information in each image block, adaptively capture context information;S5, to image block is carried out several times up sampling operation, and according to the context information captured is connected by jumping, finally obtain segmentation result.The application realizes higher accuracy, stability and efficiency in heart blood vessel segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medical image segmentation, and particularly relates to a cardiac vessel image segmentation method and system. BACKGROUND

[0002] Deep learning has the potential to improve accuracy, stability and efficiency in cardiac vessel segmentation. Cardiac vessel segmentation is a complex task that requires accurate segmentation of vessels with varying shapes, sizes and directions. Traditional methods require a lot of time and effort to manually select features and parameters for segmentation, which is easily affected by subjective factors and image noise, resulting in inaccurate and unstable segmentation results. Deep learning can automatically learn and extract features from images, reducing the need for human intervention and improving the accuracy, stability and efficiency of segmentation. Deep learning can achieve automatic segmentation, reduce the workload of doctors and improve the consistency and repeatability of segmentation.

[0003] Currently, cardiac vessel image segmentation often uses Unet, TransUent and UCaps models. However, the complex morphological structure of cardiac vessels makes the vessel texture and other tissue texture interlaced in space, which is easy to cause over-segmentation. Traditional network models often ignore the spatial position information and semantic information between features when processing images, resulting in discontinuous segmentation results, such as under-segmentation and unclear edge segmentation. At the same time, in order to achieve a certain accuracy, the current network model has a large model size and a long training time, with high training cost.

[0004] And the application number is 202211126578.2 Chinese invention patent: "a retinal blood vessel segmentation method based on UNet and Transformer fusion", the technical scheme provided by the patent is as follows: including the following steps: step 1, the preprocessed image is obtained by preprocessing the image to be trained; step 2, the preprocessed image is input into the retinal blood vessel segmentation model based on UNet and Transformer fusion to obtain the weight file, the model includes an encoder, a decoder and a fusion attention mechanism, the encoder includes a multi-stream cascaded convolution layer, a plurality of pooling layers and a plurality of residual modules, each convolution layer uses a residual module, and the pooling layer is arranged between two convolution layer units; the decoder includes a plurality of improved residual modules based on extended convolution, a plurality of up sampling modules and a deconvolution layer, and the up sampling module is arranged between two adjacent improved residual units based on extended convolution; the fusion attention mechanism takes the output of the pooling layer in the encoder and the output of the adjacent stream pooling layer as the low layer feature input and the high layer feature input respectively, and the high layer feature input of the fusion attention mechanism in the third layer is the feature map formed by the Transformer module corresponding to the encoder pooling layer; step 3, load the weight file, input the test fundus image into the model to obtain the retinal blood vessel segmentation result. The technical scheme aims to improve the sensitivity of the retinal blood vessel segmentation model to small blood vessels, while the present application aims to solve the problem of heart blood vessel segmentation, and the heart blood vessel has complex morphological structure and variable texture information, so the technical scheme provided by the comparison file cannot realize the accurate segmentation of heart blood vessels. SUMMARY

[0005] In order to solve the problem that the existing technology heart blood vessel image segmentation ignores the spatial position information and semantic information between features, resulting in discontinuous segmentation results and unclear edge segmentation, the present application provides a heart blood vessel image segmentation method and system, and the technical scheme adopted by the present application is:

[0006] The first aspect of the present application provides a heart blood vessel image segmentation method, including the following steps:

[0007] S1, normalizing the heart image to obtain a preprocessed image;

[0008] S2, performing convolution operation on the preprocessed image by convolution coding network to extract feature information of different scales and levels, to obtain a multi-scale feature map;

[0009] S3, performing multiple feature extraction and multiple dimension transformation on the multi-scale feature map by a plurality of layers of capsule coding network, to obtain a high-level semantic feature map;

[0010] S4, cutting the high-level semantic feature map into a plurality of image blocks through a Patch Embedding layer operation, and performing feature extraction on the plurality of image blocks using a Transformer encoder, and fusing feature information in each image block through a self-attention mechanism to adaptively capture context information;

[0011] S5, performing a plurality of upsampling operations on the plurality of image blocks, and performing a plurality of skip connections according to the captured context information, and finally obtaining a segmentation result.

[0012] Compared with the prior art, the method introduces a self-attention mechanism and a capsule encoding network, performs a plurality of feature extractions and dimension transformations on an image through the capsule encoding network to extract higher-level semantic features, and adaptively captures context information in the image through the self-attention mechanism to better preserve spatial position relationships between feature entities, thereby solving the problem of discontinuous segmentation results.

[0013] As a preferred solution, the step S1 specifically comprises:

[0014] The heart image is adjusted to a preset size, pixel values of the heart image are scaled to a preset range, and finally the heart image is converted into a tensor form to obtain a preprocessed image.

[0015] As a preferred solution, in the step S2, the convolutional encoding network comprises a first convolutional layer, a second convolutional layer, a third convolutional layer and an activation layer.

[0016] The first convolutional layer is configured to perform a convolution operation and edge padding on the preprocessed image simultaneously to obtain a first output image.

[0017] The second convolutional layer is configured to perform further edge padding and expand the number of output channels on the first output image to obtain a second output image.

[0018] The third convolutional layer performs further edge padding on the second output, and increases the ability of nonlinear transformation at the network level through a tanh activation function in the activation layer, and finally obtains a multi-scale feature map.

[0019] As a preferred solution, in the step S2, the convolution operation can be represented as:

[0020] out(x)=conv3(conv2(Conv1(x,k1,s1,p1,c1),k2,s2,p2,c2),k3,s3,p3,c3)

[0021] wherein x represents an input, k i , s i , pi , c i respectively represent the convolution kernel size, the convolution step, the edge padding size and the output channel number.

[0022] As a preferred solution, in the step S3, the capsule encoding network of each layer includes a plurality of capsules, the dynamic routing between the capsules is used to calculate the weight distribution between the capsules, the dynamic routing between the capsules is calculated based on the dot product of vectors and a dynamically adjusted manner to calculate the weight of the output capsule, the input vector of the jth capsule of the Ith layer capsule encoding network relative to the ith capsule u i of the (I-1)th layer capsule encoding network can be expressed as:

[0023]

[0024] wherein s ij represents the input vector of the jth capsule v j of the Ith layer capsule encoding network relative to the ith capsule u i of the (I-1)th layer capsule encoding network, W ij represents the weight matrix; then, the weight distribution between the capsules is calculated by the dynamic routing algorithm:

[0025]

[0026] wherein c ij represents the weight of the dynamic routing, b ij represents the iteration variable in the dynamic routing algorithm; finally, the output vector of the jth capsule of the Ith layer capsule encoding network is calculated in a weighted summation manner:

[0027]

[0028] The length of the capsule in the capsule encoding network represents the existence probability or confidence of a certain feature, and the norm of the output vector of the capsule is usually used as the length of the capsule. For the jth unit of the Ith layer capsule encoding network, the length calculation formula is:

[0029]

[0030] wherein d j represents the length of the jth capsule of the Ith layer capsule encoding network, v j represents the output vector of the jth capsule of the Ith layer capsule encoding network.

[0031] As a preferred solution, in the step S3, the multi-scale feature map is subjected to multiple feature extraction and dimension transformation by a 4-layer capsule encoding network, which is expressed as:

[0032] cap out(x) = cap4(cap3(cap2(cap1(x, k1, t1), k2, t2), k3, t3), k4, t4)

[0033] wherein x represents the input of the capsule encoding network, k represents the size of the convolution kernel, and t represents the dimension size of the output capsule feature.

[0034] As a preferred solution, in the step S4, the Patch Embedding layer operation is specifically:

[0035] The high-level semantic feature map is cut into several small blocks of a preset size, and each small block is mapped into a vector space to obtain an initial patch embedding matrix x e R(N 2 ) x d, wherein d represents the dimension size of the vector, and N represents the number of small blocks into which the image is cut; after each small block is expanded into a vector, a d-dimensional vector is obtained through a linear layer, i.e.

[0036] e i = W p *path i +b p

[0037] wherein e i e R d represents the feature vector of the i-th small block, W p e R d x ((patch_size) 2 x c) and b p e R d , respectively, represent the weight and bias of the Patch Embedding layer.

[0038] As a preferred solution, in the step S4, the Transformer encoder is composed of several Transformer block cycles, each Transformer block is composed of several layers of self-attention and MLP layers, and finally each Transformer block outputs a matrix Z with the same dimension as the input X; the k-th Transformer block takes the input matrix X and the matrix Z h -1 output by the last layer as input to obtain an output matrix Z h e R(N 2 ) x d after processing by the self-attention module self_attention and the fully connected layer network MLP:

[0039] Q h = X*W h *Q

[0040] K h =X*W h *K

[0041] V h =X*W h *V

[0042]

[0043] Z h =LayerNorm(X+Dropout(MLP(LayerNorm(A h )+X)))

[0044] Wherein, Q h , K h , V h respectively represent query, key, value matrix, the three matrixes are obtained through linear transformation of matrix X and weight matrix W h *Q, W h *K, W h *V;A h It represents the attention weight matrix obtained by self-attention processing;LayerNorm represents the normalization operation of this layer, Dropout represents the random inactivation operation, and MLP represents the fully connected layer network.

[0045] As a preferred scheme, in the step S5, five upsampling operations are performed on the plurality of image blocks;The first upsampling operation is performed on the result obtained after the low-resolution capsule encoding and the result obtained by the self-attention encoding, and then the result obtained by the upsampling is obtained;The second, third and fourth upsampling operations are respectively performed on the output result of the last upsampling and the third, second and first capsule encoding output results, and then the upsampling is performed;The fifth upsampling operation is performed on the result of the fourth upsampling and the result obtained by the feature extractor, and the final segmentation result is obtained by the jump connection and the upsampling.

[0046] The second aspect of the application also provides a heart blood vessel image segmentation system, comprising a preprocessing module, a convolutional encoding module, a capsule encoding module, a self-attention encoding module and a connection encoding module;

[0047] The preprocessing module is used for normalizing the heart image to obtain a preprocessed image;

[0048] The convolutional encoding module is used for performing convolution operation on the preprocessed image through a convolutional encoding network to extract feature information of different scales and levels, and obtain a multi-scale feature map;

[0049] The capsule encoding module is configured to perform multiple feature extraction and multiple dimension transformation on the multi-scale feature map through a plurality of layers of capsule encoding networks to obtain a high-level semantic feature map.

[0050] The self-attention encoding module is configured to cut the high-level semantic feature map into a plurality of image blocks through a Patch Embedding layer operation, and perform feature extraction on the plurality of image blocks using a Transformer encoder, and fuse feature information in each image block through a self-attention mechanism to adaptively capture context information.

[0051] The connection encoding module is configured to perform a plurality of upsampling operations on the plurality of image blocks, and perform a plurality of skip connections according to the captured context information to finally obtain a segmentation result.

[0052] Compared with the prior art, the present application has the beneficial effects that:

[0053] (1) The convolutional encoding network is used as a feature extractor to preliminarily extract the features of the original image. The convolution has the characteristic of parameter sharing, which can effectively reduce the data volume while extracting the main features. The feature extractor can extract feature information of different scales and levels to provide effective input feature representation for subsequent tasks, which is beneficial to the convergence and training of the network.

[0054] (2) The capsule encoding network is introduced to extract features multiple times and perform dimension transformation through a plurality of layers of capsule encoding stacking to output higher-level semantic features. The feature vector of the capsule encoding retains the spatial position relationship between feature entities, which solves the problem of lack of spatial position information between feature entities to a certain extent.

[0055] (3) The idea of the Transformer is borrowed to divide the image output by the capsule encoding layer into image blocks of a certain size as the input of the self-attention encoding layer to adaptively capture the context information of the image and further improve the accuracy of segmentation. This composite structure not only improves the accuracy, but also reduces the computational amount during the training of the capsule encoding network. BRIEF DESCRIPTION OF DRAWINGS

[0056] Figure 1 A flow chart of a cardiac vascular image segmentation method is provided for the embodiments of the present application.

[0057] Figure 2 A structural schematic diagram of a cardiac vascular image segmentation system is provided for the embodiments of the present application.

[0058] Figure 3 A cardiac vascular segmentation result image is provided for the embodiments of the present application.

[0059] Figure 4The heart blood vessel segmentation prediction label map provided for the embodiment of the present application. DETAILED DESCRIPTION

[0060] The accompanying drawings are only intended to illustrate the application, and should not be construed as limiting the patent;

[0061] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0062] The terms used in the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein means and includes any or all possible combinations of one or more associated listed items.

[0063] The following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not necessarily describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0064] In addition, in the description of the present application, "multiple" means two or more, unless otherwise specified. The association between the associated objects is described, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents a "or" relationship between the associated objects. The present application is further described below in conjunction with the drawings and examples.

[0065] The present application is further described below in conjunction with the drawings and examples.

[0066] Embodiment 1

[0067] Please refer to Figure 1 A heart blood vessel image segmentation method, comprising the following steps:

[0068] S1, normalizing the heart image to obtain a preprocessed image;

[0069] S2, performing convolution operation on the preprocessed image through a convolutional coding network to extract feature information of different scales and levels to obtain a multi-scale feature map;

[0070] S3, performing multiple feature extraction and multiple dimension transformation on the multi-scale feature map through a plurality of layers of capsule coding network to obtain a high-level semantic feature map;

[0071] S4, cutting the high-level semantic feature map into a plurality of image blocks through a Patch Embedding layer operation, and extracting features of the plurality of image blocks using a Transformer encoder, and fusing feature information in each image block through a self-attention mechanism to adaptively capture context information;

[0072] S5, performing a plurality of upsampling operations on the plurality of image blocks, and performing a plurality of skip connections according to the captured context information, and finally obtaining a segmentation result.

[0073] Compared with the prior art, the method of the present application introduces a self-attention mechanism and a capsule coding network, extracts higher-level semantic features through the capsule coding network for multiple feature extraction and dimension transformation of the image, and adaptively captures context information in the image through the self-attention mechanism to better preserve the spatial position relationship between feature entities, thereby solving the problem of discontinuity of the segmentation result.

[0074] Embodiment 2

[0075] Please refer to Figure 1 On the basis of Embodiment 1, the following contents are further disclosed in the present embodiment:

[0076] A heart blood vessel image segmentation method, comprising the following steps:

[0077] S1, normalizing the heart image to obtain a preprocessed image;

[0078] In one specific embodiment, the step S1 is specifically:

[0079] The heart image is adjusted to a preset size, and the pixel value of the heart image is scaled to a preset range, and finally the heart image is converted into a tensor form to obtain a preprocessed image.

[0080] Specifically, the heart image is adjusted to an image of size (128, 128, 128). Then, the image is converted to the direction of left, front and up, and the pixel value range is scaled to 0 to 1. Finally, each image is converted into a tensor form.

[0081] S2, performing convolution operation on the preprocessed image by a convolutional coding network to extract feature information of different scales and levels to obtain a multi-scale feature map;

[0082] In one specific embodiment, in the step S2, the convolutional coding network comprises a first convolutional layer, a second convolutional layer, a third convolutional layer and an activation layer.

[0083] The first convolutional layer is configured to perform convolution operation and edge padding on the preprocessed image simultaneously to obtain a first output image.

[0084] The second convolutional layer is configured to perform further edge padding and expand the number of output channels on the first output image to obtain a second output image.

[0085] The third convolutional layer performs further edge padding on the second output, and increases the ability of non-linear transformation in network level by a tanh activation function in the activation layer, and finally obtains the multi-scale feature map.

[0086] In one specific embodiment, in the step S2, the convolution operation can be represented as:

[0087] out(x) = conv3(conv2(conv1(x, k1, s1, p1, c1), k2, s2, p2, c2), k3, s3, p3, c3) 3, s 3, p3, c3)

[0088] Wherein x represents input, k i , s i , p i , c i respectively represent convolution kernel size, convolution step, edge padding size and output channel number.

[0089] Specifically, the first convolutional layer (conv1) uses a convolution kernel size of 5*5*5, performs convolution operation on the input image with a step of 1, performs edge padding to reduce the influence of image size reduction on information, the padding number is 2, and the output channel number is set to 16. In the second convolutional layer (conv2), a convolution kernel of 5*5*5 is also used, and the output of the previous layer is padded by 4 to achieve better feature extraction, and the output channel number is expanded to 32 to increase the expression ability of the network. Finally, in the third convolutional layer (conv3), a convolution kernel of 5*5*5 is used, and the output of the previous layer is padded by 4, and the output channel is set to 64 to further extract the features of the image. The tanh activation function can be represented as:

[0090]

[0091] Through the above improvements, different scale and hierarchical feature information can be extracted from the original image to provide effective input feature representation for subsequent tasks. The convolution method can greatly reduce the data volume while extracting the main features, which is beneficial to the convergence and training of the network.

[0092] S3, performing multiple feature extractions and multiple dimension transformations on the multi-scale feature map through a plurality of layers of capsule coding networks to obtain a high-level semantic feature map;

[0093] In one specific embodiment, in the step S3, each layer of the capsule coding network includes a plurality of capsules, and a dynamic routing between the capsules is used to calculate a weight distribution between the capsules. The dynamic routing between the capsules calculates the weight of the output capsule based on a dot product of vectors and a dynamically adjusted manner. An input vector of a jth capsule of an Ith layer of the capsule coding network relative to an ith capsule u i of an (I-1)th layer of the capsule coding network can be expressed as:

[0094]

[0095] where s ij represents an input vector of a jth capsule v j of an Ith layer of the capsule coding network relative to an ith capsule u i of an (I-1)th layer of the capsule coding network, and W ij represents a weight matrix. Then, a dynamic routing algorithm is used to calculate the weight distribution between the capsules:

[0096]

[0097] where c ij represents a weight of the dynamic routing, and b ij represents an iteration variable in the dynamic routing algorithm. Finally, an output vector of the jth capsule of the Ith layer of the capsule coding network is calculated by a weighted summation:

[0098]

[0099] The length of the capsule in the capsule coding network represents the existence probability or confidence of a certain feature. Usually, the norm of the output vector of the capsule is used as the length of the capsule. For the jth unit of the Ith layer of the capsule coding network, the length calculation formula is:

[0100]

[0101] where d j represents the length of the jth capsule of the Ith layer of the capsule coding network, and v j represents the output vector of the jth capsule of the Ith layer of the capsule coding network.

[0102] In one specific embodiment, in the step S3, the multi-scale feature maps are subjected to multiple feature extraction and dimension transformation by 4-layer capsule encoding networks, denoted as:

[0103] cap out (x) = cap4(cap3(cap2(cap1(x, k1, t1), k2, t2), k3, t3), k4, t4) 4, t4)

[0104] wherein x represents the input of the capsule encoding network, k represents the convolution kernel size, and t represents the output capsule feature dimension size.

[0105] Specifically, the feature map output by the convolutional encoding network is expanded into a capsule with 64 feature dimensions as the input of the first capsule encoding network. The convolution kernel size of the first layer of the capsule encoding network is 3*3*3, the convolution step size is 1, the boundary padding size is 1, and the number of iterations of the routing algorithm is 3. Finally, 16 capsule feature maps with 4 feature dimensions are output. The input of the second layer of the capsule encoding network is the output of the first layer of the capsule encoding network. The convolution kernel size of the second layer of the capsule encoding network is 3*3*3, the convolution step size is 2, the boundary padding size is 1, and the number of iterations of the routing algorithm is 3. Finally, 16 capsule feature maps with 5 feature dimensions are output. The input of the third layer of the capsule encoding network is the output of the second layer of the capsule encoding network. The convolution kernel size of the third layer of the capsule encoding network is 3*3*3, the convolution step size is 2, the boundary padding size is 1, and the number of iterations of the routing algorithm is 3. Finally, 8 capsule feature maps with 16 feature dimensions are output. The input of the fourth layer of the capsule encoding network is the output of the third layer of the capsule encoding network. The convolution kernel size of the third layer of the capsule encoding network is 3*3*3, the convolution step size is 2, the boundary padding size is 1, and the number of iterations of the routing algorithm is 3. Finally, 3 capsule feature maps with 32 feature dimensions are output.

[0106] These capsule encoding networks convert a series of input capsules (with different dimensions) into output capsules with different numbers and dimensions for learning the feature representation of the input data. Each capsule encoding network uses convolution operation for feature extraction and applies the routing algorithm in the capsule encoding network to aggregate features and generate output capsules. The dynamic routing mechanism used in the capsule encoding network can adaptively determine the weight between each pair of input and output capsules in multiple passes, providing strong feature support for the subsequent Transformer encoder.

[0107] S4, cutting the high-level semantic feature map into a plurality of image blocks by a Patch Embedding layer operation, and performing feature extraction on the plurality of image blocks using a Transformer encoder, and fusing feature information in each image block through a self-attention mechanism to adaptively capture context information;

[0108] In one specific embodiment, in the step S4, the Patch Embedding layer operation specifically comprises:

[0109] cutting the high-level semantic feature map into a plurality of small blocks of a preset size, and mapping each small block into a vector space to obtain an initial patch embedding matrix X e R(N 2 ) x d, where d represents a dimension size of the vector, and N represents a number of small blocks into which the image is cut; after each small block is expanded into a vector, a d-dimensional vector is obtained through a linear layer, i.e.,

[0110] e i =W p *path i +b p

[0111] where e i e R d represents a feature vector of the i-th small block, W p e R d x ((patch_size) 2 x c) and b p e R d , respectively, represent a weight and a bias of the Patch Embedding layer.

[0112] Specifically, the Patch Embedding layer cuts the high-level semantic feature map into a plurality of small blocks of patch_size (8*8*8) size.

[0113] In one specific embodiment, in the step S4, the Transformer encoder is composed of a plurality of Transformer block cycles, each Transformer block is composed of a plurality of self-attention and MLP layers, and finally each Transformer block outputs a matrix Z of the same dimension as the input X; the k-th Transformer block takes the input matrix X and the matrix Z h -1 output by the last layer as input to obtain an output matrix Z h∈ R(N 2 ) x d:

[0114] Q h = X * W h * Q

[0115] K h = X * W h * K

[0116] V h = X * W h * V

[0117]

[0118] Z h = LayerNorm(X + Dropout(MLP(LayerNorm(A h ) + X)))

[0119] where Q h , K h , V h represent query, key, value matrices respectively, which are obtained by linear transformation of matrix X and weight matrix W h * Q, W h * K, W h * V; A h represents the attention weight matrix obtained by self-attention processing; LayerNorm represents the normalization operation of this layer, Dropout represents the random inactivation operation, and MLP represents the fully connected layer network.

[0120] The attention weight is calculated by self-attention encoding, each element in the sequence interacts with other elements, and the representation of itself is calculated according to the interaction result. In this way, each element can obtain global information, and can be weighted according to the importance of other elements in the sequence, so as to better capture the context relationship in the sequence.

[0121] S5, a plurality of upsampling operations are performed on the plurality of image blocks, and a plurality of jump connections are performed according to the captured context information, and finally a segmentation result is obtained.

[0122] In a specific embodiment, in the step S5, specifically, five up-sampling operations are performed on the plurality of image blocks; the first up-sampling operation is performed on the result obtained after the capsule encoding of the low-resolution capsule and the result obtained through the self-attention encoding, and then the up-sampled result is obtained; the second, third and fourth up-sampling operations are respectively performed on the output result of the last up-sampling operation and the third, second and first capsule encoding output results, and then the up-sampling is performed; the fifth up-sampling operation is performed on the result of the fourth up-sampling operation and the result obtained by the feature extractor, and then the up-sampling is performed to obtain the final segmentation result.

[0123] Embodiment 3

[0124] Please refer to Figure 2 A cardiac vascular image segmentation system, comprising a preprocessing module, a convolutional encoding module, a capsule encoding module, a self-attention encoding module and a connection encoding module.

[0125] The preprocessing module is configured to normalize the cardiac image to obtain a preprocessed image.

[0126] The convolutional encoding module is configured to perform convolutional operation on the preprocessed image through a convolutional encoding network to extract feature information of different scales and levels, and obtain a multi-scale feature map.

[0127] The capsule encoding module is configured to perform multiple feature extractions and multiple dimension transformations on the multi-scale feature map through a plurality of layers of capsule encoding networks to obtain a high-level semantic feature map.

[0128] The self-attention encoding module is configured to cut the high-level semantic feature map into a plurality of image blocks through a Patch Embedding layer operation, and extract features of the plurality of image blocks using a Transformer encoder, and fuse feature information in each image block through a self-attention mechanism to adaptively capture context information.

[0129] The connection encoding module is configured to perform a plurality of up-sampling operations on the plurality of image blocks, and perform a plurality of jump connections according to the captured context information to finally obtain a segmentation result.

[0130] It should be noted that the convolutional encoding module comprises three 3D convolutional layers (conv1, conv2, conv3) and an activation layer.

[0131] It should be noted that the capsule encoding module comprises four layers of capsule encoding networks, and each layer of capsule encoding network comprises a plurality of capsules.

[0132] It should be noted that the self-attention encoding module includes a Patch Embedding layer and a Transformer encoder. The Transformer encoder is composed of several Transformer blocks in a loop, and each Transformer block consists of several layers of self_attention and MLP layers.

[0133] It should be noted that the connection coding module includes a skip connection layer and an upsampling module, and the upsampling module specifically includes five upsampling layers.

[0134] Example 4

[0135] This embodiment verifies and analyzes the method, more specifically:

[0136] Please refer to Figure 3 as well as Figure 4 In this invention, 807 cases were randomly selected from a collected dataset of 982 heart samples as the training set, and 175 cases were selected as the validation set. As shown in Table 1, compared to the UCAPs network, the proposed novel segmentation algorithm CapsTransUnet, based on a capsule network model and self-attention mechanism, not only has fewer parameters and shorter training time, but also exhibits superior Dice scores. This demonstrates that incorporating a self-attention mechanism into the capsule network model can significantly improve the segmentation performance with fewer capsule encodings. While CapsTransUnet's Dice coefficient is slightly lower than that of U-Net, it has advantages in terms of fewer parameters and shorter training time.

[0137] Table 1. Experimental results of different cardiac and vascular image segmentation methods

[0138]

[0139] The experimental results show that the method of the present invention, by introducing a self-attention mechanism and a capsule network, achieves end-to-end segmentation of the original images of the heart and blood vessels. Compared with other network models, it has higher segmentation accuracy and better segmentation efficiency with fewer model parameters, which can provide strong support for subsequent clinical diagnosis and treatment.

[0140] Obviously, the above embodiments of the present application are merely exemplary but not intended to limit the embodiments of the present application. Based on the above description, any other variations or changes can be made by those skilled in the art without departing from the spirit and principles of the present application. It is not necessary to list all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall fall within the scope of the claims of the present application.

Claims

1. A method of cardiac vessel image segmentation, characterized by, The method comprises the following steps: S1, normalizing the heart image to obtain a preprocessed image; S2, performing convolution operation on the preprocessed image through a convolutional coding network to extract feature information of different scales and levels, and obtaining a multi-scale feature map; S3, performing multiple feature extraction and dimension transformation on the multi-scale feature map through a plurality of capsule coding networks to obtain a high-level semantic feature map; S4, cutting the high-level semantic feature map into a plurality of image blocks through a Patch Embedding layer operation, and extracting features of the plurality of image blocks using a Transformer encoder, and fusing feature information in each image block through a self-attention mechanism to adaptively capture context information; S5, performing a plurality of upsampling operations on the plurality of image blocks, and performing a plurality of skip connections according to the captured context information to finally obtain a segmentation result; In step S3, each layer of the capsule coding network includes several capsules. Dynamic routing between capsules is used to calculate the weight allocation between capsules. The dynamic routing between capsules calculates the weight of the output capsule based on the dot product of vectors and dynamic adjustment. The j-th capsule of the I-th layer capsule coding network is relative to the i-th capsule of the (I-1)-th layer capsule coding network. The input vector can be represented as: wherein, represents the jthcapsule of the Ithlayer capsule encoding network with respect to the ithcapsule of the I-1thlayer capsule encoding network input vector of the capsule, represents the weight matrix; then, the weight distribution between capsules is computed by a dynamic routing algorithm: wherein, denotes the weight of the dynamic routing, denotes the iteration variable in the dynamic routing algorithm; finally, the output vector of the jth capsule of the Ith layer capsule encoding network is calculated in the manner of over-weighted summation: The capsule length in the capsule coding network represents the existence probability or confidence of a certain feature, and the norm of the output vector of the capsule is usually used as the capsule length. For the jth unit of the Ith layer capsule coding network, the length calculation formula is: wherein denotes the length of the jth capsule of the first layer capsule encoding network, denotes the output vector of the jth capsule of the first layer capsule encoding network.

2. The method of claim 1, wherein, The step S1 is specifically: The heart image is adjusted to a preset size, the pixel value of the heart image is scaled to a preset range, and finally the heart image is converted into a tensor form to obtain a preprocessed image.

3. The method of claim 1, wherein, In the step S2, the convolutional coding network comprises a first convolutional layer, a second convolutional layer, a third convolutional layer, and an activation layer; The first convolutional layer is used for simultaneously performing convolution operation and edge padding on the preprocessed image to obtain a first output image; The second convolutional layer is used for further edge padding and expanding the output channel number of the first output image to obtain a second output image; The third convolutional layer further performs edge padding on the second output, and increases the ability of nonlinear transformation in the network level through the tanh activation function in the activation layer, and finally obtains a multi-scale feature map.

4. The method of claim 3, wherein, In the step S2, the convolution operation can be represented as: where x represents an input, respectively represent the convolution kernel size, the convolution stride, the edge padding size, and the output channel number of the i-th layer.

5. The method of claim 1, wherein, In the step S3, the multi-scale feature map is subjected to multiple feature extraction and dimension transformation through a 4-layer capsule coding network, which is represented as: Wherein, x represents the input of the capsule encoding network, 、 、 、 Respectively represent the convolution kernel size of the 1st, 2nd, 3rd and 4th layer capsule network, 、 、 、 Respectively represent the output capsule feature dimension size of the 1st, 2nd, 3rd and 4th layer capsule network.

6. The method of claim 1, wherein, In the step S4, the Patch Embedding layer operation is specifically: cutting the high-level semantic feature map into several small blocks of a preset size, and mapping each small block into a vector space to obtain an initial patch embedding matrix where d represents the dimension size of the vector, and N represents the number of small blocks into which the image is cut; after each small block is unfolded into a vector, a d-dimensional vector is obtained through a linear layer, that is: wherein, representing the feature vector of the i-th patch, and , respectively, represent the weights and bias of the Patch Embedding layer.

7. The method of claim 6, wherein, In the step S4, the Transformer encoder is composed of several Transformer block cycles, each Transformer block is composed of several layers of self_attention and MLP layers, and finally each Transformer block outputs a matrix Z with the same dimension as the input X; the kth Transformer block, the input matrix X and the matrix output by the last layer As input, an output matrix processed by the self-attention module self_attention and the fully connected layer network MLP is obtained : wherein, , , respectively represent the query, key, value matrices, which are obtained by linear transformation of the matrix X and the weight matrix , , ; represents the attention weight matrix obtained by self-attention processing; LayerNorm indicates the normalization operation of this layer, Dropout represents the random inactivation operation, and MLP represents the fully connected layer network.

8. The method of claim 1, wherein, In the step S5, five upsampling operations are performed on the plurality of image blocks; the first upsampling operation is performed on the results obtained after low-resolution capsule coding and the results obtained through self-attention coding, and then the results obtained through upsampling are connected; The second, third, and fourth upsampling operations are respectively performed on the output results of the previous upsampling, and the third, second, and first capsule coding output results are connected through skip connection, and then upsampling is performed; The fifth upsampling operation is performed on the results of the fourth upsampling and the results obtained by the feature extractor to obtain the final segmentation result.

9. A cardiac vascular image segmentation system, characterized by, The method comprises a preprocessing module, a convolutional coding module, a capsule coding module, a self-attention coding module, and a connection coding module; The preprocessing module is used for normalizing the heart image to obtain a preprocessed image; The convolution coding module is configured to perform convolution operation on the preprocessed image through a convolution coding network to extract feature information of different scales and levels, and obtain a multi-scale feature map; The capsule coding module is configured to perform feature extraction and dimension transformation multiple times on the multi-scale feature map through a plurality of layers of capsule coding networks, and obtain a high-level semantic feature map; The self-attention coding module is configured to cut the high-level semantic feature map into a plurality of image blocks through a Patch Embedding layer operation, perform feature extraction on the plurality of image blocks using a Transformer encoder, fuse feature information in each image block through a self-attention mechanism, and adaptively capture context information; The connection coding module is configured to perform upsampling operation on the plurality of image blocks multiple times, and perform multiple jump connections according to the captured context information, and finally obtain a segmentation result; The capsule encoding network of each layer includes a plurality of capsules, dynamic routing between the capsules is used to calculate the weight distribution between the capsules, the dynamic routing between the capsules calculates the weight of the output capsule based on the dot product of vectors and a dynamically adjusted manner, the jth capsule of the Ith layer capsule encoding network is relative to the ith capsule of the (I-1)th layer capsule encoding network The input vector of the ith capsule of the (I-1)th layer capsule encoding network can be represented as: wherein, represents the jthcapsule of the Ithlayer capsule encoding network with respect to the ithcapsule of the I-1thlayer capsule encoding network input vector of the capsule, represents the weight matrix; then, the weight distribution between capsules is computed by a dynamic routing algorithm: wherein, denotes the weight of the dynamic routing, denotes the iteration variable in the dynamic routing algorithm; finally, the output vector of the jth capsule of the Ith layer capsule encoding network is calculated in the manner of over-weighted summation: The capsule length in the capsule coding network represents the existence probability or confidence of a certain feature, and the norm of the output vector of the capsule is usually used as the capsule length. For the jth unit of the Ith layer of the capsule coding network, the length calculation formula is: wherein denotes the length of the jth capsule of the first layer of capsule encoding network, denotes the output vector of the jth capsule of the first layer of capsule encoding network.

Citation Information

Patent Citations

  • UNet and Transform fusion-based retinal vessel segmentation method

    CN115908241A

  • Eye fundus blood vessel image segmentation method and system based on cavity convolution and semantic fusion

    CN115205300A

  • Medical image segmentation using an integrated edge guidance module and object segmentation network

    US10482603B1