Myocardial segmentation method, system and equipment based on myocardial contrast echocardiogram
By using the expansion convolution module and the Transformer block in the myocardial segmentation model, the problem of limited receptive field in the prior art is solved, multi-scale feature extraction and long-distance dependency capture of myocardial images are achieved, and the accuracy and integrity of myocardial segmentation are improved.
Patent Information
- Application Number
- CN202411871519.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The existing myocardial segmentation method has limited receptive field, making it difficult to capture the local characteristics of the myocardial and surrounding tissues, as well as the correlation characteristics of myocardial tissue in a larger range of images, resulting in poor myocardial segmentation accuracy.
The myocardial segmentation model based on the expansion convolution module and the Transformer block is adopted to expand the receptive field by stacking expansion convolution layers with different expansion rates, and the Transformer block is introduced in the bottleneck layer to capture the long-distance spatial dependence.
Effectively extract multi-scale features, improve the perception of tissue structures at different scales in myocardial comparison images, improve the prediction accuracy of myocardial boundary areas, and achieve more complete and accurate myocardial segmentation.
Smart Images

Figure CN120013956A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ultrasonic image processing, and in particular to a myocardial segmentation method, system and device based on myocardial contrast echocardiography. Background Art
[0002] With the widespread application of deep learning technology, the development of medical image analysis will surely go further. In myocardial contrast echocardiography (MCE), accurate segmentation and identification of the myocardium is of great significance for the diagnosis and treatment of coronary heart disease. However, in the actual operation process, due to factors such as individual differences in patients, imaging quality and noise interference, artifacts or abnormal signals will inevitably appear in the image. If not handled properly, it will often lead to inaccurate diagnostic results and affect the patient's treatment plan.
[0003] Traditional image segmentation methods, such as threshold-based, edge detection, and region growing methods, are often used in the analysis of myocardial contrast echocardiography (MCE). The threshold-based method is to set a specific grayscale threshold and divide the parts of the image with pixel values above or below the threshold into different regions to try to distinguish the myocardium from other tissues; edge detection uses the sudden change of pixel grayscale in the image to determine the boundary of myocardial tissue; region growing starts from a seed point in the image and gradually merges adjacent pixels according to certain similarity criteria until a complete myocardial region is formed. However, the structure of myocardial tissue itself is complex, and its texture, grayscale and other characteristics are not uniform. In addition, the quality of myocardial contrast echocardiography (MCE) is often interfered by various factors, such as noise and uneven distribution of contrast agents. These factors make it difficult for traditional methods to accurately determine the myocardial boundary, and the segmentation results are often not accurate enough to meet the clinical requirements for high-precision myocardial segmentation. In view of this, the U-Net network was introduced into the myocardial segmentation task. The U-Net network adopts a unique encoder-decoder structure. The encoder downsamples the image and extracts features at different levels. It can capture the global information of the image and overcome the problem that traditional methods are insufficient in extracting complex myocardial tissue features. The decoder performs upsampling and fuses the features of the corresponding levels of the encoder, so that the network can better restore image details, thereby improving the accuracy of segmentation and making up for the shortcomings of traditional methods in accurately segmenting the myocardium.
[0004] However, although the U-Net network has achieved certain results in the field of image segmentation, it still has obvious defects in the application of myocardial segmentation in myocardial contrast echocardiography. When processing myocardial images, the U-Net network, as an architecture mainly based on convolution operations, is limited by the characteristics of ordinary convolutions and has a relatively limited receptive field. When dealing with complex and finely structured objects such as myocardial tissue, the network is difficult to fully capture the local features of each part of the myocardium and its relationship with surrounding tissues; and the U-Net network has limited modeling capabilities for some long-distance spatial dependencies, making it difficult to fully capture the associated features of myocardial tissue in a larger range of images, resulting in poor myocardial segmentation accuracy. Summary of the invention
[0005] To this end, the technical problem to be solved by the present invention is to overcome the defects of existing myocardial segmentation methods, such as limited receptive field, difficulty in capturing local features of the myocardium and surrounding tissues, and difficulty in fully capturing the associated features of myocardial tissue in larger range images, resulting in poor myocardial segmentation accuracy.
[0006] In order to solve the above technical problems, the present invention provides a myocardial segmentation method based on myocardial contrast echocardiography, comprising the following steps:
[0007] Constructing a myocardial segmentation model, wherein the myocardial segmentation model includes: an encoder, a bottleneck layer and a decoder;
[0008] The encoder includes a plurality of layers connected in sequence along the forward propagation direction, each layer includes a plurality of dilated convolution modules connected in sequence, each dilated convolution module includes a plurality of dilated convolution layers with different dilation rates connected in sequence, each dilated convolution layer includes a dilated convolution, a batch normalization layer and a ReLU activation function connected in sequence, and the output features of all the dilated convolution layers in each dilated convolution module are stacked in the channel dimension to obtain an output feature map of each dilated convolution module;
[0009] Input the myocardial contrast echocardiogram into the encoder to extract features and obtain output feature maps of each layer of the encoder;
[0010] The output feature map of the last layer of the encoder is input into the bottleneck layer for feature fusion to obtain fused features;
[0011] The output feature maps of each layer of the encoder and the fused features are input into the decoder to obtain the myocardial segmentation results of myocardial contrast echocardiography.
[0012] Preferably, when the number of layers of the encoder is 5, the first layer and the second layer of the encoder each include two sequentially connected dilated convolution modules, and the third layer, the fourth layer and the fifth layer of the encoder each include three sequentially connected dilated convolution modules;
[0013] When the number of layers of the encoder is greater than 5, the first layer and the second layer of the encoder each include two sequentially connected dilated convolution modules, the third layer, the fourth layer, and the fifth layer of the encoder each include three sequentially connected dilated convolution modules, and the remaining layers of the encoder each include multiple sequentially connected dilated convolution modules.
[0014] Preferably, inputting the input features of the dilated convolution module into the dilated convolution module, and outputting the output feature map of the dilated convolution module, comprises:
[0015] The input features of the i-th dilated convolution layer in the dilated convolution module are passed through the dilated convolution of the i-th dilated convolution layer to obtain the output features of the dilated convolution of the i-th dilated convolution layer. The formula is:
[0016] Z i =ω i *F i + b,
[0017] The output features of the dilated convolution of the i-th dilated convolution layer are processed by the batch normalization layer to obtain the output features of the dilated convolution of the i-th dilated convolution layer after batch normalization. The formula is:
[0018]
[0019] The output features of the dilated convolution of the i-th dilated convolution layer after batch normalization are processed by the ReLU activation function to obtain the output features of the i-th dilated convolution layer, and the formula is:
[0020]
[0021] The output features of all dilated convolutional layers are stacked in the channel dimension to obtain the output feature map of the dilated convolution module. The formula is:
[0022]
[0023] Among them, Z i is the output feature of the dilated convolution of the i-th dilated convolutional layer, ω i is the convolution kernel of the dilated convolution of the i-th dilated convolution layer in the dilated convolution module, F i is the input feature of the i-th dilated convolutional layer in the dilated convolutional module, b is the bias term, * is the convolution operation, i=1,2…I, I is the number of dilated convolutional layers in the dilated convolutional module, is the output feature of the dilated convolution of the i-th dilated convolution layer after batch normalization, · is the dot product, μ B is the mean of the output features of the dilated convolution of the ith dilated convolutional layer corresponding to the myocardial contrast echocardiogram, is the variance of the output feature of the dilated convolution of the ith dilated convolutional layer corresponding to the myocardial contrast echocardiogram, ε is a constant, γ is the first learnable parameter, β is the second learnable parameter, is the output feature of the i-th dilated convolutional layer, max(.) is the maximum value, is the output feature map of the expanded convolution module, and ⊕ is stacking in the channel dimension.
[0024] Preferably, the input features of the dilated convolution module are input into the dilated convolution module, and the output feature map of the dilated convolution module is output, and the formula is:
[0025]
[0026] in, is the output feature map of the dilated convolution module, DCM(.) is the dilated convolution module, ⊕ is stacking in the channel dimension, ReLU(.) is the ReLU activation function, BN(.) is the batch normalization layer, DilatedConv i (.) is the dilated convolution of the i-th dilated convolution layer in the dilated convolution module, I is the number of dilated convolution layers in the dilated convolution module, and F i is the input feature of the i-th dilated convolutional layer in the dilated convolutional module, F 1 is the input feature of the dilated convolution module.
[0027] Preferably, the step of inputting the output feature map of the last layer of the encoder into the bottleneck layer to perform feature fusion to obtain fused features includes:
[0028] The output feature map of the last layer of the encoder is reshaped to obtain a reshaped feature map, which is flattened into N patches. The set of all patches forms a shape of A matrix of is the height of the output feature map of the last layer of the encoder, is the width of the output feature map of the last layer of the encoder, and P is the size of the Patch block;
[0029] The matrix is linearly mapped to the D-dimensional space, and each patch block is converted into its corresponding D-dimensional vector;
[0030] After adding position encoding to the D-dimensional vector corresponding to each patch block of the output feature map of the last layer of the encoder, it is input into the Transformer block for feature fusion to obtain the initial fused feature;
[0031] Reshape the initial fusion feature to make the size of the initial fusion feature consistent with the reshaped feature. Figure 1 To obtain the fusion features.
[0032] Preferably, the Transformer block comprises a plurality of Transformer layers connected in sequence, and each Transformer layer comprises: layer normalization, a multi-head attention mechanism and a multi-layer perceptron connected in sequence.
[0033] Preferably, the decoder comprises: a plurality of upsampling blocks, a 3×3 convolution layer and a segmentation head, each upsampling block consists of a 3×3 convolution layer and a 2×2 up-convolution, and each 3×3 convolution layer comprises a 3×3 convolution and a ReLu activation function.
[0034] Preferably, the step of inputting the output feature maps of each layer of the encoder and the fusion features into a decoder to obtain a myocardial segmentation result of myocardial contrast echocardiography includes:
[0035] Process the fused features through the first upsampling block to obtain the output feature map of the first upsampling block;
[0036] The output feature map of the jth upsampling block is stacked with the output feature map of the encoder's M-j+1th layer in the channel dimension and used as the input of the j+1th upsampling block; where j=1,2…M, M is the number of upsampling blocks;
[0037] The output feature map of the Mth upsampling block and the output feature map of the first layer of the encoder are stacked in the channel dimension and processed by a 3×3 convolutional layer and a segmentation head in sequence to obtain the myocardial segmentation result of myocardial contrast echocardiography.
[0038] The present invention also provides a myocardial segmentation system based on myocardial contrast echocardiography, comprising:
[0039] A model building module, used to build a myocardial segmentation model, wherein the myocardial segmentation model includes: an encoder, a bottleneck layer and a decoder;
[0040] The encoder includes a plurality of layers connected in sequence along the forward propagation direction, each layer includes a plurality of dilated convolution modules connected in sequence, each dilated convolution module includes a plurality of dilated convolution layers with different dilation rates connected in sequence, each dilated convolution layer includes a dilated convolution, a batch normalization layer and a ReLU activation function connected in sequence, and the output features of all the dilated convolution layers in each dilated convolution module are stacked in the channel dimension to obtain an output feature map of each dilated convolution module;
[0041] A feature extraction module, for extracting features by inputting the myocardial contrast echocardiogram into an encoder, and obtaining output feature maps of each layer of the encoder;
[0042] The fusion module is used to input the output feature map of the last layer of the encoder into the bottleneck layer to perform feature fusion and obtain fused features;
[0043] The segmentation module is used to input the output feature maps of each layer of the encoder and the fusion features into the decoder to obtain the myocardial segmentation result of myocardial contrast echocardiography.
[0044] The present invention also provides a myocardial segmentation device based on myocardial contrast echocardiography, comprising:
[0045] The memory is used to store the computer program; the processor is used to implement the steps of the above-mentioned myocardial segmentation method based on myocardial contrast echocardiography when executing the computer program.
[0046] The above technical solution of the present invention has the following beneficial effects compared with the prior art:
[0047] The myocardial segmentation method, system and device based on myocardial contrast echocardiography described in the present invention set different numbers of dilated convolution modules in different layers of the encoder. The dilated convolution modules expand the receptive field during feature extraction by stacking convolutions with different dilation rates, thereby effectively extracting multi-scale feature information, thereby improving the perception of tissue structures of different scales in myocardial contrast images. At the same time, this structure is particularly helpful in capturing a wide range of low-frequency information, such as the overall structure and background of the myocardium, making up for the defect that the traditional U-Net network is difficult to fully perceive local details due to the limited receptive field of ordinary convolution; in addition, the present invention flattens the output feature map of the last layer into patch blocks and constructs a matrix, and reintegrates the local features that were originally scattered and difficult to fully explore the association. Each patch block becomes a unit containing local information, which can more finely reflect the local details of the myocardial tissue and surrounding tissues. By introducing a Transformer block in the bottleneck layer, the Transformer block includes a plurality of Transformer layers connected in sequence, and each Transformer layer includes: layer normalization connected in sequence, a multi-head attention mechanism and a multi-layer perceptron. The multi-head attention mechanism can break through the local limitations of traditional convolution, and fully capture the long-distance spatial dependencies between various parts of the myocardium and the surrounding tissues. Through multiple Transformer layers, spatial dependencies of different scales or types can be paid attention to, and the irregular shape and complex structure of the myocardium in myocardial contrast echocardiogram can be adapted. The extracted features can be interactively fused to make the segmentation of the myocardial contour more complete, thereby improving the prediction accuracy of the myocardial boundary area. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below according to specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:
[0049] Figure 1is a flow chart of a myocardial segmentation method based on myocardial contrast echocardiography according to the present invention, Figure 1 (a) is the structural diagram of the myocardial segmentation model. Figure 1 (b) in the figure is the structural diagram of the Transformer layer.
[0050] Figure 2 It is the structural diagram of the dilated convolution module.
[0051] Figure 3 It is a schematic diagram of the structure of a single dilated convolution in the dilated convolution layer.
[0052] Figure 4 It is the receptive field map of the dilated convolution module.
[0053] Figure 5 It is a simulation result diagram after using myocardial contrast echocardiography from three apical angles to test the myocardial segmentation model proposed in the present invention and other contrast models. DETAILED DESCRIPTION
[0054] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0055] Embodiment 1 of the present invention provides a myocardial segmentation method based on myocardial contrast echocardiography, comprising the following steps:
[0056] Figure 1 is a flow chart of a myocardial segmentation method based on myocardial contrast echocardiography according to the present invention, Figure 1 (a) is the structural diagram of the myocardial segmentation model. Figure 1 (b) in the figure is the structural diagram of the Transformer layer.
[0057] Step S1: constructing a myocardial segmentation model (DillateUNet), wherein the myocardial segmentation model includes an encoder, a bottleneck layer and a decoder;
[0058] The encoder includes a plurality of layers connected in sequence along the forward propagation direction, each layer includes a plurality of dilated convolution modules connected in sequence, each dilated convolution module includes a plurality of dilated convolution layers with different dilation rates connected in sequence; each dilated convolution layer includes a dilated convolution, a batch normalization layer and a ReLU activation function connected in sequence, and the output features of all dilated convolution layers in each dilated convolution module are stacked in the channel dimension to obtain an output feature map of each dilated convolution module;
[0059] In this embodiment, specifically, the output feature maps of the layers except the last layer are downsampled and used as the input of the next layer, and the downsampling method is any one of maximum pooling and average pooling.
[0060] In this embodiment, preferably, when the number of layers of the encoder is 5, the first layer and the second layer of the encoder each include two dilated convolution modules connected in sequence, and the third layer, the fourth layer and the fifth layer of the encoder each include three dilated convolution modules connected in sequence; the five-layer encoder structure can extract features step by step and gradually enrich the feature expression. As the number of layers increases, from the two dilated convolution modules of the first and second layers to the three dilated convolution modules of the third, fourth and fifth layers, the receptive field can be gradually expanded to capture a wider range of contextual information, while avoiding problems such as gradient vanishing or gradient explosion due to the network being too deep, so that the network can effectively learn features at different scales, which not only ensures the ability to extract local detail features, but also can fully obtain global semantic information, thereby improving the model's overall representation ability and processing effect for complex data.
[0061] When the number of layers of the encoder is greater than 5, the first layer and the second layer of the encoder each include two sequentially connected dilated convolution modules, the third layer, the fourth layer, and the fifth layer of the encoder each include three sequentially connected dilated convolution modules, and the remaining layers of the encoder each include multiple sequentially connected dilated convolution modules.
[0062] like Figure 2 As shown, Figure 2 This is the structural diagram of the dilated convolution module.
[0063] In this embodiment, specifically, inputting the input features of the dilated convolution module into the dilated convolution module, and outputting the output feature map of the dilated convolution module, includes:
[0064] like Figure 3 As shown, Figure 3 This is a schematic diagram of the structure of a single dilated convolution in the dilated convolution layer. Each dilated convolution includes sequentially connected convolution, batch normalization, and ReLU activation functions.
[0065] The input features of the i-th dilated convolution layer in the dilated convolution module are passed through the dilated convolution of the i-th dilated convolution layer to obtain the output features of the dilated convolution of the i-th dilated convolution layer. The formula is:
[0066] Z i =ω i *F i + b,
[0067] The output features of the dilated convolution of the i-th dilated convolution layer are processed by the batch normalization layer to obtain the output features of the dilated convolution of the i-th dilated convolution layer after batch normalization. The formula is:
[0068]
[0069] The output features of the dilated convolution of the i-th dilated convolution layer after batch normalization are processed by the ReLU activation function to obtain the output features of the i-th dilated convolution layer, and the formula is:
[0070]
[0071] The output features of all dilated convolutional layers are stacked in the channel dimension to obtain the output feature map of the dilated convolution module. The formula is:
[0072]
[0073] Among them, Z i is the output feature of the dilated convolution of the i-th dilated convolutional layer, ω i is the convolution kernel of the dilated convolution of the i-th dilated convolution layer in the dilated convolution module, F i is the input feature of the i-th dilated convolutional layer in the dilated convolutional module, b is the bias term, * is the convolution operation, i=1,2…I, I is the number of dilated convolutional layers in the dilated convolutional module, is the output feature of the dilated convolution of the i-th dilated convolution layer after batch normalization, · is the dot product, μ B is the mean of the output features of the dilated convolution of the ith dilated convolutional layer corresponding to the myocardial contrast echocardiogram, is the variance of the output feature of the dilated convolution of the ith dilated convolutional layer corresponding to the myocardial contrast echocardiogram, ε is a constant, γ is the first learnable parameter, β is the second learnable parameter, is the output feature of the i-th dilated convolutional layer, max(.) is the maximum value, is the output feature map of the expanded convolution module, and ⊕ is stacking in the channel dimension.
[0074] The input features of the dilated convolution module are input into the dilated convolution module, and the output feature map of the dilated convolution module is output. The formula is:
[0075]
[0076] in, is the output feature map of the dilated convolution module, DCM(.) is the dilated convolution module, ⊕ is stacking in the channel dimension, ReLU(.) is the ReLU activation function, BN(.) is the batch normalization layer, DilatedConv i (.) is the dilated convolution of the i-th dilated convolution layer in the dilated convolution module, I is the number of dilated convolution layers in the dilated convolution module, and F i is the input feature of the i-th dilated convolutional layer in the dilated convolutional module, F1 is the input feature of the dilated convolution module.
[0077] like Figure 4 As shown, Figure 4 The receptive field diagram of the dilated convolution module. The present invention introduces a dilated convolution module (DCM), which expands the receptive field during feature extraction by stacking convolutions with different dilation rates, thereby effectively extracting multi-scale feature information, thereby improving the perception of tissue structures of different scales in myocardial contrast images. At the same time, this structure is also particularly helpful in capturing a wide range of low-frequency information, such as the overall structure and background of the myocardium.
[0078] Step S2: inputting the myocardial contrast echocardiogram into the encoder to extract features, and obtaining output feature maps of each layer of the encoder;
[0079] Step S3: Input the output feature map of the last layer of the encoder into the bottleneck layer to perform feature fusion to obtain fused features;
[0080] In this embodiment, specifically, the step of inputting the output feature map of the last layer of the encoder into the bottleneck layer to perform feature fusion to obtain fused features includes:
[0081] The output feature map of the last layer of the encoder is reshaped to obtain a reshaped feature map, which is flattened into N patches. The set of all patches forms a shape of A matrix of is the height of the output feature map of the last layer of the encoder, is the width of the output feature map of the last layer of the encoder, and P is the size of the Patch block;
[0082] The matrix is linearly mapped to the D-dimensional space, and each patch block is converted into its corresponding D-dimensional vector;
[0083] After adding position encoding to the D-dimensional vector corresponding to each patch block of the output feature map of the last layer of the encoder, it is input into the Transformer block for feature fusion to obtain the initial fused feature;
[0084] Reshape the initial fusion feature to make the size of the initial fusion feature consistent with the reshaped feature. Figure 1 To obtain the fusion features.
[0085] In this embodiment, preferably, the Transformer block includes a plurality of Transformer layers connected in sequence, and each Transformer layer includes: layer normalization, a multi-head attention mechanism and a multi-layer perceptron connected in sequence.
[0086] By introducing the Transformer block, the present invention can efficiently capture the long-distance dependencies between different regions in myocardial contrast ultrasound images, can adapt to the irregular shape and complex structure of the myocardium in myocardial contrast ultrasound images, and the extracted features can be interactively fused to make the segmentation of the myocardial contour more complete, thereby improving the prediction accuracy of the myocardial boundary area.
[0087] Step S4: Input the output feature maps of each layer of the encoder and the fused features into the decoder to obtain the myocardial segmentation result of the myocardial contrast echocardiography.
[0088] In this embodiment, preferably, the decoder includes: multiple upsampling blocks, 3×3 convolution layers and a segmentation head, each upsampling block is composed of a 3×3 convolution layer and a 2×2 up-convolution, and each 3×3 convolution layer includes a 3×3 convolution and ReLu activation function. The number of layers of the decoder is consistent with the number of layers of the encoder. The 3×3 convolution layer can further extract and optimize the feature map after upsampling, slide the convolution kernel on the feature map, and perform operations such as fusion and screening on the features again, and adjust the expression form of the features to make it more conducive to the subsequent segmentation head to make category judgments.
[0089] In this embodiment, specifically, the step of inputting the output feature maps of each layer of the encoder and the fusion features into the decoder to obtain the myocardial segmentation result of the myocardial contrast echocardiogram includes:
[0090] Process the fused features through the first upsampling block to obtain the output feature map of the first upsampling block;
[0091] The output feature map of the jth upsampling block is stacked with the output feature map of the encoder's M-j+1th layer in the channel dimension and used as the input of the j+1th upsampling block; where j=1,2…M, M is the number of upsampling blocks;
[0092] The output feature map of the Mth upsampling block and the output feature map of the first layer of the encoder are stacked in the channel dimension and processed by a 3×3 convolutional layer and a segmentation head in sequence to obtain the myocardial segmentation result of myocardial contrast echocardiography.
[0093] In this embodiment, specifically, the segmentation head is composed of 3×3 convolutions, and the number of output channels of the segmentation head is the total number of categories. The segmentation head is composed of 3×3 convolutions and the number of output channels is the total number of categories. This can accurately divide different myocardial structures and possible background categories, attribute each pixel in the feature map to the corresponding category, and finally output a clear and definite myocardial segmentation result.
[0094] This embodiment uses the classic model in semantic segmentation and the myocardial segmentation model (DillateUNet) proposed in the present invention to perform myocardial segmentation on myocardial contrast echocardiograms (including three types, two-chamber heart (A2C), three-chamber heart (A3C), and four-chamber heart (A4C)). The segmentation results are as follows: Figure 5 As shown, Figure 5 The figure is a simulation result diagram after using myocardial contrast echocardiography from three apical angles to test the myocardial segmentation model proposed in the present invention and other contrast models. Figure 5 The first row in the figure shows, from left to right, the original images of the two-chamber heart (A2C) myocardial contrast echocardiogram, the original images of the three-chamber heart (A3C) myocardial contrast echocardiogram, and the original images of the four-chamber heart (A4C) myocardial contrast echocardiogram. Figure 5 The second row in the figure shows the true labels corresponding to the original image of the two-chamber heart (A2C) myocardial contrast echocardiogram, the true labels corresponding to the original image of the three-chamber heart (A3C) myocardial contrast echocardiogram, and the true labels corresponding to the original image of the four-chamber heart (A4C) myocardial contrast echocardiogram from left to right. Figure 5 The third to the last row in the figure are respectively the simulation results of myocardial segmentation of three types of myocardial contrast echocardiography using the U-Net model, U-Net++ model, U-Net+++ model, U2-Net model, AttentionUNet model, Deeplabv3+ model, Segformer model, PSPNet model and the myocardial segmentation model (DillateUNet) proposed in the present invention. Figure 5 It can be seen that the myocardial segmentation model (DillateUNet) proposed in the present invention provides the most accurate myocardial segmentation, can well adapt to the fluctuation of shape and position of the myocardium due to relaxation and contraction, and can perform complete and accurate segmentation of the myocardium.
[0095] Based on the first embodiment, the input of the myocardial segmentation model (DillateUNet) can be any one of the three categories of myocardial contrast echocardiograms and myocardial contrast echocardiogram videos. The second embodiment uses a batch of different myocardial contrast echocardiograms. For example, B represents the batch size, C represents the number of channels, H represents the image height, and W represents the image width; the initial input image is reshaped, and its resolution is adjusted to 512×512. The original number of channels of each input image is 3 (R, G, and B channels). The batch size B is set according to the computing power of the experimental equipment. In this embodiment 2, 4 Nvidia Geforce RTX 4090GPUs are used, and the batch size B is set to 16.
[0096] In the second embodiment, the number of layers of the encoder is set to 5, the number of upsampling blocks M in the decoder is 4, the number of dilated convolution layers in each dilated convolution module is set to 4, P is set to 1, the number of Transformer layers in the Transformer block is 12, the downsampling method is maximum pooling, and the kth myocardial contrast echocardiogram is As an input image, it is processed by the myocardial segmentation model, including:
[0097] Take the kth myocardial contrast echocardiogram x k Input the encoder to extract features and obtain the output feature maps of each layer of the encoder, including:
[0098] Will After the first layer of the encoder, which consists of two dilated convolution modules, the output feature map of the first layer of the encoder is obtained. After that, downsampling is performed to obtain the output feature map of the first layer after downsampling
[0099] Output feature map of the first layer after downsampling After the second layer of the encoder, the second layer of the encoder consists of two dilated convolution modules, and the output feature map of the second layer of the encoder is obtained. After that, downsampling is performed to obtain the output feature map of the second layer after downsampling
[0100] Output feature map of the second layer after downsampling After the third layer of the encoder, the third layer of the encoder consists of three dilated convolution modules, and the output feature map of the third layer of the encoder is obtained. After that, downsampling is performed to obtain the output feature map of the third layer after downsampling
[0101] Output feature map of the third layer after downsampling After the fourth layer of the encoder, the fourth layer of the encoder consists of three dilated convolution modules, and the output feature map of the fourth layer of the encoder is obtained. After that, downsampling is performed to obtain the output feature map of the fourth layer after downsampling
[0102]
[0103] Output feature map of the fourth layer after downsampling After the fifth layer of the encoder, the fifth layer of the encoder consists of three dilated convolution modules, and the output feature map of the fifth layer of the encoder is obtained.
[0104] The output feature map of the fifth layer of the encoder Input the bottleneck layer and perform feature fusion to obtain fused features, including:
[0105] In this embodiment 2, the output feature map of the fifth layer of the encoder is Reshape the output feature map of the fifth layer of the encoder The channel dimension is converted from 1024 to 768 to obtain the reshaped feature map Reshape the feature map Flatten into N patches The collection of all patches forms a shape The matrix y p ;in,
[0106] The matrix y p Linearly mapped to D-dimensional space, each patch is from The dimension is mapped to the D dimension, and the formula is:
[0107]
[0108] in, is the D-dimensional vector corresponding to the e-th patch block, is the mapping matrix, y p e is the e-th patch block, b p It is the bias term of linear mapping. In this embodiment, linear mapping is performed through ordinary convolution, and the ordinary convolution kernel size and step size are both set to P.
[0109] After adding position encoding to the D-dimensional vector corresponding to each patch block of the output feature map of the last layer of the encoder, a two-dimensional sequence is obtained. The formula is:
[0110]
[0111] Among them, X Feat is a two-dimensional sequence, is the D-dimensional vector corresponding to the Nth patch block, E pos is the position encoding matrix.
[0112] The two-dimensional sequence Input to the Transformer block for feature fusion to obtain the initial fusion feature t 0 , the size and shape of the initial fusion feature t 0 The shape is still (1024, 768);
[0113] The initial fusion feature t 0Input to the first Transformer layer in the Transformer block to obtain the output features of the first Transformer layer, including:
[0114] The initial fusion features Through the layer normalization (LN layer) of the first Transformer layer, the output features of the LN layer of the first Transformer layer are obtained, and the formula is:
[0115]
[0116] Among them, t 1 ′ is the output feature of the LN layer of the first Transformer layer, μ is X Feat The mean value in the feature dimension, σ 2 For X Feat The variance in the feature dimension, ε is a constant, ⊙ is the Hadamard product, γ 1 is a learnable scaling parameter used to adjust the range of normalized eigenvalues, β 1 is a learnable offset parameter used to adjust the offset of the normalized eigenvalue.
[0117] The output feature t of the LN layer of the first Transformer layer 1 ′ Through the multi-head attention mechanism of the first Transformer layer, the output feature S of the multi-head attention mechanism of the first Transformer layer is obtained 1 ′ ,include:
[0118] The output feature t of the LN layer of the first Transformer layer 1 ′ Apply three different linear projections and get t 1 ′ The query vector Q = W Q t 1 ′ , key vector K = W K t 1 ′ , value vector V = W V t 1 ′ , where W Q is the learnable query projection matrix, W K is the learnable key projection matrix, W V is the learnable value projection matrix, W Q , W K , h is the number of attention heads;
[0119] Based on t 1 ′ The query vector Q of the attention mechanism of the τth head τ , key vector K τ , value vector V τ , calculate t 1 ′ Get the self-attention score S of the τth head τ , the formula is:
[0120]
[0121] Among them, S τ t 1 ′ The self-attention score of the τth head is obtained. Softmax(.) is the Softmax function. The Softmax function normalizes the attention score into a probability distribution. (.) T For transposition, τ=1,2…h.
[0122] The self-attention scores of the h heads are concatenated after the final linear projection to obtain the output feature S of the multi-head attention mechanism of the first Transformer layer 1 ′ , the formula is:
[0123]
[0124] Among them, S 1 ′ is the output feature of the multi-head attention mechanism of the first Transformer layer, W o is the final linear projection.
[0125] The output feature t of the LN layer of the first Transformer layer 1 ′ The output feature S of the multi-head attention mechanism of the first Transformer layer 1 ′ Add together to get the fused attention feature t of the first Transformer layer 1 ″;
[0126] The fused attention feature t of the first Transformer layer 1 Through the multi-layer perceptron, the output feature t of the multi-layer perceptron of the first Transformer layer is obtained 1 ″′, the multilayer perceptron consists of two linear layers and one nonlinear activation function (GELU), and the formula is:
[0127] t 1 ″′=ω fc1 (GELU(ω fc2 (LN(S 1 ″))+b fc2 ))+b fc1 ,
[0128] Among them, t 1 ″′ is the output feature of the multi-layer perceptron of the first Transformer layer, is the first projection matrix, is the second projection matrix, M is the dimension of the intermediate layer projection, GELU(.) is a nonlinear activation function that scales the input according to the probability of the normal distribution, and b fc1 is the first bias term, b fc2 is the second bias term.
[0129] The output feature t of the multi-layer perceptron of the first Transformer layer 1 ″′ and fused attention feature t 1 ″Add together to get the output features of the first Transformer layer
[0130] Similarly, the formula for the output features of other Transformer layers is:
[0131] t ′ θ ′ =MAS(LN(t θ;1 ))+t θ;1 , θ=1,2…12
[0132]
[0133] Among them, t ′ θ ′ is the fused attention feature of the θth Transformer layer, MAS(.) is the multi-head self-attention mechanism, LN(.) is layer normalization, MLP(.) is the multi-layer perceptron, t θ;1 is the input feature of the θth Transformer layer, is the output feature of the θth Transformer layer.
[0134] Reshape the initial fusion features so that the size of the initial fusion features is consistent with the reshaped feature map. Consistent, get fusion features
[0135] The output feature maps of each layer of the encoder are combined with the fusion features Input the decoder to get the kth myocardial contrast echocardiogram x k The myocardial segmentation results include:
[0136] The fusion features Through the 3×3 convolution layer of the first upsampling block, the output feature map of the 3×3 convolution layer of the first upsampling block is obtained The output feature map of the 3×3 convolutional layer of the first upsampling block Through the 2×2 up-convolution of the first up-sampling block, the output feature map of the first up-sampling block is obtained
[0137] The output feature map of the first upsampling block Output feature map of the fourth layer of the encoder After stacking in the channel dimension, the output feature map of the first upsampling block after stacking is obtained
[0138] The output feature map of the first upsampling block after stacking Through the 3×3 convolution layer of the second upsampling block, the output feature map of the 3×3 convolution layer of the second upsampling block is obtained The output feature map of the 3×3 convolutional layer of the second upsampling block Through the 2×2 up-convolution of the second up-sampling block, the output feature map of the second up-sampling block is obtained.
[0139] The output feature map of the second upsampling block Output feature map of the third layer of the encoder After stacking in the channel dimension, the output feature map of the second upsampling block after the stacking is obtained
[0140] The output feature map of the second upsampling block after stacking Through the 3×3 convolution layer of the third upsampling block, the output feature map of the 3×3 convolution layer of the third upsampling block is obtained The output feature map of the 3×3 convolutional layer of the third upsampling block Through the 2×2 up-convolution of the third up-sampling block, the output feature map of the third up-sampling block is obtained.
[0141] The output feature map of the third upsampling block The output feature map of the second layer of the encoder After stacking in the channel dimension, the output feature map of the third upsampling block after stacking is obtained
[0142] The output feature map of the third upsampling block after stacking Through the 3×3 convolution layer of the fourth upsampling block, the output feature map of the 3×3 convolution layer of the fourth upsampling block is obtained. The output feature map of the 3×3 convolutional layer of the fourth upsampling block Through the 2×2 up-convolution of the fourth up-sampling block, the output feature map of the fourth up-sampling block is obtained
[0143] The output feature map of the fourth upsampling block The output feature map of the first layer of the encoder After stacking in the channel dimension, the output feature map of the fourth upsampling block after stacking is obtained
[0144] The output feature map of the fourth upsampling block after stacking After a 3×3 convolutional layer, we get the output feature map of the 3×3 convolutional layer.
[0145] The output feature map of the 3×3 convolutional layer Input to the segmentation head, the kth myocardial contrast echocardiogram x k All pixels in are divided into two categories, and the final output is the kth myocardial contrast echocardiogram x k Myocardial segmentation results
[0146] In the second embodiment, by setting the number of layers of the encoder to 5, the rich and hierarchical feature information of myocardial contrast echocardiography can be gradually extracted at different depth levels, making the feature extraction more comprehensive and precise. The number of dilated convolution layers in each dilated convolution module is set to 4, and the receptive field can be effectively expanded by combining dilated convolution layers with different dilation rates without increasing too much computation, so as to better capture the contextual information and detail features in the image. P is set to 1, so that each pixel in the output feature map of the fifth layer of the encoder becomes a patch, and each pixel in the feature map can be interacted to maximize the fusion of features. While the parameters of Transformer are relatively large, the number of layers of Transformer layers in the Transformer block is set to 12 layers, which can fully utilize the powerful feature fusion and representation capabilities of the Transformer architecture while reducing the model parameters and the amount of computation, and perform deep and comprehensive feature interaction and fusion on the patch block vector after position encoding, so as to obtain more representative and accurate initial fusion features, and finally obtain more accurate myocardial segmentation results of myocardial contrast echocardiography, which improves the performance and effect of the myocardial segmentation model as a whole.
[0147] The third embodiment provides a myocardial segmentation system based on myocardial contrast echocardiography, including:
[0148] A model building module, used to build a myocardial segmentation model, wherein the myocardial segmentation model includes: an encoder, a bottleneck layer and a decoder;
[0149] The encoder includes a plurality of layers connected in sequence along the forward propagation direction, each layer includes a plurality of dilated convolution modules connected in sequence, each dilated convolution module includes a plurality of dilated convolution layers with different dilation rates connected in sequence, each dilated convolution layer includes a dilated convolution, a batch normalization layer and a ReLU activation function connected in sequence, and the output features of all the dilated convolution layers in each dilated convolution module are stacked in the channel dimension to obtain an output feature map of each dilated convolution module;
[0150] A feature extraction module, for extracting features by inputting the myocardial contrast echocardiogram into an encoder, and obtaining output feature maps of each layer of the encoder;
[0151] The fusion module is used to input the output feature map of the last layer of the encoder into the bottleneck layer to perform feature fusion and obtain fused features;
[0152] The segmentation module is used to input the output feature maps of each layer of the encoder and the fusion features into the decoder to obtain the myocardial segmentation result of myocardial contrast echocardiography.
[0153] The fourth embodiment provides a myocardial segmentation device based on myocardial contrast echocardiography, including:
[0154] The memory is used to store the computer program; the processor is used to implement the steps of the above-mentioned myocardial segmentation method based on myocardial contrast echocardiography when executing the computer program.
[0155] In this embodiment, specifically, a myocardial segmentation device based on myocardial contrast echocardiography includes: a storage device for storing data sets and intermediate results in the training process; a central processing unit (CPU) for controlling processes and processing auxiliary tasks, including data preprocessing; an image processing unit (GPU) for accelerating deep learning training; a memory (RAM) for temporarily storing data and model parameters; a power supply and cooling system for ensuring stable operation of the system and avoiding overheating and power instability; and a network device for inter-node communication in large-scale distributed training.
[0156] The GPU is any one of brands such as NVIDIA, RTX, Tesla, etc.
[0157] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0158] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0159] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0161] Obviously, the above embodiments are merely examples for the purpose of clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the present invention.
Claims
1. A myocardial segmentation method based on myocardial contrast echocardiography, characterized in that: The following steps are involved: Constructing a myocardial segmentation model, wherein the myocardial segmentation model includes: an encoder, a bottleneck layer and a decoder; The encoder includes a plurality of layers connected in sequence along the forward propagation direction, each layer includes a plurality of dilated convolution modules connected in sequence, each dilated convolution module includes a plurality of dilated convolution layers with different dilation rates connected in sequence, each dilated convolution layer includes a dilated convolution, a batch normalization layer and a ReLU activation function connected in sequence, and the output features of all the dilated convolution layers in each dilated convolution module are stacked in the channel dimension to obtain an output feature map of each dilated convolution module; Input the myocardial contrast echocardiogram into the encoder to extract features and obtain output feature maps of each layer of the encoder; The output feature map of the last layer of the encoder is input into the bottleneck layer for feature fusion to obtain fused features; The output feature maps of each layer of the encoder and the fused features are input into the decoder to obtain the myocardial segmentation results of myocardial contrast echocardiography.
2. The myocardial segmentation method based on myocardial contrast echocardiography according to claim 1, characterized in that: When the number of layers of the encoder is 5, the first layer and the second layer of the encoder each include two sequentially connected dilated convolution modules, and the third layer, the fourth layer, and the fifth layer of the encoder each include three sequentially connected dilated convolution modules; When the number of layers of the encoder is greater than 5, the first layer and the second layer of the encoder each include two sequentially connected dilated convolution modules, the third layer, the fourth layer, and the fifth layer of the encoder each include three sequentially connected dilated convolution modules, and the remaining layers of the encoder each include multiple sequentially connected dilated convolution modules.
3. The myocardial segmentation method based on myocardial contrast echocardiography according to claim 1, characterized in that: The input features of the dilated convolution module are input into the dilated convolution module, and the output feature map of the dilated convolution module is output, including: The input features of the i-th dilated convolution layer in the dilated convolution module are passed through the dilated convolution of the i-th dilated convolution layer to obtain the output features of the dilated convolution of the i-th dilated convolution layer. The formula is: Z i =ω i *F i +b, The output features of the dilated convolution of the i-th dilated convolution layer are processed by the batch normalization layer to obtain the output features of the dilated convolution of the i-th dilated convolution layer after batch normalization. The formula is: The output features of the dilated convolution of the i-th dilated convolution layer after batch normalization are processed by the ReLU activation function to obtain the output features of the i-th dilated convolution layer, and the formula is: The output features of all dilated convolutional layers are stacked in the channel dimension to obtain the output feature map of the dilated convolution module. The formula is: Among them, Z i is the output feature of the dilated convolution of the i-th dilated convolutional layer, ω i is the convolution kernel of the dilated convolution of the i-th dilated convolution layer in the dilated convolution module, F i is the input feature of the i-th dilated convolutional layer in the dilated convolutional module, b is the bias term, * is the convolution operation, i=1,2…I, I is the number of dilated convolutional layers in the dilated convolutional module, is the output feature of the dilated convolution of the i-th dilated convolution layer after batch normalization, · is the dot product, μ B is the mean of the output features of the dilated convolution of the ith dilated convolutional layer corresponding to the myocardial contrast echocardiogram, is the variance of the output feature of the dilated convolution of the ith dilated convolutional layer corresponding to the myocardial contrast echocardiogram, ε is a constant, γ is the first learnable parameter, β is the second learnable parameter, is the output feature of the i-th dilated convolutional layer, max(.) is the maximum value, is the output feature map of the dilated convolution module, To stack in the channel dimension.
4. The myocardial segmentation method based on myocardial contrast echocardiography according to claim 3, characterized in that: The input features of the dilated convolution module are input into the dilated convolution module, and the output feature map of the dilated convolution module is output. The formula is: in, is the output feature map of the dilated convolution module, DCM(.) is the dilated convolution module, To stack on the channel dimension, ReLU(.) is the ReLU activation function, BN(.) is the batch normalization layer, and DilatedConv i (.) is the dilated convolution of the i-th dilated convolution layer in the dilated convolution module, I is the number of dilated convolution layers in the dilated convolution module, and F i is the input feature of the i-th dilated convolution layer in the dilated convolution module, and F1 is the input feature of the dilated convolution module.
5. The myocardial segmentation method based on myocardial contrast echocardiography according to claim 1, characterized in that: The output feature map of the last layer of the encoder is input into the bottleneck layer to perform feature fusion to obtain fused features, including: The output feature map of the last layer of the encoder is reshaped to obtain a reshaped feature map, which is flattened into N patches. The set of all patches forms a shape of A matrix of is the height of the output feature map of the last layer of the encoder, is the width of the output feature map of the last layer of the encoder, and P is the size of the Patch block; The matrix is linearly mapped to the D-dimensional space, and each patch block is converted into its corresponding D-dimensional vector; After adding position encoding to the D-dimensional vector corresponding to each patch block of the output feature map of the last layer of the encoder, it is input into the Transformer block for feature fusion to obtain the initial fused feature; The initial fusion feature is reshaped to make the size of the initial fusion feature consistent with the reshaped feature map to obtain the fusion feature.
6. The myocardial segmentation method based on myocardial contrast echocardiography according to claim 5, characterized in that: The Transformer block includes multiple Transformer layers connected in sequence, and each Transformer layer includes: layer normalization, multi-head attention mechanism and multi-layer perceptron connected in sequence.
7. The myocardial segmentation method based on myocardial contrast echocardiography according to claim 1, characterized in that: The decoder includes: Multiple upsampling blocks, 3×3 convolutional layers and a segmentation head. Each upsampling block consists of a 3×3 convolutional layer and a 2×2 upconvolution. Each 3×3 convolutional layer includes a 3×3 convolution and a ReLu activation function.
8. The myocardial segmentation method based on myocardial contrast echocardiography according to claim 7, characterized in that: The output feature maps of each layer of the encoder and the fusion features are input into the decoder to obtain the myocardial segmentation result of the myocardial contrast echocardiogram, including: Process the fused features through the first upsampling block to obtain the output feature map of the first upsampling block; The output feature map of the jth upsampling block is stacked with the output feature map of the encoder's M-j+1th layer in the channel dimension and used as the input of the j+1th upsampling block; where j=1,2…M, M is the number of upsampling blocks; The output feature map of the Mth upsampling block and the output feature map of the first layer of the encoder are stacked in the channel dimension and processed by a 3×3 convolutional layer and a segmentation head in sequence to obtain the myocardial segmentation result of myocardial contrast echocardiography.
9. A myocardial segmentation system based on myocardial contrast echocardiography, characterized in that: include: A model building module, used to build a myocardial segmentation model, wherein the myocardial segmentation model includes: an encoder, a bottleneck layer and a decoder; The encoder includes a plurality of layers connected in sequence along the forward propagation direction, each layer includes a plurality of dilated convolution modules connected in sequence, each dilated convolution module includes a plurality of dilated convolution layers with different dilation rates connected in sequence, each dilated convolution layer includes a dilated convolution, a batch normalization layer and a ReLU activation function connected in sequence, and the output features of all the dilated convolution layers in each dilated convolution module are stacked in the channel dimension to obtain an output feature map of each dilated convolution module; A feature extraction module, for extracting features by inputting the myocardial contrast echocardiogram into an encoder, and obtaining output feature maps of each layer of the encoder; The fusion module is used to input the output feature map of the last layer of the encoder into the bottleneck layer to perform feature fusion and obtain fused features; The segmentation module is used to input the output feature maps of each layer of the encoder and the fusion features into the decoder to obtain the myocardial segmentation result of myocardial contrast echocardiography.
10. A myocardial segmentation device based on myocardial contrast echocardiography, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of a myocardial segmentation method based on myocardial contrast echocardiography in claims 1 to 8 when executing the computer program.
Citation Information
Patent Citations
Retinal macular edema multi-lesion image segmentation method
CN110349162A
Method and system for segmenting echocardiogram based on recursive aggregation deep learning
CN111368899A
Image segmentation method based on hole heterogeneous convolution
CN115631137A
Retinal blood vessel image segmentation method based on multi-scale expansion convolution residual network
CN117593317A
Ultrasonic diaphragm parameter automatic measurement method and device based on multi-scale expansion convolution, medium and product
CN118505596A
Cited By
Key structure image segmentation method and system for obstetrical ultrasound
CN122049379A