Segmentation method for assisting in nesting UNet model in echocardiography image based on SAM

By adopting a nested UNet model based on SAM assist in echocardiography image segmentation, problems such as high noise, low contrast and blurred boundaries are solved, and high-precision left ventricular segmentation is achieved, reducing the dependence on manual labeled data, and improving the robustness and adaptability of the model.

CN120014273APending Publication Date: 2025-05-16SHENYANG INSTITUTE OF CHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510096866.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing echocardiographic image segmentation methods face challenges such as high noise, low contrast and blurred boundaries. They rely on a large amount of high-quality manual annotation data, which is costly and time-consuming to obtain. It is difficult for traditional models to make full use of the potential information in the existing annotation data.

Method used

Using a nested UNet model based on SAM assist, the model segmentation performance is optimized by building a DA-UNet model, introducing SCBAM attention mechanism, using task tokens and integrating SAM assisted segmentation networks, combining Dice loss, BCE loss and SAM loss.

Benefits of technology

It significantly improves the accuracy of automated left ventricular segmentation, reduces dependence on high-quality manual annotation data, enhances the model's understanding and learning ability of the target area, and improves segmentation accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014273A_ABST
    Figure CN120014273A_ABST
Patent Text Reader

Abstract

The invention discloses a segmentation method of an SAM-assisted nested UNet model in an echocardiography image, and relates to a medical image processing method. The method comprises the following steps: constructing a DA-UNet model; a SimAM attention mechanism and a CBAM attention mechanism are fused, a task token (TaskToken) is introduced, and an SAM auxiliary segmentation network is integrated. Designing a loss function fusion strategy; dice loss, binary cross entropy loss (BCELoss) and SAM loss are combined, model segmentation performance is optimized in a weighted summation mode, and precision and stability are improved while a comprehensive loss function is ensured. According to the method, the limitation of the traditional UNet in the aspects of feature expression, global context capture, over-fitting prevention and control, multi-scale feature fusion and the like is overcome, the accuracy and robustness of medical image segmentation are remarkably improved, and an efficient and reliable solution is provided for the segmentation task of the left ventricle in the echocardiogram.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a medical image processing method, in particular to a segmentation method in ultrasonic cardiogram based on SAM-assisted nested UNet model. Background Art

[0002] According to the World Health Organization, heart disease is one of the leading causes of death worldwide. Accurate and effective early diagnosis is of great significance for the treatment of heart disease and reducing mortality. The left ventricular ejection fraction (EF value) reflects hemodynamics and myocardial contractility. The lower the EF value, the more severe the ventricular remodeling and the worse the myocardial contractile function. If the EF value is not calculated accurately due to unclear echocardiographic images, it may delay the treatment of heart disease patients and even be fatal in many cases. Therefore, in current medical research, automatic and accurate assessment of left ventricular ejection fraction is crucial. However, most existing automated left ventricular segmentation methods (the core step in calculating EF values) rely on static images, and these methods rely on a large amount of high-quality manually annotated data, which is costly and time-consuming to obtain, while traditional models usually find it difficult to fully utilize the potential information in existing annotated data.

[0003] A key problem facing current medical decision-making support systems is the lack of high-quality manually annotated data, which limits their practical clinical applications. Therefore, providing task-related contextual information to the model and enhancing the model's understanding and learning ability of the target area have become urgent issues to be solved in this field. The research can not only accurately segment the left ventricle, but also achieve accurate and effective early diagnosis of heart disease.

[0004] There are several main problems that need to be solved in left ventricular segmentation optimization technology:

[0005] 1) Echocardiographic images often face challenges such as high noise, low contrast, and blurred boundaries. Traditional medical image segmentation methods are prone to missed detection or false detection when dealing with complex cardiac structures.

[0006] 2) Existing methods rely on a large amount of high-quality manually annotated data, which is costly and time-consuming to obtain;

[0007] 3) Traditional models usually find it difficult to fully utilize the potential information in existing annotated data. Summary of the invention

[0008] The purpose of the present invention is to provide a segmentation method based on the SAM-assisted nested UNet model in ultrasound cardiac images. The method is tested on the EchoNet-Dynamic dataset, and the results show that the model is significantly superior to the advanced models in recent years in key evaluation indicators such as the Dice coefficient and intersection over union (IoU). This improvement not only improves the accuracy of automated left ventricular segmentation, but also provides a more lightweight and efficient solution for automated assisted diagnosis in practical applications.

[0009] The present invention is achieved through the following technical solutions:

[0010] In a first aspect, the present invention provides a method for segmenting echocardiographic images based on a nested UNet model assisted by SAM, which specifically comprises the following steps:

[0011] S1: Construct DA-UNet model; optimize the structure based on UNet++, introduce deep supervision and structural pruning, reduce unnecessary computing and storage overhead through the independent output of sub-networks, thereby improving the efficiency and accuracy of medical image segmentation. At the same time, adopt the nested UNet structure (NestedUNet) and multi-level pooling and upsampling operations, combined with the VGG module to strengthen multi-scale feature fusion.

[0012] S2: Construct the SCBAM attention mechanism; integrate the SimAM and CBAM attention mechanisms. CBAM includes the channel attention module (CAM) and the spatial attention module (SAM). SimAM optimizes local features through parameter-free energy functions to improve computational efficiency and feature expression capabilities.

[0013] S3: Introducing TaskToken; By fusing the learnable task token with the input feature map, it provides task-related contextual information without adding additional parameters, thereby enhancing segmentation performance and adaptability.

[0014] S4: Integrated SAM-assisted segmentation network; by dividing the input image into patches of fixed size, combining position embedding and pre-trained SAM model, it provides additional supervision signals to the main model during training, improving segmentation accuracy and robustness.

[0015] S5: Design a loss function fusion strategy; combine Dice loss, binary cross entropy loss (BCELoss) and SAM loss, and use a weighted summation method to optimize the model segmentation performance to ensure that the comprehensive loss function improves both accuracy and stability.

[0016] As a preferred solution, step S1 specifically includes the following steps:

[0017] S1-1: Construct the DA-UNet model; optimize the structure based on UNet++, introduce deep supervision and structural pruning, reduce unnecessary computing and storage overhead through the independent output of sub-networks, and thus improve the efficiency and accuracy of medical image segmentation.

[0018] S1-2: Introduce the nested UNet structure (NestedUNet); through the nested design of deep feature expression, use multi-level pooling and upsampling operations, combined with the VGG module, the VGG network includes 3×3 convolution of the input image, passing through the BN layer and then the ReLU function, the image is further convolved 3×3, passing through the BN layer and then the ReLU function, the image is passed through the SCBAM attention mechanism to obtain the feature map, and the feature map is output to the VGG network. Strengthen multi-scale feature fusion to further improve segmentation accuracy.

[0019] S1-3: Optimize network lightweight; reduce redundant parts through structural pruning, introduce SCBAM module to optimize network performance, significantly reduce computing and storage overhead while ensuring segmentation accuracy, and enhance the practical application potential of the model in medical environments.

[0020] As a preferred solution, step S2 specifically includes the following steps:

[0021] S2-1: Channel attention mechanism operation: CAM captures the relationship between channels through global average pooling and global maximum pooling, generates channel attention, and gives a given input feature map The calculation process of global average pooling and global maximum pooling of the channel attention module is as shown in formulas (1) to (2).

[0022]

[0023] F max (c) = max i,j F(c,i,j) (2)

[0024] Where H and W represent the height and width of the feature map, F(c,i,j) represents the input feature map, and F avg is the global average pooling result, F max is the global maximum pooling result.

[0025] S2-2: Shared network: pass these two descriptors through a shared multi-layer perceptron (MLP) to generate a channel attention map The calculation process is as shown in formula (3).

[0026] M c =σ(MLP(F avg )+MLP(F max )) (3)

[0027] Where σ is the Sigmoid activation function, and the MLP consists of two fully connected layers. The output dimension of the first fully connected layer is C / r (r is the reduction rate), and the output dimension of the second fully connected layer is C.

[0028] S2-3: Attention weighting: Multiply the input feature map with the channel attention map. The calculation process is as shown in formula (4).

[0029] F″=M c ⊙ F′ (4)

[0030] Among them, F′ is the input feature map, M c is the channel attention map, F″ is the weighted output feature map, and ⊙ represents element-by-element multiplication.

[0031] S2-4: The SimAM module calculates the attention weight by optimizing an energy function. The energy function is based on the spatial inhibition phenomenon in neuroscience. An energy function is defined to measure the importance of each neuron. The calculation process is shown in formula (5).

[0032]

[0033] Among them, μ and σ are the mean and variance of the channel respectively, λ is a hyperparameter, and t is a neuron.

[0034] S2-5: Calculate the importance weight of each neuron by solving the closed-form solution of the energy function. The calculation process is as shown in formula (6).

[0035]

[0036] Among them, d is the sum of squared deviations of the feature map, v is the channel variance, and λ is a hyperparameter.

[0037] S2-6: Use the attention weight to reweight the input feature map to obtain the output feature map. The calculation process is as shown in formula (7).

[0038] X att =X×σ(E inv ) (7)

[0039] Among them, σ is the Sigmoid function and X is the input feature map.

[0040] S2-7: Spatial attention mechanism operation: The SAM module captures the relationship between spatial positions through global average pooling and global maximum pooling along the channel direction to generate a spatial attention map. Given an input feature map,

[0041] Global average pooling and global maximum pooling in the channel direction obtain two descriptors and The calculation process is as shown in formulas (8) to (9).

[0042]

[0043] F m ' ax (i,j)=max c F′(c,i,j) (9)

[0044] Where F′(c,i,j) represents the input feature map, C is the number of channels, and F a ' vg is the global average pooling result, F m ' ax is the global maximum pooling result.

[0045] S2-8: Connection and convolution operation: connect the two descriptors in the channel dimension and generate a spatial attention map through the convolution layer The calculation process is as shown in formula (10).

[0046] M s =σ(Conv([F a ' vg ; F′ max ])) (10)

[0047] Among them, σ is the Sigmoid activation function, [;] represents the connection operation in the channel dimension, and Conv is a 7×7 convolutional layer.

[0048] S2-9: Attention weighting: Multiply the input feature map with the spatial attention map. The calculation process is as shown in formula (11).

[0049] F″=M s ⊙F′ (11)

[0050] Among them, F′ is the input feature map, M s is the spatial attention map, and F″ is the weighted output feature map.

[0051] S2-10: When entering the attention mechanism, the feature map first passes through the CAM module, then the SimAM module, and then the SAM module.

[0052] As a preferred solution, step S3 specifically includes the following steps:

[0053] S3-1: Introduce the task token mechanism; design learnable task token parameters, and capture the contextual information of specific segmentation tasks through back-propagation optimization, so as to provide task-specific features for the model and enhance its segmentation accuracy.

[0054] S3-2: Expansion and concatenation operations: The expansion operation adjusts the task token parameters in the spatial dimension and batch size to generate a tensor that matches the input feature map, and concatenates it with the feature map in the channel dimension to provide more task-related information.

[0055] S3-3: Convolution processing and performance improvement; The concatenated feature map restores the number of channels to the original dimension through convolution operation, enhancing the model's feature representation ability and segmentation effect, while not significantly increasing the number of model parameters, improving segmentation accuracy and efficiency.

[0056] As a preferred solution, step S4 specifically includes the following steps:

[0057] S4-1: Introduce the SAM auxiliary segmentation module; convert the input image into an embedding vector through patch embedding and position embedding, and combine the pre-trained SAM model to perform auxiliary mask generation to improve the segmentation performance and accuracy of the main model.

[0058] S4-2: Generate auxiliary masks using the SAM model; during training, bounding boxes extracted based on the true mask are passed to the SAM predictor as hints to generate accurate segmentation masks, which are further processed by the mask decoder.

[0059] S4-3: Collaboratively optimize the main model and SAM model; combine the mask generated by SAM with the prediction result of the main model, and jointly optimize the binary cross entropy loss (BCE loss) and the main loss (Dice loss and BCE loss) to significantly improve the performance and robustness of the overall segmentation system.

[0060] As a preferred solution, step S5 specifically includes the following steps:

[0061] S5-1: Introduce a comprehensive loss function; By combining Dice loss (DiceLoss), binary cross entropy loss (BCELoss) and SAM loss (SAMLoss), the segmentation performance of the DA-UNet model is optimized, and the similarity between the model prediction and the true mask, the classification error, and the difference between the auxiliary mask generated by SAM and the prediction result of the main model are comprehensively measured.

[0062] S5-2: Define Dice loss; calculate the overlap between the predicted mask and the true mask based on the Dice similarity coefficient (DSC), use formula (12) to measure the accuracy of the segmentation result, and use a smoothing factor with a denominator of zero to stabilize the calculation. Calculate BCE loss; use formula (13) to measure the difference between the predicted probability and the true label of each pixel, and use the Sigmoid activation function to convert the predicted value into a probability to optimize the classification performance of the model.

[0063]

[0064] Where P and T are the predicted mask (prediction value) and the true mask (label), respectively, |P∩T| represents the number of pixels of the intersection of the predicted mask and the true mask, |P| and |T| represent the total number of pixels of the predicted mask and the true mask, respectively, and smooth is a smoothing factor to prevent the denominator from being zero. N is the total number of pixels, T i is the true label of the i-th pixel, P i is the predicted value of the i-th pixel, which is converted into probability through the Sigmoid activation function σ.

[0065] S5-3: Define SAM loss; calculate the binary cross entropy loss between the auxiliary mask generated by SAM and the prediction result of the main model through formula (14), providing additional supervision signals for the model to further improve the segmentation accuracy and robustness.

[0066] SAM Loss=BCE Loss(P main ,P SAM ) (14)

[0067] Where: P main is the predicted output of the main model, P SAM is the auxiliary mask generated by SAM.

[0068] S5-4: Combining three loss functions; Dice loss, BCE loss and SAM loss are combined by weighted summation to form a comprehensive loss function. The weight coefficients (sum) of different losses in formula (15) can be adjusted according to task requirements to optimize the overall segmentation performance.

[0069] Total Loss = α Dice Dice Loss + α BCE BCE Loss+α SAM ·SAM Loss (15)

[0070] Where: α Dice is the weight coefficient of DICE loss, α BCE is the weight coefficient α of BCE loss SAM is the weight coefficient of SAM loss.

[0071] In a second aspect, the present invention provides an application of a method for segmenting an echocardiographic image based on a nested UNet model assisted by SAM, comprising the following units:

[0072] 1) Optimize the unit and merge CBAM and SimAM into SCBAM;

[0073] 2) Building a unit that integrates BN, RELU, Conv, and SCBAM into the VGG module;

[0074] 3) The first combining unit (Nestedunet), used to combine the VGG module with up- and down-sampling and skip connections;

[0075] 4) Optimization unit, optimize the UNet structure to L3 pruning;

[0076] 5) Construction unit, integrating expansion operation, concatenation operation and convolution operation into task token;

[0077] 6) The second combined unit (DA-UNet) replaces the pruned UNet internal convolution with the first combined unit (Nestedunet), and fuses the output convolution Conv3 with the task token;

[0078] 7) Construction unit, integrating PatchEmbedding, PositionEmbedding, and MaskDecoder into the SAM assisted segmentation module;

[0079] 8) The third combined unit (SAM-DAUNet), the SAM auxiliary segmentation module provides additional supervision signals for the second combined unit through the SAM loss (SAMLoss).

[0080] In a third aspect, the present invention provides an electronic device for using the method, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the application of the SAM-assisted nested UNet model in echocardiographic image segmentation is performed.

[0081] In a fourth aspect, the present invention provides a computer-readable storage medium in the method, on which a computer program is stored, and when the computer program is executed by a processor, the application of the SAM-assisted nested UNet model in echocardiographic image segmentation is executed.

[0082] The advantages of the present invention are:

[0083] The present invention aims at the common noise interference, low contrast and blurred boundary problems in the current echocardiographic image segmentation, and reduces the dependence on a large amount of high-quality manually annotated data. By extracting the bounding box from the real mask and inputting it into the SAM model as a prompt, an auxiliary mask is generated, and then compared with the prediction result of the nested UNet, the auxiliary loss is calculated and integrated into the total loss, thereby providing an additional supervision signal for the model and enhancing the model's understanding and learning ability of the target area. In addition, the model also introduces a learnable task token to provide it with task-related contextual information, further improving the generalization ability and adaptability of the model. In terms of loss function design, the traditional DiceLoss and BCE Loss are combined, as well as the auxiliary loss based on the SAM generated mask. These losses are integrated into the main loss through weighted coefficients, effectively mining the potential information in the annotated data, thereby improving the learning effect of the model. In view of the noise, low contrast and blurred boundary characteristics of echocardiographic images, the method proposed by the present invention significantly reduces false detection and missed detection. Through testing on the EchoNet-Dynamic dataset, the results show that the model is significantly better than the advanced models in recent years in key evaluation indicators such as the Dice coefficient and intersection over union (IoU), especially in comparison with the traditional UNet model, the Dice coefficient has increased by at least 1.14%, proving the effectiveness of the model in the left ventricular volume (ESV, EDV) segmentation task. This improvement not only improves the accuracy of automated left ventricular segmentation, but also provides a more lightweight and efficient solution for automated auxiliary diagnosis in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0085] Figure 1 Schematic diagram of the SAM-DAUNet network model structure in the present invention;

[0086] Figure 2 This is the structural diagram of the DA-UNet++ network model in the present invention;

[0087] Figure 3 This is the structure diagram of the SCNAM attention mechanism;

[0088] Figure 4 This is the structural diagram of the CBAM attention mechanism;

[0089] Figure 5 This is the structure diagram of the task token module. DETAILED DESCRIPTION

[0090] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0091] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0092] Example 1

[0093] Combine the following Figures 1 to 5 The application of a SAM-assisted nested UNet model implemented in the present invention in echocardiographic image segmentation is specifically described.

[0094] S1: Construct DA-UNet model; optimize the structure based on UNet++, introduce deep supervision and structural pruning, reduce unnecessary computing and storage overhead through the independent output of sub-networks, thereby improving the efficiency and accuracy of medical image segmentation. At the same time, adopt the nested UNet structure (NestedUNet) and multi-level pooling and upsampling operations, combined with the VGG module to strengthen multi-scale feature fusion.

[0095] S2: Construct the SCBAM attention mechanism; integrate the SimAM and CBAM attention mechanisms. CBAM includes the channel attention module (CAM) and the spatial attention module (SAM). SimAM optimizes local features through parameter-free energy functions to improve computational efficiency and feature expression capabilities.

[0096] S3: Introducing TaskToken; By fusing the learnable task token with the input feature map, it provides task-related contextual information without adding additional parameters, thereby enhancing segmentation performance and adaptability.

[0097] S4: Integrated SAM-assisted segmentation network; by dividing the input image into patches of fixed size, combining position embedding and pre-trained SAM model, it provides additional supervision signals to the main model during training, improving segmentation accuracy and robustness.

[0098] S5: Design a loss function fusion strategy; combine Dice loss, binary cross entropy loss (BCELoss) and SAM loss, and use a weighted summation method to optimize the model segmentation performance to ensure that the comprehensive loss function improves both accuracy and stability.

[0099] According to the application of the SAM-assisted nested UNet model in echocardiographic image segmentation according to claim 1, S1 includes the following specific processes:

[0100] S1-1: Construct the DA-UNet model; optimize the structure based on UNet++, introduce deep supervision and structural pruning, reduce unnecessary computing and storage overhead through the independent output of sub-networks, and thus improve the efficiency and accuracy of medical image segmentation.

[0101] S1-2: Introduce the nested UNet structure (NestedUNet); through the nested design of deep feature expression, use multi-level pooling and upsampling operations, combined with the VGG module, the VGG network includes 3×3 convolution of the input image, passing through the BN layer and then the ReLU function, the image is further convolved 3×3, passing through the BN layer and then the ReLU function, the image is passed through the SCBAM attention mechanism to obtain the feature map, and the feature map is output to the VGG network. Strengthen multi-scale feature fusion to further improve segmentation accuracy.

[0102] S1-3: Optimize network lightweight; reduce redundant parts through structural pruning, introduce SCBAM module to optimize network performance, significantly reduce computing and storage overhead while ensuring segmentation accuracy, and enhance the practical application potential of the model in medical environments.

[0103] The application of the SAM-assisted nested UNet model in echocardiographic image segmentation, S2 includes the following specific processes:

[0104] S2-1: Channel attention mechanism operation: CAM captures the relationship between channels through global average pooling and global maximum pooling, generates channel attention, and gives a given input feature map The calculation process of global average pooling and global maximum pooling of the channel attention module is as shown in formulas (1) to (2).

[0105]

[0106] F max (c) = max i,j F(c,i,j) (2)

[0107] Where H and W represent the height and width of the feature map, F(c,i,j) represents the input feature map, and F avg is the global average pooling result, F max is the global maximum pooling result.

[0108] S2-2: Shared network: pass these two descriptors through a shared multi-layer perceptron (MLP) to generate a channel attention map The calculation process is as shown in formula (3).

[0109] M c =σ(MLP(F avg )+MLP(F max )) (3)

[0110] Where σ is the Sigmoid activation function, and the MLP consists of two fully connected layers. The output dimension of the first fully connected layer is C / r (r is the reduction rate), and the output dimension of the second fully connected layer is C.

[0111] S2-3: Attention weighting: Multiply the input feature map with the channel attention map. The calculation process is as shown in formula (4).

[0112] F″=M c ⊙ F′ (4)

[0113] Among them, F′ is the input feature map, M c is the channel attention map, F″ is the weighted output feature map, and ⊙ represents element-by-element multiplication.

[0114] S2-4: The SimAM module calculates the attention weight by optimizing an energy function. The energy function is based on the spatial inhibition phenomenon in neuroscience. An energy function is defined to measure the importance of each neuron. The calculation process is shown in formula (5).

[0115]

[0116] Among them, μ and σ are the mean and variance of the channel respectively, λ is a hyperparameter, and t is a neuron.

[0117] S2-5: Calculate the importance weight of each neuron by solving the closed-form solution of the energy function. The calculation process is as shown in formula (6).

[0118]

[0119] Among them, d is the sum of squared deviations of the feature map, v is the channel variance, and λ is a hyperparameter.

[0120] S2-6: Use the attention weight to reweight the input feature map to obtain the output feature map. The calculation process is as shown in formula (7).

[0121] X att =X×σ(E inv ) (7)

[0122] Among them, σ is the Sigmoid function and X is the input feature map.

[0123] S2-7: Spatial attention mechanism operation: The SAM module captures the relationship between spatial positions through global average pooling and global maximum pooling along the channel direction to generate a spatial attention map. Given an input feature map,

[0124] Global average pooling and global maximum pooling in the channel direction obtain two descriptors and The calculation process is as shown in formulas (8) to (9).

[0125]

[0126] F m ' ax (i,j)=max c F′(c,i,j) (9)

[0127] Where F′(c,i,j) represents the input feature map, C is the number of channels, and F a ' vg is the global average pooling result, F m ' ax is the global maximum pooling result.

[0128] S2-8: Connection and convolution operation: connect the two descriptors in the channel dimension and generate a spatial attention map through the convolution layer The calculation process is as shown in formula (10).

[0129] M s =σ(Conv([F a ' vg ; F′ max ])) (10)

[0130] Among them, σ is the Sigmoid activation function, [;] represents the connection operation in the channel dimension, and Conv is a 7×7 convolutional layer.

[0131] S2-9: Attention weighting: Multiply the input feature map with the spatial attention map. The calculation process is as shown in formula (11).

[0132] F″=M s ⊙F′ (11)

[0133] Among them, F′ is the input feature map, M s is the spatial attention map, and F″ is the weighted output feature map.

[0134] S2-10: When entering the attention mechanism, the feature map first passes through the CAM module, then the SimAM module, and then the SAM module.

[0135] The application of the SAM-assisted nested UNet model in echocardiographic image segmentation, S3 includes the following specific processes:

[0136] S3-1: Introduce the task token mechanism; design learnable task token parameters, and capture the contextual information of specific segmentation tasks through back-propagation optimization, so as to provide task-specific features for the model and enhance its segmentation accuracy.

[0137] S3-2: Expansion and concatenation operations: The expansion operation adjusts the task token parameters in the spatial dimension and batch size to generate a tensor that matches the input feature map, and concatenates it with the feature map in the channel dimension to provide more task-related information.

[0138] S3-3: Convolution processing and performance improvement; The concatenated feature map restores the number of channels to the original dimension through convolution operation, enhancing the model's feature representation ability and segmentation effect, while not significantly increasing the number of model parameters, improving segmentation accuracy and efficiency.

[0139] The application of the SAM-assisted nested UNet model in echocardiographic image segmentation, S4 includes the following specific processes:

[0140] S4-1: Introduce the SAM auxiliary segmentation module; convert the input image into an embedding vector through patch embedding and position embedding, and combine the pre-trained SAM model to perform auxiliary mask generation to improve the segmentation performance and accuracy of the main model.

[0141] S4-2: Generate auxiliary masks using the SAM model; during training, bounding boxes extracted based on the true mask are passed to the SAM predictor as hints to generate accurate segmentation masks, which are further processed by the mask decoder.

[0142] S4-3: Collaboratively optimize the main model and SAM model; combine the mask generated by SAM with the prediction result of the main model, and jointly optimize the binary cross entropy loss (BCE loss) and the main loss (Dice loss and BCE loss) to significantly improve the performance and robustness of the overall segmentation system.

[0143] The application of the SAM-assisted nested UNet model in echocardiographic image segmentation, S5 includes the following specific processes:

[0144] S5-1: Introduce a comprehensive loss function; By combining Dice loss (DiceLoss), binary cross entropy loss (BCELoss) and SAM loss (SAMLoss), the segmentation performance of the DA-UNet model is optimized, and the similarity between the model prediction and the true mask, the classification error, and the difference between the auxiliary mask generated by SAM and the prediction result of the main model are comprehensively measured.

[0145] S5-2: Define Dice loss; calculate the overlap between the predicted mask and the true mask based on the Dice similarity coefficient (DSC), use formula (12) to measure the accuracy of the segmentation result, and use a smoothing factor with a denominator of zero to stabilize the calculation. Calculate BCE loss; use formula (13) to measure the difference between the predicted probability and the true label of each pixel, and use the Sigmoid activation function to convert the predicted value into a probability to optimize the classification performance of the model.

[0146]

[0147] Where P and T are the predicted mask (prediction value) and the true mask (label), respectively, |P∩T| represents the number of pixels of the intersection of the predicted mask and the true mask, |P| and |T| represent the total number of pixels of the predicted mask and the true mask, respectively, and smooth is a smoothing factor to prevent the denominator from being zero. N is the total number of pixels, T i is the true label of the i-th pixel, P i is the predicted value of the i-th pixel, which is converted into probability through the Sigmoid activation function σ.

[0148] S5-3: Define SAM loss; calculate the binary cross entropy loss between the auxiliary mask generated by SAM and the prediction result of the main model through formula (14), providing additional supervision signals for the model to further improve the segmentation accuracy and robustness.

[0149] SAM Loss=BCE Loss(P main ,P SAM ) (14)

[0150] Where: P main is the predicted output of the main model, P SAM is the auxiliary mask generated by SAM.

[0151] S5-4: Combining three loss functions; Dice loss, BCE loss and SAM loss are combined by weighted summation to form a comprehensive loss function. The weight coefficients (sum) of different losses in formula (15) can be adjusted according to task requirements to optimize the overall segmentation performance.

[0152] Total Loss = α Dice Dice Loss + α BCE BCE Loss+α SAM ·SAM Loss (15)

[0153] Where: α Dice is the weight coefficient of DICE loss, α BCE is the weight coefficient α of BCE lossSAM is the weight coefficient of SAM loss.

[0154] In this embodiment, a new echocardiographic image segmentation model is proposed, which combines the pruned nested (UNet) architecture and the "SegmentAnythingModel" (SAM) auxiliary segmentation mechanism for accurate segmentation of the left ventricular end-diastolic volume (EDV) and end-systolic volume (ESV) of the echocardiogram in the video dataset. It is mainly composed of a DA-UNet++ network, a task token, a SAM-assisted segmentation network, and a jointly optimized loss function. The DA-UNet++ network adopts a nested UNet (NestedUNet) design based on an encoding-decoding structure. The network compresses image information by encoding layer by layer and restores image details by decoding to achieve efficient image processing. The nested NestedUNet processes multi-resolution feature maps through multi-layer pooling and upsampling operations, while retaining high-resolution detail information through jump connections. In order to enhance the feature expression capability, DA-UNet++ introduces the VGG module, and uses stacked convolutional layers and batch normalization for feature extraction. In addition, we combine SCBAM (Channel and Spatial Attention Module), first performing channel attention operation, and then performing spatial attention processing through SimAM to optimize feature representation. On the last convolution layer of the network, we introduce a learnable task token (TaskToken), which provides task-related contextual information to the model by splicing with the feature map in the channel dimension, thereby improving the expressiveness of features in multiple dimensions. To further improve the segmentation performance, this model integrates a pre-trained SAM model. SAM generates auxiliary masks by accepting the bounding box extracted from the input image and the true mask. These masks act as additional supervision signals and work together with the loss function of the main model. During training, the auxiliary masks generated by SAM are compared with the probability map output by the main model to calculate an additional loss term. This loss term is jointly optimized with the Dice loss of the main model through the binary cross entropy loss (BCELoss) and is weighted proportionally and integrated into the total loss. Through this strategy, the model can effectively use the auxiliary information generated by SAM during the optimization process, provide additional supervision signals for DA-UNet++, enhance the model's understanding and learning ability of the target area, and improve segmentation accuracy and robustness. In general, this model fully utilizes the advantages of different loss functions by weighted fusion of the main loss module (Dice loss and BCE loss) and the auxiliary loss module (SAM-assisted BCE loss), thereby achieving more accurate and robust medical image segmentation. This multi-level, multi-objective loss fusion strategy not only improves segmentation performance, but also enhances the model's adaptability to complex and variable data, providing the possibility for deployment in actual medical environments.

[0155] In order to verify the effectiveness of DA-UNet++ network, task token, and SAM-assisted segmentation network, this paper conducted an ablation study on DA-UNet++ network, task token, and SAM-assisted segmentation network. In order to exclude the influence of other factors, all experiments were run in the Pytorch environment on NVIDIA GeForce RTX4090Ti GPU. The learning rate was set to 1e-4, batch_size was 2, and epoch was 50. As each module was added in turn, the segmentation results improved. The complete method achieved the best DSC and IoU results, proving the effectiveness of each module in improving segmentation performance. As shown in Table 1.

[0156] Table 1 The role of each module in the entire framework

[0157]

[0158] In order to verify the segmentation performance of the DA-UNet++ network, the following six models were compared (see Table 1 for the comparison results), namely the reference model Deeplabv3, Deeplabv3+, UNet, ResNet34, FCN, PSPNet, and SAM-DAUNet. In order to exclude the influence of other factors, all experiments were run in the Pytorch environment on the NVIDIA GeForce RTX4090Ti GPU. The learning rate was set to 1e-4, the batch_size was 2, and the epoch was 50. Among the seven models on the EchoNet-Dynamic dataset, the SAM-DAUNet network performed best in the DSC and IoU indicators on the EchoNet-Dynamic dataset. Compared with the UNet model, the Dice score increased by 1.14% and the IoU score increased by 0.8%, which further demonstrated the effectiveness of integrating SAM auxiliary supervision, which greatly helped guide the model to improve the segmentation performance and provided the possibility for deployment in actual medical environments. As shown in Table 2.

[0159] Table 2 Model comparison

[0160]

[0161] Example 2

[0162] The present invention proposes an application of a SAM-assisted nested UNet model in echocardiographic image segmentation, comprising the following units:

[0163] 1) Optimize the unit and merge CBAM and SimAM into SCBAM;

[0164] 2) Building a unit that integrates BN, RELU, Conv, and SCBAM into the VGG module;

[0165] 3) The first combining unit (Nestedunet), used to combine the VGG module with up- and down-sampling and skip connections;

[0166] 4) Optimization unit, optimize the UNet structure to L3 pruning;

[0167] 5) Construction unit, integrating expansion operation, concatenation operation and convolution operation into task token;

[0168] 6) The second combined unit (DA-UNet) replaces the pruned UNet internal convolution with the first combined unit (Nestedunet), and fuses the output convolution Conv3 with the task token;

[0169] 7) Construction unit, integrating PatchEmbedding, PositionEmbedding, and MaskDecoder into the SAM assisted segmentation module;

[0170] 8) The third combined unit (SAM-DAUNet), the SAM auxiliary segmentation module provides additional supervision signals for the second combined unit through the SAM loss (SAMLoss).

[0171] Example 3

[0172] The present invention proposes an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory through the bus, and when the machine-readable instructions are executed by the processor, an application of a SAM-assisted nested UNet model in echocardiographic image segmentation is executed.

[0173] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0174] Example 4

[0175] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the application of a SAM-assisted nested UNet model in echocardiographic image segmentation is executed.

[0176] Storage media include: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks or optical disks, and other media that can store programs.

[0177] Finally, it should be emphasized that the above content is only a preferred embodiment of the present invention, which is intended to help better understand and implement the technical solution of the present invention and does not constitute a limitation of the present invention. Although we have described the present invention in detail through specific embodiments, professionals familiar with the technology in the field can still make appropriate modifications and adjustments to their technical solutions based on these embodiments, or make equivalent replacements for certain technical features. Any modification, replacement, optimization or improvement made without departing from the spirit and basic principles of the present invention should be deemed to fall within the scope of protection of the present invention. We encourage and expect technicians in related fields to further develop and improve the present invention according to specific application requirements, in order to promote progress and innovation in this technical field.

Claims

1. A segmentation method based on SAM-assisted nested UNet model in ultrasound cardiac images, characterized in that: The method comprises the following steps: S1: Construct DA-UNet model; optimize the structure based on UNet++, introduce deep supervision and structure pruning, and use nested UNet structure (NestedUNet) and multi-level pooling and upsampling operations, combined with VGG module to strengthen multi-scale feature fusion; S2: Construct the SCBAM attention mechanism; integrate the SimAM and CBAM attention mechanisms. CBAM includes the channel attention module (CAM) and the spatial attention module (SAM). SimAM optimizes local features through parameter-free energy functions to improve computational efficiency and feature expression capabilities. S3: Introducing TaskToken; By fusing the learnable task token with the input feature map, it provides task-related context information without adding additional parameters, thereby enhancing segmentation performance and adaptability; S4: Integrated SAM-assisted segmentation network; by dividing the input image into fixed-size patches, combined with position embedding and pre-trained SAM model, it provides additional supervision signals to the main model during training, improving segmentation accuracy and robustness; S5: Design a loss function fusion strategy; combine Dice loss, binary cross entropy loss (BCELoss) and SAM loss, and use a weighted summation method to optimize the model segmentation performance to ensure that the comprehensive loss function improves both accuracy and stability.

2. According to claim 1, a segmentation method based on SAM-assisted nested UNet model in ultrasound cardiac images is characterized in that: In step S1, specifically The process includes: S1-1: Construct DA-UNet model; optimize the structure based on UNet++, introduce deep supervision and structure pruning, reduce unnecessary computing and storage overhead through the independent output of sub-networks, thereby improving the efficiency and accuracy of medical image segmentation; S1-2: Introduce the nested UNet structure (NestedUNet); through the nested design of deep feature expression, use multi-level pooling and upsampling operations, combined with the VGG module, the VGG network includes: 3×3 convolution of the input image, passing through the BN layer and then the ReLU function, the image is further convolved with 3×3, passing through the BN layer and then the ReLU function, the image is passed through the SCBAM attention mechanism to obtain the feature map, and the feature map is output to the VGG network; strengthen the multi-scale feature fusion to further improve the segmentation accuracy; S1-3: Optimize network lightweight; reduce redundant parts through structural pruning, introduce SCBAM module to optimize network performance, significantly reduce computing and storage overhead while ensuring segmentation accuracy, and enhance the practical application potential of the model in medical environments.

3. The segmentation method based on SAM-assisted nested UNet model in ultrasound cardiac images according to claim 1, characterized in that: The step S2 specifically includes the following process: S2-1: Channel attention mechanism operation: CAM captures the relationship between channels through global average pooling and global maximum pooling, generates channel attention, and gives a given input feature map The calculation process of global average pooling and global maximum pooling of the channel attention module is as shown in formulas (1) to (2). F max (c)=max i,j F(c,i,j) (2) Where H and W represent the height and width of the feature map, F(c,i,j) represents the input feature map, and F avg is the global average pooling result, F max is the global maximum pooling result; S2-2: Shared network: pass these two descriptors through a shared multi-layer perceptron (MLP) to generate a channel attention map The calculation process is as shown in formula (3). M c =σ(MLP(F avg )+MLP(F max )) (3) Where σ is the Sigmoid activation function, and the MLP consists of two fully connected layers. The output dimension of the first fully connected layer is C / r (r is the reduction rate), and the output dimension of the second fully connected layer is C; S2-3: Attention weighting: Multiply the input feature map with the channel attention map. The calculation process is as shown in formula (4). F″=M c ⊙ F′ (4) Among them, F′ is the input feature map, M c is the channel attention map, F″ is the weighted output feature map, and ⊙ represents element-by-element multiplication; S2-4: The SimAM module calculates the attention weight by optimizing an energy function. The energy function is based on the spatial inhibition phenomenon in neuroscience and defines an energy function to measure the importance of each neuron. The calculation process is shown in formula (5). Among them, μ and σ are the mean and variance of the channel respectively, λ is a hyperparameter, and t is a neuron; S2-5: Calculate the importance weight of each neuron by solving the closed-form solution of the energy function. The calculation process is as shown in formula (6). Where d is the sum of squared deviations of the feature map, v is the channel variance, and λ is a hyperparameter; S2-6: Use the attention weight to reweight the input feature map to obtain the output feature map. The calculation process is as shown in formula (7), X att =X×σ(E inv ) (7) Among them, σ is the Sigmoid function, X is the input feature map; S2-7: Spatial attention mechanism operation: The SAM module captures the relationship between spatial positions through global average pooling and global maximum pooling along the channel direction to generate a spatial attention map; given an input feature map, global average pooling and global maximum pooling in the channel direction obtain two descriptors and The calculation process is as shown in formulas (8) to (9). F m ′ ax (i,j)=max c F′(c,i,j) (9) Where F′(c,i,j) represents the input feature map, C is the number of channels, and F a ' vg is the global average pooling result, F m ' ax is the global maximum pooling result; S2-8: Connection and convolution operation: connect the two descriptors in the channel dimension and generate a spatial attention map through the convolution layer The calculation process is as shown in formula (10). M s =σ(Conv([F a ′ vg ;F′ max ])) (10) Among them, σ is the Sigmoid activation function, [;] represents the connection operation in the channel dimension, and Conv is a 7×7 convolution layer; S2-9: Attention weighting: Multiply the input feature map with the spatial attention map. The calculation process is as shown in formula (11). F″=M s ⊙F′ (11) Among them, F′ is the input feature map, M s is the spatial attention map, and F″ is the weighted output feature map; S2-10: When entering the attention mechanism, the feature map first passes through the CAM module, then the SimAM module, and then the SAM module.

4. The segmentation method based on SAM-assisted nested UNet model in ultrasound cardiac images according to claim 1, characterized in that: In step S3, specifically The process includes: S3-1: Introduce the task token mechanism; design learnable task token parameters, and capture the context information of specific segmentation tasks through back-propagation optimization, so as to provide task-specific features for the model and enhance its segmentation accuracy; S3-2: Expansion and concatenation operations: The expansion operation adjusts the task token parameters in the spatial dimension and batch size to generate a tensor that matches the input feature map, and concatenates it with the feature map in the channel dimension to provide more task-related information. S3-3: Convolution processing and performance improvement; The concatenated feature map restores the number of channels to the original dimension through convolution operation, enhancing the model's feature representation ability and segmentation effect, while not significantly increasing the number of model parameters, improving segmentation accuracy and efficiency.

5. The segmentation method based on SAM-assisted nested UNet model in ultrasound cardiac images according to claim 1, characterized in that: In step S4, specifically The process includes: S4-1: Introduce the SAM auxiliary segmentation module; convert the input image into an embedding vector through patch embedding and position embedding, and combine the pre-trained SAM model to generate auxiliary masks to improve the segmentation performance and accuracy of the main model; S4-2: Generate auxiliary mask using SAM model; During the training process, bounding boxes extracted based on the ground-truth masks are passed as hints to the SAM predictor to generate accurate segmentation masks, which are further processed by the mask decoder. S4-3: Collaboratively optimize the main model and SAM model; combine the mask generated by SAM with the prediction result of the main model, and jointly optimize the binary cross entropy loss (BCE loss) and the main loss (Dice loss and BCE loss) to improve the performance and robustness of the overall segmentation system.

6. The segmentation method based on SAM-assisted nested UNet model in ultrasound cardiac images according to claim 1, characterized in that: The step S5 specifically includes the following process: S5-1: Introduce a comprehensive loss function; By combining Dice loss (DiceLoss), binary cross entropy loss (BCELoss) and SAM loss (SAMLoss), the segmentation performance of the DA-UNet model is optimized, and the similarity between the model prediction and the true mask, the classification error, and the difference between the auxiliary mask generated by SAM and the prediction result of the main model are comprehensively measured; S5-2: Define Dice loss; calculate the overlap between the predicted mask and the true mask based on the Dice similarity coefficient (DSC), use formula (12) to measure the accuracy of the segmentation result, and prevent the smoothing factor with a denominator of zero from being used for stable calculation; calculate BCE loss; use formula (13) to measure the difference between the predicted probability and the true label of each pixel, and use the Sigmoid activation function to convert the predicted value into probability to optimize the classification performance of the model; Where P and T are the predicted mask (prediction value) and the true mask (label), respectively. |P∩T| represents the number of pixels of the intersection of the predicted mask and the true mask. |P| and |T| represent the total number of pixels of the predicted mask and the true mask, respectively. Smooth is a smoothing factor to prevent the denominator from being zero. N is the total number of pixels. T i is the true label of the i-th pixel, P i is the predicted value of the i-th pixel, which is converted into probability through the Sigmoid activation function σ; S5-3: Define SAM loss; calculate the binary cross entropy loss between the auxiliary mask generated by SAM and the prediction result of the main model through formula (14), providing additional supervision signals for the model to further improve the segmentation accuracy and robustness; SAM Loss=BCE Loss(P main ,P SAM ) (14) Where: P main is the predicted output of the main model, P SAM is the auxiliary mask generated by SAM; S5-4: Combining three loss functions; Dice loss, BCE loss and SAM loss are combined by weighted summation to form a comprehensive loss function. The weight coefficients (sum) of different losses in formula (15) can be adjusted according to task requirements to optimize the overall segmentation performance; Total Loss=a Dice ·Dice Loss+a BCE ·BCE Loss+a SAM ·SAM Loss (15) Where: α Dice is the weight coefficient of DICE loss, α BCE is the weight coefficient α of BCE loss SAM is the weight coefficient of SAM loss.

7. The segmentation method based on SAM-assisted nested UNet model in ultrasound cardiac images according to claim 1, characterized in that: The application of this method in echocardiographic image segmentation includes the following units: Optimize the unit and merge CBAM and SimAM into SCBAM; Build a unit that integrates BN, RELU, Conv, and SCBAM into the VGG module; The first combination unit (Nestedunet) is used to combine the VGG module with up and down sampling and skip connection optimization units, and optimize the UNet structure to L3 pruning; the construction unit integrates the expansion operation, splicing operation, and convolution operation into the task token (TaskToken); The second combination unit (DA-UNet) replaces the pruned UNet internal convolution with the first combination unit (Nestedunet), and fuses the output convolution Conv3 with the task token; the construction unit fuses the patch embedding (PatchEmbedding), position embedding (PositionEmbedding), and mask decoder (MaskDecoder) into the SAM auxiliary segmentation module; In the third combined unit (SAM-DAUNet), the SAM auxiliary segmentation module provides additional supervision signals for the second combined unit through the SAM loss (SAMLoss).

8. The segmentation method based on SAM-assisted nested UNet model in ultrasound cardiac images according to claim 1, characterized in that: The electronic device of the method includes: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the application of the SAM-assisted nested UNet model in echocardiographic image segmentation is performed.

9. The segmentation method based on SAM-assisted nested UNet model in ultrasound cardiac images according to claim 1, characterized in that: The method stores a computer program on a computer-readable storage medium, and when the computer program is executed by a processor, the computer program executes the application of the SAM-assisted nested UNet model in the segmentation of echocardiographic images.