SAM-based automatic prompt ultrasonic image segmentation method and system
Through the automatic prompt ultrasound image segmentation method of dual-branch encoder and two-stage decoding process, the dependence of SAM on manual prompts in ultrasound image segmentation is solved, the segmentation accuracy and generalization ability are improved, the influence of noise and shadow is reduced, and more efficient ultrasound image segmentation is achieved.
Patent Information
- Application Number
- CN202510644139.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-05-19
AI Technical Summary
Existing SAM-based segmentation methods rely on manual cues and have poor adaptability to ultrasound images, resulting in inaccurate segmentation results, especially in the presence of noise and shadows.
An automatic hinting ultrasound image segmentation method based on SAM was designed. It adopted a dual-branch encoder and a two-stage decoding process, combined with CNN classification knowledge and Transformer encoder, and automatically generated positive and negative point hints and mask hints to reduce the influence of noise and shadows and improve segmentation performance.
It effectively solves the problem of SAM's dependence on manual prompts, improves the accuracy and generalization ability of ultrasound image segmentation, reduces erroneous segmentation, and improves information capture efficiency and segmentation performance.
Smart Images

Figure CN120689610A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ultrasound (US) image processing, and in particular to a SAM-based automatic prompting ultrasound image segmentation method and system. Background Art
[0002] Ultrasound imaging is a non-invasive medical imaging technique that plays a key role in the diagnosis of a wide range of diseases due to its advantages such as high safety, low cost, and real-time performance. Currently, ultrasound imaging has been widely used in areas such as breast tumors, nerves, thyroid nodules, and fetal examination. However, speckle noise, shadowing artifacts, and signal attenuation are common in ultrasound images, causing many segmentation methods to misclassify these regions as target regions, severely impacting the accuracy of segmentation results. In recent years, with the rapid development of deep learning technology, its application in medical image analysis has become a research hotspot, and numerous deep learning-based ultrasound image segmentation methods have emerged. These methods, through the design of appropriate models and training strategies, have demonstrated excellent performance in specific segmentation tasks, significantly advancing research in the field of ultrasound image segmentation. However, these methods are often tailored to specific tasks and have limited generalization capabilities to other ultrasound image segmentation tasks. Recently, the Segment Anything Model (SAM) has attracted widespread attention and significantly advanced the development of foundational computer vision models. SAM has demonstrated excellent performance in natural image segmentation tasks and possesses zero-shot generalization capabilities to new image distributions and tasks. However, due to the scarcity of medical images in SAM's training data, its zero-shot generalization ability in medical images is poor. Furthermore, SAM's reliance on manual prompts in its semi-automatic segmentation mode not only interrupts the automatic segmentation process but also requires expert guidance, which is time-consuming and impractical in clinical settings. Therefore, further research and innovation are needed to address SAM's reliance on manual prompts and adapt it to ultrasound image segmentation tasks efficiently. Summary of the Invention
[0003] The technical problem to be solved by the present invention is as follows: In response to the above-mentioned problems in the prior art, a method and system for automatic prompting of ultrasound image segmentation based on SAM is provided. The present invention aims to solve the problem that most existing SAM-based segmentation methods rely on manual prompting and have poor adaptability to ultrasound images, reduce erroneous segmentation affected by ultrasonic noise and shadows, and improve the segmentation performance and generalization performance of the network.
[0004] In order to solve the above technical problems, the technical solution adopted by the present invention is: A method for automatic prompt ultrasound image segmentation based on SAM comprises the following steps: performing ultrasound image segmentation on an input ultrasound image using an ultrasound image segmentation network model to obtain a segmentation prediction map, wherein the ultrasound image segmentation network model comprises a CNN encoder, a classification head module, a low-level feature refinement module, a SAM Transformer encoder, a cross-branch attention module, a high-level feature enhancement module, a pixel decoder, a prompt encoder, and a mask decoder; performing ultrasound image segmentation on the input ultrasound image to obtain a segmentation prediction map comprises: extracting local detail features from the input ultrasound image through the CNN encoder and performing target existence classification through the classification head module, optimizing and enhancing the local detail features using the low-level feature refinement module and then outputting them to the pixel decoder; extracting global features from the input ultrasound image through the SAM Transformer encoder, and performing target existence classification through the classification head module. After supplementing local detail features through the cross-branch attention module, the image embedding is obtained, and then the advanced feature enhancement module is used to realize semantic feature enhancement through multi-scale feature extraction and then output to the pixel decoder; the first stage of decoding is performed by the pixel decoder to generate a preliminary segmentation prediction map, and the preliminary segmentation prediction map is thresholded to generate a mask prompt. If the result of the target existence classification indicates that there is a target, the top n points with the highest confidence in the preliminary segmentation prediction map are selected as positive point prompts, otherwise the top n points with the lowest confidence in the preliminary segmentation prediction map are selected as negative point prompts, and the positive point prompts or negative point prompts are encoded by the prompt encoder to generate a point embedding vector, and the mask prompts are encoded by the prompt encoder to generate a mask embedding vector. The point embedding vector, the mask embedding vector and the image obtained after supplementing the local detail features are embedded into the mask decoder to perform the second stage of decoding to obtain the final segmentation prediction map.
[0005] Optionally, the CNN encoder includes a total of five convolution modules, namely, convolution modules 1 to 5, which are connected in cascade. The local detail features output by the first three convolution modules of convolution modules 1 to 3 are not only used as the input of the next-level convolution module in the CNN encoder, but also as the input of a corresponding low-level feature refinement module. The local detail features output by convolution module 5 are used as the input of the classification head module to realize target existence classification. The Transformer encoder of the SAM includes a multi-level window Transformer layer. The input ultrasound image is first divided into blocks by the image block embedding layer and then combined with the position encoding as the input of the multi-level window Transformer layer. The image embedding obtained by supplementing the local detail features by the cross-branch attention module includes: the global features extracted by the multi-level window Transformer layer are normalized and then input into the cross-branch attention module. The cross-branch attention module uses the local detail features output by the CNN encoder as keys and values, and the global features as queries to perform cross attention. force operation, and then extract the attention features through a multi-head attention mechanism, add the extracted attention features and global features to obtain fused features, extract features from the local detail features output by the CNN encoder through multiple single convolution modules, normalize the fused features through layers, and then add the original fused features through a multi-layer perceptron, and then add the features extracted by multiple single convolution modules to obtain image embedding after supplementing the local detail features; the Transformer encoder of the SAM adopts a low-rank adaptive fine-tuning method to fine-tune the Transformer encoder during training, including: introducing a trainable low-rank bypass into the query and value projection layers in the multi-level window Transformer layer and the multi-head attention mechanism, the low-rank bypass consists of two low-rank matrices and an activation function to perform low-rank decomposition on the parameter update, and the low-rank matrix is a linear layer; during the training process, only the parameters of the low-rank bypass are optimized, and the weights of the original query and value projection layers are frozen; in the inference stage, the low-rank matrix is added to the original weights of the query and value projection layers to merge into new weights.
[0006] Optionally, the method of using a low-level feature refinement module to optimize and enhance local detail features and then outputting them to a pixel decoder includes: taking the local detail features as input feature maps, extracting features from the input feature maps using a channel attention mechanism, and subjecting the input feature maps to parallel average pooling operations and maximum pooling operations with a kernel size of 2×2, and average pooling operations and maximum pooling operations with a kernel size of 8×8, respectively, and then upsampling the results of the four pooling operations to the same resolution and then performing channel dimension splicing, and then generating a refined feature map through a 1×1 convolution module, residually connecting the refined feature map and the original input feature map, and then performing channel dimension splicing with the features extracted by the channel attention mechanism, and then generating a final output feature map through a 1×1 convolution module and outputting it to the pixel decoder.
[0007] Optionally, the use of the advanced feature enhancement module to achieve semantic feature enhancement through multi-scale feature extraction and then output to the pixel decoder includes: embedding the image as the input feature map, extracting features from the input feature map using the channel attention mechanism, and passing the input feature map in parallel through four dilated convolution branches for capturing features of different scales in the spatial dimension; among the four dilated convolution branches: the first dilated convolution branch consists of three 3×3 dilated convolutions with dilated rates of 1, 3, and 5, respectively, and one 1×1 convolution; the second dilated convolution branch consists of two 3×3 dilated convolutions with dilated rates of 1 and 3, respectively, and one 1×1 convolution; the third dilated convolution branch consists of a 3×3 dilated convolution with dilated rates of 3, respectively, and a 1×1 convolution; the fourth dilated convolution branch consists of a 3×3 dilated convolution with a dilated rate of 1; the features extracted by the channel attention mechanism and the features extracted by the four dilated convolution branches are spliced in the channel dimension to generate the final output feature map and output it to the pixel decoder.
[0008] Optionally, the pixel decoder includes three layers of pixel decoding modules, each layer of pixel decoding modules consists of a transposed convolution, a channel dimension connection module, and two convolution modules; the convolution module is composed of a 3×3 convolution layer, a batch normalization layer and a GELU activation function connected in sequence. The output features of the high-level feature enhancement module are input into the transposed convolution of the first layer as the input feature map of the pixel decoder to obtain a twice-upsampled feature map. The shallow features output by the first low-level feature refinement module are connected with the twice-upsampled feature map through the channel dimension connection module and then pass through the convolution module to obtain the output feature map of the first layer; the output feature map of the first layer is subjected to the transposed convolution of the second layer to obtain a four-fold upsampled feature map. The shallow features output by the second low-level feature refinement module are connected with the four-fold upsampled feature map through the channel dimension connection module and then pass through the convolution module to obtain the output feature map of the second layer; the output feature map of the second layer is subjected to the transposed convolution of the third layer to obtain an eight-fold upsampled feature map. The shallow features output by the third low-level feature refinement module are connected with the eight-fold upsampled feature map through the channel dimension connection module and then pass through the convolution module to obtain the output feature map of the third layer; the output feature map of the third layer is then subjected to a 1×1 convolution layer and a sigmoid activation function to obtain a preliminary segmentation prediction map.
[0009] Optionally, the mask decoder includes two mask decoding units, a multi-level transposed convolution module, an anti-cross attention module, a multi-layer perceptron and a dot product module, the mask decoding unit includes a self-attention module, an anti-cross attention module, a perceptron and a cross attention module connected in sequence, the point embedding vector is used as the input of the self-attention module in the first mask decoding unit, the image embedding obtained after supplementing the local detail features and the mask embedding vector are spliced as the input of the anti-cross attention module and the cross attention module in the first mask decoding unit, the output of the perceptron in the first mask decoding unit is used as the input of the self-attention module in the second mask decoding unit, and the output of the cross attention module in the first mask decoding unit is used as the input of the self-attention module It is the input of the anti-cross attention module and the cross attention module in the second mask decoding unit. The output of the perceptron in the second mask decoding unit is used as the first input of the anti-cross attention module outside the two mask decoding units. The output of the cross attention module in the second mask decoding unit is used as the second input of the anti-cross attention module outside the two mask decoding units and the input of the multi-stage transposed convolution module. The anti-cross attention module outside the two mask decoding units performs anti-cross attention on the two inputs and extracts features through the multi-layer perceptron as the first input of the dot product module. The output of the multi-stage transposed convolution module is used as the second input of the dot product module. The dot product module performs dot product operation on the two inputs to obtain the final segmentation prediction map.
[0010] Optionally, when training the ultrasound image segmentation network model, the loss function adopted by the pixel decoder is binary cross entropy loss, and the functional expression of the binary cross entropy loss is: , in, is the binary cross entropy loss, is the sample size, is the label of the i-th pixel value, is the preliminary segmentation prediction map of the i-th pixel value; the loss function adopted by the mask decoder is the weighted sum of the focus loss and the dice loss, and the function expression of the weighted sum of the focus loss and the dice loss is: , ,
[0011] in, is the weighted sum of focal loss and dice loss, is the balance factor between focus loss and dice loss, is the focal loss, is the dice loss, is the positive and negative sample balance factor, is the adjusted predicted probability of the i-th pixel value, is the final predicted probability of the i-th pixel value.
[0012] In addition, the present invention also provides a SAM-based automatic prompt ultrasound image segmentation system, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the SAM-based automatic prompt ultrasound image segmentation method.
[0013] In addition, the present invention also provides a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is programmed or configured to execute the SAM-based automatic prompt ultrasound image segmentation method through a processor.
[0014] In addition, the present invention also provides a computer program product, including a computer program or instructions, which is programmed or configured to execute the SAM-based automatic prompt ultrasound image segmentation method through a processor.
[0015] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: 1. The proposed SAM-based automatic hinting ultrasound image segmentation network model features a dual-branch encoder and a two-stage decoding process. The dual-branch encoder effectively extracts local detail features and global dependencies, while the two-stage decoding process automatically generates hint embeddings to guide the model for accurate segmentation, effectively applying SAM to the field of ultrasound image segmentation.
[0016] 2. The automatic prompting design of the SAM-based ultrasound image segmentation network model in this invention solves the problem of SAM's reliance on manual prompting. Furthermore, the positive and negative point prompting based on CNN classification knowledge effectively reduces missegmentation caused by ultrasound noise and shadows. The mask prompting also provides more detailed contextual information. The combination of these two types of automatic prompting improves the efficiency of information capture from the input image and enhances the ultrasound image segmentation performance of the ultrasound image segmentation network model.
[0017] 3. The high-level feature enhancement module and low-level feature refinement module of this invention effectively improve the quality of preliminary segmentation predictions and the accuracy of automatic suggestion. The high-level feature enhancement module is designed to capture multi-scale semantic features, while the low-level feature refinement module plays a key role at jump connections by refining shallow features of CNN branches. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Schematic diagram of the basic process of the method of the embodiment of the present invention.
[0019] Figure 2 Schematic diagram of the network structure of the ultrasound image segmentation network model in an embodiment of the present invention.
[0020] Figure 3 Schematic diagram of the network structure of the advanced feature enhancement module in an embodiment of the present invention.
[0021] Figure 4 Schematic diagram of the network structure of the low-level feature refinement module in an embodiment of the present invention.
[0022] Figure 5 Schematic diagram of the network structure of the first stage pixel decoder in an embodiment of the present invention.
[0023] Figure 6 Schematic diagram of the network structure of the second stage decoding process in an embodiment of the present invention.
[0024] Figure 7Figure 3 shows the segmentation results of BUSI ultrasound images using different segmentation methods in the embodiments of the present invention, where (a-1) to (e-1) are the segmentation prediction maps and label images obtained for the first sample using the U-Net method, TransUNet method, H-SAM method, and the method of this embodiment, respectively; (a-2) to (e-2) are the segmentation prediction maps and label images obtained for the second sample using the U-Net method, TransUNet method, H-SAM method, and the method of this embodiment, respectively; (a-3) to (e-3) are the segmentation prediction maps and label images obtained for the third sample using the U-Net method, TransUNet method, H-SAM method, and the method of this embodiment, respectively.
[0025] Figure 8 These are the incorrect segmentation results of normal images (without lesions) in the BUSI dataset using different existing segmentation methods in an embodiment of the present invention, where (a1) to (e1) are the first set of segmentation prediction maps obtained using the U-Net method, AttU-Net method, SAMed method, SAMUS method, and H-SAM method, respectively; and (a2) to (e2) are the second set of segmentation prediction maps obtained using the U-Net method, AttU-Net method, SAMed method, SAMUS method, and H-SAM method, respectively.
[0026] Figure 9 These are the segmentation results of UNS ultrasound images using different segmentation methods in the embodiments of the present invention, where (a-1) to (e-1) are the segmentation prediction maps and label images obtained for the first sample using the U-Net method, TransUNet method, H-SAM method, and the method of this embodiment, respectively; (a-2) to (e-2) are the segmentation prediction maps and label images obtained for the second sample using the U-Net method, TransUNet method, H-SAM method, and the method of this embodiment, respectively; (a-3) to (e-3) are the segmentation prediction maps and label images obtained for the third sample using the U-Net method, TransUNet method, H-SAM method, and the method of this embodiment, respectively.
[0027] Figure 10These are the segmentation results of STU ultrasound images using different segmentation methods in the embodiments of the present invention, where (a-1) to (e-1) are the segmentation prediction maps and label images obtained for the first sample using the U-Net method, TransUNet method, H-SAM method, and the method of this embodiment, respectively; (a-2) to (e-2) are the segmentation prediction maps and label images obtained for the second sample using the U-Net method, TransUNet method, H-SAM method, and the method of this embodiment, respectively; (a-3) to (e-3) are the segmentation prediction maps and label images obtained for the third sample using the U-Net method, TransUNet method, H-SAM method, and the method of this embodiment, respectively. DETAILED DESCRIPTION
[0028] The present invention will be described in further detail below with reference to the accompanying drawings and examples.
[0029] like Figure 1 and Figure 2As shown, the automatic prompt ultrasound image segmentation method based on SAM in this embodiment includes the following steps: using an ultrasound image segmentation network model to perform ultrasound image segmentation on the input ultrasound image to obtain a segmentation prediction map, the ultrasound image segmentation network model includes a CNN encoder, a classification head module, a low-level feature refinement module, a SAM Transformer encoder, a cross-branch attention module, a high-level feature enhancement module, a pixel decoder, a prompt encoder and a mask decoder, and performing ultrasound image segmentation on the input ultrasound image to obtain a segmentation prediction map includes: extracting local detail features from the input ultrasound image through the CNN encoder and performing target existence classification through the classification head module, optimizing and enhancing the local detail features using the low-level feature refinement module and then outputting them to the pixel decoder; extracting global features from the input ultrasound image through the SAM Transformer encoder , after supplementing the local detail features through the cross-branch attention module, the image embedding is obtained, and then the advanced feature enhancement module is used to achieve semantic feature enhancement through multi-scale feature extraction and then output to the pixel decoder; the pixel decoder performs the first stage of decoding to generate a preliminary segmentation prediction map, and the preliminary segmentation prediction map is thresholded to generate a mask prompt. If the result of the target existence classification indicates that there is a target, the top n points with the highest confidence in the preliminary segmentation prediction map are selected as positive point prompts, otherwise the top n points with the lowest confidence in the preliminary segmentation prediction map are selected as negative point prompts, and the positive point prompts or negative point prompts are encoded by the prompt encoder to generate a point embedding vector, and the mask prompts are encoded by the prompt encoder to generate a mask embedding vector. The point embedding vector, the mask embedding vector and the image obtained after supplementing the local detail features are embedded into the mask decoder to perform the second stage of decoding to obtain the final segmentation prediction map. In this embodiment, the CNN encoder and the Transformer encoder of SAM serve as the CNN encoding branch and the Transformer encoding branch respectively, forming a dual-branch encoder, and the pixel decoder and the mask decoder constitute a two-stage decoding process. The decoding stage includes: inputting the semantic features processed by the high-level feature enhancement module and the shallow features processed by the low-level feature refinement module into the pixel decoder to generate a preliminary segmentation prediction map, and generating automatic hints based on this; inputting the hint embedding processed by the hint encoder of the automatic hint together with the image embedding output by the dual-branch encoder into the mask decoder to generate the final segmentation prediction map.
[0030] like Figure 2As shown, the CNN encoder in this embodiment includes 5 convolution modules, namely convolution module 1 to convolution module 5, which are connected in cascade. The local detail features output by the first three convolution modules, namely convolution module 1 to convolution module 3, are not only used as the input of the next level convolution module in the CNN encoder, but also as the input of a corresponding low-level feature refinement module. The local detail features output by convolution module 5 are used as the input of the classification head module to realize target existence classification. As an optional implementation, the CNN encoder in this embodiment adopts the EfficientNetV2-S network, wherein: the EfficientNetV2-S network includes 7 stages, namely stage 0 to stage 7, and stages 0-1 correspond to Figure 2 Convolutional module 1 in stage 2 corresponds to Figure 2 Convolutional module 2 in stage 3 corresponds to Figure 2 In the convolution module 3, stages 4-5 correspond to Figure 2 Convolutional module 4 in stage 6 corresponds to Figure 2 Convolutional module 5 in stage 7 corresponds to Figure 2 In addition, CNN encoders with other structures can also be used as needed. The CNN encoder in this embodiment is pre-trained on the target existence classification task. During the segmentation training process, in order to improve training efficiency and ensure the stability of the classification results, the CNN encoder in this embodiment remains in a frozen state.
[0031] like Figure 2As shown, the Transformer encoder of the SAM in this embodiment includes a multi-level window Transformer layer. The input ultrasound image is first divided into blocks by the image block embedding layer and then combined with the position encoding as the input of the multi-level window Transformer layer. The image embedding is obtained after supplementing the local detail features by the cross-branch attention module, including: the global features extracted by the multi-level window Transformer layer are layer-normalized and then input into the cross-branch attention module. The cross-branch attention module is used to inject the local detail features extracted by the CNN encoding branch into the Transformer encoding branch to solve the problem of insufficient local features in the Transformer branch. The cross-branch attention module realizes feature fusion in the following ways: the local detail features output by the CNN encoder are used as keys and values through the cross-branch attention module, the global features are used as queries to perform cross-attention operations, and the attention features are extracted through the multi-head attention mechanism. The extracted attention features and global features are added to obtain fused features. The local detail features output by the CNN encoder are extracted through multiple single convolution modules. The fused features are normalized through layers and multi-layer perceptrons, and then added to the original fused features, and then added to the features extracted by multiple single convolution modules to obtain image embedding after supplementing the local detail features. The excellent segmentation performance of the SAM model is mainly due to its powerful image encoder architecture. The large image encoder enables it to have excellent feature extraction capabilities and cross-domain generalization capabilities. However, directly training the image encoder has the problem of high computational cost. In order to achieve efficient fine-tuning of the Transformer encoder of SAM, such as Figure 2 As shown, the Transformer encoder of SAM in this embodiment adopts a low-rank adaptive fine-tuning method to fine-tune the Transformer encoder during training, including: introducing a trainable low-rank bypass into the query and value projection layers in the multi-level window Transformer layer and the multi-head attention mechanism, wherein the low-rank bypass consists of two low-rank matrices and an activation function to perform low-rank decomposition on the parameter update, and the low-rank matrix is a linear layer; during the training process, only the parameters of the low-rank bypass are optimized, and the weights of the original query and value projection layers are frozen; in the inference stage, the low-rank matrix is added to the original weights of the query and value projection layers to form a new weight.
[0032] like Figure 2 As shown, the multiple single convolution modules in this embodiment are specifically 4 single convolution modules, and their number can be configured as needed. Among them, the single convolution module is a trainable network module, and each single convolution module includes a 3×3 convolution layer, a layer normalization layer, and a GELU activation function connected in sequence.
[0033] The target areas in ultrasound images usually have different sizes and irregular shapes. To address these problems, this embodiment designs a high-level feature enhancement module to effectively capture high-level semantic features from multiple scales. Considering the particularity and extraction difficulty of features in medical images, this embodiment uses 3×3 dilated convolutions with different dilation rates (1, 3, 5) to effectively capture features of different scales in the spatial dimension. Figure 3 As shown, the advanced feature enhancement module is used to achieve semantic feature enhancement through multi-scale feature extraction and then output to the pixel decoder, including: embedding the image as the input feature map, extracting features from the input feature map using the channel attention mechanism (enhancing the channel features by capturing the importance of different channels), and passing the input feature map in parallel through four dilated convolution branches for capturing features of different scales in the spatial dimension; among the four dilated convolution branches: the first dilated convolution branch consists of three 3×3 dilated convolutions with dilation rates of 1, 3, and 5, respectively, and one 1×1 convolution; the second dilated convolution branch consists of two 3×3 dilated convolutions with dilation rates of 1 and 3, respectively, and one 1×1 convolution; the third dilated convolution branch consists of a 3×3 dilated convolution with dilation rates of 3, respectively, and a 1×1 convolution; the fourth dilated convolution branch consists of a 3×3 dilated convolution with dilation rate of 1; the features extracted by the channel attention mechanism and the features extracted by the four dilated convolution branches are spliced in the channel dimension to generate the final output feature map and output it to the pixel decoder. The channel attention mechanism and four parallel dilated convolution branches constitute five branches. The feature maps obtained from the five branches are connected and refined through a 1×1 convolution to generate the final output feature map and output it to the pixel decoder.
[0034] In the encoder-decoder structure, skip connections play a key role in improving the accuracy and robustness of the segmentation network. In the ultrasound image segmentation task, the low-level features carried by the skip connection contain both edge detail information and interference noise. Direct connection may introduce noise to subsequent decoding. Therefore, this embodiment designs a low-level feature refinement module to optimize the shallow features at the skip connection. Figure 4As shown, the low-level feature refinement module is used to optimize and enhance the local detail features and then output them to the pixel decoder, including: taking the local detail features as the input feature map, extracting features from the input feature map using the channel attention mechanism, and passing the input feature map in parallel through the average pooling operation and maximum pooling operation with a kernel size of 2×2, and the average pooling operation and maximum pooling operation with a kernel size of 8×8, wherein the average pooling operation retains the overall average information, and the maximum pooling operation highlights the prominent information. By combining these pooling operations with different kernel sizes, the module captures the diversity and importance of the input feature map and suppresses irrelevant features; then the results of the four pooling operations are upsampled to the same resolution and spliced in the channel dimension, and then a 1×1 convolution module is used to generate a refined feature map, the refined feature map and the original input feature map are residually connected, and then spliced in the channel dimension with the features extracted by the channel attention mechanism (channel attention further enhances the module's processing capability in the channel dimension and suppresses unimportant or redundant features), and then a 1×1 convolution module is used to generate the final output feature map and output it to the pixel decoder.
[0035] The first stage decoding process is implemented using a pixel decoder, such as Figure 5 As shown in the figure, the pixel decoder includes three layers of pixel decoding modules. Each layer of pixel decoding module consists of a transposed convolution, a channel dimension connection module, and two convolution modules; the convolution module is composed of a 3×3 convolution layer, a batch normalization layer, and a GELU activation function connected in sequence. The output features of the high-level feature enhancement module are input into the transposed convolution of the first layer as the input feature map of the pixel decoder to obtain a twice-upsampled feature map. The shallow features output by the first low-level feature refinement module are connected with the twice-upsampled feature map through the channel dimension connection module and then passed through the convolution module to obtain the output feature map of the first layer; the output feature map of the first layer is subjected to the transposed convolution of the second layer to obtain a four-fold upsampled feature map. The shallow features output by the second low-level feature refinement module are connected with the four-fold upsampled feature map through the channel dimension connection module and then passed through the convolution module to obtain the output feature map of the second layer; the output feature map of the second layer is subjected to the transposed convolution of the third layer to obtain an eight-fold upsampled feature map. The shallow features output by the third low-level feature refinement module are connected with the eight-fold upsampled feature map through the channel dimension connection module and then passed through the convolution module to obtain the output feature map of the third layer; the output feature map of the third layer is processed by a 1×1 convolution layer and a sigmoid activation function to generate a preliminary segmentation prediction map, which will serve as the automatic prompt basis for the second-stage decoding process.
[0036] The second-stage decoding process consists of a hint encoder and a mask decoder, which optimize the final segmentation result by integrating automatic hint information. In this embodiment, point hints and mask hints are automatically generated during the second-stage decoding process using the preliminary segmentation prediction map generated by the first-stage decoding. The hint encoder generates point embedding vectors for the point hints. The positional encodings of the positive sample points are added to a learnable foreground embedding vector, while the positional encodings of the negative sample points are added to a learnable background embedding vector. The foreground and background embedding vectors serve as trainable parameters to assist the model in accurately identifying target and background regions. Because point hint labels extracted directly from the preliminary segmentation mask are difficult to guarantee accuracy, the present invention pre-trains the CNN branch on the object presence classification task. By utilizing the classification network to identify objects, the accuracy of point hints can be improved. Point hints are selected from the preliminary segmentation prediction map. If the classification result indicates the presence of an object, the top n points with the highest confidence in the preliminary segmentation prediction map are selected as positive point hints. Otherwise, the top n points with the lowest confidence in the preliminary segmentation prediction map are selected as negative point hints. The preliminary segmentation prediction map is thresholded to generate mask hints, thereby filtering out low-confidence areas that may lead to errors.
[0037] like Figure 6As shown, the mask decoder in this embodiment includes two mask decoding units, a multi-level transposed convolution module, an anti-cross attention module, a multi-layer perceptron and a dot product module, wherein the mask decoding unit includes a self-attention module, an anti-cross attention module, a perceptron and a cross attention module connected in sequence, and the point embedding vector is respectively used as the input of the self-attention module in the first mask decoding unit, and the image embedding and mask embedding vectors obtained after supplementing the local detail features are spliced as the input of the anti-cross attention module and the cross attention module in the first mask decoding unit respectively, and the output of the perceptron in the first mask decoding unit is used as the input of the self-attention module in the second mask decoding unit, and the cross attention module in the first mask decoding unit is used as the input of the self-attention module in the second mask decoding unit. The output of is used as the input of the anti-cross attention module and the cross attention module in the second mask decoding unit, the output of the perceptron in the second mask decoding unit is used as the first input of the anti-cross attention module outside the two mask decoding units, the output of the cross attention module in the second mask decoding unit is used as the second input of the anti-cross attention module outside the two mask decoding units and the input of the multi-stage transposed convolution module, the anti-cross attention module outside the two mask decoding units performs anti-cross attention on the two inputs and extracts features through the multi-layer perceptron as the first input of the dot product module, the output of the multi-stage transposed convolution module is used as the second input of the dot product module, and the dot product module performs dot product operation on the two inputs to obtain the final segmentation prediction map. Automatic point hints and automatic mask hints jointly guide model segmentation through a hint encoder and a mask decoder. The hint encoder encodes point hints and mask hints into point embedding vectors and mask embedding vectors, respectively. The mask embedding vectors are feature fused with the image embedding vectors generated by the dual-branch encoder to enhance the foreground context information of the original image embedding. After the point embedding vectors are processed by the self-attention mechanism, they undergo cross-attention interaction with the enhanced image embedding vectors, enabling the model to better identify target areas and reduce the segmentation of non-target areas.
[0038] In this embodiment, the parameters of the dual-branch encoder, high-level feature enhancement module, low-level feature refinement module, pixel decoder, hint encoder, and mask decoder networks are determined by optimizing a loss function during training using a training set, such that the loss function is minimized; wherein the training set includes an ultrasound image and a corresponding labeled image; the loss function refers to the error between a segmentation prediction map obtained by using an ultrasound image segmentation network model and the corresponding labeled image. As an optional embodiment, in this embodiment, when training the ultrasound image segmentation network model, the loss function used by the pixel decoder is a binary cross-entropy loss, and the functional expression of the binary cross-entropy loss is: , in, is the binary cross entropy loss, is the sample size, is the label of the i-th pixel value, is the preliminary segmentation prediction map of the i-th pixel value; the loss function adopted by the mask decoder is the weighted sum of the focus loss and the dice loss, and the function expression of the weighted sum of the focus loss and the dice loss is: , ,
[0039] in, is the weighted sum of focal loss and dice loss, is the balance factor between focus loss and dice loss, is the focal loss, is the dice loss, is the positive and negative sample balance factor, is the adjusted predicted probability of the i-th pixel value, is the predicted probability of the i-th pixel value. Loss function Used to obtain high-quality preliminary segmentation prediction maps in the first stage of decoding, the loss function Used for decoding in the second stage to obtain the final accurate segmentation prediction map to achieve excellent segmentation performance.
[0040] To validate the effectiveness of this method, the BUSI breast ultrasound image dataset and the UNS neural ultrasound image dataset were used. The STU breast ultrasound image dataset was also used to verify the zero-shot generalization performance of this method. The experiment used an ultrasound image dataset with a one-to-one correspondence between original images and labels. The number of point cues, n, in the example method was set to 10, and the mask cue threshold was set to 0.95. The positive and negative sample balance factor in the training loss function was set to 0.8, the easy and difficult sample balance factor was set to 2, and the balance factor between the focal loss and the dice loss was set to 0.2. The method of this embodiment is mainly compared with three classic segmentation methods, UNet, UNet++, and Attention-UNet; two Transformer-based segmentation methods, TransUNet and TransFuse; two advanced ultrasound image segmentation methods, AAU-Net and CMU-Net; three SAM-based fine-tuning methods, SAM, MSA, and SAMed; and three SAM-based automatic segmentation methods, AutoSAM, UN-SAM, and H-SAM. The objective evaluation indicators used include Dice coefficient, Hausdorff distance, intersection over union (IOU), accuracy, recall, and precision. The results are shown in Tables 1 and 2.
[0041] Table 1: Objective evaluation indicators of ultrasound image segmentation results of BUSI dataset
[0042] Table 2: Objective evaluation indicators of ultrasound image segmentation results of UNS dataset
[0043] Referring to the experimental data shown in Tables 1 and 2, the method of this embodiment demonstrated excellent image segmentation performance on both benchmark datasets. Specifically, on the BUSI breast ultrasound dataset, the method of this embodiment achieved the following results: a Dice coefficient of 81.82%, a Hausdorff distance of 29.71, an IoU of 75.84%, an accuracy of 96.49%, a recall of 84.13%, and a precision of 85.52%. Among them, the four evaluation metrics of Dice coefficient, IoU, precision, and recall all reached optimal values. It is worth noting that in terms of the Dice coefficient and IoU, two of the most representative metrics in the field of image segmentation, the method of this embodiment achieved significant improvements of 3.37% and 3.49% respectively over the next best method, demonstrating a significant advantage. The Hausdorff distance metric ranked second with a result of 29.71, demonstrating the method's excellent performance in ultrasound image edge segmentation. Furthermore, while achieving the best results in terms of precision and recall, the method also achieved third place in terms of precision.
[0044] On the more challenging UNS neural ultrasound dataset, the method in this embodiment achieved evaluation metrics of: Dice coefficient 74.27%, Hausdorff distance 10.16, IoU 71.10%, precision 99.15%, recall 83.36%, and accuracy 86.53%. With the exception of recall, which was slightly lower than the optimal value, the method achieved optimal results in all five metrics (Dice coefficient, Hausdorff distance, IoU, precision, and accuracy), with improvements of 0.21%, 0.04%, 0.15%, 0.02%, and 0.04%, respectively, over the suboptimal model. Compared to the BUSI dataset, ultrasound images in the UNS dataset exhibit more significant noise interference and more complex surrounding tissue features. Existing SAM-based models perform poorly in such complex environments. The improvements to SAM in this embodiment effectively enhance the model's ability to capture local features and identify target tissue, significantly reducing the missegmentation rate of non-target tissue, resulting in segmentation performance that surpasses both traditional methods and existing SAM-based methods.
[0045] To further validate the generalization capability of this embodiment's method in ultrasound image segmentation tasks, this embodiment evaluated the zero-shot generalization capabilities of various segmentation methods on the STU breast ultrasound dataset. This experiment directly used the model weights trained on the BUSI dataset without any fine-tuning or additional training. The results are shown in Table 3.
[0046] Table 3: Objective evaluation indicators of ultrasound image segmentation results of STU dataset
[0047] As shown in Table 3, the method in this embodiment demonstrates excellent zero-shot generalization performance. Specifically, among the six evaluation metrics, four—Dice coefficient, intersection over union, precision, and recall—achieved optimal values, respectively improving by 0.07%, 0.23%, 0.22%, and 0.14% compared to the suboptimal model. In zero-shot generalization experiments, the SAM-based segmentation method significantly outperformed traditional segmentation models. This further demonstrates the SAM model's powerful generalization capabilities and its significant potential in medical image segmentation.
[0048] The comparison results of some samples of the method of this embodiment and the existing method on the BUSI dataset are as follows: Figure 7 As shown. Figure 7In the figure, the ultrasound images in the first to third rows are from benign and malignant breast lesion images from the BUSI dataset. As can be clearly seen from the three different lesion regions shown in this figure, breast lesions in ultrasound images exhibit significant differences in size and shape. Non-SAM-based segmentation methods (e.g., U-Net, TransUNet) exhibited missed detections and false positives in all three lesion regions. The automatic SAM (H-SAM) segmentation method performed well in the second lesion region, but suffered from a small number of missed detections in the first lesion region and a large number of missed detections in the third lesion region. The method in this embodiment achieved perfect segmentation results in the first and second lesion regions. Although the segmentation result for the third lesion region was incomplete, the missed detection rate of the method in this embodiment was significantly reduced compared to other methods. Overall, the proposed method outperformed other methods in breast tumor segmentation. Figure 8 The results of some existing methods on missegmentation of normal images in the BUSI dataset are presented. Due to the influence of ultrasound noise, some irrelevant background areas are mistakenly segmented as lesions. The method in this embodiment introduces a CNN branch to provide classification knowledge of the presence or absence of lesions. Automatic point prompts combine this classification knowledge with the initial segmentation prediction map to guide the model to avoid segmenting non-target areas. The probability of missegmentation of normal images in this embodiment of the method on the BUSI dataset is zero.
[0049] The comparison results of some samples of the method of this embodiment and the existing method on the UNS dataset are as follows: Figure 9 As shown. The ultrasound images in the first to third rows are from the brachial plexus images of the UNS dataset. In the first sample, the traditional segmentation methods U-Net, TransUNet and the H-SAM model improved based on SAM all have the technical defect of incomplete target area segmentation. The method of this embodiment achieves complete target area segmentation; in the second sample, the segmentation performance of each comparison method did not achieve the ideal effect, among which the error rate of the method of this embodiment was the lowest; in the third sample, there is no brachial plexus in the image, and the existing comparison methods are all interfered by irrelevant tissues and produce erroneous segmentation. The method of this embodiment effectively avoids erroneous segmentation by virtue of its unique ultrasound image adaptability mechanism. Comprehensive experimental data proves that the segmentation accuracy of the method of this embodiment in ultrasound nerve images is significantly better than that of existing technical solutions.
[0050] The comparison results of some samples of the method of this embodiment and the existing method on the STU dataset are as follows: Figure 10As shown in the figure, some classic segmentation methods exhibit poor zero-shot generalization capabilities, with their experimental results showing widespread mis-segmentation or incomplete segmentation. Comparison methods based on SAM can avoid widespread mis-segmentation, but also suffer from incomplete segmentation and insufficient edge detail. In contrast, the segmentation predictions of the method in this embodiment are closer to the true labels, demonstrating its advantages in capturing edge details and target region size during segmentation.
[0051] In order to evaluate the effectiveness of the high-level feature enhancement module and the low-level feature refinement module in the ultrasound image segmentation network model of the method of this embodiment, this embodiment performs an ablation experiment on the BUSI dataset for verification. The method of this embodiment without these two modules is used as the baseline network. The results are shown in Table 4.
[0052] Table 4: Objective evaluation indicators of the ablation experiment segmentation results of the high-level feature enhancement module and the low-level feature refinement module
[0053] As shown in the experimental data in Table 4, the model's segmentation performance was effectively improved by sequentially introducing the low-level feature refinement module (+LFR), the high-level feature enhancement module (+HFE), and the high-level feature enhancement module and low-level feature refinement module (+LFR + HFE). Specifically, the results in the second row show that the low-level feature refinement module significantly improves segmentation performance, increasing the Dice coefficient and intersection-over-union ratio by 1.30% and 1.32%, respectively. Furthermore, the low-level feature refinement module effectively improves the segmentation results of edge details, improving the Hausdorff distance by 1.26. The results in the third row show that the addition of the high-level feature enhancement module improves the Dice coefficient, Hausdorff distance, intersection-over-union ratio, accuracy, and precision by 1.36%, 0.12%, 1.37%, 0.34%, and 2.88%, respectively. This fully demonstrates the effectiveness of the high-level feature enhancement module in enhancing semantic features. Finally, compared with the baseline network, after adding the two designed modules, the six evaluation indicators (Dice coefficient, Hausdorff distance, intersection over union, accuracy, recall and precision) are improved by 2.06%, 1.6, 2.34%, 0.46%, 0.81% and 2.38% respectively.
[0054] To verify the contributions of the automatic point hinting and automatic mask hinting strategies to the overall method, detailed ablation experiments were conducted. First, for the automatic point hinting strategy, experimental results were compared with different numbers of point hints (0, 1, 3, 5, 8, 10, 15, and 20). The model achieved optimal segmentation performance when 10 automatic point hints were used. The results are shown in Table 5.
[0055] Table 5: Objective evaluation indicators of the segmentation results of the ablation experiment on the number of automatic point prompts
[0056] As shown in Table 5, compared with the baseline without point hints, the configuration with 10 point hints improves the Dice coefficient, Hausdorff distance, intersection-over-union ratio, accuracy, recall rate, and precision by 2.8%, 2.4%, 2.66%, 0.44%, 2.07%, and 0.67%, respectively, demonstrating that this strategy can effectively suppress the interference of noise and shadows in ultrasound images.
[0057] Furthermore, this experiment combined the optimal number of point prompts (10 points) with the automatic mask prompt (threshold 0.95) for testing, and the results are shown in Table 6.
[0058] Table 6: Objective evaluation indicators of the segmentation results of the ablation experiment with automatic point hints and automatic mask hints
[0059] As shown in Table 6, the dual-prompt strategy is effectively combined, achieving significant improvements of 3.71%, 2.71%, 3.91%, 0.48%, 1.86% and 3.28% in six indicators (Dice coefficient, Hausdorff distance, intersection over union ratio, accuracy, recall and precision) compared to the unprompted baseline.
[0060] In summary, the method of this embodiment utilizes a dual-branch encoder and a two-stage decoding architecture. The dual-branch encoder combines the Transformer branch of the SAM for extracting global features with the CNN branch for extracting local features. During the first-stage decoding process, a high-level feature enhancement module strengthens the multi-scale semantic feature representation, while a low-level feature refinement module optimizes shallow feature details, jointly improving the quality of the preliminary segmentation prediction map. Regarding hint generation, an innovative hinting strategy combining automatic point hints and automatic mask hints is designed. Automatic point hints are generated by combining the classification knowledge of the CNN branch with the spatial information of the preliminary segmentation prediction map to effectively suppress the interference of noise and shadows in the ultrasound image. Automatic mask hints are generated by thresholding the preliminary segmentation prediction map, providing detailed contextual guidance to the model. Extensive experiments demonstrate that the method of this embodiment outperforms other segmentation models. Without the need for manual hinting, the SAM-based automatic hinting ultrasound image segmentation method achieves outstanding performance in the field of ultrasound image segmentation and demonstrates strong generalization capability. Compared with existing technologies, the method of this embodiment has the following significant advantages: 1. The SAM-based automatic hinting ultrasound image segmentation network model of this embodiment features a dual-branch encoder and a two-stage decoding process. The dual-branch encoder effectively extracts local detail features and global dependencies, while the two-stage decoding process automatically generates hints embedded in the guidance model to achieve accurate segmentation, effectively applying SAM to the field of ultrasound image segmentation. 2. The automatic hinting design of the SAM-based ultrasound image segmentation network model of this embodiment solves the problem of SAM's reliance on manual hinting. The positive and negative point hints based on CNN classification knowledge effectively reduce missegmentation caused by ultrasound noise and shadows. The mask hints also provide more detailed contextual information. The combination of these two types of automatic hints improves the information capture efficiency of the input image and enhances the ultrasound image segmentation performance of the ultrasound image segmentation network model. 3. The high-level feature enhancement module and low-level feature refinement module of this embodiment effectively improve the quality of the preliminary segmentation prediction image and the accuracy of automatic hinting. The high-level feature enhancement module aims to capture multi-scale semantic features, while the low-level feature refinement module plays a key role at jump connections by refining the shallow features of the CNN branches.
[0061] In addition, this embodiment also provides a SAM-based automatic prompt ultrasound image segmentation system, comprising an interconnected microprocessor and memory, wherein the microprocessor is programmed or configured to execute the SAM-based automatic prompt ultrasound image segmentation method. In addition, this embodiment also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program or instructions, wherein the computer program or instructions are programmed or configured to execute the SAM-based automatic prompt ultrasound image segmentation method via a processor. In addition, this embodiment also provides a computer program product, comprising a computer program or instructions, wherein the computer program or instructions are programmed or configured to execute the SAM-based automatic prompt ultrasound image segmentation method via a processor.
[0062] Those skilled in the art should understand that the technical solution provided by the present invention may be in the form of a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the functions described in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0063] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A SAM-based automatic prompt ultrasound image segmentation method, characterized in that: The invention comprises the following steps: using an ultrasonic image segmentation network model to perform ultrasonic image segmentation on an input ultrasonic image to obtain a segmentation prediction map, wherein the ultrasonic image segmentation network model comprises a CNN encoder, a classification head module, a low-level feature refinement module, a SAM Transformer encoder, a cross-branch attention module, a high-level feature enhancement module, a pixel decoder, a prompt encoder and a mask decoder, and performing ultrasonic image segmentation on the input ultrasonic image to obtain a segmentation prediction map comprises: extracting local detail features from the input ultrasonic image through the CNN encoder and performing target existence classification through the classification head module, optimizing and enhancing the local detail features through the low-level feature refinement module and then outputting them to the pixel decoder; extracting global features from the input ultrasonic image through the SAM Transformer encoder, supplementing the global features through the cross-branch attention module, and After filling in the local detail features, the image embedding is obtained, and then the advanced feature enhancement module is used to realize semantic feature enhancement through multi-scale feature extraction and then output to the pixel decoder; the first stage of decoding is performed by the pixel decoder to generate a preliminary segmentation prediction map, and the preliminary segmentation prediction map is thresholded and then a mask prompt is generated. If the result of the target existence classification indicates that there is a target, the top n points with the highest confidence in the preliminary segmentation prediction map are selected as positive point prompts, otherwise the top n points with the lowest confidence in the preliminary segmentation prediction map are selected as negative point prompts, and the positive point prompts or negative point prompts are encoded by the prompt encoder to generate a point embedding vector, and the mask prompts are encoded by the prompt encoder to generate a mask embedding vector. The point embedding vector, the mask embedding vector and the image obtained after supplementing the local detail features are embedded into the mask decoder to perform the second stage of decoding to obtain the final segmentation prediction map.
2. The automatic prompt ultrasound image segmentation method based on SAM according to claim 1, characterized in that: The CNN encoder includes a total of five convolution modules, namely, convolution modules 1 to 5, which are connected in cascade. The local detail features output by the first three convolution modules of convolution modules 1 to 3 are not only used as the input of the next-level convolution module in the CNN encoder, but also as the input of a corresponding low-level feature refinement module. The local detail features output by convolution module 5 are used as the input of the classification head module to realize target existence classification. The Transformer encoder of the SAM includes a multi-level window Transformer layer. The input ultrasound image is first divided into blocks by the image block embedding layer and then combined with the position encoding as the input of the multi-level window Transformer layer. The image embedding obtained by supplementing the local detail features by the cross-branch attention module includes: the global features extracted by the multi-level window Transformer layer are normalized and then input into the cross-branch attention module. The cross-branch attention module uses the local detail features output by the CNN encoder as keys and values and the global features as queries to perform cross-attention operations. The method further comprises the following steps: first, extracting the attention features through a multi-head attention mechanism, adding the extracted attention features to the global features to obtain fused features, extracting the local detail features output by the CNN encoder through multiple single convolution modules, normalizing the fused features through layers, performing multi-layer perceptron, adding the fused features to the original fused features, and then adding the features extracted by multiple single convolution modules to obtain image embedding after supplementing the local detail features; the Transformer encoder of the SAM adopts a low-rank adaptive fine-tuning method to fine-tune the Transformer encoder during training, including: introducing a trainable low-rank bypass into the query and value projection layers in the multi-level window Transformer layer and the multi-head attention mechanism, wherein the low-rank bypass is composed of two low-rank matrices and an activation function to perform low-rank decomposition on the parameter update, and the low-rank matrix is a linear layer; during the training process, only the parameters of the low-rank bypass are optimized, and the weights of the original query and value projection layers are frozen; in the inference stage, the low-rank matrix is added to the original weights of the query and value projection layers to form new weights.
3. The automatic prompt ultrasound image segmentation method based on SAM according to claim 1, characterized in that: The method of using a low-level feature refinement module to optimize and enhance local detail features and then outputting them to a pixel decoder includes: using the local detail features as input feature maps, extracting features from the input feature maps using a channel attention mechanism, and subjecting the input feature maps to parallel average pooling operations and maximum pooling operations with a kernel size of 2×2, and average pooling operations and maximum pooling operations with a kernel size of 8×8, respectively; then upsampling the results of the four pooling operations to the same resolution and performing channel dimension splicing; then generating a refined feature map through a 1×1 convolution module; residually connecting the refined feature map and the original input feature map; then performing channel dimension splicing with the features extracted by the channel attention mechanism; and then generating a final output feature map through a 1×1 convolution module and outputting it to the pixel decoder.
4. The automatic prompt ultrasound image segmentation method based on SAM according to claim 1, characterized in that: The method uses the advanced feature enhancement module to achieve semantic feature enhancement through multi-scale feature extraction and then outputs it to the pixel decoder, including: embedding the image as the input feature map, extracting features from the input feature map using the channel attention mechanism, and passing the input feature map in parallel through four dilated convolution branches for capturing features of different scales in the spatial dimension; among the four dilated convolution branches: the first dilated convolution branch consists of three 3×3 dilated convolutions with dilation rates of 1, 3, and 5, respectively, and one 1×1 convolution; the second dilated convolution branch consists of two 3×3 dilated convolutions with dilation rates of 1 and 3, respectively, and one 1×1 convolution; the third dilated convolution branch consists of a 3×3 dilated convolution with dilation rates of 3, respectively, and a 1×1 convolution; the fourth dilated convolution branch consists of a 3×3 dilated convolution with dilation rate of 1; the features extracted by the channel attention mechanism and the features extracted by the four dilated convolution branches are spliced in the channel dimension to generate the final output feature map and output it to the pixel decoder.
5. The automatic prompt ultrasound image segmentation method based on SAM according to claim 1, characterized in that: The pixel decoder includes three layers of pixel decoding modules, each layer of pixel decoding modules consists of a transposed convolution, a channel dimension connection module, and two convolution modules; the convolution module is composed of a 3×3 convolution layer, a batch normalization layer, and a GELU activation function connected in sequence, the output features of the high-level feature enhancement module are used as the input feature map of the pixel decoder and input into the transposed convolution of the first layer to obtain a twice up-sampled feature map, the shallow features output by the first low-level feature refinement module are connected with the twice up-sampled feature map through the channel dimension connection module and then passed through the convolution module to obtain the output feature map of the first layer; the output feature map of the first layer is passed through the convolution module. The second layer is subjected to transposed convolution to obtain a four-fold upsampled feature map. The shallow features output by the second low-level feature refinement module are connected with the four-fold upsampled feature map through the channel dimension connection module and then passed through the convolution module to obtain the output feature map of the second layer. The output feature map of the second layer is subjected to transposed convolution of the third layer to obtain an eight-fold upsampled feature map. The shallow features output by the third low-level feature refinement module are connected with the eight-fold upsampled feature map through the channel dimension connection module and then passed through the convolution module to obtain the output feature map of the third layer. The output feature map of the third layer is then subjected to a 1×1 convolution layer and a sigmoid activation function to obtain a preliminary segmentation prediction map.
6. The automatic prompt ultrasound image segmentation method based on SAM according to claim 1, characterized in that: The mask decoder includes two mask decoding units, a multi-level transposed convolution module, an anti-cross attention module, a multi-layer perceptron and a dot product module. The mask decoding unit includes a self-attention module, an anti-cross attention module, a perceptron and a cross attention module connected in sequence. The point embedding vector is used as the input of the self-attention module in the first mask decoding unit. The image embedding obtained after supplementing the local detail features and the mask embedding vector are spliced as the input of the anti-cross attention module and the cross attention module in the first mask decoding unit. The output of the perceptron in the first mask decoding unit is used as the input of the self-attention module in the second mask decoding unit, and the output of the cross attention module in the first mask decoding unit is used as the input of the The anti-cross attention module and the input of the cross attention module in the two mask decoding units, the output of the perceptron in the second mask decoding unit is used as the first input of the anti-cross attention module outside the two mask decoding units, the output of the cross attention module in the second mask decoding unit is used as the second input of the anti-cross attention module outside the two mask decoding units and the input of the multi-stage transposed convolution module. The anti-cross attention module outside the two mask decoding units performs anti-cross attention on the two inputs and extracts features through the multi-layer perceptron as the first input of the dot product module. The output of the multi-stage transposed convolution module is used as the second input of the dot product module. The dot product module performs dot product operation on the two inputs to obtain the final segmentation prediction map.
7. The automatic prompt ultrasound image segmentation method based on SAM according to claim 2, characterized in that: When training the ultrasound image segmentation network model, the loss function used by the pixel decoder is binary cross entropy loss, and the functional expression of the binary cross entropy loss is: , in, is the binary cross entropy loss, is the sample size, is the label of the i-th pixel value, is the preliminary segmentation prediction map of the i-th pixel value; the loss function adopted by the mask decoder is the weighted sum of the focus loss and the dice loss, and the function expression of the weighted sum of the focus loss and the dice loss is: , , in, is the weighted sum of focal loss and dice loss, is the balance factor between focus loss and dice loss, is the focal loss, is the dice loss, is the positive and negative sample balance factor, is the adjusted predicted probability of the i-th pixel value, is the predicted probability of the i-th pixel value.
8. A SAM-based automatic prompt ultrasound image segmentation system, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the SAM-based automatic prompt ultrasound image segmentation method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the SAM-based automatic prompt ultrasound image segmentation method according to any one of claims 1 to 7 through a processor.
10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the SAM-based automatic prompt ultrasound image segmentation method according to any one of claims 1 to 7 through a processor.
Citation Information
Patent Citations
Self-adaptive medical image segmentation method and device based on prompt learning, and medium
CN118521595A
Endoscopic polyp segmentation method based on cross-branch feature enhancement and uncertainty guidance
CN119359741A
Ultrasonic image segmentation method and system based on multistage feature extraction
CN119693383A
Medical image segmentation method based on SAM and prompter
CN120014262A
Medical image segmentation method based on multi-scale feature fusion
US20250095828A1
Cited By
Camouflage target detection method
CN121213900A
Large model image segmentation method and system based on multi-scale fusion, and medium
CN121527427A
High-dimensional coupling segmentation method for target fine semantics based on SAM
CN121861293A
New intermediate phase segmentation method for microstructure image of zinc alloy material
CN122265299A