Anatomical-level ultrasonic medical image analysis method based on SAM guidance

By introducing SAM-based automatic prompt generator and cross-view prediction technology in ultrasonic medical image analysis, combined with comparative learning to train the backbone model, the problem of insufficient analysis accuracy and robustness under data scarcity is solved, and more efficient image analysis effect is achieved.

CN120147273APending Publication Date: 2025-06-13HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510233630.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing ultrasonic medical image analysis technology has poor analysis accuracy and robustness in the case of scarce data, and is affected by artifacts, blurred boundaries and irrelevant backgrounds.

Method used

Using anatomical ultrasound medical image analysis method based on SAM guidance, automatic prompt generator is integrated into the SAM model, automatic prompts are generated and ultrasound medical image segmentation is performed; cross-view prediction is performed based on segmentation maps and image features of different resolutions; and through comparative learning, training backbone models, the similarity between features of the same anatomical structure is maximized.

Benefits of technology

Improves analysis accuracy and robustness, can generate relatively robust results when data is scarce, and adapts to a variety of downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147273A_ABST
    Figure CN120147273A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an anatomical-level ultrasonic medical image analysis method based on SAM guidance, and belongs to the technical field of image processing, and the method specifically comprises the steps: integrating an automatic prompt generator into an SAM model, carrying out the fine tuning of the SAM model through an existing ultrasonic medical image-mask pair, generating an automatic prompt, and carrying out the segmentation of an ultrasonic medical image according to the automatic prompt; performing cross-view prediction on an anatomical structure in the ultrasonic medical image; comparing a prediction result with a true value to realize anatomical-level region comparison of different resolutions and different views, taking the same anatomical structure among different views as a positive sample, taking different anatomical structures and views of different images as negative samples, maximizing similarity among features of the same anatomical structure, and training a trunk model through comparison learning; and retaining a residual network of the online updated network in the trained trunk model, and carrying out an image analysis task according to the residual network. Through the scheme disclosed by the invention, the analysis accuracy and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of image processing technology, and in particular, to a method for analyzing anatomical ultrasound medical images guided by SAM. Background Art

[0002] Currently, ultrasound medical image analysis technology is crucial in computer-aided diagnosis and has become a commonly used technology in clinical imaging. For example, it can help detect the presence and size of tumors, observe the structure of soft tissues, etc., for clinical auxiliary screening, diagnosis, and evaluation, thereby reducing the workload of clinicians and at the same time reducing the subjectivity of diagnosis, making the diagnosis more objective and accurate.

[0003] When analyzing ultrasound images based on a deep neural network architecture, training models often face the challenge of scarce data. Existing solutions mainly involve transfer learning, which adapts to downstream tasks through two stages of pre-training and fine-tuning. The general visual representations learned through pre-training are combined with fine-tuning to adapt to specific tasks. Moreover, due to the excellent zero-shot performance of SAM, it has received extensive attention in medical image segmentation tasks, and many models for analyzing medical images based on SAM have been proposed. Although these methods have improved the dilemma of scarce data, there are still problems such as being affected by artifacts, blurred boundaries, and irrelevant backgrounds existing in medical images during analysis, resulting in poor robustness of the obtained image representations.

[0004] It can be seen that there is an urgent need for a method for analyzing anatomical ultrasound medical images guided by SAM with strong analysis accuracy and robustness. Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide a method for analyzing anatomical ultrasound medical images guided by SAM, which at least partially solves the problem of poor analysis accuracy and robustness in the prior art.

[0006] Embodiments of the present disclosure provide a method for analyzing anatomical ultrasound medical images guided by SAM, including:

[0007] Step 1: Integrate an auto-prompter into the SAM model, fine-tune the SAM model using existing ultrasound medical image-mask pairs to generate auto-prompts, and segment the ultrasound medical images accordingly;

[0008] Step 2: Rely on the masks and image features of different resolutions in the segmentation map to perform cross-view prediction on the anatomical structures in the ultrasound medical images;

[0009] Step 3: By comparing the prediction results with the ground truth, perform anatomical region comparison for different resolutions and different views. Consider the same anatomical structures between different views as positive samples, and different anatomical structures and views of different images as negative samples. Maximize the similarity between the features of the same anatomical structures, and train the backbone model through contrastive learning. The backbone model includes a residual network and a mapping module;

[0010] Step 4: Retain the residual network of the online update network in the trained backbone model and perform image analysis tasks based on it. The image analysis tasks include image classification and image segmentation.

[0011] According to a specific implementation manner of the embodiments of the present disclosure, the specific steps of Step 1 include:

[0012] Step 1.1: Integrate the auto prompt generator into the SAM model, extract the global features and local features of the image, and then use multi-head cross-attention to fuse the global features and local features to obtain mixed features. The expression of the auto prompt generator is

[0013] F vc = LN(MHCA(F v , F c , F c ) + F v )

[0014] F′ v = MLP(LN(MHCA(F vc , T, T) + F vc ))

[0015] T′ = MLP(LN(MHCA(T, F vc , F vc )) + T)

[0016] where MHCA(·) represents multi-head cross-attention, LN represents normalization, MLP represents multi-layer perceptron, F v is the global feature, F c is the local feature, F vc is the mixed feature of the global and local parts of the image, T represents the initial token, F′ v and T' are obtained by performing cross-attention calculation on F vc and T. F′ v is the feature in the image that corresponds to the initial token, and T′ is the feature in the initial token that corresponds to the image region;

[0017] Step 1.2: Use F′ v as the new F v, taking T' as the new T, then performing the two multi-head cross-attention calculations in the previous step according to the new hybrid features and new tokens to obtain the new F' v and T', continuously repeating the operations in this step, and obtaining the finally information-fused token as the auto-suggestion;

[0018] Step 1.3, inputting the auto-suggestion and the initial global feature F which is used as the image embedding v into the masked decoder to implement the segmentation of the ultrasound medical image and generate the mask corresponding to the segmentation map.

[0019] According to a specific implementation manner of the embodiment of the present disclosure, the step 2 specifically includes:

[0020] Step 2.1, for the ultrasound medical image x, using random preprocessing to generate two enhanced ultrasound images with different perspectives, respectively denoted as view v 1 and view v 2 and respectively inputting them into the residual network to extract image features at different resolutions, screening out the mask groups corresponding to the image feature content from the mask of the segmentation map, respectively denoted as m and m', and calculating the potential vector representation of the mask corresponding to view v 1 according to a preset formula, where the preset formula is

[0021]

[0022] where, f s [i,j] represents the feature vector at (i,j) of the feature map with resolution s, m s [i,j] represents the mask value corresponding to (i,j) after the mask group m is adjusted to resolution s, and h s represents the multi-granularity potential vector representation of the view v 1 of the ultrasound medical image;

[0023] Step 2.2, substituting the feature vector of view v 2 and the mask group m' into the preset formula to calculate the multi-granularity potential vector representation h' 2 of view v s ;

[0024] Step 2.3, for the multi-granularity potential vectors h 1 and h' 2 of view v s and view v s , respectively using two mapping heads to map them to z θ,s and z ξ,s , making the overall network architecture for processing v 1 become an online update network, and processing v 2The overall network architecture becomes the target network, where θ is the parameter of the online update network, ξ is the parameter of the target network. By using the prediction head of the online update network, the result of predicting the content of one view from another view is obtained, and then the error is calculated with the true result output by the target network.

[0025] According to a specific implementation manner of the embodiments of the present disclosure, step 3 specifically includes:

[0026] Step 3.1, define the similarity formula between the structures of view v 1 and view v 2 as

[0027]

[0028] where τ is the temperature hyperparameter, q θ,s (·) represents the prediction mapping of the online update network, represents the distance between the result of predicting view v 1 from view v 2 and the actual view v 2 ;

[0029] Step 3.2, after obtaining n negative samples by random sampling, design a multi-granularity contrast loss function according to the similarity formula

[0030]

[0031] where n represents the number of negative samples, is the distance between v 1 and the negative sample i;

[0032] Step 3.3, based on the multi-granularity contrast loss function, use view v 2 as the input of the online update network for prediction, and integrate the contrast losses at all granularities to obtain the overall loss function of the final backbone model

[0033]

[0034] where K represents the number of samples, s represents different granularities, and S represents the total number of granularities;

[0035] Step 3.4, based on the overall loss function, update the parameter θ in the online update network of the backbone model by reducing the loss, and accordingly update the parameter ξ in the target network of the backbone model by using the exponential moving average method.

[0036] According to a specific implementation manner of the embodiments of the present disclosure, when the analysis task is image classification, step 4 specifically includes:

[0037] Use the trained residual network f θ As the encoder part of the classification model, extract the multi-level image features of the ultrasonic medical image. Use the feature pyramid network to fuse the image features of different levels, and finally use the fully connected layer to flatten the features to complete the classification task.

[0038] According to a specific implementation manner of the embodiment of the present disclosure, when the analysis task is image segmentation, step 4 specifically includes:

[0039] Replace the CNN network in the Mask-RCNN model with the trained residual network f θ Then use the SGD optimizer to optimize, fine-tune the model and segment the target area in the ultrasonic medical image to obtain the corresponding segmentation mask.

[0040] The anatomical-level ultrasonic medical image analysis solution based on SAM guidance in the embodiment of the present disclosure includes: Step 1, integrate the automatic prompt generator into the SAM model, use the existing ultrasonic medical image-mask pair to fine-tune the SAM model and generate automatic prompts, and segment the ultrasonic medical image accordingly; Step 2, rely on the masks and image features of different resolutions in the segmentation map to perform cross-view prediction on the anatomical structures in the ultrasonic medical image; Step 3, realize the anatomical-level area comparison of different resolutions and different views by comparing the prediction results with the ground truth. Take the same anatomical structures between different views as positive samples, and different anatomical structures and views of different images as negative samples, maximize the similarity between the features of the same anatomical structures, and train the backbone model through contrastive learning, where the backbone model includes a residual network and a mapping module; Step 4, retain the residual network of the online update network in the trained backbone model and perform image analysis tasks accordingly, where the image analysis tasks include image classification and image segmentation.

[0041] The beneficial effects of the embodiment of the present disclosure are: Through the solution of the present disclosure, use SAM containing an auto-prompter for effective object-level segmentation, and adopt cross-view contrastive learning on the pre-trained model, so as to produce relatively robust results and adapt to various downstream tasks in the case of scarce data, improving the analysis accuracy and robustness. Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1Schematic flow chart of an anatomical ultrasound medical image analysis method provided by an embodiment of the present disclosure. Detailed implementation manners

[0044] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0045] The following uses specific specific examples to illustrate the implementation manners of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0046] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.

[0047] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure schematically. The diagrams only show the components related to the present disclosure, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.

[0048] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0049] The embodiments of the present disclosure provide an anatomical ultrasound medical image analysis method based on SAM guidance, and the method can be applied to the medical image analysis process in a medical scenario.

[0050] SeeFigure 1 , which is a schematic flow chart of a SAM-guided anatomical ultrasound medical image analysis method provided by an embodiment of the present disclosure. As Figure 1 shown, the method mainly includes the following steps:

[0051] Step 1: Integrate an auto-prompter into the SAM model, fine-tune the SAM model using existing ultrasound medical image-mask pairs to generate auto-prompts, and segment the ultrasound medical image based on this;

[0052] Further, step 1 specifically includes:

[0053] Step 1.1: Integrate an auto-prompter into the SAM model, extract the global features and local features of the image, and then use multi-head cross-attention to fuse the global features and local features to obtain mixed features. Among them, the expression of the auto-prompter is

[0054] F vc = LN(MHCA(F v , F c , F c ) + F v )

[0055] F' v = MLP(LN(MHCA(F vc , T, T) + F vc ))

[0056] T' = MLP(LN(MHCA(t, F vc , F vc ) + T))

[0057] Among them, MHCA(·) represents multi-head cross-attention, LN represents regularization, MLP represents a multi-layer perceptron, F v is the global feature, F c is the local feature, F vc is the mixed feature of the global and local parts of the image, T represents the initial token, F' v and T' are obtained by performing cross-attention calculations on F vc and T. F' v is the feature in the image that corresponds to the initial token, and T' is the feature in the initial token that corresponds to the image region;

[0058] Step 1.2: Use F' v as the new F v , T' as the new T, and then perform the two multi-head cross-attention calculations in the previous step according to the new mixed features and new tokens to obtain the new F' vWith T and T', repeat the operation of this step continuously to obtain the token of the final information fusion as the auto - suggestion;

[0059] Step 1.3, input the auto - suggestion and the initial global feature F as the image embedding v into the mask decoder to achieve the segmentation of the ultrasonic medical image and generate the mask corresponding to the segmentation map.

[0060] Specifically, the auto - suggestion generator can be integrated into the SAM model, and the existing image - mask pairs are used to fine - tune the model to generate suggestions without manual intervention and complete the segmentation of medical images.

[0061] In the embodiment of this application, the hint - processing part of the SAM model itself is modified, and a module that can automatically generate hints is designed, which can automatically generate masks from ultrasonic images to achieve the segmentation of ultrasonic medical images.

[0062] This module retains the ViT encoder in the SAM model to obtain an embedding of the image in the latent space as the global feature F of the image v , and at the same time adds a convolutional neural network encoder to extract image features as the local feature F of the image c , and then uses cross - aware attention to fuse the two features to obtain the feature F vc :

[0063] F vc = LN(MHCA(F v , F c , F c ) + F v )

[0064] Among them, MHCA(·) represents multi - head cross - attention, LN represents regularization, and in this attention calculation, F v is the query, and F c is the key and value.

[0065] In the embodiment of this application, the initial learnable vector token T and the fusion feature F of the ultrasonic medical image vc are respectively used as the queries in the attention mechanism to perform two different cross - attention calculations:

[0066] F' v = MLP(LN(MHCA(F vc , T, T) + F vc ))

[0067] T' = MLP(LN(MHCA(T, F vc , F vc )) + T)

[0068] Among them, MLP is a multi-layer perceptron that uses the cross-attention mechanism twice to fuse the information of tokens and features F vc with each other, thereby realizing the interaction of information. Then, the calculated F′ v is used as the new F v , and T′ is used as the new T. Repeat the above operations, where F c remains unchanged.

[0069] After several rounds of iterative calculations, the final token T′ is obtained as the prompt embedding, and the initial global feature F v as the image embedding are input into the masked decoder together to achieve the segmentation of the image, automatically generate a segmentation map, and obtain the masks of each part.

[0070] Step 2: Relying on the masks in the segmentation map and the image features at different resolutions, perform cross-view prediction on the anatomical structures in the ultrasound medical image;

[0071] On the basis of the above embodiments, the specific steps of step 2 include:

[0072] Step 2.1: For the ultrasound medical image x, use random preprocessing to generate two enhanced ultrasound images with different perspectives, denoted as view v 1 and view v 2 respectively, and input them into the residual network to extract image features at different resolutions. Select the corresponding mask groups m and m′ from the masks in the segmentation map, and calculate the latent vector representation of the mask corresponding to view v 1 according to the preset formula, where the preset formula is

[0073]

[0074] where, f s [i, j] represents the feature vector at (i, j) of the feature map with resolution s, m s [i, j] represents the mask value at (i, j) after adjusting the mask group m to resolution s, and h s represents the multi-granularity latent vector representation corresponding to view v 1 of the ultrasound medical image;

[0075] Step 2.2: Substitute the feature vector of view v 2 and the mask group m′ into the preset formula to calculate the multi-granularity latent vector representation h′ 2 of view v s ;

[0076] Step 2.3: For view v 1 and view v 2Multi-granularity latent vectors h s and h′ s are respectively mapped to z θ,s and z ξ,s by using two mapping heads, and the overall network architecture for processing v 1 is called the online update network, and the overall network architecture for processing v 2 is called the target network, where θ is the parameter of the online update network and ξ is the parameter of the target network. By using the prediction head of the online update network, the result of predicting the content of one view from another view is obtained, and then the error is calculated with the true result output by the target network.

[0077] Specifically, when implemented, after obtaining the segmentation map and each part mask of the ultrasound image, the model performs contrast detection on the ultrasound images of different perspectives to achieve cross-perspective anatomical-level feature recognition.

[0078] For a given ultrasound medical image x, two corresponding ultrasound images with different perspectives are generated by using random preprocessing, and are respectively denoted as and are respectively input into two different residual networks for subsequent feature extraction.

[0079] Process the mask of each part in the ultrasound image obtained in the previous step. Through operations such as cropping, flipping, and adjustment, it is respectively aligned with one or more local parts of the objects included in v 1 and v 2 to screen out all masks corresponding to the content of the image features of the two perspectives, and are respectively denoted as mask groups m and m′. In the embodiments of the present application, it is necessary to ensure the consistency of the resolution between the masks.

[0080] In each iteration, K masks need to be sampled from m and m′ respectively, and image features are extracted from different layers of the two residual networks at the same time, thereby obtaining two groups of multi-granularity feature map groups containing various resolutions; v 1 and v 2 For the feature maps with the current resolution s, they are respectively f s and f′ s . In order to make the resolution of the feature map and the mask consistent, downsampling is used to adjust the resolution of the mask to s as well. The corresponding mask groups after adjustment are denoted as m s and m′ s ; further, mask pooling is applied to the feature map with the resolution s to calculate the multi-granularity latent vector representation:

[0081]

[0082] where f s [i,j] represents the feature vector of the feature map with the resolution s at (i,j), and m s[i,j] represents the corresponding mask value at (i,j), h s represents the multi-granularity latent vector representation corresponding to the ultrasound medical image view v 1 Accordingly, according to the above formula, m s is replaced by m′ s , f s is replaced by f′ s After that, the multi-granularity latent vector representation h′ 2 of v s can be obtained.

[0083] Using two projection heads, respectively map the multi-granularity latent vector representations h s and h′ s :

[0084]

[0085] where θ is a trainable parameter, s is the corresponding resolution, and the mapped z θ,s will be used as the input of the prediction head for anatomical-level structure prediction to obtain the result q θ,s (z θ,s ), enabling the anatomical-level features of one view to predict the same anatomical structure in another view. We call the overall network architecture for processing v 1 the online update network, with the corresponding network parameter θ, and the overall network architecture for processing v 2 the target network, with the corresponding network parameter ξ.

[0086] Step 3, realize the comparison of anatomical-level regions with different resolutions and different views by comparing the prediction results and the ground truth. Take the same anatomical structures between different views as positive samples, and different anatomical structures and views of different images as negative samples to maximize the similarity between the features of the same anatomical structure. Train the backbone model through contrastive learning, where the backbone model includes a residual network and a mapping module;

[0087] Based on the above embodiments, the specific steps of Step 3 include:

[0088] Step 3.1, define the similarity formula between the structures of view v 1 and view v 2 at resolution s as

[0089]

[0090] where τ is the temperature hyperparameter, q θ,s (·) represents the prediction mapping of the online update network, represents predicting the result of view v 1 from view v 2 and the actual view v2 The distance between;

[0091] Step 3.2, after obtaining n negative samples through random sampling, design a multi-granularity contrast loss function according to the similarity formula

[0092]

[0093] where n represents the number of negative samples, is the distance between v 1 and negative sample i;

[0094] Step 3.3, based on the multi-granularity contrast loss function, use view v 2 as the input of the online update network for prediction, and integrate the contrast losses at all granularities to obtain the overall loss function of the final backbone model

[0095]

[0096] where K represents the number of samples, s represents different granularities, and S represents the total number of granularities;

[0097] Step 3.4, based on the overall loss function, update the parameter θ in the online update network of the backbone model by reducing the loss, and accordingly update the parameter ξ in the target network of the backbone model using the exponential moving average method.

[0098] In specific implementation, add anatomical-level region contrast at different granularities and different views extracted between multiple layers of the backbone model. The same anatomical structure between different views is used as the positive sample, and different structure masks or different image features are used as the negative sample, aiming to maximize the similarity between the features of the same anatomical structure, train the model well, and achieve contrastive learning.

[0099] To achieve contrastive learning at the anatomical structure level, consider comparing the same structural content in ultrasonic medical images from different perspectives, and enhance the consistency of the same anatomical structure by maximizing the similarity degree of the feature vectors within the same mask across views; correspondingly, the similarity between the structures of the feature maps corresponding to v 1 and v 2 at resolution s is expressed as follows:

[0100]

[0101] where τ is the temperature hyperparameter, represents the gap between the structural features predicted according to v 1 for v 2 and the actual structural features of v 2 ;

[0102] After obtaining n negative samples through random sampling, according to the calculated similarity above, the following multi-granularity contrast loss function is designed, and the error is reduced by continuously updating the parameters to improve the model's ability to identify anatomical structures using negative samples:

[0103]

[0104] Among them The images of view 1 and view 2 in are obtained by processing the current ultrasound medical image x, that is, the above-mentioned v 1 and v 2 Therefore, the same masked part in v 2 as v 1 is used as the positive sample; The view i in is a negative sample for contrastive learning. In the embodiments of the present application, the selection of negative samples includes the features of different masks in the same image, or the features of views of images different from x in a batch. For example, if an image contains masks of multiple objects, the features of a certain mask can be used as the negative sample of another mask. Such a selection can increase the diversity of negative samples and enable the model to learn better.

[0105] Then, the exchange of v 1 and v 2 is performed, that is, the original v 1 is used as the input of the online update network for prediction. After the views are exchanged, v 2 is used as the input of the online update network, and the final loss function of the entire model is obtained as:

[0106]

[0107] where K represents the number of samples, s represents different granularities, and S represents the total number of granularities.

[0108] It can be known from the above loss function that the model updates the parameters of the two sub-networks. For the parameter θ in the online update network, it is updated according to the loss function L(θ; ξ); while for the parameter ξ in the target network, it depends on θ of the online update network and uses exponential moving average to update the parameters regularly.

[0109] Step 4, retain the residual network of the online update network in the trained backbone model and perform image analysis tasks accordingly, where the image analysis tasks include image classification and image segmentation.

[0110] Optionally, when the analysis task is image classification, step 4 specifically includes:

[0111] The trained residual network f θAs the encoder part of the classification model, it extracts multi-level image features of ultrasonic medical images, uses a feature pyramid network to fuse image features at different levels, and finally flattens the features with a fully connected layer to complete the classification task.

[0112] Optionally, when the analysis task is image segmentation, step 4 specifically includes:

[0113] Replace the CNN network in the Mask-RCNN model with the trained residual network f θ Then use the SGD optimizer to optimize, fine-tune the model, and segment the target area in the ultrasonic medical image to obtain the corresponding segmentation mask.

[0114] Specifically in implementation, only keep the residual network f in the online update network for the trained network θ And apply it to downstream tasks for operations such as image classification and image segmentation.

[0115] When processing downstream tasks, for the trained model, we only keep the residual network f in the online update network θ Part, and apply this encoder part according to different tasks. The following mainly mentions the applications of two downstream tasks:

[0116] The first one: For classifying lung ultrasound images and breast ultrasound images to diagnose whether the samples have diseases, we use the trained residual network f θ As the encoder part of the model to extract image features, then use a feature pyramid network to fuse features at different levels, and finally flatten the features with a fully connected layer to complete the classification task. After fine-tuning, it realizes the discrimination of healthy and diseased ultrasound images.

[0117] The second one: For segmentation tasks such as breast tumors and thyroid nodules, we consider modifying the Mask-RCNN model applied to the image segmentation task, and replace the CNN network for feature extraction in it with the trained residual network f θ Then use the SGD optimizer to optimize, fine-tune the model, and use it to segment tumors or nodules in medical images to obtain the corresponding segmentation mask.

[0118] Attributed to SAM-guided pre-training, the residual network f θIt enhances the model's ability to learn anatomical structure-level features from ultrasound images, enabling superior and stable performance in a variety of downstream tasks. At the same time, the model can be applied to various downstream tasks and used as a feature extraction module for module replacement, combination, etc. For example, for ultrasound medical images of the lungs, the contrastive learning adopted by the model can distinguish between lung images with pneumonia or normal classes; for example, in dealing with the breast tumor segmentation task, the improved SAM module of the model can achieve detailed discrimination for pixel-level prediction tasks.

[0119] The SAM-guided anatomical ultrasound medical image analysis method provided in this embodiment effectively performs object-level segmentation by using SAM containing an auto-prompter, and adopts cross-view contrastive learning on a pre-trained model, so as to produce relatively robust results and adapt to various downstream tasks in the case of scarce data, improving the analysis accuracy and robustness.

[0120] It should be understood that each part of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof.

[0121] As described above, the above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present disclosure should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for analyzing anatomical-level ultrasonic medical images based on SAM guidance, characterized in that: include: Step 1, integrating the automatic hint generator into the SAM model, using the existing ultrasound medical image-mask pairs to fine-tune the SAM model and generate automatic hints and segment the ultrasound medical image accordingly; Step 2, relying on the mask in the segmentation map and the image features of different resolutions, the anatomical structure in the ultrasound medical image is predicted across views; Step 3, by comparing the predicted results with the true values, anatomical-level regional comparison of different resolutions and different views is achieved, the same anatomical structure between different views is used as a positive sample, and the views of different anatomical structures and different images are used as negative samples, the similarity between the features of the same anatomical structure is maximized, and the backbone model is trained by contrastive learning, wherein the backbone model includes a residual network and a mapping module; Step 4: retain the residual network of the online update network in the trained backbone model and perform image analysis tasks based on it, wherein the image analysis tasks include image classification and image segmentation.

2. The method according to claim 1, characterized in that The step 1 specifically includes: Step 1.1, integrate the automatic hint generator into the SAM model, extract the global features and local features of the image, and then use multi-head cross attention to fuse the global features and local features to obtain mixed features, where the expression of the automatic hint generator is F vc =LN(MHCA(F v ,F c ,F c )+F v ) F′ v =MLP(LN(MHCA(F vc ,T,T)+F vc )) T′=MLP(LN(MHCA(T,F vc ,F vc )+T)) Among them, MHCA(·) represents multi-head cross attention, LN represents regularization, MLP represents multi-layer perceptron, and F v is the global feature, F c is the local feature, F vc is a mixed feature of the global and local features of the image, T represents the initial token, and F′ v and T′ is determined by F vc And T are cross-attention calculated to obtain F′ v is the feature found in the image that corresponds to the initial token, and T′ is the feature found in the initial token that corresponds to the image area; Step 1.2, F′ v As a new F v , T′ is used as the new T, and then the two multi-head cross attention calculations in the previous step are performed based on the new mixed features and the new token to obtain the new F′ v and T', and repeat this step to obtain the final information fusion token as an automatic prompt; Step 1.3: Use the automatic hint and the initial global feature F as the image embedding v Input into the mask decoder to segment the ultrasonic medical image and generate the mask corresponding to the segmentation map.

3. The method according to claim 2, characterized in that The step 2 specifically includes: Step 2.1, for the ultrasound medical image x, two enhanced ultrasound images of different viewing angles are generated by random preprocessing, respectively denoted as view v1 and view v2, and are respectively input into the residual network to extract image features at different resolutions, and the mask groups corresponding to the image feature content are selected from the mask of the segmentation map and denoted as m and m′, respectively. The latent vector representation of the mask corresponding to view v1 is calculated according to the preset formula, where the preset formula is: Among them, f s [i,j] represents the feature vector of the feature map with resolution s at (i,j), m s [i,j] represents the mask value corresponding to (i,j) after the mask group m is adjusted to the resolution s, h s A multi-granular latent vector representation corresponding to view v1 of an ultrasound medical image; Step 2.2, substitute the feature vector and mask group m' of view v2 into the preset formula to calculate the multi-granularity latent vector representation h' of view v2 s ; Step 2.3, for the multi-granularity latent vector h of view v1 and view v2 s and h′ s , using two mapping heads to map it into z θ,s and z ξ,s , the overall network architecture for processing v1 is called the online update network, and the overall network architecture for processing v2 is called the target network, where θ is the parameter of the online update network, ξ is the parameter of the target network, and by using the prediction head of the online update network, the result of one view predicting the content of another view is obtained, and then the error is calculated between it and the actual result output by the target network.

4. The method according to claim 3, characterized in that: The step 3 specifically includes: Step 3.1, define the similarity formula between the structures of view v1 and view v2 at resolution s as Among them, τ is the temperature hyperparameter, q θ,s (·) represents the prediction map of the online update network, Indicates the distance between the result of predicting view v2 based on view v1 and the actual view v2; Step 3.2: After obtaining n negative samples through random sampling, design a multi-granularity contrast loss function based on the similarity formula Among them, n represents the number of negative samples, is the distance between v1 and negative sample i; Step 3.3: Based on the multi-granularity contrast loss function, view v2 is used as the input of the online update network for prediction, and the contrast loss at all granularities is integrated to obtain the overall loss function of the final backbone model. Where K represents the number of samples, s represents different granularities, and S represents the total number of granularities; In step 3.4, based on the overall loss function, the parameters θ in the online update network of the backbone model are updated by reducing the loss, and accordingly, the parameters ξ in the target network of the backbone model are updated using the exponential moving average method.

5. The method according to claim 4, characterized in that When the analysis task is image classification, step 4 specifically includes: The trained residual network f θ As the encoder part of the classification model, it extracts multi-level image features of ultrasound medical images, uses the feature pyramid network to fuse image features at different levels, and finally uses the fully connected layer to flatten the features to complete the classification task.

6. The method according to claim 5, characterized in that When the analysis task is image segmentation, step 4 specifically includes: Replace the CNN network in the Mask-RCNN model with the trained residual network f θ , and then use the SGD optimizer to optimize and fine-tune the model to segment the target area in the ultrasound medical image, thereby obtaining the corresponding segmentation mask.