Medical ultrasound image data set construction method and system and storage medium
Through a multimodal segmentation model and a similarity-driven mask propagation mechanism, a medical ultrasound image dataset is automatically generated, which solves the problems of low construction efficiency and high professional level requirements in existing technologies and achieves efficient and accurate dataset construction.
Patent Information
- Application Number
- CN202510857379.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-25
AI Technical Summary
In the existing technology, the construction efficiency of medical ultrasound image datasets is low and the professional level of annotation personnel is required to be high, making it difficult to quickly build high-quality large-scale datasets.
A multimodal segmentation model and a similarity-driven mask propagation mechanism are adopted. By calculating the similarity between the ultrasound image to be annotated and the annotated dataset, reference images are screened and weighted averaged, and the predicted mask is automatically generated by combining text labels and box selection prompts. Confidence evaluation is used to ensure the accuracy of annotation.
It significantly improves the efficiency of dataset annotation, reduces the requirements for the professional level of annotation personnel, and quickly builds a high-quality large-scale medical ultrasound image dataset.
Smart Images

Figure CN120766056A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ultrasonic image processing, and in particular to a method, system and storage medium for constructing a medical ultrasonic image data set. Background Art
[0002] During ultrasound-guided regional nerve blocks, obtaining and interpreting optimal ultrasound images is a fundamental skill, which involves identifying key ultrasound anatomical structures. However, the anatomical structures can be complex and varied, and despite the continuous improvement in ultrasound image quality, the ability to obtain or interpret the correct ultrasound view remains a limiting factor. Therefore, clinicians need not only a solid theoretical foundation but also extensive practical clinical experience. This limits the promotion of ultrasound-guided nerve blocks in hospitals at all levels and also means that it takes a long time to train an anesthesiologist who is proficient in ultrasound-guided nerve blocks.
[0003] In recent years, the emergence of artificial intelligence (AI), particularly deep learning-based AI, has made a significant impact in medical imaging. This technology provides doctors with powerful tools for accurately segmenting target tissues, significantly improving the efficiency and accuracy of medical image analysis and enabling applications in teaching, training, and even clinical practice. However, deep learning models typically require large, high-quality datasets for training, while the labeling of medical images, particularly ultrasound images, is complex and time-consuming, requiring a high level of expertise from the labelers. This has become a significant factor limiting the widespread application of AI in medical imaging.
[0004] In the past, ultrasound image datasets mostly used polygonal methods to annotate different tissues. This approach is not only time-consuming but also requires years of experience from physicians. This manual annotation process places high demands on the annotators, leading to low data acquisition efficiency. With the recent development of semi-automatic segmentation models, such as the Segment Anything Model (SAM), a cue-based semi-automatic segmentation model, it has been able to semi-automatically annotate untrained images. This technology has been rapidly applied to the construction of medical image datasets, effectively improving dataset construction efficiency and reducing the time cost of physician annotation. However, existing medical semi-automatic segmentation models often rely on visual cues (bounding boxes or positive and negative sample points) and textual cue input, obtaining the desired segmentation results by selecting the bounding box of the segmented region or positive sample points. While this improves annotation efficiency, it still requires physicians with extensive clinical experience to complete the annotation, and the data acquisition cost remains very high. Therefore, how to improve the efficiency and data quality of ultrasound image dataset construction while reducing the professional requirements for annotators has become an urgent problem to be solved. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a method, system and storage medium for constructing a medical ultrasound image dataset, thereby improving the construction efficiency and data quality of the ultrasound image dataset and reducing the requirements on the professional level of the annotation personnel.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A method for constructing a medical ultrasound image dataset comprises the following steps:
[0008] S1: Acquire ultrasound images, annotate the anatomical structures in the ultrasound images with masks and text labels, and construct the initial dataset;
[0009] S2: training a pre-built multimodal segmentation model based on the initial dataset;
[0010] S3: Obtain the ultrasound image to be annotated, calculate the similarity between the ultrasound image to be annotated and the annotated ultrasound images in the initial data set, and select the most similar set of annotated ultrasound images as reference images;
[0011] S4: Based on the calculated similarity, the mask corresponding to the reference image is weighted averaged and converted into a binary mask, which is used as the fusion mask of the ultrasound image to be annotated; based on the calculated similarity, the bounding rectangle area corresponding to the mask of the reference image is weighted fused to obtain a weighted fused bounding rectangle, which is used as a selection prompt for the ultrasound image to be annotated;
[0012] S5: The ultrasound image to be annotated, the corresponding fusion mask, the text label, and the box selection prompt are used as inputs of the trained multimodal segmentation model to generate a segmentation result of the ultrasound image to be annotated;
[0013] S6: For the segmentation results of the ultrasound image to be annotated output by the multimodal segmentation model, select a confidence indicator to perform confidence calculation, and set a confidence threshold for quality assessment and review;
[0014] S7: After adding the ultrasound image segmentation results approved in S6 to the initial data set, the new initial data set is used to train the multimodal segmentation model and select reference images for subsequent ultrasound images to be annotated.
[0015] Furthermore, the similarity calculation process includes:
[0016] Adjusting the image resolution of the ultrasound image to be annotated and the annotated ultrasound images in the initial data set to the same size;
[0017] Using a pre-trained image encoder to obtain feature vectors of each image in the ultrasound image to be annotated and the annotated ultrasound image;
[0018] The cosine similarity calculation method is used to calculate the similarity between the feature vector of the ultrasound image to be annotated and the feature vectors of each annotated ultrasound image. The corresponding calculation expression is:
[0019]
[0020] Where s i is the similarity between the ultrasound image to be annotated and the i-th annotated ultrasound image, A is the feature vector of the ultrasound image to be annotated, B is the feature vector of the annotated ultrasound image, · is the vector dot product, ‖·‖ is the Euclidean distance of the feature vectors.
[0021] Furthermore, in the similarity calculation process, the method also assigns higher weights to temporally close frames based on temporal coherence to obtain weighted cosine similarity as the final similarity calculation result. The corresponding calculation expression is:
[0022] s i =ω t s i
[0023] In the formula, ω t With s i The product amplitude is given by s i ,ω t is the time weight, which is assigned according to the similarity between the time sequence of the annotated ultrasound image and the time sequence of the ultrasound image to be annotated.
[0024] Furthermore, the binary mask generation process is specifically as follows:
[0025] The mask of the reference image is weighted averaged based on the calculated similarity. The corresponding calculation expression is:
[0026]
[0027] Where M fused (p) is the weighted average mask of pixel position p, M i is the mask of the i-th reference image, M i (p) is the mask value of the i-th reference image at pixel position p, s i is the similarity between the ultrasound image to be annotated and the i-th reference image;
[0028] The weighted average mask is converted into a binary mask by setting a threshold. The corresponding calculation expression is:
[0029]
[0030] Where M binary (p) is the binary mask of pixel position p, and T is the threshold of the binary mask.
[0031] Furthermore, the process of generating the weighted fusion bounding rectangle is specifically as follows:
[0032] For each reference image mask, calculate the corresponding bounding rectangle area B i =[x min ,y min ,x max ,y max ], where x min 、y min 、x max 、y max Represent the minimum and maximum values of the bounding rectangle on the x-axis and y-axis respectively, so as to calculate the weighted fusion bounding rectangle by combining the calculated similarity. The corresponding calculation expression is:
[0033]
[0034] Where B fused is the weighted fusion bounding rectangle, s i is the similarity between the ultrasound image to be annotated and the i-th reference image, n is the number of reference images, x min,i ,y min,i ,x max,i ,y max,i Indicates the minimum and maximum values of the bounding rectangle of the i-th reference image on the x-axis and y-axis.
[0035] Furthermore, the method uses the predicted intersection-over-union ratio as a confidence indicator, and the quality assessment and review process is specifically as follows: if the prediction result of the multimodal segmentation model is higher than the confidence threshold, it is considered that the prediction result does not need manual correction and is directly added to the initial data set; if the prediction result is lower than the confidence threshold, the prediction result is manually corrected and then added to the initial data set.
[0036] Furthermore, in step S1, a polygonal annotation method is used to complete the annotation of the anatomical structure of the ultrasound image.
[0037] Furthermore, the multimodal segmentation model is an improved SAM model, which is provided with a prompt encoding module for processing masks, text labels and box selection prompts.
[0038] This embodiment further provides a system for implementing the above-mentioned method for constructing a medical ultrasound image dataset, including:
[0039] Image import module, used to load ultrasound images;
[0040] The manual annotation module is used to annotate the anatomical structures in the loaded ultrasound images with masks and text labels to construct the initial dataset;
[0041] A multimodal segmentation model training module is used to train a pre-built multimodal segmentation model based on the initial dataset;
[0042] An image similarity calculation module is used to obtain an ultrasound image to be annotated, calculate the similarity between the ultrasound image to be annotated and the annotated ultrasound images in the initial data set, and select the most similar set of annotated ultrasound images as a reference image;
[0043] a mask propagation module, configured to perform weighted averaging of the masks corresponding to the reference images based on the calculated similarity, and convert the result into a binary mask, which serves as a fusion mask for the ultrasound image to be annotated; perform weighted fusion of the bounding rectangles corresponding to the masks of the reference images based on the calculated similarity, and obtain a weighted fusion bounding rectangle, which serves as a selection prompt for the ultrasound image to be annotated; and use the ultrasound image to be annotated, the corresponding fusion mask, the text label, and the selection prompt as input to the trained multimodal segmentation model to generate a segmentation result for the ultrasound image to be annotated, and add the result to the initial dataset;
[0044] The dataset management module is used to organize and maintain datasets.
[0045] This embodiment further provides a computer-readable storage medium, on which a computer program is stored. The computer program is used by a processor to execute the above-mentioned method.
[0046] Compared with the prior art, the present invention has the following advantages:
[0047] The present invention introduces a multimodal segmentation model and a similarity-driven mask propagation mechanism. On the one hand, the multimodal segmentation model is trained using the manually annotated initial dataset. On the other hand, the reference image is screened by calculating the similarity between the ultrasound image to be annotated and the annotated dataset. The mask and the circumscribed rectangular area of each reference image are weighted averaged according to the similarity to obtain a fused mask and a box selection prompt. Together with the text label, it serves as the input of the multimodal segmentation model to automatically generate a predicted mask. The accuracy of the annotation is guaranteed based on confidence assessment, which significantly reduces the workload of manual annotation and lowers the professional level requirements for annotators. The efficiency of dataset annotation and the speed of manual correction are greatly improved, thereby realizing the rapid construction of a high-quality, large-scale medical ultrasound image dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic diagram of a flow chart of a method for constructing a medical ultrasound image dataset provided in an embodiment of the present invention;
[0049] Figure 2 A schematic diagram of obtaining a reference image based on image similarity provided in an embodiment of the present invention;
[0050] Figure 3 The figure is a flow chart of a confidence assessment process provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0052] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0053] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0054] Example 1
[0055] like Figure 1 As shown, this embodiment provides a method for constructing a medical ultrasound image dataset, comprising the following steps:
[0056] S1: Acquire ultrasound images, annotate the anatomical structures in the ultrasound images with masks and text labels, and construct the initial dataset;
[0057] S2: Train the pre-built multimodal segmentation model based on the initial dataset;
[0058] S3: Obtain the ultrasound image to be annotated, calculate the similarity between the ultrasound image to be annotated and the annotated ultrasound images in the initial data set, and select the most similar set of annotated ultrasound images as reference images;
[0059] S4: Based on the calculated similarity, the mask corresponding to the reference image is weighted averaged and converted into a binary mask, which is used as the fusion mask of the ultrasound image to be annotated; based on the calculated similarity, the bounding rectangle area corresponding to the mask of the reference image is weighted fused to obtain a weighted fused bounding rectangle, which is used as a selection prompt for the ultrasound image to be annotated;
[0060] S5: The ultrasound image to be annotated, the corresponding fusion mask, the text label and the box selection prompt are used as inputs of the trained multimodal segmentation model to generate the segmentation result of the ultrasound image to be annotated.
[0061] The fused mask, text label, and bounding box hint are derived from a weighted fusion of the reference image and serve as inputs to the multimodal segmentation model. Using only one or two of these hints is susceptible to data quality and annotation errors. For example, using only the bounding box hint may lead to model misjudgment due to inaccurate annotation, while using only the mask hint may affect the segmentation results due to noise or artifacts. By fusing multiple hints, the model can form complementary relationships between different hints, reducing its reliance on a single hint and thus improving its adaptability to diverse and complex scenes.
[0062] S6: For the prediction results of the ultrasound images to be annotated output by the multimodal segmentation model, the predicted intersection-over-union ratio is used as the confidence indicator, and a confidence threshold is set. If the prediction result of the multimodal segmentation model is higher than the confidence threshold, it is considered that the prediction result does not need manual correction and is directly added to the initial dataset; if the prediction result is lower than the confidence threshold, the prediction result is manually corrected and added to the initial dataset.
[0063] This is equivalent to evaluating the quality of the generated segmentation mask through a confidence assessment mechanism, and judging whether the user needs to provide additional visual cues or manual correction based on the intersection-over-union score.
[0064] S7: After adding the ultrasound image segmentation results approved in S6 to the initial dataset, the new initial dataset is used to train the multimodal segmentation model and select reference images for subsequent ultrasound images to be annotated.
[0065] This is equivalent to adding the reviewed and revised labeled data to the labeled dataset for iterative updating.
[0066] Preferably, in step S1, a polygonal annotation method is used to accurately annotate key anatomical structures of the ultrasound image.
[0067] In this embodiment, the process of constructing the initial data set specifically includes:
[0068] Import the ultrasound image to be annotated;
[0069] Provides annotation tools such as polygons, brushes, and erasers to adapt to the complexity and variability of different anatomical structures;
[0070] Supports real-time preview and modification of annotation results with different tags;
[0071] Supports exporting the true grayscale image of the annotation results;
[0072] Among them, the manual annotation is completed by doctors with at least 5 years of clinical experience, ensuring high-quality standards of the initial data set, and finally obtaining the mask and text label of each ultrasound image.
[0073] In step S2, the multi-modal segmentation model is an improved Segment Anything Model (SAM) model, which is provided with a prompt encoding module for processing multi-modal inputs such as masks, text labels and frame selection prompts, and outputs a segmentation mask and a predicted intersection over union.
[0074] Since the prompt encoder of the SAM model does not provide a specific implementation method for using text prompts as input, the present application generates a text feature vector for different ultrasound tissue labels based on a CLIP Text Encoder, and projects it to the prompt encoder as input through a two-layer MLP, and freezes the CLIP Text Encoder during the model training process.
[0075] In step S3, the similarity calculation process includes:
[0076] Adjust the image resolution of the to-be-labeled ultrasound image and the labeled ultrasound image in the initial data set to the same size MxN;
[0077] Obtain the feature vectors of each image in the to-be-labeled ultrasound image and the labeled ultrasound image using the pre-trained image encoder;
[0078] Calculate the similarity between the feature vectors of the to-be-labeled ultrasound image and the labeled ultrasound image using the cosine similarity calculation method; it is particularly suitable for ultrasound images of the same anatomical site, different patients or different time points of the same patient. If the anatomical structures of the two images are similar, the cosine similarity will be closer to 1, and if the content difference between the two images is large, the cosine similarity will be closer to 0 or negative. The cosine similarity calculation formula is:
[0079]
[0080] In the formula, s i is the similarity of the to-be-labeled ultrasound image and the i-th labeled ultrasound image, A is the feature vector of the to-be-labeled ultrasound image, B is the feature vector (such as the 512-dimensional vector extracted by the CLIP text encoder) of the labeled ultrasound image, A.B represents the dot product of the vectors, and ‖A‖ and ‖B‖ represent the Euclidean distance of the vectors.
[0081] Preferably, for consecutive frames collected by a video stream, considering that images close in time are more likely to be similar, the temporal coherence is increased, and higher weights are given to frames close in time. The weighted cosine similarity calculation formula is:
[0082] s i=ω t s i
[0083] In the formula, ω t With s i The product amplitude is given by s i ,ω t is a time weight (eg, 1.2), which is assigned according to the degree of similarity between the time sequence of the labeled ultrasound image and the time sequence of the ultrasound image to be labeled, and a higher weight is given to several frames near the current frame.
[0084] A similarity threshold (e.g., 0.9) is set to filter out images with too low similarity, and only a set of highly similar candidate images is retained as reference images.
[0085] In step S4, the binary mask generation process is specifically as follows:
[0086] For each text label, the reference mask of the reference image is weighted averaged based on the similarity, and a binary mask is generated by thresholding; for each pixel position p, the formula for calculating the weighted average mask is:
[0087]
[0088] Where M fused (p) is the weighted average mask of pixel position p, M i is the mask of the i-th reference image, M i (p) is the mask value (0 or 1) of the i-th reference image at pixel position p, s i is the similarity between the ultrasound image to be annotated and the i-th reference image;
[0089] For the weighted average mask M fused It is a floating point matrix between [0,1], which is converted into a binary mask by setting a threshold. For each pixel position p, the threshold calculation formula is:
[0090]
[0091] Where M binary (p) is the binary mask of pixel position p, and T is the threshold of the binary mask (e.g., 0.6), which can be dynamically adjusted to ensure that the generated binary mask is neither too large nor too small.
[0092] The process of generating a weighted fusion bounding rectangle includes extracting the bounding rectangle area of the reference mask and generating a more stable box selection prompt through weighted fusion. Specifically:
[0093] For each reference image mask M i , calculate the corresponding circumscribed rectangular area B i =[xmin ,y min ,x max ,y max ], where x min 、y min 、x max 、y max Represent the minimum and maximum values of the bounding rectangle on the x-axis and y-axis respectively, so as to calculate the weighted fusion bounding rectangle by combining the calculated similarity. The corresponding calculation expression is:
[0094]
[0095] Where B fused is the weighted fusion bounding rectangle, s i is the similarity between the ultrasound image to be annotated and the i-th reference image, n is the number of reference images, x min,i ,y min,i ,x max,i ,y max,i Indicates the minimum and maximum values of the bounding rectangle of the i-th reference image on the x-axis and y-axis.
[0096] Align and match the fused binary mask and box selection hint with the image to be annotated;
[0097] In step S5, the text label, fusion mask and selection prompt are used as inputs of the multimodal segmentation model to generate the segmentation result of the image to be annotated.
[0098] In this embodiment, the confidence evaluation in step S6 specifically includes:
[0099] The predicted IoU value output by the multimodal segmentation model is used as a confidence indicator. If it is above a preset threshold (for example, 0.8), the predicted mask is considered reliable and does not require manual correction. The predicted mask is directly output. If it is below the threshold, the user needs to input and click prompts to obtain a corrected predicted mask or directly use a brush to obtain a corrected predicted mask. Users can efficiently correct the generated mask by clicking positive sample points in false negative areas or negative sample points in false positive areas.
[0100] In this embodiment, the data iterative update process in step S7 specifically includes:
[0101] The reviewed and corrected annotated data are added to the annotated dataset. On the one hand, it is used to iteratively optimize the multimodal segmentation model to improve the segmentation and annotation accuracy, and on the other hand, it is used to select reference images for the ultrasound images to be annotated.
[0102] Example 2
[0103] This embodiment provides a system for implementing a method for constructing a medical ultrasound image dataset as described in Example 1, including:
[0104] Image import module, used to load ultrasound images;
[0105] The manual annotation module is used to annotate the anatomical structures in the loaded ultrasound images with masks and text labels to construct the initial dataset;
[0106] A multimodal segmentation model training module is used to train a pre-built multimodal segmentation model based on the initial dataset;
[0107] An image similarity calculation module is used to obtain an ultrasound image to be annotated, calculate the similarity between the ultrasound image to be annotated and the annotated ultrasound images in the initial data set, and select the most similar set of annotated ultrasound images as a reference image;
[0108] The mask propagation module is used to perform weighted averaging of the masks corresponding to the reference image based on the calculated similarity and convert them into a binary mask, which serves as the fusion mask of the ultrasound image to be annotated. The module also performs weighted fusion of the bounding rectangle area corresponding to the mask of the reference image based on the calculated similarity to obtain a weighted fused bounding rectangle, which serves as the selection prompt for the ultrasound image to be annotated. The ultrasound image to be annotated, the corresponding fusion mask, the text label, and the selection prompt are used as input to the trained multimodal segmentation model to generate the segmentation result of the ultrasound image to be annotated and add it to the initial dataset.
[0109] Confidence Assessment Module: Calculates the quality score of the segmentation results, evaluates the quality of the generated segmentation mask, and provides manual correction function;
[0110] The dataset management module is used to organize and maintain datasets.
[0111] It should be noted that the specific content and beneficial effects of the system of this application can be found in the above-mentioned method embodiment and will not be repeated here.
[0112] The functions described above in this document may be implemented as a computer software program that is tangibly contained in a computer-readable storage medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the CPU, the functions described above may be performed.
[0113] Computer program code for carrying out operations of the methods of the present application can be written in any combination of one or more programming languages. The computer program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program code, when executed by the processor or controller, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer program code can execute entirely on a machine, partly on a machine, as a stand-alone software package, partly on a machine and partly on a remote machine or entirely on a remote machine or server.
[0114] The embodiment also provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the medical ultrasound image dataset construction method in embodiment 1.
[0115] In the context of the present application, a computer readable storage medium can be a tangible medium which can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer readable storage medium can be a machine-readable signal medium or a machine-readable storage medium. The computer readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the machine readable storage medium will include one or more of an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0116] The preferred embodiments of the present application have been described in detail. It should be understood that modifications and variations can be made by those skilled in the art without creating spurious logical relationships, based on the concept of the present application. Therefore, any technical solutions obtained by logical analysis, reasoning or limited experiments based on the concept of the present application in the prior art should be within the scope of protection defined by the claims.
Claims
1. A method for constructing a medical ultrasound image dataset, characterized in that: The following steps are involved: S1: Acquire ultrasound images, annotate the anatomical structures in the ultrasound images with masks and text labels, and construct the initial dataset; S2: training a pre-built multimodal segmentation model based on the initial dataset; S3: Obtain the ultrasound image to be annotated, calculate the similarity between the ultrasound image to be annotated and the annotated ultrasound images in the initial data set, and select the most similar set of annotated ultrasound images as reference images; S4: Based on the calculated similarity, the mask corresponding to the reference image is weighted averaged and converted into a binary mask, which is used as the fusion mask of the ultrasound image to be annotated; based on the calculated similarity, the bounding rectangle area corresponding to the mask of the reference image is weighted fused to obtain a weighted fused bounding rectangle, which is used as a selection prompt for the ultrasound image to be annotated; S5: The ultrasound image to be annotated, the corresponding fusion mask, the text label, and the box selection prompt are used as inputs of the trained multimodal segmentation model to generate a segmentation result of the ultrasound image to be annotated; S6: For the segmentation results of the ultrasound image to be annotated output by the multimodal segmentation model, select a confidence indicator to perform confidence calculation, and set a confidence threshold for quality assessment and review; S7: After adding the ultrasound image segmentation results approved in S6 to the initial data set, the new initial data set is used to train the multimodal segmentation model and select reference images for subsequent ultrasound images to be annotated.
2. The method for constructing a medical ultrasound image dataset according to claim 1, wherein: The similarity calculation process includes: Adjusting the image resolution of the ultrasound image to be annotated and the annotated ultrasound images in the initial data set to the same size; Using a pre-trained image encoder to obtain feature vectors of each image in the ultrasound image to be annotated and the annotated ultrasound image; The cosine similarity calculation method is used to calculate the similarity between the feature vector of the ultrasound image to be annotated and the feature vectors of each annotated ultrasound image. The corresponding calculation expression is: Where s i is the similarity between the ultrasound image to be annotated and the i-th annotated ultrasound image, A is the feature vector of the ultrasound image to be annotated, B is the feature vector of the annotated ultrasound image, · is the vector dot product, ‖·‖ is the Euclidean distance of the feature vectors.
3. The method for constructing a medical ultrasound image dataset according to claim 2, wherein: In the similarity calculation process, the method also assigns higher weights to temporally close frames based on temporal coherence, and obtains weighted cosine similarity as the final similarity calculation result. The corresponding calculation expression is: s i =ω t s i In the formula, ω t With s i The product amplitude is given by s i ,ω t is the time weight, which is assigned according to the similarity between the time sequence of the annotated ultrasound image and the time sequence of the ultrasound image to be annotated.
4. The method for constructing a medical ultrasound image dataset according to claim 1, wherein: The generation process of the binary mask is specifically as follows: The mask of the reference image is weighted averaged based on the calculated similarity. The corresponding calculation expression is: Where M fused (p) is the weighted average mask of pixel position p, M i is the mask of the i-th reference image, M i (p) is the mask value of the i-th reference image at pixel position p, s i is the similarity between the ultrasound image to be annotated and the i-th reference image; The weighted average mask is converted into a binary mask by setting a threshold. The corresponding calculation expression is: Where M binary (p) is the binary mask of pixel position p, and T is the threshold of the binary mask.
5. The method for constructing a medical ultrasound image dataset according to claim 1, wherein: The generation process of the weighted fusion bounding rectangle is specifically as follows: For each reference image mask, calculate the corresponding bounding rectangle area B i =[x min ,y min ,x max ,y max ], where x min 、y min 、x max 、y max Represent the minimum and maximum values of the bounding rectangle on the x-axis and y-axis respectively, so as to calculate the weighted fusion bounding rectangle by combining the calculated similarity. The corresponding calculation expression is: Where B fused is the weighted fusion bounding rectangle, s i is the similarity between the ultrasound image to be annotated and the i-th reference image, n is the number of reference images, x min,i ,y min,i ,x max,i ,y max,i Indicates the minimum and maximum values of the bounding rectangle of the i-th reference image on the x-axis and y-axis.
6. The method for constructing a medical ultrasound image dataset according to claim 1, wherein: The method uses the predicted intersection-over-union ratio as a confidence indicator. The quality assessment and review process is specifically as follows: if the prediction result of the multimodal segmentation model is higher than the confidence threshold, it is considered that the prediction result does not need manual correction and is directly added to the initial data set; if the prediction result is lower than the confidence threshold, the prediction result is manually corrected and then added to the initial data set.
7. The method for constructing a medical ultrasound image dataset according to claim 1, wherein: In step S1, a polygonal annotation method is used to complete the annotation of the anatomical structure of the ultrasound image.
8. The method for constructing a medical ultrasound image dataset according to claim 1, wherein: The multimodal segmentation model is an improved SAM model, which is provided with a prompt encoding module for processing masks, text labels and box selection prompts.
9. A system for implementing the method for constructing a medical ultrasound image dataset according to any one of claims 1 to 8, characterized in that: include: Image import module, used to load ultrasound images; The manual annotation module is used to annotate the anatomical structures in the loaded ultrasound images with masks and text labels to construct the initial dataset; A multimodal segmentation model training module is used to train a pre-built multimodal segmentation model based on the initial dataset; An image similarity calculation module is used to obtain an ultrasound image to be annotated, calculate the similarity between the ultrasound image to be annotated and the annotated ultrasound images in the initial data set, and select the most similar set of annotated ultrasound images as a reference image; a mask propagation module, configured to perform weighted averaging of the masks corresponding to the reference images based on the calculated similarity, and convert the result into a binary mask, which serves as a fusion mask for the ultrasound image to be annotated; perform weighted fusion of the bounding rectangles corresponding to the masks of the reference images based on the calculated similarity, and obtain a weighted fusion bounding rectangle, which serves as a selection prompt for the ultrasound image to be annotated; and use the ultrasound image to be annotated, the corresponding fusion mask, the text label, and the selection prompt as input to the trained multimodal segmentation model to generate a segmentation result for the ultrasound image to be annotated, and add the result to the initial dataset; The dataset management module is used to organize and maintain datasets.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is used by a processor to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Interactive segmentation intelligent labeling method applied to ultrasonic image
CN117218653A
Masking pre-training method for ultrasound image segmentation
CN118537348A
Zero sample reference image segmentation method based on hierarchical prompt and directional clue
CN119049057A
Training method of medical image segmentation model, image segmentation method and related device
CN119131522A
Method and system for efficiently marking and segmenting lesion area in medical image
CN120163981A