Ultrasonic image segmentation method and device, electronic equipment and storage medium
By constructing an image segmentation model that includes a visual encoder, a text encoder, and an affine transformation-guided module, the problem of high-precision segmentation of multi-organ and multi-view ultrasound images was solved, achieving accurate identification and segmentation of multi-organ regions and improving segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202511060507.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-28
AI Technical Summary
Existing medical ultrasound image segmentation methods struggle to achieve high-precision, low-cost target extraction in multi-organ, multi-view images, especially lacking sufficient semantic accuracy for segmented regions.
An image segmentation model is constructed, including a visual encoder module, a text encoder module, and an affine transformation guidance module. By extracting multi-scale visual and semantic features, and fusing them using affine transformation, a segmentation mask map is generated.
It improves the clarity and segmentation accuracy of category semantics in medical ultrasound images, enables accurate identification and segmentation of multi-organ regions, and enhances the robustness of the model in multi-category, complex structures, and multi-angle images.
Smart Images

Figure CN121033069A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical image processing and segmentation, and particularly relates to an ultrasound image segmentation method and device, electronic equipment and a storage medium. BACKGROUND
[0002] Medical ultrasound image segmentation is a core task in computer-aided diagnosis systems, and plays a crucial role in the detection, diagnosis and treatment planning process of clinical diseases. Its goal is to accurately identify and extract the contours of organs, tissues or lesion regions from ultrasound images to support doctors in quantitative analysis and accurate diagnosis and treatment. In recent years, deep learning technology has made breakthrough progress in the field of image processing, and has promoted the evolution of medical ultrasound image segmentation methods from traditional image processing algorithms to end-to-end learning models. Models based on U-Net, FCN, Transformers and other structures perform well on multiple public datasets, and even exceed the performance of human experts on some tasks. Despite this, the medical ultrasound image segmentation task still faces a series of unique challenges.
[0003] In related technologies, a multi-organ general model specifically for ultrasound images such as SAMUS, although point or frame is introduced as guide information, it is difficult to assign accurate class semantics to the segmentation region, and the practicality still needs to be improved. Therefore, designing a medical ultrasound image general segmentation method with strong generalization ability, which can realize high-precision and low-cost target extraction in multi-organ and multi-view ultrasound images, is an important problem that needs to be solved at present. SUMMARY
[0004] Therefore, the present application provides an ultrasound image segmentation method, device, electronic equipment and storage medium, which can realize accurate identification and segmentation of ultrasound images.
[0005] A first aspect of the embodiment of the present application provides an ultrasound image segmentation method, comprising: constructing an image segmentation model, wherein the image segmentation model comprises a visual encoder module, a text encoder module and an affine transformation guide module; obtaining a medical ultrasound image to be segmented and text auxiliary information corresponding to the medical ultrasound image, the text auxiliary information representing medical prior knowledge; inputting the medical ultrasound image into the image segmentation model to obtain a segmentation mask graph of the medical ultrasound image, the segmentation mask graph being used to segment different organs in the medical ultrasound image into different regions; wherein the visual encoder module is used to extract multi-scale visual features of the medical ultrasound image, the text encoder module is used to extract semantic features of the text auxiliary information, and the affine transformation guide module is used to fuse the multi-scale visual features and the semantic features through affine transformation to generate the segmentation mask graph.
[0006] In a possible implementation, the affine transformation guiding module comprises a projection unit, a modulation unit and a fusion unit; the inputting the medical ultrasound image into the image segmentation model to obtain a segmentation mask map of the medical ultrasound image comprises: inputting the medical ultrasound image into the image segmentation model; the visual encoder module extracts the multi-scale visual features, and the text encoder module extracts the semantic features; the projection unit performs dimension alignment and compression dimension reduction on the multi-scale visual features and the semantic features; the modulation unit converts the compressed and dimension-reduced semantic features into affine parameters, and performs affine transformation on the compressed and dimension-reduced multi-scale visual features; the fusion unit performs upsampling on the connection of the affine parameters and the affine-transformed multi-scale visual features, and performs skip connection fusion with the encoder features to obtain the segmentation mask map.
[0007] In a possible implementation, the visual encoder module is a ConvNeXt-tiny convolutional neural network model; the ConvNeXt-tiny convolutional neural network model comprises a hierarchical convolution design for extracting multi-scale features, a large convolution kernel for capturing global features, and an inverted bottleneck structure for reducing the amount of model calculation.
[0008] In a possible implementation, the text encoder module is a CXR-BERT pre-training language model; the CXR-BERT pre-training language model is based on a Transformer neural network architecture and is obtained by pre-training on a historical chest X-ray report dataset; wherein the Transformer neural network architecture is used to capture long-distance dependencies in the text according to a self-attention mechanism.
[0009] In a possible implementation, the image segmentation model is trained according to a sample dataset, and the sample dataset comprises a plurality of historical multi-organ multi-view medical ultrasound images and historical text auxiliary information corresponding to each of the historical multi-organ multi-view medical ultrasound images, the historical text auxiliary information representing medical prior knowledge.
[0010] In a possible implementation, before training the image segmentation model according to the sample dataset, the method further comprises: pre-processing the sample dataset to obtain a target sample dataset, wherein the pre-processing at least comprises image uniform size normalization, image noise suppression and image enhancement transformation; and the image segmentation model is trained according to the sample dataset, comprising: the image segmentation model is trained according to the target sample dataset.
[0011] In a possible implementation, the historical multi-organ multi-view medical ultrasound images at least include: a thyroid nodule multi-view medical ultrasound image, a thyroid gland multi-view medical ultrasound image, a breast cancer multi-view medical ultrasound image, a fetal head multi-view medical ultrasound image, a uterine tumor multi-view medical ultrasound image, a left atrium multi-view medical ultrasound image, a left ventricle multi-view medical ultrasound image, and a left ventricular wall multi-view medical ultrasound image.
[0012] In a second aspect, the embodiments of the present application further provide an ultrasound image segmentation device, comprising: a construction module, an acquisition module and an input module; the construction module is configured to construct an image segmentation model, wherein the image segmentation model comprises a visual encoder module, a text encoder module and an affine transformation guide module; the acquisition module is configured to acquire a medical ultrasound image to be segmented and text auxiliary information corresponding to the medical ultrasound image, the text auxiliary information representing medical prior knowledge; the input module is configured to input the medical ultrasound image into the image segmentation model to obtain a segmentation mask graph of the medical ultrasound image, the segmentation mask graph being configured to segment different organs in the medical ultrasound image in different regions; wherein the visual encoder module is configured to extract multi-scale visual features of the medical ultrasound image, the text encoder module is configured to extract semantic features of the text auxiliary information, and the affine transformation guide module is configured to fuse the multi-scale visual features and the semantic features through affine transformation to generate the segmentation mask graph.
[0013] In a third aspect, the embodiments of the present application further provide an electronic device, comprising a processor and a memory, the memory is configured to store instructions, and the processor is configured to call the instructions in the memory, so that the electronic device executes the ultrasound image segmentation method as described in the first aspect.
[0014] In a fourth aspect, the embodiments of the present application further provide a storage medium, which stores computer instructions, when the computer instructions run on an electronic device, the electronic device executes the ultrasound image segmentation method as described in the first aspect.
[0015] Compared with the related art, the embodiments of the present application have at least the following advantages: by constructing an image segmentation model, since the image segmentation model comprises a visual encoder module, a text encoder module and an affine transformation guide module, the semantic features of the text auxiliary information of the medical ultrasound image are extracted through the text encoder module, which can improve the clarity of the class semantics in the medical ultrasound image; and then the multi-scale visual features and the semantic features are fused through affine transformation, which realizes the adaptive regulation of the multi-view image structure, thereby effectively improving the segmentation accuracy and robustness of the image segmentation model in multi-class, complex structure and multi-angle images. Therefore, after the medical ultrasound image is input into the image segmentation model, a segmentation mask graph for precise recognition and segmentation of multi-organ regions can be obtained.
[0016] The technical effects achieved by the second, third, and fourth aspects mentioned above are similar to those achieved by the corresponding technical means in the first aspect, and will not be repeated here. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the steps of an ultrasound image segmentation method provided in an embodiment of this application.
[0018] Figure 2 This is a schematic diagram of the segmentation network structure of an image segmentation model provided in an embodiment of this application.
[0019] Figure 3 This is a schematic diagram of the structure of an affine transformation guiding module provided in an embodiment of this application.
[0020] Figure 4 This is a functional block diagram of an ultrasound image segmentation device provided in an embodiment of this application.
[0021] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0022] To better understand the above-mentioned objectives, features, and advantages of this application, the application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0023] The following description sets forth many specific details to provide a full understanding of this application. The described embodiments are only some, not all, of the embodiments of this application.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.
[0025] It should be further noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0026] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects, not to describe a specific order or sequence.
[0027] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0028] For ease of understanding, some concepts related to the embodiments of this application are illustrated and explained by way of example for reference.
[0029] The ConvNeXt-tiny convolutional neural network model is a lightweight convolutional neural network model in the ConvNeXt series, designed for efficient computing and resource-constrained scenarios, while maintaining strong feature extraction capabilities. ConvNeXt-tiny achieves lightweighting by reducing the number of channels and network depth, significantly reducing computational cost and parameter count while maintaining high accuracy, making it suitable for deployment on mobile or edge devices.
[0030] CXR-BERT pre-trained language model: This is a pre-trained language model optimized for the biomedical imaging field, specifically designed for processing chest X-ray (CXR) radiology reports. To address the unique characteristics of medical texts, such as negative sentences (e.g., "no pneumonia"), uncertain expressions (e.g., "possibly infected"), and technical terms (e.g., "pulmonary consolidation"), CXR-BERT uses biomedical corpora (PubMed abstracts, MIMIC-III clinical notes, and MIMIC-CXR reports) to construct a dedicated 30K-word WordPiece segmenter, enhancing its terminology parsing capabilities.
[0031] Textual auxiliary information refers to medical prior knowledge (labels, parameters). For example, benign or malignant labels can indirectly reflect nodule characteristics (such as malignancy often accompanied by microcalcifications and irregular edges). Functional parameters: EF value indicates cardiac function status (normal >55%).
[0032] Please refer to Figure 1 , Figure 1This is a flowchart illustrating the steps of an embodiment of the ultrasound image segmentation method of this application. Depending on different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted.
[0033] It should be noted that the ultrasound image segmentation method of this application embodiment can be applied to medical image processing scenarios, and its execution subject can be an ultrasound image segmentation device. For example, in a medical image processing scenario, medical ultrasound images can be segmented using an ultrasound image segmentation device. Of course, the ultrasound image segmentation method can also be applied to other scenarios that require image segmentation, and this application does not specifically limit it in this regard.
[0034] The specific process of this embodiment is as follows: Figure 1 As shown, it includes the following steps: S101, Construct an image segmentation model, which includes a visual encoder module, a text encoder module, and an affine transformation guidance module.
[0035] In some embodiments, the visual encoder module is a ConvNeXt-tiny convolutional neural network model; the ConvNeXt-tiny convolutional neural network model includes a hierarchical convolutional design for extracting multi-scale features, a large convolutional kernel for capturing global features, and an inverted bottleneck structure for reducing the computational cost of the model.
[0036] In some embodiments, the text encoder module is a CXR-BERT pre-trained language model; the CXR-BERT pre-trained language model is based on the Transformer neural network architecture and is obtained by pre-training on a historical chest X-ray report dataset; wherein, the Transformer neural network architecture is used to capture long-distance dependencies in the text based on a self-attention mechanism.
[0037] In some embodiments, the image segmentation model is trained based on a sample dataset, which includes multiple historical multi-organ multi-view medical ultrasound images and historical textual auxiliary information representing prior medical knowledge corresponding to each historical multi-organ multi-view medical ultrasound image.
[0038] Specifically, the sample dataset covers 8 organs and integrates 9 public datasets including TN3K, DDTI, TG3K, Dataset B, BUSI, HC18, MMOTU, CAMUS and HMC-QU.
[0039] More specifically, historical multi-organ multi-view medical ultrasound images include at least: multi-view medical ultrasound images of thyroid nodules, multi-view medical ultrasound images of thyroid glands, multi-view medical ultrasound images of breast cancer, multi-view medical ultrasound images of the fetal head, multi-view medical ultrasound images of uterine tumors, multi-view medical ultrasound images of the left atrium, multi-view medical ultrasound images of the left ventricle, and multi-view medical ultrasound images of the left ventricular wall.
[0040] In some embodiments, before training the image segmentation model based on the sample dataset, the method further includes: preprocessing the sample dataset to obtain a target sample dataset, wherein the preprocessing includes at least image uniform size normalization, image noise suppression, and image enhancement transformation; the image segmentation model is trained based on the sample dataset, including: the image segmentation model is trained based on the target sample dataset.
[0041] Specifically, by setting a supervised segmentation loss function, the image segmentation model is trained end-to-end using the backpropagation algorithm, enabling the image segmentation model to learn robust representations under different organs and different viewpoints.
[0042] S102, Obtain the medical ultrasound image to be segmented and the corresponding textual auxiliary information representing prior medical knowledge.
[0043] S103, input the medical ultrasound image into the image segmentation model to obtain the segmentation mask of the medical ultrasound image. The segmentation mask is used to segment different organs in the medical ultrasound image into different regions.
[0044] Specifically, the visual encoder module is used to extract multi-scale visual features from medical ultrasound images, the text encoder module is used to extract semantic features from text auxiliary information, and the affine transformation guidance module is used to fuse multi-scale visual features and semantic features through affine transformation to generate a segmentation mask.
[0045] In some embodiments, the affine transformation guidance module includes a projection unit, a modulation unit, and a fusion unit; inputting a medical ultrasound image into an image segmentation model to obtain a segmentation mask of the medical ultrasound image includes: inputting the medical ultrasound image into the image segmentation model; a visual encoder module extracting multi-scale visual features, and a text encoder module extracting semantic features; the projection unit aligning the dimensions of the multi-scale visual features and semantic features and performing compression and dimensionality reduction; the modulation unit converting the compressed and dimensionality-reduced semantic features into affine parameters and performing affine transformation on the compressed and dimensionality-reduced multi-scale visual features; and the fusion unit concatenating the affine parameters with the affine-transformed multi-scale visual features, upsampling them, and performing skip connections to fuse them with the encoder features to obtain a segmentation mask.
[0046] Specifically, the visual encoder uses a ConvNeXt-tiny network to extract visual features of the input medical ultrasound image at multiple scales; the text encoder uses the CXR-BERT model to embed and encode the category names of the text auxiliary information to obtain high-quality semantic feature representations; the affine transformation guidance module first projects and linearly maps the semantic features, then transforms them into channel-level scaling and offset parameters, and performs channel-by-channel affine modulation on the visual features to achieve cross-modal fusion of semantic guidance. This module dynamically injects text modulation information at multiple levels of the decoder to gradually strengthen category guidance; subsequently, the fusion unit concatenates the modulated semantic guidance features with the corresponding scale visual features in each decoding layer, and completes multimodal feature fusion through operations such as channel alignment, convolutional fusion, and residual connections. Finally, it constructs a spatial resolution recovery path through upsampling and skip connections to generate a segmentation mask map.
[0047] To facilitate understanding, the following will be combined with... Figure 2 and Figure 3 This embodiment provides a detailed explanation of how medical ultrasound image segmentation is achieved: Please refer to Figure 2 This is a schematic diagram of the segmentation network structure of the image segmentation model provided in an embodiment of this application. For the input image... via visual encoder Four feature maps of different sizes were obtained. .in, , , , .
[0048] The input text is processed through the CXR-BERT network to obtain the overall feature representation of the sentence. ). Subsequently, It is fused with T through the guidance module and through upsampling operation. Perform a jump connection to obtain As shown below: ;in, Indicates an upsampling operation. The instruction module indicates that, for and It is obtained in the same way, as shown below: ; Finally, The data is fed into the segmentation head to obtain the segmentation result, which is the segmentation mask.
[0049] Specifically, Figure 2The visual encoder shown uses ConvNeXt-tiny as its backbone network. ConvNeXt is a pure convolutional neural network architecture that modernizes traditional convolutional neural networks by borrowing design principles from Transformer, achieving excellent performance in tasks such as image classification, object detection, and semantic segmentation. ConvNeXt-tiny is a lightweight version of the ConvNeXt series, achieving a good balance between model size and computational efficiency. The main features of the ConvNeXt-tiny convolutional neural network are as follows: (1) Hierarchical design: ConvNeXt-tiny adopts a hierarchical design, which includes four stages, each of which is composed of multiple ConvNeXt blocks stacked together. This design can effectively extract multi-scale features.
[0050] (2) Large convolution kernel: ConvNeXt-tiny uses The large convolutional kernel, similar to the large receptive field in the Vision Transformer, can capture more global feature information.
[0051] (3) Inverted bottleneck structure: ConvNeXt-tiny introduces an inverted bottleneck structure in each ConvNeXt block, first using Convolution increases the channel dimension, then spatial convolution is performed, and finally... Convolution reduces channel dimensionality. This structure effectively reduces computational cost while maintaining the model's expressive power. Specifically, for an input feature map... The inverted bottleneck structure can be represented as: ; ; ;in, and They represent Convolution and Convolution operation, , and These represent the corresponding convolution kernel weights.
[0052] (4) GELU activation function: ConvNeXt-tiny uses the GELU activation function instead of the traditional ReLU activation function, which can improve the nonlinear expressive power of the model.
[0053] The GELU activation function is defined as follows: ;in, It is the cumulative distribution function of the standard normal distribution. The GELU activation function has smooth nonlinear properties near zero, which can better adapt to the training of deep neural networks.
[0054] Specifically, Figure 2 The text encoder shown employs CXR-BERT. CXR-BERT is a pre-trained language model based on the Transformer architecture, specifically designed and optimized for text reporting of chest X-ray images. By pre-training on a large-scale CXR report dataset, it learns rich medical knowledge and can effectively extract semantic information from CXR reports. Its main features are as follows: (1) Transformer architecture: CXR-BERT is based on the Transformer architecture and uses the self-attention mechanism to capture long-distance dependencies in text. Figure 2 The text encoder shown includes an embedding layer, multiple hidden layers, and a last hidden layer. Inputting the text "Segment the breastgland in the ultrasound image" into the embedding layer yields the final encoded result.
[0055] Specifically, for an input text sequence ,in The embedding representation of the i-th word can be represented as follows: Each layer of the Transformer encoder can be represented as: ;in, These represent the query, key, and value matrices, respectively. This represents the dimension of the key vector. Through a multi-head attention mechanism, CXR-BERT can learn different semantic representations in parallel, thereby gaining a more comprehensive understanding of textual information.
[0056] (2) Pre-training tasks: CXR-BERT was pre-trained on a large-scale CXR report dataset, employing two pre-training tasks. The first was a masked language model, where a portion of the input text was randomly masked, and the model was asked to predict the masked words. This task forces the model to learn the contextual relationships between words, thereby better understanding the semantics of the text. Specifically, for an input text sequence... The model randomly selects 15% of the words for masking, with 80% replaced by a special marker, 10% replaced by a random word, and 10% left unchanged. The goal of the model is to predict the masked words. ;in, This represents the text sequence after the mask.
[0057] The next step is sentence prediction. Given two sentences, the model is asked to determine whether they are consecutive. This task helps the model learn the relationships between sentences, thus better understanding the discourse structure of the text. Specifically, for two sentences... and The model needs to predict Is it? The next sentence: .
[0058] (3) Domain Adaptation: CXR-BERT employs domain adaptation technology during pre-training to better adapt to the textual characteristics of CXR reports. Specifically, CXR-BERT uses a medical vocabulary and performs special processing on common medical terms in CXR reports, such as encoding anatomical locations and disease names as a whole, thereby better capturing the semantic information of medical texts.
[0059] Please refer to Figure 3 This is a schematic diagram of the affine transformation guidance module provided in the embodiments of this application.
[0060] Specifically, the input text features first pass through a projection module, which aligns the dimensions of the text features with the dimensions of the image features and reduces the number of text features. The projection process is as follows: ;in, It is a learnable matrix. express Convolutional layer Represents the ReLU activation function. Given input features The output projection features are ,in, It is the number of features after projection. It is the dimension of the projection feature, which is consistent with the dimension of the image feature.
[0061] Subsequently, the mean method is used to obtain a vector that can represent the overall text features. Then, a linear layer and a ReLU activation function are applied to obtain a vector used to modulate the visual features. The overall process is shown below: ;in, It is the ReLU activation function. It is a linear layer. It is the vector obtained by taking the mean. , Number of image feature channels Twice as much.
[0062] Subsequently, the expanded text features are divided into two parts of equal length according to the channel dimension, as shown below: ;in, It is a cutting operation. , .
[0063] Then, for the corresponding visual features The visual feature segmentation is guided by linear affine transformation, as shown below: .
[0064] Finally, after performing a residual upsampling, the result is compared with the features from the encoder. The residue is added together to obtain the output, which is then input into the upper-level guidance module, as shown below: ;in, For upsampling operation, The weighting parameters are for the residuals.
[0065] Compared with related technologies, the embodiments of this application have at least the following advantages: By constructing an image segmentation model, which includes a visual encoder module, a text encoder module, and an affine transformation guidance module, the semantic features of the text auxiliary information of medical ultrasound images can be extracted through the text encoder module, thereby improving the clarity of category semantics in medical ultrasound images. Furthermore, by fusing multi-scale visual features and semantic features through affine transformation, adaptive control of multi-view image structures is achieved, effectively improving the segmentation accuracy and robustness of the image segmentation model in multi-category, complex structure, and multi-angle images. Therefore, after inputting medical ultrasound images into the image segmentation model, a segmentation mask map that enables accurate identification and segmentation of multi-organ regions can be obtained.
[0066] Based on the same idea as the ultrasound image segmentation method in the above embodiments, this application also provides an ultrasound image segmentation apparatus, which can be used to perform the above ultrasound image segmentation method. For ease of explanation, the structural schematic diagram of the ultrasound image segmentation apparatus embodiment only shows the parts related to the embodiments of this application. Those skilled in the art will understand that the illustrated structure does not constitute a limitation on the apparatus, and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0067] like Figure 4 As shown, the ultrasound image segmentation device 40 includes a construction module 401, an acquisition module 402, and an input module 403. In some embodiments, these modules can be programmable software instructions stored in memory and executable by a processor. It is understood that in other embodiments, these modules can also be program instructions or firmware embedded in a processor.
[0068] The construction module 401 is used to construct an image segmentation model, wherein the image segmentation model includes a visual encoder module, a text encoder module, and an affine transformation guidance module; The acquisition module 402 is used to acquire the medical ultrasound image to be segmented and the text auxiliary information representing prior medical knowledge corresponding to the medical ultrasound image; The input module 403 is used to input the medical ultrasound image into the image segmentation model to obtain a segmentation mask of the medical ultrasound image. The segmentation mask is used to segment different organs in the medical ultrasound image into different regions. The visual encoder module is used to extract multi-scale visual features of the medical ultrasound image, the text encoder module is used to extract semantic features of the text auxiliary information, and the affine transformation guidance module is used to fuse the multi-scale visual features and the semantic features through affine transformation to generate the segmentation mask.
[0069] The ultrasound image segmentation device 40 provided in the above embodiments can realize the technical solutions described in the ultrasound image segmentation method embodiments. The specific implementation principles of each module or unit can be found in the corresponding content in the ultrasound image segmentation method embodiments, and will not be repeated here.
[0070] Please refer to point 5. Figure 5 This is a schematic diagram of an embodiment of the electronic device of this application.
[0071] In some embodiments, processor 501 may be a central processing unit (CPU), microprocessor, or other data processing chip, used to run program code stored in memory 502 or process data, such as the ultrasound image segmentation method of the present invention.
[0072] In some embodiments, processor 501 may be a single server or a group of servers. The server group may be centralized or distributed. In some embodiments, processor 501 may be local or remote. In some embodiments, processor 501 may be implemented on a cloud platform. In one embodiment, the cloud platform may include a private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud, multi-cloud, or any combination thereof.
[0073] In some embodiments, memory 502 may be an internal storage unit of electronic device 500, such as a hard disk or memory of electronic device 500. In other embodiments, memory 502 may also be an external storage device of electronic device 500, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on electronic device 500.
[0074] Furthermore, the memory 502 may include both internal storage units of the electronic device 500 and external storage devices. The memory 502 is used to store application software and various types of data installed on the electronic device 500.
[0075] In some embodiments, display 503 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 503 is used to display information from electronic device 500 and to display a visual user interface. Components 501-503 of electronic device 500 communicate with each other via a system bus.
[0076] In one embodiment, when processor 501 executes the ultrasound image segmentation program in memory 502, the following steps can be performed: Construct an image segmentation model, wherein the image segmentation model includes a visual encoder module, a text encoder module, and an affine transformation guidance module; Obtain the medical ultrasound image to be segmented and the corresponding textual auxiliary information representing prior medical knowledge. The medical ultrasound image is input into the image segmentation model to obtain a segmentation mask of the medical ultrasound image. The segmentation mask is used to segment different organs in the medical ultrasound image into different regions. The visual encoder module is used to extract multi-scale visual features from the medical ultrasound image, the text encoder module is used to extract semantic features from the text auxiliary information, and the affine transformation guidance module is used to fuse the multi-scale visual features and the semantic features through affine transformation to generate the segmentation mask.
[0077] It should be understood that when the processor 501 executes the ultrasound image segmentation program in the memory 502, in addition to the functions mentioned above, it can also perform other functions, as detailed in the description of the corresponding method embodiments above.
[0078] Furthermore, this embodiment of the invention does not specifically limit the type of electronic device 500 mentioned. Electronic device 500 can be a mobile phone, tablet computer, personal digital assistant (PDA), wearable device, laptop computer, or other portable electronic device. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices running iOS, Android, Microsoft, or other operating systems. The aforementioned portable electronic device can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the invention, electronic device 500 may not be a portable electronic device, but rather a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0079] Accordingly, this application also provides a storage medium for storing computer-readable programs or instructions. When the programs or instructions are executed by a processor, they can implement the steps or functions of the ultrasound image segmentation methods provided in the above-described method embodiments.
[0080] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.), and the computer program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0081] The above provides a detailed description of the ultrasound image segmentation method, apparatus, electronic device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An ultrasound image segmentation method, characterized in that, include: Construct an image segmentation model, wherein the image segmentation model includes a visual encoder module, a text encoder module, and an affine transformation guidance module; Obtain the medical ultrasound image to be segmented and the corresponding textual auxiliary information representing prior medical knowledge. The medical ultrasound image is input into the image segmentation model to obtain a segmentation mask of the medical ultrasound image. The segmentation mask is used to segment different organs in the medical ultrasound image into different regions. The visual encoder module is used to extract multi-scale visual features from the medical ultrasound image, the text encoder module is used to extract semantic features from the text auxiliary information, and the affine transformation guidance module is used to fuse the multi-scale visual features and the semantic features through affine transformation to generate the segmentation mask.
2. The ultrasound image segmentation method according to claim 1, characterized in that, The affine transformation guidance module includes a projection unit, a modulation unit, and a fusion unit; The step of inputting the medical ultrasound image into the image segmentation model to obtain a segmentation mask of the medical ultrasound image includes: The medical ultrasound image is input into the image segmentation model; The visual encoder module extracts the multi-scale visual features, and the text encoder module extracts the semantic features; The projection unit aligns the multi-scale visual features with the semantic features in terms of dimensions and performs compression and dimensionality reduction. The modulation unit converts the compressed and dimensionality-reduced semantic features into affine parameters and performs affine transformations on the compressed and dimensionality-reduced multi-scale visual features. The fusion unit upsamples the affine parameters and the multi-scale visual features after affine transformation, and then performs skip connections to fuse them with the encoder features to obtain the segmentation mask.
3. The ultrasound image segmentation method according to claim 2, characterized in that, The visual encoder module is a ConvNeXt-tiny convolutional neural network model. The ConvNeXt-tiny convolutional neural network model includes a hierarchical convolutional design for extracting multi-scale features, a large convolutional kernel for capturing global features, and an inverted bottleneck structure for reducing the computational cost of the model.
4. The ultrasound image segmentation method according to claim 2, characterized in that, The text encoder module is a CXR-BERT pre-trained language model; The CXR-BERT pre-trained language model is based on the Transformer neural network architecture and is obtained by pre-training on a historical chest X-ray report dataset; wherein, the Transformer neural network architecture is used to capture long-distance dependencies in the text based on a self-attention mechanism.
5. The ultrasound image segmentation method according to claim 1, characterized in that, The image segmentation model is trained based on a sample dataset, which includes multiple historical multi-organ multi-view medical ultrasound images and historical textual auxiliary information representing prior medical knowledge corresponding to each of the historical multi-organ multi-view medical ultrasound images.
6. The ultrasound image segmentation method according to claim 5, characterized in that, Before training the image segmentation model based on the sample dataset, the method further includes: The sample dataset is preprocessed to obtain the target sample dataset, wherein the preprocessing includes at least image size normalization, image noise suppression, and image enhancement transformation; The image segmentation model is trained based on a sample dataset and includes: The image segmentation model is trained based on the target sample dataset.
7. The ultrasound image segmentation method according to claim 5, characterized in that, The historical multi-organ, multi-view medical ultrasound images include at least: Multi-view medical ultrasound images of thyroid nodules, thyroid glands, breast cancer, fetal head, uterine tumors, left atrium, left ventricle, and left ventricular wall.
8. An ultrasonic image segmentation device, characterized in that, include: Build modules, get modules, and input modules; The building module is used to build an image segmentation model, wherein the image segmentation model includes a visual encoder module, a text encoder module, and an affine transformation guidance module; The acquisition module is used to acquire the medical ultrasound image to be segmented and the corresponding textual auxiliary information representing prior medical knowledge. The input module is used to input the medical ultrasound image into the image segmentation model to obtain a segmentation mask of the medical ultrasound image. The segmentation mask is used to segment different organs in the medical ultrasound image into different regions. The visual encoder module is used to extract multi-scale visual features from the medical ultrasound image, the text encoder module is used to extract semantic features from the text auxiliary information, and the affine transformation guidance module is used to fuse the multi-scale visual features and the semantic features through affine transformation to generate the segmentation mask.
9. An electronic device, the electronic device comprising a processor and a memory, characterized in that, The memory is used to store instructions, and the processor is used to call the instructions in the memory to cause the electronic device to perform the ultrasound image segmentation method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores computer instructions that, when executed on an electronic device, cause the electronic device to perform the ultrasound image segmentation method as described in any one of claims 1 to 7.
Citation Information
Cited By
Medical image segmentation method and device, equipment and medium
CN121564003A