Multi-modal segmentation method, system and equipment for fetal ultrasound image
By fusing text detection and image features in fetal ultrasound image segmentation and using multimodal segmentation method, the problem of difficulty in accurately segmenting fetal ultrasound images in the prior art is solved, and more efficient and accurate image segmentation is achieved.
Patent Information
- Application Number
- CN202411978682.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to accurately capture and segment subtle, highly variable structural features when processing fetal ultrasound images, resulting in unsatisfactory segmentation quality.
The multimodal segmentation method is adopted to detect the fetal ultrasound image text, extract text information, and fuse image features and text features for segmentation, optimize by aligning the feature map, and adjust the initial segmentation boundary.
It improves the accuracy and efficiency of image segmentation, can more accurately capture and segment the subtle structural features in fetal ultrasound images, and improves the segmentation quality.
Smart Images

Figure CN119941650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a multimodal segmentation method, system and device for fetal ultrasound images. Background Art
[0002] In modern medical technology, fetal ultrasound images play an important role in fetal health monitoring. Doctors can detect the fetal morphology, organ development and activity status by observing ultrasound images.
[0003] When using ultrasound images to monitor the health status of different organs or body structures of the fetus, the existing technology usually relies on professionals to manually segment and interpret the ultrasound images or segment the ultrasound images by using automatic image analysis techniques such as threshold-based segmentation methods, edge detection algorithms, region growing algorithms, etc. to obtain the regions of various organ structures of the fetus in the ultrasound image.
[0004] In the process of segmenting fetal ultrasound images using the above-mentioned ultrasound image segmentation methods, ultrasound images are affected by many factors, such as equipment differences, operator skill levels, and patient body shapes, which lead to large variability in ultrasound images. The above-mentioned methods rely on specific features and preset parameters for image segmentation. Therefore, when processing fetal ultrasound images, these methods are difficult to accurately capture and segment subtle, highly variable structural features, resulting in unsatisfactory final segmentation quality. Summary of the invention
[0005] In order to solve the above technical problems, the present invention discloses a multimodal segmentation method, system and device for fetal ultrasound images, which are used to improve the accuracy of image segmentation.
[0006] In order to achieve the above objectives, in a first aspect, the present invention discloses a multimodal segmentation method for fetal ultrasound images, comprising:
[0007] Performing text detection on the fetal ultrasound image obtained after preprocessing to extract text information in the fetal ultrasound image;
[0008] Extracting a first image feature corresponding to the fetal ultrasound image and fusing the image feature with the text feature to segment the fetal ultrasound image to obtain a segmented image including an initial segmentation boundary;
[0009] Extracting a second image feature corresponding to the segmented image, and spatially aligning the second image feature and the text feature to generate a feature map corresponding to the segmented image;
[0010] The weight of each feature point on the feature map is calculated to adjust the initial segmentation boundary according to the weight to obtain an optimized segmented image.
[0011] The present invention discloses a multimodal segmentation method for fetal ultrasound images. When segmenting a fetal ultrasound image, text recognition is first performed on the fetal ultrasound image to extract text information in the fetal ultrasound image, so as to use the text information to enhance the accuracy of subsequent image segmentation. When performing preliminary image segmentation, image segmentation is performed based on the extracted text information and the first image feature of the fetal ultrasound image to improve the efficiency and accuracy of segmentation. Then, after obtaining a segmented image containing an initial segmentation boundary, the image features of the segmented image are extracted, and the segmentation boundary of the segmented image is optimized by aligning the image features and the text features, thereby further improving the accuracy of image segmentation.
[0012] As a preferred example, the performing text detection on the fetal ultrasound image obtained after preprocessing to extract text features in the fetal ultrasound image includes:
[0013] Performing text location on the fetal ultrasound image by using a preset text detection algorithm to obtain a plurality of text regions containing text in the fetal ultrasound image and a plurality of characters in each of the text regions;
[0014] According to the text area and the characters, the text features corresponding to each of the text areas are extracted through a pre-trained language recognition model, so as to output the structured text features corresponding to each of the text areas according to the text features; wherein the structured text features include the text content in each of the text areas, the position of the text area in the image, and the text type of the text content.
[0015] In the above scheme, by detecting the text content in the fetal ultrasound image, the text semantic information for segmentation is provided for subsequent image segmentation, thereby improving the accuracy of image segmentation.
[0016] As a preferred example, the extracting the first image feature corresponding to the fetal ultrasound image and fusing the image feature and the text feature includes:
[0017] Inputting the fetal ultrasound image into a trained image segmentation model to extract image features corresponding to the fetal ultrasound image;
[0018] The image features and the structured text features are fused by a self-attention mechanism preset in the image segmentation model to obtain a joint feature map corresponding to the fetal ultrasound image.
[0019] In the above scheme, the model is used for image segmentation, and image segmentation can be adaptively performed according to the characteristics of the image itself. It can capture subtle image features in the image, thereby improving the effect of image segmentation. Furthermore, when the model is used for image segmentation, the text information in the ultrasound image is matched one-to-one with the image content in combination with the text features and image features, which can improve the accuracy of image segmentation.
[0020] As a preferred example, the segmenting of the fetal ultrasound image to obtain a segmented image containing an initial segmentation boundary includes:
[0021] Processing the joint feature map based on a contraction path preset in the image segmentation model to obtain a plurality of feature maps corresponding to the joint feature map at different resolutions;
[0022] Calculating the attention weight of each position in the feature map, weighting the feature map according to the weight, and generating a weighted feature map corresponding to each resolution;
[0023] The weighted feature maps corresponding to different resolutions are spliced by skip connection, so as to generate a segmentation mask corresponding to the fetal ultrasound image according to the spliced reshaped feature maps;
[0024] The fetal ultrasound image is segmented according to the segmentation mask to generate a segmented image including an initial segmentation boundary.
[0025] In the above scheme, when the model uses the fused features to segment the image, the feature performance of the fused features at different levels is extracted through the contraction path and the jump connection method, thereby improving the accuracy of feature extraction, and then using the extracted fine features to segment the image to ensure the accuracy of the segmented image.
[0026] As a preferred example, the extracting the second image feature corresponding to the segmented image includes:
[0027] Acquire a plurality of key text contents in the fetal image and a key text region of each key text content in the segmented image according to the structured text feature;
[0028] Matching a key image region corresponding to the key text region from the segmented image;
[0029] The key image region is input into a trained image segmentation model to extract a second image feature corresponding to the key image region.
[0030] In the above scheme, the key text area is first extracted from the text content and the key image area where the key text area is located is matched, so as to subsequently optimize the segmentation boundary of the key area in the segmented image by corresponding to the key text area and the key image area, thereby improving the accuracy of image segmentation.
[0031] As a preferred example, the step of spatially aligning the second image feature and the text feature to generate a feature map corresponding to the segmented image includes:
[0032] Extracting key text features corresponding to the key text area, and mapping the key text features and the second image features to the same feature space to obtain mapped text features and mapped image features;
[0033] The mapped text features and the mapped image features are aligned through geometric transformation to generate a feature map corresponding to the key image area.
[0034] In the above scheme, by mapping and aligning the corresponding text content and the features of the key image, the boundaries of the key areas in the segmented image are further optimized according to the text content, thereby improving the accuracy of image segmentation.
[0035] As a preferred example, the calculating the weight of each feature point on the feature map to adjust the initial segmentation boundary according to the weight includes:
[0036] Calculate the attention weight of each feature point on the feature map corresponding to the key image area based on the attention mechanism;
[0037] The feature map is weighted according to the attention weight, and an initial segmentation boundary of the key image area is adjusted according to the weighted feature map.
[0038] In the above scheme, the attention mechanism is introduced to calculate the attention weight of each point on the feature map after feature alignment, highlight the important features, and apply the weights to the original features to generate weighted feature representations, which not only improves the accuracy of image segmentation, but also enhances the attention to key areas and improves the effect of image segmentation.
[0039] In a second aspect, the present invention discloses a multimodal segmentation system for fetal ultrasound images, including a text detection module, an image segmentation module, a multimodal fusion module and a segmentation optimization module;
[0040] The text detection module is used to perform text detection on the fetal ultrasound image obtained after preprocessing to extract text features in the fetal ultrasound image;
[0041] The image segmentation module is used to extract the first image feature corresponding to the fetal ultrasound image and fuse the image feature and the text feature to segment the fetal ultrasound image to obtain a segmented image containing an initial segmentation boundary;
[0042] The multimodal fusion module is used to extract a second image feature corresponding to the segmented image, and align the second image feature and the text feature in space to generate a feature map corresponding to the segmented image;
[0043] The segmentation optimization module is used to calculate the weight of each feature point on the feature map, so as to adjust the initial segmentation boundary according to the weight and obtain an optimized segmented image.
[0044] The present invention discloses a multimodal segmentation system for fetal ultrasound images. When segmenting a fetal ultrasound image, text recognition is first performed on the fetal ultrasound image to extract text information in the fetal ultrasound image, so as to use the text information to enhance the accuracy of subsequent image segmentation. When performing preliminary image segmentation, image segmentation is performed based on the extracted text information and the first image feature of the fetal ultrasound image to improve the efficiency and accuracy of segmentation. Then, after obtaining a segmented image containing an initial segmentation boundary, the image features of the segmented image are extracted, and the segmentation boundary of the segmented image is optimized by aligning the image features and the text features, thereby further improving the accuracy of image segmentation.
[0045] As a preferred example, the text detection module includes a text recognition unit and a feature extraction unit;
[0046] The text recognition unit is used to locate text on the fetal ultrasound image by using a preset text detection algorithm to obtain a plurality of text regions containing text in the fetal ultrasound image and a plurality of characters in each of the text regions;
[0047] The feature extraction unit is used to extract text features corresponding to each text area through a pre-trained language recognition model based on the text area and the characters, so as to output structured text features corresponding to each text area based on the text features; wherein the structured text features include text content in each text area, the position of the text area in the image, and the text type of the text content.
[0048] In the above scheme, by detecting the text content in the fetal ultrasound image, the text semantic information for segmentation is provided for subsequent image segmentation, thereby improving the accuracy of image segmentation.
[0049] In a third aspect, the present invention discloses an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement a multimodal segmentation method for fetal ultrasound images as described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 : is a flow chart of a multimodal segmentation method for fetal ultrasound images disclosed in an embodiment of the present invention;
[0051] Figure 2 : is a schematic diagram of the structure of a multimodal segmentation system for fetal ultrasound images disclosed in an embodiment of the present invention;
[0052] Figure 3 : A flowchart of a multimodal segmentation method for fetal ultrasound images disclosed in another embodiment of the present invention;
[0053] Figure 4 : A schematic diagram of a model training process disclosed in another embodiment of the present invention;
[0054] Figure 5 :A schematic diagram of a text recognition process disclosed in another embodiment of the present invention:
[0055] Figure 6 : A schematic diagram of an image segmentation process disclosed in another embodiment of the present invention;
[0056] Figure 7 : A schematic diagram of an image segmentation optimization process disclosed in another embodiment of the present invention. DETAILED DESCRIPTION
[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] Embodiment 1
[0059] Reference Figure 1 , is a multimodal segmentation method for fetal ultrasound images disclosed in an embodiment of the present invention, which is used to improve the accuracy of image segmentation. Specifically, please refer to the specific implementation process of the multimodal segmentation method. Figure 1 , mainly including step 101 to step 104, the steps are:
[0060] Step 101: Perform text detection on the fetal ultrasound image obtained after preprocessing to extract text features in the fetal ultrasound image.
[0061] Step 102: extracting a first image feature corresponding to the fetal ultrasound image and fusing the image feature with the text feature to segment the fetal ultrasound image to obtain a segmented image including an initial segmentation boundary.
[0062] Step 103: extracting a second image feature corresponding to the segmented image, and aligning the second image feature and the text feature in space to generate a feature map corresponding to the segmented image.
[0063] Step 104: Calculate the weight of each feature point on the feature map to adjust the initial segmentation boundary according to the weight to obtain an optimized segmented image.
[0064] In the disclosed embodiment, after obtaining a fetal ultrasound image through ultrasound scanning, in order to obtain the health status of the fetal organs, body structures, etc. according to the fetal ultrasound image, it is necessary to perform image segmentation on the fetal ultrasound image to generate a segmentation boundary corresponding to the fetal ultrasound image, and then perform health detection on the corresponding structural image in each of the thousands of separate areas enclosed by the segmentation boundary. Among them, when generating a fetal ultrasound image based on the existing device for generating ultrasound images, it will receive user input or automatically mark each area in the ultrasound image according to the scanning process, so as to facilitate the distinction of different areas of the image according to the marking. In this regard, when performing image segmentation, text detection can be performed on the received fetal ultrasound image first, and the text information in the ultrasound image can be identified, so as to perform image segmentation according to the text information, thereby improving the accuracy of image segmentation.
[0065] Specifically, when extracting text information from the fetal ultrasound image, step 101 may extract text features of the image, and then identify corresponding text information according to the text features. Preferably, step 101 may identify text features through the following steps:
[0066] Step 1011: locating text in the fetal ultrasound image using a preset text detection algorithm to obtain a plurality of text regions containing text in the fetal ultrasound image and a plurality of characters in each of the text regions;
[0067] Step 1012 extracts text features corresponding to each text area through a pre-trained language recognition model based on the text area and the characters, so as to output structured text features corresponding to each text area based on the text features; wherein the structured text features include the text content in each text area, the position of the text area in the image, and the text type of the text content.
[0068] In the disclosed embodiment, the above steps detect the text content in the fetal ultrasound image to provide text semantic information for subsequent image segmentation, thereby improving the accuracy of image segmentation.
[0069] Furthermore, after extracting the corresponding text features, in order to adaptively perform image segmentation on different ultrasound images, a model can be constructed based on deep learning to adaptively learn the subtle structural features in different images, and then use the trained model to perform image segmentation. Specifically, in the process of using the model to perform image segmentation, step 102 can perform image segmentation by fusing the image features and the text features through the following steps, preferably:
[0070] Step 1021: inputting the fetal ultrasound image into a trained image segmentation model to extract image features corresponding to the fetal ultrasound image;
[0071] Step 1022: The image features and the structured text features are fused through a self-attention mechanism preset in the image segmentation model to obtain a joint feature map corresponding to the fetal ultrasound image.
[0072] In the embodiment of the present disclosure, the above steps utilize the model to perform image segmentation, and can adaptively perform image segmentation based on the characteristics of the image itself, which can capture subtle image features in the image, thereby improving the image segmentation effect. Furthermore, when utilizing the model to perform image segmentation, the text information in the ultrasound image is matched one-to-one with the image content in combination with the text features and the image features, thereby improving the accuracy of image segmentation.
[0073] Specifically, when performing image segmentation by using the fused features in step 1022, in order to ensure the accuracy of fusion, the step 1022 performs image segmentation by using the following steps, including:
[0074] Step 10221: Processing the joint feature map based on a contraction path preset in the image segmentation model to obtain a plurality of feature maps corresponding to the joint feature map at different resolutions;
[0075] Step 10222: Calculate the attention weight of each position in the feature map, weight the feature map according to the weight, and generate a weighted feature map corresponding to each resolution;
[0076] Step 10223: splicing weighted feature maps corresponding to different resolutions by means of skip connections, so as to generate a segmentation mask corresponding to the fetal ultrasound image according to the spliced reshaped feature maps;
[0077] Step 10224: Segment the fetal ultrasound image according to the segmentation mask to generate a segmented image including an initial segmentation boundary.
[0078] In the embodiment of the present disclosure, when the above steps use the fused features to segment the image, the feature performance of the fused features at different levels is extracted through the contraction path and the jump connection method, thereby improving the accuracy of feature extraction, and then using the extracted fine features to segment the image to ensure the accuracy of the segmented image.
[0079] Furthermore, after the segmented image including the initial segmentation boundary is outputted by the image segmentation model, in order to further improve the accuracy of image segmentation, the segmented image can be further optimized by matching the text content and the segmentation result in the segmented image. Specifically, step 103 matches the text with the image by extracting the image features corresponding to the segmented image and the text features of the text content. Step 103 mainly includes:
[0080] Step 1031: acquiring a plurality of key text contents in the fetal image and a key text region of each key text content in the segmented image according to the structured text feature;
[0081] Step 1032: matching the key image area corresponding to the key text area from the segmented image;
[0082] Step 1033: inputting the key image region into a trained image segmentation model to extract a second image feature corresponding to the key image region;
[0083] Step 1034: extracting key text features corresponding to the key text area, and mapping the key text features and the second image features to the same feature space to obtain mapped text features and mapped image features;
[0084] Step 1035: Align the mapped text features and the mapped image features through geometric transformation to generate a feature map corresponding to the key image area.
[0085] In the disclosed embodiment, the above steps first extract the key text area from the text content and match the key image area where the key text area is located, so as to subsequently optimize the segmentation boundaries of the key areas in the segmented image by corresponding to the key text area and the key image area, thereby improving the accuracy of image segmentation. Furthermore, by mapping and aligning the features of the corresponding text content and the key image, the boundaries of the key areas in the segmented image are further optimized according to the text content, thereby improving the accuracy of image segmentation.
[0086] Further, after aligning the features of the text content and the image, the initial segmentation boundary is optimized according to the feature map generated after the alignment. Specifically, the step 104 optimizes the segmentation boundary through the following steps, including:
[0087] Step 1041: Calculate the attention weight of each feature point on the feature map corresponding to the key image area based on the attention mechanism;
[0088] Step 1042: weighting the feature map according to the attention weight, and adjusting the initial segmentation boundary of the key image area according to the weighted feature map.
[0089] In the disclosed embodiment, the above steps introduce an attention mechanism, calculate the attention weight of each point on the feature map after feature alignment, highlight important features, and apply the weights to the original features to generate weighted feature representations, which not only improves the accuracy of image segmentation, but also enhances the attention to key areas, thereby improving the effect of image segmentation.
[0090] On the other hand, this embodiment also discloses a multimodal segmentation system for fetal ultrasound images. For the specific structure of the segmentation system, please refer to Figure 2 , including a text detection module 201, an image segmentation module 202, a multimodal fusion module 203 and a segmentation optimization module 204.
[0091] The text detection module 201 is used to perform text detection on the fetal ultrasound image obtained after preprocessing, so as to extract text features in the fetal ultrasound image.
[0092] The image segmentation module 202 is used to extract the first image feature corresponding to the fetal ultrasound image and fuse the image feature and the text feature to segment the fetal ultrasound image to obtain a segmented image containing an initial segmentation boundary.
[0093] The multimodal fusion module 203 is used to extract the second image features corresponding to the segmented image, and align the second image features and the text features in space to generate a feature map corresponding to the segmented image.
[0094] The segmentation optimization module 204 is used to calculate the weight of each feature point on the feature map, so as to adjust the initial segmentation boundary according to the weight and obtain an optimized segmented image.
[0095] In this embodiment, the text detection module 201 includes a text recognition unit and a feature extraction unit.
[0096] The text recognition unit is used to locate text in the fetal ultrasound image by using a preset text detection algorithm to obtain a plurality of text regions containing text in the fetal ultrasound image and a plurality of characters in each of the text regions.
[0097] The feature extraction unit is used to extract text features corresponding to each text area through a pre-trained language recognition model based on the text area and the characters, so as to output structured text features corresponding to each text area based on the text features; wherein the structured text features include text content in each text area, the position of the text area in the image, and the text type of the text content.
[0098] According to Figure 1 A multimodal segmentation method for fetal ultrasound images is shown. Accordingly, this embodiment provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the following is implemented Figure 1The multimodal segmentation method for fetal ultrasound images shown in the figure. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal, and various interfaces and lines are used to connect various parts of the entire terminal. The memory may be used to store the computer program, and the processor implements various functions of the terminal by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Med ia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0099] Compared with the prior art, the multimodal segmentation method, system and device for fetal ultrasound images provided in this embodiment have the following beneficial effects:
[0100] The present embodiment provides a multimodal segmentation method, system and device for fetal ultrasound images. When segmenting a fetal ultrasound image, text recognition is first performed on the fetal ultrasound image to extract text information in the fetal ultrasound image, so as to use the text information to enhance the accuracy of subsequent image segmentation. When performing preliminary image segmentation, image segmentation is performed based on the extracted text information and the first image feature of the fetal ultrasound image to improve the efficiency and accuracy of segmentation. Then, after obtaining a segmented image containing an initial segmentation boundary, the image features of the segmented image are extracted, and the segmentation boundary of the segmented image is optimized by aligning the image features and the text features, so as to further improve the accuracy of image segmentation.
[0101] Embodiment 2
[0102] In the prior art, when segmenting fetal ultrasound images to detect the development of fetal organs, structures, etc., ultrasound images are often segmented by manual segmentation or automatic image analysis technology. Among them, when using manual segmentation to segment ultrasound images, not only is the efficiency of image segmentation relatively low due to the professional knowledge of image segmentation personnel, but the accuracy of manual segmentation is also low. When using automatic image analysis technology to segment images, automatic image analysis technology relies on specific features and preset parameters and cannot use ultrasound images with extremely strong variability, resulting in low image segmentation effects.
[0103] In this regard, in order to solve the technical problems caused by the above-mentioned image segmentation method, this embodiment discloses a multimodal segmentation method for fetal ultrasound images, which uses deep learning to build an image segmentation model to adaptively extract image features of different ultrasound images for image segmentation, thereby improving the segmentation effect. At the same time, the text information in the image is combined to provide a basis for image segmentation and optimize the segmented image to improve the accuracy of image segmentation. Specifically, please refer to the specific implementation process of the multimodal segmentation method disclosed in this embodiment. Figure 3 , mainly including steps 301 to 304:
[0104] Step 301: Train the pre-constructed deep learning model based on the pre-constructed data training set to obtain an image segmentation model and a text recognition model.
[0105] Specifically, in certain implementations of this embodiment, ultrasound image data is first acquired, and then a training data set is constructed to train the constructed deep learning model to generate a text recognition model and an image segmentation model for text recognition.
[0106] Preferably, in some implementations of this embodiment, the training process of the text recognition model and the image segmentation model can refer to Figure 4 .like Figure 4As shown, ultrasound image data is first collected, wherein a number of historical fetal ultrasound images, a segmented image containing a segmentation boundary corresponding to each of the historical fetal ultrasound images, and text information corresponding to each of the historical fetal ultrasound images can be extracted from a database storing ultrasound images or an existing image database. The text information includes the text content in the historical fetal ultrasound image, the text area of the text content in the image, the text type (such as title, paragraph, comment, etc.), special marks and symbols, and possible semantic tags.
[0107] Furthermore, in the process of model training, in order to ensure the efficiency and accuracy of training, each of the historical fetal ultrasound images can be preprocessed, and the preprocessing includes necessary processing of the historical fetal ultrasound images, such as image size unification, contrast adjustment, denoising, etc., to optimize the input quality of subsequent models, and to convert and enhance the images by using advanced image processing algorithms and tools (such as OpenCV, PI L).
[0108] Then, each of the pre-processed historical fetal ultrasound images is input into an initial text recognition model constructed based on a deep learning network to obtain initial text information output by the initial text recognition model. The initial text recognition model can be constructed using an algorithm specifically used for image text detection, such as an OCR technology based on deep learning and a CTPN optical character detection algorithm.
[0109] Then, during the training process, the loss of the initial text information and the corresponding text information in the training set is obtained to adjust the model parameters of the initial text recognition model according to the loss until the loss is less than a preset loss threshold, and then the current model parameters are retained to obtain the text recognition model.
[0110] When training the image segmentation model, each of the preprocessed historical fetal ultrasound images and the text information is input into the initial image segmentation model constructed based on the deep learning network to obtain the initial segmented image output by the initial image segmentation model. In order to improve the accuracy of fetal ultrasound image segmentation, the parameters of deep learning networks specifically used for medical image segmentation, such as U-NET network, FCN network, DeepLab v3+ network, etc., can be used to construct the initial segmented image.
[0111] During the training process, the loss of the initial segmented image and the corresponding segmented image in the training set is obtained to adjust the model parameters of the initial image segmentation model according to the loss until the loss is less than a preset loss threshold, and then the current model parameters are retained to obtain the image segmentation model.
[0112] Step 302: pre-processing the fetal ultrasound image to be segmented, and inputting the pre-processed fetal ultrasound image into the text recognition model to extract structured text features of the fetal ultrasound image.
[0113] Specifically, in this embodiment, the image is preprocessed during the model training process. In this embodiment, when the structured text features of the fetal ultrasound image are extracted by the trained text recognition model, the fetal ultrasound image needs to be preprocessed in the same manner, including necessary processing of the fetal ultrasound image, including image size unification, contrast adjustment, denoising, etc., to optimize the input quality of the subsequent model. The image is converted and enhanced by using advanced image processing algorithms and tools (such as OpenCV, PIL).
[0114] Furthermore, the process of inputting the preprocessed fetal ultrasound image into the text recognition model to output the structured text features of the fetal ultrasound image through the text recognition model can refer to Figure 5 .
[0115] like Figure 5 As shown, after the image is input into the text recognition model, the image is first processed by the deep convolutional neural network in the text recognition model to identify the area containing the text. These text areas are further refined and segmented into multiple independent text blocks. At the same time, the convolutional neural network is used to segment and recognize the characters in each text block one by one to ensure that each character can be accurately recognized and classified.
[0116] Next, in the text recognition model process, since the data set used to train the model is text data in the field of fetal ultrasound images, these text data enable the model to better understand and recognize medical terms and text. Therefore, after locating the text in the image, the text recognition model processes these data and extracts the features of the text.
[0117] Furthermore, when extracting features of the text, the text recognition model can optimize the extracted features through a multi-step correction process. First, errors that may occur during the recognition process are corrected through spelling checking and grammar correction. Then, the text recognition model uses contextual information to further optimize the recognition results to ensure that the output text is accurate. Finally, the text recognition model converts the extracted features into structured text features. These structured features include the recognized text content, the location of the text in the image, the text type (such as title, paragraph, comment, etc.), special marks and symbols, and possible semantic tags.
[0118] The structured text features provide a solid foundation for subsequent text analysis and information extraction, allowing users to further process and utilize the extracted text information. For example, in the application scenario of medical images, reports can be automatically generated based on the recognized text content, key information can be extracted, and even doctors can be assisted in diagnosis and decision-making. Through these detailed and accurate processing steps.
[0119] Step 303: Input the structured text features and the preprocessed fetal ultrasound image into the trained image segmentation model to output an initial segmentation image containing an initial segmentation boundary.
[0120] Specifically, in this embodiment, when performing image segmentation, the fetal ultrasound image and the structured text features are first loaded into the image segmentation model to extract the image features of the fetal ultrasound image, and the image features and the structured text features are fused, and then the image is segmented according to the fused features to generate a segmented image containing an initial segmentation boundary.
[0121] Preferably, the process of fusing image features and text features to perform image segmentation by the image segmentation model can refer to Figure 6 .like Figure 6 As shown, the image segmentation model first uses a convolutional neural network to perform multi-level feature extraction on the input fetal image to generate image features of the fetal ultrasound image.
[0122] Secondly, the text features extracted by the text recognition model are deeply fused with the extracted image features. The fusion of the image features and the text features can use self-attention or fully connected layers to integrate the feature representations of the image and text, so as to combine the information such as the text content, text position, and text type included in the text features with the image features, and generate joint features including image features and text features corresponding to the fetal ultrasound image, so as to utilize the context in the text information and the context of the pixels in the image to enhance the accuracy and robustness of the image segmentation. For example, if the text information mentions specific anatomical structures, the model can pay special attention to these areas, thereby improving the accuracy of the segmentation.
[0123] Next, refer to Figure 6 , the fused features are input into the contraction path in the image segmentation model, and the contraction path gradually extracts the feature maps of the joint features at different levels through a series of convolutional layers and pooling layers. Each convolution operation in the contraction path extracts features, and reduces the size of the feature map through the pooling layer to obtain feature representations of different resolutions, thereby generating feature maps corresponding to different resolutions.
[0124] Furthermore, the feature maps of different resolutions are transmitted to the attention module in the image segmentation model to enhance the attention to important areas in the segmentation task by calculating the attention weight of each position in the feature map. The attention mechanism can be channel attention or spatial attention. By weighting different parts of the feature map, the attention module calculates the weight of each feature point to achieve weighting of features at different positions, usually using channel attention or spatial attention mechanism. Channel attention evaluates the importance of each channel by aggregating the spatial information of the feature map, while spatial attention directly evaluates the importance of each pixel position in the feature map. These weights are then applied to the original feature map to generate a weighted feature representation to highlight the key areas in the image. For each input feature map, the attention module outputs a corresponding weighted feature map, which reflects the model's attention to different areas in the image, and is then used to improve the accuracy of the segmentation task.
[0125] Furthermore, each weighted feature map is upsampled and feature concatenated to restore the feature map to the size of the original image while retaining key contextual information, so that the multiple feature maps output by the contraction path are integrated through the feature reshaping process. In feature reshaping, these multiple feature maps are not simply reshaped into one, but through upsampling and feature concatenation, the feature map is restored to the size of the original image while retaining key contextual information. Specifically, this process involves gradually upsampling the low-resolution feature map of the contraction path and concatenating it with the high-resolution feature map from the previous layer through skip connections to enrich the feature representation and retain details. The skip connection feature is obtained by directly connecting feature maps of the same resolution between the encoder and decoder of the U-Net model. This allows the detailed information in the encoder to be directly utilized in the decoder, thereby improving the accuracy and robustness of the segmentation.
[0126] Finally, the feature map after feature reshaping is input into the expansion path in the image segmentation model to restore the reshaped feature map to the size of the original image through step-by-step upsampling and convolution operations. After each upsampling, the skip connection features from the contraction path are combined to ensure that more details and edge information are retained during the reconstruction process, and finally a high-quality segmentation mask is generated.
[0127] The final segmentation results can accurately locate and identify the key structures and features in the image through the mask reconstructed by the expansion path. These segmentation results not only include the precise positioning and segmentation of the target area, but also clearly identify the key structures and features in the image, providing a basis for subsequent analysis and processing, thereby achieving efficient and accurate segmentation of medical images.
[0128] Step 304: extracting image features corresponding to the initial segmented image, optimizing the initial segmented boundary by aligning the image features with the text features, and outputting an optimized segmented image.
[0129] Specifically, in this embodiment, a plurality of key text contents in the fetal image and a key text region of each key text content in the segmented image are obtained according to the text features; a key image region corresponding to the key text region is matched from the segmented image. The initial segmentation boundary is optimized by aligning the text features corresponding to the key text region and the image features corresponding to the key image region.
[0130] Preferably, the process of segmentation image optimization can refer to Figure 7 .like Figure 7 As shown, firstly, the key text area is extracted according to the text features, and the key image area corresponding to the key text area is extracted from the segmented ultrasound image. Then, the text features of the key text area and the image features of the key image area are extracted using the text recognition model and the image segmentation model, and these features are mapped to the same feature space using a convolutional neural network (CNN) and a fully connected layer. Feature alignment ensures that the text features and image features correspond precisely in space through geometric transformation.
[0131] Next, the attention mechanism is introduced to calculate the attention weight of each feature point, highlight the important features, and apply the weight to the original features through the dot product attention or additive attention method to generate a weighted feature representation. Finally, the segmented image is optimized according to the weighted feature representation.
[0132] Furthermore, after image optimization, an interactive interface can be provided to users based on the constructed user feedback mechanism so that users can evaluate and adjust the segmentation results. This ensures the accuracy and usability of the final output. Among them, the user feedback mechanism mainly enables users to intuitively view the segmentation results through the interface and make necessary adjustments, such as marking erroneous areas or adjusting segmentation boundaries. The final segmentation results and text information will be displayed and output, which can be directly used for clinical diagnosis or as part of medical records. In addition, the results can be verified, that is, by comparing with the annotation results of experts, the accuracy and consistency of the segmentation results can be ensured.
[0133] A multimodal segmentation method for fetal ultrasound images disclosed in this embodiment adopts multimodal fusion technology, combines the recognized text information with the image segmentation capability, and can use the rich context contained in the text information to guide the image segmentation process, thereby improving the accuracy and robustness of the segmentation. Compared with the prior art, this method not only improves the segmentation quality, but also can handle more complex image scenes, especially in fetal ultrasound images, it can more accurately identify and locate the organs and structures of the fetus. Further, the powerful feature extraction and learning capabilities of the deep learning model are used to significantly improve the speed and accuracy of image segmentation. Compared with traditional image processing methods that rely on manual feature extraction, the deep learning model can automatically learn complex patterns in the image, so that when processing a large amount of data, it can quickly provide high-quality segmentation results, ensuring rapid response and high-accuracy output when facing large-scale data sets. Finally, when the results are output, the user experience is fully considered, a feedback mechanism is provided, and the dependence on professional and technical personnel is reduced, so that more medical workers can directly use the present invention to perform effective medical image analysis, improve work efficiency, and also reduce the complexity and error rate of operation.
[0134] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multimodal segmentation method for fetal ultrasound images, characterized in that: include: Performing text detection on the fetal ultrasound image obtained after preprocessing to extract text features in the fetal ultrasound image; Extracting a first image feature corresponding to the fetal ultrasound image and fusing the image feature with the text feature to segment the fetal ultrasound image to obtain a segmented image including an initial segmentation boundary; Extracting a second image feature corresponding to the segmented image, and spatially aligning the second image feature and the text feature to generate a feature map corresponding to the segmented image; The weight of each feature point on the feature map is calculated to adjust the initial segmentation boundary according to the weight to obtain an optimized segmented image.
2. A multimodal segmentation method for fetal ultrasound images according to claim 1, characterized in that: The performing text detection on the fetal ultrasound image obtained after preprocessing to extract text features in the fetal ultrasound image includes: Performing text location on the fetal ultrasound image by using a preset text detection algorithm to obtain a plurality of text regions containing text in the fetal ultrasound image and a plurality of characters in each of the text regions; According to the text area and the characters, the text features corresponding to each of the text areas are extracted through a pre-trained language recognition model, so as to output the structured text features corresponding to each of the text areas according to the text features; wherein the structured text features include the text content in each of the text areas, the position of the text area in the image, and the text type of the text content.
3. A multimodal segmentation method for fetal ultrasound images according to claim 2, characterized in that: The step of extracting the first image feature corresponding to the fetal ultrasound image and fusing the image feature and the text feature includes: Inputting the fetal ultrasound image into a trained image segmentation model to extract image features corresponding to the fetal ultrasound image; The image features and the structured text features are fused by a self-attention mechanism preset in the image segmentation model to obtain a joint feature map corresponding to the fetal ultrasound image.
4. The multimodal segmentation method for fetal ultrasound images according to claim 3, characterized in that: The step of segmenting the fetal ultrasound image to obtain a segmented image including an initial segmentation boundary includes: Processing the joint feature map based on a contraction path preset in the image segmentation model to obtain a plurality of feature maps corresponding to the joint feature map at different resolutions; Calculating the attention weight of each position in the feature map, weighting the feature map according to the weight, and generating a weighted feature map corresponding to each resolution; The weighted feature maps corresponding to different resolutions are spliced by skip connection, so as to generate a segmentation mask corresponding to the fetal ultrasound image according to the spliced reshaped feature maps; The fetal ultrasound image is segmented according to the segmentation mask to generate a segmented image including an initial segmentation boundary.
5. A multimodal segmentation method for fetal ultrasound images according to any one of claims 2-3, characterized in that: The extracting the second image feature corresponding to the segmented image includes: Acquire a plurality of key text contents in the fetal image and a key text region of each key text content in the segmented image according to the structured text feature; Matching a key image region corresponding to the key text region from the segmented image; The key image region is input into a trained image segmentation model to extract a second image feature corresponding to the key image region.
6. The multimodal segmentation method for fetal ultrasound images according to claim 5, characterized in that: The step of spatially aligning the second image feature and the text feature to generate a feature map corresponding to the segmented image includes: Extracting key text features corresponding to the key text area, and mapping the key text features and the second image features to the same feature space to obtain mapped text features and mapped image features; The mapped text features and the mapped image features are aligned through geometric transformation to generate a feature map corresponding to the key image area.
7. A multimodal segmentation method for fetal ultrasound images according to claim 6, characterized in that: The calculating the weight of each feature point on the feature map to adjust the initial segmentation boundary according to the weight includes: Calculate the attention weight of each feature point on the feature map corresponding to the key image area based on the attention mechanism; The feature map is weighted according to the attention weight, and an initial segmentation boundary of the key image area is adjusted according to the weighted feature map.
8. A multimodal segmentation system for fetal ultrasound images, characterized in that: Including text detection module, image segmentation module, multimodal fusion module and segmentation optimization module; The text detection module is used to perform text detection on the fetal ultrasound image obtained after preprocessing to extract text features in the fetal ultrasound image; The image segmentation module is used to extract the first image feature corresponding to the fetal ultrasound image and fuse the image feature and the text feature to segment the fetal ultrasound image to obtain a segmented image containing an initial segmentation boundary; The multimodal fusion module is used to extract a second image feature corresponding to the segmented image, and align the second image feature and the text feature in space to generate a feature map corresponding to the segmented image; The segmentation optimization module is used to calculate the weight of each feature point on the feature map, so as to adjust the initial segmentation boundary according to the weight and obtain an optimized segmented image.
9. The multimodal segmentation system for fetal ultrasound images according to claim 8, characterized in that: The text detection module includes a text recognition unit and a feature extraction unit; The text recognition unit is used to locate text on the fetal ultrasound image by using a preset text detection algorithm to obtain a plurality of text regions containing text in the fetal ultrasound image and a plurality of characters in each of the text regions; The feature extraction unit is used to extract text features corresponding to each text area through a pre-trained language recognition model based on the text area and the characters, so as to output structured text features corresponding to each text area based on the text features; wherein the structured text features include text content in each text area, the position of the text area in the image, and the text type of the text content.
10. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement a multimodal segmentation method for fetal ultrasound images as described in any one of claims 1 to 7.
Citation Information
Cited By
Placenta implantation evaluation system based on multi-modal sign recognition and construction method and construction device thereof
CN120147758A