Standard frame identification and multi-target segmentation method and equipment in fetal ultrasonic video
By combining a diffusion model with a pyramidal vision Transformer, standard frames are extracted from ultrasound videos and multi-target segmentation is performed. This solves the problems of inaccurate detection of small targets and blurred boundaries in fetal ultrasound images, achieving higher segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202511127797.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-07
AI Technical Summary
Existing fetal ultrasound image segmentation methods suffer from problems such as inaccurate detection of small targets, blurred boundaries, and severe background noise interference when processing complex ultrasound images, which affect diagnostic accuracy.
We adopt a method that combines a diffusion model with a pyramid visual Transformer. By extracting standard frames from ultrasound videos and performing multi-target segmentation, we use a pyramid visual Transformer feature extraction network to extract multi-scale semantic features. We combine a diffusion probability model-based denoising encoder and a cross-attention fusion module to enhance feature fusion capability and anti-interference capability.
It improves the ability to detect small targets, alleviates the problem of blurred boundaries, enhances the model's anti-interference ability, and improves the accuracy and robustness of fetal ultrasound image segmentation.
Smart Images

Figure CN120912889A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of deep learning and medical image processing, in particular to fetal ultrasound image standard body position section recognition and fetal early pregnancy ultrasound image automatic analysis technology, and particularly to a standard body position recognition method combining diffusion model and Transformer and a multi-target segmentation method, belonging to the application of deep learning in the medical field. BACKGROUND
[0002] Medical images play an important role in clinical medicine. Ultrasound has become the main tool for prenatal imaging diagnosis due to its excellent performance and advantages such as small radiation, real-time display, and low price. It assesses the growth status of the fetus by imaging the fetus and its accessories, discovers congenital defects, and assists clinicians in diagnosis. In particular, in the early pregnancy stage of the fetus, accurate segmentation and recognition of key anatomical parts are crucial for prenatal screening and diagnosis.
[0003] Due to the subjectivity of ultrasound, the common method for fetal early pregnancy structure screening is to collect ultrasound images by experienced ultrasound doctors for judgment. The operation method highly depends on the experience of ultrasound doctors, and there are inter-observer and intra-observer differences in diagnosis. It takes years of training and extensive fetal anatomy knowledge to guide the ultrasound probe to obtain the correct scanning plane through the complex fetal anatomy structure, evaluate each anatomical structure of the fetus, and make correct diagnosis. This process is prone to error, time-consuming, and heavily dependent on ultrasound doctors. The contradiction between the huge detection demand and the scarcity of doctor resources may lead to missed screening. Therefore, there is a need for an objective, non-invasive, and more reliable auxiliary screening system to reduce the misdiagnosis rate and missed diagnosis rate of ultrasound examination. It is of great significance to develop an accurate and robust method to automatically evaluate fetal biological feature parameters.
[0004] Early studies on computer vision recognition of prenatal ultrasound mainly rely on traditional segmentation methods that depend on artificial features. However, traditional edge detection and region growing methods often require a large amount of contrast and are less adaptable to ultrasound image segmentation with strong noise interference. They also lack the use of global and local information, and perform poorly in image segmentation tasks involving complex spatial information. Compared with traditional methods, deep neural networks can capture more comprehensive information and have obvious advantages in scene understanding and object recognition, which can further improve the accuracy, real-time performance and robustness of the network. In recent years, artificial intelligence algorithms based on deep learning have made significant progress in prenatal diagnosis. Currently, fetal ultrasound still faces challenges in clinical practical application due to the complexity of ultrasound features and the subjectivity of ultrasound examination. In the past few decades, with the development of deep learning technology, neural network-based models, from convolutional neural networks (CNN) to the latest Vision Transformer (ViT), have achieved excellent results in medical image processing tasks. Due to the excellent performance of U-Net network in medical images, U-shaped network has become the first choice for researchers in original ultrasound segmentation. Many researchers have introduced modifications to the UNET structure to enhance segmentation efficiency and combined attention mechanisms to improve detail capture.
[0005] However, ultrasound images are the most difficult to automatically segment in medical images. The presence of acoustic shadows, speckle noise, motion blur and boundary loss caused by the complex interaction between ultrasound waves and maternal and fetal biological tissues makes it challenging to analyze fetal ultrasound images. Therefore, current deep learning-based fetal ultrasound image classification and segmentation methods still have some problems, including inaccurate detection of small targets, blurred boundaries, and severe background noise interference. Previous methods often struggle to effectively distinguish between diseased tissue and background, resulting in insufficient segmentation accuracy and affecting subsequent diagnostic accuracy. Therefore, it is necessary to explore an efficient and accurate intelligent processing method for fetal ultrasound images. SUMMARY
[0006] To solve the above problems, the present application proposes an image classification method for extracting standard frames from ultrasound video and an intelligent processing method combining diffusion model and pyramid visual Transformer, including a standard frame extraction method for fetal ultrasound video and a multi-target segmentation method based on diffusion model and Transformer. This method has multi-scale feature extraction capability, detail enhancement capability and robust anti-interference capability, which can effectively improve the shortcomings of the above traditional methods.
[0007] The above technical problems of the present application are mainly solved by the following technical solutions: The first aspect provides a standard frame identification and multi-target segmentation method for fetal ultrasound video, comprising: extracting a fetal ultrasound image sequence from a dynamic ultrasound video; processing the ultrasound image sequence by using a pre-trained classification model to obtain a classification result and an image frame of a standard body position, wherein the classification result is used to indicate whether the fetal body position is a standard body position, and the image frame of the standard body position contains fetal anatomical structure information; segmenting the image frame of the standard body position by using a pre-trained segmentation model to obtain a multi-target segmentation result, wherein the multi-target segmentation result is used to indicate the position and boundary of a key anatomical part, and the pre-trained segmentation model comprises a pyramid vision Transformer feature extraction network, a denoising encoder based on a diffusion probability model, and a cross-attention fusion module, wherein the pyramid vision Transformer feature extraction network is used to perform multi-scale semantic feature extraction on an original image through a four-level pyramid structure, the denoising encoder based on the diffusion probability model is used to obtain enhanced features according to noisy image features and multi-scale semantic features, and the cross-attention fusion module is used to fuse the multi-scale semantic features and the enhanced features.
[0008] In an embodiment, the pre-trained classification model is a VGG-19 network, which is trained by using a fetal ultrasound image dataset containing standard body positions and non-standard body positions.
[0009] In an embodiment, the pyramid vision Transformer feature extraction network comprises a plurality of pyramid levels, each pyramid level comprising a plurality of convolution blocks and attention blocks, and the multi-scale semantic features of the fetal ultrasound image are extracted through the operations of the convolution blocks and the attention blocks.
[0010] In an embodiment, the denoising encoder based on the diffusion probability model comprises T denoising U-Net structures, T being a diffusion step length, and each denoising U-Net structure comprises four adaptive attention enhancement modules.
[0011] In an embodiment, the processing process of the adaptive attention enhancement module comprises: calculating two features of different sources at the input end: semantic features extracted by the pyramid vision Transformer feature extraction network and noisy image features; extracting features through parallel convolution layers; calculating a spatial attention map and a channel attention map, respectively, wherein the spatial attention map is generated by calculating the spatial correlation of a feature map, and the channel attention map is generated by calculating the channel correlation of the feature map; multiplying the two attention maps with the original features to obtain enhanced features.
[0012] In an implementation, the training process of the pre-trained segmentation model comprises: In the first training stage, the pre-training of the pyramid vision Transformer feature extraction network; In the second training stage, the joint training of the pyramid vision Transformer feature extraction network, the denoising encoder based on the diffusion probability model and the cross-attention fusion module.
[0013] Based on the same inventive concept, the second aspect of the present application provides a standard frame identification and multi-target segmentation device in a fetal ultrasound video, comprising: An image sequence extraction module is configured to extract a fetal ultrasound image sequence from a dynamic ultrasound video; An image classification module is configured to process the ultrasound image sequence by using a pre-trained classification model to obtain a classification result and an image frame of a standard body position, wherein the classification result is used to indicate whether the fetal body position is a standard body position, and the image frame of the standard body position contains fetal anatomical structure information; An image segmentation module is configured to segment the image frame of the standard body position by using a pre-trained segmentation model to obtain a multi-target segmentation result, wherein the multi-target segmentation result is used to indicate the position and boundary of the key anatomical part, and the pre-trained segmentation model comprises a pyramid vision Transformer feature extraction network, a denoising encoder based on a diffusion probability model and a cross-attention fusion module, wherein the pyramid vision Transformer feature extraction network is used to perform multi-scale semantic feature extraction on the original image through a four-level pyramid structure, the denoising encoder based on the diffusion probability model is used to obtain enhanced features according to the noisy image features and the multi-scale semantic features, and the cross-attention fusion module is used to fuse the multi-scale semantic features and the enhanced features.
[0014] Based on the same inventive concept, the third aspect of the present application provides an electronic device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to realize the standard frame identification and multi-target segmentation method in a fetal ultrasound video according to the first aspect.
[0015] Based on the same inventive concept, the fourth aspect of the present application provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to realize the standard frame identification and multi-target segmentation method in a fetal ultrasound video according to the first aspect.
[0016] Based on the same inventive concept, the fifth aspect of the present application provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to realize the efficient fetal ultrasound image segmentation method based on the diffusion model according to the first aspect.
[0017] Compared with the prior art, the advantages and beneficial technical effects of the present application are as follows: For an ultrasound video file, on the one hand, by introducing a pyramid vision Transformer (PVT), multi-scale semantic feature extraction of fetal ultrasound images is realized, and the detection capability for small targets is effectively improved; on the other hand, through an adaptive attention enhancement module (ARFE), the detailed features in the ultrasound image are enhanced, and the problem of boundary blur is improved. In addition, the multi-head cross-attention mechanism adopted by the feature fusion module can effectively fuse the features of different encoders, and improve the anti-interference ability of the model. Overall, the technical solution solves the problems of inaccurate detection of small targets, blurred boundaries and serious background noise interference in the previous fetal ultrasound image classification and segmentation method by combining the diffusion model with the improved Transformer.
[0018] The present application has the advantages of simple implementation, strong practicability, and the ability to improve user experience, and has important market value. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0020] Figure 1 The flowchart of the fetal ultrasound video standard frame identification and multi-target segmentation method provided by the embodiment of the present application is shown in the figure. Figure 2 The architecture diagram of the image classification model (pre-trained classification model) of the embodiment of the present application is shown in the figure. Figure 3 The architecture diagram of the segmentation model (pre-trained segmentation model) in the embodiment of the present application is shown in the figure. Figure 4 The diffusion process diagram of the diffusion model in the embodiment of the present application is shown in the figure. Figure 5 The ultrasound image used in the embodiment of the present application is shown in the figure, wherein (a) is the original ultrasound image, (b) is the corresponding real segmentation mask, and (c) is the segmentation mask predicted by the segmentation model proposed in the present application. DETAILED DESCRIPTION
[0021] The concept, specific structure and technical effects of the present application will be further described in combination with the drawings and embodiments to fully understand the purpose, features and effects of the present application.
[0022] Embodiment 1 The embodiment provides a standard frame identification and multi-target segmentation method in a fetal ultrasound video, and refers to Figure 1 , which comprises the following steps. S1: extracting a fetal ultrasound image sequence from a dynamic ultrasound video.
[0023] Specifically, the dynamic ultrasound video can be stored in a fetal ultrasound video file, from which the fetal ultrasound image sequence is extracted as the basis for subsequent image classification.
[0024] S2: processing the ultrasound image sequence by using a pre-trained classification model to obtain a classification result and a standard body position image frame, wherein the classification result is used to indicate whether the fetal body position is a standard body position, and the standard body position image frame contains fetal anatomical structure information.
[0025] Among them, the pre-trained classification model is a VGG-19 network, which is trained by using a fetal ultrasound image dataset containing standard body positions and non-standard body positions.
[0026] In the specific implementation process, Figure 2 is a schematic diagram of an image classification model (pre-trained classification model) architecture of the embodiment of the present application, convolution+ReLU represents the combination of convolution and ReLU activation function, max pooling represents maximum pooling, full connected+ReLU represents the combination of full connection layer and ReLU activation function, and softmax represents the activation function.
[0027] The VGG-19 network contains 19 layers with weight parameters, including 16 convolutional layers and the last 3 fully connected layers. The model uses a uniform 3x3 convolution kernel and a 2x2 maximum pooling operation. In the manner of transfer learning, the VGG-19 model is first trained on a natural dataset, and then the pre-trained model is used. When training on the natural training set, the last layer of the fully connected layer is 1x1x1000. When migrating to the ultrasound image picture classification of 2 classes, the last layer of the project should be 1x1x2. Therefore, the last layer parameters need to be frozen during training, and the last layer needs to be retrained.
[0028] S3: segmenting the image frame of the standard body position by using a pre-trained segmentation model to obtain a multi-target segmentation result, wherein the multi-target segmentation result is used to indicate the position and boundary of the key anatomical part, and the pre-trained segmentation model comprises a pyramid vision Transformer feature extraction network, a denoising encoder based on a diffusion probability model, and a cross-attention fusion module, wherein the pyramid vision Transformer feature extraction network is used to perform multi-scale semantic feature extraction on the original image through a 4-level pyramid structure, the denoising encoder based on the diffusion probability model is used to obtain enhanced features according to the noisy image features and the multi-scale semantic features, and the cross-attention fusion module is used to fuse the multi-scale semantic features and the enhanced features.
[0029] The pyramid vision Transformer feature extraction network comprises a plurality of pyramid levels, and each pyramid level comprises a plurality of convolution blocks and attention blocks, and the multi-scale semantic features of the fetal ultrasound image are extracted through the operations of the convolution blocks and the attention blocks.
[0030] Specifically, the PVT (pyramid vision Transformer) feature extraction network comprises a plurality of pyramid levels, and each level reduces the size of the feature map and increases the number of channels through convolution operation to realize multi-scale feature extraction. Specifically, each level comprises a plurality of convolution blocks and attention blocks, specifically convolution layers, attention layers and pooling layers, by gradually reducing the size of the feature map and increasing the number of channels, the PVT can capture semantic information of different scales and improve the detection ability of small targets.
[0031] In the specific implementation process, the PVT comprises 4 pyramid levels (Stage 1-4), and each level gradually reduces the resolution of the feature map and increases the channel dimension: Input image size: H x W x 3 Stage 1 output: H / 4 x W / 4 x C1 (C1=32) Stage 2 output: H / 8 x W / 8 x C2 (C2=64) Stage 3 output: H / 16 x W / 16 x C3 (C3=160) Stage 4 output: H / 32 x W / 32 x C4 (C4=256).
[0032] Wherein H and W are height and width, and C1-C4 are the number of channels.
[0033] The denoising encoder based on the diffusion probability model comprises T denoising U-Net structures, T is the diffusion step length, and each denoising U-Net structure comprises 4 adaptive attention enhancement modules.
[0034] The processing procedure of the adaptive attention enhancement module comprises: Two different source features are calculated at the input end: semantic features extracted by the pyramid vision Transformer feature extraction network and noisy image features; Feature extraction is performed through parallel convolution layers; Spatial attention maps and channel attention maps are respectively calculated, wherein the spatial attention maps are generated by calculating the spatial correlation of feature maps, and the channel attention maps are generated by calculating the channel correlation of feature maps; The two attention maps are multiplied with the original features to obtain enhanced features.
[0035] Specifically, the denoising encoder is constructed based on a DDPM framework and comprises T denoising blocks, T being a diffusion step length. Each denoising block comprises a residual connection and an ARFE module (adaptive attention enhancement module), and the ARFE module enhances the detailed features in the ultrasound image through spatial attention mechanisms and channel attention mechanisms. In the denoising process, the model gradually recovers the real features of the image and reduces the influence of noise and blur.
[0036] The adaptive attention enhancement module (ARFE) calculates two different source features at the input end: semantic features extracted by the PVT and noisy image features, and feature extraction is performed through parallel convolution layers. Then, spatial attention maps and channel attention maps are respectively calculated, and finally, the two attention maps are multiplied with the original features to obtain enhanced features.
[0037] See Figure 4 is a diffusion process schematic diagram of the diffusion model in the embodiment of the present application.
[0038] Each denoising block of the denoising encoder comprises a residual connection and an ARFE module, the original feature information is retained through the residual connection, and the detailed features are enhanced through the ARFE module. In the denoising process, the model gradually recovers from a noisy image to a clear image, and the image features are updated by predicting noise at each time step.
[0039] In the specific implementation process, given the input feature of the i-th layer After two convolution operations, residual connection is performed, and channel attention and spatial attention modeling are sequentially performed:
[0040]
[0041]
[0042]
[0043]
[0044] where BN denotes batch normalization: , is the batch mean, is the batch variance, γ, β are learnable parameters, is the activation function, ~ denotes the obtained intermediate feature map, denotes concatenation, denotes channel attention, denotes spatial attention.
[0045] The final output of the ith layer encoder is:
[0046] where, denotes max-pooling.
[0047] Let the encoder output feature , PVT corresponds to the hierarchical feature , first unify the channel dimension by 1x1 convolution:
[0048]
[0049] where, denotes the feature after convolution on the output feature , denotes the feature after convolution on the hierarchical feature .
[0050] The multi-head cross attention mechanism of the feature fusion module realizes feature fusion by calculating the QKV matrix,
[0051] where is the input feature, , , are the query, key, value matrices in the attention mechanism, respectively, , , are the , , weight matrices to be learned, and the combination of the normalization layer Norm and the fully connected layer Linear.
[0052] The attention map is calculated and denoted as where, is the normalized exponential function, is the learnable parameter matrix.
[0053] The fusion result is enhanced by a residual connection:
[0054]
[0055] wherein represents a multi-head self-attention calculation function, represents channel and size adjustment on the current feature fusion result, represents the current fusion result, represents the feature after the channel and size adjustment on the current feature fusion result.
[0056] The final output fusion feature map while retaining the original feature's skip connection: .
[0057] See Figure 3 , which is a schematic diagram of a segmentation model (a pre-training segmentation model) architecture in the embodiment of the present application, XT represents a noisy image feature, Encoder is an encoder, Cross-attention fusion represents cross-attention fusion, Decoder represents a decoder, and X0 is a decoded feature.
[0058] The processing model (including a pre-training classification model and a segmentation model) of the present application has high efficiency and accuracy; the efficiency refers to the resources and time required by the above processing network in the training and inference process being less than that of a traditional model of the same scale; the accuracy refers to the ability of the above processing network to accurately identify standard body positions and key anatomical parts when processing fetal ultrasound image classification and segmentation tasks. The efficiency also lies in the fact that the PVT feature extraction network in the model performs dimension reduction processing on the original input image, and at the same time, multi-scale feature extraction is achieved through a pyramid structure, thereby reducing the amount of calculation. In addition, the progressive denoising process of the denoising encoder also improves the inference efficiency of the model. The accuracy is due to the fact that the network model enhances the detailed features in the ultrasound image through the ARFE module, and at the same time, effectively fuses the features of different encoders through the multi-head cross-attention mechanism, thereby improving the detection ability of small targets and the accuracy of boundary segmentation of the model.
[0059] The training process of the pre-training segmentation model includes: In the first training stage, a pre-training pyramid visual Transformer feature extraction network is trained. In the second training stage, the pyramid visual Transformer feature extraction network, the denoising encoder based on the diffusion probability model, and the cross-attention fusion module are jointly trained.
[0060] In the first training stage, the PVT feature extraction network is pre-trained to enable it to extract multi-scale semantic features of the original fetal ultrasound image.
[0061] In the second training stage, the PVT feature extraction network, the denoising encoder and the feature fusion module are jointly trained. The denoising encoder is trained to learn the denoising process of the noisy ultrasound image segmentation mask, and the feature fusion module is trained to learn the fusion of features of different encoders. Specifically, the output of the PVT feature extraction network is input into the U-Net model as an input combined with features extracted by encoders of different sizes, and is input into the cross-attention feature fusion module. In the inference verification sampling stage, the original ultrasound image is input into the PVT feature extraction network to obtain multi-scale original image features, the noisy ultrasound image is input into the denoising encoder, and the input features are fused in the feature fusion module to finally output the classification result and the multi-target segmentation result.
[0062] Specifically, for the pre-trained classification model, an image classification training data set is first created, which includes standard fetal ultrasound images and non-standard fetal ultrasound images. For the pre-trained segmentation model, an image segmentation training data set is first created, which includes a set of fetal ultrasound images of different qualities and corresponding annotation information; the annotation information includes fetal position labels and segmentation masks of key anatomical parts.
[0063] For the pre-trained classification model, the VGG-19 image classification network is pre-trained on a large-scale natural data set. During training, the last layer of parameters is frozen, and then retraining is performed on the proposed training data set.
[0064] For the pre-trained segmentation model, the training includes two stages. In the first training stage, the PVT feature extraction network is pre-trained to enable it to extract multi-scale semantic features of the original fetal ultrasound image. In the second training stage, the PVT feature extraction network, the denoising encoder and the feature fusion module are jointly trained. The denoising encoder is trained to learn the denoising process of the noisy ultrasound image segmentation mask, and the feature fusion module is trained to learn the fusion of features of different encoders. Specifically, the output of the PVT feature extraction network is input into the U-Net model as an input combined with features extracted by encoders of different sizes, and is input into the cross-attention feature fusion module. In the inference verification sampling stage, the original ultrasound image is input into the PVT feature extraction network to obtain enhanced features, the noisy ultrasound image is input into the denoising encoder, and the input features are fused in the feature fusion module to finally output the classification result and the multi-target segmentation result.
[0065] In the first training stage, the PVT feature extraction network is pre-trained.
[0066] Based on the image segmentation training data set, multiple fetal ultrasound images are input into the PVT feature extraction network to be trained, and multiple scale semantic features are generated by the PVT feature extraction network to be trained. Based on the multi-scale semantic features and the labeled information, the PVT feature extraction network to be trained is trained by using the cross-entropy loss function until the preset training condition is met.
[0067] The second training stage: jointly training the PVT feature extraction network, the denoising encoder and the feature fusion module.
[0068] Based on the image segmentation training data set, multiple fetal ultrasound images are input into the PVT feature extraction network to be trained, and multiple scale semantic features are generated by the PVT feature extraction network to be trained. Based on the multi-scale semantic features and the labeled information, the PVT feature extraction network to be trained is trained by using the cross-entropy loss function until the preset training condition is met.
[0069] Based on the classification result, the multi-target segmentation result and the labeled information, the function value of the preset loss function is obtained by using the preset weighted cross-entropy and Dice mixed loss function. Based on the function value of the preset loss function, the denoising encoder and the feature fusion module to be trained are trained until the preset training condition is met.
[0070] Please refer to Figure 5 , which is an ultrasound image used in the embodiment of the present application, wherein the (a) part is an original ultrasound image, the (b) part is a corresponding real segmentation mask, and the (c) part is a segmentation mask predicted by the segmentation model proposed in the present application.
[0071] In general, the fetal ultrasound video standard frame recognition and multi-target segmentation method of the present application, on the one hand, by introducing the pyramid visual transformer (PVT), realizes the multi-scale semantic feature extraction of the fetal ultrasound image, effectively improves the detection ability of small targets; on the other hand, through the adaptive attention enhancement module (ARFE), the detailed features in the ultrasound image are enhanced, and the problem of blurred boundary is improved. In addition, the multi-head cross attention mechanism adopted by the feature fusion module can effectively fuse the features of different encoders, and improve the anti-interference ability of the model. Overall, the method solves the problems of inaccurate small target detection, blurred boundary and serious background noise interference in the previous fetal ultrasound image classification and segmentation method by combining the diffusion model with the improved transformer.
[0072] In specific implementation, the method provided by the technical scheme of the present application can be automatically run by a computer software technology, and a system device of the method, such as a computer readable storage medium storing a computer program of the technical scheme of the present application and a computer device including the computer program, should also be within the protection scope of the present application.
[0073] The electronic device for processing an ultrasound image of a fetus based on a diffusion model provided by the present application is described below, and the electronic device for processing an ultrasound image of a fetus based on a diffusion model described below can be correspondingly referred to the method for processing an ultrasound image of a fetus based on a diffusion model described above.
[0074] The electronic device can include a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory can communicate with each other through the communications bus. The processor can call logical instructions in the memory to execute the method for processing an ultrasound image of a fetus based on a diffusion model, mainly including software processing in the above steps.
[0075] In addition, the logical instructions in the memory described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical scheme of the present application or the part of the technical scheme that essentially contributes to the prior art or the part of the technical scheme can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0076] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor, and the computer can execute the software processing part of the method for processing an ultrasound image of a fetus based on a diffusion model provided by the above method.
[0077] In yet another aspect, the present application also provides a non-transitory computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the software processing part of the method for intelligent processing of fetal ultrasound images based on diffusion model provided by the above method.
[0078] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems, or computer program products. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can be embodied in the form of a computer program product on one or more computer readable storage media (including, but not limited to, disk memory, CD-ROMs, optical storage media, etc.) having computer usable program code embodied thereon.
[0079] The present application is described in reference to the flowchart and / or block diagrams of the method, apparatus (system) and computer program product according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagrams, as well as a combination of flows and / or blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, a special purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions, which are executed via the processor of the computer or other programmable data processing apparatus, generate a means for implementing the functions specified in the flowchart and / or block diagrams. Figure 1 The means for carrying out each one or more of the functions specified in the flowchart and / or block diagrams can be embodied in software, firmware, hardware, or any combination thereof. Figure 1 The means for carrying out each one or more of the functions specified in the flowchart and / or block diagrams can be embodied in software, firmware, hardware, or any combination thereof.
[0080] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments by those skilled in the art once they learn of the basic inventive concepts. Therefore, the appended claims are intended to encompass within their scope all possible variations and modifications of the preferred embodiments. It is apparent that those skilled in the art can, without departing from the spirit and scope of the embodiments, make various changes and modifications of the embodiments to adapt it to various usages and conditions. Thus, if these modifications and variations do not depart from the scope of the claims and their equivalents, they are intended to be included within the scope of the application.
Claims
1. A method for standard frame identification and multi-object segmentation in fetal ultrasound video, characterized in that, The method comprises the following steps: extracting a fetal ultrasound image sequence from a dynamic ultrasound video; processing the ultrasound image sequence by using a pre-trained classification model to obtain a classification result and an image frame of a standard body position, wherein the classification result is used to indicate whether the fetal body position is a standard body position, and the image frame of the standard body position contains fetal anatomical structure information; segmenting the image frame of the standard body position by using a pre-trained segmentation model to obtain a multi-target segmentation result, wherein the multi-target segmentation result is used to indicate the positions and boundaries of key anatomical parts, and the pre-trained segmentation model comprises a pyramid vision Transformer feature extraction network, a denoising encoder based on a diffusion probability model, and a cross-attention fusion module, wherein the pyramid vision Transformer feature extraction network is used to perform multi-scale semantic feature extraction on an original image through a four-level pyramid structure, the denoising encoder based on the diffusion probability model is used to obtain enhanced features according to noisy image features and multi-scale semantic features, and the cross-attention fusion module is used to fuse the multi-scale semantic features and the enhanced features.
2. The method of standard frame identification and multi-object segmentation in fetal ultrasound videos of claim 1, wherein, The pre-trained classification model is a VGG-19 network, which is trained by using a fetal ultrasound image dataset containing standard body positions and non-standard body positions.
3. The method of standard frame identification and multi-object segmentation in fetal ultrasound videos of claim 1, wherein, The pyramid vision Transformer feature extraction network comprises a plurality of pyramid levels, each pyramid level comprising a plurality of convolution blocks and attention blocks, and the multi-scale semantic features of the fetal ultrasound image are extracted through the operations of the convolution blocks and the attention blocks.
4. The method of standard frame identification and multi-object segmentation in fetal ultrasound videos of claim 1, wherein, The denoising encoder based on the diffusion probability model comprises T denoising U-Net structures, T being a diffusion step length, and each denoising U-Net structure comprises four adaptive attention enhancement modules.
5. The method of standard frame identification and multi-object segmentation in fetal ultrasound videos of claim 1, wherein, The processing process of the adaptive attention enhancement module comprises the following steps: calculating two kinds of features of different sources at the input end: semantic features extracted by the pyramid vision Transformer feature extraction network and noisy image features; extracting features through parallel convolution layers; calculating a spatial attention map and a channel attention map respectively, wherein the spatial attention map is generated by calculating the spatial correlation of a feature map, and the channel attention map is generated by calculating the channel correlation of the feature map; multiplying the two kinds of attention maps with the original features to obtain enhanced features.
6. The method for standard frame identification and multi-object segmentation in fetal ultrasound videos of claim 1, wherein, The training process of the pre-trained segmentation model comprises the following steps: in a first training stage, pre-training the pyramid vision Transformer feature extraction network; in a second training stage, jointly training the pyramid vision Transformer feature extraction network, the denoising encoder based on the diffusion probability model, and the cross-attention fusion module.
7. A device for standard frame identification and multi-object segmentation in fetal ultrasound videos, characterized in that, The method comprises the following steps: an image sequence extraction module for extracting a fetal ultrasound image sequence from a dynamic ultrasound video; an image classification module for processing the ultrasound image sequence by using a pre-trained classification model to obtain a classification result and an image frame of a standard body position, wherein the classification result is used to indicate whether the fetal body position is a standard body position, and the image frame of the standard body position contains fetal anatomical structure information; An image segmentation module is configured to segment image frames of a standard body position using a pre-trained segmentation model to obtain a multi-target segmentation result, wherein the multi-target segmentation result is used to indicate the position and boundary of a key anatomical site, and the pre-trained segmentation model comprises a pyramid vision Transformer feature extraction network, a denoising encoder based on a diffusion probability model, and a cross-attention fusion module, wherein the pyramid vision Transformer feature extraction network is configured to perform multi-scale semantic feature extraction on an original image through a four-level pyramid structure, the denoising encoder based on the diffusion probability model is configured to obtain enhanced features according to noisy image features and multi-scale semantic features, and the cross-attention fusion module is configured to fuse the multi-scale semantic features and the enhanced features.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the standard frame identification and multi-target segmentation method in the fetal ultrasound video according to any one of claims 1 to 7 when executing the program.
9. A non-transitory computer-readable storage medium, comprising: A computer program is stored thereon, and the program is executed by a processor to implement the standard frame identification and multi-target segmentation method in the fetal ultrasound video according to any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processor to implement the efficient fetal ultrasound image segmentation method based on a diffusion model according to any one of claims 1 to 6. The computer program is executed by a processor to implement the efficient fetal ultrasound image segmentation method based on a diffusion model according to any one of claims 1 to 6.