Hybrid network method for pubis symphysis and fetal head segmentation

Through the DSSAU-Net hybrid network method, combined with the sparse attention mechanism and multi-scale feature fusion, the problem of insufficient segmentation accuracy of fetal head and pubic joint is solved, achieving higher segmentation accuracy and lower computational complexity.

CN120014266APending Publication Date: 2025-05-16CHONGQING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510079796.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is inaccurate in segmentation of fetal head and pubic joints, especially in the face of speckle noise, ultrasound artifacts and blurred target boundaries.

Method used

A hybrid network method, called DSSAU-Net, is adopted, combining sparse attention mechanisms and convolutions, and through multi-scale feature fusion and U-type encoder-decoder structure, more efficient segmentation of the fetal head and pubic joint is achieved.

Benefits of technology

The segmentation accuracy of fetal head and pubic joint was improved, and a Dice similarity coefficient of 86.43%, a Hausdorff distance of 31.08 and an average surface distance of 8.39 were obtained, while reducing the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014266A_ABST
    Figure CN120014266A_ABST
Patent Text Reader

Abstract

The invention relates to a hybrid network method DSSAU-Net for pubis symphysis and fetal head segmentation, which is used for FH and PS segmentation. In each stage, a different number of dual sparse selection attention (DSSA) blocks are stacked to form a symmetric U-shaped encoder-decoder network architecture. For a given query, the DSSA is designed to explicitly perform one sparse token selection at the region and pixel levels, respectively, which facilitates further reduction of computational complexity while extracting the most relevant features. In order to compensate for information loss in the up-sampling process, jump connection with convolution is designed. In addition, a multi-scale feature fusion method is adopted to enrich global information and local information of the model. Results on a data set show that the medical image segmentation method can realize accurate and efficient medical image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of image processing, and in particular to a hybrid network method for segmenting pubic symphysis and fetal head. Background Art

[0002] Dysfunctional labor is a leading cause of maternal mortality and morbidity and refers to the inability of part of the fetus to pass through the birth canal despite strong uterine contractions. Therefore, during labor, repeated monitoring of the fetal position is necessary to prevent this from happening. Traditional vaginal examination is not only subjective but also potentially invasive, and some related methods are difficult to implement with high reliability. The advent of ultrasound-assisted diagnosis has provided a new noninvasive and accurate method for assessing fetal position and cervical dilation and has gradually gained acceptance in the field of obstetrics. Some studies have shown that intrapartum ultrasound assessment can help more accurately assess the position of the fetal head and the position of the pubic symphysis. The International Society of Ultrasound in Obstetrics and Gynecology recommends ultrasound assessment before considering instrumental delivery or suspected delayed delivery. Two reliable ultrasound parameters are used to predict the outcome of instrumental-assisted vaginal delivery: the angle of progression (AoP) and the head-pubic distance (HSD).

[0003] In ultrasound-assisted diagnosis, AoP and HSD are key parameters for assessing the progress of labor. The acquisition of these parameters depends on the accurate segmentation of the fetal head (FH) and pubic symphysis (PS), which helps clinicians monitor the fetal labor status in real time. Therefore, many methods have begun to focus on improving the segmentation performance of the fetal head (FH) and pubic symphysis (PS). For example, the fetal head-pubic symphysis segmentation network (FH-PSSNet) adopts an encoder-decoder framework and integrates a dual attention module, a multi-scale feature screening module, and a direction guidance block to achieve automatic angle of progression (AoP) measurement. The dual-path boundary-guided residual network (DBRN) integrates a multi-scale weighted module (MWM), an enhanced boundary module (EBM), and a boundary-guided dual attention residual module (BDRM) to address the challenges of achieving fully automatic and accurate segmentation of the fetal head-pubic symphysis (FH-PS) in low-contrast or anatomical boundary ambiguity. In addition, BRAU-Net performs well in region-level sparse tokens segmentation, but its segmentation robustness for small targets is still insufficient in the face of complex situations such as speckle noise, ultrasonic artifacts, and blurred target boundaries. Summary of the invention

[0004] In view of the above problems existing in the prior art, the technical problem to be solved by the present invention is: how to improve the accuracy of segmentation.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solution: a hybrid network method for pubic symphysis and fetal head segmentation, comprising the following steps:

[0006] S1: Acquire several ultrasound images during production as a data set;

[0007] S2: Construct a U-shaped hybrid network DSSAU-Net for pubic symphysis and fetal head segmentation, and use a supervised method to train DSSAU-Net to obtain a trained DSSAU-Net;

[0008] The DSSAU-Net includes an encoder, a pyramid pool module, a decoder, a 1×1 convolution layer, a 4× upsampling layer, and a 3×3 convolution layer. The encoder includes four stages from stage 1 to stage 4, and stages 1 to stage 4 include 2, 2, 8, and 2 DSSA blocks respectively, which are denoted as DSSAblock; the decoder also includes four stages from stage 5 to stage 8, and stages 5 to stage 8 include 2, 2, 8, and 2 DSSA blocks respectively, and DSSA blocks are denoted as DSSAblock.

[0009] S3: For a delivery ultrasound image, the delivery ultrasound image is input into the trained DSSAU-Net, and the output is an image with the pubic symphysis and fetal head segmented.

[0010] Further, the DSSA block comprises the following steps:

[0011] 1) For each image in the dataset, a 3×3 deep convolution is used to encode the relative position information to obtain the encoded image, and then layer normalization is performed.

[0012] 2) Apply DSSA to compute attention on the encoded image, followed by another layer of normalization;

[0013] 3) Input the normalized result of 2) into a two-layer multi-layer perceptron MLP module.

[0014] The DSSA comprises the following steps: for a given input image X of size H×W×C, X is divided into S×S non-overlapping regions to obtain a region image The region tokens are obtained by linear projection: the query, key and value are denoted as Q, K,

[0015] Q=X r W q

[0016] K=X r W k (1)

[0017] V=X r W v

[0018] Among them, S represents the number of regions, r represents the token is regional level, and W qq , Wkk , W vv ∈R C×C are the projection weight matrices of the corresponding query, key, and value, respectively.

[0019] For region-level sparse token selection, each region-level token query and key It is obtained by averaging all pixel-level tokens on the region token Q and key K; Q r and K r Attention mapping between Represents the semantic relevance between regions; a top-k1 operation is used to select the k1 most relevant regions for each region containing a given pixel-level token and record their indices in In , the process is described as follows:

[0020] A r =Q r (K r ) Τ (2)

[0021] I r =topkIndex(A r ) (3)

[0022] Aggregate the filtered region-level tokens into a matrix.

[0023] K g =gather(K,I r ) (4)

[0024] V g =gather(V,I r )

[0025] Where K g , It is the key tensor matrix and value tensor matrix obtained by aggregation.

[0026] Regarding pixel-level sparse token selection, for any pixel-level query in Q, calculate the difference between the query and K g The correlation between the pixel-level keys aggregated in is the attention matrix Expressed as The specific calculation method is as follows:

[0027]

[0028] Where p means that the token is at the pixel level.

[0029] exist In the second round of pixel-level sparse token selection, the top-k2 operation is used, and the value and index storage corresponding to the selected token are stored and represented as and

[0030]

[0031] Among them, k2 is determined by the proportional factor λ:

[0032]

[0033] According to I pp From V gg In the polymerization

[0034] V gg =gather(V g ,I p ) (9)

[0035] V gg Use A p After weighting, a local context enhancement term LCE(V) is added to the output matrix O:

[0036] O = Attention (A p ,V gg )+LCE(V) (10).

[0037] The DSSA block uses three residual connections to help alleviate the gradient disappearance. The first residual connection is to add the feature map before deep convolution encoding to the feature map after deep convolution encoding element by element; the second residual connection is to add the feature map before the first layer normalization to the feature map after DSSA element by element; the third residual connection is to add the feature map before the second layer normalization to the feature map after passing through two layers of multi-layer perceptron MLP modules element by element.

[0038]

[0039] Among them, z ll-1 represents the output after the l-1th DSSA block, and It represents the result of element-by-element addition of the feature map output by DSSA and the feature map before the first layer normalization.

[0040] Furthermore, the process of S2 training DSSAU-Net is as follows:

[0041] S2-1: For stage 1, it consists of an overlapping block embedding layer with two 3×3 convolutions and two DSSA blocks to transform the input image of size H×W×3 into a The feature image X1:

[0042] X1=DSSAblock ×2 (embedding(X)) (13)

[0043] In stage 2, a block merging layer is used to reduce the number of tokens, and a linear embedding layer is used to increase the dimension C to 2C. The size of the feature map X2 output by stage 2 containing two DSSA blocks is

[0044] X2=DSSAblock ×2 (Merging(X1)) (14)

[0045] The sizes of the corresponding feature maps X3 and X4 generated in stage 3 and stage 4 are and They are respectively expressed as:

[0046] X3=DSSAblock ×8 (Merging(X2)) (15)

[0047] X4=DSSAblock ×2 (Merging(X3)) (16)

[0048] Among them, DSSAblock ×2 and DSSAblock ×8 They represent stacking 2 DSSA blocks and 8 DSSA blocks, the embedding (X) block embedding layer, and the merging (·) block merging layer respectively.

[0049] S2-2: Use the pyramid pool module PPM to X4 Mapped to size The feature map of

[0050] S2-3: Size The feature maps are input to the decoder constructed hierarchically from stage 5 to stage 8; in this process, the feature maps of each stage are processed by the DSSA block and fused with the feature maps of the same resolution from the encoder. The resulting feature maps are then upsampled to the output size The whole process is described as:

[0051] X5=DSSAblock ×2 (PPM(X4)) (17)

[0052]

[0053] Among them, ⊕ represents element-level addition, PPM(·), Con 1×1 、Up 2× They represent the pyramid pooling module, 1x1 convolution operation, and double upsampling respectively.

[0054] S2-4: X5, X6, and X7 are upsampled to the same resolution as X8. All these upsampled feature maps are concatenated along the channel dimension, and then a 1×1 convolution layer is applied, followed by a 4× upsampling, and then a 3×3 convolution layer is applied. Finally, a feature map with a resolution of H×W×class is output to predict pixel-level segmentation. The process is expressed as:

[0055] Cat=Concat(Up 8× (X5),Up 4× (X6),Up 2× (X7),X8) (21)

[0056] output=Cov 3×3 (Up 4× (Cov 1×1 (Cat))) (22)

[0057] Among them, Cat is the result of splicing the upsampled feature maps; 8x upsampling; 4x upsampling.

[0058] S2-5: Define the loss function L as follows:

[0059]

[0060] Among them, pi is the predicted probability of the i-th pixel, gi is the true value of the i-th pixel, and N is the number of pixels.

[0061] The parameters of DSSAU-Net are updated in reverse according to the L value. When the L value no longer decreases, the trained DSSAU-Net is obtained.

[0062] Compared with the prior art, the present invention has at least the following advantages:

[0063] 1. Compared with the existing methods, the present invention considers combining the sparse attention mechanism with convolution and adopts a multi-scale feature fusion method to achieve more effective FH and PS segmentation.

[0064] 2. A novel U-network architecture combined with sparse self-attention, called DSSAU-Net, is used for FH and PS segmentation. The adopted sparse self-attention, derived from the idea of ​​focusing on tokens at the region and pixel levels, is a content-aware dynamic mechanism that is explicitly designed as a dual-selective operation. We call this sparse mechanism Dual Sparse Selective Attention (DSSA). DSSA significantly reduces the computational complexity while extracting the most relevant features.

[0065] 3. In order to compensate for the information loss in the upsampling process, a skip connection with convolution is designed. A multi-scale feature fusion method is also used to enrich the global and local information of the model. The performance of DSSAU-Net has been verified. DSSAU-Net has achieved good segmentation results, with DSC, HD and ASD reaching 86.43, 31.08 and 8.39 respectively, while the number of parameters and floating-point operations (FLOPs) are low. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 Illustration of Dual Sparse Selective Attention (DSSA).

[0067] Figure 2 It is a simplified flow chart of the present invention, wherein (a) DSSAU-Net is a U-type hybrid network, and (b) is the detailed information of the DSSA block.

[0068] Figure 3 Visualization of the prediction results on the DSSAU-Net validation set. Red and green represent the segmentation results of the pubic symphysis (PS) and fetal head (FH), respectively. DETAILED DESCRIPTION

[0069] The present invention is described in further detail below.

[0070] The present invention combines sparse attention mechanism with convolution and adopts a multi-scale feature fusion method to achieve more effective segmentation of fetal head (FH) and pubic symphysis (PS). To this end, a novel U-shaped network architecture combined with sparse self-attention, called DSSAU-Net, is proposed for the segmentation of FH and PS. Specifically, the adopted sparse self-attention, which originates from the idea of ​​paying attention to tokens at the regional and pixel levels, is a content-aware dynamic mechanism and is explicitly designed as a dual selection operation. The present invention calls this sparse mechanism Dual Sparse Selective Attention (DSSA). DSSA performs two sparse token selections at the regional and pixel levels, which can reduce the computational complexity while extracting the most relevant features; DSSA significantly reduces the computational complexity while extracting more accurate features. In addition, in order to effectively extract multi-scale features, the basic blocks built on the attention mechanism are stacked into a symmetrical U-shaped encoder-decoder structure, in which the feature maps of different resolutions of the decoder are fused to produce better segmentation results. Finally, we conducted experiments on the Intrapartum Ultrasound Challenge 2024 dataset to verify the superiority of the proposed method.

[0071] DSSAU-Net is a U-shaped hybrid network that uses a sparse attention mechanism: dual sparse selective attention (DSSA) as the core construction idea, and hierarchically designs the encoder-decoder structure. In addition, we also use a pyramid pooling module (PPM) to fuse multi-scale features, which is beneficial to improve segmentation performance.

[0072] A hybrid network method for pubic symphysis and fetal head segmentation comprises the following steps:

[0073] S1: Acquire several ultrasound images during production as a data set;

[0074] S2: Construct a U-shaped hybrid network DSSAU-Net for pubic symphysis and fetal head segmentation, and use a supervised method to train DSSAU-Net to obtain a trained DSSAU-Net;

[0075] The DSSAU-Net includes an encoder, a pyramid pool module, a decoder, a 1×1 convolution layer, a 4× upsampling layer, and a 3×3 convolution layer. The encoder includes four stages from stage 1 to stage 4, and stages 1 to stage 4 include 2, 2, 8, and 2 DSSA blocks respectively, which are denoted as DSSAblock; the decoder also includes four stages from stage 5 to stage 8, and stages 5 to stage 8 include 2, 2, 8, and 2 DSSA blocks respectively, and DSSA blocks are denoted as DSSAblock.

[0076] S3: For a delivery ultrasound image, the delivery ultrasound image is input into the trained DSSAU-Net, and the output is an image with the pubic symphysis and fetal head segmented.

[0077] Specifically, the DSSA block includes the following steps:

[0078] 1) For each image in the dataset, a 3×3 deep convolution is used to encode the relative position information to obtain the encoded image, and then layer normalization is performed.

[0079] 2) Apply DSSA to compute attention on the encoded image, followed by another layer of normalization;

[0080] 3) Input the normalized result of 2) into a two-layer multi-layer perceptron MLP module.

[0081] The core idea of ​​DSSA is to perform sparse token selection explicitly at the region and pixel levels. It consists of three main steps. First, for a given region-level token where a pixel-level token is located, the region-level query Q r and key K r The attention score matrix A between r , select the most relevant areas and filter out irrelevant areas. Secondly, for a given pixel-level token, according to the pixel-level query Q p and key K p The attention score matrix between , select the most relevant pixels and filter out irrelevant pixels. Finally, the normalized attention matrix A p and pixel-level V p The matrix multiplication between , calculates the output matrix O. The DSSA comprises the following steps: The conceptual diagram of DSSA is as follows: Figure 1 According to the BiFormer method, for a given input image X of size H×W×C, X is divided into S×S non-overlapping regions to obtain a regional image The region tokens are obtained by linear projection: the query, key and value are denoted as Q, K,

[0082] Q=X r W q

[0083] K=X r W k (1)

[0084] V=X r W v

[0085] Among them, S represents the number of regions, r represents the token is regional level, and W q , W k , W v ∈R C×Care the projection weight matrices of the corresponding query, key, and value, respectively.

[0086] For region-level sparse token selection, each region-level token query and key It is obtained by averaging all pixel-level tokens on the region token Q and key K; Q r and K r Attention mapping between Represents the semantic relevance between regions; a top-k1 operation is used to select the k1 most relevant regions for each region containing a given pixel-level token and record their indices in In , the process is described as follows:

[0087] A r =Q r (K r ) Τ (2)

[0088] I r =topkIndex(A r ) (3)

[0089] In order to fully utilize the acceleration capability of GPU, the filtered region-level tokens need to be aggregated into a matrix.

[0090] K g =gather(K,I r ) (4)

[0091] V g =gather(V,I r )

[0092] Where K g , The key tensor matrix and value tensor matrix obtained by aggregation are further used for pixel-level sparse token selection.

[0093] Regarding pixel-level sparse token selection, for any pixel-level query in Q, calculate the difference between the query and K g The correlation between the pixel-level keys aggregated in is the attention matrix Expressed as The specific calculation method is as follows:

[0094]

[0095] Where p indicates that the token is pixel-level. In addition, since the first round of region-level sparse token selection involves an average operation within a region, noise features may be retained due to the overall high correlation. This may have a negative impact on the model's ability to effectively extract features.

[0096] exist In , the second round of pixel-level sparse token selection is performed using the top-k2 operation, which not only selects the k2 most relevant tokens but also implicitly removes noise. The values ​​and indices corresponding to the selected tokens are stored and represented as and

[0097]

[0098] Among them, k2 is determined by the proportional factor λ:

[0099]

[0100] Similarly, in order to take advantage of GPU acceleration, we p From V g In the polymerization

[0101] V gg =gather(V g ,I p ) (9)

[0102] V gg use A p In addition, in order to retain finer-grained information that is beneficial to pixel-level segmentation, a local context enhancement term LCE(V) is added to the output matrix O, which is a 5×5 depth convolution:

[0103] O = Attention (A p ,V gg )+LCE(V) (10);

[0104] The output matrix O is the normalized attention score matrix A p and the matrix V after linear transformation gg The final result after matrix operations.

[0105] Specifically, the DSSA block uses three residual connections to help alleviate the gradient disappearance. The first residual connection is to add the feature map before deep convolution encoding and the feature map after deep convolution encoding element by element; the second residual connection is to add the feature map before the first layer normalization and the feature map after DSSA element by element; the third residual connection is to add the feature map before the second layer normalization and the feature map after two layers of multi-layer perceptron MLP modules element by element to improve the stability of the model.

[0106]

[0107]

[0108] Among them, z l-1 represents the output after the l-1th DSSA block, and It represents the result of element-by-element addition of the feature map output by DSSA and the feature map before the first layer normalization.

[0109] Specifically, the process of S2 training DSSAU-Net is:

[0110] S2-1: DSSAU-Net is a hybrid CNN-Transformer architecture, where the encoder-decoder structure is built based on DSSA blocks, such as Figure 2 As shown in part (a) of the figure. For stage 1, it consists of an overlapping block embedding layer with two 3×3 convolutions and two DSSA blocks to transform the input image of size H×W×3 into a The characteristic image X1 :

[0111] X1=DSSAblock ×2 (embedding(X)) (13)

[0112] In stage 2, a block merging layer is used to reduce the number of tokens, and a linear embedding layer is used to increase the dimension C to 2C. The size of the feature map X2 output by stage 2 containing two DSSA blocks is

[0113] X2=DSSAblock ×2 (Merging(X1)) (14)

[0114] The sizes of the corresponding feature maps X3 and X4 generated in stage 3 and stage 4 are and They are respectively expressed as:

[0115] X3=DSSAblock ×8(Merging(X2)) (15)

[0116] X4=DSSAblock ×2 (Merging(X3)) (16)

[0117] Among them, DSSAblock ×2 and DSSAblock ×8 They represent stacking 2 DSSA blocks and 8 DSSA blocks, respectively, the embedding (X) block embedding layer, which consists of two 3*3 convolutions, and the merging (·) block merging layer, which consists of a 3*3 convolution.

[0118] S2-2: Since X4 aggregates contextual information from different stages, a pyramid pooling module PPM is used to fully utilize the rich semantic information contained in X4.

[0119] Use the pyramid pooling module PPM to map X4 into a The dimension of the feature map is fixed to C d This setting helps reduce the number of parameters and ensures that features at different scales have equal importance.

[0120] S2-3: size The feature maps are input to the decoder constructed hierarchically from stage 5 to stage 8; in this process, the feature maps of each stage are processed by the DSSA block and fused with the feature maps of the same resolution from the encoder. The resulting feature maps are then upsampled to the output size The whole process is described as:

[0121] X5=DSSAblock ×2 (PPM(X4)) (17)

[0122]

[0123]

[0124] Among them, ⊕ represents element-level addition, PPM(·), Con 1×1 、Up 2× They represent pyramid pooling modules, that is, using pooling operations of different scales to generate multiple feature maps of different sizes, then using 1x1 convolution for dimensionality reduction, and then restoring the feature map dimension by upsampling, and then concatenating these feature maps with the original feature maps in the channel dimension to form a composite feature map containing information of multiple scales; 1x1 convolution operation, double upsampling.

[0125] S2-4: In addition, in order to effectively fuse multi-scale features at different stages, X5, X6, and X7 are upsampled to the same resolution as X8. All these upsampled feature maps are concatenated along the channel dimension, and then a 1×1 convolution layer is applied, followed by a 4× upsampling, and then a 3×3 convolution layer is applied. Finally, a feature map with a resolution of H×W×class is output to predict pixel-level segmentation. The process is expressed as:

[0126] Cat=Concat(Up 8× (X5),Up 4× (X6),Up 2× (X7),X8) (21)

[0127] output=Cov 3×3 (Up 4× (Cov 1×1 (Cat))) (22)

[0128] Among them, Cat is the result of splicing the upsampled feature maps; 8x upsampling; 4x upsampling.

[0129] S2-5: We use a hybrid loss function to train DSSAU-Net, which combines the dice loss (L dice ) and cross entropy loss (L ce ). Dice loss helps alleviate the problem of class imbalance, while cross entropy loss ensures the accuracy of pixel classification. The loss function L is defined as follows:

[0130]

[0131] Among them, pi is the predicted probability of the i-th pixel, gi is the true value of the i-th pixel, and N is the number of pixels.

[0132] The parameters of DSSAU-Net are updated in reverse according to the L value. When the L value no longer decreases, the trained DSSAU-Net is obtained.

[0133] Experimental setup

[0134] 1. Dataset

[0135] The original dataset contains 2575 training images and 40 validation images. During training, all images are resized to 256×256.

[0136] 2. Evaluation Metrics

[0137] Three main evaluation indicators are used to measure the segmentation performance of the model, namely Dice similarity coefficient (DSC), Hausdorff distance (HD) and average surface distance (ASD). The specific calculation method is as follows:

[0138]

[0139] 3. Experimental Setup

[0140] The experiments were conducted on a GeForce RTX 3090 GPU with 24GB memory. The details of the experiments are carefully designed. Specifically, the input image is resized to 256×256 and the number of regions in each DSSA stage is set to 8×8. In the encoder, the number of channels in each stage is set to [96, 192, 384, 768] and the C d is set to 64. In addition, the initial learning rate is set to 1e-4. In order to improve the generalization ability of the model, several data augmentation strategies such as rotation, flipping and contrast adjustment are introduced. The backbone network uses weights pre-trained on the ImageNet dataset. The scaling factor λ is set to 1 / 8.

[0141] 4. Experimental Results

[0142] We trained DSSAU-Net on the training set and evaluated the performance of DSSAU-Net on the validation set. The evaluation used three segmentation metrics: Dice similarity coefficient (DSC), Hausdorff distance (HD), and average surface distance (ASD), as well as two key parameters related to ultrasound-assisted diagnosis: angle of labor progression (AoP) and head-pubic distance (HSD). The results are shown in Table 1. It can be seen that DSSAU-Net achieved good segmentation results, with DSC, HD, and ASD reaching 86.43, 31.08, and 8.39, respectively, while the number of parameters and floating-point operations (FLOPs) were low. This reflects the effectiveness of our designed U-shaped hierarchical network in the pubic symphysis and fetal head segmentation tasks, as well as the effectiveness of the dual sparse selective attention (DSSA) mechanism in reducing the consumption of computational resources.

[0143] Table 1: Segmentation performance of DSSAU-Net on the validation set of the Intrapartum Ultrasound Challenge 2024

[0144]

[0145] 5. Ablation Experiment

[0146] In order to deeply explore the specific impact of each component of DSSAU-Net on the overall performance, we follow the above experimental settings and conduct a series of ablation studies on the provided training and validation sets. Specifically, we systematically analyze the impact of different factors, including the number of skip connections, the multi-scale feature fusion module, and the choice of key hyperparameters in the proposed dual sparse selective attention.

[0147] Effect of the number of skip connections: Skip connections help compensate for information loss during downsampling, which has been proven in previous studies. However, different network structures have different complexities, and too many skip connections will not only be detrimental to segmentation performance, but also increase the complexity of the network. Therefore, ablation experiments were performed at resolutions of 1 / 4, 1 / 8, and 1 / 16 to explore the impact of different numbers of skip connections on the performance of DSSAU-Net. The results are shown in Table 2. When no skip connections are used, the present invention performs the worst on all segmentation metrics and biometric parameters. As the number of skip connections increases, the segmentation performance of the present invention gradually improves. The best performance can be obtained when all skip connections are used at three resolutions. This is because multi-scale skip connections are beneficial to compensate for the loss of spatial information caused by downsampling and integrate high-level semantic information from deeper levels into the corresponding decoder layers, thereby improving segmentation performance. Therefore, we report the final results based on using all skip connections.

[0148] Table 2: Ablation experiments on the number of skip connections.

[0149]

[0150] Effect of Multi-scale Feature Fusion Module: The multi-scale feature fusion (MFF) module is the core component of DSSAU-Net and is used to achieve accurate segmentation. Therefore, an ablation experiment is designed to analyze its effectiveness. The results are shown in Table 3. It is obvious that compared with DSSAU-Net without the MFF module, the performance of DSSAU-Net with the MFF module is improved by 0.37, 1.2, and 0.81 in DSC, HD, and ASD, respectively. There are also consistent improvements in AoP and HSD. The reason is that the shallow feature map has a higher resolution and contains more local spatial information, while the deep feature map has a lower resolution and covers rich global semantic information. The MFF module can take advantage of these two advantages to achieve more accurate segmentation.

[0151] Table 3: Ablation study of the multi-scale feature fusion (MFF) module.

[0152]

[0153] The influence of the number of Top-k Tokens: The dual sparse selective attention mechanism in DSSA not only significantly reduces the computational complexity, but also can extract accurate features. However, selecting different numbers of region-level tokens and pixel-level tokens may affect the performance of DSSAU-Net. Therefore, an ablation study is designed to select appropriate k1 and λ values ​​for region-level tokens and pixel-level tokens. In order to make a fair comparison, this experiment does not use pre-trained weights. The results are shown in Table 4. It can be seen that when k1 is set to [1, 4, 16, 64], it has better segmentation performance than [2, 8, 32, 64]. We analyze that this is because selecting too many regions may introduce more noise tokens, which is not conducive to learning effective features. In addition, we fix k1 to [1, 4, 16, 64] and further study the influence of λ value. It can be seen that when λ is set to 1 / 8, the model can achieve the best performance. When k1 is set to [2, 8, 32, 64] and λ is set to 1 / 8, there is a similar phenomenon, and the model performs better than other λ values. This is easy to understand, because when fewer tokens are selected during the pixel-level sparse selection process, the model learns fewer features. On the contrary, selecting more tokens will introduce more noise. In addition, it can be seen from Table 5 that the number of floating-point operations (FLOPs) varies significantly with different λ values, which proves the advantage of DSSA in reducing computational complexity. Therefore, we choose a compromise λ value as the hyperparameter of the model.

[0154] Table 4: Ablation study of the number of top-k tokens

[0155]

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.

Claims

1. A hybrid network method for pubic symphysis and fetal head segmentation, characterized in that: The steps include: S1: Acquire several ultrasound images during production as a data set; S2: Construct a U-shaped hybrid network DSSAU-Net for pubic symphysis and fetal head segmentation, and use a supervised method to train DSSAU-Net to obtain a trained DSSAU-Net; The DSSAU-Net includes an encoder, a pyramid pooling module, a decoder, a 1×1 convolutional layer, a 4× upsampling layer, and a 3×3 convolutional layer; wherein the encoder includes four stages from stage 1 to stage 4, and stages 1 to stage 4 include 2, 2, 8, and 2 DSSA blocks respectively, which are denoted as DSSAblock; the decoder also includes four stages from stage 5 to stage 8, and stages 5 to stage 8 include 2, 2, 8, and 2 DSSA blocks respectively, and DSSA blocks are denoted as DSSAblock; S3: For a delivery ultrasound image, the delivery ultrasound image is input into the trained DSSAU-Net, and the output is an image with the pubic symphysis and fetal head segmented.

2. A hybrid network method for pubic symphysis and fetal head segmentation as claimed in claim 1, characterized in that: The DSSA block includes the following steps: 1) For each image in the dataset, use 3×3 deep convolution to encode the relative position information to obtain the encoded image, and then perform layer normalization; 2) Apply DSSA to compute attention on the encoded image, followed by another layer of normalization; 3) Input the normalized result of 2) into a two-layer multi-layer perceptron MLP module; The DSSA comprises the following steps: for a given input image X of size H×W×C, X is divided into S×S non-overlapping regions to obtain a region image The region tokens are obtained by linear projection: the query, key and value are denoted as Q, K, Among them, S represents the number of regions, r represents the token is regional level, and W q , W k , W v ∈R C×C are the projection weight matrices corresponding to the query, key, and value respectively; For region-level sparse token selection, each region-level token query and key It is obtained by averaging all pixel-level tokens on the region token Q and key K; Q r and K r Attention mapping between Represents the semantic relevance between regions; a top-k1 operation is used to select the k1 most relevant regions for each region containing a given pixel-level token and record their indices in In , the process is described as follows: A r =Q r (K r ) Τ (2) I r =topkIndex(A r ) (3) Aggregate the filtered region-level tokens into a matrix; Where K g , is the key tensor matrix and value tensor matrix obtained by aggregation; Regarding pixel-level sparse token selection, for any pixel-level query in Q, calculate the difference between the query and K g The correlation between the pixel-level keys aggregated in is the attention matrix Expressed as The specific calculation method is as follows: Where p means that the token is pixel-level; exist In the second round of pixel-level sparse token selection, the top-k2 operation is used, and the value and index storage corresponding to the selected token are stored and represented as and Among them, k2 is determined by the proportional factor λ: According to I p From V g In the polymerization V gg =gather(V g ,I p ) (9) V gg Use A p After weighting, a local context enhancement term LCE(V) is added to the output matrix O: O=Attention(A p ,V gg )+LCE(V) (10)。 3. A hybrid network method for pubic symphysis and fetal head segmentation as claimed in claim 2, characterized in that: The DSSA block uses three residual connections to help alleviate the gradient vanishing problem. The first residual connection is to add the feature map before deep convolutional coding to the feature map after deep convolutional coding element by element; the second residual connection is to add the feature map before the first layer normalization to the feature map after DSSA element by element; the third residual connection is to add the feature map before the second layer normalization to the feature map after two layers of multi-layer perceptron MLP modules element by element; Among them, z l-1 represents the output after the l-1th DSSA block, and It represents the result of element-by-element addition of the feature map output by DSSA and the feature map before the first layer normalization.

4. A hybrid network method for pubic symphysis and fetal head segmentation as claimed in claim 3, characterized in that: The process of S2 training DSSAU-Net is as follows: S2-1: For stage 1, it consists of an overlapping block embedding layer with two 3×3 convolutions and two DSSA blocks to transform the input image of size H×W×3 into a The feature image X1: X1=DSSAblock ×2 (embedding(X)) (13) In stage 2, a block merging layer is used to reduce the number of tokens, and a linear embedding layer is used to increase the dimension C to 2C; the size of the feature map X2 output by stage 2 containing two DSSA blocks is X2=DSSAblock ×2 (Merging(X1)) (14) The sizes of the corresponding feature maps X3 and X4 generated in stage 3 and stage 4 are and They are respectively expressed as: X3=DSSAblock ×8 (Merging(X2)) (15) X4=DSSAblock ×2 (Merging(X3)) (16) Among them, DSSAblock ×2 and DSSAblock ×8 They represent stacking 2 DSSA blocks and 8 DSSA blocks, embedding (X) block embedding layer, and merging (·) block merging layer respectively; S2-2: Use the pyramid pooling module PPM to map X4 into a The feature map of S2-3: Size The feature maps are input to the decoder constructed hierarchically from stage 5 to stage 8; in this process, the feature maps of each stage are processed by the DSSA block and fused with the feature maps of the same resolution from the encoder; the resulting feature maps are then upsampled to the output size The whole process is described as: X5=DSSAblock ×2 (PPM(X4)) (17) in, Indicates element level addition, PPM(·), Con 1×1 、Up 2× They represent the pyramid pooling module, 1x1 convolution operation, and double upsampling respectively; S2-4: X5, X6, and X7 are upsampled to the same resolution as X8; all these upsampled feature maps are concatenated along the channel dimension, and then a 1×1 convolution layer is applied, followed by a 4× upsampling, and then a 3×3 convolution layer is applied; finally, a feature map with a resolution of H×W×class is output to predict pixel-level segmentation; the process is expressed as: Cat=Concat(Up 8× (X5),Up 4× (X6),Up 2× (X7),X8) (21) output=Number 3×3 (Up 4× (The 1×1 (Cat))) (22) Among them, Cat is the result of concatenating the upsampled feature maps; 8x upsampling; 4x upsampling; S2-5: Define the loss function L as follows: Among them, p i is the predicted probability of the i-th pixel, g i is the true value of the i-th pixel, N is the number of pixels; The parameters of DSSAU-Net are updated in reverse according to the L value. When the L value no longer decreases, the trained DSSAU-Net is obtained.

Citation Information

Cited By

  • Key structure image segmentation method and system for obstetrical ultrasound

    CN122049379A