Double-path hybrid network method for ultrasonic boundary perception segmentation and progress angle measurement during production

Through the dual-path hybrid network method, combined with convolutional neural network and lightweight Transformer branch, the attention collapse caused by insufficient ultrasound image data during production is solved, the segmentation accuracy of the fetal head and pubic junction boundary is improved, and more accurate progress angle measurement is achieved.

CN120510171APending Publication Date: 2025-08-19GUANGZHOU LIAN MED TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510314336.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Traditional visual Transformers are prone to attention collapse when there is insufficient ultrasound image data during delivery, resulting in inaccurate separation of the fetal head and pubic junction boundary, affecting the accuracy of progression angle measurement.

Method used

The dual-path hybrid network method is adopted, combining convolutional neural networks and lightweight CNN-style Transformer branches, feature fusion and refinement are performed through the T2C module and the BARM module, and the training data set is enhanced by using a random cutting algorithm of self-written boundary areas to improve the feature capture and segmentation accuracy of boundary areas.

Benefits of technology

It improves the accuracy of boundary prediction of ultrasound images in production, solves the problem of attention collapse, enhances the fine detail recovery ability of the image, and improves the accuracy of boundary segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510171A_ABST
    Figure CN120510171A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of ultrasonic image segmentation, and relates to a double-path hybrid network method for ultrasonic boundary perception segmentation and progress angle measurement during production, which comprises the following steps: S1, enhancing a training data set by using a random cutting hybrid algorithm for marking a boundary region; s2, a convolutional neural network branch is adopted to capture local features, a lightweight CNN style Transform branch is adopted to capture a long-range dependency relationship, and after the convolutional neural network branch and the lightweight CNN style Transform branch are subjected to parallel calculation, a T2C module is adopted to perform fusion so as to generate fusion features; s3, performing jump connection of branches of the convolutional neural network by adopting a context enrichment module; s4, using a T2C module to fuse the CNN branch features and the Transform branch features; and S5, refining the features by using the foreground information, the background information and a residual structure by using a BARM module. According to the method and the device, the problem of attention collapse easily caused by a traditional visual Transform structure under the condition of insufficient data is solved, and the boundary prediction precision is improved, so that the fine details of the image are better recovered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of ultrasound image segmentation, and in particular to a dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement. Background Art

[0002] During intrapartum care, assessing labor progress is crucial to ensuring a smooth delivery and timely intervention, helping to reduce risks to the mother and fetus. Fetal head (FH) positioning is a key parameter in labor assessment. Although the World Health Organization recommends vaginal examination for this purpose, traditional digital examinations are subjective, imprecise, and invasive, often causing discomfort. However, intrapartum ultrasound (IU) offers greater accuracy. The International Society of Ultrasound in Obstetrics and Gynecology and the World Association of Midwifery both emphasize the objectivity and accuracy of intrapartum ultrasound in assessing fetal position and recommend the angle of progression (AoP) as a key indicator. Clinically, an AoP > 120° generally indicates that vaginal delivery is possible, while an AoP < 120° may necessitate a cesarean section. Accurate AoP measurement relies on precise segmentation of the fetal head and pubic symphysis (PS).

[0003] Currently, intrapartum AoP measurement relies primarily on manual assessment of ultrasound image composition, echo patterns, and edge features. This process is time-consuming, subjective, and susceptible to interobserver variability. Unlike prenatal ultrasound imaging, which is performed by trained sonographers and offers high image clarity, intrapartum ultrasound imaging is performed during labor by midwives or nurses in emergency situations. Therefore, intrapartum ultrasound images are susceptible to artifacts and noise, especially as the boundaries between the fetal head and pubic symphysis blur due to fetal movement and physiological changes. Accurate segmentation is crucial for automated AoP measurement. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to propose a dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progress angle measurement to solve the attention collapse problem that is prone to occur in traditional vision transformers (ViT) when there is insufficient data, while further improving the accuracy of boundary prediction, thereby better restoring the fine details of the image.

[0005] To solve the above technical problems, the present invention provides a dual-path hybrid network method for intrapartum ultrasound boundary perception and segmentation and progression angle measurement, which adopts the following technical solutions:

[0006] A dual-path hybrid network approach for intrapartum ultrasound boundary-aware segmentation and progression angle measurement follows an encoder and decoder framework, where the decoder consists of a convolutional neural network branch and a lightweight CNN-style Transformer branch, including the following steps:

[0007] S1, using random cut hybrid algorithm to mark boundary areas to enhance the training dataset;

[0008] S2: Use the convolutional neural network branch to capture local features, and use the lightweight CNN-style Transformer branch to capture long-range dependencies. After the convolutional neural network branch and the lightweight CNN-style Transformer branch are calculated in parallel, they are fused using the T2C module to generate fused features;

[0009] S3, using context enrichment module to perform skip connections between branches of convolutional neural network;

[0010] S4, using T2C module to fuse CNN branch features and ViT branch features;

[0011] S5. Use the BARM module to utilize foreground and background information, and use the residual structure to refine features.

[0012] Furthermore, the convolutional neural network branch uses Res2Net as the backbone network structure, which consists of four residual blocks to capture the local features of the image. The generated features are recorded as C i .

[0013] Furthermore, the lightweight CNN-style Transformer branch adopts a pooling mechanism to replace the tokenization process in ViT. For the input image I:

[0014] First, a 3×3 convolution is performed to extract local features. To match the patch size S in ViT, the input features are downsampled to the corresponding resolution through pooling. The number of downsampling operations t is calculated as follows:

[0015] t=log2S

[0016] Among them, a 2×2 maximum pooling layer is used, and each pooling layer is followed by a 3×3 convolution operation. After a series of pooling operations, the size of the input I becomes I p ∈R c ×H / 2 t ×W / 2 t , where c corresponds to the embedding dimension in ViT;

[0017] Among them, a CNN-style self-attention mechanism is constructed, and query, key, and value features are extracted respectively through three parallel 3×3 convolutional layers. Customized convolution operations are used to adaptively adjust the receptive field, so that the model can better capture long-distance dependencies. For each pixel p in the input feature map P, ij , query Q i,j and key K i,j The calculation formula is as follows:

[0018]

[0019] Among them, L q and L k is a learnable projection matrix, the index l and g range from -1 to 1, covering the p ij The 3×3 local area centered on the initial customized convolution kernel M i,j The calculation is done by the cosine similarity formula:

[0020]

[0021] in, Corresponding to the attention distribution in ViT, E d Represent the embedding dimensions of Q, K, and V. Then, a scientific Gaussian distance map G is used to dynamically determine the size of the custom convolution kernel:

[0022]

[0023] Among them, α is used to control the trend of the receptive field and needs to be set in advance. θ is used to control the variable receptive field and is proportional to the receptive field. When training is performed in the network, when α is enlarged, the size of the custom convolution kernel will gradually approach the global receptive field as the receptive field expands. The convolution kernel A i,j Obtained by multiplying matrices G and M, ensuring that each pixel p i,j It has an adaptively sized convolution kernel A, which is multiplied by V to capture long-range dependencies, where the calculation process of V is similar to that of Q:

[0024]

[0025] Secondly, the generated result undergoes a 1×1 convolution operation to generate the learned features;

[0026] Finally, the generated features are further optimized by another 1×1 convolution. Four parallel lightweight ViT structures are used in this branch, corresponding to t=2, 3, 4, 5, with patch sizes of 4, 8, 16, and 32 respectively. The outputs of the four branches are aggregated in the channel dimension and the final output F is obtained by a 1×1 convolution. t .

[0027] Furthermore, the T2C module uses spatial attention and channel attention to fuse features.

[0028] Furthermore, in the channel dimension, the T2C module first applies a global average pooling layer to assign weight information to the channel, and then two dot convolutions serve as channel context aggregators to further process features. In the spatial dimension, the T2C module applies 5×5 and 7×7 depth-separable convolutions to extract multi-scale feature information. Finally, the features extracted from the two branches are aggregated, and the Sigmoid function is used to generate weights to adjust the features of the CNN branch.

[0029] Furthermore, the BARM module uses an inverse additive attention mechanism to analyze and fuse the foreground and background features of the global features of the high-resolution details from the encoder, and uses a residual structure to gradually optimize the segmentation details. The BARM module requires two inputs, including the mixed feature T5 output from the T2C module and the side output C from each layer of the CNN branch of the encoder. i , global feature T i First, it is processed by the sigmoid function and rounded, and then 1 is subtracted to obtain the background feature B i The specific process is as follows:

[0030] B i =1-Sigmoid(T i )

[0031] The output features BAi of the boundary attention block are as follows:

[0032] BA i =C i ⊙B i +T i+1

[0033] Where ⊙ represents element-wise multiplication

[0034] Use a residual structure to combine the generated boundary attention feature BAi with the previous module output T i+1 Aggregate to form updated features T i , the process is as follows:

[0035] T i =BA i +up(T i+1 )

[0036] Here, up(·) represents upsampling.

[0037] Furthermore, the random cutting algorithm of the self-edited boundary area includes the following steps:

[0038] The input image is divided into N×N blocks, and each block i is assigned a weight w i , to reflect its boundary importance, the weight calculation is defined as follows:

[0039]

[0040] Among them, p i,j represents the activation probability of the jth pixel in the i-th block, H and W represent the height and width of the image respectively, C is the total number of segmentation targets, and p i,j (c) is the segmentation probability of the jth pixel in the i-th block for the c-th target.

[0041] Furthermore, the loss function of the random cutting algorithm of the self-written boundary area uses a combination of the cross entropy loss function and the Dice loss function as the total loss function of the model, and the cross entropy loss L ce The definition is as follows:

[0042]

[0043] Among them, y i is the ground truth label, is the predicted probability, N is the number of samples;

[0044] The Dice loss function Ldice is defined as follows:

[0045]

[0046] Among them, y i represents the ground truth label, Represents the predicted probability, N is the number of samples, and the Dice loss function optimizes the segmentation performance of the model by measuring the overlap between the predicted results and the ground truth.

[0047] Compared with the existing technology, the embodiments of the present application have the following main beneficial effects: the present application solves the attention collapse problem that is prone to occur in traditional visual Transformers when there is insufficient data, and at the same time further improves the accuracy of boundary prediction, thereby better restoring the fine details of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0049] Figure 1A flowchart of an embodiment of a dual-path hybrid network method for intrapartum ultrasound boundary-aware segmentation and progression angle measurement according to the present application;

[0050] Figure 2 It is the overall structure diagram of this application model;

[0051] Figure 3 This is the flow chart of the data enhancement algorithm of this application;

[0052] Figure 4 This is a visual comparison chart of the dataset used in this application. DETAILED DESCRIPTION

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0054] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0055] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0056] Continue to refer Figure 1-4 As shown, the present invention provides a dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement. The dual-path hybrid follows an encoder and decoder framework, where the decoder consists of a convolutional neural network branch and a lightweight CNN-style Transformer branch, including the following steps:

[0057] Step S1: Enhance the training dataset using a random cut hybrid algorithm that marks boundary areas.

[0058] Step S2: Use the convolutional neural network branch to capture local features, and use the lightweight CNN-style Transformer branch to capture long-range dependencies. After the convolutional neural network branch and the lightweight CNN-style Transformer branch are calculated in parallel, they are fused using the T2C module to generate fused features.

[0059] Step S3: Use the context enrichment module to perform skip connections of the convolutional neural network branches.

[0060] Step S4: Use the T2C module to fuse CNN branch features and ViT branch features.

[0061] Step S5: Use the BARM module to utilize foreground information and background information, and use the residual structure to refine features.

[0062] In this example, the model's training dataset consisted of 1,175 intrapartum ultrasound images from 110 pregnant women of varying ages. These images were acquired via transperineal ultrasound and documented the position of the fetal head at the onset of the second stage of labor. Data collection was conducted at three medical institutions: Nanfang Hospital, Zhujiang Hospital, and the First Affiliated Hospital of Jinan University.

[0063] All experiments were conducted using Python on an NVIDIA RTX A6000 GPU. This example uses the Adam optimizer with an initial learning rate of 1e-4. Furthermore, an early stopping mechanism was implemented during model training; training ceases when the model's loss function fails to decrease for 30 consecutive epochs. The size of all images was set to 256 × 256 pixels. To objectively evaluate the proposed model, four different evaluation metrics were used: average symmetric surface distance (ASSD), volumetric overlap error (VOE), Dice coefficient (DC), and Hausdorff distance (HD).

[0064] In this embodiment, the convolutional neural network branch (CNN) uses Res2Net as the backbone network structure, which consists of four residual blocks to capture the local features of the image. The generated features are recorded as C i ,These features have multi-scale representation capabilities, providing high-quality input for subsequent ,feature fusion.

[0065] Furthermore, the lightweight CNN-style Transformer branch

[0066] To mitigate attention collapse and accelerate model convergence, the Transformer branch was modified to introduce a CNN-like structure. Specifically, a pooling mechanism was used to replace the tokenization process in ViT. Since the input is a two-dimensional image, this approach enables us to preserve position information without the need for the additional position embedding required in traditional ViT.

[0067] For the input image I, a 3×3 convolution is first performed to extract local features. To match the patch size S in ViT, the input features are downsampled to the corresponding resolution through a pooling operation. The number of downsampling operations t is calculated as follows:

[0068] t=log2S

[0069] Among them, a 2×2 maximum pooling layer is used, and each pooling layer is followed by a 3×3 convolution operation. After a series of pooling operations, the size of the input I becomes I p ∈R c ×H / 2 t ×W / 2 t , where c corresponds to the embedding dimension in ViT.

[0070] Among them, a CNN-style self-attention mechanism is constructed, and query (Query, Q), key (Key, K) and value (Value, V) features are extracted respectively through three parallel 3×3 convolutional layers. Customized convolution operations are used to adaptively adjust the receptive field, so that the model can better capture long-distance dependencies. For each pixel p in the input feature map P, ij , query Q i,j and key K i,j The calculation formula is as follows:

[0071]

[0072] Among them, L q and L k is a learnable projection matrix, the index l and g range from -1 to 1, covering the p ij The 3×3 local area centered on the initial customized convolution kernel M i,j The calculation is done by the cosine similarity formula:

[0073]

[0074] in, Corresponding to the attention distribution in ViT, E d Represent the embedding dimensions of Q, K, and V. Then, a scientific Gaussian distance map G is used to dynamically determine the size of the custom convolution kernel:

[0075]

[0076] Among them, α is used to control the trend of the receptive field and needs to be set in advance. θ is used to control the variable receptive field and is proportional to the receptive field. When training is performed in the network, when α is enlarged, the size of the custom convolution kernel will gradually approach the global receptive field as the receptive field expands. The convolution kernel A i,j Obtained by multiplying matrices G and M, ensuring that each pixel p i,j It has an adaptively sized convolution kernel A, which is multiplied by V to capture long-range dependencies, where the calculation process of V is similar to that of Q:

[0077]

[0078] Secondly, the generated result undergoes a 1×1 convolution operation to generate the learned features;

[0079] Finally, the generated features are further optimized by another 1×1 convolution. Four parallel lightweight ViT structures are used in this branch, corresponding to t=2, 3, 4, 5, with patch sizes of 4, 8, 16, and 32 respectively. The outputs of the four branches are aggregated in the channel dimension and the final output F is obtained by a 1×1 convolution. t Because of its lightweight ViT structure, it does not increase the computational burden. The entire branching process is composed of CNNs, which effectively avoids the attention collapse problem caused by the lack of medical data, the small Transformer ratio, and slow convergence during training.

[0080] In this embodiment, the T2C module uses spatial attention and channel attention to fuse features. Since the decoder uses two branches to capture local features and long-range dependencies, it is crucial to effectively combine these two types of information for complementary enhancement. Since the ViT branch is essentially composed of a CNN structure, this embodiment proposes a T2C module to refine and enhance the local features extracted by the CNN branch using the long-range dependency information captured by the ViT branch.

[0081] Specifically, in this embodiment, in the channel dimension, the T2C module first applies a global average pooling layer to assign weight information to the channel, and then two dot convolutions are used as channel context aggregators to further process features. In the spatial dimension, the T2C module applies 5×5 and 7×7 depth-separable convolutions to extract multi-scale feature information. Finally, the features extracted from the two branches are aggregated, and the Sigmoid function is used to generate weights to adjust the features of the CNN branch.

[0082] In this embodiment, the BARM module uses an inverse additive attention mechanism to analyze and fuse the foreground and background features of the global features of the high-resolution details from the encoder, and uses a residual structure to gradually optimize the segmentation details. The BARM module requires two inputs, including the mixed feature T5 output from the T2C module and the side output C from each layer of the CNN branch of the encoder. i , global feature T i First, it is processed by the sigmoid function and rounded, and then 1 is subtracted to obtain the background feature B i The specific process is as follows:

[0083] B i =1-Sigmoid(T i )

[0084] The output features BAi of the boundary attention block are as follows:

[0085] BA i =C i ⊙B i +T i+1

[0086] Here, ⊙ represents element-wise multiplication. This process enables the model to capture and analyze the difference between foreground and background, as some foreground information may be mistakenly predicted as background. By leveraging this difference, the model can more accurately understand boundary details, thereby improving segmentation accuracy.

[0087] Finally, a residual structure is used to combine the generated boundary attention features BAi with the output of the previous module T i+1 Aggregate to form updated features T i , the process is as follows:

[0088] T i =BA i +up(T i+1 )

[0089] Here, up(·) represents upsampling. This residual connection ensures that information is effectively integrated and optimized throughout the network, enhancing the overall feature expression capability.

[0090] In this embodiment, the random cutting algorithm of the self-edited boundary area includes the following steps:

[0091] The input image is divided into N×N blocks, and each block i is assigned a weight w i , to reflect its boundary importance, the weight calculation is defined as follows:

[0092]

[0093] Among them, p i,jrepresents the activation probability of the jth pixel in the i-th block, H and W represent the height and width of the image respectively, C is the total number of segmentation targets, and p i,j (c) is the segmentation probability of the jth pixel in the i-th block for the c-th object. This formulation ensures that blocks with uncertain or ambiguous predictions are given higher weights, thereby prioritizing their contribution during training. Conversely, blocks with accurate and confident predictions are given lower weights. By emphasizing more challenging regions within boundary regions, this method can improve the model's ability to learn and accurately segment complex structures in medical images. Subsequently, blocks with weights greater than or equal to 0.3 are marked as high-difficulty regions to ensure that these regions are not cropped or overwritten in the subsequent random CutMix step. It is important to note that this process only operates on blocks containing boundaries, as even internal target regions or background blocks containing significant artifacts or noise can achieve relatively good segmentation performance. Therefore, this embodiment sets the default weight of all non-boundary blocks to less than 0.3. After ensuring that blocks with weights greater than 0.3 are not cropped or overwritten, this embodiment applies the CutMix method to randomly crop and overwrite the remaining blocks. This method can introduce richer and more complex boundary information into the model training process. When used in combination with BARM, the ability to extract boundary features is further enhanced.

[0094] In this embodiment, the loss function of the random cutting algorithm of the self-edited boundary area uses a combination of the cross entropy loss function and the Dice loss function as the total loss function of the model. The cross entropy loss L ce The definition is as follows:

[0095]

[0096] Among them, y i is the ground truth label, is the predicted probability, N is the number of samples;

[0097] The Dice loss function Ldice is defined as follows:

[0098]

[0099] Among them, y i represents the ground truth label, Represents the predicted probability, N is the number of samples, and the Dice loss function optimizes the segmentation performance of the model by measuring the overlap between the predicted results and the ground truth.

[0100] To demonstrate the model's superior performance, this example compares the proposed model with 10 methods. These include five CNN-based models: UNet, Attention UNet, ACCUNet, MTANet, and Rolling-UNet; and five Transformer-based models: SwinUNet, CTO, H2Former, ScribFormer, and DBRN. To ensure a fair comparison, the same preprocessing steps were applied to all methods.

[0101] The segmentation performance of the proposed method in this example was evaluated on an intrapartum ultrasound image dataset and compared with the results of existing methods, as shown in Table 1. Experimental results show that, without adaptive boundary enhancement (ABE), the proposed method achieved an average Dice score of 0.911, an average HD of 3.401, an average ASSD of 0.686, and an average VOE of 0.177. After introducing ABE, model performance was further improved, with the Dice score, HD, ASSD, and VOE reaching 0.915, 3.353, 0.654, and 0.165, respectively. These results demonstrate that the proposed method outperforms CNN- and Transformer-based methods, with P-values all less than 0.05. Furthermore, the ASSD metric reached 3.401, significantly outperforming other methods. Because ASSD directly measures the average symmetric distance between the predicted and true boundaries, it effectively captures boundary deviations while also accounting for bidirectional errors, making it a reliable reference for evaluating boundary accuracy. Therefore, the improved ASSD metric demonstrates that the proposed method is more effective in recovering boundary details. Furthermore, the last row of the table shows the parameter size of each model. Compared with other models, the proposed model achieves a better balance between performance and parameter efficiency.

[0102] Table 1

[0103]

[0104] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. A dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement, characterized by: This dual-path hybrid follows an encoder and decoder framework, where the decoder consists of a convolutional neural network branch and a lightweight CNN-style Transformer branch, and includes the following steps: S1, using random cut hybrid algorithm to mark boundary areas to enhance the training dataset; S2: Use the convolutional neural network branch to capture local features, and use the lightweight CNN-style Transformer branch to capture long-range dependencies. After the convolutional neural network branch and the lightweight CNN-style Transformer branch are calculated in parallel, they are fused using the T2C module to generate fused features; S3, using context enrichment module to perform skip connections between branches of convolutional neural network; S4, using the T2C module to fuse CNN branch features and Transformer branch features; S5. Use the BARM module to utilize foreground and background information, and use the residual structure to refine features.

2. The dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement according to claim 1, characterized in that: The convolutional neural network branch uses Res2Net as the backbone network structure, which consists of four residual blocks to capture the local features of the image. The generated features are recorded as C i .

3. The dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement according to claim 1, characterized in that: The lightweight CNN-style Transformer branch adopts a pooling mechanism to replace the tokenization process in the traditional Transformer. For the input image I: First, a 3×3 convolution is performed to extract local features. To match the patch size S in the Transformer, the input features are downsampled to the corresponding resolution through pooling. The number of downsampling operations t is calculated as follows: t=log2S Among them, a 2×2 maximum pooling layer is used, and each pooling layer is followed by a 3×3 convolution operation. After a series of pooling operations, the size of the input I becomes I p ∈R c ×H / 2 t ×W / 2 t , where c corresponds to the embedding dimension in Transformer; Among them, a CNN-style self-attention mechanism is constructed, and query, key, and value features are extracted respectively through three parallel 3×3 convolutional layers. Customized convolution operations are used to adaptively adjust the receptive field, so that the model can better capture long-distance dependencies. For each pixel p in the input feature map P, ij , query Q i,j and key K i,j The calculation formula is as follows: Among them, L q and L k is a learnable projection matrix, the index l and g range from -1 to 1, covering the p ij The 3×3 local area centered on the initial customized convolution kernel M i,j The calculation is done by the cosine similarity formula: in, Corresponding to the attention distribution in Transformer, E d Represent the embedding dimensions of Q, K, and V. Then, a scientific Gaussian distance map G is used to dynamically determine the size of the custom convolution kernel: Among them, α is used to control the trend of the receptive field and needs to be set in advance. θ is used to control the variable receptive field and is proportional to the receptive field. When training is performed in the network, when α is enlarged, the size of the custom convolution kernel will gradually approach the global receptive field as the receptive field expands. The convolution kernel A i,j Obtained by multiplying matrices G and M, ensuring that each pixel p i,j It has an adaptively sized convolution kernel A, which is multiplied by V to capture long-range dependencies, where the calculation process of V is similar to that of Q: Secondly, the generated result undergoes a 1×1 convolution operation to generate the learned features; Finally, the generated features are further optimized by another 1×1 convolution. Four parallel lightweight Transformer structures are used in this branch, corresponding to t=2, 3, 4, 5, with patch sizes of 4, 8, 16, and 32 respectively. The outputs of the four branches are aggregated in the channel dimension and the final output F is obtained by a 1×1 convolution. t .

4. The dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement according to claim 1, characterized in that: The T2C module uses spatial attention and channel attention to fuse features.

5. The dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement according to claim 4, characterized in that: In the channel dimension, the T2C module first applies a global average pooling layer to assign weight information to the channel, and then two dot convolutions are used as channel context aggregators to further process the features. In the spatial dimension, the T2C module applies 5×5 and 7×7 depth-separable convolutions to extract multi-scale feature information. Finally, the features extracted from the two branches are aggregated, and the Sigmoid function is used to generate weights to adjust the features of the CNN branch.

6. The dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement according to claim 3, characterized in that: The BARM module uses an inverse additive attention mechanism to analyze and fuse the foreground and background features of the global features of the high-resolution details from the encoder, and uses a residual structure to gradually optimize the segmentation details. The BARM module requires two inputs, including the mixed feature T5 output from the T2C module and the side output C from each layer of the CNN branch of the encoder. i , global feature T i First, it is processed by the sigmoid function and rounded, and then 1 is subtracted to obtain the background feature B i The specific process is as follows: B i =1-Sigmoid(T i )The output features BAi of the boundary attention block are as follows: BA i =C i ⊙B i +T i+1 Where ⊙ represents element-wise multiplication Use a residual structure to combine the generated boundary attention feature BAi with the previous module output T i+1 Aggregate to form updated features T i , the process is as follows: T i =BA i +up(T i+1 ) Here, up(·) represents upsampling.

7. The dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement according to claim 1, characterized in that: The random cutting algorithm of the self-written boundary area includes the following steps: The input image is divided into N×N blocks, and each block i is assigned a weight w i , to reflect its boundary importance, the weight calculation is defined as follows: Among them, p i,j represents the activation probability of the jth pixel in the i-th block, H and W represent the height and width of the image respectively, C is the total number of segmentation targets, and p i,j (c) is the segmentation probability of the jth pixel in the i-th block for the c-th target.

8. The dual-path hybrid network method for intrapartum ultrasound boundary perception segmentation and progression angle measurement according to claim 7, characterized in that: The loss function of the random cutting algorithm of the self-written boundary area uses a combination of the cross entropy loss function and the Dice loss function as the total loss function of the model. The cross entropy loss L ce The definition is as follows: Among them, y i is the ground truth label, is the predicted probability, N is the number of samples; The Dice loss function Ldice is defined as follows: Among them, y i represents the ground truth label, Represents the predicted probability, N is the number of samples, and the Dice loss function optimizes the segmentation performance of the model by measuring the overlap between the predicted results and the ground truth.