Endoscope medical image segmentation method based on progressive fusion network

The progressive fusion network addresses the issues of fine-grained feature loss and local relationship capture in endoscopic image segmentation by integrating noise filtering, boundary awareness, and feature fusion, resulting in improved segmentation accuracy and robustness.

CN120318254APending Publication Date: 2025-07-15ANHUI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510490237.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing medical image segmentation models are prone to miss detailed information in endoscopic images, and traditional convolutional neural networks affect segmentation accuracy when they lose too many fine-grained features. Vision Transformer's self-attention mechanism limits the learning of local relationships, resulting in unclear segmentation boundaries.

Method used

The method based on the progressive fusion network is adopted, including Pvtv2 encoder, noise filtering attention module NFAM, boundary and position perception module BLAM, auxiliary information embedding module AIEM and feature fusion module FFM, and the local continuity of features and global information capture capabilities are enhanced through pyramid structure, overlapping block embedding strategy and improved Attention Gate module.

Benefits of technology

It improves the accuracy of endoscopic medical image segmentation, and the segmentation results are clearer. It can better process the boundary and position information of organs and tissues in endoscopic images. It has stronger robustness and is adapted to various light interferences and organ morphological changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318254A_ABST
    Figure CN120318254A_ABST
Patent Text Reader

Abstract

The invention discloses an endoscope medical image segmentation method based on a progressive fusion network, and relates to the technical field of medical image segmentation. Comprising the following steps: acquiring an endoscope medical image data set, and dividing into a training set and a verification set in proportion; constructing an endoscope medical image segmentation model based on a progressive fusion network; inputting the training set into an endoscope medical image segmentation model based on a progressive fusion network for model training; inputting the verification set into a trained endoscopic medical image segmentation model based on a progressive fusion network for model evaluation; and performing primary processing on a to-be-segmented medical image, and inputting the to-be-segmented medical image into the verified endoscopic medical image segmentation model based on the progressive fusion network for image segmentation. The method can effectively solve the problem that a current medical image segmentation model is easy to omit detailed information, and improves the capability of capturing global information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and more specifically, to an endoscopic medical image segmentation method based on a progressive fusion network. Background Art

[0002] Medical Image Segmentation (MIS) is a medical image analysis technology based on computer vision and plays a crucial role in clinical diagnosis. Medical image segmentation can accurately identify target regions in medical images, such as tissues, lesions, and organs, etc., to assist in diagnosis and treatment plan formulation. Endoscopic images are also a common type of medical images. Doctors can directly observe the changes in the diseased parts of patients' bodies through endoscopic images. Currently, endoscopes have been widely used in general surgery. Automatic segmentation of endoscopic images can enable doctors to more accurately locate surgical instruments and diseased areas, improving the success rate of surgeries.

[0003] In recent years, deep learning has promoted significant progress and optimization in medical image segmentation, thus greatly improving the segmentation accuracy of various tasks. The U-Net based on Convolutional Neural Network (CNN) has been widely adopted as a standard framework for medical image segmentation. This framework consists of an encoder and a decoder, both of which are composed of multiple convolutional layers. Shallow and deep semantic information is fused through skip connections. The structure of U-Net has strong scalability, and based on this, various methods such as ResUNet, ResUNet++, UNet++, and Attention-UNet have been established to further improve the performance of U-Net. Convolutional operations can effectively extract local features of images, and the size of the convolutional kernel determines the coverage range of local feature information. Therefore, CNN encounters challenges in containing comprehensive global information existing in images and learning long-range dependencies between pixels. Although consecutive downsampling operations can expand the receptive field and help CNN extract more spatial features, this deep network structure also has certain disadvantages, such as losing fine-grained features and increasing computational complexity. However, in medical image segmentation, endoscopic images are generally high-resolution images. Therefore, traditional Convolutional Neural Network (CNN) will have a greater impact on the segmentation accuracy when losing too many fine-grained features.

[0004] Vision Transformer (ViT) integrates Transformer from natural language processing into computer vision. By utilizing the multi-head self-attention mechanism, ViT can capture global relationships between any two positions. However, the self-attention mechanism in Transformer limits their ability to learn local relationships between pixels, resulting in problems such as unclear segmentation boundaries.

[0005] Therefore, it is an urgent problem for those skilled in the art to propose an endoscopic medical image segmentation method based on a progressive fusion network to solve the difficulties existing in the prior art. Summary of the Invention

[0006] In view of this, the present invention provides an endoscopic medical image segmentation method based on a progressive fusion network, which can effectively solve the problem that the current medical image segmentation model is prone to missing detailed information and improve the ability to capture global information.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] An endoscopic medical image segmentation method based on a progressive fusion network, comprising:

[0009] S1. Obtain an endoscopic medical image data set, and divide the endoscopic medical image data set into a training set and a validation set according to a ratio;

[0010] S2. Construct an endoscopic medical image segmentation model based on a progressive fusion network;

[0011] S3. Input the training set into the endoscopic medical image segmentation model based on a progressive fusion network for model training to obtain a trained endoscopic medical image segmentation model based on a progressive fusion network;

[0012] S4. Input the validation set into the trained endoscopic medical image segmentation model based on a progressive fusion network for model evaluation to obtain a validated endoscopic medical image segmentation model based on a progressive fusion network;

[0013] S5. Input the medical image to be segmented into the validated endoscopic medical image segmentation model based on a progressive fusion network for image segmentation after preliminary processing.

[0014] In the above method, optionally, when constructing the endoscopic medical image segmentation model based on a progressive fusion network in S2, the model includes:

[0015] Pvtv2 encoder, noise filtering attention module NFAM, boundary and position awareness module BLAM, auxiliary information embedding module AIEM, and feature fusion module FFM;

[0016] The Pvtv2 encoder is used to extract multi-layer semantic features of the input image and output feature layers X of different scales i , where i = 1, 2, 3, 4;

[0017] The noise filtering attention module NFAM is used to filter out the background noise of the feature layers X of different scales i and enhance the semantic information to obtain high-quality multi-scale feature layers Ni , where \(i = 1, 2, 3, 4\). Among them, \(N_4\) is a high-level feature containing the position information of organs and tissues, and \(N_1\) is a low-level feature containing texture, color, and boundary features;

[0018] The Boundary and Location Awareness Module (BLAM) is used to obtain the complete boundary and location information of the input image, fuse the position information of organs and tissues in the high-level feature and the boundary information in the low-level feature to obtain a high-quality boundary and location information feature layer \(O\). BLAM ;

[0019] The Auxiliary Information Embedding Module (AIEM) is used to supplement the high-quality boundary and location information feature layer \(O\) BLAM into \(N\) i to enhance the semantic information of each layer of features, obtaining an enhanced feature layer \(B\) i , where \(i = 1, 2, 3, 4\); Upsample the enhanced feature layer \(B\) i so that each feature layer has the same size as \(O\) BLAM ;

[0020] The Feature Fusion Module (FFM) is used to fuse the high-quality boundary and location information feature layer \(O\) BLAM into the enhanced feature layer \(B\) i to gradually restore the global information of organs and tissues starting from the high-level feature, generating the final segmentation result image.

[0021] In the above method, optionally, the Pvtv2 encoder optimizes the structure through a pyramid structure and a linear spatial reduction attention mechanism, and introduces an overlapping block embedding strategy to improve the local continuity of features and the perception ability of boundaries.

[0022] In the above method, optionally, the Noise Filtering Attention Module (NFAM) filters out the background noise of the four layers of feature layers \(X\) with different scales i to enhance the semantic information and obtain a high-quality multi-scale feature layer \(N\) i , specifically:

[0023] Perform multi-scale pyramid pooling operations on the input feature layer \(X\) i , including: using 1×1, 3×3, and 5×5 pyramid pooling for downsampling operations on the feature layer;

[0024] Perform reshaping and connection operations: Obtained by directly reshaping the input feature layer \(X\) i , where \(C\) i is the number of channels of the feature layer \(X\) i ; \(K\) Obtained by reshaping the feature layer after pyramid pooling operations;

[0025] Reshape Q, K, and V after performing attention mechanism calculations;

[0026] The implementation formula of the NFAM module is as follows:

[0027]

[0028] where, represents linear projection, n is the number of heads, and d k is the dimension of each head, equal to

[0029] For the above method, optionally, the boundary and location awareness module BLAM fuses the location information of organs and tissues in the high-level features and the boundary information in the low-level features, specifically:

[0030] Use two 1×1 convolutional layers to reduce the number of channels of the high-level feature T to 64 to obtain T * , and reduce the number of channels of the low-level feature L to 32 to obtain L * ;

[0031] Upsample the feature layer T * to the same size as L * , and then perform a concatenation operation with L * ;

[0032] Through a 3×3 convolutional layer and a 1×1 convolutional layer, after passing through the Sigmoid function, obtain the high-quality boundary and location information feature layer O BLAM ;

[0033] The implementation formula of the BLAM module is as follows:

[0034] BLAM(T, L) = Sigmoid(Conv 3×3 (Conv 1×1 (Connect(L * , Up ×2 (T * )))))

[0035] T * = Conv 1×1 (T)

[0036] L * = Conv 1×1 (L).

[0037] In the above method, optionally, the Auxiliary Information Embedding Module (AIEM) improves the Attention Gate attention module. Group convolution is used instead of conventional convolution for intra-group feature fusion. After convolving the input features, a BatchNorm layer and a ReLU layer are added to improve the structure of the Attention Gate, resulting in the NAG attention module.

[0038] In the above method, optionally, the Auxiliary Information Embedding Module (AIEM) is used to supplement the high-quality boundary and position information feature layer O BLAM into N i to obtain the enhanced feature layer B i , specifically:

[0039] Restore the high-quality boundary and position information feature layer O BLAM to the same number of channels as the feature layer N to which the auxiliary information is to be embedded through a 3×3 convolutional layer;

[0040] Input the two feature layers into the modified NAG attention module to embed the boundary and position information into the feature layer to which the auxiliary information is to be embedded;

[0041] Pass the output four feature layers through the ECA attention module, the SA attention module, and the GHAP attention module respectively to further enhance the fused features;

[0042] The implementation formula of the AIEM module is as follows:

[0043] AIEM(O BLAM , N) = GHPA(Fusion(N, Conv 3×3 (O BLAM )))

[0044] Fusion(N, O BLAM * ) = SA(ECA(NAG(N, O BLAM * ) + N))

[0045] NAG(N, O BLAM * ) = N + (N × (Sigmoid(Conv 1×1 (GBR(N) + GBR(O BLAM * )))))

[0046] GBR(N) = ReLU(BatchNorm(GroupConv 1×1 (N)))

[0047] GBR(O BLAM* ) = ReLU(BatchNorm(GroupConv 1×1 (O BLAM * )))

[0048] Wherein, O BLAM * is obtained by the output of a 3×3 convolutional layer for O BLAM .

[0049] In the above method, optionally, the Feature Fusion Module FFM fuses the high-quality boundary and position information feature layer O BLAM into each layer of feature B i , and gradually restores the global information of organs and tissues starting from the high-level features. Specifically:

[0050] The high-quality boundary and position information feature layer O BLAM is restored to the same number of channels as the input feature layer through a 1×1 convolution to obtain O BLAM ** . The high-level feature layer T1 and the low-level feature layer T2 to be fused are cascaded with it. Among them, T1 is each enhanced layer of feature B i , and T2 is the output of B4 and the previous two FFM modules;

[0051] The cascaded feature layer F is respectively input into the channel attention module CA and the spatial attention module SA to obtain F CA and F SA ;

[0052] The two output feature layers F CA and F SA are added together, and together with the F feature layer, they pass through the pixel attention module PA to identify the boundaries of organs and tissues, and at the same time assign a weight representing its importance to each pixel inside the organs and tissues to obtain a weight layer P;

[0053] The weight P is multiplied by the high-level feature layer T1 containing more global information, and the weight (1 - P) is multiplied by the low-level feature layer T2 containing less global information. The two are added together and output through a 1×1 convolutional layer;

[0054] The implementation formula of the FFM module is as follows:

[0055] FFM(O BLAM , T1, T2) = Conv 1×1 ((T1 × p) + (T2 × (1 - p)))

[0056] p = PA((CA(F) + SA(F)), F)

[0057] F = Cat(O BLAM ** , T1, T2)

[0058] O BLAM ** = Conv 1×1 (O BLAM ).

[0059] Optionally, in the above method, the progressive fusion network adopts weighted IoU loss and weighted binary cross-entropy BCE loss to restrict the prediction map from the global structure and local details;

[0060] The calculation formula of the loss function L between the final segmentation result Predication and the Ground truth is:

[0061]

[0062] where, is the weighted IoU loss, is the weighted binary cross-entropy BCE loss.

[0063] It can be seen from the above technical solutions that, compared with the prior art, the present invention provides an endoscopic medical image segmentation method based on a progressive fusion network, and its beneficial effects are:

[0064] The present invention proposes an endoscopic medical image segmentation method based on a progressive fusion network, which can effectively solve the problem that the current medical image segmentation model is prone to missing detail information, enhance local detail features and improve the ability to capture global information; in the endoscopic image segmentation task of the present invention, it can segment a result with clearer and more complete boundary information, and has stronger robustness to various tissue and organ size and shape changes and various light interferences that may appear in endoscopic images. Description of the Drawings

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0066] Figure 1 is a flowchart of an endoscopic medical image segmentation method based on a progressive fusion network provided by the present invention;

[0067] Figure 2 is an overall view of an endoscopic medical image segmentation method based on a progressive fusion network provided by the present invention;

[0068] Figure 3 The structural diagram of the noise filtering attention module for an endoscopic medical image segmentation method based on a progressive fusion network provided by the present invention;

[0069] Figure 4 The structural diagram of the boundary and position perception module for an endoscopic medical image segmentation method based on a progressive fusion network provided by the present invention;

[0070] Figure 5 The structural diagram of the auxiliary information embedding module for an endoscopic medical image segmentation method based on a progressive fusion network provided by the present invention;

[0071] Figure 6 The structural diagram of the feature fusion module for an endoscopic medical image segmentation method based on a progressive fusion network provided by the present invention. Detailed implementation manners

[0072] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0073] Referring to Figure 1 As shown, the present invention discloses an endoscopic medical image segmentation method based on a progressive fusion network, including;

[0074] S1. Obtain an endoscopic medical image data set, and divide the endoscopic medical image data set into a training set and a validation set according to a ratio;

[0075] S2. Construct an endoscopic medical image segmentation model based on a progressive fusion network;

[0076] S3. Input the training set into the endoscopic medical image segmentation model based on a progressive fusion network for model training to obtain a trained endoscopic medical image segmentation model based on a progressive fusion network;

[0077] S4. Input the validation set into the trained endoscopic medical image segmentation model based on a progressive fusion network for model evaluation to obtain a validated endoscopic medical image segmentation model based on a progressive fusion network;

[0078] S5. Input the medical image to be segmented after preliminary processing into the validated endoscopic medical image segmentation model based on a progressive fusion network for image segmentation.

[0079] Furthermore, referring toFigure 2 As shown in the figure, in S2, an endoscopic medical image segmentation model based on a progressive fusion network is constructed. The model includes:

[0080] Pvtv2 encoder, noise filtering attention module NFAM, boundary and position awareness module BLAM, auxiliary information embedding module AIEM, and feature fusion module FFM;

[0081] The Pvtv2 encoder is used to extract multi-layer semantic features of the input image and output feature layers X of different scales i , where i = 1, 2, 3, 4;

[0082] The noise filtering attention module NFAM is used to filter out the background noise of the feature layers X of different scales i , enhance the semantic information, and obtain high-quality multi-scale feature layers N i , where i = 1, 2, 3, 4. Among them, N4 is the high-level feature, containing the position information of organs and tissues, and N1 is the low-level feature, containing texture, color, and boundary features;

[0083] The boundary and position awareness module BLAM is used to obtain the complete boundary and position information of the input image, fuse the position information of organs and tissues in the high-level feature and the boundary information in the low-level feature, and obtain a high-quality boundary and position information feature layer O BLAM ;

[0084] The auxiliary information embedding module AIEM is used to supplement the high-quality boundary and position information feature layer O BLAM into N i , enhance the semantic information of each layer of features, and obtain the enhanced feature layer B i , where i = 1, 2, 3, 4; Upsample the enhanced feature layer B i so that each feature layer has the same size as O BLAM ;

[0085] The feature fusion module FFM is used to fuse the high-quality boundary and position information feature layer O BLAM into the enhanced feature layer B i , start from the high-level feature, gradually restore the global information of organs and tissues, and generate the final segmentation result image.

[0086] Furthermore, the Pvtv2 encoder optimizes its structure through a pyramid structure and a linear spatial reduction attention mechanism, greatly reducing the resources required for computing attention; an overlapping block embedding strategy is introduced to improve the local continuity of features and the perception ability of boundaries, that is, adjacent image blocks have half of their areas overlapping, so that each block can contain more context information, thereby improving the local continuity of features and the perception ability of boundaries. By optimizing the structure and introducing the overlapping strategy, Pvtv2 not only saves computing resources but also improves the quality of feature extraction.

[0087] Furthermore, inspired by the powerful long-range modeling ability, good enhanced model expression ability, and excellent context understanding ability of the self-attention mechanism in Transformer, in order to enable the model to better capture complete semantic information and filter out background noise, a noise filtering attention module NFAM is proposed. Referring to Figure 3 As shown, the noise filtering attention module NFAM filters out the background noise of four different-scale feature layers X i (i = 1, 2, 3, 4), enhances semantic information, and obtains high-quality multi-scale feature layers N i (i = 1, 2, 3, 4). Specifically:

[0088] Perform multi-scale pyramid pooling operations on the input feature layer X i (i = 1, 2, 3, 4), including: using 1×1, 3×3, and 5×5 pyramid pooling for downsampling operations of the feature layer; while reducing the amount of computation, multi-scale pooling can ensure obtaining receptive fields of different sizes, and can identify organ tissues of different sizes in endoscopic images;

[0089] Perform reshaping and concatenation operations: Obtained directly by reshaping the input feature layer X i , where C i is the number of channels of the feature layer X i ; K, Obtained by reshaping the feature layer after the pyramid pooling operation;

[0090] Reshape after performing attention mechanism calculations on Q, K, and V;

[0091] The implementation formula of the NFAM module is as follows:

[0092]

[0093] Among them, represents linear projection, n is the number of heads, and d k is the dimension of each head, equal to

[0094] Furthermore, although low-level features contain rich edge detail information, they lack global position information at the same time. To effectively extract the boundary and position features of organs and tissues, they need to be fused with the global semantic information contained in high-level features; refer to Figure 4 As shown, the boundary and location awareness module BLAM fuses the location information of organs and tissues in high-level features with the boundary information in low-level features, specifically as follows:

[0095] Use two 1×1 convolutional layers to reduce the number of channels of the high-level feature T to 64 to obtain T * and reduce the number of channels of the low-level feature L to 32 to obtain L * ; Without affecting feature extraction, reduce the computational amount;

[0096] Upsample the feature layer T * to the same size as L * , and then perform a concatenation operation with L * ;

[0097] Through a 3×3 convolutional layer and a 1×1 convolutional layer, and passing through the Sigmoid function, obtain the high-quality boundary and position information feature layer O BLAM ;

[0098] The implementation formula of the BLAM module is as follows:

[0099] BLAM(T,L) = Sigmoid(Conv 3×3 (Conv 1×1 (Connect(L * ,Up ×2 (T * )))))

[0100] T * =Conv 1×1 (T)

[0101] L * =Conv 1×1 (L).

[0102] Furthermore, the Attention Gate module dynamically adjusts weights through the soft-attention mechanism to focus on the regions of interest. However, endoscopic images are high-resolution images, and when the Attention Gate is applied to high-resolution images, it significantly increases the computational burden. In addition, there must be a strict data dependency between the two inputs of the Attention Gate to accurately capture important features. Therefore, the Auxiliary Information Embedding Module (AIEM) improves the Attention Gate attention module by using grouped convolution instead of conventional convolution for intra-group feature fusion, adding a BatchNorm layer and a ReLU layer after convolving the input features to improve the structure of the Attention Gate, resulting in the NAG attention module.

[0103] Furthermore, referring to Figure 5 As shown, the Auxiliary Information Embedding Module (AIEM) is used to supplement the high-quality boundary and position information feature layer O BLAM to N i (i = 1, 2, 3, 4) to obtain the enhanced feature layer B i (i = 1, 2, 3, 4), specifically:

[0104] Restore the high-quality boundary and position information feature layer O BLAM to the same number of channels as the feature layer N where the auxiliary information is to be embedded through a 3×3 convolutional layer;

[0105] Input the two feature layers into the modified NAG attention module to embed the boundary and position information into the feature layer where the auxiliary information is to be embedded;

[0106] Pass the output four feature layers through the ECA attention module, SA attention module, and GHAP attention module respectively to further enhance the fused features;

[0107] The implementation formula of the AIEM module is as follows:

[0108] AIEM(O BLAM , N) = GHPA(Fusion(N, Conv 3×3 (O BLAM )))

[0109] Fusion(N, O BLAM * ) = SA(ECA(NAG(N, O BLAM * ) + N))

[0110] NAG(N, O BLAM *) = N + (N × (Sigmoid(Conv 1×1 (GBR(N) + GBR(O BLAM * )))))

[0111] GBR(N) = ReLU(BatchNorm(GroupConv 1×1 (N)))

[0112] GBR(O BLAM * ) = ReLU(BatchNorm(GroupConv 1×1 (O BLAM * )))

[0113] Wherein, O BLAM * is the output obtained by passing O BLAM through a 3×3 convolutional layer.

[0114] Furthermore, as shown in Figure 6 , the feature fusion module FFM fuses the high-quality boundary and position information feature layer O BLAM into each layer of feature B i , and gradually restores the global information of organs and tissues starting from the high-level features. Specifically:

[0115] Restore the high-quality boundary and position information feature layer O BLAM to the same number of channels as the input feature layer through a 1×1 convolution to obtain O BLAM ** . Concatenate the high-level feature layer T1 and the low-level feature layer T2 that need to be fused with it. Among them, T1 is each enhanced layer of feature B i , and T2 is the output of B4 and the previous two FFM modules;

[0116] Pass the concatenated feature layer F into the channel attention module CA and the spatial attention module SA respectively to obtain F CA and F SA ; Among them, the feature channels representing the global information of organs and tissues are screened out through the channel attention CA, and the feature channels representing the background or noise are ignored; through the spatial attention SA, the approximate positions of organs and tissues are identified, and the blank areas or irrelevant organs and tissues in the image are ignored

[0117] Pass the two output feature layers F CA and F SAAdd them together, and pass through the pixel attention module PA together with the F feature layer to identify the boundaries of organs and tissues, and at the same time assign a weight representing its importance to each pixel inside the organs and tissues, obtaining a weight layer P;

[0118] Multiply the weight P by the high-level feature layer T1 containing more global information, multiply the weight (1 - P) by the low-level feature layer T2 containing less information, add the two together, and output through a 1×1 convolutional layer;

[0119] The implementation formula of the FFM module is as follows:

[0120] FFM(O BLAM ,T1,T2) = Conv 1×1 ((T1×p)+(T2×(1 - p)))

[0121] p = PA((CA(F)+SA(F)),F)

[0122] F = Cat(O BLAM ** ,T1,T2)

[0123] O BLAM ** = Conv 1×1 (O BLAM )。

[0124] Furthermore, the progressive fusion network adopts a weighted IoU loss and a weighted binary cross-entropy BCE loss to restrict the prediction map from the global structure and local details; different from the standard BCE loss function that treats all pixels equally, it considers the importance of each pixel and assigns a higher weight to hard pixels; in addition, compared with the standard IoU loss, it pays more attention to hard pixels;

[0125] The calculation formula of the loss function L between the final segmentation result Predication and the Ground truth is:

[0126]

[0127] Among them, is the weighted IoU loss, is the weighted binary cross-entropy BCE loss.

[0128] In a specific embodiment, a proposed endoscopic medical image segmentation method based on a progressive fusion network is verified and compared on three endoscopic image datasets: the Ureter dataset, the Re-TMRS dataset, and the polyp datasets (Kvasir, CVC-ClinicDB, CVC-ColonDB, ETIS, and CVC-300) to demonstrate the superiority of the endoscopic medical image segmentation method based on a progressive fusion network in endoscopic image segmentation.

[0129] Randomly extract 2292 image data points from Ureter for training, and conduct ureter segmentation test experiments on the remaining 572 image data points; similarly, for the Re-TMRS dataset, we randomly extract 2258 image data points for training and conduct kidney tumor segmentation test experiments on the remaining 565 image data points; in the polyp segmentation experiment, randomly select 900 and 550 polyp images from Kvasir and CVC-ClinicDB for model training, and the remaining 100 and 62 images of Kvasir and CVC-ClinicDB are respectively used to test the learning ability of the model; in addition, all images of CVC-ColonDB, ETIS, and CVC-300 do not participate in model training and are only used to test the generalization ability of the model.

[0130] The network of a proposed endoscopic medical image segmentation method based on a progressive fusion network is implemented based on the Pytorch 2.0.0 and Python 3.8.0 frameworks; all models are trained on an NVIDIA A800 GPU with 80G of memory; the learning rate and weight decay of the AdamW optimizer are adopted; a unified training strategy is used. First, we resize the input image size, and then use a multi-scale strategy of [0.75, 1.0, 1.25] to enable the network to process tissue organs of different sizes. In addition, the batch size of the network is set to 16, and the maximum number of training epochs is set to 100.

[0131] Experiments are conducted on seven datasets of Ureter, Re-TMRS, Kvasir, CVC-ClinicDB, CVC-ColonDB, ETIS, and CVC-300 to verify the effectiveness of the network of the present invention and compare them with previous SOTA methods, including U-Net, PraNet, TransUNet, SwinUnet, TGDAUNet, DA-TransUNet, MSA2Net, NPD-Net, CIFG-Net, and VMUNet.

[0132] A network of an endoscopic medical image segmentation method based on a progressive fusion network according to the present invention outperforms all previous SOTA methods in terms of performance on the kidney dataset and the ureter dataset. At the same time, the mDice, mIoU, and HD evaluation metrics on four publicly available polyp datasets of Kvasir, CVC-ColonDB, CVC-ClinicDB, and CVC-300 are also better than previous SOTA methods. Although the network of the present invention does not outperform the previous SOTA method on the ETIS dataset, the gap is small. Generally speaking, a network of an endoscopic medical image segmentation method based on a progressive fusion network according to the present invention has an overall better performance than existing endoscopic image segmentation networks on 7 endoscopic image datasets.

[0133] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other.

[0134] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An endoscopic medical image segmentation method based on a progressive fusion network, characterized in that, including; S1. Obtain an endoscopic medical image dataset, and divide the endoscopic medical image dataset into a training set and a validation set according to a ratio; S2. Construct an endoscopic medical image segmentation model based on a progressive fusion network; S3. Input the training set into the endoscopic medical image segmentation model based on the progressive fusion network for model training to obtain a trained endoscopic medical image segmentation model based on the progressive fusion network; S4. Input the validation set into the trained endoscopic medical image segmentation model based on the progressive fusion network for model evaluation to obtain a validated endoscopic medical image segmentation model based on the progressive fusion network; S5. Input the medical image to be segmented into the validated endoscopic medical image segmentation model based on the progressive fusion network for image segmentation after preliminary processing.

2. The endoscopic medical image segmentation method based on a progressive fusion network according to claim 1, wherein in S2, an endoscopic medical image segmentation model based on a progressive fusion network is constructed, and the model includes: Pvtv2 encoder, noise filtering attention module NFAM, boundary and position awareness module BLAM, auxiliary information embedding module AIEM, and feature fusion module FFM; The Pvtv2 encoder is used to extract multi-layer semantic features of the input image and output feature layers X of different scales i , where i = 1, 2, 3, 4; The Noise Filtering Attention Module (NFAM) is used to filter out the background noise of the feature layers X at different scales, enhance the semantic information, and obtain the high-quality multi-scale feature layers N i so as to enhance the semantic information and obtain high-quality multi-scale feature layers N i where i = 1, 2, 3, 4. Among them, N4 is the high-level feature, containing the position information of organs and tissues, and N1 is the low-level feature, containing texture, color, and boundary features; The boundary and location awareness module BLAM is used to obtain the complete boundary and location information of the input image, fuse the location information of organs and tissues in the high-level features with the boundary information in the low-level features, and obtain a high-quality boundary and location information feature layer O BLAM ; The Auxiliary Information Embedding Module AIEM is used to supplement the high-quality boundary and position information feature layer O BLAM to N i to enhance the semantic information of each layer of features and obtain the enhanced feature layer B i , where i = 1, 2, 3, 4; upsample the enhanced feature layer B i so that each feature layer has the same size as O BLAM ; The Feature Fusion Module (FFM) is used to fuse the high-quality boundary and position information feature layer O BLAM into the enhanced feature layer B i and gradually recover the global information of organs and tissues starting from the high-level features to generate the final segmentation result image.

3. The endoscopic medical image segmentation method based on a progressive fusion network according to claim 2, wherein the Pvtv2 encoder optimizes the structure through a pyramid structure and a linear spatial reduction attention mechanism, and introduces an overlapping block embedding strategy to improve the local continuity of features and the perception ability of boundaries.

4. The endoscopic medical image segmentation method based on a progressive fusion network according to claim 2, wherein The Noise Filtering Attention Module (NFAM) filters out the background noise of four different-scale feature layers X i to enhance semantic information and obtain high-quality multi-scale feature layers N i Specifically, For the input feature layer X i perform multi-scale pyramid pooling operations, including: performing downsampling operations on the feature layer using 1×1, 3×3, and 5×5 pyramid pooling; Perform reshaping and concatenation operations: Obtained directly by reshaping the input feature layer X i where C i is the number of channels of the feature layer X i ; K, Obtained by reshaping the feature layer after performing pyramid pooling operation; perform attention mechanism calculation on Q, K, and V and then reshape; The implementation formula of the NFAM module is as follows: Among them, represents a linear projection, n is the number of heads, and d k is the dimension of each head, equal to 5. The endoscopic medical image segmentation method based on a progressive fusion network according to claim 2, wherein the boundary and position awareness module BLAM fuses the position information of organs and tissues in the high-level features and the boundary information in the low-level features, specifically: Reduce the number of channels of the high-level feature T to 64 using two 1×1 convolutional layers to obtain T * , reduce the number of channels of the low-level feature L to 32 to obtain L * ; Upsample the feature layer T * to the same size as L * and then perform a concatenation operation with L * ; Through a 3×3 convolutional layer and a 1×1 convolutional layer, and passing through the Sigmoid function, a high-quality boundary and position information feature layer O is obtained BLAM ; The implementation formula of the BLAM module is as follows: BLAM(T,L) = Sigmoid(Conv 3×3 (Conv 1×1 (Connect(L * ,Up ×2 (T * ))))) T * = Conv 1×1 (T) L * = Conv 1×1 (L).

6. The endoscopic medical image segmentation method based on a progressive fusion network according to claim 2, wherein the auxiliary information embedding module AIEM improves the Attention Gate attention module, uses grouped convolution instead of conventional convolution for intra-group feature fusion, adds a BatchNorm layer and a ReLU layer after convolving the input features to improve the structure of the Attention Gate, and obtains the NAG attention module.

7. The endoscopic medical image segmentation method based on a progressive fusion network according to claim 6, wherein The Auxiliary Information Embedding Module (AIEM) is used to supplement the high-quality boundary and position information feature layer O BLAM to N i to obtain the enhanced feature layer B i Specifically, it is as follows: Restore the high-quality boundary and position information feature layer O BLAM to the same number of channels as the feature layer N to which the auxiliary information is to be embedded through a 3×3 convolutional layer; input two feature layers into the modified NAG attention module, and embed the boundary and position information into the feature layer to which the auxiliary information is to be embedded; respectively pass the output four feature layers through the ECA attention module, the SA attention module, and the GHAP attention module to further enhance the fused features; The implementation formula of the AIEM module is as follows: AIEM(O BLAM ,N) = GHPA(Fusion(N,Conv 3×3 (O BLAM ))) Fusion(N,O BLAM * ) = SA(ECA(NAG(N,O BLAM * ) + N)) NAG(N,O BLAM * ) = N + (N × (Sigmoid(Conv 1×1 (GBR(N) + GBR(O BLAM * )))) GBR(N) = ReLU(BatchNorm(GroupConv 1×1 (N))) GBR(O BLAM * ) = ReLU(BatchNorm(GroupConv 1×1 (O BLAM * ))) Among them, O BLAM * is O BLAM obtained by output of a 3×3 convolutional layer.

8. A method for endoscopic medical image segmentation based on a progressive fusion network according to claim 2, characterized in that The Feature Fusion Module (FFM) fuses the high-quality boundary and position information feature layer O BLAM into each layer of feature B i and gradually restores the global information of organs and tissues starting from the high-level features. Specifically: Restore the high-quality boundary and position information feature layer O BLAM to the same number of channels as the input feature layer through a 1×1 convolution O BLAM ** , and perform a concatenation operation on the high-level feature layer T1 and the low-level feature layer T2 that need to be fused with it, where T1 is each enhanced feature layer B i , and T2 is the output of B4 and the first two FFM modules; The cascaded feature layer F is respectively input into the channel attention module CA and the spatial attention module SA to obtain F CA and F SA ; Add the feature layers F of the two outputs CA and F SA together, pass them through the pixel attention module PA together with the F feature layer to identify the boundaries of organs and tissues, and at the same time assign a weight representing its importance to each pixel inside the organs and tissues to obtain a weight layer P; Multiply the weight P by the high-level feature layer T1 containing more global information, multiply the weight (1 - P) by the low-level feature layer T2 containing less information, add the two, and output through a 1×1 convolutional layer; The implementation formula of the FFM module is as follows: FFM(O BLAM ,T1,T2) = Conv 1×1 ((T1 × p)+(T2 × (1 - p))) p = PA((CA(F)+SA(F)),F) F = Cat(O BLAM ** , T1, T2) O BLAM ** = Conv 1×1 (O BLAM )。 9. A method for endoscopic medical image segmentation based on a progressive fusion network according to claim 2, characterized in that The progressive fusion network uses weighted IoU loss and weighted binary cross-entropy BCE loss to restrict the prediction map from the global structure and local details; The calculation formula of the loss function L between the final segmentation result Predication and Ground truth is: Among them, is the weighted IoU loss, is the weighted binary cross-entropy BCE loss.

Citation Information

Cited By

  • Multi-modal data-based reply method, electronic equipment and readable storage medium

    CN121811423A