Laparoscopic image kidney tumor segmentation method and system based on deep learning
By designing a parallel structure of CNN-Transformer hybrid network PEEA-Net, combined with PVTv2 encoder, CNN assisted encoder and feature enhancement module, the shortcomings of the renal tumor segmentation algorithm in the prior art in capturing fine-grained features and processing brightness unevenness are solved, and more efficient renal tumor segmentation performance and lower surgical complexity are achieved.
Patent Information
- Application Number
- CN202510218873.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The existing deep learning-based renal tumor segmentation algorithm has shortcomings in capturing the fine-grained features of renal tumors and processing the problem of uneven brightness of laparoscopic images, resulting in higher complexity when assisting doctors in treating renal cancer during partial nephrectomy.
A parallel structure of CNN-Transformer hybrid network PEEA-Net is designed to extract global information through the PVTv2 encoder, and the CNN-assisted encoder extracts local detailed information, and combines PEA, ESPA and LDISF modules for feature enhancement and fusion to achieve finer-grained kidney tumor prediction.
PEEA-Net shows higher segmentation performance in the renal tumor segmentation task, and can effectively deal with the problems of different sizes and types of kidney tumors, low contrast between tumors and normal healthy tissues, small kidney tumors, and uneven brightness of laparoscopic images, reducing the complexity of doctors in surgical operations.
Smart Images

Figure CN120147337A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and more specifically, to a method and system for laparoscopic image kidney tumor segmentation based on deep learning. Background Art
[0002] In recent years, renal cancer has become one of the top ten most common cancers globally. With the popularization of the concept of minimally invasive treatment and the development of surgical techniques and surgical instruments, partial nephrectomy (PN) has become the preferred surgical treatment for small renal cancers (maximum diameter ≤ 4 cm). Laparoscopic partial nephrectomy (LPN) and robotic assisted laparoscopic partial nephrectomy (RAPN) have been widely carried out and applied, and LPN and RAPN have become important PN surgical methods globally. In LPN and RAPN, doctors only remove the tumor part of the kidney and retain the normal tissue part of the kidney.
[0003] The automatic segmentation technology of kidney tumors based on deep learning can assist doctors in quickly and accurately identifying the tumor location and improving the clinical treatment process. Therefore, in order to reduce the unknown factors and uncertainties during the operation and reduce the complexity of doctors' operation during the surgery, using deep learning technology to automatically and accurately segment kidney tumors in laparoscopic images to assist doctors in treating renal cancer has important clinical significance. However, well-annotated laparoscopic kidney tumor datasets are scarce, and the acquisition of laparoscopic image data and the annotation work of kidney tumor GT are very expensive and time-consuming; in addition, as Figure 1 shown (the first row is the original laparoscopic image in the RT-AMU dataset, including the tumor area and healthy area of the kidney; the second row is the corresponding GT image, the white area is the kidney tumor, and the black area is the healthy tissue; a, b, c, d - different sizes and types, b - low contrast, c - low brightness, d - small kidney tumor), the different sizes and types of kidney tumors, the low contrast between the tumor and normal healthy tissue, the uneven or low brightness of laparoscopic images, etc. are still challenges faced by deep learning-based kidney tumor segmentation algorithms.
[0004] Currently, deep learning algorithms have been widely used in the field of medical image segmentation and have gradually replaced traditional medical image segmentation algorithms. Existing deep learning-based medical image segmentation algorithms can be roughly divided into four categories: CNN-based methods, Transformer-based methods, CNN-Transformer hybrid methods, and Mamba-based methods. Specifically:
[0005] 1) Medical image segmentation methods based on convolutional neural networks (CNNs) typically use CNN backbone networks (such as Res2Net, HardNet68) as encoders and combine some carefully designed structures and operations to enhance the features extracted by the encoders. Although CNN-based methods have achieved initial success in medical image segmentation tasks, however, due to the limitations of convolutional operations, CNNs have a weak ability to capture global information and are difficult to segment kidney tumors with different shapes, sizes, and types. In addition, the inherent inductive bias of convolution limits the effectiveness of CNNs in more complex and variable laparoscopic image segmentation scenarios.
[0006] 2) To overcome the deficiencies of CNNs in capturing global feature information, Vision Transformer (ViT) divides the input image into a series of fixed-size patches and establishes the relationships between the patches through the self-attention mechanism to capture global context information. ViT is the first pioneering backbone network to adapt the Transformer architecture to computer vision tasks. Subsequently, the emergence of Swin-Transformer, CSwin-Transformer, PVT, and PVT v2 has reduced the computational complexity of ViT and overcome the defect that ViT can only output single-scale feature maps. The core of these Transformer architectures is the self-attention mechanism, which processes the information embedded in all positions in the input sequence in parallel rather than sequentially. This mechanism enables Transformers to proficiently manage long-range information dependencies, adapt to different input sequence lengths, and reduce the inductive bias of the model. Compared with CNN-based methods, medical image segmentation methods based on the Transformer architecture have a stronger ability to capture long-range dependencies, greatly improving the segmentation performance of the model. However, due to the lack of spatial inductive bias in modeling local information, there are still limitations in capturing local detailed information, resulting in the possible loss of small kidney tumors and partial boundary regions during the decoding stage. In addition, existing Transformer-based methods have limited capabilities in feature expression and multi-scale adaptation.
[0007] 3) To comprehensively extract the global semantic features and local detail information in medical images, the medical image segmentation method based on the CNN-Transformer hybrid architecture combines CNN and Transformer to improve the performance of semantic segmentation. The existing structural forms of the CNN-Transformer hybrid methods are mainly divided into serial structures and parallel structures: The hybrid methods with serial structures mainly use Transformer blocks to replace the deep convolutional layers in the encoding stage or stack the CNN backbone and the Transformer backbone in a sequential manner. However, this method does not fully utilize the advantages brought by Transformer and ignores introducing Transformer into feature maps at various scales; The hybrid methods with parallel structures run the CNN backbone and the Transformer backbone in parallel to generate complementary feature information, and then integrate the local detail information and the global context information, which can give full play to the advantages of CNN and Transformer. However, the existing dual-backbone models have a large number of parameters, and the fusion effect of global features and local features is not ideal. The irrelevant background noise introduced in the fusion process also limits the network performance.
[0008] 4) Recently, state space models (SSMs), such as Mamba, have performed excellently in computer vision tasks. VMamba is the first Mamba backbone network applicable to computer vision tasks. It can not only capture long-range dependencies but also has linear complexity. On this basis, Vision MambaUNet (VM-UNet) and VM-UNET-V2 introduce the Visual State Space (VSS) Block derived from VMamba to capture a wide range of global context information and show good performance in medical image tasks. However, these Mamba-based methods are still restricted by the characteristics of specific medical image datasets and have poor generalization ability.
[0009] The attention mechanism plays an important role in medical image segmentation networks. It can effectively guide the model to focus on key feature regions, reduce background interference, and thus improve the segmentation performance of the network. Currently, the widely used attention mechanisms are mainly divided into weight-based activation attention mechanisms and non-local interaction attention mechanisms: Weight-based activation attention mechanisms usually obtain the distribution weights of data by performing global compression in the spatial or channel dimension to effectively enhance significant features and suppress invalid information. However, this method lacks selective induction, resulting in the network overly focusing on the most significant global features, thus suppressing secondary features with slightly lower weights but equally important. Compared with the weight-based activation attention mechanism, the non-local interaction attention mechanism represented by the self-attention mechanism can obtain context features through global information interaction and is more sensitive to fine-grained features, but the computational complexity is slightly larger. Existing medical image segmentation networks often use the above two types of attention mechanisms to force the network to focus on important channels or important spatial regions. The structural forms are mainly divided into single attention, serial attention, and parallel attention. Specifically: The single attention structure is simple and easy to implement, but the ability to focus on important feature information is insufficient; The serial attention uses two attention modules in series, enhancing the network's ability to adaptively focus on important features, but at the same time significantly increasing the computational complexity and reducing the training efficiency; The parallel attention uses two attention modules in parallel. The two modules are calculated independently, reducing the latency caused by sequential operations, accelerating the processing speed, and improving the computational efficiency. However, there is a lack of effective communication between the modules, which may lead to redundant calculations (of features), thereby affecting the model's segmentation performance.
[0010] Therefore, how to achieve more fine-grained prediction when accurately locating the position of kidney tumors, so as to assist doctors in treating renal cancer during partial nephrectomy and reduce the complexity of doctors' surgical operations is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0011] In view of this, the present invention provides a laparoscopic image kidney tumor segmentation method and system based on deep learning, which solves the problems existing in the background technology.
[0012] In order to achieve the above object, the present invention provides the following technical solutions:
[0013] A laparoscopic image kidney tumor segmentation method based on deep learning, comprising the following steps:
[0014] Construct a laparoscopic kidney tumor dataset RT-AMU, and divide the RT-AMU dataset into a training set and a test set according to a preset ratio;
[0015] Design a parallel-structured CNN-Transformer hybrid network PEEA-Net as the initial kidney tumor segmentation model;
[0016] Input the training set into the initial kidney tumor segmentation model for model training until the loss function converges to obtain the optimized kidney tumor segmentation model;
[0017] Input the test set into the optimized kidney tumor segmentation model to obtain the kidney tumor segmentation results of laparoscopic images.
[0018] Optionally, the construction method of the laparoscopic kidney tumor dataset RT-AMU is specifically as follows:
[0019] Obtain several robot-assisted partial nephrectomy videos, extract laparoscopic images from the videos; and manually annotate the laparoscopic images by professional urologists to generate the corresponding ground truth GT.
[0020] Optionally, the parallel-structured CNN-Transformer hybrid network PEEA-Net includes: a PVTv2 encoder, a CNN auxiliary encoder, a PEA module, an ESPA module, an AG module, and an LDISF module;
[0021] The PVTv2 encoder is used to extract global information and obtain four different-level pyramid features T i (i = 1, 2, 3, 4);
[0022] The CNN auxiliary encoder is used to extract local detail information and obtain three different-scale feature maps C i (i = 1, 2, 3);
[0023] The PEA module is used to perform grouped extraction and progressive fusion on the pyramid features T i (i = 1, 2, 3, 4) to obtain enhanced features T pi (i = 1, 2, 3, 4);
[0024] The AG module is used to fuse the enhanced low-level features T pi (i = 1, 2, 3) with the upsampled high-level features UpConv(E i+1 )(i = 1, 2, 3) output by the previous decoder to obtain fused features;
[0025] The ESPA module is used to refine the spatial and channel features in the feature map E' i+1 (i = 1, 2, 3) generated by connecting the fused features with the upsampled high-level features UpConv(E i (i = 1, 2, 3) output by the previous decoder, and output the feature map E i (i = 1, 2, 3) containing important feature information;
[0026] The LDISF module is used to supplement the local detail feature information contained in the feature map C obtained by the CNN auxiliary encoder to the global feature map E output by the Transformer branch decoder 3 to generate a feature map E that can locate the position of the kidney tumor and outline the boundary of the kidney tumor 1 0
[0027] Optionally, the CNN auxiliary encoder includes three CNN blocks. The first CNN block contains a group of 3×3 convolutional units, and the latter two CNN blocks consist of a max-pooling downsampling operation and two groups of 3×3 convolutional units;
[0028] Among them, the convolutional units are used to capture local structure and texture features; the pooling window size of the max-pooling downsampling operation is 2×2, which is used to halve the size of the feature map and double the dimension.
[0029] Optionally, the PEA module uses grouped convolution operations with different receptive field sizes and a progressive fusion strategy to enrich and enhance features. Specifically:
[0030] The multi-layer features T output by the PVT v2 encoder i (i = 1, 2, 3, 4) are split along the channel dimension to generate four groups of features T ij (j = 1, 2, 3, 4), j = 1, 2, 3, 4;
[0031] Each group of features is respectively subjected to customized convolution operations. Among them: for the first group of features T i1 , a 1×1 convolutional unit is used to extract features; for the second group of features T i2 , a 3×3 convolutional unit with an expansion rate of 3 is used to capture multi-scale information; for the third group of features T i3 , an asymmetric convolution composed of a 1×3 horizontal kernel, a 3×1 vertical kernel, and a 3×3 convolutional kernel is used to capture complex asymmetric features in the laparoscopic image; for the fourth group of features T i4 , a self-calibrating convolution that can automatically adjust the convolution kernel weights according to the input features is used to capture rich detail information;
[0032] The adaptive average pooling layer is respectively concatenated with the results of the first and second branches and the third and fourth branches in the channel dimension, and two kinds of features A 1 , A 2 are output and then fused with the global information P i3 . After passing through a 1×1 convolutional unit to restore the original number of channels, and finally performing a residual connection with the original feature T i , the enhanced feature T pi is output.
[0033] Optionally, the ESPA module adopts a non-local interaction attention mechanism with parameter sharing, specifically as follows:
[0034] Combine the reshaped input feature r(E′ i ) with the position information, and then obtain the serialized tokens through layer normalization; the serialized tokens generate the channel value V C , the shared query Q, the shared key K, and the spatial value V S respectively through three parallel linear fully connected layers, where r(·) represents the feature reshaping operation;
[0035] Calculate the similarity QK between the shared query Q and the key K T , and then obtain the weight in the spatial direction through the Softmax function Multiply it with the spatial value V S to get the output spatial feature map E spatial ; Similarly, calculate the channel attention weight through the shared query Q and the key K and multiply it with the transposed matrix of the channel value V C to get the output channel feature E channel ; where, represents the dimension of the feature embedding;
[0036] spatial channel SRA i 1 3 tc tc tc tc SRA
[0037] Optionally, the LDISF module effectively fuses local features and global features, specifically as follows:
[0038] Connect the global feature map E 1 output by the Transformer branch decoder and the third-layer local feature map C 3 obtained by the CNN auxiliary encoder in the channel dimension to obtain the initial fusion feature F tc ;
[0039] tc tc tc tc tc tc Half of it; in the second branch, a convolutional branch consisting of a 1×3 convolution, a 3×1 convolution, and a 3×3 convolution with an expansion rate of 3 is used to expand the receptive field and mine the fine-grained features of laparoscopic images in different directions; in the third branch, a third convolutional branch consisting of a 3×3 convolution and a 3×3 convolution with an expansion rate of 5 is used to deepen the exploration and mining of features; in the fourth branch, deformable convolution is used to capture the boundary features in irregularly shaped kidney tumors;
[0040] Connect the outputs of the four branches in the channel dimension, use a 3×3 convolution for denoising and restore the original number of channels, and then connect with the initial fusion feature F tc Perform a residual connection, and then pass through an ECSA module to remove background noise, and output the global reliable feature E containing fine-grained detail information 0 。
[0041] Optionally, it further includes:
[0042] Use a 1×1 convolution to extract features from the feature map T output by the PEA module p4 to obtain the feature map E' 4 , and obtain the feature map E through the ESPA module 4 ;
[0043] E i (i = 0, 1, 2, 3, 4) passes through a 1×1 convolution and an upsampling operation to generate a prediction feature map D with 1 channel and the same size as the image to be segmented i (i = 0, 1, 2, 3, 4), D i (i = 0, 1, 2, 3, 4) then passes through a Sigmoid function to generate the prediction mask P i (i = 0, 1, 2, 3, 4);
[0044] For the prediction feature maps D i (i = 0, 1, 2, 3, 4) output in 5 stages, the final prediction feature map D is obtained by using the method of additive aggregation:
[0045]
[0046] Optionally, the loss function is a multi-stage loss function, and the expression is:
[0047]
[0048] L i =L wiou (D i , G)+L wbce (D i , G), (i = 0, 1, 2, 3, 4)
[0049] In the formula: L overall represents the total loss function, and L i (i = 0, 1, 2, 3, 4) respectively represent the loss functions of 5 stages. The wiou loss constrains the prediction results of kidney tumors in laparoscopic images from a global perspective, and the wbce loss constrains the prediction results of kidney tumors in laparoscopic images from a local perspective; G represents the ground truth.
[0050] A system applying the deep learning-based laparoscopic image kidney tumor segmentation method described in any one of the above, comprising:
[0051] A dataset construction module, used to construct a laparoscopic kidney tumor dataset RT-AMU, and divide the RT-AMU dataset into a training set and a test set according to a preset ratio;
[0052] A model building module, used to design a parallel-structured CNN-Transformer hybrid network PEEA-Net as the initial kidney tumor segmentation model;
[0053] A model training module, used to input the training set into the initial kidney tumor segmentation model for model training until the loss function converges, and obtain an optimized kidney tumor segmentation model;
[0054] An image segmentation module, used to input the test set into the optimized kidney tumor segmentation model to obtain the kidney tumor segmentation result of the laparoscopic image.
[0055] It can be seen from the above technical solutions that compared with the prior art, the present invention discloses a deep learning-based laparoscopic image kidney tumor segmentation method and system, having the following beneficial effects:
[0056] The present invention proposes a new network PEEA-Net for automatically and accurately segmenting kidney tumors in laparoscopic images. This network is a parallel-structured CNN-Transformer hybrid network, which can achieve more fine-grained predictions when accurately locating the position of kidney tumors. Among them, the PEA module solves the problems of limited feature expression ability and multi-scale adaptation ability of the Transformer encoder; the ESCD module solves the problem of insufficient ability of the network to adaptively select key feature maps and key regions in the feature maps; the LDISF module solves the problem that small kidney tumors and partial boundary regions may be lost in the decoding stage; in addition, the auxiliary CNN encoder also reduces the computational complexity of the double backbone encoder network.
[0057] In summary, compared with the prior art, the PEEA-Net provided by the present invention has better segmentation performance and can more effectively address the challenges brought about by the different sizes and types of renal tumors in laparoscopy, the low contrast between tumors and normal healthy tissues, small renal tumors, and the uneven or low brightness of laparoscopic images. In addition, the network has stronger generalization ability. As a deep learning-based automatic renal tumor segmentation technology, it can assist doctors in treating renal cancer during partial nephrectomy, reduce the complexity of doctors' surgical operations, improve the clinical treatment process, and has important clinical significance. Description of the Drawings
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the provided drawings.
[0059] Figure 1 The laparoscopic image provided by the present invention;
[0060] Figure 2 The flowchart of the method for segmenting renal tumors in laparoscopic images based on deep learning provided by the present invention;
[0061] Figure 3 The overall architecture diagram of the PEEA-Net provided by the present invention;
[0062] Figure 4 The structural diagram of the PEA module provided by the present invention;
[0063] Figure 5 The structural diagram of the ESPA module provided by the present invention;
[0064] Figure 6 The structural diagram of the LDISF module provided by the present invention;
[0065] Figure 7 The acquisition process of the renal tumor dataset RT-AMU provided by the present invention;
[0066] Figure 8 The comparison diagram of the visualization qualitative prediction mask results of some methods provided by the present invention on the RT-AMU dataset;
[0067] Figure 9 The comparison diagram of the visualization qualitative prediction mask results of some other methods provided by the present invention on the RT-AMU dataset. Detailed Embodiments
[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0069] An embodiment of the present invention discloses a method for segmenting renal tumors in laparoscopic images based on deep learning, as Figure 2 shown, which includes the following steps:
[0070] Construct a laparoscopic renal tumor dataset RT-AMU, and divide the RT-AMU dataset into a training set and a test set according to a preset ratio;
[0071] Design a CNN-Transformer hybrid network PEEA-Net with a parallel structure as the initial segmentation model for renal tumors;
[0072] Input the training set into the initial segmentation model for renal tumors for model training until the loss function converges to obtain an optimized segmentation model for renal tumors;
[0073] Input the test set into the optimized segmentation model for renal tumors to obtain the segmentation result of renal tumors in laparoscopic images.
[0074] To solve the problems existing in the background technology, this embodiment proposes a progressive exploration aggregation and enhanced shared parallel attention hybrid network PEEA-Net for accurately segmenting renal tumors in laparoscopic images. Next, the Figure 2 steps shown will be elaborated in detail to further understand the technical solutions of the present invention.
[0075] I. Dataset construction
[0076] The construction method of the laparoscopic renal tumor dataset RT-AMU is specifically as follows:
[0077] Obtain a number of robot-assisted partial nephrectomy videos, extract laparoscopic images from the videos; manually annotate the laparoscopic images by professional urological experts to generate corresponding ground truth GT.
[0078] II. Progressive exploration aggregation and enhanced shared parallel attention hybrid network PEEA-Net
[0079] As Figure 3As shown, the parallel-structured CNN-Transformer hybrid network PEEA-Net includes: a PVT v2 encoder, a CNN auxiliary encoder, a PEA (Progressive Exploration Aggregation) module, an ESPA (Enhanced Shared Parallel Attention) module, an AG (Attention gate) module, and an LDISF (Local Detail Information Supplementary Fusion) module; specifically, given a laparoscopic image X ∈ R H×W×3 , the operations are as follows:
[0080] The PVTv2 encoder is used to extract global information and obtain four different levels of pyramid features T i (i = 1, 2, 3, 4);
[0081] The CNN auxiliary encoder is used to extract local detail information and obtain three different-scale feature maps C i (i = 1, 2, 3);
[0082] The PEA module is used to perform grouped extraction and progressive fusion on the pyramid features T i (i = 1, 2, 3, 4), enhance the feature expression ability and multi-scale adaptation ability of T i , and obtain enhanced features T pi (i = 1, 2, 3, 4);
[0083] For T p4 , perform a 1×1 convolution operation on it for feature extraction to obtain the feature map E' 4 ;
[0084] The AG module is used to fuse the enhanced low-level features T pi (i = 1, 2, 3) with the upsampled high-level features UpConv(E i+1 )(i = 1, 2, 3) output by the previous decoder, obtain the fused features, realize the fusion of high- and low-level feature information, and filter the redundant information in the skip connection;
[0085] The ESPA module is used to refine the spatial and channel features in the feature map E′ i+1 (i = 1, 2, 3) generated by connecting the fused features with the upsampled high-level features UpConv(E i (i = 1, 2, 3) output by the previous decoder, and output the feature map E i (i = 1, 2, 3) containing important feature information;
[0086] The LDISF module is used to fuse the feature map C obtained by the CNN auxiliary encoder 3The local detailed feature information contained is supplemented to the global feature map E output by the Transformer branch decoder 1 to generate a feature map E that can locate the position of the renal tumor and outline the boundary of the renal tumor 0 . Among them, E 1 is the output of the last ESPA module and also the final output of the decoder of the Transformer branch. Here, the expression "the global feature map E output by the Transformer branch decoder 1 " can reflect the role of the LDISF module in fusing the local features extracted by the CNN and the global features restored by the Transformer
[0087] E i (i = 0, 1, 2, 3, 4) generates a predicted feature map D with 1 channel and the same size as the image to be segmented through a 1×1 convolution and upsampling operation i (i = 0, 1, 2, 3, 4), and D i (i = 0, 1, 2, 3, 4) then generates a predicted mask P i (i = 0, 1, 2, 3, 4) through a Sigmoid function. During the training process, deep supervision is performed on the predicted masks P i (i = 0, 1, 2, 3, 4) output in 5 stages
[0088] The overall network structure of PEEA-Net is defined as follows
[0089] T i = PVT v2(X), (i = 1, 2, 3, 4)
[0090] C 3 = CNN3(CNN2(CNN1(X)))
[0091] T pi = PEA(T i )(i = 1, 2, 3, 4)
[0092] E' 4 = Conv1(T p4 )
[0093] E′ i = Cat(AG(UpConv(E i+1 ), T pi ), UpConv(E i+1 ))(i = 1, 2, 3)
[0094] E i = ESPA(E′ i )(i = 1, 2, 3, 4)
[0095] E 0 = LDISF(E 1 , C 3 )
[0096] D i = Conv1(E i ), (i = 0, 1, 2, 3, 4)
[0097]
[0098] P i = Sigmoid(D i ), (i = 0, 1, 2, 3, 4)
[0099] 2.1 Encoder
[0100] Since the existing single encoder network cannot effectively extract global semantic and local detail information simultaneously, therefore, similar to the existing parallel-structured CNN-transformer hybrid method, the present invention uses Transformer and CNN to capture long-range dependencies and local context information between pixels simultaneously. Specifically, in this embodiment, the Transformer backbone network PVT v2 is used as the main encoder to extract global information, and four different levels of pyramid features T i (i = 1, 2, 3, 4); in order to reduce the number of parameters of the encoder network with a parallel structure and reduce the model complexity, different from the existing parallel hybrid network that usually uses Res2Net as the CNN encoder, this embodiment adopts a CNN auxiliary encoder to extract feature maps of three different scales from the input image
[0101] The CNN auxiliary encoder includes three CNN blocks. The first CNN block contains a group of 3×3 convolutional units, and the latter two CNN blocks are composed of a max-pooling downsampling operation and two groups of 3×3 convolutional units;
[0102] C 1 = (BConv3(X))
[0103] C i = BConv3(BConv3(MaxPool2(C i-1 )))(i = 2, 3)
[0104] In the above formula, BConv3(·) represents a 3×3 convolution kernel, batch normalization, and ReLU activation function; the convolutional unit operation can effectively capture local structure and texture features, and has robustness and fault tolerance; MaxPool2(·) represents a max-pooling downsampling operation with a pooling window size of 2×2, which is used to halve the size of the feature map and double the dimension, and can maximize the feature expression ability while maintaining computational efficiency. It can extract the most significant features from local regions, retain important information such as edges, and discard some unimportant detail information. In this way, the third-layer feature map C extracted by the CNN auxiliary encoder 3 contains rich local detail information, such as the texture and boundary of kidney tumors. Using C 3 as the input of the local detail information supplement and fusion module (LDISF), it is used to optimize the feature representation output by the Transformer branch decoder and help the network obtain the best segmentation performance.
[0105] 2.2 Progressive Exploration Aggregation Module (PEA)
[0106] The sizes, shapes, types, and numbers of kidney tumors in laparoscopic images vary. Existing Transformer-based methods are insufficient in distinguishing kidney tumors from normal healthy tissues in these laparoscopic images and have limited multi-scale adaptability to kidney tumors. (In addition, due to possible feature redundancy in the four-layer pyramid features extracted by the PVT v2 encoder, directly using the unprocessed features extracted by the encoder may lead to poor prediction results.) Therefore, this embodiment proposes a new Progressive Exploration Aggregation Module (PEA), aiming to make full use of the features extracted by the PVT v2 encoder, enhance the feature expression ability and multi-scale adaptability of the encoder, and improve the flexibility and accuracy of the network in segmenting kidney tumors of different sizes, shapes, types, and numbers.
[0107] As Figure 4 shown, the PEA module of this embodiment uses grouped convolution operations with different receptive field sizes and a progressive fusion strategy to enrich and enhance (extract and fuse) features. Specifically:
[0108] The multi-layer features T output by the PVT v2 encoder i (i = 1, 2, 3, 4) are segmented along the channel dimension to generate four groups of features T ij (j = 1, 2, 3, 4), j = 1, 2, 3, 4;;
[0109] Each group of features is respectively subjected to customized convolution operations. Among them: for the first group of features T i1 , a 1×1 convolution unit is used to extract features; for the second group of features T i2, To expand the receptive field, a 3×3 convolutional unit with an expansion rate of 3 is used to capture multi-scale information to assist in identifying large kidney tumors; since asymmetric convolution allows the network to use different weights when extracting features in certain directions, which can help the network learn more complex features, therefore, to capture the complex asymmetric features in laparoscopic images, for the third group of features T i3 , it is processed using an asymmetric convolution composed of a 1×3 horizontal kernel, a 3×1 vertical kernel, and a 3×3 convolutional kernel; to enhance the generalization ability of the model when facing different types of kidney tumors, for the fourth group of features T i4 , a self-calibrating convolution that can automatically adjust the convolutional kernel weights according to the input features is used to capture rich detail information and enhance the recognition ability for small kidney tumors;
[0110] To supplement the global information of the input feature T i , the adaptive average pooling layer is concatenated with the results of the first and second branches and the third and fourth branches respectively in the channel dimension, and two types of features A 1 、A 2 are output and then fused with the global information P i3 . Then, the original number of channels is restored through a 1×1 convolutional unit, and finally, a residual connection is made with the original feature T i to output the enhanced feature T pi :
[0111] T ij = Split(T i )(j = 1,2,3,4)
[0112] P ij = Upsample(Conv1(AvgPool(T i ))),(j = 1,2,3,4)
[0113] A 1 = Conv1(Cat(BConv1(T i1 ),DConv3(T i2 ),P i1 ))
[0114] A 2 = Conv1(Cat(Conv sc (T i3 ),Conv as (T i4 ),P i2 ))
[0115] T pi = BConv1(Cat(A 1 ,A 2 ,Pi3 )) + T i
[0116] In the above formula, Split represents splitting along the channel dimension, AvgPool(·) represents the AdaptiveAvgPool (Adaptive Average Pooling) operation, Conv1(·) represents a 1×1 convolution, Upsample(·) represents the interpolate operation, BConv1 represents a 1×1 convolution kernel, batch normalization, and ReLU activation function, DConv3(·) represents a 3×3 convolution with a dilation rate of 3, batch normalization, and PReLU activation function, Cat(·) represents a concatenation operation in the channel dimension, Conv as (·) represents an asymmetric convolution, Conv sc (·) represents a self-calibrating convolution.
[0117] 2.3 Enhanced Shared Parallel Attention Module (ESPA)
[0118] The contrast between kidney tumors and normal healthy tissues in laparoscopic images is low. In addition, the image brightness is uneven. Therefore, it is crucial for the network to adaptively focus on the important parts of laparoscopic image data. However, existing Transformer-based methods lack an attention mechanism that considers the specific channel and spatial features of laparoscopic images, and the adaptive attention ability of existing attention mechanisms is insufficient. Therefore, the present invention proposes a new Enhanced Shared Parallel Attention Module (ESPA) to perform effective feature extraction for the unique attributes of laparoscopic images, enabling the model to focus on important feature information. Different from most attention mechanisms that use global squeezing, this embodiment utilizes a self-attention mechanism (non-local interaction attention mechanism) that is more sensitive to fine-grained features. The difference is that we enhance the internal connection between space and channels through a parameter-sharing parallel attention structure, enhancing the network's ability to adaptively select key feature maps and key regions in the feature maps, while reducing the number of model parameters and computational complexity.
[0119] As Figure 5 shown, the ESPA module of this embodiment adopts a parameter-sharing non-local interaction attention mechanism, specifically:
[0120] Combine the reshaped input feature r(E′ i ) with the position information so that the model can understand the relative positions of the elements in the input data, and then obtain serialized tokens through layer normalization to improve the stability of training and accelerate the model convergence speed; the serialized tokens generate channel values V C , shared query Q, shared key K, and spatial values V S; among them, the input of the first ESPA module is the last layer of features E' extracted by the Transformer encoder processed by the PEA module and 1*1 convolution 4 ; the inputs of the second, third, and fourth ESPA modules are the feature maps E′ generated by connecting the fused features output by the AG module and the upsampled high-level features output by the previous decoder i (i = 1, 2, 3); r(·) represents the feature reshaping operation
[0121] Specifically, for spatial attention, calculate the similarity QK between the shared query Q and key K T , and then obtain the weights in the spatial direction through the Softmax function Then multiply with the spatial value V S to get the output spatial feature map E spatial ; similarly, calculate the channel attention weights through the shared query Q and key K Then multiply with the transposed matrix of the channel value V C to get the output channel feature E channel ; among them,
[0122]
[0123]
[0124] Add E spatial and E channel to the input features respectively, concatenate them in the channel dimension, and use a 1×1 convolution to restore the original number of channels to get O SRA , then add O SRA to the input features for feature enhancement and noise suppression, input them into the Feed-Forwrad module (FF), and at the same time use the residual connection operation. Finally, use a local enhancement module (LE) to reduce attention dispersion and optimize the feature representation to get the output feature E i .
[0124] O SRA = Conv1(Cat(E spatial + l(r(E′ i ))), (r(E channel ) + l(r(E′ i ))))
[0125] E i = LE(r(O SRA + r(E′ i )) + FF(r(O SRA + r(E′ i ))))
[0126] FF(x) = Conv1(R(Conv1(x)))
[0127] LE(x) = BConv3(BConv3(x))
[0128] In the above formula, r(·) represents the feature reshaping operation, l(·) represents layer normalization, Cat(·) represents the concatenation operation in the channel dimension, Conv1(·) represents the 1×1 convolution, R(·) represents the ReLU activation function, and BConv3(·) represents the 3×3 convolution kernel, batch normalization, and ReLU activation function.
[0129] Since the shared Q and K carry the inherent context statistical information, E spatial and E channel effectively refine the spatial and channel features while retaining the global context information, weaken the irrelevant information in the spatial and channel dimensions, and enhance the Transformer's extraction of image-specific channel and spatial features. Compared with the input feature E′ i , the output feature E i after passing through the ESPA module can help the decoder improve the ability to recover the global semantics and local information, and achieve more accurate reconstruction of the kidney tumor prediction feature map.
[0130] Compared with the single attention based on the non-local interaction attention mechanism, the dual attention mechanism of ESPA can better capture the important features in the image and is more sensitive to the fine-grained features; compared with the serial attention, the parallel attention structure of ESPA reduces the latency caused by sequential operations and speeds up the processing speed; compared with the parallel attention, the parameter sharing mechanism of ESPA not only strengthens the communication between the channel and spatial features, improves the model performance, but also solves the problem of redundant calculations that may occur in parallel processing, and reduces the number of model parameters.
[0131] 2.4 Local Detail Information Supplementary Fusion Module (LDISF)
[0132] To fully unleash the potential of the dual-encoder network, it is necessary to fully integrate the global coarse-grained high-level features captured by the main Transformer encoder PVTv2 and the local fine-grained low-level features extracted by the proposed auxiliary CNN encoder. The former focuses on more object-level semantic information and can help locate the specific position of kidney tumors, while the latter focuses on more detailed information and is crucial for the localization of tiny kidney tumors and more accurate boundary delineation. These two types of cross-domain feature information are (naturally) complementary to each other to a certain extent. Therefore, the present invention proposes a new Local Detail Information Supplementary Fusion Module (LDISF) to fully and effectively fuse local features with global features and eliminate background noise, enabling the network to achieve more fine-grained predictions when accurately locating the position of kidney tumors.
[0133] As Figure 6 shown, the LDISF module of this embodiment effectively fuses local features with global features, specifically as follows:
[0134] Connect the global feature map E 1 output by the Transformer branch decoder and the third-layer local feature map C 3 obtained by the CNN auxiliary encoder in the channel dimension to obtain the initial fusion feature F tc to facilitate the subsequent extraction of mutually significant features therein;
[0135] Divide the initial fusion feature F tc into four branches and use different convolutional operations to mine the significant information in the initial fusion feature F tc wherein: use a 1×1 convolution to reduce the number of feature channels of each branch to half of F tc to reduce computational consumption; in the second branch, use a convolutional branch composed of a 1×3 convolution, a 3×1 convolution, and a 3×3 convolution with a dilation rate of 3 to expand the receptive field and mine the fine-grained features of laparoscopic images in different directions, such as edge detail information, etc.; in the third branch, use a third convolutional branch composed of a 3×3 convolution and a 3×3 convolution with a dilation rate of 5 to deepen the exploration and mining of features; in the fourth branch, use deformable convolution to capture the boundary features in irregularly shaped kidney tumors, and this deformable convolution can adaptively adjust the shape and position of the convolution kernel according to the input features;
[0136] Connect the outputs of the four branches in the channel dimension, use a 3×3 convolution for denoising and restoring the original number of channels, and then perform a residual connection with the initial fusion feature F tc , and then pass through an (ECSA module with low computational cost based on weighted activation attention) to remove background noise and strengthen significant features. The LDISF module helps the input features E 1 and C3 The complementary feature information is realized, the feature information irrelevant to kidney tumor segmentation is filtered out, and the global reliable feature E containing fine-grained detail information is output 0 , which improves the ability of the network to locate small kidney tumors and refines the boundaries of the predicted masks output by the network
[0137] 2.5 Loss function
[0138] As Figure 3 shown, in this embodiment, for the predicted feature maps D output by the five stages in the proposed PEEA-Net i (i = 0, 1, 2, 3, 4), the final predicted feature map D is obtained by using the method of additive aggregation
[0139]
[0140] A multi-stage loss function is designed to supervise the predicted outputs of the five stages of the model, and the expression is
[0141]
[0142] L i = L wiou (D i , G) + L wbce (D i , G), (i = 0, 1, 2, 3, 4)
[0143] In the formula: L overall represents the total loss function, and L i (i = 0, 1, 2, 3, 4) respectively represent the loss functions of the five stages. The wiou loss constrains the prediction results of kidney tumors in laparoscopic images from a global perspective, and the wbce loss constrains the prediction results of kidney tumors in laparoscopic images from a local perspective; G represents the ground truth
[0144] III. Experiments
[0145] 3.1 Dataset
[0146] (1) Kidney tumor dataset
[0147] Currently, in the task of kidney tumor segmentation, there is a lack of a laparoscopic image dataset that can be used for comparative evaluation. Therefore, the laparoscopic kidney tumor dataset RT-AMU was created for this experiment Figure 7Shows the source of the RT-AMU dataset. From left to right, they are: robotic surgery system, robot-assisted partial nephrectomy, original laparoscopic images provided by the robotic system console during the operation, and kidney tumor segmentation masks manually annotated by urological experts. Specifically, RT-AMU is derived from 63 robot-assisted partial nephrectomy videos provided by the Department of Urology, the First Affiliated Hospital of Anhui Medical University. 2258 kidney tumor images with a resolution of 1920×1080 were extracted from these videos, and these images were carefully manually annotated by professional urological experts to generate the corresponding ground truth (GT).
[0148] In the experiment, 2258 laparoscopic image data were randomly selected from the RT-AMU dataset for training, and the remaining 565 image data were used for testing to evaluate the effectiveness and superiority of the proposed PEEA-Net in the laparoscopic kidney tumor segmentation task.
[0149] (2) Public datasets
[0150] To evaluate the performance of the proposed PEEA-Net in terms of robustness and generalization, five publicly available challenging polyp datasets were selected for this experiment. These datasets are all composed of polyp images extracted from different colonoscopy videos. Specifically, Kvasir-SEG contains 1000 polyp image data with a resolution range from 332×487 to 1920×1072; CVC-ClinicDB contains 612 polyp images with a resolution of 384×288; CVC-ColonDB contains 380 polyp images with a resolution of 574×500; ETIS contains 196 polyp images with a resolution of 1225×996; and CVC-300 contains 60 polyp images with a resolution of 574×500.
[0151] In the experiment, we used the same data distribution setting as PraNet. 90% of the polyp images were randomly selected from the Kvasir-SEG dataset and the CVC-ClinicDB dataset as the training set, and the remaining 10% of the polyp images were used as part of the test set to evaluate the learning ability of the model. In addition, all the polyp images in CVC-ColonDB, ETIS, and CVC-300 were included in the test set to verify the generalization ability of the model.
[0152] 3.2 Evaluation metrics
[0153] To quantitatively evaluate the segmentation performance of the model, we used mDice, mIoU, MAE, Accuracy, HD is used as an evaluation metric. Specifically, mDice and mIoU represent the degree of agreement between the model's predicted region and the ground truth mask; MAE is used to compare the per-pixel absolute value difference between the predicted map (P) and the ground truth (G); Accuracy represents the proportion of pixels correctly predicted by the network to the total number of pixels in the image; is used to calculate the weighted sum average of Precision and Recall; HD is used to measure the segmentation accuracy of the boundary. Among them, mDice, mIoU, Accuracy and The closer the value is to 1, the better the segmentation effect; while the closer the value of MAE is to 0, the better the segmentation effect; the lower the value of HD also means the more ideal the segmentation effect. The specific definitions of these metrics are as follows:
[0154]
[0155]
[0156] In the above formula, TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative. W and H are the width and height of the image. β is the weight coefficient, and in this paper, β is set to 1.
[0157] 3.3 Experimental Details
[0158] The PEEA-Net proposed in this embodiment is implemented using Python and the PyTorch framework, and is trained, predicted, and evaluated on an NVIDIA GeForce RTX 3090 GPU with 24GB of video memory. During training, the pre-trained transformer weights from ImageNet are used to initialize the PVT v2 backbone, the epochs are set to 100, the learning rate is set to 10 -4 , the optimizer is set to AdamW, the batch size is set to 16, the decay rate is set to 0.1. In addition, the size of the input image is adjusted to 352×352, and a multi-scale {0.75, 1.0, 1.25} training strategy is used, and the gradient clipping is limited to 0.5. In addition, translation, mirroring, and rotation are also used as data augmentation techniques to alleviate the overfitting phenomenon.
[0159] 3.4 Experimental Results on the RT-AMU Dataset
[0160] (1) Quantification Results
[0161] To verify the effectiveness and superiority of the proposed PEEA-Net in laparoscopic kidney tumor segmentation tasks, it was compared with various state-of-the-art (SOTA) methods. The comparison models included CNN-based methods: UNet, PraNet, HarDNetMSEG, CaraNet, DCRNet, NPD-Net-Res2Net; Transformer-based methods: Swin-Unet, SwinPA-Net, MSRAformer, SSFormer, NPD-Net-PVT v2, MSGAT; CNN-Transformer hybrid methods: TransUNet, TGDAUNet, CASF-Net, DBMIA-Net; Mamba-based methods: VM-UNet, VM-UNET-V2. Table 1 shows the quantitative experimental results of the proposed PEEA-Net and various SOTA methods on the laparoscopic kidney tumor dataset RT-AMU created in this invention. As shown in Table 1, compared with other SOTA models, PEEA-Net performs optimally in all evaluation metrics, achieving state-of-the-art segmentation performance. Specifically, NPD-Net-Res2Net performs optimally among all CNN-based methods, while PEEA-Net leads NPD-Net-Res2Net by 1.61% and 2.08% in the mDice and mIoU metrics respectively. This is because in the kidney tumor segmentation task, compared with the CNN backbone Res2Net used by NPD-Net-Res2Net, the main encoder network PVT v2 adopted by PEEA-Net has a stronger global context modeling ability. At the same time, the mDice and mIoU metric values of PEEA-Net are also 1.05% and 1.33% higher respectively than those of NPD-Net-PVT v2, which performs optimally among the Transformer-based methods. The main reason is that the pure Transformer architecture used by NPD-Net-PVT v2 is insufficient in capturing local detailed information, while the LDISF module of PEEA-Net can optimize the high-level feature representation output by the Transformer branch decoder and supplement the local low-level features extracted by the proposed auxiliary CNN encoder into the global high-level features. In addition, the quantitative results of PEEA-Net are much better than those of the Mamba-based method VM-UNET-V2. It leads VM-UNET-V2 by 8.6% and 10.05% in the mDice and mIoU metrics respectively, and there are similar leading results in other metrics. Overall, PEEA-Net is in mDice, Accuracy and In terms of the indicators, it is improved by 0.97%, 0.1%, and 0.5% respectively compared with the second-best DBMIA-Net among all methods. In terms of the MAE indicator, PEEA-Net leads the second-best DBMIA-Net by 0.11%, which indicates that PEEA-Net can better distinguish kidney tumors and normal backgrounds in laparoscopic images. In terms of the HD indicator, PEEA-Net is reduced by 0.3424 compared with the second-best NPD-Net-PVT v2 among all methods, which indicates that PEEA-Net improves the boundary segmentation performance of kidney tumors. These quantitative results prove the effectiveness and superiority of PEEA-Net in the laparoscopic kidney tumor segmentation task.
[0162] Table 1 Comparison of quantitative experimental results of different methods on the RT-AMU dataset (↑ indicates the higher the better, ↓ indicates the lower the better, and bold indicates the best performance)
[0163]
[0164]
[0165] (2) Qualitative results
[0166] To more intuitively see the effectiveness and superiority of the PEEA-Net of the present invention in the kidney tumor segmentation task, the qualitative prediction results of PEEA-Net on the kidney tumor dataset RT-AMU are compared with the results of various current state-of-the-art (SOTA) methods. The comparison methods include: UNet, PraNet, HarDNet MSEG, CaraNet, DCRNet, NPD-Net-Res2Net; Swin-Unet, SwinPA-Net, MSRAformer, SSFormer, NPD-Net-PVT v2, MSGAT; TransUNet, TGDAUNet, CASF-Net, DBMIA-Net; VM-UNet, VM-UNET-V2. Figure 8 、 Figure 9 shows the prediction masks of different methods on the RT-AMU dataset, where white represents the predicted kidney tumor and black represents the normal background tissue. As Figure 8 、 Figure 9As shown, the PEEA-Net of this embodiment can more accurately segment kidney tumors in laparoscopic images. Specifically, in the first and second columns, there is a large amount of blood vessels and fat occlusion in the kidney tumor area. The PEEA-Net can not only accurately locate the tumor position, but also achieve a more accurate segmentation of the kidney tumor boundary. In the third column, the kidney tumor and the normal background tissue are highly similar in appearance and have a low contrast, making it difficult to distinguish with the naked eye. Based on the Mamba-based methods (VM-UNet, VM-UNET-V2), the prediction results of Transunet and CaraNet all produced a large number of false positives. DBMIA-NET, TGDAUNet, MSRAformer, swinpa, and NPD-Net-res2net also all produced a small amount of over-segmentation. Swinunet, HarDNet-MSEG, and UNet all failed to successfully predict the correct kidney tumor area, while the prediction mask of our method is extremely close to the true GT. In the fourth column, the laparoscopic image contains multiple discontinuous small kidney tumors. Except for the proposed PEEA-Net, MSRAformer, and NPD-Net-res2net that can identify two small tumors, the prediction masks of other methods wrongly regarded them as a whole. In the small tumor image of the fifth column, swinunet, DCRNet, and UNet predicted failure; the prediction results of VMUNET, TGDAUNet, Transunet, ssformer, MSRAformer, swinpa, NPD-Net-res2net, CaraNet, HarDNet-MSEG, and PraNet all had varying degrees of over-segmentation or under-segmentation. And our PEEA-Net can more sensitively capture these (in the 4th - 5th columns) small kidney tumors and remove false positives and false negatives. In the low-brightness laparoscopic image of the sixth column, the prediction mask of the PEEA-Net of the present invention is still the most similar to the GT, achieving a more accurate segmentation of the kidney tumor. Generally speaking, our method can effectively cope with the challenges brought by the different sizes and types of kidney tumors in laparoscopy (as shown in (a - f) in Figure 8 ), the low contrast between the tumor and normal healthy tissue (as shown in (c) in Figure 8 , Figure 9 ), small kidney tumors (as shown in (d - e) in Figure 8 , Figure 9 ), and the uneven or low brightness of laparoscopic images (as shown in (f) in Figure 8 , Figure 9 ), and achieve a more accurate segmentation. These visual qualitative results once again prove the effectiveness and superiority of the PEEA-Net of the present invention in dealing with complex and variable clinical laparoscopic image data.
[0167] 3.5 Experimental Results of the Polyp Dataset
[0168] To verify the robustness and generalization ability of the proposed PEEA-Net in this invention, the mDice and mIoU metrics are used to quantitatively evaluate the experimental results of PEEA-Net and 18 state-of-the-art (SOTA) methods on 5 publicly available colonoscopy polyp datasets. Table 2 shows the quantitative experimental results of different methods on the Kvasir-SEG dataset and the CVC-ClinicDB dataset. As shown in Table 2, the quantitative results of the proposed PEEA-Net on these two datasets are better than those of the compared SOTA models. Among them, on the Kvasir-SEG dataset, the mDice and mIoU metric values of PEEA-Net are increased by 0.26% and 0.75% respectively compared with the second-best DBMIA-Net among all methods. On the CVC-ClinicDB dataset, PEEA-Net also shows a similar leading advantage, with its mDice and mIoU values reaching 0.9439 and 0.9027 respectively. This proves that PEEA-Net has stronger polyp feature learning ability and excellent robustness. Table 3 shows the comparison of the quantitative results of different methods on 3 polyp datasets CVC-ColonDB, ETIS, and CVC-300 that were not involved in the training. As shown in Table 3, the mDice and mIoU metric values of the proposed PEEA-Net method on these three unseen datasets are better than those of other methods, achieving better segmentation performance and showing stronger generalization ability. In particular, on the most challenging ETIS dataset, compared with the top-ranked MSGAT among the comparison methods, the mDice and mIoU values of PEEA-Net are increased by 2.64% and 2.69% respectively. In addition, on the CVC-ColonDB dataset, the mDice and mIoU values of PEEA-Net are 0.69% and 0.89% higher than those of the second-best CASF-Net among all methods respectively; on the CVC-300 dataset, our evaluation metric values also have a similar leading advantage. Generally speaking, the experimental results of the PEEA-Ne of this invention have achieved the best segmentation performance on 5 publicly available colonoscopy datasets, verifying that PEEA-Net has good robustness and generalization ability.
[0169] Table 2 Quantitative Experimental Results of Different Methods on the Kvasir-SEG Dataset and the CVC-ClinicDB Dataset
[0170]
[0171]
[0172] Table 3 Comparison of Quantitative Results of Different Methods on 3 Unused Polyp Datasets for Training: CVC-ColonDB, ETIS, and CVC-300
[0173]
[0174]
[0175] 3.6 Comparative Analysis of Model Computational Complexity
[0176] In this experiment, the number of parameters (in millions, M) and the number of floating-point operations (GFLOPs) are used to measure the computational complexity of PEEA-Net, and it is compared with 18 state-of-the-art (SOTA) models. Table 4 shows the comparison of the number of parameters and GFLOPs of different methods. In Table 4, SwinPA-Net and MSRAformer use serial channel-spatial attention mechanism and reverse-spatial attention mechanism respectively. Compared with these serial attention models, PEEA-Net has fewer parameters (Params) and lower GFLOPs. CASF-Net uses PVTv2 and Res2net as encoder networks and combines with a parallel cross-attention mechanism. Compared with CASF-Net, the number of parameters and GFLOPs of PEEA-Net are significantly reduced by 11.11 (M) and 9.15 (G) respectively. On the one hand, it is because the CNN-assisted encoder proposed by us can reduce the number of model parameters without degrading the model performance. On the other hand, it is because the parameter-sharing mechanism of the proposed ESPA not only solves the problem of redundant calculations that may occur in parallel processing but also reduces the number of model parameters. In addition, compared with other CNN-Transformer hybrid methods, such as TransUNet and TGDAUNet, PEEA-Net has significant advantages in terms of the number of parameters and GFLOPs. Although the Params and GFLOPs of other single-encoder architecture networks are lower than those of PEEA-Net, especially, VMUNet2 performs optimally among all methods, however, compared with these methods, PEEA-Net has obvious advantages in segmentation performance. In the laparoscopic kidney tumor segmentation task, it is more important to improve the segmentation performance while ensuring real-time performance (without significantly increasing the computational complexity). Generally speaking, the PEEA-Net of the present invention achieves a good balance between computational complexity and segmentation performance.
[0177] Table 4 Comparison of the Number of Parameters and GFLOPs of Different Methods
[0178]
[0179] 3.7 Ablation Study
[0180] In the PEEA-Net of the present invention, we use PVT v2 as the main encoder and propose a new CNN encoder as the auxiliary encoder. In addition, we also propose three new modules: PEA, ESPA, and LDISF. To verify their effectiveness, we conducted ablation experiments on the laparoscopic kidney tumor dataset RT-AMU.
[0181] Specifically, for the main encoder network, we selected 4 of the most popular current backbone networks for comparative analysis, including Res2Net50, Swintransformer-B, Swintransformer-S, and PVT v2. For the verification of the PEA module and the ESPA module, we directly removed them from the PEEA-Net; for the LDISF module, we used channel connection operations and 1 1×1 convolutional unit for replacement; for the CNN auxiliary encoder, we used Res2Net50 to replace the proposed CNN auxiliary encoder for comparative verification; for the entire CNN branch, we directly removed it to verify its effectiveness. We labeled the above ablation experiments as "w / o PEA", "w / o ESPA", "w / o LDISF", "w / o CNN", "w / o CNN-LDISF" in sequence. Table 5 shows the quantitative results of different models on the kidney tumor dataset RT-AMU in the ablation experiments and their comparison of computational complexity. As shown in Table 5, compared with other backbone networks, PVT v2 performs best in all indicators, and its mDice and mIoU values exceed those of the second-place Swin Transformer-S by 0.97% and 0.52% respectively. In addition, the number of parameters and GFLOPs of PVT v2 are also significantly lower than those of Swin Transformer-S. This proves that PVT v2 has a stronger ability to extract global feature information of kidney tumors, verifying the effectiveness of the main encoder network.Compared with PEEA-Net, the mDice and mIoU values of "w / o PEA" decreased by 0.47% and 0.62% respectively, indicating that the PEA module improved the feature expression ability and multi-scale adaptation ability of the Transformer, enabling the network to more fully capture the features of kidney tumors of various sizes, shapes, types, and quantities; in terms of the mDice and mIoU metrics, the performance of "w / o ESPA" decreased by 1.08% and 1.67% compared to PEEA-Net, which proves that ESPA enhanced the ability of the Transformer to capture global features and image-specific features, improving the model segmentation accuracy; the number of parameters and GFLOPs of "w / o CNN" increased by 10.65 (M) and 2.82 (G) respectively compared to PEEA-Net, while the segmentation performance remained the same, which proves that the proposed auxiliary CNN encoder can reduce the computational complexity without reducing the segmentation performance; compared with PEEA-Net, the mDice and mIoU values of "w / o CNN-LDISF" decreased by 0.60% and 0.86% respectively, while the HD value increased by 0.148, which proves that the CNN branch can optimize the feature representation output by the Tranformer branch decoder and refine the boundary of the predicted mask output by the network; compared with PEEA-Net and "w / o CNN-LDISF", the mDice value of "w / o LDISF" decreased by 0.79% and 0.19% respectively, which proves that LDISF can not only fully and effectively fuse local features with global features, but also eliminate the decline in model performance caused by irrelevant background noise. Generally speaking, these quantitative experimental results verify the effectiveness of the modules we proposed.
[0182] Table 5 Quantitative results of different models on the kidney tumor dataset RT-AMU in the ablation experiment and comparison of their computational complexity
[0183]
[0184]
[0185] IV. Conclusion
[0186] The present invention proposes a new network for automatically and accurately segmenting renal tumors in laparoscopic images: PEEA-Net. This network is a CNN-Transformer hybrid network with a parallel structure, capable of achieving more fine-grained predictions while accurately locating the position of renal tumors. Among them, the PEA module solves the problems of limited feature expression ability and multi-scale adaptation ability of the Transformer encoder; the ESCD module solves the problem of insufficient ability of the network to adaptively select key feature maps and key regions in the feature maps; the LDISF module solves the problem that small renal tumors and partial boundary regions may be lost in the decoding stage; in addition, the auxiliary CNN encoder also reduces the computational complexity of the dual-backbone encoder network.
[0187] A large number of quantitative and qualitative experimental results on the laparoscopic renal tumor dataset RT-AMU created in this embodiment show that, compared with other SOTA methods, the PEEA-Net proposed by the present invention has better segmentation performance and can more effectively cope with the challenges brought by the different sizes and types of renal tumors in laparoscopy, low contrast between tumors and normal healthy tissues, small renal tumors, and uneven or low brightness of laparoscopic images; at the same time, a large number of quantitative experimental results on the publicly available polyp dataset also verify that PEEA-Net has stronger generalization ability. As a deep learning-based automatic renal tumor segmentation technology, PEEA-Net can assist doctors in treating renal cancer during partial nephrectomy, reduce the complexity of doctors' surgical operations, improve the clinical treatment process, and has important clinical significance. In a further implementation, more effective laparoscopic renal tumor images can be collected to enhance the dataset and further improve the computational efficiency of the model.
[0188] In addition, corresponding to Figure 2 the method described above, the embodiment of the present invention also provides a deep learning-based laparoscopic image renal tumor segmentation system for Figure 2 the specific implementation of the method. The deep learning-based laparoscopic image renal tumor segmentation system provided by the embodiment of the present invention can be applied to computer terminals or various mobile devices, and specifically includes:
[0189] A dataset construction module, used to construct the laparoscopic renal tumor dataset RT-AMU, and divide the RT-AMU dataset into a training set and a test set according to a preset ratio;
[0190] A model building module, used to design a parallel structure CNN-Transformer hybrid network PEEA-Net as the initial renal tumor segmentation model;
[0191] A model training module, configured to input a training set into an initial kidney tumor segmentation model for model training until the loss function converges, so as to obtain an optimized kidney tumor segmentation model;
[0192] An image segmentation module, configured to input a test set into the optimized kidney tumor segmentation model to obtain the kidney tumor segmentation result of a laparoscopic image.
[0193] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference may be made to the description in the method part for the relevant parts.
[0194] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A deep learning-based laparoscopic image renal tumor segmentation method, characterized in that: The following steps are involved: Construct the laparoscopic renal tumor dataset RT-AMU and divide the RT-AMU dataset into training set and test set according to the preset ratio; Design a parallel CNN-Transformer hybrid network PEEA-Net as the initial segmentation model for kidney tumors; The training set is input into the initial segmentation model of kidney tumor for model training until the loss function converges to obtain the optimized kidney tumor segmentation model; The test set was input into the optimized renal tumor segmentation model to obtain the renal tumor segmentation results of laparoscopic images.
2. The method for segmenting kidney tumors in laparoscopic images based on deep learning according to claim 1, characterized in that: The laparoscopic renal tumor dataset RT-AMU is constructed as follows: Several robot-assisted partial nephrectomy videos were obtained and laparoscopic images were extracted from the videos. Professional urology experts manually annotated the laparoscopic images and generated the corresponding true value GT.
3. The method for segmenting kidney tumors in laparoscopic images based on deep learning according to claim 1, characterized in that: The parallel structured CNN-Transformer hybrid network PEEA-Net includes: PVTv2 encoder, CNN auxiliary encoder, PEA module, ESPA module, AG module and LDISF module; PVTv2 encoder, used to extract global information, obtains four different levels of pyramid features Ti from laparoscopic images, where i = 1, 2, 3, 4; CNN auxiliary encoder, used to extract local detail information, obtains three feature maps Ci of different scales from laparoscopic images, where i = 1, 2, 3; PEA module, used to extract and gradually fuse pyramid features Ti in groups to obtain enhanced features Tpi, where i = 1, 2, 3, 4; AG module, which is used to combine the enhanced low-level features Tpi with the upsampled high-level features UpConv (E i+1 ) to obtain fusion features, where i = 1, 2, 3; The ESPA module is used to upsample the high-level features UpConv (E i+1 ) are connected to generate the feature map E i ' to refine the spatial and channel features in the output feature map Ei containing important feature information, where i = 1, 2, 3; The LDISF module is used to supplement the local detail feature information contained in the feature map C3 obtained by the CNN auxiliary encoder to the global feature map E1 output by the Transformer branch decoder, and generate a feature map E0 that can locate the position of the kidney tumor and outline the boundary of the kidney tumor.
4. The method for segmenting kidney tumors in laparoscopic images based on deep learning according to claim 3, characterized in that: The CNN-assisted encoder consists of three CNN blocks. The first CNN block contains a group of 3×3 convolutional units, and the last two CNN blocks consist of a maximum pooling downsampling operation and two groups of 3×3 convolutional units. Among them, the convolution unit is used to capture local structure and texture features; the pooling window size of the maximum pooling downsampling operation is 2×2, which is used to halve the size of the feature map and double the dimension.
5. The method for segmenting kidney tumors in laparoscopic images based on deep learning according to claim 3, characterized in that: The PEA module uses grouped convolution operations with different receptive field sizes and a progressive fusion strategy to enrich and enhance features, specifically: The multi-layer features Ti output by the PVTv2 encoder are split along the channel dimension to generate four groups of features Tij, where i = 1, 2, 3, 4; j = 1, 2, 3, 4; Each set of features is subjected to a customized convolution operation, where: for the first set of features Ti1, a 1×1 convolution unit is used to extract features; for the second set of features Ti2, a 3×3 convolution unit with an expansion rate of 3 is used to capture multi-scale information; for the third set of features Ti3, an asymmetric convolution consisting of a 1×3 horizontal kernel, a 3×1 vertical kernel, and a 3×3 convolution kernel is used to capture the complex asymmetric features in the laparoscopic image; for the fourth set of features Ti4, a self-calibrated convolution that can automatically adjust the convolution kernel weights according to the input features is used to capture rich detail information; The adaptive average pooling layer is concatenated with the results of the first and second branches and the third and fourth branches in the channel dimension, and the two features A1 and A2 are output and then fused with the global information Pi3. Then, a 1×1 convolution unit is used to restore the original number of channels, and finally a residual connection is made with the original feature Ti to output the enhanced feature Tpi.
6. The method for segmenting kidney tumors in laparoscopic images based on deep learning according to claim 3, characterized in that: The ESPA module adopts a parameter-sharing non-local interactive attention mechanism, specifically: The reshaped input features r(Ei') and position information are combined together, and then the serialized tokens are obtained through layer normalization; the serialized tokens are respectively generated into channel values V through three parallel linear fully connected layers C , shared query Q, shared key K and space value V S , where r(·) represents the feature reshaping operation, i=1,2,3,4; Calculate the similarity QK between the shared query Q and key K T , and then the weight of the spatial direction is obtained through the Softmax function Then with the space value V S Multiply them together to get the output spatial feature map Espatial; similarly, the channel attention weight is calculated by sharing the query Q and key K Then with the channel value V C The transposed matrix of Multiply to get the output channel feature Echannel; where, Represents the dimension of feature embedding; Espatial and Echannel are added to the input features and concatenated in the channel dimension, and a 1×1 convolution is used to restore the original number of channels to obtain O SRA , and then O SRA After being added to the input features for feature enhancement and noise suppression, it is input into the Feed-Forwrad module and a residual connection operation is used at the same time. Finally, a local enhancement module is used to obtain the output feature Ei.
7. The method for segmenting kidney tumors in laparoscopic images based on deep learning according to claim 3, characterized in that: The LDISF module effectively integrates local features with global features, specifically: Connect the global feature map E1 output by the Transformer branch decoder and the third-layer local feature map C3 obtained by the CNN auxiliary encoder in the channel dimension to obtain the initial fusion feature Ftc; The initial fusion feature Ftc is divided into four branches, and different convolution operations are used to mine the significant information in the initial fusion feature Ftc, where: a 1×1 convolution is used to reduce the number of feature channels of each branch to half of Ftc; in the second branch, a convolution branch consisting of a 1×3 convolution, a 3×1 convolution, and a 3×3 convolution with an expansion rate of 3 is used to expand the receptive field and mine the fine-grained features of the laparoscopic image in different directions; in the third branch, a third convolution branch consisting of a 3×3 convolution and a 3×3 convolution with an expansion rate of 5 is used to deepen the exploration and mining of features; in the fourth branch, variable convolution is used to capture the boundary features in irregularly shaped kidney tumors; The outputs of the four branches are connected in the channel dimension, and a 3×3 convolution is used for denoising and restoring the original number of channels, and then a residual connection is performed with the initial fusion feature Ftc. Then, an ECSA module is used to remove background noise and output a global reliable feature E0 containing fine-grained detail information.
8. The method for segmenting kidney tumors in laparoscopic images based on deep learning according to claim 3, characterized in that: Also includes: A 1×1 convolution is used to extract the feature map Tp4 output by the PEA module to obtain the feature map E'4, and the feature map E4 is obtained through the ESPA module; Ei undergoes a 1×1 convolution and upsampling operation to generate a prediction feature map Di with a channel number of 1 and the same size as the image to be segmented. Di then passes through a Sigmoid function to generate a prediction mask Pi; For the predicted feature map Di output by the five stages, the final predicted feature map D is obtained by additive aggregation: Among them, i=0,1,2,3,4.
9. The method for segmenting kidney tumors in laparoscopic images based on deep learning according to claim 8, characterized in that: The loss function is a multi-stage loss function, expressed as: L i =L wiou (D i ,G)+L wbce (D i ,G),(i=0,1,2,3,4) Where: Loveall represents the total loss function, Li represents the loss functions of the five stages respectively, wiou loss constrains the prediction results of kidney tumors in laparoscopic images from a global perspective, and wbce loss constrains the prediction results of kidney tumors in laparoscopic images from a local perspective; G represents the ground truth.
10. A system using the deep learning-based laparoscopic image renal tumor segmentation method according to any one of claims 1 to 9, characterized in that: include: A data set construction module is used to construct the laparoscopic renal tumor data set RT-AMU, and divide the RT-AMU data set into a training set and a test set according to a preset ratio; The model building module is used to design the parallel structured CNN-Transformer hybrid network PEEA-Net as the initial segmentation model for kidney tumors; A model training module is used to input the training set into the initial segmentation model of kidney tumors for model training until the loss function converges to obtain an optimized kidney tumor segmentation model; The image segmentation module is used to input the test set into the optimized kidney tumor segmentation model to obtain the kidney tumor segmentation results of the laparoscopic image.
Citation Information
Cited By
2D medical image segmentation method and system based on Mama and UNet
CN120997233A
A 2D medical image segmentation method and system based on Mamba and UNet
CN120997233B