Transform-CNN medical image segmentation method and system based on multi-scale fusion semantic enhancement

The dual-encoder architecture with enhanced attention mechanisms addresses the limitations of Transformer-based methods by integrating deep and shallow features, enhancing segmentation performance and adaptability in medical image analysis.

CN120318256APending Publication Date: 2025-07-15ANHUI UNIV
View PDF 0 Cites 14 Cited by

Patent Information

Application Number
CN202510490634.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing medical image segmentation method based on Transformer performs excellently in global context information acquisition, but there is a problem of insufficient modeling of local feature dependencies, resulting in limited capture of fine-grained detailed information, especially in medical images, which is prone to spatial position deviations.

Method used

The dual encoder structure is adopted, combined with the large-core packet transform attention module and the multi-layer cross attention fusion module, global and local information are extracted through the PVTv2 and Res2Net50 encoders, and feature fusion is used to enhance attention modules to optimize the cross-scale feature transmission and local detail recovery of the decoder.

Benefits of technology

It improves the performance of medical image segmentation, enhances the multi-scale adaptability and feature characterization effect to the lesions, optimizes the recovery ability of local detailed features, and significantly improves the accuracy and robustness of tumor segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318256A_ABST
    Figure CN120318256A_ABST
Patent Text Reader

Abstract

The invention discloses a Transform-CNN medical image segmentation method and system based on multi-scale fusion semantic enhancement, and relates to the technical field of image segmentation, and the method comprises the steps: collecting medical image data, and constructing a medical image data set; constructing a medical image segmentation network model; training the medical image segmentation network model through the medical image data set; and inputting real-time acquired data into the trained medical image segmentation network model to obtain a medical image segmentation result. According to the invention, double encoders, LGDA modules and the like are adopted to capture multi-scale global features of kidney tumors and enhance representation, and the LGDA module is used for adapting to size and form changes of the kidney tumors; the MLCF module is used for supplementing information of the main encoder; the PSE module captures multi-scale global semantic information and integrates local context information, local features are fused through the SAM module, and kidney tumor positioning in an endoscope image is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image segmentation, and more specifically, to a Transformer-CNN medical image segmentation method and system based on multi-scale fusion semantic enhancement. Background Art

[0002] As a clinical routine diagnosis and treatment technique, the image visualization ability of endoscopy provides doctors with an intuitive way to observe the morphological structure of organs and tissues. Achieving accurate segmentation of organs and diseased tissues in endoscopic images is the key technical basis for assisting clinical decision-making (such as tumor boundary definition, polyp size measurement). Medical image segmentation, as an important part of medical image analysis, is based on relevant knowledge in the field of computer vision in deep learning and plays an important role in assisting doctors in clinical diagnosis. However, the recognition of medical images often faces various challenges, whether it is the low resolution of the images, the recognition of the structure of the target tissue, or the changes in various positions, shapes, and sizes. Under so many uncontrollable conditions, the method of manual recognition makes doctors feel time-consuming and laborious, and is prone to certain misdiagnosis and missed diagnosis, and such results will lead to a significant decrease in the survival rate of patients. The purpose of MIS is to accurately identify tissues, lesions, and organs in medical images, such as kidney tumor segmentation in endoscopy, polyp segmentation in colonoscopy, etc. This process is crucial for clinicians to conduct qualitative and quantitative evaluations of various anatomical analysis structures or pathological conditions and to formulate subsequent treatment plans.

[0003] In recent years, deep learning has been continuously developing, which has promoted the continuous progress and optimization of medical image segmentation, thus greatly improving the segmentation accuracy of many visual tasks. Currently, UNet based on convolutional neural network (CNN) is a classic framework for medical image segmentation. It constructs layers through local receptive fields and pooling operations to capture the spatial hierarchical structure in the image, and effectively retains the features of the image at different scales through a series of operations such as skip connections. Later, many variants have emerged on the basis of this framework, such as unet++, res-unet, trans-unet, caranet, vm-unet, u-kan, etc. Although CNN can expand the receptive field and obtain deeper spatial features through continuous downsampling operations, the problems of extracting global information and long-range dependence relationships between images have not been solved.

[0004] It wasn't until the emergence of the Transformer, especially the Vision Transformer (ViT) designed for vision tasks, that the way of processing images differed from traditional convolutional neural networks (CNNs). In contrast, the Vision Transformer (ViT) treats an image as a sequence of patches, similar to the way text is processed in natural language processing (NLP) tasks. Each image is segmented into a smaller, flattened grid of patches, which are then linearly embedded and located through learned positional embeddings. The self-attention mechanism of the Transformer allows each patch to interact with all other patches in the sequence, enabling the model to capture the global dependencies and context information of the entire image. Nowadays, ViT and its variants have achieved good results in multiple medical image segmentation tasks. For example, the moving window mechanism was introduced in the Swin-Transformer proposed by Liu et al. Self-attention calculations are performed in non-overlapping local windows at each stage, and window shifting occurs between adjacent stages to ensure the dependency relationship between windows. The H2Former model proposed by He et al. adopts a hierarchical hybrid structure, integrating the advantages of the local information of CNNs, the multi-scale channel attention features (MSRA), and the long-range features of the Transformer, effectively improving the accuracy of image segmentation. The spatial inverse attention module proposed by Wu et al. in the MSRAFormer can further identify the edge features and detailed information of the target region, etc.

[0005] Although current Transformer-based medical image segmentation methods can effectively model global context associations and the Transformer is also excellent in obtaining global context information, however, there is a bottleneck problem of insufficient modeling of local feature dependencies, and there are still limitations in capturing fine-grained detailed information. Especially for medical images, although the Transformer establishes long-range dependencies between global features at each layer, due to some lack of local information, the results may have certain spatial position deviations. These methods often ignore the rich spatial information in the shallow network, and they all model the context at the same scale, also ignoring certain cross-scale dependencies and consistencies.

[0006] Therefore, how to propose a Transformer-CNN medical image segmentation method and system based on multi-scale fusion semantic enhancement to retain the semantic information of deep features, while continuously fusing the tumor details and edge information of different dimensions of shallow features, improving the segmentation performance, the multi-scale adaptability of the network to lesions, the characterization effect of features, and optimizing the ability to recover local detailed features is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a Transformer-CNN medical image segmentation method and system based on multi-scale fusion semantic enhancement, which preserves the semantic information of deep features, and at the same time continuously fuses the tumor details and edge information of different dimensions of shallow features, improving the segmentation performance, the multi-scale adaptability of the network to lesions, the characterization effect of features, and optimizing the problem of insufficient ability to restore local detail features. To achieve the above objectives, the present invention adopts the following technical solutions:

[0008] A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement, comprising:

[0009] Collect medical image data and construct a medical image data set;

[0010] Construct a medical image segmentation network model, adopt a dual-encoder structure to couple a large-kernel grouped transformation attention module and a multi-layer cross-attention fusion module respectively, connect the large-kernel grouped transformation attention module to a parallel semantic enhancement attention module, and use a part of the output of the multi-layer cross-attention fusion module to optimize the parallel semantic enhancement attention module, and the other part is subjected to similarity aggregation output with the final output of the parallel semantic enhancement attention module;

[0011] Train the medical image segmentation network model through the medical image data set;

[0012] Input the real-time collected data into the trained medical image segmentation network model to obtain a medical image segmentation result.

[0013] Optionally, the adoption of the dual-encoder structure includes: using pvtv2 as the main encoder to capture global information, and using three-layer res2net50 as the auxiliary encoder to provide supplementary local information.

[0014] Optionally, the adoption of the dual-encoder structure to couple a large-kernel grouped transformation attention module and a multi-layer cross-attention fusion module respectively includes:

[0015] Adopt a dual-branch encoder to process the input image X∈R H×W×3 , and obtain feature maps from the CNN encoder and the PV Tv2 encoder. In the CNN encoder part, use Res2Net50 to perform three residual convolution operations on the input image, so as to obtain features T i (i = 1, 2, 3) at three scales, and then perform multi-scale feature fusion on these three layers of features through a multi-layer cross-attention fusion module to obtain an initial segmentation prediction map F1. In the PVTv2 encoder part, use the PVTv2 encoder part to obtain four different levels of features X i(i = 1, 2, 3, 4), the large kernel grouped transformation attention module is used for feature extraction and edge information enhancement X' i (i = 1, 2, 3, 4).

[0016] Optionally, connecting the large kernel grouped transformation attention module to the parallel semantic enhancement attention module, and using a part of the output of the multi-layer cross-attention fusion module to optimize the parallel semantic enhancement attention module includes: in the decoder part, using the parallel semantic enhancement attention module to retain and layer-by-layer transfer the feature information of each layer, and obtaining the decoder feature D of each layer i (i = 1, 2, 3, 4), where, at the deepest layer, 1×1 convolution and sampling operations are performed to adjust the channels and size of the F1 feature to align and fuse with D4, and some local information not available in the deep features is supplemented.

[0017] Optionally, the other part performs similarity aggregation output with the final output of the parallel semantic enhancement attention module, including: using a residual operation to restore the feature map through top-down cross-scale transfer, introducing a spatial attention module, and performing feature aggregation on the predicted feature map D1 generated by the decoder and the CNN fusion feature F1 to generate the final predicted feature map P1.

[0018] Optionally, the structure of the large kernel grouped transformation attention module is as follows:

[0019] Extract multi-scale features and enhance local detail information through two branches respectively, and optimize the feature representation by combining the attention mechanism;

[0020] The first branch uses three different sizes of convolutional kernels to capture multi-scale features: 3×3 depthwise separable convolution S 11 For efficient extraction of local features; 5×5 depthwise separable convolution S 12 Capture more extensive context information; DCN deformable convolution S 13 Then, by adaptively adjusting the position and shape of the convolutional kernel, detect the target edge and refine the features;

[0021] The output feature maps of the three branches are fused by pixel-by-pixel addition, denoted as S1. The second branch S2 focuses on enhancing local detail information and uses a receptive field attention convolution module. The output feature maps of the two branches are fused by element-wise addition, and the ReLU activation function is applied to combine global semantic information and local detail features;

[0022] Apply the ESA attention module to the fused feature map to obtain O LGDA , by learning the importance weights in the spatial dimension, enhance the model's attention ability to the target area and reduce the influence of background noise.

[0023] Optionally, the multi - layer cross - attention fusion module includes:

[0024] Receives three - layer feature information from the CNN backbone branch, then re - aggregates the information to generate fused high - level semantic feature information for supplementing the local information of the main network;

[0025] Performs a path alignment mechanism on these three layers of information T i (i = 1, 2, 3), where a differential processing strategy is adopted. Through sampling and convolution operations, the features of different input layers are aligned. For the deep feature T3, up - sampling is performed through pooling operations and combined with 1×1 depth convolution to restore the spatial resolution and maintain channel consistency; for the shallow feature T1, down - sampling is performed through pooling operations and 1×1 convolution to reduce the channel dimension to 256, making all features consistent in both spatial and channel dimensions;

[0026] Adopts a multiplicative fusion framework to perform cross - level multiplicative fusion between the aligned features. First, multiplicative fusion of T2 and T3 is performed through the EFF module to obtain T 21 , then a concatenation operation is performed on T 21 and T2 to enhance the intermediate - level features and promote cross - layer information fusion. Finally, cross - attention fusion is performed on T 21 and T 12 using non - linear operations and attention mechanisms to dynamically adjust the weights of the cross - fused features. By introducing weight parameters α and β, the fusion effect of high - level semantic information and low - level details is balanced to generate the final output O MLCF .

[0027] Optionally, the parallel semantic enhancement attention module includes:

[0028] Constructs a parallel semantic enhancement attention module composed of two branches to model these two different semantics respectively;

[0029] In branch 1, in the manner of PVT, the input features are embedded into overlapping patches using convolutional layers. Subsequently, the obtained serialized tokens are used to calculate the query Q, key K, and value V respectively through three parallel fully - connected layers. The relationship between each pixel point in the feature map and the pixels in the same row and column is calculated using Q and K through the affinity matrix, and then a normalization operation is performed through the softmax function to obtain the output D. The Aggregation operation is performed on D, V values, and H to obtain the output C. Finally, residual operations and multiplicative operations are used to enrich the long - connected context information to obtain the final output H';

[0030] In Branch 2, the local fusion information Z output by the CNN encoder is subjected to position matching and embedding, and DCNv2 and adaptive weighted fusion operations are adopted to adapt the local information to the irregular structure, and learnable weight parameters are used to fuse the local features with the decoder part. Finally, for dimension alignment and subsequent processing, the outputs of Branch 1 and Branch 2 are fused using a product operation.

[0031] Optionally, the loss of the medical image segmentation network model includes:

[0032] A multi-loss joint supervision strategy is proposed, and the overall objective function L is constructed by weighted combination of multiple sub-loss functions total :

[0033] L total = λ1L seg + λ2L dice ;

[0034] L seg = L miou = +L bce ;

[0035]

[0036] where L bce is the weighted binary cross-entropy loss, L miou is the mean intersection over union loss, y i is the ground truth label, o(y i ) is the probability that the model predicts a certain class label, L miou loss and L bce loss jointly constrain the segmentation result from two dimensions of global region matching degree and local pixel classification accuracy respectively, and L dice is the Dice loss function, which is used to emphasize the spatial overlap between the predicted boundary and the ground truth boundary.

[0037] Optionally, a Transformer-CNN medical image segmentation system based on multi-scale fusion semantic enhancement includes:

[0038] An acquisition module: used to acquire medical image data and construct a medical image dataset;

[0039] A model construction module: used to construct a medical image segmentation network model, which adopts a dual-encoder structure to couple a large kernel grouped transformation attention module and a multi-layer cross-attention fusion module respectively, connects the large kernel grouped transformation attention module to a parallel semantic enhancement attention module, and uses part of the output of the multi-layer cross-attention fusion module to optimize the parallel semantic enhancement attention module, and the other part is used for similarity aggregation output with the final output of the parallel semantic enhancement attention module;

[0040] Training module: used to train the medical image segmentation network model through the medical image dataset;

[0041] Image segmentation module: used to input the real-time acquired data into the trained medical image segmentation network model to obtain the medical image segmentation result.

[0042] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a Transformer-CNN medical image segmentation method and system based on multi-scale fusion semantic enhancement, having the following beneficial effects:

[0043] The present invention proposes a Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement, including: collecting medical image data to construct a medical image dataset; constructing a medical image segmentation network model, adopting a dual-encoder structure to couple a large-kernel grouped transformation attention module and a multi-layer cross-attention fusion module respectively, connecting the large-kernel grouped transformation attention module to a parallel semantic enhancement attention module, using a part of the output of the multi-layer cross-attention fusion module to optimize the parallel semantic enhancement attention module, and performing similarity aggregation output on the other part and the final output of the parallel semantic enhancement attention module; training the medical image segmentation network model through the medical image dataset; inputting the real-time acquired data into the trained medical image segmentation network model to obtain the medical image segmentation result.

[0044] The present invention proposes a medical image segmentation network MFSE-Net based on the CNN-Transformer structure. First, in the encoder part, a dual-encoder design idea is adopted. The pvtv2 is used as the main encoder to capture global information, and the three-layer res2net50 is used as the auxiliary encoder to provide supplementary local information. Then, a large-kernel grouped deformable attention module (LGDA) is proposed. In the pvtv2 part, the shape of the target region at each scale is feature-extracted respectively, and the ability to understand the semantic information of the edge structure is enhanced. In the res2net50 part, the multi-layer cross-attention fusion module (MLCF) is used to perform partial upsampling and downsampling on the local information first, and then perform hierarchical aggregation to form an initial segmentation prediction map, which is used for the context information guidance and optimization of the global model in the subsequent pvtv2 part. In the decoder part, the key is to pay attention to the retention and cross-scale transmission of high-level feature information. Therefore, a parallel semantic enhancement attention (PSE) module is proposed, and the preliminary prediction map generated by the CNN part is used for semantic supplementation. At the end of the shallow decoder, by introducing the SAM module, the final similarity aggregation of local pixel information and global semantic clues is performed. Finally, the accurate localization of kidney tumors is achieved. On the kidney tumor dataset RE-TMRS established by the present invention, the present invention is superior to other state-of-the-art SOTA segmentation methods in terms of mDice, mIoU, MAE, accuracy, and HD metrics. The mDice and mIoU reach 91.02% and 84.13% respectively; meanwhile, the visual segmentation effect of this method is also better than that of other SOTA methods. On five publicly available polyp datasets (Kvasir-SEG, CVC-ClinicDB, CVC-colondb, ETIS, and CVC-300), the present invention consistently achieves SOTA segmentation performance in terms of mDice, mIoU, and HD metrics. Especially on the CVC-300 dataset, the mdice and mIoU results of MFSE-Net are improved by 1.01% and 1.15% respectively compared with CIFG-Net. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0046] Figure 1 Schematic flow chart of a Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement provided by the present invention.

[0047] Figure 2 Schematic structural diagram of LGDA provided by the present invention.

[0048] Figure 3 Schematic structural diagram of MLCF provided by the present invention.

[0049] Figure 4 Schematic structural diagram of PSE provided by the present invention.

[0050] Figure 5 Visualization result diagram of the segmentation of the kidney tumor dataset provided by the present invention.

[0051] Figure 6 Visualization result diagram of the ablation experiment of the kidney tumor dataset provided by the present invention. Detailed implementation manners

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0053] The embodiments of the present invention disclose a Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement, including:

[0054] Collect medical image data and construct a medical image dataset;

[0055] Construct a medical image segmentation network model, adopt a dual-encoder structure to couple a large-kernel grouped transformation attention module and a multi-layer cross-attention fusion module respectively, connect the large-kernel grouped transformation attention module to a parallel semantic enhancement attention module, use a part of the output of the multi-layer cross-attention fusion module to optimize the parallel semantic enhancement attention module, and perform similarity aggregation output on the other part and the final output of the parallel semantic enhancement attention module;

[0056] Train the medical image segmentation network model with the medical image data set;

[0057] Input the real-time acquired data into the trained medical image segmentation network model to obtain the medical image segmentation result.

[0058] In a specific embodiment, a Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement proposes a MFSE-Net (Multi-Fusion and semantic enhancement Network) network structure, including a PVTv2 encoder, a CNN encoder, a large-kernel grouped deformable attention (LGDA) module, a multi-layer cross-attention fusion (MLCF) module, and a parallel semantic enhancement (PSE) module, as well as a loss function for training the network. The specific steps are as follows:

[0059] The overall network architecture of MFSE-Net is as Figure 1 shown. The entire model adopts a dual-encoder single-decoder architecture. For the encoder, the PVTv2 based on the swin transformer is designed as the main encoder part, and the res2net50 based on CNN is used as the auxiliary encoder to supplement the local spatial information. Then, in the PVTv2 part, a large-kernel grouped deformable attention module (LGDA, large-kernel grouped deformable attention) is designed to extract the image pixel features at each scale of the PVT encoder part. In the resnet50 part, multi-scale feature fusion is performed through a multi-layer cross-attention fusion module (MLCF, Multi-layered cross-attention fusion) to supplement the local information. In the decoder part, a parallel semantic enhancement attention (PSE, parallel semantic enhancements) module is used to guide the decoding process to retain information during cross-scale transmission. Finally, the SAM module in Polyp PVT is used as the supplement of semantic information. Finally, an MFSE-Net model is proposed, which models in an end-to-end manner based on the PVTv2 architecture and combines the extracted information of resnet50 for complementarity.

[0060] In a specific embodiment, first process the input image X ∈ R H×W×3 , and obtain the feature maps from the CNN encoder and the PVTv2 encoder. In the CNN encoder part, Res2Net50 is used to perform three residual convolution operations on the input image to obtain the features T at three scalesi (i = 1, 2, 3), and then perform multi-scale feature fusion on these three layers of features through MLCF to obtain F1. In the PVTv2 encoder part, first use the PVTv2 encoder part to obtain four different levels of features X i (i = 1, 2, 3, 4). Subsequently, use the LGDA module for feature extraction and edge information enhancement X' i (i = 1, 2, 3, 4), and then use the PSE module in the decoder part to retain and layer-by-layer transfer the feature information of each layer to obtain the decoder feature D of each layer i (i = 1, 2, 3, 4). Among them, at the deepest layer, perform 1×1 convolution and sampling operations to adjust the channels and dimensions of the F1 features to align and fuse with D4, supplement some local information not available in the deep features, and then use the residual operation to perform cross-scale transfer in a top-down manner to restore the feature map. Finally, the SAM module is introduced to aggregate the predicted feature map D1 generated by the decoder and the CNN fusion feature F1 to prevent the loss of some semantic information during the information transfer between layers, thereby generating the final predicted feature map P1. The process is as follows:

[0061] X i = PVTv2(X), (i = 1, 2, 3, 4) (1)

[0062] X i′ = LGDA(X i ), (i = 1, 2, 3, 4) (2)

[0063] T j = Res2Net50(X j ), (j = 1, 2, 3) (3)

[0064] F1 = MLCF(T1, T2, T3) (4)

[0065] D4 = PSE(X4', Conv1(F1)) (5)

[0066] D i = PSE(X i ', upsample(D i+1 + Conv1(I)), (i = 1(I = D3'), 2(I = D4'), 3(I = F1))(6)

[0067] P1 = SAM(Conv(D1), Conv1(F1)) (7)

[0068] In the specific implementation, PVTv2 is used as the encoder to extract global information, and four different levels of pyramid features are obtained from PVTv2. Among them, X1 is regarded as the low-level feature, which contains rich texture, color and edge details, as well as more noise and irrelevant information; X2, X3 and X4 are regarded as high-level features, which contain more feature information capable of locating kidney tumors. In addition, Res2Net50 is used as the auxiliary encoder to supplement local information. In the CNN network, three different levels of features T i (i = 1, 2, 3) are obtained, which respectively contain edge, texture information in the low-level local features and semantic information in the high-level feature layers, etc.

[0069] In the specific implementation, in order to further improve the model's perception ability of complex structures and irregular boundaries in medical images, a new large kernel grouped transformation attention module (LGDA) is designed to supplement the output feature map of the basic network. As Figure 2 shown, this module extracts multi-scale features and enhances local detail information through two branches respectively, and further optimizes the feature representation by combining the attention mechanism. The first branch uses three different sizes of convolutional kernels to capture multi-scale features: 3×3 depthwise separable convolution S 11 is used to efficiently extract local features; 5×5 depthwise separable convolution S 12 captures a wider range of context information; DCN deformable convolution S 13 precisely detects the target edge and refines the features by adaptively adjusting the position and shape of the convolutional kernel. Subsequently, the output feature maps of these three branches are fused by pixel-wise addition, denoted as S1. The second branch S2 focuses on enhancing local detail information and uses the receptive field attention convolutional module (RFA). Different from the existing spatial attention mechanisms (such as the convolutional block attention module CBAM and the coordinated attention CA) that only focus on spatial features and fail to fully solve the problem of convolutional kernel parameter sharing, the RFA Conv not only focuses on the receptive field spatial features, but also provides effective attention weights for large-size convolutional kernels. This design enables the second branch to effectively capture long-range dependencies while enhancing local detail information, significantly improving the extraction ability of texture and detail features. The output feature maps of the two branches are fused by element-wise addition and the ReLU activation function is applied to ensure the effective combination of global semantic information and local detail features. Finally, an efficient ESA (Efficient Spatial Attention) attention module is applied to the fused feature map to obtain O LGDA By learning the importance weights in the spatial dimension, the model's attention ability to the target region is enhanced, while the influence of background noise is reduced.

[0070] The design of the LGDA module aims to integrate the advantages of multi-scale feature extraction, local detail enhancement, and attention optimization, thereby significantly improving the model's ability to process complex medical images. The specific operation process includes: first, reducing the channel dimension of the input feature map to 64 through 1×1 convolution to reduce the computational burden. Then, in the first branch, 3×3 and 5×5 depthwise separable convolutions and DNC deformable convolutions are used to extract multi-scale features and complete the fusion through pixel-wise addition. At the same time, in the second branch, RFAConv convolution is adopted to enhance local detail information. Finally, the output feature maps of the two branches are fused through element-wise addition and the ESA attention module is applied to optimize the feature representation. This module design not only improves the model's detection ability for hidden or small targets but also effectively improves the refined expression of the target semantic structure. The process is as follows:

[0071] S 11 = BConv3(Conv1(X i ),(i = 1,2,3,4) (8)

[0072] S 12 = BConv5(Conv1(X i ),(i = 1,2,3,4) (9)

[0073] S 13 = DCNConv3(Conv1(X i ),(i = 1,2,3,4) (10)

[0074] S1 = Conv1(Cat(S 11 ,S 12 ,S 13 )) (11)

[0075] S2 = RFAConv(Avgpool(X i ),(i = 1,2,3,4) (12)

[0076] O LGDAi = BConv3(ESA(S1+S2)) (13)

[0077] In the above formulas, BConvi(i = 3,5) represents 3×3 convolution and 5×5 convolution, as well as batch normalization processing and ReLU activation function respectively. DCNConv3 represents 3×3 deformable convolution, as well as batch normalization processing and ReLU activation function. Cat(·) represents the concatenation operation in the channel dimension. Conv1(·) represents 1×1 convolution. Avgpool represents the average pooling operation. RFAConv represents the receptive field expansion convolution. BConv3 represents 3×3 convolution, batch normalization, and ReLU activation function.

[0078] In the specific implementation, the designed MLCF module mainly receives three-layer feature information from the CNN backbone branch, and then re-aggregates the information to generate fused high-level semantic feature information for supplementing the local information of the main network. First, a path alignment mechanism is performed on these three layers of information T i (i = 1, 2, 3), where a differential processing strategy is adopted. Through sampling and convolution operations, the features of different input layers are aligned. For the deep feature T3, upsampling is performed through pooling operations and combined with 1×1 depth convolution to restore the spatial resolution and maintain channel consistency; for the shallow feature T1, downsampling is performed through pooling operations and 1×1 convolution to reduce the channel dimension to 256. This alignment strategy ensures that all features are consistent in both spatial and channel dimensions. Then, a multiplicative fusion framework is adopted to perform cross-level multiplicative fusion between the aligned features to ensure high-quality interaction between the features of each layer, promote information integration, and thus reduce unnecessary redundancy. First, multiplicative fusion of T2 and T3 is performed through the EFF module, and the formula is as follows:

[0079] T 21 = EFF(T2,T3) (14)

[0080] EFF(T2, T3) = Sigmoidonv1(ReLU(T2′ + T3′))) (15)

[0081] T i ′ = GAP(or)GMP(GConv(ReLU(Conv1(T i )))),(i = 2,3) (16)

[0082] Among them, GAP / GMP represents global average pooling or global max pooling, GConv represents grouped convolution, Sigmoid(·) and ReLU(·) represent the Sigmoid activation function and the ReLU activation function, and Conv1(·) represents 1×1 convolution.

[0083] Obtain T 21 Then, a concatenation operation is performed on T 21 and T2 to enhance the intermediate features and promote cross-layer information fusion, so that the fused features of T 22 and T1 can be more complete. Finally, cross-attention fusion is performed on T 21 and T 12 through the CAM module. Using non-linear operations and attention mechanisms, a dynamic weight adjustment is made to the cross-fused features. By introducing weight parameters α and β to balance the fusion effect of high-level semantic information and low-level details, semantic deviation caused by inherent rules is avoided, and the final output O MLCF is generated, asFigure 3 As shown, the process is as follows:

[0084] T 22 = Cat(T2, T 21 ) (17)

[0085] T 12 = Multi(T1, T 22 ) (18)

[0086] F1 = CAM(T 12 , T 21 ) (19)

[0087] CAM(T 12 , T 21 ) = Cat(T′ 12 , T′ 21 ) (20)

[0088]

[0089] In the above formulas, Conv1(·) represents a 1×1 convolution. Cat(·) represents a concatenation operation in the channel dimension, Multi(·) represents an element-wise multiplication operation, RFAConv represents a receptive field expansion convolution, R&P and R&T represent reconstruction and transpose operations on the input image to adapt it to the input of the convolutional layer. The weights α and β are initialized to 0 and are gradually learned.

[0090] In the specific implementation, in the decoder part, it is necessary to use the input global semantic clues and local feature information to gradually restore the feature detail map of each layer. To achieve this goal, as Figure 4 shown, a PSE module composed of two branches is designed to model these two different semantics respectively. In branch 1, an improved self-attention module is proposed to learn long-range dependencies. First, in the manner of PVT, the input features are embedded into overlapping patches using convolutional layers. Subsequently, the obtained serialized tokens are passed through three parallel fully connected layers to calculate the query Q, key K, and value V respectively. Then, Q and K are used to calculate through the affinity matrix to obtain the relationship between each pixel point in the feature map and the pixels in the same row and column. Then, normalization is performed through the softmax function to obtain the output D. Then, the D, V values, and H are subjected to an Aggregation operation to obtain the output C. Finally, residual operations and some multiplicative operations are used to further enrich the context information of this long connection to obtain the final output H′. The process is as follows:

[0091]

[0092] H′ = γ(Softmax(Dh )·V h + Softmax(D w )·V w ) + H(24)

[0093] Among them, d(i, u) represents the spatial position coordinates of each element in the feature map, Q u ∈R C′ represents the corresponding query vector, Ω u represents the set of key vectors at all positions in the same row and column as k and u, D ∈ R (H+W-1)×(H×W) represents the energy matrix, V represents the value matrix, H is the input feature map, and γ is a learnable scaling parameter used to control the intensity of the attention contribution.

[0094] In this branch, Q and K, which are used to calculate the affinity matrix, carry the inherent context statistical information, which is the key dependence on this pixel-level long connection relationship pursued by the attention mechanism.

[0095] In another branch, the local fusion information Z output by the CNN encoder needs to be subjected to certain position matching and embedding. Therefore, DCNv2 and adaptive weighted fusion operations are adopted, so that the local information can adapt to some irregular structures, prevent feature confusion caused by excessive offset, and use learnable weight parameters to make the local features better fuse with the decoder part. At the end of the module, for dimension alignment and subsequent processing, the two are fused using a product operation, and the process is as follows:[[]]

[0096] Z′ = α☉DCNv2(Z) + (1 - α)☉CA(Z)(25)

[0097] D i = AdoptiveFusion(BConv1(H′ + Z′)), (i = 1, 2, 3, 4)(26)

[0098] The PSE module efficiently fuses the feature information of Transformer and CNN, and can be used as the main building block of the decoder of the image segmentation network. In branch 1, due to the inherent similarity of Q and K, the channel attention descriptor and spatial attention descriptor calculated by them can filter out more valuable context key information. The adaptive learning mechanism in branch 2 can focus on the perception of relatively local information positions, so as to better fuse and match. This dual attention mechanism of global and local features enables the proposed PSE module to be used as the basic building block of the decoder, gradually parsing out a reliable prediction mask.

[0099] In the specific implementation, in order to further optimize the effect of image segmentation, the weighted binary cross-entropy loss L bce and the mean intersection over union loss L mio u are respectively used in the designed MFSE-net for mask supervision, which can assign higher weights to pixels that are difficult to segment. To enhance the edge information of the LG DA module and constrain the boundary error, the Dice loss is adopted. Therefore, a multi-loss joint supervision strategy is proposed, and an overall objective function L total :

[0100] L total = λ1L seg + λ2L dice (27)

[0101] L seg = L miou + L bce (28)

[0102]

[0103] where the true label is denoted as y i , and the probability that the model predicts a certain class label is denoted as o(y i ). In formula (28), the L miou loss and the L bce loss jointly constrain the kidney tumor segmentation results from two dimensions of global region matching degree and local pixel classification accuracy. Specifically, the L miou loss improves the overall segmentation effect by optimizing the overlap rate between the predicted region and the true annotation, while the L bce loss alleviates the common foreground-background class imbalance problem in medical images by assigning higher weights to tumor pixels. To address the special challenges of kidney tumor boundary segmentation, the Dice loss function is introduced to emphasize the spatial overlap between the predicted boundary and the true boundary, effectively overcoming the problem of inaccurate boundary segmentation caused by blurred tumor edges and unbalanced pixel numbers.

[0104] In the specific implementation, the above-mentioned Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement is experimentally verified, and the results are as follows:

[0105] MFSE-Net is implemented using PyTorch. After 150 training iterations, the model finally converges. To prevent overfitting, a variety of methods are adopted, including data augmentation (such as rotation, scaling, cropping, and color variation), adding regularization terms, and batch normalization.

[0106] (1) To comprehensively evaluate the performance of the proposed MFSE-Net network model in kidney tumor segmentation tasks, the RE-TMRS dataset containing kidney tumor images extracted from endoscopic videos was selected. In addition, five publicly available polyp datasets, all of which were derived from endoscopic videos, were selected to verify the robustness and generalization ability of MFSE-Net. The specific dataset information is shown in Table 1. A total of 2258 image data points were randomly selected from RE-TMRS for training, and the remaining 565 image data points were used for kidney tumor segmentation experimental tests. In the polyp segmentation experiment, the same data settings as PraNet were adopted. 900 polyp images were randomly selected from Kvasir-SEG and 550 polyp images were randomly selected from CVC-ClinicDB for model training. The remaining 100 and 62 images from Kvasir-SEG and CVC-ClinicDB were used to test the learning ability of the model, respectively. In addition, all images in CVC-ColonDB, ETIS, and CVC-300 were not involved in model training and were only used to test the generalization ability of the model.

[0107] Table 1 Details of the experimental datasets and the configuration of training and test data.

[0108]

[0109] (2) In this kidney tumor segmentation experiment, six widely used medical image segmentation evaluation metrics were selected to quantitatively evaluate the performance of the proposed MFSE-Net and 11 other state-of-the-art (SOTA) models, including the mean Dice coefficient (mDice), mean intersection over union (mIoU), mean absolute error (MAE), accuracy (Accuracy), weighted F-measure (Fb), and Hausdorff distance (HD). Specifically, mDice and mIoU are used to quantify the degree of overlap between the predicted region and the ground truth region; MAE is used to compare the pixel-level absolute value difference between the predicted map (P) and the ground truth map (G); is used to evaluate the structural similarity between the predicted result and the actual result; accuracy represents the proportion of pixels correctly predicted by the network in the total pixels of the image; is used to calculate the weighted average of precision and recall; and HD is used to measure the accuracy of boundary segmentation evaluation. Among these metrics, the closer the values of mDice, mIoU, Accuracy, and Fb are to 1, the better the segmentation effect; the closer the MAE value is to 0, the better the segmentation effect; and the lower the HD value, the better the segmentation effect. The definitions of these metrics are as follows:

[0110]

[0111] Among them, TP represents true positive, TN represents true negative, FP represents false positive, FN represents false negative, and N represents the total number of test images. In formula (34), W and H represent the width and height of the image respectively. In formula (36), β represents the weight coefficient, and ∥-b∥ in formula (37) represents the distance function.

[0112] (3) MFSE-Net is implemented using the Python 3.10 and PyTorch 1.7.1 frameworks. The training and evaluation of the model are carried out on an NVIDIA A800 GPU with 80GB of memory. The PVTv2 backbone network is initialized with pre-trained Transformer weights from ImageNet to accelerate the network convergence process. During the training process, the batch size is set to 16, the optimizer is set to AdamW, and the initial learning rate is set to 5×10 -5 , with a decay rate of 0.1, and the model is trained for 150 epochs. To accelerate the model's convergence process, after the 15th and 30th epochs, the learning rate is gradually reduced with a decay rate of 0.5. Additionally, the image resolution is adjusted to be segmented into 352×352, and a multi-scale training strategy of [0.75, 1, 1.25] is adopted to reduce the model's sensitivity to scale changes.

[0113] (4) To verify the effectiveness of the proposed MFSE-Net in kidney tumor segmentation, its quantitative results are compared with those of UNet, Swin-UNet, MD-ViT, TransUNet, PraNet, TGDA-Net, ACC-Unet, HardNet-SEG, Polyp-PVT, and CIFG-Net, and six medical image segmentation evaluation metrics: mDice, mIoU, MAE, accuracy, and HD are used to quantitatively evaluate the segmentation performance of all models on the kidney tumor dataset Re-TMRS. Table 2 shows the comparison of the quantitative results of different algorithms on the Re-TMRS dataset. As can be seen from Table 2, compared with other models, the proposed MFSE-Net achieves state-of-the-art segmentation performance on the Re-TMRS dataset in most metrics. In terms of the mDice and mIoU metrics, MFSE-Net improves by 0.79% and 0.89% respectively compared to the second-best CIFG-Net, indicating that MFSE-Net can better distinguish kidney tumors from the normal tissue background. The accuracy and of MFSE-Net also improve by 0.19% and 0.32% respectively compared to the second-best CIFG-Net. Additionally, compared with the second-best Polyp-PVT, the MAE of MFSE-Net is reduced by 0.07%, and compared with the second-best CIFG-Net, the HD of MFSE-Net is reduced by 0.09.

[0114] Table 2 Comparison of Quantitative Results of Different Methods on the Re-TMRS[xx] Dataset

[0115]

[0116]

[0117] Among them, ↑ indicates that the higher the value, the better; ↓ indicates that the lower the value, the better; the bold numerical values represent the best performance results.

[0118] (5) To more clearly and intuitively demonstrate the advantages of the proposed MFSE-Net, the qualitative results of MFSE-Net were compared with the results of other models, including UNet, Swin-UNet, MD-ViT, TransUNet, PraNet, TGDA-Net, ACC-Unet, HardNet-SEG, Polyp-PVT, and CIFG-Net. Figure 5 The results comparison of different models on the Re-TMRS dataset is shown respectively. It can be seen from the figure that compared with other models, the proposed MFSE-Net has more accurate prediction results. The proposed MFSE-Net is more sensitive to the boundary features of kidney tumors and can better outline the boundaries of kidney tumors and remove noise regions compared with previous model methods. This is because the CNN part can effectively restore local boundary information for the backbone network PVTv2 module, and the LGDA module also reduces the influence of noise and irrelevant information. Compared with other comparison algorithms, its prediction results are more robust because the proposed PSE module improves the multi-scale adaptability of the network to lesions, enabling it to adapt to various changes in the size and shape of kidney tumors. In addition, MFSE-Net can remove noise in images with low brightness and achieve more accurate kidney tumor segmentation. The comparison of qualitative visual results proves that MFSE-Net can better handle the challenges brought by different shapes and sizes of kidney tumors and uneven brightness. At the same time, this also verifies again the effectiveness and robustness of the proposed MFSE-Net in kidney tumor segmentation.

[0119] (6) To verify the robustness and generalization ability of the proposed MFSE-Net, the results of MFSE-Net on polyp segmentation were compared with those of UNet, Swin-UNet, MD-ViT, TransUNet, PraNet, TGDA-Net, ACC-Unet, HardNet-SEG, Polyp-PVT, and CIFG-Net. The segmentation performance of all models was quantitatively evaluated using three evaluation metrics: mDice, mIoU, and HD on polyp tumor datasets Kvasi r-SEG, CVC-ClinicDB, CVC-ColonDB, ETIS, and CVC-300. Tables 3 to 5 respectively show the quantitative result comparisons of different algorithms on the Kvasir-SEG dataset, C VC-ClinicDB dataset, CVC-ColonDB dataset, ETIS dataset, and CVC-300 dataset. As can be seen from Tables 3 to 5, the results of MFSE-Net on the five polyp datasets are better than those of other comparison methods, and most of the metrics reach the state-of-the-art segmentation performance. MF SE-Net has stronger feature learning ability than other comparison models. Among them, the mDice metric results of MFSE-Net on the Kvasi r-SEG dataset and CVC-ClinicDB dataset reached 91.02% and 91.62% respectively. On the CVC-ColonDB dataset, the mDice metric result of MFSE-Net is 1.01% higher than that of the second-ranked CIFG-Net and 2.12% higher than that of the third-ranked Polyp-PVT. On the ETIS dataset, the mDice results are 0.27% and 1.88% higher than those of the second-ranked Polyp-PVT and the third-ranked CIFG-Net respectively. On the CVC-300 dataset, the mDice and mIoU results of MFSE-Net are 1.06% and 0.75% higher than those of the second-ranked CIFG-Net respectively, while the HD decreased by 0.0099. From these results, it can be seen that MFSE-Net has stronger generalization ability than other models on the three unseen polyp datasets, thus verifying the robustness and generalization ability of MFSE-Net.

[0120] Table 3 Quantitative Results and Comparisons of Segmentations on CVC-ColonDB and CVC-ClinicDB Datasets

[0121]

[0122] ↑ indicates better when higher, ↓ indicates better when lower. Bold indicates the best result.

[0123] Table 4 Quantitative Results and Comparisons of Segmentations on Kvasir-SEG and ETIS Datasets

[0124]

[0125]

[0126] ↑ indicates the higher the better, and ↓ indicates the lower the better. Bold indicates the best result.

[0127] Table 5 Quantitative results and comparison of the CVC-300 dataset segmentation

[0128]

[0129] ↑ indicates the higher the better, and ↓ indicates the lower the better. Bold indicates the best result.

[0130] (7) In the proposed MFSE-Net, three new modules (LGDA, MLCF, and PSE) are proposed to improve the performance of kidney tumor segmentation. To verify the effectiveness of each module, ablation experiments are conducted on the Re-TMRS dataset to explore the impact of each module on the performance of kidney tumor segmentation. Specifically, the backbone network is PVTv2. To verify the effectiveness of the LGDA module, it is replaced with a 1×1 convolutional layer for comparison; for the verification of the MLCF module, it is replaced with 1x1 convolution and concat operations for multi-level feature fusion; for the PSE module, the decoder is reconstructed through a deconvolution module to verify its effectiveness. The ablation experiments of the LGDA, MLCF, and PSE modules are labeled as "w / oLGDA", "w / oMLCF", and "w / oPSE", respectively. The experimental results are shown in Table 6. Specifically, in terms of the m Dice and mIoU metrics, "w / oLGDA" decreased by 0.9% and 0.11% respectively compared to MFSE-Net, indicating that compared with the multi-scale features processed by the LGDA module, the original features extracted by the encoder have limited adaptability to the changes in tumor shape and size. Compared with MFSE-Net, the m Dice and mIoU metrics of "w / oMLCF" decreased by 1.28% and 0.59% respectively, and the HD increased by 0.0049, because the lack of local detail information led to a decline in segmentation performance. Compared with MFSE-Net, the mDice and mIoU metrics of "w / oPSE" decreased by 1.05% and 0.17% respectively, indicating that the lack of the PSE module will cause over-extraction of noise and partial loss of high-level semantic information in the layer-by-layer transmission from high-level features to low-level features. From Table 6 and Figure 6 It can be seen that replacing or removing any of the proposed modules will significantly reduce the performance of kidney tumor segmentation, demonstrating the effectiveness of the proposed modules. The segmentation performance of "w / oLGDA", "w / o MLCF", and "w / oPSE" is higher than the baseline, reflecting the effectiveness of the cooperation between the various modules.

[0131] Table 6 Quantitative results of ablation experiments

[0132]

[0133] ↑ indicates the higher the better, and ↓ indicates the lower the better. Bold indicates the best result.

[0134] The present invention proposes a new multi-scale fusion semantic enhancement efficient attention network (MFSE-Net) for endoscopic image segmentation of kidney tumors. The network adopts a PVTv2 encoder and four modules (LGDA, MLCF, PSE, SAM) to capture multi-scale global features of kidney tumors and enhance the representation. By fusing local features, the kidney tumors can be more effectively located in endoscopic images. Three new modules are proposed: LGDA, MLCF, and PSE. The LGDA module is used to obtain multi-scale features with channel weight information and partial local detail features to facilitate adaptation to the size and morphological changes of kidney tumors; the MLCF module integrates the detail information and boundary information in the local features at all levels of the CNNBlock for supplementing the information of the main encoder; the PSE module uses a dual attention mechanism to capture multi-scale global semantic information and also integrates local context information. A new kidney tumor dataset Re-TMRS is also established for comparative evaluation. The quantitative and qualitative results of MFSE-Net on the kidney tumor dataset Re-TMRS are better than those of 10 state-of-the-art (SOTA) methods, demonstrating the superiority of MFSE-Net in the kidney tumor segmentation task. At the same time, the visualized qualitative results also prove that MFSE-Net can better address the challenges brought about by the shape and size changes of kidney tumors and uneven brightness. The experimental results of MFSE-Net on five publicly available polyp datasets are better than the compared SOTA models, verifying the robustness and generalization ability of the proposed MFSE-Net.

[0135] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.

[0136] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement, characterized in that Including: Collect medical image data and construct a medical image dataset; Construct a medical image segmentation network model, adopt a dual-encoder structure to couple a large-kernel grouped transformation attention module and a multi-layer cross-attention fusion module respectively, connect the large-kernel grouped transformation attention module to a parallel semantic enhancement attention module, and use part of the output of the multi-layer cross-attention fusion module to optimize the parallel semantic enhancement attention module, and the other part is used for similarity aggregation output with the final output of the parallel semantic enhancement attention module; Train the medical image segmentation network model with the medical image dataset; Input the real-time collected data into the trained medical image segmentation network model to obtain the medical image segmentation result.

2. A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement according to claim 1, characterized in that The adoption of the dual-encoder structure includes: using PVTv2 as the main encoder to capture global information, and using three-layer res2net50 as the auxiliary encoder to provide supplementary local information.

3. A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement according to claim 2, characterized in that, The adoption of the dual-encoder structure to couple a large-kernel grouped transformation attention module and a multi-layer cross-attention fusion module respectively includes: The input image \(X\in\mathbb{R}\) is processed by a dual-branch encoder H×W×3 , and feature maps from the CNN encoder and the PVTv2 encoder are obtained. In the CNN encoder part, Res2Net50 is used to perform three residual convolution operations on the input image, thereby obtaining features \(T\) at three scales i \((i = 1,2,3)\). Then, multi-scale feature fusion is performed on these three layers of features through a multi-layer cross-attention fusion module to obtain an initial segmentation prediction map \(F_1\). In the PVTv2 encoder part, four different levels of features \(X\) are obtained using the PVTv2 encoder part i \((i = 1,2,3,4)\), and a large kernel grouped transformation attention module is used for feature extraction and edge information enhancement \(X'\) i (i = 1,2,3,4).

4. A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement according to claim 1, characterized in that The large-core grouped transformation attention module is connected to the parallel semantic enhancement attention module, and a part of the output of the multi-layer cross-attention fusion module is used to optimize the parallel semantic enhancement attention module, including: in the decoder part, the parallel semantic enhancement attention module is used to retain and layer-by-layer transmit the feature information of each layer, and the decoder feature D of each layer is obtained i (i = 1, 2, 3, 4), where a 1×1 convolution and sampling operation are performed at the deepest layer to adjust the channels and dimensions of the F1 feature to align with and fuse with D4, and some local information not available in the deep features is supplemented.

5. A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement according to claim 1, characterized in that, The other part is used for similarity aggregation output with the final output of the parallel semantic enhancement attention module includes: using a residual operation to restore the feature map through top-down cross-scale transmission, introducing a spatial attention module, and performing feature aggregation on the decoder-generated prediction feature map D1 and the CNN fusion feature F1 to generate the final prediction feature map P1.

6. A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement according to claim 1, characterized in that The structure of the large-kernel grouped transformation attention module is: Extract multi-scale features and enhance local detail information through two branches respectively, and optimize the feature representation by combining the attention mechanism; The first branch uses three different-sized convolutional kernels to capture multi-scale features: a 3×3 depthwise separable convolution S 11 for efficiently extracting local features; a 5×5 depthwise separable convolution S 12 for capturing broader context information; a deformable convolutional network (DCN) S 13 which detects object edges and refines features by adaptively adjusting the position and shape of the convolutional kernel; The output feature maps of the three branches are fused by pixel-wise addition, denoted as S1. The second branch S2 focuses on enhancing local detail information and adopts a receptive field attention convolution module. The output feature maps of the two branches are fused by element-wise addition, and the ReLU activation function is applied to combine global semantic information and local detail features; Apply the ESA attention module to the fused feature map to obtain O LGDA , by learning the importance weights in the spatial dimension, enhance the model's attention ability to the target region and reduce the influence of background noise.

7. A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement according to claim 1, characterized in that, The multi-layer cross-attention fusion module includes: Receive three-layer feature information from the CNN backbone branch, and then re-aggregate its information to generate fused high-level semantic feature information for supplementing the local information of the main network; For these three layers of information T i (i = 1, 2, 3), a path alignment mechanism is performed, where a differentiated processing strategy is adopted. Through sampling and convolution operations, the features of different input layers are aligned. For the deep feature T3, upsampling is performed through pooling operations and combined with 1×1 depth convolution to restore the spatial resolution and maintain channel consistency; for the shallow feature T1, downsampling is performed through pooling operations and 1×1 convolution to reduce the channel dimension to 256, making all features consistent in both spatial and channel dimensions; Adopt a multiplicative fusion framework to perform cross-level multiplicative fusion between the aligned features. First, perform multiplicative fusion on T2 and T3 through the EFF module to obtain T 21 , then perform a concatenation operation on T 21 and T2 to enhance the intermediate-level features and promote cross-layer information fusion. Finally, perform cross-attention fusion on T 21 and T 12 using a non-linear operation and an attention mechanism to dynamically adjust the weights of the cross-fused features. By introducing weight parameters α and β to balance the fusion effect of high-level semantic information and low-level details, the final output O MLCF is generated.

8. A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement according to claim 1, characterized in that, The parallel semantic enhancement attention module includes: Construct a parallel semantic enhancement attention module composed of two branches to model these two different semantics respectively; In branch 1, in the manner of PVT, the input features are embedded into overlapping patches using a convolutional layer. Subsequently, the obtained serialized tokens are used to calculate the query Q, key K, and value V respectively through three parallel fully connected layers. Calculate using Q and K through the affinity matrix to obtain the relationship between each pixel point in the feature map and the pixels in the same row and the same column. Then, perform a normalization operation through the softmax function to obtain the output D. Aggregate the D, V values, and H to obtain the output C. Finally, use a residual operation and a multiplicative operation to enrich the context information of the long connection to obtain the final output H'; In branch 2, the local fusion information Z output by the CNN encoder is subjected to position matching and embedding, and DCNv2 and adaptive weighted fusion operations are adopted to adapt the local information to the irregular structure, and learnable weight parameters are used to fuse the local features with the decoder part. Finally, for dimension alignment and subsequent processing, the outputs of branch 1 and branch 2 are fused using a product operation.

9. A Transformer-CNN medical image segmentation method based on multi-scale fusion semantic enhancement according to claim 1, characterized in that, The loss of the medical image segmentation network model includes: A multi-loss joint supervision strategy is proposed, and the overall objective function L is constructed by weighted combination of multiple sub-loss functions total : L total = λ1L seg + λ2L dice ; L seg = L miou + L bce ; Among them, L bce is the weighted binary cross-entropy loss, and L miou is the mean intersection over union loss. y i is the ground truth label, o(y i ) is the probability that the model predicts a certain class label. The L miou loss and the L bce loss jointly constrain the segmentation results from two dimensions: the global region matching degree and the local pixel classification accuracy. L dice is the Dice loss function, which is used to emphasize the spatial overlap between the predicted boundary and the ground truth boundary.

10. A Transformer-CNN medical image segmentation system based on multi-scale fusion semantic enhancement, characterized in that, including: Acquisition module: used to acquire medical image data and construct a medical image dataset; Model construction module: used to construct a medical image segmentation network model, adopting a dual-encoder structure to couple a large-kernel grouped transformation attention module and a multi-layer cross-attention fusion module respectively, connecting the large-kernel grouped transformation attention module to a parallel semantic enhancement attention module, using part of the output of the multi-layer cross-attention fusion module to optimize the parallel semantic enhancement attention module, and aggregating the other part with the final output of the parallel semantic enhancement attention module for similarity output; Training module: used to train the medical image segmentation network model through the medical image dataset; Image segmentation module: used to input the real-time acquired data into the trained medical image segmentation network model to obtain the medical image segmentation result.

Citation Information

Cited By

  • Heart image segmentation method fusing multi-receptive-field convolution and distracting attention

    CN120783058A

  • 2D medical image segmentation method and system based on Mama and UNet

    CN120997233A

  • Image region analysis method based on entropy driving feature enhancement

    CN121074346A

  • Lithium battery health state prediction method

    CN121186616A

  • Laparoscopic medical image segmentation method based on mixed attention enhancement and multi-scale feature fusion

    CN121366172A