Medical image segmentation method based on Bottleneck and multi-scale features

The method addresses U-Net's limitations by using Bottleneck and multi-scale features with BMSE and CCA modules to improve medical image segmentation precision and efficiency without increasing model complexity.

CN120318508APending Publication Date: 2025-07-15NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325525.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing U-Net networks are difficult to capture multi-scale information simultaneously in medical image segmentation without increasing the amount of model parameters and complexity, and jump connections may lead to improper feature mixing, affecting segmentation accuracy, especially in the case of blurred target boundaries or complex structures.

Method used

Using a medical image segmentation method based on Bottleneck and multi-scale features, the BMSE module is used to extract features through Bottleneck structure and multi-scale filter, and the semantic gap is alleviated through the CCA module, combining channel attention suppression and unrelated channels to build a lightweight segmentation network.

Benefits of technology

It improves the accuracy and efficiency of medical image segmentation, especially in complex tasks, maintains the lightweight model, and reduces the computational volume and memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318508A_ABST
    Figure CN120318508A_ABST
Patent Text Reader

Abstract

The invention provides a medical image segmentation method based on Bottleneck and multi-scale features, and belongs to the technical field of medical image segmentation. The technical problem that Unet is low in precision when processing targets of different scales and complex image structures is solved. Meanwhile, the technical problems that direct simple splicing or adding of jump connection may cause improper feature mixing, and particularly, segmentation precision may be affected under the condition that a target boundary is fuzzy or a structure is complex are solved. According to the technical scheme, the method comprises the following steps of 1, dividing a data set into a training set, a verification set and a test set; 2, carrying out preprocessing and data enhancement operation on the training set; step 3, constructing a segmentation network model based on Bottleneck and multi-scale features; 4, training the model and evaluating the model; the beneficial effects of the invention are that the method improves the segmentation precision, accuracy and robustness when facing a complex segmentation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present invention belongs to the technical field of medical image segmentation, and particularly relates to a medical image segmentation method using Bottleneck and multi-scale convolutional filters. Background Art

[0002] Medical image segmentation technology is a core technology in medical image processing. Its purpose is to extract the regions of interest in the image, and usually these regions represent different tissues, organs, lesions or abnormal areas. Medical image segmentation is widely used in medical image analysis, such as automatic diagnosis, preoperative planning, lesion area detection, etc. Its purpose is to help doctors identify the lesion area more quickly and accurately, thereby improving the efficiency and accuracy of diagnosis.

[0003] Medical image segmentation technology has experienced a transformation from traditional methods mainly based on edge detection, template matching, etc. to deep learning methods. The application of deep learning, especially convolutional neural networks (CNNs), has greatly improved the accuracy and automation of segmentation, and has achieved breakthroughs especially in processing complex lesion areas, low-contrast images, etc. With the development of deep learning models such as U-Net, Mask R-CNN, GANs, etc., medical image segmentation has entered a new era.

[0004] U-Net has been verified as a mature network architecture in the field of medical image segmentation, but there are still some problems that can be further optimized. The standard U-Net network extracts the features of the image by stacking convolutional layers and pooling layers layer by layer, thereby gradually reducing the spatial resolution and enhancing the abstraction of features. Although this structure can capture high-level global information, since it processes the image sequentially and the receptive field of each layer is usually fixed, for targets or details of different scales, U-Net may not be able to capture all scale information simultaneously. This may lead to inaccurate segmentation of small-scale targets and ineffective modeling of the overall structure of large-scale targets. U-Net aims to combine the low-resolution features in the encoder part with the high-resolution features in the decoder part in order to better restore the details of the image. However, although the skip connection plays an important role in U-Net, in some complex tasks, a simple connection may not be able to fully utilize the useful features extracted in the encoder. Especially when the semantic levels of the features are different, a simple skip connection may introduce noise or irrelevant features.

[0005] Although significant progress has been made in medical image segmentation in the past few years, it still faces challenges such as data scarcity, class imbalance, noise influence, target morphological complexity, and computational resource consumption. In the future, the improvement directions of medical image segmentation will focus on enhancing data utilization, dealing with class imbalance, denoising and image enhancement, improving model architectures, multi-modal data fusion, improving computational efficiency, and improving interpretability. With the continuous development of deep learning technology and the improvement of computational power, medical image segmentation will achieve further breakthroughs in terms of accuracy, efficiency, and interpretability, promoting the optimization of clinical diagnosis and treatment.

[0006] How to solve the above technical problems is the difficult problem faced by the present invention. Summary of the Invention

[0007] The technical problems to be solved by the technology of the present invention are how to increase the ability of the network to capture multi-scale information without excessively increasing the number of parameters and complexity of the model. And how to solve the problem that direct simple splicing or addition of skip connections may lead to improper feature mixing, especially in the case of blurred object boundaries or complex structures, which may affect the segmentation accuracy.

[0008] In order to achieve the above invention purpose, the technical solution adopted by the present invention is specifically as follows: A medical image segmentation method based on Bottleneck and multi-scale features, including the following steps:

[0009] Step 1: Divide the dataset into a training set, a validation set, and a test set;

[0010] Step 2: Perform preprocessing and data augmentation operations on the training set to improve the robustness and generalization ability of the model;

[0011] Step 3: Construct a segmentation network based on Bottleneck and multi-scale features. Select the relatively mature and advanced Unet network in the field of medical image segmentation as the basic framework, and use the BMSE module constructed based on Bottleneck and multi-scale filters and channel attention as the convolutional block in the network to significantly improve the feature expression ability while maintaining the computational complexity. Then, use the CCA module to alleviate the semantic gap, suppress irrelevant channels at the same time, and enhance the features of relevant regions;

[0012] Step 4: Use the training set that has been preprocessed and data-augmented in Step 2 to train the segmentation model constructed in Step 3, and save the weights of the best model. After training is completed, load the weights of the best model obtained in Step 4, and use the test set to test the best model, and evaluate the segmentation effect according to the evaluation metrics;

[0013] The specific process in the above Step 1 is as follows:

[0014] Step 1.1: Divide the dataset into a training set, a validation set, and a test set in the ratio of 8:1:1;

[0015] The specific process in Step 2 is as follows:

[0016] Step 2.1: Load the images and masks in the dataset, and uniformly resize all the images and masks from the original 384×288 to 256×256 for convenient subsequent processing;

[0017] Step 2.2: OpenCV by default reads images in the BGR format, while Pillow and some other image processing libraries usually use the RGB format. To ensure the normal display of colors, convert the images from the BGR color space to the RGB color space and convert the images to grayscale images;

[0018] Step 2.3: Normalize the images to improve the numerical stability during the training process and make it easier for the network to converge. After normalization, the pixel values of the images range from [0,1]; divide each pixel value of the masks by 255 for normalization. For the background in the masks (pixel value is 0), it remains 0 after normalization, and for the target regions in the masks (pixel value is 255), it becomes 1 after normalization; The process formulas are as follows:

[0019]

[0020] where I is the image, (x,y) are the image coordinates, and I(x,y) is the pixel value at the corresponding coordinate position of the image;

[0021]

[0022] where M is the mask image, (x,y) are the image coordinates, and M(x,y) is the pixel value at the corresponding coordinate position of the mask;

[0023] Step 2.4: Standardize the images to transform the distribution of the input data 1 into a standard normal distribution, avoiding the overly large or small gradients that may occur during the weight update process and eliminating the dimensional differences between different features at the same time; The standardization formula is as follows:

[0024]

[0025] where μ is the mean of all pixel values of the image, σ is the standard deviation of all pixel values of the image, and I standardized is the standardized image, and the standardized image has a distribution with a mean of 0 and a standard deviation of 1;

[0026]

[0027] where W and H are the width and height of the image;

[0028]

[0029] Step 2.5: Perform data transformation on the training set data using operations such as random rotation, random horizontal flipping, random vertical flipping, and random cropping to generate more diverse training samples and improve the generalization ability of the model; the rotation formula is as follows:

[0030]

[0031] where θ is the rotation angle, sin is the sine function, cos is the cosine function, (x, y) are the coordinates of the original image, and (x′, y′) are the pixel coordinates after rotation; the horizontal flipping formula is as follows:

[0032] x′ = W - x

[0033] y′ = y

[0034] where W is the image width, and horizontal flipping is to flip the image along the vertical central axis; the vertical flipping formula is as follows:

[0035] x′ = x

[0036] y′ = H - y

[0037] where H is the image height, and vertical flipping is to flip the image along the horizontal central axis; the random cropping formula is as follows:

[0038] x′ = x - x0

[0039] y′ = y - y0

[0040] where the position of the cropping area is (x0, y0), x0 and y0 are the coordinates of the upper left corner of the cropping area, and cropping is to randomly crop a region from the image. Assume the size of the original image is H×W, and the size after cropping is W′×H′;

[0041] The specific process in Step 3 is as follows:

[0042] Step 3.1: BMSE module;

[0043] The BMSE module consists of a Bottleneck bottleneck structure, multi-scale filter branches, channel attention, etc. In the BMSE module, first, the bottleneck structure is used to compress the channels of the image, reducing the computational amount and the number of parameters; then, features are initially extracted through a 3×3 convolutional filter, and then features with different scale receptive fields are extracted through multi-scale convolutional filter branches with sizes of 5×5, 7×7, and 11×11 respectively, which can effectively learn information from different scales; the SE module is used to enable the network to automatically focus on important channels and suppress unimportant channels, thereby effectively improving the performance of the model; finally, the multi-scale feature maps after adjusting the channel weights are fused, and the number of channels is restored to the original dimension using a 1×1 convolution. The process is formulated as follows:

[0044] F 1,1 =Conv 1×1 (F input )

[0045] F 1,2 =BN(F 1,1 )

[0046] F branch1 =Conv 5×5 (F 1,2 )

[0047] F branch2 =Conv 7×7 (F 1,2 )

[0048] F branch3 =Conv 11×11 (F 1,2 )

[0049] F cat =Concate(F branch1 ,F branch2 ,F branch3 )

[0050] F 1,3 =BN(F cat )

[0051] F 1,4 =SE(F 1,3 )

[0052] F fused =Conv 1×1 (F 1,4 )

[0053] F 1,5 =BN(F fused )

[0054] F 1,6= Conv 1×1 (F 1,5 )

[0055] F 1,7 = BN(F 1,6 )

[0056] F residual = F input + F 1,7

[0057] F output = Relu(F residual )

[0058] Among them, F input is the input image, Conv is convolution, BN is batch normalization, Concate is concatenation, SE is the Squeeze-and-Excitation attention mechanism, and Relu is the activation function;

[0059] Step 3.2: CCA module;

[0060] The CCA module is designed based on channel attention. It extracts the attention weight coefficients of each channel of the output feature maps of the encoder and decoder at the same level respectively, and uses the method of integrating the attention weight coefficients of the encoder and decoder features

[0061] to alleviate the semantic gap of features in the same stage. The process is formulated as follows:

[0062] D1 = Upsampling(D)

[0063] C1 = CA(D1)

[0064] C2 = CA(E)

[0065]

[0066] E1 = C3 ⊙ E

[0067]

[0068] E output = Relu(Conv 3×3 (E2))

[0069] Among them, D is the decoder feature, Upsampling is upsampling, E is the encoder feature, is Hadamard addition, ⊙ is Hadamard multiplication, and CA is attention weight extraction; Among them, the CA formula is as follows:

[0070]

[0071] e1 = ReLU(W1·z c + b1)

[0072] e2 = Sigmoid(W2·e1 + b2)

[0073] s c = σ(e2)

[0074] where x i,j,c is the value of the c-th channel of the input feature map at position (i, j), H is the height of the feature map, W is the width of the feature map, z c is the global descriptor of the c-th channel, W1 and W2 are the weight matrices of the first and second fully connected layers respectively, b1 and b2 are the bias terms, and e1 and e2 are the outputs of the first layer and the second layer respectively, that is, the channel weights, is the Sigmoid activation function, s c is the attention coefficient of the c-th channel;

[0075] Step 3.3: Network architecture;

[0076] The medical image segmentation network based on Bottleneck and multi-scale features adopts the same encoder-decoder architecture as Unet, uses the self-constructed BMSE module to replace the original double 3×3 convolutional block, and at the same time uses the self-constructed CCA module to replace the simple splicing and fusion scheme of Unet at the skip connection.

[0077] The specific process in Step 4 is as follows:

[0078] Step 4.1: Initialize the model parameters;

[0079] Step 4.2: Set the total number of training epochs to 200, the Batchsize to 16, the initial learning rate to 0.001, and adopt the Adam optimizer;

[0080] Step 4.3: Use the cross-entropy loss and the Dice loss as the loss functions. The loss weight of the background and foreground in the cross-entropy is set to 1:2, and the weight ratio of the cross-entropy loss to the Dice loss is set to 0.8 to 1.2;

[0081] Step 4.4: Use DSC, mIoU, and Accuracy as the evaluation metrics for the model performance. The functional expressions of each evaluation metric are as follows:

[0082]

[0083] where TP is the true positive class, TN is the true negative class, FP is the false positive class, and FN is the false negative class,

[0084] Step 4.5: Use the model to train on the training set and save the optimal model parameters;

[0085] Step 4.6: After the model training is completed, use the optimal model during training to evaluate the performance of the model on the test set;

[0086] Compared with the existing technologies, the beneficial effects of the technical solution of the present invention are as follows:

[0087] 1. The present invention designs a relatively lightweight medical image segmentation network, which improves the segmentation accuracy of the model without increasing a large number of parameters and computational amounts, keeping the model lightweight.

[0088] 2. The BMSE module is constructed by combining the Bottleneck structure with multi-branch convolutional filtering paths, etc., enabling the network to simultaneously learn features of different scales. Small convolutional kernels focus on feature extraction in local regions, while large convolutional kernels can capture more extensive context information. By using these convolutional kernels in parallel, the network simultaneously learns local and global features, thereby improving the performance of the model, especially being particularly prominent in complex tasks. The introduction of the Bottleneck structure reduces the computational amount and memory usage.

[0089] 3. The present invention method proposes the CCA module, which, in the channel dimension, adds the weight coefficients dynamically extracted from the encoder and decoder features, and uses the neutralized attention weights to weight the encoder channels to reduce the semantic gap between the two. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention.

[0091] Figure 1 It is a schematic diagram of the overall flow of the method embodiment of the present invention.

[0092] Figure 2 It is a schematic diagram of the medical image segmentation network structure based on Bottleneck and multi-scale features in the method embodiment of the present invention.

[0093] Figure 3 It is a schematic diagram of the BMSE convolutional block structure of the segmentation network in the method embodiment of the present invention.

[0094] Figure 4 It is a schematic diagram of the structure of the SE module used in the BMSE convolutional block in the method embodiment of the present invention.

[0095] Figure 5 It is a schematic diagram of the structure of the spatial attention module in the CCA module in the method embodiment of the present invention. Specific Embodiments

[0096] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0097] Embodiment 1

[0098] Figure 1 It is a schematic diagram of the overall process of the method of the present invention. First, divide the CVC-ClinicDB dataset; then perform preprocessing and data augmentation operations on the dataset; subsequently, initialize the network training parameters and train on the training set; finally, load the optimal model parameters and evaluate the segmentation performance of the model on the CVC-ClinicDB validation set.

[0099] The specific implementation steps of the above method are as follows:

[0100] Step 1.1: Divide the dataset into a training set, a validation set, and a test set in a ratio of 8:1:1;

[0101] Step 2.1: Load the images and masks in the dataset, and uniformly resize all the images and masks from the original 384×288 to 256×256 for subsequent processing;

[0102] Step 2.2: OpenCV defaults to using the BGR format to read images, while Pillow and some other image processing libraries usually use the RGB format. To ensure the normal display of colors, convert the images from the BGR color space to the RGB color space and convert the images to grayscale images;

[0103] Step 2.3: Normalize the images to improve the numerical stability during training and make it easier for the network to converge. After normalization, the pixel values of the images range from [0,1]; divide each pixel value of the masks by 255 for normalization. For the background in the masks (pixel value is 0), it remains 0 after normalization, and for the target areas in the masks (pixel value is 255), it becomes 1 after normalization; the process formula is as follows:

[0104]

[0105] Where I is the image, (x,y) are the image coordinates, and I(x,y) is the pixel value at the corresponding coordinate position of the image;

[0106]

[0107] Where M is the mask image, (x,y) are the image coordinates, and M(x,y) is the pixel value at the corresponding coordinate position of the mask;

[0108] Step 2.4: Normalize the image to transform the distribution of the input data 1 into a standard normal distribution, which avoids the excessively large or small gradients that may occur during the weight update process and eliminates the dimensional differences between different features at the same time. The normalization formula is as follows:

[0109]

[0110] where μ is the mean of all pixel values of the image, σ is the standard deviation of all pixel values of the image, and I standardized is the normalized image, and the normalized image has a distribution with a mean of 0 and a standard deviation of 1;

[0111]

[0112] where W and H are the width and height of the image;

[0113]

[0114] Step 2.5: Perform data transformation on the training set data by means of operations such as random rotation, random horizontal flipping, random vertical flipping, and random cropping to generate more diverse training samples to improve the generalization ability of the model. The rotation formula is as follows:

[0115]

[0116] where θ is the rotation angle, sin is the sine function, cos is the cosine function, (x, y) are the coordinates of the original image, and (x′, y′) are the pixel coordinates after rotation. The horizontal flipping formula is as follows:

[0117] x′ = W - x

[0118] y′ = y

[0119] where W is the image width, and horizontal flipping is to flip the image along the vertical central axis. The vertical flipping formula is as follows:

[0120] x′ = x

[0121] y′ = H - y

[0122] where H is the image height, and vertical flipping is to flip the image along the horizontal central axis. The random cropping formula is as follows:

[0123] x′ = x - x0

[0124] y′ = y - y0

[0125] Among them, the position of the cropping region is (x0, y0), where x0 and y0 are the coordinates of the upper left corner of the cropping region. The cropping randomly crops a region from the image. Assume the size of the original image is H×W, and the size after cropping is W′×H′;

[0126] Step 3.1: BMSE module;

[0127] The BMSE module consists of a Bottleneck bottleneck structure, multi-scale filter branches, channel attention, etc. In the BMSE module, first, the bottleneck structure is used to compress the channels of the image to reduce the amount of computation and the number of parameters; then, a 3×3 convolutional filter is used to initially extract features, and then multi-scale convolutional filter branches with sizes of 5×5, 7×7, and 11×11 are used to extract features with different scale receptive fields, which can effectively learn information from different scales; the SE module is used to enable the network to automatically focus on important channels and suppress unimportant channels, thereby effectively improving the performance of the model; finally, the multi-scale feature maps after adjusting the channel weights are fused, and a 1×1 convolution is used to restore the number of channels to the original dimension. The process is formulated as follows:

[0128] F 1,1 =Conv 1×1 (F input )

[0129] F 1,2 =BN(F 1,1 )

[0130] F branch1 =Conv 5×5 (F 1,2 )

[0131] F branch2 =Conv 7×7 (F 1,2 )

[0132] F branch3 =Conv 11×11 (F 1,2 )

[0133] F cat =Concate(F branch1 ,F branch2 ,F branch3 )

[0134] F 1,3 =BN(F cat )

[0135] F 1,4 =SE(F 1,3 )

[0136] Ffused = Conv 1×1 (F 1,4 )

[0137] F 1,5 = BN(F fused )

[0138] F 1,6 = Conv 1×1 (F 1,5 )

[0139] F 1,7 = BN(F 1,6 )

[0140] F residual = F input + F 1,7

[0141] F output = Relu(F residual )

[0142] Among them, F input is the input image, Conv is convolution, BN is batch normalization, Concate is concatenation, SE is the Squeeze-and-Excitation attention mechanism, and Relu is the activation function;

[0143] Step 3.2: CCA module;

[0144] The CCA module is designed based on channel attention, and the attention weight coefficients of each channel of the output feature maps of the same-level encoder and decoder are extracted respectively, using the method of integrating the attention weight coefficients of the encoder and decoder features

[0145] to alleviate the semantic gap of features in the same stage. The process is formulated as follows:

[0146] D1 = Upsampling(D)

[0147] C1 = CA(D1)

[0148] C2 = CA(E)

[0149]

[0150] E1 = C3 ⊙ E

[0151]

[0152] E output = Relu(Conv 3×3 (E2))

[0153] Among them, D is the decoder feature, Upsampling is upsampling, E is the encoder feature, is Hadamard addition, ⊙ is Hadamard multiplication, and CA is the attention weight extraction; among them, the CA formula is as follows:

[0154]

[0155] e1 = ReLU(W1·z c +b1)

[0156] e2 = Sigmoid(W2·e1 + b2)

[0157] s c = σ(e2)

[0158] Among them, x i,j,c is the value of the c-th channel at the position (i, j) of the input feature map, H is the height of the feature map, W is the width of the feature map, z c is the global descriptor of the c-th channel, W1 and W2 are the weight matrices of the first and second fully connected layers respectively, b1 and b2 are the bias terms, e1 and e2 are the outputs of the first and second layers respectively, that is, the channel weights, is the Sigmoid activation function, s c is the attention coefficient of the c-th channel;

[0159] Step 3.3: Network architecture;

[0160] The medical image segmentation network based on Bottleneck and multi-scale features adopts the same encoder-decoder architecture as Unet, uses the self-constructed BMSE module to replace the original double 3×3 convolution block, and at the same time uses the self-constructed CCA module to replace the simple splicing fusion scheme of Unet at the skip connection.

[0161] Step 4.1: Initialize the model parameters;

[0162] Step 4.2: Set the total number of training epochs to 200, the Batchsize to 16, the initial learning rate to 0.001, and adopt the Adam optimizer;

[0163] Step 4.3: Use the cross-entropy loss and Dice loss for the loss function. The loss weight of the background and foreground in the cross-entropy is set to 1:2, and the weight ratio of the cross-entropy loss to the Dice loss is set to 0.8 to 1.2;

[0164] Step 4.4: Use DSC, mIoU, and Accuracy as the evaluation metrics for the model performance. The function expressions of each evaluation metric are as follows:

[0165]

[0166]

[0167] Among them, TP is the true positive class, TN is the true negative class, FP is the false positive class, and FN is the false negative class.

[0168] Step 4.5: Use this model to train on the training set and save the optimal model parameters.

[0169] Step 4.6: After the model training is completed, use the optimal model during training to evaluate the performance of the model on the test set.

[0170] Compare the method of the present invention with the current mainstream polyp lesion segmentation models (U-Net, DCSAU-Net, ConvUNeXt, AMSUnet) on the CVC-ClinicDB dataset, and the results are shown in Table 1.

[0171] In Example 1, the evaluation metrics are the Dice similarity coefficient, mean intersection over union (mIoU), accuracy (Accuracy), recall (Recall), and precision (Precision).

[0172] Table 1 Performance comparison of segmentation models on the CVC-ClinicDB dataset

[0173]

[0174] Specifically, look at the comparison experiment data in Table 1.

[0175] The Dice coefficient of the method of the present invention is 0.9073, which is about 3.74% higher than that of U-Net (0.8746), about 3.07% higher than that of UNet3+ (0.8803), about 4.03% higher than that of DCSAU-Net (0.8722), about 1.62% higher than that of ConvUNeXt (0.8928), and about 1.30% higher than that of PDFUnet (0.8956). This indicates that the overlap degree between the target area and the true annotation of this method has been significantly improved. The Accuracy of the method of the present invention is 0.9895, which is 0.74% higher than that of U-Net, 0.33% higher than that of DCSAU-Net, 0.44% higher than that of ConvUNeXt, and 0.35% higher than that of AMSUnet.

[0176] In terms of the mIoU metric, the method of the present invention achieved 0.8491, representing an improvement of approximately 4.10% compared to U-Net (0.8157), 3.76% compared to UNet3+ (0.8183), 4.40% compared to DCSAU-Net (0.8133), 2.10% compared to ConvUNeXt (0.8316), and 1.12% compared to PDFUnet (0.8397). This further demonstrates the advantage of the present method in accurately capturing the boundaries of the segmented regions.

[0177] In terms of Accuracy, the method of the present invention reached 0.9892, representing an improvement of approximately 0.24% compared to U-Net (0.9868), 0.21% compared to UNet3+ (0.9871), and 0.19% compared to DCSAU-Net (0.9873). It was basically on par with ConvUNeXt (0.9891) and PDFUnet (0.9891), but still maintained extremely high pixel-level segmentation accuracy overall.

[0178] Overall, through effective improvement strategies, the method of the present invention has achieved significant improvements in key metrics such as Dice and mIoU, and the segmentation results are more accurate in terms of the overlap degree and detail capture in the polyp lesion area. Although the improvement in Accuracy is relatively small, the overall improvement in various metrics fully verifies the superiority and practical application value of the method of the present invention in the polyp image segmentation task.

[0179] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A medical image segmentation method based on Bottleneck and multi-scale features, characterized in that, It includes the following steps: Step 1: Divide the dataset into a training set, a validation set, and a test set; Step 2: Perform preprocessing and data augmentation operations on the training set to improve the robustness and generalization ability of the model; Step 3: Construct a segmentation network based on Bottleneck and multi-scale features. Select the relatively mature and advanced Unet network in the field of medical image segmentation as the basic framework. Use the BMSE module constructed based on Bottleneck, multi-scale filters, and channel attention as the convolutional block in the network. Then, use the CCA module to alleviate the semantic gap, suppress irrelevant channels, and enhance the features of relevant regions; Step 4: Use the training set that has been preprocessed and data-augmented in Step 2 to train the segmentation model constructed in Step 3, and save the weights of the best model. After training is completed, load the weights of the best model obtained in Step 4, and use the test set to test on the best model. Evaluate the segmentation effect according to the evaluation metrics.

2. The medical image segmentation method based on Bottleneck and multi-scale features according to claim 1, wherein The said Step 1 includes the following steps: Step 1.1: Divide the dataset into a training set, a validation set, and a test set, with a ratio of 8:1:

1.

3. A medical image segmentation method based on Bottleneck and multi-scale features according to claim 1, characterized in that The said Step 2 includes the following steps: Step 2.1: Load the images and masks in the dataset, and uniformly resize all the images and masks from the original 384×288 to 256×256; Step 2.2: OpenCV defaults to reading images in the BGR format, while Pillow and some other image processing libraries use the RGB format. To ensure the normal display of colors, convert the images from the BGR color space to the RGB color space and convert the images to grayscale images; Step 2.3: Normalize the images to improve the numerical stability during the training process and make the network converge. After normalization, the pixel value range of the images is [0,1]; Divide each pixel value of the masks by 255 through normalization. For the background in the masks, the pixel value is 0, and it remains 0 after normalization. For the target regions in the masks, the pixel value is 255, and it becomes 1 after normalization; The process formulas are as follows: where I is the image, (x,y) are the image coordinates, and I(x,y) is the pixel value at the corresponding coordinate position of the image; where M is the mask image, (x,y) are the image coordinates, and M(x,y) is the pixel value at the corresponding coordinate position of the mask; Step 2.4: Standardize the images to transform the distribution of the input data I into a standard normal distribution, avoiding overly large or overly small gradients during the weight update process and eliminating the dimensional differences between different features at the same time; The standardization formula is as follows: Among them, μ is the mean of all pixel values of the image, σ is the standard deviation of all pixel values of the image, and I standardized is the standardized image, and the standardized image has a distribution with a mean of 0 and a standard deviation of 1; where W and H are the width and height of the image; Step 2.5: Perform data transformation on the training set data using random rotation, random horizontal flipping, random vertical flipping, and random cropping operations to generate more diverse training samples; The rotation formula is as follows: where θ is the rotation angle, sin is the sine function, cos is the cosine function, (x, y) are the coordinates of the original image, and (x ′ , y ′ ) are the pixel coordinates after rotation; the formula for horizontal flipping is as follows: x ′ = W - x y ′ = y where W is the image width, and horizontal flipping flips the image along the vertical central axis; The vertical flipping formula is as follows: x ′ = x y ′ = H - y where H is the image height, and vertical flipping flips the image along the horizontal central axis; The random cropping formula is as follows: x ′ = x - x0 y ′ = y - y0 Among them, the position of the cropping area is (x0, y0), where x0 and y0 are the coordinates of the upper left corner of the cropping area. Cropping is to randomly crop a region from the image. Assume that the size of the original image is H×W, and the size after cropping is W ′ ×H′.

4. A medical image segmentation method based on Bottleneck and multi-scale features according to claim 1, characterized in that The said Step 3 includes the following steps: Step 3.1: BMSE module; The BMSE module consists of a Bottleneck bottleneck structure, multi-scale filter branches, and channel attention. In the BMSE module, first, the bottleneck structure is used to compress the channels of the image, reducing the computational amount and the number of parameters. Then, features are initially extracted through a 3×3 convolutional filter, and then features with different scale receptive fields are extracted through multi-scale convolutional filter branches with sizes of 5×5, 7×7, and 11×11 respectively. The SE module is used to enable the network to automatically focus on important channels and suppress unimportant channels. Finally, the multi-scale feature maps after adjusting the channel weights are fused, and the number of channels is restored to the original dimension using a 1×1 convolution. The process is formulated as follows: F 1,1 = Conv 1×1 (F input ) F 1,2 = BN(F 1,1 ) F branch1 = Conv 5×5 (F 1,2 ) F branch2 = Conv 7×7 (F 1,2 ) F branch3 = Conv 11×11 (F 1,2 ) F cat = Concatenate(F branch1 , F branch2 , F branch3 ) F 1,3 = BN(F cat ) F 1,4 = SE(F 1,3 ) F fused = Conv 1×1 (F 1,4 ) F 1,5 = BN(F fused ) F 1,6 = Conv 1×1 (F 1,5 ) F 1,7 = BN(F 1,6 ) F residual = F input + F 1,7 F output = Relu(F residual ) Among them, F input is the input image, Conv is convolution, BN is batch normalization, Concate is concatenation, SE is the Squeeze-and-Excitation attention mechanism, and Relu is the activation function; Step 3.2: CCA module; The CCA module is designed based on channel attention. The attention weight coefficients of each channel of the output feature maps of the same-level encoder and decoder are extracted respectively, and the semantic gap of features in the same stage is alleviated by integrating the attention weight coefficients of the encoder and decoder features. The process is formulated as follows: D1 = Upsampling(D) C1 = CA(D1) C2 = CA(E) E1 = C3 ⊙ E E output = Relu(Conv 3×3 (E2)) Among them, D is the decoder feature, Upsampling is upsampling, and E is the encoder feature. is the Hadamard addition, ⊙ is the Hadamard multiplication, and CA is the attention weight extraction; among them, the CA formula is as follows: e1 = ReLU(W1·z c + b1) e2 = Sigmoid(W2·e1 + b2) s c = σ(e2) where x i,j,c is the value of the c-th channel at the position (i, j) of the input feature map, H is the height of the feature map, W is the width of the feature map, and z c is the global descriptor of the c-th channel, W1 and W2 are the weight matrices of the first and second fully connected layers respectively, b1 and b2 are the bias terms, and e1 and e2 are the outputs of the first and second layers respectively, that is, the channel weights, is the Sigmoid activation function, and s c is the attention coefficient of the c-th channel; Step 3.3: Network architecture; The medical image segmentation network based on Bottleneck and multi-scale features adopts the same encoder-decoder architecture as Unet, uses the self-constructed BMSE module to replace the original double 3×3 convolutional block, and at the same time uses the self-constructed CCA module to replace the simple splicing and fusion scheme of Unet at the skip connection.

5. A medical image segmentation method based on Bottleneck and multi-scale features according to claim 1, characterized in that, Step 4 includes the following steps: Step 4.1: Initialize the model parameters; Step 4.2: Set the total number of training epochs to 200, the Batchsize to 16, the initial learning rate to 0.001, and use the Adam optimizer; Step 4.3: The loss function uses cross-entropy loss and Dice loss. The loss weight ratio of background to foreground in cross-entropy is set to 1:2, and the weight ratio of cross-entropy loss to Dice loss is set to 0.8:1.2; Step 4.4: Use DSC, mIoU, and Accuracy as the evaluation metrics for the model performance. The functional expressions of each evaluation metric are as follows: Among them, TP is the true positive class, TN is the true negative class, FP is the false positive class, and FN is the false negative class; Step 4.5: Use this model to train on the training set and save the optimal model parameters; Step 4.6: After the model training is completed, use the optimal model during training to evaluate the performance of the model on the test set.