An Image Semantic Segmentation Method Based on Snake Convolution

Through the serpentine convolutional encoder and hybrid pooling module of the DSM-Net model, the problems of low segmentation accuracy and loss of subtle information in the existing technology are solved, and efficient and accurate image segmentation effect is achieved, which is suitable for segmentation tasks of complex scenes and small target objects.

CN119107453BActive Publication Date: 2025-07-22CHONGQING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411068041.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2025-07-22
Estimated Expiration
2044-08-06

AI Technical Summary

Technical Problem

When existing image segmentation algorithms deal with objects with large differences in different morphological distributions and sizes, the segmentation accuracy is low, especially the segmentation effect of small target objects is poor, and the pooling operation leads to the loss of subtle information, affecting the segmentation accuracy.

Method used

The DSM-Net model based on serpentine convolution is adopted, combined with the serpentine convolution encoder, a hybrid attention module and a fusion pooling module, and adaptive learning of different object morphology is achieved by adjusting the convolution kernel position and feature extraction, and fine feature information is retained through hybrid pooling, and a symmetric encoding-decoding architecture is designed for end-to-end pixel-level prediction.

Benefits of technology

It improves the accuracy and scope of application of image segmentation, especially in small-objective segmentation tasks, can effectively handle complex scenes and subtle features, reduce training costs, and expand model applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119107453B_ABST
    Figure CN119107453B_ABST
Patent Text Reader

Abstract

The present invention discloses an image semantic segmentation method based on serpentine convolution, which relates to the technical fields of image processing and computer vision in artificial intelligence. Through the design of the method, a new fast, efficient, and accurate image segmentation model DSM-Net is obtained; DSM-Net uses a symmetric encoding-decoding architecture to effectively perform feature extraction and upsampling, and realizes end-to-end pixel-level prediction; among them, in the encoding stage, serpentine convolution (EncoderDS) is combined, which can linearly adjust the position of the convolution kernel to fully learn object features; fusion pooling is adopted to efficiently transfer the effective features extracted by the encoder by combining convolution and pooling; a hybrid attention module is used to model space and channels, prompting the network to focus on different local structural features and realizing efficient feature fitting ability; the reasonable construction of the overall architecture greatly reduces the training cost, and the reasonable network depth and internal module design solve the problem of low segmentation accuracy caused by the large morphological distribution and size gap of the segmentation target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and computer vision in artificial intelligence, and specifically to an image semantic segmentation method based on serpentine convolution. Background Art

[0002] Image processing is a collection of techniques and methods for operating on images to enhance their visibility or extract useful information. It has a wide range of applications, including medical imaging, remote sensing, computer vision, photography, entertainment, etc. Its process usually includes image acquisition, preprocessing, feature extraction, representation and description, recognition and classification, and postprocessing. Subsequently, a specific task is completed through the designed model structure.

[0003] Semantic segmentation is a technique in image processing and computer vision, aiming to divide an image into several regions and assign a semantic label to each pixel to represent the object category to which the pixel belongs. It has a wide range of applications, including autonomous driving, medical image analysis, remote sensing image processing, and intelligent monitoring, etc. Its process usually includes steps such as image preprocessing, feature extraction, classification and annotation. Using a deep learning training model can extract multi-scale features while maintaining high resolution, achieving fine pixel-level classification. Semantic segmentation technology plays an important role in applications such as road scene understanding, tumor detection, and ground object classification.

[0004] In recent years, the rapid development of deep learning has demonstrated powerful performance in the field of image segmentation. Deep learning models, with automatic feature learning, convenient operation, and high efficiency, can achieve high-precision segmentation while taking into account time and space efficiency. The classic fully convolutional network (FCN) was proposed by Long et al. in 2015. Its main feature is to accept input images of any size and generate output images of the same size, assigning semantic labels to each pixel, thus realizing end-to-end pixel-level semantic segmentation. However, due to multiple pooling operations, FCN loses attention to image details. Subsequently, to solve the problem of accuracy loss, U-Net was proposed.

[0005] Due to its symmetric encoder-decoder structure, which provides a connection path for the original features to each layer of the decoder, the overall network can fully retain the original detailed information, thus obtaining a more accurate segmentation result. However, the U-Net network still has deficiencies. Due to its own structural limitations, it has a poor effect in special scenarios or for special tasks. Therefore, by utilizing the global feature fitting ability of the Transformer structure and combining it with U-Net, TransUNet was designed. By combining the ability of convolutional neural networks to efficiently extract local features with the global attention perception ability of the Transformer, a more accurate segmentation effect is achieved. For different scenarios, using different convolutional kernels to adapt to the distribution patterns of different objects can also effectively improve the segmentation effect. For example, CE-Net combines dilated convolution as a context feature extraction module, aiming to fully focus on the context feature information between objects, enabling the model to learn deeper features. While TP-Net uses dilated convolution as the processing module for the final image output. By using different-sized dilated convolution kernels to integrate information of multi-scale features, the robustness of segmentation is enhanced.

[0006] In recent years, deep learning has made remarkable progress in the field of image segmentation. Deep learning models can automatically learn features, are easy to operate, and have high performance. Under the premise of considering both time and space efficiency, they can obtain high-precision segmentation results. The classic U-Net network was designed by Ronneberger et al., adopting a symmetric encoder-decoder structure and using skip connections for primitive feature transfer, thereby improving the network's attention to subtle features. These characteristics enable it to well complete the task of retinal vessel segmentation. However, due to the extensive distribution of thin vessels and small target vessels in fundus images, it is still difficult to accurately capture vessel features. In addition, the information loss in the pooling operation exacerbates the difficulty of thin vessel segmentation. To solve these problems, Gu et al. designed a multi-kernel pooling fusion multi-attribute convolution module to fully integrate context information for vessel feature extraction, thus solving the problem of information loss caused by pooling. Yin et al. adopted multi-source image input to ensure the transfer of retinal vessel features, using feature fusion at different scales to provide the original feature information of vessels for each layer of the network, supplementing the information lost in the pooling operation to some extent. In addition, they used the Hessian matrix to obtain weak information in vessels, which, although generating some noise, also enhanced the attention to vessels. Li et al. used dual CNN and RNN encoders for feature extraction and fused thin vessel features, using multiple context information fusion modules to fully extract vessel features, alleviating the problem of information loss. In summary, despite many improvements, image segmentation, especially in the segmentation of subtle features and small targets, still faces many challenges and technical bottlenecks.

[0007] Deficiencies of the prior art:

[0008] (1) Most segmentation algorithms are limited to the traditional convolution mode. Due to the spatial invariance constraint, they have limitations in segmenting different objects in different scenarios. Especially for objects with different morphological distributions and large size differences, the segmentation accuracy is relatively low.

[0009] (2) Since there are quite a number of small target objects in the actual scenario, most networks cannot effectively perceive them. It is easy to ignore small targets, while small target objects are often widely distributed in the accurate segmentation effect. For example, vehicles, trees and other objects photographed by remote sensing satellite photos, and the segmentation effect of objects such as leaves on trees photographed at low altitude is poor and difficult to process.

[0010] (3) In order to achieve a high-precision segmentation effect, feature transfer is a very important step. The existing pooling methods are prone to losing fine information, making the deeper network structure unable to perceive small target features, and thus unable to perform segmentation. The existing models are prone to ignoring this problem of fine features, resulting in low accuracy in the segmentation task of small targets.

[0011] Therefore, a new solution needs to be proposed for the above problems. Summary of the Invention

[0012] The purpose of the present invention is to provide an image semantic segmentation method based on serpentine convolution to solve the technical problems raised in the background art.

[0013] To achieve the above purpose, the present invention provides the following technical solutions: An image semantic segmentation method based on serpentine convolution, at least including the following steps:

[0014] S1: Batch preprocess the collected color fundus data;

[0015] S2: Divide the data into a non-crossing data set;

[0016] S3: Design a loss function for accurate segmentation;

[0017] S4: Design a segmentation network based on serpentine convolution;

[0018] S5: Train the segmentation model and conduct segmentation verification. After the segmentation model is trained, complete the semantic segmentation of the image through the trained segmentation model.

[0019] Further, the S1 at least includes the following steps: Process the fundus images in JPG format, and normalize the data of different channels to input into the model.

[0020] Further, the data set division in S2 at least includes the following steps:

[0021] Randomly shuffle and divide according to fundus data;

[0022] Divide into training, validation, and test sets in a ratio of 8:1:1;

[0023] Among them, the data of any one fundus cannot cross between these sets.

[0024] Furthermore, the loss function of S3 includes DiceLoss, BCELoss, and mixed loss;

[0025] The DiceLoss is as follows:

[0026]

[0027] The BCELoss is as follows:

[0028]

[0029] The mixed loss is as follows:

[0030] L seg = α·L dice + β·L bce

[0031] Where α and β are the weight hyperparameters of the two losses respectively, and their sum is 1;

[0032] In the formula, N represents the total number of pixel points, and C represents the total number of categories; represents the category value corresponding to the i-th pixel in the c-th category of the binary GroundTruth; represents the predicted probability value of the corresponding category;

[0033] ε = 1e-5, which represents the smoothing exponent, used to prevent the denominator prediction from being 0 and thus avoid some extreme situations. In addition, it can also play a role in smoothing the loss and gradient.

[0034] Furthermore, designing a segmentation network based on serpentine convolution in S4 includes at least the following steps:

[0035] Train through the mixed loss function formulated in S3;

[0036] Use the gradient backpropagation during the training of the neural network to update the model weights. Judge whether the model is better according to the segmentation effect of the model on the validation set after training, and update the saved model weights;

[0037] Finally, after the model training is completed, evaluate the segmentation effect on the test set.

[0038] Furthermore, five common evaluation metrics are used for the segmentation verification in S5. The evaluation metrics include Acc (accuracy), Se (sensitivity), Sp (specificity), F1 (F1 score), and AUC (area under the curve).

[0039] The decision rule for the segmentation verification in S5 is that the threshold of the probability map generated by the network is 0.5. Pixel points greater than 0.5 are predicted as blood vessels, and vice versa are predicted as the background.

[0040] Furthermore, the segmentation model is the DSM-Net model, and the segmentation model includes a serpentine convolutional encoder, a hybrid attention module, and a fusion pooling module.

[0041] Furthermore, the serpentine convolutional encoder includes conventional convolution and serpentine convolution. The conventional convolution is used to capture conventional targets, and the serpentine convolution is used to further extract the morphological features of different objects, thereby extracting deeper semantic features.

[0042] The serpentine convolutional encoder tracks different structures by adjusting the shape of the convolutional kernel.

[0043] The operations of the serpentine convolutional encoder include two 3×3 convolution operations and serpentine convolution operations in two directions.

[0044] The applications of the serpentine convolutional encoder at least include the following steps:

[0045] First, ton passes through a convolutional layer with a 3×3 kernel to generate a smoothed feature map.

[0046] Meanwhile, the original feature map is linearly offset and subjected to feature extraction convolution operations in the x and y directions through serpentine convolution. The obtained smoothed feature map is fused with the serpentine convolution response maps in two directions.

[0047] The response maps include the original features of the object and the topological information extracted by the serpentine convolutional layer.

[0048] Finally, the fused features pass through another convolutional layer with a 3×3 kernel for further fusion and channel number adjustment.

[0049] The final output includes two basic components: the sliding convolution information extracted by two smoothed convolutional kernels and the fused adaptive transformation serpentine convolution, thereby obtaining the blood vessel structure features. Refer to the following formula:

[0050] F out =Conv(DSC_x(F in )⊙DSC_y(F in )⊙Conv(F in ))

[0051] In the formula, Fin is the input feature map, Fout is the output feature map, ⊙ represents the connection, Conv is the batch normalization and ReLU activation after the 3×3 convolution operation, and DSC x and DSC y respectively represent the dynamic snake-shaped convolution on the x-axis and y-axis.

[0052] Furthermore, the application of the hybrid attention module at least includes the following steps:

[0053] Create a hybrid block through the sliding window mechanism to establish a local mapping of the input feature map.

[0054] Subsequently, apply the self-attention operation to establish local-global co-dependency relationships, promoting the association of information in the hybrid group within the window. This not only preserves the spatial structure at different scales but also fuses the hybrid channel information. Refer to the following formula:

[0055] F out = Hardswish(BN(SeparConv(F in )))

[0056] In the formula, Fin is the input feature, Fout is the output feature, SeparConv is the depthwise separable convolution with different kernel sizes, the number of channels remains unchanged during the convolution process, BN is the batch normalization, and Hardswish is the Hardswish activation function. Local mapping is performed through different sliding window methods.

[0057] Furthermore, the fusion pooling module at least includes a 2×2 convolution and pooling module;

[0058] The application of the fusion pooling module at least includes the following steps:

[0059] First, pass the input feature map through a 2×2 convolution to keep the number of channels unchanged to preserve the original fine feature information;

[0060] Subsequently, use the result of the convolution as the input for max pooling and average pooling. The pooling operation effectively reduces the feature dimension while retaining the valid information;

[0061] For the feature information caused by pooling, the downsampling operation is divided into two branches. One branch uses a fusion pool to combine max pooling and average pooling to ensure effective feature extraction, and a convolution operation is introduced in the other branch;

[0062] The convolution operation uses a kernel size of 2 and a stride of 2, thereby increasing the learnability of the pooling operation. Moreover, this convolution operation can be learned through parameter adjustment, significantly retaining the fine feature information and generating a feature map of the same size as the pooling operation;

[0063] Finally, equal-weight fusion is performed on the two feature maps, which greatly preserves the feature information of thin blood vessels.

[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0065] 1. A new fast, efficient, and accurate image segmentation model DSM-Net is obtained through the design of the method of the present invention; DSM-Net uses a symmetric encoding-decoding architecture to effectively perform feature extraction and upsampling, realizing end-to-end pixel-level prediction; among them, in the encoding stage, serpentine convolution (EncoderDS) is combined, which can linearly adjust the position of the convolution kernel to fully learn the object features; fusion pooling is used to efficiently transfer the effective features extracted by the encoder by combining convolution and pooling; a hybrid attention module is used to model the space and channels, prompting the network to focus on different local structural features and realizing efficient feature fitting ability; the reasonable construction of the overall architecture greatly reduces the training cost, and the reasonable network depth and internal module design solve the problem of low segmentation accuracy caused by the large morphological distribution and size gap of the segmentation target; in order to reflect the excellent performance of the proposed segmentation model in different segmentation tasks, especially in the tubular segmentation task, multiple datasets are used for training, expanding the scope of application of the model.

[0066] 2. Compared with other segmentation algorithms, DSM-Net of the present invention combines the serpentine convolution mode, which can adaptively adjust the relative position of the convolution kernel according to the scene and the object, so as to learn features more deeply, master discrete morphological data, and achieve a more accurate segmentation effect;

[0067] 3. The overall model architecture of DSM-Net of the present invention is implemented by a symmetric encoder-decoder, and it has strong feature fitting ability and feature scale reduction ability. Training on this basis can effectively perform gradient backpropagation to reduce the loss value. Combining the hybrid attention mechanism and the fusion pooling module can effectively utilize the encoder features for spatial scale modeling, focus on the spatial representation of the object, and can be effectively applied to various scenarios BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0069] Figure 1 It is the overall image segmentation algorithm flowchart provided by the present invention;

[0070] Figure 2 It is the schematic diagram of the DSM-Net network model provided by the present invention;

[0071] Figure 3 Schematic diagram of the EncoderDS module provided by the present invention;

[0072] Figure 4 Schematic diagram of the hybrid attention framework provided by the present invention;

[0073] Figure 5 Schematic diagram of the local feature mapping framework provided by the present invention;

[0074] Figure 6 Schematic diagram of the global feature mapping framework provided by the present invention;

[0075] Figure 7 Schematic diagram of the fusion pooling module provided by the present invention;

[0076] Figure 8 Visual comparison chart of segmentation masks of different models on dataset A provided by the present invention;

[0077] Figure 9 Visual comparison chart of segmentation masks of different models on dataset B provided by the present invention;

[0078] Figure 10 Model parameter quantity comparison chart of different models on dataset C provided by the present invention. Detailed implementation manners

[0079] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0080] Refer to Figure 1 , an image semantic segmentation method based on serpentine convolution, at least including the following steps:

[0081] S1: Batch preprocess the collected color fundus data;

[0082] S2: Divide the data into a non-crossing dataset;

[0083] S3: Design a loss function for accurate segmentation;

[0084] S4: Design a segmentation network based on serpentine convolution;

[0085] S5: Train the segmentation model and perform segmentation verification. After the segmentation model is trained, complete the segmentation of the image semantics through the trained segmentation model.

[0086] S1 at least includes the following steps: Process the fundus images in JPG format and normalize the data of different channels to input into the model.

[0087] The dataset division in S2 includes at least the following steps:

[0088] Randomly shuffle and divide according to fundus data;

[0089] Divide it into training, validation, and test sets in a ratio of 8:1:1;

[0090] Among them, the data of any one fundus cannot cross between these sets.

[0091] The loss functions in S3 include DiceLoss, BCELoss, and hybrid loss;

[0092] DiceLoss is as follows:

[0093]

[0094] BCELoss is as follows:

[0095]

[0096] The hybrid loss is as follows:

[0097] L seg = α·L dice + β·L bce

[0098] Among them, α and β are the weight hyperparameters of the two losses respectively, and their sum is 1. In the experiment, they are set to [0.5, 0.5] to make the model training more balanced and stable;

[0099] In the formula, N represents the total number of pixel points, and C represents the total number of categories; represents the category value corresponding to the i-th pixel in the c-th category in the binary GroundTruth; represents the predicted probability value of the corresponding category;

[0100] ε = 1e-5, which represents the smoothing exponent, used to prevent the denominator prediction from being 0 and thus avoiding some extreme situations. In addition, it can also smooth the loss and gradient.

[0101] The design of the segmentation network based on serpentine convolution in S4 includes at least the following steps:

[0102] Train through the hybrid loss function formulated in S3;

[0103] Use the gradient backpropagation during the training of the neural network to update the model weights, and judge whether the model is better according to the segmentation effect of the model on the validation set after training, and update the saved model weights;

[0104] Finally, after the model training is completed, the segmentation effect is evaluated on the test set.

[0105] In S5, five common evaluation metrics are used for segmentation validation. The evaluation metrics include Acc (accuracy), Se (sensitivity), Sp (specificity), F1 (F1 score), and AUC (area under the curve).

[0106] The decision rule for segmentation validation in S5 is that the threshold of the probability map generated by the network is 0.5. Pixels greater than 0.5 are predicted as blood vessels, and vice versa are predicted as the background.

[0107] For Acc, Se, Sp, and F1, the definitions are as follows:

[0108]

[0109]

[0110]

[0111]

[0112] The decision rule for segmentation validation in S5 includes using the ground truth as a benchmark. Correctly classified retinal blood vessels are classified as the TP class, those that are actually retinal blood vessels but are classified as the background by the network are classified as the FN class, correctly classified background is classified as the TN class, and misclassified background classified as retinal blood vessels is classified as the FP class.

[0113] The test results of all models trained on the dataset are composed of the distribution of the mean and standard deviation, and all metrics exceed the 95% confidence level.

[0114] Refer to Figure 2 , the segmentation model is the DSM-Net model, and the segmentation model includes a serpentine convolutional encoder, a hybrid attention module, and a fusion pooling module.

[0115] Participate in Figure 3 , the serpentine convolutional encoder includes traditional convolution and serpentine convolution. Traditional convolution is used to capture conventional targets, and serpentine convolution is used to further extract the morphological features of different objects, thereby extracting deeper semantic features.

[0116] The serpentine convolutional encoder tracks different structures by adjusting the shape of the convolutional kernel.

[0117] Track different structures by adjusting the shape of the convolutional kernel. However, in the small target segmentation task, due to the complex topological structure, cluttered background information, and small morphological size, the current methods are difficult to adapt to the complex tubular segmentation scenario. Therefore, segmenting small targets and ensuring their segmentation continuity remains a major research challenge.

[0118] The position transformation of the snake-shaped kernel is achieved through iterative offset transformations in the horizontal and vertical directions, and then continuous linear offsets are performed through bilinear interpolation. The offset of the convolutional kernel is constrained by the position of the previous convolutional kernel, preventing the convolutional kernel from being attracted by background noise, thus ensuring more effective extraction of object features.

[0119] Due to the advantages of linear offset in dynamic snake convolution, it adaptively adapts to the narrow and elongated distribution patterns of the target, effectively capturing the morphological features of curved and tiny objects. This can extract richer vascular information, enabling the network to detect the smallest blood vessels.

[0120]

[0121] Among them, the linear offset of the convolutional kernel position is achieved by accumulating the iterative offsets in two directions.

[0122] The operations of the snake-shaped convolution encoder include two 3×3 convolution operations and snake-shaped convolution operations in two directions;

[0123] The application of the snake-shaped convolution encoder includes at least the following steps:

[0124] First, pass through a convolutional layer with a 3×3 kernel to generate a smooth feature map;

[0125] At the same time, linearly offset and perform feature extraction convolution operations on the original feature map in the x and y directions through snake convolution, and fuse the obtained smooth feature map with the snake convolution response maps in two directions;

[0126] The response maps include the original features of the object and the topological information extracted by the snake-shaped convolutional layer;

[0127] Finally, further fuse the fused features through another convolutional layer with a 3×3 kernel and adjust the number of channels;

[0128] The final output includes two basic components: the sliding convolution information extracted by two smooth convolutional kernels and the fused adaptive transformation snake convolution, and then the vascular structure features are obtained. See the following formula:

[0129] F out =Conv(DSC_x(F in )⊙DSC_y(F in )⊙Conv(F in ))

[0130] In the formula, Fin is the input feature map, Fout is the output feature map, ⊙ is concatenation, Conv is batch normalization and ReLU activation after a 3×3 convolution operation, and DSC x and DSC y respectively represent dynamic snake convolutions on the x-axis and y-axis.

[0131] CSA is composed of different local feature maps and global feature maps, including depthwise separable convolutions with different kernel sizes. Although the attention mechanism has been proven effective in segmentation tasks, most models focus on high-dimensional semantic feature maps. However, ignoring the attention operation in low-scale feature maps means missing key spatial scale information. There are a large number of tiny objects in different segmentation task scenarios, and each object exhibits rich distribution patterns, curvatures, and other spatial features. For example, Mou et al. designed CS-Net to perform spatial and channel attention calculations at the highest feature layer. However, after multiple pooling operations, a large amount of spatial scale information of blood vessels is lost in the high-dimensional feature map. This loss ultimately affects the effectiveness of attention calculation. To extract various distribution patterns present in images in different scenarios and establish robust spatial dependencies, a hybrid local-global response paradigm is applied.

[0132] Regarding the spatial morphological information of objects within a local region is the key to effective segmentation. Therefore, it becomes crucial to establish effective local blood vessel spatial relationships.

[0133] See Figure 4 , the application of the hybrid attention module includes at least the following steps:

[0134] Create hybrid blocks through a sliding window mechanism to establish local mappings of the input feature map;

[0135] Subsequently, apply self-attention operations to establish local-global co-dependencies, promoting the association of hybrid group information within the window, which not only preserves the spatial structures at different scales but also fuses the hybrid channel information. See the following formula:

[0136] F out = Hardswish(BN(SeparCon(F in )))

[0137] In the formula, Fin is the input feature, Fout is the output feature, SeparConv is the depthwise separable convolution with different kernel sizes, the number of channels remains unchanged during the convolution process, BN is batch normalization, and Hardswish is the Hardswish activation function. Local mappings are performed through different sliding window methods.

[0138] For local mapping, see Figure 5 For global mapping, see Figure 6 .

[0139] See Figure 7 , the fusion pooling module includes at least one 2×2 convolution (2x2 Conv) and a pooling module;

[0140] The application of the fusion pooling module at least includes the following steps:

[0141] First, pass the input feature map through a 2×2 convolution to keep the number of channels unchanged to preserve the original fine feature information;

[0142] Subsequently, use the result of the convolution as the input for max pooling (Max-Pool) and average pooling (Avg-Pool). The pooling operation effectively reduces the feature dimension while retaining the effective information;

[0143] However, the loss of feature information caused by pooling cannot be ignored, especially in the task of segmenting a large number of relatively small objects. Therefore, minimizing the loss of detailed information caused by pooling while fully retaining the effective features is the key to achieving accurate segmentation. For current mainstream techniques, such as Max-Pooling and Avg-Pooling, they tend to discard a large amount of detailed information. Max-Pooling can actively retain the overall image features and maintain details, especially in tasks at the single-digit pixel level. Therefore, the present invention proposes a fusion pooling method for precise segmentation specifically for different scenarios. Different from existing models that only fuse Max-Pooling and Avg-Pooling together.

[0144] Regarding the feature information caused by pooling, divide the downsampling operation into two branches. One branch uses a fusion pool to combine max pooling and average pooling to ensure effective feature extraction, and a convolution operation is introduced in the other branch;

[0145] The convolution operation uses a kernel size of 2 and a stride of 2, which increases the learnability of the pooling operation. This convolution operation can be learned through parameter adjustment, significantly retaining the fine feature information and generating a feature map of the same size as the pooling operation;

[0146] Finally, perform equal-weight fusion on the two feature maps, greatly retaining the feature information of thin blood vessels.

[0147] Based on the above content, further propose:

[0148] Train and test the proposed DSM-Net network on three public datasets: DRIVE, STARE, and CHASE DB1. Compare the performance of DSM-Net with other existing models.

[0149] DRIVE: The DRIVE dataset contains 40 images, 7 of which are pathological images. We use the official dataset division, that is, the first 20 images are used for training and the last 20 images are used for testing. The entire dataset is screened from the Dutch DR screening set, with a size of 584×565, acquired by a Canon CR5 non-expansive 3CCD camera, and the field of view is 45°. The first manually marked image is used as the ground truth.

[0150] CHASE DB1: The CHASE DB1 dataset includes 28 fundus images of the left and right eyes of 14 children, with a size of 999×960, taken by a handheld Nidek NM200D fundus camera in the London Cardiovascular Health Survey, and the field of view is 30°. We use the first 20 images for training and the last 8 images for testing. We use the first manually marked image as the ground truth.

[0151] STARE: The STARE dataset contains 20 images, 10 of which are pathological images, with a size of 700×605, acquired by a TopCon TRV-50 fundus camera, and the field of view is 35°. We randomly select 20% of them as test images, and the remaining images are used for network training. The first manually marked image is used as the ground truth.

[0152] Compared with the state-of-the-art methods, including 8 specialized retinal vessel segmentation models and 2 general medical image segmentation models, the specific results of the three datasets are shown in Tables 1, 2, and 3. Among them, there are 5 evaluation metrics: Acc, AUC, Se, Sp, F1. The highest metrics are shown in bold black, and "-" indicates that the experimental data are not provided in the original literature.

[0153] And further, the cross-validation results on two datasets are presented as shown in Table 4.

[0154] See Figures 8 - 10 , Figure 8 From left to right in

[0155] Figure 9 are fundus images, CS-Net segmentation results, ResDO-UNet segmentation results, DSM-Net segmentation results, and labeled images.

[0156] Figure 10 From left to right in

[0157] Table 1 Comparison Results of Models on the DRIVE Dataset

[0158]

[0159] Table 2 Comparison Results of Models on the STARE Dataset

[0160]

[0161] Table 3 Comparison Results of Models on the CHASE_DB1 Dataset

[0162]

[0163] Table 4 Cross-Validation Results on Two Datasets

[0164]

[0165] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

Claims

1. An image semantic segmentation method based on serpentine convolution, characterized in that: At least include the following steps: S1: Batch preprocess the collected color fundus data; S2: Divide the data into a non-overlapping dataset; S3: Design a loss function for precise segmentation; S4: Design a segmentation model based on serpentine convolution; S5: Train the segmentation model and conduct segmentation verification. After the segmentation model is trained, complete the semantic segmentation of the image through the trained segmentation model; The segmentation model is the DSM-Net model, and the segmentation model includes a serpentine convolution encoder, a hybrid attention module, and a fusion pooling module; The serpentine convolution encoder includes traditional convolution and serpentine convolution. The traditional convolution is used to capture conventional targets, and the serpentine convolution is used to further extract the morphological features of different objects, thereby extracting deeper semantic features; The serpentine convolution encoder tracks different structures by adjusting the shape of the convolution kernel; The operations of the serpentine convolution encoder include two 3×3 convolution operations and serpentine convolution operations in two directions; The application of the serpentine convolution encoder at least includes the following steps: First, pass ton through a convolutional layer with a 3×3 kernel to generate a smooth feature map; At the same time, linearly shift and perform feature extraction convolution operations on the original feature map in the x and y directions through serpentine convolution, and fuse the obtained smooth feature map with the serpentine convolution response maps in two directions; The response maps include the original features of the object and the topological information extracted by the serpentine convolution layer; Finally, pass the fused features through another convolutional layer with a 3×3 kernel for further fusion and channel number adjustment; The final output includes two basic components: the sliding convolution information extracted by two smooth convolution kernels and the fused adaptive transformation serpentine convolution, thereby obtaining the vascular structure features. See the following formula: F out = Conv(DSC_x(F in ) ⊙ DSC_y(F in ) ⊙ Conv(F in )) In the formula, Fin is the input feature map, Fout is the output feature map, ⊙ is the connection, Conv is the batch normalization and ReLU activation after the 3×3 convolution operation, and DSC x and DSC y respectively represent the dynamic serpentine convolution on the x-axis and y-axis; The fusion pooling module includes at least one 2×2 convolution and pooling module; The application of the fusion pooling module at least includes the following steps: First, pass the input feature map through a 2×2 convolution to keep the channel number unchanged to retain the original subtle feature information; Subsequently, use the result of the convolution as the input for max pooling and average pooling. The pooling operation effectively reduces the feature dimension while retaining the effective information; For the feature information caused by pooling, divide the downsampling operation into two branches. One branch uses a fusion pool to combine max pooling and average pooling to ensure effective feature extraction, and a convolutional operation is introduced in the other branch; The convolutional operation uses a kernel size of 2 and a stride of 2, thereby increasing the learnability of the pooling operation. Moreover, this convolutional operation can be learned through parameter adjustment, significantly retaining the subtle feature information and generating a feature map of the same size as the pooling operation; Finally, equally weight fuse the two feature maps, greatly retaining the feature information of thin blood vessels.

2. The image semantic segmentation method based on serpentine convolution according to claim 1, characterized in that: The S1 at least includes the following steps: Process the fundus images in JPG format, and normalize the data of different channels to input into the model.

3. A method for image semantic segmentation based on serpentine convolution according to claim 1, characterized in that: The dataset division in the S2 at least includes the following steps: Randomly shuffle and divide according to the fundus data; Divide it into training, validation, and test sets in the ratio of 8:1:1; The data of any one fundus cannot cross between these sets.

4. A method for image semantic segmentation based on serpentine convolution according to claim 1, characterized in that: The loss function of the S3 includes DiceLoss, BCELoss, and hybrid loss; The DiceLoss is as follows: The BCELoss is as follows: The hybrid loss is as follows: L seg = α·L dice + β·L bce Where α and β are the weight hyperparameters of the two losses respectively, and their sum is 1; In the formula, N represents the total number of pixel points, and C represents the total number of categories; represents the category value corresponding to the i-th pixel in the c-th category in the binary GroundTruth; represents the predicted probability value of the corresponding category; ε = 1e-5, representing the smoothing exponent, which is used to prevent the denominator prediction from being 0 and thus avoid some extreme situations. In addition, it can also play a role in smoothing the loss and gradient.

5. A method for image semantic segmentation based on serpentine convolution according to claim 1, characterized in that: The design of the segmentation model based on serpentine convolution in the S4 at least includes the following steps: Train through the hybrid loss function formulated in S3; Use the gradient backpropagation in the neural network during training to update the model weights, and judge whether the model is better according to the segmentation effect of the model on the validation set after training, and update the saved model weights; Finally, after the model training is completed, evaluate the segmentation effect on the test set.

6. A method for image semantic segmentation based on serpentine convolution according to claim 1, characterized in that: In the S5, five commonly used evaluation metrics are used for segmentation verification. The evaluation metrics include Acc (accuracy), Se (sensitivity), Sp (specificity), F1 (F1 score), and AUC (area under the curve); The decision rule for segmentation verification in the S5 is that the probability map threshold generated by the network is 0.

5. The pixel points greater than 0.5 are predicted as blood vessels, and vice versa are predicted as the background.

7. A method for image semantic segmentation based on serpentine convolution according to claim 1, characterized in that: The application of the hybrid attention module at least includes the following steps: Create a hybrid block through the sliding window mechanism to establish a local mapping of the input feature map; Subsequently, apply the self-attention operation to establish the local-global co-dependency relationship, promote the association of the hybrid group information within the window, which not only retains the spatial structure of different scales but also fuses the hybrid channel information. Refer to the following formula: F out = Hardswish(BN(SeparConv(F in ))) In the formula, Fin is the input feature, Fout is the output feature, SeparConv is the depthwise separable convolution with different kernel sizes, the number of channels remains unchanged during the convolution process, BN is the batch normalization, Hardswish is the Hardswish activation function, and local mapping is performed through different sliding window methods.

Citation Information

Patent Citations

  • Image semantic segmentation model and segmentation method

    CN116468740A

  • Pathological image hash retrieval method based on snakelike convolution multi-scale pooling and norm attention mechanism

    CN117789934A