Medical image segmentation lightweight method based on convolutional neural network

By constructing a lightweight UNet variant model of large core convolution blocks, jump connection blocks and self-attention layers, the dilemma of computing resources and accuracy in medical image segmentation is solved, and efficient medical image segmentation on mobile devices is achieved.

CN120259642APending Publication Date: 2025-07-04SOUTHWEAT UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510135269.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has a dilemma of high computing resource requirements and insufficient accuracy in medical image segmentation, making it difficult to efficiently perform medical image segmentation on mobile devices.

Method used

A lightweight UNet variant model is constructed using large-core convolution blocks, jump connection blocks and self-attention layers. The global features are captured by large-core convolution blocks, and the jump connection blocks are realized with feature fusion. The self-attention layer captures global context and channel interactions, reducing the computational complexity.

Benefits of technology

It improves the accuracy and efficiency of medical image segmentation, realizes efficient deployment on devices with limited computing resources, and improves the practicality and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259642A_ABST
    Figure CN120259642A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image segmentation lightweight method based on a convolutional neural network, and relates to the technical field of medical image analysis. Comprising the following steps: acquiring a medical image data set, and dividing the medical image data set into a training set and a verification set in proportion; constructing a medical image segmentation lightweight model based on the convolutional neural network; inputting the training set into a convolutional neural network-based medical image segmentation lightweight model for model training to obtain a trained convolutional neural network-based medical image segmentation lightweight model; inputting the verification set into a trained medical image segmentation lightweight model based on the convolutional neural network for model evaluation; and performing primary processing on a to-be-segmented medical image, and inputting the to-be-segmented medical image into the verified convolutional neural network-based medical image segmentation lightweight model for image segmentation. According to the method, the medical image segmentation performance is optimized, the requirement for computing resources is reduced, and the method better adapts to the application scene of medical image segmentation on mobile terminal equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image analysis, and more specifically, to a lightweight method for medical image segmentation based on a convolutional neural network. Background Art

[0002] Medical image segmentation algorithms based on deep learning have been widely used in many medical fields. However, in practical applications, especially on devices with limited computing resources (such as mobile devices), the existing technologies still have some limitations and deficiencies. The UNet variant network based on conventional convolution is one of the earliest convolutional neural networks designed specifically for medical image segmentation. This type of network uses small convolutional kernels and multi-level convolutional operations, resulting in a large number of network layers and parameters, and limited effectiveness in capturing global features of images. Due to the small receptive field of its local convolutional operations, it cannot effectively handle long-range dependencies in complex medical images; models based on the attention mechanism apply the attention mechanism to medical image segmentation, aiming to utilize its ability to capture long-range dependencies and global features to make up for the deficiencies of convolutional neural networks in local feature extraction. However, its main drawback is that the computational cost is very high, making it difficult for these models to be deployed on mobile devices; lightweight networks, in order to address the problem of excessive computational cost of traditional deep learning models, a variety of lightweight neural network structures have been proposed one after another and are suitable for deployment on mobile devices. However, in medical image segmentation tasks, simply applying these lightweight networks often leads to a decrease in segmentation accuracy.

[0003] The existing technologies often face a dilemma when dealing with complex medical images: either improving the segmentation accuracy but resulting in excessive computing resource requirements, or reducing the computational cost at the expense of accuracy. Therefore, in devices with limited resources, the existing technologies have problems such as difficulty in balancing performance and efficiency, large number of parameters, high computational complexity, inability to effectively capture global features, and insufficient accuracy of lightweight models.

[0004] Therefore, proposing a lightweight method for medical image segmentation based on a convolutional neural network to solve the difficulties existing in the existing technologies is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a lightweight method for medical image segmentation based on a convolutional neural network, which optimizes the performance of medical image segmentation, reduces the demand for computing resources, and better adapts to the application scenario of medical image segmentation on mobile devices.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A lightweight method for medical image segmentation based on a convolutional neural network, comprising;

[0008] S1. Obtain a medical image dataset, and divide the medical image dataset into a training set and a validation set according to a ratio;

[0009] S2. Construct a lightweight medical image segmentation model based on a convolutional neural network;

[0010] S3. Input the training set into the lightweight medical image segmentation model based on the convolutional neural network for model training to obtain a trained lightweight medical image segmentation model based on the convolutional neural network;

[0011] S4. Input the validation set into the trained lightweight medical image segmentation model based on the convolutional neural network for model evaluation to obtain a validated lightweight medical image segmentation model based on the convolutional neural network;

[0012] S5. After preliminarily processing the medical image to be segmented, input it into the validated lightweight medical image segmentation model based on the convolutional neural network for image segmentation.

[0013] For the above method, optionally, the medical image dataset in S1 includes: BUSI dataset and ISIC dataset.

[0014] For the above method, optionally, in S2, follow the U-shaped architecture design to construct a lightweight medical image segmentation model based on the convolutional neural network;

[0015] The model consists of an encoder, a decoder, and a bottleneck fusion block with skip connections. The whole model includes two stages: the first stage is the convolutional operation stage of the encoder and the decoder, and the second stage is the self-attention mechanism stage.

[0016] For the above method, optionally, the encoder is divided into four levels, and the decoder is divided into four levels.

[0017] For the above method, optionally, the encoder is a large kernel convolutional block, which consists of depthwise separable convolutions with large kernels, and replaces the complete ordinary convolution operation by combining depth convolution and pointwise convolution; after each convolution operation, the GELU activation function and the batch normalization layer are used to optimize the non-linear expression of the model, and the batch normalization layer is used after activation by the GELU activation function after each convolution; during the encoding process, spatial downsampling of the image is performed by a convolution operation with a stride of 2;

[0018] The decoder is a skip fusion block that achieves smooth skip connections through grouped convolutions, replacing the full fusion of semantic features between the encoder and decoder with ordinary convolutions; the convolution operation is divided into two groups, one group performs step-by-step extraction on the encoder features, and the other group performs step-by-step extraction on the upsampled decoder features. After grouped convolutions, depthwise convolution is performed, and pointwise convolution is used for feature fusion. After each convolution, there is a GELU activation function and a normalization layer; the upsampling block consists of an upsampling layer, a convolutional layer, a batch normalization layer, and a ReLU activation function, and bilinear interpolation is used to upsample the feature map by a factor of two;

[0019] The bottleneck fusion block consists of a depthwise convolutional layer and a single-head self-attention layer. The depthwise convolutional layer is used to aggregate local features, and the single-head self-attention layer is used to model global context and the interaction between channels.

[0020] In the above method, optionally, in S4, four metrics, Intersection over Union, F1 Score, Params, and GFLOPs, are used to evaluate the model;

[0021] Intersection over Union and F1 Score are used to evaluate the overlap degree between the model prediction result and the true result. Among them, Intersection over Union represents the intersection-over-union ratio of the predicted region and the true region, and F1 Score represents the comprehensive score of precision and recall between the prediction and the truth;

[0022] Params and GFLOPs are used to measure the complexity and computational overhead of the model. Params represents the number of model parameters, and GFLOPs represents the number of floating-point operations per second in billions.

[0023] In the above method, optionally, the specific content of the preliminary processing of the medical image to be segmented in S5 is to use a convolutional block to extract the top-level original features of the medical image to be segmented. The convolutional block includes a convolutional layer, a batch normalization layer, and a ReLU activation layer, with a kernel size of 3×3, a stride of 1, and a padding of 1.

[0024] From the above technical solutions, it can be seen that compared with the prior art, the present invention provides a lightweight method for medical image segmentation based on a convolutional neural network, and its beneficial effects are:

[0025] The present invention proposes a lightweight UNet variant model based on a convolutional neural network, which can efficiently complete medical image segmentation tasks; a large-kernel convolutional block and a skip connection block are designed to guide the network to focus on the global features and detailed information of the image, solve the problem of insufficient perception of global features in the traditional UNet model for medical image segmentation tasks, and improve the processing ability for blurred edges and detailed regions; enhance the model's ability to capture complex features in medical images; and achieve the efficient deployment of the lightweight model on devices with limited computing resources, improving the practicality and adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0027] Figure 1 It is a flowchart of a lightweight method for medical image segmentation based on a convolutional neural network provided by the present invention;

[0028] Figure 2 It is a structural diagram of a lightweight model for medical image segmentation based on a convolutional neural network provided by the present invention;

[0029] Figure 3 It is a detailed diagram of the encoder of a lightweight model for medical image segmentation based on a convolutional neural network provided by the present invention;

[0030] Figure 4 It is a detailed diagram of the decoder of a lightweight model for medical image segmentation based on a convolutional neural network provided by the present invention;

[0031] Figure 5 It is a detailed diagram of the bottleneck fusion block of a lightweight model for medical image segmentation based on a convolutional neural network provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0033] Refer to Figure 1 As shown, the present invention discloses a lightweight method for medical image segmentation based on a convolutional neural network, including;

[0034] S1. Obtain a medical image dataset, and divide the medical image dataset into a training set and a validation set according to a ratio;

[0035] S2. Construct a lightweight medical image segmentation model based on a convolutional neural network;

[0036] S3. Input the training set into the lightweight medical image segmentation model based on the convolutional neural network for model training to obtain a trained lightweight medical image segmentation model based on the convolutional neural network;

[0037] S4. Input the validation set into the trained lightweight medical image segmentation model based on the convolutional neural network for model evaluation to obtain a verified lightweight medical image segmentation model based on the convolutional neural network;

[0038] S5. After preprocessing the medical image to be segmented, input it into the verified lightweight medical image segmentation model based on the convolutional neural network for image segmentation.

[0039] Further, the medical image dataset in S1 includes: BUSI dataset and ISIC dataset.

[0040] Further, referring to Figure 2 As shown, in S2, following the U-shaped architecture design, the U-shaped structure has become an important paradigm in the design of medical image segmentation networks, and a lightweight medical image segmentation model based on a convolutional neural network is constructed;

[0041] The model consists of an encoder, a decoder with skip connections, and a bottleneck fusion block. The whole model includes two stages: the first stage is the convolutional operation stage of the encoder and decoder, and the second stage is the self-attention mechanism stage.

[0042] Further, the encoder is divided into four levels, and the decoder is divided into four levels.

[0043] The model consists of a four-level encoder-decoder structure with skip connections and a bottleneck fusion block. The whole model includes two stages: the first stage is the convolutional operation stage of the encoder-decoder, and the second stage is the self-attention mechanism stage;

[0044] Further, referring to Figure 3 As shown, the encoder is a large kernel convolutional block, which consists of depthwise separable convolutions with large kernels to mix long-range spatial location information. The convolution with large kernels can utilize the convolution inductive bias and obtain larger receptive field information at the same time;

[0045] By combining the use of depthwise convolution (i.e., the number of groups is equal to the number of channels) and pointwise convolution (kernel size of 1×1) to replace the full ordinary convolution operation, the depthwise convolution focuses on extracting spatial dimension information, and the subsequent pointwise convolution then completes the mixing of space and channels; the depthwise convolution not only reduces the network parameter quantity and computational complexity, but also can extract global information within each channel through the configuration of large kernel convolution, while capturing local details. After that, a residual connection is applied to maintain feature consistency;

[0046] To more fully mix spatial and channel information, pointwise convolution is further applied after the depthwise convolution. In addition, the GELU activation function and batch normalization layer are used after each convolution operation to optimize the non-linear expression of the model. The algorithm can be expressed as the following formula:

[0047] f′ l = BN(GELU{DepthwiseConv(f l-1 )}) + f l-1

[0048] f″ l = BN(GELU{PointwiseConv(f l )})

[0049] f l = BN(GELU{PointwiseConv(f″ l )})

[0050] where f l represents the output feature map of the l-th layer in the large kernel convolution block, and BN represents the batch normalization layer; f' l represents the intermediate feature map that, after depthwise convolution extraction of information from f l , undergoes GELU activation and batch normalization (BN) processing, and retains the input features through a residual connection; f″ l represents the intermediate feature map after pointwise convolution, GELU activation, and batch normalization (BN) processing of f' l ;

[0051] During the encoding process, spatial downsampling of the image is performed through a convolution operation with a stride of 2 to enhance the stability of the model during training and improve performance at the same time; however, medical images usually exhibit low resolution and small local edge variations. Compared with using convolution for downsampling, conventional pooling operations can effectively remove the noise present in medical images while maintaining the minimum computational overhead. Therefore, the adopted downsampling strategy is max pooling, with a filter window of 2×2 and a stride of 2;

[0052] Furthermore, referring to Figure 4As shown, the decoder is a skip fusion block that achieves smooth skip connections through grouped convolution, replacing the full fusion of semantic features between the encoder and decoder with ordinary convolutions; traditional skip connections usually use ordinary convolution operations for feature fusion, which increases the burden on the encoder and decoder. In the skip fusion block of the present invention, grouped convolution is used as the core component to solve these problems. The features before fusion are adaptively assigned to the grouped convolution, and the convolution operation is divided into two groups, and step-by-step extraction is performed on the encoder features and the upsampled decoder features respectively; the kernel size of the grouped convolution is 3×3, the stride is 1, and the padding is 1;

[0053] To achieve full feature fusion, depth convolution is performed after the grouped convolution, and efficient and dense pointwise convolution undertakes the heavy feature fusion work. After each convolution in the skip fusion block is a GELU activation and a normalization layer. The algorithm can be expressed as the following formula:

[0054]

[0055] f skip = BN(GELU{PointwiseConv(f concat )})

[0056] where f concat represents the channel concatenation result after convolution extraction of the encoder features and the decoder features, which is used for subsequent further fusion; f skip represents the output fusion feature map in the skip fusion block; f a and f b represent the encoder and decoder features respectively;

[0057] The upsampling block consists of an upsampling layer, a convolutional layer, a batch normalization layer, and a ReLU activation function. Bilinear interpolation is used to upsample the feature map by two times. The kernel size of the convolutional layer is 3×3, the stride is 1, and the padding is 1, which can effectively improve the resolution of the feature map while retaining important functions;

[0058] Furthermore, as shown in Figure 5 , the bottleneck fusion block consists of a depth convolutional layer and a single-head self-attention layer. The depth convolutional layer is used to aggregate local features, and the single-head self-attention layer is used to model the global context and the interaction between channels. The features are further refined through a conventional 3×3 convolution operation to enhance feature expression. The combination of depth convolution and single-head self-attention captures local and global dependencies in an efficient manner, effectively supporting the detailed processing requirements of the segmentation task. In order to solve the problem of feature redundancy in the later stage in a more cost-effective way, a simplified single-head self-attention module is adopted to avoid the computational redundancy of multi-head self-attention and ensure that the training and inference processes are more efficient and streamlined;

[0059] In the bottleneck fusion block, a single-head attention layer is applied only to a part of the input channels (assumed to be C p = rC) to achieve efficient spatial feature aggregation and keep the remaining channels unchanged. The algorithm can be expressed by the following formula:

[0060]

[0061] X att ,X res = Split(X, [C p , C - C p )

[0062] where W Q , W K , W V and W O are projection weights, d qk is the dimension of the query and key (default is 16), Concat(.) is the concatenation operation; to maintain the consistency of memory access, the initial C p channels are selected as the representatives of the entire feature map. Finally, the output projection of the single-head self-attention layer is applied to all channels, not just the initial C p channels, so as to ensure that the attention features can be effectively propagated to the remaining channels, eliminate redundant calculations, and the single-head self-attention layer can be interpreted as sequentially unfolding the redundant heads of the previous parallel calculations along the block axis.

[0063] Furthermore, in S4, four metrics, Intersection over Union, F1 Score, Params, and GFLOPs, are used for model evaluation;

[0064] Intersection over Union and F1 Score are used to evaluate the overlap between the model prediction results and the true results, where Intersection over Union represents the intersection over union ratio of the predicted region and the true region, and F1 Score represents the comprehensive score of precision and recall between the prediction and the truth;

[0065] Params and GFLOPs are used to measure the complexity and computational overhead of the model. Params represents the number of model parameters, and GFLOPs represents the number of floating-point operations per second in billions;

[0066] The higher the IoU and F1, the better the segmentation performance of the model; the lower the Params and GFLOPs, the more lightweight the model and the higher the computational efficiency;

[0067] The Params and GFLOPs of the present invention are 2.19 and 4.41 respectively. The IoU and F1 on the BUSI validation set are 74.56 and 82.90 respectively, and the IoU and F1 on the ISIC validation set are 83.56 and 90.22 respectively. After verification, it can evaluate the generalization performance of the model proposed by the present invention in different datasets and random splits.

[0068] Further, the specific content of the preliminary processing of the medical image to be segmented in S5 is to use a convolutional block to extract the top-level original features of the medical image to be segmented, avoiding reducing the output resolution and causing inconsistency with the top-level skip connection. The convolutional block includes a convolutional layer, a batch normalization layer, and a ReLU activation layer, with a kernel size of 3×3, a stride of 1, and a padding of 1. By using this method, the consistency of the skip connection with the top level can be maintained, and it is ensured that important features will not be lost during the encoding process.

[0069] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0070] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A lightweight method for medical image segmentation based on convolutional neural network, characterized in that, including; S1. Obtain a medical image dataset, and divide the medical image dataset into a training set and a validation set according to a ratio; S2. Construct a lightweight medical image segmentation model based on a convolutional neural network; S3. Input the training set into the lightweight medical image segmentation model based on the convolutional neural network for model training to obtain a trained lightweight medical image segmentation model based on the convolutional neural network; S4. Input the validation set into the trained lightweight medical image segmentation model based on the convolutional neural network for model evaluation to obtain a verified lightweight medical image segmentation model based on the convolutional neural network; S5. Perform preliminary processing on the medical image to be segmented and then input it into the verified lightweight medical image segmentation model based on the convolutional neural network for image segmentation.

2. The lightweight medical image segmentation method based on a convolutional neural network according to claim 1, wherein the medical image dataset in S1 includes: the BUSI dataset and the ISIC dataset.

3. The lightweight medical image segmentation method based on a convolutional neural network according to claim 1, wherein in S2, following the U-shaped architecture design, construct a lightweight medical image segmentation model based on a convolutional neural network; The model consists of an encoder, a decoder, and a bottleneck fusion block with skip connections. The whole model includes two stages: the first stage is the convolutional operation stage of the encoder and the decoder, and the second stage is the self-attention mechanism stage.

4. The lightweight medical image segmentation method based on a convolutional neural network according to claim 3, wherein the encoder is divided into four levels, and the decoder is divided into four levels.

5. The lightweight medical image segmentation method based on a convolutional neural network according to claim 3, wherein The encoder is a large-kernel convolutional block, which consists of depthwise separable convolutions with large kernels, and replaces the complete ordinary convolution operation by combining depth convolution and pointwise convolution; after each convolution operation, the GELU activation function and the batch normalization layer are used to optimize the nonlinear expression of the model, and the batch normalization layer is used after activation by the GELU activation function; during the encoding process, spatial downsampling of the image is performed through a convolution operation with a stride of 2; The decoder is a skip fusion block, which realizes smooth skip connections through grouped convolutions, replacing the full fusion of semantic features between the encoder and the decoder of ordinary convolutions; the convolution operation is divided into two groups, one group performs step-by-step extraction on the encoder features, and the other group performs step-by-step extraction on the upsampled decoder features. After grouped convolutions, depth convolution processing is performed, and pointwise convolution is used for feature fusion. After each convolution, there is a GELU activation function and a normalization layer; the upsampling block consists of an upsampling layer, a convolutional layer, a batch normalization layer, and a ReLU activation function, and bilinear interpolation is used to upsample the feature map by two times; The bottleneck fusion block consists of a depth convolutional layer and a single-head self-attention layer. The depth convolutional layer is used to aggregate local features, and the single-head self-attention layer is used to model the global context and the interaction between channels.

6. A lightweight method for medical image segmentation based on convolutional neural network according to claim 1, characterized in that In S4, four metrics, Intersection over Union, F1 Score, Params, and GFLOPs, are used for model evaluation; Intersection over Union and F1 Score are used to evaluate the overlap degree between the model prediction result and the true result, where Intersection over Union represents the intersection-over-union ratio of the predicted region and the true region, and F1 Score represents the comprehensive score of precision and recall between the prediction and the truth; Params and GFLOPs are used to measure the complexity and computational overhead of the model. Params represents the number of model parameters, and GFLOPs represents the number of floating-point operations per second in billions.

7. A lightweight method for medical image segmentation based on convolutional neural network according to claim 1, characterized in that The specific content of the preliminary processing of the medical image to be segmented in S5 is to extract the top-level original features of the medical image to be segmented by a convolutional block. The convolutional block includes a convolutional layer, a batch normalization layer, and a ReLU activation layer, with a kernel size of 3×3, a stride of 1, and a padding of 1.

Citation Information

Patent Citations

  • Medical image segmentation method based on heterogeneous phantom convolution

    CN117994276A