Lightweight intracranial hematoma segmentation method based on multi-receptive-field MSF-Deeplab network

By introducing multi-receptive field MSF-DeepLab network and point rendering method into medical image segmentation model, the problems of large amount of calculation and many parameters of the existing model are solved, and efficient intracranial hematoma segmentation performance is achieved.

CN119963583APending Publication Date: 2025-05-09CHANGCHUN UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510046304.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing medical image segmentation model has large calculation volume and many parameters, and it is difficult to achieve high-performance intracranial hematoma segmentation under limited resource environments.

Method used

A lightweight intracranial hematoma segmentation method based on multi-receptive field MSF-DeepLab network is proposed. Through the multi-scale adaptive spatial feature fusion module (MSFBlock) and point rendering (PointRend) method, the calculation amount is reduced and the feature representation ability is improved.

Benefits of technology

It achieves excellent intracranial hematoma segmentation performance under the use of a small number of parameters and low computational complexity, and improves the segmentation accuracy and efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963583A_ABST
    Figure CN119963583A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight intracranial hematoma segmentation method based on a multi-receptive-field MSF-Deeplab network, and relates to the field of medical image segmentation. The method comprises the following steps: preprocessing an intracranial hematoma segmentation data set to obtain a feature map of each intracranial hematoma and encoding the feature map; constructing a lightweight feature extractor, and extracting feature information of the feature map; constructing a lightweight spatial pyramid structure, and extracting local feature information of the feature map; constructing a channel regulator, adjusting the channel output dimension of the feature map, proposing a fusion multi-receptive-field strategy, and converting the feature information into depth feature information with multi-receptive fields; constructing a feature coding branch to perform information coding on the image to obtain a depth feature map, gradually reducing the size of the feature map and gradually enlarging a feature map channel; and constructing a feature decoding branch to perform information decoding on the depth feature map of the image, gradually reducing the size of the feature map and gradually reducing a feature map channel. The method can achieve excellent segmentation performance under the conditions of using a small number of parameters and low computational complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation in the field of image processing, and in particular to a lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-Deeplab network. Background Art

[0002] In the field of medical image segmentation, many deep learning-based image segmentation models have been proposed and have achieved great success in various visual tasks. Intracranial hematoma is a serious neurological disease in medicine. This disease has a high initial mortality rate, so early rapid diagnosis is very important. A large number of experiments have shown that the accuracy of image segmentation methods based on deep learning can reach the level of professional doctors. Compared with manual segmentation methods, the use of deep learning methods can save a lot of time.

[0003] Among them, complex models such as the U-Net series, vision transformer (ViT), and DeepLab series have shown remarkable results, but their impressive performance is accompanied by heavy parameters and computational burdens, which limits their widespread adoption and implementation in limited resource environments.

[0004] In recent years, how to make lightweight medical image segmentation methods perform better under limited resource conditions has become a focus. The MobileNet series uses deep separable convolution to decompose standard convolution into deep convolution and point-by-point convolution, which greatly reduces the amount of calculation and parameters. In UNeXt, multi-layer perceptrons and deep separable convolutions are used to generate parameters and floating-point operations suitable for limited resource environments. ConvUNeXt further improves U-Net by integrating lightweight attention mechanisms and utilizing large kernel convolutions to reduce parameters while maintaining the segmentation advantage.

[0005] The Chinese patent publication number is "CN110298843A", and the name is "Two-dimensional image component segmentation method and application based on improved DeepLab". This method improves the DeepLab network including an encoder and a jump decoder. The encoder includes a multi-convolutional layer unit and a multi-scale adaptive morphological feature extraction unit. The multi-scale adaptive morphological feature extraction unit is connected to the output end of the multi-convolutional layer unit, and the jump decoder simultaneously obtains deep features and shallow features. This method focuses on improving the model segmentation structure and edge clarity, but does not fully consider the limited receptive field, too large model parameters, and in real environments, small targets and blurred morphology.

[0006] Therefore, we propose a lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-Deeplab network to solve the above problems. Summary of the invention

[0007] In order to solve the problems of large amount of computation, many parameters and insufficient feature representation in lightweight models in the network model existing in the prior art, and to improve the practicality of lightweight network models in medical image segmentation, the present invention aims to provide a medical image segmentation method with lighter weight and better performance. The present invention provides a lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-DeepLab network, which proposes a new feature extractor, a multi-scale adaptive spatial feature fusion module (Multi-Scale-ASFF, MSFBlock). Specifically, the MSF module uses different sizes of convolution kernels to obtain features extracted from different receptive fields by fusing a multi-receptive field strategy in the feature extraction stage. The MSF module uses parallel depth separable convolutions to replace standard convolutions with depth convolutions and point-by-point convolutions, which greatly reduces the amount of computation of the model and keeps the model lightweight. At the same time, the ASFF module is used to adaptively fuse features of different scales to avoid information loss or noise interference caused by direct superposition of multi-scale features, thereby improving the performance of the network. In addition, based on the MSF module, by introducing the PointRend point rendering method and the cross-attention method, using the fusion multi-receptive field strategy, and improving the DeepLab model framework, the present invention constructs a new lightweight DeepLab network for intracranial hematoma segmentation, a lightweight and high-performance MSF-DeepLab network.

[0008] The solution to the technical problem of the present invention is:

[0009] A lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-DeepLab network includes the following steps:

[0010] Step S1: prepare the data set, preprocess the intracranial hematoma segmentation data set, obtain each feature map of the segmented intracranial hematoma, encode it according to the segmentation category, and use the encoding result as the true label;

[0011] Step S2: construct a lightweight feature extractor, construct a multi-receptive field module, and extract feature information of the feature map after nonlinear processing;

[0012] Step S3: construct a lightweight spatial pyramid structure, obtain multi-scale object information, extract local features in the image, and fuse feature information of different scales by constructing a cross-attention module;

[0013] Step S4: Based on the lightweight feature extractor and the lightweight spatial pyramid structure, a fusion multi-receptive field strategy is constructed to convert feature information into deep feature information with multiple receptive fields;

[0014] Step S5: construct a channel regulator, adjust the channel output dimension of the feature map extracted by steps S2 and S3, and construct a feature encoding branch and a decoding branch.

[0015] Step S6: Based on the fusion multi-receptive field strategy, decode the depth feature map processed by step S4, restore the resolution size of the depth feature map and reduce the channel output dimension.

[0016] Step S7: The segmentation result obtained in step S6 and the rough prediction mask of the shallow feature composite insulator are generated by a lightweight prediction head, and sent to a multi-layer perceptron MLP to perform sampling optimization of fine-grained features and rough prediction masks.

[0017] Step S8: Select a suitable loss function and common medical image segmentation result evaluation indicators: by minimizing the loss value between the network output result and the data label, until the number of training times reaches the set threshold or the value of the loss function reaches the set range and tends to balance, the model parameter pre-training is considered completed and the model results are saved.

[0018] The beneficial effects of the present invention are as follows:

[0019] 1. The present invention proposes a lightweight feature extractor, introduces an inverted residual connection, and increases the receptive fields of different scales.

[0020] 2. The present invention proposes a fusion multi-receptive field strategy, which improves the spatial pyramid structure, enriches the network structure without increasing the computational complexity of the network model, increases the network receptive field, and improves the segmentation accuracy and efficiency of the model.

[0021] 3. The present invention introduces a point rendering PointRend method, and the final composite insulator segmentation edge obtained will not be overly smoothed.

[0022] 4. The technical solution provided by the present invention can achieve excellent segmentation performance under conditions of using a small number of parameters and low computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic diagram of the framework of the lightweight feature extractor of the present invention;

[0024] Figure 2 Schematic diagram of the cross attention framework of the present invention;

[0025] Figure 3 It is a schematic diagram of the framework of the channel regulator of the present invention;

[0026] Figure 4 It is a schematic diagram of the framework of the lightweight intracranial hematoma segmentation method based on the multi-receptive field MSF-DeepLab network of the present invention; DETAILED DESCRIPTION

[0027] The technical solution proposed by the present invention will be further elaborated in detail below in combination with embodiments and drawings.

[0028] The embodiment of the invention provides a lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-DeepLab network, and the method specifically includes the following steps:

[0029] Step 1: Select a public medical image segmentation dataset, select the intracranial hematoma CT image dataset for screening, divide the dataset, and annotate it according to four types: epidural hematoma (EDH), subdural hematoma (SDH), intraparenchymal hematoma (IPH), and subarachnoid hematoma (SAH).

[0030] Step 2, such as Figure 1 As shown, a lightweight feature extractor (MSF) module is constructed to generate a feature map with multi-scale receptive field information during feature extraction.

[0031] Step 3: construct a lightweight spatial pyramid structure to obtain multi-scale object information and extract local features in the image.

[0032] Step 4, such as Figure 2 As shown in Figure 2, feature information of different scales captured by a lightweight spatial pyramid module is fused by constructing a cross-attention module.

[0033] Step 5, such as Figure 3 As shown, a channel regulator is constructed. Based on the channel regulator and lightweight feature extractor, a fusion multi-receptive field strategy is designed to achieve multi-branch information mining of image local information and context information by fusing shallow multi-receptive fields and deep multi-receptive fields.

[0034] Step 6, such as Figure 4 As shown, a feature encoding branch is constructed to encode information of the image to obtain a deep feature map, gradually reduce the size of the feature map and gradually enlarge the feature map channel. A feature decoding branch is constructed to decode the information of the deep feature map of the image, gradually restore the size of the feature map and gradually reduce the feature map channel.

[0035] Step 7: Select appropriate loss functions and common evaluation indicators for medical image segmentation results. During the training process, the loss functions selected are cross entropy loss function and Dice loss function. In the field of medical segmentation, commonly used evaluation indicators include DSC, TPR, mAP and mIoU.

[0036] Specifically, in step 1, the medical image segmentation dataset selected is CQ500, which contains 491 scans, including head CT scans to identify bleeding, fractures, and mass effects. We filtered the subfiles of the CQ500 dataset and only selected case slices containing four types of bleeding (EDH, SDH, SAH, IPH), while excluding slices without bleeding and containing contrast agents, and merged the qualified case slices into the training dataset. In order to ensure that the model can identify different types of hematomas, we selected 4188 slices to train the neural network model.

[0037] In the experiment, the intracranial hematoma image segmentation data was divided into training set, validation set and test set in a ratio of 8:1:1. The training set in the data set was preprocessed as follows:

[0038] Step 1.1, Image normalization: The format of the dataset is DICOM. For ease of processing, we save the original DICOM format data in NIFTI format. Set the window width and window height to 135 and 40 respectively, extract the slices in the Z direction (horizontal part), and convert them to PNG format. The CT slices are saved as 512×512 grayscale images.

[0039] Step 1.2, image normalization: In order to make the model more robust during training, the input data is also normalized before training, scaling the input data to the range of 0 to 1. This helps to make the distribution range of the data more uniform, thus helping the model converge to the optimal solution faster.

[0040] Optionally, in step 2, the process of constructing a lightweight feature extractor (MSF) module is as follows:

[0041] Step 2.1 uses a 3×3 standard convolution to increase the number of channels of the network while keeping the feature map resolution unchanged, to make up for the problem that the depthwise separable convolution does not work well on low channels. This 3×3 standard convolution is only involved in the feature extraction of the first two layers of the backbone network.

[0042] Step 2.2 uses parallel depth-wise separable convolutions to perform 3×3 depth-wise convolution and 7×7 depth-wise convolution on the feature map to extract features, and uses a residual structure to optimize the network to avoid network degradation, thereby maintaining accuracy and reducing computational effort.

[0043] Step 2.3 uses the ASFF (Adaptive Structure Feature Fusion) spatial adaptive network to fuse the multi-receptive field information extracted in step 2.2. It learns the method of spatially filtering conflicting information to suppress inconsistencies, thereby improving the scale invariance of features and adding almost no additional overhead. The feature fusion formula is as follows:

[0044]

[0045] We use represents the feature vector at position (i, j) on the feature map resized from layer n to layer l. We perform fusion on layer l according to the following formula. Represents the output feature map y l The (i,j)th vector along the channel. Represents the spatial importance weights of the three different levels to the levell feature map, which are learned by network adaptive learning. can be a scalar variable, shared by all channels. Using this approach, features at each level can be adaptively aggregated at each scale.

[0046] Optionally, the process of constructing the lightweight spatial pyramid structure in step 3 is as follows:

[0047] Step 3.1 The feature pyramid structure consists of a 1×1 convolution plus pooling pyramid and ASPP Pooling. The expansion factor of each layer of the pooling pyramid can be customized to achieve free multi-scale feature extraction. Since the number of parameters in the ASPP part is too large, we use spatially separable convolution to decompose all 3×3 dilated convolutions in ASPP into 3×1 and 1×3 convolutions in two dimensions, effectively reducing the number of parameters and calculations, and further lightweighting the network so that the image processing speed reaches the real-time standard.

[0048] Optionally, in step 4, the process of fusing feature information of different scales captured by the lightweight spatial pyramid module by constructing a cross attention module is as follows:

[0049] Step 4.1 uses cross attention to fuse the features extracted from the feature pyramid in step 3.1. The cross attention module can better capture the dependencies between different features. ASPP itself focuses on information fusion at different scales, but it mainly relies on different dilated convolutions to capture contexts of different ranges. The cross attention module can more flexibly weight different features according to context information, enhance the expression of important features, and suppress irrelevant features, thereby making feature fusion more effective and improving the expressiveness of the model. The calculation formula for cross attention is as follows:

[0050]

[0051] Where Q is the query matrix, K is the key matrix, and V is the value matrix. Calculate the inner product QK of the query matrix Q and the transpose of all key matrices K T, the purpose is to measure the similarity between each query and each key. The larger the inner product value, the stronger the correlation between the query and the key. k is the dimension of the vector in the key matrix, and the inner product result is divided by This is to prevent the dimension d of the vector in the key matrix from k When it is large, the inner product is too large, causing the gradient to vanish or explode in the subsequent softmax operation. The output is the sum of the weighted value vectors of each element in the query matrix, where the weights are determined by the attention scores of the query and key matrices.

[0052] Optionally, in step 5, a channel regulator is constructed and a multi-receptive field fusion strategy is designed. The channel regulator module maintains the resolution size of the feature map, uses point convolution to adjust the channel output dimension and the number of channels, and obtains the feature map after channel adjustment. The calculation formula of point convolution is as follows:

[0053] F final =Conv1×1(F final )

[0054] where f final Represents the input feature map, Conv1×1(f final ) represents the feature map f final Perform a 1×1 convolution operation to obtain a new feature output F fainal .

[0055] Step 5.1: After adjusting the channel output dimension, the feature map obtained is batch normalized, which is a mathematical operation of subtracting the mean and dividing by the variance of the batch data. The batch normalized feature map is nonlinearly processed using the Gaussian Error Linear Unit (GELU) activation function.

[0056] Step 5.2 designs a fusion multi-receptive field strategy based on the channel regulator and lightweight feature extractor. The fusion multi-receptive field strategy mainly consists of a lightweight feature extractor and two channel regulators. The feature extractor uses deep convolution with different convolution kernel sizes to extract features of different receptive fields, while the channel regulator is responsible for adjusting the dimension of the feature map output by the backbone network and feature pyramid to perform feature fusion of different receptive fields. The calculation formulas for the 3×3 and 7×7 deep convolutions are as follows:

[0057]

[0058] Where X(i+m-1,j+n-1) is the two-dimensional matrix of the input feature map, which contains all pixels or feature values ​​of the input data, and W 3 (m,n) is the (m,n)th element of the 3×3 depth convolution kernel, W 7(m,n) is the (m,n)th element of the 7×7 depthwise convolution kernel. 3 , b 7 It is a constant term added after the convolution operation, usually obtained through learning, used to adjust the result of the convolution output. k1 (i,j),Y k2 (i, j) is the value of the kth channel of the output feature map at position (i, j). The convolution kernel slides on the input data and performs weighted summation on the corresponding values ​​of each sliding position to obtain a new feature map.

[0059] The calculation formula of lightweight feature pyramid is as follows:

[0060]

[0061] Where Conv(X,r i ) represents different void ratios r i The dilated convolution operation, represents the connection operation, and GAP(X) represents the global average pooling operation.

[0062] The feature fusion formula is as follows:

[0063] F fused (i,j)=[Y k1 (i,j)+Y k2 (i,j)+F ASPP ]

[0064] Among them, F fused (i, j) combines the feature information extracted by 3×3 depth convolution, 7×7 depth convolution and lightweight feature pyramid to obtain a new feature map.

[0065] Optionally, in step 6, a feature encoding branch is constructed to implement information encoding of the image to obtain a depth feature map, gradually reduce the feature map size and gradually enlarge the feature map channel. A feature decoding branch is constructed to implement information decoding of the depth feature map of the image, gradually restore the feature map size and gradually reduce the feature map channel.

[0066] Step 6.1 constructs the feature encoding branch, which is mainly composed of the first encoder, the second encoder, the third encoder, the fourth encoder and ASPP (Atrous Spatial Pyramid Pooling); the encoding branch mainly draws on ResNet34. In order to keep the encoding branch lightweight, one layer of the encoder of ResNet34 is deleted. ASPP also uses spatially separable convolution (SSConv) to decompose all 3×3 dilated convolutions in ASPP into 3×1 and 1×3 convolutions in two dimensions.

[0067] Step 6.2 The first four encoders are connected in series in ascending order to form an encoder branch. The first encoder, the second encoder, the third encoder, and the fourth encoder gradually expand the channel output dimension, which are 64, 128, 256, and 512 respectively. In addition, the first encoder, the second encoder, the third encoder, and the fourth encoder gradually reduce the resolution size of the feature map, which are (H / 2, H / 2), (H / 4, H / 4), (H / 8, H / 8), and (H / 16, H / 16), respectively. H and W are the height and width of the image, respectively. ASPP uses dilated convolutions of different magnifications to extract multi-scale features without changing the resolution size of the feature map.

[0068] Step 6.3 constructs the feature decoding branch, which mainly consists of three parts. The first branch directly sends the shallow features of the second encoder in step 6.2 to the decoder; another branch fuses the multi-scale deep features in ASPP with the deep features of the third and fourth encoders using a fusion multi-receptive field strategy and sends them to the decoder; another branch obtains the shallow features of the first encoder and uses a lightweight prediction head to generate a rough prediction mask for each composite insulator, which is then sent to a multi-layer perceptron (MLP) for sampling optimization of fine-grained features and rough prediction masks.

[0069] Step 6.4 The upsampling modules used in the above decoding branch are two 2x bilinear interpolation modules and one 4x bilinear interpolation module. Finally, the segmentation structure is sent to the MLP so that the final composite insulator segmentation edge will not be over-smoothed.

[0070] Optionally, in step 7, a suitable loss function and a common evaluation index for medical image segmentation results are selected.

[0071] Step 7.1 Calculate the loss function of the network output and label, and achieve better fusion effect by minimizing the loss function. The loss function selection selects the CE loss function and the Dice loss function. The loss function calculation formulas are as follows:

[0072]

[0073] In the CE loss function, i represents the sample, y represents the actual value, a represents the output value of the network, and n represents the total number of samples; in the Dice loss function, y i and They represent the label value and predicted value of pixel i respectively, and N represents the total number of pixels.

[0074] The evaluation indicators used in step 7.2 are mAP, TPR, DSC, mIoU and Specificity. AP refers to the proportion of correctly predicted positive samples in the predicted positive samples and the average. TPR refers to the proportion of correctly predicted positive samples in the actual positive samples. In medicine, it refers to the probability of diagnosing an actual patient as sick. DSC is used to measure the similarity between two sets, and the value range is [0,1]. The larger the value, the more similar the two sets are, and it is often used to calculate the similarity of closed areas. IoU refers to calculating the ratio of the intersection and union of two sets, which are the true value and the predicted value. Specificity refers to the probability that the test result is negative under the condition of true negative. The evaluation indicators mAP, TPR, DSC, mIoU and Specificity used in this article are the results of averaging the above indicators according to the number of categories. The calculation formulas for each indicator are as follows:

[0075]

[0076]

[0077] The confusion matrix is ​​used when calculating AP, TPR, DSC, and Specificity. TP refers to the number of pixels whose true labels and network outputs are the same; FP refers to the number of pixels whose true labels are negative and whose network outputs are positive; FN refers to the number of pixels whose true labels are positive and whose network outputs are negative. TN refers to the number of pixels whose true labels are negative and whose network outputs are positive. When calculating mIoU, k represents the number of pixel classes, k+1 represents the total number of samples, and P represents the number of pixels whose true labels are negative and whose network outputs are positive. ii Indicates the correct number of pixels, P ij and P ji Indicates the number of false positives and false negatives for pixels.

[0078] This embodiment implements a lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-DeepLab network based on the PyTorch programming language on an NVIDIA 3090 GPU with 24GB of memory.

[0079] Among them, the Adam optimizer is used, and the learning rate is 1×10 -4 , set the number of training times to 100, of which the first 40 times are frozen training and the last 60 times are unfrozen training. The number of images input to the network each time in the frozen training is 8, while the number of images input to the network each time in the unfrozen training is 4. The threshold of the loss function value is set to about 0.0005. If it is less than 0.0005, it can be considered that the training of the entire network has been basically completed.

[0080] The technical solution provided by the present invention can achieve excellent segmentation performance under the conditions of using a small number of parameters and low computational complexity. Based on the fusion multi-receptive field strategy, the information of multiple receptive fields in a network layer can be fused to improve the feature representation and achieve excellent segmentation performance.

Claims

1. A lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-Deeplab network, characterized in that: The specific steps are: Step S1: preprocessing the intracranial hematoma segmentation dataset, obtaining and encoding each feature map of the segmented intracranial hematoma, and using the encoding result as the true label; Step S2: construct a lightweight feature extractor, construct a multi-receptive field module, and extract feature information of the feature map after nonlinear processing; Step S3: construct a lightweight spatial pyramid structure, obtain multi-scale object information, extract local features in the feature map, and fuse feature information of different scales by constructing a cross-attention module; Step S4: Based on the lightweight feature extractor and the lightweight spatial pyramid structure, a fusion multi-receptive field strategy is constructed to convert feature information into deep feature information with multiple receptive fields; Step S5: construct a channel regulator, adjust the channel output dimension of the feature map extracted by steps S2 and S3, and construct a feature encoding branch and a decoding branch. Step S6: Based on the fusion multi-receptive field strategy, decode the depth feature map processed by step S4, restore the resolution size of the depth feature map and reduce the channel output dimension. Step S7: The segmentation result obtained in step S6 and the rough prediction mask of the shallow feature composite insulator are generated by a lightweight prediction head, and sent to a multi-layer perceptron MLP to perform sampling optimization of fine-grained features and rough prediction masks. Step S8: Select appropriate loss function and common medical image segmentation result evaluation indicators.

2. According to claim 1, a lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-Deeplab network is characterized in that: The image standardization in S1 is in DICOM format. For ease of processing, we save the original DICOM format data in NIFTI format. Image normalization, in order to make the model more robust during training, the input data is also normalized before training, and the input data is scaled to the range of 0 to 1.

3. According to claim 1, a lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-Deeplab network is characterized in that: In S2, the lightweight feature extractor first uses a 3×3 standard convolution to increase the number of channels of the network to make up for the problem that the depth-separable convolution does not work well on low channels. Then, parallel depth-separable convolution is used to extract features of the feature map using 3×3 depth convolution and 7×7 depth convolution, and the network is optimized using a residual structure. Finally, the ASFF spatial adaptive network is used to fuse the multi-receptive field information extracted by the depth convolution. Using this method, the features of each level can be adaptively aggregated at each scale.

4. The lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-Deeplab network according to claim 1, characterized in that: In S3, spatially separable convolution is used to convert all 3×3 dilated convolutions in ASPP (Atrous Spatial Pyramid Pooling) into 3×1 and 1×3 convolutions in two dimensions, effectively reducing the number of parameters and calculations, and further lightweighting the network. Cross-attention is used to fuse the features extracted from the feature pyramid, and the cross-attention module can better capture the dependencies between different features.

5. The lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-Deeplab network according to claim 1, characterized in that: In S4, the fusion multi-receptive field strategy is mainly composed of a lightweight feature extractor and two channel regulators. The feature extractor uses deep convolution with different convolution kernel sizes to extract features of different receptive fields, while the channel regulator is responsible for adjusting the dimension of the feature map output by the backbone network and the feature pyramid to perform feature fusion of different receptive fields.

6. The lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-Deeplab network according to claim 1, characterized in that: In the S5, the feature encoding branch is constructed by the first encoder, the second encoder, the third encoder, the fourth encoder and ASPP. The output dimensions of the first four encoders are 64, 128, 256 and 512 respectively. ASPP extracts multi-scale features by using dilated convolutions of different magnifications without changing the resolution of the feature map. The feature decoding branch is constructed by two 2x bilinear interpolation modules and one 4x bilinear interpolation module, and the output dimensions are 512, 256 and 64 respectively. Finally, the segmentation structure is sent to the MLP so that the segmentation edge of the final composite insulator will not be over-smoothed.

7. The lightweight intracranial hematoma segmentation method based on a multi-receptive field MSF-Deeplab network according to claim 1, characterized in that: In S8, the loss function is selected from the CE loss function and the Dice loss function. The evaluation indicators are mAP, TPR, DSC and mIoU.

Citation Information

Patent Citations

  • Two-dimensional image component segmentation method based on improved DeepLab, and application

    CN110298843A