A lightweight portrait segmentation method based on deep learning
The lightweight portrait segmentation method that combines SegFormer with MobileNetV3 network solves the problem of low accuracy in existing portrait segmentation, achieving higher accuracy and faster segmentation results, and is suitable for applications such as portrait matting and background replacement.
Patent Information
- Application Number
- CN202210998603.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2042-08-19
AI Technical Summary
Existing portrait segmentation methods are not very accurate, and traditional methods suffer from problems such as segmentation noise and rough edges, making them unsuitable for fine segmentation tasks.
A lightweight human image segmentation method based on deep learning is adopted, which uses SegFormer and MobileNetV3 network with coordinate attention mechanism to replace SE module with CA module, thereby improving segmentation accuracy and reducing model size.
It achieves more accurate semantic segmentation results and faster segmentation efficiency, and the model performs well in applications such as portrait matting and online meetings.
Smart Images

Figure CN115457263B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image segmentation, and particularly relates to a lightweight portrait segmentation method based on deep learning. BACKGROUND
[0002] Image segmentation is an important topic in the field of computer vision, and has been widely concerned in many fields such as medicine, agriculture and military. Image segmentation is to divide an image into several sub-image blocks according to the internal law of the image itself, such as the gray value of the pixel, so that each block has a specific feature and forms a significant contrast with other blocks, so as to achieve image segmentation. Portrait segmentation is an important research content of image segmentation. Portrait segmentation is widely used in automatic driving, pedestrian detection, intelligent search and rescue, medical image and other fields, and plays an extremely important role. The traditional portrait segmentation method mainly uses the low-level semantic information of the visual layer of the image, such as the shape and color of the image, to segment the image. The threshold-based segmentation method and the edge-based segmentation method are typical traditional segmentation methods. The threshold-based segmentation method is to classify the pixel points according to the size relationship between the gray threshold value calculated and the gray value of the image pixel. If the difference between the portrait pixels and the background pixels to be separated is small, the edge boundary information will be lost. The edge-based segmentation method uses different edge detection operators for edge detection, but the noise in the image has a greater influence on the detection operator, so this method is only suitable for low-noise images. Because the traditional method has the problems of segmentation noise and rough edge, it is difficult to perform fine segmentation task, and the researchers put forward the portrait semantic segmentation technology based on deep learning. SUMMARY
[0003] The present application solves the problem of low precision of the existing portrait segmentation method, and provides a lightweight portrait segmentation method based on deep learning.
[0004] To solve the above problems, the present application is realized by the following technical scheme:
[0005] A lightweight portrait segmentation method based on deep learning includes the following steps:
[0006] Step 1, constructing a lightweight portrait segmentation model based on deep learning;
[0007] The deep learning-based lightweight portrait segmentation model is composed of an input layer, a convolutional coordinate attention layer, three block units, four multilayer perception layers, a convolutional normalization and activation layer, a convolutional layer, and an output layer; the input of the input layer is taken as the input of the deep learning-based lightweight portrait segmentation model; the output of the input layer is connected to the input of the convolutional coordinate attention layer; one output of the convolutional coordinate attention layer is connected to the input of the first block unit, and the other output is connected to the input of the fourth multilayer perception layer; one output of the first block unit is connected to the input of the second block unit, and the other output is connected to the input of the third multilayer perception layer; one output of the second block unit is connected to the input of the third block unit, and the other output is connected to the input of the second multilayer perception layer; the output of the third block unit is connected to the input of the first multilayer perception layer; the outputs of the four multilayer perception layers are simultaneously connected to the input of the convolutional normalization and activation layer, and the output of the convolutional normalization and activation layer is connected to the input of the convolutional layer; the output of the convolutional layer is connected to the input of the output layer; and the output of the output layer is taken as the output of the deep learning-based lightweight portrait segmentation model.
[0008] Step 2: training the deep learning-based lightweight portrait segmentation model constructed in step 1 by using the segmented sample image set to obtain a trained deep learning-based lightweight portrait segmentation model.
[0009] Step 3: sending an image to be segmented into the trained deep learning-based lightweight portrait segmentation model obtained in step 2, and the trained deep learning-based lightweight portrait segmentation model outputs a segmented image.
[0010] In the above scheme, the convolutional coordinate attention layer is composed of a convolutional layer and a coordinate attention layer; the input of the convolutional layer is taken as the input of the convolutional coordinate attention layer, the output of the convolutional layer is connected to the input of the coordinate attention layer, and the output of the coordinate attention layer is taken as the output of the convolutional coordinate attention layer.
[0011] In the scheme, the first block unit is composed of 9 convolution normalization layers, 3 fusion layers and 1 coordinate attention layer; the input of the first convolution normalization layer is as the input of the first block unit, the output of the first convolution normalization layer is connected to the input of the second convolution normalization layer, the output of the second convolution normalization layer is connected to the input of the third convolution normalization layer, the output of the third convolution normalization layer is connected to one input of the first fusion layer, the other input of the first fusion layer is connected to the input of the first convolution normalization layer, the output of the first fusion layer is connected to the input of the fourth convolution normalization layer; the output of the fourth convolution normalization layer is connected to the input of the fifth convolution normalization layer, the output of the fifth convolution normalization layer is connected to the input of the sixth convolution normalization layer, the output of the sixth convolution normalization layer is connected to one input of the second fusion layer, the other input of the second fusion layer is connected to the input of the fourth convolution normalization layer, the output of the second fusion layer is connected to the input of the seventh convolution normalization layer; the output of the seventh convolution normalization layer is connected to the input of the eighth convolution normalization layer, one output of the eighth convolution normalization layer is connected to one input of the ninth convolution normalization layer, the other output of the eighth convolution normalization layer is connected to the input of the coordinate attention layer, the output of the coordinate attention layer is connected to the other input of the ninth convolution normalization layer, the output of the ninth convolution normalization layer is connected to one input of the third fusion layer, the other input of the third fusion layer is connected to the input of the seventh convolution normalization layer, the output of the third fusion layer is as the output of the first block unit.
[0012] In the scheme, the second block unit is composed of 12 convolution normalization layers, 4 fusion layers and 1 coordinate attention layer; the input of the first convolution normalization layer is as the input of the second block unit, the output of the first convolution normalization layer is connected to the input of the second convolution normalization layer, the output of the second convolution normalization layer is connected to the input of the third convolution normalization layer, the output of the third convolution normalization layer is connected to one input of the first fusion layer, the other input of the first fusion layer is connected to the input of the first convolution normalization layer, the output of the first fusion layer is connected to the input of the fourth convolution normalization layer; the output of the fourth convolution normalization layer is connected to the input of the fifth convolution normalization layer, the output of the fifth convolution normalization layer is connected to the input of the sixth convolution normalization layer, the output of the sixth convolution normalization layer is connected to one input of the second fusion layer, the other input of the second fusion layer is connected to the input of the fourth convolution normalization layer, the output of the second fusion layer is connected to the input of the seventh convolution normalization layer; the output of the seventh convolution normalization layer is connected to the input of the eighth convolution normalization layer, the output of the eighth convolution normalization layer is connected to the input of the ninth convolution normalization layer, the output of the ninth convolution normalization layer is connected to one input of the third fusion layer, the other input of the third fusion layer is connected to the input of the seventh convolution normalization layer, the output of the third fusion layer is connected to the input of the tenth convolution normalization layer; the output of the tenth convolution normalization layer is connected to the input of the eleventh convolution normalization layer, one output of the eleventh convolution normalization layer is connected to one input of the twelfth convolution normalization layer, the other output of the eleventh convolution normalization layer is connected to the input of the coordinate attention layer, the output of the coordinate attention layer is connected to the other input of the twelfth convolution normalization layer, the output of the twelfth convolution normalization layer is connected to one input of the fourth fusion layer, the other input of the fourth fusion layer is connected to the input of the tenth convolution normalization layer, the output of the fourth fusion layer is as the output of the second block unit.
[0013] In the above scheme, the third block unit is composed of 9 convolution normalization layers, 3 fusion layers and 1 coordinate attention layer; the input of the third convolution normalization layer is taken as the input of the third block unit, the output of the first convolution normalization layer is connected to the input of the second convolution normalization layer, the output of the second convolution normalization layer is connected to the input of the third convolution normalization layer, the output of the third convolution normalization layer is connected to one input of the first fusion layer, the other input of the first fusion layer is connected to the input of the first convolution normalization layer, the output of the first fusion layer is connected to the input of the fourth convolution normalization layer; the output of the fourth convolution normalization layer is connected to the input of the fifth convolution normalization layer, the output of the fifth convolution normalization layer is connected to the input of the sixth convolution normalization layer, the output of the sixth convolution normalization layer is connected to one input of the second fusion layer, the other input of the second fusion layer is connected to the input of the fourth convolution normalization layer, the output of the second fusion layer is connected to the input of the seventh convolution normalization layer; the output of the seventh convolution normalization layer is connected to the input of the eighth convolution normalization layer, one output of the eighth convolution normalization layer is connected to one input of the ninth convolution normalization layer, the other output of the eighth convolution normalization layer is connected to the input of the coordinate attention layer, the output of the coordinate attention layer is connected to the other input of the ninth convolution normalization layer, the output of the ninth convolution normalization layer is connected to one input of the third fusion layer, the other input of the third fusion layer is connected to the input of the seventh convolution normalization layer, and the output of the third fusion layer is taken as the output of the third block unit.
[0014] In the above scheme, the convolution normalization layer is composed of a convolution layer and a BN normalization layer; the input of the convolution layer is taken as the input of the convolution normalization layer, the output of the convolution layer is connected to the input of the BN normalization layer, and the output of the BN normalization layer is taken as the output of the convolution normalization layer.
[0015] In the above scheme, each multi-layer perception layer is composed of a linear layer and an LN normalization layer; the input of the linear layer is taken as the input of the multi-layer perception layer, the output of the linear layer is connected to the input of the LN normalization layer, and the output of the LN normalization layer is taken as the output of the multi-layer perception layer.
[0016] In the above scheme, the convolution normalization and activation layer is composed of a convolution layer, a BN normalization layer and a ReLu activation layer; the input of the convolution layer is taken as the input of the convolution normalization and activation layer, the output of the convolution layer is connected to the input of the BN normalization layer, the output of the BN normalization layer is connected to the input of the ReLu activation layer, and the output of the ReLu activation layer is taken as the output of the convolution normalization and activation layer.
[0017] Compared with existing technologies, this invention proposes a lightweight portrait segmentation algorithm based on SegFormer and a fused coordinate attention mechanism using MobileNetV3 to achieve portrait segmentation. The SegFormer encoder is replaced with a MobileNetV3 network; simultaneously, the MobileNetV3 network is improved by incorporating a multi-layer Coordinate Attention (CA) mechanism to enhance the accuracy of portrait segmentation. Furthermore, SEModul in the original network is replaced with CA, accelerating the fitting speed during training, reducing model size, suppressing redundant features, and improving model accuracy, ultimately outputting an accurate semantically segmented image. Experimental comparisons show that our proposed SegFormer-MobileNetV3 network achieves more accurate segmentation results and faster segmentation efficiency compared to other lightweight networks. This invention can be applied to portrait image processing applications such as portrait matting, online meetings, and background replacement. Attached Figure Description
[0018] Figure 1 Here is a diagram of the SegFormer structure.
[0019] Figure 2 This is a diagram of the MobileNetV3 block structure.
[0020] Figure 3 This is a diagram of the Coordinate Attention structure.
[0021] Figure 4 This is a structural diagram of SegFormer-MobileNetV3, a lightweight portrait segmentation model based on deep learning.
[0022] Figure 5 This is the structure diagram of the first unit.
[0023] Figure 6 This is the structure diagram of the second unit.
[0024] Figure 7 This is the structural diagram of the third unit.
[0025] Figure 8 The first set of results (Matting) for segmentation masks using different algorithms.
[0026] Figure 9 The second set of effect diagrams (EG1800) for segmentation masks using different algorithms.
[0027] Figure 10 The third set of effect diagrams for segmentation masks using different algorithms (P3M-10K). Detailed Implementation
[0028] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application is further described in detail below with specific examples.
[0029] A lightweight portrait segmentation method based on deep learning, specifically comprising the following steps:
[0030] Step 1, constructing a lightweight portrait segmentation model based on deep learning.
[0031] The SegFormer network is composed of an encoder (Encoder) and a decoder (Decoder), as shown in Figure 1 The Transformer module in the encoder MiT uses an overlapping patch embedding (OPE) structure to extract features and downsample the input image. The obtained features are input to the Efficient Multihead Self-Attention (EMSA) layer and the Mix Feed Forward (MixFFN) layer. Standard convolutional layers will be used for overlapping patch embedding calculation. After flattening the two-dimensional features into one-dimensional features in space, the features are input to the EMSA layer for self-attention calculation and feature enhancement. In order to replace the position encoding information in the ordinary Transformer, a 3x3 two-dimensional convolutional layer is added between the two linear transformation layers of the ordinary feedforward layer to fuse the position information in space. The Transformer Block uses multiple stacked EMSA and mixed FFN to deepen the network to extract rich detailed and semantic features. Self-attention is calculated in the EMSA of each scale. Compared with other previous convolutional neural network-based networks, although the SegFormer network can integrate information at all scales and make the self-attention mechanism at each scale more pure by calculating self-attention, the existing SegFormer network still has the problems of large model and insufficient segmentation accuracy.
[0032] The MobileNet series is widely used in image processing downstream tasks because of its fast detection and high accuracy, and it has become a representative of lightweight networks. MobileNet uses depthwise separable convolution, which replaces part of the fully connected layer compared to the classic CNN model to reduce the amount of calculation. MoblieNetV3 was published in 2019, which combines the depthwise separable convolution of MoblieNetV1, the Inverted Residuals and Linear Bottleneck of MoblieNetV2, and the SE module, and uses NAS (Neural Architecture Search) to search for the configuration and parameters of the network. The main innovation of MoblieNetV3 is to introduce the SE module on the block of MoblieNetV2. The SE module is a lightweight channel attention module. After depthwise, it goes through the pooling layer, then the first fc layer, the channel number is reduced by 4 times, and then the second fc layer, the channel number is transformed back (expanded by 4 times), and then multiplied with depthwise bit by bit, as shown in Figure 2 .
[0033] Coordinate Attention (CA) is a new attention mechanism module designed for lightweight networks proposed by Qibin Hou et al. of the National University of Singapore, as shown in Figure 3 . It encodes the channel relationship and long-range dependency through precise position information, which is divided into two steps: coordinate information embedding and coordinate attention generation. In order to enable the attention module to capture long-range spatial interactions with precise position information, Qibin Hou et al. decomposed the global pooling into a pair of one-dimensional feature encoding operations according to the following formula:
[0034]
[0035] Given the input X, first use the pooling kernel with size (H, 1) or (1, W) to encode each channel along the horizontal and vertical coordinates respectively. Therefore, the output of the c-th channel with height h can be expressed as:
[0036]
[0037] Similarly, the output of the c-th channel with width w can be written as
[0038]
[0039] The two transformations above aggregate features along two spatial directions respectively, obtaining a pair of direction-aware feature maps. These two transformations also allow the attention module to capture long-range dependencies along one spatial direction and preserve precise location information along the other spatial direction, which helps the network to locate the subject in the image more accurately. Through experiments, it is verified that the CA module has almost no computational overhead on the model, and does not increase the final size of the model.
[0040] Based on the above analysis, the lightweight portrait segmentation model SegFormer-MobileNetV3 based on deep learning proposed in the application has an overall network model architecture as shown in Figure 4 The input of the input layer is used as the input of the lightweight portrait segmentation model based on deep learning; the output of the input layer is connected to the input of the convolution coordinate attention layer; one output of the convolution coordinate attention layer is connected to the input of the first block unit, and the other output is connected to the input of the fourth multi-layer perception layer; one output of the first block unit is connected to the input of the second block unit, and the other output is connected to the input of the third multi-layer perception layer; one output of the second block unit is connected to the input of the third block unit, and the other output is connected to the input of the second multi-layer perception layer; the output of the third block unit is connected to the input of the first multi-layer perception layer; the outputs of the four multi-layer perception layers are simultaneously connected to the inputs of the convolution normalization and activation layer; the output of the convolution normalization and activation layer is connected to the input of the convolution layer; the output of the convolution layer is connected to the input of the output layer; and the output of the output layer is used as the output of the lightweight portrait segmentation model based on deep learning.
[0041] The convolution coordinate attention layer is composed of a convolution layer and a coordinate attention layer; the input of the convolution layer is used as the input of the convolution coordinate attention layer, the output of the convolution layer is connected to the input of the coordinate attention layer, and the output of the coordinate attention layer is used as the output of the convolution coordinate attention layer.
[0042] The first block unit is composed of nine convolution normalization layers, three fusion layers and one coordinate attention layer; as Figure 5The input of the first convolution normalization layer is as the input of the first block unit, the output of the first convolution normalization layer is connected to the input of the second convolution normalization layer, the output of the second convolution normalization layer is connected to the input of the third convolution normalization layer, the output of the third convolution normalization layer is connected to one input of the first fusion layer, the other input of the first fusion layer is connected to the input of the first convolution normalization layer, and the output of the first fusion layer is connected to the input of the fourth convolution normalization layer. The output of the fourth convolution normalization layer is connected to the input of the fifth convolution normalization layer, the output of the fifth convolution normalization layer is connected to the input of the sixth convolution normalization layer, the output of the sixth convolution normalization layer is connected to one input of the second fusion layer, the other input of the second fusion layer is connected to the input of the fourth convolution normalization layer, and the output of the second fusion layer is connected to the input of the seventh convolution normalization layer. The output of the seventh convolution normalization layer is connected to the input of the eighth convolution normalization layer, one output of the eighth convolution normalization layer is connected to one input of the ninth convolution normalization layer, the other output of the eighth convolution normalization layer is connected to the input of the coordinate attention layer, the output of the coordinate attention layer is connected to the other input of the ninth convolution normalization layer, the output of the ninth convolution normalization layer is connected to one input of the third fusion layer, the other input of the third fusion layer is connected to the input of the seventh convolution normalization layer, and the output of the third fusion layer is as the output of the first block unit.
[0043] The second block unit is composed of 12 convolution normalization layers, 4 fusion layers and 1 coordinate attention layer; as Figure 6The output of the first convolutional normalization layer is connected to the input of the second convolutional normalization layer, the output of the second convolutional normalization layer is connected to the input of the third convolutional normalization layer, the output of the third convolutional normalization layer is connected to one input of the first fusion layer, the other input of the first fusion layer is connected to the input of the first convolutional normalization layer, and the output of the first fusion layer is connected to the input of the fourth convolutional normalization layer. The output of the fourth convolutional normalization layer is connected to the input of the fifth convolutional normalization layer, the output of the fifth convolutional normalization layer is connected to the input of the sixth convolutional normalization layer, the output of the sixth convolutional normalization layer is connected to one input of the second fusion layer, the other input of the second fusion layer is connected to the input of the fourth convolutional normalization layer, and the output of the second fusion layer is connected to the input of the seventh convolutional normalization layer. The output of the seventh convolutional normalization layer is connected to the input of the eighth convolutional normalization layer, the output of the eighth convolutional normalization layer is connected to the input of the ninth convolutional normalization layer, the output of the ninth convolutional normalization layer is connected to one input of the third fusion layer, the other input of the third fusion layer is connected to the input of the seventh convolutional normalization layer, and the output of the third fusion layer is connected to the input of the tenth convolutional normalization layer. The output of the tenth convolutional normalization layer is connected to the input of the eleventh convolutional normalization layer, one output of the eleventh convolutional normalization layer is connected to one input of the twelfth convolutional normalization layer, the other output of the eleventh convolutional normalization layer is connected to the input of the coordinate attention layer, the output of the coordinate attention layer is connected to the other input of the twelfth convolutional normalization layer, the output of the twelfth convolutional normalization layer is connected to one input of the fourth fusion layer, the other input of the fourth fusion layer is connected to the input of the tenth convolutional normalization layer, and the output of the fourth fusion layer is the output of the second block unit.
[0044] The third block unit is composed of 9 convolutional normalization layers, 3 fusion layers and 1 coordinate attention layer; as Figure 7The output of the first convolution normalization layer is connected to the input of the second convolution normalization layer, the output of the second convolution normalization layer is connected to the input of the third convolution normalization layer, the output of the third convolution normalization layer is connected to one input of the first fusion layer, the other input of the first fusion layer is connected to the input of the first convolution normalization layer, the output of the first fusion layer is connected to the input of the fourth convolution normalization layer. The output of the fourth convolution normalization layer is connected to the input of the fifth convolution normalization layer, the output of the fifth convolution normalization layer is connected to the input of the sixth convolution normalization layer, the output of the sixth convolution normalization layer is connected to one input of the second fusion layer, the other input of the second fusion layer is connected to the input of the fourth convolution normalization layer, the output of the second fusion layer is connected to the input of the seventh convolution normalization layer. The output of the seventh convolution normalization layer is connected to the input of the eighth convolution normalization layer, one output of the eighth convolution normalization layer is connected to one input of the ninth convolution normalization layer, the other output of the eighth convolution normalization layer is connected to the input of the coordinate attention layer, the output of the coordinate attention layer is connected to the other input of the ninth convolution normalization layer, the output of the ninth convolution normalization layer is connected to one input of the third fusion layer, the other input of the third fusion layer is connected to the input of the seventh convolution normalization layer, and the output of the third fusion layer is the output of the third block unit.
[0045] The convolution normalization layer is composed of a convolution layer and a BN normalization layer; the input of the convolution layer is the input of the convolution normalization layer, the output of the convolution layer is connected to the input of the BN normalization layer, and the output of the BN normalization layer is the output of the convolution normalization layer.
[0046] Each multi-layer perception layer is composed of a linear layer and an LN normalization layer; the input of the linear layer is the input of the multi-layer perception layer, the output of the linear layer is connected to the input of the LN normalization layer, and the output of the LN normalization layer is the output of the multi-layer perception layer.
[0047] The convolution normalization and activation layer is composed of a convolution layer, a BN normalization layer and a ReLu activation layer; the input of the convolution layer is the input of the convolution normalization and activation layer, the output of the convolution layer is connected to the input of the BN normalization layer, the output of the BN normalization layer is connected to the input of the ReLu activation layer, and the output of the ReLu activation layer is the output of the convolution normalization and activation layer.
[0048] The overall architecture of the lightweight portrait segmentation model SegFormer-MobileNetV3 based on deep learning is divided into two parts: Encoder and Decoder. The Encoder uses ConvCAblock to preprocess the image, and then the feature map is transmitted into the fourth layer of MobileNetV3 and Feature Pyramid. MobileNetV3 is divided into three blocks, each of which is composed of several residual units, and the output of each block becomes the input of the feature pyramid and the next block. In the feature pyramid, we use an improved MLP layer as the basic module of the network. After the network output of each pyramid layer, we use linear interpolation to unify the pyramid of different layers, upsample it to 1 / 4 of the original image size, and mix the outputs together through a ConvBNReLu layer, and then output a feature map with the same size as the input image through a conv1x1 layer using bilinear interpolation as the prediction result.
[0049] After inputting the image, ConvCAblock passes through a conv3x3 layer to enable the model to capture global feature information at the beginning, and at the same time, it can better determine the main feature information that the model needs to focus on. In order to make the module have a stronger ability to capture global features, compared with the previous MobileNetV3, we replace the channel attention (SEModule) with a coordinate attention mechanism (Coordinate Attention). This unit constitutes the most basic MobileNetV3 unit, and finally through experimental verification, the size of the model can be further compressed, thereby ensuring the lightweight of the model. In order to better fuse the features, each MLP layer of the feature pyramid uses Linear to extract features, and LayerNorm layer (formula 4) is used to normalize the output of the feature map. The LayerNorm layer flattens the size relationship between different samples, while preserving the size relationship between different features. To reduce the ICS phenomenon, GELU function (formula 5) is used for activation. The Gaussian Error Linear Unit activation function adds a random regularity to the activation, making up for the shortcomings of ReLU6 in nonlinearity, and can effectively avoid the problem of gradient disappearance.
[0050]
[0051]
[0052] The SegFormer-MobileNetV3 adopts a cross entropy loss function commonly used in the field of deep learning image semantic segmentation as a training loss function, and the cross entropy loss function curve is monotonous as a whole, the greater the loss, the greater the gradient, which is conducive to the rapid optimization of the back propagation, so as to ensure that a better model evaluation result can be achieved. For a binary classification problem, the corresponding loss formula is as follows:
[0053] L = -(ylogP + (1-y)log(1-P)) (6)
[0054] Wherein, y is a label, and P is a probability of correct prediction.
[0055] Step 2, using the segmented sample image set to train the lightweight portrait segmentation model based on deep learning constructed in step 1 to obtain the trained lightweight portrait segmentation model based on deep learning.
[0056] Step 3, the image to be segmented is sent into the trained lightweight portrait segmentation model based on deep learning obtained in step 2, and the trained lightweight portrait segmentation model based on deep learning outputs the segmented picture.
[0057] The performance of the present application will be described below through specific examples:
[0058] The experimental hardware platform is a Tesla V100 GPU with a total of 32GB of video memory, an Intel processor with a total of 24 cores. The software environment is ubuntu16.04, python3.7.4, and PaddlePaddle2.3.1. The best hyperparameters of the algorithm of the present application are shown in Table 1. In Table 1, the variable Epoch represents the number of training iterations; Batchsize represents the number of pictures input into the model during each small round of training; optimizer represents the network parameter used to adjust the influence of model training; Lr represents the learning rate; Factor represents the ratio of learning rate decay when the loss function no longer decreases; Embed_dim is the output channel number, which affects the channel size of the intermediate layer.
[0059] Table 1 Best hyperparameters of the algorithm of the present application
[0060]
[0061] The training data set is collected by crawling network portrait half-length photos and manually labeling mask tags. The collected portraits are 18509. In order to meet the data quantity requirement of MixVisionTransformer in SegFormer, by collecting Matting, Supervisely, P3M-10K, EG1800 and other public portrait data sets, the final portrait training set reaches 88229 images, the verification set is composed of 5811 Matting images, 425 EG1800 images and 1000 P3M-10K, a total of 7236 images. Since the collected pictures are of different sizes, the reshape function is used to train and verify the pictures according to the size of [512, 512].
[0062] Random horizontal flipping, random image distortion, random Gaussian blur and other methods are used for data augmentation on the training set. It can increase the diversity of the training set, make the input image more complex, so as to ensure that the model can adapt to more complex environment and improve the robustness of the model.
[0063] In order to correctly evaluate the method of the present application, the commonly used measurement indicators in the field of image semantic segmentation are used as evaluation indicators: mean intersection over union (MioU), accuracy (Acc), model consistency (Kappa) and image segmentation coefficient (Dice).
[0064] Intersection over Union (MioU): In the problem of semantic segmentation, one set is the true value and the other set is the predicted value. The IoU of each class is calculated, and then the average is taken, that is:
[0065]
[0066]
[0067] Where k is the number of classes, i represents the true value, j represents the predicted value, pij represents that i is predicted as j, and pji represents that j is predicted as i.
[0068] Accuracy (Acc): when performing binary classification, the calculation formula is as follows:
[0069]
[0070] Where TP, TN, FP and FN belong to the concept of confusion matrix, and their meanings are shown in Table 2.
[0071] Table 2 Confusion matrix
[0072]
[0073] Model consistency (Kappa): For classification problems, consistency is whether the model prediction result is consistent with the actual classification result.
[0074] The Kappa coefficient based on the confusion matrix is calculated as follows:
[0075]
[0076]
[0077]
[0078] Image segmentation coefficient (Dice): applied to the pixel level index, the value range is [0, 1], when Dice is closer to 1, the model effect is better.
[0079]
[0080] Where X is the predicted pixel set, and Y is the real value pixel set.
[0081] In order to test the effectiveness of the method proposed in this paper, the method proposed in this paper and FCN-HRNet-W18 and SegFormer using different Encoders are verified on three test sets, and the results are shown in Tables 3-5.
[0082] (1) Matting dataset
[0083] Table 3 Validation results of different methods on Matting dataset
[0084]
[0085] As shown in Table 3, on the 5811 Matting test set, our method not only maintains a smaller model size, but also is 11.1M and 1.4M smaller than SegFormer-MiT-B0 and SegFormer-MobileNetv3 respectively, and the four evaluation indexes of MIoU, Acc, Kappa and Dice on the 5811 Matting dataset are better than those of the compared models, which are 1.30% higher than the MIOU index and 0.66% higher than the Acc of the SegFormer-MiT-B0 model. Compared with SegFormer-MiT-B0 and SegFormer-MiT-B0-OurMLP, we found that our MLP does not reduce the running speed of the model. Compared with SegFormer-MobileNetv3 and our model, we found that our model not only reduces the size of the model, but also improves the maximum FPS, and through experimental verification, our model can maintain a stable frame rate of more than 30FPS on the 5811 pictures.
[0086] From the input image of Figure 8 It can be seen that the portrait and the background have a high degree of color similarity, resulting in a large misjudgment phenomenon in all algorithms, mainly manifested in the misjudgment of the clothes at the lower left corner of the portrait and the background, so that the gap between the background and the portrait is completely segmented out, and the hair part on the left of the figure is missing. The main reason is that the portrait mask provided by Matting is not very accurate annotation, but is relatively rough, and in the performance of Matting dataset, Figure 8 (d) is relatively Figure 8 (c) poor, but our model is relatively Figure 8 (e) and other models, are better.
[0087] (2) EG1800 dataset
[0088] Table 4 Different methods on EG1800 dataset validation results
[0089]
[0090]
[0091] From Table 4, on the 425 EG1800 test set, our model exceeds the SegFormer-MiT-B0 model by 1.83% MIOU and 0.90% Acc. Since the EG1800 dataset itself has a small amount of data, the feature information that the model can obtain during training is relatively small. This also highlights the advantage of our model, which is lightweight and does not require a large training set to better complete the training of the model. Larger models will consume more memory during training and migration. At the same time, we found that SegFormer-MiT-B0-OursMLP and SegFormer-MiT-B0, when the data volume is insufficient, SegFormer-MiT-B0-OursMLP performs better. At the same time, we compare our model with SegFormer-MobileNetv3, and the modified MobileNetv3 has a higher MIOU value when the data volume is small, exceeding the unmodified model by 0.85% MIOU value, and obtaining a faster model running frame rate.
[0092] From Figure 9 (c) and Figure 9 (d), we find that SegFormer-MiT-B0-OursMLP performs better, again proving that our improved MLP is more suitable for small data volume. Compared with Figure 9 (b) and other models, Figure 9(b) Although high-resolution HRNet is used, there are still misjudgments when separating the portrait and the background. Our model has fewer misjudgments and smaller misjudgment areas.
Claims
1. A lightweight portrait segmentation method based on deep learning, characterized in that, The method comprises the following steps: Step 1, constructing a deep learning-based lightweight portrait segmentation model; The deep learning-based lightweight portrait segmentation model is composed of an input layer, a convolution coordinate attention layer, three block units, four multilayer perception layers, a convolution normalization and activation layer, a convolution layer and an output layer; the input of the input layer serves as the input of the deep learning-based lightweight portrait segmentation model; the output of the input layer is connected with the input of the convolution coordinate attention layer; one output of the convolution coordinate attention layer is connected with the input of the first block unit, and the other output is connected with the input of the fourth multilayer perception layer; one output of the first block unit is connected with the input of the second block unit, and the other output is connected with the input of the third multilayer perception layer; one output of the second block unit is connected with the input of the third block unit, and the other output is connected with the input of the second multilayer perception layer; the output of the third block unit is connected with the input of the first multilayer perception layer; the outputs of the four multilayer perception layers are simultaneously connected with the input of the convolution normalization and activation layer, and the output of the convolution normalization and activation layer is connected with the input of the convolution layer; the output of the convolution layer is connected with the input of the output layer; the output of the output layer serves as the output of the deep learning-based lightweight portrait segmentation model; The first block unit is composed of nine convolution normalization layers, three fusion layers and one coordinate attention layer; the input of the first convolution normalization layer serves as the input of the first block unit, the output of the first convolution normalization layer is connected with the input of the second convolution normalization layer, the output of the second convolution normalization layer is connected with the input of the third convolution normalization layer, the output of the third convolution normalization layer is connected with one input of the first fusion layer, the other input of the first fusion layer is connected with the input of the first convolution normalization layer, and the output of the first fusion layer is connected with the input of the fourth convolution normalization layer; The output of the fourth convolution normalization layer is connected with the input of the fifth convolution normalization layer, the output of the fifth convolution normalization layer is connected with the input of the sixth convolution normalization layer, the output of the sixth convolution normalization layer is connected with one input of the second fusion layer, the other input of the second fusion layer is connected with the input of the fourth convolution normalization layer, and the output of the second fusion layer is connected with the input of the seventh convolution normalization layer; the output of the seventh convolution normalization layer is connected with the input of the eighth convolution normalization layer, one output of the eighth convolution normalization layer is connected with one input of the ninth convolution normalization layer, the other output of the eighth convolution normalization layer is connected with the input of the coordinate attention layer, the output of the coordinate attention layer is connected with the other input of the ninth convolution normalization layer, the output of the ninth convolution normalization layer is connected with one input of the third fusion layer, the other input of the third fusion layer is connected with the input of the seventh convolution normalization layer, and the output of the third fusion layer serves as the output of the first block unit; The second block unit is composed of 12 convolution normalization layers, 4 fusion layers and 1 coordinate attention layer; the input of the first convolution normalization layer is taken as the input of the second block unit, the output of the first convolution normalization layer is connected with the input of the second convolution normalization layer, the output of the second convolution normalization layer is connected with the input of the third convolution normalization layer, the output of the third convolution normalization layer is connected with one input of the first fusion layer, the other input of the first fusion layer is connected with the input of the first convolution normalization layer, and the output of the first fusion layer is connected with the input of the fourth convolution normalization layer; The output of the fourth convolution normalization layer is connected with the input of the fifth convolution normalization layer, the output of the fifth convolution normalization layer is connected with the input of the sixth convolution normalization layer, the output of the sixth convolution normalization layer is connected with one input of the second fusion layer, the other input of the second fusion layer is connected with the input of the fourth convolution normalization layer, and the output of the second fusion layer is connected with the input of the seventh convolution normalization layer; the output of the seventh convolution normalization layer is connected with the input of the eighth convolution normalization layer, the output of the eighth convolution normalization layer is connected with the input of the ninth convolution normalization layer, the output of the ninth convolution normalization layer is connected with one input of the third fusion layer, the other input of the third fusion layer is connected with the input of the seventh convolution normalization layer, and the output of the third fusion layer is connected with the input of the tenth convolution normalization layer; the output of the tenth convolution normalization layer is connected with the input of the eleventh convolution normalization layer, one output of the eleventh convolution normalization layer is connected with one input of the twelfth convolution normalization layer, the other output of the eleventh convolution normalization layer is connected with the input of the coordinate attention layer, the output of the coordinate attention layer is connected with the other input of the twelfth convolution normalization layer, the output of the twelfth convolution normalization layer is connected with one input of the fourth fusion layer, the other input of the fourth fusion layer is connected with the input of the tenth convolution normalization layer, and the output of the fourth fusion layer is taken as the output of the second block unit; The third block unit is composed of 9 convolution normalization layers, 3 fusion layers and 1 coordinate attention layer; the input of the third convolution normalization layer is taken as the input of the third block unit, the output of the first convolution normalization layer is connected with the input of the second convolution normalization layer, the output of the second convolution normalization layer is connected with the input of the third convolution normalization layer, the output of the third convolution normalization layer is connected with one input of the first fusion layer, the other input of the first fusion layer is connected with the input of the first convolution normalization layer, and the output of the first fusion layer is connected with the input of the fourth convolution normalization layer; The output of the fourth convolution normalization layer is connected to the input of the fifth convolution normalization layer, the output of the fifth convolution normalization layer is connected to the input of the sixth convolution normalization layer, the output of the sixth convolution normalization layer is connected to one input of the second fusion layer, the other input of the second fusion layer is connected to the input of the fourth convolution normalization layer, and the output of the second fusion layer is connected to the input of the seventh convolution normalization layer; the output of the seventh convolution normalization layer is connected to the input of the eighth convolution normalization layer, one output of the eighth convolution normalization layer is connected to one input of the ninth convolution normalization layer, the other output of the eighth convolution normalization layer is connected to the input of the coordinate attention layer, the output of the coordinate attention layer is connected to the other input of the ninth convolution normalization layer, the output of the ninth convolution normalization layer is connected to one input of the third fusion layer, the other input of the third fusion layer is connected to the input of the seventh convolution normalization layer, and the output of the third fusion layer is the output of the third block unit. Step 2: training the lightweight portrait segmentation model based on deep learning constructed in step 1 by using the segmented sample image set to obtain a trained lightweight portrait segmentation model based on deep learning. Step 3: sending an image to be segmented into the trained lightweight portrait segmentation model based on deep learning obtained in step 2, and outputting a segmented image by the trained lightweight portrait segmentation model based on deep learning.
2. The lightweight portrait segmentation method based on deep learning according to claim 1, characterized in that, The convolution coordinate attention layer is composed of a convolution layer and a coordinate attention layer; the input of the convolution layer is the input of the convolution coordinate attention layer, the output of the convolution layer is connected to the input of the coordinate attention layer, and the output of the coordinate attention layer is the output of the convolution coordinate attention layer.
3. The lightweight portrait segmentation method based on deep learning according to claim 1, characterized in that the convolution The normalization layer is composed of a convolution layer and BN a normalization layer. The input of the convolution layer is taken as the input of the convolution normalization layer, and the output of the convolution layer is connected BN the input of the normalization layer, BN the output of the normalization layer is taken as the output of the convolution normalization layer.
4. The lightweight portrait segmentation method based on deep learning according to claim 1, characterized in that, Each multi-layer perception layer consists of a linear layer and LN a normalization layer; The input of the linear layer is connected as the input of the multi-layer perception layer, and the output of the linear layer is connected LN The input of the normalization layer, LN The output of the normalization layer is connected as the output of the multi-layer perception layer.
5. The lightweight portrait segmentation method based on deep learning according to claim 1, characterized in that the convolution Normalization and activation layers consist of convolutional layers, BN Normalized layers and ReLu The activation layer consists of two parts; the input of the convolutional layer serves as the input to the convolutional normalization layer and the activation layer, and the output of the convolutional layer is connected to... BN The input to the normalization layer, BN Output connection of normalization layer ReLu Input to the activation layer, ReLu The output of the activation layer is used as the output of the convolution normalization and the activation layer.
Citation Information
Patent Citations
Lightweight semantic segmentation method for high-resolution remote sensing image
CN112183360A
RGB-T image salient target detection method combining independent decoding and joint decoding
CN113822855A