Portrait matting method based on double-layer framework and full-scale feature fusion
By constructing a portrait cutout method based on the fusion of two-layer framework and full-scale features, combining full-scale jump connection and attention mechanism, the problems of insufficient prediction of unknown areas in the three-point map and insufficient processing of edge details are solved, and a more efficient cutout effect is achieved.
Patent Information
- Application Number
- CN202411817805.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-22
AI Technical Summary
The existing portrait cutout method handles insufficient prediction of unknown areas in the three-point graph and insufficient prediction of edge details and complexity, and has a high computational burden.
A portrait cutout method based on a two-layer framework and full-scale feature fusion is constructed, including SNet, PNet and Fusion-Net modules, combining full-scale jump connection, receptive field attention and channel attention mechanisms, generate a three-point graph through SNet, PNet predicts rough alpha, Fusion-Net provides the final cutout result, and builds a loss function for optimization.
Improve the accuracy and efficiency of portrait cutouts, especially performance on real images and synthetic images, reduce dependence on three-point images, and reduce computing burden.
Smart Images

Figure CN120355908A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and relates to a human portrait matting method based on a double-layer framework and full-scale feature fusion. Background Art
[0002] Image matting is a key technology in computer vision and graphics, aiming to accurately extract foreground objects in an image. It has important applications in fields such as film production, webcasting, and virtual reality. The extraction of the foreground object is achieved by estimating alpha, which represents the transparency of the foreground object at each pixel. All image matting methods use a trimap to divide the image into three regions, namely the foreground, background, and unknown regions, and the quality of the trimap significantly affects the matting result.
[0003] Human portrait matting is a specialized subtask of image matting, where the foreground subject is a human portrait. Compared with other matting methods, human portrait matting usually needs to process very rich edge details in the foreground, such as hair, clothes, nails, etc., which increases the difficulty of the task.
[0004] Semantic Human Matting (SHM) is a specialized technology for automatic image matting that uses semantic information to accurately and automatically extract high-quality human foregrounds from images.
[0005] With the rise of deep learning-based methods, image matting methods combined with deep learning have gradually demonstrated their superiority. Although using deep learning to complete image matting work has certain advantages, some deep learning-based image matting methods still require users to input a trimap. In many other deep learning-based image matting methods, users no longer need to provide a trimap. Forte and Pitié proposed a deep learning method for simultaneously estimating the foreground, background, and alpha. Their method adopts the encoder-decoder architecture of UNet and uses ResNet-50 as the encoder. Li et al. utilized a novel end-to-end network to solve the automatic matting problem. The network backbone uses ResNet-34 and employs an SE attention module to assist the decoder in learning semantic features. Additionally, they publicly released the AIM-500 dataset, which contains 500 natural images and manually labeled alphas. Chen et al. used semantic information to guide portrait matting. Their method, called Semantic-Guided Human Matting (SGHM), consists of a shared encoder, a segmentation decoder, and a matting decoder. The shared encoder uses ResNet-50 as the core network and uses ASM to combine the features from the encoder and the output from the segmentation decoder for the matting decoder. Through their design, the matting task reduces its dependence on high-quality annotations. Chen et al. proposed a high-precision image matting method. Their network includes a semantic context branch for the semantic segmentation sub-task and a high-resolution detail branch for extracting fine details from the foreground.
[0006] Jin et al. proposed a lightweight encoder-decoder network, where the encoder uses MobileNetV2, which can simultaneously predict the foreground and alpha. And the green spill problem is solved by using bilinear upsampling and skip connections in the convolutional layers. Liu et al. used three sub-networks to generate the matting result, and it was the first to use coarsely annotated data for human matting.
[0007] The Semantic Human Matting (SHM) technology proposed by Chen et al. is the first instance of automatic portrait matting. The architecture involves three networks. The Trimap Generation Network (TNet) predicts the trimap based on semantic information, while the Matting Network (MNet) predicts the rough alpha. Finally, the fusion module generates a high-quality alpha matting. Yaman et al. proposed a portrait matting method without additional input and introduced GAN. Their model is divided into two networks, a segmentation network for generating a rough portrait segmentation and a predicted alpha network for generating alpha, which improves the final prediction accuracy.
[0008] However, the insufficient prediction of the unknown region in the trimap and the insufficient prediction of the edge details and complexity in the subsequent alpha prediction are the main challenges of SHM. In addition, learning and processing rough semantic information and details may bring a considerable computational burden. Summary of the Invention
[0009] To solve the above technical problems, the present invention proposes a portrait matting method based on a double-layer framework and full-scale feature fusion, including:
[0010] S1: Construct an SPT model including SNet, PNet, and Fusion-Net modules, input the original image into the Snet module, and output the final predicted matting image after being processed by the PNet module and the Fusion-Net module in sequence;
[0011] S2: Construct a loss function to calculate the loss of the matting image in S1.
[0012] Preferably, S1 includes:
[0013] S1.1: Input the original image into the Snet module to generate a trimap, and the probabilities of different regions in the trimap are expressed as follows:
[0014]
[0015] Among them, Fs is the foreground probability, Bs is the background probability, and Us is the unknown region probability; F is the portrait subject in portrait matting, B is the background pixel, and U is the unknown region pixel;
[0016] S1.2: If a pixel is outside the unknown region, the conditional probability that the pixel belongs to the foreground is the estimate of the matting:
[0017]
[0018] Since U s is the probability that each pixel belongs to the unknown region, then the probability representation of alpha for all pixels is expressed by the following formula:
[0019]
[0020] Among them, α r is the rough alpha output by PNet, and α p is the output of the Fusion-Net, that is, the final predicted matting image;
[0021] S1.3: Because F S +B S =1-U S , then the formula in S1.2 is simplified to:
[0022] α p = F S + U S α r 。
[0023] Preferably, S2 includes:
[0024] S2.1: Construct a loss function L p :
[0025] L p = γ||α p - α g ||1 + (1 - γ)||c p - c g ||
[0026] where L p is the overall prediction loss, γ is a proportionality parameter, the alpha loss is the difference between the alpha predicted value α p and the alpha true value α g and the fusion loss is the difference between the true image c g and the predicted image c p ;
[0027] S2.2: Based on S2.1, the overall loss L is obtained as:
[0028] L = L P + λL t
[0029] where L is the overall loss, λ = 0.01 is the decomposition constraint, and L t is the loss of the tripartite graph.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] The present invention studies a portrait matting model based on full-scale skip connections and receptive field attention, called the SPT model, which consists of three sub-networks, namely SNet, PNet, and Fusion-Net. SNet uses the SEBlock combined with the channel attention mechanism, together with Mobilenetv4 and Unet to construct a lightweight semantic segmentation module, thereby predicting a trimap. PNet uses the predicted trimap and the original image to predict a rough alpha. In the encoder stage of PNet, ASPP and receptive field attention are combined. In the decoder stage, feature maps of different scales are combined through full-scale skip connections. Fusion-Net provides an accurate matting result. The SPT model is trained on synthetic images and tested using real images and synthetic images, all of which are collated from public datasets. Compared with other methods without a trimap, the SPT model has better performance in terms of real images and synthetic images. The present invention also conducts ablation experiments to evaluate the effectiveness of model components. Description of the Drawings
[0032] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention.
[0033] In the drawings:
[0034] Figure 1 is the workflow of the SPT model;
[0035] Figure 2 are the three modules in the SPT model;
[0036] Figure 3 is a schematic diagram of the full-scale skip connection in PNet;
[0037] Figure 4 is a comparison of the matting results of typical images in the PPM-100 dataset;
[0038] Figure 5 are the test results of 3000 synthetic images;
[0039] Figure 6 is a test result graph of the supplementary test set of 1500 synthetic images;
[0040] Figure 7 is a comparison result of the trimap prediction;
[0041] Figure 8 is a comparison after inputting trimaps from different sources into DIM. Detailed Embodiments
[0042] The following combines the attached Figure 1-Appendix Figure 8 The preferred embodiments of the present invention will be described. It should be understood that the preferred embodiments described herein are only for illustrating and explaining the present invention, and are not used to limit the present invention.
[0043] Embodiment:
[0044] A portrait matting method based on a double-layer framework and full-scale feature fusion, comprising:
[0045] S1: The present invention combines a full-scale skip connection, a receptive field attention, and a channel attention mechanism to construct an SPT model including SNet, PNet, and Fusion-Net modules. SNet is a semantic segmentation network that can generate a trimap from the original image. PNet is an alpha prediction network that combines the trimap generated by SNet and the original image to generate a rough alpha. Fusion-Net combines the outputs of SNet and PNet to give the final matting result ( Figure 1 ).
[0046] SNet includes MobileNetV4, an SE module, and UNet ( Figure 2 ). The SE module between MobileNetV4 and UNet can adaptively adjust the feature values of each channel, thereby improving the network's sensitivity to image features. SNet inputs a 320×320 RGB image. After four layers of downsampling by the encoder of MobileNetV4, it then passes through the SE module. During this period, the number of channels increases from 3 to 320. After the SE module, the size of the feature map is gradually restored through four layers of upsampling in the UNet decoder, and the number of channels decreases from 320 to 3. Increasing the number of channels can provide richer features, but if the number of channels is too high, there will be no significant improvement in the final matting result.
[0047] PNet takes the trimap (with three channels) generated by SNet and the original 320×320 image (also with three channels) as inputs, that is, it receives a 6-channel input. By combining the receptive field attention mechanism in the encoder stage, PNet can more effectively identify and process local details in the image, providing a solid foundation for subsequent processing. During the encoder, the number of channels increases from 6 to 128. Then, after four layers of downsampling, with each layer combined with the receptive field attention, the maximum pooling stride is 2, and the padding is 1. The image is reduced from 320×320 to 20×20. Finally, a total of five different-scale feature maps are obtained. After that, the last feature map passes through ASPP to capture context information. After ASPP, four layers of upsampling restore the 320×320 image. Then, Fusion-Net obtains the foreground, the unknown region in the trimap generated by SNet, and the rough alpha generated by PNet to give an accurate alpha.
[0048] The present invention uses a full-scale skip connection as shown in Figure 3 in the PNet decoder of the SPT model. Specifically, feature maps of five different scales from the encoder enter a four-stage upsampling process, in which each feature map is combined with other feature maps using skip connections. The last feature map does not participate in the downward connection because it is the last layer itself. Through this full-scale skip connection method, low-level details are associated with high-level semantics. The high-level semantic information from the decoder layer performs bilinear interpolation on the feature maps of different scales to complement each other and work together in the matting task.
[0049] In this way, the SPT model uses fine-grained detail features and coarse-grained semantic features to obtain more accurate matting results, and the present invention combines information of all scales to achieve more precise matting results.
[0050] The original image is input into the Snet module and processed successively through the PNet module and the Fusion-Net module to output the final predicted matting image;
[0051] S1.1: Input the original image into the Snet module to generate a trimap. The probabilities of different regions in the trimap are represented as follows:
[0052]
[0053] Among them, Fs is the foreground probability, Bs is the background probability, and Us is the unknown region probability; F is the human body main body in portrait matting, B is the background pixel, and U is the unknown region pixel;
[0054] F in the formula S , B S , and U S must be all-ones matrices with the same width and height as the input image. The trimap generated by SNet finds the probability that each pixel in the image falls into the foreground, background, or unknown region. Pixels belonging to the unknown region represent the approximate contour of the human body and contain complex details such as hair, clothes, and nails.
[0055] S1.2: If a pixel is outside the unknown region, the conditional probability that the pixel belongs to the foreground is the estimate of the matting:
[0056]
[0057] Based on U s being the probability that each pixel belongs to the unknown region, the probability expression of alpha for all pixels at this time is as follows:
[0058]
[0059] Among them, α r is the rough alpha output by PNet, and α p is the output of Fusion-Net, that is, the final matte result;
[0060] S1.3: Since F S +B S =1-U S , the formula in S1.2 is simplified to:
[0061] α p =F S +U S α r .
[0062] It can be seen that when U S is close to 1 and F S is close to 0, α p is approximately α r . When U S is close to 0, α p is approximately F S . In this way, the rough semantics and fine details can be combined to generate the final predicted matte image.
[0063] S2: Construct a loss function to calculate the loss of the matte image in S1. It includes:
[0064] S2.1: Construct the loss function L p :
[0065] L p =γ||α p -α g ||1+(1-γ)||c p -c g ||1
[0066] Among them, L p is the overall prediction loss, γ is a proportionality parameter, the alpha loss is the difference between the alpha predicted value α p and the alpha true value α g , and the fusion loss is the difference between the true image c g and the predicted image c p ;
[0067] S2.2: Based on S2.1, the overall loss L is obtained as:
[0068] L = L P +λL t
[0069] Among them, L is the overall loss, λ = 0.01 is the decomposition constraint, Lt It is the loss of the tripartite graph.
[0070] dataset
[0071] Due to the scarcity of large-scale high-quality portrait datasets, most published articles have used synthetic datasets, combining foreground and background images to provide synthetic images required for the matting task. In this paper, we use several open-source image datasets to synthesize the large-scale portrait dataset needed for model training and evaluation. A total of 3348 foreground images in this invention come from the following public datasets: 212 portraits in the public dataset published by Xu et al. (N. Xu, B. Price, S. Cohen, and T. Huang, “Deep image matting,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2970-2979, 2017.), 500 portraits in the P3M-500-NP dataset published by Ma et al. (S. Ma, J. Li, J. Zhang, H. Zhang, and D. Tao, “Rethinking portrait matting with privacy preserving,” International journal of computer vision, vol. 131, no. 8, pp. 2172-2197, 2023.), 636 portraits in the RealWorldPortrait-636 dataset published by Yu et al. (Q. Yu, J. Zhang, H. Zhang, Y. Wang, Z. Lin, N. Xu, Y. Bai, and A. Yuille, “Mask guided matting via progressive refinement network,” in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pp. 1154-1163, 2021.), and 2000 portraits in the Human-2K dataset published by Liu et al. (Y. Liu, J. Xie, X. Shi, Y. Qiao, Y. Huang, Y. Tang, and X. Yang, “Tripartite information mining and integration for image matting,” in Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 7555-7564, 2021.).From this set, 300 portrait images were randomly selected for testing, and the remaining 3048 were used for training. 45720 synthetic images were obtained by synthesizing 15 high-quality background images from Google's Landmark dataset with 3048 portrait images. The test set containing 3000 images was created by synthesizing 10 background images from Google's Landmark dataset with 300 test portrait images. The PPM-100 dataset published by Ke et al. was used as the test set for real-world images (Z. Ke, J. Sun, K. Li, Q. Yan, and R. W. Lau, “Modnet: Realtime trimap-free portrait matting via objective decomposition,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 1140-1147, 2022.).
[0072] In addition, by synthesizing 375 portrait images from the Distinctions-646 dataset published by Qiao et al. (Y. Qiao, Y. Liu, X. Yang, D. Zhou, M. Xu, Q. Zhang, and X. Wei, “Attention-guided hierarchical structure aggregation for image matting,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 13676-13685, 2020.) and 4 background images from Google's Landmark dataset, the present invention created a supplementary test set including 1500 images. By synthesizing 75 portrait images from the AIM-500 dataset published by Li et al. (J. Li, J. Zhang, and D. Tao, “Deep automatic natural image matting,” arXiv preprint arXiv:2107.07235, 2021.) and the same 4 background images used to prepare the supplementary test set, a test set was created to evaluate the trimap prediction effect of the SPT model.
[0073] Verification experiment:
[0074] 1. Experimental environment and configuration
[0075] The experiment was conducted under the Windows 10 system. Both the training process and the testing process were carried out in the Pytorch 2.1.0 environment. The CPU used was the Intel(R) Xeon(R) Platinum 8255C CPU, and the GPU used was the NVIDIA GeForce RTX-3080.
[0076] 2. Portrait Matting Results
[0077] The methods used for comparison in this paper include KNN, DIM, AlphaGAN, MGmatting (MG), SHM, MODNet, and HAttMatting. The proposed SPT model and the methods for comparison all use the same training set and test set. For the tripartite graph-based methods, the same dilation and erosion methods are used to obtain the tripartite graphs for training, and the foreground is binarized by a threshold and the obtained tripartite tables are randomly dilated to generate the corresponding test tripartite graphs. The SPT model adopts a phased training procedure because the results of direct training are poor. That is, first train SNet, and then jointly train SNet and PNet to fine-tune the results. The number of training rounds for both independent SNet training and subsequent joint training is 120 rounds. The learning rate is set to decrease as the number of rounds increases. For independent SNet training, the initial learning rate is 0.001, and the learning rate decreases by 0.1 every 40 rounds. For joint training, the initial learning rate is 0.0001, and the learning rate decreases by 0.1 every 30 rounds. Four metrics are used to evaluate portrait matting of different methods, namely mean squared error (MSE), sum of absolute differences (SAD), slope gradient (Grad), and connectivity (Conn). The lower these metrics are, the better the matting effect.
[0078] Table 1 shows the matting results of the PPM-100 dataset, which contains 100 images from the real world and can well simulate the generalization ability of matting in real life testing. Figure 4 Shows the matting results of some representative images. According to Table 1, the SPT model is significantly better than other non-tripartite graph methods and slightly weaker than the tripartite graph-based methods.
[0079] Table 1 Test Results of PPM-100 Dataset
[0080]
[0081] Table 2 and Figure 5 Shows the test results of 3000 synthetic images.
[0082] Table 2 Test Results of 3000 Synthetic Images
[0083]
[0084] Table 3 andFigure 6 Shows the matte results of the supplementary test set. The SPT model is significantly stronger than other non-trigraph methods, but still inferior to the trigraph-based methods.
[0085] Table 3 Test results of 1500 supplementary test sets
[0086]
[0087] In all three tests, namely PPM-100, the test set of 3000 synthetic images, and the supplementary test set composed of 1500 synthetic images, the method proposed in this paper is superior to other non-trigraph methods and slightly weaker than the trigraph-based methods.
[0088] 3. Trigraph Prediction
[0089] The SNet in the SPT model is compared with the SHM method to evaluate the performance of trigraph prediction. SPT only trains SNet, and SHM only trains T-Net. Both are trained for 40 rounds, with an initial learning rate of 0.001, which decreases by 0.1 every 10 rounds. Other experimental settings are the same as above. The test set is the 300 synthetic images based on AIM-500 described in Section 3.2. The predicted trigraph is compared with the ground truth trigraph, and three metrics are used to evaluate the trigraph prediction, namely mean squared error (MSE), sum of absolute differences (SAD), and intersection over union (IOU). IOU refers to the overlap between the predicted value and the ground truth value, calculated as the ratio of the cross-sectional area of these two regions to the union area:
[0090]
[0091] where A and B represent the unknown regions of the predicted trigraph and the actual trigraph, respectively. When the IOU is large and the MSE and SAD are small, the predicted trigraph is closer to the actual trigraph. Table 4 and Figure 7 show that the trigraph generated by the SPT model is superior to the SHM model, although there is a certain deviation between the generated trigraph and the actual trigraph.
[0092] Table 4 Comparison of predicted results of trigraphs
[0093]
[0094] To further evaluate the trigraph prediction, 300 synthetic images are used to test the matte results of DIM. Since DIM is a trigraph-based method, an additional input of the ground truth trigraph, or the trigraph generated by SPT, or the trigraph generated by SHM is required. The results are shown in Table 5 and Figure 8 show that all four metrics indicate that the ground truth trigraph is superior to the trigraph predicted by SPT, and the prediction of SPT is superior to that of SHM.
[0095] Table 5 Matte extraction results of the tripartite graphs received from different sources in DIM
[0096]
[0097] 4. Ablation Experiments
[0098] Ablation experiments are used to evaluate the effects of each component of the SPT model. The operations performed before the evaluation include:
[0099] (1) Delete the receptive field attention mechanism in PNet and replace RFAConv with ordinary convolution.
[0100] (2) Delete the full-scale skip connections in the PNet decoder.
[0101] (3) Delete ASPP.
[0102] (4) Delete the channel attention in SNet and remove the SE module.
[0103] All other experimental configurations, such as the number of training epochs, the training batch size, and the image size, remain unchanged. The 1500 images in the above section are used as the test set to verify the matte extraction effect. SNet and PNet in Table 6 are after deleting the modules through the above four steps. The test results are given by gradually adding components in the table.
[0104] Table 6 Test results of ablation experiments
[0105]
[0106] It can be seen from Table 6 that combining RFAConv and full-scale connections can improve the matte extraction effect.
[0107] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A method for portrait matting based on a double-layer framework and full-scale feature fusion, characterized in that: Including: S1: Construct an SPT model including SNet, PNet, and Fusion-Net modules. Input the original image into the Snet module, and after being processed by the PNet module and the Fusion-Net module in sequence, output the final predicted matte image. S2: Construct a loss function to calculate the loss of the matte image in S1.
2. The method for portrait matting based on a double-layer framework and full-scale feature fusion according to claim 1, wherein: S1 includes: S1.1: Input the original image into the Snet module to generate a trimap. The probabilities of different regions in the trimap are expressed as follows: Among them, Fs is the foreground probability, Bs is the background probability, and Us is the unknown region probability; F is the human figure main body in the human figure matte, B is the background pixel, and U is the unknown region pixel. S1.2: If a pixel is outside the unknown region, the conditional probability that the pixel belongs to the foreground is the estimate of the matte: Since U s is the probability that each pixel belongs to the unknown region, the probability of alpha for all pixels is expressed by the following formula: Among them, α r is the rough alpha output by PNet, and α p is the output of Fusion-Net, that is, the final predicted matte image; S1.3: Because F S +B S = 1 - U S , the formula of S1.2 is simplified to: α p = F S + U s α r .
3. The method for portrait matting based on a double-layer framework and full-scale feature fusion according to claim 2, wherein: S2 Including: S2.1: Construct the loss function L p : L p = γ || α p -α g || 1+(1 - γ) || c p -c g || 1 where L p is the overall prediction loss, γ is a proportionality parameter, and the alpha loss is the difference between the alpha prediction value α p and the true alpha value α g . The fusion loss is the difference between the true image c g and the predicted image c p . S2.2: On the basis of S2.1, the overall loss L is obtained as: L = L P + λL t Among them, L is the overall loss, λ = 0.01 is the decomposition constraint, and L t is the loss of the tripartite graph.