Portrait Matting Method Based on Multi-Task Learning
Through the portrait clutch method based on multi-task learning, a basic and refined picture cutting model was established, and the problem of three-point picture reference and high-resolution picture cutting in the existing technology was solved, and efficient and accurate portrait picture cutting was achieved.
Patent Information
- Application Number
- CN202210260220.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-16
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-16
AI Technical Summary
The existing portrait cutout method requires the cutout personnel to provide additional three-point images as reference, and the precise cutout processing for high-resolution portraits cannot be done, resulting in defects or flaws in the portrait after the cutout.
The portrait clamping method based on multi-task learning is adopted. By establishing a basic cutout model and a refined cutout model, the encoder module, Segment module, Attention module and MattingDecoder module are used for multi-task learning, and a refined masking image is generated to achieve accurate cutout of high-resolution portraits.
No additional three-point reference is required, precise portrait cutouts can be performed at high resolution, avoiding blurred edges of portraits and incomplete body parts, significantly improving the quality and consistency of cutouts.
Smart Images

Figure CN114627293B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of portrait processing, and particularly to a portrait matting method based on multi-task learning. Background Art
[0002] In the process of portrait processing, it is often necessary to perform matting on the portrait in the picture to achieve various post-processing effects. However, for some current portrait matting methods, firstly, it is necessary for the matting personnel to provide an additional Trimap, that is, a tri-map, as a reference. The actual matting process is rather cumbersome, and there are various methods for generating the tri-map, and it is impossible to ensure the consistency of the provided tri-map. Secondly, the existing portrait matting methods cannot perform precise matting on high-resolution (such as 1080p, 2K, 4K, etc.) portraits, and cannot well separate the portrait from the background, resulting in defects or flaws in the obtained portrait after matting, such as blurred portrait edges, incomplete portrait body parts, etc. Therefore, improvements are needed. Summary of the Invention
[0003] In view of the defects in the prior art that the existing methods require the matting personnel to provide an additional tri-map as a reference and cannot perform precise matting on high-resolution (such as 1080p, 2K, 4K, etc.) portraits, the present invention provides a new portrait matting method based on multi-task learning.
[0004] To solve the above technical problems, the present invention is realized through the following technical solutions:
[0005] A portrait matting method based on multi-task learning, comprising the following steps:
[0006] Phase 1: Establish a basic matting model:
[0007] S1. Prepare a sample set, where the sample set includes three types of pictures: RGB original images, original transparent masks corresponding to the RGB original images, and a background data set with portraits removed from the pictures, and divide the sample set into a training set and a validation set;
[0008] S2. Randomly batch extract sample combinations from the training set, and after randomly augmenting these sample combinations, obtain the first-phase RGB image I, the first-phase target transparent mask α gt and the first-phase target segmentation map s gt , input the first-phase RGB image I into the encoder module for encoding to obtain shallow texture information, middle-level feature information, and high-level semantic representation information;
[0009] S3. Input the middle-level feature information and high-level semantic representation information into the Segment module to obtain a rough semantic segmentation map m, and use cross-entropy loss to supervise the training of the rough semantic segmentation map m to obtain a segmentation loss Lseg ;
[0010] S4. Input the high-level semantic representation information and the rough semantic segmentation map m into the Attention module for joint processing to obtain the attention representation information;
[0011] S5. Input the first-stage RGB image I, the attention representation information, and the shallow texture information into the MattingDecoder module to generate the first-stage rough transparency mask α p , the first-stage residual probability map r, and the first-stage hidden feature h. Use the L1 loss and the gradient loss to supervise the training of the first-stage rough transparency mask α p to obtain the transparency mask loss L alp . Use the L2 loss to supervise the training of the first-stage residual probability map r to obtain the residual loss L res ;
[0012] S6. By performing weighted summation on the segmentation loss L seg , the transparency mask loss L alp , and the residual loss L res to obtain the total loss L total ;
[0013] S7. Repeat steps S2 - S6 using the training set and use the validation set for validation to obtain the Loss curve graph of the validation set. Take the training result obtained at the lowest point in the Loss curve graph of the validation set as the basic matting model;
[0014] Stage two: Establish a refined matting model:
[0015] S8. Take the basic matting model and add the Refiner module to form an untrained refined matting model;
[0016] S9. Randomly batch-sample combinations from the training set and perform random augmentation on these sample combinations to obtain the second-stage RGB image I, the second-stage target transparency mask α gt and the second-stage target segmentation map s gt . After scaling the second-stage RGB image I to a size of 512 * 512 pixels, input it into the basic matting model to obtain the second-stage rough transparency mask α p , the second-stage residual probability map r, and the second-stage hidden feature h;
[0017] S10. Input the second-stage RGB image I, the second-stage rough transparency mask α p , the second-stage residual probability map r, and the second-stage hidden feature h into the Refiner module to obtain a refined mask graph α p . Use the transparency mask loss Lalp to supervise the refined matte α p during training;
[0018] S11. Repeat steps S9 - S10 using the training set to obtain a trained refined matting model;
[0019] Phase Three: Perform matting using the refined matting model:
[0020] S12. Read a high - resolution image I to be matted and a target background image B. Scale the image I to be matted to a size of 2048 * 2048 pixels and input it into the trained refined matting model to generate a refined matte α p , then scale both the refined matte α p and the target background image B to the same pixel size as the image I to be matted, denoted as α p′ , B′. Finally, use the formula C = α p′ ·I+(1 - α p′ )·B′ to obtain the image C with the background changed.
[0021] In step S1, the sample set is divided into a training set and a validation set. Training can be carried out using the training set, and the model can be verified and selected using the validation set. The RGB original images are used for training after augmentation by inputting them into the model, and the original transparent mattes corresponding to the RGB original images are used to supervise the training process. The background dataset without portrait images is used to augment the sample set. On the one hand, this can exclude the ambiguity of the semantic understanding of the model caused by photos with portrait backgrounds, and on the other hand, it also improves the generalization performance of the model. The background dataset without portrait images can select the COCO dataset.
[0022] In steps S2 and S3, the high - level semantic representation information is used to represent the categories of objects in the image, such as people, cats, dogs, etc. However, the spatial information in the high - level semantic representation information is lost relatively seriously, and the prediction effect is poor when predicting the rough semantic segmentation map m subsequently. The middle - level feature information can supplement the missing spatial information of the high - level semantic representation information at this time, so as to obtain a relatively better - predicted rough semantic segmentation map m. The shallow - layer texture information contains rich texture detail information, which can help predict the details of the original transparent matte (such as hair) during subsequent processing. The encoder module used to obtain the shallow - layer texture information, middle - level feature information, and high - level semantic representation information can select ResNet.
[0023] In step S4, the rough semantic segmentation map m contains basic category information and approximate contour information. Therefore, using the rough semantic segmentation map m as a guide enables the network to pay more attention to the high-level semantic representation information in the area of the rough semantic segmentation map m. Compared with directly using the high-level semantic representation information, stronger attention representation information can be obtained.
[0024] In step S5, since the enhanced attention representation information has been obtained in step S4, which provides strong semantic information for predicting the transparent matte, the attention representation information is input into the MattingDecoder module. This ensures that the predicted transparent matte does not have unnecessary holes. The shallow texture information in step S2 contains rich texture detail information, which helps to restore the edges and hair details of the transparent matte. At the same time, the L1 loss and gradient loss are used to supervise the training, thereby obtaining a relatively good first-stage rough transparent matte α in terms of both semantics and details. p On the other hand, to enable the model to not only learn to predict the first-stage rough transparent matte α p , but also learn the difference between the currently predicted first-stage rough transparent matte α p and the target transparent matte, the present invention further predicts a first-stage residual probability map r, and this first-stage residual probability map r also provides a reference for the training of the Refiner module in the second stage. The first-stage hidden feature h is used to provide semantic and edge detail information for the Refiner module in the second stage.
[0025] In step S6, the segmentation loss L seg enables the model to pay more attention to semantic understanding, while the transparent matte loss L alp enables the model to focus on the prediction of the transparent matte on the basis of semantic understanding. The residual loss L res enables the model to focus on the difference between the predicted transparent matte and the target transparent matte. By weighted summing the three, the total loss L total is obtained, enabling the model to simultaneously pay attention to semantic understanding and target texture detail issues.
[0026] In step S7, by repeating steps S2 - S6, the training effect can be continuously improved, thereby obtaining a basic matting model with better prediction effect, preparing for the establishment of the subsequent stage-two refined matting model.
[0027] In steps S8 to S11, by adding the Refiner module on the basis of the basic matting model obtained in stage one and performing further repeated training, a refined matting model with better quality can be further obtained, thereby preparing for subsequent matting applications.
[0028] In step S12, the refined matting model obtained in the second stage is used for matting, which can perform precise matting processing in the matting requirements of high-resolution portraits (such as 1080p, 2K, 4K, etc.). The finally matted image will not have defects or flaws, and can better meet people's matting needs.
[0029] Preferably, in the above-mentioned portrait matting method based on multi-task learning, step S2 specifically includes the following steps:
[0030] S21. Randomly read the same number of RGB original images, the original transparent masks corresponding to the RGB original images, and the background dataset with the portraits removed from the training set. The read RGB original images and the original transparent masks corresponding to the RGB original images are combined into training pairs, and the training pairs are augmented with the background dataset with the portraits removed. Then, these augmented training pairs are randomly augmented to obtain the augmented RGB images and the corresponding augmented transparent masks.
[0031] S22. Scale the augmented RGB images and the corresponding augmented transparent masks to a size of 512 * 512 pixels to obtain the first-stage RGB image I and the first-stage target transparent mask α for training. gt and perform morphological operations on the first-stage target transparent mask α gt to obtain the first-stage target segmentation map s gt ;
[0032] S23. Input the first-stage RGB image I into the encoder module to obtain shallow texture information, middle-level feature information, and high-level semantic representation information.
[0033] Since the annotation cost of high-resolution portrait transparent masks is high, the training sample size in high-resolution portrait matting tasks is often limited. Therefore, it is necessary to adopt a suitable data augmentation strategy. The present invention effectively improves the problem of insufficient training sample size through the above steps. Among them, random augmentation can comprehensively use random augmentation strategies such as color, space, and background replacement, such as cropping, affine transformation, Gamma transformation, hue adjustment, saturation adjustment, contrast adjustment, etc.
[0034] Preferably, in the above-mentioned portrait matting method based on multi-task learning, the network structure of the Segment module in step S3 is [conv-bn-relu-conv], and the calculation formula of the segmentation loss L seg is: where x i represents the first-stage RGB image I, and p[(x i )] represents the segmentation prediction result of x i , and yi Represents x i The corresponding label is the first-stage target segmentation map s gt .
[0035] By adding a Segment module to the matting task for multi-task learning, where the Segment module undertakes the segmentation auxiliary task. On the one hand, this can guide the encoder to learn deep semantic information, enabling the subsequent MattingDecoder module to perform matting based on the semantics understood by the encoder, and this makes the entire matting task of the present invention not require additional reference inputs (such as the trimap, etc.); on the other hand, the prediction result of the segmentation auxiliary task will also provide guidance for the generation of subsequent attention representation information.
[0036] Preferably, for the above-mentioned portrait matting method based on multi-task learning, the processing steps of the Attention module in step S4 specifically include the following steps:
[0037] S41. Use a convolution operation with a convolution kernel of 1×1 to compress the channels of the high-level semantic representation information to obtain a pixel-level semantic representation information;
[0038] S42. Perform matrix multiplication on the rough semantic segmentation map m and the pixel-level semantic representation information to obtain class representation information;
[0039] S43. Use convolution and deformation operations to project the pixel-level semantic representation information as the Query feature and project the class representation information as the Key feature and Value feature. The structure of the convolution and deformation operations is [conv-bn-relu]. The Query feature, Key feature, and Value feature are all two-dimensional feature maps. Among them, the Query feature depicts the representation of each pixel, while the Key feature and Value feature depict the representation of each class;
[0040] S44. Substitute the Query feature, Key feature, and Value feature into the formula Softmax(Query T Key)Value T That is, the enhanced attention representation information is obtained, where T is the transpose operation;
[0041] S45. Input the enhanced attention representation information into a convolution operation with a structure of [conv-bn-relu], and then add it to the high-level semantic representation information to obtain the final attention representation information.
[0042] If the high-level semantic representation information generated by the encoder module is directly input into the MattingDecoder module, the MattingDecoder module can also converge, but its convergence effect is not ideal. Therefore, the present invention proposes to use the Attention module to strengthen the high-level semantic representation information under the rough semantic segmentation map m. The specific approach is to use the rough semantic segmentation map m generated by the Segment module as a guide as shown in steps S41 to S45, convert the high-level semantic representation information into enhanced attention representation information through convolution and matrix multiplication operations, and finally add the enhanced attention representation information and the original high-level semantic representation information to obtain the final attention representation information. Using such attention representation information as the input of the MattingDecoder module enables the MattingDecoder module to pay more attention to the high-level semantic representation information under the rough semantic segmentation map m while learning global matting.
[0043] Preferably, in the above-mentioned portrait matting method based on multi-task learning, in step S5, the MattingDecoder module first performs channel compression on the attention representation information in turn using a convolution operation with a convolution kernel of 1×1 to reduce the computational amount, and then interpolates the compressed attention representation information to the feature scale of the shallow texture information and splices it with the shallow texture information and sends it to a convolution operation with a structure of [conv-bn-relu]. Then, the result of this convolution operation is interpolated to the scale of the first-stage RGB image I and spliced with the first-stage RGB image I and sent into a final convolution operation with a structure of [conv]. Then, the final convolution operation outputs the first-stage rough transparency mask α p , the first-stage residual probability map r, and the first-stage hidden feature h.
[0044] The present invention here adopts the idea of layer-by-layer recovery from low to high resolution to combine the attention representation information with the shallow texture information and the first-stage RGB image I in turn to obtain the final prediction output, which enables the MattingDecoder module of the present invention to pay attention to both high-level semantics and low-level texture detail information, thus achieving excellent results in both global and detail prediction.
[0045] Preferably, in the above-mentioned portrait matting method based on multi-task learning, in step S5, the transparency mask loss L alp = L a + L grad , where the definitions of L a and L grad are as follows:
[0046]
[0047] In the above formula, αp represents the first-stage rough transparent mask, α gt represents the first-stage target transparent mask, and ▽ represents the Sobel gradient operator.
[0048] The L1 loss has the characteristics of fast convergence and good prediction effect for details. However, due to the derivative mutation near the extreme points, the model oscillates back and forth near the extreme points and cannot converge well. Therefore, in the present invention, the L1 loss is smoothed, and its formula form is the above L a 、L grad , so that the model can converge quickly in the initial stage and be more stable near the extreme points during the training process, and finally achieve better global and detail prediction effects. The transparent mask loss L alp The L a loss focuses on making the predicted original transparent mask pay attention to the global information, while the L grad loss focuses on making the prediction result have better details.
[0049] Preferably, in the above-mentioned portrait matting method based on multi-task learning, in the step S5, the residual loss L res is defined as follows:
[0050] L res =(r - |α p - α gt |) 2
[0051] In the above formula, α p represents the first-stage rough transparent mask, α gt represents the first-stage target transparent mask, and r represents the first-stage residual probability map.
[0052] The residual loss L res makes the model focus on the gap between the predicted transparent mask and the target transparent mask, and the first-stage residual probability map r also provides a reference for the training of the Refiner module in the second stage.
[0053] Preferably, in the above-mentioned portrait matting method based on multi-task learning, in the step S6, the total loss L total is defined as follows:
[0054] L total =λ·L seg +(1 - λ)·(L alp + L res )
[0055] In the above formula, L seg represents the segmentation loss, L alp represents the transparent mask loss, L resDenote the residual loss as, and denote the weight occupied by the segmentation loss \(L\) as \(\lambda\). \((L + L)\) represents the matting loss, and \((1 - \lambda)\) represents the weight occupied by the matting loss \((L + L)\). seg where \(\lambda\) is the weight of the segmentation loss \(L\), alp and \((L + L)\) represents the matting loss, and \((1 - \lambda)\) represents the weight of the matting loss \((L + L)\). res alp res
[0056] Among them, \((L + L)\) is the loss established for the entire matting target. In this way, through the multi-task learning of segmentation and matting, the model is guided to establish both semantic understanding of the target and pay attention to the texture details of the target. On the other hand, since the main task of the present invention is matting and segmentation is only used as an auxiliary task here to help the convergence of the entire model, a relatively small weight (such as 0.04) can be given to the segmentation loss \(L\) during actual training, and a relatively large weight is given to the matting loss \((L + L)\) accordingly, so that the model finally converges towards a target more conducive to matting. By using multi-task learning, the model trained by the present invention has a certain semantic understanding ability, thus eliminating the need for additional reference inputs (such as Trimap, etc.). alp res seg alp res
[0057] Preferably, for the above-mentioned portrait matting method based on multi-task learning, step S9 specifically includes the following steps:
[0058] S91. Randomly read the same number of RGB original images, the original transparent masks corresponding to the RGB original images, and the background dataset with the portraits removed from the training set. Combine the read RGB original images and the original transparent masks corresponding to the RGB original images to form training pairs, perform background augmentation on the training pairs using the background dataset with the portraits removed, and then perform random augmentation on these background-augmented training pairs to obtain the augmented RGB images and the corresponding augmented transparent masks;
[0059] S92. Scale the augmented RGB images and the corresponding augmented transparent masks to a pixel size of 2048 * 2048 to obtain the second-stage RGB image \(I\) and the second-stage target transparent mask \(\alpha\) for training, and obtain the second-stage target segmentation map \(s\) by performing morphological operations on the second-stage target transparent mask \(\alpha\); gt gt gt
[0060] S93. After scaling the second-stage RGB image \(I\) to a pixel size of 512 * 512, input it into the basic matting model to obtain the second-stage rough transparent mask \(\alpha\). p , the second-stage residual probability map r and the second-stage hidden feature h.
[0061] After the first-stage training, the basic matting model of the present invention has established an efficient semantic understanding ability. Therefore, when establishing the refined matting model in the second stage, its key objective is to focus on the prediction of target edge details. Because if a high-resolution image with a pixel size of 2048*2048 is directly input into the basic matting model, the computational cost is high, and matting, a low-level task, requires texture details at high resolution. Therefore, after obtaining the augmented second-stage RGB image I, the present invention first scales the second-stage RGB image I to a pixel size of 512*512 and then inputs it into the basic matting model to obtain the second-stage rough transparency mask α p , the second-stage residual probability map r and the second-stage hidden feature h. The high-resolution second-stage RGB image I is used in the Refiner module to predict a more refined transparency mask, which not only saves the computational cost of the basic matting model but also ensures the refined prediction requirements for high-resolution matting. Among them, random augmentation can comprehensively use random augmentation strategies such as color, space, and background replacement, such as cropping, affine transformation, Gamma transformation, hue adjustment, saturation adjustment, contrast adjustment, etc.
[0062] Preferably, in the above-mentioned portrait matting method based on multi-task learning, the step S10 specifically includes the following steps:
[0063] S101. Scale the second-stage RGB image I, the second-stage rough transparency mask α p , the second-stage residual probability map r, and the second-stage hidden feature h to 1 / 2 scale of the second-stage RGB image I, and splice them together in the channel dimension. Then, at this scale, crop a batch of crop blocks at a ratio of 1 / 16 in both the horizontal and vertical directions. Then, take out the crop blocks with an average residual probability greater than 0 from the second-stage residual probability map r and denote them as crop blocks P1. Then, expand these crop blocks P1 outward by 16 pixels;
[0064] S102. Input the crop blocks P1 into a two-layer convolutional network and then magnify them by 1 times to obtain the Refine intermediate feature. Then, crop the second-stage RGB image I at a ratio of 1 / 16 in both the horizontal and vertical directions to obtain raw crop blocks. Then, expand these raw crop blocks outward by 32 pixels and splice them together with the Refine intermediate feature in the channel dimension, and then pass through two more layers of convolutional networks to obtain the refined transparency mask blocks α s , and then these refined transparency mask blocks α s Crop off 32 pixels at the edge to obtain the refined transparency mask blocks α q;
[0065] S103. Enlarge the second-stage rough transparent mask α generated by the basic matting model directly to the size of the second-stage RGB image I as the second-stage rough transparent mask α p and paste the refined transparent mask block α obtained in step S102 b back onto the second-stage rough transparent mask α q to obtain the final refined mask α b ; p ;
[0066] S104. Use the transparent mask loss L alp to supervise the training of the refined mask α p .
[0067] In steps S101 - S104, when establishing the refined matting model in the second stage, the present invention simultaneously refers to the global semantic information in the second-stage hidden feature h and the texture details of the second-stage RGB image I. However, for high-resolution images, if direct convolution operations are performed on the entire image, the computational complexity is very high. Therefore, in the Refiner module, the present invention uses the block-by-block Refine method and finally pastes the cropped blocks back into the transparent mask predicted by the enlarged MattingDecoder module. Since some defects may occur at the edges during the prediction of the cropped blocks, the present invention first expands the cropped block P1 by some pixels and then trims its redundant edges after Refine. In the Refine stage, only the transparent mask loss L alp is required for supervised training.
[0068] Through the above operations, a refined mask α with better overall and detailed prediction effects can be obtained for the target image p . BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] The following further describes the present invention in detail in conjunction with the attached Figure 1 drawings and specific embodiments, but they are not limitations to the present invention:
[0071] Embodiment 1
[0072] A portrait matting method based on multi-task learning, including the following steps:
[0073] Stage 1: Establish a basic matting model:
[0074] S1. Prepare a sample set, where the sample set includes three types of pictures: RGB original images, original transparent masks corresponding to the RGB original images, and a background dataset with portrait pictures removed, and divide the sample set into a training set and a validation set;
[0075] S2. Randomly batch extract sample combinations from the training set, and perform random augmentation on these sample combinations to obtain the first-stage RGB image I, the first-stage target transparent mask α gt and the first-stage target segmentation map s gt , input the first-stage RGB image I into the encoder module for encoding to obtain shallow texture information, middle-level feature information, and high-level semantic representation information;
[0076] S3. Input the middle-level feature information and high-level semantic representation information into the Segment module to obtain a rough semantic segmentation map m, and use cross-entropy loss to supervise the training of the rough semantic segmentation map m to obtain the segmentation loss L seg ;
[0077] S4. Input the high-level semantic representation information and the rough semantic segmentation map m into the Attention module for joint processing to obtain attention representation information;
[0078] S5. Input the first-stage RGB image I, the attention representation information, and the shallow texture information into the MattingDecoder module to generate the first-stage rough transparent mask α p , the first-stage residual probability map r, and the first-stage hidden feature h, use L1 loss and gradient loss to supervise the training of the first-stage rough transparent mask α p to obtain the transparent mask loss L alp , use L2 loss to supervise the training of the first-stage residual probability map r to obtain the residual loss L res ;
[0079] S6. By performing weighted summation on the segmentation loss L seg , the transparent mask loss L alp , and the residual loss L res to obtain the total loss L total ;
[0080] S7. Repeat steps S2 - S6 using the training set, and use the validation set for verification to obtain the Loss curve graph of the validation set, and take the training result obtained at the lowest point of the Loss curve graph of the validation set as the basic matting model;
[0081] Stage two: Establish a refined matting model:
[0082] S8. Take the basic matting model and add the Refiner module to form an untrained refined matting model;
[0083] S9. Randomly batch-sample combinations from the training set, and perform random augmentation on these sample combinations to obtain the second-stage RGB image I, the second-stage target transparent mask α gt and the second-stage target segmentation map s gt , scale the second-stage RGB image I to a size of 512*512 pixels and input it into the basic matting model to obtain the second-stage rough transparent mask α p , the second-stage residual probability map r and the second-stage hidden feature h;
[0084] S10. Input the second-stage RGB image I, the second-stage rough transparent mask α p , the second-stage residual probability map r and the second-stage hidden feature h into the Refiner module, so as to obtain a refined mask map α p , and use the transparent mask loss L alp to supervise the training of the refined mask map α p ;
[0085] S11. Repeat steps S9 - S10 using the training set to obtain a trained refined matting model;
[0086] Stage three: Perform matting using the refined matting model:
[0087] S12. Read a high-resolution image I to be matted and a target background image B, scale the image I to be matted to a size of 2048*2048 pixels and input it into the trained refined matting model to generate a refined mask map α p , then scale the refined mask map α p and the target background image B to the same pixel size as the image I to be matted, denoted as α p′ , B', and finally use the formula C = α p′ ·I + (1 - α p′ )·B' to obtain the image C with the background changed.
[0088] Preferably, the step S2 specifically includes the following steps:
[0089] S21. Randomly read the same number of original RGB images, the original transparent masks corresponding to the original RGB images, and the background dataset with portrait-containing images removed from the training set. Combine the read original RGB images and the original transparent masks corresponding to the original RGB images to form training pairs, and perform background augmentation on the training pairs using the background dataset with portrait-containing images removed. Then perform random augmentation on these background-augmented training pairs to obtain the augmented RGB images and the corresponding augmented transparent masks.
[0090] S22. Resize the augmented RGB images and the corresponding augmented transparent masks to a size of 512 * 512 pixels to obtain the first-stage RGB image I for training and the first-stage target transparent mask α. gt ,and perform morphological operations on the first-stage target transparent mask α gt to obtain the first-stage target segmentation map s. gt ;
[0091] S23. Input the first-stage RGB image I into the encoder module to obtain shallow texture information, middle-level feature information, and high-level semantic representation information.
[0092] Preferably, the network structure of the Segment module in step S3 is [conv-bn-relu-conv], and the calculation formula of the segmentation loss L seg is: where x i represents the first-stage RGB image I, p[(x i )] represents the segmentation prediction result of x i , y i represents the label corresponding to x i , that is, the first-stage target segmentation map s gt .
[0093] Preferably, the processing steps of the Attention module in step S4 specifically include the following steps:
[0094] S41. Perform channel compression on the high-level semantic representation information using a convolution operation with a convolution kernel of 1×1 to obtain a pixel-level semantic representation information.
[0095] S42. Perform matrix multiplication on the rough semantic segmentation map m and the pixel-level semantic representation information to obtain category representation information.
[0096] S43. Use convolution and deformation operations to project pixel-level semantic representation information as Query features, and project category representation information as Key features and Value features. The structure of the convolution and deformation operations is [conv-bn-relu]. The Query feature, Key feature, and Value feature are all two-dimensional feature maps. Among them, the Query feature depicts the representation of each pixel, while the Key feature and Value feature depict the representation of each category;
[0097] S44. Substitute the Query feature, Key feature, and Value feature into the formula Softmax(Query T Key)Value T That is, the enhanced attention representation information is obtained, where T is the transpose operation;
[0098] S45. Input the enhanced attention representation information into a convolution operation with the structure of [conv-bn-relu], and then add it to the high-level semantic representation information to obtain the final attention representation information.
[0099] Preferably, in step S5, the MattingDecoder module first uses a convolution operation with a convolution kernel of 1×1 to perform channel compression on the attention representation information in turn to reduce the computational amount. Then, the compressed attention representation information is interpolated to the feature scale of the shallow texture information, concatenated with the shallow texture information, and sent to a convolution operation with the structure of [conv-bn-relu]. Then, the result of this convolution operation is interpolated to the scale of the first-stage RGB image I, concatenated with the first-stage RGB image I, and sent into a final convolution operation with the structure of [conv]. Then, the final convolution operation outputs the first-stage rough transparent mask α p , the first-stage residual probability map r, and the first-stage hidden feature h.
[0100] Preferably, in step S5, the transparency mask loss L alp =L a +L grad , where the definitions of L a and L grad are as follows:
[0101]
[0102] In the above formula, α p represents the first-stage rough transparent mask, α gt represents the first-stage target transparent mask, and ▽ represents the Sobel gradient operator.
[0103] Preferably, in step S5, the residual loss L res is defined as follows:
[0104] L res = (r - |α p - α gt |) 2
[0105] In the above formula, α p represents the first-stage rough transparent mask, and α gt represents the first-stage target transparent mask, and r represents the first-stage residual probability map.
[0106] Preferably, in step S6, the total loss L total is defined as follows:
[0107] L total = λ · L seg + (1 - λ) · (L alp + L res )
[0108] In the above formula, L seg represents the segmentation loss, L alp represents the transparent mask loss, L res represents the residual loss, λ represents the weight of the segmentation loss L seg , and (L alp + L res ) represents the matting loss, and (1 - λ) represents the weight of the matting loss (L alp + L res ).
[0109] Preferably, step S9 specifically includes the following steps:
[0110] S91. Randomly read the same number of RGB original images, the original transparent masks corresponding to the RGB original images, and the background dataset with portrait-containing images removed from the training set. Combine the read RGB original images and the original transparent masks corresponding to the RGB original images to form training pairs, perform background augmentation on the training pairs using the background dataset with portrait-containing images removed, and then perform random augmentation on these background-augmented training pairs to obtain the augmented RGB images and the corresponding augmented transparent masks;
[0111] S92. Scale the augmented RGB images and the corresponding augmented transparent masks to a pixel size of 2048 * 2048 to obtain the second-stage RGB images I for training, the second-stage target transparent masks α gt , and obtain the second-stage target segmentation map s gt by performing morphological operations on the second-stage target transparent masks α gt ;
[0112] S93. Scale the second-stage RGB image I to a size of 512 * 512 pixels and input it into the basic matting model to obtain the second-stage rough transparent mask α. p , the second-stage residual probability map r, and the second-stage hidden feature h.
[0113] Preferably, the step S10 specifically includes the following steps:
[0114] S101. Scale the second-stage RGB image I, the second-stage rough transparent mask α generated by the basic matting model p , the second-stage residual probability map r, and the second-stage hidden feature h to 1 / 2 scale of the second-stage RGB image I, splice them together in the channel dimension, then crop a batch of crop blocks at a ratio of 1 / 16 in both the horizontal and vertical directions at this scale, then take out the crop blocks with an average residual probability greater than 0 through the second-stage residual probability map r and denote them as crop blocks P1, and then expand these crop blocks P1 outward by 16 pixels;
[0115] S102. Input the crop blocks P1 into a two-layer convolutional network and then magnify them by 1 times to obtain the Refine intermediate feature. Then crop the second-stage RGB image I at a ratio of 1 / 16 in both the horizontal and vertical directions to obtain raw crop blocks. Then expand these raw crop blocks outward by 32 pixels and splice them together with the Refine intermediate feature in the channel dimension and then pass through two more convolutional networks to obtain the refined transparent mask blocks α s , and then cut off 32 pixels at the edges of these refined transparent mask blocks α s to obtain the refined transparent mask blocks α q ;
[0116] S103. Directly magnify the second-stage rough transparent mask α generated by the basic matting model p to the size of the second-stage RGB image I as the second-stage rough transparent mask α b , paste the refined transparent mask blocks α obtained in step S102 q back into the second-stage rough transparent mask α b to obtain the final refined mask map α p ;
[0117] S104. Use the transparent mask loss L alp to supervise the training of the refined mask map α p .
[0118] In summary, the above are only the preferred embodiments of the present invention. All equivalent changes and modifications made within the scope of the patent application of the present invention shall fall within the scope covered by the present invention.
Claims
1. Portrait matting method based on multi-task learning, characterized in that: It includes the following steps: Phase 1: Establish a basic matting model: S1. Prepare a sample set, where the sample set includes three types of pictures: RGB original images, original transparent masks corresponding to the RGB original images, and a background data set with portrait pictures removed, and divide the sample set into a training set and a validation set; S2. Randomly batch-sample combinations from the training set, and perform random augmentation on these sample combinations to obtain the first-stage RGB image I and the first-stage target transparent mask α gt , by performing morphological operations on the first-stage target transparent mask α gt to obtain the first-stage target segmentation map s gt . Input the first-stage RGB image I into the encoder module for encoding to obtain shallow texture information, middle-level feature information, and high-level semantic representation information; S3. Input the middle-level feature information and high-level semantic representation information into the Segment module to obtain a rough semantic segmentation map m, and use cross-entropy loss to supervise the training of the rough semantic segmentation map m to obtain the segmentation loss L seg ; The segmentation loss L seg The calculation formula is: where x i represents the first-stage RGB image I, p[(x i )] represents the segmentation prediction result of x i , y i represents the label corresponding to x i , that is, the first-stage target segmentation map s gt ; S4. Input the high-level semantic representation information and the rough semantic segmentation map m into the Attention module for joint processing to obtain attention representation information; S5. Input the first-stage RGB image I, the attention representation information, and the shallow texture information into the MattingDecoder module to generate the first-stage rough transparency mask α p , the first-stage residual probability map r, and the first-stage hidden feature h. Use the L1 loss and the gradient loss to supervise the training of the first-stage rough transparency mask α p to obtain the transparency mask loss L alp . Use the L2 loss to supervise the training of the first-stage residual probability map r to obtain the residual loss L res ; S6. By performing weighted summation on the segmentation loss L seg , the transparency mask loss L alp , and the residual loss L res , the total loss L total is obtained; S7. Repeat steps S2-S6 using the training set, and use the validation set for verification to obtain the Loss curve graph of the validation set. Take the training result obtained at the lowest point of the Loss curve graph of the validation set as the basic matting model; Phase 2: Establish a refined matting model: S8. Take the basic matting model and add the Refiner module to form an untrained refined matting model; S9. Randomly batch-sample combinations from the training set, and perform random augmentation on these sample combinations to obtain the second-stage RGB image I, the second-stage target transparent mask α gt and the second-stage target segmentation map s gt . After scaling the second-stage RGB image I to a size of 512 * 512 pixels, input it into the basic matting model to obtain the second-stage rough transparent mask α p , the second-stage residual probability map r, and the second-stage hidden feature h; S10. Input the second-stage RGB image I, the second-stage rough transparent mask α p , the second-stage residual probability map r, and the second-stage hidden feature h into the Refiner module to obtain a refined mask α p , and use the transparent mask loss L alp to supervise the training of the refined mask α p ; S11. Repeat steps S9-S10 using the training set to obtain a trained refined matting model; Phase 3: Use the refined matting model for matting: S12. Read a high-resolution image I to be matte-extracted and a target background image B. After scaling the image I to be matte-extracted to a pixel size of 2048*2048, input it into the trained refined matte extraction model to generate a refined matte map α p , then scale the refined matte map α p and the target background image B to the same pixel size as the image I to be matte-extracted, denoted as α p′ , B′. Finally, use the formula C = α p′ ·I + (1 - α p′ )·B′ to obtain the image C with the background replaced.
2. The portrait matting method based on multi-task learning according to claim 1, characterized in that: The step S2 specifically includes the following steps: S21. Randomly read the same number of RGB original images, original transparent masks corresponding to the RGB original images, and a background data set with portrait pictures removed from the training set. Combine the read RGB original images and the original transparent masks corresponding to the RGB original images to form a training pair, and perform background augmentation on the training pair through the background data set with portrait pictures removed, and then perform random augmentation on these background-augmented training pairs to obtain augmented RGB images and corresponding augmented transparent masks; S22. Scale the augmented RGB image and the corresponding augmented transparency mask to a size of 512 * 512 pixels to obtain the first-stage RGB image I for training and the first-stage target transparency mask α gt , and perform morphological operations on the first-stage target transparency mask α gt to obtain the first-stage target segmentation map s gt ; S23. Input the RGB image I in the first phase into the encoder module to obtain shallow texture information, middle-level feature information, and high-level semantic representation information.
3. The portrait matting method based on multi-task learning according to claim 1, characterized in that: The network structure of the Segment module in the step S3 is [conv-bn-relu-conv].
4. The portrait matting method based on multi-task learning according to claim 1, characterized in that: The processing steps of the Attention module in the step S4 specifically include the following steps: S41. Perform channel compression on the high-level semantic representation information using a convolution operation with a convolution kernel of 1×1 to obtain a pixel-level semantic representation information; S42. Perform matrix multiplication on the rough semantic segmentation map m and the pixel-level semantic representation information to obtain category representation information; S43. Use convolution and deformation operations to project pixel-level semantic representation information as Query features and project category representation information as Key features and Value features. The structure of the convolution and deformation operations is [conv-bn-relu]. The Query features, Key features, and Value features are all two-dimensional feature maps. Among them, the Query features depict the representation of each pixel, while the Key features and Value features depict the representation of each category; S44. Substitute the Query feature, Key feature, and Value feature into the formula Softmax(Query T Key)Value T That is, the enhanced attention representation information is obtained, where T is the transpose operation; S45. Input the enhanced attention representation information into a convolution operation with the structure of [conv-bn-relu], and then add it to the high-level semantic representation information to obtain the final attention representation information.
5. The multi-task learning-based portrait matting method according to claim 1, characterized in that: In the step S5, the MattingDecoder module first performs a channel compression on the attention representation information using a convolution operation with a convolution kernel of 1×1 to reduce the computational complexity. Then, the compressed attention representation information is interpolated to the feature scale of the shallow texture information and concatenated with the shallow texture information and sent to a convolution operation with the structure of [conv-bn-relu]. Then, the result of this convolution operation is interpolated to the scale of the first-stage RGB image I and concatenated with the first-stage RGB image I and sent into a final convolution operation with the structure of [conv]. Then, the final convolution operation outputs the first-stage rough transparent mask α p , the first-stage residual probability map r, and the first-stage hidden feature h.
6. The multi-task learning-based portrait matting method according to claim 1, characterized in that: In the step S5, the transparent mask loss L alp = L a + L grad , where the definitions of L a and l grad are as follows: In the above formula, α p represents the first-stage rough transparent mask, and α gt represents the first-stage target transparent mask, and ▽ represents the Sobel gradient operator.
7. The multi-task learning-based portrait matting method according to claim 1, characterized in that: In the step S5, the residual loss L res is defined as follows: L res =(r-|α p -α gt |) 2 In the above formula, α p represents the first-stage rough transparent mask, and α gt represents the first-stage target transparent mask, and r represents the first-stage residual probability map.
8. The multi-task learning-based portrait matting method according to claim 1, characterized in that: In the step S6, the total loss L total is defined as follows: L total = λ·L seg +(1 - λ)·(L alp + L res ) In the above formula, L seg represents the segmentation loss, L alp represents the transparency mask loss, L res represents the residual loss, and λ represents the weight of the segmentation loss L seg . (L alp +L res ) represents the matting loss, and (1 - λ) represents the weight of the matting loss (L alp +L res ).
9. The multi-task learning-based portrait matting method according to claim 1, characterized in that: The step S9 specifically includes the following steps: S91. Randomly read the same number of RGB original images, the original transparent masks corresponding to the RGB original images, and the background dataset with the portraits removed from the training set. Combine the read RGB original images and the original transparent masks corresponding to the RGB original images to form training pairs, and perform background augmentation on the training pairs through the background dataset with the portraits removed. Then perform random augmentation on these background-augmented training pairs to obtain the augmented RGB images and the corresponding augmented transparent masks; S92. Scale the augmented RGB image and the corresponding augmented transparency mask to a size of 2048 * 2048 pixels to obtain the second-stage RGB image I for training and the second-stage target transparency mask α gt , and perform morphological operations on the second-stage target transparency mask α gt to obtain the second-stage target segmentation map s gt ; S93. After scaling the second-stage RGB image I to a size of 512 * 512 pixels, input it into the basic matte extraction model to obtain the second-stage rough transparency mask α p , the second-stage residual probability map r and the second-stage hidden feature h.
10. The multi-task learning-based portrait matting method according to claim 1, characterized in that: The step S10 specifically includes the following steps: S101. Scale the second-stage RGB image I, the second-stage rough transparency mask α generated by the basic matte model, p p the second-stage residual probability map r, and the second-stage hidden feature h to 1 / 2 scale of the second-stage RGB image I, splice them together in the channel dimension, then crop a batch of crop blocks at a ratio of 1 / 16 in both the horizontal and vertical directions at this scale, then extract the crop blocks with an average residual probability greater than 0 from the second-stage residual probability map r and denote them as crop blocks P1, and then expand these crop blocks P1 outward by 16 pixels; S102. Input the cropped block P1 into a two-layer convolutional network, then magnify it by 1 times to obtain the Refine intermediate feature. Then, crop the second-stage RGB image I at a ratio of 1 / 16 both horizontally and vertically to obtain the raw cropped block. Then, expand these raw cropped blocks outward by 32 pixels and splice them together with the Refine intermediate feature along the channel dimension, and then pass through two more convolutional networks to obtain the refined transparent mask block α s , and then these refined transparent mask blocks α s . Crop off 32 pixels at the edge to obtain the refined transparent mask block α q ; S103. Enlarge the second-stage rough transparent mask α generated by the basic matte model directly to the size of the second-stage RGB image I as the second-stage rough transparent mask α p and paste the refined transparent mask block α obtained in step S102 b back onto the second-stage rough transparent mask α q to obtain the final refined mask image α b ; p ; S104. Use the transparency mask loss L alp to supervise the training of the refined mask graph α p .
Citation Information
Patent Citations
A depth-learning automatic image matting method guided by semantic segmentation information
CN109035253A
Natural image matting method based on deep learning
CN111161277A