Character matting method for virtual production based on multi-scale receptive field feature enhancement fusion

By introducing a receptive field module combined with attention and feature semantic enhancement module in the MODNet model, the problem of unsatisfactory cutout effect in complex backgrounds is solved, and higher accuracy and stability are achieved.

CN119515899BActive Publication Date: 2025-05-13LIANZIXIN INTELLIGENT TECHNOLOGY (HANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510084279.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing character cutout model is not ideal when dealing with complex backgrounds or images with rich details, especially in complex details such as hair and tulle, which makes it difficult to achieve accurate separation, resulting in loss or unnatural performance, affecting the reality of subsequent synthetic pictures.

Method used

Based on the MODNet structure, a receptive field module and feature semantic enhancement module combining attention are introduced to expand the receptive field and enhance the extraction and fusion of feature semantic information, ensuring that detailed information is captured and semantic integrity is maintained in complex backgrounds.

Benefits of technology

It significantly improves the model's performance in feature extraction, processing and fusion, can capture image details and semantic information more accurately, improves the accuracy and stability of the cutout effect, and performs well in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119515899B_ABST
    Figure CN119515899B_ABST
Patent Text Reader

Abstract

The invention discloses a virtual production character matting method based on multi-scale receptive field feature enhancement and fusion, and belongs to the field of image processing technology. The method first constructs a MODNet model including a semantic estimation module, a detail prediction module and a semantic-detail fusion module, and introduces a receptive field module combined with attention and a feature semantic enhancement module. Among them, the semantic estimation module extracts features from the input image through a backbone network, and then adjusts the weights of the feature maps of the latter two stages through the receptive field module combined with attention, and expands the receptive field of the feature map to obtain semantic features. The detail prediction module encodes the fusion result of the input image and the intermediate scale feature map, fuses the encoding result with the semantic feature through the feature semantic enhancement module, and outputs the detail prediction result after decoding. Finally, the semantic-detail fusion module fuses the semantic features and the decoding results of the detail prediction module, and outputs the predicted portrait mask image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and relates to character matting in virtual production, and specifically to a virtual production character matting method based on multi-scale receptive field feature enhancement fusion. Background Art

[0002] The goal of cutout is to accurately separate the characters from the complex background so that they can be seamlessly synthesized in the virtual scene later. Traditional character cutout methods mainly rely on manually designed features, which can achieve good results when facing simple backgrounds, but when facing complex backgrounds or images with rich details, there are problems such as the need for manual intervention, difficulty in processing complex backgrounds, and strong data dependence. With the advancement of deep learning technology, character cutout models based on deep learning have gradually become a research hotspot. By training the model with a large amount of image data, automatic feature extraction and segmentation can be achieved, significantly improving the accuracy and efficiency of cutouts.

[0003] During the testing phase, traditional character cutout models need to use a three-dimensional map as an auxiliary input. High-quality three-dimensional maps not only require a lot of manpower and time, but may also affect the model cutout effect due to human subjective factors. Secondly, when dealing with complex backgrounds or foregrounds with rich details, the model effect is still not ideal, especially for parts with complex details such as hair and tulle, which are often difficult to separate accurately, resulting in loss or unnatural performance, which in turn affects the realism of the subsequent composite picture.

[0004] In addition, deep learning models are highly dependent on training data. In order to improve the accuracy and robustness of the model, a large amount of diverse data is required for training. Collecting and labeling this data is time-consuming and expensive. Even with sufficient data, the model may still have insufficient generalization problems when facing unseen scenes, resulting in unstable matting effects. In the actual project operation process, high computing costs and hardware requirements may also become limitations for its widespread application.

[0005] In summary, the accuracy and real-time performance of the existing character cutout model still need to be improved. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention proposes a virtual production character matting method based on multi-scale receptive field feature enhancement fusion. On the basis of the MODNet structure, the receptive field is expanded and the feature semantics is enhanced. While maintaining the semantic integrity, the detail information in the image is accurately captured to ensure high-quality matting effects under complex backgrounds, and the lightweight characteristics of the model are maintained to improve the matting efficiency.

[0007] The virtual production character matting method based on multi-scale receptive field feature enhancement fusion specifically includes the following steps:

[0008] Step 1: Collect virtual production scene images including people as training samples. Label the people in the training samples and generate real portrait masks as corresponding training labels.

[0009] Step 2: Build a MODNet model including a semantic estimation module, a detail prediction module, and a semantic-detail fusion module.

[0010] Step 3: Input the sample image in step 1 into the MODNet model constructed in step 2, and extract features through the backbone network in the semantic estimation module to obtain a series of feature maps of different scales. Select the feature maps of the last two stages, first adjust the weights on the channel through the SE (Squeeze-and-Excitation) module, and then input them into multiple parallel dilated convolution branches after fusion to expand the receptive field and capture the multi-level features in the image. Then fuse multiple branches to obtain the semantic feature map. .

[0011] Preferably, the backbone network of the semantic estimation module is any one of MobileNetV2, ResNet, VGG, and SegNet.

[0012] Step 4: Fuse the sample image in step 1 with a feature map of a certain stage output by the backbone network of the semantic estimation module as the encoder input of the detail prediction module. , perform feature semantic enhancement on the encoding result D of the detail prediction module:

[0013]

[0014] Among them, Concat() represents the connection of feature maps in the channel dimension, Normal() represents normalization, ReLU() represents the ReLU function, and Conv represents a 1*1 convolutional layer.

[0015] Semantic enhancement results Decode and output detailed prediction image d p .

[0016] Preferably, the detail prediction module performs no less than 4 convolution operations on the input and outputs an encoding result D.

[0017] Step 5: The semantic-detail fusion module fuses the features output by the semantic estimation module and the detail prediction module to obtain the predicted portrait cutout mask α p .

[0018] Step 6: Predict the details of the prediction p , Portrait cutout mask αp The same as the real portrait cutout mask in step 1 Compare, calculate loss value, and train model parameters.

[0019] Step 7: Input the virtual production image frame containing the person into the model trained in step 6 to obtain the predicted portrait matting mask, thereby completing the person matting.

[0020] The present invention has the following beneficial effects:

[0021] By introducing the receptive field module combined with attention and the feature semantic enhancement module on the basis of traditional MODNet, the improved model has achieved significant performance improvements in feature extraction, processing and fusion, is more accurate in capturing image details and semantic information, and achieves more efficient flow in the interaction of branch information, thus demonstrating excellent performance in various application scenarios, and is significantly better than traditional models in the effect of portrait cutout in virtual production.

[0022] On the one hand, the design concept of the feature semantic enhancement module is to improve the selectivity and expressiveness of features. Through parallel processing and fusion strategies, it can capture more useful feature information without significantly increasing the computational complexity, so that the model can better retain and enhance image details when processing images with rich details, ensuring the accuracy of the final output.

[0023] On the other hand, the receptive field module combined with attention can enhance the multi-scale feature extraction capability of convolutional neural networks, and is particularly suitable for complex image segmentation tasks. By introducing hole convolutions of different scales, this module can expand the receptive field without significantly increasing the amount of computation, and fuse information at different scales to capture rich contextual information. It performs particularly well in complex scenes and can effectively process images with complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Schematic diagram of the improved MODNet model in the embodiment;

[0025] Figure 2 Schematic diagram of a receptive field module combined with attention in an embodiment;

[0026] Figure 3 It is the schematic diagram of SE module;

[0027] Figure 4 Schematic diagram of a feature semantic enhancement module in an embodiment. DETAILED DESCRIPTION

[0028] The present invention will be further explained below with reference to the accompanying drawings;

[0029] The virtual production character matting method based on multi-scale receptive field feature enhancement fusion specifically includes the following steps:

[0030] Step 1. This embodiment uses 9421 images from the P3M-10K dataset and 50 images manually annotated in a virtual production scenario as training set data, and uses images from three subsets of the P3M-10K dataset, P3M-500-P, P3M-500-NP, and PPM-100, as well as 20 images manually annotated in a virtual production scenario as test sets.

[0031] The P3M-10K dataset is a large-scale dataset designed for portrait cutout tasks. It aims to solve the problem of portrait foreground segmentation in complex backgrounds. Each image provides an accurate alpha channel mask to mark the boundary between the portrait and the background. The dataset contains more than 10,000 high-resolution images covering various scenes and lighting conditions such as indoors, outdoors, natural light, and artificial light. It includes not only single portraits but also multi-person portraits, which can increase the complexity of the task and effectively test and verify the robustness and adaptability of the model in different scenarios.

[0032] In addition, manually annotated images from some virtual production scenes are added to provide the model with richer training materials, which can help the model better understand the cutout requirements in a virtual production environment.

[0033] Step 2: Build a MODNet model including a semantic estimation module, a detail prediction module, and a semantic-detail fusion module. The MODNet model is a character matting model that does not require a tripartite map as an auxiliary prediction input. The overall architecture uses multi-level feature extraction and fusion to ensure that while maintaining semantic integrity, details are also accurately retained, taking into account the effectiveness, real-time performance, and lightweight of the model.

[0034] First, the semantic estimation module performs multi-scale feature extraction on the input image I through the backbone network to obtain a series of intermediate feature maps and semantic features of different scales.

[0035] Then, the detail prediction module uses the fusion result of the input image and the intermediate feature map output by the backbone network as the input of the encoder to extract the detail information in the image. The encoding result is then fused with the semantic features and decoded to output the detail prediction result.

[0036] Finally, the semantic-detail fusion module fuses the semantic features and the decoding results of the detail prediction module, outputs the predicted portrait mask image, and completes the foreground separation.

[0037] Step 3: Based on the MODNet model, this application introduces a receptive field module combined with attention and a feature semantic enhancement module in the semantic estimation module and detail prediction module, respectively. Figure 1 As shown:

[0038] s3.1. In this embodiment, MobileNetV2 is used as the backbone network of the semantic estimation module to extract features from the input image I and output a series of feature maps I1~I5 with successively decreasing sizes.

[0039] MobileNetV2 is a lightweight convolutional neural network architecture that greatly reduces the number of parameters and computational complexity through depthwise separable convolutions while maintaining efficient feature extraction capabilities. It is very suitable for application scenarios such as virtual production that require real-time processing and efficient computing, and can run on mobile devices and low-computing environments.

[0040] For the feature maps I4 and I5 output by the backbone network, use Figure 2 As shown in the figure, the receptive field module combined with attention is processed to obtain the semantic feature S (I):

[0041]

[0042]

[0043] Among them, Conv() represents a 1*1 convolutional layer, and Concat() represents the linking of feature maps in the channel dimension. () represents the dilated convolution, M is , C is the number of channels of the input image I. () indicates flattening the feature map according to the channel dimension, () represents upsampling, () represents the global pooling operation, () represents passing through SE module. Figure 3 As shown:

[0044]

[0045] represents global average pooling, , represents the fully connected layer, represents the ReLU activation function, represents the Sigmoid activation function, Represents an element-wise multiplication fully connected layer.

[0046] The receptive field module combined with attention takes the feature maps of the last two stages of the backbone network output as input, first sends the two feature maps to the SE module, performs global average pooling in the SE module, compresses the spatial information of each channel into a single global feature, and summarizes the global information of each channel. Then, the global features are reduced in dimension through a fully connected layer, and the ReLU activation function is applied to introduce nonlinear features. Another fully connected layer is used to restore the features to the original number of channels, and the weight coefficient of each channel is generated by the Sigmoid activation function, which is used to re-weight the channels of the original input feature map and adjust the weight relationship between the channels. In this embodiment, the value of the scaling factor r is set to 16. Finally, through a 1×1 convolution layer, the number of channels of the feature map is reduced to 1 / 5 of the original number of channels C, and the feature map is expanded to d to ensure seamless connection with the subsequent multi-scale hole convolution network layer. Before the feature map is subjected to multi-scale hole convolution processing, the module as a whole presents a right half U-shaped structure, which helps to maintain the integrity and hierarchy of information during feature extraction. In addition, a global pooling layer is introduced after the last feature map to provide richer global context information. This design strategy effectively suppresses the expression of invalid features and ensures that the model can focus on extracting meaningful information.

[0047] The multi-scale atrous convolution network layer includes C / 5 parallel atrous convolution branches, each of which includes M continuous atrous convolutions to capture different spatial information. The output of each atrous convolution branch is compressed by a 1×1 convolution layer to reduce the number of channels. Subsequently, the compressed feature map is linked in the channel dimension to form a feature map containing multi-scale information. After convolution by a 1×1 convolution layer, the semantic feature S (I) is output. By using atrous convolutions of different scales, the receptive field can be expanded without increasing the amount of computation, and the multi-level features in the image can be captured, helping the model to better understand the global structure and detail changes of the image. Since the scenes in virtual production are complex and changeable, the boundaries between the background and the foreground are not obvious in most cases. Therefore, the ability to extract multi-scale features is particularly important to ensure that the model can still maintain high-precision matting effects under complex backgrounds.

[0048] s3.2, The detail prediction module uses the semantic feature map S (I) output by the semantic estimation module through the feature semantic enhancement module to perform feature semantic enhancement on the encoding result D of the detail prediction module, ensuring that the model does not ignore any detail information while understanding the global semantics. The detail prediction module encodes the fusion result of the input image I and the feature map I3 through 13 consecutive convolution operations, focusing on the extraction and processing of high-resolution details. Compared with the semantic estimation module, the detail prediction module has fewer convolution layers, which can reduce the computational complexity. Through the jump connection between the input and output, the impact of the downsampling operation on the detail prediction is reduced, ensuring the stability of the detail prediction.

[0049] like Figure 4 As shown in the figure, the feature semantic enhancement module adopts two parallel input paths, taking the semantic feature map S (I) output by the semantic estimation module and the encoding result D of the detail prediction module as input respectively. In each path, the input feature map is firstly subjected to a 1×1 convolution and normalization process to ensure the consistency of data distribution and promote stable training of the model. Then, through the ReLU activation function, nonlinear transformation is introduced to increase the expressive power of the features. Finally, the results of the two paths are concatenated and sent to a 1*1 convolution layer to adjust the number of channels to obtain the semantic enhancement result. :

[0050]

[0051] Among them, Normal() represents normalization. The parallel processing and fusion strategy of the feature semantic enhancement module can improve the selectivity and expressiveness of features.

[0052] s3.3, the semantic-detail fusion module adopts the feature splicing method to fuse the global context information in the semantic feature map and the high-resolution edge features output by the detail prediction module, establish a synergistic relationship between semantics and details, ensure that while maintaining the global semantic understanding, the key detail information is not lost, enhance the recognition effect of subtle structures, achieve high-precision image segmentation in a variety of application scenarios, and generate the final predicted portrait cutout mask α p , achieving precise foreground separation.

[0053] Step 4: According to the output characteristics of the semantic estimation module, detail prediction module and semantic-detail fusion module in the model, a special loss function is set to achieve optimization at multiple levels, thereby improving the overall prediction ability of the model.

[0054] s4.1. In the semantic estimation module, the real portrait is masked Perform 16 times downsampling and Gaussian blur processing as supervision signal , using L2 loss The sum of squares of the errors between the predicted value and the true value is calculated to evaluate the difference, so that the semantic mask obtained by the semantic estimation module Smoother overall:

[0055]

[0056] in, is the L2 loss function. Downscaling and Gaussian blurring can reduce the influence of details on the coarse semantic mask.

[0057] s4.2, in the detail prediction module, through the L1 loss function Detailed prediction map To conduct supervision:

[0058]

[0059]

[0060] in, is the L1 loss function, Represents corrosion, Represents expansion, Represents a real portrait cutout mask The mask after dilation and corrosion, Used to supervise the detail prediction branch to highlight the boundary area.

[0061] s4.3. Use real portrait cutout mask in semantic-detail fusion module To guide the final predicted portrait cutout mask α p Generation, loss function of the semantic-detail fusion module for:

[0062]

[0063] in, Represents the predicted portrait cutout mask α p The pixel value of the i-th pixel, the real portrait cutout mask The pixel value of is the total number of pixels in the image. is a constant, which is set to 10 in this embodiment. -6 .

[0064] s4.4. Combining the loss functions of the three modules by weighted summation ensures that each branch can work together, and the model is optimized holistically and effectively at different levels, thus improving the overall performance and efficiency of the model. The final loss function for:

[0065]

[0066] in, , and is the weight value of the loss function of each module, which is used to balance the impact of the loss of each branch. In this embodiment, they are 10, 10, and 1 respectively.

[0067] Step 5: Use the training set to train the model, and then use the test set data to evaluate the cutout effect of the trained model. In the full-body scene, this method can accurately separate the person from the background. In the half-body scene, this method can also maintain the clarity of the edges under complex background conditions, and can finely process the details of the person, especially in complex areas such as the head and shoulders, the segmentation effect is still natural and smooth.

[0068] Comparing the cutout results of this method and the commonly used models DIM, GFM, MODNet, SHM, and LFM in the prior art, experiments were conducted using multiple groups of pictures with increasing fit between the person and the background and color similarity. In high-similarity scenes, the performance of SHM and LFM is relatively unsatisfactory. SHM almost completely cuts out the foreground content, and the details are roughly processed. DIM has excellent overall performance due to its reliance on additional triplicated image input, but it is slightly insufficient when processing the background of the gap between the person. GFM performs well in detail processing, but there is still much room for improvement in the overall effect and background distinction in high-similarity images. This method and the improved MODNet both perform well in cutting out the main body of the person, but MODNet is slightly insufficient in the details of the arms and feet, and still lacks accurate judgment in images with higher difficulty. After introducing the receptive field module combined with attention and the feature semantic enhancement module, this method can maintain high accuracy in complex backgrounds.

Claims

1. A virtual production character matting method based on multi-scale receptive field feature enhancement fusion is constructed. A MODNet model including a semantic estimation module, a detail prediction module and a semantic-detail fusion module is constructed. The virtual production scene images are used as training data, and the real portrait mask α of the characters in the image is used as the training data. g As the corresponding training label, the MODNet model is trained to output the predicted portrait cutout mask α p , complete the character matting of virtual production, which is characterized by: The receptive field module combined with attention and the feature semantic enhancement module are introduced into the semantic estimation module and detail prediction module of the MODNet model respectively, and the improved MODNet model is trained; The backbone network of the semantic estimation module is MobileNetV2, which performs multi-scale feature extraction on the input image I to obtain a series of feature maps I1-I5 with successively decreasing sizes. The receptive field module combined with attention processes I4 and I5 and outputs a semantic feature map S(I): S(I)=Conv(Concat(Conv(Astrous Conv(d)*M)*C / 5)); d=flat(Conv(Concat(SE(I4),Up(Concat(SE(I5),UP(I5))))))); Among them, Conv() represents a 1*1 convolution layer, Concat() represents the linking of feature maps in the channel dimension, AstrousConv() represents dilated convolution, M is the number of Astrous Convs, and C is the number of channels of the input image I; flat() represents flattening the feature map according to the channel dimension, Up() represents upsampling, GAP() represents global pooling operation, and SE() represents passing through the SE module; The detail prediction module takes the input image I and the intermediate feature map I of the i-th stage output by the backbone network i The fusion result of is used as the input of the encoder, and then the encoding result D is fused and enhanced with the semantic feature map S(I) through the feature semantic enhancement module, the enhanced result g is decoded, and the detailed prediction map d is output. p ; The receptive field module combined with attention takes the feature maps of the last two stages output by the backbone network in the semantic estimation module as input, first adjusts the attention weights of the two feature maps on the channel through the SE module, and then expands the receptive field through multi-scale hole convolution after fusion, as the semantic feature map S(I) output by the semantic estimation module; The semantic-detail fusion module fuses the semantic feature map S(I) and the decoding result of the detail prediction module, and outputs the predicted portrait mask image α p , complete foreground separation; The feature semantic enhancement module sequentially performs convolution, normalization and activation operations on the encoding result D output by the encoder in the detail prediction module, and performs the same operations on the semantic feature map S(I), and then links them in the channel dimension, and then performs convolution to complete the semantic enhancement of the semantic feature to the detail feature, and obtains the enhanced result g; the enhanced result g is used to decode and generate the detail prediction map d p .

2. The method for virtual production character matting based on multi-scale receptive field feature enhancement fusion as claimed in claim 1, characterized in that: The backbone network of the semantic estimation module is any one of MobileNetV2, ResNet, VGG, and SegNet.

3. The method for virtual production character matting based on multi-scale receptive field feature enhancement fusion as claimed in claim 1, characterized in that: The scaling factor of the SE module is set to 16. The SE module first performs global average pooling on the input feature map, then reduces the dimension of the global features through a fully connected layer, applies the ReLU activation function, and then restores the features to the original number of channels through another fully connected layer. The weight coefficient of each channel is generated through the Sigmoid activation function, which is used to reweight the channels of the original input feature map.

4. The method for virtual production character matting based on multi-scale receptive field feature enhancement fusion as claimed in claim 1, characterized in that: The detail prediction module performs no less than 4 convolution operations on the fusion result of the input image I and the feature map I3, and outputs the encoding result D.

5. The method for virtual production character matting based on multi-scale receptive field feature enhancement fusion as claimed in claim 1, characterized in that: The predicted details are predicted in Figure d p , Portrait cutout mask α p With real portrait cutout mask α g Compare, calculate loss value, and train model parameters.

6. The method for virtual production character matting based on multi-scale receptive field feature enhancement fusion as claimed in claim 5, characterized in that: Setting the total loss function for: in, Represent the loss functions of the semantic estimation module, detail prediction module and semantic-detail fusion module respectively, λ s , d and λ α is the weight value of the loss function of each module.

7. The method for virtual production character matting based on multi-scale receptive field feature enhancement fusion as claimed in claim 6, characterized in that: m d =dilate-erode(a g ); Among them, S p represents the semantic mask obtained by the semantic estimation module, G(α g ) represents the real portrait cutout mask α g The supervisory signal G(α g ); is the L2 loss function; is the L1 loss function, erode represents corrosion, dilate represents expansion, and m d Represents the real portrait cutout mask α g The mask after dilation and corrosion; Represents the predicted portrait cutout mask α p The pixel value of the i-th pixel, the real portrait cutout mask α g The pixel value of , n is the total number of pixels in the image; ∈ is a constant.

8. The method for virtual production character matting based on multi-scale receptive field feature enhancement fusion as claimed in claim 6, characterized in that: Setting up the Lambda s , d and λ α The values ​​of yes are 10, 10, 1 respectively.

Citation Information

Patent Citations

  • Image matting method based on complex background

    CN118155206A

  • Remote sensing image semantic segmentation method based on category interactive attention and perception fusion

    CN118736231A