Real-time human video extinction model training method
By introducing high and low resolution reversible branches and feature encoding modules in the video extinction model, the problem of poor fine area clamping effect in the prior art is solved, and the detailed information and semantic information in the video airspace is fully extracted, and the keying effect is improved.
Patent Information
- Application Number
- CN202510057371.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
The existing video clamping technology still has room for improvement in the clamping effect of the detailed area. Since the model architecture is designed to ensure high efficiency, the lightweight MobileNetV3-Large is used as the backbone network, resulting in insufficient shallow detail information and deep semantic information.
A training method for real-time human video extinction model is proposed. By collecting application data, a training data set is generated, and an encoder, decoder, and semantic and detailed prediction high and low resolution reversible branches are introduced into the preset initial model, and the model parameters are adjusted until the loss function value is lower than the preset threshold.
Without increasing the model inference scale, the encoder fully extracts shallow detail information and depth semantic information in the video airspace, and improves the clamping effect of the video extinction model.
Smart Images

Figure CN119990238A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a method for training a real-time human video matting model. Background Art
[0002] Matting refers to the process of predicting the alpha matting and foreground color from an image. In human portrait matting, the goal is to predict the probability of the portrait foreground from the spatial dimension of the image. In video matting (also known as video matting or video synthesis), it involves predicting the alpha matting of consecutive video frames in the time domain. This technology has a wide range of applications in video conferencing, virtual reality and other fields.
[0003] Among the existing video matting technologies, Robust Video Matting (RVM) can perform matting without auxiliary information. RVM realizes learnable time-domain information processing by introducing recursive neural networks and supports long-distance information dependency modeling, but there is still much room for improvement in the matting effect of detail areas. At the same time, in order to ensure high efficiency in the architecture design, RVM uses lightweight MobileNetV3-Large as the backbone network, which leads to insufficient shallow detail information and deep semantic information extracted by the RVM model. Summary of the invention
[0004] In order to solve the above problems, this application proposes a real-time human video matting model training method, wherein the method includes:
[0005] Collect application data and generate a training data set based on the application data; input the training data set into a preset initial model to obtain an output set corresponding to the training data set under the current model parameters; the preset initial model is provided with an encoder, a decoder, and semantic and detail prediction high- and low-resolution reversible branches, the semantic and detail prediction high- and low-resolution reversible branches are only used to train the preset initial model, and are connected to the encoder to output high-resolution detail prediction output and low-resolution semantic estimation output; based on the output set, determine the corresponding loss function value under the current model parameters; adjust the model parameters in the preset initial model until the loss function value is lower than a preset threshold.
[0006] In one example, the encoder includes five feature encoding modules and a dilated spatial convolution pooling pyramid module; the five feature encoding modules are respectively composed of two BNeck submodules, two BNeck submodules, three BNeck submodules, six BNeck submodules, and four BNeck submodules in series, and each BNeck submodule has the same structure, which is composed of an inverse residual structure with a linear bottleneck, a depth-separable convolution, a squeeze-induced attention, and an activation function in series; the semantic and detail prediction high- and low-resolution reversible branch includes two feature encoding modules, two standard convolution blocks, three CBLinear modules, two The invention relates to a novel method for decoding a multi-layer convolutional neural network (MCNN) based on the invention. The method comprises a CBFuse module, two re-parameterized convolutional networks with cross-stage local links and efficient layer aggregation networks, and two DeConV modules. The CBLinear is constructed in series by a standard convolutional layer, a batch normalization layer, a SiLU activation function layer, and a channel dimension segmentation layer; the decoder comprises a convolutional gated recurrent unit, a bilinear interpolation module, four feature decoding modules, and a convolutional batch normalization rectified linear unit activation function module; the convolutional gated recurrent unit is used to store short-term memory information; the feature decoding module comprises a convolutional layer, a batch normalization layer, a rectified linear unit activation function, a convolutional gated recurrent unit, and a bilinear interpolation module.
[0007] In one example, the training data set is input into a preset initial model to obtain an output set corresponding to the training data set under the current model parameters, specifically including: inputting the input image of the training data set into the preset initial model, so that the input image passes through a copy module and a standard convolution downsampling module to obtain a high-resolution feature map; inputting the input image into the encoder to obtain multiple intermediate feature maps and a deep feature map; inputting the input image and the multiple intermediate feature maps into the semantic and detail prediction high- and low-resolution reversible branches to obtain a high-resolution detail prediction output and a low-resolution semantic estimation output output by the semantic and detail prediction high- and low-resolution reversible branches; passing the multiple intermediate feature maps and the deep feature maps through the decoder, the convolution batch normalization corrected linear unit activation function module, and the standard convolution block in sequence to obtain a depth-guided filtering intermediate feature map output by the decoder; inputting the depth-guided filtering intermediate feature map and the high-resolution feature map into the depth-guided filtering layer to obtain the model output.
[0008] In one example, the step of inputting the input image into the encoder to obtain a plurality of intermediate feature maps and a deep feature map specifically includes: inputting the input image into the encoder so that the input image sequentially passes through five feature encoding modules and the atrous spatial convolution pooling pyramid module to obtain a plurality of intermediate feature maps and a deep feature map output by the encoder; the plurality of intermediate feature maps include a first intermediate feature map output by a first feature encoding module, a second intermediate feature map output by a second feature encoding module, a third intermediate feature map output by a third feature encoding module, a fourth intermediate feature map output by a fourth feature encoding module, and a fifth intermediate feature map output by a fifth feature encoding module.
[0009] In one example, the input image and the multiple intermediate feature maps are input into the semantic and detail prediction high- and low-resolution reversible branch to obtain a high-resolution detail prediction output and a low-resolution semantic estimation output output by the semantic and detail prediction high- and low-resolution reversible branch, specifically including: inputting the input image into the semantic and detail prediction high- and low-resolution reversible branch so that the input image passes through two feature encoding modules and a first standard convolution block in sequence to obtain a downsampled feature map; inputting the multiple intermediate feature maps into the semantic and detail prediction high- and low-resolution reversible branch so that the third intermediate feature map passes through a first CBLinear module to obtain a first feature map group, the fourth intermediate feature map passes through a second CBLinear module to obtain a second feature map group, and the fifth intermediate feature map passes through a first After three CBLinear modules, the third feature map group is obtained; the first feature map group, the third feature map group and the downsampled feature map are successively subjected to the first CBFuse module, the first heavily parameterized convolutional network cross-stage local link efficient layer aggregation network, and the first DeConV module to obtain a high-resolution detail prediction output; the first feature map group, the third feature map group and the downsampled feature map are successively subjected to the first CBFuse module, the first heavily parameterized convolutional network cross-stage local link efficient layer aggregation network, and the second standard convolution block to obtain the sixth intermediate feature map; the second feature map group, the third feature map group and the sixth intermediate feature map are successively subjected to the second CBFuse module, the second heavily parameterized convolutional network cross-stage local link efficient layer aggregation network, and the second DeConV module to obtain a low-resolution semantic estimation output.
[0010] In one example, after the multiple intermediate feature maps and the deep feature map are sequentially passed through the decoder, the convolutional batch normalization rectified linear unit activation function module, and the standard convolution block, a deep guided filtered intermediate feature map output by the decoder is obtained, which specifically includes: inputting the deep feature map into a convolutional gated recurrent unit to obtain a first intermediate deep feature map; juxtaposing the deep feature map with the first intermediate deep feature map to form an identity mapping residual module to obtain a second intermediate deep feature map with short-term memory information; passing the second intermediate deep feature map through a bilinear interpolation module and performing double upsampling to obtain a third intermediate deep feature map; juxtaposing the third intermediate deep feature map with the fifth intermediate feature map and passing through the first A feature decoding module is used to obtain a seventh intermediate feature map; the seventh intermediate feature map is juxtaposed with the fourth intermediate feature map, and passed through a second feature decoding module to obtain an eighth intermediate feature map; the eighth intermediate feature map is juxtaposed with the third intermediate feature map, and passed through a third feature decoding module to obtain a ninth intermediate feature map; the ninth intermediate feature map is juxtaposed with the second intermediate feature map, and passed through a fourth feature decoding module to obtain a tenth intermediate feature map; the tenth intermediate feature map is juxtaposed with the first intermediate feature map, and passed through two convolution batch normalization rectified linear unit activation function modules to obtain an eleventh intermediate feature map; the eleventh intermediate feature map is subjected to standard convolution to obtain a depth-guided filtering intermediate feature map.
[0011] In one example, determining the corresponding loss function value under the current model parameters based on the output set specifically includes: determining the high-resolution detail prediction loss by the following formula and the high-resolution detail prediction output: d =(dilate(a g )-eroad(a g ))||d p -a g ||1; where L d Predicting loss for high-resolution details; a g is the real supervised alpha mask; dilate and eroad are mathematical morphological processing, representing dilation and erosion respectively, and the parameters are default values; d p Predict output for high-resolution details; ||d p -a g ||1 means d p -a g The L1 norm of the low-resolution semantic estimation is determined by the following formula and the low-resolution semantic estimation output: Among them, L s is the low-resolution semantic estimation loss, sp is the low-resolution semantic estimation output; G(a g ) is for a g Gaussian blur processing is performed with default parameters. The alpha binary pixel loss and RGB color loss are determined by the following formula and the model output: in, is the alpha binary pixel loss, i represents the pixel index after Flatten, a p Indicates the Alpha mask after DGF; To predict the value of the one-dimensional array index position i after the alpha mask is flattened into a one-dimensional array according to the row; is the value of the one-dimensional array index position i after the real Alpha mask is flattened into a one-dimensional array according to the row; ∈ is a hyperparameter with a default value of 10 -6 ; is RGB color loss, j represents the channel index, c p represents the foreground image after DGF, c g represents the real image, Represents a single-channel two-dimensional array of the predicted foreground image obtained according to channel dimension j, Represents a single-channel two-dimensional array of the real foreground image obtained according to the channel dimension j; the corresponding loss function value under the current model parameters is determined by the following formula, as well as the high-resolution detail prediction loss, low-resolution semantic estimation loss, alpha binary pixel loss and RGB color loss: L = σ l ·[w l ·L a +(1-w l )L c ]+σ d ·L d +σ s ·L s ; Where L is the loss function value, w l is a hyperparameter with a default value of 0.6, which is used to balance L a and L c The loss constraint, w l ·L a +(1-w l )L c Constituent loss; σ l , σ d and σ s is a hyperparameter used to balance component loss, low-resolution semantic estimation loss, and high-resolution detail prediction loss. The default value is σ l =σ d =10,σ s =3.
[0012] In one example, the adjusting of the model parameters in the preset initial model until the loss function value is lower than a preset threshold specifically includes: using a training data set for feedforward inference, and performing reverse chain derivation and Adam parameter optimization on the model parameters based on the loss function value and the BP algorithm, until the loss of the preset initial model on the validation set fluctuates less than a preset threshold for ten consecutive generations, thereby obtaining a final training model.
[0013] In one example, the collecting of application data specifically includes: determining a first preset number of exhibition halls as background material, and a second preset number of employees as foreground material; determining a plurality of preset employee combinations among the second preset number of employees; shooting a third preset number of video materials for the plurality of preset employee combinations in the exhibition hall to construct a first application data set; shooting a fourth number of dynamic background videos without foreground in the exhibition hall to construct a second application data set; and shooting a fifth preset number of video materials for the plurality of preset employee combinations in a green screen environment to construct a third application data set.
[0014] In one example, generating a training data set based on the application data specifically includes: using an open source video extinction model to automatically extinct a first application data set to obtain a first extinction data set, wherein the first extinction data set includes foreground data and first Alpha extinction data corresponding to the first data set; using an open source video extinction model to automatically extinct a third application data set to obtain third Alpha extinction data corresponding to the third application data set; using an open source annotation tool to verify the first Alpha extinction data and the third Alpha extinction data to obtain a first verification data set and a third verification data set; synthesizing the third verification data set and the second application data set through a foreground and background synthesis data script to obtain an intermediate data set; merging the intermediate data set, the first verification data set and the PPM open source data set as an initial data set; and in the initial data set, dividing it according to a preset ratio to obtain a training data set, a verification data set and a test data set.
[0015] The method proposed in this application can bring the following beneficial effects: using reversible branch-assisted model feature extraction technology, introducing high-resolution detail prediction and low-resolution semantic estimation reversible branches in the RVM encoding stage, which only participate in the training process of the model, without increasing the scale of model inference, the encoder can fully extract shallow detail information and deep semantic information in the video spatial domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0017] Figure 1 A flowchart of a method for training a real-time human video extinction model in an embodiment of the present application;
[0018] Figure 2 Schematic diagram of the architecture of a preset initial model in an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0020] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0021] Figure 1 A flowchart of a method for training a real-time human video matting model is provided for one or more embodiments of this specification. The process can be executed by a computing device in the corresponding field, and some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.
[0022] The analysis method involved in the embodiments of the present application can be implemented by a terminal device or a server, and the present application does not impose any special restrictions on this. For the convenience of understanding and description, the following embodiments are described in detail by taking a server as an example.
[0023] It should be noted that the server can be a single device or a system composed of multiple devices, that is, a distributed server, and this application does not make any specific limitations on this.
[0024] like Figure 1 As shown, the embodiment of the present application provides a method, comprising:
[0025] S101: Collect application data, and generate a training data set based on the application data.
[0026] When collecting application data, since the open source RVM uses foreground and background synthetic data to train the model, there are inconsistencies in brightness and lighting. Although a semantic segmentation task branch is added and multi-task training is performed based on the real data set for this task, it is still impossible to effectively avoid the overfitting of the open source RVM to the synthetic scene. In order to effectively avoid the shortcomings of the RVM training data set, this application collects and annotates high-definition, high-quality application data, and the data distribution of this data is consistent with the distribution of real application data. The specific steps for collecting application data include:
[0027] S101-1: Formulate data collection rules: Use foreground and background synthetic data to create application data sets. You only need to collect a large number of background videos and a small number of videos with people (foreground). Since production environment application data is mostly distributed in scenes with dense and mobile personnel such as exhibitions and exhibition halls, a preset number (such as 10) of exhibition and exhibition hall scenes are selected as video collection environments to ensure the consistency of data distribution. A preset number (such as 20) of employees are selected as foreground materials, and there are no special requirements for employee clothing, hair accessories, etc.
[0028] S101-2: Video collection with foreground and background: Calculate the combination set of 1 to 19 people selected from 20 employees, shuffle the set and then randomly select 10 combinations as foreground materials. Combined with 10 backgrounds, a total of 100 videos with foreground and background need to be shot, each video is 10 to 20 seconds long, to construct the first application data set.
[0029] S101-3: Dynamic background video acquisition: 20 dynamic background videos without foreground are shot for 10 exhibitions and exhibition halls, each video is 10 to 20 seconds long, to construct the second application data set.
[0030] S101-4: Green screen and foreground video collection: Under the green screen environment, calculate the combination set of 1 to 19 people selected from 20 employees, shuffle the set and then randomly select 20 combinations as foreground materials based on the index value. Combined with the green screen background, a total of 20 videos with foreground and background need to be shot, each video lasting 10 to 20 seconds, to construct the third application data set.
[0031] After collecting application data, it is necessary to create a training data set for training the model. In the existing technology, although the open source RVM uses motion and time data enhancement, the enhancement strategy is still insufficient. At the same time, the background image has a different data distribution from the real scene used. First, based on the foreground and alpha extinction open source data sets of VideoMatte240K, Distinctions-646, Adobe Image Matting, and the background open source data sets of DVM and BGMV2, synthesis processing is performed to generate a training data set. The specific steps include:
[0032] S101-5. Use the MODNet open source video extinction model to automatically extinct the first application data set to obtain a first extinction data set, where the first extinction data set includes foreground data and first Alpha extinction data corresponding to the first data set.
[0033] S101-6. Use the MODNet open source video extinction model to automatically perform extinction processing on the third application data set to obtain third Alpha extinction data corresponding to the third application data set.
[0034] S101-7, using the open source labeling tool labelme, manually verify the extinction data obtained in steps S101-5 and S101-6 to obtain a verified first verification data set and a third verification data set, respectively. The manual verification here is mainly used to verify the correctness of the extinction, which is a prior art.
[0035] S101-8. Use Python to write a foreground and background synthesis data script to synthesize the third verification data set and the second application data set. The synthesized data set, the first verification data set and the PPM open source data set together constitute the initial data set. After the initial data set is generated, a fine-tuning data set can be generated based on the processed initial data set and the PPM open source data set to avoid overfitting of the model to the synthetic scene. Then, the initial data set is split according to a preset ratio (such as 8:1:1) to obtain a training data set, a verification data set and a test data set.
[0036] S102: Input the training data set into a preset initial model to obtain an output set corresponding to the training data set under current model parameters.
[0037] like Figure 2 As shown in the figure, the preset initial model to be trained mainly includes an encoder, a decoder, and semantic and detail prediction high and low resolution reversible branches. The semantic and detail prediction high and low resolution reversible branches are only used to train the preset initial model, which can enable the encoder to fully extract shallow detail information and deep semantic information in the video spatial domain without increasing the model inference scale. The encoder includes five feature encoding modules ( Figure 2 Numbered 2, 3, 4, 5, 6) and a dilated spatial convolutional pooling pyramid module ( Figure 2 The five feature encoding modules are respectively composed of two BNeck sub-modules, two BNeck sub-modules, three BNeck sub-modules, six BNeck sub-modules, and four BNeck sub-modules in series, and each BNeck sub-module has the same structure, which is composed of an inverse residual structure with a linear bottleneck, a depth-wise separable convolution, a squeeze-stimulated attention, and an activation function in series.
[0038] The semantic and detail prediction high- and low-resolution reversible branches include two feature encoding modules: Figure 2 Numbered 22, 23), two standard convolutional blocks ( Figure 2 Numbered 24, 31), three CBLinear modules ( Figure 2 Numbered 25, 26, 27), two CBFuse modules ( Figure 2 28, 32), two heavily parameterized convolutional networks with cross-stage local links and efficient layer aggregation networks ( Figure 2 29, 33), two DeConV modules ( Figure 2 CBLinear is composed of a standard convolutional layer, a batch normalization layer, a SiLU activation function layer, and a channel dimension segmentation layer connected in series.
[0039] The decoder consists of a convolutional gated recurrent unit ( Figure 2 Numbered 8), bilinear interpolation module ( Figure 2 9), four feature decoding modules ( Figure 2 10, 12, 14, 16) and a convolutional batch normalization rectified linear unit activation function module ( Figure 2 The convolutional gated recurrent unit is used to store short-term memory information. The feature decoding module is as follows Figure 2 As shown in the lower left corner, it includes convolutional layer, batch normalization layer, rectified linear unit activation function, convolutional gated recurrent unit, and bilinear interpolation module.
[0040] In one embodiment, in order to obtain an output set, the input image of the training data set needs to be input into a preset initial model so that the input image passes through a copy module and a standard convolution downsampling module to obtain a high-resolution feature map. Then the input image is input into the encoder to obtain multiple intermediate feature maps and deep feature maps. Then the input image and multiple intermediate feature maps are input into the semantic and detail prediction high- and low-resolution reversible branches to obtain the high-resolution detail prediction output and low-resolution semantic estimation output output of the semantic and detail prediction high- and low-resolution reversible branches. At the same time, multiple intermediate feature maps and deep feature maps are sequentially passed through the decoder, the convolution batch normalization corrected linear unit activation function module, and the standard convolution block to obtain the deep guided filtering intermediate feature map output by the decoder. Finally, the deep guided filtering intermediate feature map and the high-resolution feature map are input into the deep guided filtering layer to obtain the model output.
[0041] The input image of the training data set is input into the preset initial model, so that the input image passes through the replication module and the standard convolution downsampling module to obtain a high-resolution feature map. Figure 2 The specific manifestations are:
[0042] S102-1: The input image Img passes through the copy module Silence numbered 0 and the standard convolution down-sampling module Down Sample numbered 1 in sequence to obtain a high-resolution feature map
[0043] Specifically, when an image is input into the encoder, the input image needs to be input into the encoder so that the input image passes through five feature encoding modules and a dilated spatial convolutional pooling pyramid module in sequence to obtain a plurality of intermediate feature maps and a deep feature map output by the encoder. The plurality of intermediate feature maps here include a first intermediate feature map output by the first feature encoding module, a second intermediate feature map output by the second feature encoding module, a third intermediate feature map output by the third feature encoding module, a fourth intermediate feature map output by the fourth feature encoding module, and a fifth intermediate feature map output by the fifth feature encoding module.
[0044] The above process of inputting the image into the encoder is Figure 2 It is manifested as:
[0045] S102-2: Img passes through the feature encoding modules numbered 2, 3, 4, 5, and 6 in sequence to obtain feature maps from shallow to deep (first intermediate feature map), (Second intermediate feature map), (Third intermediate feature map), (Fourth intermediate characteristic map), (Fifth intermediate feature map), in which the shallow high-resolution features are rich in detail information, while the deep low-resolution features are rich in semantic information.
[0046] S102-3: FM6 passes through the hole spatial convolution pooling pyramid module ASPP numbered 7 to obtain a deep feature map (Deep feature map). Specifically, the Encoder Blocks numbered 2, 3, 4, 5, and 6 in step S102-2 are composed of 2, 2, 3, 6, and 4 BNeck modules in series, and each BNeck module has a consistent network structure, which is composed of an inverse residual structure with a linear bottleneck, a depth-separable convolution, squeezed attention, and an activation function in series. The ASPP here is a combination of a dilated convolution and a spatial pooling pyramid. The overall feedforward process of step S102-2 includes:
[0047] S102-2-1: Img is used as the input of the first BNeck module. On the inverse residual structure with a linear bottleneck, a 1×1 convolution is used, and the number of channels remains unchanged at 16, resulting in With residual edges, jump links to activation function layers. FM 2-1 As the input of the depthwise separable convolution, a 1x1 convolution is performed to keep the channel dimension unchanged, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 2-2 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 2-4 and FM 2-2 The product of FM 2-5 As the input of the h-swish activation function, we get Juxtaposition FM 2-6-1 And Img and bring it into 1x1 convolution to reduce the channel dimension, and calculate the output of the first BNeck module
[0048] S102-2-2: FM 2-6 As the input of the second BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 16 to 64, resulting in With residual edges, jump links to activation function layers. FM 2-7 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 24, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 2-8 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 2-10 and FM 2-8 The product of FM 2-11 As the input of the h-swish activation function, we get Juxtaposition FM 2-12 and FM 2-6 And bring in 3x3 convolution to reduce the spatial and channel dimensions, and calculate the output of the second BNeck module
[0049] S102-2-3: FM2 is used as the input of the third BNeck module. On the inverse residual structure with a linear bottleneck, the dimension is expanded using 1×1 convolution, and the number of channels is expanded from 24 to 72, resulting in With residual edges, jump links to activation function layers. FM 3-1 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 24, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 3-2 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 3-4 and FM 3-2 The product of FM 3-5 As the input of the h-swish activation function, we get Juxtaposition FM 3-6-1 And FM2 and bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the third BNeck module
[0050] S102-2-4: FM 3-6 As the input of the fourth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 16 to 72, resulting in With residual edges, jump links to activation function layers. FM 3-7 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 40, and then a 3x3 convolution is performed to keep the spatial dimension unchanged. The calculation is FM 3-8 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 3-10 and FM 3-8 The product of FM 3-11 As the input of the h-swish activation function, we get Juxtaposition FM 3-12 and FM 3-6 And bring in 3x3 convolution to reduce the spatial and channel dimensions, and calculate the output of the fourth BNeck module
[0051] S102-2-5: FM3 is used as the input of the fifth BNeck module. On the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 40 to 120, resulting in With residual edges, jump links to activation function layers. FM 4-1 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 40, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 4-2 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 4-4 and FM 4-2 The product of FM 4-5 As the input of the h-swish activation function, we get Juxtaposition FM 4-6-1 And FM3 and bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the fifth BNeck module
[0052] S102-2-6: FM 4-6 As the input of the sixth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 40 to 120, resulting in With residual edges, jump links to activation function layers. FM 4-7 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 40, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 4-8 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 4-10 and FM 4-8 The product of FM 4-11 As the input of the h-swish activation function, we get Juxtaposition FM 4-12-1 and FM 4-6 And bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the sixth BNeck module
[0053] S102-2-7: FM 4-12As the input of the seventh BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 40 to 240, resulting in With residual edges, jump links to activation function layers. FM 4-13 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 80, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 4-14 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 4-16 and FM 4-14 The product of FM 4-17 As the input of the h-swish activation function, we get Juxtaposition FM 4-18 and FM 4-12 And bring in 3x3 convolution to reduce the spatial and channel dimensions, and calculate the output of the seventh BNeck module
[0054] S102-2-8: FM4 is used as the input of the eighth BNeck module. On the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 80 to 200, resulting in With residual edges, jump links to activation function layers. FM 5-1 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 80, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 5-2 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 5-4 and FM 5-2 The product of FM 5-5 As the input of the h-swish activation function, we get Juxtaposition FM 5-6-1 And FM4 and bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the eighth BNeck module
[0055] S102-2-9: FM 5-6As the input of the ninth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 80 to 184, resulting in With residual edges, jump links to activation function layers. FM 5-7 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 80, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 4-8 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 5-10 and FM 5-8 The product of FM 5-11 As the input of the h-swish activation function, we get Juxtaposition FM 5-12-1 and FM 5-6 And bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the ninth BNeck module
[0056] S102-2-10: FM 5-12 As the input of the tenth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 80 to 184, resulting in With residual edges, jump links to activation function layers. FM 5-13 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 80, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 5-14 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 5-16 and FM 5-14 The product of FM 5-17 As the input of the h-swish activation function, we get Juxtaposition FM 5-18-1 and FM 5-12 And bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the tenth BNeck module
[0057] S102-2-11: FM 5-18As the input of the eleventh BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 80 to 480, resulting in With residual edges, jump links to activation function layers. FM 5-19 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 112, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 5-20 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 5-22 and FM 5-20 The product of FM 5-23 As the input of the h-swish activation function, we get Juxtaposition FM 5-24-1 and FM 5-18 And bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the eleventh BNeck module
[0058] S102-2-12: FM 5-24 As the input of the twelfth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 112 to 672, resulting in With residual edges, jump links to activation function layers. FM 5-25 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 112, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 5-26 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 5-28 and FM 5-26 The product of FM 5-29 As the input of the h-swish activation function, we get Juxtaposition FM 5-30-1 and FM 5-24 And bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the twelfth BNeck module
[0059] S102-2-13: FM 5-30As the input of the thirteenth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 112 to 672, resulting in With residual edges, jump links to activation function layers. FM 5-31 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 160, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 5-32 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 5-34 and FM 5-32 The product of FM 5-35 As the input of the h-swish activation function, we get Juxtaposition FM 5-36 and FM 5-30 And bring in 3x3 convolution to reduce the spatial and channel dimensions, and calculate the output of the thirteenth BNeck module
[0060] S102-2-14: FM5 is used as the input of the fourteenth BNeck module. On the inverse residual structure with a linear bottleneck, the dimension is expanded using 1×1 convolution, and the number of channels is expanded from 160 to 960, resulting in With residual edges, jump links to activation function layers. FM 6-1 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 160, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 6-2 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 6-4 and FM 6-2 The product of FM 6-5 As the input of the h-swish activation function, we get Juxtaposition FM 6-6-1 And FM4 and bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the fourteenth BNeck module
[0061] S102-2-15: FM 6-6As the input of the fifteenth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 160 to 960, resulting in With residual edges, jump links to activation function layers. FM 6-7 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 160, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 6-8 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 6-10 and FM 6-8 The product of FM 6-11 As the input of the h-swish activation function, we get Juxtaposition FM 6-12-1 and FM 6-6 And bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the fifteenth BNeck module
[0062] S102-2-16: FM 6-12 As the input of the sixteenth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 160 to 1920, resulting in With residual edges, jump links to activation function layers. FM 6-13 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 960, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 6-14 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 6-16 and FM 6-14 The product of FM 6-17 As the input of the h-swish activation function, we get Juxtaposition FM 6-18-1 and FM 6-12 And bring in 1x1 convolution to reduce the channel dimension, and calculate the output of the sixteenth BNeck module
[0063] S102-2-17: FM 6-18As the input of the seventeenth BNeck module, on the inverse residual structure with a linear bottleneck, a 1×1 convolution is used to expand the dimension, and the number of channels is expanded from 960 to 3840, resulting in With residual edges, jump links to activation function layers. FM 6-19 As the input of the depthwise separable convolution, a 1x1 convolution is first performed to reduce the channel dimension to 960, and then a 3x3 convolution is performed to keep the spatial dimension unchanged, resulting in FM 6-20 As the input of the squeeze-stimulate attention module, we first perform global average pooling in the spatial dimension to obtain Then bring it into the 2-layer fully connected network to get Calculate FM 6-22 and FM 6-20 The product of FM 6-23 As the input of the h-swish activation function, we get Juxtaposition FM 6-24 and FM 6-28 And bring in 3x3 convolution to reduce the spatial and channel dimensions, and calculate the output of the seventeenth BNeck module
[0064] In one embodiment, when the input image and the multiple intermediate feature maps are input into the semantic and detail prediction high- and low-resolution reversible branch to obtain a high-resolution detail prediction output and a low-resolution semantic estimation output of the semantic and detail prediction high- and low-resolution reversible branch, the input image needs to be input into the semantic and detail prediction high- and low-resolution reversible branch so that the input image passes through two feature encoding modules and the first standard convolution block in sequence to obtain a downsampled feature map.
[0065] A plurality of intermediate feature maps are input into the high- and low-resolution reversible branches of semantic and detail prediction, so that the third intermediate feature map passes through the first CBLinear module to obtain the first feature map group, the fourth intermediate feature map passes through the second CBLinear module to obtain the second feature map group, and the fifth intermediate feature map passes through the third CBLinear module to obtain the third feature map group; the first feature map group, the third feature map group and the downsampled feature map pass through the first CBFuse module, the first heavily parameterized convolutional network cross-stage local link efficient layer aggregation network, and the first DeConV module in turn to obtain a high-resolution detail prediction output; the first feature map group, the third feature map group and the downsampled feature map pass through the first CBFuse module, the first heavily parameterized convolutional network cross-stage local link efficient layer aggregation network, and the second standard convolution block in turn to obtain the sixth intermediate feature map; the second feature map group, the third feature map group and the sixth intermediate feature map pass through the second CBFuse module, the second heavily parameterized convolutional network cross-stage local link efficient layer aggregation network, and the second DeConV module in turn to obtain a low-resolution semantic estimation output.
[0066] The above process Figure 2 It is reflected in:
[0067] S102-4: Img is sequentially subjected to Encoder Blocks 22 and 23 for feature extraction to obtain a feature map The network structures of Encoder Blocks 22 and 23 are consistent with those of Encoder Blocks 2 and 3, respectively. 23 After being processed by the standard convolution block Conv numbered 24, the downsampled feature map is obtained.
[0068] S102-5: The CB Linear modules numbered 25, 26, and 27 have the same structure, and use FM4, FM5, and FM6 as the input of CBLinear, respectively, to obtain the first feature map group The second feature map group The third feature map group FM fg27 Perform split operation to obtain 2 sub-feature map groups and CB Linear is constructed by connecting a standard convolutional layer, a batch normalization layer, a SiLU activation function layer, and a channel dimension segmentation layer in series.
[0069] S102-6: The CBFuse modules No. 28 and 32 have the same structure. The CBFuse No. 28 receives FM 24 ,FM fg25and FM fg27-1 As input, we get a high-resolution detail prediction feature map
[0070] S102-7: FM 28 The feature maps are obtained by passing through the reparameterized convolutional network No. 29, the cross-stage local link efficient layer aggregation network RepNCSPELAN, and the DeConv No. 30. and detail prediction segmentation matting. 29 The sixth intermediate feature map is obtained by downsampling the Conv numbered 31
[0071] S102-8: CBFuse No. 32 receives FM fg26 ,FM fg27-2 and FM 31 As input, we get a low-resolution semantic estimation feature map FM 32 After passing through RepNCSPELAN numbered 33 and DeConv numbered 34 in sequence, a semantic estimation segmentation matting is obtained.
[0072] Among them, the network structures of RepNCSPELAN and DeConv in step S102-7 and step S102-8 are consistent.
[0073] In one embodiment, after the multiple intermediate feature maps and the deep feature map are sequentially passed through the decoder, the convolutional batch normalization rectified linear unit activation function module, and the standard convolution block, when the deep guided filtering intermediate feature map output by the decoder is obtained, the deep feature map needs to be input into the convolutional gated recurrent unit to obtain a first intermediate deep feature map; the deep feature map and the first intermediate deep feature map are juxtaposed to form an identity mapping residual module to obtain a second intermediate deep feature map with short-term memory information; the second intermediate deep feature map is passed through a bilinear interpolation module and doubled upsampled to obtain a third intermediate deep feature map; the third intermediate deep feature map is juxtaposed with the fifth intermediate feature map and passed through the first feature map. The seventh intermediate feature map is obtained by a feature decoding module; the seventh intermediate feature map is juxtaposed with the fourth intermediate feature map, and passed through a second feature decoding module to obtain an eighth intermediate feature map; the eighth intermediate feature map is juxtaposed with the third intermediate feature map, and passed through a third feature decoding module to obtain a ninth intermediate feature map; the ninth intermediate feature map is juxtaposed with the second intermediate feature map, and passed through a fourth feature decoding module to obtain a tenth intermediate feature map; the tenth intermediate feature map is juxtaposed with the first intermediate feature map, and passed through two convolution batch normalization rectified linear unit activation function modules to obtain an eleventh intermediate feature map; the eleventh intermediate feature map is subjected to standard convolution to obtain a depth-guided filtering intermediate feature map.
[0074] The above process Figure 2 It is reflected in:
[0075] S102-9: FM7 passes through the convolutional gated recurrent unit ConvGRU numbered 8 to store short-term memory information and obtain a deep feature map Then, FM7 and FM8 are juxtaposed to form an identity mapping residual module to avoid gradient vanishing and enable the network to learn more useful information, thus obtaining the second intermediate deep feature map with short-term memory information.
[0076] S102-10: FM input9 After the bilinear interpolation module Bilinear2X numbered 9, the third intermediate deep feature map is obtained by upsampling by 2 times
[0077] S102-11: Concatenate FM9 and FM6 to obtain a characteristic spectrum And input it to the feature decoding module Decoder Block numbered 10 to obtain the seventh intermediate feature map
[0078] S102-12: The network structures of the decoder blocks numbered 10, 12, 14, and 16 are the same, and their input and output processing steps are consistent with step S102-11. The output of the decoder block numbered n∈{10, 12, 14, 16} is the feature map FM n (including the eighth intermediate characteristic map, the ninth intermediate characteristic map, the tenth intermediate characteristic map, and the eleventh intermediate characteristic map).
[0079] S102-13: FM 16 After passing through the convolution batch normalization corrected linear unit activation function modules CBR numbered 18 and 19, the feature map is obtained. About FM input20 Perform standard convolution Conv processing to obtain the depth-guided filtering intermediate feature map It will be used as one of the inputs of the depth guided filter DGF.
[0080] Specifically, the overall structure of the Decoder Block in step S5-3 is as follows: Figure 1 As shown in the example in the lower left corner, it includes:
[0081] S102-11-1: Input feature map FM input10 After passing through the convolution layer Conv, the batch normalization layer Batch Norm, and the rectified linear unit activation function ReLU in sequence, the feature map FM is obtained. d1 .
[0082] S102-11-2: FM for characteristic spectrum d1 Perform channel splitting operation to obtain two feature maps FM respectively d1-1 and FM d1-2 , FM d1-1 After the convolutional gated recurrent unit ConvGRU stores short-term and long-term memory information, the feature map FM is obtained d2-1 . Then FM d2-1 With FM d1-2 Perform the juxtaposition operation to obtain the feature map FM d2 , the entire structure is a residual unit to avoid gradient disappearance.
[0083] S102-11-3: Feature Spectrum FM d2 After the bilinear interpolation module Bilinear2X, the feature map FM is obtained by upsampling by 2 times 10 .
[0084] S103: Based on the output set, determine the corresponding loss function value under the current model parameters.
[0085] In one embodiment, when determining the loss function value, the high-resolution detail prediction loss may be determined by the following formula and the high-resolution detail prediction output:
[0086] L d =(dilate(a g )-eroad(a g ))||d p -a g ||1
[0087] Among them, L d Predicting loss for high-resolution details; a g is the real supervised alpha mask; dilate and eroad are mathematical morphological processing, representing dilation and erosion respectively, and the parameters are default values; d p Predict output for high-resolution details; ||d p -a g ||1 means d p -a g The L1 norm of .
[0088] The low-resolution semantic estimation loss can be determined by the following formula and the low-resolution semantic estimation output:
[0089]
[0090] Among them, L s is the low-resolution semantic estimation loss, s p is the low-resolution semantic estimation output; G(a g ) is for a g Perform Gaussian blur processing with default parameters.
[0091] The alpha binary pixel loss and RGB color loss are determined by the following formula and the model output:
[0092]
[0093] in, is the alpha binary pixel loss, i represents the pixel index after Flatten, a p Indicates the Alpha mask after DGF; To predict the value of the alpha mask (two-dimensional array) at index position i after flattening it into a one-dimensional array according to the rows; is the value of the one-dimensional array index position i after the real Alpha mask (two-dimensional array) is flattened into a one-dimensional array according to the row; ∈ is a hyperparameter with a default value of 10 -6 ; is RGB color loss, j represents the channel index, c prepresents the foreground image after DGF, c g represents the real image, Represents the predicted foreground image (RGB, three-dimensional array) obtained by channel dimension j as a single-channel two-dimensional array, Represents the single-channel two-dimensional array obtained according to the channel dimension j of the real foreground image (RGB, three-dimensional array).
[0094] The corresponding loss function value under the current model parameters is determined by the following formula, as well as the high-resolution detail prediction loss, low-resolution semantic estimation loss, alpha binary pixel loss, and RGB color loss:
[0095] L=σ l ·[w l ·L a +(1-w l )L c ]+σ d ·L d +σ s ·L s
[0096] Among them, L is the loss function value, w l is a hyperparameter with a default value of 0.6, which is used to balance L a and L c The loss constraint, w l ·L a +(1-w l )L c Constituent loss; σ l , σ d and σ s is a hyperparameter used to balance component loss, low-resolution semantic estimation loss, and high-resolution detail prediction loss. The default value is σ l =σ d =10,σ s =3.
[0097] S104: Adjusting the model parameters in the preset initial model until the loss function value is lower than a preset threshold.
[0098] After obtaining the loss function value, when adjusting the model parameters, feedforward reasoning can be performed based on the above process, and the model parameters can be reverse chained and optimized according to the loss function and BP algorithm until the model converges, that is, the loss of the model on the validation set fluctuates less than 0.002 for 10 consecutive generations, and the final training model is obtained.
[0099] When verifying the model, you can verify the evaluation indicators of the model itself on the test set, and apply the model to the production environment to verify the model effect. The specific evaluation indicators are: SAD absolute error, MAD mean absolute difference, MSE mean square error, Gradient Error, and Connectivity Error.
[0100] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0101] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0102] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0103] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0104] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0106] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0107] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0108] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0109] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0110] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for training a real-time human video matting model, characterized in that: include: Collecting application data, and generating a training data set based on the application data; Input the training data set into a preset initial model to obtain an output set corresponding to the training data set under the current model parameters; The preset initial model is provided with an encoder, a decoder, and a semantic and detail prediction high- and low-resolution reversible branch, the semantic and detail prediction high- and low-resolution reversible branch is only used to train the preset initial model, and is connected to the encoder to output a high-resolution detail prediction output and a low-resolution semantic estimation output; Based on the output set, determine the corresponding loss function value under the current model parameters; The model parameters in the preset initial model are adjusted until the loss function value is lower than a preset threshold.
2. The method according to claim 1, characterized in that The encoder includes five feature encoding modules and a dilated spatial convolutional pooling pyramid module; the five feature encoding modules are respectively composed of two BNeck submodules, two BNeck submodules, three BNeck submodules, six BNeck submodules, and four BNeck submodules in series, and each BNeck submodule has a consistent structure, which is sequentially composed of an inverse residual structure with a linear bottleneck, a depth-separable convolution, a squeeze-excited attention, and an activation function in series; The semantic and detail prediction high and low resolution reversible branch includes two feature encoding modules, two standard convolution blocks, three CBLinear modules, two CBFuse modules, two re-parameterized convolutional network cross-stage local link efficient layer aggregation networks, and two DeConV modules. The CBLinear is constructed by connecting a standard convolution layer, a batch normalization layer, a SiLU activation function layer, and a channel dimension segmentation layer in series. The decoder includes a convolutional gated recurrent unit, a bilinear interpolation module, four feature decoding modules and a convolutional batch normalization rectified linear unit activation function module; the convolutional gated recurrent unit is used to store short-term memory information; The feature decoding module includes a convolution layer, a batch normalization layer, a rectified linear unit activation function, a convolutional gated recurrent unit, and a bilinear interpolation module.
3. The method according to claim 2, characterized in that Input the training data set into the preset initial model to obtain the output set corresponding to the training data set under the current model parameters, specifically including: Inputting the input image of the training data set into the preset initial model, so that the input image passes through a replication module and a standard convolution downsampling module to obtain a high-resolution feature map; Inputting the input image into the encoder to obtain a plurality of intermediate feature maps and a deep feature map; Inputting the input image and the plurality of intermediate feature maps into the semantic and detail prediction high- and low-resolution reversible branch to obtain a high-resolution detail prediction output and a low-resolution semantic estimation output output by the semantic and detail prediction high- and low-resolution reversible branch; After the multiple intermediate feature maps and the deep feature map are sequentially passed through the decoder, the convolutional batch normalization rectified linear unit activation function module, and the standard convolution block, a deep guided filtering intermediate feature map output by the decoder is obtained; The depth-guided filtering intermediate feature map and the high-resolution feature map are input into the depth-guided filtering layer to obtain the model output.
4. The method according to claim 3, characterized in that The step of inputting the input image into the encoder to obtain a plurality of intermediate feature maps and a deep feature map specifically includes: Inputting the input image into the encoder, so that the input image passes through five feature encoding modules and the dilated spatial convolutional pooling pyramid module in sequence to obtain a plurality of intermediate feature maps and a deep feature map output by the encoder; The multiple intermediate feature maps include a first intermediate feature map output by the first feature encoding module, a second intermediate feature map output by the second feature encoding module, a third intermediate feature map output by the third feature encoding module, a fourth intermediate feature map output by the fourth feature encoding module, and a fifth intermediate feature map output by the fifth feature encoding module.
5. The method according to claim 4, characterized in that The step of inputting the input image and the plurality of intermediate feature maps into the semantic and detail prediction high- and low-resolution reversible branch to obtain a high-resolution detail prediction output and a low-resolution semantic estimation output output of the semantic and detail prediction high- and low-resolution reversible branch specifically includes: Input the input image into the semantic and detail prediction high and low resolution reversible branch, so that the input image passes through two feature encoding modules and the first standard convolution block in sequence to obtain a downsampled feature map; Input the multiple intermediate feature maps into the semantic and detail prediction high and low resolution reversible branch, so that the third intermediate feature map passes through the first CBLinear module to obtain the first feature map group, the fourth intermediate feature map passes through the second CBLinear module to obtain the second feature map group, and the fifth intermediate feature map passes through the third CBLinear module to obtain the third feature map group; The first feature map group, the third feature map group and the downsampled feature map are sequentially passed through a first CBFuse module, a first heavily parameterized convolutional network cross-stage local link efficient layer aggregation network, and a first DeConV module to obtain a high-resolution detail prediction output; The first feature map group, the third feature map group and the downsampled feature map are sequentially passed through a first CBFuse module, a first re-parameterized convolutional network cross-stage local link efficient layer aggregation network, and a second standard convolutional block to obtain a sixth intermediate feature map; The second feature map group, the third feature map group, and the sixth intermediate feature map are sequentially passed through the second CBFuse module, the second parameterized convolutional network cross-stage local link efficient layer aggregation network, and the second DeConV module to obtain a low-resolution semantic estimation output.
6. The method according to claim 4, characterized in that After the multiple intermediate feature maps and the deep feature map are sequentially passed through the decoder, the convolution batch normalization rectified linear unit activation function module, and the standard convolution block, a deep guided filtering intermediate feature map output by the decoder is obtained, specifically including: Inputting the deep feature map into a convolutional gated recurrent unit to obtain a first intermediate deep feature map; Concatenate the deep feature map and the first intermediate deep feature map to form an identity mapping residual module, and obtain a second intermediate deep feature map with short-term memory information; The second intermediate deep feature map is subjected to a bilinear interpolation module and doubled upsampled to obtain a third intermediate deep feature map; Concatenate the third intermediate deep feature map and the fifth intermediate feature map, and pass them through the first feature decoding module to obtain a seventh intermediate feature map; Concatenate the seventh intermediate feature map with the fourth intermediate feature map, and pass the result through a second feature decoding module to obtain an eighth intermediate feature map; Concatenate the eighth intermediate feature map with the third intermediate feature map, and pass the third feature decoding module to obtain a ninth intermediate feature map; Concatenate the ninth intermediate feature map with the second intermediate feature map, and pass the result through a fourth feature decoding module to obtain a tenth intermediate feature map; The tenth intermediate feature map is concatenated with the first intermediate feature map, and passed through two convolution batch normalization rectified linear unit activation function modules to obtain an eleventh intermediate feature map; Performing standard convolution on the eleventh intermediate feature map to obtain a depth-guided filtered intermediate feature map.
7. The method according to claim 3, characterized in that Determining the corresponding loss function value under the current model parameters based on the output set specifically includes: The high-resolution detail prediction loss is determined by the following formula and the high-resolution detail prediction output: L d =(dilate(a g )-eroad(a g ))||d p -a g ||1 Among them, L d Predicting loss for high-resolution details; a g is the real supervised alpha mask; dilate and eroad are mathematical morphological processing, representing dilation and erosion respectively, and the parameters are default values; d p Predict output for high-resolution details; ||d p -a g ||1 means d p -a g The L1 norm of ; The low-resolution semantic estimation loss is determined by the following formula and the low-resolution semantic estimation output: Among them, L s is the low-resolution semantic estimation loss, s p is the low-resolution semantic estimation output; G(a g ) is for a g Perform Gaussian blur processing with default parameters; The alpha binary pixel loss and RGB color loss are determined by the following formula and the model output: in, is the alpha binary pixel loss, i represents the pixel index after Flatten, a p Indicates the Alpha mask after DGF; To predict the value of the one-dimensional array index position i after the alpha mask is flattened into a one-dimensional array according to the row; is the value of the one-dimensional array index position i after the real Alpha mask is flattened into a one-dimensional array according to the row; ∈ is a hyperparameter with a default value of 10 -6 ; is RGB color loss, j represents the channel index, c p represents the foreground image after DGF, c g represents the real image, Represents a single-channel two-dimensional array of the predicted foreground image obtained according to channel dimension j, Represents a single-channel two-dimensional array of the real foreground image obtained according to the channel dimension j; The corresponding loss function value under the current model parameters is determined by the following formula, as well as the high-resolution detail prediction loss, low-resolution semantic estimation loss, alpha binary pixel loss, and RGB color loss: L=σ l ·[w l ·L a +(1-w l )L c ]+s d ·L d +s s ·L s Among them, L is the loss function value, w l is a hyperparameter with a default value of 0.6, which is used to balance L a and L c The loss constraint, w l ·L a +(1-w l )L c Constituent loss; σ l , σ d and σ s is a hyperparameter used to balance component loss, low-resolution semantic estimation loss, and high-resolution detail prediction loss. The default value is σ l =σ d =10,σ s =3.
8. The method according to claim 1, characterized in that The adjusting the model parameters in the preset initial model until the loss function value is lower than a preset threshold value specifically includes: The training data set is used for feedforward reasoning, and the model parameters are reverse chain derived and Adam parameter optimized according to the loss function value and the BP algorithm, until the loss of the preset initial model on the validation set fluctuates less than the preset threshold for ten consecutive generations, thereby obtaining the final training model.
9. The method according to claim 1, characterized in that: The collected application data specifically includes: Determining a first preset number of exhibition halls as background material and a second preset number of employees as foreground material; Determining a plurality of preset employee combinations among the second preset number of employees; Shooting a third preset number of video materials of the plurality of preset employee groups in the exhibition hall to construct a first application data set; Shooting a fourth number of dynamic background videos without foreground in the exhibition hall to construct a second application data set; In a green screen environment, a fifth preset number of video materials are shot for the plurality of preset employee combinations to construct a third application data set.
10. The method according to claim 9, characterized in that Generating a training data set based on the application data specifically includes: Using an open source video extinction model, automatically extincting the first application data set to obtain a first extinction data set, where the first extinction data set includes foreground data and first Alpha extinction data corresponding to the first data set; Using an open source video extinction model, automatically extinct the third application data set to obtain third Alpha extinction data corresponding to the third application data set; Using an open source annotation tool, verifying the first Alpha extinction data and the third Alpha extinction data to obtain a first verification data set and a third verification data set; The third verification data set and the second application data set are synthesized by a foreground and background synthesis data script to obtain an intermediate data set; The intermediate data set, the first verification data set and the PPM open source data set are combined as an initial data set; The initial data set is divided according to a preset ratio to obtain a training data set, a verification data set and a test data set.