Image matting method and device for foreground and background similar areas and storage medium

By introducing ESEM semantic feature enhancement module and NPRM non-reconstructed area asymptotic refinement module in the U-Net network architecture, the problem of mistakenly pinching of the existing depth cutout method in the front background similar areas is solved, and a more efficient image cutout effect is achieved.

CN120235892APending Publication Date: 2025-07-01GUIZHOU MINZU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150512.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing depth cutout method is difficult to accurately distinguish the foreground and background when the front background color or texture is highly similar, resulting in missed cutout.

Method used

The cutout model is built based on the U-Net network architecture, combined with the ESEM semantic feature enhancement module and the NPRM non-reconstructed area asymptotic refinement module, forming a hierarchical feature enhancement and refinement mechanism, and enhancing the model's adaptability to complex scenarios by synthesising a large-scale training set.

Benefits of technology

It effectively solves the problem of missed cutting in the previous background similar areas, and improves the overall performance of the cutout model and its adaptability to complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235892A_ABST
    Figure CN120235892A_ABST
Patent Text Reader

Abstract

The invention provides an image matting method and device for a foreground and background similar region and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: constructing a matting model based on a U-Net architecture, connecting a semantic feature enhancement module with a non-reconstruction region asymptotic refinement module, connecting the semantic feature enhancement module with an encoder and a decoder, and enhancing feature extraction; and the image matting module is connected with a multi-stride output point of the decoder, synthesizes the foreground image to a plurality of background images to form a training set, trains the image matting model, outputs features to the decoder to recover resolution, and the non-reconstruction region asymptotic refinement module refines the non-reconstruction region and outputs an image matting result. According to the method, a new matting model is constructed, a semantic feature enhancement module is connected with an encoder, a non-reconstruction region asymptotic refining module is connected with different stride output points of a decoder, a hierarchical feature enhancement and refining mechanism is formed, and the problem of wrong matting of foreground and background similar regions is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of image processing, and particularly relates to an image matting method, device and storage medium for similar foreground and background regions. Background Art

[0002] Traditional matting methods require a well-annotated trimap as an auxiliary guiding input. Although such annotation makes the matting problem easier to handle and has achieved very good accuracy, traditional matting methods, especially in portrait matting, rely too much on the trimap. However, for users without any prior knowledge, obtaining a trimap may be quite burdensome and costly.

[0003] In recent years, with the development of deep learning technology, matting methods based on convolutional neural networks (CNNs) have made significant progress. These methods directly learn the semantic features of the foreground and background from RGB images through end-to-end training, reducing the dependence on the trimap. However, existing deep matting methods still have defects. For example, when the foreground and background colors or textures are highly similar, the model is difficult to accurately distinguish the foreground and background, resulting in mis-matting situations. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an image matting method, device and storage medium for similar foreground and background regions in view of the deficiencies of the prior art.

[0005] The technical solution of the present invention to solve the above technical problems is as follows: An image matting method for similar foreground and background regions includes the following steps:

[0006] Construct a matting model based on the U-Net network architecture. Among them, construct an ESEM semantic feature enhancement module, connect the encoder of the U-Net network architecture with the input of the ESEM semantic feature enhancement module, connect the output of the ESEM semantic feature enhancement module with the decoder of the U-Net network architecture, and construct multiple NPRM non-reconstruction region asymptotic refinement modules, connect the multiple NPRM non-reconstruction region asymptotic refinement modules with the output points at multiple strides of the decoder respectively, and use the last NPRM non-reconstruction region asymptotic refinement module as the output of the matting model;

[0007] Construct a foreground image set and a background image set, synthesize each foreground image in the foreground image set onto multiple background images in the background image set, and form a training set with the obtained multiple synthesized images;

[0008] The foreground region prediction training is performed on the ESEM semantic feature enhancement network model through the training set, and the obtained foreground region prediction features are output to the decoder. The decoder restores the resolution of the image, and the NPRM non-reconstruction region asymptotic refinement module at each stride respectively performs refinement training on the non-reconstruction region of the image with the restored resolution output at the corresponding stride, and outputs the matting result.

[0009] Another technical solution for the present invention to solve the above technical problems is as follows: An image matting device for foreground and background similar regions, comprising:

[0010] A model construction module, configured to construct a matting model based on the U-Net network architecture. Among them, an ESEM semantic feature enhancement module is constructed, the encoder of the U-Net network architecture is connected to the input of the ESEM semantic feature enhancement module, the output of the ESEM semantic feature enhancement module is connected to the decoder of the U-Net network architecture, and multiple NPRM non-reconstruction region asymptotic refinement modules are constructed. The multiple NPRM non-reconstruction region asymptotic refinement modules are respectively connected to the output points at multiple strides of the decoder, and the last NPRM non-reconstruction region asymptotic refinement module is used as the output of the matting model;

[0011] A training set construction module, configured to construct a foreground image set and a background image set, synthesize each foreground image in the foreground image set onto multiple background images in the background image set, and form a training set with the obtained multiple synthesized images;

[0012] A model training module, configured to perform foreground region prediction training on the ESEM semantic feature enhancement network model through the training set, output the obtained foreground region prediction features to the decoder, restore the resolution of the image through the decoder, and respectively perform refinement training on the non-reconstruction regions of the images with the restored resolution output at each stride by the NPRM non-reconstruction region asymptotic refinement module at each stride, and output the matting result.

[0013] Another technical solution for the present invention to solve the above technical problems is as follows: An image matting device for foreground and background similar regions, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the image matting method for foreground and background similar regions as described above is implemented.

[0014] Another technical solution for the present invention to solve the above technical problems is as follows: A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image matting method for foreground and background similar regions as described above is implemented.

[0015] The beneficial effects of the present invention are as follows: Based on the U-Net network architecture, an ESEM semantic feature enhancement module and an NPRM non-reconstructed region asymptotic refinement module are connected to construct a new matte extraction model. By connecting the ESEM module to the encoder and the NPRM module to the output points of different strides of the decoder, a hierarchical feature enhancement and refinement mechanism is formed, effectively solving the problem of incorrect extraction in regions where the foreground and background are similar; by synthesizing a large-scale training set, the adaptability of the model to complex scenes is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flowchart of the image matte extraction method provided by an embodiment of the present invention;

[0017] Figure 2 It is a schematic structural diagram of the NPRM non-reconstructed region asymptotic refinement module provided by an embodiment of the present invention;

[0018] Figure 3 It is a schematic diagram of the definition description of the non-reconstructed region for simulating information loss provided by an embodiment of the present invention;

[0019] Figure 4 It is a schematic structural diagram of the ESEM semantic feature enhancement module provided by an embodiment of the present invention;

[0020] Figure 5 It is a block diagram of the modules of the image matte extraction device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0022] Embodiment 1: As Figure 1 、 Figure 2 shown, an embodiment of the present invention provides a matte extraction method for regions where the foreground and background are similar, including the following steps:

[0023] S1. Construct a matte extraction model based on the U-Net network architecture. Among them, an ESEM semantic feature enhancement module is constructed, the encoder of the U-Net network architecture is connected to the input of the ESEM semantic feature enhancement module, and the output of the ESEM semantic feature enhancement module is connected to the decoder of the U-Net network architecture. Multiple NPRM non-reconstructed region asymptotic refinement modules are constructed and respectively connected to the output points at multiple strides of the decoder. The last NPRM non-reconstructed region asymptotic refinement module is used as the output of the matte extraction model;

[0024] S2. Construct a foreground image set and a background image set, synthesize each foreground image in the foreground image set onto multiple background images in the background image set, and form a training set with the obtained multiple synthesized images;

[0025] S3. Perform foreground region prediction training on the ESEM semantic feature enhancement network model through the training set, output the predicted foreground region features to the decoder, restore the resolution of the image through the decoder, and perform refinement training on the non-reconstructed regions of the images with restored resolution output at each stride through the NPRM non-reconstructed region asymptotic refinement module at each stride, and output the matte result.

[0026] In Embodiment 1, based on the U-Net network architecture, an ESEM semantic feature enhancement module and an NPRM non-reconstructed region asymptotic refinement module are connected to construct a new matte model, realizing the full-process optimization from feature extraction to detail restoration, and improving the overall performance of the matte model; by connecting the ESEM module to the encoder and the NPRM module to different stride output points of the decoder, a hierarchical feature enhancement and refinement mechanism is formed, effectively solving the problem of incorrect matte extraction in the foreground-background similar regions; by synthesizing a large-scale training set, the adaptability of the model to complex scenes is enhanced.

[0027] Embodiment 2: Network training requires a large amount of labeled data. However, there are few publicly available labeled data sets, and more seriously, the color labels may be inaccurate and noisy. In the prior art, synthetic techniques are used to obtain a large amount of training data, and some subsequent works [5, 15, 16, 20] are all trained on this data. However, since the foreground and background images are usually sampled from different distributions, there may be synthetic artifacts in the data augmentation process, resulting in a large domain gap between the synthetic images and natural images, making the learning unstable. To solve this problem, this embodiment uses a foreground consistency learning strategy to randomly select transparency masks to blend with background images to generate synthetic training data. Although the synthesized images may be semantically meaningless, they can provide accurate and unbiased foreground color labels in the transparent regions, improving the generalization of foreground color prediction.

[0028] In the step S2, synthesizing each foreground image in the foreground image set onto multiple background images in the background image set specifically includes:

[0029] Let F i ={f1, f2, f3,..., f p} represent that there are P foreground images, and B ij ={b i1 , b i2 , b i3 ,..., b iq}(indicating there are Q background images, G) ij ={g i1 , g i2 , g i3 , …, g iq} is the synthesized image. Each foreground image in the foreground image set is synthesized onto multiple background images in the background image set through the foreground consistency learning strategy expression. The foreground consistency learning strategy expression is as follows:

[0030]

[0031] where L con represents the consistency loss, P represents the number of foreground images, Q represents the number of background images, e ij and e ik represent the output features of the encoder. JS(e ij , e ik ) represents the JS divergence between e ij and e ik , with a value range from 0 to 1. 0 indicates that e ij and e ik are exactly the same, and 1 indicates that e ij and e ik are completely different.

[0032] In Embodiment 2, the foreground consistency learning strategy (FCLS) is used to generate synthetic data. The JS divergence is used to constrain the encoder feature consistency, reducing the impact of synthetic artifacts on the model; by randomly mixing the foreground and background, diverse training samples are generated, making the model more robust in real scenarios; the JS divergence is introduced to quantify the feature differences, avoiding inaccurate color labels in unknown regions and providing a clear optimization goal for the training process.

[0033] Embodiment 3: As shown in Figure 4 , the ESEM semantic feature enhancement module includes an ASPP atrous spatial pyramid pooling module, an EMA multi-scale attention module, and a CBR convolutional output module;

[0034] In step S3, the ESEM semantic feature enhancement network model is trained for foreground region prediction using the training set. Specifically:

[0035] The processing processes of the ASPP atrous spatial pyramid pooling module, the EMA multi-scale attention module, and the CBR convolutional output module are represented by the first formula:

[0036]

[0037] where It represents the image features obtained by convolving the image features output by the encoder with the 3×3 dilated convolutional layer of the ASPP (Atrous Spatial Pyramid Pooling) module. dil2, dil4, and dil8 represent dilation rates of 2, 4, and 8 respectively. It represents the image features obtained by convolving the image features output by the dilated convolutional layer with the 1×1 convolutional layer of the ASPP module. CAT represents the concatenation of each image feature by the CAT concatenation layer to obtain multi-scale features. EMA represents that the EMA multi-scale attention module enhances the multi-scale features through a multi-scale attention mechanism to obtain enhanced image features. CBR represents that the CBR convolutional output module normalizes the enhanced image features and then performs non-linear processing on the normalized image features to obtain foreground region prediction features.

[0038] Specifically, for an image with an input resolution adjusted to 512×512, the finally extracted feature dimension in the encoder is (512, 32, 32), where 512 is the number of channels, and the two 32s respectively represent the length and width. Then, the output tensor of the last block in the encoder is input into the ESEM, as Figure 4 shown. The ESEM consists of two parts: 1) One 1×1 convolution, three 3×3 dilated convolutions with dilation rates of (2, 4, 8) respectively, and a global average pooling layer; 2) Two parallel 1×1 convolution path branches and one 3×3 convolution branch. The input feature tensor received by the ESEM first passes through one 1×1 convolution and three 3×3 dilated convolutions, and the output channel number of each convolution is 256. At the same time, an image-level feature is obtained through a global average pooling layer, the output channel number is adjusted to 256 through a 1×1 convolution, and the resolution is adjusted to 32×32 using bilinear interpolation. Then, the feature maps of different scales obtained in the above steps are concatenated together along the channel dimension. At this time, the dimension of the feature tensor is (1280, 32, 32). Then, the concatenated feature is transmitted to the grouped feature layer. At this time, the grouping factor factors is set to 16, and the dimension of the feature tensor is (32, 32, 32). Through the one-dimensional global average pooling operations in the two parallel 1×1 convolution path branches to encode channel information in two spatial directions respectively, and the single 3×3 convolution branch is used to capture multi-scale feature representations. Finally, through a CBR operation, a feature tensor of (512, 32, 32) is finally obtained.

[0039] In Embodiment 3, multi-scale features are extracted through the ASPP Atrous Spatial Pyramid Pooling module, combined with the EMA multi-scale attention mechanism, significantly enhancing the model's ability to capture semantic information at different scales. The high-level semantic features are used to guide the low-level features to correctly predict the details of the foreground region, generating more recognizable feature representations. The CAT concatenation layer can combine features of different scales and characteristics to form a feature tensor containing richer information, providing a multi-scale and multi-faceted information basis for the subsequent EMA module and CBR module processing, helping the model better learn image features and improving the accuracy of portrait matting.

[0040] The CBR Convolutional Output module includes a 1×1 convolution, a BN layer, and a ReLU layer, which perform normalization and non-linear processing on the features, enhancing the discriminability of the features and the convergence speed of the model.

[0041] The EMA multi-scale attention module, through grouped feature layers and cross-channel attention mechanism, maintains the expressive ability of multi-scale features while reducing the computational amount.

[0042] Embodiment 4: As Figure 2 shown, the NPRM Non-Reconstruction Region Asymptotic Refinement module at each stride separately refines the non-reconstruction region of the resolution-restored image output at the corresponding stride. Specifically:

[0043] In step S3, the NPRM Non-Reconstruction Region Asymptotic Refinement module determines the non-reconstruction region of the resolution-restored image output at the corresponding stride, including:

[0044] Let M h be the input label image of the current layer h, S ↑ and S ↓ respectively represent 2x nearest neighbor downsampling and upsampling operations. A non-reconstruction region is obtained from the decoder output of the previous layer at the corresponding stride through the second formula:

[0045]

[0046] where, N h represents the non-reconstruction region mask, ⊕ represents the logical exclusive OR operation, O ↑ represents 2x downsampling by performing a logical OR operation within each 2x2 neighborhood. When at least one pixel of the input label image M h-1 is different from its reconstructed label image, then the pixel (x, y) is a pixel of the non-reconstruction region. When the pixel point belongs to the non-reconstruction region, i.e., N h (x, y) = 1, when the pixel point does not belong to the non-reconstruction region, N h (x, y) = 0;

[0047] Define the transparency region of the non-reconstructed region as an unknown region, that is, 1 < α < 0, and replace the unknown region in the input of the previous layer with the original input of the current layer to update the output of the current layer, as expressed by the third formula:

[0048]

[0049] Among them, represents the matte output transparency value after the update of the current layer h, represents the transparency matte prediction of the current layer h, which is the output result without processing the non-reconstructed region, and α h-1 represents the transparency matte prediction of the previous layer h-1, and (1 - N h ) represents the determined region of the previous layer, which is the logical negation of N h , that is, when the pixel belongs to the non-reconstructed region, 1 - N h (x, y) = 0, and when the pixel does not belong to the non-reconstructed region, 1 - N h (x, y) = 1.

[0050] In Embodiment 4, the reduction of the image spatial resolution in the downsampling operation of the matte network encoder to extract features will cause the loss of image detail features to a certain extent, and these lost detail features cannot be restored during the decoder upsampling process, resulting in difficulty in extracting the relatively fine features in the foreground edge region. In this embodiment, the input label image itself is downsampled to simulate the information loss caused by downsampling in the network, and the pixel points that cannot restore features during the upsampling process are defined as non-reconstructed points, and the region composed of non-reconstructed points is defined as the non-reconstructed region (Non-reconstructed regions). Intuitively, the non-reconstructed regions are mostly distributed in the foreground boundary or high-frequency regions, and are composed of points missing from the real label image or with additional predicted wrong labels. The process is as Figure 3 shown.

[0051] Input mask: Figure 3 The upper left corner in

[0052] is the input mask, which is the initial image mask state, containing many small squares, representing different attributes or states.

[0053] Down-sampling: Starting from the input mask, a compressed mask is obtained through the down-sampling operation, the number of squares decreases, and the mask size becomes smaller, which simulates the information loss caused by downsampling in the network.

[0054] Determine different regions: Compare the input mask with the recovery mask to determine different regions. These regions are the parts that are different from the original input during the recovery process.

[0055] Determine non-reconstructed regions: Based on the different regions, further determine non-reconstructed regions. Non-reconstructed regions are regions composed of pixel points where features cannot be recovered during the upsampling process, and are of great significance in subsequent matte extraction and other operations, guiding the model's processing of specific regions.

[0056] Specifically, as Figure 2 shown, an NPRM Non-Reconstructed Region Asymptotic Refinement Module is attached at the output strides 1, 4, and 8 of the decoder respectively to selectively fuse the matte outputs from the previous layer and the current layer.

[0057] The advantages of Embodiment 4 are: The NPRM Non-Reconstructed Region Asymptotic Refinement Module can accurately detect non-reconstructed regions, simulate information loss through downsampling-upsampling, combine the logical exclusive OR operation (⊕) and the neighborhood OR operation (O↑) to accurately identify regions where high-frequency details are lost (such as the hair edges of a portrait), use the asymptotic refinement mechanism to gradually update the transparency values of non-reconstructed regions at different decoder strides (1 / 4 / 8), retain the determined regions (1 - N h ) and only refine the unknown regions (N h ), avoiding global repeated calculations; through hierarchical progressive refinement, gradually narrow the range of non-reconstructed regions, making the model more focused on improving the accuracy of local similar regions; provide self-guiding information features of non-reconstructed regions through learning during the decoding process, gradually improve the unknown regions of the image, so as to achieve the robustness of matte extraction for local similar regions of a portrait.

[0058] Specifically, as Figure 2As shown, using mask guidance can only provide rough spatial prior information of the foreground region. Therefore, the NPRN model of this embodiment needs to perform a higher-level semantic understanding of the input mask in order to robustly detect the foreground, background, and unknown regions. At the same time, the proposed model must capture low-level information such as the edges and textures of the image to generate fine details of the transparency mask. Therefore, the matting model is based on the U-Net architecture and uses ResNet-34 with an ASPP module as the backbone network. Specifically, the first convolutional layer is adjusted to adopt a 4-channel input composed of an RGB image and a single-channel mask, which are jointly transmitted to the encoder for feature extraction. To enhance the network's ability to learn high-level semantic features, the output of the last block of the encoder is input into the proposed Efficient Semantic Feature Enhancement Module (ESEM) to learn features under different receptive fields of the backbone network. After that, the features processed by ESEM are transmitted to the decoder part for decoding operations. At stride 1, stride 4, and stride 8 of the decoder output, a Progressive Refinement Module for Non-reconstructed regions (NPRM) is attached respectively to improve the robustness of matting the local regions of the portrait. At the same time, during the entire network training process, a Foreground Consistency Learning Strategy (FCLS) is used to randomly generate synthetic training data from the Alpha image and the RGB image to estimate inaccurate color labels of the unknown regions of the image. In addition, there are shortcut connections between the encoder and the decoder. The shortcut block consists of two convolutional layers with a stride of 1, a batch normalization layer, and a Relu activation function, which are used to process low-level texture features.

[0059] Embodiment 5: It further includes step S4. After the refinement training, the matting model is optimized using a loss function:

[0060] The loss function is:

[0061]

[0062] where λ h represents the loss weight assigned to the outputs of different levels, h represents the hierarchical index variable, N h represents the non-reconstructed region mask, α h represents the value of the h-th layer transparency mask layer, represents the value of the transparency mask predicted by the model at the h-th layer.

[0063] where L is the regression loss function L l1, the combined loss function \(L\) com and the Laplacian loss function \(L\) lap , that is

[0064] In Embodiment 5, weights are dynamically allocated to the losses of different hierarchical outputs, such as \(\lambda_0:\lambda_1:\lambda_2 = 1:2:3\). The attention of the enhanced model to the detail layer is strengthened; the regression loss function \(L\) is integrated l1 , the combined loss function \(L\) com and the Laplacian loss function \(L\) lap , which can take into account the global consistency and local gradient smoothness of the transparency mask; by adjusting the loss calculation through the non-reconstruction region mask (N h ), the interference of noise pixels to training is reduced, and the robustness of the model is improved.

[0065] Embodiment 6: As Figure 5 shown, the embodiment of the present invention also provides a matte extraction device for foreground-background similar regions, including:

[0066] A model construction module, configured to construct a matte extraction model based on the U-Net network architecture. Among them, an ESEM semantic feature enhancement module is constructed, the encoder of the U-Net network architecture is connected to the input of the ESEM semantic feature enhancement module, and the output of the ESEM semantic feature enhancement module is connected to the decoder of the U-Net network architecture, and multiple NPRM non-reconstruction region asymptotic refinement modules are constructed, and the multiple NPRM non-reconstruction region asymptotic refinement modules are respectively connected to the output points at multiple strides of the decoder, and the last NPRM non-reconstruction region asymptotic refinement module is used as the output of the matte extraction model;

[0067] A training set construction module, configured to construct a foreground image set and a background image set, synthesize each foreground image in the foreground image set onto multiple background images in the background image set, and form a training set from the obtained multiple synthesized images;

[0068] A model training module, configured to perform foreground region prediction training on the ESEM semantic feature enhancement network model through the training set, output the obtained foreground region prediction features to the decoder, restore the resolution of the image through the decoder, and respectively perform refinement training on the non-reconstruction regions of the images with restored resolution output at each stride through the NPRM non-reconstruction region asymptotic refinement modules at each stride, and output the matte extraction result.

[0069] Embodiment 7: The ESEM semantic feature enhancement module includes an ASPP atrous spatial pyramid pooling module, an EMA multi-scale attention module, and a CBR convolutional output module;

[0070] In the model training module, the ESEM semantic feature enhancement network model is trained for foreground region prediction using the training set. Specifically:

[0071] The processing processes of the ASPP atrous spatial pyramid pooling module, EMA multi-scale attention module, and CBR convolutional output module are represented by the first formula:

[0072]

[0073] Among them, represents the image features obtained by convolving the image features output by the encoder with the 3×3 atrous convolutional layer of the ASPP atrous spatial pyramid pooling module. dil2, dil4, and dil8 represent atrous rates of 2, 4, and 8 respectively. represents the image features obtained by convolving the image features output by the atrous convolutional layer with the 1×1 convolutional layer of the ASPP atrous spatial pyramid pooling module. CAT represents the concatenation of various image features by the CAT concatenation layer to obtain multi-scale features; EMA represents that the EMA multi-scale attention module enhances the multi-scale features through a multi-scale attention mechanism to obtain enhanced image features; CBR represents that the CBR convolutional output module normalizes the enhanced image features and then performs non-linear processing on the normalized image features to obtain foreground region prediction features.

[0074] Example 8: In the model training module, the NPRM non-reconstructed region asymptotic refinement module at each stride is used to refine the non-reconstructed region of the image after resolution recovery output at the corresponding stride. Specifically:

[0075] The NPRM non-reconstructed region asymptotic refinement module determines the non-reconstructed region of the image after resolution recovery output at the corresponding stride, including:

[0076] Let M h be the input label image of the current layer h, S ↑ and S ↓ represent 2x nearest neighbor downsampling and upsampling operations respectively. A non-reconstructed region is obtained from the decoder output of the previous layer at the corresponding stride through the second formula:

[0077]

[0078] Among them, N h represents the non-reconstructed region mask, ⊕ represents the logical exclusive OR operation, O ↑ represents 2x downsampling by performing a logical OR operation within each 2x2 neighborhood. When the input label image M h-1When at least one pixel of the tag image after reconstruction is different, the pixel (x, y) is a pixel in the non-reconstructed area. When the pixel belongs to the non-reconstructed area, i.e., N h (x, y) = 1. When the pixel does not belong to the non-reconstructed area, N h (x, y) = 0;

[0079] Let the transparency area of the non-reconstructed area be defined as the unknown area, i.e., 1 < α < 0, and use the original input of the current layer to replace the unknown area in the input of the previous layer to update the output of the current layer, which is expressed by the third formula:

[0080]

[0081] Among them, represents the matte output transparency value after the update of the current layer h, represents the matte mask prediction of the current layer h, which is the output result before processing the non-reconstructed area, α h-1 represents the matte mask prediction of the previous layer h - 1, (1 - N h ) represents the determined area of the previous layer, which is the logical negation of N h , that is, when the pixel belongs to the non-reconstructed area, 1 - N h (x, y) = 0. When the pixel does not belong to the non-reconstructed area, 1 - N h (x, y) = 1.

[0082] Example 9: The embodiment of the present invention also provides a matte extraction device for foreground and background similar areas, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the image matte extraction method for foreground and background similar areas as described above is implemented.

[0083] Example 10: The embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the image matte extraction method for foreground and background similar areas as described above is implemented.

[0084] Analysis of comparative experiment results:

[0085] The portrait matting data is composed of a mixture of portrait images collected from existing public datasets, specifically including AdobeMatting, Real-world Portrait-636, PPM-100, and Distinctions-646. A total of 1323 portrait foregrounds were collected. According to the synthesis route in the literature, the collected data was randomly divided into a training set of 1190 images and a test set of 133 images at a ratio of 9:1. Then, the foreground images were randomly synthesized into the BG-20k background dataset, where 20 backgrounds and 10 backgrounds were randomly selected for each training and test foreground respectively. In this way, a training dataset with 1190 unique foreground objects and 23800 images was created, named MCP-20k, while the test dataset has 133 unique objects and 1330 images, named MCP-1k dataset.

[0086] To verify the performance of the NPRN model in the foreground mask extraction problem when the foreground and background information of the portrait is similar, it was compared with the depth matting models based on tripartite graphs on the MCP-1k dataset and the P3M-500-NP real public dataset, specifically including DIM, MGM, and FGI. The results are shown in Table 1. Table 1 shows the error comparison results of different matting methods for each model on the MCP-1k dataset.

[0087] Table 1

[0088] Method MSE SAD Grad Conn DIM 0.0071 42.9001 31.9893 45.3273 FGI 0.0204 31.1850 27.9905 41.2200 MGM 0.0178 73.1839 28.3804 35.1534 MGM* 0.0149 67.6554 26.1308 32.7752 NPRN* 0.0066 39.3047 24.0353 28.5328 NPRN 0.0062 37.1809 23.4635 30.6407

[0089] Among them, the * sign indicates the test results with the segmentation map of the SOLOv2 model as the input.

[0090] Table 1 shows the error results of the NPRN* model, the NPRN model, and other matting models on the synthetic dataset MCP-1k. Compared with the matting models DIM, FGI, and MGM based on the tripartite graph, the NPRN* model outperforms these models in terms of the MSE, Grad, and Conn metrics, and is only lower than the FGI model in terms of the SAD metric, i.e., 39.3047 VS 31.1850. However, the NPRN* model is tested with the segmentation map as the auxiliary input instead of the costly tripartite graph. In contrast, the NPRN* model has a lower testing cost and generally maintains a lower matting error. At the same time, compared with the MGM* model, the errors of the NPRN* model in the MSE, SAD, Grad, and Conn metrics are much lower than those of the MGM* model. This shows that compared with the models using the segmentation map as the test auxiliary input, the NPRN* model of the present invention has better performance in the foreground mask extraction problem when the foreground and background information of the portrait is extremely similar. Meanwhile, the present invention uses the tripartite graph as the test auxiliary input of the NPRN model. By comparing with other matting methods based on the tripartite graph and the segmentation map, it is found that the errors of the NPRN model in the four metrics are mostly lower than those of other models, and the error in the SAD metric is only slightly higher than that of the FGI model, i.e., 37.1809 VS 31.1850. The main reason for the higher error of the NPRN model in the SAD metric is that although the NPRN model uses the tripartite graph as the test input, the model trained in the present invention has a 4-channel input, while the tripartite graph has a 6-channel input when used as the test input. Therefore, data processing is performed on the tripartite graph during input, and what is actually tested is the segmentation map obtained by processing the tripartite graph, rather than the original annotated tripartite graph. So the error in SAD is higher than that of the FGI model. Although the error of the NPRN model in the SAD metric is slightly higher than that of the FGI model, the NPRN model designed in the present invention performs the best in the other three metrics, which proves that the improvement made by the present invention for the matting problem when the foreground and background information of the portrait is extremely similar is effective. In addition, compared with the NPRN model, the error differences of the NPRN* model in the four metrics are not significant, and only the difference in SAD is relatively large, i.e., 39.3047 VS 37.1809. Moreover, the NPRN* model performs better than the NPRN model in Grad, which shows that different test inputs have a greater impact on SAD and a relatively weaker impact on other metrics. This proves that the model designed in the present invention has stronger robustness to different test inputs.

[0091] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0092] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0093] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0094] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0095] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0096] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0097] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for cutting out an image of a foreground-background similar region, characterized in that: The steps include: A cutout model is constructed based on a U-Net network architecture, wherein an ESEM semantic feature enhancement module is constructed, an encoder of the U-Net network architecture is connected to an input of the ESEM semantic feature enhancement module, and an output of the ESEM semantic feature enhancement module is connected to a decoder of the U-Net network architecture, and multiple NPRM non-reconstructed region asymptotic refinement modules are constructed, multiple NPRM non-reconstructed region asymptotic refinement modules are respectively connected to output points at multiple strides of the decoder, and the last NPRM non-reconstructed region asymptotic refinement module is used as the output of the cutout model; Constructing a foreground image set and a background image set, synthesizing each foreground image in the foreground image set onto multiple background images in the background image set, and forming a training set with the obtained multiple synthesized images; The ESEM semantic feature enhancement network model is trained for foreground area prediction using the training set, and the obtained foreground area prediction features are output to the decoder. The resolution of the image is restored by the decoder, and the non-reconstructed area of ​​the image after resolution restoration output at the corresponding step is refined and trained by the NPRM non-reconstructed area asymptotic refinement module at each step, and the cutout result is output.

2. The image cutout method according to claim 1, characterized in that: The step of synthesizing each foreground image in the foreground image set onto a plurality of background images in the background image set is specifically as follows: Let F i ={f1, f2, f3, ..., f p } indicates that there are P foreground images, B ij = {b i1 , b i2 , b i3 , …, b iq } indicates that there are Q background images, G ij = {g i1 , g i2 , g i3 , …, g iq } is a synthetic image, and each foreground image in the foreground image set is synthesized onto multiple background images in the background image set by using a foreground consistency learning strategy expression, and the foreground consistency learning strategy expression is: Among them, L con represents the consistency loss, P represents the number of foreground images, Q represents the number of background images, and e ij and e ik represents the output feature of the encoder, JS(e ij ,e ik ) means e ij and e ik The JS divergence between the two values ​​ranges from 0 to 1, where 0 means e ij and e ik Exactly the same, 1 means e ij and e ik Completely different.

3. The image cutout method according to claim 1, characterized in that: The ESEM semantic feature enhancement module includes an ASPP void space pyramid pooling module, an EMA multi-scale attention module and a CBR convolution output module; The foreground area prediction training of the ESEM semantic feature enhancement network model is performed by using the training set, specifically: The processing process of the ASPP atrous spatial pyramid pooling module, the EMA multi-scale attention module, and the CBR convolution output module is represented by the first formula: in, The 3×3 atrous convolution layer of the ASPP atrous spatial pyramid pooling module convolves the image features output by the encoder. dil2, dil4, and dil8 represent the atrous rates of 2, 4, and 8, respectively. It represents the image features obtained by convolving the image features output by the dilated convolution layer with the 1×1 convolution layer of the ASPP dilated spatial pyramid pooling module. CAT represents the CAT cascade layer cascading each image feature to obtain a multi-scale feature. EMA represents the EMA multi-scale attention module enhancing the multi-scale features through a multi-scale attention mechanism to obtain enhanced image features. CBR represents the CBR convolution output module normalizing the enhanced image features, and then performing nonlinear processing on the standardized image features to obtain foreground area prediction features.

4. The image cutout method according to claim 1, characterized in that: The NPRM non-reconstructed region asymptotic refinement module at each step performs refinement training on the non-reconstructed region of the resolution restored image output at the corresponding step, specifically: The NPRM non-reconstructed region asymptotic refinement module determines the non-reconstructed region of the resolution restored image output at the corresponding step, including: Assume M h is the input label image of the current layer h, S ↑ and S ↓ Represent the 2x nearest neighbor down-sampling and up-sampling operations respectively, and a non-reconstructed region is obtained from the decoder output of the previous layer at the corresponding stride through the second formula: Among them, N h represents the non-reconstructed region mask, Represents logical XOR operation, O ↑ represents a 2x downsampling by performing a logical OR operation in each 2×2 neighborhood. When the input label image M h-1 If at least one pixel of the reconstructed label image is different from that of the reconstructed label image, the pixel (x, y) is a pixel in the non-reconstructed area. When the pixel belongs to the non-reconstructed area, that is, N h (x, y) = 1, when the pixel does not belong to the non-reconstructed area, N h (x,y)=0; Assume that the transparency area of ​​the non-reconstructed area is defined as the unknown area, that is, 1<α<0, and replace the unknown area in the previous layer input with the original input of the current layer to update the output of the current layer, which is expressed by the third formula: in, Indicates the transparency value of the cutout output after the current layer h is updated. Represents the transparency mask prediction of the current layer h, which is the output result without processing the non-reconstructed area, α h-1 Represents the transparency mask prediction of the previous layer h-1, (1-N h ) represents the determined area of ​​the previous layer, and N h The logical inversion of , that is, when the pixel belongs to the non-reconstructed area, 1-N h (x, y) = 0, when the pixel does not belong to the non-reconstructed area, 1-N h (x,y)=1.

5. The image cutout method according to claim 4, characterized in that: After the refinement training, the step of optimizing the cutout model using a loss function is also included: The loss function is: Among them, λ h represents the loss weights assigned to different levels of output, h represents the level index variable, N h represents the non-reconstructed region mask, α h Indicates the value of the transparent mask layer of layer h. represents the transparency mask value predicted by the h-th layer model, and L is the regression loss function L l1 , synthetic loss function L com And the Laplace loss function L lap The sum of 6. An image cutout device for similar foreground and background regions, characterized in that: include: A model construction module is used to construct a cutout model based on a U-Net network architecture, wherein an ESEM semantic feature enhancement module is constructed, an encoder of the U-Net network architecture is connected to an input of the ESEM semantic feature enhancement module, and an output of the ESEM semantic feature enhancement module is connected to a decoder of the U-Net network architecture, and multiple NPRM non-reconstructed region asymptotic refinement modules are constructed, multiple NPRM non-reconstructed region asymptotic refinement modules are respectively connected to output points at multiple strides of the decoder, and the last NPRM non-reconstructed region asymptotic refinement module is used as the output of the cutout model; A training set construction module, used to construct a foreground image set and a background image set, synthesize each foreground image in the foreground image set onto a plurality of background images in the background image set, and form a training set with the obtained plurality of synthesized images; The model training module is used to perform foreground area prediction training on the ESEM semantic feature enhancement network model through the training set, output the obtained foreground area prediction features to the decoder, restore the resolution of the image through the decoder, and perform refinement training on the non-reconstructed area of ​​the image after resolution restoration output at the corresponding step through the NPRM non-reconstructed area asymptotic refinement module at each step, and output the cutout result.

7. The image cutout device according to claim 6, characterized in that: The ESEM semantic feature enhancement module includes an ASPP void space pyramid pooling module, an EMA multi-scale attention module and a CBR convolution output module; In the model training module, the ESEM semantic feature enhancement network model is trained for foreground area prediction using the training set, specifically: The processing process of the ASPP atrous spatial pyramid pooling module, the EMA multi-scale attention module, and the CBR convolution output module is represented by the first formula: in, The 3×3 atrous convolution layer of the ASPP atrous spatial pyramid pooling module convolves the image features output by the encoder. dil2, dil4, and dil8 represent the atrous rates of 2, 4, and 8, respectively. It represents the image features obtained by convolving the image features output by the dilated convolution layer with the 1×1 convolution layer of the ASPP dilated spatial pyramid pooling module. CAT represents the CAT cascade layer cascading each image feature to obtain a multi-scale feature. EMA represents the EMA multi-scale attention module enhancing the multi-scale features through a multi-scale attention mechanism to obtain enhanced image features. CBR represents the CBR convolution output module normalizing the enhanced image features, and then performing nonlinear processing on the standardized image features to obtain foreground area prediction features.

8. The image cutout device according to claim 6, characterized in that: In the model training module, the non-reconstructed regions of the resolution-restored images output at the corresponding strides are respectively refined and trained through the NPRM non-reconstructed region asymptotic refinement modules at each stride, specifically: The NPRM non-reconstructed region asymptotic refinement module determines the non-reconstructed region of the resolution restored image output at the corresponding step, including: Assume M h is the input label image of the current layer h, S ↑ and S ↓ Represent the 2x nearest neighbor down-sampling and up-sampling operations respectively, and a non-reconstructed region is obtained from the decoder output of the previous layer at the corresponding stride through the second formula: Among them, N h represents the non-reconstructed region mask, Represents logical XOR operation, O ↑ represents a 2x downsampling by performing a logical OR operation in each 2×2 neighborhood. When the input label image M h-1 If at least one pixel of the reconstructed label image is different from that of the reconstructed label image, the pixel (x, y) is a pixel in the non-reconstructed area. When the pixel belongs to the non-reconstructed area, that is, N h (x, y) = 1, when the pixel does not belong to the non-reconstructed area, N h (x,y)=0; Assume that the transparency area of ​​the non-reconstructed area is defined as the unknown area, that is, 1<α<0, and replace the unknown area in the previous layer input with the original input of the current layer to update the output of the current layer, which is expressed by the third formula: in, Indicates the transparency value of the cutout output after the current layer h is updated. Represents the transparency mask prediction of the current layer h, which is the output result without processing the non-reconstructed area, α h-1 Represents the transparency mask prediction of the previous layer h-1, (1-N h ) represents the determined area of ​​the previous layer, and N h The logical inversion of , that is, when the pixel belongs to the non-reconstructed area, 1-N h (x, y) = 0, when the pixel does not belong to the non-reconstructed area, 1-N h (x,y)=1.

9. An image cutout device for similar areas of foreground and background, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the image cutout method for foreground-background similar areas as described in any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image cutout method of foreground-background similar areas as described in any one of claims 1 to 5 is implemented.