Neural network training and matting method, device, and storage medium

By training the neural network on the user-specified area of interest, generating prediction mask diagrams and conducting supervision and training, the existing cutout algorithms are solved, and efficient and accurate cutout effects are achieved.

CN114742839BActive Publication Date: 2025-07-22SHANGHAI SENSETIME TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210080369.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-07-22
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

The existing cutout algorithm requires users to provide accurate three-point annotation, which results in a long time-consuming process. The method without a three-point image cannot accurately cut out the picture in complex scenarios, especially for transparent objects.

Method used

By determining the object area that the user is concerned about as the area of interest, the initial neural network is used to generate a predictive mask map, and the neural network is trained based on the differences between the predictive mask map and the real mask to obtain the target neural network for cutting.

Benefits of technology

It effectively saves labeling time and improves the quality of cutouts, especially for complex scenes and transparent objects, expands the application scenarios of image cutouts and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114742839B_ABST
    Figure CN114742839B_ABST
Patent Text Reader

Abstract

The present disclosure provides a neural network training and matting method and apparatus, and a storage medium. The method includes: determining a region of interest on a sample image, where the region of interest is an image region where an object of interest to a user is located; inputting the sample image marked with the region of interest into an initial neural network, and the initial neural network determines a predicted mask image corresponding to the region of interest, where the predicted mask image is used to identify foreground pixel points and / or background pixel points in the region of interest; determining a predicted matting mask corresponding to the region of interest output by the initial neural network based on the sample image marked with the region of interest and the predicted mask image; training the initial neural network based on the difference between the predicted matting mask and the true mask corresponding to the region of interest, and after the training is completed, obtaining a target neural network for matting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision, and in particular, to a neural network training and matting method and apparatus, and a storage medium. Background Art

[0002] Image matting is a fundamental and challenging task in computer vision. Given one or more natural images, it is to predict the foreground and background parts therein and estimate the transparency values of the pixels in the mixed regions. Matting technology has important research value and a wide range of application scenarios. For example, video conferencing, live scene replacement, movie processing, image editing, video editing, etc.

[0003] Currently, matting algorithms can be divided into two categories:

[0004] The first category is the Trimap-based method.

[0005] This type of method requires the user to provide an accurate trimap of the object to be matted while inputting a color mode (Red Green Blue, RGB) image, that is, to mark the definite foreground, the definite background, and the uncertain region. Although this type of method can obtain relatively fine prediction values, the annotation of the trimap itself is very time-consuming.

[0006] The second category is the Trimap-free method.

[0007] Although this type of method saves a large amount of initial interaction time, due to the complex and variable matting objects, when there are multiple objects in an image, it is impossible to select the correct object for matting without interaction information guidance. When the object to be matted is a transparent object, usually a good matting effect cannot be obtained either. Summary of the Invention

[0008] The present disclosure provides a neural network training and matting method and apparatus, and a storage medium.

[0009] According to a first aspect of the embodiments of the present disclosure, a neural network training method is provided. The method includes: determining a region of interest on a sample image, where the region of interest is an image region where an object of interest to the user is located; inputting the sample image with the region of interest marked into an initial neural network, and determining, by the initial neural network, a predicted mask map corresponding to the region of interest, where the predicted mask map is used to identify foreground pixel points and / or background pixel points in the region of interest; determining a predicted matte mask corresponding to the region of interest output by the initial neural network based on the sample image with the region of interest marked and the predicted mask map; training the initial neural network based on the difference between the predicted matte mask and a ground truth matte corresponding to the region of interest, and after the training is completed, obtaining a target neural network for matte extraction.

[0010] In some optional embodiments, the step of determining a region of interest on a sample image includes: determining the object marked by the user based on a specified annotation method on the sample image; and taking the image region within the bounding box of the object on the sample image as the region of interest.

[0011] In some optional embodiments, the method further includes: outputting a prompt message, where the prompt message is used to prompt the user to mark the object of interest on the sample image based on the specified annotation method; and the step of determining the object marked by the user based on a specified annotation method on the sample image includes: determining the object marked by the user based on the prompt message on the sample image.

[0012] In some optional embodiments, the initial neural network includes a backbone network, a first branch network, and a second branch network. The backbone network is used to extract features from the input image, the first branch network is used to determine the predicted mask map, and the second branch network is used to determine the predicted matte mask.

[0013] In some optional embodiments, the method further includes: determining a ground truth mask map corresponding to the region of interest based on the ground truth matte; training the backbone network and the first branch network based on the difference between the predicted mask map and the ground truth mask map, and after the training is completed, performing the step of determining a predicted matte mask corresponding to the region of interest output by the initial neural network based on the sample image with the region of interest marked and the predicted mask map; and the step of training the initial neural network based on the difference between the predicted matte mask and a ground truth matte image corresponding to the region of interest includes: training the backbone network and the second branch network based on the difference between the predicted matte mask and the ground truth matte.

[0014] In some alternative embodiments, determining the predicted matte corresponding to the region of interest in the initial neural network output based on the sample image with the region of interest marked and the predicted mask map includes: predicting the transparency of the foreground pixel points in the region of interest based on the sample image with the region of interest marked and the predicted mask map, and determining the predicted matte.

[0015] In some alternative embodiments, the method further includes: performing at least one feature information extraction in the backbone network and the second branch network respectively; wherein the feature information extraction is used to extract the feature information corresponding to multiple different convolutional kernel sizes in the feature map corresponding to the region of interest; and fusing the feature information corresponding to multiple different convolutional kernel sizes extracted each time to determine the fused feature information.

[0016] In some alternative embodiments, fusing the feature information corresponding to multiple different convolutional kernel sizes extracted each time to determine the fused feature information includes: multiplying the feature information corresponding to multiple different convolutional kernel sizes extracted each time by multiple adaptive weight values respectively and then adding them to obtain the fused feature information.

[0017] According to a second aspect of the embodiments of the present disclosure, a matte extraction method is provided. The method includes: determining a region of interest on the image to be processed; wherein the region of interest is the image region where the object of interest to the user is located; inputting the image to be processed with the region of interest marked into a target neural network for matte extraction to obtain a predicted matte for matting the object output by the target neural network; wherein the target neural network is trained by using the method according to any one of the first aspects described above.

[0018] According to a third aspect of the embodiments of the present disclosure, a neural network training device is provided, including: a first region determination module, configured to determine a region of interest on a sample image; wherein the region of interest is the image region where the object of interest to the user is located; a predicted mask map determination module, configured to input the sample image with the region of interest marked into an initial neural network, and determine a predicted mask map corresponding to the region of interest by the initial neural network; wherein the predicted mask map is used to identify the foreground pixel points and / or background pixel points in the region of interest; a predicted matte determination module, configured to determine a predicted matte corresponding to the region of interest output by the initial neural network based on the sample image with the region of interest marked and the predicted mask map; and a first training module, configured to train the initial neural network based on the difference between the predicted matte and the true matte corresponding to the region of interest, and obtain a target neural network for matte extraction after the training is completed.

[0019] In some alternative embodiments, the region determination module includes: an object determination sub-module, configured to determine, on the sample image, the object annotated by the user based on a specified annotation method; and a region determination sub-module, configured to use the image region within the bounding box of the object on the sample image as the region of interest.

[0020] In some alternative embodiments, the device further includes: an output module, configured to output a prompt message; wherein the prompt message is used to prompt the user to annotate the object of interest on the sample image based on the specified annotation method; the object determination sub-module includes: an object determination unit, configured to determine, on the sample image, the object annotated by the user based on the prompt message.

[0021] In some alternative embodiments, the initial neural network includes a backbone network, a first branch network, and a second branch network; wherein the backbone network is configured to perform feature extraction on the input image, the first branch network is configured to determine the predicted mask map, and the second branch network is configured to determine the predicted matte.

[0022] In some alternative embodiments, the device further includes: a ground-truth mask map determination module, configured to determine a ground-truth mask map corresponding to the region of interest based on the ground-truth matte; a second training module, configured to train the backbone network and the first branch network based on the difference between the predicted mask map and the ground-truth mask map, and after the training is completed, control the predicted matte determination module to determine, based on the sample image annotated with the region of interest and the predicted mask map, the predicted matte corresponding to the region of interest output by the initial neural network; the first training module includes: a training sub-module, configured to train the backbone network and the second branch network based on the difference between the predicted matte and the ground-truth matte.

[0023] In some alternative embodiments, the predicted matte determination module includes: a predicted matte determination sub-module, configured to perform transparency prediction on the foreground pixel points in the region of interest based on the sample image annotated with the region of interest and the predicted mask map, and determine the predicted matte.

[0024] In some alternative embodiments, the device further includes: a feature information extraction module, configured to perform at least one feature information extraction respectively in the backbone network and the second branch network; wherein the feature information extraction is used to extract the feature information corresponding to multiple different convolutional kernel sizes in the feature map corresponding to the region of interest; and a fusion module, configured to fuse the feature information corresponding to multiple different convolutional kernel sizes extracted each time to determine the fused feature information.

[0025] In some alternative embodiments, the fusion module includes: a fusion sub-module configured to multiply each time the extracted feature information corresponding to a plurality of different convolutional kernel sizes by a plurality of adaptive weight values respectively and then sum them to obtain the fused feature information.

[0026] According to a fourth aspect of the embodiments of the present disclosure, there is provided a matting device, including: a second region determination module configured to determine a region of interest on an image to be processed; wherein the region of interest is an image region where an object of interest to the user is located; a matting module configured to input the image to be processed with the region of interest marked into a target neural network for matting to obtain a predicted matting mask for matting the object output by the target neural network; wherein the target neural network is trained by using the method according to any one of the first aspects described above.

[0027] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program for executing the neural network training method according to any one of the first aspects or the matting method according to the second aspect described above.

[0028] According to a sixth aspect of the embodiments of the present disclosure, there is provided a neural network training device, including: a processor; a memory for storing executable instructions executable by the processor; wherein the processor is configured to call the executable instructions stored in the memory to implement the neural network training method according to any one of the first aspects.

[0029] According to a seventh aspect of the embodiments of the present disclosure, there is provided a matting device, including: a processor; a memory for storing executable instructions executable by the processor; wherein the processor is configured to call the executable instructions stored in the memory to implement the matting method according to the second aspect.

[0030] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0031] In the embodiments of the present disclosure, an image region where an object of interest to the user is located, i.e., a region of interest, can be determined on a sample image. By using the sample image with the region of interest marked, a predicted mask image corresponding to the region of interest is determined, and then based on the sample image with the region of interest marked and the predicted mask image, a predicted matting mask for the region of interest is jointly determined. Based on the true mask corresponding to the region of interest, an initial neural network is supervised and trained to obtain a target neural network for matting. The present disclosure can select the region of interest where the object of interest to the user is located as prior information, thereby effectively saving the time for annotating the sample image and improving the matting quality of the trained target neural network.

[0032] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0034] Figure 1 is a flowchart of a neural network training method shown according to an exemplary embodiment of the present disclosure;

[0035] Figure 2A is a flowchart of another neural network training method shown according to an exemplary embodiment of the present disclosure;

[0036] Figure 2B is a schematic diagram of a labeling method shown according to an exemplary embodiment of the present disclosure;

[0037] Figure 3 is a flowchart of another neural network training method shown according to an exemplary embodiment of the present disclosure;

[0038] Figure 4 is a schematic diagram of a neural network structure shown according to an exemplary embodiment of the present disclosure;

[0039] Figure 5 is a flowchart of another neural network training method shown according to an exemplary embodiment of the present disclosure;

[0040] Figure 6A is a schematic diagram of a network structure during multi-scale feature extraction shown according to an exemplary embodiment of the present disclosure;

[0041] Figure 6B is a flowchart of another neural network training method shown according to an exemplary embodiment of the present disclosure;

[0042] Figure 7A is a schematic diagram of an initial neural network structure shown according to an exemplary embodiment of the present disclosure;

[0043] Figure 7B is a schematic diagram of another initial neural network structure shown according to an exemplary embodiment of the present disclosure;

[0044] Figure 8 is a flowchart of a matting method shown according to an exemplary embodiment of the present disclosure;

[0045] Figure 9 is a block diagram of a neural network training apparatus shown according to an exemplary embodiment of the present disclosure;

[0046] Figure 10 is a block diagram of a matting device shown according to an exemplary embodiment of the present disclosure;

[0047] Figure 11 is a schematic structural diagram of a neural network training device shown according to an exemplary embodiment of the present disclosure;

[0048] Figure 12 is a schematic structural diagram of a matting device shown according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0049] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0050] The terms used in the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0051] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0052] In the current matting method, an accurate trimap of the object to be matted needs to be provided during the training stage, resulting in a long annotation process. If no trimap is provided, it may lead to an inability to accurately determine the object to be matted, and the matting effect will also decline. To solve this technical problem, the present disclosure provides the following neural network training method and matting method.

[0053] For example Figure 1 as shown Figure 1 is a neural network training method shown according to an exemplary embodiment. This method can be used for a training platform and includes the following steps 101 to 104:

[0054] In step 101, on the sample image, determine the region of interest.

[0055] In the embodiments of the present disclosure, the region of interest refers to the image region where the object of interest to the user is located. Among them, the object of interest to the user may refer to the object that the user expects to perform matte extraction on the image. Here, the object may include, but is not limited to, buildings, people, animals, various items, plants, vehicles, etc.

[0056] In the embodiments of the present disclosure, the number of sample images may be one or more, and the number of regions of interest on each sample image may be one or more. The present disclosure does not limit this.

[0057] In step 102, input the sample image marked with the region of interest into the initial neural network, and the initial neural network determines the predicted matte map corresponding to the region of interest.

[0058] In the embodiments of the present disclosure, the predicted matte map is used to identify foreground pixel points and / or background pixel points in the region of interest. Among them, foreground pixel points are pixel points belonging to the foreground part, and background pixel points are pixel points belonging to the background part.

[0059] In step 103, based on the sample image marked with the region of interest and the predicted matte map, determine the predicted matte extraction mask corresponding to the region of interest output by the initial neural network.

[0060] In the embodiments of the present disclosure, the predicted matte extraction mask refers to an image mask that can perform matte extraction on the object of interest to the user in the region of interest.

[0061] In step 104, based on the difference between the predicted matte extraction mask and the true mask corresponding to the region of interest, train the initial neural network. After the training is completed, obtain the target neural network for matte extraction.

[0062] In the embodiments of the present disclosure, the initial neural network can be supervised and trained based on the true mask corresponding to the region of interest. When the difference between the predicted matte extraction mask and the true mask satisfies the specified difference range, it is determined that the training of the initial neural network is completed, and at this time, the target neural network is obtained.

[0063] In the above embodiments, the region of interest where the object of interest to the user is located can be selected as the prior information, thereby effectively saving the time for annotating the sample image and improving the matte extraction quality of the trained target neural network.

[0064] In some optional embodiments, for example Figure 2A as shown, the above step 101 may include the following steps 201 to 202:

[0065] In step 201, on the sample image, determine the object annotated by the user based on the specified annotation method.

[0066] In the embodiments of the present disclosure, the specified annotation method includes but is not limited to at least one of the following: foreground point annotation method, line drawing annotation method, bounding box annotation method, foreground point-background point annotation method, and trisecting icon annotation method.

[0067] Among them, the foreground point annotation method refers to the user marking the foreground pixel points belonging to the object of interest on the sample image. The line drawing annotation method refers to the user drawing a line on the sample image for the object of interest. The bounding box annotation method refers to the user marking the bounding box of the object of interest on the sample image. The foreground point-background point annotation method refers to the user marking the foreground pixel points belonging to the object of interest, and / or marking the background pixel points corresponding to the object on the sample image. The trisecting icon annotation method refers to marking the trisecting icon corresponding to the object of interest on the sample image.

[0068] Specifically, the bounding box annotation method, the trisecting icon annotation method, and the line drawing annotation method can be referred to Figure 2B as shown.

[0069] In step 202, use the image area within the bounding box of the object on the sample image as the region of interest.

[0070] In the embodiments of the present disclosure, the training platform can determine the bounding box of the object based on the object annotated by the user, and then use the image area within the bounding box of the object on the sample image as the region of interest.

[0071] In the above embodiments, the object of interest to the user can be accurately determined, avoiding the problem that when there are multiple objects in the sample image, the object to be matte cannot be accurately determined, resulting in the training result not meeting the training requirements, and the usability is high.

[0072] In some alternative embodiments, for example Figure 3 as shown, the above neural network training method may further include the following step 100:

[0073] In step 100, output a prompt message.

[0074] In the embodiments of the present disclosure, the prompt message is used to prompt the user to annotate the object of interest on the sample image based on the specified annotation method.

[0075] Correspondingly, the above step 201 can be specifically:

[0076] On the sample image, determine the object annotated by the user based on the prompt message.

[0077] In the above embodiments, the object of interest to the user can be determined by interacting with the user, effectively solving the problem of long time consumption in the annotation process, expanding the time-domain scenario of image matting, improving the matting performance, and providing a better user experience.

[0078] In some alternative embodiments, the network architecture of the initial neural network, for example Figure 4 as shown, may include: a backbone network, a first branch network, and a second branch network.

[0079] Among them, the backbone network may adopt but is not limited to network architectures such as GoogleNet, Visual Geometry Group (VGG) network, and Residual network.

[0080] In the embodiments of the present disclosure, the backbone network is used to extract features from the input image, here referring to the input sample image annotated with the region of interest, to determine the feature map. The first branch network is used to perform binary mask prediction on the foreground region to determine the predicted mask map, and the second branch network is used to perform transparency prediction, boundary correction, etc. on the foreground region to determine the predicted matting mask.

[0081] Based on Figure 4 the network architecture shown, for example Figure 5 as shown, the above neural network training method may further include the following steps 105 to 106:

[0082] In step 105, based on the true mask, determine the true mask map corresponding to the region of interest.

[0083] In the embodiments of the present disclosure, the foreground pixel points on the true mask, that is, the pixel points with non-zero pixel values, can be determined, and the pixel values of the foreground pixel points are uniformly set to a preset value, which can be 1, and the pixel values of the background pixel points on the true mask are 0, so that the true mask map can be obtained.

[0084] In step 106, based on the difference between the predicted mask map and the true mask map, train the backbone network and the first branch network.

[0085] In the embodiments of the present disclosure, the backbone network and the first branch network can be trained with the true mask map as the supervision. Adjust the network parameters of the backbone network and the first branch network. When the difference between the predicted mask map and the true mask map is less than or equal to the preset difference, it is determined that the training of the backbone network and the first branch network is completed.

[0086] Among them, the cross entropy between the predicted mask map and the ground truth mask map can be calculated, and the cross entropy is used as the difference between the predicted mask map and the ground truth mask map.

[0087] Further, after the backbone network and the first branch network are trained, the above step 103 can be continued.

[0088] Correspondingly, the above step 104 can be specifically:

[0089] Based on the difference between the predicted matte and the ground truth matte, the backbone network and the second branch network are trained.

[0090] In the above embodiment, the matte extraction process can be decoupled, split into a foreground region division and a subsequent transparency prediction process, and trained step by step, so as to improve the matte extraction performance of the finally trained target neural network, and the usability is high.

[0091] In some alternative embodiments, the above step 103 can be specifically:

[0092] Based on the sample image with the region of interest annotated and the predicted mask map, the transparency of the foreground pixel points in the region of interest is predicted to determine the predicted matte.

[0093] In the embodiments of the present disclosure, considering that there are large distribution differences in the pixel values of the foreground pixel points of transparent objects and non-transparent objects, among which, the pixel value distribution of the foreground pixel points belonging to transparent objects is relatively uniform, while the pixel value distribution of the foreground pixel points belonging to non-transparent objects is concentrated near 0 or 255. In view of this distribution difference, the technical solution decouples the matte extraction task, and splits the matte extraction task into the above foreground region segmentation and transparency prediction tasks.

[0094] After obtaining the transparency prediction result, the transparency prediction result is used as the predicted matte.

[0095] In the above embodiment, the matte extraction process is decoupled, split into a foreground region division and a subsequent transparency prediction process, and different tasks are solved step by step, so as to improve the matte extraction performance of the finally trained target neural network, and the usability is high.

[0096] In some alternative embodiments, considering the problem that the sizes of the objects concerned by users are inconsistent, the present disclosure proposes that multi-scale feature fusion can be used to dynamically perform multi-scale feature selection. When performing multi-scale feature fusion, the corresponding neural network structure can be, for example Figure 6A As shown above, the above neural network training method, for example Figure 6B As shown, may further include the following steps 107 to 108:

[0097] In step 107, at least one feature information extraction is respectively performed in the backbone network and the second branch network.

[0098] In the embodiments of the present disclosure, the feature information extraction is used to extract the feature information corresponding to multiple different convolutional kernel sizes in the feature map corresponding to the region of interest. Specifically, pooling operations with different convolutional kernel sizes can be performed to obtain the feature information corresponding to multiple different convolutional kernel sizes. Among them, the pooling operation includes but is not limited to the maximum pooling operation or the average pooling operation.

[0099] For example Figure 6A As shown, pooling operations with convolutional kernel sizes of the entire region, one-fourth region, one-ninth region, and one-sixteenth region are respectively performed on the feature map corresponding to the region of interest.

[0100] In step 108, the feature information corresponding to multiple different convolutional kernel sizes extracted each time is fused to determine the fused feature information.

[0101] In the embodiments of the present disclosure, the feature information corresponding to multiple different convolutional kernel sizes can be multiplied by multiple adaptive weight values respectively and then added to obtain the fused feature information.

[0102] Among them, the adaptive weight value is the weight value determined by the initial neural network through adaptive learning and can be determined by the softmax activation function. Referring to Figure 6A As shown, the feature information corresponding to multiple different convolutional kernel sizes can be multiplied by the adaptive weight values w0, w1, w2, and w3 respectively and then added to obtain the fused feature information of this region of interest.

[0103] In the above embodiments, feature fusion can be adaptively performed at different scales, effectively focusing on global information and local detail information, improving the prediction ability of the final trained target neural network for the pixel transparency of the region where the object to be matte is located, and at the same time improving the matte prediction ability of the boundary and local details of the object to be matte.

[0104] In some alternative embodiments, for example Figure 7A As shown, an overall flowchart of the neural network training process is provided.

[0105] First, prompt information can be output to allow the user to label the object of interest on the sample image based on the specified annotation method, eliminating the ambiguity when there are multiple objects in the sample image.

[0106] Secondly, foreground region segmentation is performed through the backbone network and the first branch network, and binary mask prediction is performed on the foreground region to determine the predicted mask map corresponding to the region of interest.

[0107] Further, based on the sample image with the region of interest annotated and the predicted mask image, predict the transparency values of the foreground pixel points in the region of interest, complete the boundary modification, and determine the predicted matte mask.

[0108] Based on the difference between the predicted matte mask and the ground truth mask corresponding to the region of interest, train the initial neural network. After the training is completed, obtain the target neural network for matte extraction.

[0109] The network architecture of the initial neural network can be, for example Figure 7B as shown.

[0110] First, annotate the region of interest.

[0111] Input the RGB sample image and output the RGB sample image with the region of interest annotated by the user.

[0112] Specifically, the user determines the region of interest based on the RGB sample image and the matte extraction requirement and the prompt information. Among them, the user can use a specified annotation method to annotate the object of interest, and the specified annotation method includes but is not limited to at least one of the following: foreground point annotation method, line drawing annotation method, bounding box annotation method, foreground point - background point annotation method, and trisect icon annotation method.

[0113] Secondly, input the RGB sample image with the region of interest annotated into the initial neural network to obtain the predicted mask image corresponding to the region of interest.

[0114] Based on the RGB sample image and the region of interest, divide the foreground pixel points and background pixel points in the region of interest through the initial neural network to obtain the predicted mask image.

[0115] In this process, the ground truth mask corresponding to the region of interest can be determined, and the ground truth mask image can be calculated. Calculate the cross - entropy between the predicted mask image and the ground truth mask image, so as to train the backbone network and the first branch network. When the cross - entropy is less than or equal to the preset difference, it is determined that the training of the backbone network and the first branch network is completed.

[0116] Finally, input the RGB sample image with the region of interest annotated and the predicted mask image into the initial neural network to determine the output predicted matte mask.

[0117] In the embodiments of the present disclosure, based on the predicted mask image and combined with the RGB sample image with the region of interest annotated, predict the transparency of the foreground pixel points in the region of interest, and construct the regression prediction of the transparency value of each foreground pixel point.

[0118] Meanwhile, to address the issue of inconsistent object scales of interest to users, a multi-scale feature fusion module is proposed to dynamically perform multi-scale feature selection. Combining Figure 6A As shown, in the decoder part, before sampling at each stage, pooling operations based on different convolutional kernel sizes are used to determine feature information at different scales, and then adaptive weights are used for feature fusion. Based on the fused feature information, the transparency prediction result is finally determined to obtain the predicted matte.

[0119] In the embodiments of the present disclosure, the initial neural network can be supervised and trained based on the ground truth matte corresponding to the region of interest. When the difference between the predicted matte and the ground truth matte satisfies the specified difference range, it is determined that the training of the initial neural network is completed, and at this time, the target neural network is obtained.

[0120] In the above embodiments, the target neural network obtained based on the above training method adopts a dual decoder model, and effective matte results can be obtained for both opaque objects and transparent objects. In addition, the interactive annotation method provided by the present disclosure expands the temporal scenario of image matting and improves the matting performance of the target neural network. And through the multi-scale feature fusion module, global information and local detail information are effectively concerned, thereby ensuring high-quality matte results.

[0121] In some alternative embodiments, for example Figure 8 As shown, Figure 8 is a matting method shown according to an exemplary embodiment, which can be used in a machine device and includes the following steps 301 to 302:

[0122] In step 301, on the image to be processed, a region of interest is determined.

[0123] In the embodiments of the present disclosure, the region of interest is the image region where the object of interest to the user is located. The method for determining the region of interest is similar to the method for determining the region of interest in the above steps 201 to 202 and will not be elaborated here.

[0124] In step 302, the image to be processed with the region of interest marked is input into the target neural network for matting, and the predicted matte for matting the object output by the target neural network is obtained.

[0125] In the embodiments of the present disclosure, the target neural network is obtained by using any of the above neural network training methods.

[0126] In the above embodiments, after the target neural network is trained in the above manner, in the actual application process, the region of interest can also be determined first, and then the to-be-processed image marked with the region of interest is input into the trained target neural network to obtain the corresponding predicted matte. It can provide an efficient and high-quality solution for a variety of matte extraction tasks, including but not limited to portrait matte extraction, animal matte extraction, annotation work in processes such as movie features, etc., and provide a high-quality matte extraction tool for image and / or video editing software. Moreover, it can provide a high-quality foreground matte for the image marked with the bounding box, with high usability.

[0127] Corresponding to the foregoing method embodiments, the present disclosure also provides embodiments of an apparatus.

[0128] As Figure 9 shown, Figure 9 FIG. is a block diagram of a neural network training apparatus according to an exemplary embodiment of the present disclosure, including: a first region determination module 401, configured to determine a region of interest on a sample image; wherein, the region of interest is an image region where an object concerned by a user is located; a predicted mask map determination module 402, configured to input the sample image marked with the region of interest into an initial neural network, and the initial neural network determines a predicted mask map corresponding to the region of interest; wherein, the predicted mask map is used to identify foreground pixel points and / or background pixel points in the region of interest; a predicted matte determination module 403, configured to determine a predicted matte corresponding to the region of interest output by the initial neural network based on the sample image marked with the region of interest and the predicted mask map; a first training module 404, configured to train the initial neural network based on the difference between the predicted matte and the true matte corresponding to the region of interest, and after the training is completed, obtain a target neural network for matte extraction.

[0129] In some alternative embodiments, the region determination module includes: an object determination sub-module, configured to determine the object marked by the user based on a specified annotation method on the sample image; a region determination sub-module, configured to use the image region within the bounding box of the object on the sample image as the region of interest.

[0130] In some alternative embodiments, the apparatus further includes: an output module, configured to output a prompt message; wherein, the prompt message is used to prompt the user to mark the object of interest on the sample image based on the specified annotation method; the object determination sub-module includes: an object determination unit, configured to determine the object marked by the user based on the prompt message on the sample image.

[0131] In some alternative embodiments, the initial neural network includes a backbone network, a first branch network, and a second branch network; wherein, the backbone network is used to extract features from the input image, the first branch network is used to determine the predicted mask map, and the second branch network is used to determine the predicted matte.

[0132] In some alternative embodiments, the device further includes: a ground-truth mask map determination module, configured to determine a ground-truth mask map corresponding to the region of interest based on the ground-truth matte; a second training module, configured to train the backbone network and the first branch network based on the difference between the predicted mask map and the ground-truth mask map, and after the training is completed, control the predicted air conditioner matte determination module to determine the predicted matte corresponding to the region of interest output by the initial neural network based on the sample image with the region of interest annotated and the predicted mask map; the first training module includes: a training sub-module, configured to train the backbone network and the second branch network based on the difference between the predicted matte and the ground-truth matte.

[0133] In some alternative embodiments, the predicted matte determination module includes: a predicted matte determination sub-module, configured to predict the transparency of the foreground pixel points in the region of interest based on the sample image with the region of interest annotated and the predicted mask map, and determine the predicted matte.

[0134] In some alternative embodiments, the device further includes: a feature information extraction module, configured to perform at least one feature information extraction in the backbone network and the second branch network respectively; wherein, the feature information extraction is used to extract the feature information corresponding to multiple different convolutional kernel sizes in the feature map corresponding to the region of interest; a fusion module, configured to fuse the feature information corresponding to multiple different convolutional kernel sizes extracted each time to determine the fused feature information.

[0135] In some alternative embodiments, the fusion module includes: a fusion sub-module, configured to multiply the feature information corresponding to multiple different convolutional kernel sizes extracted each time by multiple adaptive weight values respectively and then sum them to obtain the fused feature information.

[0136] As Figure 10 shown, Figure 10The following is a block diagram of a matte extraction device shown according to an exemplary embodiment of the present disclosure, including: a second region determination module 501, configured to determine a region of interest on an image to be processed; wherein the region of interest is an image region where an object of interest to the user is located; a matte extraction module 502, configured to input the image to be processed with the region of interest marked into a target neural network for matte extraction, and obtain a predicted matte mask for extracting the object output by the target neural network; wherein the target neural network is trained by using the neural network training method described in any one of the above.

[0137] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present disclosure. A person of ordinary skill in the art can understand and implement it without creative work.

[0138] The embodiment of the present disclosure also provides a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is used to execute the neural network training method or the matte extraction method described in any one of the above.

[0139] In some optional embodiments, the embodiment of the present disclosure provides a computer program product, including computer-readable code. When the computer-readable code runs on a device, a processor in the device executes instructions for implementing the neural network training method or the matte extraction method provided in any one of the above embodiments.

[0140] In some optional embodiments, the embodiment of the present disclosure also provides another computer program product, used to store computer-readable instructions, and when the instructions are executed, the computer executes the neural network training method or the matte extraction method provided in any one of the above embodiments.

[0141] The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0142] An embodiment of the present disclosure further provides a neural network training device, including: a processor; a memory for storing executable instructions executable by the processor; wherein, the processor is configured to call the executable instructions stored in the memory to implement the neural network training method described in any one of the above.

[0143] An embodiment of the present disclosure further provides a matte extraction device, including: a processor; a memory for storing executable instructions executable by the processor; wherein, the processor is configured to call the executable instructions stored in the memory to implement the matte extraction method described in any one of the above.

[0144] Figure 11 FIG. 6 is a schematic hardware structure diagram of a neural network training device provided by an embodiment of the present disclosure. The neural network training device 610 includes a processor 611, and may further include an input device 612, an output device 613, and a memory 614. The input device 612, the output device 613, the memory 614, and the processor 611 are interconnected through a bus.

[0145] The memory includes, but is not limited to, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM). The memory is used for storing relevant instructions and data. The input device is used for inputting data and / or signals, and the output device is used for outputting data and / or signals. The output device and the input device may be independent devices or an integrated device.

[0146] The processor may include one or more processors, for example, including one or more central processing units (CPUs). In the case where the processor is a single CPU, the CPU may be a single-core CPU or a multi-core CPU.

[0147] The memory is used for storing program codes and data of the network device.

[0148] The processor is used for calling the program codes and data in the memory to execute the neural network training steps in the above method embodiments. For specific reference, see the description in the method embodiments, which will not be elaborated here.

[0149] It can be understood that Figure 11Only a simplified design of a neural network training device is shown. In practical applications, the neural network training device may also separately include other necessary components, including but not limited to any number of input / output devices, processors, controllers, memories, etc., and all neural network training devices that can implement the embodiments of the present disclosure are within the protection scope of the present disclosure.

[0150] Figure 12 The following is a schematic diagram of the hardware structure of a matte extraction device provided by an embodiment of the present disclosure. The matte extraction device 710 includes a processor 711, and may also include an input device 712, an output device 713, and a memory 714. The input device 712, the output device 713, the memory 714, and the processor 711 are interconnected with each other through a bus.

[0151] The memory includes but is not limited to a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM). The memory is used for storing relevant instructions and data. The input device is used for inputting data and / or signals, and the output device is used for outputting data and / or signals. The output device and the input device may be independent devices or an integrated device.

[0152] The processor may include one or more processors, for example, including one or more central processing units (CPUs). In the case where the processor is a single CPU, the CPU may be a single-core CPU or a multi-core CPU.

[0153] The memory is used for storing the program code and data of the network device.

[0154] The processor is used for calling the program code and data in the memory and executing the matte extraction steps in the above method embodiments. For specific details, reference may be made to the descriptions in the method embodiments, which will not be elaborated herein.

[0155] It can be understood that Figure 12 Only a simplified design of a matte extraction device is shown. In practical applications, the matte extraction device may also separately include other necessary components, including but not limited to any number of input / output devices, processors, controllers, memories, etc., and all matte extraction devices that can implement the embodiments of the present disclosure are within the protection scope of the present disclosure.

[0156] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known or customary technical means in the art not disclosed herein. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0157] The foregoing is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A neural network training method, characterized in that, The method includes: On a sample image, determine a region of interest; wherein, the region of interest is an image region where the object of interest to the user is located; Input the sample image marked with the region of interest into an initial neural network, and the initial neural network determines a predicted mask map corresponding to the region of interest; wherein, the predicted mask map is used to identify foreground pixel points and / or background pixel points in the region of interest; Based on the sample image marked with the region of interest and the predicted mask map, determine a predicted matte mask corresponding to the region of interest output by the initial neural network; Based on the difference between the predicted matte mask and the true matte corresponding to the region of interest, train the initial neural network, and after the training is completed, obtain a target neural network for matte extraction; The initial neural network includes a backbone network, a first branch network, and a second branch network; wherein, the backbone network is used to extract features from the input image, the first branch network is used to determine the predicted mask map, and the second branch network is used to determine the predicted matte mask; Based on the true matte, determine a true mask map corresponding to the region of interest; Based on the difference between the predicted mask map and the true mask map, train the backbone network and the first branch network, and after the training is completed, perform the step of determining the predicted matte mask corresponding to the region of interest output by the initial neural network based on the sample image marked with the region of interest and the predicted mask map; The training of the initial neural network based on the difference between the predicted matte mask and the true matte image corresponding to the region of interest includes: Based on the difference between the predicted matte mask and the true matte, train the backbone network and the second branch network.

2. The method according to claim 1, characterized in that The determining of the region of interest on the sample image includes: On the sample image, determine the object marked by the user based on a specified annotation method; Take the image region within the bounding box of the object on the sample image as the region of interest.

3. The method according to claim 2, wherein The method further includes: Output a prompt message; wherein, the prompt message is used to prompt the user to mark the object of interest on the sample image based on the specified annotation method; The determining of the object marked by the user based on the specified annotation method on the sample image includes: On the sample image, determine the object marked by the user based on the prompt message.

4. The method according to claim 1, wherein The determining of the predicted matte mask corresponding to the region of interest output by the initial neural network based on the sample image marked with the region of interest and the predicted mask map includes: Based on the sample image marked with the region of interest and the predicted mask map, predict the transparency of the foreground pixel points in the region of interest to determine the predicted matte mask.

5. The method according to claim 1 or 4, characterized in that, The method further includes: At least one feature information extraction is respectively performed in the backbone network and the second branch network; wherein, the feature information extraction is used to extract the feature information corresponding to multiple different convolutional kernel sizes in the feature map corresponding to the region of interest. Fuse the feature information corresponding to multiple different convolutional kernel sizes extracted each time to determine the fused feature information.

6. The method according to claim 5, characterized in that, The fusing the feature information corresponding to multiple different convolutional kernel sizes extracted each time to determine the fused feature information includes: Multiply the feature information corresponding to multiple different convolutional kernel sizes extracted each time by multiple adaptive weight values respectively and then sum them to obtain the fused feature information.

7. A matte extraction method, characterized in that, The method includes: On the image to be processed, determine the region of interest; wherein, the region of interest is the image region where the object concerned by the user is located. Input the image to be processed marked with the region of interest into the target neural network for matte extraction to obtain the predicted matte mask for matte extraction of the object output by the target neural network; wherein, the target neural network is trained by using the method according to any one of claims 1-6.

8. A neural network training device, characterized in that, Includes: A first region determination module, configured to determine a region of interest on a sample image; wherein, the region of interest is the image region where the object concerned by the user is located. A predicted mask map determination module, configured to input the sample image marked with the region of interest into an initial neural network, and the initial neural network determines a predicted mask map corresponding to the region of interest; wherein, the predicted mask map is used to identify foreground pixel points and / or background pixel points in the region of interest. A predicted matte mask determination module, configured to determine a predicted matte mask corresponding to the region of interest output by the initial neural network based on the sample image marked with the region of interest and the predicted mask map. A first training module, configured to train the initial neural network based on the difference between the predicted matte mask and the true matte corresponding to the region of interest, and after the training is completed, obtain a target neural network for matte extraction. The initial neural network includes a backbone network, a first branch network and a second branch network; wherein, the backbone network is used to extract features from the input image, the first branch network is used to determine the predicted mask map, and the second branch network is used to determine the predicted matte mask. A true mask map determination module, configured to determine a true mask map corresponding to the region of interest based on the true matte. A second training module, configured to train the backbone network and the first branch network based on the difference between the predicted mask map and the true mask map, and after the training is completed, control the predicted matte mask determination module to determine a predicted matte mask corresponding to the region of interest output by the initial neural network based on the sample image marked with the region of interest and the predicted mask map. The first training module includes: a training sub-module, configured to train the backbone network and the second branch network based on the difference between the predicted matte mask and the true matte.

9. A matte extraction device, characterized in that, Includes: A second region determination module, configured to determine a region of interest on an image to be processed; wherein the region of interest is an image region where an object of interest to the user is located; A matte extraction module, configured to input the image to be processed with the region of interest marked into a target neural network for matte extraction, and obtain a predicted matte for extracting the object output by the target neural network; wherein the target neural network is trained by using the method according to any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the neural network training method according to any one of claims 1-6 or the matte extraction method according to claim 7.

11. A neural network training device, characterized in that, Comprising: A processor; A memory for storing executable instructions executable by the processor; Wherein the processor is configured to call the executable instructions stored in the memory to implement the neural network training method according to any one of claims 1-6.

12. A matte extraction device, characterized in that, Comprising: A processor; A memory for storing executable instructions executable by the processor; Wherein the processor is configured to call the executable instructions stored in the memory to implement the matte extraction method according to claim 7.

Citation Information

Patent Citations

  • Image instance segmentation method and device

    CN110705558A

  • Full-automatic portrait mask matting method and system

    CN111223106A