Foreground object acquisition method and apparatus, electronic device, and storage medium
By determining and selecting the appropriate foreground mask image, the problems of missing segmentation and missegment in the cutout method are solved, and the cutout quality is improved, especially for sticker images with pixel style or stroke style.
Patent Information
- Application Number
- PCT/CN2024/128635
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
The existing cutout method has problems such as missing segmentation, poor segmentation quality, and missed segmentation when processing sticker images, resulting in poor cutout effect.
By determining the first foreground mask image, the second foreground mask image and the third foreground mask image corresponding to the original image, the target foreground mask image is selected according to the mask image attributes and edge features to obtain the foreground objects in the original image to avoid the situation where the edge lines are not closed and missegmented.
Improve the quality of cutouts in the foreground object with the same color as the background, reduce the problem of mask prediction disbelief at edges, corners, etc., and improve the cutout effect.
Smart Images

Figure CN2024128635_08052025_PF_FP_ABST
Abstract
Description
Method, device, electronic device and storage medium for acquiring foreground object
[0001] This application claims priority to the Chinese invention patent application filed on October 31, 2023, with application number 202311435089.X and titled “A method, device, electronic device and storage medium for acquiring foreground objects”. The entire contents of the Chinese patent application are incorporated into this application by reference. Technical Field
[0002] The embodiments of the present disclosure relate to data processing technology, and more particularly to a method, device, electronic device, and storage medium for acquiring a foreground object. Background Art
[0003] With the development of mobile terminals and Internet technology, more and more application services are installed on mobile terminals to provide users with rich interactive experiences. For example, users can enrich the display of pictures or video frames through stickers.
[0004] Sticker images can be RGB images generated using image generation algorithms. Because the foreground object in sticker images is blended into a solid background, they cannot be used directly as stickers. Typically, sticker images are processed using a cutout method to obtain a mask of the foreground object. This mask can then be used to cut out the sticker image to create a sticker, which can then be overlaid on other materials. However, existing cutout methods suffer from issues such as missed segmentation, poor segmentation quality, and mis-segmentation, which can compromise the effectiveness of cutouts.
[0005] Summary of the Invention
[0006] The embodiments of the present disclosure provide a foreground object acquisition method, device, electronic device, and storage medium, which can solve the problem of poor cutout effect in related cutout methods.
[0007] In a first aspect, an embodiment of the present disclosure provides a foreground object acquisition method, comprising: determining a first foreground mask map, a second foreground mask map, and a third foreground mask map corresponding to an original image, wherein the first foreground mask map and the second foreground mask map represent masks of the foreground area of the original image generated using different models, and the third foreground mask map represents a mask of the foreground area of the predicted mask map of the original image; determining the first foreground mask map or the third foreground mask map as an alternative foreground mask map according to mask map attributes; determining the alternative foreground mask map or the second foreground mask map as a target foreground mask map according to edge features of the alternative foreground mask map and difference features between the alternative foreground mask map and the second foreground mask map; and acquiring the foreground object in the original image according to the target foreground mask map.
[0008] In the second aspect, an embodiment of the present disclosure also provides a foreground object acquisition device, which includes: a mask image determination module, used to determine a first foreground mask image, a second foreground mask image and a third foreground mask image corresponding to the original image, wherein the first foreground mask image and the second foreground mask image represent the mask of the foreground area of the original image generated using different models, and the third foreground mask image represents the mask of the foreground area of the predicted mask image of the original image; an alternative mask image determination module, used to determine the first foreground mask image or the third foreground mask image as the alternative foreground mask image according to the mask image attributes; a target mask image determination module, used to determine the alternative foreground mask image or the second foreground mask image as the target foreground mask image according to the edge features of the alternative foreground mask image and the difference features between the alternative foreground mask image and the second foreground mask image; an object acquisition module, used to acquire the foreground object in the original image according to the target foreground mask image.
[0009] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the foreground object acquisition method as described in any embodiment of the present disclosure.
[0010] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the foreground object acquisition method as described in any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0012] FIG1 is a schematic flow chart of a method for acquiring a foreground object according to an embodiment of the present disclosure;
[0013] FIG2 is a schematic diagram of an edge feature map provided by an embodiment of the present disclosure;
[0014] FIG3 is a schematic diagram of a difference area provided by an embodiment of the present disclosure;
[0015] FIG4 is a schematic diagram of a difference region edge map provided by an embodiment of the present disclosure;
[0016] FIG5 is a schematic diagram of a third image provided by an embodiment of the present disclosure;
[0017] FIG6 is a schematic diagram of a fourth image provided by an embodiment of the present disclosure;
[0018] FIG7 is a flow chart of another method for acquiring a foreground object provided by an embodiment of the present disclosure;
[0019] FIG8 is a schematic diagram of a process for generating a foreground mask for an object with a hollow structure, provided by an embodiment of the present disclosure;
[0020] FIG9 is a schematic structural diagram of a foreground object acquisition device provided by an embodiment of the present disclosure; and
[0021] FIG10 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0023] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0024] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0025] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0026] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0027] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0028] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0029] Figure 1 is a flow chart of a foreground object acquisition method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation of clipping stickers generated in pixel style or stroke style. The method can be executed by a foreground object acquisition device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, PC or server, etc.
[0030] As shown in FIG1 , the method includes the following steps.
[0031] S110 : Determine a first foreground mask image, a second foreground mask image, and a third foreground mask image corresponding to the original image.
[0032] The first foreground mask map and the second foreground mask map represent masks of the foreground area of the original image generated using different models, the third foreground mask map represents the mask of the foreground area of the predicted mask map of the original image, and the predicted mask map represents the mask of the foreground object of the original image predicted by the subject segmentation model. The subject segmentation model is a pre-trained model that obtains the foreground mask of the original image.
[0033] The original image may be an image to be cutout. In the disclosed embodiment, the original image may be a sticker image generated in pixel style or stroke style. The sticker image in pixel style or stroke style has the following characteristics: First, the sticker image in pixel style or stroke style has a solid background. For example, the background color is white. Second, the foreground area of the sticker will not cling to the edge of the sticker image. Third, the sticker image can be divided into color block areas according to the color blocks, and the color values in each color block area are relatively uniform and fixed, and there is almost no gradient color on the screen.
[0034] Exemplarily, an erosion operation is performed on the foreground object of the original image to obtain a first image, and a first model is used to segment the foreground object and background area in the first image to obtain a first foreground mask image. A foreground selection box is determined in the original image, and a second model is used to segment the foreground object and background area in the foreground selection box to obtain a second foreground mask image. The original image is adjusted according to the binarized predicted mask image, and an erosion operation is performed on the foreground object in the adjusted original image to obtain a second image, and the first model is used to segment the foreground object and background area in the second image to obtain a third foreground mask image.
[0035] Among them, the edge line of the foreground object can be the outline of the foreground object in the original image. For pixel-style or stroke-style sticker images, the edge line of the foreground object is composed of multiple line segments, and some edge lines may not be closed, which may cause the color block area corresponding to the unclosed area to be missed. By performing an erosion operation on the foreground object to obtain the first image, the situation where the edge line of the foreground object is not closed can be optimized to a certain extent. The difference between the first image and the original image is that some edge lines of the foreground object are changed from an unclosed state to a closed state. The first foreground mask image represents the foreground mask of the first image. By fusing the first foreground mask image to the original image, the foreground object in the original image can be obtained.
[0036] In some embodiments, the foreground object of the original image is subjected to an erosion operation to obtain a first image, including: performing an erosion operation on the foreground object of the original image to thicken the edge line of the foreground object, thereby causing the unclosed area in the edge line of the foreground object to become closed, thereby obtaining the first image. For example, the erosion operation can be implemented by setting a convolution kernel to perform a convolution operation on the foreground object of the original image. Since the original image is a pixel-style or stroke-style sticker image, there may be some poor closure at the edge stroke. For the sticker generation image generated for a white background, the foreground color brightness is generally lower than the background color. An erosion operation can be performed on the foreground object, and the erosion size can be a set number of pixels, so as to improve the situation where the foreground edge is not closed to a certain extent. By performing an erosion operation on the foreground object, the black dark area near the edge line is enlarged, and the white bright area near the edge line is reduced, thereby causing the unclosed area of the edge line of the foreground object to become closed.
[0037] It should be noted that for areas with sufficiently small edge stroke intervals, the areas not well enclosed by the edge strokes can be enclosed by the erosion operation. If the edge stroke intervals are relatively large, the unenclosed areas may not be enclosed by the erosion operation.
[0038] In some embodiments, using a first model to segment a foreground object and a background area in the first image to obtain a first foreground mask image includes: determining a first reference point in the first image and obtaining adjacent pixels corresponding to the first reference point in the first image. The reference point is a starting point for traversal of the first image. For example, each of the four vertices of the first image can be used as the first reference point. Starting from the first reference point, the first image can be traversed along a set direction to obtain adjacent pixels of the first reference point. If the brightness difference between the adjacent pixel and the first reference point is less than a set threshold, the adjacent pixel is used as the new first reference point, and the first image is traversed along the corresponding direction until the brightness difference between the adjacent pixel and the reference point is equal to or greater than the set threshold. The pixel traversal operation in the corresponding direction is then stopped. It should be noted that the set direction can be any of several directions extending outward from the first reference point. For example, the set direction can be any of the four directions, up, down, left, and right, centered on the first reference point. This disclosure does not specifically limit the meaning of the set direction. The set threshold serves as a traversal stopping condition and can be related to the outline of the foreground object in a practical application, so that pixel traversal in each direction stops at the edge of the foreground object, thereby obtaining a background mask.
[0039] A first background mask image is determined based on adjacent pixels whose brightness difference with the first reference point is less than a set threshold. The first background mask image is color-inverted to obtain a first foreground mask image. Since the background area in the first background mask image is white and the foreground object area is black, the color inversion causes the background area of the first background mask image to be black and the foreground object area to be white. Thus, the white area is removed from the original image to obtain the first foreground mask image.
[0040] In some embodiments, determining a foreground selection frame in the original image and segmenting the foreground object and background area within the foreground selection frame using a second model to obtain a second foreground mask image includes: reducing the size of the original image to reduce computational complexity, thereby accelerating model calculations and improving the efficiency of generating the second foreground mask image. The foreground selection frame is determined based on the edges of the reduced original image. For example, the foreground selection frame can be obtained by shrinking a reference frame formed by the outermost pixels of the reduced original image inward by a set number of pixels. The area outside the foreground selection frame is defined as the background area, and the area within the foreground selection frame is defined as the type-unknown area. A Gaussian mixture model is used to model pixel clustering, and an image segmentation algorithm is used to iterate the model a set number of times, optimizing classification with the goal of minimizing the iterative energy function. After node connection is completed, the maximum flow minimum cut theorem is used to segment the foreground object and background area within the foreground selection frame to obtain a result mask image. The result mask image is then scaled up based on the size of the original image to restore the size of the scaled result mask image to the size of the original image, thereby obtaining the second foreground mask image.
[0041] In some embodiments, adjusting the original image based on the predicted mask image includes: binarizing the predicted mask image based on a set threshold value, such that pixels in the predicted mask image having pixel values greater than the preset threshold value are set to 1, and pixels in the predicted mask image having pixel values equal to or less than the preset threshold value are set to 0, thereby obtaining a binarized result of the predicted mask image. Pixels in the original image corresponding to regions in the binarized result where the pixels are greater than 0 are set to a color that is significantly different from the background color of the original image, thereby obtaining an adjusted original image.
[0042] In some embodiments, performing an erosion operation on the foreground object in the adjusted original image to obtain a second image, and segmenting the foreground object and background area in the second image using the first model to obtain a third foreground mask image, includes: performing an erosion operation on the foreground object in the adjusted original image using a set convolution kernel to close the unclosed area of the edge line of the foreground object, thereby obtaining the second image. For example, erosion can be achieved by performing a convolution operation on the foreground object in the adjusted original image using a set convolution kernel. The erosion size can be a set number of pixels, thereby improving the situation where the foreground edge line is not closed to a certain extent.
[0043] In some embodiments, using the first model to segment the foreground object and background area in the second image to obtain a third foreground mask image includes: determining a second reference point in the second image and obtaining adjacent pixels corresponding to the second reference point in the second image. For example, the four vertices of the second image can be used as second reference points. Starting from the second reference point, the second image can be traversed along a set direction to obtain adjacent pixels of the second reference point. If the brightness difference between the adjacent pixel and the second reference point is less than a set threshold, the adjacent pixel is used as the new second reference point, and the second image is traversed in the corresponding direction until the brightness difference between the adjacent pixel and the reference point is equal to or greater than the set threshold. The pixel traversal operation in the corresponding direction is stopped. When the pixel traversal in all directions is stopped, a background mask in the second image is obtained. A second background mask image is determined based on the adjacent pixels whose brightness difference with the second reference point is less than the set threshold. The second background mask image is color-inverted to obtain the third foreground mask image.
[0044] S120: Determine, according to mask image attributes, whether the first foreground mask image or the third foreground mask image is a candidate foreground mask image.
[0045] The mask image attributes represent the attributes of the foreground mask region in the foreground mask image. For example, the mask image attributes may include the foreground mask area, etc.
[0046] Exemplarily, for the first foreground mask image or the third foreground mask image, a foreground mask area corresponding to the foreground mask region is determined as a mask image attribute; and the first foreground mask image or the third foreground mask image is determined as an alternative foreground mask image based on the foreground mask area.
[0047] For example, determine the area difference between the foreground mask area corresponding to the first foreground mask image and the foreground mask area corresponding to the third foreground mask image, determine the ratio of the area difference to the foreground mask area corresponding to the first foreground mask image, compare the ratio with a set threshold, and determine the first foreground mask image or the third foreground mask image as the alternative foreground mask image based on the comparison result.
[0048] The area increase ratio of the foreground mask of the third foreground mask image compared to the first foreground mask image can be calculated, and the degree of improvement of the overall background removal effect by the third foreground mask image can be represented by the area increase ratio. Then, the quality of the foreground mask can be improved by improving the overall background removal effect.
[0049] Since the above ratio represents the ratio of the increase in the area of the foreground mask of the third foreground mask image compared to the first foreground mask image, if the ratio is less than a set threshold, it is determined that combining the predicted mask image does not significantly improve background removal. In this case, there is a high risk that the segmentation quality of the third foreground mask image is poor, and the first foreground mask image is adopted as the alternative foreground mask image. Conversely, if the ratio is equal to or greater than the set threshold, it is determined that combining the predicted mask image significantly improves background removal. In this case, there is a low risk that the segmentation quality of the third foreground mask image is poor, and the third foreground mask image is adopted as the alternative foreground mask image.
[0050] S130 : Determine the candidate foreground mask image or the second foreground mask image as a target foreground mask image according to edge features of the candidate foreground mask image and difference features between the candidate foreground mask image and the second foreground mask image.
[0051] The edge feature represents the result of an opening operation on the candidate foreground mask image. The opening operation can be performed by first performing an erosion operation on the candidate foreground mask image and then performing a dilation operation on the erosion result. The difference feature represents the edge of a difference region where the pixel value in the candidate foreground mask image is smaller than that in the second foreground mask image.
[0052] Exemplarily, an opening operation is performed on the alternative foreground mask image to obtain an edge feature map of the alternative foreground mask image; a difference area in the alternative foreground mask image whose brightness is less than that of the second foreground mask image is determined, and a difference area edge map is determined based on the difference area; based on the alternative foreground mask image, the edge feature map and the difference area edge map, the alternative foreground mask image or the second foreground mask image is determined to be the target foreground mask image.
[0053] FIG2 is a schematic diagram of an edge feature map provided by an embodiment of the present disclosure. As shown in FIG2 , an opening operation is performed on a candidate foreground mask map 210 to obtain an edge feature map 220 corresponding to the candidate foreground mask map 210 .
[0054] Figure 3 is a schematic diagram of a difference region provided by an embodiment of the present disclosure. As shown in Figure 3, the region in the candidate foreground mask image 310 where the pixel value is smaller than the third foreground mask image 320 is calculated as a difference region image 330, and the white region in the difference region image 330 represents the difference region.
[0055] In some embodiments, determining the difference region edge map based on the difference region includes: performing a dilation operation on the difference region to obtain a difference dilation region; and determining the difference region edge map based on a difference between the difference dilation region and the difference region.
[0056] FIG4 is a schematic diagram of a difference region edge map provided by an embodiment of the present disclosure. As shown in FIG4 , a dilation operation is performed on a difference region in a difference region map 410 to expand the difference region to a set size, thereby obtaining a dilated difference region. A difference region edge map 420 is determined based on the difference between the dilated difference region and the difference region.
[0057] In some embodiments, the alternative foreground mask image or the second foreground mask image is determined to be the target foreground mask image based on the alternative foreground mask image, the edge feature map and the difference area edge map, including: determining a third image based on the intersection of the alternative foreground mask image and the difference area edge map; determining a fourth image based on the intersection of the edge feature map and the difference area edge map; and determining the alternative foreground mask image or the second foreground mask image to be the target foreground mask image based on the area of the third image and the area of the fourth image.
[0058] FIG5 is a schematic diagram of a third image provided by an embodiment of the present disclosure. As shown in FIG5 , the intersection of the candidate foreground mask image 510 and the difference area edge image 520 is used as the third image 530 .
[0059] FIG6 is a schematic diagram of a fourth image provided by an embodiment of the present disclosure. As shown in FIG6 , the intersection of the edge feature map 610 and the difference region edge map 620 is used as the fourth image 630 .
[0060] The area of the foreground object in the third image is calculated as the third image area, and the area of the foreground object in the fourth image is calculated as the fourth image area. Then, the area ratio between the third image area and the fourth image area is calculated, and based on the comparison result of the area ratio with a preset threshold, either the alternative foreground mask image or the second foreground mask image is determined as the target foreground mask image. For situations where the edges of the original image are not well closed or the subject segmentation has obvious hollowing problems, by selecting the target foreground mask image from the alternative foreground mask image and the second foreground mask image, and then using the target foreground mask image to obtain the foreground object in the original image, the problem of poor clipping effect when using the alternative foreground mask image to obtain the foreground object in the original image can be avoided.
[0061] S140: Acquire a foreground object in the original image according to the target foreground mask image.
[0062] Exemplarily, the target foreground mask image is superimposed on the original image so that the foreground mask in the target foreground mask image covers the foreground object in the original image, and then the foreground object is subtracted from the original image using the target foreground mask image.
[0063] The technical solution of the disclosed embodiment determines the first foreground mask image, the second foreground mask image and the third foreground mask image corresponding to the original image, and selects the first foreground mask image or the third foreground mask image as the alternative foreground mask image according to the mask image attributes, and then selects the alternative foreground mask image or the second foreground mask image as the target foreground mask image, and obtains the foreground object in the original image according to the target foreground mask image. It can avoid the situation where some foreground areas are missed due to the edge lines of the original image not being well closed, and can also avoid the situation where color block areas with different colors from the background are mistakenly segmented, thereby improving the clipping quality of the color block areas with the same color as the background in the foreground object, and reducing the problem of unreliable mask prediction at edges, corners and other positions.
[0064] FIG7 is a flow chart of another foreground object acquisition method provided by an embodiment of the present disclosure. Based on the above embodiment, the present disclosure embodiment additionally defines a method for processing an original image with a hollow structure. As shown in FIG7 , the method includes the following steps.
[0065] S710: Determine a first foreground mask image, a second foreground mask image, and a third foreground mask image corresponding to the original image.
[0066] S720: Determine, according to mask image attributes, whether the first foreground mask image or the third foreground mask image is a candidate foreground mask image.
[0067] S730 : Determine the candidate foreground mask image or the second foreground mask image as a target foreground mask image according to edge features of the candidate foreground mask image and difference features between the candidate foreground mask image and the second foreground mask image.
[0068] S740: Determine the background area of the original image.
[0069] S750: Acquire candidate hollowed-out areas in the first foreground mask image according to the attribute information of the background area.
[0070] The candidate hollowed-out area can be a region in the foreground consisting of blocks of color close to the background color. This can include target hollowed-out areas corresponding to the hollowed-out structure of the foreground object, or areas of the foreground object with a color close to the background color. For example, if the foreground object is a ring, the inner portion of the ring is the target hollowed-out area.
[0071] Exemplarily, the background color is estimated by the mean value based on the first background mask image and the original image to avoid the situation where the background colors of each pixel of the original image are not strictly consistent and thus affect the screening of candidate hollow areas.
[0072] Because the colors of the foreground's hollowed-out areas and parts of the foreground may be similar to the background color, it's impossible to simply calculate the color difference between the entire image and the background color and then identify the foreground object based on the color difference. In view of this problem, the disclosed embodiment uses the original image and background color to calculate the color difference, and then uses a color difference threshold to find all target color blocks in the original image that are similar to the background color. Thus, candidate hollowed-out areas are determined based on the target color blocks.
[0073] Exemplarily, the color difference between each color block region in the original image and the background color is determined; a target color block region is determined based on the color difference and a set color difference threshold, and the region corresponding to the target color block region within the first foreground mask image is selected as a candidate hollowing region. It should be noted that the set color difference threshold is used to determine whether the color block region is similar in color to the background region and can be set according to actual application requirements. Optionally, the set color difference threshold can reuse the set threshold used for pixel traversal in the above-mentioned embodiment.
[0074] S760 , filtering the candidate hollow regions according to the dilation result of the background region to obtain filtered candidate hollow regions.
[0075] The background area of the original image is expanded to a set size to obtain an expanded background area. Then, the intersection of the candidate hollow area and the expanded background area is calculated. If the intersection is empty, the candidate hollow area is used as the filtered candidate hollow area. If the intersection is not empty, the candidate hollow area is discarded.
[0076] S770 . For the filtered candidate hollowed-out area, determine whether the filtered candidate hollowed-out area is a target hollowed-out area according to the confidence of the target area corresponding to the filtered candidate hollowed-out area in the predicted mask image.
[0077] The confidence level can represent the degree of credibility of the target area corresponding to the filtered candidate hollowed-out area in the predicted mask image being predicted as the background area.
[0078] For example, the ring band part of a ring is a hollow structure. If the ring band is accurately predicted as the background area in the predicted mask image, it is determined that the confidence level of the subject segmentation model in predicting the ring band as the background is high.
[0079] Exemplarily, for the candidate filtered hollowed-out region, a target region corresponding to the candidate filtered hollowed-out region in the predicted mask image is obtained, and it is determined whether the prediction of the target region as the background is reliable. For example, whether the candidate filtered hollowed-out region is the target hollowed-out region is determined based on the proportion of pixels greater than 0 in the target region to the total pixels in the target region. For example, if the proportion of pixels greater than 0 in the target region is less than a set threshold, then the prediction result that the candidate filtered hollowed-out region is the background is sufficiently reliable, and the candidate filtered hollowed-out region can be determined as the target hollowed-out region.
[0080] Figure 8 is a schematic diagram of the process for generating a foreground mask for an object with a hollow structure, provided by an embodiment of the present disclosure. Referring to Figure 8 , for a first foreground mask image 810 corresponding to the original image, three candidate hollow regions can be identified. Because a reasonable candidate hollow region is separated from the background region by at least the distance of the inner and outer stroke layers, after the background region of the original image is expanded to a set size, candidate hollow regions that intersect with the expanded background region are discarded, resulting in filtered candidate hollow regions 830.
[0081] For either of the two smaller candidate hollowed-out regions 830, a first target region corresponding to that region in the binary prediction mask image 820, i.e., the ring region, is obtained. Since the ring region in the binary prediction mask image 820 is white, the proportion of pixels greater than 0 in the first target region is determined to be greater than a set threshold, indicating that the prediction result that the region is background is unreliable and the region is not the target hollowed-out region 840. For the larger candidate hollowed-out region 830, a second target region corresponding to the candidate hollowed-out region 830 in the binary prediction mask image 820, i.e., the ring region, is obtained. Since the ring region in the binary prediction mask image 820 is black, the proportion of pixels greater than 0 in the second target region is determined to be less than a set threshold, indicating that the prediction result that the candidate hollowed-out region 830 is background is sufficiently reliable and the candidate hollowed-out region 830 is the target hollowed-out region 840.
[0082] S780: Delete the target hollowed-out area from the target foreground mask image to obtain a new target foreground mask image.
[0083] As shown in Figure 8, the target hollow area 840 is deleted from the target foreground mask image, so that the target hollow area 840 in the target foreground mask image 850 becomes the background, thereby deleting the target hollow area 840 that was not accurately removed in the background removal method, and obtaining a new target foreground mask image 850.
[0084] S790: Acquire a foreground object in the original image according to the target foreground mask image.
[0085] The technical solution of the disclosed embodiment is as follows: after the target foreground mask map is determined, the color block connected domain is determined according to the color difference between the original image and the background color, and the candidate hollow area corresponding to the color block connected domain in the first foreground mask map is determined. Then, the candidate hollow area is filtered according to the expansion result of the background area of the original image to obtain filtered candidate hollow areas. According to the confidence of the target area corresponding to each filtered candidate hollow area in the predicted mask map, it is determined whether the filtered candidate hollow area is the target hollow area, and the target hollow area is deleted from the target foreground mask map to obtain a new target foreground mask map, so as to improve the situation where the hollow structure of the foreground object is mis-segmented due to inaccurate background removal, thereby improving the segmentation effect.
[0086] Figure 9 is a schematic structural diagram of a foreground object acquisition device provided in an embodiment of the present disclosure. The device can execute the foreground object acquisition method provided in any embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, a PC or a server, etc.
[0087] As shown in FIG. 9 , the apparatus includes: a mask image determining module 910 , an alternative mask image determining module 920 , a target mask image determining module 930 , and an object acquiring module 940 .
[0088] The mask map determination module 910 is used to determine a first foreground mask map, a second foreground mask map and a third foreground mask map corresponding to the original image, wherein the first foreground mask map and the second foreground mask map represent the mask of the foreground area of the original image generated using different models, and the third foreground mask map represents the mask of the foreground area of the predicted mask map of the original image.
[0089] The candidate mask image determining module 920 is configured to determine, according to mask image attributes, whether the first foreground mask image or the third foreground mask image is a candidate foreground mask image.
[0090] The target mask image determining module 930 is configured to determine the candidate foreground mask image or the second foreground mask image as a target foreground mask image based on edge features of the candidate foreground mask image and difference features between the candidate foreground mask image and the second foreground mask image.
[0091] The object acquisition module 940 is configured to acquire the foreground object in the original image according to the target foreground mask image.
[0092] Optionally, the target mask image determination module 930 includes: an edge feature determination unit, used to perform an opening operation on the alternative foreground mask image to obtain an edge feature map of the alternative foreground mask image; a difference area determination unit, used to determine a difference area in the alternative foreground mask image whose brightness is less than that of the second foreground mask image, and determine a difference area edge map based on the difference area; a target mask image determination unit, used to determine that the alternative foreground mask image or the second foreground mask image is the target foreground mask image based on the alternative foreground mask image, the edge feature map and the difference area edge map.
[0093] Furthermore, the difference region determining unit is specifically configured to: perform a dilation operation on the difference region to obtain a difference dilation region; and determine a difference region edge map according to a difference between the difference dilation region and the difference region.
[0094] Furthermore, the target mask image determination unit is specifically used to: determine the third image based on the intersection of the alternative foreground mask image and the difference area edge image; determine the fourth image based on the intersection of the edge feature map and the difference area edge map; and determine the alternative foreground mask image or the second foreground mask image as the target foreground mask image based on the area of the third image and the area of the fourth image.
[0095] Optionally, it also includes a hollow deletion module for: determining the background area of the original image before obtaining the foreground object in the original image according to the target foreground mask map; obtaining the candidate hollow area in the first foreground mask map according to the attribute information of the background area; filtering the candidate hollow area according to the expansion result of the background area to obtain a filtered candidate hollow area; for the filtered candidate hollow area, determining whether the filtered candidate hollow area is the target hollow area according to the confidence of the target area corresponding to the filtered candidate hollow area in the predicted mask map; deleting the target hollow area from the target foreground mask map to obtain a new target foreground mask map.
[0096] Optionally, the mask image determination module 910 is specifically used to: perform an erosion operation on the foreground object of the original image to obtain a first image, and use a first model to segment the foreground object and the background area in the first image to obtain a first foreground mask image; determine a foreground selection box in the original image, and use a second model to segment the foreground object and the background area in the foreground selection box to obtain a second foreground mask image; adjust the original image according to the predicted mask image, perform an erosion operation on the foreground object in the adjusted original image to obtain a second image, and use the first model to segment the foreground object and the background area in the second image to obtain a third foreground mask image.
[0097] Optionally, the alternative mask image determination module 920 is specifically used to: for the first foreground mask image or the third foreground mask image, determine the foreground mask area corresponding to the foreground mask area as a mask image attribute; and determine that the first foreground mask image or the third foreground mask image is an alternative foreground mask image based on the foreground mask area.
[0098] The foreground object acquisition device provided by the embodiment of the present disclosure can execute the foreground object acquisition method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0099] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0100] The embodiments of the present disclosure provide a foreground object acquisition method, apparatus, electronic device and storage medium. The method determines a first foreground mask image, a second foreground mask image and a third foreground mask image corresponding to an original image, selects the first foreground mask image or the third foreground mask image as an alternative foreground mask image according to the mask image attributes, and then selects the alternative foreground mask image or the second foreground mask image as a target foreground mask image. The foreground object in the original image is acquired according to the target foreground mask image. This can avoid the situation where some foreground areas are missed due to poor closure of the edge lines of the original image, and can also avoid the situation where color block areas with different colors from the background are mistakenly segmented. This improves the clipping quality of the color block areas with the same color as the background in the foreground object, and reduces the problem of unreliable mask prediction at edges, corners and other positions.
[0101] Figure 10 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to Figure 10 below, it shows a schematic diagram of the structure of an electronic device (such as the terminal device or server in Figure 10) 1000 suitable for implementing an embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in Figure 10 is only an example and should not bring any limitations to the functions and scope of use of the embodiments of the present disclosure.
[0102] As shown in FIG10 , the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the electronic device 1000 are also stored in the RAM 1003. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An edit / output (I / O) interface 1005 is also connected to the bus 1004.
[0103] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 10 illustrates the electronic device 1000 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0104] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0105] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0106] The electronic device provided by the embodiment of the present disclosure and the foreground object acquisition method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0107] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the foreground object acquisition method provided in the above embodiment is implemented.
[0108] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0109] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0110] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0111] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0112] Determining a first foreground mask map, a second foreground mask map, and a third foreground mask map corresponding to the original image, wherein the first foreground mask map and the second foreground mask map represent masks of a foreground area of the original image generated using different models, and the third foreground mask map represents a mask of a foreground area of a predicted mask map of the original image;
[0113] Determining, according to mask image attributes, the first foreground mask image or the third foreground mask image as a candidate foreground mask image;
[0114] determining, based on edge features of the candidate foreground mask image and difference features between the candidate foreground mask image and the second foreground mask image, that the candidate foreground mask image or the second foreground mask image is a target foreground mask image;
[0115] The foreground object in the original image is obtained according to the target foreground mask image.
[0116] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0118] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0119] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0120] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0121] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0122] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0123] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for acquiring a foreground object, comprising: Determine a first foreground mask map, a second foreground mask map, and a third foreground mask map corresponding to the original image, wherein the first foreground mask map and the second foreground mask map represent masks of a foreground area of the original image generated using different models, and the third foreground mask map represents a mask of a foreground area of a predicted mask map of the original image; Determining, according to the mask image attribute, the first foreground mask image or the third foreground mask image as a candidate foreground mask image; Determining the candidate foreground mask image or the second foreground mask image as a target foreground mask image according to edge features of the candidate foreground mask image and difference features between the candidate foreground mask image and the second foreground mask image; The foreground object in the original image is obtained according to the target foreground mask image.
2. The method according to claim 1, wherein determining the candidate foreground mask image or the second foreground mask image as the target foreground mask image according to edge features of the candidate foreground mask image and difference features between the candidate foreground mask image and the second foreground mask image comprises: Performing an opening operation on the candidate foreground mask image to obtain an edge feature map of the candidate foreground mask image; Determine a difference region in the candidate foreground mask image whose brightness is less than that of the second foreground mask image, and determine a difference region edge map according to the difference region; According to the candidate foreground mask image, the edge feature image and the difference area edge image, the candidate foreground mask image or the second foreground mask image is determined to be the target foreground mask image.
3. The method according to claim 2, wherein determining a difference region edge map according to the difference region comprises: Performing a dilation operation on the difference region to obtain a difference dilation region; A difference region edge map is determined according to the difference between the difference expansion region and the difference region.
4. The method according to claim 2, wherein determining the candidate foreground mask image or the second foreground mask image as the target foreground mask image according to the candidate foreground mask image, the edge feature map and the difference area edge map comprises: Determine a third image according to the intersection of the candidate foreground mask image and the difference area edge image; Determining a fourth image according to the intersection of the edge feature map and the difference area edge map; According to the third image area and the fourth image area, the candidate foreground mask image or the second foreground mask image is determined as the target foreground mask image.
5. The method according to claim 1, wherein before obtaining the foreground object in the original image according to the target foreground mask map, it also includes: Determine the background area of the original image; Acquire a candidate hollowed-out area in the first foreground mask image according to the attribute information of the background area; Filtering the candidate hollow regions according to the expansion result of the background region to obtain filtered candidate hollow regions; For the filtered candidate hollowed-out area, according to the predicted mask image corresponding to the filtered candidate hollowed-out area The confidence of the target area is used to determine whether the filtered candidate hollowed-out area is the target hollowed-out area; The target hollowed-out area is deleted from the target foreground mask image to obtain a new target foreground mask image.
6. The method according to claim 1, wherein the determining the first foreground mask image, the second foreground mask image and the third foreground mask image corresponding to the original image comprises: Performing an erosion operation on the foreground object of the original image to obtain a first image, and using a first model to segment the foreground object and the background area in the first image to obtain a first foreground mask image; Determine a foreground selection box in the original image, and use a second model to segment the foreground object and the background area in the foreground selection box to obtain a second foreground mask image; The original image is adjusted according to the predicted mask image, an erosion operation is performed on the foreground object in the adjusted original image to obtain a second image, and the foreground object and background area in the second image are segmented using the first model to obtain a third foreground mask image.
7. The method according to claim 1, wherein determining the first foreground mask image or the third foreground mask image as a candidate foreground mask image according to the mask image attribute comprises: For the first foreground mask image or the third foreground mask image, determining a foreground mask area corresponding to the foreground mask region as a mask image attribute; The first foreground mask image or the third foreground mask image is determined as a candidate foreground mask image according to the foreground mask area.
8. A foreground object acquisition device, comprising: A mask image determination module, used to determine a first foreground mask image, a second foreground mask image and a third foreground mask image corresponding to the original image, wherein the first foreground mask image and the second foreground mask image represent masks of the foreground area of the original image generated by different models, and the third foreground mask image represents a mask of the foreground area of the predicted mask image of the original image; A candidate mask image determining module, configured to determine the first foreground mask image or the third foreground mask image as a candidate foreground mask image according to mask image attributes; a target mask image determining module, configured to determine the candidate foreground mask image or the second foreground mask image as a target foreground mask image according to edge features of the candidate foreground mask image and difference features between the candidate foreground mask image and the second foreground mask image; The object acquisition module is used to acquire the foreground object in the original image according to the target foreground mask image.
9. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the foreground object acquisition method as described in any one of claims 1-7.
10. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the foreground object acquisition method according to any one of claims 1 to 7 when executed by a computer processor.
11. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Target image acquisition method and device, electronic equipment and storage medium
CN112634314A
Video matting method and system based on prediction foreground mask prediction and storage medium
CN113259605A
Foreground object acquisition method and device, electronic equipment and storage medium
CN117372466A
Methods and apparatus for soft edge masking
US8406566B1