Image processing method and apparatus, electronic device, and storage medium

By acquiring and adjusting the foreground mask diagram of the image and combining the prediction mask diagram, the problem of disbelieving mask prediction in the main cutout algorithm is solved, and the cutout effect and segmentation quality are improved.

WO2025092858A1PCT designated stage expired Publication Date: 2025-05-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/128664
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-10-30
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

When determining the prediction mask, the existing main body cutout algorithm has problems with some location disbelief, which affects the quality of segmentation and leads to poor cutout effect.

Method used

By acquiring the foreground area of ​​the original image, the first foreground mask map and the second foreground mask map are determined, combined with the predicted mask map, the foreground area is adjusted to optimize the mask quality, and the result mask map is finally determined for image processing.

Benefits of technology

The quality of the cutout of the color block areas in the foreground area with the same color as the background is improved, and the situation where the color block areas with different colors are misdivided is avoided, and the problem of mask prediction distrust in the edges, corners and other positions is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128664_08052025_PF_FP_ABST
    Figure CN2024128664_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an image processing method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring a foreground area of an original image, and determining a first foreground mask image on the basis of the foreground area; acquiring a predicted mask image corresponding to the original image, and obtaining a first target image on the basis of the predicted mask image; adjusting a foreground area of the first target image to obtain a second target image, and determining a second foreground mask image on the basis of the second target image; and determining a result mask image on the basis of the first foreground mask image and the second foreground mask image, and processing the original image on the basis of the result mask image. The embodiments of the present disclosure solve the problem in the solution of the prior art of poor image matting effect, and improves the image matting effect of the foreground area of the original image.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, device, electronic device and storage medium

[0001] This application claims priority to the Chinese invention patent application entitled “An image processing method, device, electronic device and storage medium” and application number 202311435122.9 filed on October 31, 2023. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to data processing technology, and more particularly to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0003] Image cutout technology is a technology that separates a part of an image from the original image.

[0004] Currently, a subject-based cutout algorithm is commonly used to determine a predicted mask, which is then used to perform cutout processing on the image to be processed. However, the predicted mask determined by a common subject-based cutout algorithm may have some unreliable locations, which affects the segmentation quality and leads to poor cutout results.

[0005] Summary of the Invention

[0006] The embodiments of the present disclosure provide an image processing method, device, electronic device, and storage medium, which can solve the problem of poor image cutout effect in related technologies.

[0007] In a first aspect, an embodiment of the present disclosure provides an image processing method, including: obtaining a foreground area of ​​an original image, and determining a first foreground mask map based on the foreground area; obtaining a predicted mask map corresponding to the original image, and obtaining a first target image according to the predicted mask map; adjusting the foreground area of ​​the first target image to obtain a second target image, and determining a second foreground mask map according to the second target image; determining a result mask map according to the first foreground mask map and the second foreground mask map, and processing the original image according to the result mask map.

[0008] In the second aspect, the embodiments of the present disclosure also provide an image processing device, including: a first foreground mask map determination module, used to obtain the foreground area of ​​the original image, and determine a first foreground mask map based on the foreground area; a first target image determination module, used to obtain a predicted mask map corresponding to the original image, and obtain a first target image according to the predicted mask map; a second foreground mask map determination module, used to adjust the foreground area of ​​the first target image to obtain a second target image, and determine a second foreground mask map according to the second target image; an image processing module, used to determine a result mask map according to the first foreground mask map and the second foreground mask map, and process the original image according to the result mask map.

[0009] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as described in any embodiment of the present disclosure.

[0010] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the image processing method as described in any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0012] FIG1 is a schematic flow chart of an image processing method provided by an embodiment of the present disclosure;

[0013] FIG2 is a schematic diagram of an edge stroke state in a foreground area provided by an embodiment of the present disclosure;

[0014] FIG3 is a schematic diagram of determining a foreground mask based on an original image according to an embodiment of the present disclosure;

[0015] FIG4 is a schematic diagram showing an effect of determining a foreground mask based on a predicted mask image according to an embodiment of the present disclosure;

[0016] FIG5 is a flow chart of another image processing method provided by an embodiment of the present disclosure;

[0017] FIG6 is a schematic diagram of a process for generating a foreground mask for an object with a hollow structure, provided by an embodiment of the present disclosure;

[0018] FIG7 is a schematic diagram of an overall cutout process provided by an embodiment of the present disclosure;

[0019] FIG8 is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure; and

[0020] FIG9 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0022] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0023] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0024] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0025] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0026] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0027] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0028] Because existing image generation algorithms generate pixel-style or stroke-style sticker images, the foreground area is blended into a pure white background, making them unusable as stickers. Using a general subject-cutting algorithm to generate a predicted mask for the sticker image results in poor cutout quality for foreground areas that share the same color as the background, and sometimes mis-segmenting areas with a different color. Furthermore, the foreground mask prediction may be unreliable at edges and corners of the foreground area.

[0029] Figure 1 is a flow chart illustrating an image processing method provided by an embodiment of the present disclosure. This embodiment of the present disclosure is applicable to the case of performing cutouts on pixel-style or stroke-style sticker images. The method can be performed by an image processing device, which can be implemented in software and / or hardware. Alternatively, the method can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.

[0030] As shown in FIG1 , the method includes: S110 , acquiring a foreground area of ​​an original image, and determining a first foreground mask image based on the foreground area.

[0031] Among them, the original image can be an image to be cutout processed. In the embodiment of the present disclosure, the original image can generate stickers in pixel style or stroke style. The sticker images in pixel style or stroke style have the following characteristics: First, the sticker images in pixel style or stroke style have a solid color background. For example, the background color of pixel style and stroke style is white. Second, the foreground area of ​​the sticker will not be close to the edge of the sticker image. Third, the sticker image can be divided into color block areas according to color blocks, and the color values ​​in each color block area are relatively uniform and fixed, and there is almost no gradient color on the screen.

[0032] The foreground area is the area corresponding to the foreground object in the original image. The edge area of ​​the foreground area represents the contour line of the foreground object, etc. Adjusting the edge area can be to adjust the thickness of the contour line. Since the contour line of the foreground area is composed of multiple line segments, there may be a situation where the contour of some areas is not closed, which leads to the foreground mask missing segmentation of the color block area corresponding to the unclosed area. By adjusting the edge area to make the contour line thicker, so that some unclosed areas become closed, the situation where the edge of the foreground area is not closed is optimized to a certain extent, and the third target image is obtained. The difference between the third target image and the original image is that some edge areas of the foreground area are changed from unclosed to closed. The first foreground mask image represents the foreground mask of the third target image. By fusing the first foreground mask image to the original image, the foreground area in the original image can be obtained.

[0033] Exemplarily, determining a first foreground mask image based on the foreground area includes: performing an erosion operation on the foreground area of ​​the original image to obtain a third target image; determining a first reference point in the third target image, and obtaining adjacent pixel points corresponding to the first reference point in the third target image; and determining the first foreground mask image based on the adjacent pixel points.

[0034] Among them, the erosion operation can be to take the minimum brightness value in a small area of ​​the original image as the brightness value of the small area. The erosion operation can be achieved by setting a convolution kernel to perform a convolution operation on the foreground area of ​​the original image. Since the original image is a pixel-style or stroke-style sticker image, the edge stroke of its foreground area may not be well closed. Figure 2 is a schematic diagram of the edge stroke state in a foreground area provided by an embodiment of the present disclosure. As shown in Figure 2, there is an area 210 at the bottom of the foreground object in the foreground area where the edge stroke is not well closed. Since the sticker image generated for the white background generally has a lower foreground color brightness than the background color, an erosion operation can be performed on the foreground area, and the erosion size can be a set number of pixels to improve the situation where the foreground edge is not closed to a certain extent. By performing an erosion operation on the foreground area, the darker black area in the edge area is enlarged, and the brighter white area in the edge area is reduced, thereby making the set unclosed area of ​​the edge area of ​​the foreground area become closed, see the area 220 with a well-closed edge stroke in Figure 2.

[0035] It should be noted that the unclosed area may be an area with a small enough edge stroke interval so that the area not well closed by the edge stroke can be closed by the erosion operation. If the edge stroke interval is large, it may not be closed by the erosion operation.

[0036] Among them, the reference point is the starting point of the traversal of the third target image. For example, the four vertices of the third target image can be used as the first reference point respectively. Starting from the first reference point, the third target image can be traversed along the set direction to obtain the adjacent pixel points of the first reference point. If the brightness difference between the adjacent pixel point and the first reference point is less than the set threshold, the adjacent pixel point is used as the new first reference point, and the third target image is continued to be traversed along the corresponding direction until the brightness difference between the traversed adjacent pixel point and the brightness of the reference point is equal to or greater than the set threshold, and the pixel traversal operation in the corresponding direction is stopped. It should be noted that the set direction can be a number of directions extending outward with the first reference point as the center. For example, the set direction can be the four directions of up, down, left and right with the first reference point as the center. The present disclosure does not specifically limit the meaning of the set direction. The threshold is set as the traversal stop condition, which can be related to the outline of the foreground area in actual applications, so that the pixel traversal along each direction stops at the edge of the foreground area, thereby obtaining a background mask.

[0037] In some embodiments, determining the first foreground mask image based on adjacent pixels includes: determining a first background mask image based on adjacent pixels whose brightness difference with the first reference point is less than a set threshold; and performing color inversion on the first background mask image to obtain the first foreground mask image. Because the background in the first background mask image is white and the foreground area is black, color inversion causes the background area of ​​the first background mask image to become black and the foreground area to become white, thereby removing the white area from the original image to obtain the first foreground mask image.

[0038] Figure 3 is a schematic diagram of determining a foreground mask based on an original image, according to an embodiment of the present disclosure. As shown in Figure 3, the above steps can effectively achieve a basic scene cutout without hollowed-out areas and with closed edges. However, for scenes with internal hollowed-out areas or imperfectly closed edges, the foreground mask quality is poor, and the foreground mask may miss segmenting certain color blocks.

[0039] S120: Obtain a predicted mask image corresponding to the original image, and obtain a first target image according to the predicted mask image.

[0040] Wherein, the predicted mask image represents the mask of the foreground area of ​​the original image predicted by the subject segmentation model. The subject segmentation model is a pre-trained model for obtaining the foreground mask of the original image. The predicted mask image is binarized based on a set threshold, and the target pixel point can be a pixel point in the original image corresponding to the reference area in the binarized predicted mask image. The reference area can be an area composed of pixels whose brightness meets the set conditions in the binarized predicted mask image, and the set conditions are used to determine the pixels whose brightness is greater than 0 in the binarized predicted mask image. For example, the reference area can be an area composed of pixels whose brightness is greater than 0 in the binarized predicted mask image. The attribute information includes the color of the target pixel point, etc.

[0041] Exemplarily, obtaining a predicted mask image corresponding to the original image and obtaining a first target image according to the predicted mask image includes:

[0042] The original image is input into the subject segmentation model to obtain a predicted mask image output by the subject segmentation model. A reference region in the predicted mask image whose brightness satisfies a set condition is determined; and the color of target pixels in the original image corresponding to the reference region is set to a set color to obtain a first target image.

[0043] FIG4 is a schematic diagram of an effect of determining a foreground mask based on a predicted mask image provided by an embodiment of the present disclosure. The original image 410 is input into the subject segmentation model to obtain a predicted mask image 420 output by the subject segmentation model. Since the foreground area has unclosed edges, when the predicted mask image 420 is determined by the subject segmentation model, there is a color block area in the predicted mask image 420 that is missed, as shown in the black area in the foreground mask in the predicted mask image 420 in FIG4 . The predicted mask image 420 is binarized based on a set threshold to obtain a binarized predicted mask image 430. Determine a reference area in the binarized predicted mask image 430 whose brightness meets the set conditions; set the color of the target pixel corresponding to the reference area in the original image to the set color to obtain a first target image 440.

[0044] For example, a reference region corresponding to pixels with brightness greater than 0 in the binarized prediction mask image is determined. Based on the reference region, the corresponding target pixel is searched for in the original image. The color of the target pixel is set to a color that differs significantly from the background color to obtain the first target image. For example, if the background color is white, the color of the target pixel can be set to black.

[0045] S130: Adjust the foreground area of ​​the first target image to obtain a second target image, and determine a second foreground mask image according to the second target image.

[0046] Exemplarily, a set convolution kernel is used to perform an erosion operation on the foreground area of ​​the first target image, so that the set unclosed area of ​​the edge area of ​​the foreground area becomes closed, thereby obtaining a second target image. For example, the erosion operation can be implemented by performing a convolution operation on the foreground area of ​​the first target image with a set convolution kernel. The erosion size can be a set number of pixels, which can improve the situation where the foreground edge is not closed to a certain extent. By performing an erosion operation on the foreground area of ​​the first target image, the black dark area in the edge area is enlarged and the white bright area in the edge area is reduced, thereby making the set unclosed area of ​​the edge area of ​​the foreground area closed.

[0047] Determine a second reference point in the second target image and obtain adjacent pixel points corresponding to the second reference point in the second target image. For example, the four vertices of the second target image can be used as second reference points. Starting from the second reference point, the second target image can be traversed along a set direction to obtain adjacent pixel points of the second reference point. If the brightness difference between the adjacent pixel point and the second reference point is less than a set threshold, the adjacent pixel point is used as a new first reference point, and the second target image is continued to be traversed along the corresponding direction until the brightness difference between the adjacent pixel point and the reference point is equal to or greater than the set threshold. Then, the pixel traversal operation in the corresponding direction is stopped. When the pixel traversal in all directions is stopped, the background mask in the second target image is obtained.

[0048] A second background mask is determined based on adjacent pixels whose brightness difference with the second reference point is less than a set threshold. The second background mask is then color-inverted to obtain a second foreground mask. Figure 4 also illustrates first foreground mask 450 and second foreground mask 460. Second foreground mask 460 can address the issue of missing segmentation of a certain color block in first foreground mask 450.

[0049] S140: Determine a result mask image according to the first foreground mask image and the second foreground mask image, and process the original image according to the result mask image.

[0050] The result mask image may be a mask image obtained by performing a matting process on the original image, and the one with better mask quality between the first foreground mask image and the second foreground mask image is used as the result mask image.

[0051] Exemplarily, determining a result mask image based on the first foreground mask image and the second foreground mask image includes: determining the area difference between the foreground mask area in the first foreground mask image and the foreground mask area in the second foreground mask image; and determining the first foreground mask image or the second foreground mask image as the result mask image based on the area difference.

[0052] As can be seen from Figure 4, the area of ​​the foreground mask in the second foreground mask image is larger than the area of ​​the foreground mask in the first foreground mask image. The area increase ratio of the foreground mask in the second foreground mask image compared to the first foreground mask image can be calculated, and the degree of improvement of the overall background removal effect by the second foreground mask image can be represented by the area increase ratio. Then, the quality of the foreground mask can be improved by improving the overall background removal effect.

[0053] If S is used C1 Characterize the area of ​​the foreground mask in the first foreground mask image, using S C2 The area of ​​the foreground mask in the second foreground mask image is represented, and the ratio is used to represent the area increase ratio. Then, the following relationship exists:

[0054] When the ratio is less than the set threshold, the risk of poor quality of the foreground mask in the second foreground mask image is determined to be higher, and the first foreground mask image is preferentially selected as the result mask image. Conversely, when the ratio is equal to or greater than the set threshold, the second foreground mask image is preferentially selected as the result mask image.

[0055] The resulting mask is fused to the original image to obtain the foreground area in the original image. This foreground area can then be overlaid on other images, enriching the editing methods of image assets and improving the user experience. When generating stickers from a pixel or stroke style original image, the foreground area can be subtracted from the sticker using the resulting mask, allowing the sticker to be overlaid on other images.

[0056] The technical solution of the disclosed embodiment determines a first foreground mask image for removing the background based on the original image, and determines a second foreground mask image for removing the background based on the original image and its corresponding predicted mask image. The mask image with the better mask quality between the first and second foreground masks is then used as the result mask image, and the original image is subjected to a cutout process based on the result mask image. This solves the problem of poor cutout effects in the main cutout algorithms of the related art, improves the cutout quality of color blocks in the foreground area that are the same color as the background, avoids the mis-segmentation of color blocks that are different from the background color, and reduces the problem of unreliable mask predictions at edges and corners.

[0057] FIG5 is a flow chart of another image processing method provided by an embodiment of the present disclosure. Based on the above embodiment, the present disclosure embodiment additionally defines a method for processing an original image with a hollow structure. As shown in FIG5 , the method includes:

[0058] S510: Acquire a foreground area of ​​an original image, and determine a first foreground mask image based on the foreground area.

[0059] In some embodiments, a set convolution kernel is used to perform an erosion operation on the foreground area of ​​the original image, so that the set unclosed area of ​​the edge area of ​​the foreground area becomes a closed state, thereby obtaining a third target image; a first reference point in the third target image is determined, and adjacent pixel points corresponding to the first reference point in the third target image are obtained; a first background mask image is determined based on adjacent pixel points whose brightness difference with the first reference point is less than a set threshold, and the color of the first background mask image is inverted to obtain a first foreground mask image.

[0060] S520: Obtain a predicted mask image corresponding to the original image, and obtain a first target image according to the predicted mask image.

[0061] In some embodiments, an original image is input into a subject segmentation model, a predicted mask image output by the subject segmentation model is obtained, the predicted mask image is binarized based on a set threshold, and a reference region in the binarized predicted mask image whose brightness meets a set condition is determined. The color of target pixels in the original image corresponding to the reference region is set to a color that is significantly different from the background color, thereby obtaining a first target image.

[0062] S530: Adjust the foreground area of ​​the first target image to obtain a second target image, and determine a second foreground mask image according to the second target image.

[0063] In some embodiments, a set convolution kernel is used to perform an erosion operation on the foreground area of ​​the first target image, so that the set unclosed area of ​​the edge area of ​​the foreground area becomes a closed state, thereby obtaining a second target image; a second reference point in the second target image is determined, and the adjacent pixel points corresponding to the second reference point in the second target image are obtained; a second background mask image is determined based on the adjacent pixel points whose brightness difference with the second reference point is less than a set threshold, and the color of the second background mask image is inverted to obtain a second foreground mask image.

[0064] S540: Determine a result mask image according to the first foreground mask image and the second foreground mask image.

[0065] In some embodiments, an area difference between a foreground mask area in the first foreground mask image and a foreground mask area in the second foreground mask image is determined, and the first foreground mask image or the second foreground mask image is determined as a result mask image based on the area difference.

[0066] S550: Determine the background area of ​​the original image.

[0067] S560: Acquire candidate hollowed-out areas in the first foreground mask image according to the attribute information of the background area.

[0068] The candidate hollow regions may be regions in the foreground consisting of color blocks close to the background color. The candidate hollow regions may include target hollow regions corresponding to the hollow structures of objects in the foreground region, or regions corresponding to structures of objects in the foreground region close to the background color. The target hollow region may represent the hollow structures of objects in the foreground region of the original image, as shown in FIG3 , for example, the inner region of the ring band of a ring. The attribute information of the background region may include the color and brightness of the background region, or a combination of color and brightness, etc., which is not specifically limited in the present embodiment.

[0069] In some embodiments, the background color is estimated by a mean value based on the original image, the first mask image, and the original image to avoid the situation where the background colors of each pixel of the original image are not strictly consistent and thus affect the screening of candidate hollow areas.

[0070] Because the colors of the foreground's hollowed-out areas and parts of the foreground may be close to the background color, it's impossible to simply calculate the color difference between the entire image and the background color and then determine the foreground area based on the color difference. In view of this problem, in the disclosed embodiment, the color difference is calculated using the original image and the background color. A color difference threshold is then used to determine all target color block areas in the original image that are close to the background color. Thus, candidate hollowed-out areas are determined based on the target color block areas.

[0071] Exemplarily, the color difference between each color block region and the background region in the original image is determined; a target color block region is determined based on the color difference and a set color difference threshold, and an area corresponding to the target color block region within the first foreground mask image is obtained as a candidate hollowed-out region. It should be noted that the set color difference threshold is used to determine whether the color block region is similar in color to the background region and can be set according to actual application requirements. Optionally, the set color difference threshold can reuse the set threshold used for pixel traversal in the above-mentioned embodiment.

[0072] S570 . For the candidate hollowed-out area, determine whether the candidate hollowed-out area is a target hollowed-out area according to the confidence of the target area corresponding to the candidate hollowed-out area in the predicted mask image.

[0073] The confidence level can represent the degree of confidence that the target area corresponding to the candidate hollowed-out area in the predicted mask image is predicted as the background area. Figure 6 illustrates a schematic diagram of the process for generating a foreground mask for an object with a hollowed-out structure, as provided by an embodiment of the present disclosure. As shown in Figure 6, the predicted mask image of a ring has a hollowed-out structure. If the predicted mask image accurately predicts the inner portion of the ring as a hollowed-out area, then the confidence level of the subject segmentation model predicting the inner portion of the ring as the background area is high.

[0074] For example, for a candidate hollowed-out region, if the ratio of the number of pixels greater than 0 to the total number of pixels in the target region corresponding to the candidate hollowed-out region in the predicted mask image is less than a set threshold, then the result of predicting the target region as the background region is sufficiently confident, and thus, the candidate hollowed-out region is determined to be the target hollowed-out region. Alternatively, if the ratio of the area greater than 0 to the target region is less than a set threshold, then the result of predicting the target region as the background region is sufficiently confident, and the candidate hollowed-out region can be determined as the target hollowed-out region.

[0075] Referring to Figure 6 , for the first foreground mask image 620 corresponding to the original image 610, three candidate cutout regions 640 can be identified. For the first candidate cutout region, which is positioned slightly above, a first target region corresponding to the first candidate cutout region is obtained from the predicted mask image 630, namely, the ring base region. Since the first target region is white in the predicted mask image 630, the proportion of pixels greater than 0 in the first target region is determined to be greater than a set threshold. This indicates that the prediction result that the first candidate cutout region is the background is unreliable and the first candidate cutout region is not the target cutout region 650. For the second candidate cutout region, which is positioned in the middle, a second target region corresponding to the second candidate cutout region is obtained from the predicted mask image 630, namely, the inner region of the ring. Since the second target region is black in the predicted mask image 630, the proportion of pixels greater than 0 in the second target region is determined to be less than a set threshold. This indicates that the prediction result that the second candidate cutout region is the background is sufficiently reliable and the second candidate cutout region is the target cutout region 650. For the third candidate cutout region at the bottom, the third target region corresponding to the third candidate cutout region in predicted mask image 630, i.e., the shank body region, is obtained. Since the third target region in predicted mask image 630 includes both white and gray, the proportion of pixels greater than 0 in the third target region is determined to be greater than a set threshold. This indicates that the prediction result that the third candidate cutout region is the background is not reliable, and the third candidate cutout region is not target cutout region 650.

[0076] S580: Delete the target hollowed-out area from the result mask image to obtain a new result mask image.

[0077] In some embodiments, the target hollow area 650 is deleted from the result mask image, so that the target hollow area in the result mask image becomes the background, and a new result mask image 660 is obtained, which can solve the problem of mis-segmentation of the hollow structure in the original result mask image.

[0078] S590: Process the original image according to the result mask image.

[0079] The technical solution of the embodiment of the present disclosure is to obtain the candidate hollow area in the first foreground mask image according to the attribute information of the background area of ​​the original image after the result mask image is determined, and for the candidate hollow area in the first foreground mask image, determine whether the candidate hollow area is the target hollow area according to the confidence of the target area corresponding to the candidate hollow area in the predicted mask image, delete the target hollow area from the result mask image, and obtain a new result mask image, thereby improving the situation of mis-segmentation of the internal hollow area and improving the segmentation effect of the image segmented by using the result mask image.

[0080] FIG7 is a schematic diagram of an overall cutout process provided by an embodiment of the present disclosure. As shown in FIG7 , a sticker image in pixel style or stroke style is obtained, and an erosion operation is performed on the foreground area of ​​the sticker image. The four vertices of the eroded sticker image are used as seed points, and the sticker image is traversed along a set direction to obtain adjacent pixel points. If the brightness difference between the adjacent pixel point and the seed point meets the set conditions, the adjacent pixel point is used as the new seed point to continue traversing the sticker image until the brightness of the adjacent pixel point and the seed point does not meet the set conditions, thereby obtaining a first background mask. The color of the first background mask is inverted to obtain a first foreground mask image.

[0081] For scenes where the edges of the sticker image are not well closed, obtain the predicted mask image corresponding to the sticker image, binarize the predicted mask image based on a set threshold, set the areas in the sticker image corresponding to pixels greater than 0 in the binarized predicted mask image to a color that is significantly different from the background color, and perform the erosion and pixel traversal steps above on the color-adjusted sticker image to obtain a second background mask. Invert the color of the second background mask image to obtain a second foreground mask image, thereby preventing the second foreground mask image from missing any color block areas. For the determined first and second foreground mask images, determine the difference between the foreground mask area in the first foreground mask image and the foreground mask area in the second foreground mask image. Based on the area difference, determine the first foreground mask image or the second foreground mask image as the result mask image.

[0082] For scenes where the foreground area of ​​the sticker image includes hollowed-out regions, meaning that part of the foreground object in the sticker image has hollowed-out structures, the aforementioned background removal method cannot accurately identify the hollowed-out regions in the foreground area. Furthermore, the colors of the hollowed-out foreground objects and parts of the foreground may be close to the background color, making it impossible to simply calculate the color difference between the full image and the background color and then determine the foreground area based on the color difference. To eliminate the foreground hollowed-out regions, a mean estimate of the background color can be performed based on the first background mask and the sticker image. The color difference is calculated using the sticker image and the background color, and a color difference threshold is used to calculate all target color blocks that are close to the background color. The regions corresponding to the target color blocks within the foreground area of ​​the first foreground mask image are selected as candidate hollowed-out regions. For each candidate hollowed-out region, the percentage of pixels with a brightness greater than 0 in the target area corresponding to the hollowed-out region in the binarized predicted mask image is determined. If the percentage of pixels with a brightness greater than 0 is sufficiently small, the prediction that the candidate hollowed-out region is background is sufficiently confident, and the candidate hollowed-out region is determined as the target hollowed-out region. The target hollowed-out region is then deleted from the result mask image to obtain a new result mask image.

[0083] FIG8 is a schematic diagram of the structure of an image processing device provided in an embodiment of the present disclosure. The device can execute the image processing method provided in any embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware. Optionally, the device can be implemented in an electronic device, such as a mobile terminal, a PC, or a server.

[0084] As shown in Figure 8, the apparatus includes: a first foreground mask image determination module 810, a first target image determination module 820, a second foreground mask image determination module 830, and an image processing module 840. The first edge adjustment module 810 is used to obtain a foreground area of ​​an original image and determine a first foreground mask image based on the foreground area; the brightness adjustment module 820 is used to obtain a predicted mask image corresponding to the original image and obtain a first target image based on the predicted mask image; the second edge adjustment module 830 is used to adjust the foreground area of ​​the first target image to obtain a second target image and determine a second foreground mask image based on the second target image; and the image processing module 840 is used to determine a result mask image based on the first foreground mask image and the second foreground mask image and process the original image based on the result mask image.

[0085] The technical solution provided by the embodiments of the present disclosure solves the problem of poor cutout effect of the main cutout algorithm in the related art, improves the cutout quality of the color block area with the same color as the background in the foreground area, avoids the situation where the color block area with a different color from the background is mis-segmented, and reduces the problem of unreliable mask prediction at edges, corners and other locations.

[0086] Optionally, the first foreground mask image determination module 810 is specifically used to: perform an erosion operation on the foreground area of ​​the original image to obtain a third target image; determine a first reference point in the third target image, and obtain adjacent pixel points corresponding to the first reference point in the third target image; and determine the first foreground mask image based on the adjacent pixel points.

[0087] Furthermore, determining the first foreground mask image based on adjacent pixel points includes: determining the first background mask image based on adjacent pixel points whose brightness difference with the first reference point is less than a set threshold; and inverting the color of the first background mask image to obtain the first foreground mask image.

[0088] Optionally, the first target image determination module 820 is specifically used to: determine a reference area in the predicted mask image whose brightness meets set conditions; set the color of the target pixel point corresponding to the reference area in the original image to a set color to obtain the first target image.

[0089] Optionally, the image processing module 840 is specifically used to: determine the area difference between the foreground mask area in the first foreground mask image and the foreground mask area in the second foreground mask image; and determine the first foreground mask image or the second foreground mask image as the result mask image based on the area difference.

[0090] Optionally, the device also includes a hollow area deletion module, which is used to: determine the background area of ​​the original image before processing the original image according to the result mask image; obtain the candidate hollow area in the first foreground mask image according to the attribute information of the background area; for the candidate hollow area, determine whether the candidate hollow area is the target hollow area according to the confidence of the target area corresponding to the candidate hollow area in the predicted mask image; delete the target hollow area from the result mask image to obtain a new result mask image.

[0091] Furthermore, the method of obtaining a candidate hollowed-out area in the first foreground mask image based on the attribute information of the background area includes: determining the color difference between each color block area in the original image and the background area; determining a target color block area based on the color difference and a set color difference threshold, and obtaining an area in the first foreground mask image corresponding to the target color block area as a candidate hollowed-out area.

[0092] The image processing device provided by the embodiments of the present disclosure can execute the image processing method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0093] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0094] FIG9 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to FIG9 , a schematic diagram of the structure of an electronic device (such as a terminal device or server in FIG9 ) 900 suitable for implementing an embodiment of the present disclosure is shown below. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG9 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0095] As shown in Figure 9, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An edit / output (I / O) interface 905 is also connected to the bus 904.

[0096] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although FIG9 shows the electronic device 900 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0097] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0098] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0099] The electronic device provided by the embodiment of the present disclosure and the image processing method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0100] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the image processing method provided by the above embodiment is implemented.

[0101] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0102] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0103] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0104] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to: obtain the foreground area of ​​the original image, and determine a first foreground mask map based on the foreground area; obtain a predicted mask map corresponding to the original image, and obtain a first target image according to the predicted mask map; adjust the foreground area of ​​the first target image to obtain a second target image, and determine a second foreground mask map according to the second target image; determine a result mask map according to the first foreground mask map and the second foreground mask map, and process the original image according to the result mask map.

[0105] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0107] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0108] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0109] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0111] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0112] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An image processing method, comprising: Acquire a foreground area of ​​the original image, and determine a first foreground mask image based on the foreground area; Acquire a predicted mask image corresponding to the original image, and obtain a first target image according to the predicted mask image; Adjusting the foreground area of ​​the first target image to obtain a second target image, and determining a second foreground mask image according to the second target image; A result mask image is determined according to the first foreground mask image and the second foreground mask image, and the original image is processed according to the result mask image.

2. The method according to claim 1, wherein determining a first foreground mask map based on the foreground area comprises: Performing an erosion operation on the foreground area of ​​the original image to obtain a third target image; Determine a first reference point in the third target image, and obtain adjacent pixel points corresponding to the first reference point in the third target image; A first foreground mask image is determined according to adjacent pixels.

3. The method according to claim 2, wherein determining the first foreground mask image according to adjacent pixels comprises: Determine a first background mask image based on adjacent pixel points whose brightness difference with the first reference point is less than a set threshold; The first background mask image is inverted in color to obtain a first foreground mask image.

4. The method according to claim 1, wherein obtaining the first target image according to the predicted mask image comprises: Determine a reference area in the prediction mask image whose brightness meets a set condition; The color of the target pixel points corresponding to the reference area in the original image is set to a set color to obtain a first target image.

5. The method according to claim 1, wherein determining a result mask image according to the first foreground mask image and the second foreground mask image comprises: Determine an area difference between a foreground mask area in the first foreground mask image and a foreground mask area in the second foreground mask image; The first foreground mask image or the second foreground mask image is determined as a result mask image according to the area difference.

6. The method according to claim 1, wherein before processing the original image according to the result mask image, it further comprises: Determine the background area of ​​the original image; Acquire a candidate hollowed-out area in the first foreground mask image according to the attribute information of the background area; For the candidate hollowed-out region, determining whether the candidate hollowed-out region is a target hollowed-out region according to the confidence of the target region corresponding to the candidate hollowed-out region in the predicted mask image; The target hollowed-out area is deleted from the result mask image to obtain a new result mask image.

7. The method according to claim 6, wherein the step of obtaining the candidate hollowed-out region in the first foreground mask image according to the attribute information of the background region comprises: Determine the color difference between each color block area and the background area in the original image; A target color block area is determined according to the color difference and a set color difference threshold, and an area in the first foreground mask image corresponding to the target color block area is obtained as a candidate hollowing area.

8. An image processing device, comprising: A first foreground mask image determining module, used to obtain a foreground area of ​​the original image, and determine a first foreground mask image based on the foreground area; A first target image determination module, used to obtain a predicted mask image corresponding to the original image, and obtain a first target image according to the predicted mask image; A second foreground mask image determining module, configured to adjust the foreground area of ​​the first target image to obtain a second target image, and determine a second foreground mask image according to the second target image; The image processing module is used to determine a result mask image according to the first foreground mask image and the second foreground mask image, and process the original image according to the result mask image.

9. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method as described in any one of claims 1 to 7.

10. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the image processing method according to any one of claims 1 to 7 when executed by a computer processor.

11. A computer program product, the computer program product being tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the image processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image processing method and device, terminal and storage medium

    CN110458921A

  • Target image acquisition method and device, electronic equipment and storage medium

    CN112634314A

  • Image processing method and device

    CN115880197A

  • Image processing method and device, electronic equipment and storage medium

    CN117455944A

  • Method and apparatus for detecting a click on an icon, device, and storage medium

    US20230195288A1