Head portrait matting method and device based on Ali cloud visual intelligent open platform
By performing cutout operations on Alibaba Cloud Vision Intelligent Open Platform and converting avatar cutouts into mask images, the platform's limitation on image size and the inability to return mask images are solved, and efficient and good quality background-free avatar cutout processing is achieved, reducing user costs.
Patent Information
- Application Number
- CN202510108361.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, Alibaba Cloud Vision Intelligent Open Platform has image size limitations when processing high-definition pictures, resulting in users needing to compress images to meet platform requirements, thereby losing image quality. At the same time, the platform cannot return mask images, limiting users' flexibility in avatar clamp processing.
By performing the cutout operation on Alibaba Cloud Vision Intelligent Open Platform, the portrait cutout and avatar cutout are obtained, and the avatar cutout is converted into mask, subtracting the portrait cutout to obtain avatarless cutout, and then performing the opening operation to remove noise, and finally obtain the required background-free avatarless cutout.
It solves the problem of size limitations for users when processing high-definition pictures, improves the efficiency and quality of image processing, reduces user costs, and provides more flexible avatar cutting processing capabilities.
Smart Images

Figure CN120070487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method and device for avatar matting based on the Alibaba Cloud Vision Intelligence Open Platform. Background Art
[0002] In today's e-commerce field, when online sellers promote products, they often need to go through a series of complex steps to ensure the online display effect of the goods. First of all, professional photographers and models must be hired to shoot at the corresponding venue. This process is not only time-consuming but also costly. Photographers need to have professional skills and equipment to capture the best angles and details of the products to ensure that the pictures can attract consumers' attention.
[0003] Existing users cannot do the above-mentioned way when selling products across borders. Therefore, after shooting locally, the background image and the human face are replaced to facilitate promotion in the corresponding country. In the existing technology, the Alibaba Cloud Vision Intelligence Open Platform provides a segmentation and matting function, but this service does have certain limitations when processing pictures.
[0004] Specifically, the platform requires that the size of the input picture must be less than 2000×2000 pixels, that is, the longest side does not exceed 1999 pixels. This limitation has led to many users being unable to directly use this function to process the high-definition pictures in their hands because most existing picture resolutions exceed this limit; users have to take some compromise measures, such as first compressing the picture to a size that meets the platform requirements and then using the platform's segmentation and matting function for processing. However, this approach will inevitably sacrifice the quality of the picture, especially some details and clarity may be lost during the compression process.
[0005] Moreover, the Alibaba Cloud Vision Intelligence Open Platform only provides clothing matting, portrait matting (without background) and avatar matting. Among them, clothing matting and portrait matting can return the mask image, while avatar matting cannot return the mask image; and users need to obtain the matting of the removed background and avatar during use, which results in users having to perform matting by means of PS, greatly reducing the work efficiency and increasing the cost of merchants. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method and device for avatar matting based on the Alibaba Cloud Vision Intelligence Open Platform, which greatly improves the efficiency of users and reduces the cost of users.
[0007] In a first aspect, the present invention provides a method for avatar matting based on the Alibaba Cloud Vision Intelligence Open Platform, including the following steps:
[0008] Step 1: Perform matte extraction on the image uploaded by the user through the Visual Intelligence Open Platform to obtain a portrait matte and a head portrait matte. Among them, the portrait matte is a mask image, and the head portrait matte is not a mask image;
[0009] Step 2: Convert the head portrait matte into a head portrait mask image;
[0010] Step 3: Subtract the head portrait mask image from the portrait matte to obtain a first headless matte;
[0011] Step 4: Perform an opening operation on the first headless matte to remove the set noise, obtain a second headless matte, and restore the second headless matte to obtain the required image.
[0012] In a second aspect, the present invention provides a head portrait removal device based on the Alibaba Cloud Visual Intelligence Open Platform, including:
[0013] A matte extraction module that performs matte extraction on the image uploaded by the user through the Visual Intelligence Open Platform to obtain a portrait matte and a head portrait matte. Among them, the portrait matte is a mask image, and the head portrait matte is not a mask image;
[0014] A conversion module that converts the head portrait matte into a head portrait mask image;
[0015] An acquisition module that subtracts the head portrait mask image from the portrait matte to obtain a first headless matte;
[0016] A restoration module that performs an opening operation on the first headless matte to remove the set noise, obtain a second headless matte, and restore the second headless matte to obtain the required image.
[0017] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0018] Through the method of the present invention, users can process images through the Visual Intelligence Open Platform to obtain the headless and backgroundless matte they need, and then optimize the matte according to their own needs, which greatly saves the time of users, improves work efficiency, and greatly reduces the cost of users.
[0019] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.
[0021] Figure 1It is the flowchart of the method in Embodiment 1 of the present invention;
[0022] Figure 2 It is the structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed implementation manners
[0023] The overall idea of the technical solution in the embodiments of this application is as follows:
[0024] Due to the cropping size limit of the Alibaba Cloud Vision Intelligence Open Platform, when the longest side exceeds 2000 pixels, it is necessary to perform proportional scaling. Call Imgproc.resize and use the Imgproc.INTER_LANCZOS4 algorithm to perform image proportional scaling.
[0025] Call Imgproc.resize and use the Imgproc.INTER_LANCZOS4 algorithm to perform image proportional scaling. This algorithm can reduce the influence of artifacts while maintaining the edge sharpness.
[0026] Obtain the cropping result according to the mask image + original image:
[0027] a. Take out the alpha channel of the original image. Call Core.min, which will take the minimum value of the mask image and the alpha channel of the original image to ensure that the original information of the alpha channel is retained. Through this step of processing, the lines of the obtained cropped image can be made more perfect;
[0028] b. Remove the alpha channel of the original image and add the mask image as the new alpha channel for layer merging;
[0029] c. Set the color value of the transparent area to black. This step is to reduce the size of the cropping result image and save storage costs.
[0030] d. Portrait removal head cropping processing
[0031] i. The portrait cropping result returned by Alibaba is not a mask image, so it needs to be converted into a mask image first (traverse the pixel points to generate a mask image with black for transparent and white for opaque);
[0032] ii. Subtract the two mask images. Call Core.subtract (a method in OpenCV, subtract the portrait cropping from the head cropping);
[0033] iii. Perform an opening operation on the subtraction result (remove small noises). The purpose is to remove small pixel blocks such as hair. The specific steps are as follows:
[0034] 1. First, create a morphological operation point Imgproc.getStructuringElement(Imgproc.MORPH_ELLIPSE, new Size(3, 3)). Here, a 3x3 pixel point is created. That is, if there is an independent pixel block within 3x3, it will be removed.
[0035] 2. Call Imgproc.morphologyEx for opening operation to remove pixel blocks within 3x3.
[0036] 3. Principle of opening operation: First perform erosion operation, and then perform dilation operation to remove small white areas in the image.
[0037] Function of opening operation:
[0038] The opening operation can remove bright areas smaller than the structuring element (in a white foreground image) while maintaining the general outline of the overall shape.
[0039] Removing noise: The opening operation can effectively remove small noise points or isolated bright points in the image because these noise points are usually smaller than the structuring element and will be removed during the erosion operation.
[0040] Separating connected objects: If there is only a thin connecting bridge between two objects, the opening operation can disconnect them.
[0041] Smoothing object contours: The opening operation can smooth the burrs or spikes on the edges of objects.
[0042] The core code is as follows:
[0043] / / Create a fully transparent matrix with the same size as the original image for comparison
[0044] Mat compareAlpha = new Mat(outImg.size(), CvType.CV_8UC1, Scalar.all(0.0));
[0045] / / Used to save the comparison result
[0046] Mat compareResult = new Mat();
[0047] / / Compare the alpha channel of the original image with alpha. The positions with a value of 0 in the obtained mask are the transparent positions in the original image.
[0048] Core.compare(outPlanes.get(3), compareAlpha, compareResult, Core.CMP_EQ);
[0049] / / Create a completely black matrix with the same size as the original image
[0050] Mat black = new Mat(outImg.size(), outImg.type(), Scalar.all(0));
[0051] / / Copy the black color in black to the corresponding positions in outImg according to the mask
[0052] Core.bitwise_and(black, outImg, outImg, compareResult);
[0053] / / Portrait matting mask image
[0054] Mat humanImageMat = toMat(humanImageBytes);
[0055] mats.add(humanImageMat);
[0056] / / The portrait matting result returned by Alibaba is not a mask image, so it needs to be converted into a mask image
[0057] Mat subImageMat;
[0058] try {
[0059] / / Traverse the pixel points to generate a mask image with black for transparent and white for opaque
[0060] subImageMat = toMat(toGray(subImageBytes));
[0061] } catch (IOException e) {
[0062] throw new RuntimeException(e);
[0063] }
[0064] mats.add(subImageMat);
[0065] / / Subtract the two mask images
[0066] Mat result = new Mat();
[0067] Core.subtract(humanImageMat, subImageMat, result);
[0068] mats.add(result);
[0069] / / Perform opening operation to remove small white areas
[0070] Mat kernel = Imgproc.getStructuringElement(Imgproc.MORPH_ELLIPSE, new Size(3, 3));
[0071] Imgproc.morphologyEx(result, result, Imgproc.MORPH_OPEN, kernel);
[0072] mats.add(kernel);
[0073] Example 1
[0074] As Figure 1 shown, this embodiment provides a method for removing the avatar based on the Alibaba Cloud Vision Intelligence Open Platform, including the following steps:
[0075] Step 1: Perform matte extraction on the image uploaded by the user through the Vision Intelligence Open Platform to obtain a portrait matte and an avatar matte. Among them, the portrait matte is a mask image, and the avatar matte is not a mask image;
[0076] Step 2: Convert the avatar matte into an avatar mask image;
[0077] Step 3: Subtract the avatar mask image from the portrait matte to obtain a first headless matte;
[0078] Step 4: Perform an opening operation on the first headless matte to remove the set noise, obtain a second headless matte, and restore the second headless matte to obtain the required image.
[0079] In this embodiment, preferably, step 1 is specifically as follows: Determine the pixels of the image uploaded by the user. If the pixels of the image are less than 2000×2000, directly perform a matting operation through the Visual Intelligence Open Platform to obtain a portrait matting and a head matting, where the portrait matting is a mask image and the head matting is not a mask image; if the pixels of the image are greater than or equal to 2000×2000, scale the image proportionally through a scaling algorithm to obtain a scaled image; perform a matting operation on the scaled image through the Visual Intelligence Open Platform to obtain a portrait scaled matting and a head scaled matting, where the portrait scaled matting is a mask image and the head scaled matting is not a mask image, and convert the head scaled matting into a head mask scaled image; use the scaling algorithm to restore the portrait scaled matting and the head mask scaled image to the original-sized portrait matting and head mask image, and enter step 3.
[0080] In this embodiment, preferably, step 1 is specifically as follows: Determine the pixels of the image uploaded by the user. If the pixels of the image are less than 2000×2000, directly perform a matting operation through the Visual Intelligence Open Platform to obtain a portrait matting and a head matting, where the portrait matting is a mask image and the head matting is not a mask image; if the pixels of the image are greater than or equal to 2000×2000, use the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to scale the image proportionally to obtain a scaled image; perform a matting operation on the scaled image through the Visual Intelligence Open Platform to obtain a portrait scaled matting and a head scaled matting, where the portrait scaled matting is a mask image and the head scaled matting is not a mask image, and convert the head scaled matting into a head mask scaled image; use the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to proportionally restore the portrait scaled matting and the head mask scaled image to the original-sized portrait matting and head mask image, and enter step 3.
[0081] In this embodiment, preferably, step 3 is specifically as follows: Call the Core.subtract method in OpenCV to subtract the head mask image from the portrait matting to obtain a first headless matting;
[0082] In this embodiment, preferably, step 4 is specifically as follows: First, create an opening operation point Imgproc.getStructuringElement. The opening operation point is a pixel block with a set threshold. Perform an opening operation on the first headless image cutout, remove the independent pixel blocks less than or equal to the set threshold, and obtain a second headless image cutout. Read the alpha channel of each pixel point in the image uploaded by the user to obtain a first matrix. Read the alpha channel of each pixel point in the second headless image cutout to obtain a second matrix. Call Core.min to merge the first matrix and the second matrix to obtain a third matrix; replace the alpha channel in the image uploaded by the user with the third matrix to obtain the required image.
[0083] Based on the same inventive concept, the present application also provides an apparatus corresponding to the method in Embodiment 1. For details, see Embodiment 2.
[0084] Embodiment 2
[0085] As Figure 2 shown, in this embodiment, an apparatus for head removal based on the Alibaba Cloud Vision Intelligence Open Platform is provided, including:
[0086] An image cutout module that performs an image cutout operation on the image uploaded by the user through the Vision Intelligence Open Platform to obtain a human figure cutout and a head cutout. Among them, the human figure cutout is a mask image, and the head cutout is not a mask image;
[0087] A conversion module that converts the head cutout into a head mask image;
[0088] An acquisition module that subtracts the head mask image from the human figure cutout to obtain a first headless image cutout;
[0089] A restoration module that performs an opening operation on the first headless image cutout to remove the set noise, obtains a second headless image cutout, and restores the second headless image cutout to obtain the required image.
[0090] In this embodiment, preferably, the matte extraction module is specifically configured to: determine the pixels of the image uploaded by the user. If the pixels of the image are less than 2000×2000, directly perform matte extraction operations through the Visual Intelligence Open Platform to obtain a portrait matte and a head portrait matte, where the portrait matte is a mask image and the head portrait matte is not a mask image; if the pixels of the image are greater than or equal to 2000×2000, scale the image proportionally through a scaling algorithm to obtain a scaled image; perform matte extraction operations on the scaled image through the Visual Intelligence Open Platform to obtain a scaled portrait matte and a scaled head portrait matte, where the scaled portrait matte is a mask image and the scaled head portrait matte is not a mask image, and convert the scaled head portrait matte into a scaled head portrait mask image; use the scaling algorithm to restore the scaled portrait matte and the scaled head portrait mask image to the original-sized portrait matte and head portrait mask image, and enter the acquisition module.
[0091] In this embodiment, preferably, the matte extraction module is specifically configured to: determine the pixels of the image uploaded by the user. If the pixels of the image are less than 2000×2000, directly perform matte extraction operations through the Visual Intelligence Open Platform to obtain a portrait matte and a head portrait matte, where the portrait matte is a mask image and the head portrait matte is not a mask image; if the pixels of the image are greater than or equal to 2000×2000, use the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to scale the image proportionally to obtain a scaled image; perform matte extraction operations on the scaled image through the Visual Intelligence Open Platform to obtain a scaled portrait matte and a scaled head portrait matte, where the scaled portrait matte is a mask image and the scaled head portrait matte is not a mask image, and convert the scaled head portrait matte into a scaled head portrait mask image; use the Imgproc.resize in OpenCV with the Imgproc.INTER_LANCZ0S4 algorithm to scale the scaled portrait matte and the scaled head portrait mask image proportionally back to the original-sized portrait matte and head portrait mask image, and enter the acquisition module.
[0092] In this embodiment, preferably, the acquisition module is specifically configured to: subtract the head portrait mask image from the portrait matte by calling the Core.subtract method in OpenCV to obtain a first headless matte.
[0093] In this embodiment, preferably, the reduction module is specifically: first, create an opening operation point Imgproc.getStructuringElement, where the opening operation point is a pixel block with a set threshold. Perform an opening operation on the first headless matte, remove the independent pixel blocks less than or equal to the set threshold, and obtain a second headless matte. Read the alpha channel of each pixel point in the user-uploaded image to obtain a first matrix, read the alpha channel of each pixel point in the second headless matte to obtain a second matrix, and call Core.min to merge the first matrix and the second matrix to obtain a third matrix; replace the alpha channel in the user-uploaded image with the third matrix to obtain the required image.
[0094] Since the device introduced in the second embodiment of the present invention is the device used to implement the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device used in the method of the first embodiment of the present invention falls within the scope of protection of the present invention.
[0095] Although the specific implementation manners of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A head removal method based on Alibaba Cloud Visual Intelligence Open Platform, characterized in that: The steps include: Step 1: Perform a cutout operation on the image uploaded by the user through the visual intelligence open platform to obtain a portrait cutout and an avatar cutout, wherein the portrait cutout is a mask image, and the avatar cutout is not a mask image; Step 2: Convert the avatar cutout into an avatar mask image; Step 3, subtract the head mask image from the portrait cutout image to obtain the first headless cutout image; Step 4: Open the first headless image to remove the set noise to obtain the second headless image, and restore the second headless image to obtain the desired image.
2. According to claim 1, a head image removal method based on Alibaba Cloud Visual Intelligence Open Platform is characterized in that: The step 1 is specifically as follows: judging the pixels of the picture uploaded by the user, if the pixels of the picture are less than 2000×2000, directly performing a cutout operation through the visual intelligent open platform to obtain a portrait cutout and an avatar cutout, wherein the portrait cutout is a mask image, and the avatar cutout is not a mask image; if the pixels of the picture are greater than or equal to 2000×2000, scaling the picture in proportion through a scaling algorithm to obtain a scaled image; performing a cutout operation on the scaled image through the visual intelligent open platform to obtain a portrait scaled cutout and an avatar scaled cutout, wherein the portrait scaled cutout is a mask image, and the avatar scaled cutout is not a mask image, and converting the avatar scaled cutout into an avatar mask scaled image; restoring the portrait scaled cutout and the avatar mask scaled image to the original size portrait cutout and avatar mask image using a scaling algorithm, and entering step 3.
3. The head image removal method based on Alibaba Cloud Visual Intelligence Open Platform according to claim 1 is characterized in that: The step 1 is specifically as follows: the pixels of the picture uploaded by the user are judged. If the pixels of the picture are less than 2000×2000, the cutout operation is directly performed through the visual intelligence open platform to obtain the portrait cutout and the avatar cutout, wherein the portrait cutout is a mask image, and the avatar cutout is not a mask image; if the pixels of the picture are greater than or equal to 2000×2000, the Imgproc.resize in OpenCV is used to use the Imgproc.INTER_LANCZ0S4 algorithm to scale the picture in proportion to obtain a scaled image. ; Perform a cutout operation on the scaled image through the visual intelligence open platform to obtain a portrait scaled cutout and an avatar scaled cutout, wherein the portrait scaled cutout is a mask image, and the avatar scaled cutout is not a mask image, and the avatar scaled cutout is converted into an avatar mask scaled image; Use the Imgproc.INTER_LANCZ0S4 algorithm in OpenCV to proportionally restore the portrait scaled cutout and the avatar scaled cutout to the original size portrait cutout and avatar mask image, and proceed to step 3.
4. The head image removal method based on Alibaba Cloud Visual Intelligence Open Platform according to claim 1 is characterized in that: The step 3 is specifically as follows: by calling the Core.subtract method in OpenCV, the portrait cutout image is subtracted from the head mask image to obtain a first headless cutout image.
5. The head image removal method based on Alibaba Cloud Visual Intelligence Open Platform according to claim 1 is characterized in that: The step 4 is specifically as follows: firstly, create an opening operation point Imgproc.getStructuringElement, where the opening operation point is a pixel block with a set threshold, perform an opening operation on the first headless cutout, remove independent pixel blocks less than or equal to the set threshold, and obtain a second headless cutout, read out the transparent channel of each pixel in the picture uploaded by the user, and obtain a first matrix, read out the transparent channel of each pixel in the second headless cutout, and obtain a second matrix, and call Core.min to merge the first matrix with the second matrix, and obtain a third matrix; The transparent channel in the picture uploaded by the user is replaced by the third matrix to obtain the desired picture.
6. A head removal device based on Alibaba Cloud Visual Intelligence Open Platform, characterized in that: include: The cutout module cuts out the image uploaded by the user through the visual intelligence open platform to obtain the portrait cutout and the head portrait cutout. The portrait cutout is a mask image, while the head portrait cutout is not a mask image. The conversion module converts the avatar cutout image into the avatar mask image; The acquisition module subtracts the head mask image from the portrait cutout image to obtain the first headless cutout image; The restoration module performs an opening operation on the first headless image to remove the set noise to obtain a second headless image, and restores the second headless image to obtain the desired image.
7. The head removal device based on Alibaba Cloud Visual Intelligence Open Platform according to claim 6 is characterized in that: The cutout module specifically comprises: judging the pixels of the picture uploaded by the user, if the pixels of the picture are less than 2000×2000, directly performing the cutout operation through the visual intelligent open platform to obtain the portrait cutout and the avatar cutout, wherein the portrait cutout is a mask image, and the avatar cutout is not a mask image; if the pixels of the picture are greater than or equal to 2000×2000, scaling the picture in proportion through a scaling algorithm to obtain a scaled image; performing the cutout operation on the scaled image through the visual intelligent open platform to obtain the portrait scaled cutout and the avatar scaled cutout, wherein the portrait scaled cutout is a mask image, and the avatar scaled cutout is not a mask image, and converting the avatar scaled cutout into an avatar mask scaled image; restoring the portrait scaled cutout and the avatar mask scaled image to the original size portrait cutout and avatar mask image using a scaling algorithm, and entering the acquisition module.
8. The head portrait removal device based on Alibaba Cloud Visual Intelligence Open Platform according to claim 7, characterized in that: The cutout module specifically: judges the pixels of the picture uploaded by the user. If the pixels of the picture are less than 2000×2000, the cutout operation is directly performed through the visual intelligence open platform to obtain the portrait cutout and the avatar cutout, where the portrait cutout is a mask image, and the avatar cutout is not a mask image; if the pixels of the picture are greater than or equal to 2000×2000, the Imgproc.resize in OpenCV is used to use the Imgproc.INTER_LANCZ0S4 algorithm to scale the picture in proportion to obtain a scaled image. ; The scaled image is cut out through the visual intelligence open platform to obtain the portrait scaled cutout and the avatar scaled cutout, wherein the portrait scaled cutout is a mask image, and the avatar scaled cutout is not a mask image, and the avatar scaled cutout is converted into an avatar mask scaled image; the portrait scaled cutout and the avatar scaled cutout are restored to the original size of the portrait cutout and the avatar mask image in proportion through Imgproc.resize in OpenCV using the Imgproc.INTER_LANCZ0S4 algorithm, and enter the acquisition module.
9. The head portrait removal device based on Alibaba Cloud Visual Intelligence Open Platform according to claim 6, characterized in that: The acquisition module specifically calls the Core.subtract method in OpenCV to subtract the head mask image from the portrait image to obtain a first headless image.
10. The head removal device based on Alibaba Cloud Visual Intelligence Open Platform according to claim 6, characterized in that: The restoration module specifically comprises: firstly creating an opening operation point Imgproc.getStructuringElement, wherein the opening operation point is a pixel block with a set threshold, performing an opening operation on the first headless cutout, removing independent pixel blocks less than or equal to the set threshold, obtaining a second headless cutout, reading out the transparent channel of each pixel in the picture uploaded by the user, obtaining a first matrix, reading out the transparent channel of each pixel in the second headless cutout, obtaining a second matrix, calling Core.min to merge the first matrix with the second matrix, obtaining a third matrix; The transparent channel in the picture uploaded by the user is replaced by the third matrix to obtain the desired picture.